Version 2.20 has reached end of support. Upgrade to a newer supported version.
LogoLogo
⌘Ctrlk
Contact support
  • Home
  • SaaS
  • Self-hosted
  • Multi-tenant
AI Assistant
Good afternoon

I'm here to help you with the docs.

⌘Ctrli
AI Based on your context
LogoLogo
    • Overview
    • What's New
    • Installation
    • Authentication and Authorization
    • Advanced Setup
    • Infrastructure Procedures
    • Manage AI Initiatives
    • Scheduling and Resource Optimization
    • Policies
    • Monitor Performance and Health
    • Introduction to Workloads
    • NVIDIA Run:ai Workload Types
    • Workloads
    • Workload Assets
    • Workload Templates
    • Experiment Using Workspaces
    • Train Models Using Training
    • Deploy Models Using Inference
      • Deploy a Custom Inference Workload
      • Deploy Inference Workloads from Hugging Face
      • Deploy Inference Workloads with NVIDIA NIM
      • Quick Starts
        • Run Your First Custom Inference Workload
    • CLI Reference
    • Product Support Policy
    • Product Version Life Cycle
For the complete documentation index, see llms.txt. This page is also available as Markdown.
  1. Self-hosted
  2. v2.20
  3. Workloads in NVIDIA Run:ai
  4. Deploy Models Using Inference

Quick Starts

Run Your First Custom Inference Workload
PreviousDeploy Inference Workloads with NVIDIA NIMNextRun Your First Custom Inference Workload

Last updated 11 months ago

LogoLogo

Corporate Info

  • NVIDIA.com Home
  • About NVIDIA
  • Privacy Policy
  • Your Privacy Choices
  • Terms of Service

NVIDIA Developer

  • Developer Home
  • Blog

Resources

  • Contact Us
  • Developer Program

Copyright © 2026, NVIDIA Corporation.