# Welcome to NVIDIA Run:ai Documentation

NVIDIA Run:ai accelerates AI operations with dynamic orchestration across the AI life cycle, maximizing GPU efficiency, scaling workloads, and integrating seamlessly into hybrid AI infrastructure with zero manual effort.

Find all the product information, step-by-step guides, and references you need.

<table data-view="cards" data-full-width="false"><thead><tr><th></th><th data-hidden data-card-cover data-type="files"></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><h4>SaaS Documentation</h4><p>For customers using NVIDIA Run:ai’s fully managed, cloud-hosted platform. Always kept up to date with the latest features.</p></td><td><a href="/files/EbycmX885Gel7Gvxswhn">/files/EbycmX885Gel7Gvxswhn</a></td><td><a href="/spaces/LiY1aIqfxD3a58ufUYOM">/spaces/LiY1aIqfxD3a58ufUYOM</a></td></tr><tr><td><h4>Self-hosted Documentation</h4><p>For on-prem and private cloud deployments. Versioned and aligned with your cluster releases.</p></td><td><a href="/files/9362AEHGsY3Vpt2Obcdl">/files/9362AEHGsY3Vpt2Obcdl</a></td><td><a href="/spaces/ZNGVYs8RM1KF89eNrzzE/pages/59XZXEn6cJiHUIhnmHXv">/spaces/ZNGVYs8RM1KF89eNrzzE/pages/59XZXEn6cJiHUIhnmHXv</a></td></tr><tr><td><h4>Multi-tenant Documentation</h4><p>For on-prem and private cloud deployments that use a centralized control plane to serve multiple isolated organizations. Versioned and aligned with your cluster releases.</p></td><td><a href="/files/eNjhwRllOvGHnAZ0Ni7G">/files/eNjhwRllOvGHnAZ0Ni7G</a></td><td><a href="/spaces/5GMWvizCFSWw09pjtlbu/pages/R9o1V77Px0dh0bjSXbke">/spaces/5GMWvizCFSWw09pjtlbu/pages/R9o1V77Px0dh0bjSXbke</a></td></tr></tbody></table>

### Resources

Explore videos, technical blogs, and NVIDIA On-Demand sessions to learn more about NVIDIA Run:ai capabilities, AI infrastructure, and GPU orchestration.

<table data-view="cards"><thead><tr><th></th><th data-hidden data-card-cover data-type="image">Cover image</th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><h4>Videos</h4></td><td><a href="/files/GM1WVpjGpskhlq1Qs7D3">/files/GM1WVpjGpskhlq1Qs7D3</a></td><td><a href="/spaces/ZNGVYs8RM1KF89eNrzzE/pages/vtaIx5eMk3Do9bPnTnOU">/spaces/ZNGVYs8RM1KF89eNrzzE/pages/vtaIx5eMk3Do9bPnTnOU</a></td></tr><tr><td><h4>Blogs and Articles</h4></td><td><a href="/files/LKPExCf9sOkFpwGJ0W19">/files/LKPExCf9sOkFpwGJ0W19</a></td><td><a href="/spaces/ZNGVYs8RM1KF89eNrzzE/pages/0sx4sX8y3wsswrwi5ldr">/spaces/ZNGVYs8RM1KF89eNrzzE/pages/0sx4sX8y3wsswrwi5ldr</a></td></tr><tr><td><h4>NVIDIA On-Demand Sessions</h4></td><td><a href="/files/2Hmbr3Jz6hBhrNABXkkN">/files/2Hmbr3Jz6hBhrNABXkkN</a></td><td><a href="/spaces/ZNGVYs8RM1KF89eNrzzE/pages/HIHMiSKQhiE1ZSryvGzt">/spaces/ZNGVYs8RM1KF89eNrzzE/pages/HIHMiSKQhiE1ZSryvGzt</a></td></tr></tbody></table>


# Overview

NVIDIA Run:ai is a GPU orchestration and optimization platform that helps organizations maximize compute utilization for AI workloads. By optimizing the use of expensive compute resources, NVIDIA Run:ai accelerates AI development cycles, and drives faster time-to-market for AI-powered innovations.

Built on Kubernetes, NVIDIA Run:ai supports dynamic GPU allocation, workload submission, workload scheduling, and resource sharing, ensuring that AI teams get the compute power they need while IT teams maintain control over infrastructure efficiency.

## Getting Started with NVIDIA Run:ai

NVIDIA Run:ai includes onboarding flows that appear directly in the user interface when logging in for the first time:

* **Administrators** are guided to install the cluster, set up SSO and invite the first research team.
* **Researchers** are guided to log in for the first time and create their initial workspace.

These onboarding flows provide a step-by-step UI experience that helps new users get set up quickly. Once onboarding is complete, administrators and AI practitioners can expand their use of the platform as shown below.

## How NVIDIA Run:ai Helps Your Organization

### For Infrastructure Administrators

NVIDIA Run:ai centralizes cluster management and optimizes infrastructure control by offering:

* [**Centralized cluster management**](/saas/infrastructure-setup/procedures/clusters) - Manage all clusters from a single platform, ensuring consistency and control across environments.
* [**Usage monitoring and capacity planning**](/saas/platform-management/monitor-performance/before-you-start) - Gain real-time and historical insights into GPU consumption across clusters to optimize resource allocation and plan future capacity needs efficiently.
* [**Policy enforcement**](/saas/platform-management/policies/policies-and-rules) - Define and enforce security and usage policies to align GPU consumption with business and compliance requirements.
* [**Enterprise-grade authentication**](/saas/infrastructure-setup/authentication/overview) - Integrate with your organization's identity provider for streamlined authentication (Single Sign On) and role-based access control (RBAC).
* **Kubernetes-native application** - Install as a Kubernetes-native application, seamlessly extending Kubernetes for native cloud experience and operational standards (install, upgrade, configure).

### For Platform Administrators

NVIDIA Run:ai simplifies AI infrastructure management by providing a structured approach to managing AI initiatives, resources, and user access. It enables platform administrators maintain control, efficiency, and scalability across their infrastructure:

* [**AI Initiative structuring and management**](/saas/platform-management/aiinitiatives/adapting-ai-initiatives#mapping-your-organization) - Map and set up AI initiatives according to your organization's structure, ensuring clear resource allocation.
* [**Centralized GPU resource management**](/saas/platform-management/aiinitiatives/adapting-ai-initiatives#mapping-your-resources) - Enable seamless sharing and pooling of GPUs across multiple users, reducing idle time and optimizing utilization.
* [**User and access control**](/saas/platform-management/aiinitiatives/adapting-ai-initiatives#assigning-users-to-projects-and-departments) - Assign users (AI practitioners, ML engineers) to specific projects and departments to manage access and enforce security policies, utilizing role-based access control (RBAC) to ensure permissions align with user roles.
* [**Workload scheduling**](/saas/platform-management/runai-scheduler/scheduling/how-the-scheduler-works) - Use scheduling to prioritize and allocate GPUs based on workload needs.
* [**Monitoring and insights**](/saas/platform-management/monitor-performance/before-you-start) - Track real-time and historical data on GPU usage to help track resource consumption and optimize costs.

### For AI Practitioners

NVIDIA Run:ai empowers data scientists and ML engineers by providing:

* [**Optimized workload scheduling**](/saas/platform-management/runai-scheduler/scheduling/how-the-scheduler-works) - Ensure high-priority jobs get GPU resources. Workloads dynamically receive resources based on demand.
* [**Fractional GPU usage**](/saas/platform-management/runai-scheduler/resource-optimization/fractions) - Request and utilize only a fraction of a GPU's memory, ensuring efficient resource allocation and leaving room for other workloads.
* [**AI initiatives lifecycle support**](/saas/workloads-in-nvidia-run-ai/introduction-to-workloads) - Run your entire AI initiatives lifecycle – Jupyter Notebooks, training jobs, and inference workloads efficiently.
* [**Interactive session** ](/saas/workloads-in-nvidia-run-ai/workload-types)- Ensure an uninterrupted experience when working on Jupyter Notebooks without taking away GPUs.
* [**Scalability for training and inference**](/saas/workloads-in-nvidia-run-ai/workload-types) - Support for distributed training across multiple GPUs and auto-scales inference workloads.
* [**Integrations**](/saas/infrastructure-setup/advanced-setup/integrations) - Integrate with popular ML frameworks - PyTorch, TensorFlow, XGBoost, Knative, Spark, Kubeflow Pipelines, Apache Airflow, Argo workloads, Ray and more.
* [**Flexible workload submission**](/saas/workloads-in-nvidia-run-ai/introduction-to-workloads) - Submit workloads using the NVIDIA Run:ai UI, API, CLI or run third-party workloads.

## NVIDIA Run:ai System Components

NVIDIA Run:ai is made up of two components both installed over a [Kubernetes](https://kubernetes.io) cluster:

* **NVIDIA Run:ai cluster** - Provides scheduling and workload management, extending Kubernetes native capabilities.
* **NVIDIA Run:ai control plane** - Provides resource management, handles workload submission and provides cluster monitoring and analytics.

<figure><img src="/files/b7zWGK87MH1w2LQNhH10" alt=""><figcaption></figcaption></figure>

### NVIDIA Run:ai Cluster

The NVIDIA Run:ai cluster is responsible for scheduling AI workloads and efficiently allocating GPU resources across users and projects:

* [**NVIDIA** **Run:ai Scheduler**](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles) - Applies AI-aware rules to efficiently schedule workloads submitted by AI practitioners.
* [**Workload management**](/saas/workloads-in-nvidia-run-ai/introduction-to-workloads) - Handles workload management which includes the researcher code running as a Kubernetes container and the system resources required to run the code, such as storage, credentials, network endpoints to access the container and so on.
* [**Kubernetes operator-based deployment**](https://kubernetes.io/docs/concepts/extend-kubernetes/operator/) - Installed as a Kubernetes Operator to automate deployment, upgrades and configuration of NVIDIA Run:ai cluster services.
* **Storage** - Supports Kubernetes-native storage using [Storage Classes](https://kubernetes.io/docs/concepts/storage/storage-classes/), allowing organizations to bring their own storage solutions. Additionally, it also integrates with [external storage solutions](/saas/workloads-in-nvidia-run-ai/assets/datasources) such as Git, S3, and NFS to support various data requirements.
* **Secured communication** - Uses an outbound-only, secured (SSL) connection to synchronize with the NVIDIA Run:ai control plane.
* **Private** - NVIDIA Run:ai only synchronizes metadata and operational metrics (e.g., workloads, nodes) with the control plane. No proprietary data, model artifacts, or user data sets are ever transmitted, ensuring full data privacy and security.

### NVIDIA Run:ai Control Plane

The NVIDIA Run:ai control plane provides a centralized management interface for organizations to oversee their GPU infrastructure across multiple locations/subnets, accessible via Web UI, [API](https://run-ai-docs.nvidia.com/api/) and [CLI](/saas/reference/cli/install-cli). The control plane can be deployed on the cloud or on-premise for organizations that require local control over their infrastructure (self-hosted).

* [**Multi-cluster management**](/saas/infrastructure-setup/procedures/clusters) - Manages multiple NVIDIA Run:ai clusters for a single tenant across different locations and subnets from a single unified interface.
* [**Resource and access management**](/saas/platform-management/aiinitiatives/adapting-ai-initiatives) - Allows administrators to define Projects, Departments and user roles, enforcing policies for fair resource distribution.
* [**Workload submission and monitoring**](/saas/workloads-in-nvidia-run-ai/workloads) - Allows teams to submit workloads, track usage, and monitor GPU performance in real time.

## Installation types <a href="#installation-types" id="installation-types"></a>

There are two main installation options:

| Installation Type | Description                                                                                                                                                                                                                                                                                                            |
| ----------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| SaaS              | <p>NVIDIA Run:ai is installed on the customer's data science GPU clusters. The cluster connects to the NVIDIA Run:ai control plane on the cloud (https\://<code>\<tenant-name></code>.run.ai).<br>With this installation, the cluster requires an <strong>outbound</strong> connection to the NVIDIA Run:ai cloud.</p> |
| Self-hosted       | The NVIDIA Run:ai control plane is also installed in the customer's data center                                                                                                                                                                                                                                        |

<figure><img src="/files/CMVXlbGP99KL7zlgq5dQ" alt=""><figcaption></figcaption></figure>


# What's New

The what's new provides transparency into the latest changes and improvements to NVIDIA Run:ai’s SaaS platform. The updates include new features, optimizations, and fixes aimed at improving performance and user experience.

{% hint style="info" %}
**Important**

For a complete list of deprecations, see [Deprecation notifications](#deprecation-notifications). Deprecated features, APIs, and capabilities remain available for **six months** from the time of the deprecation notice, after which they may be removed.
{% endhint %}

## Gradual Rollout

SaaS features and bug fixes are gradually rolled out to customers to ensure a smooth transition and minimize any potential disruption. SaaS releases follow a scheduled rollout cadence, typically every two weeks, allowing us to introduce new functionalities in a controlled and predictable manner. **All customers receive the changes within 10 days of the initial release.**

In contrast, hotfixes are deployed as needed to address urgent issues and are released immediately to ensure the stability and security of the service.

## DGX Cloud

Certain features are first made available in fully managed cloud-based deployments provisioned through DGX Cloud. These features are labeled as `DGX Cloud only` and will become available to all customers in future releases.

## Feature Life Cycle

NVIDIA Run:ai uses life cycle labels to indicate the maturity and stability of features across releases:

* <mark style="color:green;">`Experimental`</mark> - This feature is in early development. It may not be stable and could be removed or changed significantly in future versions. Use with caution.
* <mark style="color:orange;">`Beta`</mark> - This feature is still being developed for official release in a future version and may have some limitations. Use with caution.
* `Legacy` - This feature is scheduled to be removed in future versions. We recommend using alternatives if available. Use only if necessary.

## July 2026 Releases

### July 05

#### Product Enhancements

* **Enhanced workload structure and inspection** - Supported workload types include enhanced visibility into workload structure and details directly from NVIDIA Run:ai. `From cluster v2.26 onward`
  * **Workload element YAML** - Displays the live Kubernetes custom resource for a workload, workload element, or pod, including the current spec, status, and metadata. The YAML can be copied or downloaded for further investigation. See [Structure](/saas/workloads-in-nvidia-run-ai/workloads#structure) for more details.
  * **Workload details for supported workload types** - Extends the workload Details experience to supported workload types, providing access to configuration, runtime information, resources, and workload structure. See [Show/Hide Details](/saas/workloads-in-nvidia-run-ai/workloads#show-hide-details) for more details.
* **Private NGC repository support for AI applications** - AI applications support private NGC Helm repositories. When a private repository is selected as the chart source, users are prompted to provide an NGC API key credential to authenticate with the repository before browsing and selecting charts. See [AI Applications](/saas/ai-applications/ai-applications) for more details. `From cluster v2.26 onward`
* **Workloads v2 API enhancements** - The Workloads v2 API supports PersistentVolumeClaims and user credentials as part of the workload specification, allowing these resources to be created and managed alongside workloads. See [Workloads v2](https://run-ai-docs.nvidia.com/api/workloads/workloads-v2) API for more details. `From cluster v2.26 onward`
  * **PersistentVolumeClaims (PVCs)** - Create PersistentVolumeClaims as part of the workload definition, enabling workloads to provision and use storage without requiring pre-existing PVCs.
  * **User credentials** - Reference user credentials directly from workload definitions without exposing secret values in the workload manifest.
* **Suspend and resume workloads** - NVIDIA Run:ai supports suspend and resume operations for any supported workload type that defines suspend behavior, extending workload lifecycle management beyond native NVIDIA Run:ai workloads. Users can suspend and resume supported workloads through both the UI and API, providing a consistent management experience across workload types. Supported workload types include CronJob, Job, PyTorchJob, RayJob, and TFJob. See [Workloads](/saas/workloads-in-nvidia-run-ai/workloads) for more details. `From cluster v2.26 onward`

{% hint style="warning" %}
**Known Kubernetes Issue**

A known Kubernetes issue can cause batch job workload types (such as Job, PyTorchJob, and RayJob) to get stuck in the **Resuming** state. The issue is fixed in Kubernetes v1.32.9, v1.33.6, v1.34.2, and v1.35.0 (or newer). Upgrade to a fixed Kubernetes release to resolve this.
{% endhint %}

* **NUMA-aware scheduling for node pools** - Administrators can now enable NUMA-aware scheduling per node pool. When enabled, the Scheduler places workloads so that GPU, CPU, and CPU memory are allocated within the same NUMA node where possible, minimizing cross-NUMA fragmentation and improving performance on multi-NUMA servers such as DGX H100. When disabled, the Scheduler places workloads using aggregate node capacity only, without considering NUMA topology. Requires NFD deployed with the topology updater enabled and a Topology Manager policy other than `none` on the pool's nodes. Disabled by default. See [Configuring NUMA-Aware Scheduling](/saas/platform-management/aiinitiatives/resources/numa-aware-scheduling) for more details. `From cluster v2.26 onward`
* **GPU device-level placement strategy for node pools** - Node pools support a separate device placement strategy for GPU pods in addition to the node placement strategy. The device placement strategy (Bin-pack or Spread) determines how GPU workloads are distributed across GPU devices within each node. This helps fractional distributed workloads - for example, bin-pack at the node level to use fewer nodes while spreading at the device level so pods land on different GPUs within a node. See [Node pools](/saas/platform-management/aiinitiatives/resources/node-pools#adding-a-new-node-pool) for more details. `From cluster v2.26 onward`
* **Grove topology integration (API)** - NVIDIA Run:ai network topologies can be linked to a corresponding Grove ClusterTopology. Once linked, NVIDIA Run:ai surfaces the Grove topology name and levels through the [Network Topologies](https://run-ai-docs.nvidia.com/api/organizations/network-topologies) API, allowing users to identify the correct topology for their target node pool without needing to understand the underlying Grove topology configuration. See [Accelerating workloads with network topology-aware scheduling](/saas/platform-management/aiinitiatives/resources/topology-aware-scheduling#topology-aware-scheduling-for-dynamo-over-grove) for more details. `From cluster v2.26 onward`
* **Workload policies for any supported workload type** - AWorkload policies support any supported workload type, extending policy governance beyond native NVIDIA Run:ai workloads. NVIDIA Run:ai provides built-in policy mappings for a wide range of supported workload types, so administrators can create policies immediately without additional configuration. Policies can define workload requirements, constraints, defaults, and enforcement actions. See [Supported workload types policies](/saas/platform-management/policies/supported-workload-type-policies) for more details. To extend policy support to additional workload types introduced to the platform, administrators can define mappings using the Policy Mappings API. See [Workload types policy mapping](/saas/platform-management/policies/supported-workload-type-policies/workload-type-policy-mapping) for more details. <mark style="color:green;">`Experimental`</mark> `From cluster v2.26 onward`
* **Kubernetes 1.36 support** - NVIDIA Run:ai now supports Kubernetes version 1.36. `From cluster v2.26 onward`

#### Resolved Bugs

<table><thead><tr><th width="253.79296875">ID</th><th>Description</th></tr></thead><tbody><tr><td>RUN-40942</td><td>Fixed an issue where the applications list endpoint returned an empty creation date.</td></tr><tr><td>RUN-40506</td><td>Fixed an issue where the distributed workload submission form incorrectly set GPU fields for the master replica when the MPI framework was selected, resulting in a submission error.</td></tr><tr><td>RUN-40357</td><td>Fixed an issue where the Users minimal read API returned incorrect search results by including results matched on the creator field instead of filtering by the user's name or username.</td></tr><tr><td>RUN-40354</td><td>Fixed an issue where a workload policy with <code>canAdd: false</code> for all storage data source types still displayed all storage options in the data source dropdown.</td></tr><tr><td>RUN-39865</td><td>Fixed an issue where the <code>runai nodepool list</code> command crashed when the cluster had nodes excluded using the <code>managedNodePoolNames</code> configuration.</td></tr><tr><td>RUN-39472</td><td>Fixed an issue where deleting a DataVolume left stale labels on the origin PVC, preventing it from being re-created.</td></tr><tr><td>RUN-38957</td><td>Fixed an issue where NIMService workloads using fractional GPU annotations did not display GPU allocation in the NVIDIA Run:ai UI.</td></tr><tr><td>RUN-38857</td><td>Fixed an issue where the NodePort address was not correctly displayed for workloads configured with the NodePort networking option.</td></tr><tr><td>RUN-37078</td><td>Fixed an issue where the access rules dialog did not display existing users and service accounts in the selection component on the first click.</td></tr><tr><td>RUN-40745</td><td>Fixed a security vulnerability related to GHSA-x744-4wpc-v9h2 with severity HIGH.</td></tr><tr><td>RUN-40448</td><td>Fixed a security vulnerability related to CVE-2026-42151 with severity HIGH.</td></tr><tr><td>RUN-39873</td><td>Fixed security vulnerabilities related to CVE-2026-42586 and CVE-2026-42579 with severity HIGH.</td></tr></tbody></table>

## June 2026 Releases

### June 09

#### Product Enhancements

* **Full name display for users** - Users now appear by their full name next to their email, in the format **Full Name (email)**, wherever they show up in the platform. First and last names can be set for local users when they are created or edited; names for SSO users are sourced from the identity provider, and the email is shown on its own when no name is available. A new User Minimal Read Access permission (`users-minimal`) provides the minimal access needed to display user names. It is included by default in some predefined roles and can be added to custom roles. See [Users](/saas/infrastructure-setup/authentication/users) and [Roles](/saas/infrastructure-setup/authentication/roles) for more details.
* **Workload connection endpoints in the Structure tab** - The Structure tab in the workload Details pane now displays a workload's connection endpoints, making it easier to find and access the tools and services a workload exposes. See [Structure](/saas/workloads-in-nvidia-run-ai/workloads#structure) for more details. `From cluster v2.25 onward`

#### Resolved Bugs

<table><thead><tr><th width="253.79296875">ID</th><th>Description</th></tr></thead><tbody><tr><td>RUN-39755</td><td>Fixed an issue where the workload Connections dialog displayed the node IP instead of the user-defined tool name.</td></tr><tr><td>RUN-38842</td><td>Fixed an issue where a missing runai-go-operator-tls-secret prevented Helm upgrades from completing.</td></tr><tr><td>RUN-39222</td><td>Fixed an issue where the CLI accepted unsupported storage flags for distributed inference submissions.</td></tr><tr><td>RUN-39453</td><td>Fixed a security vulnerability related to CVE-2026-42945 with severity CRITICAL.</td></tr><tr><td>RUN-39456</td><td>Fixed a security vulnerability related to CVE-2026-42154 with severity HIGH.</td></tr><tr><td>RUN-39868</td><td>Fixed security vulnerabilities related to CVE-2026-39832, CVE-2026-46595, CVE-2026-39831, CVE-2026-39828, CVE-2026-39833, and CVE-2026-42508 with severity HIGH.</td></tr><tr><td>RUN-39869</td><td>Fixed a security vulnerability related to CVE-2025-14459 with severity HIGH.</td></tr><tr><td>RUN-39311</td><td>Fixed a security vulnerability related to GHSA-389r-gv7p-r3rp with severity HIGH.</td></tr><tr><td>RUN-39203</td><td>Fixed a security vulnerability related to CVE-2026-42501 with severity HIGH.</td></tr><tr><td>RUN-39658</td><td>Fixed a security vulnerability related to GHSA-fqw6-gf59-qr4w with severity HIGH.</td></tr><tr><td>RUN-39874</td><td>Fixed a security vulnerability related to CVE-2026-39852 with severity HIGH.</td></tr></tbody></table>

## May 2026 Releases

### May 25

#### Product Enhancements

* **Access rule transition for legacy roles** - Due to the deprecation of some roles, referred to as legacy roles, access rules using them should be transitioned to supported roles. The Access Rules page now surfaces a banner when rules still use legacy roles, and a new dialog lets admins map each legacy role to its replacement. Selected rules are recreated under the new role with the same subject and scope, and the legacy rules are deleted. See [Access Rules](/saas/infrastructure-setup/authentication/accessrules) for more details. `From cluster v2.24 onward`
* **Workload element metrics** - The Structure tab now lets you view metrics per workload element. In the expanded view, selecting a workload element scopes the Metrics tab to that element. See [Structure](/saas/workloads-in-nvidia-run-ai/workloads#structure) for more details. `From cluster v2.25 onward`

#### Resolved Bugs

<table><thead><tr><th width="253.79296875">ID</th><th>Description</th></tr></thead><tbody><tr><td>RUN-39479</td><td>Fixed an issue where some filters on the Clusters table did not work.</td></tr><tr><td>RUN-39038</td><td>Fixed an issue where the GPU profiling metrics toggle in General Settings > Analytics displayed without a label.</td></tr><tr><td>RUN-39194</td><td>Fixed a security vulnerability related to CVE-2026-41080 with severity HIGH.</td></tr><tr><td>RUN-39159</td><td>Fixed a security vulnerability related to CVE-2026-33811 with severity HIGH.</td></tr><tr><td>RUN-39014</td><td>Fixed a security vulnerability related to GHSA-78h2-9frx-2jm8 with severity HIGH.</td></tr></tbody></table>

### May 14

#### Product Enhancements

* **Actual topology placement per workload element (API)** - For multi-component workloads (such as Dynamo over Grove), the Workloads API now returns the actual topology placement of each workload element, so you can verify where each element's pods landed in the network hierarchy after scheduling. Workload-level actual placement is also now returned for any distributed workload whose node pool is attached to a topology, regardless of whether topology constraints were applied. See [Accelerating Workloads with Network Topology-Aware Scheduling](/saas/platform-management/aiinitiatives/resources/topology-aware-scheduling#topology-constraints-visibility) for more details. `From cluster v2.25 onward`
* **Permission-aware General settings** - The General settings now reflects the user's settings permissions. The page is visible only when the user has at least one settings permission with CRUD (Create, Read, Update, Delete) access; if all of the user's settings permissions are read-only, the page is not visible (matching the existing behavior). See [General Settings](/saas/settings/general-settings) and [Roles](/saas/infrastructure-setup/authentication/roles) for more details.
* **Workload structure expanded view** - The Structure tab in the workload Details pane now supports an expanded view, giving you a larger canvas to inspect a workload's hierarchy of workload elements, resource allocations, and pod status. This makes it easier to trace dependencies in workloads with many child elements - for example, multi-service Dynamo or LeaderWorkerSet deployments. See [Structure](/saas/workloads-in-nvidia-run-ai/workloads#structure) for more details. `From cluster v2.25 onward`

#### Resolved Bugs

<table><thead><tr><th width="253.79296875">ID</th><th>Description</th></tr></thead><tbody><tr><td>RUN-37894</td><td>Fixed a security vulnerability related to GHSA-xw7x-h9fj-p2c7 with severity CRITICAL.</td></tr><tr><td>RUN-38147</td><td>Fixed a security vulnerability related to GHSA-78h2-9frx-2jm8 with severity HIGH.</td></tr><tr><td>RUN-38467</td><td>Fixed a security vulnerability related to GHSA-pc3f-x583-g7j2 with severity HIGH.</td></tr><tr><td>RUN-38975</td><td>Fixed a security vulnerability related to GHSA-vmg3-7v43-9g23 with severity HIGH.</td></tr><tr><td>RUN-38895</td><td>Fixed a security vulnerability related to CVE-2026-4878 with severity HIGH.</td></tr><tr><td>RUN-38667</td><td>Fixed a security vulnerability related to GHSA-5jv8-h7qh-rf5p with severity HIGH.</td></tr><tr><td>RUN-38680</td><td>Fixed a security vulnerability related to GHSA-mh2q-q3fh-2475 with severity HIGH.</td></tr><tr><td>RUN-38224</td><td>Fixed an issue where NIMService external URLs were intermittently unavailable under load.</td></tr><tr><td>RUN-38670</td><td>Fixed a security vulnerability related to CVE-2026-22016 with severity HIGH.</td></tr><tr><td>RUN-38604</td><td>Fixed an issue where the Event history "Download to CSV/JSON" exported only the current page instead of all records.</td></tr><tr><td>RUN-38641</td><td>Fixed an issue where the Event history "Subject Type" filter did not always return matching results.</td></tr><tr><td>RUN-38774</td><td>Fixed an issue where the Terminal tab blocked connections to a running pod when the workload phase was not Running.</td></tr><tr><td>RUN-38850</td><td>Fixed an issue where the Project telemetry endpoint returned a 500 error for project-scoped users.</td></tr></tbody></table>

## April 2026 Releases

### April 28

* **Deploy AI applications directly from NGC catalog** - Blueprints (and other Helm charts) from the NVIDIA NGC catalog or via direct URL can be deployed as AI applications through the UI and API, without requiring direct cluster access. Deploying them as AI applications enables fast assembly of the building blocks that power agentic pipelines in a single workflow. You can override Helm values before deployment (for example, request GPU fractions), enabling flexible configuration and faster deployment of complex, multi-component AI applications. See [AI applications](/saas/ai-applications/ai-applications) for more details. <mark style="color:green;">`Experimental`</mark> `From cluster v2.25 onward`
* **Topology placement visibility for workloads** - NVIDIA Run:ai exposes the actual topology placement of a workload, making it easy to validate scheduling decisions and troubleshoot performance issues for multi-node workloads. For each workload, you can see the topology name, the actual placement (topology level and value), the requested constraints, and whether those constraints were met. This information is available for both native and supported workload types via the [Workloads](https://run-ai-docs.nvidia.com/api/workloads/workloads) API, and for native workloads in the workload [Details](/saas/workloads-in-nvidia-run-ai/workloads#show-hide-details) view in the UI. See [Accelerating workloads with network topology-aware scheduling](/saas/platform-management/aiinitiatives/resources/topology-aware-scheduling) for more details. `From cluster v2.25 onward`
* **Enhanced workload structure visibility** - NVIDIA Run:ai now surfaces the full internal structure of [supported workload types](/saas/workloads-in-nvidia-run-ai/submit-via-yaml), allowing users to navigate from the top-level workload down to individual elements and pods. For each element, you can inspect its name, type, and parent-child relationships both via the API and in the UI through a new hierarchical **Structure** tab in the workload [Details](/saas/workloads-in-nvidia-run-ai/workloads#show-hide-details) panel. Administrators can use Karta to define the structure of additional workload types. This enables newly introduced workloads to expose the same hierarchy and relationships, making them visible and consistent across the API and UI. See [Defining a Karta](/saas/workloads-in-nvidia-run-ai/workload-types/extending-workload-support/defining-a-karta) for more details. `From cluster v2.25 onward`
* **DRA support for GPU devices** - NVIDIA Run:ai supports scheduling supported workload types submitted [via YAML](/saas/workloads-in-nvidia-run-ai/submit-via-yaml) using Kubernetes Dynamic Resource Allocation (DRA) ResourceClaims for GPU devices, in addition to the existing extended resources method (`nvidia.com/gpus`). This enables more flexible and expressive GPU resource requests, improving scheduling accuracy and laying the foundation for advanced capabilities. DRA-based workloads are fully supported across the NVIDIA Run:ai UI, API, and CLI. To avoid conflicts when mixing DRA and extended resources, it is recommended to use separate node pools for DRA workloads. See [Dynamic resource allocation (DRA)](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles#dynamic-resource-allocation-dra) for more details. `From cluster v2.25 onward`
* **Automatic external endpoint discovery for workloads and AI applications** - After deploying a workload, its externally reachable endpoints are automatically discovered and displayed in the **Connections** column of the [Workloads](/saas/workloads-in-nvidia-run-ai/workloads#connections-associated-with-the-workload) and [AI Applications](/saas/ai-applications/ai-applications#connections-associated-with-the-ai-application) tables. Endpoints are also exposed as part of the workload object via the [Workloads](https://run-ai-docs.nvidia.com/api/workloads/workloads) API. For AI applications composed of multiple workloads, endpoints from all workloads are aggregated and can be viewed in the UI or queried at the application level via the [AI Applications](https://run-ai-docs.nvidia.com/api/ai-applications/ai-applications) API. Endpoint state is kept in sync automatically as networking resources are created, updated, or deleted. `From cluster v2.25 onward`
* **Hierarchical topology-aware scheduling for Dynamo workloads** - Dynamo workloads support hierarchical, multi-component topology-aware scheduling using Grove topology definitions. Administrators define and link Grove and NVIDIA Run:ai topologies, which are then exposed and used to apply topology constraints at both the workload (Deployment) and component (Service) levels. The NVIDIA Run:ai scheduler enforces these constraints, enabling precise placement across clusters with multiple topologies. This improves performance and avoids over-constrained scheduling compared to previous workload-level topology constraints. See [Accelerating workloads with network topology-aware scheduling](/saas/platform-management/aiinitiatives/resources/topology-aware-scheduling#topology-aware-scheduling-for-dynamo-over-grove) for more details. `From cluster v2.25 onward`
* **Endpoint access control for NIM services API** - NVIDIA NIM services deployed via the NIM Operator support endpoint access control for NVIDIA Run:ai users, groups, and service accounts through the API. This enables secure access to inference endpoints and aligns NIM services with enterprise security requirements. See [NVIDIA NIM](https://run-ai-docs.nvidia.com/api/workloads/nvidia-nim) API for more details. `From cluster v2.25 onward`
* **CLI support for NVIDIA NIM services** - NVIDIA NIM services deployed via the NIM Operator can be managed directly from the NVIDIA Run:ai CLI, bringing a full command-line experience without requiring API access. See [CLI command reference](/saas/reference/cli/runai/runai-inference-nim) for more details. `From cluster v2.25 onward`
  * Supports the full service lifecycle: submit, update, list, delete, exec, logs, port-forward, and describe.
  * Applies to existing NIM capabilities including autoscaling, fractional GPU, multi-node deployments, and Multi-LLM workloads.
* **Notifications API and Slack integration updates** - The notifications API was redesigned to provide a clearer and more structured model for managing notification state and channels. Email and Slack integrations are handled through dedicated APIs, replacing the previous unified notification channel endpoints. Slack integration is now fully supported through the UI across all deployment types, with the platform managing app creation, connection, and permissions. Legacy notification channel and Slack-specific endpoints are [deprecated](#api-deprecation-notifications) as part of this update. `From cluster v2.25 onward`
* **NVIDIA Run:ai artifacts now available on NGC** - NVIDIA Run:ai artifacts are now published to NVIDIA NGC, including Helm charts, container images, CLI packages, air-gapped packages, and documentation. This provides a unified experience with other NVIDIA products, allowing customers to access and manage NVIDIA Run:ai software using the same NGC tools and credentials. JFrog support is now [deprecated](#jfrog-artifacts). Existing deployments will continue to work during the deprecation period, and customers are encouraged to migrate to NGC. See [Installation](/saas/getting-started/installation) for more details. `From cluster v2.25 onward`
* **Host-based routing is now the default** - To support tools such as RStudio and Visual Studio Code without requiring additional configuration after installation, host-based routing is enabled by default. Workload URLs are exposed as subdomains, allowing workloads to run at the root path and avoiding file path issues. Kubernetes clusters require additional setup as part of system requirements configuration. OpenShift clusters require no additional configuration. Clusters upgrading from path-based routing are not affected. For setup instructions, see [System requirements](/saas/getting-started/installation/install-using-helm/system-requirements#host-based-routing-default). For more details, see [External access to containers](/saas/infrastructure-setup/advanced-setup/container-access/external-access-to-containers). `From cluster v2.25 onward`
* **Kubernetes Gateway API support** - NVIDIA Run:ai supports the Kubernetes Gateway API as a modern routing infrastructure, providing an alternative to Kubernetes Ingress. The feature is available as an opt-in option in v2.25 and can run alongside Ingress to enable zero-downtime migration. Ingress remains the default, with Gateway API expected to become the default in a future release. See [Kubernetes Gateway API](/saas/infrastructure-setup/advanced-setup/kubernetes-gateway-api) for more details. `From cluster v2.25 onward`
* **Vertical Pod Autoscaling (VPA) for cluster services** - NVIDIA Run:ai now supports configuring Vertical Pod Autoscaling (VPA) on cluster-level services to automatically right-size CPU and memory resources based on observed usage. VPA can be configured globally, per scaling group, or per component via `runaiconfig`, with multiple update modes including `Off` (recommendations only), `Initial`, `Auto`, and `InPlaceOrRecreate`. See [Vertical pod autoscaling](/saas/infrastructure-setup/advanced-setup/vpa) for more details. `From cluster v2.25 onward`
* **Component version updates** - NVIDIA Run:ai now supports the following. See the [Support matrix](/saas/getting-started/installation/support-matrix) for more details. `From cluster v2.25 onward`
  * OpenShift version 4.21
  * GPU Operator version 26.3
  * Network Operator 26.1
  * DRA driver 25.12
  * Support for Kubernetes version 1.32 and OpenShift version 4.17 has been removed.

#### Resolved Bugs

<table><thead><tr><th width="253.79296875">ID</th><th>Description</th></tr></thead><tbody><tr><td>RUN-38510</td><td>Fixed a security vulnerability related to GHSA-9jj7-4m8r-rfcm with severity HIGH.</td></tr><tr><td>RUN-38474</td><td>Fixed a security vulnerability related to CVE-2026-27144 with severity HIGH.</td></tr><tr><td>RUN-38314</td><td>Fixed a security vulnerability related to CVE-2026-33810 with severity HIGH.</td></tr><tr><td>RUN-38290</td><td>Fixed a security vulnerability related to CVE-2026-32281 with severity HIGH.</td></tr><tr><td>RUN-38240</td><td>Fixed a security vulnerability related to GHSA-6v2p-p543-phr9 with severity HIGH.</td></tr><tr><td>RUN-38175</td><td>Fixed a security vulnerability related to CVE-2026-32280 with severity HIGH.</td></tr><tr><td>RUN-38094</td><td>Fixed a security vulnerability related to CVE-2026-27654 with severity HIGH.</td></tr><tr><td>RUN-38084</td><td>Fixed a security vulnerability related to GHSA-hfvc-g4fc-pqhx with severity HIGH.</td></tr><tr><td>RUN-37701</td><td>Fixed a security vulnerability related to GHSA-p77j-4mvh-x3m3 with severity CRITICAL.</td></tr><tr><td>RUN-36615</td><td>Fixed a security vulnerability related to CVE-2024-12797 with severity HIGH.</td></tr><tr><td>RUN-34128</td><td>Fixed a security vulnerability in the OAuth/OpenID login flow.</td></tr><tr><td>RUN-38260</td><td>Fixed an issue where the Admin > Event history page returned a 500 error when filtering by Event ID.</td></tr><tr><td>RUN-38030</td><td>Fixed an issue where the CLI v2 <code>--preemptible</code> flag did not set preemptibility on workspace workloads.</td></tr><tr><td>RUN-38029</td><td>Fixed an issue where the <code>GET /api/v1/workloads/pods</code> endpoint ignored project scope for user API tokens.</td></tr><tr><td>RUN-37972</td><td>Fixed an issue where <code>ingressClass</code> was missing from the Minimal Cluster resource.</td></tr><tr><td>RUN-36456</td><td>Fixed an issue where enabling <code>enableWorkloadOwnershipProtection</code> prevented the Workload Overseer from enforcing project scheduling rules such as Idle GPU timeout and Workload Duration limits.</td></tr><tr><td>RUN-34639</td><td>Fixed an issue where the UI displayed impossible Free GPU values (allocated + free exceeding the node's total capacity) on nodes with fractional GPU allocations.</td></tr></tbody></table>

### April 15

#### Product Enhancements

* **MNNVL acceleration for YAML-based workload submission** - When submitting YAML-based workloads, you can specify whether Multi-Node NVLink (MNNVL) acceleration is required on a per-workload basis. When required, workloads are scheduled only on MNNVL-capable nodes; otherwise, they can run on any compatible nodes. This provides more precise control over placement and performance for multi-GPU workloads. See [Submit supported workload types via YAML](/saas/workloads-in-nvidia-run-ai/submit-via-yaml) for more details. `From cluster v2.24 onward`
* **Distributed inference deployment from the UI** - AI practitioners can now submit and manage distributed (multi-node) inference workloads directly from the NVIDIA Run:ai UI. This enables deployment of large models that exceed the capacity of a single node. Each distributed replica consists of a leader and multiple workers, both configured as part of the submission form. Distributed mode is supported for custom inference only. See [Deploy a distributed inference workload](/saas/workloads-in-nvidia-run-ai/using-inference/distributed-inference) for more details. `From cluster v2.22 onward`
* **Deployment of models from Hugging Face** - Models from Hugging Face can now be deployed directly through the NIM serving path using Multi-LLM NIM — a single container that enables deployment of a broad range of models. Previously, only NGC catalog models with dedicated NIM images could use the NIM serving flow; Hugging Face models were restricted to generic serving options. This removes the one-image-per-model dependency and allows a wider range of models to be served through the NIM path, directly from the Run:ai UI. See [Deploy inference workloads with NVIDIA NIM](/saas/workloads-in-nvidia-run-ai/using-inference/nim-inference) for more details.
* **Model descriptions in NIM model selection** - The NIM model selection dropdown displays a short description beneath each model name, sourced from the NGC model card. AI practitioners can evaluate available models directly in the NVIDIA Run:ai UI without needing to look them up externally.
* **Trending model indicators in the Hugging Face model catalog** - The Hugging Face model catalog highlights trending models directly in the model selection dropdown. The top 20 trending models are marked with a trending indicator, making it easier to discover popular and emerging models alongside the most downloaded ones.
* **User credentials support in the CLI** - Credentials configured in User settings (My Credentials) can now be managed through the CLI including creating, listing, and deleting credentials. Supported credential types include Generic, Docker registry, and NGC API keys. Credentials can also be used during workload submission as image pull secrets or environment variables, enabling a more streamlined and consistent submission workflow across the CLI, UI, and API. See [CLI command reference](/saas/reference/cli/runai/runai-my-credential) for more details. `From cluster v2.22 onward`
* **Extended CLI support for YAML-based workloads -** The CLI supports additional operations for supported workload types submitted via YAML, including `runai workload-type list` and `runai workload-type describe`. This improves visibility and control over supported workload types without requiring access to the UI or direct interaction with Kubernetes resources. See [CLI command reference](/saas/reference/cli/runai/runai-workload-type) for more details. `From cluster v2.23 onward`
* **Time-based fairshare UI parameters** - Administrators can now configure key time-based fairshare parameters directly in the UI, including Historical usage weight and Historical usage window. Previously, these settings were only available through the node pool API. This makes it easier to tune how resource fairness is calculated based on historical and current usage, improving control over fairshare behavior without requiring API-level configuration. See [Node pools](/saas/platform-management/aiinitiatives/resources/node-pools#adding-a-new-node-pool) for more details. `From cluster v2.24 onward`

#### Resolved Bugs

<table><thead><tr><th width="253.79296875">ID</th><th>Description</th></tr></thead><tbody><tr><td>RUN-37482</td><td>Fixed an issue where workspaces using the TensorBoard environment preset failed to start due to the tensorboard executable not being found in $PATH.</td></tr><tr><td>RUN-37170</td><td>Fixed a security vulnerability related to GHSA-23c5-xmqv-rm74 with severity HIGH.</td></tr><tr><td>RUN-37932</td><td>Fixed a security vulnerability related to GHSA-p77j-4mvh-x3m3 with severity HIGH.</td></tr><tr><td>RUN-38055</td><td>Fixed an issue where the Access Rules API accepted invalid <code>subjectType</code> values without returning an error.</td></tr><tr><td>RUN-37565</td><td>Fixed an issue where the emptyDir storage option was not applied to inference workloads even when selected.</td></tr><tr><td>RUN-36678</td><td>Fixed an issue where the workload count displayed in the General settings was incorrect when filtering by enabled or disabled.</td></tr><tr><td>RUN-37959</td><td>Fixed an issue where the automatic topology constraint applied to a workload targeted the farthest topology level instead of the closest.</td></tr><tr><td>RUN-36615</td><td>Fixed a security vulnerability related to CVE-2024-12797 with severity HIGH.</td></tr></tbody></table>

## March 2026 Releases

### March 22

#### Product Enhancements

* **NGC SaaS tenant creation** - NVIDIA Run:ai SaaS tenants can now be provisioned directly from the NVIDIA NGC portal through a new workflow. Provisioning is initiated by NGC organization admins, with a direct link to the NVIDIA Run:ai tenant console provided upon completion. See [Create and access the control plane tenant](/saas/getting-started/installation/install-using-helm/control-plane-tenant) for more details.
* **Scope-aware filtering for over-time widgets** - Over-time widgets in the Overview dashboard now support scope-aware filtering by department or project. Users without full cluster permissions can select a single department or project within their allowed scope, and the widgets update accordingly. If only one department or project is available to the user, it is selected automatically. The Overview dashboard is now fully aware of the user’s scope and displays only relevant data.

#### Resolved Bugs

<table><thead><tr><th width="253.79296875">ID</th><th>Description</th></tr></thead><tbody><tr><td>RUN-37375</td><td>Fixed a security vulnerability related to GHSA-m297-3jv9-m927 with severity HIGH.</td></tr><tr><td>RUN-37341</td><td>Fixed a security vulnerability related to CVE-2025-61732 with severity HIGH.</td></tr><tr><td>RUN-37278</td><td>Fixed a security vulnerability related to CVE-2024-1013 with severity HIGH.</td></tr><tr><td>RUN-37174</td><td>Fixed a security vulnerability related to GHSA-72hv-8253-57qq with severity HIGH.</td></tr><tr><td>RUN-37169</td><td>Fixed a security vulnerability related to GHSA-5rq4-664w-9x2c with severity HIGH.</td></tr><tr><td>RUN-37168</td><td>Fixed a security vulnerability related to GHSA-9h8m-3fm2-qjrq with severity HIGH.</td></tr><tr><td>RUN-37497</td><td>Fixed a security vulnerability related to CVE-2026-27142 with severity HIGH.</td></tr><tr><td>RUN-36413</td><td>Fixed a security vulnerability related to CVE-2024-41110 with severity HIGH.</td></tr><tr><td>RUN-36407</td><td>Fixed an issue where workspace workload submissions intermittently failed with a "Workload failed due to a Network issue" error.</td></tr><tr><td>RUN-34564</td><td>Fixed an issue where updating a project with node type names in the payload did not return the node type names in the 200 success response.</td></tr></tbody></table>

### March 08

#### Product Enhancements

* **Sort by pod name in the Pods modal** - The Pods modal on both the Nodes view and the Workloads view now supports sorting by pod name.
* **Assign access rules during user and service account creation** - When creating a local [user](/saas/infrastructure-setup/authentication/users#creating-a-local-user) or [service account](/saas/infrastructure-setup/authentication/service-accounts#creating-a-service-account), you can now optionally assign access rules directly within the creation flow. This step is only shown to users with permissions to manage access rules.
* **View logs from previous container instances** - The [Logs](/saas/workloads-in-nvidia-run-ai/workloads#logs) view now includes a **Container's logs from** dropdown, allowing you to switch between logs from the current running container instance and the previous one. This makes it easier to investigate restarts and container failures without leaving the workload view.
* **Extended metrics time range for workloads** - The Metrics view now includes a **Since first run** option in the time range selector, displaying metrics from the first time a workload transitioned to the Running phase until the current time.
* **Early removal of legacy Grafana dashboards** - The legacy Grafana-based dashboards have been removed ahead of the originally planned deprecation timeline (April 2026). Security vulnerabilities (CVEs) were identified in the Grafana version in use that could not be remediated through an upgrade at this stage. The new dashboards provide equivalent functionality and are the recommended replacement.

#### Resolved Bugs

<table><thead><tr><th width="253.79296875">ID</th><th>Description</th></tr></thead><tbody><tr><td>RUN-37113</td><td>Fixed an issue where image strings that included a port number in the registry URL were not parsed correctly.</td></tr><tr><td>RUN-37060</td><td>Fixed an issue where the NVLink total bytes per pod metric was labeled with GPU metrics labels instead of the expected pod labels.</td></tr><tr><td>RUN-36734</td><td>Fixed an issue where the Analytics table displayed incorrect GPU Compute Utilization values for Training and Interactive workloads.</td></tr><tr><td>RUN-36598</td><td>Fixed an issue where department data was not synced to the cluster, affecting both department creation and updates.</td></tr><tr><td>RUN-36560</td><td>Fixed an issue where the Connect button did not open the workspace URL for workloads submitted through YAML.</td></tr><tr><td>RUN-36555</td><td>Fixed a security vulnerability related to CVE-2024-56171 with severity HIGH.</td></tr><tr><td>RUN-36443</td><td>Fixed an issue where the dashboard returned a 500 error instead of an informative error message.</td></tr><tr><td>RUN-36370</td><td>Fixed an issue where NIM and HuggingFace inference templates failed to submit when a policy defined locked storage instances.</td></tr><tr><td>RUN-35612</td><td>Fixed a security vulnerability related to CVE-2025-64756 with severity HIGH.</td></tr><tr><td>RUN-34564</td><td>Fixed an issue where updating a project with node type names in the payload did not return the node type names in the 200 success response.</td></tr><tr><td>RUN-36732</td><td>Fixed a security vulnerability related to GHSA-5vv4-hvf7-2h46 with severity HIGH.</td></tr></tbody></table>

## February 2026 Releases

### February 23

#### Product Enhancements

* **New blocked rule for workload policies** - A new `blocked` rule was added to workload policies, allowing administrators to prevent AI practitioners from specifying a value for a field. This can be used to lock security-related configurations from user modification without enforcing a specific default value (for example, supplemental groups). See [Policy YAML reference](/saas/platform-management/policies/native-workload-policies/policy-yaml-reference#rule-types) for more details.
* **UI inactivity timeout updates** - The **Session timeout** setting was renamed to **UI inactivity timeout**. If left blank, users are logged out after 24 hours of inactivity by default. CLI and API access remain unaffected. See [General settings](/saas/settings/general-settings#security) for more details.
* **Permission-based access to the Overview dashboard** - The Overview dashboard is now available to roles that have at least one of the relevant READ permissions listed below. Dashboard widgets and data are displayed according to each user’s permissions, and widgets that are not applicable to the user’s permissions are automatically hidden. This enhancement also supports custom roles with different permission sets.
  * Clusters READ
  * Node pools READ
  * Nodes READ
  * Projects READ
  * Departments READ
  * Workloads READ

#### Resolved Bugs

<table><thead><tr><th width="253.79296875">ID</th><th>Description</th></tr></thead><tbody><tr><td>RUN-34472</td><td>Fixed an issue where the "Allocation ratio by node pool" widget in the Overview dashboard aggregated unlimited quotas together with other quotas, resulting in incorrect data.</td></tr><tr><td>RUN-33566</td><td>Fixed an issue where, after the <code>runai upgrade</code> command completed successfully, the CLI incorrectly prompted the user to run the upgrade again.</td></tr><tr><td>RUN-36045</td><td>Fixed an issue where inference workload metrics were not being refreshed correctly.</td></tr><tr><td>RUn-36257</td><td>Fixed an issue in the flexible workload submission form where image pull secret section would present shared credentials instead of shared secrets resulting in a failure to submit the workload.</td></tr><tr><td>RUN-36381</td><td>Fixed a security vulnerability related to GHSA-jmp9-x22r-554x with severity HIGH.</td></tr><tr><td>RUN-36598</td><td>Fixed an issue where department data was not synced to the cluster, affecting both department creation and updates.</td></tr><tr><td>RUN-34624</td><td>Fixed an issue in Projects and Departments where GPU utilization/allocation metrics were not displayed if only partial data was available.</td></tr><tr><td>RUN-36382</td><td>Fixed a security vulnerability related to GHSA-cv78-6m8q-ph82 with severity HIGH.</td></tr><tr><td>RUN-36414</td><td>Fixed a security vulnerability related to CVE-2025-14459 and CVE-2025-64324 with severity HIGH.</td></tr><tr><td>RUN-36451</td><td>Fixed an issue where users with the appropriate permissions could not delete system templates in the UI.</td></tr><tr><td>RUN-36457</td><td>Fixed an issue where, on rare occasions, "Allocation ratio by node pool" widget would show incorrect data.</td></tr><tr><td>RUN-36501</td><td>Fixed an issue where a node pool that included nodes without required topology labels became stuck in Updating after a topology was attached.</td></tr><tr><td>RUN-36506</td><td>Fixed an issue where the UI shows the wrong GPU quotas for node pools associated with the “Default” department.</td></tr><tr><td>RUN-36505</td><td>Fixed an issue where, on rare occasions, there was a race condition in some of the metrics causing the average GPU utilization to be above 100%.</td></tr></tbody></table>

### February 10

#### Product Enhancements

* **Redesigned Projects and Departments management** - NVIDIA Run:ai introduces an improved organization management experience that provides better visibility into resource distribution and clearer explainability for how resources are prioritized and allocated across the organization. This update simplifies large-scale organizational management while maintaining full compatibility with NVIDIA Run:ai’s advanced scheduling capabilities. See [Projects](/saas/platform-management/aiinitiatives/organization/projects) and [Departments](/saas/platform-management/aiinitiatives/organization/departments) for more details. `From cluster v2.20 onward`
  * **Improved organizational visibility** - A clearer, “big picture” view of projects and departments, making it easier to understand how GPU resources are distributed and prioritized.
  * **Bulk management operations** - Administrators can perform bulk actions across multiple organizational units directly from the UI and API, reducing operational overhead.
  * **Clearer resource explainability** - Improved transparency into resource contention and ordering, helping align scheduling behavior with business needs.
* **Increased initialization timeout for inference workloads** - The maximum initialization timeout for inference workloads and templates in the UI has been increased to 720 minutes, allowing workloads with longer startup times, such as large models, to complete successfully without premature failure.
* **Delete predefined environment assets** - Users can now delete predefined environment assets for inference workloads, `chatbot-ui`, `gpt2`, and `llm-server`, giving greater control over environment configuration.
* **UI adjustments for distributed training** - The distributed training workflow (workloads and templates) in the UI now includes a third step for mutual workload setup. The "Allow different setup for the master" toggle has been removed. By default, the master and workers use the same setup unless a policy defines different behavior. This applies to Flexible submission only. See [Train models using a distributed training workload](/saas/workloads-in-nvidia-run-ai/using-training/distributed-training-models) for more details.
* **New guided tour for projects and departments** - A built-in tour guides administrators through the projects and departments experience, highlighting key areas and workflows to help them get started quickly.

#### Resolved Bugs

<table><thead><tr><th width="253.79296875">ID</th><th>Description</th></tr></thead><tbody><tr><td>RUN-36122</td><td>Fixed an issue where credentials assets were not displayed in the Credentials table.</td></tr><tr><td>RUN-36010</td><td>Fixed an issue where navigating back to the root level in dashboard widgets caused the dashboard to crash.</td></tr><tr><td>RUN-35976</td><td>Fixed an issue where workloads submitted with names longer than 63 characters failed to schedule.</td></tr><tr><td>RUN-35922</td><td>Fixed a security vulnerability related to CVE-2026-0861 with severity HIGH.</td></tr><tr><td>RUN-35637</td><td>Fixed an issue where, when CPU quota and Limit projects from exceeding department quota were both enabled, updating department or project memory quotas to very large values failed with incorrect validation errors, even though the values were valid.</td></tr><tr><td>RUN-35620</td><td>Fixed an issue where providing an invalid admin password during installation caused the tenant to become permanently stuck.</td></tr><tr><td>RUN-35594</td><td>Fixed an issue where the <code>workload describe</code> command did not display the master specification for distributed workloads.</td></tr><tr><td>RUN-35511</td><td>Fixed an issue where an incorrect FQDN used during certificate generation caused errors.</td></tr><tr><td>RUN-35834</td><td>Fixed an issue where AI practitioner role did not have read access to policies granted through workload submission permission sets (for example, <code>workspaceEditAccess</code>).</td></tr><tr><td>RUN-36254</td><td>Fixed an issue where a race condition during webhook certificate generation caused failures.</td></tr><tr><td>RUN-35443</td><td>Fixed a security vulnerability related to CVE-2025-68973 with severity HIGH.</td></tr><tr><td>RUN-35326</td><td>Fixed an issue where the Projects/Departments table in the Overview dashboard sometimes showed fewer than 15 projects/departments when their workloads did not have allocated GPUs or were not in Running or Pending status.</td></tr><tr><td>RUN-35169</td><td>Fixed an issue where distributed inference workloads could be submitted successfully with an invalid workers value</td></tr><tr><td>RUN-34593</td><td>Fixed an issue in the Overview dashboard where the Node pool filter did not work for the Idle workloads table.</td></tr><tr><td>RUN-34017</td><td>Fixed an issue where <code>runai template list</code> returned incorrect output when using <code>--page-size</code> and <code>--max-items</code> together.</td></tr></tbody></table>

## January 2026 Releases

### January 26

#### Product Enhancements

**Pod logs and terminal access** - Accessing pod logs and interactive shells is now faster and more consistent across the Workloads experience. You can open logs or connect to running pods directly from multiple entry points, with pod selection and status kept in sync as you move between views. See [Workloads](/saas/workloads-in-nvidia-run-ai/workloads#show-hide-details) for more details. `From cluster v2.24 onward`

* One-click access to logs and terminals from the Pods view and Logs view, with the selected pod opened automatically.
* New Terminal tab for interactive access to pods and containers, including automatic connection when launching from the Workload grid.
* Synchronized pod selection and status across Pods, Logs, and Terminal views, while preserving existing responsive pod name behavior.

#### Resolved Bugs

<table><thead><tr><th width="253.79296875">ID</th><th>Description</th></tr></thead><tbody><tr><td>RUN-35623</td><td>Fixed an issue where running <code>runai logout</code> returned 404 Not Found when the session token had already expired. The logout command now completes successfully and returns a clear message.</td></tr><tr><td>RUN-35583</td><td>Fixed an issue where the <code>template describe</code> command did not display the master specification for distributed templates when the master and worker configurations differed.</td></tr><tr><td>RUN-35566</td><td>Fixed an issue where image pull secrets marked with <code>exclude=true</code> were not excluded from the workload.</td></tr><tr><td>RUN-35460</td><td>Fixed an issue where during password change, the wrong current password logged the user out and redirected them to the login page.</td></tr><tr><td>RUN-35769</td><td>Fixed an issue on OpenShift clusters where missing permissions to manage finalizers caused all workloads to remain stuck in Creating state.</td></tr><tr><td>RUN-35421</td><td>Fixed a security vulnerability related to CVE-2025-15284 with severity HIGH.</td></tr><tr><td>RUN-35388</td><td>Fixed an issue where distributed training workloads were not blocked when master and worker roles used different node pools.</td></tr><tr><td>RUN-35148</td><td>Fixed an issue where charts in the Overview dashboard did not render data after the node pool filter was changed.</td></tr><tr><td>RUN-32181</td><td>Fixed a security vulnerability related to CVE-2025-32988 with severity HIGH.</td></tr><tr><td>RUN-34875</td><td>Fixed an issue where enabling authentication and authorization prevented user metrics from being collected for inference workloads running on Knative and NIM.</td></tr></tbody></table>

### January 15

#### Product Enhancements

* **Custom roles using permission sets (API)** - Administrators can now create custom roles by combining predefined permission sets using the [Roles](https://run-ai-docs.nvidia.com/api/authentication-and-authorization/roles) API. Permission sets are predefined, supported groupings of permissions that represent all required dependencies for a specific operation (for example, workload submission). Custom roles can then be assigned to users or groups through access rules in the UI or API, alongside the existing NVIDIA Run:ai predefined roles. This allows organizations to tailor access control to their operational needs while maintaining compatibility with the platform’s supported permission model. See [Roles](/saas/infrastructure-setup/authentication/roles#custom-roles-api-only) for more details. `From cluster v2.21 onward`
* **NVIDIA NIM service API enhancements** - NVIDIA Run:ai expands support for deploying and managing NIMs through the NVIDIA NIM Operator, providing a standardized, operator-based deployment flow aligned with NIM-native configurations. NIM services are fully managed through the NVIDIA Run:ai API, with UI and CLI support planned for a future release. This capability does not replace the current [NIM deployment flow](/saas/workloads-in-nvidia-run-ai/using-inference/nim-inference) and is available as an additional option. See [NVIDIA NIM](https://run-ai-docs.nvidia.com/api/workloads/nvidia-nim) API for more details. `From cluster v2.23 onward`
  * Autoscaling allowing NIM services to scale dynamically based on demand
  * Fractional GPU support, enabling NIM services to request and use partial GPUs for more efficient GPU utilization
  * Multi-node NIM deployments, enabling distributed NIM workloads across multiple nodes
  * Policy enforcement through a dedicated [NVIDIA NIM Policy](https://run-ai-docs.nvidia.com/api/policies/policy) API for consistent governance of NIM services
  * Partial updates via a new PATCH endpoint, allowing targeted changes without resubmitting the full specification
  * [NIM Cache](https://docs.nvidia.com/nim-operator/latest/cache.html) support for model stores, enabling caching of specific LLM or multi-LLM model artifacts to improve startup time and reuse across deployments
* **MNNVL acceleration for supported workload types** - NVIDIA Run:ai now enables running supported workload types on Multi-Node NVLink (MNNVL) domains, including GB200 NVL72 systems. NVIDIA Run:ai applies the appropriate compute domain configuration to ensure workloads are placed and scaled within the same NVLink domain. AI practitioners can submit supported workload types using the [Workloads V2](https://run-ai-docs.nvidia.com/api/workloads/workloads-v2) API and configure their MNNVL preference as part of the workload submission. `From cluster v2.24 onward`
* **Separate priority and preemptibility controls** - Workload priority and preemptibility are configured as two independent parameters across the UI, CLI, and API for native and supported workload types. If no preemptibility value is specified, the existing behavior based on priority is applied automatically. See [Workload priority and preemption](/saas/platform-management/runai-scheduler/scheduling/workload-priority-control) for more details. `From cluster v2.24 onward`
* **Authenticated browsing for the NGC catalog** - Browse the NGC catalog and private NGC registries as an authenticated user by selecting your NGC API key credentials during [workload](/saas/workloads-in-nvidia-run-ai/workloads#adding-a-new-workload) submission or [template](/saas/workloads-in-nvidia-run-ai/workload-templates#adding-a-new-template) creation. This provides access to models and containers that require authentication while preserving the option to browse the public container registry. Private NGC registries require administrator configuration in the [General settings](/saas/settings/general-settings#workloads). <mark style="color:orange;">`Beta`</mark> `From cluster v2.23 onward`
* **NGC API key support for NVIDIA NIM workloads** - NVIDIA Run:ai supports using an NGC API key when deploying NIM workloads to handle both image access and model runtime authentication. A single NGC API key is automatically applied for pulling NIM images from the NGC catalog and injected as a runtime environment variable required for downloading model weights. This streamlines NIM deployment by removing the need for separate pull secrets and runtime credentials while enabling full user self-service for authenticated NIM workloads. See See [Deploy inference workloads from NVIDIA NIM](/saas/workloads-in-nvidia-run-ai/using-inference/nim-inference) for more details. `From cluster v2.23 onward`
* **Updated predefined roles** - Predefined roles in NVIDIA Run:ai have been updated to better align with common organizational responsibilities and workflows. See [Roles](/saas/infrastructure-setup/authentication/roles#roles-in-nvidia-runai) for more details:
  * Added new predefined roles - AI practitioner, Data and storage administrator, and Project administrator
  * Some existing predefined roles have been deprecated. See [Deprecation notifications](#january-2026) for more details.
* **Reduced access to clusters and node pools** - New APIs and permissions are now available to support reduced access to clusters and node pools - Clusters minimal and Node pools minimal. These APIs allow roles to perform actions such as workload submission while exposing only the minimal required cluster and node pool information (for example, names and IDs), rather than full read access. This improves role design by aligning the visible data with what is actually required for the action being performed. Roles that rely on full read access remain unchanged. Some predefined roles are planned to transition to the new minimal access as described in the [Deprecation notifications](#january-2026).
* **Visibility into workload topology constraints** - Workloads now expose the topology constraints requested during scheduling, providing clear visibility into how network topology influences placement decisions. In the UI, NVIDIA Run:ai native workloads display the requested topology constraints in the workload [Details](/saas/workloads-in-nvidia-run-ai/workloads#show-hide-details) view, while the [Workloads](https://run-ai-docs.nvidia.com/api/workloads/workloads) API exposes these fields across native and supported workload types. `From cluster v2.24 onward`
* **Control access scope for inference serving endpoints** - Set whether an inference serving endpoint is accessible externally or restricted to internal cluster traffic when submitting [workloads](/saas/workloads-in-nvidia-run-ai/using-inference) or creating [templates](/saas/workloads-in-nvidia-run-ai/workload-templates/inference-templates). Endpoints can be configured as External (public access), if your administrator has configured Knative to support external access, or Internal only, limiting access to in-cluster traffic. `From cluster v2.24 onward`
* **Asset-based workload submission in the CLI** - The NVIDIA Run:ai CLI supports submitting native workloads using workload assets, such as compute resources, environments, and data sources. This allows AI practitioners to reuse the same predefined configurations available in the UI and API, reducing the need for long, flag-heavy CLI commands. Assets can be browsed and inspected directly from the CLI to support consistent and reliable workload submission. See [CLI command reference](/saas/reference/cli/runai). `From cluster v2.23 onward`
* **Improved fractional GPU support for multi-container pods** - Fractional GPUs are no longer limited to the first container in a pod. You can explicitly specify which container should receive fractional GPU resources using an annotation. If no container is specified, fractional GPUs continue to be associated with the first container by default. See [GPU fractions](/saas/platform-management/runai-scheduler/resource-optimization/fractions#setting-gpu-fractions-via-yaml) and [Dynamic GPU fractions](/saas/platform-management/runai-scheduler/resource-optimization/dynamic-fractions#setting-dynamic-gpu-fractions-via-yaml) for more details. `From cluster v2.24 onward`
* **Support for elastic distributed workloads on NVLink domains** - Elastic distributed workloads, including auto-scaling and dynamically sized deployments, are fully supported on GB200 NVL72 and Multi-Node NVLink (MNNVL) domains using NVIDIA DRA driver version 25.8 and later. NVIDIA Run:ai automatically applies ComputeDomain configuration and topology-aware scheduling to ensure workloads scale within the same NVLink domain. See [Using GB200 NVL72 and Multi-Node NVLink domains](/saas/platform-management/aiinitiatives/resources/using-gb200) for more details. `From cluster v2.24 onward`
* **Native Load Balancer support** - NVIDIA Run:ai exposes LoadBalancer connectivity directly in the UI and CLI when submitting [workloads](/saas/workloads-in-nvidia-run-ai/workloads#adding-a-new-workload) or creating [templates](/saas/workloads-in-nvidia-run-ai/workload-templates#adding-a-new-template) (assuming a load balancer is already installed in the cluster). Configure service ports explicitly and view clearer port configuration and connectivity status. `From cluster v2.24 onward`
* **Time-based fairshare configuration per node pool** - NVIDIA Run:ai supports time-based fairshare to improve long-term fairness in over-quota resource allocation. Instead of relying only on momentary demand, the Scheduler factors in historical GPU usage over time, ensuring that projects with lower recent consumption are given fair access to resources. Usage is tracked continuously, and each project’s GPU-hour consumption is evaluated against its configured weight to balance resource distribution more effectively across projects. Time-based fairshare can be enabled and configured per node pool using the [Node pools](/saas/platform-management/aiinitiatives/resources/node-pools#adding-a-new-node-pool) form, with advanced customization available through the [Node pools](https://run-ai-docs.nvidia.com/api/organizations/nodepools) API. `From cluster v2.24 onward`
* **Extended storage visibility in CLI describe commands** - The `describe` command for native workloads supports `--storage`, showing storage resources such as PVCs, ConfigMaps, and Secrets.
* **Default pod anti-affinity** - A new cluster configuration, `global.requireDefaultPodAntiAffinity`, applies a default pod anti-affinity rule to prevent pods from the same service from being scheduled on the same node when possible. This setting is enabled by default. See [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config) for more details. `From cluster v2.24 onward`
* **Ingress controller recommendation update** - Due to an announced deprecation by the upstream [NGINX Ingress Controller](https://kubernetes.io/blog/2025/11/11/ingress-nginx-retirement/) project, NVIDIA Run:ai is updating its recommended ingress controller to HAProxy Ingress for supported environments. The Kubernetes Ingress standard remains fully supported. This change affects only the underlying ingress controller implementation and is intended to ensure long-term security, stability, and maintainability. For fresh installations, see [Installation](/saas/getting-started/installation/install-using-helm/system-requirements#kubernetes-ingress-controller). To upgrade from earlier versions, see [Migrate from NGINX to HAProxy Ingress](/saas/getting-started/installation/install-using-helm/upgrade#migrate-from-nginx-to-haproxy-ingress). `From cluster v2.24 onward`
* **Component version updates** - NVIDIA Run:ai now supports Kubernetes version 1.35, OpenShift version 4.20, and GPU Operator version 25.10. Support for Kubernetes version 1.31 and OpenShift version 4.16 has been removed. `From cluster v2.24 onward`
* **Rancher Kubernetes Engine (RKE1)** - RKE1 is no longer supported due to reaching end of life (EOL). RKE2 is the recommended Rancher distribution. See the Rancher [migration guide](https://support.scc.suse.com/s/kb/RKE-to-RKE2-replatforming-instructions-and-FAQs?language=en_US) for more details. `From cluster v2.24 onward`

#### Resolved Bugs

<table><thead><tr><th width="253.79296875">ID</th><th>Description</th></tr></thead><tbody><tr><td>RUN-35189</td><td>Fixed an issue where the <code>--working-dir</code> parameter was ignored for Knative-based inference workloads, causing containers to start in <code>/</code> instead of the specified directory.</td></tr><tr><td>RUN-34639</td><td>Fixed an issue where the Fully free GPU devices column displayed <code>-</code> instead of <code>0</code> when no fully free GPU devices were available under fractional GPU allocations.</td></tr><tr><td>RUN-34867</td><td>Fixed an issue where projects created or updated during node pool deletion could reference a non-existent node pool and remain NotReady.</td></tr><tr><td>RUN-34607</td><td>Fixed issues where readiness probes did not work correctly with serving port authorization in single-node Knative inference workloads.</td></tr><tr><td>RUN-35348</td><td>Fixed an issue where the <code>/v1/k8s/setting</code> endpoint returned a 500 error for tenants without clusters, causing the UI to hang instead of redirecting to cluster creation.</td></tr><tr><td>RUN-34381</td><td>Fixed an issue where the Node column displayed a sort icon but did not actually sort results in the Running / Requested Pods modal.</td></tr><tr><td>RUN-34379</td><td>Fixed an issue where image names longer than the display limit were truncated without providing access to the full name.</td></tr><tr><td>RUN-35206</td><td>Fixed an issue causing a delay before newly created clusters could be deleted, leaving the Remove option temporarily unavailable after creation.</td></tr><tr><td>RUN-34611</td><td>Fixed an issue where Overview widgets did not update correctly when navigating from departments to projects.</td></tr><tr><td>RUN-35290</td><td>Fixed an issue where copying a workload that included node affinity settings caused the re-submission to fail.</td></tr><tr><td>RUN-34721</td><td>Fixed a security vulnerability related to CVE-2024-25621 with severity HIGH.</td></tr><tr><td>RUN-32181</td><td>Fixed a security vulnerability related to CVE-2025-32988 with severity HIGH.</td></tr><tr><td>RUN-34720</td><td>Fixed a security vulnerability related to CVE-2025-65637 with severity HIGH.</td></tr><tr><td>RUN-34680</td><td>Fixed a security vulnerability related to CVE-2025-58183 with severity HIGH.</td></tr><tr><td>RUN-35089</td><td>Fixed a security vulnerability related to CVE-2025-64756 with severity HIGH.</td></tr><tr><td>RUN-34620</td><td>Fixed an issue where, in rare cases, sessions could disconnect due to token refresh handling.</td></tr></tbody></table>

### January 05

#### Product Enhancements

* **YAML-based workload submission in the UI** - Submit supported workload types defined in YAML directly from the UI. This brings YAML-based submission, previously available through the API, into interactive workflows, allowing you to submit existing Kubernetes or framework-specific manifests while still benefiting from NVIDIA Run:ai scheduling, resource management, and monitoring. See [Submit supported workload types via YAML](/saas/workloads-in-nvidia-run-ai/submit-via-yaml) for more details. `From cluster v2.23 onward`
* **Automatic network topology acceleration for supported workloads** - Network topology–aware scheduling is applied automatically to supported distributed workloads submitted [via YAML](/saas/workloads-in-nvidia-run-ai/submit-via-yaml). Once a topology is attached to a node pool, NVIDIA Run:ai automatically applies Preferred topology constraints at the lowest available level for the entire workload, optimizing pod placement without additional user configuration. This expands the topology acceleration beyond NVIDIA Run:ai native distributed workloads to additional workload types. See [Accelerating workloads with network topology-aware scheduling](/saas/platform-management/aiinitiatives/resources/topology-aware-scheduling) for more details. `From cluster v2.23 onward`
* **AI application–based workload grouping** **in the UI** - NVIDIA Run:ai provides a dedicated AI applications view. This view automatically groups Kubernetes resources deployed via Helm charts into a single logical application, allowing you to list, sort, and filter AI applications. You can also inspect aggregated resource requests and allocations (GPU, CPU, memory) and view the underlying workloads through the Details pane, making it easier to understand and manage complex, multi-component solutions. See [AI applications](/saas/ai-applications/ai-applications) for more details. `From cluster v2.23 onward`
* **Dynamo as a supported workload type** - NVIDIA Run:ai supports Dynamo-based inference workloads through the DynamoGraphDeployment workload type. This allows Dynamo workloads to be deployed, scheduled, and monitored using the same platform capabilities and operational model as native workloads. See [Supported workload types](/saas/workloads-in-nvidia-run-ai/workload-types/supported-workload-types) for more details. `From cluster v2.23 onward`

  Key capabilities include:

  * **YAML-based deployment and management** - Dynamo workloads can be submitted using [YAML](/saas/workloads-in-nvidia-run-ai/submit-via-yaml) from the UI, API, or CLI, without requiring direct cluster access.
  * **Hierarchical gang scheduling** - NVIDIA Run:ai supports hierarchical (multi-level) gang scheduling for Dynamo workloads. Replica groups are scheduled together as sub-gangs, and the entire workload is then scheduled as a single unit. This ensures coordinated placement and execution across all components of the Dynamo workload.
  * **Topology-aware scheduling** - NVIDIA Run:ai applies [topology-aware scheduling](/saas/platform-management/aiinitiatives/resources/topology-aware-scheduling) at the workload level to ensure Dynamo workload components are placed according to the underlying cluster topology, improving communication efficiency and execution consistency.
  * **Automatic discovery of Dynamo frontend endpoints** - NVIDIA Run:ai automatically detects Dynamo frontend endpoints and exposes them for access and monitoring.
  * **Unified workload lifecycle and status visibility** - Dynamo workloads are managed, monitored, and tracked with a unified lifecycle and status view.
* **Distributed inference support in the CLI** - Native distributed inference workloads can be submitted and managed directly from the NVIDIA Run:ai CLI. AI practitioners can use familiar NVIDIA Run:ai commands to work with distributed inference workloads, such as list, describe, logs, exec, port-forward, update, and delete. See [CLI command reference](/saas/reference/cli/runai/runai-inference-distributed) for more details. `From cluster v2.23 onward`
* **Template-based workload submission in the CLI** - The NVIDIA Run:ai CLI now supports submitting native workloads using existing templates. This allows AI practitioners to reuse the same predefined configurations available in the UI and API, reducing the need for long, flag-heavy CLI commands. Templates can be browsed and inspected directly from the CLI to support consistent and reliable workload submission. See [CLI command reference](/saas/reference/cli/runai/runai-template) for more details. `From cluster v2.23 onward`

#### Resolved Bugs

<table><thead><tr><th width="253.79296875">ID</th><th>Description</th></tr></thead><tbody><tr><td>RUN-34613</td><td>Fixed an issue where the Project GET API returned missing limit fields instead of an explicit unlimited value when CPU quotas were enabled.</td></tr><tr><td>RUN-30979</td><td>Fixed an issue where the PVC API did not validate <code>claimName</code> uniqueness.</td></tr><tr><td>RUN-34203</td><td>Fixed an issue where workloads using multiple GPU fractions were missing GPU utilization and memory metrics.</td></tr><tr><td>RUN-34631</td><td>Fixed an issue where the identity manager failed to start when the notification service was disabled.</td></tr><tr><td>RUN-34684</td><td>Fixed a security vulnerability related to CVE-2025-58183 with severity HIGH.</td></tr><tr><td>RUN-34694</td><td>Fixed a security vulnerability related to CVE-2025-58186 with severity HIGH.</td></tr><tr><td>RUN-34703</td><td>Fixed a security vulnerability related to CVE-2025-58187 with severity HIGH.</td></tr><tr><td>RUN-34712</td><td>Fixed a security vulnerability related to CVE-2025-61729 with severity HIGH.</td></tr><tr><td>RUN-34758</td><td>Fixed an issue where setting a GPU memory limit caused workload creation to fail.</td></tr></tbody></table>

## December 2025 Releases

### December 14

#### Product Enhancements

* **LeaderWorkerSet (LWS) as a new workload type** - LeaderWorkerSet is now available as a supported workload type. LWS workloads can be deployed and managed [YAML submission](/saas/workloads-in-nvidia-run-ai/submit-via-yaml) from the UI, API, or CLI, providing a standardized way to run leader–worker and multi-process workloads across the platform without direct cluster access. See [Supported workload types](/saas/workloads-in-nvidia-run-ai/workload-types/supported-workload-types) for more details. `From cluster v2.23 onward`
* **Submit workloads from YAML via the CLI** - The CLI now supports submitting workloads directly from a YAML definition using a new `runai workload submit -f` command. This enables declarative workload creation while still allowing key fields to be overridden at submission time. Workloads created from YAML can also be deleted through the CLI, providing a simple way to manage YAML-defined workloads. See [CLI commands reference](/saas/reference/cli/runai/runai-workload-submit) for more details.
* **Workloads v2 API update** - The [Workloads v2](https://run-ai-docs.nvidia.com/api/workloads/workloads-v2) API now includes a new PUT endpoint for updating workloads. This endpoint requires submitting a complete workload manifest, which fully replaces the existing workload configuration. `From cluster v2.23 onward`
* **Update NVIDIA NIM services** **API** - The [NVIDIA NIM](https://run-ai-docs.nvidia.com/api/workloads/nvidia-nim) API now includes a new PATCH endpoint for modifying NIM service workloads. This endpoint supports partial updates, allowing you to update only the fields you need without submitting the full workload definition. `From cluster v2.23 onward`
* **New NVIDIA NIM performance histograms added to Metrics** - The Metrics pane now includes two new histograms for NVIDIA NIM metrics. See [NVIDIA NIM metrics](/saas/workloads-in-nvidia-run-ai/workloads#nvidia-nim) for more details. `From cluster v2.23 onward`
  * **End-to-end request latency** - Displays request distribution across latency buckets, helping you identify performance patterns and outliers over time.
  * **Time to first token (TTFT)** - Shows the distribution of TTFT across requests, enabling faster detection of model responsiveness issues.
* **Updated** `runai login` **CLI command** - The `runai login` CLI command has been updated to streamline authentication options. These changes align the CLI login modes with the current authentication model and improve clarity in how each method is used. See [CLI commands reference](/saas/reference/cli/runai/runai-login-access-key) for more details.
  * `application` login is deprecated and replaced with `access-key`, which is now the supported method for logging in with a service account or user access key.
  * The previous `user` login mode is deprecated and renamed to `password` to more accurately reflect username-and-password authentication.
* **Email invitations for local users** - Administrators can now choose to automatically send an email invitation when creating a local user.
* **New EmptyDir data source for ephemeral storage** - A new EmptyDir data source is now available for one-time configuration during [workload](/saas/workloads-in-nvidia-run-ai/workloads#adding-a-new-workload) submission or [template](/saas/workloads-in-nvidia-run-ai/workload-templates#adding-a-new-template) creation. EmptyDir provides temporary, node-local storage that exists only for the lifetime of the workload.

#### Resolved Bugs

<table><thead><tr><th width="253.79296875">ID</th><th>Description</th></tr></thead><tbody><tr><td>RUN-30979</td><td>Fixed an issue where the PVC API did not validate <code>claimName</code> uniqueness.</td></tr><tr><td>RUN-33516</td><td>Fixed an issue so each access rule created or deleted in a batch action is now audited in the events history.</td></tr><tr><td>RUN-33806</td><td>Fixed an issue where containers ran as root instead of a non-privileged user.</td></tr><tr><td>RUN-33971</td><td>Fixed a permissions issue that allowed users with write-settings permissions to edit a centralized channel.</td></tr><tr><td>RUN-34048</td><td>Fixed an issue where inference workload URLs were generated as <code>http</code> instead of <code>https</code>.</td></tr><tr><td>RUN-34196</td><td>Fixed an issue where users with L1 Researcher and L2 Researcher roles could not list node pools using the NVIDIA Run:ai CLI.</td></tr><tr><td>RUN-34420</td><td>Fixed an issue where the NGC API key asset <code>getById</code> API response was missing the <code>status</code> field.</td></tr><tr><td>RUN-34429</td><td>Fixed an issue where users with the correct project permissions could create templates but were blocked from saving edits due to incorrect permission checks.</td></tr></tbody></table>

### December 03

#### Product Enhancements

* **Service accounts replacing applications in the UI** - The Applications feature has been renamed to Service accounts throughout the UI. All existing functionality remains the same, and existing application records continue to appear unchanged. See [Service accounts](/saas/infrastructure-setup/authentication/service-accounts) for more details.
* **Connections column enabled by default in the Workloads grid** - The Connections column is now selected by default in the [Workloads](/saas/workloads-in-nvidia-run-ai/workloads#workloads-table) table. When a workload has a single connection its URL is displayed directly, with long URLs automatically shortened using an ellipsis. The URL is clickable and opens the Connections dialog. When multiple connections exist, the table displays the total count.
* **Enhanced workload details view** - The Workload Details tab now provides an enriched and clearer view of workload configuration data. The updated design improves readability and makes it easier to understand how a workload was submitted and configured. Key enhancements include:
  * **Improved layout and data presentation** - Configuration fields are now grouped and displayed more intuitively, helping users quickly find the information they need.
  * **Specification selector** - When a workload contains multiple specs, a new dropdown allows you to easily switch between them.
* **Overview dashboard enhancements** - We’ve made several improvements to the Overview dashboard to strengthen visibility and better support key monitoring workflows. These enhancements also support deprecating the legacy Grafana dashboards. See [Deprecation notifications](#september-2025) for more details:
  * **Enhancements to the Projects/Departments tables** - Timeframe controls for GPU allocation, utilization, and memory utilization are now located within each column. The tables now display up to 15 entries, include GPU quota, separate pending and running workloads, and provide direct links to each project or department.
  * **New pending-time widget** - Introduced a new widget that displays pending workloads count by pending time, helping admins understand how much time their workloads are waiting and also identify the projects/departments experiencing extended pending times.
  * **New guided tour for the Overview dashboard** - A built-in tour now walks administrators and AI practitioners through the key areas of the Overview dashboard, helping them navigate the interface and become familiar with the functionalities that enable them to get the most out of the dashboard.
  * **Additional dashboard improvements:**
    * Added a top-stats **Failed workloads** widget.
    * Added numeric counters to bar graphs, displaying values directly on each bar rather than only in tooltips.
    * Updated the **Workloads by category/type** widget to count **running** workloads only.
    * Added an **Idle time** column to the idle workloads table.
    * Updated widget ordering to separate **current time widgets** from **over time widgets**, improving the analysis process from identifying issue to over time investigation.
* **Consumption report enhancements for GPU hour breakdown** - The Consumption report now includes two new columns, GPU deserved quota hours and GPU over-quota hours. These metrics fully support all existing grouping options, including cluster, node pool, department, and project. This change also supports deprecating the legacy Consumption dashboard. See [Deprecation notifications](#september-2025) for more details.
* **Network topology visibility in clusters and node pools** - The Network topologies modal in the [Clusters](/saas/infrastructure-setup/procedures/clusters#network-topologies-associated-with-the-cluster) page displays a new column showing which node pools each topology is associated with. This information is also available in the [Network topologies](https://run-ai-docs.nvidia.com/api/organizations/network-topologies) API. In addition, the node pools list command in the CLI now includes a network topology column, showing the name of the topology assigned to each node pool. `From cluster v2.23 onward`
* **Policy-aware behavior for templates and assets** - Templates and assets that do not fully comply with policy are no longer blocked outright when submitting a workload. Instead, NVIDIA Run:ai now evaluates non-compliance on a case-by-case basis:
  * **Fixable non-compliance** - If compliance can be achieved by adjusting settings during submission, the template or asset can be loaded. The UI highlights what needs to be updated to meet policy requirements.
  * **Non-fixable non-compliance** - If the non-compliant configuration cannot be changed, the template or asset cannot be used, and the relevant policy is displayed to explain the restriction.
  * **Quick workload submit behavior** - Templates with any non-compliance can now be loaded, but quick submit is automatically blocked. The full workload submission flow opens by default, where the UI highlights what needs to be updated to meet policy requirements.
* **Authenticated browsing for the NGC catalog** - You can now browse the NGC catalog as an authenticated user by selecting your own NGC API key credentials during workload submission or template creation. This enables access to models and containers that require authentication while preserving the existing option to browse the public container registry. <mark style="color:orange;">`Beta`</mark> `From cluster v2.23 onward`
* **New PVC events** - NVIDIA Run:ai now emits new PVC asset lifecycle events - Creating, Deleting, and Syncing. These events appear in the PVC’s Event history, extending the visibility introduced in previous releases and giving administrators clearer insight into PVC asset changes and activity over time.

#### Resolved Bugs

<table><thead><tr><th width="253.79296875">ID</th><th>Description</th></tr></thead><tbody><tr><td>RUN-34252</td><td>Fixed an issue that caused charts to remain in a loading state on every data refresh instead of only during the initial load.</td></tr><tr><td>RUN-33902</td><td>Fixed an issue where the workloads service could enter a CrashLoopBackOff during upgrade.</td></tr><tr><td>RUN-31856</td><td>Fixed a security vulnerability related to CVE-2025-47907 with severity HIGH.</td></tr><tr><td>RUN-33841</td><td>Fixed an issue that caused session disconnections.</td></tr><tr><td>RUN-33802</td><td>Fixed an issue that caused distributed inference workloads to become unsynchronized.</td></tr><tr><td>RUN-33642</td><td>Fixed an issue where the external-workload-integrator on OpenShift entered a constant reconcile loop, causing high CPU utilization.</td></tr><tr><td>RUN-33613</td><td>Fixed missing validations for CPU resources when the CPU quota feature flag was disabled, which caused project and department updates to skip required CPU checks.</td></tr><tr><td>RUN-33526</td><td>Fixed an issue that could cause the operator to crash during installation due to a race condition in ingress initialization.</td></tr><tr><td>RUN-33519</td><td>Fixed an issue where the UI incorrectly prevented creating templates with the same name across different scopes.</td></tr><tr><td>RUN-32889</td><td>Fixed an issue where idle GPU timeout rules were incorrectly applied to preemptible workspaces.</td></tr></tbody></table>

## November 2025 Releases

### November 18

#### Product Enhancements

* **Service accounts replacing applications** - The applications feature has been renamed to service accounts in the API. Service accounts provide the same functionality for programmatic authentication and management but with updated terminology. The deprecation of applications begins with version 2.24 and will continue for two additional releases before removal. Existing application records and endpoints will remain functional during this period to ensure backward compatibility. See the [Applications](https://run-ai-docs.nvidia.com/api/authentication-and-authorization/applications) API (`/api/v1/apps`) for more details.
* **Distributed inference templates (API)** - Distributed inference templates allow you to save workload configurations that can be reused across distributed inference submissions. These templates simplify the submission process and promote standardization across distributed inference workloads. `From cluster v2.22 onward`
* **Policy API for NIM services (API)** - A new Policy API is now available for NVIDIA NIM services, enabling administrators to define and enforce policies that control the behavior of NIM service workloads. These policies help ensure consistent configurations across deployments, improve governance, and simplify management of NIM service workloads. `From cluster v2.23 onward`
* **Autoscaling support in NVIDIA NIM service API** - The NVIDIA NIM service API now supports autoscaling for inference workloads deployed through the NIM Operator. When enabled, NIM services automatically adjust the number of active replicas based on defined metrics, allowing deployments to scale up or down dynamically as traffic changes. `From cluster v2.23 onward`
* **Multi-node NIM support in the NVIDIA NIM service API** - The NVIDIA NIM service API now supports deploying multi-node NVIDIA NIM workloads. `From cluster v2.23 onward`
* **Hugging Face model catalog browsing** - You can now browse and search the Hugging Face model catalog directly from the NVIDIA Run:ai UI and API when creating inference workloads. The live catalog view displays model details such as download count and gated status. For gated models, the platform prompts you to provide a Hugging Face token for access, while open models can be selected without authentication. See [Deploy inference workloads from Hugging Face](/saas/workloads-in-nvidia-run-ai/using-inference/hugging-face-inference) for more details.

#### Resolved Bugs

<table><thead><tr><th width="253.79296875">ID</th><th>Description</th></tr></thead><tbody><tr><td>RUN-33471</td><td>Fixed an issue where cluster authentication didn’t use the tenant URL.</td></tr><tr><td>RUN-33638</td><td>Fixed an issue where the DCGM metric chart was displayed even when the cluster did not support DCGM metrics.</td></tr><tr><td>RUN-33634</td><td>Fixed an issue where resource name validation failed for hugepage resources by enhancing validation rules to properly support hugepages.</td></tr><tr><td>RUN-33448</td><td>Fixed an issue where switching between workloads in the workload Details drawer displayed incorrect data, particularly the workload lifespan value.</td></tr><tr><td>RUN-33418</td><td>Fixed an issue where the master spec was not inherited when creating a distributed workload from a template.</td></tr><tr><td>RUN-33364</td><td>Fixed an issue where policies allowed <code>canEdit: false</code> under <code>attributes</code> without specifying a default value, which incorrectly passed validation.</td></tr><tr><td>RUN-33313</td><td>Fixed an issue where the log viewer for distributed workloads displayed only a partial and unsorted list of pods.</td></tr><tr><td>RUN-33300</td><td>Fixed an issue where the metric <code>gpu_memory_utilization_avg</code> returned a NaN value.</td></tr><tr><td>RUN-33144</td><td>Fixed a security vulnerability related to CVE-2025-62156 with severity HIGH.</td></tr><tr><td>RUN-33127</td><td>Fixed an issue where workload submission in the CLI failed when commands contained special characters.</td></tr><tr><td>RUN-33099</td><td>Fixed an issue where a mismatch between Helm schema validation and pre-hooks runtime validation code caused <code>clusterConfig.binder.resources</code> errors during upgrades</td></tr><tr><td>RUN-33091</td><td>Fixed an issue where workloads logs initially loaded older logs instead of the most recent ones.</td></tr><tr><td>RUN-33054</td><td>Fixed an issue where creating or updating a policy failed with an ‘asset Id not found’ error when specifying an imposedAsset.</td></tr><tr><td>RUN-33044</td><td>Fixed an issue where the workload controller could delete all running workloads when <code>init-ca</code> generated a new certificate (every 30 days).</td></tr><tr><td>RUN-32702</td><td>Fixed an issue where users running Red Hat OpenShift Serverless experienced “Down” status alerts in OpenShift monitoring due to NVIDIA Run:ai Knative ServiceMonitors.</td></tr><tr><td>RUN-32680</td><td>Fixed an issue where logs were not displayed in the UI for workloads submitted using the Workloads v2 submission API.</td></tr><tr><td>RUN-32673</td><td>Fixed an issue where inference workload metrics did not allow selecting a specific pod for viewing metrics.</td></tr><tr><td>RUN-32642</td><td>Fixed an issue where the UI displayed an incorrect access rule status for users with Cloud Operator roles.</td></tr><tr><td>RUN-32572</td><td>Fixed an issue where the <code>RunaiAgentPullRateLow</code> and <code>RunaiAgentClusterInfoPushRateLow</code> Prometheus alerts were firing incorrectly without cause.</td></tr><tr><td>RUN-32449</td><td>Fixed an issue where a race condition between the NVIDIA Run:ai operator and upgrade/install post hooks caused the upgrade to fail.</td></tr><tr><td>RUN-31738</td><td>Fixed an issue where GPU fraction requests were not applied when submitting distributed workloads.</td></tr><tr><td>RUN-32989</td><td>Fixed an issue where the NVIDIA Run:ai operator experienced unusually high CPU utilization after upgrade.</td></tr><tr><td>RUN-32986</td><td>Fixed an issue where PVCs appeared with the status “Issues found” after upgrading to version 2.22.</td></tr></tbody></table>

### November 02

#### Product Enhancements

* **Updated credential creation in the UI** - The Credentials page has been redesigned for improved usability. The Access key and Username & password credential types have been consolidated under Generic secret, where each secret format now opens a dedicated form with context-specific input fields. In addition, a dedicated SSH key format has been added under Generic secret for easier configuration of SSH-based authentication. This change simplifies the UI and provides a more streamlined experience for managing credentials. See [Credentials](/saas/workloads-in-nvidia-run-ai/assets/credentials#adding-new-credentials) for more details.
* **Min/max worker configuration for PyTorch distributed training** - You can now define the minimum and maximum number of workers directly from the UI when submitting PyTorch distributed training workloads. This provides greater flexibility and control over resource allocation. See [Train models using a distributed training workload](/saas/workloads-in-nvidia-run-ai/using-training/distributed-training-models#setting-up-compute-resources) for more details.
* **Audit logging for password resets** - Audit logs now capture all password reset events, including administrator-initiated resets, user-initiated resets, and password-recovery (“forgot password”) actions. This enhancement improves traceability and security visibility across user management workflows.
* **Access keys replacing user applications** - The User applications feature has been renamed to Access keys across the UI and API (`/api/v1/user-applications`). Access keys provide the same functionality for programmatic authentication and management but with updated terminology. The deprecation of User applications begins with version 2.24 and will continue for two additional releases before removal. Existing User application records and endpoints will remain functional during this period to ensure backward compatibility. See [Access keys](/saas/settings/user-settings/user-access-keys) for more details.

#### Resolved Bugs

<table><thead><tr><th width="253.79296875">ID</th><th>Description</th></tr></thead><tbody><tr><td>RUN-33365</td><td>Fixed an issue where selecting an environment asset template in the flexible workload form would not present the the capabilities field correctly.</td></tr><tr><td>RUN-33447</td><td>Fixed an issue where the API allowed creating a PVC asset without a <code>claimName</code> when <code>existingPVC=false</code>.</td></tr><tr><td>RUN-32968</td><td>Fixed an issue where users without permission to create data source assets were blocked from adding one-time data sources during workload submission.</td></tr><tr><td>RUN-33314</td><td>Fixed an issue where the NGC API key validation did not allow special characters (<code>-</code>, <code>_</code>, <code>.</code>). Validation now supports these characters as expected.</td></tr><tr><td>RUN-33177</td><td>Fixed an issue where removing the logo in Branding settings displayed an empty square.</td></tr><tr><td>RUN-33176</td><td>Fixed an issue where pagination in the Node Pool page did not respond.</td></tr><tr><td>RUN-33038</td><td>Fixed an issue where department administrators could not include cluster-scope templates in workloads due to incorrect validation of permitted scopes.</td></tr><tr><td>RUN-33036</td><td>Fixed an issue where the grace period preemption field in the UI was limited to 5 minutes, even when the workload policy allowed longer durations.</td></tr><tr><td>RUN-33006</td><td>Fixed an issue in the CLI installer where the PATH was not configured for all shells. The installer now correctly configures PATH for both zsh and bash.</td></tr><tr><td>RUN-32995</td><td>Fixed an issue where policies were not applied when submitting a workload using a template.</td></tr><tr><td>RUN-32752</td><td>Fixed an issue where the filterBy department option in the consumption report did not work as expected.</td></tr><tr><td>RUN-29375</td><td>Fixed an issue where stale department were not properly removed after deleting a cluster.</td></tr><tr><td>RUN-33053</td><td>Fixed an issue that caused conflicts with additional built-in Prometheus Operator deployments in OpenShift.</td></tr><tr><td>RUN-32876</td><td>Fixed an issue where running a NIM inference workload on a fractional GPU prevented the Triton server from starting, causing inference endpoint requests to fail.</td></tr><tr><td>RUN-32730</td><td>Fixed an issue where incorrect average GPU utilization per project and workload type was displayed in the Projects view charts and tables.</td></tr><tr><td>RUN-32159</td><td>Fixed an issue where the <code>updatedBy</code> field of a policy did not show the latest user who updated it.</td></tr><tr><td>RUN-31803</td><td>Fixed an issue where the Quota management dashboard occasionally displayed incorrect GPU quota values.</td></tr></tbody></table>

## October 2025 Releases

### October 19

#### Resolved Bugs

<table><thead><tr><th width="253.79296875">ID</th><th>Description</th></tr></thead><tbody><tr><td>RUN-33039</td><td>Fixed an issue where setting <code>uid</code> or <code>gid</code> to <code>0</code> during environment creation was not allowed.</td></tr><tr><td>RUN-33147</td><td>Fixed an issue where users with expired refresh tokens (after 24 hours) could not log in, as the token endpoint returned a 400 error.</td></tr><tr><td>RUN-33168</td><td>Fixed an issue where certain policy calls failed when at least one unconfigured cluster existed in the system.</td></tr></tbody></table>

### October 08

#### Product Enhancements

**Cluster diagnostics collection command** - Added a new CLI command, `runai diagnostics collect-logs`, which gathers diagnostic logs from the Kubernetes cluster for troubleshooting or sharing with NVIDIA Run:ai support. You can collect logs from all or specific namespaces, specify an output directory, and choose whether to include previous pod logs, simplifying cluster debugging and support workflows. See [runai diagnostics](/saas/reference/cli/runai/runai-diagnostics) command for more details.

#### Resolved Bugs

<table><thead><tr><th width="253.79296875">ID</th><th>Description</th></tr></thead><tbody><tr><td>RUN-32571</td><td>Fixed an issue where credentials that were not yet synced to the cluster appeared in the credential selection dropdown in Hugging Face and NIM inference workloads.</td></tr><tr><td>RUN-32652</td><td>Fixed an issue where YAML submitted workloads were not supported in batch deletion.</td></tr><tr><td>RUN-32605</td><td>Fixed a security vulnerability related to CVE-2025-58754 with severity HIGH.</td></tr><tr><td>RUN-32314</td><td>Fixed an issue where deleting a project did not remove access rules scoped to that project</td></tr></tbody></table>

## September 2025 Releases

### September 28

#### Product Enhancements

* **Guided onboarding for first-time admins** - A new onboarding flow helps system and platform administrators quickly get started by walking through cluster installation, setting up SSO and onboarding the first research team, reducing setup complexity and accelerating time to adoption.
* **Guided onboarding experience for new researchers** - On their first login, all new researchers are directed to the Workloads page and guided through creating their first Jupyter Notebook workspace with a short tour. A template is available for immediate launch, helping users get started quickly. The guided tour remains available anytime from the Help menu.
* **Workload extensibility with Karta** - Karta enables organizations to extend NVIDIA Run:ai with new workload types from any ML framework, tool, or Kubernetes resource using a no-code configuration through the Workload Types API. This allows organizations to incorporate emerging AI/ML tools or custom resources without platform updates or code changes. These workloads become immediately available across the organization, empowering teams to innovate and collaborate while benefiting from advanced scheduling and monitoring. See [Extending workload support with Karta](/saas/workloads-in-nvidia-run-ai/workload-types/extending-workload-support) for more details. <mark style="color:green;">`Experimental`</mark> `From cluster v2.23 onward`
  * No-code onboarding - Register new workload types instantly via the [Workload Types](https://run-ai-docs.nvidia.com/api/workloads/workload-properties#post-api-v1-workload-types) API.
  * Seamless researcher experience - Submit and run workloads using a standard YAML manifest via the [Workloads v2](https://run-ai-docs.nvidia.com/api/workloads/workloads-v2) API.
  * Unified management - Newly added workloads are available to all teams and benefit from the same orchestration and monitoring as native types.
  * Karta-powered integration - Defines how each workload is interpreted and optimized, enabling consistent support for scaling, dependencies, and advanced scheduling.
  * Newly added [workload types](/saas/workloads-in-nvidia-run-ai/workload-types/extending-workload-support#supported-workload-types) - NIM Services, KServe, and JobSet.
* **New workload template capabilities** - The new templates simplify the workload submission experience by allowing you to launch a workload in a single click, without modifying any settings. In addition, several supporting capabilities have been introduced. See [Workload templates](/saas/workloads-in-nvidia-run-ai/workload-templates) for more details. `From cluster v2.23 onward`
  * Preset templates - A set of ready-to-use workload templates for NeMo, BioNeMo and PyTorch is now available, enabling you to launch workloads quickly.
  * Linked assets - Templates can now be linked to assets such as environments and compute resources. Any changes to these assets are automatically reflected in the template, ensuring consistency across workloads.
  * Migrating legacy templates - Existing legacy templates can now be migrated into the new workload templates format, allowing teams to retain their saved configurations while taking advantage of new features. This capability is available when the Flexible workload templates setting is toggled on. You will not lose your existing templates - all legacy templates remain available.
* **NGC public registry support for environment images** - Environment images and tags can now be selected directly from the NGC public registry when creating [workloads](/saas/workloads-in-nvidia-run-ai/workloads), [environment](/saas/workloads-in-nvidia-run-ai/assets/environments) assets and [templates](/saas/workloads-in-nvidia-run-ai/workload-templates). This provides a streamlined way to access trusted NVIDIA containers without manually entering image URLs. <mark style="color:orange;">`Beta`</mark> `From cluster v2.23 onward`
* **Enhanced logging with per-container support** - Workload logs can be viewed at the container level within each pod through the UI, API and CLI, giving researchers and administrators finer control when monitoring and debugging workloads. In addition, downloaded logs are saved with unique file names that include the workload, pod, container, and timestamp, making it easier to organize and analyze logs from distributed workloads. See [Workloads](/saas/workloads-in-nvidia-run-ai/workloads#logs) for more details.
* **Networking metrics** - A new metric, NVLink bandwidth total, has been added to [Nodes](/saas/platform-management/aiinitiatives/resources/nodes#resource-utilization) and [Workloads](/saas/workloads-in-nvidia-run-ai/workloads#resource-utilization) views in the UI and is also available through the [Nodes](https://run-ai-docs.nvidia.com/api/organizations/nodes) and [Pods](https://run-ai-docs.nvidia.com/api/workloads/pods) APIs. This improves visibility into network utilization, giving teams deeper insight into consumption patterns and resource allocations. `From cluster v2.23 onward`
* **Enhanced Git credential management** - Git data sources can now be configured with Generic secret credentials through the UI or API, with support for SSH private keys. This provides a consistent and secure way to authenticate to repositories, simplifying setup for administrators and enabling users to connect to Git-based workflows more easily. See [Credentials](/saas/workloads-in-nvidia-run-ai/assets/credentials#generic-secret) for more details. `From cluster v2.23 onward`
* **Customize your CLI list views** - The new `--columns` flag allows you to tailor the output of `runai list` commands to display only the fields you need, giving you complete control over table views. See [CLI commands reference](/saas/reference/cli/runai) for more details.
  * Select and order columns - Define exactly which columns to display and in what order.
  * Discover more data - Show useful fields that are not part of the default output.
  * Autocompletion support - Use tab completion to discover and select all available columns for any list command.
* **Distributed inference API enhancements** - The inference API has been extended with support for multi-node deployments, adding autoscaling and rolling updates. These enhancements improve the robustness, scalability, and manageability of distributed inference workloads. See [Distributed inferences](https://run-ai-docs.nvidia.com/api/workloads/distributed-inferences) API for more details. `From cluster v2.22 onward`
* **Distributed inference support for GB200 and MNNVL** - Distributed inference workloads can now take advantage of NVIDIA GB200 NVL72 and other Multi-Node NVLink systems. This enables automatic infrastructure detection, domain labeling, and optimized cross-node communication for high-bandwidth, performance-optimized inference execution. See [Using GB200 NVL72 and Multi-Node NVLink domains](/saas/platform-management/aiinitiatives/resources/using-gb200) for more details. `From cluster v2.23 onward`
* **NVIDIA NIM service deployment API** - A new API is available for deploying NVIDIA NIM services, allowing programmatic creation and management of NIM service workloads for easier automation and integration. See [NVIDIA NIM](https://run-ai-docs.nvidia.com/api/workloads/nvidia-nim) API for more details. `From cluster v2.23 onward`
* **Support for Dynamo inference workloads** - Multi-node inference workloads deployed with the NVIDIA Dynamo framework can now be scheduled efficiently using gang scheduling and topology-aware scheduling. This ensures fast startup, low latency, and better resource utilization for disaggregated inference pipelines. <mark style="color:green;">`Experimental`</mark> `From cluster v2.23 onward`
* **Network topology-aware scheduling for distributed workloads** - NVIDIA Run:ai now supports topology-aware scheduling to optimize placement of distributed workloads across data center nodes. By leveraging Kubernetes node labels, the Scheduler can co-locate pods on nodes that are “closer” to each other in the network. This reduces communication overhead, improves workload efficiency, and helps maximize GPU utilization. Once administrators configure the network topology and associate it with node pools, scheduling is applied automatically for distributed workloads submitted through the platform. See [Accelerating workloads with network topology-aware scheduling](/saas/platform-management/aiinitiatives/resources/topology-aware-scheduling) for more details. `From cluster v2.23 onward`
* **Scoped access rules** - Users with permissions restricted to a specific scope are now limited to access rules within that scope. This capability is enabled via a tenant setting (`enable_scoped_authorization`) in the [Settings](https://run-ai-docs.nvidia.com/api/authentication-and-authorization/settings) API. Once enabled, the [Access rules](https://run-ai-docs.nvidia.com/api/authentication-and-authorization/access-rules) API returns only the rules within the viewer’s scope (or narrower), and the same scope filtering is applied when viewing access rules in the UI. This ensures access control is aligned with scope boundaries and prevents users from seeing or modifying rules outside their domain.
* **Cluster configuration via Helm values** - Cluster configurations can now be managed directly through the Helm values interface (`clusterConfig`). At runtime, `runaiconfig` is the actual source of truth, representing what is actively running in the cluster. When a Helm upgrade is performed, the Helm values overwrite the existing `runaiconfig`, ensuring alignment with the chart. As a result, clusters configured through a Helm chart should always be managed through Helm. This keeps configurations consistent and predictable across deployments and upgrades. See [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config) for more details. `From cluster v2.23 onward`
* **Component version updates** – NVIDIA Run:ai now supports Kubernetes version 1.34. Support for OpenShift version 4.15 has been removed. `From cluster v2.23 onward`
* **Support for ARM on OpenShift** - NVIDIA Run:ai now supports running on ARM-based nodes in OpenShift clusters, expanding deployment flexibility, and allowing organizations to leverage ARM architectures alongside existing x86 infrastructure within their OpenShift environments. `From cluster v2.23 onward`
* **Deleted workloads visible by default** - Deleted workloads are displayed by default in the UI under Workload manager. The toggle to enable this view has been removed, simplifying the experience and making it easier for users to track and review deleted workloads without extra configuration.
* **Direct tool connection** - When a workload has only one configured tool, clicking **Connect** opens the connection directly, without showing a selection menu. If multiple tools are configured, the selection menu will still appear.
* **Custom logo branding** - You can now upload a custom logo to appear in the top-right corner of the NVIDIA Run:ai platform interface. This allows organizations to personalize the platform UI with their own branding. Logos can be uploaded in SVG or PNG format (up to 128 KB) directly from the Branding settings.

#### Resolved Bugs

<table><thead><tr><th width="253.79296875">ID</th><th>Description</th></tr></thead><tbody><tr><td>RUN-32601</td><td>Fixed an issue where external token exchange failed because the API was incompatible with <code>access_tokens</code>.</td></tr><tr><td>RUN-31422</td><td>Fixed an issue where updating project resources created through the deprecated Projects API did not work correctly.</td></tr><tr><td>RUN-32551</td><td>Fixed an issue where inference workloads failed when using user credentials as an image pull secret.</td></tr><tr><td>RUN-32548</td><td>Fixed an issue where, in certain edge cases, removing an inference workload without deleting its revision caused the cluster to panic during revision sync.</td></tr><tr><td>RUN-32346</td><td>Fixed an issue where mappers could not be updated in identity providers (IdPs).</td></tr><tr><td>RUN-31993</td><td>Fixed a security vulnerability related to CVE-2025-22868 with severity HIGH.</td></tr><tr><td>RUN-31961</td><td>Fixed a security vulnerability related to CVE-2025-7425 with severity HIGH.</td></tr><tr><td>RUN-31051</td><td>Fixed a security vulnerability related to CVE-2025-49794 with severity HIGH.</td></tr><tr><td>RUN-31008</td><td>Fixed a security vulnerability related to CVE-2025-53547 with severity HIGH.</td></tr><tr><td>RUN-32123</td><td>Fixed an issue where email notifications configured through User settings were still sent after selecting and then immediately de-selecting all notification types.</td></tr><tr><td>RUN-32659</td><td>Fixed an issue where the search and filter logic for NIM models retrieved from the NGC catalog produced inconsistent results, causing some models to appear in unexpected positions in the list.</td></tr><tr><td>RUN-32699</td><td>Fixed an issue in the distributed inference policy API where some error messages displayed field names twice.</td></tr><tr><td>RUN-32789</td><td>Fixed an issue in CLI v2 where the <code>--master-extended-resource</code> flag had no effect in MPI training workloads.</td></tr><tr><td>RUN-30628</td><td>Fixed a security vulnerability related to CVE-2025-22874 with severity HIGH.</td></tr></tbody></table>

### September 16

#### Product Enhancements

* **AI Application-based workload grouping** - NVIDIA Run:ai now automatically groups related workloads into a single logical application for any workloads deployed via Helm charts. This provides a unified view of complex solutions. Using the API, you can track aggregated resource requests and allocations (GPU, CPU, memory) and monitor the overall application status. In the UI, you can filter the Workloads page by application name to easily see all components of a solution together. See [AI Applications](https://run-ai-docs.nvidia.com/api/ai-applications) API for more details.
* **Flexible inference workload templates** - Flexible workload templates allow you to save workload configurations that can be reused across workload submissions. You can create templates from scratch or base them on existing assets - environments, compute resources, or data sources. These templates simplify the submission process and promote standardization across users and teams. See [Inference templates](/saas/workloads-in-nvidia-run-ai/workload-templates/inference-templates) for more details.
* **Application access for inference serving endpoints** - All inference workloads - [custom](/saas/workloads-in-nvidia-run-ai/using-inference/custom-inference), [Hugging Face](/saas/workloads-in-nvidia-run-ai/using-inference/hugging-face-inference) and [NVIDIA NIM](/saas/workloads-in-nvidia-run-ai/using-inference/nim-inference), support authorizing applications (in addition to users and groups) when connecting to inference serving endpoints. This enables secure, programmatic access to inference endpoints when accessed externally from the cluster. To use this capability, configure the serving endpoint, authenticate using a token granted by an application, and use the token in API requests to the endpoint.
* **Credential creation during NIM and Hugging Face submissions** - You can now create My credentials of type Generic secret directly in the [NVIDIA NIM](/saas/workloads-in-nvidia-run-ai/using-inference/nim-inference) and [Hugging Face](/saas/workloads-in-nvidia-run-ai/using-inference/hugging-face-inference) inference workloads submission, avoiding the need to leave the flow to configure authentication.
* **NVIDIA NIM observability metrics** - Observability metrics are now available for NVIDIA NIM inference workloads via the UI and [Workloads](https://run-ai-docs.nvidia.com/api/workloads/workloads) / [Pods](https://run-ai-docs.nvidia.com/api/workloads/pods) APIs, giving teams better visibility into the performance of large language model (LLM) deployments. These metrics can be collected when deploying NIM through NVIDIA Run:ai, NIM operator, Helm chart, or directly via container images (with `run.ai/nim-workload: "true"` label). This enhancement enables more effective monitoring and troubleshooting of NIM-based inference workloads. See [Workloads](/saas/workloads-in-nvidia-run-ai/workloads#metrics) and [NIM observability metrics via API](https://run-ai-docs.nvidia.com/api/api-guides/nim-observability-metrics-via-api) for more details. `From cluster v2.23 onward`
* **Application access for workload tools** - Added support for authorizing applications (in addition to users and groups) when connecting to tools. This makes it easier to integrate external systems or services that need direct access to workload tools, providing more flexibility in how connections are managed.
* **PVC details view in data sources** - A new details pane is available when selecting a PVC data source from the Data sources table. The pane shows **Event History** for cluster events, as well as **Details** such as scope, request settings, and partial storage class information. This enhancement gives administrators and AI practitioners greater visibility into PVC usage history and configuration, improving monitoring and debugging. See [Data sources](/saas/workloads-in-nvidia-run-ai/assets/datasources#show-hide-details) for more details. `From cluster v2.23 onward`
* **System policies for workload governance** - By default, every NVIDIA Run:ai account is governed by system policies that establish foundational security controls across all workloads, scopes, and interfaces (UI, API and CLI). These policies ensure consistent workload behavior and prevent unauthorized escalation, and can be viewed as part of the effective policy for any scope. Administrators can create new policies to update these defaults at any desired scope. This flexibility allows easing certain API restrictions when needed, while ensuring every change is explicit and auditable. See [System policies](/saas/platform-management/policies/policies-and-rules#system-policies) for more details.
  * **Privileged parameter** - Set to `false` by default and not editable (`canEdit: false`), preventing containers from running with full host access unless explicitly enabled by an administrator.
  * **Grace period** - Defines how long a workload can continue running after a preemption request before termination. The default grace period is 30 seconds, with a system-enforced maximum of 5 minutes across UI, API and CLI submissions. This value can be updated at any scope within the policy hierarchy.
* **Policy synchronization changes** - Starting in version 2.23, control plane policies are no longer synchronized with the cluster. Policies are now stored and enforced only in the control plane, preventing conflicts with outdated cluster policies. See [Workload policies](/saas/platform-management/policies/native-workload-policies) for more details. `From cluster v2.23 onward`
* **Keyboard shortcuts for dialogs and forms** - Common keyboard actions across most UI screens and dialogs. Press **Enter** to confirm actions and **Esc** to cancel, making it quicker and easier to navigate workflows.
* **Updated General settings toggles** - The following options are now enabled by default - Flexible workload submission, Flexible workload templates, Data volumes, and Policies.
* **Metrics view updates** - The metrics view has been reorganized with new naming and grouping:
  * Renamed **Default** metrics view to **Resource utilization**
  * Renamed **Advanced** metrics view to **GPU profiling**
  * Inference metrics are shown in a dedicated **Inference** dropdown, available for all inference workloads

#### Resolved Bugs

<table><thead><tr><th width="253.79296875">ID</th><th>Description</th></tr></thead><tbody><tr><td>RUN-32656</td><td>Fixed an issue where the selected node pool was not preserved when switching sections within the workload submission form for all workloads.</td></tr><tr><td>RUN-32002</td><td>Fixed an issue where exported CSVs had misaligned columns, causing values (e.g., scope, workload type, creation time, cluster) to shift into incorrect fields.</td></tr><tr><td>RUN-32150</td><td>Fixed a security vulnerability related to CVE-2025-5914 with severity HIGH.</td></tr><tr><td>RUN-31797</td><td>Fixed a security vulnerability related to CVE-2025-53547 with severity HIGH.</td></tr></tbody></table>

## August 2025 Releases

### August 31

#### Product Enhancements

* **Expanded cluster role permissions** - Cluster roles have been updated to include `watch` permissions for all supported workload Custom Resource Definitions (CRDs) wherever `get` and `list` permissions were already present. This change ensures compatibility with Kubernetes operators that require `get`, `list`, and `watch` access for proper monitoring and integration with NVIDIA Run:ai workloads. `From cluster v2.21 onward`
* **Policy API for distributed inference** - A dedicated policy API is available for distributed inference enabling fine-grained control over distributed inference workloads. Administrators can define and enforce policies that govern scheduling, scaling, and update behavior, ensuring workloads adhere to organizational requirements and operate consistently across environments. See [Policy](https://run-ai-docs.nvidia.com/api/policies/policy) API for more details.
* **Removed General settings toggles** - The following options have been removed from the General settings page: Job submission, MPI distributed training, Weights & Biases SWEEP integration, and Docker image registry.

#### Resolved Bugs

<table><thead><tr><th width="253.79296875">ID</th><th>Description</th></tr></thead><tbody><tr><td>RUN-31860</td><td>Fixed a security vulnerability related to CVE-2025-47907 with severity HIGH.</td></tr><tr><td>RUN-31745</td><td>Fixed a bug which presented the value of the CPU memory in the wrong unit.</td></tr></tbody></table>

### August 17

#### Product Enhancements

* **Workloads by category over time** - Added a widget to the Overview dashboard that shows the number of workloads per category (e.g., Train, Build, Deploy) over time. This visualization helps identify usage trends, compare activity across categories, and track changes over specific periods. This feature is also supported in the [API](https://run-ai-docs.nvidia.com/api/workloads/workloads#get-api-v1-workloads-telemetry). `From cluster v2.22 onward`
* **Minimum guaranteed runtime for preemptible workloads** - You can now configure a minimum guaranteed runtime for preemptible workloads in node pools via the UI and [API](https://run-ai-docs.nvidia.com/api/organizations/nodepools#post-api-v1-node-pools). This setting specifies the minimum time a preemptible workload will run once scheduled and bound to a node before becoming eligible for preemption. This reduces unexpected interruptions and makes workload execution more predictable. See [Node pools](/saas/platform-management/aiinitiatives/resources/node-pools) for more details. `From cluster v2.23 onward`
* **Cluster filter enhancements for Nodes page** - The Nodes page now includes an “All” option in the clusters filter to make it easier to view and manage nodes across multiple clusters at once. When multiple clusters are selected, a Cluster column is displayed by default, showing each node’s associated cluster. Available in both the UI and API.
* **Separate admin toggles for Hugging Face and NVIDIA NIM models** - Previously, enabling Hugging Face and NVIDIA NIM models was managed through a single Models toggle in the Admin [settings](/saas/settings/general-settings). These options are now separated into distinct toggles, allowing administrators to enable or disable Hugging Face and NIM models independently for finer control over inference model availability.

#### Resolved Bugs

<table><thead><tr><th width="253.79296875">ID</th><th>Description</th></tr></thead><tbody><tr><td>RUN-31850</td><td>Fixed an issue where creating a workspace /training workload returned the error "terminationGracePeriod is not supported in this cluster.</td></tr><tr><td>RUN-31849</td><td>Fixed an issue where the non-preemptible priority over-quota warning text was missing from the inference workload creation page.</td></tr><tr><td>RUN-31579</td><td>Fixed an issue in the CLI documentation where the <code>--new-pvc</code> description did not clearly indicate that creating a new pvc means creating a new volume that is used only for the duration of the workload's lifecycle.</td></tr><tr><td>RUN-28394</td><td>Fixed an issue where the "Get Role by ID" API returned an "insufficient permissions" error for system administrator.</td></tr><tr><td>RUN-31304</td><td>Fixed a security vulnerability related to CVE-2025-22868 with severity HIGH.</td></tr><tr><td>RUN-31792</td><td>Fixed a security vulnerability related to CVE-2025-7425 with severity HIGH.</td></tr></tbody></table>

### August 03

#### Product Enhancements

* **Flexible submission form for NVIDIA NIM and Hugging Face workloads** - The flexible submission form is now supported for [NVIDIA NIM](/saas/workloads-in-nvidia-run-ai/using-inference/nim-inference) and [Hugging Face](/saas/workloads-in-nvidia-run-ai/using-inference/hugging-face-inference) inference workloads. This form allows users to submit workloads using an existing setup or provide custom settings for one-time use, enabling faster, more consistent submissions aligned with organizational policies.
* **Advanced setup form for NVIDIA NIM and Hugging Face workloads** - You can now access advanced configuration options when submitting [NVIDIA NIM](/saas/workloads-in-nvidia-run-ai/using-inference/nim-inference) and [Hugging Face](/saas/workloads-in-nvidia-run-ai/using-inference/hugging-face-inference) inference workloads, including editing the image and tag, modifying or adding environment variables, and setting workload priority. This provides greater flexibility for adapting workload configurations to specific requirements.
* **Dynamic NVIDIA NIM model list from NGC catalog** - The platform now retrieves the list of available NVIDIA NIM models directly from the NGC catalog using an API call. This ensures the model list remains current and reflects the latest offerings.

#### Resolved Bugs

<table><thead><tr><th width="253.79296875">ID</th><th>Description</th></tr></thead><tbody><tr><td>RUN-31392</td><td>Fixed an issue where the audit logs page filter converted strings to lowercase, causing filtration to fail.</td></tr><tr><td>RUN-31410</td><td>Fixed an issue where templates did not appear in the templates table.</td></tr><tr><td>RUN-31269</td><td>Fixed an issue where upgrades failed due to changes in the OpenShift monitoring stack.</td></tr><tr><td>RUN-31687</td><td>Fixed an issue where the workload flexible submission form did not load the correct default node pools for a project.</td></tr><tr><td>RUN-31504</td><td>Fixed an issue where workloads created via CLI could not be cloned in the UI when flexible submission was disabled.</td></tr><tr><td>RUN-30746</td><td>Fixed an issue where workloads could not be scheduled if the combined length of the project name and node pool name was excessively long.</td></tr><tr><td>RUN-31208</td><td>Fixed an issue where, in OpenShift environments, certain container failures caused workloads to remain in the "Pending" phase instead of transitioning to "Failed".</td></tr><tr><td>RUN-31358</td><td>Fixed an issue where enabling <code>enableWorkloadOwnershipProtection</code> for inference workloads caused newly submitted workloads to get stuck.</td></tr><tr><td>RUN-31252</td><td>Fixed an issue where the <code>terminationGracePeriodSeconds</code> field accepted values greater than 300 seconds when submitted via the API.</td></tr><tr><td>RUN-31263</td><td>Fixed an issue where setting defaults for <code>servingPort</code> fields failed and incorrectly required the container port default as well.</td></tr><tr><td>RUN-31265</td><td>Fixed a security vulnerability related to CVE-2025-30749 with severity HIGH.</td></tr><tr><td>RUN-31488</td><td>Fixed an issue where the UI logs view called an unsupported API. The fix added a cluster version check to ensure the correct API is used.</td></tr><tr><td>RUN-30918</td><td>Fixed an issue where the <code>createdAt</code> timestamp was not updated when a policy was recreated, causing the timestamp to incorrectly reflect the original creation time.</td></tr><tr><td>RUN-31039</td><td>Fixed a base image security vulnerability in <code>libxml2</code> related to CVE-2025-49796 with severity HIGH.</td></tr><tr><td>RUN-25973</td><td>Fixed an issue where some services were missing from cluster service groups.</td></tr><tr><td>RUN-31380</td><td>Fixed an issue where the SAML metadata XML redirect URL was invalid.</td></tr></tbody></table>

## July 2025 Releases

### July 20

#### Product Enhancements

**Improved status messaging for node pools with undrained nodes** - When creating a node pool or labeling nodes to add to the node pool, nodes that are not fully drained (i.e., still have running workloads) will now trigger clearer status messages in the API and UI. These messages indicate that the node pool cannot include the affected nodes until they are drained and reach a "Ready" state. This helps administrators better understand node pool readiness and identify which nodes are still in transition. `From cluster v2.23 onward`

#### Resolved Bugs

<table><thead><tr><th width="253.79296875">ID</th><th>Description</th></tr></thead><tbody><tr><td>RUN-31167</td><td>Removed Groups resource type from the available permissions in NVIDIA Run:ai roles.</td></tr><tr><td>RUN-31129</td><td>Fixed an issue where the Inference Policy View option was missing from the Project and Department pages.</td></tr><tr><td>RUN-31036</td><td>Fixed a security vulnerability in <code>runai-container-runtime-installer</code> and <code>runai-container-toolkit</code> related to CVE-2025-49794 with severity HIGH.</td></tr><tr><td>RUN-31066</td><td>Fixed an issue where the validation for the number of workers in a policy was not applied correctly.</td></tr><tr><td>RUN-30740</td><td>Fixed an issue where negative values were allowed for GPU resource optimization swap size in node pool API.</td></tr></tbody></table>

## Transition Notices

### July 2026

#### GPU Fractions: Transition to CUDA-Based GPU Sharing

NVIDIA Run:ai is transitioning GPU fractions from its current proprietary sharing engine to technology built directly into NVIDIA CUDA. This moves GPU sharing onto the standard NVIDIA stack for broader industry compatibility and long-term NVIDIA maintenance. The NVIDIA Run:ai experience, APIs, YAML definitions, and workload management workflows are not changing.

* **Current release (cluster v2.26)** - Uses the existing GPU fractions engine. No changes, no action required.
* **Future releases** - CUDA-based fractions will replace the legacy engine. This upgrade will require updated CUDA driver and GPU Operator versions on all GPU nodes. GPU fractions will be disabled until both requirements are in place; all other workloads will continue to run normally. Once the stack requirements are met, fractions re-enable automatically with no further Run:ai upgrade needed.

**What to do now**

* No action is required today; cluster v2.26 is fully supported.
* Review your current CUDA driver and GPU Operator versions to prepare for upcoming stack requirements.
* Plan a maintenance window for the future stack upgrade - GPU fractions will need to be suspended during the upgrade.
* For questions, contact your NVIDIA Run:ai account team.

## Deprecation Notifications

{% hint style="info" %}
**Note**

Deprecated features, APIs, and capabilities remain available for **six months** from the time of the deprecation notice, after which they may be removed.
{% endhint %}

### July 2026

#### GPU Memory SWAP

GPU Memory SWAP will be deprecated and removed from future releases. GPU Memory SWAP is being replaced by Snapshot, an improved NVIDIA-native capability built on CUDA Checkpoint / Restore. Snapshot captures the full state of a Kubernetes Pod, releases idle GPU and CPU memory, and reduces cold-start time when workloads resume. No immediate action is required as part of this release. GPU Memory SWAP remains fully supported in NVIDIA Run:ai cluster v2.26.

#### TensorFlow, XG Boost, and JAX

The TensorFlow, XG Boost, and JAX distributed frameworks are deprecated and will be removed in a future release. These frameworks are based on Kubeflow Training Operator v1 CRDs, which are deprecated upstream. As a result, submitting workloads using these frameworks as native NVIDIA Run:ai distributed workload types is also deprecated. See [Distributed training](/saas/getting-started/installation/install-using-helm/system-requirements#distributed-training) for supported frameworks.

#### Resource Interfaces to Kartas

The Resource Interface feature has been renamed to Karta across NVIDIA Run:ai. The `resourceInterfaces` API field is deprecated and will be removed in a future release. See [Extending workload support with Karta](/saas/workloads-in-nvidia-run-ai/workload-types/extending-workload-support) for more details.

#### Node Type

The Node Type feature is deprecated and will be removed in a future release. To control workload placement on specific nodes, use workload policies with node affinity instead of the `nodeType` field in organization unit or workload definitions. See [NVIDIA Run:ai native workload policies](/saas/platform-management/policies/native-workload-policies) for more details.

#### Legacy Workload Submission

The legacy workload submission form is deprecated and will be removed in a future release. Submitting workloads through the legacy form is being replaced by flexible workload submission, which lets you select an existing setup or start from scratch, review existing setups, and understand policy definitions. We recommend transitioning to flexible workload submission, which is enabled by default. If it has been disabled, re-enable it under **General settings** → Workloads → Flexible workload submission.

#### API Deprecation Notifications <a href="#api-deprecation-notifications-june-2026" id="api-deprecation-notifications-june-2026"></a>

**Deprecated Endpoints**

| Deprecated Endpoint          | Replacement Endpoint             |
| ---------------------------- | -------------------------------- |
| `/api/v1/logo`               | `/api/v1/branding/settings/logo` |
| `/api/v1/org-unit/node-type` | N/A                              |

**Deprecated Parameters**

| Endpoint                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                 | Deprecated Parameter                           | Replacement Parameter                                                                                       |
| ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| `/api/v2/workloads`                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      | `resourceInterfaces`                           | `kartas`                                                                                                    |
| `/api/v1/node-pools`                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     | `schedulingConfiguration.placementStrategy`    | <p><code>gpuPodsPlacementStrategy</code> + <code>cpuPodsPlacementStrategy</code><br></p>                    |
| `/api/v1/node-pools`                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     | `gpuResourceOptimization`                      | N/A                                                                                                         |
| <ul><li><code>/api/v1/workloads/distributed</code></li><li><code>/api/v2/policy/distributed</code></li><li><code>/api/v1/workload-templates/distributed</code></li></ul>                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                 | `DistributedFramework`: `TF`, `XGBoost`, `JAX` | N/A                                                                                                         |
| <ul><li><code>/api/v1/asset/workload-template</code></li><li><code>/api/v1/org-unit/departments</code></li><li><code>/api/v1/org-unit/projects</code></li><li><code>/api/v1/workload-templates</code></li><li><code>/api/v1/workload-templates/distributed</code></li><li><code>/api/v1/workload-templates/inferences</code></li><li><code>/api/v1/workload-templates/trainings</code></li><li><code>/api/v1/workload-templates/workspaces</code></li><li><code>/api/v1/workloads/distributed</code></li><li><code>/api/v1/workloads/distributed-inferences</code></li><li><code>/api/v1/workloads/inferences</code></li><li><code>/api/v1/workloads/trainings</code></li><li><code>/api/v1/workloads/workspaces</code></li><li><code>/api/v2/policy/distributed</code></li><li><code>/api/v2/policy/distributed-inferences</code></li><li><code>/api/v2/policy/inferences</code></li><li><code>/api/v2/policy/trainings</code></li><li><code>/api/v2/policy/workspaces</code></li></ul> | `nodeType`                                     | Use a workload policy node selector to target node types instead - `nodeAffinityRequired.nodeSelectorTerms` |

### April 2026

#### JFrog Artifacts

Support for JFrog as an artifact registry is now deprecated. JFrog will continue to be supported for existing deployments, but it will be removed in a future release within approximately one year. Customers are encouraged to migrate to NVIDIA NGC, which is now the recommended artifact registry for NVIDIA Run:ai.

#### TGI Server - Hugging Face

The Text Generation Inference (TGI) server used in the Hugging Face flow is deprecated. TGI is in maintenance mode and expected to reach end-of-life. Existing deployments will continue to work, but users should transition to other supported inference options.

#### Administrator CLI

The Administrator CLI (`runai-adm`) is deprecated. The two tasks it supported are now handled as follows:

* **Log collection** - Use `runai diagnostics collect-logs` to collect diagnostic logs from the Kubernetes cluster for troubleshooting or sharing with NVIDIA Run:ai Support. See [CLI command reference](/saas/reference/cli/runai/runai-diagnostics) for more details.
* **Node role management** - Use `kubectl` to configure node roles, which is already the recommended approach for node management. See [Node roles](/saas/infrastructure-setup/advanced-setup/node-roles) for more details.

#### Node Level Scheduler

The Node-level Scheduler feature is deprecated. This allows NVIDIA Run:ai to focus on more scalable and flexible GPU resource optimization capabilities. Existing deployments will continue to work, but the feature will be removed in a future release.

#### Department Administrator Role

The previous Department administrator role has been deprecated and renamed to Department administrator legacy. A new Department administrator role replaces it with the same permissions, except it no longer includes Read Clusters access. Existing role assignments are not affected, as roles are referenced by ID in all APIs.

#### API Deprecation Notifications <a href="#api-deprecation-notifications" id="api-deprecation-notifications"></a>

**Deprecated Endpoints**

| Deprecated Endpoint                                   | Replacement Endpoint                                                                           |
| ----------------------------------------------------- | ---------------------------------------------------------------------------------------------- |
| `/api/v1/token`                                       | `/api/v2/token`                                                                                |
| `POST /api/v1/notification-channels/slack/create-app` | `POST /api/v1/slack/app`                                                                       |
| `GET /api/v1/notification-channels/slack/add-app`     | `POST /api/v1/slack/workspace`                                                                 |
| `POST /api/v1/notification-channels/slack/auth-app`   | `GET /api/v1/slack/auth`                                                                       |
| `GET /api/v1/notification-state`                      | <p><code>GET /api/v1/email/state</code><br><code>GET /api/v1/slack/state</code></p>            |
| `PATCH /api/v1/notification-state`                    | <p><code>PATCH /api/v1/slack/state</code></p><p><code>PATCH /api/v1/email/state</code></p>     |
| `GET /api/v1/notification-state/detailed`             | `GET /api/v1/slack/state/detailed`                                                             |
| `POST /api/v1/validate-notification-channel`          | <p><code>POST /api/v1/slack/validate</code></p><p><code>POST /api/v1/email/validate</code></p> |
| `/api/v1/clusters/{clusterUuid}/nodes`                | `/api/v1/nodes`                                                                                |
| `/api/v1/org-unit/priorities`                         | `/api/v1/org-unit/ranks`                                                                       |
| `/api/v1/administration/user-applications`            | `/api/v1/administration/access-keys`                                                           |
| `/api/v1/administration/user-applications/{appId}`    | `/api/v1/administration/access-keys{accessKeyId}`                                              |

**Deprecated Parameters**

| Endpoint             | Deprecated Parameter    | Replacement Parameter                       |
| -------------------- | ----------------------- | ------------------------------------------- |
| `/api/v1/node-pools` | `overProvisioningRatio` | N/A                                         |
| `/api/v1/node-pools` | `placementStrategy`     | `schedulingConfiguration.placementStrategy` |
| `/api/v1/workloads`  | `externalConnections`   | `endpoints`                                 |
| `/api/v1/workloads`  | `urls`                  | `endpoints`                                 |

### January 2026

#### NVIDIA Run:ai Predefined Roles

The following predefined roles are deprecated in the UI and API. Review the new predefined roles to determine whether they meet your requirements, or create a custom role using the API. See [Roles](/saas/infrastructure-setup/authentication/roles) for more details:

* Compute resource administrator
* Credentials administrator
* Data source administrator
* Data volume administrator
* Environment administrator
* L1 researcher
* L2 researcher
* ML engineer
* Research manager
* Template administrator

During the deprecation period, the following predefined roles will be updated with minimal access to cluster and node pool data:

* Both cluster and node pool access are planned to transition to Clusters minimal and Node pools minimal - L1 researcher, L2 researcher, Research manager
* Cluster access is planned to transition to Clusters minimal while node pools access remains unchanged - ML engineer
* Cluster access is planned to transition to Clusters minimal - Compute resource administrator, Credentials administrator, Data source administrator, Data volume administrator, Environment administrator, Template administrator

#### Models Catalog

The Models catalog page is deprecated. Previously, the Models catalog provided a quick start experience for deploying a curated set of Hugging Face models. The same capability is now available through the Hugging Face inference workload flow, which integrates directly with Hugging Face and allows you to browse, select, and deploy any supported model from an open list. To deploy Hugging Face models, use the [Hugging Face inference workload](/saas/workloads-in-nvidia-run-ai/using-inference/hugging-face-inference) flow.

#### API Deprecation Notifications

**Deprecated Endpoints**

| Deprecated Endpoint                  | Replacement Endpoint          |
| ------------------------------------ | ----------------------------- |
| <p><code>/api/v1/apps</code><br></p> | `/api/v1/service-accounts`    |
| `/api/v1/user-applications`          | `/api/v1/access-keys`         |
| `/api/v1/authorization/roles`        | `/api/v2/authorization/roles` |

**Deprecated Parameters**

| Endpoint                                                                                                                                                           | Deprecated Parameter       | Replacement Parameter           |
| ------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -------------------------- | ------------------------------- |
| `api/v1/authorization/access-rules`                                                                                                                                | `subjectType: app` (enum)  | `subjectType: service-account`  |
| <ul><li><code>/api/v2/authorization/roles</code></li><li><code>/api/v1/authorization/roles</code></li><li><code>/api/v1/authorization/permissions</code></li></ul> | `resourceType: app` (enum) | `resourceType: service-account` |
| <ul><li><code>/api/v1/org-unit/projects</code></li><li><code>/api/v1/org-unit/departments</code></li></ul>                                                         | `resources: priority`      | `resources: rank`               |

#### CLI Deprecation Notifications

| CLI Command                   | Deprecated Parameter | Replacement Parameter    |
| ----------------------------- | -------------------- | ------------------------ |
| `runai login`                 | `application`        | `access-key`             |
| `runai login`                 | `user`               | `password`               |
| `runai [workloadtype] submit` | `--environment`      | `--environment-variable` |

### September 2025

#### Grafana Dashboards

The legacy Grafana dashboards - **Overview** and **Analytics** - are deprecated and will be removed in a future release. We recommend transitioning to the new dashboards available in the NVIDIA Run:ai UI, which are powered by NVIDIA Run:ai APIs. These dashboards provide improved visibility with drill-down capabilities and more flexibility for analyzing usage and performance.

{% hint style="info" %}
**Note**

The Consumption dashboard was deprecated in version [July 2025](#july-2025) and replaced with Reports.
{% endhint %}

#### CLI v1

CLI v1 was deprecated in [January 2025](#january-2025) and has now been fully removed from the platform. All command-line interactions should be performed using [CLI v2](/saas/reference/cli/runai).

{% hint style="info" %}
**Note**

CLI v1 will still be available for clusters below v2.18.
{% endhint %}

#### Jobs

The Jobs workload type was deprecated in [January 2025](#january-2025) and has now been fully removed from the platform. This means the Jobs option no longer appears in the **General settings** or in the Workloads page.

### July 2025

#### Consumption Dashboard <a href="#consumption-dashboard" id="consumption-dashboard"></a>

The Consumption dashboard is deprecated and replaced with [Reports](/saas/platform-management/monitor-performance/reports). Consumption reports provide improved visibility into resource usage with enhanced filtering and export capabilities. We recommend transitioning to consumption reports for the most up-to-date insights.

#### Templates <a href="#templates" id="templates"></a>

The Templates feature is deprecated. We recommend transitioning to [flexible workload templates](/saas/workloads-in-nvidia-run-ai/workload-templates), which offer enhanced functionality and support for flexible workload types - including workspace, standard training, and distributed training.

### April 2025

#### Cluster API for Workload Submission <a href="#cluster-api-for-workload-submission" id="cluster-api-for-workload-submission"></a>

Using the Cluster API to submit NVIDIA Run:ai workloads via YAML was [deprecated](https://docs.run.ai/v2.18/home/whats-new-2-18/#cluster-api-deprecation) starting from NVIDIA Run:ai version 2.18. For cluster version 2.18 and above, use the NVIDIA Run:ai REST API to submit workloads. The Cluster API documentation has also been removed.

### January 2025

#### Ongoing Dynamic MIG Deprecation Process <a href="#ongoing-dynamic-mig-deprecation-process" id="ongoing-dynamic-mig-deprecation-process"></a>

The Dynamic MIG deprecation process started in version 2.19. NVIDIA Run:ai supports standard MIG profiles as detailed in [Configuring NVIDIA MIG profiles](/saas/platform-management/aiinitiatives/resources/mig-profiles).

* Before upgrading to version 2.20, workloads submitted with Dynamic MIG and their associated node configurations must be removed
* In version 2.20, MIG was removed from the NVIDIA Run:ai UI under compute resources.
* In Q2/25 all ‘Dynamic MIG’ APIs and CLI commands will be fully deprecated. (it will fail)

#### CLI v1 Deprecation <a href="#cli-v1-deprecation" id="cli-v1-deprecation"></a>

CLI V1 is deprecated and no new features will be developed for it. It will remain available for use for the next two releases to ensure a smooth transition for all users. We recommend switching to **CLI v2**, which provides feature parity, backwards compatibility, and ongoing support for new enhancements. CLI v2 is designed to deliver a more robust, efficient, and user-friendly experience.

#### Legacy Jobs View Deprecation <a href="#legacy-jobs-view-deprecation" id="legacy-jobs-view-deprecation"></a>

The legacy **Jobs view** will be discontinued in favor of the more advanced **Workloads view**. The legacy submission form will still be accessible via the Workload manager view for a smoother transition.

#### appID and appSecret Deprecation <a href="#appid-and-appsecret-deprecation" id="appid-and-appsecret-deprecation"></a>

Deprecating appID and appSecret parameters used for [requesting an API token](https://run-ai-docs.nvidia.com/api/getting-started/how-to-authenticate-to-the-api). It will remain available for use for the next two releases. To create application tokens, use your client credentials - Client ID and Client secret.


# Quick Start Guides

This quick start helps you get up and running with NVIDIA Run:ai by guiding you through the main stages of platform adoption. Each guide is written for a specific role and provides step-by-step instructions.

## Install the Platform

**For infrastructure administrators**

Follow the instructions in the [Quick start for infrastructure administrators](/saas/getting-started/quick-starts/infra-admin-quick-start) to install NVIDIA Run:ai and complete the required post-installation infrastructure setup.

## Onboard Organization, Projects and Users

**For platform administrators**

Follow the instructions in the [Quick start for platform administrators](/saas/getting-started/quick-starts/platform-admin-quick-start) to complete the onboarding wizard, configure authentication and organizational structure, and prepare the platform for teams to run workloads.

## Build, Train and Deploy Models

**For AI practitioners**

Follow the instructions in the [Quick start for AI practitioners](/saas/getting-started/quick-starts/ai-practitioner-quick-start) to log in, select a project, and run experiments and production workloads on NVIDIA Run:ai.


# Quick Start for Infrastructure Administrators

This guide is for infrastructure administrators responsible for installing, configuring, and operating NVIDIA Run:ai.

The quick start walks through the initial infrastructure setup lifecycle, including platform installation and the essential post-installation configuration required to prepare the cluster for onboarding and workload execution. It focuses on infrastructure-level concerns such as cluster readiness, security boundaries, and operational stability.

## Prerequisites

Before you begin, ensure that:

* A Kubernetes cluster is up and running.
* [Helm](https://helm.sh/) 3.14 or later is installed.
* You have `kubectl` access to the cluster with admin-level permissions.

## Installation

The platform supports deployment using two primary methods, depending on your environment:

* [Install using Helm](/saas/getting-started/installation/install-using-helm) - The standard installation method using Helm charts. Provides full control and flexibility over configuration and deployment.
* [Install using Base Command Manager (BCM)](/saas/getting-started/installation/bcm-install) - A guided installation method available through NVIDIA Base Command Manager intended to simplify deployment, employing defaults meant to enable most NVIDIA Run:ai capabilities on NVIDIA DGX SuperPOD systems.

## Getting Started: The Onboarding Wizard

After installation, sign in to the NVIDIA Run:ai UI. The onboarding wizard launches automatically and guides you through the required steps.

The wizard includes both infrastructure-level and organizational steps. As an infrastructure administrator, you are responsible for completing the infrastructure-related steps and then handing off the remaining organizational setup to a [platform administrator](/saas/getting-started/quick-starts/platform-admin-quick-start#complete-the-onboarding-wizard).

{% hint style="info" %}
**Note**

Do not close the wizard before all steps are complete. The onboarding wizard cannot be reopened once dismissed.
{% endhint %}

### Connect Your Cluster

{% hint style="info" %}
**Note**

If the NVIDIA Run:ai cluster is already deployed and connected (e.g., via BCM installation or a pre-run Helm installation), the wizard will automatically detect the connection and skip this part.
{% endhint %}

A cluster is your organization’s compute infrastructure, where AI workloads are executed. The wizard first directs you to review [system](/saas/getting-started/installation/install-using-helm/system-requirements) and [network](/saas/getting-started/installation/install-using-helm/network-requirements) requirements. It then generates a Helm command that you run on your Kubernetes cluster to install the required components and prepare the cluster for workload scheduling.

The wizard displays **Waiting for cluster to connect** while the cluster is being installed and connected. Once the installation completes successfully and the cluster establishes communication with NVIDIA Run:ai, the wizard updates to **Cluster connected**. After completing the wizard flow, the cluster is added to the [Clusters](/saas/infrastructure-setup/procedures/clusters) table.

### Configure Platform Authentication

This step integrates NVIDIA Run:ai with your organization’s identity and access management system. Configure [Single Sign-On (SSO)](/saas/infrastructure-setup/authentication/overview#single-sign-on-sso) using SAML 2.0 or OpenID Connect (OIDC) to connect NVIDIA Run:ai to your corporate Identity Provider (IdP).

## Post Installation Infrastructure Setup

After installing NVIDIA Run:ai, complete the following foundational infrastructure configuration steps to ensure the platform is production-ready and can safely support organizational onboarding and workloads. These steps focus on cluster readiness and operational guardrails, rather than day-to-day platform usage:

* Validate node readiness and assign node roles as required
* Configure advanced cluster settings based on your environment requirements
* Enable required integrations and networking components
* Apply security and operational best practices
* Prepare the platform for scale, availability, and ongoing maintenance

The exact configuration required depends on your environment, scale, and operational model. Detailed procedures and advanced options are documented in the [Advanced setup](/saas/infrastructure-setup/advanced-setup) and [Infrastructure procedures](/saas/infrastructure-setup/procedures) sections.


# Quick Start for Platform Administrators

This guide is for platform administrators responsible for configuring, governing, and operating NVIDIA Run:ai across the organization.

The quick start outlines the critical, high-level setup phases you must complete immediately after NVIDIA Run:ai is installed. It focuses on establishing authentication, organizational structure, resource governance, and operational visibility required to enable teams to run workloads.

## Prerequisites

* Access to the NVIDIA Run:ai tenant - You have the NVIDIA Run:ai tenant URL for your organization (for example, `https://<your-domain>`).
* Administrator credentials - You have credentials with system administrator permissions to sign in to the NVIDIA Run:ai UI.

## Complete the Onboarding Wizard

After the infrastructure administrator completes the infrastructure-related steps of the onboarding wizard (cluster connection and authentication), responsibility is handed off to the platform administrator to complete the organizational setup.

### Onboard Your First Research Team

This final step establishes the organizational structure. Provide a team name and add the email addresses of the first team members. If SSO is configured, you can create the team using an identity provider group name instead of adding individual email addresses. A project is then created for the team, initial quotas and permissions are applied, and invitations are sent.

## Ongoing Platform Management

After completing the onboarding wizard, continue managing the platform through the following core activities.

### Configure Platform Behavior and Admin Settings

NVIDIA Run:ai provides global configuration options to control system-wide features and functionality.

Use the [General settings](/saas/settings/general-settings) in the Admin panel to tailor how the platform operates, including feature enablement and analytics behavior. These settings apply across all users and workloads and can be adjusted to align with organizational policies and operational requirements.

### Define Authorization and Access Control

Authorization ensures users can access only the features and resources required for their role. Define how users and teams access platform capabilities using [roles](/saas/infrastructure-setup/authentication/roles) and [access rules](/saas/infrastructure-setup/authentication/accessrules) to:

* Grant users the appropriate level of access
* Control who can submit workloads or manage resources
* Extend access by creating additional [custom roles](/saas/infrastructure-setup/authentication/roles#custom-roles-api-only) as platform usage evolves

### Define Organizational Structure and Quota

Create additional [departments](/saas/platform-management/aiinitiatives/organization/departments) and [projects](/saas/platform-management/aiinitiatives/organization/projects) to logically partition resources. For each project, define guaranteed and over-quota resource allocations to control how resources are shared across the organization. At a high level, these constructs often align with an organization’s internal structure, such as business units, teams, or similar groupings. You can model them to reflect how your organization plans and manages AI initiatives. See [Adapting AI initiatives to your organization](/saas/platform-management/aiinitiatives/adapting-ai-initiatives) for more details.

### Configure Node Pools

Node pools allow you to translate organizational and business priorities into infrastructure-level scheduling decisions. By grouping worker nodes based on hardware type, capabilities, or location (for example, H100 vs. A100 GPUs), [node pools](/saas/platform-management/aiinitiatives/resources/node-pools) enable you to:

* Reserve high-end or scarce GPUs for business-critical or production workloads
* Provide predictable performance and scheduling behavior guarantees for prioritized projects
* Prevent lower-priority or experimental workloads from impacting production or revenue-generating use cases

Node pools form a foundational layer for enforcing resource access, workload placement, and scheduling policies across the organization.

### Define Policies (Governance)

Policies define how workloads are governed and scheduled across the platform. By combining workload policies with scheduling rules, you can standardize workload behavior and control how resources are used and shared.

Use policies to:

* Standardize workload behavior with [workload policies](/saas/platform-management/policies/native-workload-policies), enforcing best practices and organizational limits
* Define scheduling behavior with [scheduling rules](/saas/platform-management/policies/scheduling-rules) that determine how workloads are placed or how long they can run

Scheduling rules are applied at the project or department level and affect all matching workloads for that scope.

### Monitor and Optimize the Platform

Monitor platform usage and health to ensure efficient and reliable operation.

* Monitor usage - Use analytics dashboards to track GPU utilization, identify idle resources, and review consumption by department and project.
* Review logs and system health - Monitor control plane and cluster components to proactively troubleshoot issues and manage maintenance activities.


# Quick Start for AI Practitioners

This guide is for AI practitioners responsible for running experiments and production workloads on NVIDIA Run:ai.

The quick start walks through the essential steps to begin using the platform, from initial access and project selection to launching a workspace and submitting your first workloads. The focus is on day-to-day workload execution and resource consumption, so you can experiment, train models, and deploy inference within your assigned project.

## Prerequisites

To begin, ensure you meet the following conditions set up by your platform administrator:

* You have an active user account and credentials to access the NVIDIA Run:ai UI
* You are assigned to at least one project
* Your project has available resources to run workloads

## Getting Started

Choose a quick start based on your goal. Each scenario walks through a practical example so you can validate access, confirm resource availability, and understand how workloads run in your environment.

* [Run your first workspace](/saas/workloads-in-nvidia-run-ai/using-workspaces/quick-starts/jupyter-quickstart) - Launch a Jupyter notebook workspace for interactive development and experimentation. A guided tour is also available in the UI to help you familiarize yourself with the workspace experience.
* [Run a standard training workload](/saas/workloads-in-nvidia-run-ai/using-training/quick-starts/standard-training-quickstart) - Submit a standard training job to run a model training script on a single GPU.
* [Run a distributed training workload](/saas/workloads-in-nvidia-run-ai/using-training/quick-starts/distributed-training-quickstart) - Submit a distributed PyTorch training job and launch a multi-node training workload using an example PyTorch image.
* [Run a custom inference workload](/saas/workloads-in-nvidia-run-ai/using-inference/quick-starts/inference-quickstart) - Submit an inference workload and query the inference server to verify it is serving requests correctly.

## Understand Workload Capabilities

After completing the quick starts, explore the broader workload capabilities available in NVIDIA Run:ai. This helps you move beyond basic scenarios and take advantage of advanced scheduling, scaling, and configuration options.

* [Introduction to workloads](/saas/workloads-in-nvidia-run-ai/introduction-to-workloads) - How workloads are defined, scheduled, and executed in NVIDIA Run:ai.
* [Workload types and features](/saas/workloads-in-nvidia-run-ai/workload-types) - The different supported workload types and the capabilities available for each, including scaling, resource configuration, scheduling behavior, and other advanced options.
* [Workload assets](/saas/workloads-in-nvidia-run-ai/assets) - Shared resources used by workloads, such as environments, data sources, and credentials.
* [Workload templates](/saas/workloads-in-nvidia-run-ai/workload-templates) - Reusable configurations that help standardize and simplify workload creation.

## Run Workloads for Your Use Case

Once you understand the supported workload types and configuration options, proceed to the workload-specific documentation to configure and run workloads tailored for your project. Each workload section includes complete configuration examples and step-by-step instructions for the UI, API, and CLI.

* [Workspace](/saas/workloads-in-nvidia-run-ai/using-workspaces/running-workspace) - Interactive development environment for building and testing. Recommended for lightweight experimentation and debugging.
* [Training](/saas/workloads-in-nvidia-run-ai/using-training/train-models) - Workload for standard or distributed training models. Recommended for resource-intensive model development.
* [Inference](/saas/workloads-in-nvidia-run-ai/using-inference/nvidia-run-ai-inference-overview) - Deployment of an AI model for serving via an API. Recommended for production use.
* [Via YAML](/saas/workloads-in-nvidia-run-ai/submit-via-yaml) - Submission of a range of supported workload types using a standard Kubernetes YAML.

## Tutorials for End-to-End Workflows

Full end-to-end tutorials are available for deeper learning. These guides provide complete, practical examples that walk through development, training, and deployment workflows, showing how NVIDIA Run:ai features work together in real-world scenarios. See [Tutorials](/saas/tutorials/inference-tutorials) for more details.


# Installation

Run:ai is a Kubernetes-native orchestration and management platform designed to maximize GPU utilization for AI workloads.

## NVIDIA Run:ai System Components

NVIDIA Run:ai is made up of two components both installed over a [Kubernetes](https://kubernetes.io/) cluster:

* **NVIDIA Run:ai control plane** - Provides resource management, handles workload submission and provides cluster monitoring and analytics.
* **NVIDIA Run:ai cluster** - Provides enhanced scheduling and workload management, extending Kubernetes native capabilities.

As part of the installation process, you will:

* Create an NVIDIA Run:ai control plane tenant, provisioned and managed by NVIDIA
* Install one or more NVIDIA Run:ai clusters on your Kubernetes infrastructure and connect them to the tenant

<figure><img src="/files/tfgwbkPPlwNdzLYSOqxM" alt="" width="375"><figcaption></figcaption></figure>

## SaaS Deployment Model

The SaaS option is for organizations that prefer a fully managed control plane. With this model, NVIDIA hosts and operates the NVIDIA Run:ai control plane on your behalf. You are responsible only for installing the NVIDIA Run:ai cluster on your own Kubernetes infrastructure.

| Aspect        | Description                                                                     |
| ------------- | ------------------------------------------------------------------------------- |
| Control plane | Hosted and managed by NVIDIA. Access is provisioned via the NVIDIA NGC catalog. |
| Cluster       | Installed and managed by you on your Kubernetes infrastructure.                 |
| Connectivity  | The cluster connects to the NVIDIA-hosted control plane over the internet.      |


# Support Matrix

The support matrix outlines the verified compatibility standards for NVIDIA Run:ai SaaS. To ensure a stable and performant deployment, all infrastructure components, including Kubernetes/OpenShift distributions, NVIDIA Operators, and specialized frameworks, must align with the versions specified below. Use this matrix as a validation checklist prior to performing new installations or upgrades.

## Operator and Framework Versions <a href="#operator-and-framework-versions" id="operator-and-framework-versions"></a>

{% tabs %}
{% tab title="v2.26 (latest)" %}

| Component                                                                                                                                    | Supported Versions   |
| -------------------------------------------------------------------------------------------------------------------------------------------- | -------------------- |
| [NVIDIA GPU Operator](/saas/getting-started/installation/install-using-helm/system-requirements#nvidia-gpu-operator)                         | 25.10-26.3           |
| [NVIDIA Network Operator](/saas/getting-started/installation/install-using-helm/system-requirements#nvidia-network-operator)                 | 25.10-26.1           |
| [NVIDIA DRA driver](/saas/getting-started/installation/install-using-helm/system-requirements#nvidia-dynamic-resource-allocation-dra-driver) | 25.8 - 25.12         |
| [Prometheus / Kube‑Prometheus Stack](/saas/getting-started/installation/install-using-helm/system-requirements#prometheus)                   | 3.5 / 76.0 and above |
| [Kubeflow Training Operator](/saas/getting-started/installation/install-using-helm/system-requirements#distributed-training)                 | 1.9.2                |
| [MPI Operator](/saas/getting-started/installation/install-using-helm/system-requirements#distributed-training)                               | 0.6.0 or later       |
| [Knative Serving](/saas/getting-started/installation/install-using-helm/system-requirements#inference)                                       | 1.19 – 1.21          |
| [Leader-Worker Set (LWS)](/saas/getting-started/installation/install-using-helm/system-requirements#distributed-inference)                   | 0.7.0 or higher      |
| {% endtab %}                                                                                                                                 |                      |

{% tab title="v2.25" %}

| Component                                                                                                                                    | Supported Versions   |
| -------------------------------------------------------------------------------------------------------------------------------------------- | -------------------- |
| [NVIDIA GPU Operator](/saas/getting-started/installation/install-using-helm/system-requirements#nvidia-gpu-operator)                         | 25.10-26.3           |
| [NVIDIA Network Operator](/saas/getting-started/installation/install-using-helm/system-requirements#nvidia-network-operator)                 | 25.10-26.1           |
| [NVIDIA DRA driver](/saas/getting-started/installation/install-using-helm/system-requirements#nvidia-dynamic-resource-allocation-dra-driver) | 25.8 - 25.12         |
| [Prometheus / Kube‑Prometheus Stack](/saas/getting-started/installation/install-using-helm/system-requirements#prometheus)                   | 3.5 / 76.0 and above |
| [Kubeflow Training Operator](/saas/getting-started/installation/install-using-helm/system-requirements#distributed-training)                 | 1.9.2                |
| [MPI Operator](/saas/getting-started/installation/install-using-helm/system-requirements#distributed-training)                               | 0.6.0 or later       |
| [Knative Serving](/saas/getting-started/installation/install-using-helm/system-requirements#inference)                                       | 1.19 – 1.21          |
| [Leader-Worker Set (LWS)](/saas/getting-started/installation/install-using-helm/system-requirements#distributed-inference)                   | 0.7.0 or higher      |
| {% endtab %}                                                                                                                                 |                      |

{% tab title="v2.24" %}

| Component                                                                                                                                    | Supported Versions   |
| -------------------------------------------------------------------------------------------------------------------------------------------- | -------------------- |
| [NVIDIA GPU Operator](/saas/getting-started/installation/install-using-helm/system-requirements#nvidia-gpu-operator)                         | 25.3-25.10           |
| [NVIDIA Network Operator](/saas/getting-started/installation/install-using-helm/system-requirements#nvidia-network-operator)                 | 24.4-25.10           |
| [NVIDIA DRA driver](/saas/getting-started/installation/install-using-helm/system-requirements#nvidia-dynamic-resource-allocation-dra-driver) | 25.3-25.8            |
| [Prometheus / Kube‑Prometheus Stack](/saas/getting-started/installation/install-using-helm/system-requirements#prometheus)                   | 3.5 / 76.0 and above |
| [Kubeflow Training Operator](/saas/getting-started/installation/install-using-helm/system-requirements#distributed-training)                 | 1.9.2                |
| [MPI Operator](/saas/getting-started/installation/install-using-helm/system-requirements#distributed-training)                               | 0.6.0 or later       |
| [Knative Serving](/saas/getting-started/installation/install-using-helm/system-requirements#inference)                                       | 1.11 – 1.18          |
| [Leader-Worker Set (LWS)](/saas/getting-started/installation/install-using-helm/system-requirements#distributed-inference)                   | 0.7.0 or higher      |
| {% endtab %}                                                                                                                                 |                      |

{% tab title="v2.23" %}

| Component                                                                                                                                    | Supported Versions   |
| -------------------------------------------------------------------------------------------------------------------------------------------- | -------------------- |
| [NVIDIA GPU Operator](/saas/getting-started/installation/install-using-helm/system-requirements#nvidia-gpu-operator)                         | 25.3-25.10           |
| [NVIDIA Network Operator](/saas/getting-started/installation/install-using-helm/system-requirements#nvidia-network-operator)                 | 24.4-25.10           |
| [NVIDIA DRA driver](/saas/getting-started/installation/install-using-helm/system-requirements#nvidia-dynamic-resource-allocation-dra-driver) | 25.3-25.8            |
| [Prometheus / Kube‑Prometheus Stack](/saas/getting-started/installation/install-using-helm/system-requirements#prometheus)                   | 3.5 / 76.0 and above |
| [Kubeflow Training Operator](/saas/getting-started/installation/install-using-helm/system-requirements#distributed-training)                 | 1.9.2                |
| [MPI Operator](/saas/getting-started/installation/install-using-helm/system-requirements#distributed-training)                               | 0.6.0 or later       |
| [Knative Serving](/saas/getting-started/installation/install-using-helm/system-requirements#inference)                                       | 1.11 – 1.18          |
| [Leader-Worker Set (LWS)](/saas/getting-started/installation/install-using-helm/system-requirements#distributed-inference)                   | 0.7.0 or higher      |
| {% endtab %}                                                                                                                                 |                      |

{% tab title="v2.22" %}

| Component                                                                                                                                    | Supported Versions    |
| -------------------------------------------------------------------------------------------------------------------------------------------- | --------------------- |
| [NVIDIA GPU Operator](/saas/getting-started/installation/install-using-helm/system-requirements#nvidia-gpu-operator)                         | 24.9-25.3             |
| [NVIDIA Network Operator](/saas/getting-started/installation/install-using-helm/system-requirements#nvidia-network-operator)                 | 24.4-25.10            |
| [NVIDIA DRA driver](/saas/getting-started/installation/install-using-helm/system-requirements#nvidia-dynamic-resource-allocation-dra-driver) | 25.3-25.8             |
| [Prometheus / Kube‑Prometheus Stack](/saas/getting-started/installation/install-using-helm/system-requirements#prometheus)                   | 2.53 / 61.0 and above |
| [Kubeflow Training Operator](/saas/getting-started/installation/install-using-helm/system-requirements#distributed-training)                 | 1.9.2                 |
| [MPI Operator](/saas/getting-started/installation/install-using-helm/system-requirements#distributed-training)                               | 0.6.0 or later        |
| [Knative Serving](/saas/getting-started/installation/install-using-helm/system-requirements#inference)                                       | 1.11 – 1.18           |
| {% endtab %}                                                                                                                                 |                       |
| {% endtabs %}                                                                                                                                |                       |

## Supported NVIDIA GPUs

NVIDIA Run:ai is compatible with all Data Center GPUs supported by the NVIDIA GPU Operator. Hardware compatibility is determined by the specific version of the GPU Operator deployed within your cluster.

* Supported operator range - See the [Operator and framework versions](#operator-and-framework-versions) section to determine the supported GPU Operator version per NVIDIA Run:ai cluster.
* Hardware verification - To confirm if a specific GPU model is supported, please cross-reference your Operator version with the [Supported NVIDIA Data Center GPUs and Systems](https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/platform-support.html#supported-nvidia-data-center-gpus-and-systems) documentation.

{% hint style="info" %}
**Note**

* NVIDIA DGX Spark, NVIDIA Jetson and workstations are not supported.
* vGPU is not supported. NVIDIA Run:ai currently supports GPU passthrough only.
* In addition to GPU Operator compatibility, support also depends on the framework being used (for example, CUDA, PyTorch, TensorFlow, or NVIDIA NIM). Before running a workload, verify that the selected framework version supports your target GPU architecture according to the relevant support matrix.
  {% endhint %}

## NVIDIA Run:ai Compatible Distributions

NVIDIA Run:ai supports a wide range of Kubernetes distributions across on-premises, hybrid, and public cloud environments. Use the table below to verify version requirements for your specific platform.

{% tabs %}
{% tab title="v2.26 (latest)" %}

| Orchestration Platform                 | Versions                         | Engine     | x86       | ARM       |
| -------------------------------------- | -------------------------------- | ---------- | --------- | --------- |
| Upstream Kubernetes                    | 1.34-1.36                        | Containerd | Supported | Supported |
| Red Hat OpenShift                      | 4.18-4.21                        | CRI-O      | Supported | Supported |
| Amazon Elastic Kubernetes Engine (EKS) | Based on the upstream Kubernetes | Containerd | Supported | Supported |
| Google Kubernetes Engine (GKE)         | Based on the upstream Kubernetes | Containerd | Supported | Supported |
| Azure Kubernetes Service (AKS)         | Based on the upstream Kubernetes | Containerd | Supported | Supported |
| Oracle Kubernetes Engine (OKE)         | Based on the upstream Kubernetes | Containerd | Supported | Supported |
| Rancher Kubernetes Engine 2 (RKE2)     | Based on the upstream Kubernetes | Containerd | Supported | Supported |
| {% endtab %}                           |                                  |            |           |           |

{% tab title="v2.25" %}

| Orchestration Platform                 | Versions                         | Engine     | x86       | ARM       |
| -------------------------------------- | -------------------------------- | ---------- | --------- | --------- |
| Upstream Kubernetes                    | 1.33-1.35                        | Containerd | Supported | Supported |
| Red Hat OpenShift                      | 4.18-4.21                        | CRI-O      | Supported | Supported |
| Amazon Elastic Kubernetes Engine (EKS) | Based on the upstream Kubernetes | Containerd | Supported | Supported |
| Google Kubernetes Engine (GKE)         | Based on the upstream Kubernetes | Containerd | Supported | Supported |
| Azure Kubernetes Service (AKS)         | Based on the upstream Kubernetes | Containerd | Supported | Supported |
| Oracle Kubernetes Engine (OKE)         | Based on the upstream Kubernetes | Containerd | Supported | Supported |
| Rancher Kubernetes Engine 2 (RKE2)     | Based on the upstream Kubernetes | Containerd | Supported | Supported |
| {% endtab %}                           |                                  |            |           |           |

{% tab title="v2.24" %}

| Orchestration Platform                 | Versions                         | Engine     | x86       | ARM       |
| -------------------------------------- | -------------------------------- | ---------- | --------- | --------- |
| Upstream Kubernetes                    | 1.33-1.35                        | Containerd | Supported | Supported |
| Red Hat OpenShift                      | 4.17-4.20                        | CRI-O      | Supported | Supported |
| Amazon Elastic Kubernetes Engine (EKS) | Based on the upstream Kubernetes | Containerd | Supported | Supported |
| Google Kubernetes Engine (GKE)         | Based on the upstream Kubernetes | Containerd | Supported | Supported |
| Azure Kubernetes Service (AKS)         | Based on the upstream Kubernetes | Containerd | Supported | Supported |
| Oracle Kubernetes Engine (OKE)         | Based on the upstream Kubernetes | Containerd | Supported | Supported |
| Rancher Kubernetes Engine 2 (RKE2)     | Based on the upstream Kubernetes | Containerd | Supported | Supported |
| {% endtab %}                           |                                  |            |           |           |

{% tab title="v2.23" %}

| Orchestration Platform                 | Versions                         | Engine     | x86       | ARM       |
| -------------------------------------- | -------------------------------- | ---------- | --------- | --------- |
| Upstream Kubernetes                    | 1.31-1.34                        | Containerd | Supported | Supported |
| Red Hat OpenShift                      | 4.16-4.19                        | CRI-O      | Supported | Supported |
| Amazon Elastic Kubernetes Engine (EKS) | Based on the upstream Kubernetes | Containerd | Supported | Supported |
| Google Kubernetes Engine (GKE)         | Based on the upstream Kubernetes | Containerd | Supported | Supported |
| Azure Kubernetes Service (AKS)         | Based on the upstream Kubernetes | Containerd | Supported | Supported |
| Oracle Kubernetes Engine (OKE)         | Based on the upstream Kubernetes | Containerd | Supported | Supported |
| Rancher Kubernetes Engine 2 (RKE2)     | Based on the upstream Kubernetes | Containerd | Supported | Supported |
| {% endtab %}                           |                                  |            |           |           |

{% tab title="v2.22" %}

| Orchestration Platform                 | Versions                         | Engine     | x86       | ARM           |
| -------------------------------------- | -------------------------------- | ---------- | --------- | ------------- |
| Upstream Kubernetes                    | 1.31-1.33                        | Containerd | Supported | Supported     |
| Red Hat OpenShift                      | 4.15-4.19                        | CRI-O      | Supported | Not supported |
| Amazon Elastic Kubernetes Engine (EKS) | Based on the upstream Kubernetes | Containerd | Supported | Supported     |
| Google Kubernetes Engine (GKE)         | Based on the upstream Kubernetes | Containerd | Supported | Supported     |
| Azure Kubernetes Service (AKS)         | Based on the upstream Kubernetes | Containerd | Supported | Supported     |
| Oracle Kubernetes Engine (OKE)         | Based on the upstream Kubernetes | Containerd | Supported | Supported     |
| Rancher Kubernetes Engine 2 (RKE2)     | Based on the upstream Kubernetes | Containerd | Supported | Supported     |
| {% endtab %}                           |                                  |            |           |               |
| {% endtabs %}                          |                                  |            |           |               |

For existing Kubernetes clusters, see the following Kubernetes version support matrix for the latest NVIDIA Run:ai cluster releases:

| NVIDIA Run:ai version | Supported Kubernetes versions | Supported OpenShift versions |
| --------------------- | ----------------------------- | ---------------------------- |
| 2.26 (latest)         | 1.34 to 1.36                  | 4.18 to 4.21                 |
| 2.25                  | 1.33 to 1.35                  | 4.18 to 4.21                 |
| 2.24                  | 1.33 to 1.35                  | 4.17 to 4.20                 |
| 2.23                  | 1.31 to 1.34                  | 4.16 to 4.19                 |
| 2.22                  | 1.31 to 1.33                  | 4.15 to 4.19                 |

For information on supported versions of managed Kubernetes, it's important to consult the release notes provided by your Kubernetes service provider. There, you can confirm the specific version of the underlying Kubernetes platform supported by the provider, ensuring compatibility with NVIDIA Run:ai. For an up-to-date end-of-life statement see [Kubernetes Release History](https://kubernetes.io/releases/) or [OpenShift Container Platform Life Cycle Policy](https://access.redhat.com/support/policy/updates/openshift).

## Partner-Compatible Distributions

The following Kubernetes distributions are **partner-compatible**. They are tested and validated by the partner, who is responsible for maintaining compatibility with NVIDIA Run:flag\_ai:

| Kubernetes distribution                                                                        | NVIDIA Run:ai version               | Supported Kubernetes versions                     |
| ---------------------------------------------------------------------------------------------- | ----------------------------------- | ------------------------------------------------- |
| Crusoe Managed Kubernetes (CMK)                                                                | 2.22                                | 1.33                                              |
| [Mirantis k0rdent](https://catalog.k0rdent.io/v1.7.0/apps/runai-cp/)                           | <ul><li>2.22</li><li>2.23</li></ul> | <ul><li>1.32-1.33</li><li>1.33-1.34</li></ul>     |
| [Rafay platform](https://docs.rafay.co/aiml/app_marketplace/helm_app/apps/runai/requirements/) | <ul><li>2.23</li><li>2.24</li></ul> | <ul><li>1.33 - 1.34</li><li>1.33 - 1.35</li></ul> |
| [vCluster](https://www.vcluster.com/docs/platform/integrations/certified-stacks/runai)         | 2.24                                | 1.34                                              |
| VMware vSphere Kubernetes Service (VKS)                                                        | <ul><li>2.22</li><li>2.24</li></ul> | <ul><li>1.33</li><li>1.35</li></ul>               |


# Install Using Helm


# Create and Access the Control Plane Tenant

To access NVIDIA Run:ai SaaS, create a control plane tenant through the [NVIDIA GPU Cloud (NGC](https://org.ngc.nvidia.com)) catalog. The tenant is your organization's dedicated, NVIDIA-hosted Run:ai environment. The control plane provides the central management layer for NVIDIA Run:ai, handling multi-cluster management, resource and access management, as well as workload submission and monitoring. NVIDIA hosts and maintains the control plane. Once the tenant is provisioned, you can install the NVIDIA Run:ai cluster and connect it to the control plane, making it ready to run training, inference, and other workloads.

{% hint style="info" %}
**Note**

To create a control plane tenant, your NGC account must have an active entitlement for NVIDIA Run:ai SaaS.
{% endhint %}

## Prerequisites

An active **NVIDIA Run:ai SaaS** entitlement in your NGC organization. To activate your entitlement, see [Activating an NGC Product from an NVIDIA Commercial Entitlement Certificate](https://docs.nvidia.com/ngc/latest/ngc-user-guide.html#activating-ngc-product-from-an-nvidia-commercial-entitlement-certificate).

## Provision the Control Plane Tenant

{% hint style="info" %}
**Note**

To delegate provisioning to another user, go to [org.ngc.nvidia.com/users](https://org.ngc.nvidia.com/users) and assign them the **NVIDIA Run:ai SaaS Admin** role. They can then go to [runai.ngc.nvidia.com](https://runai.ngc.nvidia.com/) to perform the provisioning steps.
{% endhint %}

1. Log in to [NGC](https://org.ngc.nvidia.com) and navigate to **Catalog** -> **NVIDIA Run:ai SaaS**.

<figure><img src="/files/fOM24hgadkqyuoijdhFD" alt="" width="161"><figcaption></figcaption></figure>

2. In the **Create Control Plane Tenant** dialog, fill in the following:
   * **Tenant Admin Email** - Pre-filled with your NGC account email. You can assign additional admins after provisioning is complete.
   * **Tenant Alias** - A unique subdomain name for your tenant URL (for example, `myorg` results in `myorg.nv.run.ai`).

<figure><img src="/files/X0R6YASu8AG4oEYUualV" alt="" width="563"><figcaption></figcaption></figure>

3. Click **Create Control Plane Tenant**. NGC provisions your tenant. This may take a few minutes.

<figure><img src="/files/0ALGo2JTZpMKjM5goqkY" alt="" width="563"><figcaption></figcaption></figure>

Once provisioning is complete, the following details are displayed:

* **Tenant Name** - An auto-generated identifier for the tenant.
* **Tenant Alias** - The alias you provided.
* **Tenant Domain** - The URL of your NVIDIA Run:ai control plane, based on your alias (for example, `myorg.nv.run.ai`).
* **Created** - The date and time the tenant was created.

{% hint style="info" %}
**Note**

This is the only time your **Tenant Alias** and **Tenant Domain** are displayed. If you want to access your tenant directly using the Tenant Domain URL, make a note of it now. Alternatively, you can always access your tenant by navigating to [runai.ngc.nvidia.com](https://runai.ngc.nvidia.com), which automatically redirects you.
{% endhint %}

An email is sent to the Tenant Admin Email address with a link to activate your account.

<figure><img src="/files/cPMYYISxaX6lJIijPtyh" alt="" width="563"><figcaption></figcaption></figure>

## Access the Control Plane Tenant

1. Open the activation email and click **SIGN IN TO NVIDIA RUN:AI**.

<figure><img src="/files/ILpkE4lHqn3tNHg6K7M5" alt="" width="563"><figcaption></figcaption></figure>

2. On the password setup page, enter a **New Password**, confirm it, and click **Reset password**.
3. Once your account is updated, click **Back to Application**.
4. On the NVIDIA Run:ai sign-in page, enter your email and new password, then click **Sign In**.
5. Review the **Run:ai End User License Agreement** and click **I Agree**.

## Next Steps

After signing in, the onboarding wizard opens automatically. Before proceeding, ensure your environment meets all [system](/saas/getting-started/installation/install-using-helm/system-requirements) and [network](/saas/getting-started/installation/install-using-helm/network-requirements) requirements. Then follow the wizard to [install](/saas/getting-started/installation/install-using-helm/helm-install) the NVIDIA Run:ai cluster.


# System Requirements

The NVIDIA Run:ai cluster is a Kubernetes application. It has specific system and Kubernetes environment requirements that must be met before installation. These requirements ensure that NVIDIA Run:ai cluster services can be deployed successfully and support AI workloads after connecting to the NVIDIA Run:ai control plane.

This section describes the minimum hardware, supported Kubernetes and OpenShift versions, and environment prerequisites required for cluster installation. Environment prerequisites include critical infrastructure components such as a properly configured Fully Qualified Domain Name (FQDN) with DNS resolution, TLS certificate configuration and ingress readiness.

The cluster requires the NVIDIA GPU Operator and other operators and frameworks to be installed in the cluster to provision and manage GPUs and support AI workload execution. For detailed version compatibility, including supported Kubernetes, GPU Operator versions and more, refer to the [Support matrix](/saas/getting-started/installation/support-matrix).

## Hardware Requirements

The following hardware requirements are for the Kubernetes cluster nodes. By default, all NVIDIA Run:ai cluster services run on all available nodes. For production deployments, you may want to [set node roles](/saas/infrastructure-setup/advanced-setup/node-roles), to separate between system and worker nodes, reduce downtime and save CPU cycles on expensive GPU machines.

### NVIDIA Run:ai Cluster - System Nodes

This configuration is the minimum requirement you need to install and use NVIDIA Run:ai cluster.

| Component  | Required Capacity |
| ---------- | ----------------- |
| CPU        | 10 cores          |
| Memory     | 20GB              |
| Disk space | 50GB              |

{% hint style="info" %}
**Note**

To designate nodes to NVIDIA Run:ai system services, follow the instructions as described in [System nodes](/saas/infrastructure-setup/advanced-setup/node-roles#system-nodes).
{% endhint %}

### NVIDIA Run:ai Cluster - Worker Nodes

The NVIDIA Run:ai cluster supports x86 and ARM CPUs, and any NVIDIA GPUs supported by the NVIDIA GPU Operator. The list of supported GPUs depends on the version of the NVIDIA GPU Operator installed in the cluster. NVIDIA Run:ai supports GPU Operator versions 25.10 to 26.3.

For the list of supported GPUs, see [Supported NVIDIA Data Center GPUs and Systems](https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/platform-support.html#supported-nvidia-data-center-gpus-and-systems). To install the GPU Operator, see [NVIDIA GPU Operator](#nvidia-gpu-operator).

{% hint style="info" %}
**Note**

* NVIDIA DGX Spark, NVIDIA Jetson and workstations are not supported.
* vGPU is not supported. NVIDIA Run:ai currently supports GPU passthrough only.
  {% endhint %}

The following configuration represents the minimum hardware requirements for installing and operating the NVIDIA Run:ai cluster on worker nodes. Each node must meet these specifications:

| Component | Required Capacity |
| --------- | ----------------- |
| CPU       | 2 cores           |
| Memory    | 4GB               |

{% hint style="info" %}
**Note**

To designate nodes to NVIDIA Run:ai workloads, follow the instructions as described in [Worker nodes](/saas/infrastructure-setup/advanced-setup/node-roles#worker-nodes).
{% endhint %}

### Shared Storage

NVIDIA Run:ai workloads must be able to access data from any worker node in a uniform way, to access training data and code as well as save checkpoints, weights, and other machine-learning-related artifacts.

Typical protocols are Network File Storage (NFS) or Network-attached storage (NAS). NVIDIA Run:ai cluster supports both, for more information see [Shared storage](/saas/infrastructure-setup/procedures/shared-storage).

## Software Requirements

The following software requirements must be fulfilled on the Kubernetes cluster.

### Operating System

The following notes apply to the operating system running on NVIDIA Run:ai cluster nodes:

* Any **Linux** operating system supported by both Kubernetes and NVIDIA GPU Operator.
* NVIDIA Run:ai cluster on Google Kubernetes Engine (GKE) supports both Ubuntu and Container Optimized OS (COS).
  * COS is supported only with NVIDIA GPU Operator 24.6 and above, and NVIDIA Run:ai cluster version 2.19 and above.
* NVIDIA Run:ai cluster on Elastic Kubernetes Service (EKS) does not support Bottlerocket or Amazon Linux.
* NVIDIA Run:ai cluster on Oracle Kubernetes Engine (OKE) supports only Ubuntu.
* Internal tests are being performed on **Ubuntu** and **CoreOS** for OpenShift.

### Kubernetes Distribution

#### NVIDIA Run:ai Compatible Distributions

NVIDIA Run:ai cluster requires Kubernetes. The following Kubernetes distributions are supported:

* Vanilla Kubernetes
* OpenShift Container Platform (OCP)
* Elastic Kubernetes Engine (EKS)
* Google Kubernetes Engine (GKE)
* Azure Kubernetes Service (AKS)
* Oracle Kubernetes Engine (OKE)
* Rancher Kubernetes Engine 2 (RKE2)

{% hint style="info" %}
**Note**

* The latest release of the NVIDIA Run:ai cluster supports **Kubernetes 1.34 to 1.36** and **OpenShift 4.18 to 4.21**.
* For [Multi-Node NVLink](/saas/platform-management/aiinitiatives/resources/using-gb200) support (e.g. GB200), Kubernetes 1.32 and above is required.
  {% endhint %}

For the full list of NVIDIA Run:ai compatible distributions and per-version Kubernetes and OpenShift compatibility, see [Support matrix](/saas/getting-started/installation/support-matrix#nvidia-run-ai-compatible-distributions).

#### Partner-Compatible Distributions

The following Kubernetes distributions are **partner-compatible**. They are tested and validated by the partner, who is responsible for maintaining compatibility with NVIDIA Run:ai:

* Crusoe Managed Kubernetes (CMK)
* Mirantis k0rdent
* Rafay platform
* vCluster
* VMware vSphere Kubernetes Service (VKS)

For the full list of partner-compatible distributions and per-partner version compatibility, see [Support matrix](/saas/getting-started/installation/support-matrix#partner-compatible-distributions).

### Container Runtime

NVIDIA Run:ai supports the following [container runtimes](https://kubernetes.io/docs/setup/production-environment/container-runtimes/). Make sure your Kubernetes cluster is configured with one of these runtimes:

* [Containerd](https://kubernetes.io/docs/setup/production-environment/container-runtimes/#containerd) (default in Kubernetes)
* [CRI-O](https://cri-o.io/) (default in OpenShift)

### Kubernetes Pod Security Admission

NVIDIA Run:ai supports `restricted` policy for [Pod Security Admission](https://kubernetes.io/docs/concepts/security/pod-security-admission/) (PSA) on OpenShift only. Other Kubernetes distributions are only supported with `privileged` policy.

For NVIDIA Run:ai on OpenShift to run with PSA `restricted` policy:

* Label the `runai` namespace as described in [Pod Security Admission](https://kubernetes.io/docs/concepts/security/pod-security-admission/) with the following labels:

  ```bash
  pod-security.kubernetes.io/audit=privileged
  pod-security.kubernetes.io/enforce=privileged
  pod-security.kubernetes.io/warn=privileged
  ```
* The workloads submitted through NVIDIA Run:ai should comply with the restrictions of PSA restricted policy. This can be enforced using [Policies](/saas/platform-management/policies/native-workload-policies).

### NVIDIA Run:ai Namespace

The NVIDIA Run:ai cluster must be installed in a namespace or project (OpenShift) called `runai`. Use the following to create the namespace/project:

{% tabs %}
{% tab title="Kubernetes" %}

```bash
kubectl create ns runai
```

{% endtab %}

{% tab title="OpenShift" %}

```bash
oc new-project runai
```

{% endtab %}
{% endtabs %}

### Kubernetes Ingress Controller

NVIDIA Run:ai cluster requires [Kubernetes Ingress Controller](https://kubernetes.io/docs/concepts/services-networking/ingress-controllers/) to be installed on the Kubernetes cluster.

* OpenShift and RKE2 come with a pre-installed ingress controller.
* Make sure that a default ingress controller, `global.ingress.ingressClass` is set. For more details, see [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config).

There are multiple ways to install and configure an ingress controller. The following example demonstrates how to install and configure the [HAProxy Kubernetes Ingress Controller](https://www.haproxy.com/documentation/kubernetes-ingress/) using [helm](https://helm.sh/).

<details>

<summary>Vanilla Kubernetes</summary>

```bash
helm repo add haproxytech https://haproxytech.github.io/helm-charts
helm repo update
helm install haproxy-kubernetes-ingress haproxytech/kubernetes-ingress \
--create-namespace \
--namespace haproxy-controller \
--set controller.kind=DaemonSet \
--set controller.service.type=NodePort
```

</details>

<details>

<summary>Managed Kubernetes (EKS, GKE, AKS)</summary>

```bash
helm repo add haproxytech https://haproxytech.github.io/helm-charts
helm repo update
helm install haproxy-kubernetes-ingress haproxytech/kubernetes-ingress \
--create-namespace \
--namespace haproxy-controller \
--set controller.service.type=LoadBalancer \
```

</details>

<details>

<summary>Oracle Kubernetes Engine (OKE)</summary>

```bash
helm repo add haproxytech https://haproxytech.github.io/helm-charts
helm repo update
helm install haproxy-kubernetes-ingress haproxytech/kubernetes-ingress \
  --create-namespace \
  --namespace haproxy-controller \
  --set controller.kind=DaemonSet \
  --set controller.service.type=LoadBalancer \
  --set controller.service.externalTrafficPolicy=Local \
  --set controller.service.annotations."oci-network-load-balancer\.oraclecloud\.com/is-preserve-source"="True" \
  --set controller.service.annotations."oci-network-load-balancer\.oraclecloud\.com/security-list-management-mode"=All \
  --set controller.service.annotations."oci\.oraclecloud\.com/load-balancer-type"=nlb
```

</details>

### NVIDIA GPU Operator

The NVIDIA Run:ai cluster requires NVIDIA GPU Operator to be installed on the Kubernetes cluster. GPU Operator versions 25.10 to 26.3 are supported.

{% hint style="info" %}
**Note**

For [Multi-Node NVLink](/saas/platform-management/aiinitiatives/resources/using-gb200) support (e.g. GB200), GPU Operator 25.3 and above is required.
{% endhint %}

* For installation instructions, see [Installing the NVIDIA GPU Operator](https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/getting-started.html).
* See the following notes below:
  * Use the default `gpu-operator` namespace. Otherwise, you must specify the target namespace using the flag `runai-operator.config.nvidiaDcgmExporter.namespace` as described in [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config).
  * Cluster 2.23 only - For NVIDIA GPU Operator **v25.10**, containerd must be explicitly set as the default runtime. Add the following flags to the installation command:

    ```bash
    --set toolkit.env[3].name=CONTAINERD_SET_AS_DEFAULT \
    --set-string toolkit.env[3].value=true
    ```

    <div data-gb-custom-block data-tag="hint" data-style="info" class="hint hint-info"><p><strong>Note</strong></p><p>OpenShift customers should skip this step. The GPU Operator automatically detects CRI-O as the container runtime on OpenShift — no manual container runtime flags are required. The <code>CONTAINERD_SET_AS_DEFAULT</code> flag is containerd-specific and does not apply to CRI-O/OpenShift environments.</p></div>
  * NVIDIA drivers may already be installed on the nodes. In such cases, use the NVIDIA GPU Operator flags `--set driver.enabled=false`. [DGX OS](https://docs.nvidia.com/dgx/dgx-os-6-user-guide/release_notes.html) is one such example as it comes bundled with NVIDIA Drivers.
* For distribution-specific requirements and additional instructions, see the sections below:

<details>

<summary>OpenShift Container Platform (OCP)</summary>

The Node Feature Discovery (NFD) Operator is a prerequisite for the NVIDIA GPU Operator in OpenShift. Install the NFD Operator using the Red Hat OperatorHub catalog in the OpenShift Container Platform web console. For more information, see [Installing the Node Feature Discovery (NFD) Operator](https://docs.nvidia.com/datacenter/cloud-native/openshift/latest/install-nfd.html).

</details>

<details>

<summary>Elastic Kubernetes Service (EKS)</summary>

* When setting-up the cluster, do **not** install the NVIDIA device plug-in (we want the NVIDIA GPU Operator to install it instead).
* When using the [eksctl](https://eksctl.io/) tool to create a cluster, use the flag `--install-nvidia-plugin=false` to disable the installation.

For GPU nodes, EKS uses an AMI which already contains the NVIDIA drivers. As such, you must use the GPU Operator flags: `--set driver.enabled=false.`

</details>

<details>

<summary>Google Kubernetes Engine (GKE)</summary>

Before installing the GPU Operator:

1. Create the `gpu-operator` namespace by running:

```bash
kubectl create ns gpu-operator
```

2. Create the following file:

<pre class="language-yaml"><code class="lang-yaml">#resourcequota.yaml

<strong>apiVersion: v1
</strong>kind: ResourceQuota
metadata:
name: gcp-critical-pods
namespace: gpu-operator
spec:
scopeSelector:
    matchExpressions:
    - operator: In
    scopeName: PriorityClass
    values:
    - system-node-critical
    - system-cluster-critical
</code></pre>

3. Run:

<pre class="language-bash"><code class="lang-bash"><strong>kubectl apply -f resourcequota.yaml
</strong></code></pre>

</details>

<details>

<summary>Rancher Kubernetes Engine 2 (RKE2)</summary>

Before installing the GPU Operator, verify the [host OS requirements](https://docs.rke2.io/add-ons/gpu_operators?GPUoperator=v25.3.x#host-os-requirements) are met. Then, install the [operator](https://docs.rke2.io/add-ons/gpu_operators#operator-installation).

When installing GPU Operator v25.3, update the Helm values file as follows:

```yaml
apiVersion: helm.cattle.io/v1
kind: HelmChart
metadata:
  name: gpu-operator
  namespace: kube-system
spec:
  repo: https://helm.ngc.nvidia.com/nvidia
  chart: gpu-operator
  version: v25.3.4
  targetNamespace: gpu-operator
  createNamespace: true
  valuesContent: |-
    toolkit:
      env:
      - name: CONTAINERD_SOCKET
        value: /run/k3s/containerd/containerd.sock
```

</details>

<details>

<summary>Oracle Kubernetes Engine (OKE)</summary>

* During cluster setup, [create a nodepool](https://docs.oracle.com/en-us/iaas/tools/python/latest/api/container_engine/models/oci.container_engine.models.NodePool.html#oci.container_engine.models.NodePool.initial_node_labels), and set `initial_node_labels` to include `oci.oraclecloud.com/disable-gpu-device-plugin=true` which disables the NVIDIA GPU device plugin.
* For GPU nodes, OKE defaults to Oracle Linux, which is incompatible with NVIDIA drivers. To resolve this, use a custom Ubuntu image instead.

</details>

For troubleshooting information, see the [NVIDIA GPU Operator Troubleshooting Guide](https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/troubleshooting.html).

### NVIDIA Network Operator

When deploying on clusters with RDMA or Multi-Node NVLink‑capable nodes (e.g. B200, GB200), the NVIDIA Network Operator is required to enable high-performance networking features such as GPUDirect RDMA in Kubernetes. Network Operator versions 25.10 to 26.1 are supported.

The Network Operator works alongside the NVIDIA GPU Operator to provide:

* NVIDIA networking drivers for advanced network capabilities.
* Kubernetes device plugins to expose high‑speed network hardware to workloads.
* Secondary network components to support network‑intensive applications.

The Network Operator must be installed and configured as follows:

1. Install the network operator as detailed in [Network Operator Deployment on Vanilla Kubernetes Cluster](https://docs.nvidia.com/networking/display/kubernetes2440/getting-started-kubernetes.html#network-operator-deployment-on-vanilla-kubernetes-cluster).
2. Configure SR-IOV InfiniBand support as detailed in [Network Operator Deployment with an SR-IOV InfiniBand Network](https://docs.nvidia.com/networking/display/kubernetes2440/getting-started-kubernetes.html#network-operator-deployment-with-an-sr-iov-infiniband-network).

### NVIDIA Dynamic Resource Allocation (DRA) Driver

When deploying on clusters with Multi-Node NVLink (e.g. GB200), the NVIDIA DRA driver is required to enable Dynamic Resource Allocation at the Kubernetes level. To install, follow the instructions in [Configure and Helm-install the driver](https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/dra-intro-install.html#configure-and-helm-install-the-driver). For air-gapped installations, the DRA driver is installed with the [GPU Operator](#nvidia-gpu-operator). DRA driver versions 25.8 to 25.12 are supported.

After the DRA driver is installed, update the cluster configuration using `GPUNetworkAccelerationEnabled` flag to enable GPU network acceleration. This triggers an update of the NVIDIA Run:ai workload controller deployment and restarts the controller. For details on how to configure this value using Helm or `runaiconfig`, see [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config).

### Prometheus

{% hint style="info" %}
**Note**

Installing Prometheus applies to Kubernetes only.
{% endhint %}

NVIDIA Run:ai requires Prometheus on the Kubernetes cluster. The setup has two parts:

* **Prerequisite: Kube-Prometheus Stack.** On Kubernetes, install the Kube-Prometheus Stack to provide the Prometheus Operator and the Prometheus Custom Resource Definitions (CRDs). OpenShift includes equivalent components through its built-in [Cluster Monitoring Operator](https://docs.redhat.com/en/documentation/openshift_container_platform/4.18/html-single/monitoring/index#about-ocp-monitoring), so no install is required.
* **NVIDIA Run:ai Prometheus instance.** During cluster installation, the NVIDIA Run:ai cluster operator creates a dedicated Prometheus Custom Resource (CR) in the `runai` namespace. This step applies to both Kubernetes and OpenShift, and produces the Prometheus instance that NVIDIA Run:ai uses for monitoring.

For RKE2, see [Enable Monitoring](https://ranchermanager.docs.rancher.com/how-to-guides/advanced-user-guides/monitoring-alerting-guides/enable-monitoring) for instructions on installing Prometheus.

The following example installs the community [Kube-Prometheus Stack](https://artifacthub.io/packages/helm/prometheus-community/kube-prometheus-stack) using [helm](https://helm.sh/):

```bash
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo update
helm install prometheus prometheus-community/kube-prometheus-stack \
    -n monitoring --create-namespace --set grafana.enabled=false
```

## Routing Traffic to and from NVIDIA Run:ai Services

This section describes how to route traffic to and from NVIDIA Run:ai services. Proper traffic routing is required to enable secure access to the NVIDIA Run:ai control plane (via the UI or API), as well as external access to development workspaces, training workloads, and inference workloads.

NVIDIA Run:ai supports two routing approaches for exposing services: host-based routing and path-based routing. While path-based routing exposes multiple services under a shared domain using URL paths, host-based routing assigns each service its own subdomain. Since many development tools and applications expect to run at the root path, host-based routing avoids common compatibility issues and is therefore used by default in NVIDIA Run:ai.

NVIDIA Run:ai uses host-based routing by default, which relies on DNS and TLS configuration to securely expose services. To support this, three key components must be configured:

* A [Fully Qualified Domain Name (FQDN)](#fully-qualified-domain-name-fqdn) structure
* [TLS certificates](#tls-certificate) for secure communication
* [Host-based routing](#host-based-routing-default) for workload exposure

Together, these components ensure that traffic is routed correctly and securely across all NVIDIA Run:ai services.

{% hint style="info" %}
**Note**

* NVIDIA Run:ai also supports path-based routing. If this approach better fits your environment, you can use it instead of the default host-based routing. In this case, the [workspace and training wildcard certificate](#workspaces-and-training-workload-certificate) is not required.
* To use path-based routing, disable host-based routing by setting `clusterConfig.global.subdomainSupport: false` during Helm installation. See [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config).
  {% endhint %}

### Fully Qualified Domain Name (FQDN)

{% hint style="info" %}
**Note**

Fully Qualified Domain Name applies to Kubernetes only.
{% endhint %}

NVIDIA Run:ai services rely on Fully Qualified Domain Names (FQDNs) to route traffic between system components and to expose workloads externally. In NVIDIA Run:ai, the FQDN settings are needed for:

* Enabling communication between the control plane and the cluster
* Exposing development workspaces and training workloads via subdomains
* Exposing inference workloads via dedicated subdomains

You must configure domain names for each of the following communication types:

* Control plane ↔ cluster communication\
  Example: `runai.mycorp.local`\
  The IP address of this domain must be resolvable within the organization’s private network.
* Workspace and training workloads (external access)\
  Example: `*.runai.mycorp.local`
* Inference workloads (external access)\
  Example: `*.runai-inference.mycorp.local`

{% hint style="info" %}
**Note**

The inference wildcard (`*.runai-inference.mycorp.local`) is only required if you plan to expose inference workloads outside the cluster. For cluster-local inference, skip the inference wildcard record below and the matching [Inference wildcard certificate](#inference-wildcard-certificate). See [Inference](#inference).
{% endhint %}

Since NVIDIA Run:ai uses host-based routing, wildcard DNS records must be configured to enable external access to workloads.

Configure the following DNS records. Both records should resolve to the same cluster public IP address, or to the cluster’s load balancer IP in on-prem environments. This ensures that each workspace or workload is assigned a unique subdomain under the wildcard domains:

```bash
*.runai.mycorp.local → <cluster IP>
*.runai-inference.mycorp.local → <cluster IP>
```

### NVIDIA Run:ai TLS Certificates

TLS certificates secure communication between NVIDIA Run:ai components and enable HTTPS access to exposed services.

NVIDIA Run:ai requires three TLS certificates, each aligned with a specific domain:

* Cluster domain certificate
* Workspaces and training workload certificate
* Inference certificate (wildcard)

#### Cluster Domain Certificates (Single-Domain)

* **Kubernetes** - To enable secure communication between the NVIDIA Run:ai control plane and the cluster, configure a TLS certificate associated with the cluster’s main domain (e.g. `runai.mycorp.local`). This certificate should be stored as a secret named `runai-cluster-domain-tls-secret` in the `runai` namespace.

  * Replace `/path/to/fullchain.pem` with the actual path to your TLS certificate.
  * Replace `/path/to/private.pem` with the actual path to your private key.

  ```bash
  kubectl create secret tls runai-cluster-domain-tls-secret -n runai \
    --cert /path/to/fullchain.pem \
    --key /path/to/private.pem
  ```
* **OpenShift** - NVIDIA Run:ai uses the OpenShift default Ingress router for serving. The TLS certificate configured for this router must be issued by a trusted CA. For more details, see the OpenShift documentation on [configuring certificates](https://docs.redhat.com/en/documentation/openshift_container_platform/4.18/html/security_and_compliance/configuring-certificates#replacing-default-ingress).

#### Workspaces & Training Workload Wildcard Certificate

{% hint style="info" %}
**Note**

For path-based routing, ignore this configuration and move to the next step.
{% endhint %}

* **Kubernetes** - To allow secure access to workspace and training workloads via subdomains, configure a wildcard TLS certificate that matches the cluster domain (e.g. `*.runai.mycorp.local`). This certificate should be stored as a secret named `runai-cluster-domain-star-tls-secret` in the `runai` namespace.
  * Replace `/path/to/fullchain.pem` with the actual path to your TLS certificate.
  * Replace `/path/to/private.pem` with the actual path to your private key.

    ```bash
    kubectl create secret tls runai-cluster-domain-star-tls-secret -n runai \
      --cert /path/to/fullchain.pem \
      --key /path/to/private.pem
    ```
* **OpenShift** - A wildcard TLS certificate for Workspace and Training workloads is not required. OpenShift Routes handle TLS termination for inference endpoints using the platform’s built-in routing and certificate management.

#### Inference Wildcard Certificate

{% hint style="info" %}
**Note**

The inference wildcard certificate is only required if inference services are exposed externally via wildcard FQDN. For cluster-local inference, skip this subsection. See [Inference](#inference).
{% endhint %}

* **Kubernetes** - To securely expose inference services over HTTPS, configure a wildcard TLS certificate for the inference domain (e.g. `*.runai-inference.mycorp.local`). This certificate should be stored as a secret named `runai-cluster-inference-tls-secret` in the `knative-serving` namespace.

  * Replace `/path/to/fullchain.pem` with the actual path to your TLS certificate.
  * Replace `/path/to/private.pem` with the actual path to your private key.

  ```bash
  kubectl create secret tls runai-cluster-inference-tls-secret -n knative-serving \
      --cert /path/to/fullchain.pem \
      --key /path/to/private.pem
  ```
* **OpenShift** - A wildcard TLS certificate for Inference workloads is not required. OpenShift Routes handle TLS termination for inference endpoints using the platform’s built-in routing and certificate management.

### Host-Based Routing (Default)

{% hint style="info" %}
**Note**

* The following steps are required for Kubernetes only. For OpenShift, no additional configuration is required.
* NVIDIA Run:ai also support path-based routing. If you prefer to use it instead of the default host-based routing, disable host-based routing by setting `clusterConfig.global.subdomainSupport: false` during the Helm installation. See [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config).
* If you choose path-based routing, skip the below steps.
  {% endhint %}

Host-based routing binds together the configured domains (FQDN), TLS certificates, and ingress rules to expose workloads externally. With NVIDIA Run:ai host-based routing, workloads are exposed using subdomains, so each workload is assigned its own URL. For example:

```bash
https://<project>-<workload>.<CLUSTER_URL>
```

Host-based routing relies on:

* The FQDN structure defined earlier
* The TLS certificates configured for those domains

This section describes how to connect these components to enable workload exposure.

1. Ensure that:
   * Wildcard DNS records are configured (see [FQDN](#fully-qualified-domain-name-fqdn) section)
   * TLS certificates are created and stored as Kubernetes secrets (see [TLS certificates](#nvidia-run-ai-tls-certificates) section)
2. Create the ingress resource

   ```yaml
   apiVersion: networking.k8s.io/v1
   kind: Ingress
   metadata:
     name: runai-cluster-domain-star-ingress
     namespace: runai
   spec:
     rules:
     - host: '*.<CLUSTER_URL>'
     tls:
     - hosts:
       - '*.<CLUSTER_URL>'
       secretName: runai-cluster-domain-star-tls-secret
   ```
3. Run the following:

   ```bash
   kubectl apply -f <filename>
   ```

### Local Certificate Authority

A local certificate authority serves as the root certificate for organizations that cannot use **publicly trusted certificate authority**. Follow the steps below to configure the local certificate authority.

In air-gapped environments, you **must** configure and install the local CA's public key in the Kubernetes cluster. This is required for the installation to succeed:

1\. Add the public key to the required namespace:

{% tabs %}
{% tab title="Kubernetes" %}

```bash
kubectl -n runai create secret generic runai-ca-cert \
    --from-file=runai-ca.pem=<ca_bundle_path>
kubectl label secret runai-ca-cert -n runai run.ai/cluster-wide=true run.ai/name=runai-ca-cert --overwrite
```

{% endtab %}

{% tab title="OpenShift" %}

```bash
oc -n runai create secret generic runai-ca-cert \
    --from-file=runai-ca.pem=<ca_bundle_path>
oc -n openshift-monitoring create secret generic runai-ca-cert \
    --from-file=runai-ca.pem=<ca_bundle_path>
oc label secret runai-ca-cert -n runai run.ai/cluster-wide=true run.ai/name=runai-ca-cert --overwrite
```

{% endtab %}
{% endtabs %}

2. When installing the cluster, make sure the following flag is added to the helm command `--set global.customCA.enabled=true`. See [Install the cluster](/saas/getting-started/installation/install-using-helm/helm-install#installation).

{% hint style="info" %}
**Note**

For Git and S3 data source integrations, NVIDIA Run:ai supports the following options:

* Use the same custom CA defined during cluster installation by setting:\
  `--set global.customCAGit.enabled=true` or\
  `--set global.customCAS3.enabled=true`.
* Use a different CA certificate specifically for Git or S3 by enabling the setting and providing a custom secret name:\
  `--set global.customCAGit.enabled=true`\
  `--set global.customCAGit.secret.name=<git-ca-cert>` or\
  `--set global.customCAS3.enabled=true`\
  `--set global.customCAS3.secret.name=<s3-ca-cert>`

For more details, see [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config).
{% endhint %}

## Additional Software Requirements

To enable NVIDIA Run:ai capabilities such as Distributed Training and Inference, additional Kubernetes applications (frameworks) must be installed on the cluster.

### Distributed Training

Distributed training enables training of AI models over multiple nodes. This requires installing a distributed training framework on the cluster. The following frameworks are supported:

* ​[TensorFlow](https://www.tensorflow.org/)​
* ​[PyTorch](https://pytorch.org/)​
* ​[XGBoost](https://xgboost.readthedocs.io/)​
* ​[MPI v2](https://docs.open-mpi.org/)​
* ​[JAX](https://docs.jax.dev/en/latest/index.html)

There are several ways to install each framework. A simple method of installation example is the [Kubeflow Training Operator](https://www.kubeflow.org/docs/components/training/installation/) which includes TensorFlow, PyTorch, XGBoost and JAX.

It is recommended to use **Kubeflow Training Operator v1.9.2**, and **MPI Operator v0.6.0 or later** for compatibility with advanced workload capabilities, such as [Stopping a workload](/saas/workloads-in-nvidia-run-ai/workloads) and [Scheduling rules](/saas/platform-management/policies/scheduling-rules).

* To install the Kubeflow Training Operator for TensorFlow, PyTorch, XGBoost and JAX frameworks, run the following command:

  ```bash
  kubectl apply --server-side -k "github.com/kubeflow/training-operator.git/manifests/overlays/standalone?ref=v1.9.2"
  ```
* To install the MPI Operator for MPI v2, run the following command:

  ```bash
  kubectl apply --server-side -f https://raw.githubusercontent.com/kubeflow/mpi-operator/v0.6.0/deploy/v2beta1/mpi-operator.yaml
  ```

{% hint style="info" %}
**Note**

If you require both the MPI Operator and Kubeflow Training Operator, follow the steps below:

* Install the Kubeflow Training Operator as described above.
* Disable and delete MPI v1 in the Kubeflow Training Operator by running:

  ```bash
  kubectl patch deployment training-operator -n kubeflow --type='json' -p='[{"op": "add", "path": "/spec/template/spec/containers/0/args", "value": ["--enable-scheme=tfjob", "--enable-scheme=pytorchjob", "--enable-scheme=xgboostjob", "--enable-scheme=jaxjob"]}]'
  kubectl delete crd mpijobs.kubeflow.org
  ```
* Install the MPI Operator as described above.
  {% endhint %}

### Inference

Inference enables serving of AI models. This requires the [Knative Serving](https://knative.dev/docs/serving/) framework to be installed on the cluster and supports Knative versions 1.19 to 1.21.

To configure inference:

* **Inference request metrics (OTLP)** - For Knative Serving 1.19 or later, add the `observability` block to your `KnativeServing` YAML to export request metrics to the NVIDIA Run:ai API and UI. See [Inference request metrics (OTLP)](#inference-request-metrics-otlp) below.
* **Internal or external access** - Choose how inference services are reached: externally via a wildcard FQDN (the default) or cluster-local. See the matching section below.

#### Inference Request Metrics (OTLP)

To surface inference request metrics in the NVIDIA Run:ai API and UI, configure Knative Serving 1.19 or later to export request metrics over OpenTelemetry Protocol (OTLP) to the NVIDIA Run:ai Prometheus instance. The `observability` block included in the `KnativeServing` YAML examples below sets the required `config-observability` keys (`metrics-protocol`, `request-metrics-protocol`, `request-metrics-endpoint`). For supported fields, see [Enabling metric collection](https://knative.dev/docs/serving/observability/metrics/collecting-metrics/#enabling-metric-collection) in the Knative documentation.

The OTLP receiver on the NVIDIA Run:ai Prometheus instance is enabled by default from NVIDIA Run:ai 2.25. On NVIDIA Run:ai 2.24 and earlier, the observability block is not required.

To disable it, set `spec.prometheus.config.enableOTLPReceiver: false`. See [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config) for more details.

#### External-Access Inference

With external access, inference services are reachable from outside the cluster. Use this mode when inference workloads must be reachable by clients outside the cluster.

{% tabs %}
{% tab title="Kubernetes" %}
On Kubernetes, external access exposes inference services through a wildcard FQDN (for example, `*.runai-inference.mycorp.local`). This requires the wildcard DNS record (see [FQDN](#fully-qualified-domain-name-fqdn)), the [Inference wildcard certificate](#inference-wildcard-certificate), and an HAProxy ingress to handle TLS termination.

Follow the [Installing Knative](https://knative.dev/docs/install/operator/knative-with-operators/) instructions or run:

```bash
helm repo add knative-operator https://knative.github.io/operator
helm install knative-operator --create-namespace --namespace knative-operator --version {VERSION} knative-operator/knative-operator
```

Once installed, follow the steps below:

1. Create the `knative-serving` namespace:

   ```bash
   kubectl create ns knative-serving
   ```
2. Create a YAML file named `knative-serving.yaml` and replace the placeholder FQDN with your wildcard inference domain (for example, `runai-inference.mycorp.local`).

   ```yaml
   apiVersion: operator.knative.dev/v1beta1
   kind: KnativeServing
   metadata:
     name: knative-serving
     namespace: knative-serving
   spec:
     config:
       config-autoscaler:
         enable-scale-to-zero: "true"
       config-features:
         kubernetes.podspec-affinity: enabled
         kubernetes.podspec-init-containers: enabled
         kubernetes.podspec-persistent-volume-claim: enabled
         kubernetes.podspec-persistent-volume-write: enabled
         kubernetes.podspec-schedulername: enabled
         kubernetes.podspec-securitycontext: enabled
         kubernetes.podspec-tolerations: enabled
         kubernetes.podspec-volumes-emptydir: enabled
         kubernetes.podspec-fieldref: enabled
         kubernetes.containerspec-addcapabilities: enabled
         kubernetes.podspec-nodeselector: enabled
         multi-container: enabled
         kubernetes.podspec-hostipc: enabled
         kubernetes.podspec-hostnetwork: enabled
       domain:
         runai-inference.mycorp.local: "" # replace with the wildcard FQDN for Inference
       network:
         domainTemplate: '{{.Name}}-{{.Namespace}}.{{.Domain}}'
         ingress-class: kourier.ingress.networking.knative.dev
         default-external-scheme: https
       observability:
         metrics-protocol: "prometheus"
         request-metrics-protocol: "http/protobuf"
         request-metrics-endpoint: "http://prometheus-operated.runai.svc.cluster.local:9090/api/v1/otlp/v1/metrics"
     high-availability:
       replicas: 2
     ingress:
       kourier:
         enabled: true
   ```
3. Apply the changes:

   ```bash
   kubectl apply -f knative-serving.yaml
   ```
4. Configure HAProxy to proxy requests to Kourier / Knative and handle TLS termination using the wildcard certificate. Create a YAML file named `knative-ingress.yaml` and replace the FQDN placeholders with your wildcard inference domain:

   ```yaml
   apiVersion: networking.k8s.io/v1
   kind: Ingress
   metadata:
     name: knative-serving
     namespace: knative-serving
   spec:
     ingressClassName: haproxy
     rules:
     - host: '*.runai-inference.mycorp.local' # replace with the wildcard FQDN for Inference
       http:
         paths:
         - backend:
             service:
               name: kourier
               port:
                 number: 80
           path: /
           pathType: Prefix
     tls:
     - hosts:
       - '*.runai-inference.mycorp.local' # replace with the wildcard FQDN for Inference
       secretName: runai-cluster-inference-tls-secret
   ```
5. Apply the changes:

   ```bash
   kubectl apply -f knative-ingress.yaml
   ```

{% endtab %}

{% tab title="OpenShift" %}
On OpenShift Serverless, external access is the platform default: each KnativeService automatically gets an OpenShift Route.

Follow the [Installing the OpenShift Serverless Operator](https://docs.redhat.com/en/documentation/red_hat_openshift_serverless/1.37/html/installing_openshift_serverless/install-serverless-operator) instructions. Once installed, follow the steps below:

1. Create the `knative-serving` project:

   ```bash
   oc new-project knative-serving
   ```
2. Create a YAML file named `knative-serving.yaml`:

   ```yaml
   apiVersion: operator.knative.dev/v1beta1
   kind: KnativeServing
   metadata:
     finalizers:
       - knative-serving-openshift
       - knativeservings.operator.knative.dev
     name: knative-serving
     namespace: knative-serving
   spec:
     config:
       config-features:
         kubernetes.podspec-tolerations: enabled
         kubernetes.podspec-volumes-emptydir: enabled
         kubernetes.podspec-persistent-volume-claim: enabled
         multi-container: enabled
         kubernetes.podspec-persistent-volume-write: enabled
         kubernetes.podspec-fieldref: enabled
         kubernetes.podspec-schedulername: enabled
         kubernetes.podspec-nodeselector: enabled
         kubernetes.podspec-init-containers: enabled
         kubernetes.podspec-securitycontext: enabled
         kubernetes.podspec-affinity: enabled
         kubernetes.containerspec-addcapabilities: enabled
       observability:
         metrics-protocol: "prometheus"
         request-metrics-protocol: "http/protobuf"
         request-metrics-endpoint: "http://prometheus-operated.runai.svc.cluster.local:9090/api/v1/otlp/v1/metrics"
     controller-custom-certs:
       name: ''
       type: ''
     registry: {}
   ```
3. Apply the changes:

   ```bash
   oc apply -f knative-serving.yaml
   ```

{% endtab %}
{% endtabs %}

#### Cluster-Local Inference

With cluster-local access, inference services are reachable only from inside the cluster through the cluster-internal domain (`<service>.<namespace>.svc.cluster.local`). This mode is recommended when inference workloads are consumed by other applications running in the same cluster.

On OpenShift Serverless, external access is the platform default; cluster-local must be enabled per inference workload. See the OpenShift tab below.

{% tabs %}
{% tab title="Kubernetes" %}
On Kubernetes, cluster-local access does not require an external DNS record, a wildcard certificate, or an HAProxy ingress because traffic does not leave the cluster.

Follow the [Installing Knative](https://knative.dev/docs/install/operator/knative-with-operators/) instructions or run:

```bash
helm repo add knative-operator https://knative.github.io/operator
helm install knative-operator --create-namespace --namespace knative-operator --version {VERSION} knative-operator/knative-operator
```

Once installed, follow the steps below:

1. Create the `knative-serving` namespace:

   ```bash
   kubectl create ns knative-serving
   ```
2. Create a YAML file named `knative-serving.yaml`:

   ```yaml
   apiVersion: operator.knative.dev/v1beta1
   kind: KnativeServing
   metadata:
     name: knative-serving
     namespace: knative-serving
   spec:
     config:
       config-autoscaler:
         enable-scale-to-zero: "true"
       config-features:
         kubernetes.podspec-affinity: enabled
         kubernetes.podspec-init-containers: enabled
         kubernetes.podspec-persistent-volume-claim: enabled
         kubernetes.podspec-persistent-volume-write: enabled
         kubernetes.podspec-schedulername: enabled
         kubernetes.podspec-securitycontext: enabled
         kubernetes.podspec-tolerations: enabled
         kubernetes.podspec-volumes-emptydir: enabled
         kubernetes.podspec-fieldref: enabled
         kubernetes.containerspec-addcapabilities: enabled
         kubernetes.podspec-nodeselector: enabled
         multi-container: enabled
         kubernetes.podspec-hostipc: enabled
         kubernetes.podspec-hostnetwork: enabled
       domain:
         svc.cluster.local: ""
       network:
         domainTemplate: '{{.Name}}.{{.Namespace}}.{{.Domain}}'
         ingress-class: kourier.ingress.networking.knative.dev
         default-external-scheme: http
       observability:
         metrics-protocol: "prometheus"
         request-metrics-protocol: "http/protobuf"
         request-metrics-endpoint: "http://prometheus-operated.runai.svc.cluster.local:9090/api/v1/otlp/v1/metrics"
     high-availability:
       replicas: 2
     ingress:
       kourier:
         enabled: true
         service-type: ClusterIP
   ```
3. Apply the changes:

   ```bash
   kubectl apply -f knative-serving.yaml
   ```

{% endtab %}

{% tab title="OpenShift" %}
On OpenShift Serverless, external access is the platform default: KnativeServices automatically get an OpenShift Route. To make an inference workload cluster-local, configure the workload's routing settings so that the `networking.knative.dev/visibility=cluster-local` label is applied to the underlying KnativeService.

Follow the [Installing the OpenShift Serverless Operator](https://docs.redhat.com/en/documentation/red_hat_openshift_serverless/1.37/html/installing_openshift_serverless/install-serverless-operator) instructions. Once installed, follow the steps below:

1. Create the `knative-serving` project:

   ```bash
   oc new-project knative-serving
   ```
2. Create a YAML file named `knative-serving.yaml`.

   ```yaml
   apiVersion: operator.knative.dev/v1beta1
   kind: KnativeServing
   metadata:
     finalizers:
       - knative-serving-openshift
       - knativeservings.operator.knative.dev
     name: knative-serving
     namespace: knative-serving
   spec:
     config:
       config-features:
         kubernetes.podspec-tolerations: enabled
         kubernetes.podspec-volumes-emptydir: enabled
         kubernetes.podspec-persistent-volume-claim: enabled
         multi-container: enabled
         kubernetes.podspec-persistent-volume-write: enabled
         kubernetes.podspec-fieldref: enabled
         kubernetes.podspec-schedulername: enabled
         kubernetes.podspec-nodeselector: enabled
         kubernetes.podspec-init-containers: enabled
         kubernetes.podspec-securitycontext: enabled
         kubernetes.podspec-affinity: enabled
         kubernetes.containerspec-addcapabilities: enabled
       observability:
         metrics-protocol: "prometheus"
         request-metrics-protocol: "http/protobuf"
         request-metrics-endpoint: "http://prometheus-operated.runai.svc.cluster.local:9090/api/v1/otlp/v1/metrics"
     controller-custom-certs:
       name: ''
       type: ''
     registry: {}
   ```
3. Apply the changes:

   ```bash
   oc apply -f knative-serving.yaml
   ```

{% endtab %}
{% endtabs %}

### Autoscaling

NVIDIA Run:ai allows for autoscaling a deployment according to the below metrics:

* Latency (milliseconds)
* Throughput (requests/sec)
* Concurrency (requests)

Using a custom metric (for example, Latency) requires installing the [Kubernetes Horizontal Pod Autoscaler (HPA)](https://knative.dev/docs/install/yaml-install/serving/install-serving-with-yaml/#install-optional-serving-extensions). Use the following command to install. Make sure to update the VERSION in the below command with a [supported Knative version](#inference).

```bash
kubectl apply -f https://github.com/knative/serving/releases/download/knative-{VERSION}/serving-hpa.yaml
```

### Distributed Inference

NVIDIA Run:ai supports distributed inference (multi-node) deployments using the Leader Worker Set (LWS). To enable this capability, install LWS on your cluster:

{% tabs %}
{% tab title="Kubernetes" %}
Install the [LWS Helm chart](https://lws.sigs.k8s.io/docs/installation/#install-by-helm) in version 0.7.0 or higher:

```bash
CHART_VERSION=0.7.0
helm install lws oci://registry.k8s.io/lws/charts/lws \
  --version=$CHART_VERSION \
  --namespace lws-system \
  --create-namespace \
  --wait --timeout 300s
```

{% endtab %}

{% tab title="OpenShift" %}

1. Install the Leader Worker Set Operator by following the [Installing the Leader Worker Set Operator](https://docs.redhat.com/en/documentation/openshift_container_platform/4.19/html/ai_workloads/leader-worker-set-operator#lws-install-operator_lws-managing) instructions in the OpenShift documentation. Only the operator installation step is required; the remaining sections can be ignored.
2. Configure the LWS namespace so NVIDIA Run:ai can detect it. The OpenShift LWS Operator installs into the `openshift-lws-operator` namespace by default. Choose one of the following methods:
   * **Helm flag:** Set `clusterConfig.lws.namespace=openshift-lws-operator` during cluster installation. For more details, see [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config).
   * **runaiconfig patch:** Apply the setting directly:

     ```bash
     oc patch runaiconfig runai -n runai --type merge \
       -p '{"spec":{"lws":{"namespace":"openshift-lws-operator"}}}'
     ```

{% endtab %}
{% endtabs %}

## Integrations

Integrations are Kubernetes components and external tools that can be used with NVIDIA Run:ai for development, training, orchestration, data access, and monitoring.

In many cases, the integration is “out of the box” from the NVIDIA Run:ai side. Once the component is installed in the cluster, you can submit its custom resource definitions (CRDs) and still benefit from NVIDIA Run:ai scheduling, resource management, and visibility. Examples include NVIDIA components such as the **NIM Operator** and **Dynamo Operator**. See [Integrations](/saas/infrastructure-setup/advanced-setup/integrations) for more details.


# Network Requirements

NVIDIA Run:ai requires certain network connectivity and access. This section outlines the network endpoints and protocols that must be reachable from your NVIDIA Run:ai control plane and cluster nodes to support installation, artifact retrieval, and ongoing platform communication.

Meeting these network requirements ensures that:

* Clusters can register with and communicate to the control plane
* The platform can access external services required for monitoring, logging, and artifact distribution

Follow the guidance below to verify and configure network access before proceeding with installation.

## External Access

Listed below are the domains to whitelist and ports to open for installation, upgrades, and usage of the application and its management.

{% hint style="info" %}
**Note**

Ensure the inbound and outbound rules are correctly applied to your firewall.
{% endhint %}

### Inbound Rules

To allow your organization’s NVIDIA Run:ai users to interact with the cluster using the [NVIDIA Run:ai Command-line interface](/saas/reference/cli), or access specific UI features, certain inbound ports need to be open.

| Name                  | Description      | Source  | Destination                | Port |
| --------------------- | ---------------- | ------- | -------------------------- | ---- |
| NVIDIA Run:ai cluster | HTTPS entrypoint | 0.0.0.0 | NVIDIA Run:ai system nodes | 443  |

### Outbound Rules

{% hint style="info" %}
**Note**

* **For IPv6-only environments** - `runai.jfrog.io` and `nvcr.io` only have IPv4 DNS records, so clients on IPv6-only networks cannot resolve them and image pulls will fail. Two options:
  * **Configure NAT64/DNS64** - Translates between IPv6 and IPv4 so the cluster reaches these registries transparently.
  * **Deploy an internal mirror registry** - Use Harbor, Artifactory, or a similar registry over IPv6, configured to pull from `runai.jfrog.io` and `nvcr.io` over IPv4. Point the cluster at the mirror through the container runtime config and NVIDIA Run:ai Helm image-registry overrides. Choose this option for air-gapped or strictly controlled networks.
* `gcr.io`, `quay.io`, and `docker.io` are reachable over IPv6 directly.
  {% endhint %}

For the NVIDIA Run:ai cluster installation and usage, certain **outbound** ports must be open:

| Name                       | Description                                                                      | Source                             | Destination                                                                                                                                                                                                                                                                           | Port |
| -------------------------- | -------------------------------------------------------------------------------- | ---------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---- |
| Cluster sync               | Sync NVIDIA Run:ai cluster with NVIDIA Run:ai control plane                      | NVIDIA Run:ai cluster system nodes | NVIDIA Run:ai control plane FQDN                                                                                                                                                                                                                                                      | 443  |
| Metric store               | Push NVIDIA Run:ai cluster metrics to NVIDIA Run:ai control plane's metric store | NVIDIA Run:ai cluster system nodes | NVIDIA Run:ai control plane FQDN                                                                                                                                                                                                                                                      | 443  |
| Container Registry         | Pull NVIDIA Run:ai images and Helm chart for installation                        | All kubernetes nodes               | runai.jfrog.io and [JFrog Cloud Storage URLs](https://jfrog.com/help/r/artifactory-what-urls-ips-should-i-add-to-an-allowlist-for-direct-cloud-storage-download-in-jfrog-saas/artifactory-what-urls/ips-should-i-add-to-an-allowlist-for-direct-cloud-storage-download-in-jfrog-saas) | 443  |
| NVIDIA Run:ai NGC Registry | Pull NVIDIA Run:ai images and Helm chart for installation                        | All Kubernetes nodes               | nvcr.io                                                                                                                                                                                                                                                                               | 443  |
| Helm repository            | NVIDIA Run:ai Helm repository for installation                                   | Installer machine                  | runai.jfrog.io                                                                                                                                                                                                                                                                        | 443  |

The NVIDIA Run:ai installation has [software requirements](/saas/getting-started/installation/install-using-helm/system-requirements#software-requirements) that require additional components to be installed on the cluster. This article includes simple installation examples which can be used optionally and require the following cluster outbound ports to be open:

| Name                       | Description                                | Source               | Destination | Port |
| -------------------------- | ------------------------------------------ | -------------------- | ----------- | ---- |
| Kubernetes Registry        | Ingress HAProxy image repository           | All kubernetes nodes | docker.io   | 443  |
| Google Container Registry  | GPU Operator, and Knative image repository | All kubernetes nodes | gcr.io      | 443  |
| Red Hat Container Registry | Prometheus Operator image repository       | All kubernetes nodes | quay.io     | 443  |
| Docker Hub Registry        | Training Operator image repository         | All kubernetes nodes | docker.io   | 443  |

## Internal Network

Ensure that all Kubernetes nodes can communicate with each other across all necessary ports. Kubernetes assumes full interconnectivity between nodes, so you must configure your network to allow this seamless communication. Specific port requirements may vary depending on your network setup.


# Install the Cluster

In this section you will install the NVIDIA Run:ai cluster on your Kubernetes environment using Helm. The cluster extends Kubernetes with NVIDIA Run:ai orchestration capabilities - scheduling and workload management - and connects to NVIDIA’s cloud-hosted control plane for centralized management.

Once you access the NVIDIA Run:ai UI for the first time, an onboarding wizard opens automatically. The wizard guides you through the cluster setup and generates a Helm installation command.

This procedure includes:

* Adding the NVIDIA Run:ai Helm repository from NGC or JFrog
* Installing the NVIDIA Run:ai cluster into the `runai` namespace
* Registering the NVIDIA Run:ai cluster with the NVIDIA Run:ai control plane using the provided connection details

By completing this process, the NVIDIA Run:ai cluster will be connected to NVIDIA’s cloud-hosted control plane and ready to run training, inference, and other workloads.

## System and Network Requirements

Before installing the NVIDIA Run:ai cluster, validate that the [system requirements](https://github.com/run-ai/runai-product-docs/blob/SaaS/getting-started/installation/system-requirements.md) and [network requirements](/saas/getting-started/installation/install-using-helm/network-requirements) are met.

Once all the requirements are met, it is highly recommended to use the NVIDIA Run:ai cluster preinstall diagnostics tool to:

* Test the below requirements in addition to failure points related to Kubernetes, NVIDIA, storage, and networking
* Look at additional components installed and analyze their relevance to a successful installation

To run the preinstall diagnostics tool, [download](https://runai.jfrog.io/ui/native/pd-cli-prod/preinstall-diagnostics-cli/) the latest version, and run:

```bash
chmod +x ./preinstall-diagnostics-<platform> && \
./preinstall-diagnostics-<platform> \
  --domain ${COMPANY_NAME}.run.ai \
  --cluster-domain ${CLUSTER_FQDN}
```

For more information, see [preinstall diagnostics](https://github.com/run-ai/preinstall-diagnostics).

## Helm

NVIDIA Run:ai cluster requires Helm 3.14 or above. To install Helm, see [Helm Install](https://helm.sh/docs/helm/helm_install/).

{% hint style="info" %}
**Note**

Helm 4 defaults to [server-side apply](https://helm.sh/docs/overview/#server-side-apply) when installing a new chart release, which can conflict with resources managed by the NVIDIA Run:ai operator. Append `--server-side=false` to your `helm upgrade` command. NVIDIA Run:ai clusters originally installed with Helm 3.x are unaffected.
{% endhint %}

## Permissions

Using a Kubernetes user with the `cluster-admin` role to ensure a successful installation is recommended. For more information, see [Using RBAC authorization](https://kubernetes.io/docs/reference/access-authn-authz/rbac/).

## Installation

Follow these instructions to install using Helm.

{% hint style="info" %}
**Note**

* To customize the installation based on your environment, see [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config).
* You can store the `clientSecret` as a Kubernetes secret within the cluster instead of using plain text. You can then configure the installation to use it by setting the `controlPlane.existingSecret` and `controlPlane.secretKeys.clientSecret` parameters as described in [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config).
  {% endhint %}

### Artifact Source <a href="#artifact-source" id="artifact-source"></a>

Starting with v2.24, NVIDIA Run:ai artifacts are available on both NVIDIA NGC and JFrog. NGC is the recommended artifact source. JFrog remains supported but will be removed in a future release.

### Adding a New Cluster

When adding a cluster for the first time, the onboarding wizard opens automatically when you log in to the NVIDIA Run:ai platform. You cannot perform other actions in the platform until the cluster is created.

1. Enter a unique name for your cluster and click **CONTINUE**.
2. Set your **Cluster URL**. Enter the Kubernetes cluster's URL. It will only be accessible within the organization network. For more information, see [Fully Qualified Domain Name (FQDN)](/saas/getting-started/installation/install-using-helm/system-requirements#fully-qualified-domain-name-fqdn).
3. Click **CONTINUE**

#### Installation Instructions

In the next section, the NVIDIA Run:ai cluster installation steps will be presented.

1. Before installing the NVIDIA Run:ai cluster, ensure that all required [system](/saas/getting-started/installation/install-using-helm/system-requirements) and [network](/saas/getting-started/installation/install-using-helm/network-requirements) requirements are met.
2. The NVIDIA Run:ai platform displays the Helm installation command in the cluster wizard. Follow the instructions for your artifact source.

{% tabs %}
{% tab title="NGC (Recommended)" %}
Modify the UI-generated command as follows:

* Add `--username='$oauthtoken'` and `--password=<NGC_API_KEY>` to the `helm repo add` command, and replace `<NGC_API_KEY>` with your NGC API key.
* If you are using a local certificate authority, add `--set global.customCA.enabled=true` to the Helm command as described in the [Local certificate authority](/saas/getting-started/installation/install-using-helm/system-requirements#local-certificate-authority) section.
* The recommended ingress controller is HAProxy. If you are using a different ingress controller, update the ingress class to match the ingress controller.

```bash
helm repo add runai https://helm.ngc.nvidia.com/nvidia/runai --force-update \
  --username='$oauthtoken' \
  --password=<NGC_API_KEY>
helm repo update
helm upgrade -i runai-cluster runai/runai-cluster -n runai \
    --set controlPlane.url=... \
    --set controlPlane.clientSecret=... \
    --set cluster.uid=... \
    --set cluster.url=... --version="<VERSION>" --create-namespace \
    --set clusterConfig.global.ingress.ingressClass=haproxy
```

{% endtab %}

{% tab title="JFrog" %}
Run the Helm commands exactly as shown in the UI.

* If you are using a local certificate authority, add `--set global.customCA.enabled=true` to the Helm command as described in the [Local certificate authority](/saas/getting-started/installation/install-using-helm/system-requirements#local-certificate-authority) section.
* The recommended ingress controller is HAProxy. If you are using a different ingress controller, update the ingress class to match the ingress controller.

```bash
helm repo add runai https://runai.jfrog.io/artifactory/api/helm/run-ai-charts --force-update
helm repo update
helm upgrade -i runai-cluster runai/runai-cluster -n runai \
  --set controlPlane.url=... \
  --set controlPlane.clientSecret=... \
  --set cluster.uid=... \
  --set cluster.url=... --version="<VERSION>" --create-namespace
  --set clusterConfig.global.ingress.ingressClass=haproxy
```

{% endtab %}
{% endtabs %}

The wizard displays **Waiting for cluster to connect** while the cluster is being installed and connected to the control plane. Once the installation completes successfully and the cluster establishes communication with the control plane, the wizard updates to **Cluster connected**. After completing the wizard flow, the cluster is added to the [Clusters](/saas/infrastructure-setup/procedures/clusters) table.

## Troubleshooting

If you encounter an issue with the installation, try the troubleshooting scenario below.

### Installation

If the NVIDIA Run:ai cluster installation failed, check the installation logs to identify the issue. Run the following script to print the installation logs:

{% file src="/files/MqckgdW2v7HB4rKj98sU" %}

### Cluster Status

If the NVIDIA Run:ai cluster installation completed, but the cluster status did not change its status to Connected, check the cluster [troubleshooting scenarios](/saas/infrastructure-setup/procedures/clusters#troubleshooting-scenarios).

## Next Steps

Once the cluster is installed and connected, the NVIDIA Run:ai UI guides you through optional post-installation configurations. These steps are optional but recommended for a production setup:

* **SSO (Single Sign-On)** - Configure SSO to allow users to log in with your organization's identity provider. See [SSO](/saas/infrastructure-setup/authentication/sso) for setup instructions using SAML or OpenID Connect.
* **Create your first research team** - Set up [projects](/saas/platform-management/aiinitiatives/organization/projects) and [permissions](/saas/infrastructure-setup/authentication/accessrules) to organize your AI practitioners and allocate GPU resources.


# Upgrade

This section explains how to upgrade NVIDIA Run:ai cluster version.

## System and Network Requirements

Before upgrading the NVIDIA Run:ai cluster, validate that the latest [system requirements](/saas/getting-started/installation/install-using-helm/system-requirements) and [network requirements](/saas/getting-started/installation/install-using-helm/network-requirements) are met, as they can change from time to time.

{% hint style="info" %}
**Note**

It is highly recommended to upgrade the Kubernetes version together with the NVIDIA Run:ai cluster version, to ensure compatibility with latest supported version of your [Kubernetes distribution](/saas/getting-started/installation/install-using-helm/system-requirements#kubernetes-distribution).
{% endhint %}

## Helm

The latest releases of the NVIDIA Run:ai cluster require [Helm 3.14](https://helm.sh/docs/helm/helm_install/) or above.

{% hint style="info" %}
**Note**

Helm 4 defaults to [server-side apply](https://helm.sh/docs/overview/#server-side-apply) when installing a new chart release, which can conflict with resources managed by the NVIDIA Run:ai operator. Append `--server-side=false` to your `helm upgrade` command. NVIDIA Run:ai clusters originally installed with Helm 3.x are unaffected.
{% endhint %}

## Upgrade

Follow the instructions to upgrade using Helm. The Helm commands to upgrade the NVIDIA Run:ai cluster version may differ between versions. The steps below describe how to get the instructions from the NVIDIA Run:ai UI.

### Getting Installation Instructions

Follow the setup and installation instructions below to get the installation instructions to upgrade the NVIDIA Run:ai cluster.

#### Setup

1. In the NVIDIA Run:ai UI, go to Resources -> Clusters
2. Select the cluster you want to upgrade
3. Click **INSTALLATION INSTRUCTIONS**
4. Optional: Select the NVIDIA Run:ai cluster version (latest, by default)
5. Click **CONTINUE**

#### Installation Instructions

In the next section, the NVIDIA Run:ai cluster installation steps will be presented.

1. Before installing the NVIDIA Run:ai cluster, ensure that all required [system](/saas/getting-started/installation/install-using-helm/system-requirements) and [network](/saas/getting-started/installation/install-using-helm/network-requirements) requirements are met.
2. The NVIDIA Run:ai platform displays the Helm installation command in the cluster wizard. Follow the instructions for your artifact source.

{% tabs %}
{% tab title="NGC" %}

1. Modify the UI-generated command as follows:
   * Replace `helm repo add` to pull from NGC instead of JFrog as shown below.
   * Add `--username='$oauthtoken'` and `--password=<NGC_API_KEY>` to the `helm repo add` command, and replace `<NGC_API_KEY>` with your NGC API key.
   * If you are using a local certificate authority, add `--set global.customCA.enabled=true` to the Helm command as described in the [Local certificate authority](/saas/getting-started/installation/install-using-helm/system-requirements#local-certificate-authority) section.
   * The recommended ingress controller is HAProxy. If you are using a different ingress controller, update the ingress class to match the ingress controller.

<pre class="language-bash"><code class="lang-bash">helm repo add runai https://helm.ngc.nvidia.com/nvidia/runai --force-update \
<strong>  --username='$oauthtoken' \
</strong>  --password=&#x3C;NGC_API_KEY>
helm repo update
helm upgrade -i runai-cluster runai/runai-cluster -n runai \
  --set controlPlane.url=... \
  --set controlPlane.clientSecret=... \
  --set cluster.uid=... \
  --set cluster.url=... --version="&#x3C;VERSION>" --create-namespace
  --set clusterConfig.global.ingress.ingressClass=haproxy
</code></pre>

2. Click **DONE**

Once installation is complete, validate the cluster is **Connected** and listed with the new cluster version. Once you have done this, the cluster is upgraded to the latest version. If the cluster does not appear as expected, see the [cluster troubleshooting scenarios](/saas/infrastructure-setup/procedures/clusters#troubleshooting-scenarios).
{% endtab %}

{% tab title="JFrog" %}

1. Run the Helm commands exactly as shown in the UI.
   * If you are using a local certificate authority, add `--set global.customCA.enabled=true` to the Helm command as described in the [Local certificate authority](/saas/getting-started/installation/install-using-helm/system-requirements#local-certificate-authority) section.
   * The recommended ingress controller is HAProxy. If you are using a different ingress controller, update the ingress class to match the ingress controller.

```bash
helm repo add runai https://runai.jfrog.io/artifactory/api/helm/run-ai-charts --force-update
helm repo update
helm upgrade -i runai-cluster runai/runai-cluster -n runai \
  --set controlPlane.url=... \
  --set controlPlane.clientSecret=... \
  --set cluster.uid=... \
  --set cluster.url=... --version="<VERSION>" --create-namespace
  --set clusterConfig.global.ingress.ingressClass=haproxy
```

2. Click **DONE**

Once installation is complete, validate the cluster is **Connected** and listed with the new cluster version. Once you have done this, the cluster is upgraded to the latest version. If the cluster does not appear as expected, see the [cluster troubleshooting scenarios](/saas/infrastructure-setup/procedures/clusters#troubleshooting-scenarios).
{% endtab %}
{% endtabs %}

## Migrate from NGINX to HAProxy Kubernetes Ingress Controller

{% hint style="info" %}
**Note**

This section applies to Kubernetes only. OpenShift includes a pre-installed ingress controller by default and does not require this migration.
{% endhint %}

Starting with v2.24, NVIDIA Run:ai recommends using [HAProxy Kubernetes Ingress Controller](https://www.haproxy.com/documentation/kubernetes-ingress/) as the ingress controller. This change aligns with the announced retirement of the upstream NGINX Ingress Controller project.

Clusters upgraded from earlier versions typically already have NGINX installed. After upgrading to v2.24, follow the steps below to migrate ingress traffic from NGINX Ingress Controller to HAProxy Kubernetes Ingress Controller.

{% hint style="info" %}
**Migration Resources**

You can utilize additional resources to streamline your migration. Visit the [HAProxy Migration Center](https://www.haproxy.com/landing/ingress-nginx-retirement) to access a dedicated migration assistant tool and a technical webinar on moving from NGINX to HAProxy.
{% endhint %}

### Check the Service Type of the Existing Ingress Controller

Before installing the HAProxy Kubernetes Ingress Controller, identify which ingress controller is currently in use. If your cluster already has an ingress controller installed, verify how it is exposed to avoid port address conflicts.

```bash
kubectl get svc -n <nginx-namespace>
```

* If the existing ingress controller uses **NodePort**, note the HTTP/HTTPS NodePort values to ensure HAProxy is configured with non-overlapping ports.
* If the existing ingress controller uses **LoadBalancer**, no additional action is required.

When running more than one ingress controller in the same cluster, port conflicts are relevant only for NodePort-based setups. LoadBalancer-based controllers automatically receive separate external IP addresses.

{% hint style="info" %}
**Note**

If your setup differs from the examples above, adjust the configuration accordingly. When using external LoadBalancer on top of Ingress with service type NodePort, you may need to update external resources to route traffic to HAProxy’s configured NodePort values.
{% endhint %}

### Install and Configure HAProxy Kubernetes Ingress Controller

Ingress controllers can be installed and configured in different ways depending on your Kubernetes distribution and how you expose services (for example, NodePort vs. LoadBalancer).

The sections below provide environment-specific Helm installation examples. Select the option that matches your deployment environment.

{% hint style="info" %}
**Note**

OpenShift and RKE2 include a pre-installed ingress controller by default.
{% endhint %}

<details>

<summary>Vanilla Kubernetes</summary>

If your cluster already has an ingress controller installed (for example, NGINX) and it is exposed via NodePort, configure HAProxy to use different NodePort values so both controllers can run simultaneously.

**Ensure the selected NodePort values do not overlap with ports already used by the existing ingress controller.**

<pre class="language-bash"><code class="lang-bash"><strong>helm repo add haproxytech https://haproxytech.github.io/helm-charts
</strong>helm repo update
helm install haproxy-kubernetes-ingress haproxytech/kubernetes-ingress \
  --create-namespace \
  --namespace haproxy-controller \
  --set controller.ingressClassResource.enabled=true \
  --set controller.service.type=NodePort \
  --set controller.service.nodePorts.http=32080 \
  --set controller.service.nodePorts.https=32443
</code></pre>

</details>

<details>

<summary>Managed Kubernetes (EKS, GKE, AKS)</summary>

When using a LoadBalancer, each ingress controller automatically receives its own external IP address from the cloud provider. This allows multiple ingress controllers to run in the same cluster without additional configuration.

```bash
helm repo add haproxytech https://haproxytech.github.io/helm-charts
helm repo update
helm install haproxy-kubernetes-ingress haproxytech/kubernetes-ingress \
--create-namespace \
--namespace haproxy-controller \
--set controller.service.type=LoadBalancer \
```

</details>

<details>

<summary>Oracle Kubernetes Engine (OKE)</summary>

When using a LoadBalancer, each ingress controller automatically receives its own external IP address from the cloud provider. This allows multiple ingress controllers to run in the same cluster without additional configuration.

```bash
helm repo add haproxytech https://haproxytech.github.io/helm-charts
helm repo update
helm install haproxy-kubernetes-ingress haproxytech/kubernetes-ingress \
  --create-namespace \
  --namespace haproxy-controller \
  --set controller.kind=DaemonSet \
  --set controller.service.type=LoadBalancer \
  --set controller.service.externalTrafficPolicy=Local \
  --set controller.service.annotations."oci-network-load-balancer\.oraclecloud\.com/is-preserve-source"="True" \
  --set controller.service.annotations."oci-network-load-balancer\.oraclecloud\.com/security-list-management-mode"=All \
  --set controller.service.annotations."oci\.oraclecloud\.com/load-balancer-type"=nlb
```

</details>

### Verify HAProxy Ingress

After installing the HAProxy Kubernetes Ingress Controller, verify that HAProxy ingresses are reachable before switching NVIDIA Run:ai components to use it. You can do this by deploying a simple hello-world application.

To run the test, identify the IP address that should reach the cluster’s nodes in your environment.

1. Create a local `haproxy-test.yml` file:

   ```yaml
   apiVersion: apps/v1
   kind: Deployment
   metadata:
     name: hello
   spec:
     replicas: 1
     selector:
       matchLabels:
         app: hello
     template:
       metadata:
         labels:
           app: hello
       spec:
         containers:
         - name: hello
           image: hashicorp/http-echo:1.0
           args:
             - "-text=hello from haproxy-ingress"
           ports:
             - containerPort: 5678
   ---
   apiVersion: v1
   kind: Service
   metadata:
     name: hello
   spec:
     selector:
       app: hello
     ports:
     - port: 80
       targetPort: 5678
   ---
   apiVersion: networking.k8s.io/v1
   kind: Ingress
   metadata:
     name: hello
   spec:
     ingressClassName: haproxy
     rules:
     - http:
         paths:
         - path: /
           pathType: Prefix
           backend:
             service:
               name: hello
               port:
                 number: 80
   ```
2. Run the following command:

   ```yaml
   kubectl apply -f ha-proxy-test.yml
   ```

Once the application is deployed, access the cluster’s IP address in a browser. If the page displays **“hello from haproxy-ingress”**, HAProxy is functioning correctly and you can proceed with upgrading NVIDIA Run:ai.

### Upgrade the Cluster

#### Setup

1. In the NVIDIA Run:ai UI, go to Resources -> Clusters
2. Select the cluster you want to upgrade
3. Click **INSTALLATION INSTRUCTIONS**
4. Click **CONTINUE**

#### Installation Instructions

1. Follow the installation instructions. Run the Helm commands provided on your Kubernetes cluster.
2. If not present, add the following flag to the helm install command:

   ```bash
   --set clusterConfig.global.ingress.ingressClass=haproxy
   ```
3. Click **DONE**
4. Once installation is complete, validate the cluster is **Connected** and listed with the new cluster version (see the [cluster troubleshooting scenarios](/saas/infrastructure-setup/procedures/clusters#troubleshooting-scenarios)). Once you have done this, the cluster is upgraded and the workloads in this cluster will now use HAProxy instead of NGINX.

## Troubleshooting

If you encounter an issue with the cluster upgrade, use the troubleshooting scenarios below.

### Installation Fails

If the NVIDIA Run:ai cluster installation failed, check the installation logs to identify the issue. Run the following script to print the installation logs:

{% file src="/files/MqckgdW2v7HB4rKj98sU" %}

### Cluster Status

If the NVIDIA Run:ai cluster upgrade completes, but the cluster status does not show as **Connected**, refer to the [cluster troubleshooting scenarios](/saas/infrastructure-setup/procedures/clusters#troubleshooting-scenarios).


# Uninstall

This section explains how to uninstall the NVIDIA Run:ai cluster from the Kubernetes cluster.

To uninstall the NVIDIA Run:ai cluster, run the following [helm](https://helm.sh/) command in your terminal:

```bash
helm uninstall runai-cluster -n runai
```

To remove the NVIDIA Run:ai cluster from the NVIDIA Run:ai platform, see [Removing a cluster](/saas/infrastructure-setup/procedures/clusters#removing-a-cluster).

{% hint style="info" %}
**Note**

Uninstall of NVIDIA Run:ai cluster from the Kubernetes cluster does **not** delete existing projects, departments or workloads submitted by users.
{% endhint %}


# Install Using Base Command Manager (BCM)

This section explains the steps required to install the NVIDIA Run:ai cluster on a DGX Kubernetes Cluster using NVIDIA [Base Command Manager (BCM)](https://docs.nvidia.com/base-command-manager/index.html).

## NVIDIA Run:ai Installer

The NVIDIA Run:ai installer is a wizard that simplifies the deployment of the NVIDIA Run:ai cluster on DGX. The NVIDIA Run:ai installer is installed via the BCM cluster wizard when the cluster is created.

{% hint style="info" %}
**Note**

For custom deployment options, check the [Install using Helm](/saas/getting-started/installation/install-using-helm/helm-install).
{% endhint %}

## System and Network Requirements

Before installing the NVIDIA Run:ai cluster on a DGX system using BCM, ensure that your [system requirements](/saas/getting-started/installation/install-using-helm/system-requirements) and [network requirements](/saas/getting-started/installation/install-using-helm/network-requirements) meets the necessary prerequisites.

The BCM cluster wizard deploys essential [software requirements](/saas/getting-started/installation/install-using-helm/system-requirements#software-requirements), such as the [Kubernetes Ingress Controller](/saas/getting-started/installation/install-using-helm/system-requirements#kubernetes-ingress-controller), [NVIDIA GPU Operator](/saas/getting-started/installation/install-using-helm/system-requirements#nvidia-gpu-operator), and [Prometheus](/saas/getting-started/installation/install-using-helm/system-requirements#prometheus), as part of the NVIDIA Run:ai Installer deployment. Additional optional software requirements for [Distributed training](/saas/getting-started/installation/install-using-helm/system-requirements#distributed-training) and [Inference](/saas/getting-started/installation/install-using-helm/system-requirements#inference) requires manual setup.

## Tenant Name

Your tenant name is predefined and supplied by NVIDIA Run:ai. Each customer is provided with a unique, dedicated URL in the format `<tenant-name>.run.ai` which includes the required tenant name.

## Application Secret Key

An application secret key is required to connect the cluster to the NVIDIA Run:ai platform. In order to get the application secret key, a new cluster must be added.

1. Follow the [Adding a new cluster](/saas/getting-started/installation/install-using-helm/helm-install) setup instructions. **Do not follow the Installation instructions**.
2. Once cluster instructions are displayed, find the `controlPlane.clientSecret` flag in the displayed Helm command, copy and save its value.

## TLS Certificate

A TLS private and public keys for the cluster’s [Fully Qualified Domain Name (FQDN)](/saas/getting-started/installation/install-using-helm/system-requirements#fully-qualified-domain-name-fqdn) are required for HTTP access to the cluster

{% hint style="info" %}
**Note**

TLS Certificate must be trusted. Self-signed certificates are not supported.
{% endhint %}

## Installation

Follow these instructions to install using BCM.

### Installing a Cluster

The cluster installer is available via the locally installed BCM landing page,

1. Go to the locally installed BCM landing page, Select the NVIDIA Run:ai tile or access directly to `http://<BCM-CLUSTER-IP>:30080/runai-installer` (HTTP only)\
   ![image(56).png](/files/6VBJoGEZDrlD5aRYnNjr)
2. Click **VERIFY** in order to check System Requirements are met.\
   ![image(57).png](/files/o5o2LdnYMcrXL8RDewm8)
3. After verification completed successfully, click **CONTINUE**.\
   ![image(58).png](/files/eWqgsD5rcu3eCMaIVDF0)
4. Enter the cluster information and click **CONTINUE**.\
   ![image(61).png](/files/oVoqUbuuCbcpC8YXycBP)
5. The NVIDIA Run:ai installation will start and should be complete within a few minutes\
   ![image(63).png](/files/bcgiLXK6sfNltPh4uTvC)
6. Once a message of **NVIDIA Run:ai was installed successfully!** is displayed, Click on **START USING NVIDIA Run:ai** to launch the login page of the tenant in a new browser tab.\
   ![image(62).png](/files/J7p84T3qQiTaiM6Biiy8)

## Troubleshooting

If you encounter an issue with the installation, try the troubleshooting scenario below.

### NVIDIA Run:ai Installer

The NVIDIA Run:ai installer is a pod in Kubernetes. The pod is responsible for the installation preparation and prerequisite gathering phase. If there is an error during the prerequisite verification process, run the following command to print the logs:

```bash
kubectl get pods -n runai | grep 'cluster-installer' #Find the cluster installer pod's name
kubectl logs <POD-NAME> -n runai #Print the cluster installer pod logs
```

### Installation

If the NVIDIA Run:ai cluster installation failed, check the installation logs to identify the issue. Run the following script to print the installation logs:

{% file src="/files/MqckgdW2v7HB4rKj98sU" %}

### Cluster Status

If the NVIDIA Run:ai cluster installation is complete but the cluster status did not change to **Connected**, check the cluster [troubleshooting scenarios](/saas/infrastructure-setup/procedures/clusters#troubleshooting-scenarios)


# Authentication and Authorization


# Authentication and Authorization

NVIDIA Run:ai authentication and authorization enables a streamlined experience for the user with precise controls covering the data each user can see and the actions each user can perform in the NVIDIA Run:ai platform.

Authentication verifies user identity during login, and authorization assigns the user with specific permissions according to the assigned [access rules](/saas/infrastructure-setup/authentication/accessrules).

Authenticated access is required to use all aspects of the NVIDIA Run:ai interfaces, including the NVIDIA Run:ai platform, the NVIDIA Run:ai Command Line Interface (CLI) and APIs.

## Authentication

There are multiple methods to authenticate and access NVIDIA Run:ai.

### Single Sign-On (SSO)

NVIDIA Run:ai supports three methods to set up SSO:

* [SAML](/saas/infrastructure-setup/authentication/sso/saml)
* [OpenID Connect (OIDC)](/saas/infrastructure-setup/authentication/sso/openidconnect)
* [OpenShift](/saas/infrastructure-setup/authentication/sso/openshift)

When using SSO, it is highly recommended to manage at least one local user, as a breakglass account (an emergency account), in case access to SSO is not possible.

### Username and Password

Username and password access can be used when SSO integration is not possible.

### Secret Key (for Application Programmatic Access)

Secret is the authentication method for [Service accounts](/saas/infrastructure-setup/authentication/service-accounts). Service accounts use the NVIDIA Run:ai APIs to perform automated tasks including scripts and pipelines based on their assigned [access rules](/saas/infrastructure-setup/authentication/accessrules).

## Authorization

The NVIDIA Run:ai platform uses Role Based Access Control (RBAC) to manage authorization. Once a user or service account is authenticated, they can perform actions according to their assigned access rules.

### Role Based Access Control (RBAC) in NVIDIA Run:ai

While Kubernetes RBAC is limited to a single cluster, NVIDIA Run:ai expands the scope of Kubernetes RBAC, making it easy for administrators to manage access rules across multiple clusters.

RBAC at NVIDIA Run:ai is configured using access rules. An access rule is the assignment of a [role](/saas/infrastructure-setup/authentication/roles) to a [subject in a scope](/saas/platform-management/aiinitiatives/adapting-ai-initiatives#scopes-in-an-organization): `<Subject>` is a `<Role>` in a `<Scope>`.

* **Subject**
  * A user, group, or service account assigned with the role
* **Role**
  * A set of permissions that can be assigned to subjects. NVIDIA Run:ai provides predefined roles that are system-defined and cannot be edited or deleted. Administrators can also create custom roles via the API to meet specific organizational requirements.
  * A permission is the actions (view, edit, create and delete) that the user can perform over a NVIDIA Run:ai entity (e.g. projects, workloads, users). For example, a role might allow a user to create and read Projects, but not update or delete them. For a reference of all available permissions, see [Permissions](/saas/infrastructure-setup/authentication/permissions)
* **Scope**
  * A scope is part of an organization in which a set of permissions (roles) is effective. Scopes include Projects, Departments, Clusters, Account (all clusters).

Below is an example of an access rule: **<username@company.com>** is a **Department admin** in **Department: A**

![](/files/LXZqZFi9yHgBkyD8d3YP)


# Users

Users can be managed locally, or via the identity provider (Idp), while assigned with [access rules](/saas/infrastructure-setup/authentication/accessrules) to manage permissions. For example, user **<user@domain.com>** is a **department admin** in **department A**.

## Users Table

The Users table can be found under **Access** in the NVIDIA Run:ai platform.

The users table provides a list of all the users in the platform.\
You can manage users and user permissions (access rules) for both local and [SSO users](/saas/infrastructure-setup/authentication/overview#single-sign-on-sso).

{% hint style="info" %}
**Single Sign-On Users**

SSO users are managed by the identity provider and appear once they have signed in to NVIDIA Run:ai.
{% endhint %}

<figure><img src="/files/t3yJ5H1pMakbt4ZMdtk6" alt=""><figcaption></figcaption></figure>

The Users table consists of the following columns:

| Column         | Description                                        |
| -------------- | -------------------------------------------------- |
| User           | The unique identity of the user (email address)    |
| Full name      | The user's first and last name                     |
| Type           | The type of the user - SSO / local                 |
| Last login     | The timestamp for the last time the user signed in |
| Access rule(s) | The access rule assigned to the user               |
| Created By     | The user who created the user                      |
| Creation time  | The timestamp for when the user was created        |
| Last updated   | The last time the user was updated                 |

### Customizing the Table View

* Filter - Click ADD FILTER, select the column to filter by, and enter the filter values
* Search - Click SEARCH and type the value to search by
* Sort - Click each column header to sort by
* Column selection - Click COLUMNS and select the columns to display in the table
* Download table - Click MORE and then Click Download as CSV. Export to CSV is limited to 20,000 rows.

## Creating a Local User

To create a local user:

1. Click **+NEW LOCAL USER**
2. Enter the user’s **Email address**
3. (Optional) Enter the user’s **First name** and **Last name**. The full name is displayed alongside the user’s email across the platform.
4. (Optional) Select the checkbox to send an email invitation to the new user. The invitation allows the user to sign in and get started. To enable email invitations, make sure email notifications are configured under [**General settings**](/saas/settings/general-settings/notifications).
5. Click **CREATE**
6. Review and copy the user’s credentials:
   * **User Email**
   * **Temporary password** to be used on first sign-in
7. Click **DONE** or click **ADD ACCESS RULES** to proceed to the **Access rules** step
8. In the **Access rules** step:
   * Select a **role**
   * Select up to 10 **scopes** where the access rule will apply
   * Click **SAVE RULE**
   * Click **CLOSE**

{% hint style="info" %}
**Note**

The temporary password is visible only at the time of user’s creation and must be changed after the first sign-in.
{% endhint %}

## Editing a User

Administrators can update the first and last name of an existing local user. Names for SSO users are sourced from the identity provider and cannot be edited in NVIDIA Run:ai.

To edit a local user:

1. Select the local user you want to edit
2. Click **EDIT**
3. Update the user’s **First name** and **Last name**
4. Click **SAVE**

## Adding an Access Rule to a User

To create an access rule:

1. Select the user you want to add an access rule for
2. Click **ACCESS RULES**
3. Click **+ACCESS RULE**
4. Select a **role**
5. Select up to 10 **scopes** where the access rule will apply
6. Click **SAVE RULE**
7. Click **CLOSE**

## Deleting User’s Access Rule

To delete an access rule:

1. Select the user you want to remove an access rule from
2. Click **ACCESS RULES**
3. Find the access rule assigned to the user you would like to delete
4. Click on the trash icon
5. Click **CLOSE**

## Resetting a User Password

To reset a user’s password:

1. Select the user you want to reset it’s password
2. Click **RESET PASSWORD**
3. Click **RESET**
4. Review and copy the user’s credentials:
   * **User Email**
   * **Temporary password** to be used on next sign-in
5. Click **DONE**

## Deleting a User

1. Select the user you want to delete
2. Click **DELETE**
3. In the dialog, click **DELETE** to confirm

{% hint style="info" %}
**Note**

To ensure administrative operations are always available, at least one local user with System Administrator role should exist.
{% endhint %}

## Using API

Go to the [Users](https://run-ai-docs.nvidia.com/api/authentication-and-authorization/users), [Access rules](https://run-ai-docs.nvidia.com/api/authentication-and-authorization/access-rules) API reference to view the available actions.


# SSO


# Set Up SSO with SAML

Single Sign-On (SSO) is an authentication scheme, allowing users to log in with a single pair of credentials to multiple, independent software systems.

This guide explains the procedure to [configure SSO to NVIDIA Run:ai](/saas/infrastructure-setup/authentication/overview#single-sign-on-sso) using the SAML 2.0 protocol.

## Prerequisites

Before you start, make sure you have the **IDP Metadata** **XML** available from your identity provider.

## Setup

### Adding the Identity Provider

1. Go to **General settings**
2. Open the Security section and click **+IDENTITY PROVIDER**
3. Select **Custom SAML 2.0**
4. Select either **From computer** or **From URL** to upload your identity provider metadata file
   * **From computer** - Click the Metadata XML file field, then select your file for upload
   * **From URL** - In the Metadata XML field, enter the URL to the IDP Metadata XML file
5. You can either copy the **Redirect URL** and **Entity ID** displayed on the screen and enter them in your identity provider, or use the **service provider metadata XML**, which contains the same information in XML format. This file becomes available after you click **SAVE** in step 7.
6. Optional: Enter the user attributes and their value in the identity provider as shown in the below table
7. Click **SAVE.** After save, click **Open service provider metadata XML** to access the metadata file. This file can be used to configure your identity provider.
8. Optional: Enable **Auto-Redirect to SSO** to automatically redirect users to your configured identity provider’s login page when accessing the platform.

| Attribute            | Default value in NVIDIA Run:ai | Description                                                                                                                                                                                                  |
| -------------------- | ------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| User role groups     | GROUPS                         | If it exists in the IDP, it allows you to assign NVIDIA Run:ai role groups via the IDP. The IDP attribute must be a list of strings.                                                                         |
| Linux User ID        | UID                            | If it exists in the IDP, it allows Researcher containers to start with the Linux User UID. Used to map access to network resources such as file systems to users. The IDP attribute must be of type integer. |
| Linux Group ID       | GID                            | If it exists in the IDP, it allows Researcher containers to start with the Linux Group GID. The IDP attribute must be of type integer.                                                                       |
| Supplementary Groups | SUPPLEMENTARYGROUPS            | If it exists in the IDP, it allows Researcher containers to start with the relevant Linux supplementary groups. The IDP attribute must be a list of integers.                                                |
| Email                | email                          | Defines the user attribute in the IDP holding the user's email address, which is the user identifier in NVIDIA Run:ai.                                                                                       |
| User first name      | firstName                      | Used as the user’s first name appearing in the NVIDIA Run:ai platform.                                                                                                                                       |
| User last name       | lastName                       | Used as the user’s last name appearing in the NVIDIA Run:ai platform.                                                                                                                                        |

### Testing the Setup

1. Open the NVIDIA Run:ai platform as an admin
2. Add [access rules](/saas/infrastructure-setup/authentication/accessrules) to an SSO user defined in the IDP
3. Open the NVIDIA Run:ai platform in an incognito browser tab
4. On the sign-in page click **CONTINUE WITH SSO.**\
   You are redirected to the identity provider sign in page
5. In the identity provider sign-in page, log in with the SSO user who you granted with access rules
6. If you are unsuccessful signing-in to the identity provider, follow the [Troubleshooting](#troubleshooting) section below

### Editing the Identity Provider

You can view the identity provider details and edit its configuration:

1. Go to **General settings**
2. Open the Security section
3. On the identity provider box, click **Edit identity provider**
4. You can edit either the metadata file or the user attributes
5. You can view the identity provider URL, identity provider entity ID, and the certificate expiration date

### Removing the Identity Provider

You can remove the identity provider configuration:

1. Go to **General settings**
2. Open the Security section
3. On the identity provider card, click **Remove identity provider**
4. In the dialog, click **REMOVE** to confirm the action

{% hint style="info" %}
**Note**

To avoid losing access, removing the identity provider must be carried out by a local user.
{% endhint %}

## Downloading the IDP Metadata XML File

You can download the XML file to view the identity provider settings:

1. Go to **General settings**
2. Open the Security section
3. On the identity provider card, click **Edit identity provider**
4. In the dialog, click **DOWNLOAD IDP METADATA XML FILE**

## Troubleshooting

If testing the setup was unsuccessful, try the different troubleshooting scenarios according to the error you received. If an error still occurs, check the [advanced troubleshooting section](#advanced-troubleshooting).

### Troubleshooting Scenarios

<details>

<summary><strong>Error:</strong> "Invalid signature in response from identity provider"</summary>

**Description**: After trying to log in, the following message is received in the NVIDIA Run:ai login page.

**Mitigation:**

1. Go to the **General settings** menu
2. Open the Security section
3. In the identity provider box, check for a "Certificate expired” error
4. If it is expired, update the SAML metadata file to include a valid certificate

</details>

<details>

<summary><strong>Error:</strong> "401 - We’re having trouble identifying your account because your email is incorrect or can’t be found."</summary>

**Description:** Authentication failed because email attribute was not found.

**Mitigation**: Validate the user’s email attribute is mapped correctly

</details>

<details>

<summary><strong>Error:</strong> "403 - Sorry, we can’t let you see this page. Something about permissions…"</summary>

**Description:** The authenticated user is missing permissions

**Mitigation**:

1. Validate either the user or its related group/s are assigned with [access rules](/saas/infrastructure-setup/authentication/accessrules)
2. Validate the user’s groups attribute is mapped correctly

**Advanced:**

1. Open the Chrome DevTools: Right-click on page → Inspect → Network tab
2. Navigate to the Clusters page to trigger an API request
3. In the Network tab, find the request to the Clusters API
4. In the request headers, copy the value of the Authorization header (the token starts with `Bearer`)
5. Paste the token value (without the `Bearer` prefix) in <https://jwt.io>
6. Under the Payload section validate the values of the user's attributes

</details>

### Advanced Troubleshooting

<details>

<summary>Validating the SAML request</summary>

The SAML login flow can be separated into two parts:

* NVIDIA Run:ai redirects to the IDP for log-ins using a SAML Request
* On successful log-in, the IDP redirects back to NVIDIA Run:ai with a SAML Response

Validate the SAML Request to ensure the SAML flow works as expected:

1. Go to the NVIDIA Run:ai login screen
2. Open the Chrome Network inspector: Right-click → Inspect on the page → Network tab
3. On the sign-in page click CONTINUE WITH SSO.
4. Once redirected to the Identity Provider, search in the Chrome network inspector for an HTTP request showing the SAML Request. Depending on the IDP url, this would be a request to the IDP domain name. For example, `accounts.google.com/idp?1234`.
5. When found, go to the Payload tab and copy the value of the SAML Request
6. Paste the value into a SAML decoder (e.g. <https://www.samltool.com/decode.php>)
7. Validate the request:
   * The content of the `<saml:Issuer>` tag is the same as `Entity ID` given when [adding the identity provider](#adding-the-identity-provider)
   * The content of the `AssertionConsumerServiceURL` is the same as the `Redirect URI` given when [adding the identity provider](#adding-the-identity-provider)
8. Validate the response:
   * The user email under the `<saml2:Subject>` tag is the same as the logged-in user
   * Make sure that under the `<saml2:AttributeStatement>` tag, there is an Attribute named `email` (lowercase). This attribute is mandatory.
   * If other, optional user attributes (`groups`, `firstName`, `lastName`, `uid`, `gid`) are mapped make sure they also exist under `<saml2:AttributeStatement>` along with their respective values.

</details>


# Set Up SSO with OpenID Connect

Single Sign-On (SSO) is an authentication scheme, allowing users to log-in with a single pair of credentials to multiple, independent software systems.

This guide explains the procedure to [configure SSO to NVIDIA Run:ai](/saas/infrastructure-setup/authentication/overview#single-sign-on-sso) using the OpenID Connect protocol.

## Prerequisites

Before you start, make sure you have the following available from your identity provider:

* Discovery URL - The OpenID server where the content discovery information is published.
* ClientID - The ID used to identify the client with the Authorization Server.
* Client Secret - A secret password that only the Client and Authorization server know.
* Optional: Scopes - A set of user attributes to be used during authentication to authorize access to a user's details.

## Setup

### Adding the Identity Provider

1. Go to **General settings**
2. Open the Security section and click **+IDENTITY PROVIDER**
3. Select **Custom OpenID Connect**
4. Enter the **Discovery URL**, **Client ID**, and **Client Secret**
5. Copy the Redirect URL to be used in your identity provider
6. Optional: Add the OIDC scopes
7. Optional: Enter the user attributes and their value in the identity provider as shown in the below table
8. Click **SAVE**
9. Optional: Enable **Auto-Redirect to SSO** to automatically redirect users to your configured identity provider’s login page when accessing the platform.

| Attribute            | Default value in NVIDIA Run:ai | Description                                                                                                                                                                                                  |
| -------------------- | ------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| User role groups     | GROUPS                         | If it exists in the IDP, it allows you to assign NVIDIA Run:ai role groups via the IDP. The IDP attribute must be a list of strings or an object where the group names are the values.                       |
| Linux User ID        | UID                            | If it exists in the IDP, it allows Researcher containers to start with the Linux User UID. Used to map access to network resources such as file systems to users. The IDP attribute must be of type integer. |
| Linux Group ID       | GID                            | If it exists in the IDP, it allows Researcher containers to start with the Linux Group GID. The IDP attribute must be of type integer.                                                                       |
| Supplementary Groups | SUPPLEMENTARYGROUPS            | If it exists in the IDP, it allows Researcher containers to start with the relevant Linux supplementary groups. The IDP attribute must be a list of integers.                                                |
| Email                | email                          | Defines the user attribute in the IDP holding the user's email address, which is the user identifier in NVIDIA Run:ai                                                                                        |
| User first name      | firstName                      | Used as the user’s first name appearing in the NVIDIA Run:ai user interface                                                                                                                                  |
| User last name       | lastName                       | Used as the user’s last name appearing in the NVIDIA Run:ai user interface                                                                                                                                   |

### Testing the Setup

1. Log in to the NVIDIA Run:ai platform as an admin
2. Add [access rules](/saas/infrastructure-setup/authentication/accessrules) to an SSO user defined in the IDP
3. Open the NVIDIA Run:ai platform in an incognito browser tab
4. On the sign-in page click **CONTINUE WITH SSO**\
   You are redirected to the identity provider sign in page
5. In the identity provider sign-in page, log in with the SSO user who you granted with access rules
6. If you are unsuccessful signing-in to the identity provider, follow the [Troubleshooting](#troubleshooting) section below

### Editing the Identity Provider

You can view the identity provider details and edit its configuration:

1. Go to **General settings**
2. Open the Security section
3. On the identity provider box, click **Edit identity provider**
4. You can edit either the **Discovery URL**, **Client ID**, **Client Secret**, **OIDC scopes**, or the **User attributes**

### Removing the Identity Provider

You can remove the identity provider configuration:

1. Go to **General settings**
2. Open the Security section
3. On the identity provider card, click **Remove identity provider**
4. In the dialog, click **REMOVE** to confirm the action

{% hint style="info" %}
**Note**

To avoid losing access, removing the identity provider must be carried out by a local user.
{% endhint %}

## Troubleshooting

If testing the setup was unsuccessful, try the different troubleshooting scenarios according to the error you received.

### Troubleshooting Scenarios

<details>

<summary><strong>Error:</strong> "403 - Sorry, we can’t let you see this page. Something about permissions…"</summary>

**Description:** The authenticated user is missing permissions

**Mitigation**:

1. Validate either the user or its related group/s are assigned with [access rules](/saas/infrastructure-setup/authentication/accessrules)
2. Validate groups attribute is available in the configured OIDC Scopes
3. Validate the user’s groups attribute is mapped correctly

**Advanced:**

1. Open the Chrome DevTools: Right-click on page → Inspect → Console tab
2. Run the following command to retrieve and paste the user’s token: `localStorage.token;`
3. Paste in [https://jwt.io](https://jwt.io/)
4. Under the Payload section validate the values of the user’s attribute

</details>

<details>

<summary><strong>Error:</strong> "401 - We’re having trouble identifying your account because your email is incorrect or can’t be found."</summary>

**Description:** Authentication failed because email attribute was not found.

**Mitigation**:

1. Validate email attribute is available in the configured OIDC Scopes
2. Validate the user’s email attribute is mapped correctly

</details>

<details>

<summary><strong>Error:</strong> "Unexpected error when authenticating with identity provider"</summary>

**Description:** User authentication failed

<div align="left"><figure><img src="/files/DI2xJ1NGogtzgJyadX49" alt="" width="375"><figcaption></figcaption></figure></div>

**Mitigation**: Validate the configured OIDC Scopes exist and match the Identity Provider’s available scopes

**Advanced:** Look for the specific error message in the URL address

</details>

<details>

<summary><strong>Error:</strong> "Unexpected error when authenticating with identity provider (SSO sign-in is not available)"</summary>

**Description:** User authentication failed

<div align="left"><figure><img src="/files/dOAHkhXU6zYrGSwP6DqO" alt="" width="375"><figcaption></figcaption></figure></div>

**Mitigation**:

1. Validate the configured OIDC scope exists in the Identity Provider
2. Validate the configured Client Secret match the Client Secret in the Identity Provider

**Advanced:** Look for the specific error message in the URL address

</details>

<details>

<summary><strong>Error:</strong> "Client not found"</summary>

**Description:** OIDC Client ID was not found in the Identity Provider

**Mitigation**: Validate the configured Client ID matches the Identity Provider Client ID

</details>


# Set Up SSO with OpenShift

Single Sign-On (SSO) is an authentication scheme, allowing users to log-in with a single pair of credentials to multiple, independent software systems.

This guide explains the procedure to [configure SSO to NVIDIA Run:ai](/saas/infrastructure-setup/authentication/overview#single-sign-on-sso) to NVIDIA Run:ai using the OpenID Connect protocol in OpenShift V4.

## Prerequisites

Before starting, make sure you have the following available from your OpenShift cluster:

* [OpenShift OAuth client](https://docs.openshift.com/container-platform/4.16/authentication/configuring-oauth-clients.html#oauth-register-additional-client_configuring-oauth-clients):
  * ClientID - The ID used to identify the client with the Authorization Server.
  * Client Secret - A secret password that only the Client and Authorization Server know.
* Base URL - The OpenShift API Server endpoint (for example, [https://api.\<cluster-url>:6443](https://api.noa-ocp.runailabs.com:6443/))

## Setup

### Adding the Identity Provider

1. Go to **General settings**
2. Open the Security section and click **+IDENTITY PROVIDER**
3. Select **OpenShift V4**
4. Enter the **Base URL**, Client ID, and **Client Secret** from your OpenShift OAuth client.
5. Copy the Redirect URL to be used in your OpenShift OAuth client
6. Optional: Enter the user attributes and their value in the identity provider as shown in the below table
7. Click **SAVE**
8. Optional: Enable **Auto-Redirect to SSO** to automatically redirect users to your configured identity provider’s login page when accessing the platform.

| Attribute            | Default value in NVIDIA Run:ai | Description                                                                                                                                                                                                  |
| -------------------- | ------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| User role groups     | GROUPS                         | If it exists in the IDP, it allows you to assign NVIDIA Run:ai role groups via the IDP. The IDP attribute must be a list of strings.                                                                         |
| Linux User ID        | UID                            | If it exists in the IDP, it allows researcher containers to start with the Linux User UID. Used to map access to network resources such as file systems to users. The IDP attribute must be of type integer. |
| Linux Group ID       | GID                            | If it exists in the IDP, it allows researcher containers to start with the Linux Group GID. The IDP attribute must be of type integer.                                                                       |
| Supplementary Groups | SUPPLEMENTARYGROUPS            | If it exists in the IDP, it allows researcher containers to start with the relevant Linux supplementary groups. The IDP attribute must be a list of integers.                                                |
| Email                | email                          | Defines the user attribute in the IDP holding the user's email address, which is the user identifier in NVIDIA Run:ai                                                                                        |
| User first name      | firstName                      | Used as the user’s first name appearing in the NVIDIA Run:ai platform                                                                                                                                        |
| User last name       | lastName                       | Used as the user’s last name appearing in the NVIDIA Run:ai platform                                                                                                                                         |

### Testing the Setup

1. Open the NVIDIA Run:ai platform as an admin
2. Add [access rules](/saas/infrastructure-setup/authentication/accessrules) to an SSO user defined in the IDP
3. Open the NVIDIA Run:ai platform in an incognito browser tab
4. On the sign-in page click **CONTINUE WITH SSO**\
   You are redirected to the OpenShift IDP sign-in page
5. In the identity provider sign-in page, log in with the SSO user who you granted with access rules
6. If you are unsuccessful signing-in to the identity provider, follow the [Troubleshooting](#troubleshooting) section below

### Editing the Identity Provider

You can view the identity provider details and edit its configuration:

1. Go to **General settings**
2. Open the Security section
3. On the identity provider box, click **Edit identity provider**
4. You can edit either the **Base URL**, **Client ID**, **Client Secret**, or the **User attributes**

### Removing the Identity Provider

You can remove the identity provider configuration:

1. Go to **General settings**
2. Open the Security section
3. On the identity provider card, click **Remove identity provider**
4. In the dialog, click **REMOVE** to confirm

{% hint style="info" %}
**Note**

To avoid losing access, removing the identity provider must be carried out by a local user.
{% endhint %}

## Troubleshooting

If testing the setup was unsuccessful, try the different troubleshooting scenarios according to the error you received.

### Troubleshooting Scenarios

<details>

<summary><strong>Error:</strong> "403 - Sorry, we can’t let you see this page. Something about permissions…"</summary>

**Description:** The authenticated user is missing permissions

**Mitigation**:

1. Validate either the user or its related group/s are assigned with [access rules](/saas/infrastructure-setup/authentication/accessrules)
2. Validate groups attribute is available in the configured OIDC Scopes
3. Validate the user’s groups attribute is mapped correctly

**Advanced:**

1. Open the Chrome DevTools: Right-click on page → Inspect → Console tab
2. Run the following command to retrieve and copy the user’s token: `localStorage.token;`
3. Paste in [https://jwt.io](https://jwt.io/)
4. Under the Payload section validate the value of the user’s attributes

</details>

<details>

<summary><strong>Error:</strong> "401 - We’re having trouble identifying your account because your email is incorrect or can’t be found."</summary>

**Description:** Authentication failed because email attribute was not found.

**Mitigation**:

1. Validate email attribute is available in the configured OIDC Scopes
2. Validate the user’s email attribute is mapped correctly

</details>

<details>

<summary><strong>Error:</strong> "Unexpected error when authenticating with identity provider"</summary>

**Description:** User authentication failed

<div align="left"><figure><img src="/files/DI2xJ1NGogtzgJyadX49" alt="" width="375"><figcaption></figcaption></figure></div>

**Mitigation**: Validate the the configured OIDC Scopes exist and match the Identity Provider’s available scopes

**Advanced:** Look for the specific error message in the URL address

</details>

<details>

<summary><strong>Error:</strong> "Unexpected error when authenticating with identity provider (SSO sign-in is not available)"</summary>

**Description:** User authentication failed

<div align="left"><figure><img src="/files/dOAHkhXU6zYrGSwP6DqO" alt="" width="375"><figcaption></figcaption></figure></div>

**Mitigation**:

1. Validate the the configured OIDC scope exists in the Identity Provider
2. Validate the configured Client Secret match the Client Secret value in the OAuthclient Kubernetes object.

**Advanced:** Look for the specific error message in the URL address

</details>

<details>

<summary><strong>Error:</strong> "unauthorized_client"</summary>

**Description:** OIDC Client ID was not found in the OpenShift IDP

<img src="/files/7NQEbFytDJz4RSmktJM9" alt="" data-size="original">

**Mitigation**: Validate the the configured Client ID matches the value in the OAuthclient Kubernetes object

</details>


# Permissions

This page provides a reference of the permissions available in the NVIDIA Run:ai platform. Use it when building [custom roles](/saas/infrastructure-setup/authentication/roles) to understand what each permission controls and how permission names in the API map to the names shown in the user interface.

## About Permissions

A permission is the actions (View, Edit, Create and Delete) that the user can perform over a NVIDIA Run:ai entity (resource type). These actions map to the API actions `read`, `update`, `create`, and `delete` respectively. Related permissions are grouped into [permission sets](#permission-sets), and a [role](/saas/infrastructure-setup/authentication/roles) is defined as a collection of permission sets assigned to a subject in a scope. NVIDIA Run:ai provides predefined roles that are system-defined and cannot be edited or deleted. Administrators can also create custom roles via the API.

When you build a custom role through the API, permissions are referenced by their **API name** (for example, `compute-resources`). The NVIDIA Run:ai user interface presents the same permission under a **display name** (for example, "Compute resources"). The tables below map the two and describe what each permission controls.

{% hint style="info" %}
**Note**

* Some permissions are system-internal and are not shown in the Roles view in the user interface. These are noted in the tables below.
* **Deprecated:** The `apps` API name is deprecated. Use `service-account` for managing service accounts.
  {% endhint %}

## Permission Reference

The following table lists all permissions, mapping each permission's API name to its user interface display name and describing what it controls.

| API name                       | Display name                                                                                                         | Description                                                                                                                                                                                                                                                                                                                                                                 |
| ------------------------------ | -------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `access-keys`                  | [Access Keys](/saas/settings/user-settings/user-access-keys)                                                         | Access keys are used for API integrations with NVIDIA Run:ai. An access key contains a client ID and a client secret. This permission is not shown in the Roles view.                                                                                                                                                                                                       |
| `access_rules`                 | [Access rules](/saas/infrastructure-setup/authentication/accessrules)                                                | Access rules provide users, groups, or service accounts privileges to system entities. An access rule is the assignment of a role to a subject in a scope.                                                                                                                                                                                                                  |
| `tenant`                       | Account                                                                                                              | The top-level organizational node headed by the account, comprised of clusters, departments, and projects.                                                                                                                                                                                                                                                                  |
| `applications`                 | [AI Applications](/saas/ai-applications/ai-applications)                                                             | An AI application represents a high-level logical grouping of all Kubernetes resources that together deliver a functional AI solution.                                                                                                                                                                                                                                      |
| `apps`                         | [Applications](/saas/infrastructure-setup/authentication/service-accounts)                                           | Deprecated. Superseded by `service-account`. Distinct from `applications` (AI Applications).                                                                                                                                                                                                                                                                                |
| `branding-settings`            | [Branding settings](/saas/settings/general-settings#branding)                                                        | Custom logo displayed in the top-right corner of the NVIDIA Run:ai platform interface.                                                                                                                                                                                                                                                                                      |
| `cluster-config`               | [Cluster configuration](/saas/infrastructure-setup/advanced-setup/cluster-config)                                    | Governs synchronization of cluster configuration between a cluster and the control plane. This permission is system-internal: it is not shown in the Roles view and is not held by tenant administrators. Do not include `cluster-config` in a custom role. A role that includes it cannot be assigned, and access rule creation fails with `403 Insufficient permissions`. |
| `cluster`                      | [Clusters](/saas/infrastructure-setup/procedures/clusters)                                                           | A registered Kubernetes cluster that the NVIDIA Run:ai control plane manages, schedules workloads on, and monitors.                                                                                                                                                                                                                                                         |
| `clusters-minimal`             | [Clusters minimal](/saas/infrastructure-setup/procedures/clusters)                                                   | Provides read-only access to a lightweight subset of cluster metadata (name, ID, status). Automatically included by the platform when a role grants cluster read access, to support resource display in the user interface.                                                                                                                                                 |
| `compute-resources`            | [Compute resources](/saas/workloads-in-nvidia-run-ai/assets/compute-resources)                                       | A compute resource is a workload asset template that simplifies how workloads are submitted and can be used by AI practitioners when they submit their workloads.                                                                                                                                                                                                           |
| `credentials`                  | [Credentials](/saas/workloads-in-nvidia-run-ai/assets/credentials)                                                   | Credentials are workload assets that mask sensitive access information, such as passwords, tokens, and access keys, necessary for accessing protected resources.                                                                                                                                                                                                            |
| `pvc-assets`                   | [Data sources](/saas/workloads-in-nvidia-run-ai/assets/datasources#pvc)                                              | A data source asset backed by a Persistent Volume Claim (PVC), a Kubernetes concept for managing cluster storage provisioned by an administrator or dynamically using a StorageClass.                                                                                                                                                                                       |
| `git-assets`                   | [Data sources](/saas/workloads-in-nvidia-run-ai/assets/datasources)                                                  | A data source asset that copies code from a Git branch into a dedicated folder in the container, providing the workload with the latest code repository.                                                                                                                                                                                                                    |
| `nfs-assets`                   | [Data sources](/saas/workloads-in-nvidia-run-ai/assets/datasources)                                                  | A data source asset backed by a Network File System (NFS), used for sharing storage among pods in the cluster; content is preserved outside the lifecycle of a single pod.                                                                                                                                                                                                  |
| `s3-assets`                    | [Data sources](/saas/workloads-in-nvidia-run-ai/assets/datasources)                                                  | A data source asset that maps a remote S3 bucket into the workload's file system, accessible across different workload executions.                                                                                                                                                                                                                                          |
| `host-path-assets`             | [Data sources](/saas/workloads-in-nvidia-run-ai/assets/datasources)                                                  | A data source asset that mounts a host path file or directory on the workload's file system; data persists across workloads.                                                                                                                                                                                                                                                |
| `cm-volume-assets`             | [Data sources](/saas/workloads-in-nvidia-run-ai/assets/datasources)                                                  | A data source asset backed by a Kubernetes ConfigMap, used to store non-confidential data such as environment variables and command-line arguments.                                                                                                                                                                                                                         |
| `secret-volume-assets`         | [Data sources](/saas/workloads-in-nvidia-run-ai/assets/datasources)                                                  | A data source asset that maps a credential into the workload's file system to provide sensitive access information such as passwords, tokens, and access keys.                                                                                                                                                                                                              |
| `datavolumes`                  | [Data volumes](/saas/workloads-in-nvidia-run-ai/assets/data-volumes)                                                 | Data volumes are workload assets that offer a solution for storing, managing, and sharing AI training data across scopes and workloads.                                                                                                                                                                                                                                     |
| `department`                   | [Departments](/saas/platform-management/aiinitiatives/organization/departments)                                      | Departments group multiple projects under a shared organizational scope for quota management, policy enforcement, and asset sharing.                                                                                                                                                                                                                                        |
| `environments`                 | [Environments](/saas/workloads-in-nvidia-run-ai/assets/environments)                                                 | An environment is a workload asset consisting of a configuration that simplifies how workloads are submitted and can be used by AI practitioners.                                                                                                                                                                                                                           |
| `events-history`               | [Event history](/saas/infrastructure-setup/procedures/event-history)                                                 | An audit log of all changes to business objects (clusters, projects, assets, and more) in the NVIDIA Run:ai platform, accessible to system administrators with tenant-wide permissions.                                                                                                                                                                                     |
| `inferences`                   | [Inferences](/saas/workloads-in-nvidia-run-ai/using-inference)                                                       | Inference workloads deploy trained models to make real-time or batch predictions in a production environment.                                                                                                                                                                                                                                                               |
| `network-topologies`           | [Network topologies](/saas/infrastructure-setup/procedures/clusters)                                                 | Named configurations of node labels that describe the physical network structure of a cluster (racks, blocks, NVLink domains), used by the scheduler to minimize latency for distributed workloads.                                                                                                                                                                         |
| `nodepools`                    | [Node pools](/saas/platform-management/aiinitiatives/resources/node-pools)                                           | A node pool is a NVIDIA Run:ai construct representing a set of nodes grouped into a bucket of resources using a predefined or administrator-defined node label.                                                                                                                                                                                                             |
| `nodepools-minimal`            | [Node pools minimal](/saas/platform-management/aiinitiatives/resources/node-pools)                                   | Provides read-only access to a lightweight subset of node pool metadata. Automatically included by the platform when a role grants node pool, department, or project read access.                                                                                                                                                                                           |
| `nodes`                        | [Nodes](/saas/platform-management/aiinitiatives/resources/nodes)                                                     | Nodes are Kubernetes elements automatically discovered by the NVIDIA Run:ai platform and registered for scheduling and administrator visibility.                                                                                                                                                                                                                            |
| `notification-settings`        | [Notification settings](/saas/settings/general-settings/notifications)                                               | Configuration for notification options that keep administrators and users informed of important events and updates.                                                                                                                                                                                                                                                         |
| `policies`                     | [Policies](https://github.com/run-ai/runai-product-docs/tree/SaaS/platform-management/policies/workload-policies.md) | Workload policies allow administrators to set best practices, enforce limitations, and standardize workload submission processes within their organization.                                                                                                                                                                                                                 |
| `project`                      | [Projects](/saas/platform-management/aiinitiatives/organization/projects)                                            | Projects are the primary organization management unit in NVIDIA Run:ai, used to control how resources are allocated and prioritized across teams and initiatives.                                                                                                                                                                                                           |
| `registries`                   | [Registries](/saas/workloads-in-nvidia-run-ai/assets/environments#adding-a-new-environment)                          | Stored connection configurations for private container image registries, used when selecting images in environment assets.                                                                                                                                                                                                                                                  |
| `roles`                        | [Roles](/saas/infrastructure-setup/authentication/roles)                                                             | A role defines a set of permissions that can be assigned to a subject in a scope, determining which actions the subject is allowed to perform on NVIDIA Run:ai entities.                                                                                                                                                                                                    |
| `security-settings`            | [Security settings](/saas/settings/general-settings#security)                                                        | Platform-level identity provider (SSO) configuration and session security settings.                                                                                                                                                                                                                                                                                         |
| `service-account`              | [Service accounts](/saas/infrastructure-setup/authentication/service-accounts)                                       | Service accounts are used for API integrations with NVIDIA Run:ai. A service account contains a client ID and a client secret.                                                                                                                                                                                                                                              |
| `settings`                     | [Settings](/saas/settings/general-settings)                                                                          | Centralized controls for enabling or disabling key NVIDIA Run:ai platform features across resources, workload types, and security.                                                                                                                                                                                                                                          |
| `storage-class-configuration`  | [Storage class configuration](https://run-ai-docs.nvidia.com/api/2.26/workload-assets/storage-class-configuration)   | Kubernetes storage classes define how storage is provisioned for AI workloads, used when creating PVC data sources and data volumes.                                                                                                                                                                                                                                        |
| `templates`                    | [Templates](/saas/workloads-in-nvidia-run-ai/workload-templates)                                                     | A workload template is a reusable setup that defines all the required configuration fields for submitting a workload, allowing users to predefine settings.                                                                                                                                                                                                                 |
| `trainings`                    | [Trainings](/saas/workloads-in-nvidia-run-ai/using-training)                                                         | Training workloads run resource-intensive, batch-mode model development jobs, often requiring distributed computing across multiple nodes.                                                                                                                                                                                                                                  |
| `users`                        | [Users](/saas/infrastructure-setup/authentication/users)                                                             | Users can be managed locally or via an identity provider (IdP) and are assigned permissions through access rules.                                                                                                                                                                                                                                                           |
| `users-minimal`                | [Users minimal](/saas/infrastructure-setup/authentication/users)                                                     | Provides read-only access to a minimal subset of user data, such as user IDs and display names. Used by system roles to resolve user identities without exposing full user management permissions.                                                                                                                                                                          |
| `workload-integration-metrics` | Workload integration metrics                                                                                         | System-internal. Used by the platform to collect workload metrics from integrations. Not intended to be assigned when building custom roles.                                                                                                                                                                                                                                |
| `workload-properties`          | [Workload properties](/saas/workloads-in-nvidia-run-ai/workloads#managing-workload-properties)                       | The priority, preemptibility, and workload category assigned to each workload type by default.                                                                                                                                                                                                                                                                              |
| `workloads`                    | [Workloads](/saas/workloads-in-nvidia-run-ai/workloads)                                                              | Workloads are the fundamental building blocks for consuming resources, enabling AI practitioners to support the entire lifecycle of an AI initiative.                                                                                                                                                                                                                       |
| `workspaces`                   | [Workspaces](/saas/workloads-in-nvidia-run-ai/using-workspaces/running-workspace)                                    | The workspace is where data scientists conduct initial research, experiment with data sets, and test algorithms — the most flexible stage in the ML lifecycle.                                                                                                                                                                                                              |

{% hint style="info" %}
**Note**

The seven `*-assets` API names above all appear under the single **Data sources** display name in the user interface. When building a custom role via the API, grant the specific asset types your users need using the relevant permission sets.
{% endhint %}

## Permission Sets

Permissions are bundled into **permission sets** - reusable building blocks that group the actions for a particular functional area (for example, viewing workloads or managing clusters). Permission sets are used through the [Roles](/saas/infrastructure-setup/authentication/roles) API to define both predefined and custom roles, and are referenced by the API names below. Most permission sets come in a **read** variant (view only) and an **edit** variant (create, read, update, and delete). The following permission sets are available:

| API name                              | Description                                                                                                                                                        |
| ------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `workspaceEditAccess`                 | Grants full permissions to create, read, update, and delete workspaces, and to access all related data and resources required for their management.                |
| `trainingEditAccess`                  | Grants full permissions to create, read, update, and delete training workloads, and to access all related data and resources required for their management.        |
| `inferenceEditAccess`                 | Grants full permissions to create, read, update, and delete inference workloads, and to access all related data and resources required for their management.       |
| `yamlWorkloadEditAccess`              | Grants full permissions to create, read, update, and delete workloads defined in YAML, and to access all related data and resources required for their management. |
| `workloadReadAccess`                  | Grants permissions to view workloads, including all related data and resources associated with them.                                                               |
| `workloadPropertiesEditAccess`        | Grants full permissions to create, read, update, and delete workload properties, and to access all related data and resources required for their management.       |
| `workloadPropertiesReadAccess`        | Grants permissions to view workload properties, including all related data and resources associated with them.                                                     |
| `templatesEditAccess`                 | Grants full permissions to create, read, update, and delete workload templates, and to access all related data and resources required for their management.        |
| `templatesReadAccess`                 | Grants permissions to view workload templates, including all related data and resources associated with them.                                                      |
| `dataAndStorageEditAccess`            | Grants full permissions to create, read, update, and delete data and storage, and to access all related data and resources required for their management.          |
| `dataAndStorageReadAccess`            | Grants permissions to view data and storage, including all related data and resources associated with them.                                                        |
| `policiesEditAccess`                  | Grants full permissions to create, read, update, and delete policies, and to access all related data and resources required for their management.                  |
| `policiesReadAccess`                  | Grants permissions to view policies, including all related data and resources associated with them.                                                                |
| `projectsEditAccess`                  | Grants full permissions to create, read, update, and delete projects, and to access all related data and resources required for their management.                  |
| `projectsReadAccess`                  | Grants permissions to view projects, including all related data and resources associated with them.                                                                |
| `departmentsEditAccess`               | Grants full permissions to create, read, update, and delete departments, and to access all related data and resources required for their management.               |
| `departmentsReadAccess`               | Grants permissions to view departments, including all related data and resources associated with them.                                                             |
| `clustersEditAccess`                  | Grants permissions to create, read, update, and delete clusters.                                                                                                   |
| `clustersReadAccess`                  | Grants permissions to view clusters.                                                                                                                               |
| `networkTopologyEditAccess`           | Grants full permissions to create, read, update, and delete network topologies, and to access all related data and resources required for their management.        |
| `networkTopologyReadAccess`           | Grants permissions to view network topologies, including all related data and resources associated with them.                                                      |
| `nodesReadAccess`                     | Grants permissions to view nodes, including all related data and resources associated with them.                                                                   |
| `nodePoolsEditAccess`                 | Grants full permissions to create, read, update, and delete node pools, and to access all related data and resources required for their management.                |
| `nodePoolsReadAccess`                 | Grants permissions to view node pools, including all related data and resources associated with them.                                                              |
| `quotaManagementDashboardReadAccess`  | Grants permissions to view the quota management dashboard, including all related data and resources associated with them.                                          |
| `usersEditAccess`                     | Grants permissions to create, read, update, and delete users.                                                                                                      |
| `usersReadAccess`                     | Grants permissions to view users.                                                                                                                                  |
| `serviceAccountsEditAccess`           | Grants permissions to create, read, update, and delete service accounts.                                                                                           |
| `serviceAccountsReadAccess`           | Grants permissions to view service accounts.                                                                                                                       |
| `accessRulesEditAccess`               | Grants full permissions to create, read, update, and delete access rules, and to access all related data and resources required for their management.              |
| `accessRulesReadAccess`               | Grants permissions to view access rules, including all related data and resources associated with them.                                                            |
| `rolesEditAccess`                     | Grants permissions to create, read, update, and delete roles.                                                                                                      |
| `rolesReadAccess`                     | Grants permissions to view roles.                                                                                                                                  |
| `settingsEditAccess`                  | Grants permissions to create, read, update, and delete settings.                                                                                                   |
| `settingsReadAccess`                  | Grants permissions to view settings.                                                                                                                               |
| `securitySettingsEditAccess`          | Grants permissions to create, read, update, and delete security settings.                                                                                          |
| `securitySettingsReadAccess`          | Grants permissions to view security settings.                                                                                                                      |
| `brandingSettingsEditAccess`          | Grants permissions to create, read, update, and delete branding settings.                                                                                          |
| `brandingSettingsReadAccess`          | Grants permissions to view branding settings.                                                                                                                      |
| `eventHistoryReadAccess`              | Grants permissions to view event history.                                                                                                                          |
| `accountEditAccess`                   | Grants permissions to create, read, update, and delete accounts.                                                                                                   |
| `accountReadAccess`                   | Grants permissions to view accounts.                                                                                                                               |
| `accessKeysEditAccess`                | Grants full permissions to create, read, update, and delete user access keys.                                                                                      |
| `storageClassConfigurationEditAccess` | Grants full permissions to create, read, update, and delete storage class configurations used by volume and PVC assets.                                            |


# Roles

A role defines a set of permissions that can be assigned to a [subject in a scope](/saas/platform-management/aiinitiatives/adapting-ai-initiatives#scopes-in-an-organization), determining which actions the subject is allowed to perform on NVIDIA Run:ai entities such as projects, workloads, and users. Each permission represents an allowed action including view, create, edit, and delete on a specific entity.

Related permissions are grouped into permission sets, which bundle actions for a particular functional area (for example, managing workloads). These permission sets act as reusable building blocks that simplify role definition and ensure consistent access control across the platform.

NVIDIA Run:ai provides [predefined roles](#roles-in-nvidia-run-ai) for common use cases and allows administrators to create [custom roles](#custom-roles-api-only) to meet specific organizational requirements. For a reference of all available permissions, including how API permission names map to their user interface display names, see [Permissions](/saas/infrastructure-setup/authentication/permissions).

## Roles Table

The Roles table can be found under **Access** in the NVIDIA Run:ai platform.

The Roles table displays a list of roles available to users in the NVIDIA Run:ai platform. Both predefined and custom roles will be displayed in the table.

<figure><img src="/files/mw31vjcSIGrdrZA7ARWk" alt=""><figcaption></figcaption></figure>

The Roles table consists of the following columns:

| Column        | Description                             |
| ------------- | --------------------------------------- |
| Role          | The name of the role                    |
| Created by    | The name of the role creator            |
| Creation time | The timestamp when the role was created |

### Customizing the Table View

* Filter - Click ADD FILTER, select the column to filter by, and enter the filter values
* Search - Click SEARCH and type the value to search by
* Sort - Click each column header to sort by
* Column selection - Click COLUMNS and select the columns to display in the table
* Download table - Click MORE and then Click Download as CSV. Export to CSV is limited to 20,000 rows.

## Reviewing a Role

1. To review a role click the role name on the table
2. In the role form review the following:
   * **Role name**\
     The name of the role
   * **Entity**\
     A system-managed object that can be viewed, edited, created or deleted by a user based on their assigned role and scope
   * **Actions**\
     The actions that the role assignee is authorized to perform for each entity
     * **View** - If checked, an assigned user with this role can view instances of this type of entity within their defined scope
     * **Edit** - If checked, an assigned user with this role can change the settings of an instance of this type of entity within their defined scope
     * **Create** - If checked, an assigned user with this role can create new instances of this type of entity within their defined scope
     * **Delete** - If checked, an assigned user with this role can delete instances of this type of entity within their defined scope

## Roles in NVIDIA Run:ai

NVIDIA Run:ai supports the following roles and their permissions. Under each role is a detailed list of the actions that the role assignee is authorized to perform for each entity.

{% hint style="info" %}
**Note**

Legacy roles are **deprecated** and will be removed in a future release. Review the new predefined roles to determine whether they meet your requirements, or create a [custom role](#custom-role-management-api-only) using the API.

* L1 researcher, L2 researcher, ML engineer → consider using AI practitioner
* Data source administrator, Data volume administrator → consider using Data & storage administrator
* Research manager → consider using Project administrator
  {% endhint %}

<details>

<summary>AI practitioner</summary>

<table><thead><tr><th width="221.51171875">Entity</th><th data-type="checkbox">View</th><th data-type="checkbox">Edit</th><th data-type="checkbox">Create</th><th data-type="checkbox">Delete</th></tr></thead><tbody><tr><td>Account</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hxy5GaNYO0MrDXL8r67p">AI Applications</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss#branding">Branding settings</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/KXk9Ii3EjuXL7rJP8Vhz">Departments</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hohfv4UUHT0dvOZ5dHbi">Event history</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Network topologies</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/EoHF94AYOoPhPYml00YS">Policies</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/6wxLpJPbjlnaHiHErmjg">Projects</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Security settings</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Clusters</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Clusters minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/64hpIzKT1fTtBZw8YSvQ">Node pools</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Node pools minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/FbftKDBouwYrCsumy9ie">Nodes</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/UPP9oksqWgXJbc9k3zdu">Access rules</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Applications</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/2CvyzKWvsYYcFSy6VpU1">Roles</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Service accounts</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/juQ16RWzgmohYWR4AbjH">Users</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Users minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VTxcwsPIfBGOXPU2NAc7">Inferences</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/aqIRhDCT9Ndcn878Dlwr">Trainings</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O#managing-workload-properties">Workload properties</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O">Workloads</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/K8gSRGHjeJLeW7fbyqbz">Workspaces</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/VpIuJWjoVervgD05b6e7">Compute resources</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/GuBN1o0IXQ2RGRmADNDd">Credentials</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS">Data sources</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/Aa4tONFSp6wI8dajbWxr">Data volumes</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4">Environments</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS#pvc">PVC</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4#adding-a-new-environment">Registries</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="https://run-ai-docs.nvidia.com/api/workload-assets/storage-class-configuration">Storage class configurations</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/j08IRnRyHoDV4ChTtT0a">Templates</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr></tbody></table>

</details>

<details>

<summary>Compute resource administrator <code>Legacy</code></summary>

<table><thead><tr><th width="222.015625">Entity</th><th data-type="checkbox">View</th><th data-type="checkbox">Edit</th><th data-type="checkbox">Create</th><th data-type="checkbox">Delete</th></tr></thead><tbody><tr><td>Account</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hxy5GaNYO0MrDXL8r67p">AI Applications</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss#branding">Branding settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/KXk9Ii3EjuXL7rJP8Vhz">Departments</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hohfv4UUHT0dvOZ5dHbi">Event history</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Network topologies</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/EoHF94AYOoPhPYml00YS">Policies</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/6wxLpJPbjlnaHiHErmjg">Projects</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss#security">Security settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Clusters</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Clusters minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/64hpIzKT1fTtBZw8YSvQ">Node pools</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Node pools minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/FbftKDBouwYrCsumy9ie">Nodes</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/UPP9oksqWgXJbc9k3zdu">Access rules</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Applications</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/2CvyzKWvsYYcFSy6VpU1">Roles</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Service accounts</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/juQ16RWzgmohYWR4AbjH">Users</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Users minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VTxcwsPIfBGOXPU2NAc7">Inferences</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/aqIRhDCT9Ndcn878Dlwr">Trainings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O#managing-workload-properties">Workload properties</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O">Workloads</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/K8gSRGHjeJLeW7fbyqbz">Workspaces</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VpIuJWjoVervgD05b6e7">Compute resources</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/GuBN1o0IXQ2RGRmADNDd">Credentials</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS">Data sources</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/Aa4tONFSp6wI8dajbWxr">Data volumes</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4">Environments</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS#pvc">PVC</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4#adding-a-new-environment">Registries</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="https://run-ai-docs.nvidia.com/api/workload-assets/storage-class-configuration">Storage class configurations</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/j08IRnRyHoDV4ChTtT0a">Templates</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr></tbody></table>

</details>

<details>

<summary>Credentials administrator <code>Legacy</code></summary>

<table><thead><tr><th width="221.51171875">Entity</th><th data-type="checkbox">View</th><th data-type="checkbox">Edit</th><th data-type="checkbox">Create</th><th data-type="checkbox">Delete</th></tr></thead><tbody><tr><td>Account</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hxy5GaNYO0MrDXL8r67p">AI Applications</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss#branding">Branding settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/KXk9Ii3EjuXL7rJP8Vhz">Departments</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hohfv4UUHT0dvOZ5dHbi">Event history</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Network topologies</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/EoHF94AYOoPhPYml00YS">Policies</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/6wxLpJPbjlnaHiHErmjg">Projects</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Security settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Clusters</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Clusters minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/64hpIzKT1fTtBZw8YSvQ">Node pools</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Node pools minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/FbftKDBouwYrCsumy9ie">Nodes</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/UPP9oksqWgXJbc9k3zdu">Access rules</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Applications</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/2CvyzKWvsYYcFSy6VpU1">Roles</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Service accounts</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/juQ16RWzgmohYWR4AbjH">Users</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Users minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VTxcwsPIfBGOXPU2NAc7">Inferences</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/aqIRhDCT9Ndcn878Dlwr">Trainings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O#managing-workload-properties">Workload properties</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O">Workloads</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/K8gSRGHjeJLeW7fbyqbz">Workspaces</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VpIuJWjoVervgD05b6e7">Compute resources</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/GuBN1o0IXQ2RGRmADNDd">Credentials</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS">Data sources</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/Aa4tONFSp6wI8dajbWxr">Data volumes</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4">Environments</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS#pvc">PVC</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4#adding-a-new-environment">Registries</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="https://run-ai-docs.nvidia.com/api/workload-assets/storage-class-configuration">Storage class configurations</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/j08IRnRyHoDV4ChTtT0a">Templates</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr></tbody></table>

</details>

<details>

<summary>Data and storage administrator</summary>

<table><thead><tr><th width="221.51171875">Entity</th><th data-type="checkbox">View</th><th data-type="checkbox">Edit</th><th data-type="checkbox">Create</th><th data-type="checkbox">Delete</th></tr></thead><tbody><tr><td>Account</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hxy5GaNYO0MrDXL8r67p">AI Applications</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss#branding">Branding settings</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/KXk9Ii3EjuXL7rJP8Vhz">Departments</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hohfv4UUHT0dvOZ5dHbi">Event history</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Network topologies</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/EoHF94AYOoPhPYml00YS">Policies</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/6wxLpJPbjlnaHiHErmjg">Projects</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Security settings</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Clusters</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Clusters minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/64hpIzKT1fTtBZw8YSvQ">Node pools</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Node pools minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/FbftKDBouwYrCsumy9ie">Nodes</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/UPP9oksqWgXJbc9k3zdu">Access rules</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Applications</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/2CvyzKWvsYYcFSy6VpU1">Roles</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Service accounts</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/juQ16RWzgmohYWR4AbjH">Users</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Users minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VTxcwsPIfBGOXPU2NAc7">Inferences</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/aqIRhDCT9Ndcn878Dlwr">Trainings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O#managing-workload-properties">Workload properties</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O">Workloads</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/K8gSRGHjeJLeW7fbyqbz">Workspaces</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VpIuJWjoVervgD05b6e7">Compute resources</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/GuBN1o0IXQ2RGRmADNDd">Credentials</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS">Data sources</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/Aa4tONFSp6wI8dajbWxr">Data volumes</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4">Environments</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS#pvc">PVC</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4#adding-a-new-environment">Registries</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="https://run-ai-docs.nvidia.com/api/workload-assets/storage-class-configuration">Storage class configurations</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/j08IRnRyHoDV4ChTtT0a">Templates</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr></tbody></table>

</details>

<details>

<summary>Data source administrator <code>Legacy</code></summary>

<table><thead><tr><th width="221.51171875">Entity</th><th data-type="checkbox">View</th><th data-type="checkbox">Edit</th><th data-type="checkbox">Create</th><th data-type="checkbox">Delete</th></tr></thead><tbody><tr><td>Account</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hxy5GaNYO0MrDXL8r67p">AI Applications</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss#branding">Branding settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/KXk9Ii3EjuXL7rJP8Vhz">Departments</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hohfv4UUHT0dvOZ5dHbi">Event history</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Network topologies</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/EoHF94AYOoPhPYml00YS">Policies</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/6wxLpJPbjlnaHiHErmjg">Projects</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Security settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Clusters</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Clusters minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/64hpIzKT1fTtBZw8YSvQ">Node pools</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Node pools minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/FbftKDBouwYrCsumy9ie">Nodes</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/UPP9oksqWgXJbc9k3zdu">Access rules</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Applications</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/2CvyzKWvsYYcFSy6VpU1">Roles</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Service accounts</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/juQ16RWzgmohYWR4AbjH">Users</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Users minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VTxcwsPIfBGOXPU2NAc7">Inferences</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/aqIRhDCT9Ndcn878Dlwr">Trainings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O#managing-workload-properties">Workload properties</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O">Workloads</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/K8gSRGHjeJLeW7fbyqbz">Workspaces</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VpIuJWjoVervgD05b6e7">Compute resources</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/GuBN1o0IXQ2RGRmADNDd">Credentials</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS">Data sources</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/Aa4tONFSp6wI8dajbWxr">Data volumes</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4">Environments</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS#pvc">PVC</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4#adding-a-new-environment">Registries</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="https://run-ai-docs.nvidia.com/api/workload-assets/storage-class-configuration">Storage class configurations</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/j08IRnRyHoDV4ChTtT0a">Templates</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr></tbody></table>

</details>

<details>

<summary>Data volume administrator <code>Legacy</code></summary>

<table><thead><tr><th width="221.51171875">Entity</th><th data-type="checkbox">View</th><th data-type="checkbox">Edit</th><th data-type="checkbox">Create</th><th data-type="checkbox">Delete</th></tr></thead><tbody><tr><td>Account</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hxy5GaNYO0MrDXL8r67p">AI Applications</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss#branding">Branding settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/KXk9Ii3EjuXL7rJP8Vhz">Departments</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hohfv4UUHT0dvOZ5dHbi">Event history</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Network topologies</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/EoHF94AYOoPhPYml00YS">Policies</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/6wxLpJPbjlnaHiHErmjg">Projects</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Security settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Clusters</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Clusters minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/64hpIzKT1fTtBZw8YSvQ">Node pools</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Node pools minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/FbftKDBouwYrCsumy9ie">Nodes</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/UPP9oksqWgXJbc9k3zdu">Access rules</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Applications</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/2CvyzKWvsYYcFSy6VpU1">Roles</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Service accounts</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/juQ16RWzgmohYWR4AbjH">Users</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Users minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VTxcwsPIfBGOXPU2NAc7">Inferences</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/aqIRhDCT9Ndcn878Dlwr">Trainings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O#managing-workload-properties">Workload properties</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O">Workloads</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/K8gSRGHjeJLeW7fbyqbz">Workspaces</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VpIuJWjoVervgD05b6e7">Compute resources</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/GuBN1o0IXQ2RGRmADNDd">Credentials</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS">Data sources</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/Aa4tONFSp6wI8dajbWxr">Data volumes</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4">Environments</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS#pvc">PVC</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4#adding-a-new-environment">Registries</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="https://run-ai-docs.nvidia.com/api/workload-assets/storage-class-configuration">Storage class configurations</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/j08IRnRyHoDV4ChTtT0a">Templates</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr></tbody></table>

</details>

<details>

<summary>Department administrator</summary>

<table><thead><tr><th width="221.51171875">Entity</th><th data-type="checkbox">View</th><th data-type="checkbox">Edit</th><th data-type="checkbox">Create</th><th data-type="checkbox">Delete</th></tr></thead><tbody><tr><td>Account</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hxy5GaNYO0MrDXL8r67p">AI Applications</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss#branding">Branding settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/KXk9Ii3EjuXL7rJP8Vhz">Departments</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hohfv4UUHT0dvOZ5dHbi">Event history</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Network topologies</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/EoHF94AYOoPhPYml00YS">Policies</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/6wxLpJPbjlnaHiHErmjg">Projects</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Security settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Clusters</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Clusters minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/64hpIzKT1fTtBZw8YSvQ">Node pools</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Node pools minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/FbftKDBouwYrCsumy9ie">Nodes</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/UPP9oksqWgXJbc9k3zdu">Access rules</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Applications</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/2CvyzKWvsYYcFSy6VpU1">Roles</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Service accounts</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/juQ16RWzgmohYWR4AbjH">Users</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td>Users minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VTxcwsPIfBGOXPU2NAc7">Inferences</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/aqIRhDCT9Ndcn878Dlwr">Trainings</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O#managing-workload-properties">Workload properties</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O">Workloads</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/K8gSRGHjeJLeW7fbyqbz">Workspaces</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/VpIuJWjoVervgD05b6e7">Compute resources</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/GuBN1o0IXQ2RGRmADNDd">Credentials</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS">Data sources</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/Aa4tONFSp6wI8dajbWxr">Data volumes</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4">Environments</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS#pvc">PVC</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4#adding-a-new-environment">Registries</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="https://run-ai-docs.nvidia.com/api/workload-assets/storage-class-configuration">Storage class configurations</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/j08IRnRyHoDV4ChTtT0a">Templates</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr></tbody></table>

</details>

<details>

<summary>Department administrator (Legacy) <code>Legacy</code></summary>

<table><thead><tr><th width="221.51171875">Entity</th><th data-type="checkbox">View</th><th data-type="checkbox">Edit</th><th data-type="checkbox">Create</th><th data-type="checkbox">Delete</th></tr></thead><tbody><tr><td>Account</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hxy5GaNYO0MrDXL8r67p">AI Applications</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss#branding">Branding settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/KXk9Ii3EjuXL7rJP8Vhz">Departments</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hohfv4UUHT0dvOZ5dHbi">Event history</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Network topologies</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/EoHF94AYOoPhPYml00YS">Policies</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/6wxLpJPbjlnaHiHErmjg">Projects</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Security settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Clusters</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Clusters minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/64hpIzKT1fTtBZw8YSvQ">Node pools</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Node pools minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/FbftKDBouwYrCsumy9ie">Nodes</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/UPP9oksqWgXJbc9k3zdu">Access rules</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Applications</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/2CvyzKWvsYYcFSy6VpU1">Roles</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Service accounts</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/juQ16RWzgmohYWR4AbjH">Users</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td>Users minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VTxcwsPIfBGOXPU2NAc7">Inferences</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/aqIRhDCT9Ndcn878Dlwr">Trainings</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O#managing-workload-properties">Workload properties</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O">Workloads</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/K8gSRGHjeJLeW7fbyqbz">Workspaces</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/VpIuJWjoVervgD05b6e7">Compute resources</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/GuBN1o0IXQ2RGRmADNDd">Credentials</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS">Data sources</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/Aa4tONFSp6wI8dajbWxr">Data volumes</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4">Environments</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS#pvc">PVC</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4#adding-a-new-environment">Registries</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="https://run-ai-docs.nvidia.com/api/workload-assets/storage-class-configuration">Storage class configurations</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/j08IRnRyHoDV4ChTtT0a">Templates</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr></tbody></table>

</details>

<details>

<summary>Department viewer</summary>

<table><thead><tr><th width="221.51171875">Entity</th><th data-type="checkbox">View</th><th data-type="checkbox">Edit</th><th data-type="checkbox">Create</th><th data-type="checkbox">Delete</th></tr></thead><tbody><tr><td>Account</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hxy5GaNYO0MrDXL8r67p">AI Applications</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss#branding">Branding settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/KXk9Ii3EjuXL7rJP8Vhz">Departments</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hohfv4UUHT0dvOZ5dHbi">Event history</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Network topologies</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/EoHF94AYOoPhPYml00YS">Policies</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/6wxLpJPbjlnaHiHErmjg">Projects</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Security settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Clusters</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Clusters minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/64hpIzKT1fTtBZw8YSvQ">Node pools</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Node pools minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/FbftKDBouwYrCsumy9ie">Nodes</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/UPP9oksqWgXJbc9k3zdu">Access rules</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Applications</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/2CvyzKWvsYYcFSy6VpU1">Roles</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Service accounts</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/juQ16RWzgmohYWR4AbjH">Users</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Users minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VTxcwsPIfBGOXPU2NAc7">Inferences</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/aqIRhDCT9Ndcn878Dlwr">Trainings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O#managing-workload-properties">Workload properties</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O">Workloads</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/K8gSRGHjeJLeW7fbyqbz">Workspaces</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VpIuJWjoVervgD05b6e7">Compute resources</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/GuBN1o0IXQ2RGRmADNDd">Credentials</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS">Data sources</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/Aa4tONFSp6wI8dajbWxr">Data volumes</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4">Environments</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS#pvc">PVC</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4#adding-a-new-environment">Registries</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="https://run-ai-docs.nvidia.com/api/workload-assets/storage-class-configuration">Storage class configurations</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/j08IRnRyHoDV4ChTtT0a">Templates</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr></tbody></table>

</details>

<details>

<summary>Editor</summary>

<table><thead><tr><th width="221.51171875">Entity</th><th data-type="checkbox">View</th><th data-type="checkbox">Edit</th><th data-type="checkbox">Create</th><th data-type="checkbox">Delete</th></tr></thead><tbody><tr><td>Account</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hxy5GaNYO0MrDXL8r67p">AI Applications</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss#branding">Branding settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/KXk9Ii3EjuXL7rJP8Vhz">Departments</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/hohfv4UUHT0dvOZ5dHbi">Event history</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Network topologies</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/EoHF94AYOoPhPYml00YS">Policies</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/6wxLpJPbjlnaHiHErmjg">Projects</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Security settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Clusters</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Clusters minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/64hpIzKT1fTtBZw8YSvQ">Node pools</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Node pools minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/FbftKDBouwYrCsumy9ie">Nodes</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/UPP9oksqWgXJbc9k3zdu">Access rules</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Applications</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/2CvyzKWvsYYcFSy6VpU1">Roles</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Service accounts</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/juQ16RWzgmohYWR4AbjH">Users</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Users minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VTxcwsPIfBGOXPU2NAc7">Inferences</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/aqIRhDCT9Ndcn878Dlwr">Trainings</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O#managing-workload-properties">Workload properties</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O">Workloads</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/K8gSRGHjeJLeW7fbyqbz">Workspaces</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/VpIuJWjoVervgD05b6e7">Compute resources</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/GuBN1o0IXQ2RGRmADNDd">Credentials</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS">Data sources</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/Aa4tONFSp6wI8dajbWxr">Data volumes</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4">Environments</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS#pvc">PVC</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4#adding-a-new-environment">Registries</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="https://run-ai-docs.nvidia.com/api/workload-assets/storage-class-configuration">Storage class configurations</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/j08IRnRyHoDV4ChTtT0a">Templates</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr></tbody></table>

</details>

<details>

<summary>Environment administrator <code>Legacy</code></summary>

<table><thead><tr><th width="221.51171875">Entity</th><th data-type="checkbox">View</th><th data-type="checkbox">Edit</th><th data-type="checkbox">Create</th><th data-type="checkbox">Delete</th></tr></thead><tbody><tr><td>Account</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hxy5GaNYO0MrDXL8r67p">AI Applications</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss#branding">Branding settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/KXk9Ii3EjuXL7rJP8Vhz">Departments</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hohfv4UUHT0dvOZ5dHbi">Event history</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Network topologies</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/EoHF94AYOoPhPYml00YS">Policies</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/6wxLpJPbjlnaHiHErmjg">Projects</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Security settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Clusters</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Clusters minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/64hpIzKT1fTtBZw8YSvQ">Node pools</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Node pools minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/FbftKDBouwYrCsumy9ie">Nodes</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/UPP9oksqWgXJbc9k3zdu">Access rules</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Applications</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/2CvyzKWvsYYcFSy6VpU1">Roles</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Service accounts</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/juQ16RWzgmohYWR4AbjH">Users</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Users minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VTxcwsPIfBGOXPU2NAc7">Inferences</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/aqIRhDCT9Ndcn878Dlwr">Trainings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O#managing-workload-properties">Workload properties</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O">Workloads</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/K8gSRGHjeJLeW7fbyqbz">Workspaces</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VpIuJWjoVervgD05b6e7">Compute resources</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/GuBN1o0IXQ2RGRmADNDd">Credentials</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS">Data sources</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/Aa4tONFSp6wI8dajbWxr">Data volumes</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4">Environments</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS#pvc">PVC</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4#adding-a-new-environment">Registries</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="https://run-ai-docs.nvidia.com/api/workload-assets/storage-class-configuration">Storage class configurations</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/j08IRnRyHoDV4ChTtT0a">Templates</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr></tbody></table>

</details>

<details>

<summary>L1 researcher <code>Legacy</code></summary>

<table><thead><tr><th width="221.51171875">Entity</th><th data-type="checkbox">View</th><th data-type="checkbox">Edit</th><th data-type="checkbox">Create</th><th data-type="checkbox">Delete</th></tr></thead><tbody><tr><td>Account</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hxy5GaNYO0MrDXL8r67p">AI Applications</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss#branding">Branding settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/KXk9Ii3EjuXL7rJP8Vhz">Departments</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hohfv4UUHT0dvOZ5dHbi">Event history</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Network topologies</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/EoHF94AYOoPhPYml00YS">Policies</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/6wxLpJPbjlnaHiHErmjg">Projects</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Security settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Clusters</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Clusters minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/64hpIzKT1fTtBZw8YSvQ">Node pools</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Node pools minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/FbftKDBouwYrCsumy9ie">Nodes</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/UPP9oksqWgXJbc9k3zdu">Access rules</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Applications</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/2CvyzKWvsYYcFSy6VpU1">Roles</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Service accounts</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/juQ16RWzgmohYWR4AbjH">Users</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Users minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VTxcwsPIfBGOXPU2NAc7">Inferences</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/aqIRhDCT9Ndcn878Dlwr">Trainings</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O#managing-workload-properties">Workload properties</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O">Workloads</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/K8gSRGHjeJLeW7fbyqbz">Workspaces</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/VpIuJWjoVervgD05b6e7">Compute resources</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/GuBN1o0IXQ2RGRmADNDd">Credentials</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS">Data sources</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/Aa4tONFSp6wI8dajbWxr">Data volumes</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4">Environments</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS#pvc">PVC</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4#adding-a-new-environment">Registries</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="https://run-ai-docs.nvidia.com/api/workload-assets/storage-class-configuration">Storage class configurations</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/j08IRnRyHoDV4ChTtT0a">Templates</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr></tbody></table>

</details>

<details>

<summary>L2 researcher <code>Legacy</code></summary>

<table><thead><tr><th width="221.51171875">Entity</th><th data-type="checkbox">View</th><th data-type="checkbox">Edit</th><th data-type="checkbox">Create</th><th data-type="checkbox">Delete</th></tr></thead><tbody><tr><td>Account</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hxy5GaNYO0MrDXL8r67p">AI Applications</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss#branding">Branding settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/KXk9Ii3EjuXL7rJP8Vhz">Departments</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hohfv4UUHT0dvOZ5dHbi">Event history</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Network topologies</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/EoHF94AYOoPhPYml00YS">Policies</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/6wxLpJPbjlnaHiHErmjg">Projects</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Security settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Clusters</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Clusters minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/64hpIzKT1fTtBZw8YSvQ">Node pools</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Node pools minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/FbftKDBouwYrCsumy9ie">Nodes</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/UPP9oksqWgXJbc9k3zdu">Access rules</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Applications</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/2CvyzKWvsYYcFSy6VpU1">Roles</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Service accounts</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/juQ16RWzgmohYWR4AbjH">Users</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Users minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VTxcwsPIfBGOXPU2NAc7">Inferences</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/aqIRhDCT9Ndcn878Dlwr">Trainings</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O#managing-workload-properties">Workload properties</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O">Workloads</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/K8gSRGHjeJLeW7fbyqbz">Workspaces</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/VpIuJWjoVervgD05b6e7">Compute resources</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/GuBN1o0IXQ2RGRmADNDd">Credentials</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS">Data sources</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/Aa4tONFSp6wI8dajbWxr">Data volumes</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4">Environments</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS#pvc">PVC</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4#adding-a-new-environment">Registries</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="https://run-ai-docs.nvidia.com/api/workload-assets/storage-class-configuration">Storage class configurations</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/j08IRnRyHoDV4ChTtT0a">Templates</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr></tbody></table>

</details>

<details>

<summary>ML engineer <code>Legacy</code></summary>

<table><thead><tr><th width="221.51171875">Entity</th><th data-type="checkbox">View</th><th data-type="checkbox">Edit</th><th data-type="checkbox">Create</th><th data-type="checkbox">Delete</th></tr></thead><tbody><tr><td>Account</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hxy5GaNYO0MrDXL8r67p">AI Applications</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss#branding">Branding settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/KXk9Ii3EjuXL7rJP8Vhz">Departments</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hohfv4UUHT0dvOZ5dHbi">Event history</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Network topologies</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/EoHF94AYOoPhPYml00YS">Policies</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/6wxLpJPbjlnaHiHErmjg">Projects</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Security settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Clusters</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Clusters minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/64hpIzKT1fTtBZw8YSvQ">Node pools</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Node pools minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/FbftKDBouwYrCsumy9ie">Nodes</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/UPP9oksqWgXJbc9k3zdu">Access rules</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Applications</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/2CvyzKWvsYYcFSy6VpU1">Roles</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Service accounts</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/juQ16RWzgmohYWR4AbjH">Users</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Users minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VTxcwsPIfBGOXPU2NAc7">Inferences</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/aqIRhDCT9Ndcn878Dlwr">Trainings</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O#managing-workload-properties">Workload properties</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O">Workloads</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/K8gSRGHjeJLeW7fbyqbz">Workspaces</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VpIuJWjoVervgD05b6e7">Compute resources</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/GuBN1o0IXQ2RGRmADNDd">Credentials</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS">Data sources</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/Aa4tONFSp6wI8dajbWxr">Data volumes</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4">Environments</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS#pvc">PVC</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4#adding-a-new-environment">Registries</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="https://run-ai-docs.nvidia.com/api/workload-assets/storage-class-configuration">Storage class configurations</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/j08IRnRyHoDV4ChTtT0a">Templates</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr></tbody></table>

</details>

<details>

<summary>Project administrator</summary>

<table><thead><tr><th width="221.51171875">Entity</th><th data-type="checkbox">View</th><th data-type="checkbox">Edit</th><th data-type="checkbox">Create</th><th data-type="checkbox">Delete</th></tr></thead><tbody><tr><td>Account</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hxy5GaNYO0MrDXL8r67p">AI Applications</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss#branding">Branding settings</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/KXk9Ii3EjuXL7rJP8Vhz">Departments</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hohfv4UUHT0dvOZ5dHbi">Event history</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Network topologies</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/EoHF94AYOoPhPYml00YS">Policies</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/6wxLpJPbjlnaHiHErmjg">Projects</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Security settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Clusters</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Clusters minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/64hpIzKT1fTtBZw8YSvQ">Node pools</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Node pools minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/FbftKDBouwYrCsumy9ie">Nodes</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/UPP9oksqWgXJbc9k3zdu">Access rules</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Applications</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/2CvyzKWvsYYcFSy6VpU1">Roles</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Service accounts</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/juQ16RWzgmohYWR4AbjH">Users</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Users minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VTxcwsPIfBGOXPU2NAc7">Inferences</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/aqIRhDCT9Ndcn878Dlwr">Trainings</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O#managing-workload-properties">Workload properties</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O">Workloads</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/K8gSRGHjeJLeW7fbyqbz">Workspaces</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/VpIuJWjoVervgD05b6e7">Compute resources</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/GuBN1o0IXQ2RGRmADNDd">Credentials</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS">Data sources</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/Aa4tONFSp6wI8dajbWxr">Data volumes</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4">Environments</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS#pvc">PVC</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4#adding-a-new-environment">Registries</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="https://run-ai-docs.nvidia.com/api/workload-assets/storage-class-configuration">Storage class configurations</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/j08IRnRyHoDV4ChTtT0a">Templates</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr></tbody></table>

</details>

<details>

<summary>Research manager <code>Legacy</code></summary>

<table><thead><tr><th width="221.51171875">Entity</th><th data-type="checkbox">View</th><th data-type="checkbox">Edit</th><th data-type="checkbox">Create</th><th data-type="checkbox">Delete</th></tr></thead><tbody><tr><td>Account</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hxy5GaNYO0MrDXL8r67p">AI Applications</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss#branding">Branding settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/KXk9Ii3EjuXL7rJP8Vhz">Departments</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hohfv4UUHT0dvOZ5dHbi">Event history</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Network topologies</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/EoHF94AYOoPhPYml00YS">Policies</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/6wxLpJPbjlnaHiHErmjg">Projects</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Security settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Clusters</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Clusters minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/64hpIzKT1fTtBZw8YSvQ">Node pools</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Node pools minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/FbftKDBouwYrCsumy9ie">Nodes</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/UPP9oksqWgXJbc9k3zdu">Access rules</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Applications</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/2CvyzKWvsYYcFSy6VpU1">Roles</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Service accounts</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/juQ16RWzgmohYWR4AbjH">Users</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Users minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VTxcwsPIfBGOXPU2NAc7">Inferences</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/aqIRhDCT9Ndcn878Dlwr">Trainings</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O#managing-workload-properties">Workload properties</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O">Workloads</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/K8gSRGHjeJLeW7fbyqbz">Workspaces</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/VpIuJWjoVervgD05b6e7">Compute resources</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/GuBN1o0IXQ2RGRmADNDd">Credentials</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS">Data sources</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/Aa4tONFSp6wI8dajbWxr">Data volumes</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4">Environments</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS#pvc">PVC</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4#adding-a-new-environment">Registries</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="https://run-ai-docs.nvidia.com/api/workload-assets/storage-class-configuration">Storage class configurations</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/j08IRnRyHoDV4ChTtT0a">Templates</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr></tbody></table>

</details>

<details>

<summary>System administrator</summary>

<table><thead><tr><th width="221.51171875">Entity</th><th data-type="checkbox">View</th><th data-type="checkbox">Edit</th><th data-type="checkbox">Create</th><th data-type="checkbox">Delete</th></tr></thead><tbody><tr><td>Account</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hxy5GaNYO0MrDXL8r67p">AI Applications</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss#branding">Branding settings</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/KXk9Ii3EjuXL7rJP8Vhz">Departments</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/hohfv4UUHT0dvOZ5dHbi">Event history</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Network topologies</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/EoHF94AYOoPhPYml00YS">Policies</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/6wxLpJPbjlnaHiHErmjg">Projects</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Security settings</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Settings</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Clusters</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td>Clusters minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/64hpIzKT1fTtBZw8YSvQ">Node pools</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td>Node pools minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/FbftKDBouwYrCsumy9ie">Nodes</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/UPP9oksqWgXJbc9k3zdu">Access rules</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Applications</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/2CvyzKWvsYYcFSy6VpU1">Roles</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Service accounts</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/juQ16RWzgmohYWR4AbjH">Users</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td>Users minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VTxcwsPIfBGOXPU2NAc7">Inferences</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/aqIRhDCT9Ndcn878Dlwr">Trainings</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O#managing-workload-properties">Workload properties</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O">Workloads</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/K8gSRGHjeJLeW7fbyqbz">Workspaces</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/VpIuJWjoVervgD05b6e7">Compute resources</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/GuBN1o0IXQ2RGRmADNDd">Credentials</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS">Data sources</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/Aa4tONFSp6wI8dajbWxr">Data volumes</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4">Environments</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS#pvc">PVC</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4#adding-a-new-environment">Registries</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="https://run-ai-docs.nvidia.com/api/workload-assets/storage-class-configuration">Storage class configurations</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr><tr><td><a href="/pages/j08IRnRyHoDV4ChTtT0a">Templates</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr></tbody></table>

</details>

<details>

<summary>Template administrator <code>Legacy</code></summary>

<table><thead><tr><th width="221.51171875">Entity</th><th data-type="checkbox">View</th><th data-type="checkbox">Edit</th><th data-type="checkbox">Create</th><th data-type="checkbox">Delete</th></tr></thead><tbody><tr><td>Account</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hxy5GaNYO0MrDXL8r67p">AI Applications</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss#branding">Branding settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/KXk9Ii3EjuXL7rJP8Vhz">Departments</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hohfv4UUHT0dvOZ5dHbi">Event history</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Network topologies</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/EoHF94AYOoPhPYml00YS">Policies</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/6wxLpJPbjlnaHiHErmjg">Projects</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Security settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Clusters</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Clusters minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/64hpIzKT1fTtBZw8YSvQ">Node pools</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Node pools minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/FbftKDBouwYrCsumy9ie">Nodes</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/UPP9oksqWgXJbc9k3zdu">Access rules</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Applications</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/2CvyzKWvsYYcFSy6VpU1">Roles</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Service accounts</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/juQ16RWzgmohYWR4AbjH">Users</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Users minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VTxcwsPIfBGOXPU2NAc7">Inferences</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/aqIRhDCT9Ndcn878Dlwr">Trainings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O#managing-workload-properties">Workload properties</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O">Workloads</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/K8gSRGHjeJLeW7fbyqbz">Workspaces</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VpIuJWjoVervgD05b6e7">Compute resources</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/GuBN1o0IXQ2RGRmADNDd">Credentials</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS">Data sources</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/Aa4tONFSp6wI8dajbWxr">Data volumes</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4">Environments</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS#pvc">PVC</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4#adding-a-new-environment">Registries</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="https://run-ai-docs.nvidia.com/api/workload-assets/storage-class-configuration">Storage class configurations</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/j08IRnRyHoDV4ChTtT0a">Templates</a></td><td>true</td><td>true</td><td>true</td><td>true</td></tr></tbody></table>

</details>

<details>

<summary>Viewer</summary>

<table><thead><tr><th width="221.51171875">Entity</th><th data-type="checkbox">View</th><th data-type="checkbox">Edit</th><th data-type="checkbox">Create</th><th data-type="checkbox">Delete</th></tr></thead><tbody><tr><td>Account</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hxy5GaNYO0MrDXL8r67p">AI Applications</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss#branding">Branding settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/KXk9Ii3EjuXL7rJP8Vhz">Departments</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/hohfv4UUHT0dvOZ5dHbi">Event history</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Network topologies</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/EoHF94AYOoPhPYml00YS">Policies</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/6wxLpJPbjlnaHiHErmjg">Projects</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Security settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/0LYytKCo5qglT1stQYss">Settings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MpGBP9Zn2tD9N1zN0A4b">Clusters</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Clusters minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/64hpIzKT1fTtBZw8YSvQ">Node pools</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Node pools minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/FbftKDBouwYrCsumy9ie">Nodes</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/UPP9oksqWgXJbc9k3zdu">Access rules</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Applications</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/2CvyzKWvsYYcFSy6VpU1">Roles</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/MZYKafmWAc94s3kd5so8">Service accounts</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/juQ16RWzgmohYWR4AbjH">Users</a></td><td>false</td><td>false</td><td>false</td><td>false</td></tr><tr><td>Users minimal</td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VTxcwsPIfBGOXPU2NAc7">Inferences</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/aqIRhDCT9Ndcn878Dlwr">Trainings</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O#managing-workload-properties">Workload properties</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/841Ux70WYGSy59VA5A9O">Workloads</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/K8gSRGHjeJLeW7fbyqbz">Workspaces</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VpIuJWjoVervgD05b6e7">Compute resources</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/GuBN1o0IXQ2RGRmADNDd">Credentials</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS">Data sources</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/Aa4tONFSp6wI8dajbWxr">Data volumes</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4">Environments</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/En4yELogvYkYI4gGTmjS#pvc">PVC</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/VvT46F94K378ZKZzzwA4#adding-a-new-environment">Registries</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="https://run-ai-docs.nvidia.com/api/workload-assets/storage-class-configuration">Storage class configurations</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr><tr><td><a href="/pages/j08IRnRyHoDV4ChTtT0a">Templates</a></td><td>true</td><td>false</td><td>false</td><td>false</td></tr></tbody></table>

</details>

## Custom Roles (API Only)

Administrators can create and manage custom roles to meet specific organizational access requirements using the [Roles](https://run-ai-docs.nvidia.com/api/authentication-and-authorization/roles) API.

A custom role defines access by combining one or more **permission sets**, which together determine the actions that can be performed across NVIDIA Run:ai resources. By selecting the appropriate permission sets, administrators can tailor roles to match their organization’s access model while maintaining consistent and controlled authorization across the platform.

{% hint style="info" %}
**Note**

Custom roles can be created and managed via API only. Once created, they are available for viewing in the Roles grid and can be selected when assigning access rules in the UI and API.
{% endhint %}

### Permission Sets

Permission sets are the building blocks of both predefined and custom roles. They represent the complete set of privileges required for a role to perform specific operations across NVIDIA Run:ai resources.

A permission set is constructed from a collection of permissions. Each permission within a permission set defines:

* a **resource type**, and
* the allowed **actions** on that resource: `create`, `read`, `update`, or `delete`

The Permission sets API provides a catalog of all available permission sets in the NVIDIA Run:ai platform. Permission sets may overlap in the resource types or actions they include. For more details, see the [Permissions](https://run-ai-docs.nvidia.com/api/authentication-and-authorization/permissions) API.

#### Permission Set Naming and Scope

Permission sets are designed to group actions logically, following a clear pattern to define scope. To view the complete list of permission sets and understand how each is defined, see the [Permissions](https://run-ai-docs.nvidia.com/api/authentication-and-authorization/permissions) API:

* Edit Access (`EditAccess`) - These sets grant full CRUD (Create, Read, Update, Delete) permissions over the specified resource type, along with Read access to any associated data or dependencies required for management. For example, **workspaceEditAccess** grants full permissions to `create`, `read`, `update`, and `delete` workspaces, and `read` access all related data and resources required for their management, such as:
  * `workloads`
  * `workload-properties`
  * `nodepools-minimal`
  * `clusters-minimal`
  * `users-minimal`
  * `projects`
  * `templates`
  * `environments`
  * `compute-resources`
  * `pvc-assets`
* Read Access (`ReadAccess`) - These sets grant only view permissions for the specific resource and its related components. For example, **workloadReadAccess** grants permissions to `read` workloads, and `read` access all related data and resources associated with them, such as:
  * `workspaces`
  * `trainings`
  * `inferences`
  * `workload-properties`
  * `nodepools-minimal`
  * `clusters-minimal`
  * `projects`

### Creating a Custom Role

#### Requirements for Custom Roles

Custom roles must meet the following requirements:

* **UI access permissions** - To allow a custom role to access the NVIDIA Run:ai user interface, include the following permission sets:
  * `settingsReadAccess`
  * `accountReadAccess`
  * `brandingSettingsReadAccess`
  * `securitySettingsReadAccess`
  * `notificationSettingsReadAccess`
* **Kubernetes (cluster) permissions** - Kubernetes permissions are **required only** if the role needs direct access to the Kubernetes cluster. If cluster access is required, the custom role must include a `kubernetesPermissions` object. This ensures the role inherits Kubernetes cluster-level permissions from an existing predefined NVIDIA Run:ai role.

#### Steps for Creating a Custom Role

1. Identify the required permission sets:
   * Determine the capabilities the new role needs (e.g., managing training jobs, viewing clusters, etc.).
   * Retrieve the full catalog of available permission set identifiers (`name` and `id`) using the `GET /v1/api/permission-sets` endpoint. This provides the list of all valid IDs (e.g., `inferenceEditAccess`, `workloadReadAccess`) to populate the `permissionSets` array in your request body.
   * Include `settingsReadAccess`, `accountReadAccess`, `brandingSettingsReadAccess`, `securitySettingsReadAccess` and `notificationSettingsReadAccess` to allow the custom role access the NVIDIA Run:ai user interface.
2. (Optional) Identify the Kubernetes predefined role ID. The custom role must inherit its underlying Kubernetes cluster permissions from an existing NVIDIA Run:ai predefined role if the role needs to access the Kubernetes cluster directly:
   * Retrieve the list of all available predefined roles using the `GET /v2/authorization/roles` endpoint.
   * Select the unique ID (a string identifier, e.g., "12") of the predefined role that grants the appropriate level of cluster access for your custom role.
   * This ID will populate the `predefinedRole` field under the `kubernetesPermissions` object.
3. Construct and send the API request. Use the `POST /v2/authorization/roles` endpoint to create the role.

**Example: MLOps Role**

This example demonstrates a role that can fully manage inference workloads and access related monitoring information.

```json
{
  "name": "MLOps",
  "description": "Fully manage inference workloads, and access other relevant information for monitoring",
  "permissionSets": [
    { "name": "inferenceEditAccess", "id": "<permissionSetId>" },
    { "name": "templatesEditAccess", "id": "<permissionSetId>" },
    { "name": "dataAndStorageReadAccess", "id": "<permissionSetId>" },
    { "name": "nodesReadAccess", "id": "<permissionSetId>" },
    { "name": "clustersReadAccess", "id": "<permissionSetId>" },
    { "name": "settingsReadAccess", "id": "<permissionSetId>" },
    { "name": "accountReadAccess", "id": "<permissionSetId>" }
  ],
  "kubernetesPermissions": {
    "predefinedRole": "12" 
  }
}
```

### Enabling / Disabling Custom Roles

Administrators can control whether it is available for use by enabling or disabling it:

* **Enabled** - Can be assigned to users and used in access rules.
* **Disabled** - Remain defined in the system but cannot be assigned to users or used in access rules.

### Troubleshooting Common Issues

<details>

<summary>Can't delete role</summary>

**Description:** The role cannot be deleted.

**Mitigation:**

1. Verify that the role is not assigned to any users.
2. If the role is currently in use, disable the role first, and then delete it.

</details>

<details>

<summary>Permissions not taking effect</summary>

**Description:** The selected permission sets are not taking effect.

**Mitigation:**

1. Ensure that the role is **enabled**.
2. Verify that the user has been **assigned the role**.
3. Check for **conflicting permissions** granted by other roles assigned to the user.

</details>

<details>

<summary>Cluster permissions not working</summary>

**Description:** Cluster-level permissions are not taking effect.

**Mitigation:**

1. Verify that the required ClusterRole exists in all relevant clusters.
2. Confirm that the cluster version is 2.21 or later.
3. Ensure that the cluster is properly connected to the tenant.

</details>


# Service Accounts (formerly applications)

Service accounts are used for API integrations with NVIDIA Run:ai. A service account contains a client ID and a client secret. With the client credentials, you can obtain an access token as detailed in [API authentication](https://run-ai-docs.nvidia.com/api/getting-started/how-to-authenticate-to-the-api) and use it within subsequent API calls.

Service accounts are assigned with [access rules ](/saas/infrastructure-setup/authentication/accessrules)to manage permissions. For example, service account **ci-pipeline-prod** is assigned with a **Researcher** role in **Cluster: A**.

{% hint style="info" %}
**Note**

You can also create your own access keys for API integrations with NVIDIA Run:ai. See [User access keys](/saas/settings/user-settings/user-access-keys) for more details.
{% endhint %}

## Service Accounts Table

The Service accounts table can be found under **Access** in the NVIDIA Run:ai platform.

The Service accounts table provides a list of all the services accounts defined in the platform, and allows you to manage them.

<figure><img src="/files/oyfHgiw9fneE8nKm95WZ" alt=""><figcaption></figcaption></figure>

The Service accounts table consists of the following columns:

| Column          | Description                                            |
| --------------- | ------------------------------------------------------ |
| Service account | The name of the service account                        |
| Client ID       | The client ID of the service account                   |
| Access rule(s)  | The access rules assigned to the service account       |
| Last login      | The timestamp for the last time the user signed in     |
| Created by      | The user who created the service account               |
| Creation time   | The timestamp for when the service account was created |
| Last updated    | The last time the service account was updated          |

### Customizing the Table View

* Filter - Click ADD FILTER, select the column to filter by, and enter the filter values
* Search - Click SEARCH and type the value to search by
* Sort - Click each column header to sort by
* Column selection - Click COLUMNS and select the columns to display in the table
* Download table - Click MORE and then Click Download as CSV. Export to CSV is limited to 20,000 rows.

## Creating a Service Account

To create a service account:

1. Click **+NEW SERVICE ACCOUNT**
2. Enter the service account's **name**
3. Click **CREATE**
4. Copy the **Client ID** and **Client secret** and store them securely
5. Click **DONE** or click **ADD ACCESS RULES** to proceed to the **Access rules** step
6. In the **Access rules** step:
   * Select a **role**
   * Select up to 10 **scopes** where the access rule will apply
   * Click **SAVE RULE**
   * Click **CLOSE**

{% hint style="info" %}
**Note**

The client secret is visible only at the time of creation. It cannot be recovered but can be regenerated.
{% endhint %}

## Adding an Access Rule to a Service Account

To create an access rule:

1. Select the service account you want to add an access rule for
2. Click **ACCESS RULES**
3. Click **+ACCESS RULE**
4. Select a **role**
5. Select up to 10 **scopes** where the access rule will apply
6. Click **SAVE RULE**
7. Click **CLOSE**

## Deleting an Access Rule from a Service Account

To delete an access rule:

1. Select the service account you want to remove an access rule from
2. Click **ACCESS RULES**
3. Find the access rule assigned to the user you would like to delete
4. Click on the trash icon
5. Click **CLOSE**

## Regenerating a Client Secret

To regenerate a client secret:

1. Locate the service account you want to regenerate its client secret
2. Click **REGENERATE CLIENT SECRET**
3. Click **REGENERATE**
4. Copy the **New client secret** and store it securely
5. Click **DONE**

{% hint style="warning" %}
**Important**

Regenerating a client secret revokes the previous one.
{% endhint %}

## Deleting a Service Account

1. Select the service account you want to delete
2. Click **DELETE**
3. On the dialog, click **DELETE** to confirm

## Using API

Go to the [Service accounts](https://run-ai-docs.nvidia.com/api/authentication-and-authorization/service-accounts), [Access rules](https://run-ai-docs.nvidia.com/api/authentication-and-authorization/access-rules) API reference to view the available actions.


# Access Rules

Access rules provide users, groups, or service accounts privileges to system entities. An access rule is the assignment of a [role ](/saas/infrastructure-setup/authentication/roles)to a [subject in a scope](/saas/platform-management/aiinitiatives/adapting-ai-initiatives#scopes-in-an-organization): `<Subject>` is a `<Role>` in a `<Scope>`. For example, user **<user@domain.com>** is a **department admin** in **department A**.

## Access Rules Table

The Access rules table can be found under **Access** in the NVIDIA Run:ai platform.

The Access rules table provides a list of all the access rules defined in the platform and allows you to manage them.

{% hint style="info" %}
**Flexible management**

It is also possible to manage access rules directly for a specific [user](/saas/infrastructure-setup/authentication/users), [service account](/saas/infrastructure-setup/authentication/service-accounts), [project](/saas/platform-management/aiinitiatives/organization/projects), or [department](/saas/platform-management/aiinitiatives/organization/departments).
{% endhint %}

<figure><img src="/files/e3s4xsIo7JCRKn9q0FDE" alt=""><figcaption></figcaption></figure>

The Access rules table consists of the following columns:

| Column        | Description                                                                                                  |
| ------------- | ------------------------------------------------------------------------------------------------------------ |
| Type          | The type of subject assigned to the access rule (user, SSO group, or service account).                       |
| Subject       | The user, SSO group, or service account assigned with the role                                               |
| Role          | The role assigned to the subject                                                                             |
| Scope         | The scope to which the subject has access. Click the name of the scope to see the scope and its subordinates |
| Status        | The access rule creation status.                                                                             |
| Authorized by | The user who granted the access rule                                                                         |
| Creation time | The timestamp for when the rule was created                                                                  |
| Last updated  | The last time the access rule was updated                                                                    |

### Access Rule Status

The following table describes the condition of access rules and whether they were successfully created for the selected scope. The status indicates whether the access rule is synced in the cluster.

| Status          | Description                                                                                                                                                        |
| --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| No issues found | No issues were found while creating the access rule                                                                                                                |
| Issues found    | Issues were found while enforcing the access rule in the cluster(s). [Contact NVIDIA Run:ai support](https://www.nvidia.com/en-eu/support/enterprise/#contact-us). |
| Creating…       | The access rule is being created in the cluster                                                                                                                    |
| No status       | Status cannot be displayed because the current version of the cluster is not up to date or unknown                                                                 |

### Customizing the Table View

* Filter - Click ADD FILTER, select the column to filter by, and enter the filter values
* Search - Click SEARCH and type the value to search by
* Sort - Click each column header to sort by
* Column selection - Click COLUMNS and select the columns to display in the table
* Download table - Click MORE and then Click Download as CSV. Export to CSV is limited to 20,000 rows.

## Adding New Access Rules

To add a new access rule:

1. Click **+NEW ACCESS RULE**
2. Select a subject - **User, SSO Group**, or **Service Account**
3. Select or enter the subject identifier. You can define up to 10 subjects of the selected type:
   * **User** for a local user created in NVIDIA Run:ai or for an SSO user as recognized by the IDP. Search by name or email, or manually enter the user's email if they do not appear in the search results.
   * **Group name** as recognized by the IDP
   * **Service account name** as created in NVIDIA Run:ai
4. Select a **role**
5. Select up to 10 **scopes** where the access rule will apply
6. Click **SAVE RULE**

## Editing an Access Rule

Access rules cannot be edited. To change an access rule, you must delete the rule, and then create a new rule to replace it.

## Deleting an Access Rule

1. Select one or more access rules you want to delete
2. Click **DELETE**
3. On the dialog, click **DELETE** to confirm

## Transitioning Access Rules to New Roles

Due to the deprecation of some roles, referred to as [legacy roles](/saas/infrastructure-setup/authentication/roles#roles-in-nvidia-run-ai), access rules using them should be transitioned to supported roles. When access rules still use legacy roles, a banner at the top of the Access rules table prompts you to transition them. Transitioning recreates each selected rule under a new role with the same subject and scope, and deletes the legacy rule.

To transition access rules:

1. On the banner, click **transition to new roles**
2. In the **Transition Access Rules To New Roles** dialog, select the access rules to transition
3. For each legacy role listed, select a **new role**. For some legacy roles, NVIDIA Run:ai pre-selects a recommended replacement, which you can change.
4. Click **TRANSITION ACCESS RULES** to confirm

{% hint style="info" %}
**Required permissions**

Your role must include the **Create**, **Read**, **Update**, and **Delete** permissions on the **Access rules** resource.
{% endhint %}

## Viewing Your User Access Rule

To view the assigned roles and scopes you have access to:

1. Click the user avatar at the top right corner, then select **Settings**
2. Click **User details**

The list of assigned roles and scopes will be displayed.

## Using API

Go to the [Access rules](https://run-ai-docs.nvidia.com/api/authentication-and-authorization/access-rules) API reference to view the available actions.


# Cluster Authentication

To allow users to securely submit workloads using `kubectl`, you must configure the Kubernetes API server to authenticate users via the NVIDIA Run:ai identity provider. This is done by adding OpenID Connect (OIDC) flags to the Kubernetes API server configuration on each cluster.

### Retrieve Required OIDC Flags

1. Go to **General settings**
2. Navigate to **Cluster authentication**

```yaml
  containers:
  - command:
    ...
    - --oidc-client-id=runai
    - --oidc-issuer-url=https://<HOST>/auth/realms/runai
    - --oidc-username-prefix=-
```

* `--oidc-client-id` - A client id that all tokens must be issued for.
* `--oidc-issuer-url` - The URL of the NVIDIA Run:ai identity provider
* `--oidc-username-prefix` - Prefix prepended to username claims to prevent clashes with existing names (e.g., `-user@example.com`).

{% hint style="info" %}
**Note**

These flags must be configured in the API server startup parameters for each cluster in your environment.
{% endhint %}

## Kubernetes Distribution-Specific Configuration

{% hint style="info" %}
**Note**

* Azure Kubernetes Service (AKS) is not supported.
* For other Kubernetes distributions, refer to specific instructions in the documentation.
  {% endhint %}

<details>

<summary>Vanilla Kubernetes</summary>

1. Locate the Kubernetes API server configuration file. For vanilla Kubernetes, the configuration file is typically located at: `/etc/kubernetes/manifests/kube-apiserver.yaml`.
2. Edit the file. Under the `command` section, add the [required OIDC flags](#retrieve-required-oidc-values).
3. Verify that the changes have been applied. After saving the file, the API server should automatically restart since it's managed as a static pod. Confirm that the `kube-apiserver-<master-node-name>` pod in the `kube-system` namespace has restarted and is running with the new configuration. You can run the following command to check the pod status:

   ```bash
   kubectl get pods -n kube-system kube-apiserver-<master-node-name> -o yaml
   ```

</details>

<details>

<summary>OpenShift Container Platform (OCP)</summary>

No additional configuration is required.

</details>

<details>

<summary>Rancher Kubernetes Engine 2 (RKE2)</summary>

If you're using the [RKE2 Quickstart](https://docs.rke2.io/install/quickstart/):

1. Edit `/etc/rancher/rke2/config.yaml`.
2. Add the [required OIDC flags](#retrieve-required-oidc-values) under `kube-apiserver-arg`, using the format shown below:

   ```yaml
   kube-apiserver-arg:
   - "oidc-client-id=runai" # 
   ...
   ```

If you're using Rancher UI:

1. Add the required flags during the cluster provisioning process.
2. Navigate to: Cluster Management > Create, select RKE2, and choose your platform.
3. In the Cluster Configuration screen, go to: Advanced > Additional API Server Args.
4. Add the [required OIDC flags](#retrieve-required-oidc-values) as `<key>=<value>` (e.g. `oidc-username-prefix=-`).

</details>

<details>

<summary>Google Kubernetes Engine (GKE)</summary>

To configure researcher authentication on GKE, use **Anthos Identity Service** and apply the appropriate OIDC configuration.

1. Install [Anthos identity service](https://cloud.google.com/kubernetes-engine/docs/how-to/oidc#enable-oidc) by running:

   ```bash
   gcloud container clusters update <gke-cluster-name> \
       --enable-identity-service --project=<gcp-project-name> --zone=<gcp-zone-name>
   ```
2. Install the [yq](https://github.com/mikefarah/yq) utility.
3. Configure the OIDC provider for username-password authentication. Make sure to use the [required OIDC flags](#retrieve-required-oidc-values):

   ```bash
   kubectl get clientconfig default -n kube-public -o yaml > login-config.yaml
   yq -i e ".spec +={\"authentication\":[{\"name\":\"oidc\",\"oidc\":{\"clientID\":\"runai\",\"issuerURI\":\"$OIDC_ISSUER_URL\",\"kubectlRedirectURI\":\"http://localhost:8000/callback\",\"userClaim\":\"sub\",\"userPrefix\":\"-\"}}]}" login-config.yaml
   kubectl apply -f login-config.yaml
   ```
4. Or, configure the OIDC provider for single-sign-on. Make sure to use the [required OIDC flags](#retrieve-required-oidc-values):

   ```bash
   kubectl get clientconfig default -n kube-public -o yaml > login-config.yaml
   yq -i e ".spec +={\"authentication\":[{\"name\":\"oidc\",\"oidc\":{\"clientID\":\"runai\",\"issuerURI\":\"$OIDC_ISSUER_URL\",\"groupsClaim\":\"groups\",\"kubectlRedirectURI\":\"http://localhost:8000/callback\",\"userClaim\":\"sub\",\"userPrefix\":\"-\"}}]}" login-config.yaml
   kubectl apply -f login-config.yaml
   ```
5. Update the `runaiconfig` with the Anthos Identity Service endpoint. First, get the external IP of the `gke-oidc-envoy` service:

   ```bash
   kubectl get svc -n anthos-identity-service
   NAME               TYPE           CLUSTER-IP    EXTERNAL-IP     PORT(S)              AGE
   gke-oidc-envoy     LoadBalancer   10.37.3.111   39.201.319.10   443:31545/TCP        12h
   ```
6. Then, patch the `runaiconfig` to use this endpoint. Replace the below with the actual IP address of the `gke-oidc-envoy` service:

   ```bash
   kubectl -n runai patch runaiconfig runai -p '{"spec": {"researcher-service": 
   {"args": {"gkeOidcEnvoyHost": "35.236.229.19"}}}}'  --type="merge"
   ```

</details>

<details>

<summary>Elastic Kubernetes Engine (EKS)</summary>

1. In the AWS Console, under EKS, find your cluster.
2. Go to `Configuration` and then to `Authentication`.
3. Associate a new `identity provider`. Use the [required OIDC flags](#retrieve-required-oidc-values).

The process can take up to 30 minutes.

</details>

<details>

<summary>NVIDIA Base Command Manager (BCM)</summary>

While it is possible to edit the manifest files directly on the control plane nodes (as described in the Vanilla Kubernetes instructions), the BCM preferred method is to make changes through the kubeadm configmap. This ensures that kubeadm can always regenerate the manifest files correctly — for example, when adding new control plane nodes or upgrading Kubernetes.

{% hint style="warning" %}
**Warning**

If you previously modified `/etc/kubernetes/manifests/kube-apiserver.yaml` directly on the nodes, those changes may be overwritten the next time `cm-kubeadm-manage` is used or when Kubernetes is upgraded. Use the configmap method below to ensure your OIDC flags persist.
{% endhint %}

1. Edit the kubeadm configmap:

   ```bash
   kubectl edit configmap -n kube-system kubeadm-config
   ```
2. Add the [required OIDC flags](#retrieve-required-oidc-values) under the `apiServer.extraArgs` block:

   ```yaml
   extraArgs:
   - name: oidc-client-id
     value: "runai"
   - name: oidc-issuer-url
     value: "https://<HOST>/auth/realms/runai"
   - name: oidc-username-prefix
     value: "-"
   ```

   A full example of the relevant configmap section:

   ```yaml
   apiVersion: v1
   data:
     ClusterConfiguration: |
       apiServer:
         certSANs:
         - <your-cluster-hostname>
         - master
         - localhost
         extraArgs:
         - name: oidc-client-id
           value: "runai"
         - name: oidc-issuer-url
           value: "https://<HOST>/auth/realms/runai"
         - name: oidc-username-prefix
           value: "-"
       apiVersion: kubeadm.k8s.io/v1beta4
   ```
3. Propagate the configmap changes to the control plane nodes:

   ```bash
   /cm/local/apps/cmd/scripts/cm-kubeadm-manage --kube-cluster default update_configmap
   ```
4. For each control plane node, regenerate the API server manifest:

   ```bash
   /cm/local/apps/cmd/scripts/cm-kubeadm-manage --kube-cluster default update_apiserver <node-name>
   ```

   Validate the changes by confirming the API server pod is using the OIDC flags:

   ```bash
   kubectl describe pod -n kube-system kube-apiserver-<node-name> | grep oidc
   ```

   Expected output:

   ```
   --oidc-client-id=runai
   --oidc-issuer-url=https://<HOST>/auth/realms/runai
   --oidc-username-prefix=-
   ```

   Once confirmed, proceed to the next control plane node and repeat until all nodes are updated.

</details>


# Advanced Setup


# Node Roles

This guide explains how to designate specific node roles in a Kubernetes cluster to ensure optimal performance and reliability in production deployments.

For optimal performance in production clusters, it is essential to avoid extensive CPU usage on GPU nodes where possible. This can be done by ensuring the following:

* NVIDIA Run:ai system-level services run on dedicated CPU-only nodes.
* Workloads that do not request GPU resources (e.g. Machine Learning jobs) are executed on CPU-only nodes.

NVIDIA Run:ai services are scheduled on the defined node roles by applying [Kubernetes Node Affinity](https://kubernetes.io/docs/concepts/scheduling-eviction/assign-pod-node/#affinity-and-anti-affinity) using node labels .

## Configure Node Roles

The following node roles can be configured on the cluster:

* **System node:** Reserved for NVIDIA Run:ai system-level services.
* **GPU Worker node:** Dedicated for GPU-based workloads.
* **CPU Worker node:** Used for CPU-only workloads.

### System Nodes

NVIDIA Run:ai system nodes run system-level services required to operate. This can be done via [Kubectl](https://kubernetes.io/docs/reference/kubectl/). By default, NVIDIA Run:ai applies a node affinity rule to prefer nodes that are labeled with `node-role.kubernetes.io/runai-system` for system services scheduling. You can modify the default node affinity rule by:

* Editing the `global.affinity` configuration parameter as detailed in [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config).
* Editing the `global.affinity` configuration as detailed in [Advanced control plane configurations](/self-hosted/2.25/infrastructure-setup/advanced-setup/control-plane-config) for self-hosted deployments.

To set a system role for a node in your Kubernetes cluster using Kubectl, follow these steps:

1. Use the `kubectl get nodes` command to list all the nodes in your cluster and identify the name of the node you want to modify.
2. Run one of the following commands to label the node with its role:

   ```bash
   kubectl label nodes <node-name> node-role.kubernetes.io/runai-system=true
   kubectl label nodes <node-name> node-role.kubernetes.io/runai-system=false
   ```

{% hint style="info" %}
**Note**

* To ensure [high availability](/saas/infrastructure-setup/procedures/high-availability) and prevent a single point of failure, it is recommended to configure at least three system nodes in your cluster.
* By default, Kubernetes master nodes are configured to prevent workloads from running on them as a best-practice measure to safeguard control plane stability. While this restriction is generally recommended, certain NVIDIA reference architectures allow adding tolerations to the NVIDIA Run:ai deployment so critical system services can run on these nodes.
  {% endhint %}

### Worker Nodes

NVIDIA Run:ai worker nodes run user-submitted workloads and system-level [DeamonSets](https://kubernetes.io/docs/concepts/workloads/controllers/daemonset/) required to operate. This can be managed via [Kubectl](https://kubernetes.io/docs/reference/kubectl/).

By default, GPU workloads are scheduled on GPU nodes based on the `nvidia.com/gpu.present` label. When `clusterConfig.global.nodeAffinity.restrictScheduling` is set to true via the [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config):

* GPU Workloads are scheduled with node affinity rule to require nodes that are labeled with `node-role.kubernetes.io/runai-gpu-worker`
* CPU-only Workloads are scheduled with node affinity rule to require nodes that are labeled with `node-role.kubernetes.io/runai-cpu-worker`

To set a worker role for a node in your Kubernetes cluster using Kubectl, follow these steps:

1. Validate the `clusterConfig.global.nodeAffinity.restrictScheduling` is set to true in the cluster’s [Configurations](/saas/infrastructure-setup/advanced-setup/cluster-config).
2. Use the `kubectl get nodes` command to list all the nodes in your cluster and identify the name of the node you want to modify.
3. Run one of the following commands to label the node with its role. Replace the label and value (`true`/`false`) to enable or disable GPU/CPU roles as needed:

   ```bash
   kubectl label nodes <node-name> node-role.kubernetes.io/runai-gpu-worker=true
   kubectl label nodes <node-name> node-role.kubernetes.io/runai-cpu-worker=false
   ```


# Advanced Cluster Configurations

Advanced cluster configurations allow you to customize your NVIDIA Run:ai cluster deployment to support your environment. Some settings may be required for deployment, while others can be fine-tuned to align with organizational policies, security requirements, or other operational preferences.

By adjusting these configurations, you can influence system behavior, including functionality, scheduling policies, and resource management, giving you greater control over how the cluster operates. This article provides guidance on configuring and managing these settings so you can adapt your NVIDIA Run:ai cluster to your organization’s needs.

## Configuration Scope

The Helm chart provides the complete set of configuration options for the NVIDIA Run:ai cluster. Some options, such as `global.affinity`, are available only through Helm. The rest of the configurable settings are grouped under `clusterConfig`. These `clusterConfig` settings can be applied through Helm as part of your deployment or upgrade process.

The `clusterConfig` subset can also be managed at runtime through the `runaiconfig` Custom Resource (under `spec`). For details, see [Modify cluster configurations at runtime](#modify-cluster-configurations-at-runtime).

At runtime, `runaiconfig` is the source of truth for the active cluster configuration. If a configuration key is defined in both Helm and `runaiconfig` and the values differ, a Helm upgrade will overwrite the `runaiconfig` value to match the chart.

{% hint style="info" %}
**Note**

The approach remains backward compatible so existing clusters configured via `runaiconfig` continue to work. However, when using Helm, you should manage configurations exclusively through Helm values. Mixing Helm values and manual `runaiconfig` edits is not recommended.
{% endhint %}

## Helm Chart Values

The NVIDIA Run:ai cluster installation can be customized to support your environment via Helm [values files](https://helm.sh/docs/chart_template_guide/values_files/) or [Helm install](https://helm.sh/docs/helm/helm_install/) flags. For example:

```yaml
# values.yaml
global:
  affinity:
    nodeAffinity:
      requiredDuringSchedulingIgnoredDuringExecution:
        nodeSelectorTerms:
          - matchExpressions:
              - key: node-role.kubernetes.io/runai-system
                operator: Exists
```

## Modify Cluster Configurations at Runtime

The `clusterConfig` subset of settings can also be managed at runtime via the `runaiconfig` [Kubernetes Custom Resource](https://kubernetes.io/docs/concepts/extend-kubernetes/api-extension/custom-resources/).

1. To edit the cluster configurations, run:

   ```bash
   kubectl edit runaiconfig runai -n runai
   ```
2. To see the full `runaiconfig` object structure, use:

   ```bash
   kubectl get crds/runaiconfigs.run.ai -n runai -o yaml
   ```

When using `runaiconfig`, the `clusterConfig` values appear under `spec`.

## Configurations

The following configurations allow you to enable or disable features, control permissions, and customize the behavior of your NVIDIA Run:ai cluster

{% hint style="info" %}
**Note**

* Keys that start with `clusterConfig` are available both in Helm and at runtime in `runaiconfig` (under `spec`). All other keys are available only through Helm.
* At runtime, `runaiconfig` reflects the active configuration. If specific `clusterConfig` values set in `runaiconfig` differ from those in Helm, a subsequent Helm upgrade will overwrite only those fields to align with the chart.
  {% endhint %}

### Helm Only Configurations

| Key                                              | Description                                                                                                                                                                                                                                                                                                                         |
| ------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `global.image.registry` *(string)*               | <p>Global Docker image registry<br>Default: <code>""</code></p>                                                                                                                                                                                                                                                                     |
| `global.additionalImagePullSecrets` *(list)*     | <p>List of image pull secrets references<br>Default: <code>\[]</code></p>                                                                                                                                                                                                                                                           |
| `global.affinity` *(object)*                     | <p>Sets the system nodes where NVIDIA Run:ai system-level services are scheduled. Using global.affinity will overwrite the <a href="/pages/3aLJWOEfQ78hvlKKBk8O">node roles</a> set using <code>kubectl</code>.<br>Default: Prefer to schedule on nodes that are labeled with <code>node-role.kubernetes.io/runai-system</code></p> |
| `global.tolerations` *(object)*                  | Configure Kubernetes tolerations for NVIDIA Run:ai system-level services                                                                                                                                                                                                                                                            |
| `global.additionalJobLabels` *(object)*          | <p>Set NVIDIA Run:ai and 3rd party services' <a href="https://kubernetes.io/docs/concepts/overview/working-with-objects/labels">Pod Labels </a>in a format of key/value pairs.<br>Default: <code>""</code></p>                                                                                                                      |
| `global.additionalJobAnnotations` *(object)*     | <p>Set NVIDIA Run:ai and 3rd party services' <a href="https://kubernetes.io/docs/concepts/overview/working-with-objects/annotations/">Annotations</a> in a format of key/value pairs.<br>Default: <code>""</code></p>                                                                                                               |
| `global.customCA.enabled`                        | Enables the use of a custom Certificate Authority (CA) in your deployment. When set to `true`, the system is configured to trust a user-provided CA certificate for secure communication.                                                                                                                                           |
| `global.customCAGit.enabled`                     | Enables the use of a custom Certificate Authority (CA) for Git data sources. When set to `true`, the system uses the global CA certificate defined at installation unless overridden using `global.customCAGit.secret.name`.                                                                                                        |
| `global.customCAGit.secret.name`                 | Specifies the name of the Kubernetes secret that contains a custom CA certificate for Git data sources. Overrides the default global CA when `global.customCAGit.enabled` is set to `true`.                                                                                                                                         |
| `global.customCAS3.enabled`                      | Enables the use of a custom Certificate Authority (CA) for S3 data sources. When set to `true`, the system uses the global CA certificate defined at installation unless overridden using `global.customS3Git.secret.name`.                                                                                                         |
| `global.customCAS3.secret.name`                  | Specifies the name of the Kubernetes secret that contains a custom CA certificate for S3 data sources. Overrides the default global CA when `global.customCAS3.enabled` is set to `true`.                                                                                                                                           |
| `openShift.securityContextConstraints.create`    | Enables the deployment of Security Context Constraints (SCC). Disable for CIS compliance. Default: `true`                                                                                                                                                                                                                           |
| `researcherService.ingress.tlsSecret` *(string)* | Existing secret key where cluster [TLS certificates](/saas/getting-started/installation/install-using-helm/system-requirements#nvidia-run-ai-tls-certificates) are stored (non-OpenShift) Default: `runai-cluster-domain-tls-secret`                                                                                                |
| `researcherService.route.tlsSecret` *(string)*   | Existing secret key where cluster [TLS certificates](/saas/getting-started/installation/install-using-helm/system-requirements#nvidia-run-ai-tls-certificates) are stored (OpenShift only) Default: `""`                                                                                                                            |
| `controlPlane.existingSecret`                    | Specifies the name of the existing Kubernetes secret where the cluster’s `clientSecret` used for secure connection with the control plane is stored.                                                                                                                                                                                |
| `controlPlane.secretKeys.clientSecret`           | Specifies the key within the `controlPlane.existingSecret` that stores the cluster’s `clientSecret` used for secure connection with the control plane.                                                                                                                                                                              |

### All Configurations (Helm and runaiconfig)

<table><thead><tr><th width="236.20703125">Helm Key</th><th width="235.8125">runaiconfig Key</th><th>Description</th></tr></thead><tbody><tr><td><code>clusterConfig.global.nodeAffinity.restrictScheduling</code> <em>(boolean)</em></td><td><code>spec.global.nodeAffinity.restrictScheduling</code> <em>(boolean)</em></td><td>Enables setting <a href="/pages/3aLJWOEfQ78hvlKKBk8O">node roles</a> and restricting workload scheduling to designated nodes<br>Default: <code>false</code></td></tr><tr><td><code>clusterConfig.global.ingress.ingressClass</code></td><td><code>spec.global.ingress.ingressClass</code></td><td>NVIDIA Run:ai uses NGINX as the default ingress controller. If your cluster has a different ingress controller, you can configure the ingress class to be created by NVIDIA Run:ai.</td></tr><tr><td><code>clusterConfig.global.subdomainSupport</code> <em>(boolean)</em></td><td><code>spec.global.subdomainSupport</code> <em>(boolean)</em></td><td>Enables host-based routing, exposing workload URLs as subdomains on the cluster's <a href="/pages/Ody4Af15R7fOat8L9zlX#fully-qualified-domain-name-fqdn">FQDN</a>. To use path-based routing instead, set to <code>false</code>. For setup instructions, see <a href="/pages/Ody4Af15R7fOat8L9zlX#host-based-routing-default">System requirements</a>. For more details, see <a href="/pages/9U3H4WiDkcbwrj4W66NS">External Access to Containers</a>.<br>Default: <code>true</code></td></tr><tr><td><code>clusterConfig.global.devicePluginBindings</code> <em>(boolean)</em></td><td><code>spec.global.devicePluginBindings</code> <em>(boolean)</em></td><td>Instruct NVIDIA Run:ai fractions to use device plugin for host mount instead of NVIDIA Run:ai fractions using explicit host path mount configuration on the pod. See <a href="/pages/Ap3Un7M2pNFM5QewTcuu">GPU fractions</a> and <a href="/pages/UgIxXMHwanFJyzcVnkte">dynamic GPU fractions</a>.<br>Default: <code>false</code></td></tr><tr><td><code>clusterConfig.global.enableWorkloadOwnershipProtection</code> <em>(boolean)</em></td><td><code>spec.global.enableWorkloadOwnershipProtection</code> <em>(boolean)</em></td><td>Prevents users within the same project from deleting workloads created by others. This enhances workload ownership security and ensures better collaboration by restricting unauthorized modifications or deletions.<br>Default: <code>false</code></td></tr><tr><td><code>clusterConfig.global.requireDefaultPodAntiAffinity</code> <em>(boolean)</em></td><td><code>spec.global.requireDefaultPodAntiAffinity</code> <em>(boolean)</em></td><td>When enabled, NVIDIA Run:ai applies a default pod anti-affinity rule that attempts to prevent pods belonging to the same service from being scheduled on the same node.<br>Default: <code>true</code></td></tr><tr><td><code>clusterConfig.project-controller.createNamespaces</code> <em>(boolean)</em></td><td><code>spec.project-controller.createNamespaces</code> <em>(boolean)</em></td><td>Allows Kubernetes namespace creation for new projects<br>Default: <code>true</code></td></tr><tr><td><code>clusterConfig.project-controller.CreateRoleBindings</code> <em>(boolean)</em></td><td><code>spec.project-controller.CreateRoleBindings</code> <em>(boolean)</em></td><td>Specifies if role bindings should be created in the project's namespace<br>Default: <code>true</code></td></tr><tr><td><code>clusterConfig.project-controller.limitRange</code> <em>(boolean)</em></td><td><code>spec.project-controller.limitRange</code> <em>(boolean)</em></td><td>Specifies if limit ranges should be defined for projects<br>Default: <code>true</code></td></tr><tr><td><code>clusterConfig.project-controller.clusterWideSecret</code> <em>(boolean)</em></td><td><code>spec.project-controller.clusterWideSecret</code> <em>(boolean)</em></td><td>Allows Kubernetes Secrets creation at the cluster scope. See <a href="/pages/GuBN1o0IXQ2RGRmADNDd#creating-secrets-in-advance">Credentials</a> for more details.<br>Default: <code>true</code></td></tr><tr><td><code>clusterConfig.workload-controller.failureResourceCleanupPolicy</code></td><td><code>spec.workload-controller.failureResourceCleanupPolicy</code></td><td><p>NVIDIA Run:ai cleans the workload's unnecessary resources:</p><ul><li><code>All</code> - Removes all resources of the failed workload</li><li><code>None</code> - Retains all resources</li><li><code>KeepFailing</code> - Removes all resources except for those that encountered issues (primarily for debugging purposes)</li></ul><p>Default: <code>All</code></p></td></tr><tr><td><code>clusterConfig.workload-controller.externalAuthUrlEnabled</code> <em>(boolean)</em></td><td><code>spec.workload-controller.externalAuthUrlEnabled</code> <em>(boolean)</em></td><td>Determines whether to use the OAuth2 proxy when submitting workloads that expose an external URL. Set to <code>false</code> to disable.<br>Default: <code>true</code></td></tr><tr><td><code>clusterConfig.inference-workload-controller.externalAuthUrlEnabled</code> <em>(boolean)</em></td><td><code>spec.inference-workload-controller.externalAuthUrlEnabled</code> <em>(boolean)</em></td><td>Determines whether to use the OAuth2 proxy when submitting inference workloads that expose an external URL. Set to <code>false</code> to disable.<br>Default: <code>true</code></td></tr><tr><td><code>clusterConfig.workload-controller.GPUNetworkAccelerationEnabled</code></td><td><code>spec.workload-controller.GPUNetworkAccelerationEnabled</code></td><td>Enables GPU network acceleration for workloads managed by the workload controller. See <a href="/pages/6qh7jEUwmQ2poP59Ty5Y">Using GB200 NVL72 and Multi-Node NVLink Domains</a> for more details.<br>Default: <code>false</code></td></tr><tr><td><code>clusterConfig.inference-workload-controller.GPUNetworkAccelerationEnabled</code></td><td><code>spec.inference-workload-controller.GPUNetworkAccelerationEnabled</code></td><td>Enables GPU network acceleration for workloads managed by the inference workload controller. See <a href="/pages/6qh7jEUwmQ2poP59Ty5Y">Using GB200 NVL72 and Multi-Node NVLink Domains</a> for more details.<br>Default: <code>false</code></td></tr><tr><td><code>clusterConfig.anyworkload-controller.GPUNetworkAccelerationEnabled</code></td><td><code>spec.anyworkload-controller.GPUNetworkAccelerationEnabled</code></td><td>Enables GPU network acceleration for <a href="/pages/ObeRivQmxJokrh8FcIfa">supported workload types</a>. See <a href="/pages/6qh7jEUwmQ2poP59Ty5Y">Using GB200 NVL72 and Multi-Node NVLink Domains</a> for more details.<br>Default: <code>false</code></td></tr><tr><td><code>clusterConfig.workload-controller.additionalPodLabels</code></td><td><code>spec.workload-controller.additionalPodLabels</code></td><td>Set additional workload's <a href="https://kubernetes.io/docs/concepts/overview/working-with-objects/labels">Pod Labels </a>in a format of key/value pairs. Default: <code>""</code></td></tr><tr><td><code>clusterConfig.mps-server.enabled</code> <em>(boolean)</em></td><td><code>spec.mps-server.enabled</code> <em>(boolean)</em></td><td>Enabled when using <a href="https://docs.nvidia.com/deploy/mps/index.html">NVIDIA MPS</a><br>Default: <code>false</code></td></tr><tr><td><code>clusterConfig.daemonSetsTolerations</code> <em>(object)</em></td><td><code>spec.daemonSetsTolerations</code> <em>(object)</em></td><td>Configure Kubernetes tolerations for NVIDIA Run:ai daemonSets / engine</td></tr><tr><td><code>clusterConfig.runai-container-toolkit.enabled</code> <em>(boolean)</em></td><td><code>spec.runai-container-toolkit.enabled</code> <em>(boolean)</em></td><td>Enables workloads to use <a href="/pages/Ap3Un7M2pNFM5QewTcuu">GPU fractions</a><br>Default: <code>true</code></td></tr><tr><td><code>clusterConfig.runai-container-toolkit.logLevel</code> <em>(boolean)</em></td><td><code>spec.runai-container-toolkit.logLevel</code> <em>(boolean)</em></td><td>Specifies the NVIDIA Run:ai-container-toolkit logging level: either 'SPAM', 'DEBUG', 'INFO', 'NOTICE', 'WARN', or 'ERROR'<br>Default: <code>INFO</code></td></tr><tr><td><code>clusterConfig.node-scale-adjuster.args.gpuMemoryToFractionRatio</code> <em>(object)</em></td><td><code>spec.node-scale-adjuster.args.gpuMemoryToFractionRatio</code> <em>(object)</em></td><td>A scaling-pod requesting a single GPU device will be created for every 1 to 10 pods requesting fractional GPU memory (1/gpuMemoryToFractionRatio). This value represents the ratio (0.1-0.9) of fractional GPU memory (any size) to GPU fraction (portion) conversion.<br>Default: <code>0.1</code></td></tr><tr><td><code>clusterConfig.global.core.dynamicFractions.enabled</code> <em>(boolean)</em></td><td><code>spec.global.core.dynamicFractions.enabled</code> <em>(boolean)</em></td><td>Enables <a href="/pages/UgIxXMHwanFJyzcVnkte">dynamic GPU fractions</a><br>Default: <code>true</code></td></tr><tr><td><code>clusterConfig.global.core.swap.enabled</code> <em>(boolean)</em></td><td><code>spec.global.core.swap.enabled</code> <em>(boolean)</em></td><td>Enables <a href="/pages/5RuFMA1wSz8EuuZsaBHE">memory swap</a> for GPU workloads<br>Default: <code>false</code></td></tr><tr><td><code>clusterConfig.global.core.swap.biDirectional</code> <em>(string)</em></td><td><code>spec.global.core.swap.biDirectional</code> <em>(string)</em></td><td>Sets the read/write memory mode of GPU memory swap to bi-directional (fully duplex). This produces higher performance (typically +80%) vs. uni-directional (simplex) read-write operations. For more details, see <a href="/pages/5RuFMA1wSz8EuuZsaBHE">GPU memory swap</a>.<br>Default: <code>false</code></td></tr><tr><td><code>clusterConfig.global.core.swap.mode</code> <em>(string)</em></td><td><code>spec.global.core.swap.mode</code> <em>(string)</em></td><td>Sets the GPU to CPU memory swap method to use UVA and optimized memory prefetch for optimized performance in some scenarios. For more details, see <a href="/pages/5RuFMA1wSz8EuuZsaBHE">GPU memory swap</a>.<br>Default: None. The parameter is not set by default. To add this parameter set <code>mode=mapped</code> .</td></tr><tr><td><code>clusterConfig.global.core.nodeScheduler.enabled</code> <em>(boolean)</em></td><td><code>spec.global.core.nodeScheduler.enabled</code> <em>(boolean)</em></td><td>Enables the <a href="/pages/t7DvA5KXoZhvW3M5S7X6">node-level scheduler</a><br>Default: <code>false</code></td></tr><tr><td><code>clusterConfig.global.core.timeSlicing.mode</code> <em>(string)</em></td><td><code>spec.global.core.timeSlicing.mode</code> <em>(string)</em></td><td><p>Sets the <a href="/pages/oaLJMYjlUN5ywE5Y67Ne">GPU time-slicing mode</a>. Possible values:</p><ul><li><code>timesharing</code> - all pods on a GPU share the GPU compute time evenly.</li><li><code>strict</code> - each pod gets an exact time slice according to its memory fraction value.</li><li><code>fair</code> - each pod gets an exact time slice according to its memory fraction value and any unused GPU compute time is split evenly between the running pods.</li></ul><p>Default: <code>timesharing</code></p></td></tr><tr><td><code>clusterConfig.runai-scheduler.args.fullHierarchyFairness</code> <em>(boolean)</em></td><td><code>spec.runai-scheduler.args.fullHierarchyFairness</code> <em>(boolean)</em></td><td>Enables fairness between departments, on top of projects fairness<br>Default: <code>true</code></td></tr><tr><td><code>clusterConfig.runai-scheduler.args.defaultStalenessGracePeriod</code></td><td><code>spec.runai-scheduler.args.defaultStalenessGracePeriod</code></td><td><p>Sets the timeout in seconds before the scheduler evicts a stale pod-group (gang) that went below its min-members in running state:</p><ul><li><code>0s</code> - Immediately (no timeout)</li><li><code>-1</code> - Never</li></ul><p>Default: <code>60s</code></p></td></tr><tr><td><code>clusterConfig.runai-scheduler.args.verbosity</code> <em>(int)</em></td><td><code>spec.runai-scheduler.args.verbosity</code> <em>(int)</em></td><td>Configures the level of detail in the logs generated by the scheduler service<br>Default: <code>4</code></td></tr><tr><td><code>clusterConfig.pod-grouper.args.gangSchedulingKnative</code> <em>(boolean)</em></td><td><code>spec.pod-grouper.args.gangSchedulingKnative</code> <em>(boolean)</em></td><td>Enables gang scheduling for inference workloads.For backward compatibility with versions earlier than v2.19, change the value to false<br>Default: <code>false</code></td></tr><tr><td><code>clusterConfig.pod-grouper.args.gangScheduleArgoWorkflow</code> <em>(boolean)</em></td><td><code>spec.pod-grouper.args.gangScheduleArgoWorkflow</code> <em>(boolean)</em></td><td>Groups all pods of a single ArgoWorkflow workload into a single Pod-Group for gang scheduling<br>Default: <code>true</code></td></tr><tr><td><code>clusterConfig.limitRange.cpuDefaultRequestCpuLimitFactorNoGpu</code> <em>(string)</em></td><td><code>spec.limitRange.cpuDefaultRequestCpuLimitFactorNoGpu</code> <em>(string)</em></td><td>Sets a default ratio between the CPU request and the limit for workloads without GPU requests<br>Default: <code>0.1</code></td></tr><tr><td><code>clusterConfig.limitRange.memoryDefaultRequestMemoryLimitFactorNoGpu</code> <em>(string)</em></td><td><code>spec.limitRange.memoryDefaultRequestMemoryLimitFactorNoGpu</code> <em>(string)</em></td><td>Sets a default ratio between the memory request and the limit for workloads without GPU requests<br>Default: <code>0.1</code></td></tr><tr><td><code>clusterConfig.limitRange.cpuDefaultRequestGpuFactor</code> <em>(string)</em></td><td><code>spec.limitRange.cpuDefaultRequestGpuFactor</code> <em>(string)</em></td><td>Sets a default amount of CPU allocated per GPU when the CPU is not specified<br>Default: <code>100</code></td></tr><tr><td><code>clusterConfig.limitRange.cpuDefaultLimitGpuFactor</code> <em>(int)</em></td><td><code>spec.limitRange.cpuDefaultLimitGpuFactor</code> <em>(int)</em></td><td>Sets a default CPU limit based on the number of GPUs requested when no CPU limit is specified<br>Default: <code>NO DEFAULT</code></td></tr><tr><td><code>clusterConfig.limitRange.memoryDefaultRequestGpuFactor</code> <em>(string)</em></td><td><code>spec.limitRange.memoryDefaultRequestGpuFactor</code> <em>(string)</em></td><td>Sets a default amount of memory allocated per GPU when the memory is not specified<br>Default: <code>100Mi</code></td></tr><tr><td><code>clusterConfig.limitRange.memoryDefaultLimitGpuFactor</code> <em>(string)</em></td><td><code>spec.limitRange.memoryDefaultLimitGpuFactor</code> <em>(string)</em></td><td>Sets a default memory limit based on the number of GPUs requested when no memory limit is specified<br>Default: <code>NO DEFAULT</code></td></tr></tbody></table>

### NVIDIA Run:ai Services Resource Management

NVIDIA Run:ai cluster includes many different services. To simplify resource management, the configuration structure allows you to configure the containers CPU / memory resources for each service individually or group of services together.

{% hint style="info" %}
**Note**

For resource recommendations, see [Vertical scaling](/saas/infrastructure-setup/procedures/scaling#vertical-scaling). To automatically adjust resources based on actual usage, see [Vertical Pod Autoscaling (VPA)](/saas/infrastructure-setup/advanced-setup/vpa). Note that static resource configuration remains the primary mechanism. VPA works alongside it by generating recommendations or dynamically applying adjustments depending on the configured update mode.
{% endhint %}

| Service Group      | Description                                                                                                      | NVIDIA Run:ai containers                                                                                                                                                                             |
| ------------------ | ---------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| SchedulingServices | Containers associated with the NVIDIA Run:ai Scheduler                                                           | Scheduler, StatusUpdater, MetricsExporter, PodGrouper, PodGroupAssigner, PodGroupController, QueueController, NodePoolController, Binder, DevicePlugin                                               |
| SyncServices       | Containers associated with syncing updates between the NVIDIA Run:ai cluster and the NVIDIA Run:ai control plane | Agent, ClusterSync, AssetsSync                                                                                                                                                                       |
| WorkloadServices   | Containers associated with submitting NVIDIA Run:ai workloads                                                    | WorkloadController, JobController, WorkloadOverseer, ExternalWorkloadIntegrator, ClusterRedis, ClusterAPI, InferenceWorkloadController, ResearcherService, SharedObjectsController, WorkloadExporter |

Apply the following configuration in order to change resources request and limit for a group of services:

{% tabs %}
{% tab title="Helm" %}

```yaml
clusterConfig:
  global:
   <service-group-name>: # schedulingServices | SyncServices | WorkloadServices
     resources:
       limits:
         cpu: 1000m
         memory: 1Gi
       requests:
         cpu: 100m
         memory: 512Mi
```

{% endtab %}

{% tab title="runaiconfig" %}

```yaml
spec:
  global:
   <service-group-name>: # schedulingServices | syncServices | workloadServices
     resources:
       limits:
         cpu: 1000m
         memory: 1Gi
       requests:
         cpu: 100m
         memory: 512Mi
```

{% endtab %}
{% endtabs %}

Or, apply the following configuration in order to change resources request and limit for each service individually:

{% tabs %}
{% tab title="Helm" %}

```yaml
clusterConfig:
  <service-name>: # for example: pod-grouper
    resources:
      limits:
        cpu: 1000m
        memory: 1Gi
      requests:
        cpu: 100m
        memory: 512Mi
```

{% endtab %}

{% tab title="runaiconfig" %}

```yaml
spec:
  <service-name>: # for example: pod-grouper
    resources:
      limits:
        cpu: 1000m
        memory: 1Gi
      requests:
        cpu: 100m
        memory: 512Mi
```

{% endtab %}
{% endtabs %}

### NVIDIA Run:ai Services Replicas

By default, all NVIDIA Run:ai containers are deployed with a single replica. Some services support multiple replicas for redundancy and performance.

To simplify configuring replicas, a global replicas configuration can be set and is applied to all supported services:

{% tabs %}
{% tab title="Helm" %}

```yaml
clusterConfig:
  global: 
    replicaCount: 1 # default
```

{% endtab %}

{% tab title="runaiconfig" %}

```yaml
spec:
  global: 
    replicaCount: 1 # default
```

{% endtab %}
{% endtabs %}

This can be overwritten for specific services (if supported). Services without the `replicas` configuration does not support replicas:

{% tabs %}
{% tab title="Helm" %}

<pre class="language-yaml"><code class="lang-yaml"><strong>clusterConfig:
</strong>  &#x3C;service-name>: # for example: pod-grouper
    replicas: 1 # default
</code></pre>

{% endtab %}

{% tab title="runaiconfig" %}

<pre class="language-yaml"><code class="lang-yaml"><strong>spec:
</strong>  &#x3C;service-name>: # for example: pod-grouper
    replicas: 1 # default
</code></pre>

{% endtab %}
{% endtabs %}

### Prometheus

NVIDIA Run:ai uses two separate Prometheus instances:

* The Prometheus instance used for metrics collection and alerting.
* A dedicated Prometheus instance used the NVIDIA Run:ai Scheduler for tracking historical resource usage and enabling time-based fairshare.

#### Metrics Collection

The Prometheus instance in NVIDIA Run:ai is used for metrics collection and alerting.

The configuration scheme follows the official [PrometheusSpec](https://prometheus-operator.dev/docs/api-reference/api/#monitoring.coreos.com/v1.PrometheusSpec) and supports additional custom configurations. The PrometheusSpec schema is available using the `spec.prometheus.spec` configuration.

A common use case using the PrometheusSpec is for metrics retention. This prevents metrics loss during potential connectivity issues and can be achieved by configuring local temporary metrics retention. For more information, see [Prometheus Storage](https://prometheus.io/docs/prometheus/latest/storage/#storage):

{% tabs %}
{% tab title="Helm" %}

```yaml
clusterConfig:  
  prometheus:
    spec: # PrometheusSpec
      retention: 2h # default 
      retentionSize: 20GB
```

{% endtab %}

{% tab title="runaiconfig" %}

```yaml
spec:  
  prometheus:
    spec: # PrometheusSpec
      retention: 2h # default 
      retentionSize: 20GB
```

{% endtab %}
{% endtabs %}

In addition to the PrometheusSpec schema, some custom NVIDIA Run:ai configurations are also available:

* Additional labels - Set additional labels for NVIDIA Run:ai's [built-in alerts](/saas/infrastructure-setup/procedures/system-monitoring#built-in-alerts) sent by Prometheus.
* Log level configuration - Configure the `logLevel` setting for the Prometheus container.
* Image override - Use `prometheus.spec.image` to manually specify the Prometheus image reference. Due to a known issue, the `imageRegistry` setting in the Prometheus Helm chart is ignored. To pull the image from a different registry, specify the full image reference. Default: `quay.io/prometheus/prometheus`.
* Image pull secrets - Use `prometheus.spec.imagePullSecrets` to list Kubernetes image pull secrets in the `runai` namespace. This is particularly relevant for air-gapped installations where pulling Prometheus images requires authentication. Default: `[]`.
* Advanced metrics (GPU profiling metrics) - Use `prometheus.spec.config.advancedMetricsEnabled` to activate GPU profiling metrics from NVIDIA DCGM. When enabled, Prometheus collects and aggregates advanced GPU performance data such as SM activity, memory bandwidth, and tensor core utilization. For setup instructions, see [GPU profiling metrics](/saas/platform-management/monitor-performance/gpu-profiling-metrics).
* OTLP receiver - Use `prometheus.config.enableOTLPReceiver` to enable or disable the OpenTelemetry Protocol (OTLP) receiver on the NVIDIA Run:ai Prometheus instance. When enabled, Knative Serving can push inference request metrics to Prometheus over OTLP, surfacing them in the NVIDIA Run:ai API and UI. Default: `true`.

{% tabs %}
{% tab title="Helm" %}

```yaml
clusterConfig:  
  prometheus:
    logLevel: info # debug | info | warn | error
    additionalAlertLabels:
      - env: prod # example
    config:
      enableOTLPReceiver: true # default: true
```

{% endtab %}

{% tab title="runaiconfig" %}

```yaml
spec:  
  prometheus:
    logLevel: info # debug | info | warn | error
    additionalAlertLabels:
      - env: prod # example
    config:
      enableOTLPReceiver: true # default: true
```

{% endtab %}
{% endtabs %}

#### Time-Based Fairshare

Time-based fairshare relies on historical resource usage metrics collected by a dedicated Prometheus instance. By default, persistent storage is not configured for this Prometheus instance. As a result, if the Prometheus pod restarts or fails, all historical usage data is lost. In this scenario, the Scheduler recalculates fairshare from scratch and makes scheduling decisions without taking previously collected data into account until new data is accumulated.

To preserve historical usage data across restarts, you can enable persistent storage for the Prometheus instance used for time-based fairshare. It is recommended to use network-based storage that is accessible from all nodes in the cluster.

To enable persistent storage, configure the following cluster settings:

{% tabs %}
{% tab title="Helm" %}

```bash
clusterConfig:
  historical-usage-prometheus:
    enablePersistentStorage: true #default false
    storageSize: 50Gi # default 50Gi
    storageClassName: fast-ssd # defaults to the default storage class configured in the cluster  
```

{% endtab %}

{% tab title="runaiconfig" %}

```bash
spec:
  historical-usage-prometheus:
    enablePersistentStorage: true #default false
    storageSize: 50Gi # default 50Gi
    storageClassName: fast-ssd # defaults to the default storage class configured in the cluster
```

{% endtab %}
{% endtabs %}

### NVIDIA Run:ai Managed Nodes

To include or exclude specific nodes from running workloads within a cluster managed by NVIDIA Run:ai, use the `nodeSelectorTerms` flag. For additional details, see [Kubernetes nodeSelector](https://kubernetes.io/docs/concepts/scheduling-eviction/assign-pod-node/#nodeselector).

Label the nodes using the below:

* key - Label key (e.g., zone, instance-type).
* operator - Operator defining the inclusion/exclusion condition (In, NotIn, Exists, DoesNotExist).
* values - List of values for the key when using In or NotIn.

The below example shows how to include NVIDIA GPUs only and exclude all other GPU types in a cluster with mixed nodes, based on product type GPU label:

{% tabs %}
{% tab title="Helm" %}

```yaml
clusterConfig:   
  global:
     managedNodes:
       inclusionCriteria:
          nodeSelectorTerms:
          - matchExpressions:
            - key: nvidia.com/gpu.product  
              operator: Exists
```

{% endtab %}

{% tab title="runaiconfig" %}

```yaml
spec:   
  global:
     managedNodes:
       inclusionCriteria:
          nodeSelectorTerms:
          - matchExpressions:
            - key: nvidia.com/gpu.product  
              operator: Exists
```

{% endtab %}
{% endtabs %}

### Custom Certificate Authority for Git and S3

To override the default global CA used by the system and inject a custom CA certificate for Git or S3 data sources, follow these steps:

1. Create a Kubernetes secret with the custom CA certificate:

   ```bash
   kubectl -n runai create secret generic runai-cluster-git-ca \
       --from-file=runai-ca-git.pem=<ca_bundle_path>
   kubectl label secret runai-cluster-git-ca -n runai run.ai/cluster-wide=true run.ai/name=runai-ca-cert --overwrite
   ```
2. When installing the cluster, make sure the following flags are added to the helm command. See [Install cluster](https://github.com/run-ai/runai-product-docs/blob/SaaS/getting-started/installation-1/helm-install.md).

{% tabs %}
{% tab title="Git" %}

```bash
--set global.customCAGit.enabled=true \
--set global.customCAGit.secret.name=<secret-name>
```

{% endtab %}

{% tab title="S3" %}

```bash
--set global.customCAS3.enabled=true \
--set global.customCAS3.secret.name=<secret-name>
```

{% endtab %}
{% endtabs %}


# Vertical Pod Autoscaling (VPA)

Vertical Pod Autoscaling (VPA) automatically adjusts CPU and memory requests and limits for pods based on observed usage. Unlike horizontal scaling which adds replicas, vertical scaling right-sizes existing pods. NVIDIA Run:ai supports configuring VPA on cluster-level services to ensure they have adequate resources as your environment grows. For more information, see [Kubernetes' documentation](https://kubernetes.io/docs/concepts/workloads/autoscaling/vertical-pod-autoscale/).

As clusters scale in nodes and workloads, NVIDIA Run:ai services may benefit from resource adjustments beyond the static defaults. VPA helps identify services that are over-provisioned or under-provisioned and can automatically apply right-sized resource recommendations - reducing the need for manual tuning.

{% hint style="info" %}
**Note**

VPA can be configured on NVIDIA Run:ai cluster components only, not the control plane. For more information about NVIDIA Run:ai service groups and manual scaling recommendations, see [NVIDIA Run:ai at scale](/saas/infrastructure-setup/procedures/scaling).
{% endhint %}

### Update Modes

VPA can operate in the following modes:

| Mode                | Description                                                                                                                                       |
| ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- |
| `Off`               | VPA only generates recommendations and does not apply them. Useful for safely observing suggested values before enabling automatic scaling.       |
| `Initial`           | VPA applies recommendations only when pods are created. Existing pods are not modified.                                                           |
| `Auto`              | VPA applies recommendations by evicting and recreating pods when needed. This may cause restarts.                                                 |
| `InPlaceOrRecreate` | VPA updates resources in-place without restarting pods when supported. If in-place updates are not possible, it falls back to recreating the pod. |

{% hint style="warning" %}
**Important**

VPA does not replace static resource configuration. Your existing static resource settings remain active and in effect regardless of the VPA mode. For static resource configuration, see [NVIDIA Run:ai services resource management](/saas/infrastructure-setup/advanced-setup/cluster-config#nvidia-run-ai-services-resource-management).
{% endhint %}

## Installing VPA

VPA is an external component that does not come preinstalled in most Kubernetes clusters. If it is not installed in your cluster, run the following:

```bash
git clone https://github.com/kubernetes/autoscaler.git
cd autoscaler/vertical-pod-autoscaler
./hack/vpa-up.sh
```

## Configuring VPA

VPA is configured via the `runaiconfig` object and supports multiple levels of configuration. This allows you to define defaults globally, override them per scaling group, and fine-tune behavior for specific components.

The configuration hierarchy is:

1. **Global** - applies to all services
2. **Scaling group** - overrides global for a group of services
3. **Component** - overrides both group and global

When the same service is configured at multiple levels, the most specific level takes precedence. Component-level configuration overrides scaling group, and scaling group overrides global.

VPA is disabled by default. The syntax for configuring VPA follows [Kubernetes' documentation](https://kubernetes.io/docs/concepts/workloads/autoscaling/vertical-pod-autoscale/#resource-policies). There is no need to set `targetRef`, as that is already configured by NVIDIA Run:ai.

### Examples and Recommendations

NVIDIA Run:ai recommends starting with a **global** configuration in **Off** mode, which generates recommendations in the background without applying any changes. After running for 1-2 days, review VPA's resource suggestions. If it recommends higher resources than expected, you can set VPA to `InPlaceOrRecreate` for that specific component to dynamically apply the suggestions.

#### Global Configuration

Use global configuration to define a default VPA policy for all NVIDIA Run:ai cluster services.

{% tabs %}
{% tab title="Helm" %}

```yaml
clusterConfig:
  global:
    vpa:
      enabled: true
      updatePolicy:
        updateMode: "Off"
      resourcePolicy:
        containerPolicies:
          - containerName: "*"
            minAllowed:
              cpu: "100m"
              memory: "128Mi"
            maxAllowed:
              cpu: "4"
              memory: "8Gi"
```

{% endtab %}

{% tab title="runaiconfig" %}

```yaml
spec:
  global:
    vpa:
      enabled: true
      updatePolicy:
        updateMode: "Off"
      resourcePolicy:
        containerPolicies:
          - containerName: "*"
            minAllowed:
              cpu: "100m"
              memory: "128Mi"
            maxAllowed:
              cpu: "4"
              memory: "8Gi"
```

{% endtab %}
{% endtabs %}

To inspect VPA objects and their recommendations, run:

```bash
kubectl get vpa -n runai
```

#### Scaling Group Configuration

Use scaling group configuration to override the global setting for one of the built-in service groups: `workloadServices`, `syncServices`, or `schedulingServices`. For a description of each group and sizing recommendations, see [NVIDIA Run:ai services resource management](/saas/infrastructure-setup/advanced-setup/cluster-config#nvidia-run-ai-services-resource-management) and [NVIDIA Run:ai at scale](/saas/infrastructure-setup/procedures/scaling).

{% tabs %}
{% tab title="Helm" %}

```yaml
clusterConfig:
  global:
    workloadServices: # Replace with the relevant service group
      vpa:
        enabled: true
        updatePolicy:
          updateMode: "Auto"
        resourcePolicy:
          containerPolicies:
            - containerName: "*"
              minAllowed:
                cpu: "100m"
                memory: "128Mi"
              maxAllowed:
                memory: "8Gi"
```

{% endtab %}

{% tab title="runaiconfig" %}

```yaml
spec:
  global:
    workloadServices: # Replace with the relevant service group
      vpa:
        enabled: true
        updatePolicy:
          updateMode: "Auto"
        resourcePolicy:
          containerPolicies:
            - containerName: "*"
              minAllowed:
                cpu: "100m"
                memory: "128Mi"
              maxAllowed:
                memory: "8Gi"
```

{% endtab %}
{% endtabs %}

#### Component Configuration

Use component configuration to override both global and scaling group settings for a specific service.

{% tabs %}
{% tab title="Helm" %}

<pre class="language-yaml"><code class="lang-yaml">clusterConfig:
  workload-controller: # Replace with the relevant component name
    vpa:
      enabled: true
      updatePolicy:
        updateMode: "Off"
      resourcePolicy:
        containerPolicies:
          - containerName: "workload-controller"
            minAllowed:
              cpu: "100m"
              memory: "128Mi"
<strong>            maxAllowed:
</strong>              cpu: "2"
              memory: "4Gi"
</code></pre>

{% endtab %}

{% tab title="runaiconfig" %}

```yaml
spec:
  workload-controller: # Replace with the relevant component name
    vpa:
      enabled: true
      updatePolicy:
        updateMode: "Off"
      resourcePolicy:
        containerPolicies:
          - containerName: "workload-controller"
            minAllowed:
              cpu: "100m"
              memory: "128Mi"
            maxAllowed:
              cpu: "2"
              memory: "4Gi"
```

{% endtab %}
{% endtabs %}


# Kubernetes Gateway API

NVIDIA Run:ai supports the Kubernetes Gateway API as an alternative to Ingress for routing external traffic. Gateway API provides a flexible and extensible model for defining how traffic is exposed and routed within the cluster.

This page builds on the concepts described in [Routing Traffic to and from NVIDIA Run:ai Services](/saas/getting-started/installation/install-using-helm/system-requirements#routing-traffic-to-and-from-nvidia-run-ai-services), including FQDN configuration, TLS certificates, and routing behavior.

{% hint style="info" %}
**Note**

Gateway API support in NVIDIA Run:ai is **optional**. If you are currently using HAProxy Ingress or other ingress controllers, no action is required. Customers who wish to adopt Gateway API may do so using the instructions on this page.
{% endhint %}

## Scope and Prerequisites

Before proceeding, ensure that the following prerequisites are already configured:

* Fully Qualified Domain Names (FQDN) for:
  * Control plane access
  * Development workspaces and training workloads
  * Inference workloads
* TLS certificates associated with the configured FQDNs

These configurations are described in the [Routing Traffic to and from NVIDIA Run:ai Services](/saas/getting-started/installation/install-using-helm/system-requirements#routing-traffic-to-and-from-nvidia-run-ai-services) section.

## Routing Modes in NVIDIA Run:ai

NVIDIA Run:ai supports two routing approaches for exposing services: **host-based routing** and **path-based routing**.

Different NVIDIA Run:ai services use these routing approaches as follows:

| Service                         | Routing Mode                     | Example                                          |
| ------------------------------- | -------------------------------- | ------------------------------------------------ |
| Control plane                   | Host-based (single domain)       | `https://runai.mycorp.local`                     |
| Inference workloads             | Host-based (wildcard subdomains) | `https://<service>.runai-inference.mycorp.local` |
| Workspaces & training workloads | Host-based or path-based         | See section below                                |

## Installing a Gateway Controller

NVIDIA Run:ai supports any [conformant Gateway API implementation](https://gateway-api.sigs.k8s.io/implementations/). The example below uses **KGateway**. If you are using a different conformant controller, follow its installation documentation and then proceed to the migration steps.

1. Install the Gateway API CRDs:

   ```bash
   kubectl apply -f https://github.com/kubernetes-sigs/gateway-api/releases/download/v1.4.0/standard-install.yaml
   ```
2. Install the KGateway CRDs:

   ```bash
   helm upgrade -i --create-namespace \
     --namespace kgateway-system \
     --version v2.2.1 \
     kgateway-crds oci://cr.kgateway.dev/kgateway-dev/charts/kgateway-crds
   ```
3. Install the KGateway controller:

   ```bash
   helm upgrade -i -n kgateway-system kgateway \
     oci://cr.kgateway.dev/kgateway-dev/charts/kgateway \
     --version v2.2.1
   ```

Ensure that the Gateway controller is running before proceeding.

## Routing Traffic for Workspaces and Training Workloads

The following sections describe how development workspaces and training workloads are exposed using Gateway API.

Depending on the selected routing mode, these workloads are accessed differently:

* In **host-based routing**, each development workspace and training workload is exposed using its own subdomain:

  ```bash
  https://<project>-<workload>.<CLUSTER_URL>
  ```
* In **path-based routing**, development workspaces and training workloads are exposed under a shared domain using URL paths:

  ```bash
  https://runai.mycorp.local/<project>/<workload>
  ```

This distinction determines the required FQDN structure, TLS certificates, and Gateway configuration described in the sections below.

## Gateway API with Host-Based Routing

In this configuration:

* Development workspaces and training workloads are exposed using subdomains
* A wildcard FQDN is required (e.g., `*.runai.mycorp.local`)
* A wildcard TLS certificate is required for those workloads
* The Gateway includes listeners for:
  * The cluster domain (control plane access)
  * Wildcard subdomains (workspace and training workloads)

### Gateway Configuration

{% hint style="info" %}
**Note**

Routing is based on hostnames (subdomains). HTTPRoute resources for the control plane and workloads are created automatically by NVIDIA Run:ai and bind to the configured Gateway.
{% endhint %}

```yaml
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
  name: runai-gateway
  namespace: runai
spec:
  gatewayClassName: kgateway
  listeners:
  - name: https-cluster-domain
    protocol: HTTPS
    port: 443
    hostname: "<CLUSTER_DOMAIN>"
    tls:
      mode: Terminate
      certificateRefs:
      - kind: Secret
        name: runai-cluster-domain-tls-secret

  - name: https-workloads
    protocol: HTTPS
    port: 443
    hostname: "*.<CLUSTER_DOMAIN>"
    tls:
      mode: Terminate
      certificateRefs:
      - kind: Secret
        name: runai-cluster-domain-star-tls-secret
```

## Gateway API with Path-Based Routing

In this configuration:

* Development workspaces and training workloads are exposed under a single domain using URL paths
* A wildcard FQDN is **not required** for workspace and training workloads
* A wildcard TLS certificate is **not required** for workspace and training workloads
* The Gateway uses a single domain listener for these services

Ensure that host-based routing is disabled using the following. For more details, see [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config).

```
clusterConfig.global.subdomainSupport: false
```

### Gateway Configuration

{% hint style="info" %}
**Note**

Routing is based on URL paths under a shared domain. HTTPRoute resources for the control plane and workloads are created automatically by NVIDIA Run:ai and bind to the configured Gateway.
{% endhint %}

```yaml
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
  name: runai-gateway
  namespace: runai
spec:
  gatewayClassName: kgateway
  listeners:
  - name: https-cluster-domain
    protocol: HTTPS
    port: 443
    hostname: "<CLUSTER_DOMAIN>"
    tls:
      mode: Terminate
      certificateRefs:
      - kind: Secret
        name: runai-cluster-domain-tls-secret
```

## Inference Routing with Gateway API

Inference workloads are always exposed using **host-based routing**, regardless of the selected routing mode for development workspaces and training workloads.

Inference endpoints are accessed via dedicated subdomains, for example:

```bash
https://<service>.runai-inference.mycorp.local
```

To enable inference routing with Gateway API:

* A wildcard FQDN must be configured for inference (e.g., `*.runai-inference.mycorp.local`)
* A wildcard TLS certificate must be configured for that domain
* The Gateway must include a listener for the inference subdomain
* Knative Serving must be configured to route inference traffic from the Gateway to inference workloads

### TLS Certificate

Ensure that a wildcard TLS certificate for inference is created:

* Secret name: `runai-cluster-inference-tls-secret`
* Namespace: `runai`

Replace `/path/to/inference-fullchain.pem` and `/path/to/inference-private.pem` with the actual paths to your certificate and private key:

```bash
kubectl create secret tls runai-cluster-inference-tls-secret -n runai \
    --cert /path/to/inference-fullchain.pem \
    --key /path/to/inference-private.pem
```

### Gateway Configuration

{% hint style="info" %}
**Note**

Routing to inference workloads is based on hostnames (subdomains).
{% endhint %}

Add an inference listener to the existing Gateway resource:

```yaml
- name: https-inference
  protocol: HTTPS
  port: 443
  hostname: "*.runai-inference.<CLUSTER_DOMAIN>"
  tls:
    mode: Terminate
    certificateRefs:
    - kind: Secret
      name: runai-cluster-inference-tls-secret
```

### Configure Knative Serving

NVIDIA Run:ai supports Knative-based inference workloads. Inference traffic arrives at the Gateway and is forwarded to Knative Serving through Kourier, which routes it to the individual inference endpoints. The following steps install Knative Serving and configure the HTTPRoute to connect the Gateway to Kourier. Knative versions 1.19 to 1.21 are supported.

{% tabs %}
{% tab title="Kubernetes" %}

1. Install Knative Serving. Follow the [Installing Knative](https://knative.dev/docs/install/operator/knative-with-operators/) instructions or run:

   ```bash
   helm repo add knative-operator https://knative.github.io/operator
   helm install knative-operator --create-namespace --namespace knativeoperator --version 1.18.2 knative-operator/knative-operator
   ```
2. Create the `knative-serving` namespace:

   ```bash
   kubectl create ns knative-serving
   ```
3. Create a YAML file named `knative-serving.yaml` and replace the placeholder FQDN with your wildcard [inference FQDN](/saas/getting-started/installation/install-using-helm/system-requirements#fully-qualified-domain-name-fqdn) (for example, `runai-inference.mycorp.local`):

   ```yaml
   apiVersion: operator.knative.dev/v1beta1
   kind: KnativeServing
   metadata:
     name: knative-serving
     namespace: knative-serving
   spec:
     config:
       config-autoscaler:
         enable-scale-to-zero: "true"
       config-features:
         kubernetes.podspec-affinity: enabled
         kubernetes.podspec-init-containers: enabled
         kubernetes.podspec-persistent-volume-claim: enabled
         kubernetes.podspec-persistent-volume-write: enabled
         kubernetes.podspec-schedulername: enabled
         kubernetes.podspec-securitycontext: enabled
         kubernetes.podspec-tolerations: enabled
         kubernetes.podspec-volumes-emptydir: enabled
         kubernetes.podspec-fieldref: enabled
         kubernetes.containerspec-addcapabilities: enabled
         kubernetes.podspec-nodeselector: enabled
         multi-container: enabled
         kubernetes.podspec-hostipc: enabled
         kubernetes.podspec-hostnetwork: enabled
       domain:
         runai-inference.mycorp.local: "" # replace with the wildcard FQDN for Inference
       network:
         domainTemplate: '{{.Name}}-{{.Namespace}}.{{.Domain}}'
         ingress-class: kourier.ingress.networking.knative.dev
         default-external-scheme: https
     high-availability:
       replicas: 2
     ingress:
       kourier:
         enabled: true
   ```
4. Apply the changes:

   ```bash
   kubectl apply -f knative-serving.yaml
   ```
5. Create a YAML file named `knative-httproute.yaml` to route inference traffic from the Gateway to the Kourier service. Replace the FQDN placeholder with your wildcard [inference FQDN](/saas/getting-started/installation/install-using-helm/system-requirements#fully-qualified-domain-name-fqdn):

   ```yaml
   apiVersion: gateway.networking.k8s.io/v1
   kind: HTTPRoute
   metadata:
     name: knative-serving-httproute
     namespace: knative-serving
   spec:
     parentRefs:
       - name: runai-gateway
         namespace: runai
     hostnames:
       - "*.runai-inference.mycorp.local" # replace with the wildcard FQDN for Inference
     rules:
       - matches:
         - path:
             type: PathPrefix
             value: /
         backendRefs:
           - name: kourier
             port: 80
   ```
6. Apply the changes:

   ```bash
   kubectl apply -f knative-httproute.yaml
   ```

{% endtab %}

{% tab title="OpenShift" %}

1. Install the OpenShift Serverless Operator. Follow the [Installing the OpenShift Serverless Operator](https://docs.redhat.com/en/documentation/red_hat_openshift_serverless/1.37/html/installing_openshift_serverless/install-serverless-operator) instructions. Once installed, follow the steps below.
2. Create the `knative-serving` project:

   ```bash
   oc new-project knative-serving
   ```
3. Create a YAML file named `knative-serving.yaml`:

   ```yaml
   apiVersion: operator.knative.dev/v1beta1
   kind: KnativeServing
   metadata:
     finalizers:
       - knative-serving-openshift
       - knativeservings.operator.knative.dev
     name: knative-serving
     namespace: knative-serving
   spec:
     config:
       config-features:
         kubernetes.podspec-tolerations: enabled
         kubernetes.podspec-volumes-emptydir: enabled
         kubernetes.podspec-persistent-volume-claim: enabled
         multi-container: enabled
         kubernetes.podspec-persistent-volume-write: enabled
         kubernetes.podspec-fieldref: enabled
         kubernetes.podspec-schedulername: enabled
         kubernetes.podspec-nodeselector: enabled
         kubernetes.podspec-init-containers: enabled
         kubernetes.podspec-securitycontext: enabled
         kubernetes.podspec-affinity: enabled
         kubernetes.containerspec-addcapabilities: enabled
     controller-custom-certs:
       name: ''
       type: ''
     registry: {}
   ```
4. Apply the changes:

   ```bash
   oc apply -f knative-serving.yaml
   ```
5. Create a YAML file named `knative-httproute.yaml` to route inference traffic from the Gateway to the Kourier service. Replace the FQDN placeholder with your wildcard [inference FQDN](/saas/getting-started/installation/install-using-helm/system-requirements#fully-qualified-domain-name-fqdn):

   <div data-gb-custom-block data-tag="hint" data-style="info" class="hint hint-info"><p><strong>Note</strong></p><p>OpenShift's Ingress Operator automatically creates DNS records for Gateway listeners, so no manual DNS configuration is required for the inference subdomain.</p></div>

   ```yaml
   apiVersion: gateway.networking.k8s.io/v1
   kind: HTTPRoute
   metadata:
     name: knative-serving-httproute
     namespace: knative-serving
   spec:
     parentRefs:
       - name: openshift-default
         namespace: openshift-ingress
     hostnames:
       - "*.runai-inference.mycorp.local" # replace with the wildcard FQDN for Inference
     rules:
       - matches:
         - path:
             type: PathPrefix
             value: /
         backendRefs:
           - name: kourier
             port: 80
   ```
6. Apply the changes:

   ```bash
   oc apply -f knative-httproute.yaml
   ```

{% endtab %}
{% endtabs %}

## Introduce the Gateway API to NVIDIA Run:ai Services

To enable NVIDIA Run:ai to route traffic through the configured Gateway, both the control plane and the cluster must be updated to reference the Gateway.

At this stage, traffic is still served through the existing Ingress. The Gateway is introduced alongside the current setup and will become active only after DNS is updated in a later step.

### Configure the NVIDIA Run:ai Cluster

```bash
helm upgrade runai-cluster -n runai runai/runai-cluster \
  --reuse-values \
  --set-string clusterConfig.global.gatewayAPI.enabled=true \
  --set-string clusterConfig.global.gatewayAPI.name=<gateway-name> \
  --set-string clusterConfig.global.gatewayAPI.namespace=<gateway-namespace>
```

## Verify Gateway Configuration

```bash
kubectl get gateway -n runai
```

```bash
kubectl get httproute -A
```

Test connectivity:

```bash
curl -H "Host: <CLUSTER_DOMAIN>" https://<gateway-ip>/
```

## Switch Traffic to Gateway

Before switching traffic, ensure that the NVIDIA Run:ai control plane and cluster are already configured to use the Gateway API.

Update DNS:

* `<CLUSTER_DOMAIN>` -> Gateway IP
* `*.<CLUSTER_DOMAIN>` -> Gateway IP

Traffic is routed based on DNS configuration. Until DNS records are updated to point to the Gateway IP, all traffic continues to be served through the existing Ingress.

```bash
nslookup <CLUSTER_DOMAIN>
```

```bash
curl https://<CLUSTER_DOMAIN>/
```

Disable Ingress on the cluster:

```bash
helm upgrade runai-cluster -n runai runai/runai-cluster \
  --reuse-values \
  --set-string clusterConfig.global.ingress.enabled=false
```

## Rollback to Ingress

To roll back to Ingress, re-enable Ingress and disable Gateway API on both the cluster and the control plane. Then update DNS records to point back to the Ingress IP.

```bash
helm upgrade runai-cluster -n runai runai/runai-cluster \
  --reuse-values \
  --set-string clusterConfig.global.ingress.enabled=true \
  --set-string clusterConfig.global.gatewayAPI.enabled=false
```

Update DNS records for `<CLUSTER_DOMAIN>` and `*.<CLUSTER_DOMAIN>` to point back to the Ingress IP.


# Container Access


# External Access to Containers

Researchers may need to access containers remotely during workload execution. Common use cases include:

* Running a Jupyter Notebook inside the container
* Connecting PyCharm for remote Python development
* Viewing machine learning visualizations using TensorBoard

To enable this access, you must expose the relevant container ports.

## Exposing Container Ports

Accessing the containers remotely requires exposing container ports. In Docker, ports are exposed by [declaring](https://docs.docker.com/reference/cli/docker/container/run/) them when launching the container. NVIDIA Run:ai provides similar functionality within a Kubernetes environment.

Since Kubernetes abstracts the container's physical location, exposing ports is more complex. Kubernetes supports multiple methods for exposing container ports. For more details, refer to the [Kubernetes services and networking documentation](https://kubernetes.io/docs/concepts/services-networking/service).

<table><thead><tr><th width="155.71875">Method</th><th width="370.2734375">Description</th><th width="288.94921875">NVIDIA Run:ai Support</th></tr></thead><tbody><tr><td>Port Forwarding</td><td>Simple port forwarding allows access to the container via local and/or remote port.</td><td>Supported natively via Kubernetes</td></tr><tr><td>NodePort</td><td>Exposes the service on each Node’s IP at a static port (the NodePort). You’ll be able to contact the NodePort service from outside the cluster by requesting <code>&#x3C;NODE-IP>:&#x3C;NODE-PORT></code> regardless of which node the container actually resides in.</td><td>Supported</td></tr><tr><td>LoadBalancer</td><td>Exposes the service externally using a cloud provider’s load balancer.</td><td>Supported</td></tr></tbody></table>

## Access to the Running Workload's Container

Many tools used by researchers, such as Jupyter, TensorBoard, or VSCode, require remote access to the running workload's container. In NVIDIA Run:ai, this access is provided through dynamically generated URLs. NVIDIA Run:ai supports two routing models for exposing workloads:

* Host-based routing (default)
* Path-based routing

### Host-Based Routing

By default, NVIDIA Run:ai uses host-based routing to expose workload URLs using subdomains. This allows all workloads to run at the root path, avoiding file path issues and ensuring proper application behavior.

```bash
https://project-name-workload-name.<CLUSTER_URL>/
```

Host-based routing is enabled by default via the `subdomainSupport: true` cluster configuration. For setup instructions, see [Host-based routing](/saas/getting-started/installation/install-using-helm/system-requirements#host-based-routing-default) in the system requirements.

### Path-Based Routing

Path-based routing uses the [Cluster URL](/saas/getting-started/installation/install-using-helm/system-requirements#fully-qualified-domain-name-fqdn) provided to dynamically create SSL-secured URLs in the following format:

```bash
https://<CLUSTER_URL>/project-name/workload-name
```

To use path-based routing instead, disable host-based routing by setting `subdomainSupport: false` in [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config).

While path-based routing works with applications such as Jupyter Notebooks, it may not be compatible with other applications. Some applications assume they are running at the root file system, so hardcoded file paths and settings within the container may become invalid when running at a path other than the root. For example, if an application expects to access `/etc/config.json` but is served at `/project-name/workspace-name`, the file will not be found. This can cause the container to fail or not function as intended.

{% hint style="info" %}
**Note**

For existing clusters, changing the routing model affects how workload URLs are generated. Administrators should plan this transition accordingly.
{% endhint %}


# User Identity in Containers

The identity of the user inside a container determines its access to various resources. For example, network file systems often rely on this identity to control access to mounted volumes. As a result, propagating the correct user identity into a container is crucial for both functionality and security.

By default, containers in both Docker and Kubernetes run as the `root` user. This means any process inside the container has full administrative privileges, capable of modifying system files, installing packages, or changing configurations.

While this level of access provides researchers with maximum flexibility, it conflicts with modern enterprise security practices. If the container’s root identity is propagated to external systems (e.g., network-attached storage), it can result in elevated permissions outside the container, increasing the risk of security breaches.

## NVIDIA Run:ai Controls for User Identity and Privileges

NVIDIA Run:ai allows you to enhance security and enforce organizational policies by:

* Controlling root access and privilege escalation within containers
* Propagating the user identity to align with enterprise access policies

## Root Access and Privilege Escalation

NVIDIA Run:ai supports security-related workload configurations to control user permissions and restrict privilege escalation. These options are available via the API and CLI during workload creation:

* `runAsNonRoot` / `--run-as-user` - Force the container to run as non-root user.
* `allowPrivilegeEscalation` / `--allow-privilege-escalation` - Allow the container to use `setuid` binaries to escalate privileges, even when running as a non-root user. This setting can increase security risk and should be disabled if elevated privileges are not required.
* `privileged` - Grant full host access to the container, bypassing isolation and security restrictions. By system policy, this parameter is always set to `false` to prevent containers from gaining privileged access. Only an administrator can override this restriction by explicitly creating a policy to override the system policy. This ensures that privileged access is granted only through a deliberate, audited action. See [System policies](/saas/platform-management/policies/policies-and-rules#system-policies) for more details.

Administrators can enforce secure defaults across the environment using [Policies](/saas/platform-management/policies/native-workload-policies/policy-yaml-reference), ensuring consistent workload behavior aligned with organizational security practices.

{% hint style="info" %}
**Note**

If both `privileged` and `allowPrivilegeEscalation` are set to `true`, the `allowPrivilegeEscalation` setting becomes redundant. Enabling `privileged` gives the container full host access, which already includes all privilege escalation capabilities.
{% endhint %}

## Passing User Identity

### Passing User Identity from Identity Provider

A best practice is to store the **User Identifier (UID)** and **Group Identifier (GID)** in the organization's directory. NVIDIA Run:ai allows you to pass these values to the container and use them as the container identity. To perform this, you must set up [single sign-on](/saas/infrastructure-setup/authentication/overview) and perform the steps for UID/GID integration.

### Passing User Identity via UI

It is possible to explicitly pass user identity when creating an [environment](/saas/workloads-in-nvidia-run-ai/assets/environments) or submitting a [workload](/saas/workloads-in-nvidia-run-ai/workloads):

* **From the image** - Use the UID/GID defined in the container image.
* **From the IdP token** - Use identity attributes provided by the SSO identity provider (available only in SSO-enabled installations).
* **Custom** - Manually set the **User ID (UID)**, **Group ID (GID)** and **supplementary groups** that can run commands in the container.

Administrators can enforce secure defaults across the environment using [Policies](/saas/platform-management/policies/native-workload-policies/policy-yaml-reference), ensuring consistent workload behavior aligned with organizational security practices.

{% hint style="info" %}
**Note**

It is also possible to set the above using the API or CLI.
{% endhint %}

## Using OpenShift or Gatekeeper to Provide Cluster Level Controls

In OpenShift, Security Context Constraints (SCCs) manage pod-level security, including root access. By default, containers are assigned a random non-root UID, and flags such as `--run-as-user` and `--allow-privilege-escalation` are disabled.

On non-OpenShift Kubernetes clusters, similar enforcement can be achieved using tools like [Gatekeeper](https://open-policy-agent.github.io/gatekeeper/website/docs/), which applies system-level policies to restrict containers from running as root.

## Enabling UID and GID on OpenShift

By default, OpenShift restricts setting specific user and group IDs (UIDs/GIDs) in workloads through its SCCs. To allow NVIDIA Run:ai workloads to run with explicitly defined UIDs and GIDs, a cluster administrator must modify the relevant SCCs.

To enable UID and GID assignment:

1. Edit the `runai-user-job` SCC:

   ```bash
   oc edit scc runai-user-job
   ```
2. Edit the `runai-jupyter-notebook` SCC (only required if using Jupyter environments):

   ```bash
   oc edit scc runai-jupyter-notebook
   ```
3. In both SCC definitions, ensure the following sections are configured:

   ```yaml
   runAsUser:
     type: RunAsAny
   supplementalGroups:
     type: RunAsAny
   ```

These settings allow NVIDIA Run:ai to pass specific UID and GID values into the container, enabling compatibility with identity-aware file systems and enterprise access controls.

## Creating a Temporary Home Directory

When containers run as a specific user, the user must have a home directory defined within the image. Otherwise, starting a shell session will fail due to the absence of a home directory.

Since pre-creating a home directory for every possible user is impractical, NVIDIA Run:ai offers the `createHomeDir` / `--create-home-dir` option. When enabled, this flag creates a temporary home directory for the user inside the container at runtime. By default, the directory is created at `/home/<username>`.

{% hint style="info" %}
**Note**

* This home directory is temporary and exists only for the duration of the container's lifecycle. Any data saved in this location will be lost when the container exits.
* By default, this flag is set to `true` when `--run-as-user` is enabled, and `false` otherwise.
  {% endhint %}


# Service Mesh

NVIDIA Run:ai supports service mesh implementations. When a service mesh is deployed with sidecar injection, specific configurations must be applied to ensure compatibility with NVIDIA Run:ai. This document outlines the required changes for the NVIDIA Run:ai control plane and cluster.

## Control Plane Configuration

{% hint style="info" %}
**Note**

This section applies to **self-hosted only**.
{% endhint %}

By default, NVIDIA Run:ai prevents Istio from injecting sidecar containers into system jobs in the control plane. For other service mesh solutions, users must manually add annotations during installation.

To disable sidecar injection in the NVIDIA Run:ai control plane, modify the Helm values file by adding the required pod labels to the following components. See [Advanced control plane configurations](/self-hosted/2.24/infrastructure-setup/advanced-setup/control-plane-config) for more details.

Example for [Open Service Mesh](https://release-v0-11.docs.openservicemesh.io/docs/guides/app_onboarding/sidecar_injection/#explicitly-disabling-automatic-sidecar-injection-on-pods):

```yaml
authorizationMigrator:
  podLabels:
    openservicemesh.io/sidecar-injection: disabled
clusterMigrator:
  podLabels:
    openservicemesh.io/sidecar-injection: disabled
identityProviderReconciler:
  podLabels:
    openservicemesh.io/sidecar-injection: disabled
keepPVC:
  podLabels:
    openservicemesh.io/sidecar-injection: disabled
orgUnitsMigrator:
  podLabels:
    openservicemesh.io/sidecar-injection: disabled
```

## Cluster Configuration

### Installation Phase

Sidecar containers injected by some service mesh solutions can prevent NVIDIA Run:ai installation hooks from completing. To avoid this, modify the Helm installation command to include the required labels or annotations:

```bash
helm upgrade -i ... 
--set global.additionalJobLabels.A=B --set global.additionalJobAnnotations.A=B
```

Example for [Istio Service Mesh](https://istio.io/latest/docs/setup/additional-setup/sidecar-injection/#controlling-the-injection-policy):

```bash
helm upgrade -i ... 
--set-json global.additionalJobLabels='{"sidecar.istio.io/inject":false}'
```

### Workloads

To prevent sidecar injection in workloads created at runtime (such as training workloads), update the `runaiconfig` resource. See [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config) for more details:

```yaml
spec:
  workload-controller:
    additionalPodLabels:
      sidecar.istio.io/inject: false
```


# Integrations

Integrations are Kubernetes components and external tools that can be used with NVIDIA Run:ai for development, training, orchestration, data access, and monitoring.

Integrations fall into two support levels:

* **Supported integrations (out of the box)** - NVIDIA Run:ai includes built-in support and documentation. You may still need cluster-side installation (for example, install an operator and its CRDs) before you can use the integration.
* **Community Support integrations** - Not supported out of the box, but commonly used with prior customer support experience and reference guides.

## Supported Integrations

### Frameworks

<table><thead><tr><th width="110.7734375">Framework</th><th width="127.54296875">Category</th><th>Supported Version</th><th width="500.12109375">Additional Information</th></tr></thead><tbody><tr><td>Dynamo operator</td><td>Distributed inference</td><td>1.2.0</td><td><p>Dynamo operator is a Kubernetes operator that simplifies the deployment, configuration, and lifecycle management of DynamoGraphs.</p><p>NVIDIA Run:ai provides out of the box support for submitting Dynamo workloads <a href="/pages/5yxvtdITa6dDsF9gsXKB">via YAML</a>. See <a href="https://docs.nvidia.com/dynamo/latest/">Dynamo operator</a> documentation for more details.</p></td></tr><tr><td>NIM operator</td><td>Model Serving</td><td>3.0.x</td><td><p>The NVIDIA NIM Operator enables Kubernetes cluster administrators to operate the software components and services necessary to deploy NVIDIA NIMs and NVIDIA NeMo microservices in Kubernetes.</p><p>NVIDIA Run:ai provides out of the box support for submitting NIM operator workloads <a href="/pages/5yxvtdITa6dDsF9gsXKB">via YAML</a>. See <a href="https://docs.nvidia.com/nim-operator/latest/">NIM operator</a> documentation for more details.</p></td></tr><tr><td>LeaderWorkerSet (LWS)</td><td>Distributed inference</td><td>0.6.0 or higher</td><td>NVIDIA Run:ai provides out of the box support for submitting <a href="https://github.com/kubernetes-sigs/lws">LWS</a> workloads <a href="/pages/5yxvtdITa6dDsF9gsXKB">via YAML</a> and NVIDIA Run:ai native <a href="/pages/MvfEUxoFFjscvdMfwx4t">distributed inference</a> workloads using LWS.</td></tr><tr><td>Kubeflow MPI</td><td>Distributed training</td><td>MPI Operator v0.6.0 or higher</td><td>NVIDIA Run:ai provides out of the box support for submitting MPI workloads via API, CLI or UI. See <a href="/pages/QefgLB9trdzzetQCpfgD">Distributed training</a> for more details.</td></tr><tr><td>PyTorch</td><td>Distributed training</td><td>Kubeflow Training Operator v1.9.2</td><td>NVIDIA Run:ai provides out of the box support for submitting PyTorch workloads via API, CLI or UI. See <a href="/pages/QefgLB9trdzzetQCpfgD">Distributed training</a> for more details.</td></tr><tr><td>TensorFlow</td><td>Distributed training</td><td>Kubeflow Training Operator v1.9.2</td><td>NVIDIA Run:ai provides out of the box support for submitting TensorFlow workloads via API, CLI or UI. See <a href="/pages/QefgLB9trdzzetQCpfgD">Distributed training</a> for more details.</td></tr><tr><td>XGBoost</td><td>Distributed training</td><td>Kubeflow Training Operator v1.9.2</td><td>NVIDIA Run:ai provides out of the box support for submitting XGBoost via API, CLI or UI. See <a href="/pages/QefgLB9trdzzetQCpfgD">Distributed training</a> for more details.</td></tr><tr><td>JAX</td><td>Distributed training</td><td>Kubeflow Training Operator v1.9.2</td><td>NVIDIA Run:ai provides out of the box support for submitting JAX workloads via API, CLI or UI. See <a href="/pages/QefgLB9trdzzetQCpfgD">Distributed training</a> for more details.</td></tr><tr><td>Triton</td><td>Orchestration</td><td>Any version</td><td>Usage via docker base image</td></tr></tbody></table>

### Development Tools

<table><thead><tr><th width="110.7734375">Tool</th><th width="127.54296875">Category</th><th width="502.9609375">Additional Information</th></tr></thead><tbody><tr><td>Jupyter Notebook</td><td>Development</td><td>NVIDIA Run:ai provides integrated support with Jupyter Notebooks. See <a href="/pages/P4TYcp0mB43oqCzeNQa5">Jupyter Notebook quick start</a> example.</td></tr><tr><td>PyCharm</td><td>Development</td><td>Containers created by NVIDIA Run:ai can be accessed via PyCharm.</td></tr><tr><td>VScode</td><td>Development</td><td>Containers created by NVIDIA Run:ai can be accessed via Visual Studio Code. You can automatically launch Visual Studio code web from the NVIDIA Run:ai console.</td></tr></tbody></table>

### Storage and Registries

<table><thead><tr><th width="110.7734375">Tool</th><th width="127.54296875">Category</th><th width="507.38671875">Additional Information</th></tr></thead><tbody><tr><td>Docker Registry</td><td>Repositories</td><td>NVIDIA Run:ai allows using a docker registry as a <a href="/pages/GuBN1o0IXQ2RGRmADNDd">Credential</a> asset</td></tr><tr><td>GitHub</td><td>Storage</td><td>NVIDIA Run:ai communicates with GitHub by defining it as a <a href="/pages/En4yELogvYkYI4gGTmjS">data source</a> asset</td></tr><tr><td>S3</td><td>Storage</td><td>NVIDIA Run:ai communicates with S3 by defining a <a href="/pages/En4yELogvYkYI4gGTmjS">data source</a> asset</td></tr></tbody></table>

### Experiment Tracking and Monitoring

<table><thead><tr><th width="110.7734375">Tool</th><th width="127.54296875">Category</th><th width="510.625">Additional Information</th></tr></thead><tbody><tr><td>TensorBoard</td><td>Experiment tracking</td><td>NVIDIA Run:ai comes with a preset TensorBoard <a href="/pages/VvT46F94K378ZKZzzwA4">Environment</a> asset</td></tr></tbody></table>

### Infrastructure and Cost Optimization

<table><thead><tr><th width="110.7734375">Tool</th><th width="127.54296875">Category</th><th width="507.65625">Additional Information</th></tr></thead><tbody><tr><td>Karpenter</td><td>Cost Optimization</td><td>NVIDIA Run:ai provides out of the box support for Karpenter to save cloud costs. Integration notes with Karpenter can be found <a href="/pages/grUKbZNqqHOcOZ6H2ah4">here</a>.</td></tr></tbody></table>

## Community Support Integrations

Our Customer Success team has prior experience assisting customers with setup. In many cases, the NVIDIA Enterprise Support Portal may include additional reference documentation provided on an as-is basis.

<table><thead><tr><th width="110.7734375">Tool</th><th width="127.54296875">Category</th><th width="509.7734375">Additional Information</th></tr></thead><tbody><tr><td>Apache Airflow</td><td>Orchestration</td><td>It is possible to schedule Airflow workflows with the NVIDIA Run:ai Scheduler. Sample code: <a href="https://enterprise-support.nvidia.com/s/article/How-to-integrate-Run-ai-with-Apache-Airflow">How to integrate NVIDIA Run:ai with Apache Airflow</a>.</td></tr><tr><td>Argo workflows</td><td>Orchestration</td><td>It is possible to schedule Argo workflows with the NVIDIA Run:ai Scheduler. Sample code: <a href="https://enterprise-support.nvidia.com/s/article/How-to-integrate-Run-ai-with-Argo-Workflows">How to integrate NVIDIA Run:ai with Argo Workflows</a>.</td></tr><tr><td>ClearML</td><td>Experiment tracking</td><td>It is possible to schedule ClearML workloads with the NVIDIA Run:ai Scheduler.</td></tr><tr><td>JupyterHub</td><td>Development</td><td>It is possible to submit NVIDIA Run:ai workloads via JupyterHub.</td></tr><tr><td>Kubeflow notebooks</td><td>Development</td><td>It is possible to launch a Kubeflow notebook with the NVIDIA Run:ai Scheduler. Sample code: <a href="https://enterprise-support.nvidia.com/s/article/How-to-integrate-Run-ai-with-Kubeflow">How to integrate NVIDIA Run:ai with Kubeflow</a>.</td></tr><tr><td>Kubeflow Pipelines</td><td>Orchestration</td><td>It is possible to schedule kubeflow pipelines with the NVIDIA Run:ai Scheduler. Sample code: <a href="https://enterprise-support.nvidia.com/s/article/How-to-integrate-Run-ai-with-Kubeflow">How to integrate NVIDIA Run:ai with Kubeflow</a>.</td></tr><tr><td>MLFlow</td><td>Model Serving</td><td>It is possible to use ML Flow together with the NVIDIA Run:ai Scheduler.</td></tr><tr><td>Ray</td><td>Training, inference, data processing</td><td>It is possible to schedule Ray jobs with the NVIDIA Run:ai Scheduler. Sample code: <a href="https://enterprise-support.nvidia.com/s/article/How-to-Integrate-Run-ai-with-Ray">How to Integrate NVIDIA Run:ai with Ray</a>.</td></tr><tr><td>SeldonX</td><td>Orchestration</td><td>It is possible to schedule Seldon Core workloads with the NVIDIA Run:ai Scheduler.</td></tr><tr><td>Spark</td><td>Orchestration</td><td>It is possible to schedule Spark workflows with the NVIDIA Run:ai Scheduler.</td></tr><tr><td>Weights &#x26; Biases</td><td>Experiment tracking</td><td>It is possible to schedule W&#x26;B workloads with the NVIDIA Run:ai Scheduler. Sample code: <a href="https://enterprise-support.nvidia.com/s/article/How-to-integrate-with-Weights-and-Biases">How to integrate with Weights and Biases</a>.</td></tr></tbody></table>


# Interworking with Karpenter

Karpenter is an open-source, Kubernetes cluster autoscaler built for cloud deployments. Karpenter optimizes the cloud cost of a customer’s cluster by moving workloads between different node types, consolidating workloads into fewer nodes, using lower-cost nodes where possible, scaling up new nodes when needed, and shutting down unused nodes.

Karpenter’s main goal is cost optimization. Unlike Karpenter, NVIDIA Run:ai’s Scheduler optimizes for fairness and resource utilization. Therefore, there are a few potential friction points when using both on the same cluster.

## Friction Points Using Karpenter with NVIDIA Run:ai

1. Karpenter looks for “unschedulable” pending workloads and may try to scale up new nodes to make those workloads schedulable. However, in some scenarios, these workloads may exceed their quota parameters, and the NVIDIA Run:ai Scheduler will put them into a pending state.
2. Karpenter is not aware of the NVIDIA Run:ai fractions mechanism and may try to interfere incorrectly.
3. Karpenter preempts any type of workload (i.e., high-priority, non-preemptible workloads will potentially be interrupted and moved to save cost).
4. Karpenter has no pod-group (i.e., workload) notion or gang scheduling awareness, meaning that Karpenter is unaware that a set of “arbitrary” pods is a single workload. This may cause Karpenter to schedule those pods into different node pools (in the case of multi-node-pool workloads) or scale up or down a mix of wrong nodes.

### Mitigating the Friction Points

NVIDIA Run:ai Scheduler mitigates the friction points using the following techniques (each numbered bullet below corresponds to the related friction point listed above):

1. Karpenter uses a “nominated node” to recommend a node for the Scheduler. The NVIDIA Run:ai Scheduler treats this as a “preferred” recommendation, meaning it will try to use this node, but it’s not required and it may choose another node.
2. Fractions - Karpenter won’t consolidate nodes with one or more pods that cannot be moved. The NVIDIA Run:ai reservation pod is marked as ‘do not evict’ to allow the NVIDIA Run:ai Scheduler to control the scheduling of fractions.
3. Non-preemptible workloads - NVIDIA Run:ai marks non-preemptible workloads as ‘do not evict’ and Karpenter respects this annotation.
4. NVIDIA Run:ai node pools (single-node-pool workloads) - Karpenter respects the ‘node affinity’ that NVIDIA Run:ai sets on a pod, so Karpenter uses the node affinity for its recommended node. For the gang-scheduling/pod-group (workload) notion, NVIDIA Run:ai Scheduler considers Karpenter directives as preferred recommendations rather than mandatory instructions and overrides Karpenter instructions where appropriate.

### Deployment Considerations

* Using multi-node-pool workloads
  * Workloads may include a list of optional node pools. Karpenter is not aware that only a single node pool should be selected out of that list for the workload. It may therefore recommend putting pods of the same workload into different node pools and may scale up nodes from different node pools to serve a “multi-node-pool” workload instead of nodes on the selected single node pool.
  * If this becomes an issue (i.e., if Karpenter scales up the wrong node types), users can set an inter-pod affinity using the node pool label or another common label as a ‘topology’ identifier. This will force Karpenter to choose nodes from a single-node pool per workload, selecting from any of the node pools listed as allowed by the workload.
  * An alternative approach is to use a single-node pool for each workload instead of multi-node pools.
* Consolidation
  * To make Karpenter more effective when using its consolidation function, users should consider separating preemptible and non-preemptible workloads, either by using node pools, node affinities, taint/tolerations, or inter-pod anti-affinity.
  * If users don’t separate preemptible and non-preemptible workloads (i.e., make them run on different nodes), Karpenter’s ability to consolidate (bin-pack) and shut down nodes will be reduced, but it is still effective.
* Conflicts between bin-packing and spread policies
  * If NVIDIA Run:ai is used with a scheduling spread policy, it will clash with Karpenter’s default bin-packs/consolidation policy, and the outcome may be a deployment that is not optimized for any of these policies.
  * Usually spread is used for Inference, which is non-preemptible and therefore not controlled by Karpenter (NVIDIA Run:ai Scheduler will mark those workloads as ‘do not evict’ for Karpenter), so this should not present a real deployment issue for customers.


# Security Best Practices

This guide provides actionable best practices for administrators to securely configure, operate, and manage NVIDIA Run:ai environments. Each section highlights both platform-native features and mapped Kubernetes security practices to maintain robust protection for workloads and resources.

| Security Area                             | Best Practice                                                                              |
| ----------------------------------------- | ------------------------------------------------------------------------------------------ |
| Access control (RBAC)                     | Enforce least privilege, segment roles by scope, audit regularly                           |
| Authentication and sessions management    | Use SSO, token-based authentication, strong passwords, limit idle time                     |
| Workload policies                         | Require non-root, set UID/GID, block overrides, use trusted images                         |
| Namespace and resource management         | Require namespace approval, limit secret propagation, apply quotas                         |
| Tools and serving endpoint access control | Control who can access tools and endpoints; restrict network exposure                      |
| Maintenance and compliance                | Follow secure install guides, perform vulnerability scans, maintain data-privacy alignment |

## Access Control (RBAC)

NVIDIA Run:ai uses Role‑Based Access Control to define what each user, group, or service account can do, and where. Roles are assigned within a scope, such as a project, department, or cluster, and permissions cover actions like viewing, creating, editing, or deleting entities. Unlike Kubernetes RBAC, NVIDIA Run:ai’s RBAC works across multiple clusters, giving you a single place to manage access rules. See [Role Based Access Control (RBAC)](/saas/infrastructure-setup/authentication/overview#role-based-access-control-rbac-in-nvidia-run-ai) for more details.

### Best Practices

* Assign the minimum required permissions to users, groups and service accounts.
* Segment duties using organizational scopes to restrict roles to specific projects or departments.
* Regularly audit access rules and remove unnecessary privileges, especially admin-level roles.

### Kubernetes Connection

NVIDIA Run:ai predefined roles are automatically mapped to Kubernetes cluster roles (also predefined by NVIDIA Run:ai). This means administrators do not need to manually configure role mappings.

These cluster roles define permissions for the entities NVIDIA Run:ai manages and displays (such as workloads) and also apply to users who access cluster data directly through Kubernetes tools (for example, `kubectl`).

## Authentication and Session Management

NVIDIA Run:ai supports several authentication methods to control platform access. You can use single sign-on (SSO) for unified enterprise logins, traditional username/password accounts if SSO isn’t an option, and API secret keys for automated application access. Authentication is mandatory for all interfaces, including the UI, CLI, and APIs, ensuring only verified users or applications can interact with your environment.

Administrators can also configure session timeout. This refers to the period of inactivity before a user is automatically logged out. Once the timeout is reached, the session ends and re‑authentication is required, helping protect against risks from unattended or abandoned sessions. See [Authentication and authorization](/saas/infrastructure-setup/authentication/overview) for more details.

### Best Practices

* Integrate corporate SSO for centralized identity management.
* Enforce strong password policies for local accounts.
* Set appropriate session timeout values to minimize idle session risk.
* Prefer SSO to eliminate password management within NVIDIA Run:ai.

### Kubernetes Connection

Configure the Kubernetes API server to validate tokens via NVIDIA Run:ai’s identity service, ensuring unified authentication across the platform. For more information, see [Cluster authentication](/saas/infrastructure-setup/authentication/cluster-authentication).

## Workload Policies: Enforcing Security at Submission

Workload policies allow administrators to define and enforce how AI workloads are submitted and controlled across projects and teams. With these policies, you can set clear rules and defaults for workload parameters such as which resources can be requested, required security settings, and which defaults should apply. Policies are enforced whether workloads are submitted via the UI, CLI, API or Kubernetes YAML, and can be scoped to specific projects, departments, or clusters for fine-grained control. See [Policies and rules](/saas/platform-management/policies/policies-and-rules) for more details.

### Best Practices

* Enforce containers to run as non-root by default. Define policies that set constraints and defaults for workload submissions, such as requiring non-root users or specifying minimum UID/GID. Example security fields in policies:
  * `security.runAsNonRoot: true`
  * `security.runAsUid: 1000`
  * Restrict `runAsUid` with `canEdit: false` to prevent users from overriding.
* Require explicit user/group IDs for all workload containers.
* Impose data source and resource usage limits through policies.
* Use policy rules to prevent users from submitting non-compliant workloads.
* Apply policies by organizational scope for nuanced control within departments or projects.

### Kubernetes Connection

Map these policies to `PodSecurityContext` settings in Kubernetes, and enforce them with Pod Security Admission or Kyverno for stricter compliance.

## Managing Namespace and Resource Creation

NVIDIA Run:ai offers flexible controls for how namespaces and resources are created and managed within your clusters. When a new project is set up, you can choose whether Kubernetes namespaces are created automatically, and whether users are auto-assigned to those projects. There are also options to manage how secrets are propagated across namespaces and to enable or disable resource limit enforcement using Kubernetes LimitRange objects. See [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config) for more details.

### Best Practices

* Require admin approval for namespace creation to avoid sprawl.
* Limit secret propagation to essential cases only.
* Use Kubernetes `LimitRanges` and `ResourceQuotas` alongside NVIDIA Run:ai policies for layered resource control.
* Regularly audit and remove unused namespaces, secrets, and workloads.

## Tools and Serving Endpoint Access Control

NVIDIA Run:ai provides flexible options to control access to tools and serving endpoints. Access can be defined during workload submission or updated later, ensuring that only the intended users or groups can interact with the resource.

When configuring an endpoint or tool, users can select from the following access levels:

* **Public** - Everyone within the network can access with no authentication (serving endpoints).
* **All authenticated users** - Access is granted to anyone in the organization who can log in (NVIDIA Run:ai or SSO).
* **Specific groups** - Access is restricted to members of designated identity provider groups.
* **Specific users** - Access is restricted to individual users by email or username.

By default, network exposure is restricted, and access must be explicitly granted. Model endpoints automatically inherit RBAC and workload policy controls, ensuring consistent enforcement of role- and scope-based permissions across the platform. Administrators can also limit who can deploy, view, or manage endpoints, and should open network access only when required.

### Best Practices

* Define explicit roles for model management/use.
* Restrict endpoint access to authorized users, groups and applications.
* Monitor and audit endpoint access logs.

### Kubernetes Connection

Use Kubernetes `NetworkPolicies` to limit inter-pod and external traffic to model-serving pods. Pair with NVIDIA Run:ai RBAC for end-to-end control.

## Secure Installation and Maintenance

A secure deployment is the foundation on which all other controls rest, and NVIDIA Run:ai’s installation procedures are built to align with organizational policies such as OpenShift Security Context Constraints (SCC). See [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config) for more details.

* Deploy NVIDIA Run:ai cluster following secure installation guides (including IT compliance mandates such as SCC for OpenShift).
* Run regular security scans and patch/update NVIDIA Run:ai deployments promptly when vulnerabilities are reported.
* Regularly review and update all security policies, both at the NVIDIA Run:ai and Kubernetes levels, to adapt to evolving risks.

## Compliance and Data Privacy

NVIDIA Run:ai supports SaaS and self-hosted modes to satisfy a range of data security needs. The self-hosted mode keeps all models, logs, and user data entirely within your infrastructure; SaaS requires careful review of what (minimal) data is transmitted for platform operations and analytics. See [Compliance](/saas/infrastructure-setup/procedures/compliance) for more details.

* Use the self-hosted mode when full control over the environment is required - including deployment and day-2 operations such as upgrades, monitoring, backup, and metadata restore.
* Ensure transmission to the NVIDIA Run:ai cloud is scoped (in SaaS mode) and aligns with organization policy.
* Encrypt secrets and sensitive resources; control secret propagation.
* Document and audit data flows for regulatory alignment.


# Infrastructure Procedures


# NVIDIA Run:ai at Scale

Operating NVIDIA Run:ai at scale ensures that the system can efficiently handle fluctuating workloads while maintaining optimal performance. As clusters grow, whether due to an increasing number of nodes or a surge in workload demand, NVIDIA Run:ai services must be appropriately tuned to support large-scale environments.

This guide outlines the best practices for optimizing NVIDIA Run:ai for high-performance deployments, including NVIDIA Run:ai system services configurations, vertical scaling (adjusting CPU and memory resources) and where applicable, horizontal scaling (replicas).

## NVIDIA Run:ai Services

### Vertical Scaling

Each of the NVIDIA Run:ai containers has default resource requirements that reflect an average customer load. With significantly larger cluster loads, certain NVIDIA Run:ai services will require more CPU and memory resources. NVIDIA Run:ai supports configuring these resources for each NVIDIA Run:ai service group separately. For instructions and more information, see [NVIDIA Run:ai services resource management](/saas/infrastructure-setup/advanced-setup/cluster-config#nvidia-run-ai-services-resource-management).

#### Scheduling Services

The scheduling services group should be scaled together with the number of [nodes](/saas/platform-management/aiinitiatives/resources/nodes) and the number of [workloads](/saas/workloads-in-nvidia-run-ai/workloads) handled by the[ Scheduler](/saas/platform-management/runai-scheduler/scheduling/how-the-scheduler-works) (running / pending). These resource recommendations are based on internal benchmarks performed on stressed environments:

<table><thead><tr><th width="246">Scale (nodes/workloads)</th><th width="154">CPU (request)</th><th>Memory (request)</th></tr></thead><tbody><tr><td>Small - 30 / 480</td><td>1</td><td>1GB</td></tr><tr><td>Medium - 100 / 1600</td><td>2</td><td>2GB</td></tr><tr><td>Large - 500 / 8500</td><td>2</td><td>7GB</td></tr></tbody></table>

#### Sync and Workload Services

The sync and workload service groups are less sensitive for scale. The recommendation for large or intensive environments is set to the following:

<table><thead><tr><th width="246">Scale (nodes/workloads)</th><th width="154">CPU (request)</th><th>Memory (request)</th></tr></thead><tbody><tr><td>Small - 30 / 480</td><td>1</td><td>2GB</td></tr><tr><td>Medium - 100 / 1600</td><td>2</td><td>10GB</td></tr><tr><td>Large - 500 / 8500</td><td>4</td><td>24GB</td></tr></tbody></table>

### Horizontal Scaling

By default, NVIDIA Run:ai cluster services are deployed with a single replica. For large scale and intensive environments it is recommended to scale the NVIDIA Run:ai services horizontally by increasing the number of replicas. For more information, see [NVIDIA Run:ai services replicas](/saas/infrastructure-setup/advanced-setup/cluster-config#nvidia-run-ai-services-replicas).

## Metrics Collection

NVIDIA Run:ai relies on [Prometheus](/saas/getting-started/installation/install-using-helm/system-requirements#prometheus) to scrape cluster metrics and forward them to the NVIDIA Run:ai control plane. The volume of metrics generated is directly proportional to the number of nodes, workloads, and projects in the system. When operating at scale—reaching hundreds, and thousands of nodes and projects—the system generates a significant volume of metrics which can place a strain on the cluster and the network bandwidth.

To mitigate this impact, it is recommended to tune the Prometheus [remote-write](https://prometheus.io/docs/specs/remote_write_spec/) configurations. See [remote write tuning](https://prometheus.io/docs/practices/remote_write/#remote-write-tuning) to read more about the tuning parameters available via the remote write configuration and refer to this [article](https://last9.io/blog/optimizing-prometheus-remote-write-performance-guide/) for optimizing Prometheus remote write performance.

You can apply the remote-write configurations required as described in [advanced cluster configurations.](/saas/infrastructure-setup/advanced-setup/cluster-config#prometheus)

The following example demonstrates the recommended approach in NVIDIA Run:ai for tuning **Prometheus remote-write** configurations:

```yaml
remoteWrite:
  queueConfig:
    capacity: 5000
    maxSamplesPerSend: 1000
    maxShards: 100
```


# High Availability

This guide outlines the best practices for configuring the NVIDIA Run:ai platform to ensure high availability and maintain service continuity during system failures or under heavy load. The goal is to reduce downtime and eliminate single points of failure by leveraging Kubernetes best practices alongside NVIDIA Run:ai specific configuration options. The NVIDIA Run:ai platform relies on two fundamental high availability strategies:

* **Use of system nodes** - Assigning multiple dedicated nodes for critical system services ensures control, resource isolation, and enables system-level scaling.
* **Replication of core and third-party services** - Configuring multiple replicas of essential services, distributes workloads and reduces single points of failure. If a component fails on one node, requests can seamlessly route to another instance.

## System Nodes

The NVIDIA Run:ai platform allows you to dedicate specific nodes (system nodes) exclusively for core platform services. This approach provides improved operational isolation and easier resource management.

Ensure that at least **three system nodes** are configured to support high availability. If you use only a single node for core services, horizontally scaled components will not be distributed, resulting in a single point of failure. See [NVIDIA Run:ai system nodes](/saas/infrastructure-setup/advanced-setup/node-roles) for more details.

## Service Replicas <a href="#undefined" id="undefined"></a>

By default, NVIDIA Run:ai cluster services are deployed with a single replica. To achieve high availability, it is recommended to configure multiple replicas for core NVIDIA Run:ai services. For more information, see [NVIDIA Run:ai services replicas](/saas/infrastructure-setup/advanced-setup/cluster-config#nvidia-run-ai-services-replicas).


# Monitoring and Maintenance

Deploying NVIDIA Run:ai in mission-critical environments requires proper monitoring and maintenance of resources to ensure workloads run and are deployed as expected.

Details on how to monitor different parts of the physical resources in your Kubernetes system, including [clusters](/saas/infrastructure-setup/procedures/clusters) and [nodes](/saas/platform-management/aiinitiatives/resources/nodes), can be found in the monitoring and maintenance section. Adjacent configuration and troubleshooting sections also cover high availability, [restoring](/saas/infrastructure-setup/procedures/cluster-restore) and [securing](/saas/infrastructure-setup/procedures/secure-your-cluster) clusters, [collecting logs](/saas/infrastructure-setup/procedures/logs-collection), and [reviewing audit logs](/saas/infrastructure-setup/procedures/event-history) to meet compliance requirements.

In addition to monitoring NVIDIA Run:ai resources, it is also highly recommended to monitor NVIDIA Run:ai runs on Kubernetes, which manages containerized applications. In particular, focus on three main layers:

## NVIDIA Run:ai Control Plane and Cluster Services

This is the highest layer and includes the parts of NVIDIA Run:ai pods, which run in containers managed by Kubernetes.

## Kubernetes Cluster

This layer includes the main Kubernetes system that runs and manages NVIDIA Run:ai components. Important elements to monitor include:

* The health of the cluster and nodes (machines in the cluster).
* The status of key Kubernetes services, such as the API server. For detailed information on managing clusters, see the [official Kubernetes documentation](https://kubernetes.io/docs/tasks/debug/debug-cluster/resource-usage-monitoring/).

## Host Infrastructure

This is the base layer, representing the actual machines (virtual or physical) that make up the cluster IT teams need to handle:

* Managing CPU, memory, and storage
* Keeping the operating system updated
* Setting up the network and balancing the load

NVIDIA Run:ai does not require any special configurations at this level.

The articles below explain how to monitor these layers, maintain system security and compliance, and ensure the reliable operation of NVIDIA Run:ai in critical environments.


# NVIDIA Run:ai System Monitoring

This guide explains how to configure NVIDIA Run:ai to generate health alerts and to connect these alerts to alert-management systems within your organization. Alerts are generated for NVIDIA Run:ai clusters.

## Alert Infrastructure

NVIDIA Run:ai uses Prometheus for externalizing metrics and providing visibility to end-users. The NVIDIA Run:ai Cluster installation includes Prometheus or can connect to an existing Prometheus instance used in your organization. The alerts are based on the Prometheus AlertManager. Once installed, it is enabled by default.

This document explains how to:

* Configure alert destinations - triggered alerts send data to specified destinations
* Understand the out-of-the-box cluster alerts, provided by NVIDIA Run:ai
* Add additional custom alerts

## Prerequisites

* A Kubernetes cluster with the necessary permissions
* Up and running NVIDIA Run:ai environment, including Prometheus Operator
* [kubectl](https://kubernetes.io/docs/reference/kubectl/) command-line tool installed and configured to interact with the cluster

## Setup

Use the steps below to set up monitoring alerts.

### Validating Prometheus Operator Installed

1. Verify that the Prometheus Operator Deployment is running. Copy the following command and paste it in your terminal, where you have access to the Kubernetes cluster. In your terminal, you can see an output indicating the deployment's status, including the number of replicas and their current state.

   ```bash
   kubectl get deployment kube-prometheus-stack-operator -n monitoring
   ```
2. Verify that Prometheus instances are running. Copy the following command and paste it in your terminal. You can see the Prometheus instance(s) listed along with their status:

   ```bash
   kubectl get prometheus -n runai
   ```

### Enabling Prometheus AlertManager

In each of the steps in this section, copy the content of the code snippet to a new YAML file (e.g., `step1.yaml`).

1. Copy the following command to your terminal, to apply the YAML file to the cluster:

   ```bash
   kubectl apply -f step1.yaml 
   ```
2. Copy the following command to your terminal to create the AlertManager CustomResource, to enable AlertManager:

   ```yaml
   apiVersion: monitoring.coreos.com/v1  
   kind: Alertmanager  
   metadata:  
      name: runai  
      namespace: runai  
   spec:  
      replicas: 1  
      alertmanagerConfigSelector:  
         matchLabels:
            alertmanagerConfig: runai 
   ```
3. Copy the following command to your terminal to validate that the AlertManager instance has started:

   ```bash
   kubectl get alertmanager -n runai
   ```
4. Copy the following command to your terminal to validate that the Prometheus operator has created a Service for AlertManager:

   ```bash
   kubectl get svc alertmanager-operated -n runai
   ```

### Configuring Prometheus to Send Alerts

1. Open the terminal on your local machine or another machine that has access to your Kubernetes cluster.
2. Copy and paste the following command in your terminal to edit the Prometheus configuration for the `runai` namespace. This command opens the Prometheus configuration file in your default text editor (usually `vi` or `nano`):

   ```bash
   kubectl edit runaiconfig -n runai
   ```
3. Copy and paste the following text to your terminal to change the configuration file:

   ```yaml
   prometheus:
     spec:
       alerting:
         alertmanagers:
         - name: alertmanager-operated
           namespace: runai
           port: web
   ```
4. Delete the prometheus pod to reset the pod's settings:

   ```bash
   kubectl delete pod prometheus-runai-0 -n runai
   ```
5. Save the changes and exit the text editor.

{% hint style="info" %}
**Note**

To save changes using `vi`, type `:wq` and press Enter.\
The changes are applied to the Prometheus configuration in the cluster.
{% endhint %}

## Alert Destinations

Set out below are the various alert destinations.

### Configuring AlertManager for Custom Email Alerts

In each step, copy the contents of the code snippets to a new file and apply it to the cluster using `kubectl apply -f`.

1. Add your smtp password as a secret:

   ```yaml
   apiVersion: v1  
   kind: Secret  
   metadata:  
      name: alertmanager-smtp-password  
      namespace: runai  
   stringData:
      password: "your_smtp_password"
   ```
2. Replace the relevant smtp details with your own, then apply the `alertmanagerconfig` using `kubectl apply`:

   ```yaml
   apiVersion: monitoring.coreos.com/v1alpha1  
    kind: AlertmanagerConfig  
    metadata:  
      name: runai  
      namespace: runai  
    labels:  
       alertmanagerConfig: runai  
    spec:  
       route:  
          continue: true  
          groupBy:   
          - alertname

          groupWait: 30s  
          groupInterval: 5m  
          repeatInterval: 1h

       matchers:  
       - matchType: =~  
         name: alertname  
         value: Runai.*

       receiver: email

    receivers:  
    - name: 'email'  
      emailConfigs:  
      - to: '<destination_email_address>'  
        from: '<from_email_address>'  
        smarthost: 'smtp.gmail.com:587'  
        authUsername: '<smtp_server_user_name>'  
        authPassword:  
          name: alertmanager-smtp-password
            key: password  
   ```
3. Save and exit the editor. The configuration is automatically reloaded.

### Third Party Alert Destinations

Prometheus AlertManager provides a structured way to connect to alert-management systems. There are built-in plugins for popular systems such as PagerDuty and OpsGenie, including a generic Webhook.

#### Example: Integrating NVIDIA Run:ai With a Webhook

1. Use [webhook.site](https://webhook.site/) to get a unique URL.
2. Use the upgrade cluster instructions to modify the values file:\
   Edit the values file to add the following, and replace `<WEB-HOOK-URL>` with the URL from [webhook.site](http://webhook.site):

   ```yaml
   codekube-prometheus-stack:  
     ...  
     alertmanager:  
       enabled: true  
       config:  
         global:  
           resolve_timeout: 5m  
         receivers:  
         - name: "null"  
         - name: webhook-notifications  
           webhook_configs:  
             - url: <WEB-HOOK-URL>  
               send_resolved: true  
         route:  
           group_by:  
           - alertname  
           group_interval: 5m  
           group_wait: 30s  
           receiver: 'null'  
           repeat_interval: 10m  
           routes:  
           - receiver: webhook-notifications
   ```
3. Verify that you are receiving alerts on the [webhook.site](https://webhook.site/), in the left pane:

![](/files/O4REjkWp0M66k3NIR4zC)

### Built-in Alerts

An NVIDIA Run:ai cluster comes with several built-in alerts. Each alert notifies on a specific functionality of a NVIDIA Run:ai’s entity. There is also a single, inclusive alert: `NVIDIA Run:ai Critical Problems`, which aggregates all component-based alerts into a single cluster health test.

#### NVIDIA Run:ai Agent Cluster Info Push Rate Low

| **Meaning**                    | The `cluster-sync` Pod in the `runai` namespace might not be functioning properly                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| ------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Impact**                     | Possible impact - no info/partial info from the cluster is being synced back to the control-plane                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| **Severity**                   | Critical                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                         |
| **Diagnosis**                  | `kubectl get pod -n runai` to see if the `cluster-sync` pod is running                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                           |
| **Troubleshooting/Mitigation** | <p>To diagnose issues with the <code>cluster-sync</code> pod, follow these steps:</p><ol><li><strong>Paste the following command to your terminal, to receive detailed information about the</strong> <code>cluster-sync</code> deployment:<code>kubectl describe deployment cluster-sync -n runai</code></li><li><strong>Check the Logs</strong>: Use the following command to view the logs of the <code>cluster-sync</code> deployment:<code>kubectl logs deployment/cluster-sync -n runai</code></li><li><strong>Analyze the Logs and Pod Details</strong>: From the information provided by the logs and the deployment details, attempt to identify the reason why the <code>cluster-sync</code> pod is not functioning correctly</li><li><strong>Check Connectivity</strong>: Ensure there is a stable network connection between the cluster and the NVIDIA Run:ai Control Plane. A connectivity issue may be the root cause of the problem.</li><li><strong>Contact Support</strong>: If the network connection is stable and you are still unable to resolve the issue, contact NVIDIA Run:ai support for further assistance</li></ol> |

#### NVIDIA Run:ai Agent Pull Rate Low

| **Meaning**                    | The `runai-agent` pod may be too loaded, is slow in processing data (possible in very big clusters), or the `runai-agent` pod itself in the `runai` namespace may not be functioning properly.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| ------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Impact**                     | Possible impact - no info/partial info from the control-plane is being synced in the cluster                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                              |
| **Severity**                   | Critical                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                  |
| **Diagnosis**                  | Run: `kubectl get pod -n runai` And see if the `runai-agent` pod is running.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                              |
| **Troubleshooting/Mitigation** | <p>To diagnose issues with the <code>runai-agent</code> pod, follow these steps:</p><ol><li><strong>Describe the Deployment</strong>: Run the following command to get detailed information about the <code>runai-agent</code> deployment:<code>kubectl describe deployment runai-agent -n runai</code></li><li><strong>Check the Logs</strong>: Use the following command to view the logs of the <code>runai-agent</code> deployment:<code>kubectl logs deployment/runai-agent -n runai</code></li><li><strong>Analyze the Logs and Pod Details</strong>: From the information provided by the logs and the deployment details, attempt to identify the reason why the <code>runai-agent</code> pod is not functioning correctly. There may be a connectivity issue with the control plane.</li><li><strong>Check Connectivity</strong>: Ensure there is a stable network connection between the <code>runai-agent</code> and the control plane. A connectivity issue may be the root cause of the problem.</li><li><strong>Consider Cluster Load</strong>: If the <code>runai-agent</code> appears to be functioning properly but the cluster is very large and heavily loaded, it may take more time for the agent to process data from the control plane.</li><li><strong>Adjust Alert Threshold</strong>: If the cluster load is causing the alert to fire, you can adjust the threshold at which the alert triggers. The default value is 0.05. You can try changing it to a lower value (e.g., 0.045 or 0.04). To edit the value, paste the following in your terminal:<code>kubectl edit runaiconfig -n runai/.</code> In the editor, navigate to: <code>spec: prometheus: agentPullPushRateMinForAlert</code> . If the <code>agentPullPushRateMinForAlert</code> value does not exist, add it under <code>spec -> prometheus</code> .</li></ol> |

#### NVIDIA Run:ai Container Memory Usage Critical

| **Meaning**                    | `Runai` container is using more than 90% of its Memory limit                                                                              |
| ------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------- |
| **Impact**                     | The container might run out of memory and crash.                                                                                          |
| **Severity**                   | Critical                                                                                                                                  |
| **Diagnosis**                  | Calculate the memory usage, this is performed by pasting the following to your terminal: `container_memory_usage_bytes{namespace=~"runai` |
| **Troubleshooting/Mitigation** | Add more memory resources to the container. If the issue persists, contact NVIDIA Run:ai                                                  |

#### NVIDIA Run:ai Container Memory Usage Warning

| **Meaning**                    | Runai container is using more than 80% of its memory limit                                                                               |
| ------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------- |
| **Impact**                     | The container might run out of memory and crash                                                                                          |
| **Severity**                   | Warning                                                                                                                                  |
| **Diagnosis**                  | Calculate the memory usage, this can be done by pasting the following to your terminal: `container_memory_usage_bytes{namespace=~"runai` |
| **Troubleshooting/Mitigation** | Add more memory resources to the container. If the issue persists, contact NVIDIA Run:ai                                                 |

#### NVIDIA Run:ai Container Restarting

| **Meaning**                    | `Runai` container has restarted more than twice in the last 10 min                                                                                                                                                                                                                                                  |
| ------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Impact**                     | The container might become unavailable and impact the NVIDIA Run:ai system                                                                                                                                                                                                                                          |
| **Severity**                   | Warning                                                                                                                                                                                                                                                                                                             |
| **Diagnosis**                  | To diagnose the issue and identify the problematic pods, paste this into your terminal: `kubectl get pods -n runai kubectl get pods -n runai-backend`One or more of the pods have a restart count >= 2.                                                                                                             |
| **Troubleshooting/Mitigation** | Paste this into your terminal:`kubectl logs -n NAMESPACE POD_NAME`Replace `NAMESPACE` and `POD_NAME` with the relevant pod information from the previous step. Check the logs for any standout issues and verify that the container has sufficient resources. If you need further assistance, contact NVIDIA Run:ai |

#### NVIDIA Run:ai CPU Usage Warning

| **Meaning**                    | `runai` container is using more than 80% of its CPU limit                                                                                  |
| ------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------ |
| **Impact**                     | This might cause slowness in the operation of certain NVIDIA Run:ai features.                                                              |
| **Severity**                   | Warning                                                                                                                                    |
| **Diagnosis**                  | Paste the following query to your terminal in order to calculate the CPU usage: `rate(container_cpu_usage_seconds_total{namespace=~"runai` |
| **Troubleshooting/Mitigation** | Add more CPU resources to the container. If the issue persists, please contact NVIDIA Run:ai.                                              |

#### NVIDIA Run:ai Critical Problem

| **Meaning**   | One of the critical NVIDIA Run:ai alerts is currently active                    |
| ------------- | ------------------------------------------------------------------------------- |
| **Impact**    | Impact is based on the active alert                                             |
| **Severity**  | Critical                                                                        |
| **Diagnosis** | Check NVIDIA Run:ai alerts in Prometheus to identify any active critical alerts |

#### Unknown State Alert for a Node

| **Meaning**   | The Kubernetes node hosting GPU workloads is in an unknown state, and its health and readiness cannot be determined.                                                                                                                                                                                                                                                         |
| ------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Impact**    | This may interrupt GPU workload scheduling and execution.                                                                                                                                                                                                                                                                                                                    |
| **Severity**  | <p><strong>Critical</strong> - Node is either unschedulable or has unknown status. The node is in one of the following states:</p><ul><li><code>Ready=Unknown</code>: The control plane cannot communicate with the node.</li><li><code>Ready=False</code>: The node is not healthy.</li><li><code>Unschedulable=True</code>: The node is marked as unschedulable.</li></ul> |
| **Diagnosis** | Check the node's status using kubectl describe node, verify Kubernetes API server connectivity, and inspect system logs for GPU-specific or node-level errors.                                                                                                                                                                                                               |

#### Low Memory Node Alert

| **Meaning**   | The Kubernetes node hosting GPU workloads has insufficient memory to support current or upcoming workloads.                                            |
| ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Impact**    | GPU workloads may fail to schedule, experience degraded performance, or crash due to memory shortages, disrupting dependent applications.              |
| **Severity**  | <p><strong>Critical</strong> - Node is using more than 90% of its memory.<br><strong>Warning</strong> - Node is using more than 80% of its memory.</p> |
| **Diagnosis** | Use `kubectl` top node to assess memory usage, identify memory-intensive pods, consider resizing the node or optimizing memory usage in affected pods. |

#### NVIDIA Run:ai DaemonSet Rollout Stuck / DaemonSet Unavailable on Nodes

| **Meaning**                    | There are currently 0 available pods for the `runai` daemonset on the relevant node                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                         |
| ------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Impact**                     | No fractional GPU workloads support                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                         |
| **Severity**                   | Critical                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| **Diagnosis**                  | Paste the following command to your terminal: `kubectl get daemonset -n runai-backend` In the result of this command, identify the daemonset(s) that don’t have any running pods                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| **Troubleshooting/Mitigation** | <p>Paste the following command to your terminal, where <code>daemonsetX</code> is the problematic daemonset from the previous step: <code>kubectl describe daemonsetX -n runai</code> on the relevant deamonset(s) from the previous step. The next step is to look for the specific error which prevents it from creating pods. Possible reasons might be:</p><ul><li><strong>Node Resource Constraints</strong>: The nodes in the cluster may lack sufficient resources (CPU, memory, etc.) to accommodate new pods from the daemonset.</li><li><strong>Node Selector or Affinity Rules</strong>: The daemonset may have node selector or affinity rules that are not matching with any nodes currently available in the cluster, thus preventing pod creation.</li></ul> |

#### NVIDIA Run:ai Deployment Insufficient Replicas /Deployment No Available Replicas /Deployment Unavailable Replicas

| **Meaning**                    | `Runai` deployment has one or more unavailable pods                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                               |
| ------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Impact**                     | When this happens, there may be scale issues. Additionally, new versions cannot be deployed, potentially resulting in missing features.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                           |
| **Severity**                   | Critical                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                          |
| **Diagnosis**                  | Paste the following commands to your terminal, in order to get the status of the deployments in the `runai` and `runai-backend` namespaces:`kubectl get deployment -n runai kubectl get deployment -n runai-backend`Identify any deployments that have missing pods. Look for discrepancies in the `DESIRED` and `AVAILABLE` columns. If the number of `AVAILABLE` pods is less than the `DESIRED` pods, it indicates that there are missing pods.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| **Troubleshooting/Mitigation** | <ul><li>Paste the following commands to your terminal, to receive detailed information about the problematic deployment:<code>kubectl describe deployment \<DEPLOYMENT\_NAME> -n runai kubectl describe deployment \<DEPLOYMENT\_NAME> -n runai-backend</code></li><li>Paste the following commands to your terminal, to check the replicaset details associated with the deployment:<code>kubectl describe replicaset \<REPLICASET\_NAME> -n runai kubectl describe replicaset \<REPLICASET\_NAME> -n runai-backend</code></li><li>Paste the following commands to your terminal to retrieve the logs for the deployment to identify any errors or issues:<code>kubectl logs deployment/\<DEPLOYMENT\_NAME> -n runai kubectl logs deployment/\<DEPLOYMENT\_NAME> -n runai-backend</code></li><li><p>From the logs and the detailed information provided by the <code>describe</code> commands, analyze the reasons why the deployment is unable to create pods. Look for common issues such as:</p><ul><li>Resource constraints (CPU, memory)</li><li>Misconfigured deployment settings or replicasets</li><li>Node selector or affinity rules preventing pod scheduling</li></ul><p>If the issue persists, contact NVIDIA Run:ai.</p></li></ul> |

#### NVIDIA Run:ai Project Controller Reconcile Failure

| Meaning                        | The `project-controller` in `runai` namespace had errors while reconciling projects                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| ------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Impact**                     | Some projects might not be in the “Ready” state. This means that they are not fully operational and may not have all the necessary components running or configured correctly.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| **Severity**                   | Critical                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                  |
| **Diagnosis**                  | Retrieve the logs for the `project-controller` deployment by pasting the following command in your terminal:`kubectl logs deployment/project-controller -n runai` Carefully examine the logs for any errors or warning messages. These logs help you understand what might be going wrong with the project controller.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| **Troubleshooting/Mitigation** | <p>Once errors in the log have been identified, follow these steps to mitigate the issue: The error messages in the logs should provide detailed information about the problem.</p><ol><li>Read through them to understand the nature of the issue. If the logs indicate which project failed to reconcile, you can further investigate by checking the status of that specific project.</li><li>Run the following command, replacing <code>\<PROJECT\_NAME></code> with the name of the problematic project:<code>kubectl get project \<PROJECT\_NAME> -o yaml</code></li><li>Review the status section in the YAML output. This section describes the current state of the project and provide insights into what might be causing the failure. If the issue persists, contact NVIDIA Run:ai.</li></ol> |

#### NVIDIA Run:ai StatefulSet Insufficient Replicas / StatefulSet No Available Replicas

| Meaning                        | `Runai` statefulset has no available pods                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| ------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Impact**                     | Absence of Metrics Database Unavailability                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                   |
| **Severity**                   | Critical                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| **Diagnosis**                  | <p>To diagnose the issue, follow these steps:</p><ol><li>Check the status of the stateful sets in the <code>runai-backend</code> namespace by running the following command:<code>kubectl get statefulset -n runai-backend</code></li><li>Identify any stateful sets that have no running pods. These are the ones that might be causing the problem.</li></ol>                                                                                                                                                                                                                                                                              |
| **Troubleshooting/Mitigation** | <p>Once you've identified the problematic stateful sets, follow these steps to mitigate the issue:</p><ol><li>Describe the stateful set to get detailed information on why it cannot create pods. Replace <code>X</code> with the name of the stateful set:<code>kubectl describe statefulset X -n runai-backend</code></li><li>Review the description output to understand the root cause of the issue. Look for events or error messages that explain why the pods are not being created.</li><li>If you're unable to resolve the issue based on the information gathered, contact NVIDIA Run:ai support for further assistance.</li></ol> |

### Adding a Custom Alert

You can add additional alerts from NVIDIA Run:ai. Alerts are triggered by using the Prometheus query language with any NVIDIA Run:ai metric.

To create an alert, follow these steps using Prometheus query language with NVIDIA Run:ai Metrics:

* **Modify Values File:** Use the upgrade cluster instructions to modify the values file.
* **Add Alert Structure:** Incorporate alerts according to the structure outlined below. Replace placeholders `<ALERT-NAME>`, `<ALERT-SUMMARY-TEXT>`, `<PROMQL-EXPRESSION>`, `<optional: duration s/m/h>`, and `<critical/warning>` with appropriate values for your alert, as described below:

  ```yaml
  kube-prometheus-stack:  
     additionalPrometheusRulesMap:  
       custom-runai:  
         groups:  
         - name: custom-runai-rules  
           rules:  
           - alert: <ALERT-NAME>  
             annotations:  
               summary: <ALERT-SUMMARY-TEXT>  
             expr:  <PROMQL-EXPRESSION>  
             for: <optional: duration s/m/h>  
             labels:  
               severity: <critical/warning>
  ```

  * `<ALERT-NAME>`: Choose a descriptive name for your alert, such as `HighCPUUsage` or `LowMemory`.
  * `<ALERT-SUMMARY-TEXT>`: Provide a brief summary of what the alert signifies, for example, `High CPU usage detected` or `Memory usage below threshold`.
  * `<PROMQL-EXPRESSION>`: Construct a Prometheus query (PROMQL) that defines the conditions under which the alert should trigger. This query should evaluate to a boolean value (`1` for alert, `0` for no alert).
  * `<optional: duration s/m/h>`: Optionally, specify a duration in seconds (`s`), minutes (`m`), or hours (`h`) that the alert condition should persist before triggering an alert. If not specified, the alert triggers as soon as the condition is met.
  * `<critical/warning>`: Assign a severity level to the alert, indicating its importance. Choose between `critical` for severe issues requiring immediate attention, or `warning` for less critical issues that still need monitoring.

You can find an example in the [Prometheus documentation](https://prometheus.io/docs/prometheus/latest/querying/examples/).


# Clusters

This guide explains the procedure to view and manage Clusters.

The Cluster table provides a quick and easy way to see the status of your cluster.

## Clusters Table

The Clusters table can be found under **Resources** in the NVIDIA Run:ai platform.

The clusters table provides a list of the clusters added to NVIDIA Run:ai platform, along with their status.

<figure><img src="/files/GxVu1Y1QmrAR17DZ0ELM" alt=""><figcaption></figcaption></figure>

The clusters table consists of the following columns:

| Column                        | Description                                                                                                                                                                                                                                                                                                                 |
| ----------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Cluster                       | The name of the cluster                                                                                                                                                                                                                                                                                                     |
| Kubernetes distribution       | The flavor of Kubernetes distribution                                                                                                                                                                                                                                                                                       |
| Kubernetes version            | The version of Kubernetes installed                                                                                                                                                                                                                                                                                         |
| Status                        | The status of the cluster. For more information see the [table below](#cluster-status). Hover over the information icon for a short description and links to troubleshooting                                                                                                                                                |
| Last connected                | <p>Indicates the most recent time the cluster successfully connected to the control plane.</p><ul><li>If the cluster is currently connected, the value is displayed as Now.</li><li>If the cluster is disconnected or has experienced issues, the exact timestamp of the last successful connection is displayed.</li></ul> |
| Network topologies            | The network topologies associated with the cluster                                                                                                                                                                                                                                                                          |
| Creation time                 | The timestamp when the cluster was created                                                                                                                                                                                                                                                                                  |
| URL                           | The URL that was given to the cluster                                                                                                                                                                                                                                                                                       |
| NVIDIA Run:ai cluster version | The NVIDIA Run:ai version installed on the cluster                                                                                                                                                                                                                                                                          |
| NVIDIA Run:ai cluster UUID    | The unique ID of the cluster                                                                                                                                                                                                                                                                                                |

### Cluster Status

| Status                | Description                                                                                                                                                                               |
| --------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Waiting to connect    | The cluster has never been connected.                                                                                                                                                     |
| Disconnected          | There is no communication from the cluster to the Control plane. This may be due to a network issue. [See troubleshooting scenarios.](#troubleshooting-scenarios)                         |
| Missing prerequisites | Some prerequisites are missing from the cluster. As a result, some features may be impacted. [See troubleshooting scenarios.](#troubleshooting-scenarios)                                 |
| Service issues        | At least one of the services is not working properly. You can view the list of nonfunctioning services for more information. [See troubleshooting scenarios.](#troubleshooting-scenarios) |
| Connected             | The NVIDIA Run:ai cluster is connected, and all NVIDIA Run:ai services are running.                                                                                                       |

### Network Topologies Associated with the Cluster

Click one of the values in the Network topologies column to view the list of network topologies and their parameters.

| Column        | Description                                                           |
| ------------- | --------------------------------------------------------------------- |
| Topology      | The name of the topology                                              |
| Labels        | The ordered set of node label keys that define the topology hierarchy |
| Node pools    | The node pool(s) the network topology is associated with              |
| Created by    | The user who created the network topology                             |
| Creation time | The timestamp of when the network topology was created                |

### Customizing the Table View

* Filter - Click ADD FILTER, select the column to filter by, and enter the filter values
* Search - Click SEARCH and type the value to search by
* Sort - Click each column header to sort by
* Column selection - Click COLUMNS and select the columns to display in the table
* Download table - Click MORE and then Click Download as CSV. Export to CSV is limited to 20,000 rows.

## Adding a New Cluster

Follow the setup and installation instructions below to get the installation instructions to install the NVIDIA Run:ai cluster.

### Setup

1. In the NVIDIA Run:ai UI, go to Resources -> Clusters
2. Enter a unique **name** for your cluster
3. Optional: Choose the NVIDIA Run:ai cluster version (latest, by default)
4. Enter the Cluster URL. For more information, see [Fully Qualified Domain Name](/saas/getting-started/installation/install-using-helm/system-requirements#fully-qualified-domain-name-fqdn) requirement.
5. Click **CONTINUE**

### Installation Instructions

1. Follow the installation instructions and run the commands provided on your Kubernetes cluster. To add a new cluster, see the [installation guide](/saas/getting-started/installation/install-using-helm/helm-install).
2. Click **DONE**

The cluster is displayed in the table with the status **Waiting to connect**. Once installation is complete, the cluster status changes to **Connected**.

## Managing Network Topologies

Network topologies optimize placement and accelerate distributed workloads by keeping pods on nodes that are as close to each other as possible in the network. For more details, see [Accelerating workloads with network topology-aware scheduling](/saas/platform-management/aiinitiatives/resources/topology-aware-scheduling).

To add topologies that represent the cluster's network:

1. Select the cluster you want to add a network topology for
2. In the top action bar, click **NETWORK TOPOLOGIES**
3. In the **Network Topologies Associated with \<Cluster Name>** modal, click **+ NETWORK TOPOLOGY**
4. Enter a unique **name** for the topology. If the name already exists, you will be requested to enter a different name.
5. Click **+ LABEL** to add the node label keys that represent the network hierarchy
   * Order labels from farthest (first) to closest (last)
   * Ensure the labels match the corresponding keys on the nodes. For example: `cloud.provider.com/topology-block`, `cloud.provider.com/topology-rack`, `kubernetes.io/hostname`
   * Drag labels to adjust their order if needed
6. Click **SAVE NETWORK TOPOLOGY**

After creating the network topology, the administrator must associate it with the relevant node pool(s). See [Node pools](/saas/platform-management/aiinitiatives/resources/node-pools) for more details:

* If you are using the default node pool (that is, no additional node pools are defined), attach the topology to the default node pool.
* If different node pools have different topologies, each node pool must be linked to its corresponding topology.
* If the entire cluster shares the same topology, link the same topology to all node pools.

## Removing a Cluster

1. Select the cluster you want to remove
2. Click **REMOVE**
3. A dialog appears: Make sure to carefully read the message before removing
4. Click **REMOVE** to confirm the removal.

## Using the API

Go to the [Clusters](https://run-ai-docs.nvidia.com/api/organizations/clusters) API reference to view the available actions

## Troubleshooting

Before starting, make sure you have access to the Kubernetes cluster where NVIDIA Run:ai is deployed with the necessary permissions

### Troubleshooting Scenarios

<details>

<summary>Cluster disconnected</summary>

**Description:** When the cluster's status is ‘disconnected’, there is no communication from the cluster services reaching the NVIDIA Run:ai Platform. This may be due to networking issues or issues with NVIDIA Run:ai services.

**Mitigation:**

1. Check NVIDIA Run:ai’s services status:
   1. Open your terminal
   2. Make sure you have access to the Kubernetes cluster with permissions to view pods
   3. Copy and paste the following command to verify that NVIDIA Run:ai’s services are running:

      ```bash
      kubectl get pods -n runai | grep -E 'runai-agent|cluster-sync|assets-sync'
      ```
   4. If any of the services are not running, see the ‘cluster has service issues’ scenario.
2. **Check the network connection**
   1. Open your terminal
   2. Make sure you have access to the Kubernetes cluster with permissions to create pods
   3. Copy and paste the following command to create a connectivity check pod:

      ```bash
      kubectl run control-plane-connectivity-check -n runai --image=wbitt/network-multitool --command -- /bin/sh -c 'curl -sSf <control-plane-endpoint> > /dev/null && echo "Connection Successful" || echo "Failed connecting to the Control Plane"'
      ```
   4. Replace `<control-plane-endpoint>` with the URL of the Control Plane in your environment. If the pod fails to connect to the Control Plane, check for potential network policies
3. **Check and modify the network policies**
   1. Open your terminal
   2. Copy and paste the following command to check the existence of network policies:

      ```bash
      kubectl get networkpolicies -n runai
      ```
   3. Review the policies to ensure that they allow traffic from the NVIDIA Run:ai namespace to the Control Plane. If necessary, update the policies to allow the required traffic

      Example of allowing traffic:

      ```yaml
      apiVersion: networking.k8s.io/v1
      kind: NetworkPolicy
      metadata:
        name: allow-control-plane-traffic
        namespace: runai
      spec:
        podSelector:
          matchLabels:
            app: runai
        policyTypes:
          - Ingress
          - Egress
        egress:
          - to:
              - ipBlock:
                  cidr: <control-plane-ip-range>
            ports:
              - protocol: TCP
                port: <control-plane-port>
        ingress:
          - from:
              - ipBlock:
                  cidr: <control-plane-ip-range>
            ports:
              - protocol: TCP
                port: <control-plane-port>
      ```
   4. Check infrastructure-level configurations:
      * Ensure that firewall rules and security groups allow traffic between your Kubernetes cluster and the Control Plane
      * Verify required ports and protocols:
        * Ensure that the necessary ports and protocols for NVIDIA Run:ai’s services are not blocked by any firewalls or security groups
4. **Check NVIDIA Run:ai services logs**
   1. Open your terminal
   2. Make sure you have access to the Kubernetes cluster with permissions to view logs
   3. Copy and paste the following commands to view the logs of the NVIDIA Run:ai services:

      ```bash
      kubectl logs deployment/runai-agent -n runai
      kubectl logs deployment/cluster-sync -n runai
      kubectl logs deployment/assets-sync -n runai
      ```
   4. Try to identify the problem from the logs. If you cannot resolve the issue, continue to the next step.
5. **Diagnosing internal network issues:**\
   NVIDIA Run:ai operates on Kubernetes, which uses its internal subnet and DNS services for communication between pods and services. If you find connectivity issues in the logs, the problem might be related to Kubernetes' internal networking.

   To diagnose DNS or connectivity issues, you can start a debugging {{glossary.Pod}} with networking utilities:

   1. Copy the following command to your terminal, to start a pod with networking tools:

      ```bash
      kubectl run -i --tty netutils --image=dersimn/netutils -- bash
      ```

      This command creates an interactive pod (`netutils`) where you can use networking commands like `ping`, `curl`, `nslookup`, etc., to troubleshoot network issues.
   2. Use this pod to perform network resolution tests and other diagnostics to identify any DNS or connectivity problems within your Kubernetes {{glossary.Cluster}}.
6. **Contact NVIDIA Run:ai’s support**
   * If the issue persists, [contact NVIDIA Run:ai’s support](https://www.nvidia.com/en-eu/support/enterprise/#contact-us) for assistance.

</details>

<details>

<summary>Cluster disconnected due to a faulty GPU node</summary>

**Description:** A single faulty GPU node can cause the entire cluster to show a **Disconnected** status in the NVIDIA Run:ai platform, even when all other nodes are healthy. This typically occurs when the node is stuck in GPU validation and the NVIDIA Run:ai system DaemonSets cannot complete initialization on it. This happens because the NVIDIA Run:ai platform evaluates cluster health based on all system DaemonSet pods being available. A single stuck pod causes the DaemonSet to report as unavailable, which cascades to the cluster-level health status.

Cordoning the node is not sufficient. A common first response is to cordon the faulty node using `kubectl cordon`. However, cordoning only prevents new pods from being scheduled on the node. It does not evict pods that already tolerate the `node.kubernetes.io/unschedulable` taint. The `runai-container-toolkit` DaemonSet tolerates this taint by default. As a result, its pod remains running (or stuck) on the cordoned node, and the cluster continues to report as Disconnected.

**Mitigation:** These steps require `cluster-admin` permissions to taint nodes. To isolate the faulty node and allow the cluster to recover:

1. Add a custom taint to the faulty node to prevent new DaemonSet pods from being scheduled on it:

   ```bash
   kubectl taint node <node-name> runai/gpu-fault=true:NoSchedule
   ```
2. Delete the stuck DaemonSet pod on the faulty node. The taint from step 1 prevents the DaemonSet controller from recreating it, which reduces the expected pod count and allows the DaemonSet to report as fully available:

   ```bash
   kubectl -n runai delete pod -l app=runai-container-toolkit --field-selector spec.nodeName=<node-name>
   ```
3. Verify the cluster status returns to **Connected** in the NVIDIA Run:ai platform UI.
4. Once the hardware issue on the node is resolved, remove the taint:

   ```bash
   kubectl taint node <node-name> runai/gpu-fault=true:NoSchedule-
   ```

{% hint style="info" %}
**Note**

If the node has no workloads worth preserving on its remaining healthy GPUs, a full drain and reboot (`kubectl drain <node-name> --ignore-daemonsets --delete-emptydir-data`) is an alternative approach. This resolves the underlying hardware fault directly and the taint workaround is not needed.
{% endhint %}

</details>

<details>

<summary>Cluster has service issues</summary>

**Description:** When a cluster's status is ‘Has service issues\`, it means that one or more NVIDIA Run:ai services running in the cluster are not available.

**Mitigation:**

1. **Verify non-functioning services**
   1. Open your terminal
   2. Make sure you have access to the Kubernetes cluster with permissions to view the `runaiconfig` resource
   3. Copy and paste the following command to determine which services are not functioning:

      ```bash
      kubectl get runaiconfig -n runai runai -ojson | jq -r '.status.conditions | map(select(.type == "Available"))'
      ```
2. **Check for Kubernetes events**
   1. Open your terminal
   2. Make sure you have access to the Kubernetes cluster with permissions to view events
   3. Copy and paste the following command to get all [Kubernetes events](https://kubernetes.io/docs/reference/kubernetes-api/cluster-resources/event-v1/):

      ```bash
      kubectl get events -A
      ```
3. **Inspect resource details**
   1. Open your terminal
   2. Make sure you have access to the Kubernetes cluster with permissions to describe resources
   3. Copy and paste the following command to check the details of the required resource:

      ```bash
      kubectl describe <resource_type> <name>
      ```
4. **Contact NVIDIA Run:ai’s Support**
   * If the issue persists, contact [contact NVIDIA Run:ai’s support](https://www.nvidia.com/en-eu/support/enterprise/#contact-us) for assistance.

</details>

<details>

<summary>Cluster is waiting to connect</summary>

**Description:** When the cluster's status is ‘waiting to connect’, it means that no communication from the cluster services reaches the NVIDIA Run:ai Platform. This may be due to networking issues or issues with NVIDIA Run:ai services.

**Mitigation:**

1. **Check NVIDIA Run:ai’s services status**
   1. Open your terminal
   2. Make sure you have access to the Kubernetes cluster with permissions to view pods
   3. Copy and paste the following command to verify that NVIDIA Run:ai’s services are running:

      ```bash
      kubectl get pods -n runai | grep -E 'runai-agent|cluster-sync|assets-sync'
      ```
   4. If any of the services are not running, see the ‘cluster has service issues’ scenario.
2. **Check the network connection**
   1. Open your terminal
   2. Make sure you have access to the Kubernetes cluster with permissions to create pods
   3. Copy and paste the following command to create a connectivity check pod:

      ```bash
      kubectl run control-plane-connectivity-check -n runai --image=wbitt/network-multitool --command -- /bin/sh -c 'curl -sSf <control-plane-endpoint> > /dev/null && echo "Connection Successful" || echo "Failed connecting to the Control Plane"'
      ```
   4. Replace `<control-plane-endpoint>` with the URL of the Control Plane in your environment. If the pod fails to connect to the Control Plane, check for potential network policies:
3. **Check and modify the network policies**
   1. Open your terminal
   2. Copy and paste the following command to check the existence of network policies:

      ```bash
      kubectl get networkpolicies -n runai
      ```
   3. Review the policies to ensure that they allow traffic from the NVIDIA Run:ai namespace to the Control Plane. If necessary, update the policies to allow the required traffic
   4. Example of allowing traffic:

      ```yaml
      apiVersion: networking.k8s.io/v1
      kind: NetworkPolicy
      metadata:
        name: allow-control-plane-traffic
        namespace: runai
      spec:
        podSelector:
          matchLabels:
            app: runai
        policyTypes:
          - Ingress
          - Egress
        egress:
          - to:
              - ipBlock:
                  cidr: <control-plane-ip-range>
            ports:
              - protocol: TCP
                port: <control-plane-port>
        ingress:
          - from:
              - ipBlock:
                  cidr: <control-plane-ip-range>
            ports:
              - protocol: TCP
                port: <control-plane-port>
      ```
   5. Check infrastructure-level configurations:
   6. Ensure that firewall rules and security groups allow traffic between your Kubernetes cluster and the Control Plane
   7. Verify required ports and protocols:
      * Ensure that the necessary ports and protocols for NVIDIA Run:ai’s services are not blocked by any firewalls or security groups
4. **Check NVIDIA Run:ai services logs**
   1. Open your terminal
   2. Make sure you have access to the Kubernetes cluster with permissions to view logs
   3. Copy and paste the following commands to view the logs of the NVIDIA Run:ai services:

      ```bash
      kubectl logs deployment/runai-agent -n runai
      kubectl logs deployment/cluster-sync -n runai
      kubectl logs deployment/assets-sync -n runai
      ```
   4. Try to identify the problem from the logs. If you cannot resolve the issue, continue to the next step
5. **Contact NVIDIA Run:ai’s support**
   * If the issue persists, [contact NVIDIA Run:ai’s support](https://www.nvidia.com/en-eu/support/enterprise/#contact-us) for assistance

</details>

<details>

<summary>Cluster is missing prerequisites</summary>

**Description:** When a cluster's status displays Missing prerequisites, it indicates that at least one of the Mandatory Prerequisites has not been fulfilled. In such cases, NVIDIA Run:ai services may not function properly.

**Mitigation:**

If you have ensured that all prerequisites are installed and the status still shows *missing prerequisites*, follow these steps:

1. Check the message in the NVIDIA Run:ai platform for further details regarding the missing prerequisites.
2. **Inspect the** `runai-public` ConfigMap:
   1. Open your terminal. In the terminal, type the following command to list all ConfigMaps in the `runai` namespace:

      ```bash
      kubectl get configmap runai-public -n runai
      ```
3. **Describe the ConfigMap**
   1. Locate the ConfigMap named `runai-public` from the list
   2. To view the detailed contents of this ConfigMap, type the following command:

      ```bash
      kubectl describe configmap runai-public -n runai
      ```
4. **Find Missing Prerequisites**
   1. In the output displayed, look for a section labeled `dependencies.required`
   2. This section provides detailed information about any missing resources or prerequisites. Review this information to identify what is needed
5. **Contact NVIDIA Run:ai’s support**
   * If the issue persists, [contact NVIDIA Run:ai’s support](https://www.nvidia.com/en-eu/support/enterprise/#contact-us) for assistance

</details>


# Shared Storage

Shared storage is a critical component in AI and machine learning workflows, particularly in scenarios involving distributed training and shared datasets. In AI and ML environments, data must be readily accessible across multiple nodes, especially when training large models or working with vast datasets. Shared storage enable seamless access to data, ensuring that all nodes in a distributed training setup can read and write to the same datasets simultaneously. This setup not only enhances efficiency but is also crucial for maintaining consistency and speed in high-performance computing environments.

While NVIDIA Run:ai Platform supports a variety of remote data sources, such as Git and S3, it is often more efficient to keep data close to the compute resources. This proximity is typically achieved through the use of shared storage, accessible to multiple nodes in your Kubernetes cluster.

## Shared Storage

When implementing shared storage in Kubernetes, there are two primary approaches:

* Utilizing the [Kubernetes Storage Classes](https://kubernetes.io/docs/concepts/storage/storage-classes/) of your storage provider (Recommended)
* Using a direct NFS (Network File System) mount

NVIDIA Run:ai [Data Sources](/saas/workloads-in-nvidia-run-ai/assets/datasources) support both direct NFS mount and Kubernetes Storage Classes.

### Kubernetes Storage Classes

Storage classes in Kubernetes defines how storage is provisioned and managed. This allows you to select storage types optimized for AI workloads. For example, you can choose storage with high IOPS (Input/Output Operations Per Second) for rapid data access during intensive training sessions, or tiered storage options to balance cost and performance-based on your organization’s requirements. This approach supports dynamic provisioning, enabling storage to be allocated on-demand as required by your applications.

NVIDIA Run:ai data sources such as [Persistent Volume Claims (PVC)](/saas/workloads-in-nvidia-run-ai/assets/datasources#pvc) and [Data Volumes](/saas/workloads-in-nvidia-run-ai/assets/data-volumes) leverage storage class to manage and allocate storage efficiently. This ensures that the most suitable storage option is always accessible, contributing to the efficiency and performance of AI workloads.

{% hint style="info" %}
**Note**

NVIDIA Run:ai lists all available storage classes in the Kubernetes cluster, making it easy for users to select the appropriate storage. Additionally, [policies](/saas/platform-management/policies/policies-and-rules) can be set to restrict or enforce the use of specific storage classes, to help maintain compliance with organizational standards and optimize resource utilization.
{% endhint %}

### Direct NFS Mount

Direct NFS allows you to mount a shared file system directly across multiple nodes in your Kubernetes cluster. This method provides a straightforward way to share data among nodes and is often used for simple setups or when a dedicated NFS server is available.

However, using NFS can present challenges related to security and control. Direct NFS setups might lack the fine-grained control and security features available with storage class.


# Nodes Maintenance

This guide provides detailed instructions on how to manage both planned and unplanned node downtimes in a Kubernetes cluster running NVIDIA Run:ai. It covers all the steps to maintain service continuity and ensure the proper handling of workloads during these events.

## Prerequisites

* **Access to Kubernetes cluster** - Administrative access to the Kubernetes cluster, including permissions to run `kubectl` commands
* **Basic knowledge of Kubernetes** - Familiarity with Kubernetes concepts such as nodes, taints, and workloads
* **NVIDIA Run:ai installation** - The [NVIDIA Run:ai software installed](/saas/getting-started/installation/install-using-helm/helm-install) and configured within your Kubernetes cluster
* **Node naming conventions** - Know the names of the nodes within your cluster, as these are required when executing the commands

## Node Types

This section distinguishes between two types of nodes within a NVIDIA Run:ai installation:

* **Worker nodes** - Nodes on which AI practitioners can submit and run workloads
* **NVIDIA Run:ai system nodes** - Nodes on which the NVIDIA Run:ai software runs, managing the cluster's operations

### Worker Nodes

Worker nodes are responsible for running workloads. When a worker node goes down, either due to planned maintenance or unexpected failure, workloads ideally migrate to other available nodes or wait in the queue to be executed when possible.

#### Training vs. Interactive Workloads

The following workload types can run on worker nodes:

* **Training workloads** - These are long-running processes that, in case of node downtime, can automatically move to another node.
* **Interactive workloads** - These are short-lived, interactive processes that require manual intervention to be relocated to another node.

{% hint style="info" %}
**Note**

While training workloads can be automatically migrated, it is recommended to plan maintenance and manually manage this process for a faster response, as it may take time for Kubernetes to detect a node failure.
{% endhint %}

#### Planned Maintenance

Before stopping a worker node for maintenance, perform the following steps:

1. **Prevent new workloads on the node**

   To stop the Kubernetes Scheduler from assigning new workloads to the node and to safely remove all existing workloads, copy the following command to your terminal:

   ```bash
   kubectl taint nodes <node-name> runai=drain:NoExecute
   ```

   * `<node-name>`\
     Replace this placeholder with the actual name of the node you want to drain
   * `kubectl taint nodes`\
     This command is used to add a taint to the node, which prevents any new pods from being scheduled on it
   * `runai=drain:NoExecute`\
     This specific taint ensures that all existing pods on the node are evicted and rescheduled on other available nodes, if possible

   **Result:** The node stops accepting new workloads, and existing workloads either migrate to other nodes or are placed in a queue for later execution.
2. **Shut down and perform maintenance**

   After draining the node, you can safely shut it down and perform the necessary maintenance tasks.
3. **Restart the node**

   Once maintenance is complete and the node is back online, remove the taint to allow the node to resume normal operations. Copy the following command to your terminal:

   ```bash
   kubectl taint nodes <node-name> runai=drain:NoExecute-
   ```

   `runai=drain:NoExecute-`\
   The `-` at the end of the command indicates the removal of the taint. This allows the node to start accepting new workloads again.

   **Result:** The node rejoins the cluster's pool of available resources, and workloads can be scheduled on it as usual.

#### Unplanned Downtime

In the event of unplanned downtime:

1. **Automatic restart**\
   If a node fails but immediately restarts, all services and workloads automatically resume.
2. **Extended downtime**

   If the node remains down for an extended period, drain the node to migrate workloads to other nodes. Copy the following command to your terminal:

   ```bash
   kubectl taint nodes <node-name> runai=drain:NoExecute
   ```

   The command works the same as in the planned maintenance section, ensuring that no workloads remain scheduled on the node while it is down.
3. **Reintegrate the node**

   Once the node is back online, remove the taint to allow it to rejoin the cluster's operations. Copy the following command to your terminal:

   ```bash
   kubectl taint nodes <node-name> runai=drain:NoExecute- 
   ```

   **Result:** This action reintegrates the node into the cluster, allowing it to accept new workloads.
4. **Permanent shutdown**

   If the node is to be permanently decommissioned, remove it from Kubernetes with the following command:

   ```bash
   kubectl delete node <node-name>
   ```

   * `kubectl delete node`\
     This command completely removes the node from the cluster
   * `<node-name>`\
     Replace this placeholder with the actual name of the node

   **Result:** The node is no longer part of the Kubernetes cluster. If you plan to bring the node back later, it must be rejoined to thegl cluster using the steps outlined in the next section.

### NVIDIA Run:ai System Nodes

In a production environment, the services responsible for scheduling, submitting and managing NVIDIA Run:ai workloads operate on one or more NVIDIA Run:ai system nodes. It is recommended to have more than one system node to ensure [high availability](/saas/infrastructure-setup/procedures/high-availability). If one system node goes down, another can take over, maintaining continuity. If a second system node does not exist, you must designate another node in the cluster as a temporary NVIDIA Run:ai system node to maintain operations.

The protocols for handling planned maintenance and unplanned downtime are identical to those for worker nodes. Refer to the above section for detailed instructions.

## Rejoining a Node Into the Kubernetes Cluster

To rejoin a node to the Kubernetes cluster, follow these steps:

1. **Generate a join command on the master node**

   On the master node, copy the following command to your terminal:

   ```bash
   kubeadm token create --print-join-command
   ```

   * `kubeadm token create`\
     This command generates a token that can be used to join a node to the Kubernetes cluster.
   * `--print-join-command`\
     This option outputs the full command that needs to be run on the worker node to rejoin it to the cluster.

   **Result:** The command outputs a `kubeadm join` command.
2. **Run the join command on the worker node**

   Copy the `kubeadm join` command generated from the previous step and run it on the worker node that needs to rejoin the cluster.

   ```bash
   kubeadm join <master-ip>:<master-port> 
   --token <token> \ --discovery-token-ca-cert-hash sha256:<hash>
   ```

   The `kubeadm join` command re-enrolls the node into the cluster, allowing it to start participating in the cluster's workload scheduling.
3. **Verify node rejoining**

   Verify that the node has successfully rejoined the cluster by running:

   ```bash
   kubectl get nodes
   ```

   * `kubectl get nodes`\
     This command lists all nodes currently part of the Kubernetes cluster, along with their status

   **Result:** The rejoined node should appear in the list with a status of Ready
4. **Re-label nodes**

   Once the node is ready, ensure it is labeled according to its role within the cluster.


# Cluster Restore

This section explains how to restore a NVIDIA Run:ai cluster on a different Kubernetes environment.

In the event of a critical Kubernetes failure or alternatively, if you want to migrate a NVIDIA Run:ai cluster to a new Kubernetes environment, simply reinstall the NVIDIA Run:ai cluster. Once you have reinstalled and reconnected the cluster, projects, workloads and other cluster data are synced automatically.

The restoration or backup of NVIDIA Run:ai [advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config) which are stored locally on the Kubernetes cluster is optional and can be restored and backed up separately.

## Back Up the Cluster

As back-up of data is not required, the backup procedure is optional for advanced deployments, as explained above.

### Save Cluster Configurations

To back up the NVIDIA Run:ai cluster configurations, you should save both the Helm values and the runtime configuration (`runaiconfig`).

1. **Back up Helm values** - Run the following command to export the Helm values used for deployment:

   ```bash
   helm get values runai-cluster -n runai > runai_cluster_values_backup.yaml
   ```
2. **Back up the runtime configuration (`runaiconfig`)** - Run the following command to export the active runtime configuration:

   ```bash
   kubectl get runaiconfig runai -n runai -o yaml -o=jsonpath='{.spec}' > runaiconfig_backup.yaml
   ```
3. Save both backup files (`runai_cluster_values_backup.yaml` and `runaiconfig_backup.yaml`) externally so they can be retrieved later if needed.

## Restore the Cluster

Follow the steps below to restore the NVIDIA Run:ai cluster on a new Kubernetes environment.

### Prerequisites

Before restoring the NVIDIA Run:ai cluster, it is essential to validate that it is both disconnected and uninstalled.

1. If the Kubernetes cluster is still available, [uninstall](/saas/getting-started/installation/install-using-helm/uninstall) the NVIDIA Run:ai cluster. Make sure not to remove the cluster from the control plane.
2. Navigate to the **Clusters** grid in the NVIDIA Run:ai UI
3. Locate the cluster and verify its status is **Disconnected**

### Re-install the Cluster

1. Follow the NVIDIA Run:ai cluster [installation](/saas/getting-started/installation/install-using-helm/helm-install) instructions and ensure all [prerequisites](/saas/getting-started/installation/install-using-helm/system-requirements) are met.
2. If you have a backup of the cluster configurations, reload it once the installation is complete:

   ```bash
   kubectl apply -f runaiconfig_backup.yaml -n runai
   ```
3. Navigate to the **Clusters** grid in the NVIDIA Run:ai UI
4. Locate the cluster and verify its status is **Connected**

### Restore Namespace and RoleBindings

If your cluster configuration disables automatic namespace creation for projects, you must manually:

* Re-create each project namespace
* Reapply the required role bindings for access control

For more information, see [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config).


# Secure Your Cluster

This guide details the security considerations for deploying NVIDIA Run:ai. It is intended to help administrators and security officers understand the specific permissions required by NVIDIA Run:ai.

## Access to the Kubernetes Cluster

NVIDIA Run:ai integrates with Kubernetes clusters and requires specific permissions to successfully operate. These are permissions are controlled with configuration flags that dictate how NVIDIA Run:ai interacts with cluster resources. Prior to installation, security teams can review the permissions and ensure it aligns with their organization’s policies.

### Permissions and their Related Use Case

NVIDIA Run:ai provides various security-related permissions that can be customized to fit specific organizational needs. Below are brief descriptions of the key use cases for these customizations:

| Permission                       | Use case                                                                                                                                                                            |
| -------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Automatic Namespace creation     | Controls whether NVIDIA Run:ai automatically creates Kubernetes namespaces when new projects are created. Useful in environments where namespace creation must be strictly managed. |
| Automatic user assignment        | Decides if users are automatically assigned to projects within NVIDIA Run:ai. Helps manage user access more tightly in certain compliance-driven environments.                      |
| Secret propagation               | Determines whether NVIDIA Run:ai should propagate secrets across the cluster. Relevant for organizations with specific security protocols for managing sensitive data.              |
| Disabling Kubernetes limit range | Chooses whether to disable the Kubernetes Limit Range feature. May be adjusted in environments with specific resource management needs.                                             |

{% hint style="info" %}
**Note**

These security customizations allow organizations to tailor NVIDIA Run:ai to their specific needs. All changes should be modified cautiously and only when necessary to meet particular security, compliance or operational requirements.
{% endhint %}

## Secure Installation

Many organizations enforce IT compliance rules for Kubernetes, with strict access control for installing and running workloads. OpenShift uses Security Context Constraints (SCC) for this purpose. NVIDIA Run:ai fully supports SCC, ensuring integration with OpenShift's security requirements.

## Security Vulnerabilities

The platform is actively monitored for security vulnerabilities, with regular scans conducted to identify and address potential issues. Necessary fixes are applied to ensure that the software remains secure and resilient against emerging threats, providing a safe and reliable experience.


# Compliance

This guide details the data privacy and compliance considerations for deploying NVIDIA Run:ai. It is intended to help administrators and compliance teams understand the data management practices involved with NVIDIA Run:ai. This ensures the permissions align with organizational policies and regulatory requirements before installation and during integration and onboarding of the various teams.

## Data Privacy

When using the NVIDIA Run:ai SaaS cluster, the Control plane operates through the NVIDIA Run:ai cloud, requiring the transmission of certain data for control and analytics. Below is a detailed breakdown of the specific data sent to the NVIDIA Run:ai cloud in the SaaS offering.

{% hint style="info" %}
**Note**

For organizations where data privacy policies do not align with this data transmission, NVIDIA Run:ai offers a self-hosted version. This version includes the control plane on premise and does not communicate with the cloud.
{% endhint %}

## Data Sent to the NVIDIA Run:ai Cloud

| Asset                  | Details                                                                                                                  |
| ---------------------- | ------------------------------------------------------------------------------------------------------------------------ |
| Workload Metrics       | Includes workload names, CPU, GPU, and memory metrics, as well as parameters provided during the `runai submit` command. |
| Workload Assets        | Covers environments, compute resources, and data resources associated with workloads.                                    |
| Resource Credentials   | Credentials for cluster resources, encrypted with a SHA-512 algorithm specific to each tenant.                           |
| Node Metrics           | Node-specific data including names, IPs, and performance metrics (CPU, GPU, memory).                                     |
| Cluster Metrics        | Cluster-wide metrics such as names, CPU, GPU, and memory usage.                                                          |
| Projects & Departments | Includes names and quota information for projects and departments.                                                       |
| Users                  | User roles within NVIDIA Run:ai, email addresses, and passwords.                                                         |

## Key Consideration

NVIDIA Run:ai ensures that no deep-learning artifacts, such as code, images, container logs, training data, models, or checkpoints, are transmitted to the cloud. These assets remain securely within your organization's firewalls, safeguarding sensitive intellectual property and data.


# Logs Collection

This guide provides instructions for IT administrators on collecting NVIDIA Run:ai logs for support, including prerequisites, CLI commands, and log file retrieval. It also covers enabling verbose logging for Prometheus and the NVIDIA Run:ai Scheduler.

## Collect Logs to Send to Support

To collect NVIDIA Run:ai logs, follow these steps:

### Prerequisites

* Ensure that you have administrator-level access to the Kubernetes cluster where NVIDIA Run:ai is installed.
* The NVIDIA Run:ai [CLI](/saas/reference/cli/install-cli) must be installed.

### Step-by-step Instructions

1. Open a terminal on any machine where the NVIDIA Run:ai CLI is installed and configured with access to the Kubernetes cluster. The user running this command typically requires system administrator [permissions](/saas/infrastructure-setup/authentication/roles), including access to pods and namespaces.
2. Collect the Logs. Execute the following command to collect the logs. See the [CLI commands reference](/saas/reference/cli/runai/runai-diagnostics) for more details:

   ```bash
   runai diagnostics collect-logs
   ```

This command gathers diagnostic logs from your Kubernetes cluster to facilitate troubleshooting or support requests with NVIDIA Run:ai Support. The following optional flags are available:

| Flag                  | Description                                                                                                                                                                         |
| --------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--output-dir <path>` | Directory where the log archive is saved. Defaults to the directory from which the command is run.                                                                                  |
| `--namespaces <list>` | Comma-separated list of namespaces to collect logs from. Defaults to `runai`, `runai-reservation`, `runai-backend`, `training-operator`, `knative`, `gpu-operator`, `nim-operator`. |
| `--no-previous`       | Excludes previous pod logs from the collection.                                                                                                                                     |

3. After the command completes, the CLI displays the path of the generated compressed log file. Send this file to NVIDIA Run:ai Support for troubleshooting.

{% hint style="info" %}
**Note**

The command collects diagnostic logs from Kubernetes namespaces in a NVIDIA Run:ai installation. By default, logs are collected from the following namespaces: `runai`, `runai-reservation`, `runai-backend`, `training-operator`, `knative`, `gpu-operator`, and `nim-operator`. Use the `--namespaces` flag to target specific namespaces.
{% endhint %}

## Logs Verbosity

Increase log verbosity to capture more detailed information, providing deeper insights into system behavior and make it easier to identify and resolve issues.

### Prerequisites

Before you begin, ensure you have the following:

* Access to the Kubernetes cluster where NVIDIA Run:ai is installed
  * Including [necessary permissions](/saas/infrastructure-setup/authentication/roles) to view and modify configurations.
* kubectl installed and configured:
  * The Kubernetes command-line tool, `kubectl`, must be installed and configured to interact with the cluster.
  * Sufficient privileges to edit configurations and view logs.
* Monitoring Disk Space
  * When enabling verbose logging, ensure adequate disk space to handle the increased log output, especially when enabling debug or high verbosity levels.

### Adding Verbosity

<details>

<summary>Adding verbosity to Prometheus</summary>

To increase the logging verbosity for Prometheus, follow these steps:

1. Edit the `RunaiConfig` to adjust Prometheus log levels. Copy the following command to your terminal:

   ```bash
   kubectl edit runaiconfig runai -n runai
   ```
2. In the configuration file that opens, add or modify the following section to set the log level to `debug`:

   ```bash
   spec:
     prometheus:
       spec:
         logLevel: debug
   ```
3. Save the changes. To view the Prometheus logs with the new verbosity level, run:

   ```bash
   kubectl logs -n runai prometheus-runai-0 
   ```

   This command streams the last 100 lines of logs from Prometheus, providing detailed information useful for debu

</details>

<details>

<summary>Adding verbosity to the Scheduler</summary>

To enable extended logging for the NVIDIA Run:ai scheduler:

1. Edit the `RunaiConfig` to adjust scheduler verbosity:

   ```bash
   kubectl edit runaiconfig runai -n runai
   ```
2. Add or modify the following section under the scheduler settings:

   ```bash
   runai-scheduler:
     args:
       verbosity: 6
   ```

   This increases the verbosity level of the scheduler logs to provide more detailed output.

**Warning:** Enabling verbose logging can significantly increase disk space usage. Monitor your storage capacity and adjust the verbosity level as necessary.

</details>


# Event History

The NVIDIA Run:ai control plane provides the audit log API and event history table in the NVIDIA Run:ai UI. Both reflect the same information regarding changes to business objects: clusters, projects and assets etc.

{% hint style="info" %}
**Note**

Only system administrator users with tenant-wide permissions can access Audit log.
{% endhint %}

## Event History Table

The Event history table can be found under Event history in the NVIDIA Run:ai UI.

<figure><img src="/files/sbWhGUQoWqOeSGAaAvg8" alt=""><figcaption></figcaption></figure>

The Event history table consists of the following columns:

| Column       | Description                                                                                                                                                                                              |
| ------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Subject      | The name of the subject                                                                                                                                                                                  |
| Subject type | The user or service account assigned with the role                                                                                                                                                       |
| Source IP    | The IP address of the subject                                                                                                                                                                            |
| Date & time  | The exact timestamp at which the event occurred. Format dd/mm/yyyy for date and hh:mm am/pm for time.                                                                                                    |
| Event        | The type of the event. Possible values: Create, Update, Delete, Login, Password reset, Password set                                                                                                      |
| Event ID     | Internal event ID, can be used for support purposes                                                                                                                                                      |
| Status       | The outcome of the logged operation. Possible values: Succeeded, Failed                                                                                                                                  |
| Entity type  | The type of the logged business object.                                                                                                                                                                  |
| Entity name  | The name of logged business object.                                                                                                                                                                      |
| Entity ID    | The system's internal id of the logged business object.                                                                                                                                                  |
| *URL*        | The endpoint or address that was accessed during the logged event.                                                                                                                                       |
| HTTP Method  | The HTTP operation method used for the request. Possible values include standard HTTP methods such as `GET`, `POST`, `PUT`, `DELETE`, indicating what kind of action was performed on the specified URL. |

### Customizing the Table View

* Filter - Click ADD FILTER, select the column to filter by, and enter the filter values
* Search - Click SEARCH and type the value to search by
* Sort - Click each column header to sort by
* Column selection - Click COLUMNS and select the columns to display in the table
* Download table - Click MORE and then Click Download as CSV or Download as JSON

## Using the Event History Date Selector

The Event history table saves events for the last 90 days. However, the table itself presents up to the last 30 days of information due to the potentially very high number of operations that might be logged during this period.

![](/files/qAmu2IiaW1BG7qQme8ZG)

To view older events, or to refine your search for more specific results or fewer results, use the time selector and change the period you search for.\
You can also refine your search by clicking and using ADD FILTER accordingly.

## Using API

Go to the [Audit log](https://run-ai-docs.nvidia.com/api/audit/auditlogs) API reference to view the available actions.\
Since the amount of data is not trivial, the API is based on paging. It retrieves a specified number of items for each API call. You can get more data by using subsequent calls.

## Limitations

Submissions of workloads are not audited. As a result, the system does not track or log details of workload submissions, such as timestamps or user activity.


# Manage AI Initiatives


# Adapting AI Initiatives to Your Organization

AI initiatives refer to advancing research, development, and implementation of AI technologies. These initiatives represent your business needs and involve collaboration between individuals, teams, and other stakeholders. AI initiatives require compute resources and a methodology to effectively and efficiently use those compute resources and split them among the different AI initiatives stakeholders. The building blocks of AI [compute resources](/saas/workloads-in-nvidia-run-ai/assets/compute-resources) are GPUs, CPUs, and memory, which are built into [nodes](/saas/platform-management/aiinitiatives/resources/nodes) (servers) and can be further grouped into [node pools](/saas/platform-management/aiinitiatives/resources/node-pools). Nodes and node pools are part of a Kubernetes cluster.

To manage AI initiatives in NVIDIA Run:ai you should:

* Map your organization and initiatives to projects and optionally departments
* Map compute resources (node pools and quotas) to projects and optionally departments
* Assign users (e.g. AI practitioners, ML engineers, Admins) to projects and departments

## Mapping Your Organization

The way you map your AI initiatives and organization into NVIDIA Run:ai [projects](/saas/platform-management/aiinitiatives/organization/projects) and [departments](/saas/platform-management/aiinitiatives/organization/departments) should reflect your organization’s structure and Project management practices. There are multiple options, and we provide you here with 3 examples of typical forms in which to map your organization, initiatives, and users into NVIDIA Run:ai, but of course, other ways that suit your requirements are also acceptable.

### Based on Individuals

A typical use case would be students (individual practitioners) within a faculty (business unit) - an individual practitioner may be involved in one or more initiatives. In this example, the resources are accounted for by the student (project) and aggregated per faculty (department).

Department = business unit / Project = individual practitioner

![](/files/934cSTZ007X3Bn628m2U)

### Based on Business Units

A typical use case would be an AI service (business unit) split into AI capabilities (initiatives) - an individual practitioner may be involved in several initiatives. In this example, the resources are accounted for by Initiative (project) and aggregated per AI service (department).

Department = business unit / Project = initiative

![](/files/UcJZJrZGU7BpSIMrZZf5)

### Based on the Organizational Structure

A typical use case would be a business unit split into teams - an individual practitioner is involved in a single team (project) but the team may be involved in several AI initiatives. In this example, the resources are accounted for by team (project) and aggregated per business unit (department).

Department = business unit / Project = team

![](/files/QaUVp1BW3fPibMN474fZ)

## Mapping Your Resources

AI initiatives require compute resources such as GPUs and CPUs to run. Compute resources in any organization are limited, either due to the number of servers (nodes) owned by the organization is limited, the budget it can spend to lease resources in the cloud or spending for in-house servers is also limited. Every organization strives to optimize the usage of its resources by maximizing their utilization and providing all users with their needs. Therefore, the organization needs to split resources according to the organization's internal priorities and budget constraints. But even after splitting the resources, the orchestration layer should still provide fairness between the resourced consumers, and allow access to unused resources to minimize scenarios of idle resources.

Another aspect of resource management is how to group your resources effectively, especially in large environments, or environments that are made of heterogeneous types of hardware, where some users need to use specific hardware types, or where other users should avoid occupying critical hardware of some users or initiatives.

NVIDIA Run:ai assists you with all of these complex issues by allowing you to map your cluster resources to node pools, then map each Project and Department a quota allocation per node pool, and set access rights to unused resources ([over quota](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles#over-quota)) per node pool.

### Grouping Your Resources

There are several reasons why you would group resources (nodes) into node pools:

* **Control the GPU type to use in heterogeneous hardware environment** - in many cases, AI models can be optimized per hardware type they will use, e.g. a training workload that is optimized for H100 does not necessarily run optimally on an A100, and vice versa. Therefore segmenting into node pools, each with a different hardware type gives the AI researcher and ML engineer better control of where to run.
* **Quota control** - splitting to node pools allows the admin to set specific quota per hardware type, e.g. give high priority project guaranteed access to advanced GPU hardware, while keeping lower priority project with a lower quota or even with no quota at all for that high-end GPU, but give it a “best-effort” access only (i.e. if the high priority guaranteed project is not using those resources).
* **Multi-region or multi-availability-zone cloud environments** - if some or all of your clusters run on the cloud (or even on-premise) but any of your clusters uses different physical locations or different topologies (e.g. racks), you probably want to segment your resources per region/zone/topology to be able to control where to run your workloads, how much quota to assign to specific environments (per project, per department), even if all those locations are all using the same hardware type. This methodology can help in optimizing the performance of your workloads because of the superior performance of local computing such as the locality of distributed workloads, local storage etc.
* **Explainability and predictability** - large environments are complex to understand, this becomes even more complex when an environment is loaded. To maintain users’ satisfaction and their understanding of the resources state, as well as to keep predictability of your workload chances to get scheduled, segmenting your cluster into smaller pools may significantly help.
* **Scale** - NVIDIA Run:ai implementation of node pools has many benefits, one of the main of them is scale. Each node pool has its own [Scheduler](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles) instance, therefore allowing the cluster to handle more nodes and schedule workloads faster when segmented into node pools vs. one large cluster. To allow your workloads to use any resource within a cluster that is split to node pools, a second-level Scheduler is in charge of scheduling workloads between node pools according to your preferences and resource availability.
* **Prevent mutual exclusion** - Some AI workloads consume CPU-only resources, to prevent those workloads from consuming the CPU resources of GPU nodes and thus block GPU workloads from using those nodes, it is recommended to group CPU-only nodes into a dedicated node pool(s) and assign a quota for CPU projects to CPU node-pools only while keeping GPU node-pools with zero quota and optionally “best-effort” over quota access for CPU-only projects.

#### Grouping Examples

Set out below are illustrations of different grouping options.

**Example: grouping nodes by topology**

![](/files/UxAiRgcsB3Dh1qpB8o1J)

**Example: grouping nodes by hardware type**

![](/files/JUf1Z0deaILz9qJ1v4Us)

### Assigning Your Resources

After the initial grouping of resources, it is time to associate resources to AI initiatives, this is performed by assigning quotas to projects and optionally to departments. Assigning GPU quota to a project, on a node pool basis, means that the workloads submitted by that project are entitled to use those GPUs as guaranteed resources and can use them for all [workload types](/saas/workloads-in-nvidia-run-ai/workload-types).

However, what happens if the project requires more resources than its quota? This depends on the type of workloads that the user wants to submit. If the user requires more resources for non-preemptible workloads, then the quota must be increased, because non-preemptible workloads require guaranteed resources. On the other hand, if the type of workload is, for example, a model Training workload that is preemptible - in this case the project can exploit unused resources of other projects, as long as the other projects don’t need them. over quota is set per project on a node-pool basis and per department.

Administrators can use quota allocations to prioritize resources between users, teams, and AI initiatives. The administrator can completely prevent the use of certain node pools by a project or department by setting the node pool quota to 0 and disabling over quota for that node pool, or it can keep the quota to 0 and enable over quota to that node pool and allow access based on resource availability only (e.g. unused GPUs). However, when a project with a non-zero quota needs to use those resources, the Scheduler reclaims those resources back and preempts the preemptible workloads of over quota projects. As an administrator, you can also have an impact on the amount of over quota resources a project or department uses.

It is essential to make sure that the sum of all projects' quota does NOT surpass that of the Department, and that the sum of all departments does not surpass the number of physical resources, per node pool and for the entire cluster (we call such behavior - ‘over-subscription’). The reason over-subscription is not recommended is that it may produce unexpected scheduling decisions, especially those that might preempt ‘non-preemptible’ workloads or fail to schedule workloads within quota, either non-preemptible or preemptible, thus quota cannot be considered anymore as ‘guaranteed’. Admins can opt-in a system flag that helps to prevent over-subscription scenarios.

**Example: assigning resources to projects**

![](/files/HG60V4huWzbL8BX34Xs4)

## Assigning Users to Projects and Departments

NVIDIA Run:ai system is using [‘Role Based Access Control’ (RBAC)](/saas/infrastructure-setup/authentication/overview#role-based-access-control-rbac-in-run-ai) to manage users’ access rights to the different objects of the system, its resources, and the set of allowed actions.\
To allow AI researchers, ML engineers, Project Admins, or any other stakeholder of your AI initiatives to access projects and use AI compute resources with their AI initiatives, the administrator needs to assign users to projects. After a user is assigned to a project with the proper role, e.g. ‘L1 Researcher’, the user can submit and monitor its workloads under that project. Assigning users to departments is usually done to assign ‘Department Admin’ to manage a specific department. Other roles, such as ‘L1 Researcher’, can also be assigned to departments, this allows the researcher access to all projects within that department.

## Scopes in an Organization

This is an example of an organization, as represented in the NVIDIA Run:ai platform:

![](/files/JQ80CI8eVsw9RnP40tcE)

The organizational tree is structured from top down under a single node headed by the account. The account is comprised of clusters, departments and projects.

After mapping and building your hierarchal structured organization as shown above, you can assign or associate various NVIDIA Run:ai components (e.g. workloads, roles, assets, policies, and more) to **different parts** of the organization - these organizational parts are the **Scopes**. The following organizational example consists of 5 optional scopes:

![](/files/5i9S3Ya16KHxptCVB8No)

{% hint style="info" %}
**Note**

When a scope is selected, the very same unit, including all of its subordinates (both existing and any future subordinates, if added), are selected as well.
{% endhint %}

## Next Steps

Now that resources are grouped into node pools, organizational units or business initiatives are mapped into projects and departments, projects’ quota parameters are set per node pool, and users are assigned to projects, you can finally [submit workloads](/saas/workloads-in-nvidia-run-ai/workloads) from a project and use compute resources to run your AI initiatives.


# Managing Your Organization


# Projects

Researchers submit AI workloads to NVIDIA Run:ai. To control how resources are allocated and prioritized across teams and initiatives, NVIDIA Run:ai uses [Projects](/saas/platform-management/aiinitiatives/adapting-ai-initiatives#mapping-your-organization) as the primary organization management unit. Projects allow administrators to define resource quotas, scheduling behavior, access controls, and policy enforcement, while logically separating workloads between different organizational efforts.

A project can represent a team, an individual user, or a specific initiative, and can be assigned dedicated GPU and CPU quotas, scheduling rules, and permissions. Projects can also be grouped into [departments](/saas/platform-management/aiinitiatives/organization/departments) to reflect the organizational hierarchy and enable centralized resource governance.

For example, a team working on a face-recognition initiative might collaborate under a shared project named `face-recognition-2024`. Alternatively, each team member can be assigned an individual project with a personal resource quota to ensure predictable and isolated resource usage.

## Projects Table

The Projects table can be found under **Organization** in the NVIDIA Run:ai platform.

The Projects table provides a list of all projects defined for a specific cluster, and allows you to manage them. You can switch between clusters by selecting your cluster using the filter at the top.

<figure><img src="/files/kpn2oMNbrcZVLYOUJmNw" alt=""><figcaption></figcaption></figure>

The Projects table consists of the following columns:

| Column                                          | Description                                                                                                                                                                                                                                                                                                                                                                                           |
| ----------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Project                                         | The name of the project                                                                                                                                                                                                                                                                                                                                                                               |
| Description                                     | A description of the project                                                                                                                                                                                                                                                                                                                                                                          |
| Department                                      | The name of the parent department. Several projects may be grouped under a department.                                                                                                                                                                                                                                                                                                                |
| Cluster                                         | The cluster that the project is associated with                                                                                                                                                                                                                                                                                                                                                       |
| Status                                          | The Project creation status. Projects are manifested as Kubernetes namespaces. The project status represents the Namespace creation status.                                                                                                                                                                                                                                                           |
| Node pool(s)                                    | The node pools associated with the project. By default, a new project is associated with all node pools within its associated cluster. Administrators can change the node pools’ quota parameters for a project. Click the values under this column to view the list of node pools with their parameters (as described below).                                                                        |
| Subject(s)                                      | The users, SSO groups, or service accounts with access to the project. Click the values under this column to view the list of subjects with their parameters (as described below). This column is only viewable if your role in the NVIDIA Run:ai platform allows you those permissions.                                                                                                              |
| GPU quota                                       | The GPU quota allocated to the project. This number represents the sum of all node pools’ GPU quota allocated to this project.                                                                                                                                                                                                                                                                        |
| Allocated GPUs                                  | The total number of GPUs allocated by successfully scheduled workloads under this project                                                                                                                                                                                                                                                                                                             |
| Avg. GPU allocation                             | The average number of GPU devices allocated by workloads submitted within this project, based on the selected time range                                                                                                                                                                                                                                                                              |
| Avg. GPU utilization                            | The average percentage of GPU utilization across all workloads submitted within this project, based on the selected time range                                                                                                                                                                                                                                                                        |
| Avg. GPU memory utilization                     | The average percentage of GPU memory usage across all workloads submitted within this project, based on the selected time range                                                                                                                                                                                                                                                                       |
| Allocated CPUs (Core)                           | The total number of CPU cores allocated by workloads submitted within this project. (This column is only available if the CPU Quota setting is enabled, as described below).                                                                                                                                                                                                                          |
| Allocated CPU Memory                            | The total number of CPUs allocated by successfully scheduled workloads under this project. (This column is only available if the CPU Quota setting is enabled, as described below).                                                                                                                                                                                                                   |
| GPU allocation ratio                            | The ratio of Allocated GPUs to GPU quota. This number reflects how well the project’s GPU quota is utilized by its descendent workloads. A number higher than 100% indicates the project is using over quota GPUs.                                                                                                                                                                                    |
| CPU allocation ratio                            | The ratio of Allocated CPUs (cores) to CPU quota (cores). This number reflects how much the project’s ‘CPU quota’ is utilized by its descendent workloads. A number higher than 100% indicates the project is using over quota CPU cores.                                                                                                                                                             |
| CPU memory allocation ratio                     | The ratio of Allocated CPU memory to CPU memory quota. This number reflects how well the project’s ‘CPU memory quota’ is utilized by its descendent workloads. A number higher than 100% indicates the project is using over quota CPU memory.                                                                                                                                                        |
| CPU quota (Cores)                               | CPU quota allocated to this project. (This column is only available if the CPU Quota setting is enabled, as described below). This number represents the sum of all node pools’ CPU quota allocated to this project. The ‘unlimited’ value means the CPU (cores) quota is not bounded and workloads using this project can use as many CPU (cores) resources as they need (if available).             |
| CPU memory quota                                | CPU memory quota allocated to this project. (This column is only available if the CPU Quota setting is enabled, as described below). This number represents the sum of all node pools’ CPU memory quota allocated to this project. The ‘unlimited’ value means the CPU memory quota is not bounded and workloads using this Project can use as much CPU memory resources as they need (if available). |
| Node type (affinity) - Training                 | The list of NVIDIA Run:ai node-affinities. Any training workload submitted within this project must specify one of those NVIDIA Run:ai node affinities, otherwise it is not submitted.                                                                                                                                                                                                                |
| Node type (affinity) - Workspaces               | The list of NVIDIA Run:ai node-affinities. Any interactive (workspace) workload submitted within this project must specify one of those NVIDIA Run:ai node affinities, otherwise it is not submitted.                                                                                                                                                                                                 |
| Idle GPU time limit - Training                  | The time in days:hours:minutes after which the project stops a training workload not using its allocated GPU resources.                                                                                                                                                                                                                                                                               |
| Idle GPU time limit - preemptible workspaces    | The time in days:hours:minutes after which the project stops a preemptible interactive (workspace) workload not using its allocated GPU resources.                                                                                                                                                                                                                                                    |
| Idle GPU time limit - non preemptible workloads | The time in days:hours:minutes after which the project stops a non-preemptible interactive (workspace) workload not using its allocated GPU resources..                                                                                                                                                                                                                                               |
| Training time limit                             | The duration in days:hours:minutes after which the project stops a training workload                                                                                                                                                                                                                                                                                                                  |
| Workspace time limit                            | The duration in days:hours:minutes after which the project stops an interactive (workspace) workload                                                                                                                                                                                                                                                                                                  |
| Creation time                                   | The timestamp for when the project was created                                                                                                                                                                                                                                                                                                                                                        |
| Workload(s)                                     | The list of workloads associated with the project. Click the values under this column to view the list of workloads with their resource parameters (as described below).                                                                                                                                                                                                                              |

### Node Pools with Quota Associated with the Project

Click one of the values of Node pool(s) with quota column, to view the list of node pools and their parameters.

| Column                     | Description                                                                                                                                                                                                                                                                                                                                                                                    |
| -------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Node pool                  | The name of the node pool is given by the administrator during node pool creation. All clusters have a default node pool created automatically by the system and named ‘default’.                                                                                                                                                                                                              |
| Order of priority          | The default order in which the Scheduler uses node pools to schedule a workload. This is used only if the order of priority of node pools is not set in the workload during submission, either by an admin policy or the user. An empty value means the node pool is not part of the project’s default list, but can still be chosen by an admin policy or the user during workload submission |
| GPU quota                  | The amount of GPU quota the administrator dedicated to the project for this node pool (floating number, e.g. 2.3 means 230% of GPU capacity).                                                                                                                                                                                                                                                  |
| Project Rank               | The project's scheduling rank compared to other projects in the same department and same node pool                                                                                                                                                                                                                                                                                             |
| Over-quota weight          | Represents the relative weight used to calculate the amount of non-guaranteed overage resources a project can get on top of its quota in this node pool. Unused resources are split between projects that require the use of overage resources.                                                                                                                                                |
| Max GPU devices allocation | Represents the maximum GPU device allocation the project can get from this node pool - the maximum sum of quota and over-quota GPUs.                                                                                                                                                                                                                                                           |
| CPU (Cores)                | The amount of CPUs (cores) quota the administrator has dedicated to the project for this node pool (floating number, e.g. 1.3 Cores = 1300 mili-cores). The ‘unlimited’ value means the CPU (Cores) quota is not bounded and workloads using this node pool can use as many CPU (Cores) resources as they require, (if available).                                                             |
| CPU memory                 | The amount of CPU memory quota the administrator has dedicated to the project for this node pool (floating number, in MB or GB). The ‘unlimited’ value means the CPU memory quota is not bounded and workloads using this node pool can use as much CPU memory resource as they need (if available).                                                                                           |
| Allocated GPUs             | The actual amount of GPUs allocated by workloads using this node pool under this project. The number of allocated GPUs may temporarily surpass the GPU quota if over quota is used.                                                                                                                                                                                                            |
| Allocated CPU (Cores)      | The actual amount of CPUs (cores) allocated by workloads using this node pool under this project. The number of allocated CPUs (cores) may temporarily surpass the CPUs (Cores) quota if over quota is used.                                                                                                                                                                                   |
| Allocated CPU memory       | The actual amount of CPU memory allocated by workloads using this node pool under this Project. The number of Allocated CPU memory may temporarily surpass the CPU memory quota if over quota is used.                                                                                                                                                                                         |

### Subjects Authorized for the Project

Click one of the values in the Subject(s) column, to view the list of subjects and their parameters. This column is only viewable, if your role in the NVIDIA Run:ai system affords you those permissions.

| Column        | Description                                                                                                                                                                                                              |
| ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Subject       | A user, SSO group, or service account assigned with a role in the scope of this Project                                                                                                                                  |
| Type          | The type of subject assigned to the access rule (user, SSO group, or service account)                                                                                                                                    |
| Scope         | The scope of this project in the organizational tree. Click the name of the scope to view the organizational tree diagram, you can only view the parts of the organizational tree for which you have permission to view. |
| Role          | The role assigned to the subject, in this project’s scope                                                                                                                                                                |
| Authorized by | The user who granted the access rule                                                                                                                                                                                     |
| Last updated  | The last time the access rule was updated                                                                                                                                                                                |

### Workloads Associated with the Project

Click one of the values of Workload(s) column, to view the list of workloads and their parameters

| Column                  | Description                                                                                                                                                                                      |
| ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Workload                | The name of the workload, given during its submission. Optionally, an icon describing the type of workload is also visible                                                                       |
| Type                    | The type of the workload, e.g. Workspace, Training, Inference                                                                                                                                    |
| Status                  | The state of the workload and time elapsed since the last status change                                                                                                                          |
| Created by              | The subject that created this workload                                                                                                                                                           |
| Running/ requested pods | The number of running pods out of the number of requested pods for this workload. e.g. a distributed workload requesting 4 pods but may be in a state where only 2 are running and 2 are pending |
| Creation time           | The date and time the workload was created                                                                                                                                                       |
| GPU compute request     | The amount of GPU compute requested (floating number, represents either a portion of the GPU compute, or the number of whole GPUs requested)                                                     |
| GPU memory request      | The amount of GPU memory requested (floating number, can either be presented as a portion of the GPU memory, an absolute memory size in MB or GB, or a MIG profile)                              |
| CPU memory request      | The amount of CPU memory requested (floating number, presented as an absolute memory size in MB or GB)                                                                                           |
| CPU compute request     | The amount of CPU compute requested (floating number, represents the number of requested Cores)                                                                                                  |

### Customizing the Table View

* Filter - Click ADD FILTER, select the column to filter by, and enter the filter values
* Search - Click SEARCH and type the value to search by
* Sort - Click each column header to sort by
* Column selection - Click COLUMNS and select the columns to display in the table
* Download table - Click MORE and then Click Download as CSV. Export to CSV is limited to 20,000 rows.

## Adding a New Project

1. Click **+NEW PROJECT**
2. Select a **scope.** Choose the department or organizational scope under which the project will be created. You can view and select only the clusters for which your assigned roles grant you permission.
3. Enter a **name** for the project. Project names must start with a letter and can only contain lower case Latin letters, numbers or a hyphen ('-’).
4. **Namespace** associated with the project. Each project is mapped to a Kubernetes namespace in the cluster. All workloads created under this project run in that namespace.
   * By default, NVIDIA Run:ai creates a new namespace in the Kubernetes cluster using the project's name with the `run:ai` prefix (`runai-<project-name>`).
   * Alternatively, you can use an existing namespace from the cluster. You must associate this namespace with the new project using the provided `kubectl` command. Copy the command shown in the UI and run it in your terminal to complete the association.
5. Choose how to proceed:

   * Click **CREATE PROJECT & CLOSE** to create the project without assigning resources.
   * Click **CREATE PROJECT & CONTINUE** to set dedicated resources and scheduling guarantees.

   <div data-gb-custom-block data-tag="hint" data-style="info" class="hint hint-info"><p><strong>Note</strong></p><ul><li>The project is created and appears in the <strong>Projects</strong> grid.</li><li><p>When you select <strong>CREATE PROJECT &#x26; CLOSE</strong>, the project is created without assigned resources. The project’s quota is set to 0, which means workloads can run only when spare GPUs are available. The following default settings are applied:</p><ul><li>The project is created with a <strong>Medium Low</strong> rank.</li><li>Over quota weight is always applied when distributing over-quota GPUs among projects in the same rank. By default, the project’s weight is derived proportionally from its assigned GPU quota. Since the quota is set to 0, the effective weight is 0. If <strong>Over quota weight</strong> is enabled via <strong>General settings</strong>, the project is assigned a default weight of 2.</li></ul></li></ul></div>

### Assigning GPU Quota

Assign GPUs to the project and define how the [Scheduler](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles) prioritizes node pools when running the project’s workloads.

1. Review the **node pools** available to the project. By default, a project inherits the node pools defined at the department level. This represents the default (suggested) configuration set by the department and can be changed per project by selecting which node pools are available and adjusting their order.
2. Add **node pools** and set the default order in which workloads will be scheduled to use them. The leftmost node pool is the first one workloads will attempt to schedule onto. The Scheduler first tries to allocate resources using the highest priority node pool, then the next in priority, until it reaches the lowest priority node pool list, then the Scheduler starts from the highest again.

   <div data-gb-custom-block data-tag="hint" data-style="info" class="hint hint-info"><p><strong>Note</strong></p><ul><li>When no node pools are configured, you can assign quota for the whole project, instead of per node pool. After node pools are created, you can assign quota for each node pool separately.</li><li>The order of node pools can be overridden during workload submission by an admin policy or by the user.</li></ul></div>
3. Review the department's GPU quota per node pool:
   * **Unassigned** - The number of GPUs not assigned as quota to subordinate projects.
   * **Total** - The total number of GPUs assigned to the department in this node pool.
4. Set the project's **GPU quota** in the selected node pool. A **Quota distribution** graph next to the project shows how GPUs are distributed across the department for that node pool. The graph reflects:
   * **Assigned quota** - The number of GPUs from the node pool that are assigned to this project.
   * **Unassigned quota** - The number of GPUs still available in the node pool that are not yet assigned for any project.
   * **Sum of other projects’ assigned quota** - The total number of GPUs already assigned to all other projects in the same node pool.

     <div data-gb-custom-block data-tag="hint" data-style="info" class="hint hint-info"><p><strong>Note</strong></p><ul><li>If the department’s quota is set to <strong>Unlimited</strong>, projects under that department can consume available GPUs up to the physical capacity of the node pool or cluster.</li><li><p>If the department’s quota is limited, project behavior depends on whether the <strong>Limit projects from exceeding department quota</strong> is enabled/disabled via <strong>General settings</strong>:</p><ul><li>If disabled, projects can be assigned a GPU quota that exceeds the department quota. In this case, the system may display warnings, but the department quota is not enforced.</li><li>If enabled, the project’s assigned GPU quota must remain within the department quota. You cannot assign more GPUs than the department has available.</li></ul></li><li>Setting a quota to 0 (GPU, CPU, or CPU memory) and disabling <strong>Over quota weight</strong> via <strong>General settings</strong>, means the project is blocked from using those resources on this node pool.</li></ul></div>
5. You can adjust **other projects' quotas** to free up additional resources. A table lists all projects and their assigned GPU quotas. Updating these values redistributes the department’s assigned and unassigned GPU quota across projects for the selected node pool, as reflected in the **Quota distribution** graph next to each project:
   * Use the **Search** icon to find a project
   * Click the **GPU** quota value and set the number of GPUs
   * Click **APPLY**
6. These steps apply to the selected node pool. If multiple node pools are configured, repeat the process for each.
7. Click **SAVE & CONTINUE**

### Defining GPU Over Quota

Define how workloads can consume GPUs beyond the project’s assigned quota when additional resources are available.

1. **Allow workloads to exceed the project’s quota** is enabled by default, allowing workloads in the project to consume GPUs beyond the assigned quota. If disabled, workloads from the project will run only within the assigned quota and will not consume additional GPUs, even if resources are available.

   <div data-gb-custom-block data-tag="hint" data-style="info" class="hint hint-info"><p><strong>Note</strong></p><p>Workloads running over quota are <strong>opportunistic</strong> and may be preempted or stopped at any time when higher priority workloads require resources.</p></div>
2. Set the **max GPU allocation**. This value defines the maximum GPU device allocation the project can get from the node pool, representing the maximum sum of assigned quota and over-quota GPUs:
   * By default, this field is set to **Unlimited**, allowing the project to consume any available over-quota GPUs
   * Enter a **number** equal to or above the assigned quota
3. Rank projects in the order in which their workloads will be scheduled. A project’s **rank** determines the scheduling priority of its workloads only when competing for over-quota GPUs within the same node pool. By default, projects are assigned a **Medium Low** rank. Projects with a higher rank are scheduled before projects with a lower rank when over-quota GPUs become available. When multiple projects share the same rank, available over-quota GPUs are distributed according to each project’s **weight**:

   * Filter the list of projects:
     * Click **Add filter**
     * Select **Rank** to filter by rank level, or **Project** to filter by project name
     * Enter or select the filter criteria and click **APPLY**
   * Locate and select projects:
     * Locate the project or projects whose rank you want to change
     * Select one or more projects
   * Change the rank:
     * Click **MOVE TO RANK**
     * In the **Move to Rank** dialog, select the destination rank
     * Click **MOVE \<n> PROJECTS**

   To learn more about project ranks, see [The NVIDIA Run:ai Scheduler: concepts and principles](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles).
4. Assign **weights** to determine the projects' share of resources. A project’s weight determines how available over-quota GPUs are distributed between projects that have the same rank within the same node pool. Weight behavior depends on whether **Over quota weight** is enabled or disabled via **General settings**. When disabled, the project’s weight is proportional to its assigned GPU quota, displayed for reference, and cannot be edited. When enabled, all projects within the same rank are assigned a **default weight of 2**, and the value can be adjusted per project:
   * Filter the list of projects:
     * Click **Add filter**
     * Select **Rank** to filter by rank level, or **Project** to filter by project name
     * Enter or select the filter criteria and click **APPLY**
   * Review weight distribution:
     * Projects are grouped by rank
     * For each project, the **Weight distribution** graph shows:
       * The project’s **weight**
       * The **total weight of other projects** in the same rank
   * Edit the project’s weight (when Over quota weight is enabled):
     * Click the **Weight** value for the project and enter a value
     * Click **APPLY**
5. These steps apply to the selected node pool. If multiple node pools are configured, repeat the process for each.
6. Click **SAVE &** **CONTINUE**

### Assigning CPU Quota

Assign CPU resources (cores and memory) to the project and define the maximum CPU usage per node pool. This form is displayed only if **CPU quota** is enabled via the **General settings**.

1. Set the number of **CPU (Cores)** the project can use in the selected node pool. By default, CPU cores are set to **Unlimited**, allowing the project to consume any available CPU resources in the node pool. Enter a **number**.
2. Set the amount of **CPU memory** available to the project in the selected node pool. By default, CPU memory is set to **Unlimited**, allowing the project to consume any available CPU resources in the node pool. Enter a **number** and select a unit (**MB**, **MiB**, or **GB**).

   <div data-gb-custom-block data-tag="hint" data-style="info" class="hint hint-info"><p><strong>Note</strong></p><ul><li><p>If the department’s quota is limited, project behavior depends on whether the <strong>Limit projects from exceeding department quota</strong> is enabled/disabled via <strong>General settings</strong>:</p><ul><li>If disabled, projects can be assigned a CPU quota that exceeds the department quota. In this case, the system may display warnings, but the department quota is not enforced.</li><li>If enabled, the project’s assigned CPU quota must remain within the department quota. You cannot assign more CPU resources than the department has available.</li></ul></li><li>When the <strong>CPU quota</strong> feature flag is disabled, any previously set CPU quotas for a project are automatically removed.</li></ul></div>
3. These steps apply to the selected node pool. If multiple node pools are configured, repeat the process for each.
4. Click **SAVE & CONTINUE**

### Setting Scheduling Rules

Scheduling rules control how compute resources are used by workloads. The restrict either the resources (nodes) on which workloads can run or the duration of the run time. Scheduling rules are set for projects and apply to specific workload types. By default, projects inherit scheduling rules from their parent department. Project-level scheduling rules can be used to add or override restrictions for workloads in that project. See [Scheduling rules](/saas/platform-management/policies/scheduling-rules) for more details.

1. Click **+RULE**
2. Select a rule type from the dropdown:
   * **+ Idle GPU timeout** - Limit the duration of a workload whose GPU is idle
   * **+ Workload time limit** - Set a time limit for workspaces regardless of their activity (e.g., stop the workspace after 1 day of work)
   * **+ Node type (Affinity**) - Limit workloads to run on specific node types
3. For each rule, select the **workload type(s)** to which the rule applies and configure the rule parameters
4. Click **SAVE & CONTINUE**

Once scheduling rules are set for a project, all matching workloads associated with the project have the restrictions applied to them, as defined, when the workload is submitted. New scheduling rules added to a project are not applied over previously created workloads associated with that project.

{% hint style="info" %}
**Note**

* When editing a scheduling rule within a project, you can only **tighten** rules inherited from the department (for example, by setting a shorter time limit).
* Scheduling rules defined at the department level **cannot be deleted** from within a project.
  {% endhint %}

### Applying Access Rules

1. Click **+ACCESS RULE**
2. Select a subject - **User, SSO group**, or **Service account**
3. Select or enter the subject identifier. You can define up to 10 subjects of the selected type:
   * **User email** for a local user created in NVIDIA Run:ai or for SSO user as recognized by the IDP
   * **Group name** as recognized by the IDP
   * **Service account name** as created in NVIDIA Run:ai
4. Select a **role**
5. Click **SAVE RULE**
6. Click **CLOSE**

## Editing a Project

To edit a project:

1. Select the project you want to edit
2. Click **EDIT**
3. From the dropdown menu, select the form you want to modify
4. Update the settings and click **SAVE & CONTINUE**

{% hint style="info" %}
**Note**

When editing a project:

* You can navigate directly to a specific form without completing the full project creation flow again.
* Changes apply only to the selected form; previously configured settings in other forms remain unchanged.
* You cannot reduce GPU and CPU quotas below the quotas currently consumed by **non-preemptible** workloads.
  {% endhint %}

## Viewing a Project’s Policy

To view the policy of a project:

1. Select the project for which you want to view its [policies](/saas/platform-management/policies/native-workload-policies). This option is only active for projects with defined policies in place.
2. Click **VIEW POLICY** and select the workload type for which you want to view the policies
3. In the Policy form, view the workload rules that are enforcing your project for the selected workload type as well as the defaults:
   * **Parameter** - The workload submission parameter that Rules and Defaults are applied to
   * **Type (applicable for data sources only)** - The data source type (Git, S3, nfs, pvc etc.)
   * **Default** - The default value of the Parameter
   * **Rule** - Set up constraints on workload policy fields
   * **Source** - The origin of the applied policy (cluster, department or project)

{% hint style="info" %}
**Note**

The policy affecting the project consists of rules and defaults. Some of these rules and defaults may be derived from policies of a parent cluster and/or department (source). You can see the source of each rule in the policy form.
{% endhint %}

## Deleting a Project

To delete a project:

1. Select the project you want to delete
2. Click **DELETE**
3. On the dialog, click **DELETE** to confirm

{% hint style="info" %}
**Note**

**Clusters < v2.20**

Deleting a project does not delete its associated namespace, any of the workloads running using this namespace, or the policies defined for this project. However, any assets created in the scope of this project such as compute resources, environments, data sources, templates and credentials, are permanently deleted from the system.

**Clusters >=v2.20**

Deleting a project does not delete its associated namespace, but will attempt to delete it’s associated workloads and assets. Any assets created in the scope of this project such as compute resources, environments, data sources, templates and credentials, are permanently deleted from the system.
{% endhint %}

## Using CLI

To view the available actions on projects, see the project [CLI v2 reference](/saas/reference/cli/runai/runai_project).

## Using API

To view the available actions, go to the [Projects](https://run-ai-docs.nvidia.com/api/organizations/projects) API reference.


# Departments

Departments group multiple projects under a shared organizational scope. By organizing projects into a department, you can manage resources and governance at scale, for example, by allocating quotas across a set of projects, applying policies at the department level, and creating assets that can be shared across all projects (or selected descendant projects) within the department.

For example, in an academic environment, a department can be the Physics Department grouping various projects (AI Initiatives) within the department, or grouping projects where each project represents a single student.

## Departments Table

The Departments table can be found under **Organization** in the NVIDIA Run:ai platform.

The Departments table lists all departments defined for a specific cluster and allows you to manage them. You can switch between clusters by selecting your cluster using the filter at the top.

<figure><img src="/files/f1RCCj2Fboh3xEvXuNka" alt=""><figcaption></figcaption></figure>

The Departments table consists of the following columns:

| Column                      | Description                                                                                                                                                                                                                                                                                                                                                                                           |
| --------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Department                  | The name of the department                                                                                                                                                                                                                                                                                                                                                                            |
| Description                 | A description of the department                                                                                                                                                                                                                                                                                                                                                                       |
| Node pool(s)                | The node pools associated with this department. By default, all node pools within a cluster are associated with each department. Administrators can change the node pools’ quota parameters for a department. Click the values under this column to view the list of node pools with their parameters (as described below)                                                                            |
| Cluster                     | The cluster that the department is associated with                                                                                                                                                                                                                                                                                                                                                    |
| Project(s)                  | List of projects associated with this department                                                                                                                                                                                                                                                                                                                                                      |
| Subject(s)                  | The users, SSO groups, or service accounts with access to the project. Click the values under this column to view the list of subjects with their parameters (as described below). This column is only viewable if your role in NVIDIA Run:ai platform allows you those permissions.                                                                                                                  |
| GPU quota                   | GPU quota associated with the department                                                                                                                                                                                                                                                                                                                                                              |
| Allocated GPUs              | The total number of GPUs allocated by successfully scheduled workloads in projects associated with this department                                                                                                                                                                                                                                                                                    |
| Avg. GPU allocation         | The average number of GPU devices allocated by workloads submitted within this department, based on the selected time range.                                                                                                                                                                                                                                                                          |
| Avg. GPU utilization        | The average percentage of GPU utilization across all workloads submitted within this department, based on the selected time range.                                                                                                                                                                                                                                                                    |
| Avg. GPU memory utilization | The average percentage of GPU memory usage across all workloads submitted within this department, based on the selected time range.                                                                                                                                                                                                                                                                   |
| Allocated CPUs (Core)       | The total number of CPU cores allocated by workloads submitted within this project. (This column is only available if the CPU Quota setting is enabled, as described below).                                                                                                                                                                                                                          |
| Allocated CPU Memory        | The total number of CPUs allocated by successfully scheduled workloads under this project. (This column is only available if the CPU Quota setting is enabled, as described below).                                                                                                                                                                                                                   |
| GPU allocation ratio        | The ratio of Allocated GPUs to GPU quota. This number reflects how well the department’s GPU quota is utilized by its descendant projects. A number higher than 100% means the department is using over quota GPUs. A number lower than 100% means not all projects are utilizing their quotas. A quota becomes allocated once a workload is successfully scheduled.                                  |
| CPU allocation ratio        | The ratio of Allocated CPUs (cores) to CPU quota (cores). This number reflects how much the project’s ‘CPU quota’ is utilized by its descendent workloads. A number higher than 100% indicates the project is using over quota CPU cores.                                                                                                                                                             |
| CPU memory allocation ratio | The ratio of Allocated CPU memory to CPU memory quota. This number reflects how well the project’s ‘CPU memory quota’ is utilized by its descendent workloads. A number higher than 100% indicates the project is using over quota CPU memory.                                                                                                                                                        |
| CPU quota (Cores)           | CPU quota allocated to this project. (This column is only available if the CPU Quota setting is enabled, as described below). This number represents the sum of all node pools’ CPU quota allocated to this project. The ‘unlimited’ value means the CPU (cores) quota is not bounded and workloads using this project can use as many CPU (cores) resources as they need (if available).             |
| CPU memory quota            | CPU memory quota allocated to this project. (This column is only available if the CPU Quota setting is enabled, as described below). This number represents the sum of all node pools’ CPU memory quota allocated to this project. The ‘unlimited’ value means the CPU memory quota is not bounded and workloads using this Project can use as much CPU memory resources as they need (if available). |
| Creation time               | The timestamp for when the department was created                                                                                                                                                                                                                                                                                                                                                     |
| Workload(s)                 | The list of workloads under projects associated with this department. Click the values under this column to view the list of workloads with their resource parameters (as described below)                                                                                                                                                                                                            |

### Projects Associated with the Department

Click one of the values of Projects(s) column, to view the list of workloads and their parameters.

| Column                | Description                                                                                                                                                                                                                                                                                                                                                                                           |
| --------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Project               | The name of the project                                                                                                                                                                                                                                                                                                                                                                               |
| GPU quota             | The GPU quota allocated to the project. This number represents the sum of all node pools’ GPU quota allocated to this project.                                                                                                                                                                                                                                                                        |
| CPU quota (Cores)     | CPU quota allocated to this project. (This column is only available if the CPU Quota setting is enabled, as described below). This number represents the sum of all node pools’ CPU quota allocated to this project. The ‘unlimited’ value means the CPU (cores) quota is not bounded and workloads using this project can use as many CPU (cores) resources as they need (if available).             |
| CPU memory quota      | CPU memory quota allocated to this project. (This column is only available if the CPU Quota setting is enabled, as described below). This number represents the sum of all node pools’ CPU memory quota allocated to this project. The ‘unlimited’ value means the CPU memory quota is not bounded and workloads using this Project can use as much CPU memory resources as they need (if available). |
| Allocated GPUs        | The total number of GPUs allocated by successfully scheduled workloads under this project                                                                                                                                                                                                                                                                                                             |
| Allocated CPUs (Core) | The total number of CPU cores allocated by workloads submitted within this project. (This column is only available if the CPU Quota setting is enabled, as described below).                                                                                                                                                                                                                          |
| Allocated CPU Memory  | The total number of CPUs allocated by successfully scheduled workloads under this project. (This column is only available if the CPU Quota setting is enabled, as described below).                                                                                                                                                                                                                   |
| GPU allocation ratio  | The ratio of Allocated GPUs to GPU quota. This number reflects how well the project’s GPU quota is utilized by its descendent workloads. A number higher than 100% indicates the project is using over quota GPUs.                                                                                                                                                                                    |
| CPU allocation ratio  | The ratio of Allocated CPUs (cores) to CPU quota (cores). This number reflects how much the project’s ‘CPU quota’ is utilized by its descendent workloads. A number higher than 100% indicates the project is using over quota CPU cores.                                                                                                                                                             |

### Node Pools with Quota Associated with the Department

Click one of the values of Node pool(s) with quota column, to view the list of node pools and their parameters.

| Column                     | Description                                                                                                                                                                                                                                                                                                                   |
| -------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Node pool                  | The name of the node pool is given by the administrator during node pool creation. All clusters have a default node pool created automatically by the system and named ‘default’.                                                                                                                                             |
| Priority                   | The node pool's order of priority for newly created projects under this department                                                                                                                                                                                                                                            |
| GPU quota                  | The amount of GPU quota the administrator dedicated to the department for this node pool (floating number, e.g. 2.3 means 230% of a GPU capacity)                                                                                                                                                                             |
| Department Rank            | The department’s scheduling rank relative to other departments sharing the same node pool                                                                                                                                                                                                                                     |
| Over-quota weight          | Represents a weight used to calculate the amount of non-guaranteed overage resources a department can get on top of its quota in this node pool. All unused resources are split between departments that require the use of overage resources.                                                                                |
| Max GPU devices allocation | The maximum GPU device allocation the department can get from this node pool - the maximum sum of quota and over-quota GPUs                                                                                                                                                                                                   |
| CPU (Cores)                | The amount of CPU (cores) quota the administrator has dedicated to the department for this node pool (floating number, e.g. 1.3 Cores = 1300 mili-cores). The ‘unlimited’ value means the CPU (Cores) quota is not bound and workloads using this node pool can use as many CPU (Cores) resources as they need (if available) |
| CPU memory                 | The amount of CPU memory quota the administrator has dedicated to the department for this node pool (floating number, in MB or GB). The ‘unlimited’ value means the CPU memory quota is not bounded and workloads using this node pool can use as much CPU memory resource as they need (if available).                       |
| Allocated GPUs             | The total amount of GPUs allocated by workloads using this node pool under projects associated with this department. The number of allocated GPUs may temporarily surpass the GPU quota of the department if over quota is used.                                                                                              |
| Allocated CPU (Cores)      | The total amount of CPUs (cores) allocated by workloads using this node pool under all projects associated with this department. The number of allocated CPUs (cores) may temporarily surpass the CPUs (Cores) quota of the department if over quota is used.                                                                 |
| Allocated CPU memory       | The actual amount of CPU memory allocated by workloads using this node pool under all projects associated with this department. The number of Allocated CPU memory may temporarily surpass the CPU memory quota if over quota is used.                                                                                        |

### Subjects Authorized for the Project

Click one of the values of the Subject(s) column, to view the list of subjects and their parameters. This column is only viewable if your role in the NVIDIA Run:ai system affords you those permissions.

| Column        | Description                                                                                                                                                                                                                     |
| ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Subject       | A user, SSO group, or service account assigned with a role in the scope of this department                                                                                                                                      |
| Type          | The type of subject assigned to the access rule (user, SSO group, or service account).                                                                                                                                          |
| Scope         | The scope of this department within the organizational tree. Click the name of the scope to view the organizational tree diagram, you can only view the parts of the organizational tree for which you have permission to view. |
| Role          | The role assigned to the subject, in this department’s scope                                                                                                                                                                    |
| Authorized by | The user who granted the access rule                                                                                                                                                                                            |
| Last updated  | The last time the access rule was updated                                                                                                                                                                                       |

{% hint style="info" %}
**Note**

A role given in a certain scope, means the role applies to this scope and any descendant scopes in the organizational tree.
{% endhint %}

### Workloads Associated with the Department

Click one of the values of Workload(s) column, to view the list of workloads and their parameters

| Column                  | Description                                                                                                                                                                                      |
| ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Workload                | The name of the workload, given during its submission. Optionally, an icon describing the type of workload is also visible                                                                       |
| Type                    | The type of the workload, e.g. Workspace, Training, Inference                                                                                                                                    |
| Status                  | The state of the workload and time elapsed since the last status change                                                                                                                          |
| Created by              | The subject that created this workload                                                                                                                                                           |
| Running/ requested pods | The number of running pods out of the number of requested pods for this workload. e.g. a distributed workload requesting 4 pods but may be in a state where only 2 are running and 2 are pending |
| Creation time           | The date and time the workload was created                                                                                                                                                       |
| GPU compute request     | The amount of GPU compute requested (floating number, represents either a portion of the GPU compute, or the number of whole GPUs requested)                                                     |
| GPU memory request      | The amount of GPU memory requested (floating number, can either be presented as a portion of the GPU memory, an absolute memory size in MB or GB, or a MIG profile)                              |
| CPU memory request      | The amount of CPU memory requested (floating number, presented as an absolute memory size in MB or GB)                                                                                           |
| CPU compute request     | The amount of CPU compute requested (floating number, represents the number of requested Cores)                                                                                                  |

### Customizing the Table View

* Filter - Click ADD FILTER, select the column to filter by, and enter the filter values
* Search - Click SEARCH and type the value to search by
* Sort - Click each column header to sort by
* Column selection - Click COLUMNS and select the columns to display in the table
* Download table - Click MORE and then Click Download as CSV. Export to CSV is limited to 20,000 rows.

## Adding a New Department

To create a new Department:

1. Click **+NEW DEPARTMENT**
2. Select a **scope**.\
   By default, the field contains the scope of the current UI context cluster, viewable at the top left side of your screen. You can change the current UI context cluster by clicking the ‘Cluster: cluster-name’ field and applying another cluster as the UI context. Alternatively, you can choose another cluster within the ‘+ New Department’ form by clicking the organizational tree icon on the right side of the scope field, opening the organizational tree and selecting one of the available clusters.
3. Enter a **name** for the department. Department names must start with a letter and can only contain lower case latin letters, numbers or a hyphen ('-’).
4. Choose how to proceed:

   * Click **CREATE DEPARTMENT & CLOSE** to create the department without assigning resources.
   * Click **CREATE DEPARTMENT & CONTINUE** to assign resources and scheduling behavior.

   <div data-gb-custom-block data-tag="hint" data-style="info" class="hint hint-info"><p><strong>Note</strong></p><ul><li>The department is created and appears in the <strong>Departments</strong> grid.</li><li><p>When you select <strong>CREATE DEPARTMENT &#x26; CLOSE</strong>, the department is created without assigned resources. The department's quota is set to 0, which means workloads can run only when spare GPUs are available in the node pool. The following default settings are applied:</p><ul><li>The department is created with a <strong>Medium Low</strong> rank.</li><li>Over quota weight is always applied when distributing over-quota GPUs among departments in the same rank. By default, the department’s weight is derived proportionally from its assigned GPU quota. Since the quota is set to 0, the effective weight is 0. If <strong>Over quota weight</strong> is enabled via <strong>General settings</strong>, the department is assigned a default weight of 2.</li></ul></li></ul></div>

### Assigning GPU Quota

Assign GPUs to the department and define how the [Scheduler](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles) prioritizes node pools when running the project’s workloads.

1. Review the **node pools** available to the department. The list includes all node pools configured in the platform and represents the department’s default (suggested) configuration. This configuration determines which node pools and priority order are applied to newly created projects in the department. Administrators can override this configuration per project by selecting different node pools and adjusting their order.
2. Add **node pools** and set the default order in which workloads will be scheduled to use them. The leftmost node pool is the first one workloads will attempt to schedule onto. The Scheduler first tries to allocate resources using the highest priority node pool, then the next in priority, until it reaches the lowest priority node pool list, then the Scheduler starts from the highest again.

   <div data-gb-custom-block data-tag="hint" data-style="info" class="hint hint-info"><p><strong>Note</strong></p><ul><li>When no node pools are configured, you can assign quota for the whole department, instead of per node pool. After node pools are created, you can assign quota for each node pool separately.</li><li>The order of node pools can be overridden during workload submission by an admin policy or by the user.</li></ul></div>
3. Review the GPU quota per node pool:
   * **Unassigned** - The number of GPUs not assigned to subordinate projects.
   * **Total** - The total number of GPUs assigned to the department in this node pool.
4. Set the department's **GPU quota** in the selected node pool. A **Quota distribution** next to the department visualizes how GPUs are distributed across the department for that node pool. The graph reflects:
   * **Assigned quota** - The number of GPUs from the node pool that are assigned to this department.
   * **Unassigned quota** - The number of GPUs still available in the node pool that are not yet assigned for any department.
   * **Sum of other departments’ assigned quota** - The total number of GPUs already assigned to all other departments in this same node pool.

     <div data-gb-custom-block data-tag="hint" data-style="info" class="hint hint-info"><p><strong>Note</strong></p><p>When assigning a department GPU quota that exceeds the node pool’s available GPUs, or setting the department quota to <strong>Unlimited</strong>, the system may display warnings, but the quota is not enforced. Actual GPU consumption is constrained by the physical capacity of the node pool or cluster, regardless of the assigned quota.</p></div>
5. You can adjust **other departments' quotas** to free up additional resources. A table lists all departments and their assigned GPU quotas. Updating these values redistributes the department’s assigned and unassigned GPU quota across projects for the selected node pool, as reflected in the **Quota distribution** graph next to each department:
   * Use the **Search** icon to find a department
   * Click the **GPU** quota value and set the number of GPUs
   * Click **APPLY**
6. These steps apply to the selected node pool. If multiple node pools are configured, repeat the process for each.
7. Click **SAVE & CONTINUE**

### Defining GPU Over Quota

Define how workloads can consume GPUs beyond the department's assigned quota when additional resources are available.

1. **Allow workloads to exceed the department’s quota** is enabled by default, allowing workloads in the department to consume GPUs beyond the assigned quota. If disabled, workloads from the department will run only within the assigned quota and will not consume additional GPUs, even if resources are available.

   <div data-gb-custom-block data-tag="hint" data-style="info" class="hint hint-info"><p><strong>Note</strong></p><p>Workloads running over quota are <strong>opportunistic</strong> and may be preempted or stopped at any time when higher priority workloads require resources.</p></div>
2. Set the **max GPU allocation**. This value defines the maximum GPU device allocation the department can get from the node pool, representing the maximum sum of assigned quota and over-quota GPUs:
   * By default, this field is set to **Unlimited**, allowing the department to consume any available over-quota GPUs
   * Enter a **number** equal to or above the assigned quota
3. Rank departments in the order in which their workloads will be scheduled. A department’s **rank** determines the scheduling priority of its workloads only when competing for over-quota GPUs within the same node pool. By default, departments are assigned a **Medium Low** rank. Departments with a higher rank are scheduled before departments with a lower rank when over-quota GPUs become available. When multiple departments share the same rank, available over-quota GPUs are distributed according to each department's **weight**:
   * Filter the list of departments:
     * Click **Add filter**
     * Select **Rank** to filter by rank level, or **Department** to filter by department name
     * Enter or select the filter criteria and click **APPLY**
   * Locate and select departments:
     * Locate the department or departments whose rank you want to change
     * Select one or more departments
   * Change the rank:

     * Click **MOVE TO RANK**
     * In the **Move to Rank** dialog, select the destination rank
     * Click **MOVE \<n> DEPARTMENTS** to confirm the change

     To learn more about department ranks, see [The NVIDIA Run:ai Scheduler: concepts and principles](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles).
4. Assign **weights** to determine the departments' share of resources. A department’s weight determines how available over-quota GPUs are distributed between departments that have the same rank within the same node pool. Weight behavior depends on whether **Over quota weight** is enabled or disabled via **General settings**. When disabled, the department’s weight is proportional to its assigned GPU quota, displayed for reference, and cannot be edited. When enabled, all departments within the same rank are assigned a **default weight of 2**, and the value can be adjusted per department:
   * Filter the list of departments:
     * Click **Add filter**
     * Select **Rank** to filter by rank level, or **Department** to filter by department name
     * Enter or select the filter criteria and click **APPLY**
   * Review weight distribution:
     * Departments are grouped by rank
     * For each department, the **Weight distribution** graph shows:
       * The departments **weight**
       * The **total weight of other departments** in the same rank
   * Edit the department's weight (when Over quota weight is enabled):
     * Click the **Weight** value for the project and enter a value
     * Click **APPLY**
5. These steps apply to the selected node pool. If multiple node pools are configured, repeat the process for each.
6. Click **SAVE &** **CONTINUE**

### Assigning CPU Quota

Assign CPU resources (cores and memory) to the project and define the maximum CPU usage per node pool. This form is displayed only if **CPU quota** is enabled via the **General settings**.

1. Set the number of **CPU (Cores)** department can use in the selected node pool. By default, CPU cores are set to **Unlimited**, allowing the department to consume any available CPU resources in the node pool. Enter a **number**.
2. Set the amount of **CPU memory** available to the department in the selected node pool. By default, CPU memory is set to **Unlimited**, allowing the department to consume any available CPU resources in the node pool. Enter a **number** and select a unit (**MB**, **MiB**, or **GB**)

   <div data-gb-custom-block data-tag="hint" data-style="info" class="hint hint-info"><p><strong>Note</strong></p><ul><li>When assigning a department CPU quota that exceeds the node pool’s available CPUs, or setting the CPU quota to <strong>Unlimited</strong>, the system may display warnings, but the quota is not enforced. Actual CPU consumption is constrained by the physical capacity of the node pool or cluster, regardless of the assigned quota.</li><li>When the <strong>CPU quota</strong> feature flag is disabled, any previously set CPU quotas for a project are automatically removed.</li></ul></div>
3. These steps apply to the selected node pool. If multiple node pools are configured, repeat the process for each.
4. Click **SAVE & CONTINUE**

### Setting Scheduling Rules

Scheduling rules control how compute resources are used by workloads. The restrict either the resources (nodes) on which workloads can run or the duration of the run time. Scheduling rules are set for departments and apply to specific workload types. See [Scheduling rules](/saas/platform-management/policies/scheduling-rules) for more details.

1. Click **+RULE**
2. Select a rule type from the dropdown:
   1. **+ Idle GPU timeout** - Limit the duration of a workload whose GPU is idle
   2. **+ Workload time limit** - Set a time limit for workspaces regardless of their activity (e.g., stop the workspace after 1 day of work)
   3. **+ Node type (Affinity**) - Limit workloads to run on specific node types
3. For each rule, select the **workload type(s)** to which the rule applies and configure the rule parameters
4. Click **SAVE & CONTINUE**

Once scheduling rules are set for a department, all matching workloads associated with the department have the restrictions applied to them, as defined, when the workload is submitted. New scheduling rules added to a department are not applied over previously created workloads associated with that department.

{% hint style="info" %}
**Note**

Setting scheduling rules in a department enforces the rules on all associated projects.
{% endhint %}

### Applying Access Rules

1. Click **+ACCESS RULE**
2. Select a subject - **User, SSO group**, or **Service account**
3. Select or enter the subject identifier. You can define up to 10 subjects of the selected type:
   * **User email** for a local user created in NVIDIA Run:ai or for SSO user as recognized by the IDP
   * **Group name** as recognized by the IDP
   * **Service account name** as created in NVIDIA Run:ai
4. Select a **role**
5. Click **SAVE RULE**
6. Click **CLOSE**

## Editing a Department

1. Select the Department you want to edit
2. Click **EDIT**
3. From the dropdown menu, select the form you want to modify
4. Update the settings and click **SAVE & CONTINUE**

{% hint style="info" %}
**Note**

When editing a department:

* You can navigate directly to a specific form without completing the full department creation flow again.
* Changes apply only to the selected form; previously configured settings in other forms remain unchanged.
* You cannot set GPU and CPU quotas below the quotas currently consumed by **non-preemptible workloads**.
* You cannot set CPU or GPU quotas that are less than the quota already assigned to its **subordinate projects**.
  {% endhint %}

## Viewing a Department’s Policy

To view the policy of a department:

1. Select the department for which you want to view its [policies](/saas/platform-management/policies/native-workload-policies).\
   This option is only active if the department has defined policies in place.
2. Click **VIEW POLICY** and select the workload type for which you want to view the policies
3. In the Policy form, view the workload rules that are enforcing your department for the selected workload type as well as the defaults:
   * **Parameter** - The workload submission parameter that Rule and Default is applied on
   * **Type (applicable for data sources only)** - The data source type (Git, S3, nfs, pvc etc.)
   * **Default** - The default value of the Parameter
   * **Rule** - Set up constraints on workload policy fields
   * **Source** - The origin of the applied policy (cluster, department or project)

{% hint style="info" %}
**Note**

* The policy affecting the department consists of rules and defaults. Some of these rules and defaults may be derived from the policies of a parent cluster (source). You can see the source of each rule in the policy form.
* A policy set for a department affects all subordinated projects and their workloads, according to the policy workload type
  {% endhint %}

## Deleting a Department

1. Select the department you want to delete
2. Click **DELETE**
3. On the dialog, click **DELETE** to confirm the deletion

{% hint style="info" %}
**Note**

Deleting a department permanently deletes its subordinated projects, any assets created in the scope of this department, and any of its subordinated projects such as compute resources, environments, data sources, templates, and credentials. However, workloads running within the department’s subordinated projects, or the policies defined for this department or its subordinated projects - remain intact and running.
{% endhint %}

## Using CLI

To view the available actions on departments, see the department [CLI v2 reference](/saas/reference/cli/runai/runai-department).

## Using API

To view the available actions, go to the [Departments](https://run-ai-docs.nvidia.com/api/organizations/departments) API reference.


# Managing Your Resources


# Nodes

Nodes are Kubernetes elements automatically discovered by the NVIDIA Run:ai platform. Once a node is discovered by the NVIDIA Run:ai platform, an associated instance is created in the Nodes table, administrators can view the Node’s relevant information, and NVIDIA Run:ai scheduler can use the node for [Scheduling](/saas/platform-management/runai-scheduler/scheduling/how-the-scheduler-works).

## Nodes Table

The Nodes table can be found under **Resources** in the NVIDIA Run:ai platform.

The Nodes table displays a list of predefined nodes available to users in the NVIDIA Run:ai platform.

{% hint style="info" %}
**Note**

* It is not possible to create additional nodes, or edit, or delete existing nodes.
* Only users with relevant permissions can view the table.
  {% endhint %}

<figure><img src="/files/49FbQ4yqWMeZqIVU86Ax" alt=""><figcaption></figcaption></figure>

The Nodes table consists of the following columns:

| Column                    | Description                                                                                                                                                                                                                                            |
| ------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Node                      | The Kubernetes name of the node                                                                                                                                                                                                                        |
| Status                    | The state of the node. Nodes in the Ready state are eligible for scheduling. If the state is Not ready then the main reason appears in parenthesis on the right side of the state field. Hovering the state lists the reasons why a node is Not ready. |
| Node pool                 | The name of the associated node pool. By default, every node in the NVIDIA Run:ai platform is associated with the default node pool, if no other node pool is associated.                                                                              |
| Cluster                   | The cluster that the node is associated with                                                                                                                                                                                                           |
| NVLink domain UID         | Indicates if the MNNVL domain ID is part of the MNNVL label value. In case the MNNVL label is not the default MNNVL label key (`nvidia.com/gpu.clique`), this field will show the whole label value.                                                   |
| MNNVL domain clique ID    | Indicates if the MNNVL clique ID is part of the MNNVL label value. In case the MNNVL label is not the default MNNVL label key (`nvidia.com/gpu.clique`), this field will show an empty value.                                                          |
| GPU type                  | The GPU model, for example, H100, or V100                                                                                                                                                                                                              |
| Ready / total GPU devices | The number of GPU devices installed on the node. Clicking this field pops up a dialog with details per GPU (described below in this article).                                                                                                          |
| GPU memory                | The total amount of GPU memory installed on this node. For example, if the number is 640GB and the number of GPU devices is 8, then each GPU is installed with 80GB of memory (assuming the node is assembled of homogenous GPU devices).              |
| Allocated GPUs            | The total allocation of GPU devices in units of GPUs (decimal number). For example, if 3 GPUs are 50% allocated, the field prints out the value 1.50. This value represents the portion of GPU memory consumed by all running pods using this node.    |
| Free GPU devices          | The current number of fully vacant GPU devices                                                                                                                                                                                                         |
| CPU (Cores)               | The number of CPU cores installed on this node                                                                                                                                                                                                         |
| CPU memory                | The total amount of CPU memory installed on this node                                                                                                                                                                                                  |
| Allocated CPU (Cores)     | The number of CPU cores allocated by pods running on this node (decimal number, e.g. a pod allocating 350 mili-cores shows an allocation of 0.35 cores).                                                                                               |
| Allocated CPU memory      | The total amount of CPU memory allocated by pods running on this node (in GB or MB)                                                                                                                                                                    |
| Pod(s)                    | List of pods running on this node, click the field to view details (described below in this article)                                                                                                                                                   |

### GPU Devices for Node

Click one of the values in the GPU devices column, to view the list of GPU devices and their parameters.

| Column              | Description                                                                                        |
| ------------------- | -------------------------------------------------------------------------------------------------- |
| Index               | The GPU index, read from the GPU hardware. The same index is used when accessing the GPU directly. |
| Allocated compute   | The total amount of GPU compute allocated                                                          |
| Allocated memory    | The total amount of GPU memory allocated                                                           |
| Used memory         | The amount of memory used by pods and drivers using the GPU (in GB or MB)                          |
| Compute utilization | The portion of time the GPU is being used by applications (percentage)                             |
| Memory utilization  | The portion of the GPU memory that is being used by applications (percentage)                      |
| Idle time           | The elapsed time since the GPU was used (i.e. the GPU is being idle for ‘Idle time’)               |

### Pods Associated with a Node

Click one of the values in the Pod(s) column, to view the list of pods and their parameters.

{% hint style="info" %}
**Note**

This column is only viewable if your role in the NVIDIA Run:ai platform gives you read access to workloads, even if you are allowed to view workloads, you can only view the workloads within your allowed scope. This means, there might be more pods running on this node then appear in the list you are viewing.
{% endhint %}

| Column                | Description                                                                                                                                                                                                                |
| --------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Pod                   | The Kubernetes name of the pod. Usually name of the pod is made of the name of the parent workload if there is one, and an index for unique for that pod instance within the workload.                                     |
| Status                | The state of the pod. In steady state this should be Running and the amount of time the pod is running.                                                                                                                    |
| Project               | The NVIDIA Run:ai project name the pod belongs to. Clicking this field takes you to the Projects table filtered by this project name.                                                                                      |
| Workload              | The workload name the pod belongs to. Clicking this field takes you to the Workloads table filtered by this workload name.                                                                                                 |
| Allocated GPUs        | The total allocation of GPU devices in units of GPUs (decimal number). For example, if 3 GPUs are 50% allocated, the field prints out the value 1.50. This value represents the portion of GPU memory consumed by the pod. |
| Allocated GPU memory  | The total amount of GPU memory allocated by the pod (in GB or MB)                                                                                                                                                          |
| Allocated CPU (Cores) | The number of CPU cores allocated by the pod (decimal number, e.g. a pod allocating 350 mili-cores shows an allocation of 0.35 cores)                                                                                      |
| Allocated CPU memory  | The total amount of CPU memory allocated by the pod (in GB or MB)                                                                                                                                                          |
| Image                 | The full path of the image used by the main container of this pod                                                                                                                                                          |
| IP                    | The unique internal IP address assigned to a pod                                                                                                                                                                           |
| Creation time         | The pod’s creation date and time                                                                                                                                                                                           |

### Customizing the Table View

* Filter - Click ADD FILTER, select the column to filter by, and enter the filter values
* Search - Click SEARCH and type the value to search by
* Sort - Click each column header to sort by
* Column selection - Click COLUMNS and select the columns to display in the table
* Download table - Click MORE and then Click Download as CSV. Export to CSV is limited to 20,000 rows.
* Show/Hide details - Click to view additional information on the selected row

### Show/Hide Details

Click a row in the Nodes table, then click the Show details button at the upper right of the action bar. The Metrics screen appears, containing a dropdown allowing you to switch between **Resource utilization** and **GPU profiling** metrics views:

* **Resource utilization** - Displays general GPU, CPU and network metrics
* **GPU profiling** - Shows additional NVIDIA-specific metrics

{% hint style="info" %}
**Note**

GPU profiling metrics are disabled by default. If unavailable, your administrator must enable it under **General settings** → Analytics → GPU profiling metrics. Before enabling, the administrator must configure GPU profiling through the DCGM Exporter and NVIDIA Run:ai Prometheus integration. For configuration steps, see [GPU profiling metrics](/saas/platform-management/monitor-performance/gpu-profiling-metrics).
{% endhint %}

#### Resource Utilization

* **GPU utilization**\
  Per GPU graph and an average of all GPUs graph, all on the same chart, along an adjustable period allows you to see the trends of all GPUs compute utilization (percentage of GPU compute) in this node.
* **GPU memory utilization**\
  Per GPU graph and an average of all GPUs graph, all on the same chart, along an adjustable period allows you to see the trends of all GPUs memory usage (percentage of the GPU memory) in this node.
* **CPU compute utilization**\
  The average of all CPUs’ cores compute utilization graph, along an adjustable period allows you to see the trends of CPU compute utilization (percentage of CPU compute) in this node.
* **CPU memory utilization**\
  The utilization of all CPUs memory in a single graph, along an adjustable period allows you to see the trends of CPU memory utilization (percentage of CPU memory) in this node.
* **CPU memory usage**\
  The usage of all CPUs memory in a single graph, along an adjustable period allows you to see the trends of CPU memory usage (in GB or MB of CPU memory) in this node.
* **NVLink bandwidth total**\
  The rate of data transmitted / received over NVLink, not including protocol headers, in bytes per second. The value represents an average over a time interval and is not an instantaneous value. The rate is averaged over the time interval. For example, if 1 GB of data is transferred over 1 second, the rate is 1 GB/s regardless of the data transferred at a constant rate or in bursts. The theoretical maximum NVLink Gen2 bandwidth is 25 GB/s per link per direction.

#### GPU Profiling

Select **GPU profiling** from the dropdown to view extended GPU device-level metrics such as memory bandwidth, SM occupancy, and other data.

* For the full list of supported metrics, see [Metrics and telemetry](/saas/platform-management/monitor-performance/metrics#gpu-profiling).
* For definitions and technical descriptions, refer to the [NVIDIA documentation](https://docs.nvidia.com/datacenter/dcgm/latest/user-guide/feature-overview.html#profiling-metrics).

#### Navigating the Graphs

* For GPUs charts - Click the GPU legend on the right-hand side of the chart, to activate or deactivate any of the GPU lines.
* You can click the date picker to change the presented period
* You can use your mouse to mark a sub-period in the graph for zooming in, and use the ‘Reset zoom’ button to go back to the preset period
* Changes in the period affect all graphs on this screen.

## Using API

To view the available actions, go to the [Nodes](https://run-ai-docs.nvidia.com/api/organizations/nodes) API reference.


# Configuring NVIDIA MIG Profiles

NVIDIA’s Multi-Instance GPU (MIG) enables splitting a GPU into multiple logical GPU devices, each with its own memory and compute portion of the physical GPU.

NVIDIA provides two MIG strategies:

* **Single** - A GPU can be divided evenly. This means all MIG profiles are the same.
* **Mixed** - A GPU can be divided into different profiles.

The NVIDIA Run:ai platform supports running workloads using NVIDIA MIG. Administrators can set the Kubernetes nodes to their preferred MIG strategy and configure the appropriate MIG profiles for researchers and MLOPS engineers to use.

This guide explains how to configure MIG in each strategy to [submit workloads](/saas/workloads-in-nvidia-run-ai/workloads). It also outlines the individual implications of each strategy and best practices for administrators.

{% hint style="info" %}
**Note**

* Starting from v2.19, Dynamic MIG feature began a [deprecation process](https://docs.run.ai/v2.19/home/whats-new-2-19/#dynamic-mig-deprecation) and is now no longer supported. With Dynamic MIG, the NVIDIA Run:ai platform automatically configured MIG profiles according to on-demand user requests for different MIG profiles or memory fractions.
* GPU fractions and memory fractions are not supported with MIG profiles.
* Single strategy supports both NVIDIA Run:ai and third-party workloads. Using mixed strategy can only be done using third-party workloads. For more details on NVIDIA Run:ai and third-party workloads, see [Introduction to workloads](/saas/workloads-in-nvidia-run-ai/introduction-to-workloads).
  {% endhint %}

## Before You Start

To use MIG single and mixed strategy effectively, make sure to familiarize yourself with the following NVIDIA resources:

* [NVIDIA Multi-Instance GPU](https://www.nvidia.com/en-eu/technologies/multi-instance-gpu/)
* [MIG User Guide](https://docs.nvidia.com/datacenter/tesla/mig-user-guide/index.html)
* [GPU Operator with MIG](https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/gpu-operator-mig.html)

## Configuring Single MIG Strategy

When deploying MIG using single strategy, all GPUs within a [node](/saas/platform-management/aiinitiatives/resources/nodes) are configured with the same profile. For example, a node might have GPUs configured with 3 MIG slices of profile type 2g.20gb, or 7 MIG slices of profile 1g.10gb. With this strategy, MIG profiles are displayed as whole GPU devices by CUDA.

The NVIDIA Run:ai platform discovers these MIG profiles as whole GPU devices as well, ensuring MIG devices are transparent to the end-user (practitioner). For example, a node that consists of 8 physical GPUs split into MIG slices, 3×2g20gb slices each, is discovered by the NVIDIA Run:ai platform as a node with 24 GPU devices.

Users can submit workloads by requesting a specific number of GPU devices (X GPU) and NVIDIA Run:ai will allocate X MIG slices (logical devices). The NVIDIA Run:ai platform deducts X GPUs from the workload’s [Project quota](/saas/platform-management/aiinitiatives/organization/projects), regardless of whether this ‘logical GPU’ represents 1/3 of a physical GPU device or 1/7 of a physical GPU device.

## Configuring Mixed MIG Strategy

When deploying MIG using mixed strategy, each GPU in a [node](/saas/platform-management/aiinitiatives/resources/nodes) can be configured with a different combination of MIG profiles such as 2×2g.20gb and 3×1g.10gb. For details on supported combinations per GPU type, refer to [Supported MIG Profiles](https://docs.nvidia.com/datacenter/tesla/mig-user-guide/index.html#supported-mig-profiles).

In mixed strategy, physical GPU devices continue to be displayed as physical GPU devices by CUDA, and each MIG profile is shown individually. The NVIDIA Run:ai platform identifies the physical GPU devices normally, however, MIG profiles are not visible in the UI or node APIs.

When submitting third-party workloads with this strategy, the user should explicitly specify the exact requested MIG profile (for example, nvidia.com/gpu.product: A100-SXM4-40GB-MIG-3g.20gb). The NVIDIA Run:ai [Scheduler](/saas/platform-management/runai-scheduler/scheduling/how-the-scheduler-works) finds a node that can provide this specific profile and binds it to the workload.

A third-party workload submitted with a MIG profile of type Xg.Ygb (e.g. 3g.40gb or 2g.20gb) is considered as consuming X GPUs. These X GPUs will be deducted from the workload’s project quota of GPUs. For example, a 3g.40gb profile deducts 3 GPUs from the associated Project’s quota, while 2g.20gb deducts 2 GPUs from the associated Project’s quota. This is done to maintain a logical ratio according to the characteristics of the MIG profile.

## Best Practices for Administrators

### Single Strategy

* Configure proper and uniform sizes of MIG slices (profiles) across all GPUs within a node.
* Set the same MIG profiles on all nodes of a single [node pool](/saas/platform-management/aiinitiatives/resources/node-pools).
* Create separate node pools with different MIG profile configurations allowing users to select the pool that best matches their workloads’ needs.
* Ensure Project quotas are allocated according to the MIG profile sizes.

### Mixed Strategy

* Use mixed strategy with workloads that require diverse resources. Make sure to evaluate the workload requirements and plan accordingly.
* Configure individual MIG profiles on each node by using a limited set of MIG profile combinations to minimize complexity. Make sure to evaluate your requirements and node configurations.
* Ensure Project quotas are allocated according to the MIG profile sizes.

{% hint style="info" %}
**Note**

Since MIG slices are a fixed size, once configured, changing MIG profiles requires administrative intervention.
{% endhint %}


# Using GB200 NVL72 and Multi-Node NVLink Domains

Multi-Node NVLink (MNNVL) systems, including NVIDIA GB200, NVIDIA GB200 NVL72 and its derivatives are fully supported by the NVIDIA Run:ai platform.

Kubernetes does not natively recognize NVIDIA’s MNNVL architecture, which makes managing and scheduling workloads across these high-performance domains more complex. The NVIDIA Run:ai platform simplifies this by abstracting the complexity of MNNVL configuration. Without this abstraction, optimal performance on a GB200 NVL72 system would require deep knowledge of NVLink domains, their hardware dependencies, and manual configuration for each distributed workload. NVIDIA Run:ai automates these steps, ensuring high performance with minimal effort. While GB200 NVL72 supports all [workload types](/saas/workloads-in-nvidia-run-ai/workload-types), distributed workloads - training and inference - benefit most from its accelerated GPU networking capabilities.

{% hint style="info" %}
**Note**

Distributed workloads include both distributed training and distributed inference.
{% endhint %}

To learn more about GB200, MNNVL and related NVIDIA technologies, refer to the following:

* [NVIDIA GB200 NVL72](https://www.nvidia.com/en-us/data-center/gb200-nvl72/)
* [NVIDIA Blackwell datasheet](https://nvdam.widen.net/s/wwnsxrhm2w/blackwell-datasheet-3384703)
* [NVIDIA Multi-Node NVLink Systems](https://docs.nvidia.com/multi-node-nvlink-systems/)

## Benefits of Using GB200 NVL72 with NVIDIA Run:ai

The NVIDIA Run:ai platform enables administrators, researchers, and MLOps engineers to fully leverage GB200 NVL72 systems and other NVLink-based domains without requiring deep knowledge of hardware configurations or NVLink topologies. Key capabilities include:

* **Automatic detection and labeling**
  * Detects GB200 NVL72 nodes and identifies MNNVL domains (e.g., GB200 NVL72 racks).
  * Automatically detects whether a node pool contains GB200 NVL72.
  * Supports manual override of GB200 MNNVL detection and label key for future compatibility and improved resiliency.
* **Simplified distributed workload submission**
  * Allows seamless submission of distributed workloads into GB200-based node pools, eliminating all the complexities involved with that operation on top of GB200 MNNVL domains.
  * Abstracts away the complexity of configuring workloads for NVL domains.
  * Automatically optimizes distributed workloads performance.
* **Flexible support for NVLink domain variants**
  * Compatible with current and future NVL domain configurations.
  * Supports any number of domains or GB200 racks.
* **Enhanced monitoring and visibility**
  * Provides detailed NVIDIA Run:ai dashboards for monitoring GB200 nodes and MNNVL domains by node pool.
* **Control and customization**
  * Offers manual override and label configuration for greater resiliency and future-proofing.
  * Enables advanced users to fine-tune GB200 scheduling behavior based on workload requirements.
* **Full support for elastic distributed workloads**
  * Elastic distributed workloads (auto-scaling or dynamically sized distributed workloads) are fully supported on GB200 NVL72 and MNNVL domains.
  * The NVIDIA Run:ai platform automatically applies ComputeDomain configuration and topology-aware scheduling rules to ensure elastic workloads scale within the same NVLink domain while maintaining optimal performance.
  * Scaling events such as adding or removing worker replicas respect NVLink domain boundaries and inherit the same high-bandwidth placement optimizations as fixed-size distributed workloads.

{% hint style="info" %}
**Note**

Elastic distributed workloads are supported on GB200 NVL72 and Multi-Node NVLink (MNNVL) domains using NVIDIA DRA driver version 25.8 and later.
{% endhint %}

## Prerequisites

* **Kubernetes version** - Requires Kubernetes 1.32 or later.
* **NVIDIA GPU Operator** - Install NVIDIA GPU Operator version 25.3 or above. See the [NVIDIA GPU Operator](/saas/getting-started/installation/install-using-helm/system-requirements#nvidia-gpu-operator) section for installation instructions. This version must include the associated **Dynamic Resource Allocation (DRA) driver**, which provides support for GB200 accelerated networking resources and the ComputeDomain feature. For detailed steps on installing the DRA driver and configuring ComputeDomain, refer to [NVIDIA Dynamic Resource Allocation (DRA) Driver](/saas/getting-started/installation/install-using-helm/system-requirements#nvidia-dynamic-resource-allocation-dra-driver).
* **NVIDIA Network Operator** - Install the NVIDIA Network Operator. See the [NVIDIA Network Operator](/saas/getting-started/installation/install-using-helm/system-requirements#nvidia-network-operator) section for installation instructions.
* **Enable GPU network acceleration** - After installation, update the cluster configuration to enable GPU network acceleration by setting the appropriate flag for the relevant controller. Enabling either flag triggers an update of the corresponding NVIDIA Run:ai controller deployment and automatically restarts the controller to apply the change. Configure the flag, `GPUNetworkAccelerationEnabled=true`, for the controller that applies to your workloads. For details on how to configure this value using Helm or `runaiconfig`, see [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config).

## Configuring and Managing GB200 NVL72 Domains

Administrators must define dedicated node pools that align with GB200 NVL72 rack topologies. These node pools ensure that workloads are isolated to nodes with NVLink interconnects and are not scheduled on incompatible hardware. Each node pool can be manually configured in the NVIDIA Run:ai platform and associated with specific node labels. Two key configurations are required for each node pool:

* **Node Labels** - Identify nodes equipped with GB200.
* **MNNVL Domain Discovery** - Specify how the platform detects whether the node pool includes NVLink-connected nodes.

To create a node pool with GPU network acceleration, see [Node pools](/saas/platform-management/aiinitiatives/resources/node-pools).

### Identifying GB200 Nodes

To enable the [NVIDIA Run:ai Scheduler](/saas/platform-management/runai-scheduler/scheduling/how-the-scheduler-works) to recognize GB200-based nodes, administrators must:

* Use the default node label provided by the NVIDIA GPU Operator - `nvidia.com/gpu.clique`.
* Or, apply a custom label that clearly marks the node as GB200/MNNVL capable.

This node label serves as the basis for identifying appropriate nodes and ensuring workloads are scheduled on the correct hardware.

### Enabling MNNVL Domain Discovery

The administrator can configure how the NVIDIA Run:ai platform detects MNNVL domains for each node pool. The available options include:

* **Auto-detect** - Uses the default label key `nvidia.com/gpu.clique`, or a custom label key specified by the administrator. The NVIDIA Run:ai platform automatically discovers MNNVL domains within node pools. If a node is labeled with the MNNVL label key, the NVIDIA Run:ai platform indicates this node pool as MNNVL detected. MNNVL detected node pools are treated differently by the NVIDIA Run:ai platform when submitting a distributed workload.
* **MNNVL is present** - Explicitly indicate that the node pool contains MNNVL nodes
* **MNNVL is not present** - Explicitly indicate that the node pool does not contain MNNVL nodes

When auto-detect is enabled, all GB200 nodes that are part of the same physical rack (NVL72 or other future topologies) are part of the same NVL Domain and automatically labeled by the GPU Operator with a common label using a unique label value per domain and sub-domain. The default label key set by the NVIDIA GPU Operator is `nvidia.com/gpu.clique` and its value consists of - `<NVL Domain ID (ClusterUUID)>.<Clique ID>` :

* The **NVL Domain ID (ClusterUUID)** is a unique identifier that represents the physical NVL domain, for example, a physical GB200 NVL72 rack.
* The **Clique ID** denotes a logical MNNVL sub-domain. A clique represents a further logical split of the MNNVL into smaller domains that enable secure, fast, and isolated communication between pods running on different GB200 nodes within the same GB200 NVL72.

The [Nodes table](/saas/platform-management/aiinitiatives/resources/nodes) provides more information on which GB200 NVL72 domain each node belongs to, and which Clique ID it is associated with.

### Submitting Distributed Workloads

When a distributed workload is submitted to an MNNVL node pool, the NVIDIA Run:ai platform automates several key configuration steps to ensure optimal workload execution:

* **ComputeDomain creation** - The NVIDIA Run:ai platform creates a ComputeDomain Custom Resource Definition (CRD), which is a proprietary resource used to manage NVLink-based domain assignments.
* **Resource Claim injection** - A reference to the ComputeDomain is automatically added to the workload specification as a resource claim, allowing the Scheduler to link the workload to a specific NVLink domain.
* **Network topology injection** - If network topology is configured on the GB200 node pool, the NVIDIA Run:ai platform adds a `NetworkTopology` request to the distributed workloads. This applies a default Preferred constraint at the lowest defined topology level, ensuring pods are placed as close as possible in the network hierarchy. The Scheduler uses the defined topology hierarchy to minimize communication overhead and improve workload performance, delivering more optimal scheduling results than the `PodAffinity` plugin. While pod affinity evaluates placement one pod at a time, network topology is aware of all pods in the workload and optimizes their placement accordingly. See [Accelerating workloads with network topology-aware scheduling](/saas/platform-management/aiinitiatives/resources/topology-aware-scheduling) for more details.
* **Pod affinity configuration** (if network topology is not configured) - Pod affinity is applied using a Preferred policy with the MNNVL label key (e.g., `nvidia.com/gpu.clique`) as the topology key. This ensures that pods within the distributed workload are located on nodes with NVLink interconnects.
* **Node affinity configuration** - Node affinity is also applied using a Preferred policy based on the same label key, further guiding the Scheduler to place workloads within the correct node group.

These additional steps are crucial for the creation of underlying HW resources (also known as IMEX channels) and stickiness of the distributed workload to MNNVL topologies and nodes. When a distributed workload is stopped or evicted, the platform automatically removes the corresponding ComputeDomain.

## Best Practices for MNNVL Node Pool Management

* Attaching a network topology to an MNNVL node pool, especially one that includes the MNNVL level label (e.g. GB200 rack) can significantly improve the Scheduler's optimal placement of distributed workloads.
* Avoid using the hostname label in network topology for MNNVL node pools. Assigning more than one pod to the same hostname will cause the request to fail, since a compute domain can schedule only one pod per node. Instead, rely on higher-level topology labels to ensure proper placement and avoid scheduling conflicts.
* When submitting a distributed workload, you should explicitly specify a list of one or more MNNVL node pools, or a list of one or more non-MNNVL node pools. A mix of MNNVL and non-MNNVL node pools is not supported. A GB200 MNNVL node pool is a pool that contains at least one node belonging to an MNNVL domain.
* Other workload types (not distributed) can include a list of mixed MNNVL and non-MNNVL node pools, from which the Scheduler will choose.
* MNNVL node pools can include any size of MNNVL domains (i.e. NVL72 and any future domain size) and support any Grace-Blackwell models (GB200 and any future models).
* To support the submission of larger distributed workloads, it is recommended to group as many GB200 racks as possible into fewer node pools. When possible, use a single GB200 node pool, unless there is a specific operational reason to divide resources across multiple node pools.
* When submitting distributed workloads with the controller pod set as a distinct non-GPU workload, the MNNVL feature should be used with the default Preferred mode as explained in the below section.

## Fine-Tuning Distributed Workload Constraints

You can influence how the Scheduler places distributed workloads into GB200 MNNVL node pools using the **Topology** field available in the [distributed training](/saas/workloads-in-nvidia-run-ai/using-training/distributed-training-models) and [distributed inference](/saas/workloads-in-nvidia-run-ai/using-inference/distributed-inference) workload submission forms.

{% hint style="info" %}
**Note**

The following options are based on [inter-pod affinity rules](https://kubernetes.io/docs/concepts/scheduling-eviction/assign-pod-node/), which define how pods are grouped based on topology.
{% endhint %}

* **Confine a workload to a single GB200 MNNVL domain** - To ensure the workload is scheduled within a single GB200 MNNVL domain (e.g., a GB200 NVL72 rack), apply a label with a Required policy using the MNNVL label key (`nvidia.com/gpu.clique`). This instructs the Scheduler to strictly place all pods within the same MNNVL domain. If the workload exceeds 18 pods (or 72 GPUs), the Scheduler will not be able to find a matching domain and will fail to schedule the workload.
* **Try to schedule a workload using a Preferred topology** - To guide the Scheduler to prioritize a specific topology without enforcing it, apply a label with a policy of Preferred. You can apply any label with a Preferred policy. These labels are treated with higher scheduling weight than the default Preferred pod affinity automatically applied by NVIDIA Run:ai for MNNVL.
* **Mandate a custom topology** - To force scheduling a workload into a custom topology, add a label with a policy of Required. This ensures the workload is strictly scheduled according to the specified topology. Keep in mind that using a Required policy can significantly constrain scheduling. If matching resources are not available, the Scheduler may fail to place the workload.

## Fine-tuning MNNVL per Workload

You can customize how the NVIDIA Run:ai platform applies the MNNVL feature to each distributed workload. This allows you to override the default behavior when needed. To configure this behavior, set the proprietary label key `run.ai/MNNVL` in the **General settings** section of the [distributed training](/saas/workloads-in-nvidia-run-ai/using-training/distributed-training-models) and [distributed inference](/saas/workloads-in-nvidia-run-ai/using-inference/distributed-inference) workload submission forms. The following values are supported:

* **None** - Disables the MNNVL feature for the workload. The platform does not create a ComputeDomain and no pod affinity or node affinity is applied by default.
* **Preferred** (default) *-* Indicates that MNNVL feature is preferred but not required. This is the default behavior when submitting a distributed workload:
  * If the workload is submitted to a non-MNNVL node pool, then the NVIDIA Run:ai platform does not add a ComputeDomain, ComputeDomain claim, pod affinity or node affinity for MNNVL nodes.
  * Otherwise, if the workload is submitted to a 'MNNVL' node pool, then the NVIDIA Run:ai platform automatically adds: ComputeDomain, ComputeDomain claim, NodeAffinity and PodAffinity both with a Preferred policy and using the MNNVL label.
  * If you manually add an additional Preferred topology label, it will be given higher scheduling weight than the default embedded pod affinity (which has weight = 1).
* **Required** - Enforces a strict use of MNNVL domains for the workload. The workload must be scheduled on MNNVL supported nodes:
  * The NVIDIA Run:ai platform creates a ComputeDomain and ComputeDomain claim.
  * The NVIDIA Run:ai platform will automatically add a node affinity rule with a Required policy using the appropriate label.
  * Pod affinity is set to Preferred by default, but you can override it manually with a Required pod affinity rule using the MNNVL label key or another custom label.
  * If any of the targeted node pools do not support MNNVL or if the workload (or any of its pods) does not request GPU resources, the workload will fail to run.

## Known Limitations and Compatibility

* If the DRA driver is not installed correctly in the cluster, particularly if the required CRDs are missing, and the MNNVL feature is enabled in the NVIDIA Run:ai platform, the workload controller will enter a crash loop. This will continue until the DRA driver is properly installed with all necessary CRDs or the MNNVL feature is disabled in the NVIDIA Run:ai platform.
* To run workloads on a GB200 node pool (i.e., a node pool MNNVL-enabled), the workload must explicitly request that node pool. To prevent unintentional use of MNNVL node pools, administrators must ensure these node pools are not included in any project's default list of node pools.
* Only one distributed workload per node can use GB200 accelerated networking resources. If GPUs remain unused on that node, other workload types may still utilize them.
* Workloads created in versions earlier than 2.21 do not include GB200 MNNVL node pools and are therefore not expected to experience compatibility issues.
* If a node pool that was previously used in a workload submission is later updated to include GB200 nodes (i.e., becomes a mixed node pool), the workload submitted before version 2.21 will not use any accelerated networking resources, although it may still run on GB200 nodes.


# Accelerating Workloads with Network Topology-Aware Scheduling

Topology-aware scheduling in NVIDIA Run:ai optimizes the placement of workloads across data center nodes by leveraging knowledge of the underlying network topology. In modern AI/ML clusters, communication between the different pods of a distributed workload can be a significant performance bottleneck. By scheduling workload pods on nodes that are “closer” to each other in the network (e.g., same rack, block, NVLink domain), NVIDIA Run:ai reduces communication overhead and improves workload efficiency.

Kubernetes represents hierarchical network structures, such as racks, blocks, or NVLink domains, using node labels. These labels describe where each node resides in the cluster's network topology. The [NVIDIA Run:ai Scheduler](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles) uses this topology information to keep workloads on nodes that minimize latency and maximize bandwidth availability.

{% hint style="info" %}
**Note**

For guidance on using topology-aware scheduling with GB200 and Multi-Node NVLink (MNNVL) systems, see [Using GB200 and Multi-Node NVLink Domains](/saas/platform-management/aiinitiatives/resources/using-gb200).
{% endhint %}

## Benefits of Network Topology-Aware Scheduling

* **Improved performance for distributed workloads** - Reduces inter-node communication latency by scheduling pods on nodes closer to each other.
* **Optimized GPU utilization** - Keeps workloads within NVLink/NVL72 domains where possible, leveraging high-bandwidth interconnects.
* **Multi-level topology** - Supports multi-level topology definitions (e.g., rack → block → node), giving administrators fine-grained control. The NVIDIA Run:ai Scheduler considers all levels when placing workloads. If scheduling cannot occur at a lower level, it automatically moves up the hierarchy, attempting placement layer by layer in order.
* **Seamless distributed workload experience** - Topology-aware scheduling for distributed workloads is applied automatically once an administrator has configured the network topology. This ensures performance gains without requiring any additional user configuration.

## Topology Labels in Kubernetes

Topology-aware scheduling relies on Kubernetes node labels that describe each node’s location in the topology. These labels are applied at the Kubernetes level and are managed outside of NVIDIA Run:ai.

How labels are created depends on the environment and user setup. Users may apply labels manually or use topology discovery tools such as [Topograph](https://github.com/NVIDIA/topograph). In cloud environments, labels are often applied automatically by the cloud provider. When NVIDIA NetQ is used in on-prem environments, Topograph can integrate with NetQ to obtain network topology and real-time network telemetry data and derive the appropriate labels based on the physical network layout. See the [NVIDIA NetQ User Guide](https://docs.nvidia.com/networking-ethernet-software/cumulus-netq/) for more details.

Once the labels are in place, administrators only need to specify the **label keys** (as described below) that represent the topology which the NVIDIA Run:ai Scheduler should reference when making placement decisions.

## Configuring Network Topology

When creating or editing a cluster, administrators assign a network topology that represents the connectivity of nodes within the cluster. This topology is defined through Kubernetes node label keys. See [Managing network topologies](/saas/infrastructure-setup/procedures/clusters#managing-network-topologies) for more details.

These labels must match those configured on the cluster nodes and are used by the NVIDIA Run:ai Scheduler to guide workload placement decisions. The order of labels defines the hierarchy:

* The **first label** represents the farthest point in the network (for example, a region).
* The **last label** represents the closest switch or node (for example, a hostname or rack).

Example:

```json
{
  "levels": [
    "topology.kubernetes.io/region",
    "topology.kubernetes.io/zone",
    "cloud.provider.com/topology-block",
    "cloud.provider.com/topology-rack",
    "kubernetes.io/hostname"
  ],
  "name": "default-topology",
  "clusterId": "<CLUSTER_ID>"
}
```

After creating a topology, the administrator must associate it with the relevant [node pool(s)](/saas/platform-management/aiinitiatives/resources/node-pools#adding-a-new-node-pool).

* If you are using the default node pool (that is, no additional node pools are defined), attach the topology to the default node pool.
* If different node pools have different topologies, each node pool must be linked to its corresponding topology.
* If the entire cluster shares the same topology, link the same topology to all node pools.

## Automatic Topology-Aware Scheduling

When a distributed workload is submitted, the platform automatically applies topology-aware scheduling based on the network topology configured on the target node pool. This behavior ensures that distributed workloads benefit from improved performance without additional user configuration.

Topology-aware scheduling in NVIDIA Run:ai is applied at the **workload level**. This means the Scheduler considers the entire distributed workload as a single unit and places all of its pods according to the same topology constraints.

NVIDIA Run:ai automatically applies a **Preferred** topology constraint at the lowest defined topology level. This co-locates pods as close as possible in the network hierarchy, reducing communication overhead. If the Scheduler cannot place pods at this level, it automatically escalates placement by moving up through the topology hierarchy (for example: from node → rack → block → zone), always seeking the closest available level to minimize latency.

This behavior applies to distributed [native workloads](/saas/workloads-in-nvidia-run-ai/workload-types/native-workloads) and [supported workload types](/saas/workloads-in-nvidia-run-ai/workload-types/supported-workload-types). To override the default behavior and define your own topology-related annotations when submitting a distributed workload, see [Fine-tuning topology-aware scheduling per workload](#fine-tuning-topology-aware-scheduling-per-workload).

#### LeaderWorkerSet (LWS) Behavior

LeaderWorkerSet (LWS) is a supported distributed workload type, receiving the same automatic topology-aware scheduling behavior described above. For LWS workloads, topology-aware scheduling is applied **per replica**:

* Each replica of a LeaderWorkerSet (as defined by `spec.replicas`) is gang scheduled.
* The same topology constraints are applied to each replica. This ensures that the leader and worker pods that belong to the same replica are placed according to the configured network topology (for example, within the same rack, block, or NVLink domain).

Example:

```yaml
apiVersion: leaderworkerset.x-k8s.io/v1
kind: LeaderWorkerSet
metadata:
  annotations:
    kai.scheduler/topology: "cluster-topology"
    kai.scheduler/topology-required-placement: "cloud.provider.com/topology-block"
    kai.scheduler/topology-preferred-placement: "cloud.provider.com/topology-rack"
  labels:
    runai/queue: test
  namespace: runai-test
  name: vllm
spec:
  replicas: 2
  leaderWorkerTemplate:
    size: 2
```

## Topology-Aware Scheduling for Dynamo over Grove

Dynamo disaggregated inference workloads require per-component topology-aware scheduling, since applying a single topology constraint across all pods can lead to suboptimal placement. NVIDIA Run:ai supports hierarchical, multi-component topology-aware scheduling for Dynamo over Grove workloads, enforcing constraints at both the workload (Deployment) and component (Service) levels.

{% hint style="info" %}
**Note**

Support for multiple topologies in the same cluster requires Dynamo 1.1.0 and Grove v0.1.0-alpha.9 or later. Starting from these versions, administrators can configure matching topologies in both NVIDIA Run:ai and Grove for each node pool, ensuring workloads are scheduled against the correct topology for the node pool where they run.

In earlier versions, only a single cluster-wide topology is supported - using multiple topologies in those versions may result in scheduling mismatches.
{% endhint %}

Grove manages topologies independently and exposes them as the mechanism for requesting topology constraints in the Dynamo workload spec.

* **Administrator setup** - For each topology defined in NVIDIA Run:ai, the administrator creates a matching topology in Grove and references it to its corresponding KAI topology. Once linked, NVIDIA Run:ai surfaces the Grove topology name and levels to users through the [Network Topologies](https://run-ai-docs.nvidia.com/api/organizations/network-topologies) API, without requiring them to understand the underlying Grove topology configuration.
* **Workload submission** - When submitting a Dynamo workload, the user selects the topology that matches the node pool on which the workload will run. If only a single topology exists in the cluster, it is always used. The Grove topology is surfaced through the [Network Topologies](https://run-ai-docs.nvidia.com/api/organizations/network-topologies) API, allowing users to identify the correct topology for their target node pool.
* **Scheduling** - The NVIDIA Run:ai Scheduler enforces the topology constraints at both the workload (Deployment) and component (Service) levels.

If there is no matching topology for the required node pool, a mismatch can occur between what the user requests, what the Scheduler places, and what is visible on the workload. See [Topology constraints visibility](#topology-constraints-visibility) for more details.

## Fine-tuning Topology-Aware Scheduling per Workload

You can customize how NVIDIA Run:ai applies topology labels for each distributed workload, allowing you to override the default scheduling behavior when needed.

* For native distributed workloads, define the topology annotations in the workload’s annotations field through the NVIDIA Run:ai UI, API, or CLI.
* For supported workload types submitted via YAML, specify the topology-related annotations directly in the YAML manifest. NVIDIA Run:ai respects the user-defined annotations and does not override them.

Use the following topology-aware scheduling annotations when submitting a distributed workload:

* `kai.scheduler/topology` - Specifies the name of the topology to use.\
  The value must match the name of a topology entity defined for the cluster.
* `kai.scheduler/topology-required` and/or `kai.scheduler/topology-preferred` - Define placement constraints using a topology label key (for example, `"rack"`).
  * Required - Enforces strict placement. All pods must be scheduled within the specified topology level.
  * Preferred - Expresses a soft preference. The Scheduler attempts to place pods within the specified topology level when possible.

You can use Required, Preferred, or both together for the same topology tree. When combining them, keep in mind that network topologies are hierarchical (tree-structured):

* A Preferred constraint defined at the same level as, or higher than, a Required constraint has no effect.
* A Preferred constraint is effective only when it is defined at a lower (more specific) topology level than the Required constraint.
* In this case, the Scheduler enforces the Required constraint while attempting to further group pods according to the Preferred constraint.

## Topology Constraints Visibility

The topology name, applied topology constraints, and actual topology placement, whether applied automatically by NVIDIA Run:ai or defined manually, are exposed through the [Workloads](https://run-ai-docs.nvidia.com/api/workloads/workloads) API for native and supported workload types. For native workloads, this information is also visible in the workload [Details](/saas/workloads-in-nvidia-run-ai/workloads#show-hide-details) view in the UI.

* **Topology name** - The name of the topology applied to the workload.
* **Topology levels** - Each configured topology level with its constraint type (Required or Preferred) and the level at which pods were actually placed. When the constraint is met, pods are placed according to the topology, ensuring optimal pod-to-pod communication and performance. When a Preferred constraint is not met, pods may be placed farther apart than preferred, which can impact communication latency and performance.
* **Actual placement** - The topology label level and value where the workload pods were actually placed (for example, `network.topology.nvidia.com/zone (us-central)`). Actual placement is reported even when no topology constraints were applied to the workload, so you can always see where pods landed in the network hierarchy. For multi-component workloads such as Dynamo over Grove, the [Workloads](https://run-ai-docs.nvidia.com/api/workloads/workloads) API also returns the actual placement per workload element, so you can verify where each element's pods were placed.

## Workloads Submitted via kubectl (Manual)

When submitting distributed workloads via `kubectl`, topology-aware scheduling is not applied automatically by NVIDIA Run:ai. Instead, you can configure the workload with annotations to ensure the Scheduler respects the desired topology.

* Add annotations to the workload manifest by specifying the topology name and constraint type.
* You can use either **Required** or **Preferred** constraints, or combine both for the same topology tree. When using both, note that network topologies are hierarchical (tree-structured). Applying a Preferred constraint at the same level as, or higher than, a Required constraint has no effect. A Preferred constraint is meaningful only when it is defined at a lower (more specific) topology level than the Required constraint. In this case, the topology-aware scheduling logic attempts to further group pods at that lower level, while still enforcing the mandatory Required constraint.

For example:

```yaml
apiVersion: batch/v1
kind: Job
metadata:
  name: topology-aware-job
  annotations:
    kai.scheduler/topology-preferred-placement: "rack"
    kai.scheduler/topology-required-placement: "zone"
    kai.scheduler/topology: "network"
```

## Pod Affinity vs. Topology-Aware Scheduling

The following example demonstrates the difference between pod affinity and topology-aware scheduling when placing distributed workloads across two GB200 racks:

<figure><img src="/files/TKQDKBpgS7KXGBfCfHty" alt=""><figcaption></figcaption></figure>

In the example, two workloads (Workload 1 requiring 12 nodes and Workload 2 requiring 15 nodes) are already running. A third workload that requires 6 nodes is submitted:

* With pod affinity, the Scheduler places pods one by one, only checking for “closeness” to existing pods without awareness of the entire workload compared to the available nodes. As a result, the workload is split across Rack A and Rack B, introducing unnecessary cross-rack communication overhead.
* In contrast, topology-aware scheduling evaluates the full node requirement in advance and uses knowledge of the hierarchy (rack, block, NVLink domains) to allocate resources. This ensures all 6 nodes are placed together in Rack A, minimizing latency and maximizing bandwidth efficiency. By avoiding fragmentation and cross-rack placement, topology-aware scheduling improves workload performance and overall cluster utilization compared to pod affinity.

## Using API

To view the available actions, go to the [Network Topologies](https://run-ai-docs.nvidia.com/api/organizations/network-topologies) API reference.

## Known Limitations

* If a topology is detached from a node pool, workloads that are already running will continue using it, as long as the topology still exists.
* If a topology is completely deleted while workloads are still using it:
  * Running workloads will continue unaffected.
  * Suspended workloads that are later resumed, or workloads not yet bound to a node, will become unschedulable and remain in Pending.
* Submitting a workload to multiple node pools that each have different topologies is not supported. Workloads submitted through NVIDIA Run:ai will fail, while external workloads may either remain pending or run if the topology matches at least one node pool.


# Configuring NUMA-Aware Scheduling

NUMA (Non-Uniform Memory Access) architecture divides the CPUs, memory, and GPUs within a physical server into multiple domains called NUMA nodes. Resources that span different NUMA nodes communicate over a higher-latency, lower-throughput interconnect, so workloads whose GPU and CPU allocations are split across NUMA boundaries can experience degraded performance.

NVIDIA Run:ai supports NUMA-aware scheduling at the node pool level. When enabled, the NVIDIA Run:ai Scheduler places workloads so that GPU, CPU, and CPU memory are allocated within the same NUMA node where possible, minimizing cross-NUMA fragmentation and improving performance. When disabled, the Scheduler places workloads using aggregate node capacity only, without considering NUMA topology.

NUMA-aware scheduling is especially important on high-density GPU servers such as DGX H100 or DGX B300, and in clusters where the Kubelet Topology Manager is configured to `restricted` or `single-numa-node` mode. These are more restrictive policies, and workloads may fail if NUMA-aware scheduling is not enabled on the corresponding node pool.

## Prerequisites

The following must be in place before enabling NUMA-aware scheduling on a node pool. These are configured outside of NVIDIA Run:ai.

{% hint style="info" %}
**Note**

Before setting up the prerequisites below, verify whether they are already configured in your cluster.
{% endhint %}

### Node Feature Discovery (NFD) Topology Updater

NFD must be deployed with the topology updater enabled. The topology updater publishes per-node NUMA topology data as Kubernetes NodeResourceTopology custom resources, including available and allocated GPUs, CPUs, and memory per NUMA node, as well as the Kubelet Topology Manager policy and scope configured on each node. The NVIDIA Run:ai Scheduler reads these resources to make NUMA-aware placement decisions. See the [NFD Topology Updater documentation](https://kubernetes-sigs.github.io/node-feature-discovery/stable/usage/nfd-topology-updater.html) for deployment instructions.

### Kubelet Topology Manager Policy

The Kubelet Topology Manager policy controls how NUMA alignment is enforced for workload resource allocations. NVIDIA Run:ai reads this policy from the NFD topology data and takes it into account when scheduling. If the policy is set to `none` on the pool's nodes, enabling NUMA-aware scheduling has no practical effect. For a description of the available policies (`none`, `best-effort`, `restricted`, `single-numa-node`), see the [Kubernetes Topology Manager documentation](https://kubernetes.io/docs/tasks/administer-cluster/topology-manager/#topology-manager-policies).

The `restricted` and `single-numa-node` policies are the more restrictive modes. If the cluster admin has configured nodes in a node pool to one of these policies but NUMA-aware scheduling is not enabled on that node pool in NVIDIA Run:ai, workload submissions may fail intermittently. Without NUMA-aware scheduling enabled, the NVIDIA Run:ai Scheduler does not predict NUMA placement, which can result in pods being assigned to nodes where the Kubelet rejects them.

### Kubelet Topology Manager Scope

The `--topology-manager-scope` Kubelet setting controls whether NUMA alignment is computed per container (the default) or for the pod as a whole. NVIDIA Run:ai reads this setting from the NFD topology data alongside the topology manager policy and takes it into account when scheduling. Both `container` and `pod` scopes are supported. See [Topology Manager scopes](https://kubernetes.io/docs/tasks/administer-cluster/topology-manager/#topology-manager-scopes) for details.

### Additional Kubelet Requirements

For NUMA alignment to take effect, the following Kubelet settings must also be configured on the pool's nodes:

* [`--cpu-manager-policy=static`](https://kubernetes.io/docs/tasks/administer-cluster/cpu-management-policies/): required for the Topology Manager to align CPU allocations to NUMA nodes
* [`--memory-manager-policy=Static`](https://kubernetes.io/docs/tasks/administer-cluster/memory-manager/): required for the Topology Manager to align memory allocations to NUMA nodes

The Topology Manager only coordinates hints from the CPU manager and memory manager; without these policies set to static, NUMA alignment for CPU and memory has no effect.

## Enabling NUMA-Aware Scheduling

NUMA-aware scheduling is configured per node pool and is disabled by default. To enable it, toggle the **NUMA-aware scheduling** option when creating or editing a node pool. See [Node pools](https://run-ai-docs.nvidia.com/api/organizations/nodepools) for details.

{% hint style="info" %}
**Note**

The NUMA-aware scheduling toggle is only available for clusters running version 2.26 or later.
{% endhint %}

## How It Works

When NUMA-aware scheduling is enabled on a node pool, the Scheduler reads each node's NUMA topology from NodeResourceTopology resources published by the NFD topology updater. This includes the Topology Manager policy and scope configured on each node, and the available and allocated GPUs, CPUs, and memory per NUMA node.

When placing a workload, the Scheduler predicts whether a node can satisfy the workload's GPU, CPU, and memory request within NUMA boundaries before the pod reaches the Kubelet. The goal is to determine placement compatibility early, minimize cross-NUMA fragmentation, and improve performance. Whether resources may span NUMA nodes depends on the Topology Manager policy configured on the node:

* **`best-effort`** - The Scheduler prefers NUMA-aligned placement but always admits the pod, even if resources must span NUMA nodes.
* **`restricted`** - Spanning across NUMA nodes is only allowed when every requested resource type (GPUs, CPUs, and memory) individually requires more than a single NUMA node. If one resource spans two nodes but another fits on one, the requests do not align and the pod is rejected. For example, on a server with 2 NUMA nodes and 4 GPUs each, a workload requesting 8 GPUs must span both nodes. If that workload also requests only 1 CPU and 100 GB of memory (both of which fit on a single NUMA node), the pod is rejected. The pod is admitted only when the CPU and memory requests are large enough to require both NUMA nodes as well, so all resources align to the same span.
* **`single-numa-node`** - All resources must fit within a single NUMA node. The workload is rejected if no single NUMA node can satisfy the full request.

If no node satisfies the workload's resource request within the policy constraints, the workload remains pending with a NUMA-specific event explaining the reason.

Enabling or disabling NUMA-aware scheduling does not affect workloads already running in the pool unless they are rescheduled.

## Best Practices

* Configure all nodes in the same node pool to the same Kubelet Topology Manager policy. Mixed policies within a pool can produce inconsistent scheduling results.
* For the lowest-risk approach to NUMA-aware scheduling, set all nodes to `best-effort`. The Kubelet attempts to align resources to a single NUMA node but always admits the pod, so workloads are never rejected due to NUMA constraints. This may produce less optimal performance than `restricted` or `single-numa-node`, but avoids placement failures.
* You can also configure all nodes to `best-effort` and keep NUMA-aware scheduling disabled in NVIDIA Run:ai. In this mode, the Kubelet still attempts to align resources to a single NUMA node on its own, but the NVIDIA Run:ai Scheduler does not predict NUMA placement. Workloads are never rejected due to NUMA constraints, though placement may be less optimal than with NUMA-aware scheduling enabled.

## Using API

To view the available actions, go to the [Node pools](https://run-ai-docs.nvidia.com/api/organizations/node-pools) API reference.


# Enabling Spectrum-X Networking for NVIDIA Run:ai Workloads

NVIDIA Spectrum-X is an AI-optimized Ethernet networking platform designed to deliver high throughput and predictable performance for large-scale, multi-node GPU workloads.

Leveraging Spectrum-X with NVIDIA Run:ai extends these benefits into day-to-day AI operations. NVIDIA Run:ai streamlines workload submission and management, enabling workloads to be scheduled in a way that takes advantage of Spectrum-X–enabled infrastructure and improves scale-out efficiency. At the same time, administrators can define policies and apply best practices that direct eligible workloads to the appropriate network configuration, helping maintain consistent operations aligned with network capabilities designed for AI at scale. See [NVIDIA Spectrum-X Ethernet Networking Platform](https://www.nvidia.com/en-eu/networking/spectrumx/) for more details.

## Prerequisites

Before using Spectrum-X with NVIDIA Run:ai, ensure the following components are installed and configured:

* NVIDIA Network Operator version 26.1.0. See the [NVIDIA Network Operator](/saas/getting-started/installation/install-using-helm/system-requirements#nvidia-network-operator) section for installation instructions.
* NVIDIA Spectrum-X Operator version 2.1, installed via the Network Operator.

{% hint style="info" %}
**Note**

If you are using an earlier version of the NVIDIA Spectrum-X Operator, contact [NVIDIA support](https://www.nvidia.com/en-eu/support/enterprise/#contact-us).
{% endhint %}

## Submitting a Workload with Spectrum-X

To leverage Spectrum-X networking with NVIDIA Run:ai, the workload must be configured to use the Spectrum-X network and request the required networking capabilities and resources.

These settings can be applied when submitting [NVIDIA Run:ai native workloads](/saas/workloads-in-nvidia-run-ai/workload-types/native-workloads) or workloads submitted [via YAML](/saas/workloads-in-nvidia-run-ai/submit-via-yaml).

### Required Configuration

* **Add the required Linux capabilities** - Configure the workload container to include the `IPC_LOCK` Linux capability. This provides the necessary networking and system permissions required for Spectrum-X, without granting full root privileges:
* **Request extended resources** - Request the appropriate NVIDIA extended resource and quantity according to your Spectrum-X configuration. The resource name is taken from the `spec.resourceName` field of the `OVSNetwork`. Use this value as the resource key when requesting resources for the workload (for example, `nvidia.com/sriov_resource`).
* **Attach the Spectrum-X network(s)** - Add the required Kubernetes annotation, `k8s.v1.cni.cncf.io/networks`, to attach one or more Spectrum-X networks to the workload. Each network name in the annotation must match the `metadata.name` of an existing `OVSNetwork` resource.

### YAML Example

The following example shows only the relevant sections of a workload manifest. It illustrates where to define the annotation, capabilities, and extended resource.

<pre class="language-yaml"><code class="lang-yaml">spec:
  template:
    metadata:
      annotations:
        k8s.v1.cni.cncf.io/networks: &#x3C;network-name>
    spec:
      containers:
        - name: example-container
          securityContext:
            capabilities:
              add:
                - IPC_LOCK
<strong>          resources:
</strong><strong>            requests:
</strong>              nvidia.com/sriov_resource: 1
            limits:
              nvidia.com/sriov_resource: 1
</code></pre>

## Enforcing Spectrum-X Requirements Using Workload Policies

Administrators can use NVIDIA Run:ai [workload policies](/saas/platform-management/policies/native-workload-policies) to ensure that workloads intended to run on Spectrum-X infrastructure are configured with the required capabilities, resources, and network attachments. For full policy structure and supported fields, see [Policy YAML Reference](/saas/platform-management/policies/native-workload-policies/policy-yaml-reference).

{% hint style="info" %}
**Note**

Workload policies are supported for [NVIDIA Run:ai native workloads](/saas/workloads-in-nvidia-run-ai/workload-types/native-workloads) only.
{% endhint %}

**Example workload policy:**

<pre class="language-yaml"><code class="lang-yaml">defaults:
  annotations:
    instances:
      - name: k8s.v1.cni.cncf.io/networks
        value: &#x3C;network-name>
  security:
    capabilities:
<strong>      - IPC_LOCK
</strong>  compute:
    extendedResources:
      instances:
        - resource: nvidia.com/sriov_resource
          quantity: "1"
rules:
  security:
    capabilities:
      canEdit: false
</code></pre>


# Node Pools

Node pools assist in managing heterogeneous resources effectively. A node pool is a NVIDIA Run:ai construct representing a set of nodes grouped into a bucket of resources using a predefined node label (e.g. NVIDIA GPU type) or an administrator-defined node label (any key/value pair).

Typically, the grouped nodes share a common feature or property, such as GPU type or other HW capability (such as Infiniband connectivity), or represent a proximity group (i.e. nodes interconnected via a local ultra-fast switch). Researchers and ML Engineers would typically use node pools to run specific workloads on specific resource types.

In the NVIDIA Run:ai Platform a user with the System administrator role can create, view, edit, and delete node pools. Creating a new node pool creates a new instance of the NVIDIA Run:ai [Scheduler](/saas/platform-management/runai-scheduler/scheduling/how-the-scheduler-works). Workloads submitted to a node pool are scheduled using the node pool’s designated scheduler instance.

Once created, the new node pool is automatically assigned to all [projects](/saas/platform-management/aiinitiatives/organization/projects) and [departments](/saas/platform-management/aiinitiatives/organization/departments) with a quota of zero GPU resources, unlimited CPU resources, and over quota enabled (medium weight if over quota weight is enabled). This allows any project and department to use any node pool when over quota is enabled, even if the administrator has not assigned a quota for a specific node pool within that project or department.

When submitting a new [workload](/saas/workloads-in-nvidia-run-ai/workloads), users can add a prioritized list of node pools. The node pool selector picks one node pool at a time (according to the prioritized list) and the designated node pool scheduler instance handles the submission request and tries to match the requested resources within that node pool. If the scheduler cannot find resources to satisfy the submitted workload, the node pool selector moves the request to the next node pool in the prioritized list, if no node pool satisfies the request, the node pool selector starts from the first node pool again until one of the node pools satisfies the request.

## Node Pools Table

The Node pools table can be found under **Resources** in the NVIDIA Run:ai platform.

The Node pools table lists all the node pools defined in the NVIDIA Run:ai platform and allows you to manage them.

{% hint style="info" %}
**Note**

By default, the NVIDIA Run:ai platform includes a single node pool named ‘default’. When no other node pool is defined, all existing and new nodes are associated with the ‘default’ node pool. When deleting a node pool, if no other node pool matches any of the nodes’ labels, the node will be included in the default node pool.
{% endhint %}

<figure><img src="/files/vzkZFiKMen3MnsN5HiPO" alt=""><figcaption></figcaption></figure>

The Node pools table consists of the following columns:

| Column                          | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| ------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Node pool                       | The node pool name, set by the administrator during its creation (the node pool name cannot be changed after its creation).                                                                                                                                                                                                                                                                                                                                                                                    |
| Status                          | <p>Node pool status: Creating, Updating, Empty, Ready, Unschedulable, Deleting, Deleted:</p><ul><li>Empty - No nodes are currently included in that node pool.</li><li>Ready - The Scheduler can use this node pool to schedule workloads.</li><li>Unschedulable - The Scheduler cannot use this node pool to schedule workloads.</li></ul><p>The status also includes a status message that provides more details.</p>                                                                                        |
| <p>Label key<br>Label value</p> | The node pool controller will use this node-label key-value pair to match nodes into this node pool.                                                                                                                                                                                                                                                                                                                                                                                                           |
| Node(s)                         | List of nodes included in this node pool. Click the field to view details (the details are in the [Nodes](/saas/platform-management/aiinitiatives/resources/nodes) article).                                                                                                                                                                                                                                                                                                                                   |
| Network topology                | The network topology associated with this node pool                                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| MNNVL                           | Indicates whether the discovery method of Multi-Node NVL nodes is done automatically or manually                                                                                                                                                                                                                                                                                                                                                                                                               |
| MNNVL label key                 | The label key that is used to automatically detect if a node is part of an MNNVL domain. The default MNNVL domain label is `nvidia.com/gpu.clique.`                                                                                                                                                                                                                                                                                                                                                            |
| MNNVL nodes                     | Indicates whether MNNVL nodes are detected - automatically or manually.                                                                                                                                                                                                                                                                                                                                                                                                                                        |
| Minimum guaranteed runtime      | The minimum guaranteed runtime that a workload will run before it can be preempted by a higher priority workload.                                                                                                                                                                                                                                                                                                                                                                                              |
| GPU placement strategy          | Sets the Scheduler node-level strategy for the assignment of pods requesting **both GPU and CPU resources** to nodes, which can be either Bin-pack or Spread. By default, Bin-Pack is used, but can be changed to Spread by editing the node pool. When set to Bin-pack the scheduler will try to fill nodes as much as possible before using empty or sparse nodes, when set to spread the scheduler will try to keep nodes as sparse as possible by spreading workloads across as many nodes as it succeeds. |
| CPU placement strategy          | Sets the Scheduler node-level strategy for the assignment of pods requesting **only CPU** **resources** to nodes, which can be either Bin-pack or Spread. By default, Bin-Pack is used, but can be changed to Spread by editing the node pool. When set to Bin-pack the scheduler will try to fill nodes as much as possible before using empty or sparse nodes, when set to spread the scheduler will try to keep nodes as sparse as possible by spreading workloads across as many nodes as it succeeds.     |
| Total GPU devices               | The total number of GPU devices installed into nodes included in this node pool. For example, a node pool that includes 12 nodes each with 8 GPU devices would show a total number of 96 GPU devices.                                                                                                                                                                                                                                                                                                          |
| Total GPU memory                | The total amount of GPU memory included in this node pool. The total amount of GPU memory installed in nodes included in this node pool. For example, a node pool that includes 12 nodes, each with 8 GPU devices, and each device with 80 GB of memory would show a total memory amount of 7.68 TB.                                                                                                                                                                                                           |
| Allocated GPUs                  | The total allocation of GPU devices in units of GPUs (decimal number). For example, if 3 GPUs are 50% allocated, the field prints out the value 1.50. This value represents the portion of GPU memory consumed by all running pods using this node pool. ‘Allocated GPUs’ can be larger than ‘Projects’ GPU quota’ if over quota is used by workloads, but not larger than GPU devices.                                                                                                                        |
| GPU resource optimization ratio | Shows the Node Level Scheduler mode                                                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| Total CPU (Cores)               | The number of CPU cores installed on nodes included in this node                                                                                                                                                                                                                                                                                                                                                                                                                                               |
| Total CPU memory                | The total amount of CPU memory installed on nodes using this node pool                                                                                                                                                                                                                                                                                                                                                                                                                                         |
| Allocated CPU (Cores)           | The total allocation of CPU compute in units of Cores (decimal number). This value represents the amount of CPU cores consumed by all running pods using this node pool. ‘Allocated CPUs’ can be larger than ‘Projects’ GPU quota’ if over quota is used by workloads, but not larger than CPUs (Cores).                                                                                                                                                                                                       |
| Allocated CPU memory            | The total allocation of CPU memory in units of TB/GB/MB (decimal number). This value represents the amount of CPU memory consumed by all running pods using this node pool. ‘Allocated CPUs’ can be larger than ‘Projects’ CPU memory quota’ if over quota is used by workloads, but not larger than CPU memory.                                                                                                                                                                                               |
| Last updated                    | The date and time when the node pool was last updated                                                                                                                                                                                                                                                                                                                                                                                                                                                          |
| Creation time                   | The date and time when the node pool was created                                                                                                                                                                                                                                                                                                                                                                                                                                                               |
| Workload(s)                     | List of workloads running on nodes included in this node pool, click the field to view details (described below in this article)                                                                                                                                                                                                                                                                                                                                                                               |

### Workloads Associated with the Node Pool

Click one of the values in the Workload(s) column, to view the list of workloads and their parameters.

{% hint style="info" %}
**Note**

This column is only viewable if your role in the NVIDIA Run:ai platform gives you read access to workloads, even if you are allowed to view workloads, you can only view the workloads within your allowed scope. This means, there might be more pods running on this node than appear in the list your are viewing.
{% endhint %}

| Column                        | Description                                                                                                                                                                                |
| ----------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Workload                      | The name of the workload. If the workloads’ type is one of the recognized types (for example: Pytorch, MPI, Jupyter, Ray, Spark, Kubeflow, and many more), an appropriate icon is printed. |
| Type                          | The NVIDIA Run:ai platform type of the workload - Workspace, Training, or Inference                                                                                                        |
| Status                        | The state of the workload. The Workloads state is described in the NVIDIA Run:ai [workloads](/saas/workloads-in-nvidia-run-ai/workloads) section                                           |
| Created by                    | The user or service account created this workload                                                                                                                                          |
| Running/requested pods        | The number of running pods out of the number of requested pods within this workload.                                                                                                       |
| Creation time                 | The workload’s creation date and time                                                                                                                                                      |
| Allocated GPU compute         | The total amount of GPU compute allocated by this workload. A workload with 3 Pods, each allocating 0.5 GPU, will show a value of 1.5 GPUs for the workload.                               |
| Allocated GPU memory          | The total amount of GPU memory allocated by this workload. A workload with 3 Pods, each allocating 20GB, will show a value of 60 GB for the workload.                                      |
| Allocated CPU compute (cores) | The total amount of CPU compute allocated by this workload. A workload with 3 Pods, each allocating 0.5 Core, will show a value of 1.5 Cores for the workload.                             |
| Allocated CPU memory          | The total amount of CPU memory allocated by this workload. A workload with 3 Pods, each allocating 5 GB of CPU memory, will show a value of 15 GB of CPU memory for the workload.          |

### Customizing the Table View

* Filter - Click ADD FILTER, select the column to filter by, and enter the filter values
* Search - Click SEARCH and type the value to search by
* Sort - Click each column header to sort by
* Column selection - Click COLUMNS and select the columns to display in the table
* Download table - Click MORE and then Click Download as CSV. Export to CSV is limited to 20,000 rows.
* Show/Hide details - Click to view additional information on the selected row

### Show/Hide Details

Select a row in the Node pools table and then click Show details in the upper-right corner of the action bar. The details window appears, presenting metrics graphs for the whole node pool:

* **Node GPU allocation** - This graph shows an overall sum of the Allocated, Unallocated, and Total number of GPUs for this node pool, over time. From observing this graph, you can learn about the occupancy of GPUs in this node pool, over time.
* **GPU Utilization Distribution** - This graph shows the distribution of GPU utilization in this node pool over time. Observing this graph, you can learn how many GPUs are utilized up to 25%, 25%-50%, 50%-75%, and 75%-100%. This information helps to understand how many available resources you have in this node pool, and how well those resources are utilized by comparing the allocation graph to the utilization graphs, over time.
* **GPU Utilization** - This graph shows the average GPU utilization in this node pool over time. Comparing this graph with the GPU Utilization Distribution helps to understand the actual distribution of GPU occupancy over time.
* **GPU Memory Utilization** - This graph shows the average GPU memory utilization in this node pool over time, for example an average of all nodes’ GPU memory utilization over time.
* **CPU Utilization** - This graph shows the average CPU utilization in this node pool over time, for example, an average of all nodes’ CPU utilization over time.
* **CPU Memory Utilization** - This graph shows the average CPU memory utilization in this node pool over time, for example an average of all nodes’ CPU memory utilization over time.

## Adding a New Node Pool

To create a new node pool:

1. Click **+NEW NODE POOL**
2. Enter a **name** for the node pool.\
   Node pools names must start with a letter and can only contain lowercase Latin letters, numbers or a hyphen ('-’)
3. Enter the **node pool label**:\
   The node pool controller will use this node-label key-value pair to match nodes into this node pool.
   * **Key** is the unique identifier of a node label.
     * The key must fit the following regular expression: `^(([A-Za-z0-9][-A-Za-z0-9_.]*)?[A-Za-z0-9])?/?([A-Za-z0-9][-A-Za-z0-9_.]*)?[A-Za-z0-9]$`
     * The administrator can put an automatically preset label such as the nvidia.com/gpu.product that labels the GPU type or any other key from a node label.
   * **Value** is the value of that label identifier (key). The same key may have different values, in this case, they are\
     considered as different labels.
     * Value must fit the following regular expression: `^(([A-Za-z0-9][-A-Za-z0-9_.]*)?[A-Za-z0-9])?$`
   * A node pool is defined by a single key-value pair. You must not use different labels that are set on the same node by different node pools, this situation may lead to unexpected results.
4. Define **scheduling** configurations:
   * Set the **minimum guaranteed runtime**. The minimum guaranteed runtime is the time that a workload will run before it can be preempted by a higher priority workload. If the workload runs for less than the time you set, it will not be preempted. You can set the value in days, hours, minutes, and seconds. Default is 0.
   * **Allow time-based fairshare** - When enabled, the Scheduler adjusts fairshare scores based on each project’s weight and historical resource usage. Historical usage data is not persisted by default. When disabled, the Scheduler uses the default fairshare, where scheduling fairness is based on the project’s current weight and resource usage. For more details, see [Time-based fairshare](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles#time-based-fairshare). To persist historical usage data across restarts, enable persistent data at the cluster level. For details, see [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config#time-based-fairshare).

     * **Historical usage weight** - Sets the weight of historical resource usage in the fairshare calculation. Default is 1.0:
       * **0.0** - Historical usage data is ignored; fairshare is based entirely on current weight.
       * **1.0** - Creates an equilibrium between historical usage and current weight. If historical usage equals the current weight, the fairshare matches exactly the weight portion.
       * **Between 0.0 and 1.0** - Historical usage has a reduced effect on fairshare, compensating for deficient usage or penalizing excess usage to a lesser degree.
       * **Greater than 1.0** - Historical usage has a greater effect on fairshare, making it the dominant factor in the fairshare calculation.
     * **Historical usage window** - Sets the duration of the sliding time window used to calculate historical resource usage. Default is 7 days. Only usage within this window is considered; anything outside the window is no longer counted.

     Administrators can customize additional time-based fairshare configuration parameters at the node pool level using the [Node pools](https://run-ai-docs.nvidia.com/api/organizations/nodepools) API.
   * **NUMA-aware scheduling** - When enabled, the Scheduler places workloads so that GPU, CPU, and CPU memory are allocated within the same NUMA node where possible, minimizing cross-NUMA fragmentation and improving performance. There are cluster prerequisites for enabling this feature. When disabled, the Scheduler places workloads using aggregate node capacity only, without considering NUMA topology. See [Configuring NUMA-Aware Scheduling](/saas/platform-management/aiinitiatives/resources/numa-aware-scheduling) for more details.
   * Configure the **GPU pods** placement strategy. GPU workloads are workloads that request both GPU and CPU resources.
     * Set the **node placement strategy** - Determines how workloads are distributed across nodes in the node pool:
       * **Bin-pack** - Place as many workloads as possible in each GPU and node to use fewer resources and maximize GPU and node vacancy.
       * **Spread** - Spread workloads across as many GPUs and nodes as possible to minimize the load and maximize the available resources per workload.
     * Set the **device placement strategy** - Determines how workloads are distributed across GPU devices within each node:
       * **Bin-pack** - Place as many workloads as possible on each GPU device to minimize the number of GPU devices used.
       * **Spread** - Spread workloads across GPU devices within each node to maximize the available resources per workload.
   * Configure the **CPU pods** placement strategy. CPU workloads are workloads that request purely CPU resources.
     * Set the **node placement strategy** - Determines how CPU workloads are distributed across nodes in the node pool:
       * **Bin-pack** - Place as many workloads as possible in each CPU and node to use fewer resources and maximize CPU and node vacancy.
       * **Spread** - Spread workloads across as many CPUs and nodes as possible to minimize the load and maximize the available resources per workload.
5. Set the **GPU network acceleration**:
   * Network topologies are defined at the cluster level and must be associated with one or more node pools. Administrators can create and manage network topologies from the [Clusters](/saas/infrastructure-setup/procedures/clusters#managing-network-topologies) page and attach them to node pools, or create a network topology directly while configuring a node pool.
     * Select the **network topology** that represents this node pool’s network. This optimizes placement and accelerates distributed workloads by keeping pods on nodes that are as close to each other as possible in the network.
     * To add a topology that represents the node pool's network:

       * Click **+ NEW NETWORK TOPOLOGY**
       * Enter a unique **name** for the topology. If the name already exists, you will be requested to enter a different name.
       * Click **+ LABEL** to add the node label keys that represent the network hierarchy
         * Order labels from farthest (first) to closest (last)
         * Ensure the labels match the corresponding keys on the nodes. For example: `cloud.provider.com/topology-block`, `cloud.provider.com/topology-rack`, `kubernetes.io/hostname`
         * Drag labels to adjust their order if needed
       * Click **SAVE NETWORK TOPOLOGY**

       See [Accelerating workloads with network topology-aware scheduling](/saas/platform-management/aiinitiatives/resources/topology-aware-scheduling) for more details.
   * Define whether MNNVL is present on this node pool. For more details, see [Using GB200 NVL72 and Multi-Node NVLink domains](/saas/platform-management/aiinitiatives/resources/using-gb200):
     * **Auto-detect** - Automatically detect whether the node pool contains any MNNVL nodes. MNNVL nodes that share the same ID are part of the same NVL rack.
     * **MNNVL is present** - Explicitly indicate that the node pool contains MNNVL nodes
     * **MNNVL is not present** - Explicitly indicate that the node pool does not contain MNNVL nodes
   * Set the node’s label used to discover GPU network acceleration (MNNVL) to `nvidia.com/gpu.clique`
6. Click **CREATE NODE POOL**

### Labeling Nodes for Node Pool Grouping

The administrator can use a preset node label, such as the `nvidia.com/gpu.product` that labels the GPU type, or configure any other node label (e.g. `faculty=physics`).

To assign a label to nodes you want to group into a node pool, set a node label on each node:

1. Obtain the list of nodes and their current labels by copying the following to your terminal:

   ```bash
   kubectl get nodes --show-labels
   ```
2. Annotate a specific node with a new label by copying the following to your terminal:

   ```bash
   kubectl label node <node-name> <key>=<value>
   ```

### Labeling Nodes via Cloud Providers

Most cloud providers allow you to configure node labels at the node pool level. You can apply labels when creating a cluster, creating a node pool, or by editing an existing node pool.

Ensure that each node is labeled using the Kubernetes label format. This label ensures that workloads are scheduled correctly based on node pool definitions:

```bash
run.ai/type=<TYPE_VALUE>
```

Refer to the provider-specific documentation below for guidance on how to configure node pool labels:

* [Google Kubernetes Engine (GKE)](https://cloud.google.com/kubernetes-engine/docs/concepts/kubernetes-engine-overview)
* [Azure Kubernetes Service (AKS)](https://learn.microsoft.com/en-us/azure/aks/)
* [Amazon Elastic Kubernetes (EKS)](https://docs.aws.amazon.com/eks/)

## Editing a Node Pool

1. Select the node pool you want to edit
2. Click **EDIT**
3. Update the node pool and click **SAVE**

## Deleting a Node Pool

1. Select the node pool you want to delete
2. Click **DELETE**
3. On the dialog, click **DELETE** to confirm the deletion

{% hint style="info" %}
**Note**

The default node pool cannot be deleted. When deleting a node pool, if no other node pool matches any of the nodes’ labels, the node will be included in the default node pool.
{% endhint %}

## Using API

To view the available actions, go to the [Node pools](https://run-ai-docs.nvidia.com/api/organizations/node-pools) API reference.


# Scheduling and Resource Optimization


# Scheduling


# The NVIDIA Run:ai Scheduler: Concepts and Principles

When a user [submits a workload](/saas/workloads-in-nvidia-run-ai/workloads), it is directed to the designated Kubernetes cluster and managed by the NVIDIA Run:ai Scheduler. The Scheduler’s primary role is to optimize placement, allocating workloads to the most suitable nodes based on specific resource requirements, workload characteristics, and organizational policies. This ensures high utilization while strictly adhering to NVIDIA Run:ai’s fairness and quota management logic.

The Scheduler is platform-agnostic, supporting native Kubernetes workloads, NVIDIA Run:ai workloads, and various third-party frameworks. For a detailed breakdown of compatible types, see [Introduction to workloads](/saas/workloads-in-nvidia-run-ai/introduction-to-workloads#types-of-workloads-in-runai).

To better understand the logic driving these allocation decisions, get to know the key concepts, resource management and scheduling principles of the Scheduler.

## Workloads, Pod Groups and Sub Groups

### Workloads and Pod Groups

Workloads can range from a single pod running on individual nodes to distributed workloads using multiple pods, each running on a node (or part of a node). A large-scale training workload could use dozens of nodes; similarly, an inference workload could use many pods (replicas) across multiple nodes.

Every newly created pod is assigned to a pod group that represents one or multiple pods belonging to a single workload. For example, a distributed PyTorch training workload with 32 workers is grouped into a single pod group.

The NVIDIA Run:ai Scheduler is a workload scheduler. This means the Scheduler always uses the pod group associated with each pod and handles the pod as part of a larger group of pods, while taking into consideration the common characteristics of the pod group, such as [gang scheduling](#gang-scheduling), minimum replicas or workers, workload priority class, workload preemptibility policy, topology information, and more. These characteristics are applied consistently across the entire pod group.

### Multi-Level Sub Groups

Some workloads use a workload structure composed of multiple functions (functionality disaggregation). In this structure, each function can have one or more replicas, and each replica is made up of one or more pods, usually structured as leader and worker sets (distributed replicas). One example of such a workload structure is the NVIDIA Dynamo inference framework, which uses Grove as a Kubernetes API (CRD) and controller for the underlying pods. A single Dynamo workload is composed of three functions: an incoming router (gateway), a prefill, and a decoder. Each of these functions may have multiple replicas, where the prefill and decoder replicas are structured as leader-worker sets.

This workload structure requires framing sets of leader-worker replicas as gang-scheduled sub-groups, which may be further grouped into sets of scale replicas, forming a second level of sub-groups. Therefore, the NVIDIA Run:ai workload group structure, handled by the Scheduler, is a hierarchy consisting of a top-level pod group with subordinate sub-groups. In this hierarchical group structure, a sub-group always points to a higher-level sub-group and ultimately to the top-level pod group. By default, each pod group has at least one sub-group under which all workers or replicas are federated.

## Scheduling Queue

A scheduling queue (or simply a queue) represents a scheduler primitive that manages the scheduling of workloads based on different parameters.

A queue is created for each [project/node pool pair](/saas/platform-management/aiinitiatives/adapting-ai-initiatives#mapping-your-organization) and [department/node pool pair](/saas/platform-management/aiinitiatives/adapting-ai-initiatives#mapping-your-organization). The NVIDIA Run:ai Scheduler supports hierarchical queueing, project queues are bound to department queues, per node pool. This allows an organization to manage per node pool quota, over quota and more parameters for projects and their associated departments.

## Resource Management

### Quota

Each project and department includes a set of deserved resource quotas, per node pool and resource type. For example, project “LLM-Train/Node Pool NV-H100” quota parameters specify the number of GPUs, CPUs(cores), and the amount of CPU memory that this project deserves to get when using this node pool. [Non-preemptible workloads](#priority-and-preemption) can only be scheduled if their requested resources are within the deserved resource quotas of their respective project/node pool and department/node pool.

#### Deserved Quota and Workload Scheduling

Deserved quota means that a project is entitled to use up to a maximum number of resources defined by its quota (for example, 10 GPUs). The Scheduler does not allow a project to exceed this limit within a given node pool for non-preemptible workloads. A workload scheduled within its deserved quota is assured to run once scheduled, particularly in the case of non-preemptible workloads.

Preemptible workloads, on the other hand, may be preempted by higher priority workloads using the same project / node pool resource quota.

#### When Deserved Quota Cannot Be Fully Satisfied

The Scheduler may not always be able to provide the full deserved quota for several reasons. For example, this may occur due to cluster fragmentation, additional workload constraints such as specific topology requirements, node failures, or situations where administrators assign projects or departments more quota than the physically available resources (that is, [over-subscription](#over-subscription)).

### Over Quota

Projects and departments can have a share in the unused resources of any node pool, beyond their quota of deserved resources. These resources are referred to as over quota resources. The administrator configures the [over quota parameters](/saas/platform-management/aiinitiatives/organization/projects) per node pool for each project and department. Over quota resources can only be used by preemptible workloads.

### Over Quota Weight

Projects can receive a share of the cluster/node pool unused resources when the over quota weight setting is enabled. The part each Project receives depends on its over quota **weight** value, and the total weights of all other projects over quota weights. The same applies to Departments. The administrator configures the [over quota weight parameters](/saas/platform-management/aiinitiatives/organization/projects) per node pool for each project and department.

### Max GPU Device Allocation

Administrators can limit the maximum number of GPUs that a project can allocate per node pool (that is, deserved quota + over quota <= max GPU device allocation). The same constraint applies at the department level.

A department’s max GPU device allocation effectively limits the total number of GPUs that can be allocated by all projects under that department. This applies even if the combined max GPU device allocations of the individual projects exceed the department’s max GPU device allocation.

### Project and Department Rank

{% hint style="info" %}
**Note**

NVIDIA Run:ai terminology for project priority and department priority has changed to project rank and department rank, across the UI and API.
{% endhint %}

Administrators can assign each project/node pool pair a rank. The rank determines the project’s relative precedence for resources compared to its siblings within the same department and node pool pair. Projects with a higher rank are allocated resources first, both within their quota and over quota, and may also reclaim (preempt) resources from lower-ranked projects within the same department/node pool pair.

The same principle applies to departments: a department’s rank is relative to its sibling departments using the same node pool. Because a project’s rank is local to its associated department/node pool pair, a high-ranked project in a low-ranked department can have lower overall resource precedence than a lower-ranked project in a higher-ranked department.

### Multi-Level Quota System

Each project has a set of deserved resource quotas (GPUs, CPUs, and CPU memory) defined per node pool. Projects can exceed their deserved quota and receive a share of unused resources in the node pool beyond that quota. The same model applies at the department level.

The Scheduler first balances over-quota resources across departments and then, within each department, distributes those resources among its projects. A department’s deserved quota and over-quota limits constrain the total amount of resources that can be allocated by all projects within that department.

If a project still has available deserved quota but the department’s deserved quota is exhausted, the Scheduler does not allocate additional deserved resources to that project. The same rule applies to over-quota resources: over-quota capacity is first allocated to the department and only then divided among its projects.

### Resource Limit Parameter

Each project and department has a Limit parameter per resource type. This parameter defines the maximum amount of that resource that a project or department can allocate, effectively placing an upper bound on the combined deserved quota + over-quota resources.

The NVIDIA Run:ai API exposes the Limit parameter for all resource types at both the project and department levels. In the UI, this parameter is exposed only for GPUs and is referred to as Max GPU Device Allocation, as GPUs are typically the most critical resource in AI clusters.

By default, all projects and departments have a Limit value of Unlimited, meaning they can theoretically allocate resources up to the cluster’s total capacity. In practice, however, resource allocation is constrained by additional factors, such as quotas, over-quota distribution, and competition between departments and projects for unused resources.

### Over-Subscription

Over-subscription is a scenario where the sum of all deserved resource quotas surpasses the physical resources of the cluster or node pool. In this case, there may be scenarios in which the Scheduler cannot find matching nodes to all workload requests, even if those requests were within the deserved resource quota of their associated projects.

## Scheduling Principles

### Fairness (Fair Resource Distribution)

[Fairness](/saas/platform-management/runai-scheduler/scheduling/how-the-scheduler-works) is a major principle within the NVIDIA Run:ai scheduling system. It means that the NVIDIA Run:ai Scheduler always respects certain resource splitting rules (fairness) between projects and between departments.

### Fairshare and Fairshare Balancing

To implement fairness, the NVIDIA Run:ai Scheduler calculates a numerical value called fairshare for each project or department, per node pool. This value represents the sum of the project’s or department’s deserved resources (quota) plus its share of unused resources in that node pool (over-quota resources).

The Scheduler then takes the minimum of:

* the calculated fairshare, and
* the total resources requested by the project’s or department’s workloads in that node pool

and uses this value as the effective fairshare for the project or department.

The Scheduler aims to provide each project or department with the resources they deserve per node pool using two main parameters: deserved quota and deserved fairshare (that is, quota plus over-quota resources). If one project’s node pool queue is below its fairshare while another project’s node pool queue exceeds its fairshare, the Scheduler shifts resources between queues to rebalance fairness. This process may result in the preemption of some over-quota, preemptible workloads.

### Time-Based Fairshare

Administrators can enable, per node pool, a fairshare mode called [time-based fairshare](/saas/platform-management/runai-scheduler/scheduling/how-the-scheduler-works#time-based-fairshare).

When this mode is enabled, the Scheduler continuously collects historical resource usage data for each project and department, per node pool. It evaluates each project’s and department’s GPU-hour consumption relative to its configured weight and uses this information to balance resource distribution more effectively across projects and departments over time.

This capability applies only to over-quota resources. Quota resources always have the highest allocation precedence. Only after all quota allocations are satisfied does the Scheduler distribute excess (over-quota) resources among departments and projects, according to their [ranks and weights](/saas/platform-management/aiinitiatives/organization/projects).

This time-based fairshare calculation ensures that each project and department receives its fair share of resource processing time (GPUs, CPUs, and CPU memory) over extended periods. As a result, over-quota resource distribution becomes more stable and fair, reducing short-term imbalances.

### Priority and Preemption

NVIDIA Run:ai supports scheduling workloads using different priority and preemption policies:

* Workload's priority and preemption are two distinct parameters set in a workload. Workload's priority sets the scheduling precedence within a project, while workload's preemption sets whether the workload is preemptible or non-preemptible.
* High priority workloads (pods) can preempt [lower priority workloads](#preemption-of-lower-priority-workloads-within-a-project) (pods) within the same scheduling queue (project), according to their preemption policy (i.e. if preemptible).
* If no explicit preemption policy parameter is set, the NVIDIA Run:ai Scheduler implicitly assumes any PriorityClass >= 100 is non-preemptible and any PriorityClass < 100 is preemptible.
* Cross project and cross department workload preemptions are referred to as [resource reclaim](#reclaim-of-resources-between-projects-and-departments) and are based on [fairness](#fairness-fair-resource-distribution) between queues rather than the priority of the workloads.

To make it easier for users to submit workloads, NVIDIA Run:ai preconfigured several Kubernetes PriorityClass objects. The NVIDIA Run:ai preset PriorityClass objects have their ‘preemptionPolicy’ always set to ‘PreemptLowerPriority’, regardless of their actual NVIDIA Run:ai preemption policy within the NVIDIA Run:ai platform.

A non-preemptible workload is only scheduled if in-quota and cannot be preempted after being scheduled, not even by a higher priority workload. To see the default priority and preemption policy of a workload and for details on how to change the priority and preemption, see [Workload priority and preemption](/saas/platform-management/runai-scheduler/scheduling/workload-priority-control).

### Preemption of Lower Priority Workloads Within a Project

Workload priority is always respected within a project. This means higher priority workloads are scheduled before lower priority workloads. It also means that higher priority workloads may preempt lower priority workloads within the same project if the lower priority workloads are preemptible.

### Reclaim of Resources Between Projects and Departments

[Reclaim](/saas/platform-management/runai-scheduler/scheduling/how-the-scheduler-works#reclaim-preemption-between-projects-and-departments) is an inter-project (and inter-department) scheduling action that takes back resources from one project (or department) that has used them as over quota, back to a project (or department) that deserves those resources as part of its deserved quota, or to balance fairness between projects, each to its fairshare (i.e. sharing fairly the portion of the unused resources).

### Gang Scheduling

Gang scheduling describes a scheduling principle in which a workload composed of multiple pods is either fully scheduled (all pods are scheduled and running) or fully pending (none of the pods are running). Gang scheduling applies to a single pod group.

#### Multi-Level Gang Scheduling

The NVIDIA Run:ai Scheduler supports scheduling workloads that use multi-level pod-group structures, as described in [Workloads, pod groups, and sub groups](#workloads-pod-groups-and-sub-groups). To support these workloads, the Scheduler implements multi-level pod grouping.

The top level of the hierarchy is always a pod group, regardless of whether the workload is flat or hierarchical. Any grouping level below the top level is a sub-group. A workload may include multiple levels of sub-groups, as required, and each sub-group points to its immediate parent in the hierarchy.

A sub-group has parameters similar to those of a pod group. For example, it includes a min-members parameter, which indicates the minimum number of pods that must be scheduled for the sub-group to be considered ganged.

For a hierarchical workload to transition to the Running state, all sub-groups must be successfully ganged, including the top-level pod group.

### Placement Strategy - Bin-Pack and Spread

The administrator can set a [placement strategy](/saas/platform-management/aiinitiatives/resources/node-pools#adding-a-new-node-pool), bin-pack or spread, of the Scheduler per node pool. For GPU based workloads, workloads can request both GPU and CPU resources. For CPU-only based workloads, workloads can request CPU resources only.

* **GPU workloads:**
  * **Node placement strategy** - Determines how workloads are distributed across nodes in the node pool:
    * **Bin-pack** - The Scheduler places as many workloads as possible in each GPU and node to use fewer resources and maximize GPU and node vacancy.
    * **Spread** - The Scheduler spreads workloads across as many GPUs and nodes as possible to minimize the load and maximize the available resources per workload.
  * **Device placement strategy** - Determines how workloads are distributed across GPU devices within each node:
    * **Bin-pack** - The Scheduler places as many workloads as possible on each GPU device to minimize the number of GPU devices used.
    * **Spread** - The Scheduler spreads workloads across GPU devices within each node to maximize the available resources per workload.
  * The Scheduler applies these strategies in order: first the node placement strategy to select nodes, then the device placement strategy to place pods across GPU devices within each selected node.
* **CPU workloads:**
  * **Bin-pack** - The Scheduler places as many workloads as possible in each CPU and node to use fewer resources and maximize CPU and node vacancy.
  * **Spread** - The Scheduler spreads workloads across as many CPUs and nodes as possible to minimize the load and maximize the available resources per workload.

## Workload Resources

The primary function of the NVIDIA Run:ai Scheduler is to match workloads with available nodes that satisfy their resource requirements and other constraints (such as node type or topology). Workload resources generally fall into three categories:

* **Kubernetes Native Resources** - CPUs and CPU memory.
* **Kubernetes Extended Resources** - Non-native Kubernetes resources such as GPUs, FPGAs, or networking adapters. Historically requested as static "extended resources" (for example, `nvidia.com/gpu: 2`). Starting with Kubernetes v1.34, workloads can leverage [Dynamic Resource Allocation (DRA)](#dynamic-resource-allocation-dra).
* **Kubernetes Storage Resources** - Persistent Volumes (PVs) and Persistent Volume Claims (PVCs), managed by the Kubernetes CSI subsystem.

{% hint style="info" %}
**Note**

With Kubernetes v1.34, DRA is GA. You can also use earlier Kubernetes version v1.33 where DRA was Beta.
{% endhint %}

### Dynamic Resource Allocation (DRA)

Dynamic Resource Allocation (DRA) is a Kubernetes-native mechanism that provides a more flexible alternative to static extended resources. Instead of requesting a fixed number of devices, workloads can express complex hardware requirements - such as GPU type, minimum GPU memory, GPU sharing, or priority lists - using `ResourceClaim` and `ResourceClaimTemplate` objects. `ResourceClaims` and `ResourceClaimTemplates` define the proprietary parameters of the requested hardware, such as GPUs or communication channels.

NVIDIA Run:ai has introduced DRA support progressively. Each supported resource type in DRA requires the respective vendor to advertise their proprietary driver:

* **v2.20** - DRA support for `ComputeDomain` `ResourceClaims`
* **v2.25** - DRA support for NVIDIA GPUs

The NVIDIA Run:ai platform supports scheduling and presentation of workloads using DRA claims in two scenarios:

* Workloads submitted directly to the cluster
* Workloads submitted [via YAML](/saas/workloads-in-nvidia-run-ai/submit-via-yaml) using NVIDIA Run:ai UI, API and CLI

In both scenarios, the platform schedules the workloads, accounts for the requested resources in quota management, and surfaces workload status and other parameters via the UI, API, and CLI.

{% hint style="info" %}
**Note**

* When submitting NVIDIA Run:ai native workloads, the platform continues using extended resources. Both extended resources and DRA are supported during the transition period.
* Using DRA and extended resources on the same underlying nodes can cause resource inconsistencies and allocation collisions. It is recommended to separate nodes by splitting them into different node pools - one for extended resources and one for DRA. This allows you to adopt DRA incrementally while maintaining the legacy 'extended resources’ workloads and nodes unaffected.
  {% endhint %}

## Next Steps

Now that you have learned the key concepts and principles of the NVIDIA Run:ai Scheduler, see [how the Scheduler works](/saas/platform-management/runai-scheduler/scheduling/how-the-scheduler-works) - allocating pods to workloads, applying preemption mechanisms, and managing resources.


# How the Scheduler Works

Efficient resource allocation is critical for managing AI and compute-intensive workloads in Kubernetes clusters. The NVIDIA Run:ai Scheduler enhances Kubernetes’ native capabilities by introducing advanced scheduling principles such as fairness, quota management, dynamic resource balancing, multi-level pod grouping, and topology-aware scheduling. It ensures that workloads, ranging from simple single-pod workloads to distributed multi-pod workloads and disaggregated workloads composed of multiple cooperating components, are allocated resources effectively while adhering to organizational policies and priorities.

This guide explores the NVIDIA Run:ai Scheduler’s allocation process, preemption mechanisms, and resource management. Through examples and detailed explanations, you'll gain insights into how the Scheduler dynamically balances workloads to optimize cluster utilization and maintain fairness across [projects and departments](/saas/platform-management/aiinitiatives/adapting-ai-initiatives#mapping-your-organization).

## Allocation Process

Resource allocation follows a hierarchical order. Workloads submitted to higher-ranked projects (and departments) are served before workloads submitted to lower-ranked projects (and departments). As a result, a lower-priority workload submitted to a higher-ranked project may be served before a higher-priority workload submitted to a lower-ranked project.

The NVIDIA Run:ai Scheduler first allocates resources to the highest-ranked departments and their projects. It prioritizes allocating all possible workload resource requests within deserved quota. Only after all in-quota allocations are satisfied does the Scheduler allocate over-quota resources, which are resources not used by any in-quota projects or departments.

Over-quota resources are also assigned according to department and project ranks. The amount of over-quota resources each project or department receives is determined by its rank and its weight within that rank, relative to other departments and projects.

### Pod Creation and Grouping

When a workload is submitted, the submitting workload controller creates a single pod or multiple pods (for distributed training workloads or deployment based inference). When the Scheduler gets a submit request with the first pod, it creates a [pod group](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles#workloads-and-pod-groups) and a [sub group](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles#sub-groups), and allocates all the relevant building blocks of that workload. The next pods of the same workload are attached to the same pod group, either to the same sub-group or to another sub-group within the top pod-group.

### Queue Management

A workload, together with its associated pod group and sub-groups, is placed in the appropriate [scheduling queue](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles#scheduling-queue). In each scheduling cycle, the Scheduler determines the order in which queues are served by calculating their scheduling precedence. Projects and departments each have a separate queue per node pool.

For each workload considered for scheduling, the Scheduler evaluates the state of the project queue and then the state of the department queue to determine whether the workload is eligible for scheduling, based on available resources within the context of the project, department, and node pool.

If the sum of all project quotas exceeds the quota of their parent department, the Scheduler does not allow non-preemptible workloads to consume resources beyond the department’s quota.

### Resource Binding

The next step in the scheduling process is resource binding. During this step, the Scheduler selects nodes for the pods, assigns each pod to a node (the bind operation), and binds additional pod resources such as storage, ingress, and other required components.

If the pod group has a [gang scheduling](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles#gang-scheduling) or [multi-level gang scheduling rule](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles#multi-level-gang-scheduling) applied, the Scheduler either allocates and binds all pods together or places all of them in a pending state. If the latter occurs, the Scheduler retries scheduling the entire pod group together in the next scheduling cycle.

During this process, the Scheduler updates the status of the pods and their associated pod group and sub-groups. Users can track the workload submission and scheduling process using either the CLI or the NVIDIA Run:ai UI. For more details on submitting and managing workloads, see [Workloads](/saas/workloads-in-nvidia-run-ai/workloads).

## Preemption

If the Scheduler cannot find sufficient resources to schedule a workload (including all of its associated pods), and the workload is entitled to resources, either because it is within its deserved quota or within its [fairshare](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles#fairshare-and-fairshare-balancing), the Scheduler attempts the following actions, in order:

1. **Consolidation (bin packing)** - The Scheduler first tries to consolidate workloads into smaller number of nodes to make room for the currently scheduled workload.
2. **Resource rebalancing across queues** - If consolidation does not succeed, the Scheduler attempts to [reclaim resources](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles#reclaim-of-resources-between-projects-and-departments) from other queues (projects) that are either over their deserved quota or over their calculated fairshare.
3. **Preemption within the same queue** - If resources are still unavailable, the Scheduler tries to preempt [lower priority preemptible workloads](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles#preemption-of-lower-priority-workloads-within-a-project) within the same queue (project).

### When a Queue Receives Additional Resources

A [scheduling queue](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles#scheduling-queue) deserves more resources from the pool of unused resources or from other queues if the following conditions are met:

1. The queue has workloads that request resources.
2. The queue has not reached its deserved quota, or, it is already over its deserved quota but below its fairshare (= deserved quota + over quota share of the free resources).

In this situation, the Scheduler attempts to rebalance resources across queues to provide the entitled queue with its deserved resources. If rebalancing does not resolve the resource shortage, the Scheduler tries to preempt [lower priority preemptible workloads](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles#preemption-of-lower-priority-workloads-within-a-project) within the same queue (project).

### Scheduling Above Fairshare

In certain scenarios, queues that are above their calculated fairshare may still receive additional resources. This can occur when:

* Other queues in the node pool are not fully using their deserved quota or fairshare, leaving unused over-quota resources.
* The combination of resources requests by workloads cannot be fulfilled simultaneously and the Scheduler has to tie break and decide which queue and which workload will be scheduled first.

For example, a cluster with 10 GPUs and two projects each with quota of 5 GPUs and over-quota allowed, both try to submit a 10 GPU workload. In this case both are over their deserved quota and fairshare and the Scheduler cannot fulfill both requests simultaneously because their total request is 20 GPUs while the total GPUs in the cluster is 10. Therefore, the Scheduler will give precedence the project (queue) that submitted the workload first.

### Consolidation Action

A consolidation action is applied when the Scheduler fails to find sufficient resources (nodes) for the workload currently being scheduled. Consolidation is a process in which the Scheduler evaluates whether re-scheduling other preemptible workloads can free enough resources to allow the current workload to be scheduled successfully.

During consolidation, the Scheduler may preempt workloads from other queues (projects or departments). However, consolidation is performed only if all preempted workloads can be immediately re-scheduled and continue running. To ensure this, the Scheduler simulates different consolidation options in memory. Only if a simulation succeeds, meaning that all workloads, including the currently scheduled workload, can be successfully re-scheduled, does the Scheduler perform the actual consolidation operation.

Administrators can control the number of consolidation evaluations the Scheduler performs when searching for a viable re-scheduling combination, or they can disable consolidation entirely. Note that disabling consolidation may increase cluster resource fragmentation and reduce overall cluster utilization.

### Reclaim Preemption Between Projects and Departments

Reclaim is an inter-project and inter-department resource balancing action that takes back resources from one project or department that has used them as an over quota. It returns the resources back to a project (or department) that deserves those resources as part of its deserved quota, or to re-balance fairshare between projects (or departments), this ensures a project (or department) does not exceed its fairshare.

This mode of operation means that a lower priority workload submitted in one project (e.g. training) can reclaim resources from a project that runs a higher priority workload (e.g. preemptive workspace) if fairness re-balancing is required.

{% hint style="info" %}
**Note**

Only preemptible workloads can consume over-quota resources, as these workloads are subject to resource reclamation through preemption across projects and departments. The amount of over-quota resources that a project or department can receive depends on its over quota weight, or on its quota when the Over quota weight setting is disabled.

When the Over quota weight setting is disabled, the weight of each project or department is derived directly from its assigned quota. In both cases, over-quota resources are distributed proportionally across scheduling queues based on their relative queue [weight](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles#over-quota-weight).

*A scheduling queue is defined per project/node pool or department/node pool.*
{% endhint %}

### Priority Preemption Within a Project

[Higher priority workloads](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles#priority-and-preemption) may preempt lower priority preemptible workloads within the same project/node pool queue. For example, in a project that runs a training workload that uses the project within quota (for a certain node pool), a newly submitted workspace within the same project/node pool may stop (preempt) the training workload if there are not enough over quota resources for the project within that node pool to run both workloads (i.e. workspace using in-quota resources and training using over quota resources).

{% hint style="info" %}
**Note**

Workload priority applies only within the same project and does not influence workloads across different projects, where fairness determines precedence.
{% endhint %}

## Quota, Over Quota, and Fairshare

The NVIDIA Run:ai Scheduler is designed to ensure fairness between projects and between departments. Each department and project is always allocated its deserved quota first, and only then may receive over-quota resources.

The Scheduler splits unused resources between projects and departments according to defined fairness rules. It first fulfills the deserved quota of departments and projects, giving precedence to higher-ranked departments and then to higher-ranked projects. If additional resources beyond the deserved quota are required, unused resources are split as over-quota resources according to fairness rules: queues with higher rank are served first, and fairshare is calculated per queue based on its rank and weight (where a higher weight results in a larger share of unused resources compared to other queues). Between queues with the same rank, queues that are more resource-starved are prioritized.

If a project requires resources beyond its calculated fairshare and the Scheduler finds unused resources that no other project requires, the project may consume resources beyond its fairshare. For example, consider two queues with equal rank and weight, each with a quota of 5 GPUs and over-quota enabled, in a cluster with 20 GPUs. In this case, the theoretical fairshare for each queue is 10 GPUs (5 deserved quota plus 5 over-quota). However, if queue #2 requests only 5 GPUs, its effective fairshare becomes 5 GPUs, while queue #1 requests 15 GPUs and its fairshare becomes 15 GPUs for the current scheduling cycle. In the next scheduling cycle, this distribution may change.

Some scenarios can prevent the Scheduler from fully providing deserved quota and fairness:

* Fragmentation or other scheduling constraints such affinities, taints, topology constraints etc.
* Some requested resources, for example as GPUs and CPU memory, are potentially available for allocation, while others, like CPU cores, are insufficient to meet the request. As a result, the Scheduler will place the workload in a pending state until all required resource becomes available for allocation.

### Fairshare Calculations Methods

The NVIDIA Run:ai Scheduler’s point-in-time fairshare calculation method allocates over-quota resources to a scheduling queue based on its weight relative to other queues at the same rank. The fairshare score is recalculated during every scheduling cycle and can therefore change frequently, potentially triggering scheduling adjustments.

For example, a workload whose queue is entitled to a certain amount of over-quota resources in one scheduling cycle may find that its queue has a lower fairshare score in the next cycle, causing the Scheduler to preempt the workload.

This method represents a short-term fairshare scoring approach. It enables departments and projects to receive a fast response when over-quota resources are required, but it may also result in frequent changes to the state of running over-quota workloads.

Administrators can control the frequency of such changes by configuring a guaranteed min-runtime per node pool. This setting ensures that once a preemptible workload starts running, it effectively behaves as non-preemptible for the duration of the guaranteed min-runtime.

{% hint style="info" %}
**Note**

Setting a guaranteed min-runtime on a node pool that also serves high-priority or interactive/real-time workloads may introduce significant delays for those workloads when they need to run. As a result, such workloads may effectively lose their interactive or real-time characteristics and break their SLA.
{% endhint %}

Below you can find a numerical example of how point-in-time fairshare scoring actually works.

### Example of Splitting Quota

The example below illustrates a split of quota between different projects and departments using several node pools:

![](/files/Ny513YRetiHCFl7L8zLw)

The example below illustrates how fairshare scoring is calculated per project/node pool for the above example:

![](/files/mG8JjoGtMrJQqTSr1n3K)

* For each Project:

  * The **over quota (OQ)** portion of each project (per node pool) is calculated as:

  \[(OQ-Weight) / (Σ Projects OQ-Weights)] x (Unused Resource per node pool)

  * **Fairshare** is calculated as the sum of quota + over quota.
* In Project 2, we assume that out of the 36 available GPUs in node pool A, 20 GPUs are currently unused. This means either these GPUs are not part of any project’s quota, or they are part of a project’s quota but not used by any workloads of that project:
  * Project 2 **over quota share**:

    \[(Project 2 OQ-Weight) / (Σ all Projects OQ-Weights)] x (Unused Resource within node pool A)

    \[(3) / (2 + 3 + 1)] x (20) = (3/6) x 20 = 10 GPUs
  * **Fairshare** = deserved quota + over quota = 6 +10 = 16 GPUs. Similarly, fairshare is also calculated for CPU and CPU memory. The Scheduler can grant a project more resources than its fairshare if the Scheduler finds resources not required by other projects that may deserve those resources.
* In Project 3, **fairshare** = deserved quota + over quota = 0 +3 = 3 GPUs. Project 3 has no guaranteed quota, but it still has a share of the excess resources in node pool A. The NVIDIA Run:ai Scheduler ensures that Project 3 receives its part of the unused resources for over quota, even if this results in reclaiming resources from other projects and preempting preemptible workloads.

## Fairshare Balancing

The Scheduler constantly re-calculates the fairshare of each project and department per node pool, represented in the scheduler as queues, resulting in the re-balancing of resources between projects and between departments. This means that a preemptible workload that was granted resources to run in one scheduling cycle, can find itself preempted and go back to pending state while waiting for resources in the next cycle.

A queue, representing a scheduler-managed object for each project or department per node pool, can be in one of 3 states:

* **In-quota**: The queue’s allocated resources ≤ queue deserved quota. The Scheduler’s first priority is to ensure each queue receives its deserved quota.
* **Over quota but below fairshare**: The queue’s deserved quota < queue’s allocated resources <= queue’s fairshare. The Scheduler tries to find and allocate more resources to queues that need resources beyond their deserved quota and up to their fairshare.
* **Over-fairshare and over quota**: The queue’s fairshare < queue’s allocated resources. The Scheduler tries to allocate resources to queues that need even more resources beyond their fairshare.

When re-balancing resources between queues of different projects and departments, the Scheduler goes in the opposite direction, i.e. first take resources from over-fairshare queues, then from over quota queues, and finally, in some scenarios, even from queues that are below their deserved quota (for example, if a workload is partially in-quota and partially over-quota, in this case it is considered over-quota and may be preempted and its resources reclaimed).

![](/files/Bu6wiPtEZvpxOqEaRzgW)

## Time-Based Fairshare

Time-based fairshare is a fairshare scoring mode in which fairshare is calculated based on historical resource usage over time, rather than solely on point-in-time scheduling demand within a single scheduling cycle.

When time-based fairshare is enabled in a node pool, the Scheduler continuously collects real-time resource usage data and uses this data to calculate the fairshare of each queue. This approach enables fairshare to reflect actual resource consumption over time, allowing resources to be distributed more evenly and reducing long-term imbalances caused by short-term demand fluctuations and optimizing resource allocation dynamically.

Resource usage data in the cluster is collected and persisted by default using Prometheus. This data is then used by the Scheduler to perform fairshare calculations. Queues that consume more resources over time relative to their fairshare receive fewer over-quota resources compared to other queues, while queues that consume less than their fairshare over time are rewarded. Over time, this will result in the queues' over-quota resources being reclaimed by more starved queues, thus achieving a more fair allocation of resources over time.

To avoid scenarios in which one queue consistently consumes its fairshare steadily while another queue attempts to consume a large amount of resources over a short period, the Scheduler regulates historical usage data by optionally applying time-based decay. This ensures that queues consuming resources regularly in line with their fairshare are not penalized by queues that attempt to consume a large amount of resources in a short time window.

The NVIDIA Run:ai platform administrator can configure how historical usage data is evaluated, including the time window over which usage is calculated, the exponential decay applied to historical data, and the influence of historical usage on the current fairshare calculation (also referred to as the **K-value**). These settings control how much weight is given to recent usage versus past usage. For more details, see [Node pools](/saas/platform-management/aiinitiatives/resources/node-pools).

## Next Steps

Now that you have gained insights into how the Scheduler dynamically balances workloads to optimize cluster utilization and maintain fairness across projects and departments, you can [submit workloads](/saas/workloads-in-nvidia-run-ai/workloads). Before submitting your workloads, it’s important to familiarize yourself with the following key topics:

* [Introduction to workloads](/saas/workloads-in-nvidia-run-ai/introduction-to-workloads) - Learn what workloads are and what is supported for both NVIDIA Run:ai and third-party workloads.
* [NVIDIA Run:ai workload types](/saas/workloads-in-nvidia-run-ai/workload-types) - Explore the various NVIDIA Run:ai workload types available and understand their specific purposes to enable you to choose the most appropriate workload type for your needs.


# Setting the Default Scheduler

By default, Kubernetes uses its own native scheduler to determine pod placement. The NVIDIA Run:ai platform provides a custom scheduler, `runai-scheduler`, which is used by default for workloads submitted using the [NVIDIA Run:ai](/saas/workloads-in-nvidia-run-ai/introduction-to-workloads) platform.

This guide outlines how to configure workloads submitted directly to Kubernetes or through external frameworks to run with the [NVIDIA Run:ai Scheduler](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles), instead of the default Kubernetes scheduler.

## Enforce the Scheduler at the Namespace Level

When submitting workloads in a given namespace (i.e., NVIDIA Run:ai [project](/saas/platform-management/aiinitiatives/organization/projects)), the parameter `enforceRunaiScheduler` is enabled (true) by default. This ensures that any workload associated with a NVIDIA Run:ai project automatically uses the `runai-scheduler`, including workloads submitted directly to Kubernetes or through external frameworks.

If this parameter is disabled, `enforceRunaiScheduler=false`, workloads will no longer default to the NVIDIA Run:ai Scheduler. In this case, you can still use the NVIDIA Run:ai Scheduler by specifying it manually in the workload YAML.

## Specify the Scheduler in the Workload YAML

To use the NVIDIA Run:ai Scheduler, specify it in the workload’s YAML file. This instructs Kubernetes to schedule the workload using the NVIDIA Run:ai Scheduler instead of the default one.

```yaml
spec:
  schedulerName: runai-scheduler
```

**For example:**

```yaml
apiVersion: v1
kind: Pod
metadata:
  annotations:
    user: test
    gpu-fraction: "0.5"
    gpu-fraction-num-devices: "2"
  labels:
    runai/queue: test
  name: multi-fractional-pod-job
  namespace: test
spec:
  containers:
  - image: gcr.io/run-ai-demo/quickstart-cuda
    imagePullPolicy: Always
    name: job
    env:
    - name: RUNAI_VERBOSE
      value: "1"
    resources:
      limits:
        cpu: 200m
        memory: 200Mi
      requests:
        cpu: 100m
        memory: 100Mi
    securityContext:
      capabilities:
        drop: ["ALL"]
  schedulerName: runai-scheduler
  serviceAccount: default
  serviceAccountName: default
  terminationGracePeriodSeconds: 5
```


# Workload Priority and Preemption

NVIDIA Run:ai defines workload priority and preemptibility to determine how workloads are scheduled within a project. These mechanisms influence scheduling order, resource allocation, and whether running workloads may be interrupted when higher-priority workloads require resources.

* **Workload priority** - Determines the workload's position in the project scheduling queue managed by the NVIDIA Run:ai [Scheduler](/saas/platform-management/runai-scheduler/scheduling/how-the-scheduler-works). By adjusting the priority, you can increase the likelihood that a workload will be scheduled and preferred over others within the same project, ensuring that critical tasks are given higher priority and resources are allocated efficiently.
* **Workload preemptibility** - Determines the workload's resource usage policy and its guarantee against interruption:
  * Non-preemptible workloads must run within the project’s deserved quota, cannot use over-quota resources, and will not be interrupted once scheduled.
  * Preemptible workloads can use opportunistic resources beyond the project’s quota and may be interrupted at any time by higher priority workload, even if running within the project's quota.

{% hint style="info" %}
**Note**

This applies only within a single project. It does not impact the scheduling queues or workloads of other projects.
{% endhint %}

## Priority Dictionary

Workload priority is defined by selecting a priority from a predefined list in the NVIDIA Run:ai priority dictionary. Each string corresponds to a specific Kubernetes [PriorityClass](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles#priority-and-preemption), which in turn determines scheduling behavior.

<table><thead><tr><th width="370.5078125">Priority Class Name</th><th>Kubernetes Priority Value</th></tr></thead><tbody><tr><td><code>very-low</code></td><td>25</td></tr><tr><td><code>low</code></td><td>40</td></tr><tr><td><code>medium-low</code></td><td>65</td></tr><tr><td><code>medium</code></td><td>80</td></tr><tr><td><code>medium-high</code></td><td>90</td></tr><tr><td><code>high</code></td><td>125</td></tr><tr><td><code>very-high</code></td><td>150</td></tr></tbody></table>

## Default Priority and Preemptibility per Workload

NVIDIA Run:ai defines the following default mappings of workload types to priorities and preemptibility. Each workload type comes with a default category that determines it default priority and preemptibility value. To retrieve the default priority and preemptibility per workload type, refer to the [List workload types](https://run-ai-docs.nvidia.com/api/workloads/workload-properties#get-api-v1-workload-types) API.

{% hint style="info" %}
**Note**

* For more information on workload support, see [Introduction to workloads](/saas/workloads-in-nvidia-run-ai/introduction-to-workloads).
* Changing the priority is not supported for NVCF workloads.
  {% endhint %}

| Category | Default Priority | Default Preemptibility |
| -------- | ---------------- | ---------------------- |
| Build    | High             | Non-preemptible        |
| Train    | Low              | Preemptible            |
| Deploy   | Very high        | Non-preemptible        |

### NVIDIA Run:ai Native Workloads

<table><thead><tr><th>Workload Type</th><th data-type="checkbox">Build</th><th data-type="checkbox">Train</th><th data-type="checkbox">Deploy</th></tr></thead><tbody><tr><td>Workspaces</td><td>true</td><td>false</td><td>false</td></tr><tr><td>Standard training</td><td>false</td><td>true</td><td>false</td></tr><tr><td>Distributed training</td><td>false</td><td>true</td><td>false</td></tr><tr><td>Custom inference</td><td>false</td><td>false</td><td>true</td></tr><tr><td>NVIDIA NIM inference</td><td>false</td><td>false</td><td>true</td></tr><tr><td>Hugging Face inference</td><td>false</td><td>false</td><td>true</td></tr><tr><td>Distributed inference</td><td>false</td><td>false</td><td>true</td></tr></tbody></table>

### Supported Workload Types

<table><thead><tr><th>Workload Type</th><th data-type="checkbox">Build</th><th data-type="checkbox">Train</th><th data-type="checkbox">Deploy</th></tr></thead><tbody><tr><td>AMLJob</td><td>false</td><td>true</td><td>false</td></tr><tr><td>CronJob</td><td>false</td><td>false</td><td>true</td></tr><tr><td>Deployment</td><td>false</td><td>false</td><td>true</td></tr><tr><td>DevWorkspace</td><td>true</td><td>false</td><td>false</td></tr><tr><td>DynamoGraphDeployment</td><td>false</td><td>false</td><td>true</td></tr><tr><td>InferenceService (KServe)</td><td>false</td><td>false</td><td>true</td></tr><tr><td>JAXJob</td><td>false</td><td>true</td><td>false</td></tr><tr><td>Job</td><td>true</td><td>false</td><td>false</td></tr><tr><td>JobSet</td><td>false</td><td>true</td><td>false</td></tr><tr><td>LeaderWorkerSet (LWS)</td><td>false</td><td>false</td><td>true</td></tr><tr><td>MPIJob</td><td>false</td><td>true</td><td>false</td></tr><tr><td>NIMCache</td><td>false</td><td>false</td><td>true</td></tr><tr><td>NIMServices</td><td>false</td><td>false</td><td>true</td></tr><tr><td>NodeSet (Slinky)</td><td>false</td><td>false</td><td>true</td></tr><tr><td>Notebook</td><td>true</td><td>false</td><td>false</td></tr><tr><td>PipelineRun</td><td>false</td><td>false</td><td>true</td></tr><tr><td>Pod</td><td>false</td><td>false</td><td>true</td></tr><tr><td>PyTorchJob</td><td>false</td><td>true</td><td>false</td></tr><tr><td>RayCluster</td><td>false</td><td>true</td><td>false</td></tr><tr><td>RayJob</td><td>false</td><td>true</td><td>false</td></tr><tr><td>RayService</td><td>false</td><td>false</td><td>true</td></tr><tr><td>ReplicaSet</td><td>false</td><td>false</td><td>true</td></tr><tr><td>ScheduledWorkflow</td><td>false</td><td>false</td><td>true</td></tr><tr><td>SeldonDeployment</td><td>false</td><td>false</td><td>true</td></tr><tr><td>Service</td><td>false</td><td>false</td><td>true</td></tr><tr><td>SPOTRequest</td><td>false</td><td>false</td><td>false</td></tr><tr><td>StatefulSet</td><td>false</td><td>false</td><td>true</td></tr><tr><td>TaskRun</td><td>true</td><td>false</td><td>false</td></tr><tr><td>TFJob</td><td>false</td><td>true</td><td>false</td></tr><tr><td>VirtualMachineInstance</td><td>false</td><td>true</td><td>false</td></tr><tr><td>Workflow</td><td>false</td><td>false</td><td>true</td></tr><tr><td>XGBoostJob</td><td>false</td><td>true</td><td>false</td></tr></tbody></table>

## Setting Priority and Preemptibility During Workload Submission

{% hint style="info" %}
**Note**

* If preemptibility is not explicitly configured, the system uses the default preemptibility behavior associated with the selected workload priority.
* Changing a workload’s priority and preemptibility may impact its ability to be scheduled. For example, switching a workload from a low priority, preemptible value (which allows over-quota usage) to high priority, non-preemptible value (which requires in-quota resources) may reduce its chances of being scheduled in cases where the required quota is unavailable.
  {% endhint %}

### NVIDIA Run:ai Native Workloads

For native NVIDIA Run:ai workloads, priority and preemptibility can be set during workload submission using one of the following methods:

* **UI** - Set workload priority and preemptibility under **General** settings
* **API** - Set using the `priorityClass` and `preemptibility` field
* **CLI** - Set using the `--priority` and `--preemptibility` flag

### Supported Workload Types

{% hint style="info" %}
**Note**

If priority or preemptibility is set through the UI, API, or CLI, those values override any values defined in the YAML manifest.
{% endhint %}

For [supported workload types](/saas/workloads-in-nvidia-run-ai/workload-types/supported-workload-types) submitted with a YAML manifest, priority and preemptibility can be set as follows:

* **UI** - Set workload priority and preemptibility under **General** settings
* **API** - Set using the `priority` and `preemptibility` fields
* **CLI** - Set using the `--priority` and `--preemptibility` flags
* **via YAML manifest** - Set by adding the following labels to your YAML manifest under the `metadata.labels` section of your workload definition.
  * Use the following values for priority - `very-low`, `low`, `medium-low`, `medium`, `medium-high`, `high`, `very-high` :
  * Use the following values for preemptibility - `preemptible` or `non-preemptible`

    ```yaml
    metadata:
      labels:
        priorityClassName: <priority>
        kai.scheduler/preemptibility: <preemptibility_value>
    ```

## Updating the Default Mapping

Administrators can change the default priority and preemptibility assigned to a workload type by updating the mapping using the [NVIDIA Run:ai API](https://run-ai-docs.nvidia.com/api/). To update the priority mapping:

1. Retrieve the list of workload types and their IDs using `GET /api/v1/workload-types`.
2. Identify the `workloadTypeId` of the workload type you want to modify.
3. Retrieve the list of available priorities and their IDs using `GET /api/v1/workload-priorities`.
4. Send a request to update the workload type with the new priority using\
   `PUT /api/v1/workload-types/{workloadTypeId}` and include the `priorityId` in the request body.

## Using API

Go to the [Workload priorities](https://run-ai-docs.nvidia.com/api/workloads/workload-properties#get-api-v1-workload-priorities) API reference to view the available actions.


# Quick Starts


# Over Quota, Fairness and Preemption

This quick start provides a step-by-step walkthrough of the core scheduling concepts - [over quota](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles#over-quota), [fairness](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles#fairness-fair-resource-distribution), and [preemption](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles#priority-and-preemption). It demonstrates the simplicity of resource provisioning and how the system eliminates bottlenecks by allowing users or teams to exceed their resource quota when free GPUs are available.

* **Over quota** - In this scenario, team-a runs two training workloads and team-b runs one. Team-a has a quota of 3 GPUs and is over quota by 1 GPU, while team-b has a quota of 1 GPU. The system allows this over quota usage as long as there are available GPUs in the cluster.
* **Fairness and preemption** - Since the cluster is already at full capacity, when team-b launches a new b2 workload requiring 1 GPU , team-a can no longer remain over quota. To maintain fairness, the [NVIDIA Run:ai Scheduler](/saas/platform-management/runai-scheduler/scheduling/how-the-scheduler-works) preempts workload a1 (1 GPU), freeing up resources for team-b.

## Prerequisites

* You have created two [projects](/saas/platform-management/aiinitiatives/organization/projects) - team-a and team-b - or have them created for you.
* Each project has an assigned quota of 2 GPUs. In this example, we have 4 GPUs on 2 machines with 2 GPUs each.

{% hint style="info" %}
**Note**

Flexible workload submission is enabled by default. If unavailable, contact your administrator to enable it under **General settings** → Workloads → Flexible workload submission.
{% endhint %}

## Step 1: Logging In

{% tabs %}
{% tab title="UI" %}
Browse to the provided NVIDIA Run:ai user interface and log in with your credentials.
{% endtab %}

{% tab title="CLI v2" %}
Run the below --help command to obtain the login options and log in according to your setup:

```sh
runai login --help
```

{% endtab %}

{% tab title="API" %}
To use the API, you will need to obtain a token as shown in [API authentication](https://run-ai-docs.nvidia.com/api/getting-started/how-to-authenticate-to-the-api).
{% endtab %}
{% endtabs %}

## Step 2: Submitting the First Training Workload (team-a) <a href="#i3c9jpfzerlq" id="i3c9jpfzerlq"></a>

{% tabs %}
{% tab title="UI - Flexible" %}

1. Go to Workload manager → Workloads
2. Click **+NEW WORKLOAD** and select **Training**
3. Select under which **cluster** to create the workload
4. Select the **project** named team-a
5. Under **Workload architecture**, select **Standard**
6. Select **Start from scratch** to launch a new training quickly
7. Enter **a1** as the workload **name**
8. Click **CONTINUE**

   In the next step:
9. Under **Environment**, enter the **Image URL** - `runai.jfrog.io/demo/quickstart`
10. Under **Compute resources**, click the **load** icon. A side pane appears, displaying a list of available compute resources. Select the **‘one-gpu’** compute resource for your workload.
    * If ‘one-gpu’ is not displayed, follow the below steps to create a one-time compute resource configuration:
      * Set **GPU devices** per pod - 1
      * Optional: set the **CPU compute** per pod - 0.1 cores (default)
      * Optional: set the **CPU memory** per pod - 100 MB (default)
11. Click **CREATE TRAINING**
    {% endtab %}

{% tab title="UI - Original" %}

1. Go to the Workload Manager → Workloads
2. Click **+NEW WORKLOAD** and select **Training**
3. Select under which **cluster** to create the workload
4. Select the **project** named team-a
5. Under **Workload architecture**, select **Standard**
6. Select **Start from scratch** to launch a new training quickly
7. Enter **a1** as the workload **name**
8. Click **CONTINUE**\
   In the next step:
9. Create a new environment:

   * Click **+NEW ENVIRONMENT**
   * Enter quick-start as the **name** for the environment. The name must be unique.
   * Enter the **Image URL** - `runai.jfrog.io/demo/quickstart`
   * Click **CREATE ENVIRONMENT**

   The newly created environment will be selected automatically
10. Select the **‘one-gpu’** compute resource for your workload

    * If ‘one-gpu’ is not displayed in the gallery, follow the below steps:
      * Click **+NEW COMPUTE RESOURCE**
      * Enter one-gpu as the **name** for the compute resource. The name must be unique.
      * Set **GPU devices** per pod - 1
      * Optional: set the **CPU compute** per pod - 0.1 cores (default)
      * Optional: set the **CPU memory** per pod - 100 MB (default)
      * Click **CREATE COMPUTE RESOURCE**

    The newly created compute resource will be selected automatically
11. Click **CREATE TRAINING**
    {% endtab %}

{% tab title="CLI v2" %}
Copy the following command to your terminal. For more details, see [CLI reference](/saas/reference/cli/runai):

```sh
runai training submit a1 -i runai.jfrog.io/demo/quickstart -g 1 -p team-a
```

{% endtab %}

{% tab title="API" %}
Copy the following command to your terminal. Make sure to update the following parameters. For more details, see [Trainings](https://run-ai-docs.nvidia.com/api/workloads/trainings) API.

```bash
curl --location 'https://<COMPANY-URL>/api/v1/workloads/trainings' \
--header 'Content-Type: application/json' \
--header 'Authorization: Bearer <TOKEN>' \ 
--data '{
  "name": "a1",
  "projectId": "<PROJECT-ID>", 
  "clusterId": "<CLUSTER-UUID>",
  "spec": {
    "image":"runai.jfrog.io/demo/quickstart",
    "compute": {
      "gpuDevicesRequest": 1
    }
  }
}'
```

* `<COMPANY-URL>` - The link to the NVIDIA Run:ai user interface
* `<TOKEN>` - The API access token obtained in [Step 1](#a13adq7eth7w)
* `<PROJECT-ID>` - The ID of the Project the workload is running on. You can get the Project ID via the [Get Projects](https://run-ai-docs.nvidia.com/api/organizations/projects#get-api-v1-org-unit-projects) API.
* `<CLUSTER-UUID>` - The unique identifier of the Cluster. You can get the Cluster UUID via the [Get Clusters](https://run-ai-docs.nvidia.com/api/organizations/clusters#get-api-v1-clusters) API.

{% hint style="info" %}
**Note**

The above API snippet runs with NVIDIA Run:ai clusters of 2.18 and above only.
{% endhint %}
{% endtab %}
{% endtabs %}

## Step 3: Submitting the Second Training Workload (team-a) <a href="#i3c9jpfzerlq" id="i3c9jpfzerlq"></a>

{% tabs %}
{% tab title="UI - Flexible" %}

1. Go to the Workload Manager → Workloads
2. Click **+NEW WORKLOAD** and select **Training**
3. Select the **cluster** where the previous training workload was created
4. Select the **project** named team-a
5. Under **Workload architecture**, select **Standard**
6. Select **Start from scratch** to launch a new training quickly
7. Enter **a2** as the workload **name**
8. Click **CONTINUE**\
   In the next step:
9. Under **Environment**, enter the **Image URL** - `runai.jfrog.io/demo/quickstart`
10. Under **Compute resources**, click the **load** icon. A side pane appears, displaying a list of available compute resources. Select the **‘two-gpus’** compute resource for your workload.
    * If ‘two-gpus’ is not displayed, follow the below steps to create a one-time compute resource configuration:
      * Set **GPU devices** per pod - 2
      * Optional: set the **CPU compute** per pod - 0.1 cores (default)
      * Optional: set the **CPU memory** per pod - 100 MB (default)
11. Click **CREATE TRAINING**
    {% endtab %}

{% tab title="UI - Original" %}

1. Go to the Workload Manager → Workloads
2. Click **+NEW WORKLOAD** and select **Training**
3. Select the **cluster** where the previous training workload was created
4. Select the **project** named team-a
5. Under **Workload architecture**, select **Standard**
6. Select **Start from scratch** to launch a new training quickly
7. Enter **a2** as the workload **name**
8. Click **CONTINUE**\
   In the next step:
9. Select the environment created in [Step 2](#i3c9jpfzerlq)
10. Select the **‘two-gpus’** compute resource for your workload

    * If ‘two-gpus’ is not displayed in the gallery, follow the below steps:
      * Click **+NEW COMPUTE RESOURCE**
      * Enter two-gpus as the **name** for the compute resource. The name must be unique.
      * Set **GPU devices** per pod - 2
      * Optional: set the **CPU compute per pod** - 0.1 cores (default)
      * Optional: set the **CPU memory per pod** - 100 MB (default)
      * Click **CREATE COMPUTE RESOURCE**

    The newly created compute resource will be selected automatically
11. Click **CREATE TRAINING**
    {% endtab %}

{% tab title="CLI v2" %}
Copy the following command to your terminal. For more details, see [CLI reference](/saas/reference/cli/runai):

```sh
runai training submit a2 -i runai.jfrog.io/demo/quickstart -g 2 -p team-a
```

{% endtab %}

{% tab title="API" %}
Copy the following command to your terminal. Make sure to update the following parameters. For more details, see [Trainings](https://run-ai-docs.nvidia.com/api/workloads/trainings) API.

```bash
curl --location 'https://<COMPANY-URL>/api/v1/workloads/trainings' \
--header 'Content-Type: application/json' \
--header 'Authorization: Bearer <TOKEN>' \ 
--data '{
  "name": "a2",
  "projectId": "<PROJECT-ID>", 
  "clusterId": "<CLUSTER-UUID>",
  "spec": {
    "image":"runai.jfrog.io/demo/quickstart",
    "compute": {
      "gpuDevicesRequest": 2
    }
  }
}'
```

* `<COMPANY-URL>` - The link to the NVIDIA Run:ai user interface
* `<TOKEN>` - The API access token obtained in [Step 1](#a13adq7eth7w)
* `<PROJECT-ID>` - The ID of the Project the workload is running on. You can get the Project ID via the [Get Projects](https://run-ai-docs.nvidia.com/api/organizations/projects#get-api-v1-org-unit-projects) API.
* `<CLUSTER-UUID>` - The unique identifier of the Cluster. You can get the Cluster UUID via the [Get Clusters](https://run-ai-docs.nvidia.com/api/organizations/clusters#get-api-v1-clusters) API.

{% hint style="info" %}
**Note**

The above API snippet runs with NVIDIA Run:ai clusters of 2.18 and above only.
{% endhint %}
{% endtab %}
{% endtabs %}

## Step 4: Submitting the First Training Workload (team-b) <a href="#i3c9jpfzerlq" id="i3c9jpfzerlq"></a>

{% tabs %}
{% tab title="UI - Flexible" %}

1. Go to the Workload Manager → Workloads
2. Click **+NEW WORKLOAD** and select **Training**
3. Select the **cluster** where the previous training was created
4. Select the **project** named team-b
5. Under **Workload architecture**, select **Standard**
6. Select **Start from scratch** to launch a new training quickly
7. Enter **b1** as the workload **name**
8. Click **CONTINUE**

   In the next step:
9. Under **Environment**, enter the **Image URL** - `runai.jfrog.io/demo/quickstart`
10. Under **Compute resources**, click the **load** icon. A side pane appears, displaying a list of available compute resources. Select the **‘one-gpu’** compute resource for your workload.
    * If ‘one-gpu’ is not displayed, follow the below steps to create a one-time compute resource configuration:
      * Set **GPU devices** per pod - 1
      * Optional: set the **CPU compute** per pod - 0.1 cores (default)
      * Optional: set the **CPU memory** per pod - 100 MB (default)
11. Click **CREATE TRAINING**
    {% endtab %}

{% tab title="UI - Original" %}

1. Go to the Workload Manager → Workloads
2. Click **+NEW WORKLOAD** and select **Training**
3. Select the **cluster** where the previous training was created
4. Select the **project** named team-b
5. Under **Workload architecture**, select **Standard**
6. Select **Start from scratch** to launch a new training quickly
7. Enter **b1** as the workload **name**
8. Click **CONTINUE**\
   In the next step:
9. Create a new environment:

   * Click **+NEW ENVIRONMENT**
   * Enter quick-start as the **name** for the environment. The name must be unique.
   * Enter the **Image URL** - `runai.jfrog.io/demo/quickstart`
   * Click **CREATE ENVIRONMENT**

   The newly created environment will be selected automatically
10. Select the **‘one-gpu’** compute resource for your workload

    * If ‘one-gpu’ is not displayed in the gallery, follow the below steps:
      * Click **+NEW COMPUTE RESOURCE**
      * Enter one-gpu as the **name** for the compute resource. The name must be unique.
      * Set **GPU devices** per pod - 1
      * Optional: set the **CPU compute** per pod - 0.1 cores (default)
      * Optional: set the **CPU memory** per pod - 100 MB (default)
      * Click **CREATE COMPUTE RESOURCE**

    The newly created compute resource will be selected automatically
11. Click **CREATE TRAINING**
    {% endtab %}

{% tab title="CLI v2" %}
Copy the following command to your terminal. For more details, see [CLI reference](/saas/reference/cli/runai):

```sh
runai training submit b1 -i runai.jfrog.io/demo/quickstart -g 1 -p team-b
```

{% endtab %}

{% tab title="API" %}
Copy the following command to your terminal. Make sure to update the following parameters. For more details, see [Trainings](https://run-ai-docs.nvidia.com/api/workloads/trainings) API.

```bash
curl --location 'https://<COMPANY-URL>/api/v1/workloads/trainings' \
--header 'Content-Type: application/json' \
--header 'Authorization: Bearer <TOKEN>' \
--data '{
  "name": "b1",
  "projectId": "<PROJECT-ID>", 
  "clusterId": "<CLUSTER-UUID>",
  "spec": {
    "image":"runai.jfrog.io/demo/quickstart",
    "compute": {
      "gpuDevicesRequest": 1
    }
  }
}'
```

* `<COMPANY-URL>` - The link to the NVIDIA Run:ai user interface
* `<TOKEN>` - The API access token obtained in [Step 1](#a13adq7eth7w)
* `<PROJECT-ID>` - The ID of the Project the workload is running on. You can get the Project ID via the [Get Projects](https://run-ai-docs.nvidia.com/api/organizations/projects#get-api-v1-org-unit-projects) API.
* `<CLUSTER-UUID>` - The unique identifier of the Cluster. You can get the Cluster UUID via the [Get Clusters](https://run-ai-docs.nvidia.com/api/organizations/clusters#get-api-v1-clusters) API.

{% hint style="info" %}
**Note**

The above API snippet runs with NVIDIA Run:ai clusters of 2.18 and above only.
{% endhint %}
{% endtab %}
{% endtabs %}

### Over Quota Status

{% tabs %}
{% tab title="UI" %}
System status after run:

![](/files/JG9zOY5v96W4x5Y73iRu)
{% endtab %}

{% tab title="CLI v2" %}
System status after run:

```sh
~ runai workload list -A
Workload  Type      Status   Project  Running/Req.Pods  GPU Alloc.
────────────────────────────────────────────────────────────────────────────
a2       Training   Running   team-a        1/1           2.00
b1       Training   Running   team-b        1/1           1.00
a1       Training.  Running   team-a        0/1           1.00
```

{% endtab %}

{% tab title="API" %}
System status after run:

```
curl --location 'https://<COMPANY-URL>/api/v1/workloads' \
--header 'Content-Type: application/json' \
--header 'Authorization: Bearer <TOKEN>' \ #<TOKEN> is the API access token obtained in Step 1.
--data ''
```

{% endtab %}
{% endtabs %}

## Step 5: Submitting the Second Training Workload (team-b) <a href="#i3c9jpfzerlq" id="i3c9jpfzerlq"></a>

{% tabs %}
{% tab title="UI - Flexible" %}

1. Go to the Workload Manager → Workloads
2. Click **+NEW WORKLOAD** and select **Training**
3. Select the **cluster** where the previous training was created
4. Select the **project** named team-b
5. Under **Workload architecture**, select **Standard**
6. Select **Start from scratch** to launch a new training quickly
7. Enter **b2** as the workload **name**
8. Click **CONTINUE**

   In the next step:
9. Under **Environment**, enter the **Image URL** - `runai.jfrog.io/demo/quickstart`
10. Under **Compute resources**, click the **load** icon. A side pane appears, displaying a list of available compute resources. Select the **‘one-gpu’** compute resource for your workload.
    * If ‘one-gpu’ is not displayed, follow the below steps to create a one-time compute resource configuration:
      * Set **GPU devices** per pod - 1
      * Optional: set the **CPU compute** per pod - 0.1 cores (default)
      * Optional: set the **CPU memory** per pod - 100 MB (default)
11. Click **CREATE TRAINING**
    {% endtab %}

{% tab title="UI - Original" %}

1. Go to the Workload Manager → Workloads
2. Click **+NEW WORKLOAD** and select **Training**
3. Select the **cluster** where the previous training was created
4. Select the **project** named team-b
5. Under **Workload architecture**, select **Standard**
6. Select **Start from scratch** to launch a new training quickly
7. Enter **b2** as the workload **name**
8. Click **CONTINUE**\
   In the next step:
9. Select the environment created in [Step 4](#i3c9jpfzerlq-2)
10. Select the compute resource created in [Step 4](#i3c9jpfzerlq-2)
11. Click **CREATE TRAINING**
    {% endtab %}

{% tab title="CLI v2" %}
Copy the following command to your terminal. For more details, see [CLI reference](/saas/reference/cli/runai):

```sh
runai training submit b2 -i runai.jfrog.io/demo/quickstart -g 1 -p team-b
```

{% endtab %}

{% tab title="API" %}
Copy the following command to your terminal. Make sure to update the following parameters. For more details, see [Trainings](https://run-ai-docs.nvidia.com/api/workloads/trainings) API.

```bash
curl --location 'https://<COMPANY-URL>/api/v1/workloads/trainings' \
--header 'Content-Type: application/json' \
--header 'Authorization: Bearer <TOKEN>' \ 
--data '{
  "name": "b2",
  "projectId": "<PROJECT-ID>", 
  "clusterId": "<CLUSTER-UUID>",
  "spec": {
    "image":"runai.jfrog.io/demo/quickstart",
    "compute": {
      "gpuDevicesRequest": 1
    }
  }
}'
```

* `<COMPANY-URL>` - The link to the NVIDIA Run:ai user interface
* `<TOKEN>` - The API access token obtained in [Step 1](#a13adq7eth7w)
* `<PROJECT-ID>` - The ID of the Project the workload is running on. You can get the Project ID via the [Get Projects](https://run-ai-docs.nvidia.com/api/organizations/projects#get-api-v1-org-unit-projects) API.
* `<CLUSTER-UUID>` - The unique identifier of the Cluster. You can get the Cluster UUID via the [Get Clusters](https://run-ai-docs.nvidia.com/api/organizations/clusters#get-api-v1-clusters) API.

{% hint style="info" %}
**Note**

The above API snippet runs with NVIDIA Run:ai clusters of 2.18 and above only.
{% endhint %}
{% endtab %}
{% endtabs %}

### Basic Fairness and Preemption Status

{% tabs %}
{% tab title="UI" %}
Workloads status after run:

![](/files/AdxXiABNvMsGshh8qQwU)
{% endtab %}

{% tab title="CLI v2" %}
Workloads status after run:

```sh
~ runai workload list -A
Workload  Type      Status   Project  Running/Req.Pods  GPU Alloc.
────────────────────────────────────────────────────────────────────────────
a2       Training   Running   team-a        1/1           2.00
b1       Training   Running   team-b        1/1           1.00
b2       Training   Running   team-b        1/1           1.00
a1       Training.  Pending   team-a        0/1           1.00
```

{% endtab %}

{% tab title="API" %}
Workloads status after run:

```bash
curl --location 'https://<COMPANY-URL>/api/v1/workloads' \
--header 'Content-Type: application/json' \
--header 'Authorization: Bearer <TOKEN>' \ #<TOKEN> is the API access token obtained in Step 1.
--data ''
```

{% endtab %}
{% endtabs %}

## Next Steps

Manage and monitor your newly created workload using the [Workloads](/saas/workloads-in-nvidia-run-ai/workloads) table.


# Resource Optimization


# GPU Fractions

To submit a [workload](/saas/workloads-in-nvidia-run-ai/workloads) with GPU resources in Kubernetes, you typically need to specify an integer number of GPUs. However, workloads often require diverse GPU memory and compute requirements or even use GPUs intermittently depending on the application (such as inference workloads, training workloads or notebooks at the model-creation phase). Additionally, GPUs are becoming increasingly powerful, offering more processing power and larger memory capacity for applications. Despite the increasing model sizes, the increasing capabilities of GPUs allow them to be effectively shared among multiple users or applications.

NVIDIA Run:ai’s GPU fractions provide an agile and easy-to-use method to share a GPU or multiple GPUs across workloads. With GPU fractions, you can divide the GPU/s memory into smaller chunks and share the GPU/s compute resources between different workloads and users, resulting in higher GPU utilization and more efficient resource allocation.

## Benefits of GPU Fractions

Utilizing GPU fractions to share GPU resources among multiple workloads provides numerous advantages for both platform administrators and practitioners, including improved efficiency, resource optimization, and enhanced user experience.

* For the AI practitioner:
  * **Reduced wait time** - Workloads with smaller GPU requests are more likely to be scheduled quickly, minimizing delays in accessing resources.
  * **Increased workload capacity** - More workloads can be run using the same admin-defined GPU [quota](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles#quota) and available unused resources - [over quota](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles#over-quota).
* For the platform administrator:
  * **Improved GPU utilization** - Sharing GPUs across workloads increases the utilization of individual GPUs, resulting in better overall platform efficiency.
  * **Higher resource availability** - More users gain access to GPU resources, ensuring better distribution.
  * **Enhanced workload throughput** - More workloads can be served per GPU, ensuring maximum output from existing hardware.
  * **Optimized scheduling** - Smaller and dynamic resource allocations gives the [Scheduler](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles) a higher chance of finding GPU resources for incoming workloads.

## Quota Planning with GPU Fractions

When planning the quota distribution for your [projects](/saas/platform-management/aiinitiatives/organization/projects) and [departments](/saas/platform-management/aiinitiatives/organization/departments), using fractions gives the platform administrator the ability to allocate more precise quota per project and department, assuming the usage of GPU fractions or enforcing it with [pre-defined policies](/saas/platform-management/policies/native-workload-policies/policy-yaml-reference) or [compute resource](/saas/workloads-in-nvidia-run-ai/assets/compute-resources) templates.

For example, in an organization with a department budgeted for **two nodes of 8×H100 GPUs** and a team of 32 researchers:

* Allocating 0.5 GPU per researcher ensures all researchers have access to GPU resources.
* Using fractions enables researchers to run smaller workloads intermittently within their quota or go over their quota by using temporary over quota resources with higher resource demanding workloads.
* Using GPUs for notebook-based model development, where GPUs are not continuously active and can be shared among multiple users.

For more details on mapping your organization and resources, see [Adapting AI initiatives to your organization](/saas/platform-management/aiinitiatives/adapting-ai-initiatives).

## How GPU Fractions Work

When a workload is submitted, the Scheduler finds a node with a GPU that can satisfy the requested GPU portion or GPU memory, then it schedules the pod to that node. The NVIDIA Run:ai GPU fractions logic, running locally on each NVIDIA Run:ai worker node, allocates the requested memory size on the selected GPU. **Each pod uses its own separate virtual memory address space.** NVIDIA Run:ai’s GPU fractions logic enforces the requested memory size, so no workload can use more than requested, and no workload can run over another workload’s memory. This gives users the experience of a ‘logical GPU’ per workload.

While [MIG](/saas/platform-management/aiinitiatives/resources/mig-profiles) requires administrative work to configure every MIG slice, where a slice is a fixed chunk of memory, GPU fractions allow dynamic and fully flexible allocation of GPU memory chunks. By default, GPU fractions use NVIDIA’s time-slicing to share the GPU compute runtime. You can also use the [NVIDIA Run:ai GPU time-slicing](/saas/platform-management/runai-scheduler/resource-optimization/time-slicing) which allows dynamic and fully flexible splitting of the GPU compute time.

NVIDIA Run:ai GPU fractions are agile and dynamic allowing a user to allocate and free GPU fractions during the runtime of the system, at any size between zero to the maximum GPU portion (100%) or memory size (up to the maximum memory size of a GPU).

The NVIDIA Run:ai Scheduler can work alongside other schedulers. In order to avoid collisions with other schedulers, the NVIDIA Run:ai Scheduler creates special reservation pods. Once a workload is submitted requesting a fraction of a GPU, NVIDIA Run:ai will create a pod in a dedicated runai-reservation namespace with the full GPU as a resource, allowing other schedulers to understand that the GPU is reserved.

{% hint style="info" %}
**Note**

* Splitting a GPU into fractions may generate some fragmentation of the GPU memory. The [Scheduler](/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles) will try to consolidate GPU resources where feasible (i.e. preemptible workloads).
* Using [bin-pack](/saas/platform-management/aiinitiatives/resources/node-pools#adding-a-new-node-pool) as a scheduling placement strategy can also reduce GPU fragmentation.
* Using [dynamic GPU fractions ](/saas/platform-management/runai-scheduler/resource-optimization/dynamic-fractions)ensures that even small unused fragments of GPU memory are utilized by workloads.
  {% endhint %}

## Multi-GPU Fractions

NVIDIA Run:ai also supports workload submission using multi-GPU fractions. Multi-GPU fractions work similarly to single-GPU fractions, however, the NVIDIA Run:ai Scheduler allocates the same fraction size on multiple GPU devices within the same node. For example, if practitioners develop a new model that uses 8 GPUs and requires 40GB of memory per GPU, they can allocate 8×40GB with multi-GPU fractions instead of reserving the full memory of each GPU (e.g. 80GB). This leaves 40GB of GPU memory available on each of the 8 GPUs for other workloads within that node.

Time sharing where single GPUs can serve multiple workloads with fractions remains unchanged, only now, it serves multiple workloads using multi-GPUs per workload, single-GPU per workload, or a mix of both.

## Deployment Considerations

* Selecting a GPU portion using percentages as units does not guarantee the exact memory size. This means 50% of an A-100-40GB is 20GB while 50% of an A-100-80 is 40GB. To have better control over the exact allocated memory, specify the exact memory size, i.e. 40GB.
* Using NVIDIA Run:ai GPU fractions controls the memory split (i.e. 0.5 GPU means 50% of the GPU memory) but not the compute (processing time). To split the compute time, see [NVIDIA Run:ai’s GPU time slicing](/saas/platform-management/runai-scheduler/resource-optimization/time-slicing).
* NVIDIA Run:ai GPU fractions and [MIG mode](/saas/platform-management/aiinitiatives/resources/mig-profiles) cannot be used on the same node.
* Some users restrict the use of Kubernetes containers with direct `hostPath` mounts due to stricter security policies and best practices. NVIDIA Run:ai offers an alternative for configuring fractions without relying on `hostPath`. Instead, you can enable device plugin–based host mounts by setting the `clusterConfig.global.devicePluginBindings` parameter to `true` (default is `false`, and uses the standard `hostPath` mount method). For details on how to configure this value using Helm or `runaiconfig`, see [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config).

## Setting GPU Fractions

Using the [compute resources](/saas/workloads-in-nvidia-run-ai/assets/compute-resources) asset, you can define the compute requirements by specifying your requested GPU portion or GPU memory, and use it with any of the [NVIDIA Run:ai workload types](/saas/workloads-in-nvidia-run-ai/workload-types) for single GPU and multi-GPU fractions.

* **Single-GPU fractions** - Define the compute requirement to run 1 GPU device, by specifying either a fraction (percentage) of the overall memory or a memory request (GB, MB).
* **Multi-GPU fractions** - Define the compute requirement to run multiple GPU devices, by specifying either a fraction (percentage) of the overall memory or a memory request (GB, MB).

## Setting GPU Fractions via YAML

To enable GPU fractions for workloads submitted via Kubernetes YAML, use the following annotations to define the GPU fraction configuration. You can configure either `gpu-fraction` or `gpu-memory`.

GPU fractions can be assigned to only one container in the pod. By default, GPU fraction allocation is applied to the first container (index 0) in the pod. Set the `gpu-fraction-container-name` annotation to to specify which container should consume the fractional GPU resources. The specified container can be the main container, a sidecar, or an init container, any container other than the default first container.

{% hint style="info" %}
**Note**

Make sure the [default scheduler](/saas/platform-management/runai-scheduler/scheduling/default-scheduler) is set to `runai-scheduler`.
{% endhint %}

<table><thead><tr><th width="207.12890625">Variable</th><th>Input Format</th><th>Where to Set</th></tr></thead><tbody><tr><td><code>gpu-fraction</code></td><td>A portion of GPU memory as a double-precision floating-point number. Example: <code>0.25</code>, <code>0.75</code>.</td><td>Pod annotation (<code>metadata.annotations</code>)</td></tr><tr><td><code>gpu-memory</code></td><td>Memory size in MiB. Example: <code>2500,</code> <code>4096</code>. The <code>gpu-memory</code> values are always in MiB.</td><td>Pod annotation (<code>metadata.annotations</code>)</td></tr><tr><td><code>gpu-fraction-num-devices</code></td><td>The number of GPU devices to allocate using the specified <code>gpu-fraction</code> or <code>gpu-memory</code> value. Set this annotation only if you want to request multiple GPU devices.</td><td>Pod annotation (<code>metadata.annotations</code>)</td></tr><tr><td><code>gpu-fraction-container-name</code></td><td>By default, GPU fraction allocation is applied to the first container (index 0) in the pod. Set this annotation to specify a different container to receive the GPU allocation.</td><td>Pod annotation (<code>metadata.annotations</code>)</td></tr></tbody></table>

The following example YAML creates a pod that requests 2 GPU devices, each requesting 50% of memory (`gpu-fraction: "0.5"`) .

```yaml
apiVersion: v1
kind: Pod
metadata:
  annotations:
    user: test
    gpu-fraction: "0.5"
    gpu-fraction-num-devices: "2"
    # Specify which container should receive the GPU fraction allocation
    # By default, the first container (index 0) receives the GPU allocation
    # Use this annotation to specify a different container by name
    gpu-fraction-container-name: "gpu-workloads"
  labels:
    runai/queue: test
  name: multi-fractional-pod-job
  namespace: test
spec:
  containers:
  - image: gcr.io/run-ai-demo/quickstart-cuda
    imagePullPolicy: Always
    name: job
    env:
    - name: RUNAI_VERBOSE
      value: "1"
    resources:
      limits:
        cpu: 200m
        memory: 200Mi
      requests:
        cpu: 100m
        memory: 100Mi
    securityContext:
      capabilities:
        drop: ["ALL"]
  schedulerName: runai-scheduler
  serviceAccount: default
  serviceAccountName: default
  terminationGracePeriodSeconds: 5
```

## Using CLI

To view the available actions, go to the [CLI v2 reference](/saas/reference/cli/runai) and run according to your workload.

## Using API

To view the available actions, go to the [API reference](https://run-ai-docs.nvidia.com/api/) and run according to your workload.


# Dynamic GPU Fractions

Many workloads utilize GPU resources intermittently, with long periods of inactivity. These workloads typically need GPU resources when they are running AI applications or debugging a model in development. Other workloads such as inference may utilize GPUs at lower rates than requested, but may demand higher resource usage during peak utilization. The disparity between resource request and actual resource utilization often leads to inefficient utilization of GPUs. This usually occurs when multiple workloads request resources based on their peak demand, despite operating below those peaks for the majority of their runtime.

To address this challenge, NVIDIA Run:ai has introduced dynamic GPU fractions. This feature optimizes GPU utilization by enabling workloads to dynamically adjust their resource usage. It allows users to specify a guaranteed fraction of GPU memory and compute resources with a higher limit that can be dynamically utilized when additional resources are requested.

## How Dynamic GPU Fractions Work

With dynamic GPU fractions, users can [submit workloads](/saas/workloads-in-nvidia-run-ai/workloads) using GPU fraction Request and Limit which is achieved by leveraging the Kubernetes Request and Limit notations. You can either:

* Request a GPU fraction (portion) using a percentage of a GPU and specify a Limit
* Request a GPU memory size (GB, MB) and specify a Limit

When setting a GPU memory limit either as GPU fraction or GPU memory size, the Limit must be equal to or greater than the GPU fractional memory request. Both GPU fraction and GPU memory are translated into the actual requested memory size of the Request (guaranteed resources) and the Limit (burstable resources - non guaranteed).

For example, a user can specify a workload with a GPU fraction request of 0.25 GPU, and add a limit of up to 0.80 GPU. The NVIDIA Run:ai [Scheduler](/saas/platform-management/runai-scheduler/scheduling/how-the-scheduler-works) schedules the workload to a node that can provide the GPU fraction request (0.25), and then assigns the workload to a GPU. The GPU scheduler monitors the workload and allows it to occupy memory between 0 to 0.80 of the GPU memory (based on the Limit), where only 0.25 of the GPU memory is guaranteed to that workload. The rest of the memory (from 0.25 to 0.8) is “loaned” to the workload, as long as it is not needed by other workloads.

NVIDIA Run:ai automatically manages the state changes between Request and Limit as well as the reverse (when the balance needs to be "returned"), updating the workloads’ utilization vs. Request and Limit parameters in the [metrics pane for each workload](/saas/workloads-in-nvidia-run-ai/workloads).

To guarantee fair quality of service between different workloads using the same GPU, NVIDIA Run:ai developed an extendable GPUOOMKiller (Out Of Memory Killer) component that guarantees the quality of service using Kubernetes semantics for resources of Request and Limit.

The OOMKiller capability requires adding CAP\_KILL capabilities to the dynamic GPU fractions and to the NVIDIA Run:ai core scheduling module (toolkit daemon). This capability is enabled by default.

{% hint style="info" %}
**Note**

Dynamic GPU fractions is enabled by default in the [cluster](/saas/infrastructure-setup/advanced-setup/cluster-config). Disabling dynamic GPU fractions removes the CAP\_KILL capability.
{% endhint %}

## Multi-GPU Dynamic Fractions

NVIDIA Run:ai also supports workload submission using multi-GPU dynamic fractions. Multi-GPU dynamic fractions work similarly to dynamic fractions on a single GPU workload, however, instead of a single GPU device, the NVIDIA Run:ai Scheduler allocates the same dynamic fraction pair (Request and Limit) on multiple GPU devices within the same node. For example, if practitioners develop a new model that uses 8 GPUs and requires 40GB of memory per GPU, but may want to burst out and consume up to the full GPU memory, they can allocate 8×40GB with multi-GPU fractions and a limit of 80GB (e.g. H100 GPU) instead of reserving the full memory of each GPU (e.g. 80GB). This leaves 40GB of GPU memory available on each of the 8 GPUs for other workloads within that node.This is useful during model development, where memory requirements are usually lower due to experimentation with smaller models or configurations.

This approach significantly improves GPU utilization and availability, enabling more precise and often smaller quota requirements for the end user. Time sharing where single GPUs can serve multiple workloads with dynamic fractions remains unchanged, only now, it serves multiple workloads using multi-GPUs per workload.

## Deployment Considerations

Some users restrict the use of Kubernetes containers with direct `hostPath` mounts due to stricter security policies and best practices. NVIDIA Run:ai offers an alternative for configuring fractions without relying on `hostPath`. Instead, you can enable device plugin–based host mounts by setting the `clusterConfig.global.devicePluginBindings` parameter to `true` (default is `false`, and uses the standard `hostPath` mount method). For details on how to configure this value using Helm or `runaiconfig`, see [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config).

## Setting Dynamic GPU Fractions

{% hint style="info" %}
**Note**

Dynamic GPU fractions is disabled by default in the NVIDIA Run:ai UI. To use dynamic GPU fractions, it must be enabled by your Administrator, under **General Settings** → Resources → GPU resource optimization.
{% endhint %}

Using the [compute resources](/saas/workloads-in-nvidia-run-ai/assets/compute-resources) asset, you can define the compute requirements by specifying your requested GPU portion or GPU memory, and set a Limit. You can then use the compute resource with any of the [NVIDIA Run:ai workload types](/saas/workloads-in-nvidia-run-ai/workload-types) for single and multi-GPU dynamic fractions. In addition, you will be able to view the workloads’ utilization vs. Request and Limit parameters in the [metrics pane for each workload](/saas/workloads-in-nvidia-run-ai/workloads).

* **Single dynamic GPU fractions** - Define the compute requirement to run 1 GPU device, by specifying either a fraction (percentage) of the overall memory or specifying the memory request (GB, MB) with a Limit. The limit must be equal to or greater than the GPU fractional memory request.
* **Multi-GPU dynamic fractions** - Define the compute requirement to run multiple GPU devices, by specifying either a fraction (percentage) of the overall memory or specifying the memory request (GB, MB) with a Limit. The limit must be equal to or greater than the GPU fractional memory request.

{% hint style="info" %}
**Note**

When setting a workload with dynamic GPU fractions, (for example, when using it with GPU Request or GPU memory Limits), you practically make the workload burstable. This means it can use memory that is not guaranteed for that workload and is susceptible to an ‘OOM Kill’ signal if the actual owner of that memory requires it back. This applies to non-preemptible workloads as well. For that reason, it is recommended that you use dynamic GPU fractions with Interactive workloads running Notebooks. Notebook pods are not evicted when their GPU process is OOM Kill’ed. This behavior is the same as standard Kubernetes burstable CPU workloads.
{% endhint %}

## Setting Dynamic GPU Fractions via YAML

To enable dynamic GPU fractions for workloads submitted via Kubernetes YAML, use the following annotations to define the GPU fraction configuration. You can configure either `gpu-fraction` or `gpu-memory`. You must also set the `RUNAI_GPU_MEMORY_LIMIT` environment variable in the container to enforce the memory limit.

GPU fractions can be assigned to only one container in the pod. By default, GPU fraction allocation is applied to the first container (index 0) in the pod. Set the `gpu-fraction-container-name` annotation to to specify which container should consume the fractional GPU resources. The specified container can be the main container, a sidecar, or an init container, any container other than the default first container.

{% hint style="info" %}
**Note**

Make sure the [default scheduler](/saas/platform-management/runai-scheduler/scheduling/default-scheduler) is set to `runai-scheduler`.
{% endhint %}

<table><thead><tr><th width="207.12890625">Variable</th><th>Input Format</th><th>Where to Set</th></tr></thead><tbody><tr><td><code>gpu-fraction</code></td><td>A portion of GPU memory as a double-precision floating-point number. Example: <code>0.25</code>, <code>0.75</code>.</td><td>Pod annotation (<code>metadata.annotations</code>)</td></tr><tr><td><code>gpu-memory</code></td><td>Memory size in MiB. Example: <code>2500,</code> <code>4096</code>. The <code>gpu-memory</code> values are always in MiB.</td><td>Pod annotation (<code>metadata.annotations</code>)</td></tr><tr><td><code>gpu-fraction-num-devices</code></td><td>The number of GPU devices to allocate using the specified <code>gpu-fraction</code> or <code>gpu-memory</code> value. Set this annotation only if you want to request multiple GPU devices.</td><td>Pod annotation (<code>metadata.annotations</code>)</td></tr><tr><td><code>gpu-fraction-container-name</code></td><td>By default, GPU fraction allocation is applied to the first container (index 0) in the pod. Set this annotation to specify a different container to receive the GPU allocation.</td><td>Pod annotation (<code>metadata.annotations</code>)</td></tr><tr><td><code>RUNAI_GPU_MEMORY_LIMIT</code></td><td><ul><li>To use for <code>gpu-fraction</code> - Specify a double-precision floating-point number. Example: <code>0.95</code></li><li>To use for <code>gpu-memory</code> - Specify a Kubernetes resource quantity format. Example: <code>500000000</code>, <code>2500M</code></li></ul><p>The limit must be equal to or greater than the GPU fractional memory request.</p></td><td>Environment variable in the container</td></tr></tbody></table>

The following example YAML creates a pod that requests 2 GPU devices, each requesting 50% of memory (`gpu-fraction: "0.5"`) and allows usage of up to 95% (`RUNAI_GPU_MEMORY_LIMIT: "0.95"`) if available.

<pre class="language-yaml"><code class="lang-yaml">apiVersion: v1
<strong>kind: Pod
</strong>metadata:
  annotations:
    user: test
    gpu-fraction: "0.5"
    gpu-fraction-num-devices: "2"
    # Specify which container should receive the GPU fraction allocation
    # By default, the first container (index 0) receives the GPU allocation
    # Use this annotation to specify a different container by name
    gpu-fraction-container-name: "gpu-workloads"
  labels:
    runai/queue: test
  name: multi-fractional-pod-job
  namespace: test
spec:
  containers:
  - image: gcr.io/run-ai-demo/quickstart-cuda
    imagePullPolicy: Always
    name: job
    env:
    - name: RUNAI_VERBOSE
      value: "1"
    - name: RUNAI_GPU_MEMORY_LIMIT
      value: "0.95"
<strong>    resources:
</strong>      limits:
        cpu: 200m
        memory: 200Mi
      requests:
        cpu: 100m
        memory: 100Mi
    securityContext:
      capabilities:
        drop: ["ALL"]
  schedulerName: runai-scheduler
  serviceAccount: default
  serviceAccountName: default
  terminationGracePeriodSeconds: 5
</code></pre>

## Using CLI

To view the available actions, go to the [CLI v2 reference](/saas/reference/cli/runai) and run according to your workload.

## Using API

To view the available actions, go to the [API reference](https://run-ai-docs.nvidia.com/api/) and run according to your workload.


# Optimize Performance with Node Level Scheduler

{% hint style="info" %}
**Note**

Node Level Scheduler has been deprecated and will be removed in a future release.
{% endhint %}

The Node Level Scheduler optimizes the performance of your pods and maximizes the utilization of GPUs by making optimal local decisions on GPU allocation to your pods. While the [NVIDIA Run:ai Scheduler](/saas/platform-management/runai-scheduler/scheduling/how-the-scheduler-works) chooses the specific node for a pod, it has no visibility to the node’s GPUs' internal state. The Node Level Scheduler is aware of the local GPUs' states and makes optimal local decisions such that it can optimize both the GPU utilization and pods’ performance running on the node’s GPUs.

This guide provides an overview of the best use cases for the Node Level Scheduler and instructions for configuring it to maximize GPU performance and pod efficiency.

## Deployment Considerations

* While the Node Level Scheduler applies to all [workload types](/saas/workloads-in-nvidia-run-ai/workload-types), it will best optimize the performance of burstable workloads. Burstable workloads are workloads that use [dynamic GPU fractions](/saas/platform-management/runai-scheduler/resource-optimization/dynamic-fractions), giving those more GPU memory than requested and up to the Limit specified.
* Burstable workloads are always susceptible to an OOM Kill signal if the owner of the excess memory requires it back. This means that using the Node Level Scheduler with inference or training workloads may cause pod preemption.
* Using interactive workloads with notebooks is the best use case for burstable workloads and Node Level Scheduler. These workloads behave differently since the OOM Kill signal will cause the notebooks' GPU process to exit but not the notebook itself. This keeps the interactive pod running and retrying to attach a GPU again.

## Interactive Notebooks Use Case

This use case is one scenario that shows how Node Level Scheduler locally optimizes and maximizes GPU utilization and workspaces’ performance.

1. The below shows a node with 2 GPUs and 2 submitted workspaces:

![Unallocated GPU nodes](/files/OfJ5l0dU39NAxOwW5n7E)

2. The Scheduler instructs the node to put the 2 workspaces on a single GPU, [bin-packing](/saas/platform-management/aiinitiatives/resources/node-pools#adding-a-new-node-pool) a single GPU and leaving the other free for a workload that requires resources. This means GPU#2 is idle while the two workspaces can only use up to half a GPU, even if they temporarily need more:

![Single allocated GPU node](/files/jjzPPVjlYj6EYTYpDiCZ)

3. With the Node Level Scheduler enabled, the local decision will be to spread those 2 workspaces on 2 GPUs and allow them to maximize both workspaces’ performance and GPUs’ utilization by bursting out up to the full GPU memory and GPU compute resources:

![Two allocated GPU nodes](/files/qPrqbBnpPA4Danjr2BcG)

4. The NVIDIA Run:ai Scheduler still sees a node with one fully empty GPU and one fully occupied GPU. When a 3rd workload is scheduled, and it requires a full GPU (or more than 0.5 GPU), the Scheduler will schedule it to that node, and the Node Level Scheduler will move one of the workspaces to run with the other in GPU#1, as was the Scheduler’s initial plan. Moving the workspace from GPU#1 back to GPU#2 maintains the workspace running while the GPU process within the Jupyter notebook is killed and re-established on GPU#2, continuing to serve the workspace:

![Node Level Scheduler locally optimized GPU nodes](/files/BtQwV1RQw2HRSx2WFSxB)

## Using Node Level Scheduler

The Node Level Scheduler can be enabled per node pool. To use Node Level Scheduler, follow the below steps.

### Enable on Your Cluster

Enable the Node Level Scheduler at the cluster level (per cluster):

1. **Using Helm** - Set the following value in your `values.yaml` file under `clusterConfig` and upgrade the chart:

   <pre class="language-yaml"><code class="lang-yaml">clusterConfig: 
   <strong>  global: 
   </strong>      core: 
           nodeScheduler:
             enabled: true
   </code></pre>
2. **Using runaiconfig at runtime** - Use the following `kubectl` patch command:

   ```bash
   kubectl patch -n runai runaiconfigs.run.ai/runai --type='merge' --patch '{"spec":{"global":{"core":{"nodeScheduler":{"enabled": true}}}}}'
   ```

See [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config) for more details.

### Enable on a Node Pool

{% hint style="info" %}
**Note**

GPU resource optimization is disabled by default. It must be enabled by your Administrator, under **General Settings** → Resources → GPU resource optimization.
{% endhint %}

Enable Node Level Scheduler on any of the node pools:

1. Select Resources → Node pools
2. [Create a new node pool](/saas/platform-management/aiinitiatives/resources/node-pools#adding-a-new-node-pool) or [edit an existing node pool](/saas/platform-management/aiinitiatives/resources/node-pools#editing-a-node-pool)
3. Under the **Resource Utilization Optimization** tab, change the **number of workloads on each GPU** to any value other than **Not Enforced** (i.e. 2, 3, 4, 5)

The Node Level Scheduler is now ready to be used on that node pool.

### Submit a Workload

In order for a workload to be considered by the Node Level Scheduler for rerouting, it must be submitted with a GPU Request and Limit where the Limit is larger than the Request:

* Enable and set [dynamic GPU fractions](/saas/platform-management/runai-scheduler/resource-optimization/dynamic-fractions)
* Then [submit a workload](/saas/workloads-in-nvidia-run-ai/workloads) using dynamic GPU fractions


# GPU Time-Slicing

NVIDIA Run:ai supports simultaneous submission of multiple workloads to single or multi-GPUs when using [GPU fractions](/saas/platform-management/runai-scheduler/resource-optimization/fractions). This is achieved by slicing the GPU memory between the different workloads according to the requested GPU fraction, and by using NVIDIA’s GPU time-slicing to share the GPU compute runtime. NVIDIA Run:ai ensures each workload receives the exact share of the GPU memory (= gpu\_memory \* requested), while the NVIDIA GPU time-slicing splits the GPU runtime evenly between the different workloads running on that GPU.

To provide customers with predictable and accurate GPU compute resource scheduling, NVIDIA Run:ai’s GPU time-slicing adds **fractional compute** capabilities on top of NVIDIA Run:ai GPU fraction capabilities.

## How GPU Time-Slicing Works

While the default NVIDIA GPU time-slicing allows for sharing the GPU compute runtime evenly without splitting or limiting the runtime of each workload, NVIDIA Run:ai’s GPU time-slicing mechanism gives each workload exclusive access to the full GPU for a **limited** amount of time, **lease time**, in each scheduling cycle, **plan time**. This cycle repeats itself for the lifetime of the workload. Using the GPU runtime this way guarantees a workload is granted its requested GPU compute resources proportionally to its requested GPU fraction, but also allows splitting GPU unused compute time up to a requested Limit.

For example, when there are 2 workloads running on the same GPU, with NVIDIA’s default GPU time slicing, each workload gets 50% of the GPU compute runtime, even if one workload requests 25% of the GPU memory, and the other workload requests 75% of the GPU memory. With the NVIDIA Run:ai GPU time-slicing, the first workload will get 25% of the GPU compute time and the second will get 75%. If one of the workloads does not use its deserved GPU compute time, the others can split that time evenly between them. As shown in the example, if one of the workloads does not request the GPU for some time, the other will get the full GPU compute time.

### GPU Time-Slicing Modes

NVIDIA Run:ai offers two GPU time-slicing modes:

* **Strict** - Each workload gets its **precise** GPU compute fraction, which equals to its requested GPU (memory) fraction. In terms of official Kubernetes resource specification, this means:

```sh
gpu-compute-request = gpu-compute-limit = gpu-(memory-)fraction
```

* **Fair** - Each workload is guaranteed at least its GPU compute fraction, but at the same time can also use additional GPU runtime compute slices that are not used by other idle workloads. Those excess time slices are divided equally between all workloads running on that GPU (after each got at least its requested GPU compute fraction). In terms of official Kubernetes resource specification, this means:

```sh
gpu-compute-request = gpu-(memory-)fraction
gpu-compute-limit = 1.0
```

The figure below illustrates how **Strict** time-slicing mode uses the GPU from Lease (slice) and Plan (cycle) perspective:

![Strict time-slicing mode](/files/IW3TPAZHlMrkA79XqysU)

The figure below illustrates how **Fair** time-slicing mode uses the GPU from Lease (slice) and Plan (cycle) perspective:

![Fair time-slicing mode](/files/Z7xT1Xghc0vS33oR55bg)

## Time-Slicing Plan and Lease Times

Each GPU scheduling cycle is a **plan**. The plan is determined by the lease time and granularity (precision). By default, basic lease time is 250ms with 5% granularity (precision), which means the plan (cycle) time is: 250 / 0.05 = 5000ms (5 Sec). Using these values, a workload that requests gpu-fraction=0.5 gets 2.5s runtime out of the 5s cycle time.

Different workloads require different SLA and precision, so it also possible to tune the lease time and precision for customizing the time-slicing capabilities to your cluster.

{% hint style="info" %}
**Note**

Decreasing the lease time makes time-slicing less accurate. Increasing the lease time makes the system more accurate, but each workload is less responsive.
{% endhint %}

Once timeSlicing is enabled in the [cluster configuration](#enabling-gpu-time-slicing), all submitted GPU fractions or GPU memory workloads will have their gpu-compute-request/limit set automatically by the system, depending on the annotation used on the time-slicing mode:

* Strict compute resources:

| **Annotation** | **Value** | **GPU Compute Request** | **GPU Compute Limit** |
| -------------- | --------- | ----------------------- | --------------------- |
| `gpu-fraction` | x         | x                       | x                     |
| `gpu-memory`   | x         | 0                       | 1.0                   |

* Fair compute resources:

| **Annotation** | **Value** | **GPU Compute Request** | **GPU Compute Limit** |
| -------------- | --------- | ----------------------- | --------------------- |
| `gpu-fraction` | x         | x                       | 1.0                   |
| `gpu-memory`   | x         | 0                       | 1.0                   |

{% hint style="info" %}
**Note**

The above tables show that when submitting a workload using gpu-memory annotation, the system will split the GPU compute time between the different workloads running on that GPU. This means the workload can get anything from very little compute time (>0) to full GPU compute time (1.0).
{% endhint %}

## Enabling GPU Time-Slicing

NVIDIA Run:ai’s GPU time-slicing is a cluster flag which changes the default NVIDIA time-slicing used by GPU fractions. For more details, see [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config).

Enable GPU time-slicing:

1. **Using Helm** - Set the following value in your `values.yaml` file under `clusterConfig` and upgrade the chart:

   ```yaml
   clusterConfig: 
     global: 
         core: 
           timeSlicing:
             mode: fair/strict
   ```
2. **Using runaiconfig at runtime** - Use the following `kubectl` patch command::

   ```bash
   kubectl patch -n runai runaiconfigs.run.ai/runai --type='merge' --patch '{"spec":{"global":{"core":{"timeSlicing":{"mode": fair/strict}}}}}'
   ```

If the `timeSlicing` flag is not set, the system continues to use the default NVIDIA GPU time-slicing to maintain backward compatibility.


# GPU Memory Swap

NVIDIA Run:ai’s GPU memory swap helps administrators and AI practitioners to further increase the utilization of their existing GPU hardware by improving GPU sharing between AI initiatives and stakeholders. This is done by expanding the GPU physical memory to the CPU memory, typically an order of magnitude larger than that of the GPU.

Expanding the GPU physical memory helps the NVIDIA Run:ai system to put more workloads on the same GPU physical hardware, and to provide a smooth workload context switching between GPU memory and CPU memory, eliminating the need to kill workloads when the memory requirement is larger than what the GPU physical memory can provide.

## Benefits of GPU Memory Swap

There are several use cases where GPU memory swap can benefit and improve the user experience and the system's overall utilization.

### Sharing a GPU Between Multiple Interactive Workloads (Notebooks)

AI practitioners use notebooks to develop and test new AI models and to improve existing AI models. While developing or testing an AI model, notebooks use GPU resources intermittently, yet, required resources of the GPUs are pre-allocated by the notebook and cannot be used by other workloads after one notebook has already reserved them. To overcome this inefficiency, NVIDIA Run:ai introduced [dynamic GPU fractions](/saas/platform-management/runai-scheduler/resource-optimization/dynamic-fractions).

When one or more workloads require more than their requested GPU resources, there’s a high probability not all workloads can run on a single GPU because the total memory required is larger than the physical size of the GPU memory.

With GPU memory swap, several workloads can run on the same GPU, even if the sum of their used memory is larger than the size of the physical GPU memory. GPU memory swap can swap in and out workloads interchangeably, allowing multiple workloads to each use the full amount of GPU memory. The most common scenario is for one workload to run on the GPU (for example, an interactive notebook), while other notebooks are either idle or using the CPU to develop new code (while not using the GPU). From a user experience point of view, the swap in and out is a smooth process since the notebooks do not notice that they are being swapped in and out of the GPU memory. On rare occasions, when multiple notebooks need to access the GPU simultaneously, slower workload execution may be experienced.

Notebooks typically use the GPU intermittently, therefore with high probability, only one workload (for example, an [interactive notebook](/saas/workloads-in-nvidia-run-ai/workload-types)), will use the GPU at a time. The more notebooks the system puts on a single GPU, the higher the chances are that there will be more than one notebook requiring the GPU resources at the same time. Admins have a significant role here in fine tuning the number of notebooks running on the same GPU, based on specific use patterns and required SLAs.

### Sharing a GPU Between Inference/Interactive Workloads and Training Workloads

A single GPU can be shared between an [interactive or inference workload](/saas/workloads-in-nvidia-run-ai/workload-types) (for example, a Jupyter notebook, image recognition services, or an LLM service), and a training workload that is not time-sensitive or delay-sensitive. At times when the inference/interactive workload uses the GPU, both training and inference/interactive workloads share the GPU resources, each running part of the time swapped-in to the GPU memory, and swapped-out into the CPU memory the rest of the time.

Whenever the inference/interactive workload stops using the GPU, the swap mechanism swaps out the inference/interactive workload GPU data to the CPU memory. Kubernetes wise, the pod is still alive and running using the CPU. This allows the training workload to run faster when the inference/interactive workload is not using the GPU, and slower when it does, thus sharing the same resource between multiple workloads, fully utilizing the GPU at all times, and maintaining uninterrupted service for both workloads.

### Serving Inference Warm Models with GPU Memory Swap

Running multiple[ inference models](/saas/workloads-in-nvidia-run-ai/workload-types) is a demanding task and you will need to ensure that your SLA is met. You need to provide high performance and low latency, while maximizing GPU utilization. This becomes even more challenging when the exact model usage patterns are unpredictable. You must plan for the agility of inference services and strive to keep models on standby in a ready state rather than an idle state.

NVIDIA Run:ai’s GPU memory swap feature enables you to load multiple models to a single GPU, where each can use up to the full amount GPU memory. Using an application load balancer, the administrator can control to which server each inference request is sent. Then the GPU can be loaded with multiple models, where the model in use is loaded into the GPU memory and the rest of the models are swapped-out to the CPU memory. The swapped models are stored as ready models to be loaded when required. GPU memory swap always maintains the context of the workload (model) on the GPU so it can easily and quickly switch between models. This is unlike industry standard model servers that load models from scratch into the GPU whenever required.

## How GPU Memory Swap Works

Swapping the workload’s GPU memory to and from the CPU is performed simultaneously and synchronously for all GPUs used by the workload. In some cases, if workloads specify a memory limit smaller than a full GPU memory size, multiple workloads can run in parallel on the same GPUs, maximizing the utilization and shortening the response times.

In other cases, workloads will run serially, with each workload running for a few seconds before the system swaps them in/out. If multiple workloads occupy more than the GPU physical memory and attempt to run simultaneously, memory swapping will occur. In this scenario, each workload will run part of the time on the GPU while being swapped out to the CPU memory the other part of the time, slowing down the execution of the workloads. Therefore, it is important to evaluate whether memory swapping is suitable for your specific use cases, weighing the benefits against the potential for slower execution time. To better understand the benefits and use cases of GPU memory swap, refer to the detailed sections below. This will help you determine how to best utilize GPU swap for your workloads and achieve optimal performance.

The workload MUST use [dynamic GPU fractions](/saas/platform-management/runai-scheduler/resource-optimization/dynamic-fractions). This means the workload’s memory Request is less than a full GPU, but it may add a GPU memory Limit to allow the workload to effectively use the full GPU memory. The NVIDIA Run:ai Scheduler allocates the dynamic fraction pair (Request and Limit) on single or multiple GPU devices in the same node.

To enable GPU memory swap, you must first configure the cluster with the required global setting, `global.core.swap.enabled`. After the cluster supports swap, you can enable GPU memory swap at the node-pool level. When creating or updating a node pool via the [Node pools](https://run-ai-docs.nvidia.com/api/organizations/nodepools) API, set the `gpuResourceOptimization.swapEnabled` parameter to `true`. NVIDIA Run:ai automatically applies the required `run.ai/swap-enabled=true` label to the nodes as part of this configuration. You can also use the `gpuResourceOptimization.cpuSwapMemorySize` field to specify the reserved CPU memory size for serving swapped GPU memory. See [Enabling and configuring GPU memory swap](#enabling-and-configuring-gpu-memory-swap).

## Multi-GPU Memory Swap

NVIDIA Run:ai also supports workload submission using multi-GPU memory swap. Multi-GPU memory swap works similarly to single GPU memory swap, but instead of swapping memory for a single GPU workload, it swaps memory for workloads across multiple GPUs simultaneously and synchronously.

The NVIDIA Run:ai Scheduler allocates the same dynamic GPU fraction pair (Request and Limit) on multiple GPU devices in the same node. For example, if you want to run two LLM models, each consuming 8 GPUs that are not used simultaneously, you can use GPU memory swap to share their GPUs. This approach allows multiple models to be stacked on the same node.

The following outlines the advantages of stacking multiple models on the same node:

* **Maximizes GPU utilization** - Efficiently uses available GPU resources by enabling multiple workloads to share GPUs.
* **Improves cold start times** - Loading large LLM models to a node and its GPUs can take several minutes during a “cold start”. Using memory swap turns this process into a “warm start” that takes only a fraction of a second to a few seconds (depending on the model size and the GPU model).
* **Increases GPU availability** - Frees up and maximizes GPU availability for additional workloads (and users), enabling better resource sharing.
* **Smaller quota requirements** - Enables more precise and often smaller quota requirements for the end user.

## Deployment Considerations

* A pod created before the GPU memory swap feature was enabled in that cluster, cannot be scheduled to a swap-enabled node. A proper event is generated in case no matching node is found. Users must re-submit those pods to make them swap-enabled.
* GPU memory swap cannot be enabled if the NVIDIA Run:ai [strict or fair time-slicing](/saas/platform-management/runai-scheduler/resource-optimization/time-slicing#gpu-time-slicing-modes) is used. GPU memory swap can only be used with the default NVIDIA time-slicing mechanism.
* CPU RAM size cannot be decreased once GPU memory swap is enabled.

## Enabling and Configuring GPU Memory Swap

Before configuring GPU memory swap, dynamic GPU fractions must be enabled. Dynamic GPU fractions enable you to make your workloads burstable as well as maximize your workloads’ performance and GPU utilization within a single node.

To enable GPU memory swap:

1. Enable the required global cluster setting. For more details, see [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config):
   * **Using Helm** - Set the following value in your `values.yaml` file under `clusterConfig` and upgrade the chart:

     ```yaml
     clusterConfig:
       global:
         core:
           swap:
             enabled: true
     ```
   * **Using runaiconfig at runtime** - Use the following `kubectl` patch command:

     ```bash
     kubectl patch -n runai runaiconfigs.run.ai/runai --type='merge' --patch '{"spec":{"global":{"core":{"swap":{"enabled": true}}}}}'
     ```
2. When creating or updating a node pool via the [Node pools](https://run-ai-docs.nvidia.com/api/organizations/nodepools) API, set the `gpuResourceOptimization.swapEnabled` parameter to `true`. The required `run.ai/swap-enabled=true` label is automatically applied to the nodes.
3. Specify the desired swap memory limit using the `gpuResourceOptimization.cpuSwapMemorySize` field. The default amount of CPU RAM reserved for Swap is **100GB**. CPU memory is shared across all GPUs on a Kubernetes node (GPU server), therefore when setting this parameter, administrators should set this value based on both the number of GPUs in the server and the number of workloads expected to use Swap. For example, consider a GPU server with 8×H100 GPUs, each with 80 GB of memory, and 4 LLM workloads sharing those GPUs. If each LLM consumes 40 GB per GPU, the required CPU memory reservation is:
   * 8 GPUs × 40 GB = 320 GB per LLM
   * 320 GB × 4 LLM workloads = 1,280 GB (1.2 TB) of CPU RAM needed for Swap

See the [Node pools](https://run-ai-docs.nvidia.com/api/organizations/nodepools) API for more details. For example:

```bash
{
  "name": "v100",
  "labelKey": "node-type",
  "labelValue": "type-x",
  "clusterId": "d73a738f-fab3-430a-8fa3-xxxx",
  "gpuResourceOptimization": {
    "swapEnabled": true,
    "cpuSwapMemorySize": "100GB"
  }
}
```

### Configuring System Reserved GPU Resources

Swappable workloads require reserving a small part of the GPU memory for non-swappable allocations like binaries and GPU context. To avoid getting out-of-memory (OOM) errors due to non-swappable memory regions, the system reserves a 2GiB of GPU RAM memory by default, effectively truncating the total size of the GPU memory. For example, a 16GiB T4 will appear as 14GiB on a swap-enabled node. The exact reserved size is application-dependent, and 2GiB is a safe assumption for 2-3 applications sharing and swapping on a GPU.

This value can be changed when creating or updating a node pool via the API using the `gpuResourceOptimization.reservedGpuMemoryForSwapOperations` parameter. See the [Node pools](https://run-ai-docs.nvidia.com/api/organizations/nodepools) API for more details. For example:

```bash
{
  "name": "v100",
  "labelKey": "node-type",
  "labelValue": "type-x",
  "clusterId": "d73a738f-fab3-430a-8fa3-xxxx",
  "gpuResourceOptimization": {
    "swapEnabled": true,
    "cpuSwapMemorySize": "100GB",
    "reservedGpuMemoryForSwapOperations": "2GB"
  }
}
```

### Performance Optimizations

#### Using Bi-Directional GPU-CPU Memory Swap Read-Write Operations

To optimize GPU-to-CPU memory swap performance, administrators can enable full-duplex bi-directional read-write operations using `clusterConfig.global.core.swap.biDirectional.enabled`. See [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config) for more details:

1. **Using Helm** - Set the following value in your `values.yaml` file under `clusterConfig` and upgrade the chart:

   ```yaml
   clusterConfig:
     global:
       core:
         swap:
           biDirectional:
             enabled: true
   ```
2. **Using runaiconfig at runtime** - Use the following `kubectl` patch command:

   ```bash
   kubectl patch -n runai runaiconfigs.run.ai/runai --type='merge' --patch '{"spec":{"global":{"core":{"swap":{"biDirectional": {"enabled": true}}}}}}'
   ```

Setting the read/write memory mode of GPU memory swap to bi-directional (full duplex) produces higher performance (typically +80%) vs. uni-directional (simplex) read-write operations.

#### Using UVA Based GPU-CPU Memory Mapped Swap Access

This setting enables the use of Unified Virtual Addressing (UVA) and early memory prefetching for GPU-to-CPU memory swap. It is especially effective when the GPU memory Request is close to the GPU memory Limit (e.g., Request = 90%, Limit = 100%). Enabling mapped mode can improve memory swap performance by +80–90% on newer GPUs such as H100 or B100 and up to +400% on GPUs such as A10 and L40.

Administrators can enable this setting using the `clusterConfig.global.core.swap.mode=mapped`. See [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config) for more details:

1. **Using Helm** - Set the following value in your `values.yaml` file under `clusterConfig` and upgrade the chart:

   ```yaml
   clusterConfig:
     global:
       core:
         swap:
           mode: mapped
   ```
2. **Using runaiconfig at runtime** - Use the following `kubectl` patch command:

   ```bash
   kubectl patch -n runai runaiconfigs.run.ai/runai --type='merge' --patch '{"spec":{"global":{"core":{"swap":{"mode": "mapped"}}}}}'
   ```

### Preventing Your Workloads from Getting Swapped

If you prefer your workloads not to be swapped into CPU memory, you can specify on the pod an anti-affinity to `run.ai/swap-enabled=true` node label when submitting your workloads and the Scheduler will ensure not to use swap-enabled nodes. An alternative way is to set swap on a dedicated node pool and not use this node pool for workloads you prefer not to swap.

### What Happens When the CPU Reserved Memory for GPU Swap Is Exhausted?

CPU memory is limited, and since a single CPU serves multiple GPUs on a node, this number is usually between 2 to 8. For example, when using 80GB of GPU memory, each swapped workload consumes up to 80GB (but may use less) assuming each GPU is shared between 2-4 workloads. In this example, you can see how the swap memory can become very large. Therefore, we give administrators a way to limit the size of the CPU reserved memory for swapped GPU memory on each swap-enabled node as shown in [enabling and configuring GPU memory swap](#enabling-and-configuring-gpu-memory-swap).


# CPU Compute and Memory Allocation

When allocating compute resources for workloads, GPUs are often seen as the main bottleneck. However, CPU compute and CPU memory are equally important:

* **CPU compute** - Essential for tasks such as data preprocessing and post-processing during training.
* **CPU memory** - Directly affects batch sizes and the volume of data your training run can handle efficiently.

Modern GPU servers typically include substantial CPU compute and memory to support these needs.

## Requesting CPU Compute and Memory

When submitting a [workload](/saas/workloads-in-nvidia-run-ai/workloads) or creating [compute resources](/saas/workloads-in-nvidia-run-ai/assets/compute-resources), you can explicitly request both CPU compute and memory resources. The system guarantees that, if the workload is scheduled, the requested resources will be available for that workload.

The number of CPUs and amount of memory your workload will receive is guaranteed to be the number requested. In practice, however, you may receive more resources than requested:

* If the workload you submitted is the only workload running on a node, it can utilize all available CPU compute and memory on that node until another workload is scheduled.
* When another workload is submitted, each workload will receive a number of compute and memory proportional to the number requested. For example, if the first workload requests 1 CPU and the second requests 3 CPUs on a node with 40 CPUs, the workloads will receive 10 and 30 CPUs respectively.
* If CPU compute and memory are not explicitly requested, resources are assigned from the [cluster default](#cluster-default-resource-settings).

{% hint style="info" %}
**Note**

If your workload temporarily uses more memory than requested and new workloads are scheduled, it may be forced to release memory, potentially resulting in an out of memory (OOM) error.
{% endhint %}

## Setting CPU Compute and Memory Limits

You can further control CPU compute and memory resource usage by specifying limits. The system will ensure your workload does not consume more than the limits set for compute or memory.

* If your workload exceeds its memory limit, it will get an out of memory error and may be terminated.
* The limit must be equal to or greater than the request.

## Cluster Default Resource Settings

If CPU compute and memory resource requests and/or limits are not explicitly provided, the system applies cluster-wide defaults.

### CPU Compute Requests

* **If GPUs are requested** - The default CPU allocation is determined by a set ratio of CPUs per GPU.\
  For example, with a default ratio of 1:6 and a request for 2 GPUs, your job will be assigned 12 CPUs (2 GPUs × 6 CPUs each).
* **If no GPUs are requested** - The default CPU allocation is determined by a ratio of CPUs per CPU limit. For example, with a default ratio of 1:0.2 and a CPU limit of 10, your job will be assigned 2 CPUs (10 × 0.2).
* **System defaults** - The out-of-the-box default is 1:1 (1 CPU per GPU requested) and 1:0.1 (0.1 CPUs per CPU limit if no GPUs are requested). These ratios can be modified in the cluster settings.

### CPU Memory Requests

* **If GPUs are requested** - The default memory allocation is set as a specific amount per GPU.\
  For example, if the default is 100MiB per GPU and your job requests 4 GPUs, it will be assigned 400MiB of memory.
* **If no GPUs are requested** - The default memory allocation is determined by a ratio of CPU memory limit to CPU memory request. By default, this ratio is 1:0.1 (your memory request will be 10% of the memory limit you specify).
* **System defaults** - The defaults are 100MiB per GPU, and 0.1 (10%) for memory requests when no GPUs are involved. These defaults can be customized in the cluster settings.

### CPU Compute and Memory Limits

By default, NVIDIA Run:ai sets the limit to Auto, meaning the workload can use up to the node’s maximum available resources unless you explicitly set a limit. Administrators can configure default limits using the cluster configuration.

## Setting Cluster Defaults

Administrators can change cluster-wide defaults for CPU and memory requests and/or limits. See [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config) for more details:

* **Using Helm** - Set the values under `clusterConfig` in your `values.yaml` file and upgrade the chart.
* **Using runaiconfig at runtime** - Edit the `runaiconfig` Custom Resource under `spec`.

The example below shows:

* A CPU request with a default ratio of 2:1 CPUs to GPUs.
* A CPU memory request with a default ratio of 200MB per GPU.
* A CPU limit with a default ratio of 4:1 CPU to GPU.
* A memory limit with a default ratio of 2GB per GPU.
* A CPU request with a default ratio of 0.1 CPUs per 1 CPU limit.
* A CPU memory request with a default ratio of 0.1:1 request per CPU memory limit.

```yaml
  clusterConfig:
    limitRange:
      cpuDefaultRequestGpuFactor: 2
      memoryDefaultRequestGpuFactor: 200Mi
      cpuDefaultLimitGpuFactor: 4
      memoryDefaultLimitGpuFactor: 2Gi
      cpuDefaultRequestCpuLimitFactorNoGpu: 0.1
      memoryDefaultRequestMemoryLimitFactorNoGpu: 0.1
```

## Validating CPU Resource Allocations

Once a workload is submitted, its resource requests and limits are reflected in the underlying Kubernetes pod. Review pod specifications in Kubernetes to check the assigned CPU and memory.

1. Get the pod name by running: `runai workload describe WORKLOAD_NAME`
2. The pod will appear under the `PODS` category. Run: `kubectl describe pod <POD_NAME>`

The information will appear under `Requests` and `Limits`. For example:

```yaml
Limits:
    nvidia.com/gpu:  2
Requests:
    cpu:             1
    memory:          104857600
    nvidia.com/gpu:  2
```


# Quick Starts


# Launching Workloads with GPU Fractions

This quick start provides a step-by-step walkthrough for running a Jupyter Notebook workspace using [GPU fractions](/saas/platform-management/runai-scheduler/resource-optimization/fractions).

NVIDIA Run:ai’s GPU fractions provides an agile and easy-to-use method to share a GPU or multiple GPUs across workloads. With GPU fractions, you can divide the GPU/s memory into smaller chunks and share the GPU/s compute resources between different workloads and users, resulting in higher GPU utilization and more efficient resource allocation.

## Prerequisites

Before you start, make sure:

* You have created a [project](/saas/platform-management/aiinitiatives/organization/projects) or have one created for you.
* The project has an assigned quota of at least 0.5 GPU.

{% hint style="info" %}
**Note**

Flexible workload submission is enabled by default. If unavailable, contact your administrator to enable it under **General settings** → Workloads → Flexible workload submission.
{% endhint %}

## Step 1: Logging In

{% tabs %}
{% tab title="UI" %}
Browse to the provided NVIDIA Run:ai user interface and log in with your credentials.
{% endtab %}

{% tab title="CLI v2" %}
Run the below --help command to obtain the login options and log in according to your setup:

```sh
runai login --help
```

{% endtab %}

{% tab title="API" %}
To use the API, you will need to obtain a token as shown in [API authentication](https://run-ai-docs.nvidia.com/api/getting-started/how-to-authenticate-to-the-api).
{% endtab %}
{% endtabs %}

## Step 2: Submitting a Workspace

{% tabs %}
{% tab title="UI - Flexible" %}

1. Go to the Workload manager → Workloads
2. Click **+NEW WORKLOAD** and select **Workspace**
3. Select under which **cluster** to create the workload
4. Select the **project** in which your workspace will run
5. Select **Start from scratch** to launch a new workspace quickly
6. Enter a **name** for the workspace (if the name already exists in the project, you will be requested to submit a different name)
7. Click **CONTINUE**

   In the next step:
8. Under **Environment**, click the **load** icon. A side pane appears, displaying a list of available environments. Select the **‘jupyter-lab’** environment for your workspace (Image URL: `jupyter/scipy-notebook)`
   * If ‘jupyter-lab’ is not displayed in the gallery, follow the below steps to create a one-time environment configuration:
     * Enter the jupyter-lab **Image URL** - `jupyter/scipy-notebook`
     * Tools - Set the connection for your tool
       * Click **+TOOL**
       * Select **Jupyter** tool from the list
     * Set the runtime settings for the environment. Click **+COMMAND & ARGUMENTS** and add the following:

       * Enter the command - `start-notebook.sh`
       * Enter the arguments - `--NotebookApp.token=''`

       **Note:** If [path-based routing](/saas/infrastructure-setup/advanced-setup/container-access/external-access-to-containers#access-to-the-running-workloads-container) is enabled on the cluster, enter `--NotebookApp.base_url=/${RUNAI_PROJECT}/${RUNAI_JOB_NAME} --NotebookApp.token=''`.
9. Under **Compute resources**, click the **load** icon. A side pane appears, displaying a list of available compute resources. Select the **‘small-fraction’** compute resource for your workspace.
   * If ‘small-fraction’ is not displayed in the gallery, follow the below steps to create a one-time compute resource configuration:
     * Set **GPU devices** per pod - 1
     * Enable **GPU fractioning** to set the GPU memory per device:
       * Select **% (of device)** - Fraction of a GPU device’s memory
       * Set the memory **Request** - 10 (the workload will allocate 10% of the GPU memory)
     * Optional: set the **CPU compute per pod** - 0.1 cores (default)
     * Optional: set the **CPU memory per pod** - 100 MB (default)
10. Click **CREATE WORKSPACE**
    {% endtab %}

{% tab title="UI - Original" %}

1. Go to the Workload manager → Workloads
2. Click **+NEW WORKLOAD** and select **Workspace**
3. Select under which **cluster** to create the workload
4. Select the **project** in which your workspace will run
5. Select **Start from scratch** to launch a new workspace quickly
6. Enter a **name** for the workspace (if the name already exists in the project, you will be requested to submit a different name)
7. Click **CONTINUE**

   In the next step:
8. Select the **‘jupyter-lab’** environment for your workspace (Image URL: `jupyter/scipy-notebook)`

   * If the ‘jupyter-lab’ is not displayed in the gallery, follow the below steps:
     * Click **+NEW ENVIRONMENT**
     * Enter jupyter-lab as the **name** for the environment. The name must be unique.
     * Enter the jupyter-lab **Image URL** - `jupyter/scipy-notebook`
     * Tools - Set the connection for your tool
       * Click **+TOOL**
       * Select **Jupyter** tool from the list
     * Set the runtime settings for the environment. Click **+COMMAND & ARGUMENTS** and add the following:

       * Enter the command - `start-notebook.sh`
       * Enter the arguments - `--NotebookApp.token=''`

       **Note:** If [path-based routing](/saas/infrastructure-setup/advanced-setup/container-access/external-access-to-containers#access-to-the-running-workloads-container) is enabled on the cluster, enter `--NotebookApp.base_url=/${RUNAI_PROJECT}/${RUNAI_JOB_NAME} --NotebookApp.token=''`.
     * Click **CREATE ENVIRONMENT**

   The newly created environment will be selected automatically
9. Select the **‘small-fraction’** compute resource for your workspace

   * If ‘small-fraction’ is not displayed in the gallery, follow the below steps:
     * Click **+NEW COMPUTE RESOURCE**
     * Enter small-fraction as the **name** for the compute resource. The name must be unique.
     * Set **GPU devices** per pod - 1
     * Enable **GPU fractioning** to set the GPU memory per device:
       * Select **% (of device)** - Fraction of a GPU device’s memory
       * Set the memory **Request** - 10 (the workload will allocate 10% of the GPU memory)
     * Optional: set the **CPU compute per pod** - 0.1 cores (default)
     * Optional: set the **CPU memory per pod** - 100 MB (default)
     * Click **CREATE COMPUTE RESOURCE**

   The newly created compute resource will be selected automatically
10. Click **CREATE WORKSPACE**
    {% endtab %}

{% tab title="CLI v2" %}
Copy the following command to your terminal. Make sure to update the below with the name of your project and workload. For more details, see [CLI reference](/saas/reference/cli/runai):

```sh
runai project set "project-name"
runai workspace submit "workload-name" --image jupyter/scipy-notebook \
--gpu-devices-request 0.1 --command --external-url container=8888 \
--name-prefix jupyter --command -- start-notebook.sh \
--NotebookApp.token=
```

{% endtab %}

{% tab title="API" %}
Copy the following command to your terminal. Make sure to update the below parameters. For more details, see [Workspaces](https://run-ai-docs.nvidia.com/api/workloads/workspaces) API:

```bash
curl -L 'https://<COMPANY-URL>/api/v1/workloads/workspaces' \ 
-H 'Content-Type: application/json' \
-H 'Authorization: Bearer <TOKEN>' \ 
-d '{ 
    "name": "workload-name", 
    "projectId": "<PROJECT-ID>",  
    "clusterId": "<CLUSTER-UUID>", 
    "spec": {
        "command" : "start-notebook.sh",
        "args" : "--NotebookApp.token=''",
        "image": "jupyter/scipy-notebook",
        "compute": {
            "gpuDevicesRequest": 1,
            "gpuRequestType": "portion",
            "gpuPortionRequest": 0.1

        },
        "exposedUrls" : [
            { 
                "container" : 8888,
                "toolType": "jupyter-notebook", 
                "toolName": "Jupyter" 
            }
        ]
    }
}
```

* `<COMPANY-URL>` - The link to the NVIDIA Run:ai user interface
* `<TOKEN>` - The API access token obtained in [Step 1](#step-1-logging-in)
* `<PROJECT-ID>` - The ID of the Project the workload is running on. You can get the Project ID via the [Get Projects](https://run-ai-docs.nvidia.com/api/organizations/projects#get-api-v1-org-unit-projects) API.
* `<CLUSTER-UUID>` - The unique identifier of the Cluster. You can get the Cluster UUID via the [Get Clusters](https://run-ai-docs.nvidia.com/api/organizations/clusters#get-api-v1-clusters) API.
* `toolType` will show the Jupyter icon when connecting to the Jupyter tool via the user interface.
* `toolName` will show when connecting to the Jupyter tool via the user interface.

{% hint style="info" %}
**Note**

The above API snippet runs with NVIDIA Run:ai clusters of 2.18 and above only.
{% endhint %}
{% endtab %}
{% endtabs %}

## Step 3: Connecting to the Jupyter Notebook

{% tabs %}
{% tab title="UI" %}

1. Select the newly created workspace with the Jupyter application that you want to connect to
2. Click **CONNECT**
3. Select the Jupyter tool. The selected tool is opened in a new tab on your browser.
   {% endtab %}

{% tab title="CLI v2" %}
To connect to the Jupyter Notebook, browse directly to <mark style="color:blue;">https\://\<COMPANY-URL>/\<PROJECT-NAME>/\<WORKLOAD-NAME></mark>
{% endtab %}

{% tab title="API" %}
To connect to the Jupyter Notebook, browse directly to <mark style="color:blue;">https\://\<COMPANY-URL>/\<PROJECT-NAME>/\<WORKLOAD-NAME></mark>
{% endtab %}
{% endtabs %}

## Next Steps

Manage and monitor your newly created workload using the [Workloads](/saas/workloads-in-nvidia-run-ai/workloads) table.


# Launching Workloads with Dynamic GPU Fractions

This quick start provides a step-by-step walkthrough for running a Jupyter Notebook with [dynamic GPU fractions](/saas/platform-management/runai-scheduler/resource-optimization/dynamic-fractions).

NVIDIA Run:ai’s dynamic GPU fractions optimizes GPU utilization by enabling workloads to dynamically adjust their resource usage. It allows users to specify a guaranteed fraction of GPU memory and compute resources with a higher limit that can be dynamically utilized when additional resources are requested.

## Prerequisites

Before you start, make sure:

* You have created a [project](/saas/platform-management/aiinitiatives/organization/projects) or have one created for you.
* The project has an assigned quota of at least 0.5 GPU.
* [Dynamic GPU fractions](/saas/platform-management/runai-scheduler/resource-optimization/dynamic-fractions) is enabled.

{% hint style="info" %}
**Note**

* Flexible workload submission is enabled by default. If unavailable, contact your administrator to enable it under **General settings** → Workloads → Flexible workload submission.
* Dynamic GPU fractions is disabled by default in the NVIDIA Run:ai UI. To use dynamic GPU fractions, it must be enabled by your administrator, under **General Settings** → Resources → GPU resource optimization.
  {% endhint %}

## Step 1: Logging In

{% tabs %}
{% tab title="UI" %}
Browse to the provided NVIDIA Run:ai user interface and log in with your credentials.
{% endtab %}

{% tab title="CLI v2" %}
Run the below --help command to obtain the login options and log in according to your setup:

```sh
runai login --help
```

{% endtab %}

{% tab title="API" %}
To use the API, you will need to obtain a token as shown in [API authentication](https://run-ai-docs.nvidia.com/api/getting-started/how-to-authenticate-to-the-api).
{% endtab %}
{% endtabs %}

## Step 2: Submitting the First Workspace

{% tabs %}
{% tab title="UI - Flexible" %}

1. Go to the Workload manager → Workloads
2. Click **+NEW WORKLOAD** and select **Workspace**
3. Select under which **cluster** to create the workload
4. Select the **project** in which your workspace will run
5. Select **Start from scratch** to launch a new workspace quickly
6. Enter a **name** for the workspace (if the name already exists in the project, you will be requested to submit a different name)
7. Click **CONTINUE**

   In the next step:
8. Under **Environment**, click the **load** icon. A side pane appears, displaying a list of available environments. To add a new environment:
   * Click the **+** icon to create a new environment
   * Enter quick-start as the **name** for the environment. The name must be unique.
   * Enter the **Image URL** - `gcr.io/run-ai-lab/pytorch-example-jupyter`
   * Tools - Set the connection for your tool:
     * Click **+TOOL**
     * Select **Jupyter** tool from the list
   * Set the runtime settings for the environment. Click **+COMMAND & ARGUMENTS** and add the following:

     * Enter the command - `start-notebook.sh`
     * Enter the arguments - `--NotebookApp.token=''`

     **Note:** If [path-based routing](/saas/infrastructure-setup/advanced-setup/container-access/external-access-to-containers#access-to-the-running-workloads-container) is enabled on the cluster, enter `--NotebookApp.base_url=/${RUNAI_PROJECT}/${RUNAI_JOB_NAME} --NotebookApp.token=''`.
   * Click **CREATE ENVIRONMENT**
   * Select the newly created environment from the side pane
9. Under **Compute resources**, click the **load** icon. A side pane appears, displaying a list of available compute resources. To add a new compute resource:
   * Click the **+** icon to create a new compute resource
   * Enter request-limit as the **name** for the compute resource. The name must be unique.
   * Set **GPU devices** per pod - 1
   * Enable **GPU fractioning** to set the GPU memory per device:
     * Select **GB** **-** Fraction of a GPU device’s memory
     * Set the memory **Request** - 4GB (the workload will allocate 4GB of the GPU memory)
     * Set the memory **Limit** - 12GB
   * Optional: set the **CPU compute per pod** - 0.1 cores (default)
   * Optional: set the **CPU memory per pod** - 100 MB (default)
   * Select **More settings** and toggle **Increase shared memory size**
   * Click **CREATE COMPUTE RESOURCE**
   * Select the newly created compute resource from the side pane
10. Click **CREATE WORKSPACE**
    {% endtab %}

{% tab title="UI - Original" %}

1. Go to the Workload manager → Workloads
2. Click **+NEW WORKLOAD** and select **Workspace**
3. Select under which **cluster** to create the workload
4. Select the **project** in which your workspace will run
5. Select **Start from scratch** to launch a new workspace quickly
6. Enter a **name** for the workspace (if the name already exists in the project, you will be requested to submit a different name)
7. Click **CONTINUE**

   In the next step:
8. Create an environment for your workspace

   * Click **+NEW ENVIRONMENT**
   * Enter quick-start as the **name** for the environment. The name must be unique.
   * Enter the **Image URL** - `gcr.io/run-ai-lab/pytorch-example-jupyter`
   * Tools - Set the connection for your tool
     * Click **+TOOL**
     * Select **Jupyter** tool from the list
   * Set the runtime settings for the environment. Click **+COMMAND & ARGUMENTS** and add the following:

     * Enter the command - `start-notebook.sh`
     * Enter the arguments - `--NotebookApp.token=''`

     **Note:** If [path-based routing](/saas/infrastructure-setup/advanced-setup/container-access/external-access-to-containers#access-to-the-running-workloads-container) is enabled on the cluster, enter `--NotebookApp.base_url=/${RUNAI_PROJECT}/${RUNAI_JOB_NAME} --NotebookApp.token=''`.
   * Click **CREATE ENVIRONMENT**

   The newly created environment will be selected automatically
9. Create a new “**request-limit**” compute resource for your workspace

   * Click **+NEW COMPUTE RESOURCE**
   * Enter request-limit as the **name** for the compute resource. The name must be unique.
   * Set **GPU devices** per pod - 1
   * Enable **GPU fractioning** to set the GPU memory per device:
     * Select **GB** **-** Fraction of a GPU device’s memory
     * Set the memory **Request** - 4GB (the workload will allocate 4GB of the GPU memory)
     * Set the memory **Limit** - 12GB
   * Optional: set the **CPU compute per pod** - 0.1 cores (default)
   * Optional: set the **CPU memory per pod** - 100 MB (default)
   * Select **More settings** and toggle **Increase shared memory size**
   * Click **CREATE COMPUTE RESOURCE**

   The newly created compute resource will be selected automatically
10. Click **CREATE WORKSPACE**
    {% endtab %}

{% tab title="CLI v2" %}
Copy the following command to your terminal. Make sure to update the below with the name of your project and workload. For more details, see [CLI reference](/saas/reference/cli/runai):

<pre class="language-sh"><code class="lang-sh">runai project set "project-name"
runai workspace submit "workload-name" \
--image gcr.io/run-ai-lab/pytorch-example-jupyter \
--gpu-memory-request 4G --gpu-memory-limit 12G --large-shm \
--external-url container=8888 --name-prefix jupyter  \
<strong>--command -- start-notebook.sh \
</strong><strong>--NotebookApp.token=
</strong></code></pre>

{% endtab %}

{% tab title="API" %}
Copy the following command to your terminal. Make sure to update the below parameters. For more details, see [Workspaces](https://run-ai-docs.nvidia.com/api/workloads/workspaces) API:

```bash
curl -L 'https://<COMPANY-URL>/api/v1/workloads/workspaces' \ 
-H 'Content-Type: application/json' \
-H 'Authorization: Bearer <TOKEN>' \ 
-d '{ 
    "name": "workload-name", 
    "projectId": "<PROJECT-ID>", 
    "clusterId": "<CLUSTER-UUID>",
    "spec": {
        "command" : "start-notebook.sh",
        "args" : "--NotebookApp.token=''",
        "image": "gcr.io/run-ai-lab/pytorch-example-jupyter",
        "compute": {
            "gpuDevicesRequest": 1,
            "gpuMemoryRequest": "4G",
            "gpuMemoryLimit": "12G",
            "largeShmRequest": true

        },
        "exposedUrls" : [
            { 
                "container" : 8888,
                "toolType": "jupyter-notebook", 
                "toolName": "Jupyter"  
            }
        ]
    }
}
```

* `<COMPANY-URL>` - The link to the NVIDIA Run:ai user interface
* `<TOKEN>` - The API access token obtained in [Step 1](#step-1-logging-in)
* `<PROJECT-ID>` - The ID of the Project the workload is running on. You can get the Project ID via the [Get Projects](https://run-ai-docs.nvidia.com/api/organizations/projects#get-api-v1-org-unit-projects) API.
* `<CLUSTER-UUID>` - The unique identifier of the Cluster. You can get the Cluster UUID via the [Get Clusters](https://run-ai-docs.nvidia.com/api/organizations/clusters#get-api-v1-clusters) API.
* `toolType` will show the Jupyter icon when connecting to the Jupyter tool via the user interface.
* `toolName` will show when connecting to the Jupyter tool via the user interface.

{% hint style="info" %}
**Note**

The above API snippet runs with NVIDIA Run:ai clusters of 2.18 and above only.
{% endhint %}
{% endtab %}
{% endtabs %}

## Step 3: Submitting the Second Workspace

{% tabs %}
{% tab title="UI - Flexible" %}

1. Go to the Workload manager → Workloads
2. Click **+NEW WORKLOAD** and select **Workspace**
3. Select the **cluster** where the previous workspace was created
4. Select the **project** where the previous workspace was created
5. Select **Start from scratch** to launch a new workspace quickly
6. Enter a **name** for the workspace (if the name already exists in the project, you will be requested to submit a different name)
7. Click **CONTINUE**

   In the next step:
8. Under **Environment**, click the **load** icon. A side pane appears, displaying a list of available environments. Select the environment created in [Step 2](#step-2-submitting-the-first-workspace).
9. Under **Compute resources**, click the **load** icon. A side pane appears, displaying a list of available compute resources. Select the compute resources created in [Step 2](#step-2-submitting-the-first-workspace).
10. Click **CREATE WORKSPACE**
    {% endtab %}

{% tab title="UI - Original" %}

1. Go to the Workload manager → Workloads
2. Click **+NEW WORKLOAD** and select **Workspace**
3. Select the **cluster** where the previous workspace was created
4. Select the **project** where the previous workspace was created
5. Select **Start from scratch** to launch a new workspace quickly
6. Enter a **name** for the workspace (if the name already exists in the project, you will be requested to submit a different name)
7. Click **CONTINUE**

   In the next step:
8. Select the environment created in [Step 2](#step-2-submitting-the-first-workspace)
9. Select the compute resource created in [Step 2](#step-2-submitting-the-first-workspace)
10. Click **CREATE WORKSPACE**
    {% endtab %}

{% tab title="CLI v2" %}
Copy the following command to your terminal. Make sure to update the below with the name of your project and workload. For more details, see [CLI reference](/saas/reference/cli/runai):

```sh
runai project set "project-name"
runai workspace submit "workload-name" \
--image gcr.io/run-ai-lab/pytorch-example-jupyter --gpu-memory-request 4G \
--gpu-memory-limit 12G --large-shm --external-url container=8888 \
--name-prefix jupyter --command -- start-notebook.sh \
--NotebookApp.token=
```

{% endtab %}

{% tab title="API" %}
Copy the following command to your terminal. Make sure to update the below parameters. For more details, see [Workspaces](https://run-ai-docs.nvidia.com/api/workloads/workspaces) API:

```bash
curl -L 'https://<COMPANY-URL>/api/v1/workloads/workspaces' \ 
-H 'Content-Type: application/json' \
-H 'Authorization: Bearer <TOKEN>' \ 
-d '{ 
    "name": "workload-name", 
    "projectId": "<PROJECT-ID>", 
    "clusterId": "<CLUSTER-UUID>",
    "spec": {
        "command" : "start-notebook.sh",
        "args" : "--NotebookApp.token=''",
        "image": "gcr.io/run-ai-lab/pytorch-example-jupyter",
        "compute": {
            "gpuDevicesRequest": 1,
            "gpuMemoryRequest": "4G",
            "gpuMemoryLimit": "12G",
            "largeShmRequest": true

        },
        "exposedUrls" : [
            { 
                "container" : 8888,
                "toolType": "jupyter-notebook",  
                "toolName": "Jupyter" 
            }
        ]
    }
}
```

* `<COMPANY-URL>` - The link to the NVIDIA Run:ai user interface
* `<TOKEN>` - The API access token obtained in [Step 1](#step-1-logging-in)
* `<PROJECT-ID>` - The ID of the Project the workload is running on. You can get the Project ID via the [Get Projects](https://run-ai-docs.nvidia.com/api/organizations/projects#get-api-v1-org-unit-projects) API.
* `<CLUSTER-UUID>` - The unique identifier of the Cluster. You can get the Cluster UUID via the [Get Clusters](https://run-ai-docs.nvidia.com/api/organizations/clusters#get-api-v1-clusters) API.
* `toolType` will show the Jupyter icon when connecting to the Jupyter tool via the user interface.
* `toolName` will show when connecting to the Jupyter tool via the user interface.

{% hint style="info" %}
**Note**

The above API snippet runs with NVIDIA Run:ai clusters of 2.18 and above only.
{% endhint %}
{% endtab %}
{% endtabs %}

## Step 4: Connecting to the Jupyter Notebook

{% tabs %}
{% tab title="UI" %}

1. Select the newly created workspace with the Jupyter application that you want to connect to
2. Click **CONNECT**
3. Select the Jupyter tool. The selected tool is opened in a new tab on your browser.
4. Open a terminal and use the `watch nvidia-smi` to get a constant reading of the memory consumed by the pod. Note that the number shown in the memory box is the Limit and not the Request or Guarantee.
5. Open the file `Untitled.ipynb` and move the frame so you can see both tabs
6. Execute both cells in `Untitled.ipynb`. This will consume about **3 GB of GPU memory** and be well below the 4GB of the **GPU Memory Request** value.
7. In the second cell, edit the value after `--image-size` from 100 to 200 and run the cell. This will increase the GPU memory utilization to about 11.5 GB which is above the Request value.
   {% endtab %}

{% tab title="CLI v2" %}

1. To connect to the Jupyter Notebook, browse directly to <mark style="color:blue;">https\://\<COMPANY-URL>/\<PROJECT-NAME>/\<WORKLOAD-NAME></mark>
2. Open a terminal and use the `watch nvidia-smi` to get a constant reading of the memory consumed by the pod. Note that the number shown in the memory box is the Limit and not the Request or Guarantee.
3. Open the file `Untitled.ipynb` and move the frame so you can see both tabs
4. Execute both cells in `Untitled.ipynb`. This will consume about **3 GB of GPU memory** and be well below the 4GB of the **GPU Memory Request** value.
5. In the second cell, edit the value after `--image-size` from 100 to 200 and run the cell. This will increase the GPU memory utilization to about 11.5 GB which is above the Request value.
   {% endtab %}

{% tab title="API" %}

1. To connect to the Jupyter Notebook, browse directly to <mark style="color:blue;">https\://\<COMPANY-URL>/\<PROJECT-NAME>/\<WORKLOAD-NAME></mark>
2. Open a terminal and use the `watch nvidia-smi` to get a constant reading of the memory consumed by the pod. Note that the number shown in the memory box is the Limit and not the Request or Guarantee.
3. Open the file `Untitled.ipynb` and move the frame so you can see both tabs
4. Execute both cells in `Untitled.ipynb`. This will consume about **3 GB of GPU memory** and be well below the 4GB of the **GPU Memory Request** value.
5. In the second cell, edit the value after `--image-size` from 100 to 200 and run the cell. This will increase the GPU memory utilization to about 11.5 GB which is above the Request value.
   {% endtab %}
   {% endtabs %}

## Next Steps

Manage and monitor your newly created workload using the [Workloads](/saas/workloads-in-nvidia-run-ai/workloads) table.


# Launching Workloads with GPU Memory Swap

This quick start provides a step-by-step walkthrough for running multiple LLMs (inference workload) on a single GPU using [GPU memory swap](/saas/platform-management/runai-scheduler/resource-optimization/memory-swap).

GPU memory swap expands the GPU physical memory to the CPU memory, allowing NVIDIA Run:ai to place and run more workloads on the same GPU physical hardware. This provides a smooth workload context switching between GPU memory and CPU memory, eliminating the need to kill workloads when the memory requirement is larger than what the GPU physical memory can provide.

## Prerequisites

Before you start, make sure:

* You have created a [project](/saas/platform-management/aiinitiatives/organization/projects) or have one created for you.
* The project has an assigned quota of at least 1 GPU.
* [Dynamic GPU fractions](/saas/platform-management/runai-scheduler/resource-optimization/dynamic-fractions) is enabled.
* GPU memory swap is enabled on at least one free node as detailed [here](/saas/platform-management/runai-scheduler/resource-optimization/memory-swap#enabling-and-configuring-gpu-memory-swap).
* [Host-based routing](https://github.com/run-ai/runai-product-docs/blob/SaaS/getting-started/installation/system-requirements.md#host-based-routing) is configured.

{% hint style="info" %}
**Note**

* Flexible workload submission is enabled by default. If unavailable, contact your administrator to enable it under **General settings** → Workloads → Flexible workload submission.
* Dynamic GPU fractions is disabled by default in the NVIDIA Run:ai UI. To use dynamic GPU fractions, it must be enabled by your administrator, under **General Settings** → Resources → GPU resource optimization.
  {% endhint %}

## Step 1: Logging In

{% tabs %}
{% tab title="UI" %}
Browse to the provided NVIDIA Run:ai user interface and log in with your credentials.
{% endtab %}

{% tab title="API" %}
To use the API, you will need to obtain a token as shown in [API authentication](https://run-ai-docs.nvidia.com/api/getting-started/how-to-authenticate-to-the-api).
{% endtab %}
{% endtabs %}

## Step 2: Submitting the First Inference Workload

{% tabs %}
{% tab title="UI - Flexible" %}

1. Go to the Workload manager → Workloads
2. Click **+NEW WORKLOAD** and select **Inference**
3. Select under which **cluster** to create the workload
4. Select the **project** in which your workload will run
5. Select **custom** inference from **Inference type** (if applicable)
6. Enter a **name** for the workload (if the name already exists in the project, you will be requested to submit a different name)
7. Click **CONTINUE**

   In the next step:
8. Under **Environment**, click the **load** icon. A side pane appears, displaying a list of available environments. To add a new environment:
   * Click the **+** icon to create a new environment
   * Enter quick-start as the **name** for the environment. The name must be unique.
   * Enter the NVIDIA Run:ai vLLM **Image URL** - `runai.jfrog.io/core-llm/runai-vllm:v0.6.4-0.10.0`
   * Set the inference **serving endpoint** to **HTTP** and the container port to `8000`
   * Set the runtime settings for the environment. Click **+ENVIRONMENT VARIABLE** and add the following:
     * **Name:** RUNAI\_MODEL **Source:** Custom **Value:** `meta-llama/Llama-3.2-1B-Instruct` (you can choose any vLLM supporting model from Hugging Face)
     * **Name:** RUNAI\_MODEL\_NAME **Source:** Custom **Value:** `Llama-3.2-1B-Instruct`
     * **Name:** HF\_TOKEN **Source:** Custom **Value:** \<Your Hugging Face token> (only needed for gated models)
     * **Name:** VLLM\_RPC\_TIMEOUT **Source:** Custom **Value:** 60000
   * Click **CREATE ENVIRONMENT**
   * Select the newly created environment from the side pane
9. Under **Compute resources**, click the **load** icon. A side pane appears, displaying a list of available compute resources. To add a new compute resource:
   * Click the **+** icon to create a new compute resource
   * Enter request-limit as the **name** for the compute resource. The name must be unique.
   * Set **GPU devices** per pod - 1
   * Enable **GPU fractioning** to set the GPU memory per device:
     * Select **% (of device)** - Fraction of a GPU device’s memory
     * Set the memory **Request** - 50 (the workload will allocate 50% of the GPU memory)
     * Set the memory **Limit** - 100%
   * Optional: set the **CPU compute** per pod - 0.1 cores (default)
   * Optional: set the **CPU memory** per pod - 100 MB (default)
   * Select **More settings** and toggle **Increase shared memory size**
   * Click **CREATE COMPUTE RESOURCE**
   * Select the newly created compute resource from the side pane
10. Click **CREATE INFERENCE**
    {% endtab %}

{% tab title="UI - Original" %}

1. Go to the Workload manager → Workloads
2. Click **+NEW WORKLOAD** and select **Inference**
3. Select under which **cluster** to create the workload
4. Select the **project** in which your workload will run
5. Select **custom** inference from **Inference type** (if applicable)
6. Enter a **name** for the workload (if the name already exists in the project, you will be requested to submit a different name)
7. Click **CONTINUE**

   In the next step:
8. Create an environment for your workload

   * Click **+NEW ENVIRONMENT**
   * Enter quick-start as the **name** for the environment. The name must be unique.
   * Enter the NVIDIA Run:ai vLLM **Image URL** - `runai.jfrog.io/core-llm/runai-vllm:v0.6.4-0.10.0`
   * Set the runtime settings for the environment. Click **+ENVIRONMENT VARIABLE** and add the following:
     * **Name:** RUNAI\_MODEL **Source:** Custom **Value:** `meta-llama/Llama-3.2-1B-Instruct` (you can choose any vLLM supporting model from Hugging Face)
     * **Name:** RUNAI\_MODEL\_NAME **Source:** Custom **Value:** `Llama-3.2-1B-Instruct`
     * **Name:** HF\_TOKEN **Source:** Custom **Value:** \<Your Hugging Face token> (only needed for gated models)
     * **Name:** VLLM\_RPC\_TIMEOUT **Source:** Custom **Value:** 60000
   * Click **CREATE ENVIRONMENT**

   The newly created environment will be selected automatically
9. Create a new “**request-limit**” compute resource

   * Click **+NEW COMPUTE RESOURCE**
   * Enter request-limit as the **name** for the compute resource. The name must be unique.
   * Set **GPU devices** per pod - 1
   * Enable **GPU fractioning** to set the GPU memory per device:
     * Select **% (of device)** - Fraction of a GPU device’s memory
     * Set the memory **Request** - 50 (the workload will allocate 50% of the GPU memory)
     * Set the memory **Limit** - 100%
   * Optional: set the **CPU compute per pod** - 0.1 cores (default)
   * Optional: set the **CPU memory per pod** - 100 MB (default)
   * Select **More settings** and toggle **Increase shared memory size**
   * Click **CREATE COMPUTE RESOURCE**

   The newly created compute resource will be selected automatically
10. Click **CREATE INFERENCE**
    {% endtab %}

{% tab title="API" %}
Copy the following command to your terminal. Make sure to update the below parameters. For more details, see [Inferences](https://run-ai-docs.nvidia.com/api/workloads/inferences) API:

```bash
curl -L 'https://<COMPANY-URL>/api/v1/workloads/inferences' \ 
-H 'Content-Type: application/json' \
-H 'Authorization: Bearer <TOKEN>' \
-d '{ 
    "name": "workload-name", 
    "useGivenNameAsPrefix": true,
    "projectId": "<PROJECT-ID>", 
    "clusterId": "<CLUSTER-UUID>", 
    "spec": {
        "image": "runai.jfrog.io/core-llm/runai-vllm:v0.6.4-0.10.0",
        "imagePullPolicy":"IfNotPresent",
        "environmentVariables": [
          {
            "name": "RUNAI_MODEL",
            "value": "meta-lama/Llama-3.2-1B-Instruct"
          },
          {
            "name": "VLLM_RPC_TIMEOUT",
            "value": "60000"
          },
          {
            "name": "HF_TOKEN",
            "value":"<INSERT HUGGINGFACE TOKEN>"
          }
        ],
        "compute": {
            "gpuDevicesRequest": 1,
            "gpuRequestType": "portion",
            "gpuPortionRequest": 0.1,
            "gpuPortionLimit": 1,
            "cpuCoreRequest":0.2,
            "cpuMemoryRequest": "200M",
            "largeShmRequest": false

        },
        "servingPort": {
            "container": 8000,
            "protocol": "http",
            "authorizationType": "public"
        }
    }
}       
```

* `<COMPANY-URL>` - The link to the NVIDIA Run:ai user interface
* `<TOKEN>` - The API access token obtained in [Step 1](#step-1-logging-in)
* `<PROJECT-ID>` - The ID of the Project the workload is running on. You can get the Project ID via the [Get Projects](https://run-ai-docs.nvidia.com/api/organizations/projects#get-api-v1-org-unit-projects) API.
* `<CLUSTER-UUID>` - The unique identifier of the Cluster. You can get the Cluster UUID via the [Get Clusters](https://run-ai-docs.nvidia.com/api/organizations/clusters#get-api-v1-clusters) API.

{% hint style="info" %}
**Note**

The above API snippet runs with NVIDIA Run:ai clusters of 2.18 and above only.
{% endhint %}
{% endtab %}
{% endtabs %}

## Step 3: Submitting the Second Inference Workload

{% tabs %}
{% tab title="UI - Flexible" %}

1. Go to the Workload manager → Workloads
2. Click **+NEW WORKLOAD** and select **Inference**
3. Select the **cluster** where the previous inference workload was created
4. Select the **project** where the previous inference workload was created
5. Select **custom** inference from **Inference type** (if applicable)
6. Enter a **name** for the workload (if the name already exists in the project, you will be requested to submit a different name)
7. Click **CONTINUE**

   In the next step:
8. Under **Environment**, click the **load** icon. A side pane appears, displaying a list of available environments. Select the environment created in [Step 2](#step-2-submitting-the-first-inference-workload).
9. Under **Compute resources**, click the **load** icon. A side pane appears, displaying a list of available compute resources. Select the compute resources created in [Step 2](#step-2-submitting-the-first-inference-workload).
10. Click **CREATE INFERENCE**
    {% endtab %}

{% tab title="UI - Original" %}

1. Go to the Workload manager → Workloads
2. Click **+NEW WORKLOAD** and select **Inference**
3. Select the **cluster** where the previous inference workload was created
4. Select the **project** where the previous inference workload was created
5. Select **custom** inference from **Inference type** (if applicable)
6. Enter a **name** for the workload (if the name already exists in the project, you will be requested to submit a different name)
7. Click **CONTINUE**

   In the next step:
8. Select the environment created in [Step 2](#step-2-submitting-the-first-inference-workload)
9. Select the compute resource created in [Step 2](#step-2-submitting-the-first-inference-workload)
10. Click **CREATE INFERENCE**
    {% endtab %}

{% tab title="API" %}
Copy the following command to your terminal. Make sure to update the below parameters. For more details, see [Inferences](https://run-ai-docs.nvidia.com/api/workloads/inferences) API:

```bash
curl -L 'https://<COMPANY-URL>/api/v1/workloads/inferences' \ 
-H 'Content-Type: application/json' \
-H 'Authorization: Bearer <TOKEN>' \ 
-d '{ 
    "name": "workload-name", 
    "useGivenNameAsPrefix": true,
    "projectId": "<PROJECT-ID>",  
    "clusterId": "<CLUSTER-UUID>",
    "spec": {
        "image": "runai.jfrog.io/core-llm/runai-vllm:v0.6.4-0.10.0",
        "imagePullPolicy":"IfNotPresent",
        "environmentVariables": [
          {
            "name": "RUNAI_MODEL",
            "value": "meta-lama/Llama-3.2-1B-Instruct"
          },
          {
            "name": "VLLM_RPC_TIMEOUT",
            "value": "60000"
          },
          {
            "name": "HF_TOKEN",
            "value":"<INSERT HUGGINGFACE TOKEN>"
          }
        ],
        "compute": {
            "gpuDevicesRequest": 1,
            "gpuRequestType": "portion",
            "gpuPortionRequest": 0.1,
            "gpuPortionLimit": 1,
            "cpuCoreRequest":0.2,
            "cpuMemoryRequest": "200M",
            "largeShmRequest": false

        },
        "servingPort": {
            "container": 8000,
            "protocol": "http",
            "authorizationType": "public"
        }
    }
}       
```

* `<COMPANY-URL>` - The link to the NVIDIA Run:ai user interface
* `<TOKEN>` - The API access token obtained in [Step 1](#step-1-logging-in)
* `<PROJECT-ID>` - The ID of the Project the workload is running on. You can get the Project ID via the [Get Projects](https://run-ai-docs.nvidia.com/api/organizations/projects#get-api-v1-org-unit-projects) API.
* `<CLUSTER-UUID>` - The unique identifier of the Cluster. You can get the Cluster UUID via the [Get Clusters](https://run-ai-docs.nvidia.com/api/organizations/clusters#get-api-v1-clusters) API.

{% hint style="info" %}
**Note**

The above API snippet runs with NVIDIA Run:ai clusters of 2.18 and above only.
{% endhint %}
{% endtab %}
{% endtabs %}

## Step 4: Submitting the First Workspace

{% tabs %}
{% tab title="UI - Flexible" %}

1. Go to the Workload manager → Workloads
2. Click COLUMNS and select **Connections**
3. Select the link under the Connections column for the first inference workload created in[ Step 2](#step-2-submitting-the-first-inference-workload)
4. In the **Connections Associated with Workload form,** copy the URL under the **Address** column
5. Click **+NEW WORKLOAD** and select **Workspace**
6. Select the **cluster** where the previous inference workloads were created
7. Select the **project** where the previous inference workloads were created
8. Select **Start from scratch** to launch a new workspace quickly
9. Enter a **name** for the workspace (if the name already exists in the project, you will be requested to submit a different name)
10. Click **CONTINUE**

    In the next step:
11. Under **Environment**, click the **load** icon. A side pane appears, displaying a list of available environments. Select the **‘chatbot-ui’** environment for your workspace (Image URL: `runai.jfrog.io/core-llm/llm-app)`
    * Set the runtime settings for the environment with the following **environment variables**:
      * **Name:** RUNAI\_MODEL\_NAME **Source:** Custom **Value:** `meta-llama/Llama-3.2-1B-Instruct`
      * **Name:** RUNAI\_MODEL\_BASE\_URL **Source:** Custom **Value:** Add the address link from Step 4
      * Delete the **PATH\_PREFIX** environment variable if you are using host-based routing.
    * If ‘chatbot-ui’ is not displayed in the gallery, follow the below steps:
      * Click the **+** icon to create a new environment
      * Enter chatbot-ui as the **name** for the environment. The name must be unique.
      * Enter the chatbot-ui **Image URL** - `runai.jfrog.io/core-llm/llm-app`
      * Tools - Set the connection for your tool
        * Click **+TOOL**
        * Select **Chatbot UI** tool from the list
      * Set the runtime settings for the environment. Click **+ENVIRONMENT VARIABLE** and add the following:
        * **Name:** RUNAI\_MODEL\_NAME **Source:** Custom **Value:** `meta-llama/Llama-3.2-1B-Instruct`
        * **Name:** RUNAI\_MODEL\_BASE\_URL **Source:** Custom **Value:** Add the **Address** link
        * **Name:** RUNAI\_MODEL\_TOKEN\_LIMIT **Source:** Custom **Value:** 8192
        * **Name:** RUNAI\_MODEL\_MAX\_LENGTH **Source:** Custom **Value:** 16384
      * Click **CREATE ENVIRONMENT**
      * Select the newly created environment from the side pane
12. Under **Compute resources**, click the **load** icon. A side pane appears, displaying a list of available compute resources. Select **‘cpu-only’** from the list.
    * If ‘cpu-only’ is not displayed, follow the below steps:
      * Click the **+** icon to create a new compute resource
      * Enter cpu-only as the **name** for the compute resource. The name must be unique.
      * Set **GPU devices** per pod - 0
      * Set **CPU compute** per pod - 0.1 cores
      * Set the **CPU memory** per pod - 100 MB (default)
      * Click **CREATE COMPUTE RESOURCE**
      * Select the newly created compute resource from the side pane
13. Click **CREATE WORKSPACE**
    {% endtab %}

{% tab title="UI - Original" %}

1. Go to the Workload manager → Workloads
2. Click COLUMNS and select **Connections**
3. Select the link under the Connections column for the first inference workload created in[ Step 2](#step-2-submitting-the-first-inference-workload)
4. In the **Connections Associated with Workload form,** copy the URL under the **Address** column
5. Click **+NEW WORKLOAD** and select **Workspace**
6. Select the **cluster** where the previous inference workloads were created
7. Select the **project** where the previous inference workloads were created
8. Select **Start from scratch** to launch a new workspace quickly
9. Enter a **name** for the workspace (if the name already exists in the project, you will be requested to submit a different name)
10. Click **CONTINUE**

    In the next step:
11. Select the **‘chatbot-ui’** environment for your workspace (Image URL: `runai.jfrog.io/core-llm/llm-app`)

    * Set the runtime settings for the environment with the following **environment variables**:
      * **Name:** RUNAI\_MODEL\_NAME **Source:** Custom **Value:** `meta-llama/Llama-3.2-1B-Instruct`
      * **Name:** RUNAI\_MODEL\_BASE\_URL **Source:** Custom **Value:** Add the address link from Step 4
      * Delete the **PATH\_PREFIX** environment variable if you are using host-based routing.
    * If ‘chatbot-ui’ is not displayed in the gallery, follow the below steps:
      * Click **+NEW ENVIRONMENT**
      * Enter chatbot-ui as the **name** for the environment. The name must be unique.
      * Enter the chatbot-ui **Image URL** - `runai.jfrog.io/core-llm/llm-app`
      * Tools - Set the connection for your tool
        * Click **+TOOL**
        * Select **Chatbot UI** tool from the list
      * Set the runtime settings for the environment. Click **+ENVIRONMENT VARIABLE** and add the following:
        * **Name:** RUNAI\_MODEL\_NAME **Source:** Custom **Value:** `meta-llama/Llama-3.2-1B-Instruct`
        * **Name:** RUNAI\_MODEL\_BASE\_URL **Source:** Custom **Value:** Add the **Address** link
        * **Name:** RUNAI\_MODEL\_TOKEN\_LIMIT **Source:** Custom **Value:** 8192
        * **Name:** RUNAI\_MODEL\_MAX\_LENGTH **Source:** Custom **Value:** 16384
      * Click **CREATE ENVIRONMENT**

    The newly created environment will be selected automatically
12. Select the **‘cpu-only’** compute resource for your workspace

    * If ‘cpu-only’ is not displayed in the gallery, follow the below steps:
      * Click **+NEW COMPUTE RESOURCE**
      * Enter cpu-only as the **name** for the compute resource. The name must be unique.
      * Set **GPU devices** per pod - 0
      * Set **CPU compute** per pod - 0.1 cores
      * Set the **CPU memory** per pod - 100 MB (default)
      * Click **CREATE COMPUTE RESOURCE**

    The newly created compute resource will be selected automatically
13. Click **CREATE WORKSPACE**
    {% endtab %}

{% tab title="API" %}
Copy the following command to your terminal. Make sure to update the below parameters. For more details, see [Workspaces](https://run-ai-docs.nvidia.com/api/workloads/workspaces) API:

```bash
curl -L 'https://<COMPANY-URL>/api/v1/workloads/workspaces' \ 
-H 'Content-Type: application/json' \
-H 'Authorization: Bearer <TOKEN>' \ 
-d '{ 
    "name": "workload-name", 
    "projectId": "<PROJECT-ID>", 
    "clusterId": "<CLUSTER-UUID>",
    "spec": {  
        "image": "runai.jfrog.io/core-llm/llm-app",
        "environmentVariables": [
          {
            "name": "RUNAI_MODEL_NAME",
            "value": "meta-llama/Llama-3.2-1B-Instruct"
          },
          {
            "name": "RUNAI_MODEL_BASE_URL",
            "value": "<URL>" 
          }
        ],
        "compute": {
            "cpuCoreRequest":0.1,
            "cpuMemoryRequest": "100M",
        }
    }
}
```

* `<COMPANY-URL>` - The link to the NVIDIA Run:ai user interface
* `<TOKEN>` - The API access token obtained in [Step 1](#step-1-logging-in)
* `<PROJECT-ID>` - The ID of the Project the workload is running on. You can get the Project ID via the [Get Projects](https://run-ai-docs.nvidia.com/api/organizations/projects#get-api-v1-org-unit-projects) API.
* `<CLUSTER-UUID>` - The unique identifier of the Cluster. You can get the Cluster UUID via the [Get Clusters](https://run-ai-docs.nvidia.com/api/organizations/clusters#get-api-v1-clusters) API.
* `<URL>` - The URL for connecting an external service related to the workload. You can get the URL via the [List Workloads](https://run-ai-docs.nvidia.com/api/workloads/workloads#get-api-v1-workloads) API.

{% hint style="info" %}
**Note**

The above API snippet runs with NVIDIA Run:ai clusters of 2.18 and above only.
{% endhint %}
{% endtab %}
{% endtabs %}

## Step 5: Submitting the Second Workspace

{% tabs %}
{% tab title="UI - Flexible" %}

1. Go to the Workload manager → Workloads
2. Click COLUMNS and select **Connections**
3. Select the link under the Connections column for the second inference workload created in [Step 3](#step-3-submitting-the-second-inference-workload)
4. In the **Connections Associated with Workload form,** copy the URL under the **Address** column
5. Click **+NEW WORKLOAD** and select **Workspace**
6. Select the **cluster** where the previous inference workloads were created
7. Select the **project** where the previous inference workloads were created
8. Select **Start from scratch** to launch a new workspace quickly
9. Enter a **name** for the workspace (if the name already exists in the project, you will be requested to submit a different name)
10. Click **CONTINUE**

    In the next step:
11. Under **Environment**, click the **load** icon. A side pane appears, displaying a list of available environments. Select the environment created in [Step 4](#step-4-submitting-the-first-workspace).
    * Set the runtime settings for the environment with the following **environment variables**:
      * **Name:** RUNAI\_MODEL\_NAME **Source:** Custom **Value:** `meta-llama/Llama-3.2-1B-Instruct`
      * **Name:** RUNAI\_MODEL\_BASE\_URL **Source:** Custom **Value:** Add the **Address** link
      * Delete the **PATH\_PREFIX** environment variable if you are using host-based routing.
12. Under **Compute resources**, click the **load** icon. A side pane appears, displaying a list of available compute resources. Select the compute resources created in [Step 4](#step-4-submitting-the-first-workspace).
13. Click **CREATE WORKSPACE**
    {% endtab %}

{% tab title="UI - Original" %}

1. Go to the Workload manager → Workloads
2. Click COLUMNS and select **Connections**
3. Select the link under the Connections column for the second inference workload created in [Step 3](#step-3-submitting-the-second-inference-workload)
4. In the **Connections Associated with Workload form,** copy the URL under the **Address** column
5. Click **+NEW WORKLOAD** and select **Workspace**
6. Select the **cluster** where the previous inference workloads were created
7. Select the **project** where the previous inference workloads were created
8. Select **Start from scratch** to launch a new workspace quickly
9. Enter a **name** for the workspace (if the name already exists in the project, you will be requested to submit a different name)
10. Click **CONTINUE**

    In the next step:
11. Select the environment created in [Step 4](#step-4-submitting-the-first-workspace)
    * Set the runtime settings for the environment with the following **environment variables**:
      * **Name:** RUNAI\_MODEL\_NAME **Source:** Custom **Value:** `meta-llama/Llama-3.2-1B-Instruct`
      * **Name:** RUNAI\_MODEL\_BASE\_URL **Source:** Custom **Value:** Add the **Address** link
      * Delete the **PATH\_PREFIX** environment variable if you are using host-based routing.
12. Select the compute resource created in [Step 4](#step-4-submitting-the-first-workspace)
13. Click **CREATE WORKSPACE**
    {% endtab %}

{% tab title="API" %}
Copy the following command to your terminal. Make sure to update the below parameters. For more details, see [Workspaces](https://run-ai-docs.nvidia.com/api/workloads/workspaces) API:

```bash
curl -L 'https://<COMPANY-URL>/api/v1/workloads/workspaces' \ 
-H 'Content-Type: application/json' \
-H 'Authorization: Bearer <TOKEN>' \ 
-d '{ 
    "name": "workload-name", 
    "projectId": "<PROJECT-ID>", '\ 
    "clusterId": "<CLUSTER-UUID>", \ 
    "spec": {  
        "image": "runai.jfrog.io/core-llm/llm-app",
        "environmentVariables": [
          {
            "name": "RUNAI_MODEL_NAME",
            "value": "meta-llama/Llama-3.2-1B-Instruct"
          },
          {
            "name": "RUNAI_MODEL_BASE_URL",
            "value": "<URL>" 
          }
        ],
        "compute": {
            "cpuCoreRequest":0.1,
            "cpuMemoryRequest": "100M",
        }
    }
}
```

* `<COMPANY-URL>` - The link to the NVIDIA Run:ai user interface
* `<TOKEN>` - The API access token obtained in [Step 1](#step-1-logging-in)
* `<PROJECT-ID>` - The ID of the Project the workload is running on. You can get the Project ID via the [Get Projects](https://run-ai-docs.nvidia.com/api/organizations/projects#get-api-v1-org-unit-projects) API.
* `<CLUSTER-UUID>` - The unique identifier of the Cluster. You can get the Cluster UUID via the [Get Clusters](https://run-ai-docs.nvidia.com/api/organizations/clusters#get-api-v1-clusters) API.
* `<URL>` - The URL for connecting an external service related to the workload. You can get the URL via the [List Workloads](https://run-ai-docs.nvidia.com/api/workloads/workloads#get-api-v1-workloads) API.

{% hint style="info" %}
**Note**

The above API snippet runs with NVIDIA Run:ai clusters of 2.18 and above only.
{% endhint %}
{% endtab %}
{% endtabs %}

## Step 6: Connecting to Chatbot-UI

{% tabs %}
{% tab title="UI" %}

1. Select the newly created workspace that you want to connect to
2. Click **CONNECT**
3. Select the ChatbotUI tool. The selected tool is opened in a new tab on your browser.
4. Query both workspaces simultaneously and see them both responding. The one on CPU RAM at the time will take longer as it switches back to the GPU and vice versa.
   {% endtab %}

{% tab title="API" %}

1. To connect to the ChatbotUI tool, browse directly to <mark style="color:blue;">https\://\<COMPANY-URL>/\<PROJECT-NAME>/\<WORKLOAD-NAME></mark>
2. Query both workspaces simultaneously and see them both responding. The one on CPU RAM at the time will take longer as it switches back to the GPU and vice versa.
   {% endtab %}
   {% endtabs %}

## Next Steps

Manage and monitor your newly created workloads using the [Workloads](/saas/workloads-in-nvidia-run-ai/workloads) table.


# Policies


# Policies and Rules

At NVIDIA Run:ai, administrators can access a suite of tools designed to facilitate efficient account management. This article focuses on two key features: workload policies and workload scheduling rules. These features empower admins to establish default values and implement restrictions allowing enhanced control, assuring compatibility with organizational policies, and optimizing resource usage and utilization.

## Workload Policies

A workload policy is an end-to-end solution for AI managers and administrators to control and simplify how workloads are submitted. This solution allows them to set best practices, enforce limitations, and standardize processes for the submission of workloads for AI projects within their organization. It acts as a key guideline for data scientists, researchers, ML & MLOps engineers by standardizing submission practices and simplifying the workload submission process.

### Why Use a Workload Policy?

Implementing workload policies is essential when managing complex AI projects within an enterprise for several reasons:

1. **Resource control and management** - Defining or limiting the use of costly resources across the enterprise via a centralized management system to ensure efficient allocation and prevent overuse.
2. **Setting best practices** - Provide managers with the ability to establish guidelines and standards to follow, reducing errors amongst AI practitioners within the organization.
3. **Security and compliance** - Define and enforce permitted and restricted actions to uphold organizational security and meet compliance requirements.
4. **Simplified setup** - Conveniently allow setting defaults and streamline the workload submission process for AI practitioners.
5. **Scalability and diversity**
   1. Multi-purpose clusters with various workload types that may have different requirements and characteristics for resource usage.
   2. The organization has multiple hierarchies, each with distinct goals, objectives, and degrees of flexibility.
   3. Manage multiple users and projects with distinct requirements and methods, ensuring appropriate utilization of resources.

### Understanding the Mechanism

The following sections provide details of how the workload policy mechanism works.

#### Cross-Interface Enforcement

The policy enforces the workloads regardless of whether they were submitted via UI, CLI, Rest APIs, or Kubernetes YAMLs.

#### Policy Types

NVIDIA Run:ai policies can target any workload type in the platform -- both NVIDIA Run:ai [native workloads](/saas/workloads-in-nvidia-run-ai/workload-types/native-workloads) and [supported workload types](/saas/workloads-in-nvidia-run-ai/workload-types/supported-workload-types).

* NVIDIA Run:ai native workloads are fully integrated into the platform. Policies for these workloads are configured using the policy YAML editor.
* Supported workload types include any other workload type registered in the platform (such as PyTorchJob, RayJob, and others). Policies for these types can be configured interactively using the Rule builder or written as YAML.

{% hint style="info" %}
**Note**

To apply policies to a new [workload type](/saas/workloads-in-nvidia-run-ai/workload-types/extending-workload-support) your organization has registered, you must first map it to the policy engine using the mapping APIs. See [Workload types policy mapping](/saas/platform-management/policies/supported-workload-type-policies/workload-type-policy-mapping).
{% endhint %}

### Policy Structure

The structure of a policy depends on the workload type it targets.

* Policies for NVIDIA Run:ai native workloads use three separate constructs:

  * **Rules** - Hard constraints on workload fields. Rules are enforced at submission time and cannot be overridden by users. For example, a rule can cap the maximum number of GPUs a workload may request, or prevent a field from being edited at all.
  * **Defaults** - Suggested values that are pre-filled when a user submits a workload. Defaults can be modified by the user before submission. For example, a default can set an initial CPU request that users are free to adjust.
  * **Imposed assets** - Assets (such as a PVC data source) that are automatically attached to any workload submitted under the policy’s scope, regardless of what the user specifies.

  See [NVIDIA Run:ai native workload policies](/saas/platform-management/policies/native-workload-policies) for more details.
* Supported workload type policies use a single rules construct. Each rule identifies a specific field and spec location within the workload, then applies one or more enforcements that control how that field behaves during submission. See [Supported workload types policies](/saas/platform-management/policies/supported-workload-type-policies) for more details.

## Scope of Effectiveness

{% hint style="info" %}
**Note**

Scope availability depends on the policy type:

* Policies for NVIDIA Run:ai native workloads support all scopes: system, cluster, department, and project.
* Policies for supported workload types support project scope only.
  {% endhint %}

Numerous teams working on various projects require the use of different tools, requirements, and safeguards. One policy may not suit all teams and their requirements. Hence, administrators can select the scope to cover the effectiveness of the policy. When a scope is selected, all of its subordinate units are also affected. As a result, all workloads submitted within the selected scope are controlled by the policy.

For example, if a policy is set for Department A, all workloads submitted by any of the projects within this department are controlled.

A scope for a policy can be:

![](/files/JQ80CI8eVsw9RnP40tcE)

{% hint style="info" %}
**Note**

The policy submission to the entire account scope is supported via API only.
{% endhint %}

The different scoping of policies also allows the breakdown of the responsibility between different administrators. This allows delegation of ownership between different levels within the organization. The policies, containing rules and defaults, propagate\* down the organizational tree, forming an “effective” policy that enforces any workload submitted by users within the project.

![](/files/hoxuJ1Ydkm4eiK0ZudW6)

If a field is used by multiple policies at different scopes, the platform applies a reconciliation mechanism to determine which policy takes effect. Defaults of the same field can still be submitted by different organizational policies, as they are considered “soft” rules. In this case, the closest scope to the workload becomes the effective default (project default “wins” vs. department default, department default “wins” vs. cluster default, etc.). For rules, precedence depends on their type: simple rules on non-security and non-compute fields follow the same order as defaults (project > department > cluster), while strict rules on security and compute fields apply in reverse order (cluster > department > project).

<details>

<summary>NVIDIA Run:ai policies vs. Kyverno policies</summary>

Kyverno runs as a dynamic admission controller in a Kubernetes cluster. Kyverno receives validating and mutating admission webhook HTTP callbacks from the Kubernetes API server and applies matching policies to return results that enforce admission policies or reject requests. Kyverno policies can match resources using the resource kind, name, label selectors, and much more. For more information, see [How Kyverno Works](https://kyverno.io/docs/introduction/#how-kyverno-works).

</details>

### System Policies

{% hint style="info" %}
**Note**

System policies apply to NVIDIA Run:ai native workloads only. They do not apply to policies for supported workload types.
{% endhint %}

By default, every account in NVIDIA Run:ai is governed by system policies that establish foundational security controls across all workloads, scopes, and interfaces (UI, CLI, API). These policies ensure consistent workload behavior and prevent unauthorized escalation. They can also be viewed as part of the effective policy for every scop&#x65;**.**

Administrators can create new policies to modify these defaults at any scope, including the project scope. A cluster- or department-scoped policy is not required to override a system policy default. This flexibility allows easing certain API restrictions that were previously enforced more strictly, while ensuring all changes are explicit and audited. Any update requires direct administrative action, adding an additional security layer.

* **Privileged parameter** - By default, the privileged parameter is set to `false` and is not editable (`canEdit: False`). This means containers are prevented from running with full host access, bypassing almost all container isolation, unless an administrator explicitly enables it.
* **Grace period** - Grace period defines how long a workload is allowed to continue running after a preemption request before it is terminated. The default grace period is 30 seconds, which may not always be sufficient for checkpointing or saving state. To support longer checkpointing operations, a maximum grace period of 5 minutes enforced by the system policy, which applies across API, CLI, and UI submissions. This period can be updated at any scope within the policy hierarchy.

## Scheduling Rules

[Scheduling rules](/saas/platform-management/policies/scheduling-rules) limit a researcher's access to resources and provides a way for the admin to control resource allocation and prevent the waste of resources. Admins should use the rules to prevent GPU idleness, prevent GPU hogging and allocate specific types of resources to different types of workloads.

Admin can limit the duration of a workload, the duration of the idle time, or the type of nodes the workload can use. Rules are defined for and apply to all workloads in the project or department. In addition, rules can be applied to a specific type of workload in a project or department (workspace, standard training, or inference). When a workload reaches the limitation of the rule, it is stopped if the rule is time-limited. The rule type prevents the workload from being scheduled on nodes that violate the rule limitation.


# NVIDIA Run:ai Native Workload Policies

This guide explains the structure and procedures for managing policies for [NVIDIA Run:ai native workloads](/saas/workloads-in-nvidia-run-ai/workload-types/native-workloads).

{% hint style="info" %}
**Note**

* Workload policies are enabled by default. If you do not see it in the menu, contact your administrator to enable it under **General settings** → Workloads → Policies.
* Starting from version 2.23, the creation and editing of control plane policies are no longer synchronized to the cluster. This ensures that new policy features function correctly and prevents conflicts with outdated cluster policies.
  * For cluster versions 2.21 and above, changes to control plane policies do not affect existing cluster policies, because workload policies are stored and enforced only in the control plane. Existing cluster policies impact only workloads submitted directly to the cluster, and administrators may choose to keep or manually delete these older cluster policies.
  * For cluster versions 2.18 to 2.20, policies in the cluster affect both control plane and cluster-submitted workloads. When a policy is created or edited in the control plane, the corresponding or effective policy in the cluster is automatically removed to avoid conflicts. Administrator acknowledgment is required before this change takes effect.
  * To enforce policies directly on the cluster in order to prevent direct workload submission, or for more information, please contact [NVIDIA Run:ai Support](https://www.nvidia.com/en-eu/support/enterprise/#contact-us).
    {% endhint %}

## Workload Policies Table

The Workload policies table can be found under **Policies** in the NVIDIA Run:ai platform.

The Workload policies table provides a list of all the policies defined in the platform, and allows you to manage them.

<figure><img src="/files/Xd3J56mPoTGBOklP31fT" alt=""><figcaption></figcaption></figure>

The Workload policies table consists of the following columns:

| Column              | Description                                                                                                                                                                                      |
| ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Policy              | The policy name which is a combination of the policy scope and the policy type                                                                                                                   |
| Type                | The policy type is per NVIDIA Run:ai workload type. This allows administrators to set different policies for each [workload type](/saas/workloads-in-nvidia-run-ai/workload-types).              |
| Policy architecture | Indicates whether the policy targets a **Standard** workload (Training, Workspace, or Inference) or a **Distributed** workload.                                                                  |
| Status              | The current state of the policy. Indicates whether the policy is Ready to enforce workloads.                                                                                                     |
| Scope               | The scope the policy affects. Click the name of the scope to view the organizational tree diagram. You can only view the parts of the organizational tree for which you have permission to view. |
| Created by          | The user who created the policy                                                                                                                                                                  |
| Creation time       | The timestamp for when the policy was created                                                                                                                                                    |
| Last updated        | The last time the policy was updated                                                                                                                                                             |

### Customizing the Table View

* Filter - Click ADD FILTER, select the column to filter by, and enter the filter values
* Search - Click SEARCH and type the value to search by
* Sort - Click each column header to sort by
* Column selection - Click COLUMNS and select the columns to display in the table
* Refresh - Click REFRESH to update the table with the latest data

## Policy Structure

A workload policy consists of three constructs:

* **Rules** - Constraints that limit or control the values of workload fields. Rules are enforced during submission and cannot be overridden by users.
* **Defaults** - Suggested values for workload fields. Defaults are pre-filled during submission but can be modified by users.
* **Imposed assets** - Assets (such as a PVC data source) that are automatically applied to any workload submitted under the policy's scope.

For more information, see [rules](/saas/platform-management/policies/native-workload-policies/policy-yaml-reference#rules), [defaults](/saas/platform-management/policies/native-workload-policies/policy-yaml-reference#defaults), and [imposed assets](/saas/platform-management/policies/policies-and-rules).

## Adding a Policy

To create a new policy:

1. Click **+NEW POLICY**
2. Select a **scope**
3. Select the **workload type**
4. Click **+POLICY YAML**
5. In the **YAML editor**, type or paste a YAML policy with defaults and rules.\
   You can use the following references and examples:
   * [Policy YAML reference](/saas/platform-management/policies/native-workload-policies/policy-yaml-reference)
   * [Policy YAML examples](/saas/platform-management/policies/native-workload-policies/policy-yaml-examples)
6. Click **SAVE POLICY**

## Editing a Policy

1. Select the policy you want to edit
2. Click **EDIT**
3. Update the policy and click **APPLY**
4. Click **SAVE POLICY**

## Troubleshooting

Listed below are issues that might occur when creating or editing a policy via the YAML Editor:

| Issue                                                                                                              | Message                                                                                                                              | Mitigation                                                                                                                                                                                                                                               |
| ------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Cluster connectivity issues                                                                                        | There's no communication from cluster “cluster\_name“. Actions may be affected, and the data may be stale.                           | Verify that you are on a network that has been allowed access to the cluster. Reach out to your cluster administrator for instructions on verifying the issue.                                                                                           |
| Policy can’t be applied due to a rule that is occupied by a different policy                                       | Field “field\_name” already has rules in cluster: “cluster\_id”                                                                      | Remove the rule from the new policy or adjust the old policy for the specific rule.                                                                                                                                                                      |
| Policy is not visible in the UI                                                                                    | -                                                                                                                                    | Check that the policy hasn’t been deleted.                                                                                                                                                                                                               |
| Policy syntax is no valid                                                                                          | Add a valid policy YAML;json: unknown field "field\_name"                                                                            | For correct syntax check the [Policy YAML reference](/saas/platform-management/policies/native-workload-policies/policy-yaml-reference) or the [Policy YAML examples](/saas/platform-management/policies/native-workload-policies/policy-yaml-examples). |
| Policy can’t be saved for some reason                                                                              | The policy couldn't be saved due to a network or other unknown issue. Download your draft and try pasting and saving it again later. | Possible cluster connectivity issues. Try updating the policy once again at a different time.                                                                                                                                                            |
| Policies were submitted before version 2.18, you upgraded to version 2.18 or above and wish to submit new policies | If you have policies and want to create a new one, first contact NVIDIA Run:ai support to prevent potential conflicts                | Contact NVIDIA Run:ai support. R\&D can migrate your old policies to the new version.                                                                                                                                                                    |

## Viewing a Policy

To view a policy:

1. Select the policy for which you want to view its [policies](/saas/platform-management/policies/policies-and-rules).
2. Click **VIEW POLICY**
3. In the Policy form per workload section, view the workload rules and defaults:
   * **Parameter**\
     The workload submission parameter that Rules and Defaults are applied to
   * **Type** (applicable for data sources only)\
     The data source type (Git, S3, nfs, pvc etc.)
   * **Default**\
     The default value of the Parameter
   * **Rule**\
     Set up constraint on workload policy field
   * **Source**\
     The origin of the applied policy (cluster, department or project)

{% hint style="info" %}
**Note**

Some of the rules and defaults may be derived from policies of a parent cluster and/or department. You can see the source of each rule in the policy form. For more information, check the [Scope of effectiveness documentation](/saas/platform-management/policies/policies-and-rules#scope-of-effectiveness).
{% endhint %}

## Deleting a Policy

1. Select the policy you want to delete
2. Click **DELETE**
3. On the dialog, click **DELETE** to confirm the deletion

## Using API

Go to the [Policies](https://run-ai-docs.nvidia.com/api/policies/policy) API reference to view the available actions.


# Policy YAML Examples

This page provides YAML examples for NVIDIA Run:ai native workload policies. These policies use a hierarchical YAML format with `defaults`, `rules`, and `imposedAssets` blocks organized by spec section.

## Creating a New Rule Within a Policy

This example shows how to add a new limitation to the GPU usage for workloads of type workspace:

1. Check the [workload API](https://run-ai-docs.nvidia.com/api/policies/policy) fields documentation and select the field(s) that are most relevant for GPU usage.

   ```json
   {
   "spec": {
       "compute": {
       "gpuDevicesRequest": 1,
       "gpuRequestType": "portion",
       "gpuPortionRequest": 0.5,
       "gpuPortionLimit": 0.5,
       "gpuMemoryRequest": "10M",
       "gpuMemoryLimit": "10M",
       "migProfile": "1g.5gb",
       "cpuCoreRequest": 0.5,
       "cpuCoreLimit": 2,
       "cpuMemoryRequest": "20M",
       "cpuMemoryLimit": "30M",
       "largeShmRequest": false,
       "extendedResources": [
           {
           "resource": "hardware-vendor.example/foo",
           "quantity": 2,
           "exclude": false
           }
       ]
       },
   }
   }
   ```
2. Search the field in the [Policy YAML fields - reference table](/saas/platform-management/policies/native-workload-policies/policy-yaml-reference). For example, gpuDevicesRequest appears under the **Compute fields** sub-table and appears as follow:

| Fields           | Description                                                                                                                           | Value type | Supported NVIDIA Run:ai workload type |
| ---------------- | ------------------------------------------------------------------------------------------------------------------------------------- | ---------- | ------------------------------------- |
| gpuDeviceRequest | Specifies the number of GPUs to allocate for the created workload. Only if `gpuDeviceRequest = 1`, the gpuRequestType can be defined. | integer    | Workspace & Training                  |

3. Use the value type of the gpuDevicesRequest field indicated in the table - “integer” and navigate to the **Value types** table to view the possible rules that can be applied to this value type -

   for integer, the options are:

   * canEdit
   * required
   * min
   * max
   * step
4. Proceed to the [Rule Type](/saas/platform-management/policies/native-workload-policies/policy-yaml-reference#rule-types) table, select the required rule for the limitation of the field - for example “max” and use the examples syntax to indicate the maximum GPU device requested.

```yaml
compute:
    gpuDevicesRequest:
        max: 2
```

## Policy YAML Best Practices

<details>

<summary>Create a policy that has multiple defaults and rules</summary>

**Best practices description**

Presentation of the syntax while adding a set of defaults and rules

**Example**

```yaml
defaults:
  createHomeDir: true
  environmentVariables:
    instances:
      - name: MY_ENV
        value: my_value
  security:
    allowPrivilegeEscalation: false

rules:
  storage:
    s3:
      attributes:
        url:
          options:
            - value: https://www.google.com
              displayed: https://www.google.com
            - value: https://www.yahoo.com
              displayed: https://www.yahoo.com
```

</details>

<details>

<summary>Allow only single selection out of many</summary>

**Best practices description**

Blocking the option to create all types of data sources except the one that is allowed is the solution

**Example**

```yaml
rules:
  storage:
    dataVolume:
      instances:
        canAdd: false
    hostPath:
      instances:
        canAdd: false
    pvc:
      instances:
        canAdd: false
    git:
      attributes:
        repository:
          required: true
        branch:
          required: true
        path:
          required: true
    nfs:
      instances:
        canAdd: false
    s3:
      instances:
        canAdd: false
```

</details>

<details>

<summary>Create a robust set of guidelines</summary>

**Best practices description**

Set rules for specific compute resource usage, addressing most relevant spec fields

**Example**

```yaml
compute:
    cpuCoreRequest:
      required: true
      min: 0
      max: 8
    cpuCoreLimit:
      min: 0
      max: 8
    cpuMemoryRequest:
      required: true
      min: '0'
      max: 16G
    cpuMemoryLimit:
      min: '0'
      max: 8G
    migProfile:
      canEdit: false
    gpuPortionRequest:
      min: 0
      max: 1
    gpuMemoryRequest:
      canEdit: false
    extendedResources:
      instances:
        canAdd: false
```

</details>

<details>

<summary>Policy for distributed training workloads</summary>

**Best practices description**

Set rules and defaults for a distributed training workload with different setting for master and workers

**Example**

```yaml
defaults:
  worker:
    command: my-command-worker-1
    environmentVariables:
      instances:
        - name: LOG_DIR
          value: policy-worker-to-be-ignored
        - name: ADDED_VAR
          value: policy-worker-added
    security:
      runAsUid: 500
    storage:
      s3:
        attributes:
          bucket: bucket1-worker
  master:
    command: my-command-master-2
    environmentVariables:
      instances:
        - name: LOG_DIR
          value: policy-master-to-be-ignored
        - name: ADDED_VAR
          value: policy-master-added
    security:
      runAsUid: 800
    storage:
      s3:
        attributes:
          bucket: bucket1-master
rules:
  worker:
    command:
      options:
        - value: my-command-worker-1
          displayed: command1
        - value: my-command-worker-2
          displayed: command2
    storage:
      nfs:
        instances:
          canAdd: false
      s3:
        attributes:
          bucket:
            options:
              - value: bucket1-worker
              - value: bucket2-worker
  master:
    command:
      options:
        - value: my-command-master-1
          displayed: command1
        - value: my-command-master-2
          displayed: command2
    storage:
      nfs:
        instances:
          canAdd: false
      s3:
        attributes:
          bucket:
            options:
              - value: bucket1-master
              - value: bucket2-maste
```

</details>

<details>

<summary>Examples for specific sections in the policy</summary>

**Best practices description**

Environment creation

**Example**

```yaml
rules:
  imagePullPolicy:
    required: true
    options:
      - value: Always
        displayed: Always
      - value: Never
        displayed: Never
  createHomeDir:
    canEdit: false
```

**Best practices description**

Setting security measures

**Example**

```yaml
rules:
  security:
    runAsUid:
      min: 1
      max: 32700
    allowPrivilegeEscalation:
      canEdit: false
```

**Best practices description**

Overriding the privileged system policy

**Example**

```yaml
defaults:
  security:
    privileged: true
rules:
  security:
    privileged:
      canEdit: true
```

**Best practices description**

Impose an asset

**Example**

<pre class="language-yaml"><code class="lang-yaml"><strong>defaults: null
</strong>rules: null
imposedAssets:
  - f12c965b-44e9-4ff6-8b43-01d8f9e630cc
</code></pre>

</details>

## Example of a Whole Policy

```yaml
defaults:
  createHomeDir: true
  imagePullPolicy: IfNotPresent
  nodePools:
    - node-pool-a
    - node-pool-b
  environmentVariables:
    instances:
      - name: WANDB_API_KEY
        value: REPLACE_ME!
      - name: WANDB_BASE_URL
        value: https://wandb.mydomain.com
  compute:
    cpuCoreRequest: 0.1
    cpuCoreLimit: 20
    cpuMemoryRequest: 10G
    cpuMemoryLimit: 40G
    largeShmRequest: true
  security:
    allowPrivilegeEscalation: false
  storage:
    git:
      attributes:
        repository: https://git-repo.my-domain.com
        branch: master
    hostPath:
      instances:
        - name: vol-data-1
          path: /data-1
          mountPath: /mount/data-1
        - name: vol-data-2
          path: /data-2
          mountPath: /mount/data-2
rules:
  createHomeDir:
    canEdit: false
  imagePullPolicy:
    canEdit: false
  environmentVariables:
    instances:
      locked:
        - WANDB_BASE_URL
  compute:
    cpuCoreRequest:
      max: 32
    cpuCoreLimit:
      max: 32
    cpuMemoryRequest:
      min: 1G
      max: 20G
    cpuMemoryLimit:
      min: 1G
      max: 40G
    largeShmRequest:
      canEdit: false
    extendedResources:
      instances:
        canAdd: false
  security:
    allowPrivilegeEscalation:
      canEdit: false
    runAsUid:
      min: 1
  storage:
    hostPath:
      instances:
        locked:
          - vol-data-1
          - vol-data-2
imposedAssets:
  - 4ba37689-f528-4eb6-9377-5e322780cc27

```


# Policy YAML Reference

A workload policy is an end-to-end solution for AI managers and administrators to control and simplify how workloads are submitted, setting best practices, enforcing limitations, and standardizing processes for AI projects within their organization.

This guide explains the policy YAML fields and the possible rules and defaults that can be set for each field.

{% hint style="info" %}
**Note**

This reference applies to NVIDIA Run:ai native workloads. To define policies for other supported workload types, see [Supported workload types policies](/saas/platform-management/policies/supported-workload-type-policies).
{% endhint %}

## Policy YAML Fields - Reference Table

The policy fields are structured in a similar format to the workload API fields. The following tables represent a structured guide designed to help you understand and configure policies in a YAML format. It provides the fields, descriptions, defaults and rules for each workload type.

Click the link to view the value type of each field.

<table><thead><tr><th width="150.5390625">Fields</th><th width="311.234375">Description</th><th width="132.86328125">Value type</th><th width="187.92578125">Supported NVIDIA Run:ai workload type</th></tr></thead><tbody><tr><td>args</td><td>When set, contains the arguments sent along with the command. These override the entry point of the image in the created workload</td><td><a href="#value-types">string</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li></ul></td></tr><tr><td>command</td><td>A command to serve as the entry point of the container running the workload</td><td><a href="#value-types">string</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li></ul></td></tr><tr><td>createHomeDir</td><td>Instructs the system to create a temporary home directory for the user within the container. Data stored in this directory is not saved when the container exists. When the runAsUser flag is set to true, this flag defaults to true as well</td><td><a href="#value-types">boolean</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li></ul></td></tr><tr><td>environmentVariables</td><td>Set of environmentVariables to populate the container running the workload</td><td><a href="#value-types">itemized</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>image</td><td>Specifies the image to use when creating the container running the workload</td><td><a href="#value-types">string</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>imagePullPolicy</td><td>Specifies the pull policy of the image when starting t a container running the created workload. Options are: Always, Never, or IfNotPresent</td><td><a href="#value-types">string</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>imagePullSecrets</td><td>Specifies a list of references to Kubernetes secrets in the same namespace used for pulling container images.</td><td><a href="#value-types">array</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>workingDir</td><td>Container’s working directory. If not specified, the container runtime default is used, which might be configured in the container image</td><td><a href="#value-types">string</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li></ul></td></tr><tr><td>nodeType</td><td>Nodes (machines) or a group of nodes on which the workload runs</td><td><a href="#value-types">string</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li></ul></td></tr><tr><td>nodePools</td><td>A prioritized list of node pools for the scheduler to run the workload on. The scheduler always tries to use the first node pool before moving to the next one when the first is not available.</td><td><a href="#value-types">array</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>annotations</td><td>Set of annotations to populate into the container running the workload</td><td><a href="#value-types">itemized</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>labels</td><td>Set of labels to populate into the container running the workload</td><td><a href="#value-types">itemized</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>terminateAfterPreemtpion</td><td>Indicates whether the job should be terminated, by the system, after it has been preempted</td><td><a href="#value-types">boolean</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li></ul></td></tr><tr><td>autoDeletionTimeAfterCompletionSeconds</td><td>Specifies the duration after which a finished workload (Completed or Failed) is automatically deleted. If this field is set to zero, the workload becomes eligible to be deleted immediately after it finishes.</td><td><a href="#value-types">integer</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li></ul></td></tr><tr><td>terminationGracePeriodSeconds</td><td>The duration, in seconds, that a workload is allowed to continue running after a preemption request before it is forcibly terminated. The grace period acts as a buffer that allows the workload to reach a safe checkpoint before termination. The default value is 30 seconds and is limited by a system policy to 300 seconds (5 minutes) across all workloads. Administrators can override the default by creating a new policy at the desired scope. See <a href="/pages/6r51yAr8J5lAAFDA67GO#system-policies">System policies</a> for more details.</td><td><a href="#value-types">integer</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li></ul></td></tr><tr><td>backoffLimit</td><td>Specifies the number of retries before marking a workload as failed</td><td><a href="#value-types">integer</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li></ul></td></tr><tr><td>restartPolicy</td><td><p>Specify the restart policy of the workload pods. Default is empty, which is determine by the framework default</p><p>Enum: "Always" "Never" "OnFailure"</p></td><td><a href="#value-types">string</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Distributed inference (API only)</li></ul></td></tr><tr><td>cleanPodPolicy</td><td><p>Specifies which pods will be deleted when the workload reaches a terminal state (completed/failed). The policy can be one of the following values:</p><ul><li><code>Running</code> - Only pods still running when a job completes (for example, parameter servers) will be deleted immediately. Completed pods will not be deleted so that the logs will be preserved. (Default for MPI)</li><li><code>All</code> - All (including completed) pods will be deleted immediately when the job finishes.</li><li><code>None</code> - No pods will be deleted when the job completes. It will keep running pods that consume GPU, CPU and memory over time. It is recommended to set to None only for debugging and obtaining logs from running pods. (Default for PyTorch)</li></ul></td><td><a href="#value-types">string</a></td><td>Distributed training</td></tr><tr><td>completions</td><td>Used with Hyperparameter Optimization. Specifies the number of successful pods the job should reach to be completed. The Job is marked as successful once the specified amount of pods has succeeded.</td><td><a href="#value-types">integer</a></td><td>Standard training</td></tr><tr><td>parallelism</td><td>Used with Hyperparameters Optimization. Specifies the maximum desired number of pods the workload should run at any given time.</td><td><a href="#value-types">itemized</a></td><td>Standard training</td></tr><tr><td>exposedUrls</td><td>Specifies a set of exported URLs (e.g. ingress) from the container running the created workload.</td><td><a href="#value-types">itemized</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li></ul></td></tr><tr><td>relatedUrls</td><td>Specifies a set of URLs related to the workload. For example, a URL to an external server providing statistics or logging about the workload.</td><td><a href="#value-types">itemized</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li></ul></td></tr><tr><td>PodAffinitySchedulingRule</td><td>Indicates if we want to use the Pod affinity rule as: the “hard” (required) or the “soft” (preferred) option.</td><td><a href="#value-types">string</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li></ul></td></tr><tr><td>podAffinityTopology</td><td>Specifies the Pod Affinity Topology to be used for scheduling the job.</td><td><a href="#value-types">string</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li></ul></td></tr><tr><td>category</td><td>Specifies the workload category assigned to the workload. Categories are used to classify and monitor different types of workloads within the NVIDIA Run:ai platform.</td><td><a href="#value-types">string</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>sshAuthMountPath</td><td>Specifies the directory where SSH keys are mounted</td><td><a href="#value-types">string</a></td><td>Distributed training (MPI only)</td></tr><tr><td>mpiLauncherCreationPolicy</td><td><p>Define s whether the MPI Launcher is created in parallel with the workers, or if its creation is postponed until all workers are in Ready state. This prevents failures when the launcher attempts to connect to workers that are not yet ready.</p><p>Enum: <code>AtStartup</code>, <code>WaitForWorkersReady</code></p></td><td><a href="#value-types">string</a></td><td>Distributed training (MPI only)</td></tr><tr><td>ports</td><td>Specifies a set of ports exposed from the container running the created workload. More information in <a href="#ports-fields">Ports fields</a> below.</td><td><a href="#value-types">itemized</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li></ul></td></tr><tr><td>preemptibility</td><td>Specifies whether the workload can be preempted by higher-priority workloads. Valid values are preemptible and non-preemptible.</td><td></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>probes</td><td>Specifies the ReadinessProbe to use to determine if the container is ready to accept traffic. More information in <a href="#probes-fields">Probes fields</a> below</td><td>-</td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>tolerations</td><td>Toleration rules which apply to the pods running the workload. Toleration rules guide (but do not require) the system to which node each pod can be scheduled to or evicted from, based on matching between those rules and the set of taints defined for each Kubernetes node.</td><td><a href="#value-types">itemized</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>priorityClass</td><td><p>Specifies the priority class of the workload. The default values are:</p><ul><li>Workspace - <code>High</code></li><li>Training / distributed training - <code>Low</code></li><li>Inference - <code>Very high</code></li></ul><p>You can change it to any of the following valid values to adjust the workload's scheduling behavior: <code>very-low</code>, <code>low</code>, <code>medium- low</code>, <code>medium</code>, <code>medium-high</code>, <code>high</code>, <code>very-high</code>.</p></td><td><a href="#value-types">string</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>nodeAffinityRequired</td><td>If the affinity requirements specified by this field are not met at scheduling time, the pod will not be scheduled onto the node. If the affinity requirements specified by this field cease to be met at some point during pod execution (e.g. due to an update), the system may or may not try to eventually evict the pod from its node.</td><td><a href="#value-types">array</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li></ul></td></tr><tr><td>startupPolicy</td><td><p>Determines when the worker pods should start during workload initialization.</p><ul><li><code>LeaderCreated</code>: Workers start after the leader pod is created.</li><li><code>LeaderReady</code>: Workers start only after the leader pod is ready.</li></ul><p>Default: "LeaderCreated"</p></td><td><a href="#value-types">string</a></td><td>Distributed inference (API only)</td></tr><tr><td>workers</td><td>Specifies the number of worker nodes to run. If set to 0, only the leader node will run, and no worker pods will be created. In this case, worker spec is not required. Default: 0</td><td><a href="#value-types">integer</a></td><td>Distributed inference (API only)</td></tr><tr><td>replicas</td><td>Specifies the number of leader-worker sets to deploy. Each replica represents a group consisting of one leader pod and multiple worker pods. For example, setting replicas: 3 will create 3 independent groups, each with its own leader and corresponding set of workers. Default: 1</td><td><a href="#value-types">integer</a></td><td>Distributed inference (API only)</td></tr><tr><td>replicas</td><td>The number of replicas to deploy. Default: 1</td><td><a href="#value-types">integer</a></td><td>NVIDIA NIM services (API only)</td></tr><tr><td>leader</td><td>Defines the pod specification for the leader. Must always be provided, regardless of the number of workers.</td><td>-</td><td>Distributed inference (API only)</td></tr><tr><td>worker</td><td>Defines the pod specification for the workers. Required only if the number of workers is greater than 0.</td><td>-</td><td>Distributed inference (API only)</td></tr><tr><td>multiNode</td><td>Defines whether the NIM service runs as a multi-node deployment. If workers is set to 1 or more, the service runs in multi-node.</td><td>-</td><td>NVIDIA NIM services (API only)</td></tr><tr><td>ngcAuthSecret</td><td>The name of a Kubernetes secret containing the NGC access credentials. The secret must contain a key named NGC_API_KEY with the API key as the value.</td><td><a href="#value-types">integer</a></td><td>NVIDIA NIM services (API only)</td></tr><tr><td>storage</td><td>Contains all the fields related to storage configurations. More information in <a href="#storage-fields">Storage fields</a> below.</td><td>-</td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li></ul></td></tr><tr><td>security</td><td>Contains all the fields related to security configurations. More information in <a href="#security-fields">Security fields</a> below.</td><td>-</td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>compute</td><td>Contains all the fields related to compute configurations. More information in <a href="#compute-fields">Compute fields </a>below.</td><td>-</td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>tty</td><td>Whether this container should allocate a TTY for itself, also requires 'stdin' to be true</td><td><a href="#value-types">boolean</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li></ul></td></tr><tr><td>stdin</td><td>Whether this container should allocate a buffer for stdin in the container runtime. If this is not set, reads from stdin in the container will always result in EOF</td><td><a href="#value-types">boolean</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li></ul></td></tr><tr><td>numWorkers</td><td>the number of workers that will be allocated for running the workload.</td><td><a href="#value-types">integer</a></td><td>Distributed training</td></tr><tr><td>distributedFramework</td><td><p>The distributed training framework used in the workload.</p><p>Enum: "MPI" "PyTorch" "TF" "XGBoost" "JAX"</p></td><td><a href="#value-types">string</a></td><td>Distributed training</td></tr><tr><td>slotsPerWorker</td><td>Specifies the number of slots per worker used in hostfile. Defaults to 1. (applicable only for MPI)</td><td><a href="#value-types">integer</a></td><td>Distributed training (MPI only)</td></tr><tr><td>minReplicas</td><td>The lower limit for the number of worker pods to which the training job can scale down. (applicable only for PyTorch)</td><td><a href="#value-types">integer</a></td><td>Distributed training (PyTorch only)</td></tr><tr><td>maxReplicas</td><td>The upper limit for the number of worker pods that can be set by the autoscaler. Cannot be smaller than MinReplicas. (applicable only for PyTorch)</td><td><a href="#value-types">integer</a></td><td>Distributed training (PyTorch only)</td></tr><tr><td>servingPort</td><td>Specifies the port for accessing the inference service. See <a href="#serving-port-fields">Serving Port Fields</a>.</td><td>-</td><td><ul><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>autoscaling</td><td>Specifies the minimum and maximum number of replicas to be scaled up and down to meet the changing demands of inference services. See <a href="#autoscaling-fields">Autoscaling Fields</a>.</td><td>-</td><td><ul><li>Inference</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>servingConfiguration</td><td>Specifies the inference workload serving configuration. See <a href="#serving-configuration-fields">Serving Configuration Fields</a>.</td><td>-</td><td>Inference</td></tr></tbody></table>

### Ports Fields

<table><thead><tr><th width="150.9375">Fields</th><th width="311.328125">Description</th><th width="132.6953125">Value type</th><th width="187.6640625">Supported NVIDIA Run:ai workload type</th></tr></thead><tbody><tr><td>container</td><td>The port that the container running the workload exposes.</td><td><a href="#value-types">string</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li></ul></td></tr><tr><td>serviceType</td><td>Specifies the default service exposure method for ports. the default shall be used for ports which do not specify service type. Options are: LoadBalancer, NodePort or ClusterIP. For more information see the <a href="/pages/9U3H4WiDkcbwrj4W66NS">External Access to Containers</a> guide.</td><td><a href="#value-types">string</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li></ul></td></tr><tr><td>external</td><td>The external port which allows a connection to the container port. If not specified, the port is auto-generated by the system.</td><td><a href="#value-types">integer</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li></ul></td></tr><tr><td>toolType</td><td>The tool type that runs on this port.</td><td><a href="#value-types">string</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li></ul></td></tr><tr><td>toolName</td><td>A name describing the tool that runs on this port.</td><td><a href="#value-types">string</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li></ul></td></tr></tbody></table>

### Probes Fields

<table><thead><tr><th width="151.17578125">Fields</th><th width="310.578125">Description</th><th width="133.0859375">Value type</th><th width="187.98828125">Supported NVIDIA Run:ai workload type</th></tr></thead><tbody><tr><td>readiness</td><td>Specifies the Readiness Probe to use to determine if the container is ready to accept traffic.</td><td>-</td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr></tbody></table>

<details>

<summary>Readiness Field Details</summary>

* **Description:** Specifies the Readiness Probe to use to determine if the container is ready to accept traffic
* **Value type:** [itemized](#value-types)
* **Example policy snippet:**

```yaml
defaults:
   probes:
     readiness:
         initialDelaySeconds: 2
```

<table><thead><tr><th>Spec readiness fields</th><th width="273">Description</th><th>Value type</th></tr></thead><tbody><tr><td>initialDelaySeconds</td><td>Number of seconds after the container has started before liveness or readiness probes are initiated.</td><td><a href="#value-types">integer</a></td></tr><tr><td>periodSeconds</td><td>How often (in seconds) to perform the probe</td><td><a href="#value-types">integer</a></td></tr><tr><td>timeoutSeconds</td><td>Number of seconds after which the probe times out</td><td><a href="#value-types">integer</a></td></tr><tr><td>successThreshold</td><td>Minimum consecutive successes for the probe to be considered successful after having failed</td><td><a href="#value-types">integer</a></td></tr><tr><td>failureThreshold</td><td>When a probe fails, the number of times to try before giving up</td><td><a href="#value-types">integer</a></td></tr></tbody></table>

</details>

### Security Fields

<table><thead><tr><th width="132.8046875">Fields</th><th width="311.32421875">Description</th><th width="133.25">Value type</th><th width="187.07421875">Supported NVIDIA Run:ai workload type</th></tr></thead><tbody><tr><td>uidGidSource</td><td><p>Indicates the way to determine the user and group ids of the container. The options are:</p><ul><li><code>fromTheImage</code> - user and group IDs are determined by the docker image that the container runs. This is the default option.</li><li><code>custom</code> - user and group IDs can be specified in the environment asset and/or the workspace creation request.</li><li><code>fromIdpToken</code> - user and group IDs are automatically taken from the identity provider (IdP) token (available only in SSO-enabled installations).</li></ul><p>For more information, see <a href="/pages/mPGUjedN8G1g4TiMme2l">User identity in containers</a>.</p></td><td><a href="#value-types">string</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li></ul></td></tr><tr><td>capabilities</td><td>The capabilities field allows adding a set of unix capabilities to the container running the workload. Capabilities are Linux distinct privileges traditionally associated with superuser which can be independently enabled and disabled</td><td><a href="#value-types">array</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li></ul></td></tr><tr><td>seccompProfileType</td><td><p>Indicates which kind of seccomp profile is applied to the container. The options are:</p><ul><li>RuntimeDefault - the container runtime default profile should be used</li><li>Unconfined - no profile should be applied</li></ul></td><td><a href="#value-types">string</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li></ul></td></tr><tr><td>runAsNonRoot</td><td>Indicates that the container must run as a non-root user.</td><td><a href="#value-types">boolean</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li></ul></td></tr><tr><td>readOnlyRootFilesystem</td><td>If true, mounts the container's root filesystem as read-only.</td><td><a href="#value-types">boolean</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li></ul></td></tr><tr><td>runAsUid</td><td>Specifies the Unix user id with which the container running the created workload should run.</td><td><a href="#value-types">integer</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>runasGid</td><td>Specifies the Unix Group ID with which the container should run.</td><td><a href="#value-types">integer</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>supplementalGroups</td><td>Comma separated list of groups that the user running the container belongs to, in addition to the group indicated by runAsGid.</td><td><a href="#value-types">string</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li></ul></td></tr><tr><td>allowPrivilegeEscalation</td><td>Allows the container running the workload and all launched processes to gain additional privileges after the workload starts</td><td><a href="#value-types">boolean</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li></ul></td></tr><tr><td>privileged</td><td>Grants the container full access to the host, bypassing almost all container isolation; the container acts like root. Default is false.<br><br>This parameter is governed by a system policy, which enforces <code>privileged: false</code> by default and marks it as non-editable (<code>canEdit: false</code>). Containers cannot run in privileged mode unless an administrator explicitly updates the system policy to allow it. See <a href="/pages/6r51yAr8J5lAAFDA67GO#system-policies">System policies</a> for more details.</td><td><a href="#value-types">boolean</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li></ul></td></tr><tr><td>hostIpc</td><td>Whether to enable hostIpc. Defaults to false.</td><td><a href="#value-types">boolean</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li></ul></td></tr><tr><td>hostNetwork</td><td>Whether to enable host network.</td><td><a href="#value-types">boolean</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li></ul></td></tr></tbody></table>

### Compute Fields

<table><thead><tr><th width="133.4375">Fields</th><th width="311.1015625">Description</th><th width="132.6640625">Value type</th><th width="187.5">Supported NVIDIA Run:ai workload type</th></tr></thead><tbody><tr><td>cpuCoreRequest</td><td>CPU units to allocate for the created workload (0.5, 1, .etc). The workload receives at least this amount of CPU. Note that the workload is not scheduled unless the system can guarantee this amount of CPUs to the workload.</td><td><a href="#value-types">number</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>cpuCoreLimit</td><td>Limitations on the number of CPUs consumed by the workload (0.5, 1, .etc). The system guarantees that this workload is not able to consume more than this amount of CPUs.</td><td><a href="#value-types">number</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>cpuMemoryRequest</td><td>The amount of CPU memory to allocate for this workload (1G, 20M, .etc). The workload receives at least this amount of memory. Note that the workload is not scheduled unless the system can guarantee this amount of memory to the workload</td><td><a href="#value-types">quantity</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>cpuMemoryLimit</td><td>Limitations on the CPU memory to allocate for this workload (1G, 20M, .etc). The system guarantees that this workload is not be able to consume more than this amount of memory. The workload receives an error when trying to allocate more memory than this limit.</td><td><a href="#value-types">quantity</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>largeShmRequest</td><td>A large /dev/shm device to mount into a container running the created workload (shm is a shared file system mounted on RAM).</td><td><a href="#value-types">boolean</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li></ul></td></tr><tr><td>gpuRequestType</td><td>Sets the unit type for GPU resources requests to either portion or memory. Only if <code>gpuDeviceRequest = 1</code>, the request type can be stated as <code>portion</code> or <code>memory</code>.</td><td><a href="#value-types">string</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>gpuPortionRequest</td><td>Specifies the fraction of GPU to be allocated to the workload, between 0 and 1. For backward compatibility, it also supports the number of gpuDevices larger than 1, currently provided using the gpuDevices field.</td><td><a href="#value-types">number</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>gpuDeviceRequest</td><td>Specifies the number of GPUs to allocate for the created workload. Only if <code>gpuDeviceRequest = 1</code>, the gpuRequestType can be defined.</td><td><a href="#value-types">integer</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>gpuPortionLimit</td><td>When a fraction of a GPU is requested, the GPU limit specifies the portion limit to allocate to the workload. The range of the value is from 0 to 1.</td><td><a href="#value-types">number</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>gpuMemoryRequest</td><td>Specifies GPU memory to allocate for the created workload. The workload receives this amount of memory. Note that the workload is not scheduled unless the system can guarantee this amount of GPU memory to the workload.</td><td><a href="#value-types">quantity</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>gpuMemoryLimit</td><td>Specifies a limit on the GPU memory to allocate for this workload. Should be no less than the gpuMemory.</td><td><a href="#value-types">quantity</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>extendedResources</td><td>Specifies values for extended resources. Extended resources are third-party devices (such as high-performance NICs, FPGAs, or InfiniBand adapters) that you want to allocate to your Job.</td><td><a href="#value-types">itemized</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li></ul></td></tr></tbody></table>

### Storage Fields

<table><thead><tr><th width="133.29296875">Fields</th><th width="310.96484375">Description</th><th width="133.31640625">Value type</th><th width="188.3984375">Supported NVIDIA Run:ai workload type</th></tr></thead><tbody><tr><td>dataVolume</td><td>Set of data volumes to use in the workload. Each data volume is mapped to a file-system mount point within the container running the workload.</td><td><a href="#value-types">itemized</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li></ul></td></tr><tr><td><a href="#hostpath-field-details">hostPath</a></td><td>Maps a folder to a file-system mount point within the container running the workload.</td><td><a href="#value-types">itemized</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li></ul></td></tr><tr><td><a href="#git-field-details">git</a></td><td>Details of the git repository and items mapped to it.</td><td><a href="#value-types">itemized</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li></ul></td></tr><tr><td><a href="#pvc-field-details">pvc</a></td><td>Specifies persistent volume claims to mount into a container running the created workload.</td><td><a href="#value-types">itemized</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li></ul></td></tr><tr><td><a href="#nfs-field-details">nfs</a></td><td>Specifies NFS volume to mount into the container running the workload.</td><td><a href="#value-types">itemized</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li></ul></td></tr><tr><td><a href="#s3-field-details">s3</a></td><td>Specifies S3 buckets to mount into the container running the workload.</td><td><a href="#value-types">itemized</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li></ul></td></tr><tr><td>configMapVolumes</td><td>Specifies ConfigMaps to mount as volumes into a container running the created workload.</td><td><a href="#value-types">itemized</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li></ul></td></tr><tr><td>secretVolume</td><td>Set of secret volumes to use in the workload. A secret volume maps a secret resource in the cluster to a file-system mount point within the container running the workload.</td><td><a href="#value-types">itemized</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li></ul></td></tr><tr><td><a href="#emptydirvolume-field-details">emptyDirVolume</a></td><td>A list of emptyDir volumes to mount in the workload.</td><td><a href="#value-types">itemized</a></td><td><ul><li>Workspace</li><li>Standard training</li><li>Distributed training</li><li>Inference</li><li>Distributed inference (API only)</li></ul></td></tr></tbody></table>

#### Storage Field Examples

<details>

<summary>hostPath Field Details</summary>

* **Description:** Maps a folder to a file system mount oint within the container running the workload
* **Value type:** [itemized](#value-types)
* **Example policy snippet:**

```yaml
defaults:
  storage:
    hostPath:
      instances:
        - path: h3-path-1
          mountPath: h3-mount-1
        - path: h3-path-2
          mountPath: h3-mount-2
      attributes:
        - readOnly: true
```

<table><thead><tr><th>hostPath fields</th><th width="264">Description</th><th>Value type</th></tr></thead><tbody><tr><td>name</td><td>Unique name to identify the instance. Primarily used for policy locked rules.</td><td><a href="#value-types">string</a></td></tr><tr><td>path</td><td>Local path within the controller to which the host volume is mapped.</td><td><a href="#value-types">string</a></td></tr><tr><td>readOnly</td><td>Force the volume to be mounted with read-only permissions. Defaults to false</td><td><a href="#value-types">boolean</a></td></tr><tr><td>mountPath</td><td><p>The path that the host volume is mounted to when in use. Enum:</p><ul><li>"None"</li><li>"HostToContainer"</li></ul></td><td><a href="#value-types">string</a></td></tr><tr><td>mountPropagation</td><td>Share this volume mount with other containers. If set to HostToContainer, this volume mount receives all subsequent mounts that are mounted to this volume or any of its subdirectories. In case of multiple hostPath entries, this field should have the same value for all of them.</td><td><a href="#value-types">string</a></td></tr></tbody></table>

</details>

<details>

<summary>Git Field Details</summary>

* **Description:** Details of the git repository and items mapped to it
* **Value type:** [itemized](#value-types)
* **Example policy snippet:**

```yaml
defaults:
  storage:
    git:
      attributes:
        Repository: https://runai.public.github.com
      instances
        - branch: "master"
          path: /container/my-repository
          passwordSecret: my-password-secret
```

<table><thead><tr><th>Git fields</th><th width="341">Description</th><th>Value type</th></tr></thead><tbody><tr><td>repository</td><td>URL to a remote git repository. The content of this repository is mapped to the container running the workload</td><td><a href="#value-types">string</a></td></tr><tr><td>revision</td><td>Specific revision to synchronize the repository from</td><td><a href="#value-types">string</a></td></tr><tr><td>path</td><td>Local path within the workspace to which the S3 bucket is mapped</td><td><a href="#value-types">string</a></td></tr><tr><td>secretName</td><td>Optional name of Kubernetes secret that holds your git username and password</td><td><a href="#value-types">string</a></td></tr><tr><td>username</td><td>If secretName is provided, this field should contain the key, within the provided Kubernetes secret, which holds the value of your git username. Otherwise, this field should specify your git username in plain text (example: myuser).</td><td><a href="#value-types">string</a></td></tr></tbody></table>

</details>

<details>

<summary>PVC Field Details</summary>

* **Description:** Specifies persistent volume claims to mount into a container running the created workload
* **Value type:** [itemized](#value-types)
* **Example policy snippet:**

```yaml
defaults:
  storage:
    pvc:
      instances:
        - claimName: pvc-staging-researcher1-home
          existingPvc: true
          path: /myhome
          readOnly: false
          claimInfo:
            accessModes:
              readWriteMany: true
```

<table><thead><tr><th>Spec PVC fields</th><th width="341">Description</th><th>Value type</th></tr></thead><tbody><tr><td>claimName (mandatory)</td><td>A given name for the PVC. Allowed referencing it across workspaces</td><td><a href="#value-types">string</a></td></tr><tr><td>ephemeral</td><td>Use <strong>true</strong> to set PVC to ephemeral. If set to <strong>true</strong>, the PVC is deleted when the workspace is stopped.</td><td><a href="#value-types">boolean</a></td></tr><tr><td>path</td><td>Local path within the workspace to which the PVC bucket is mapped</td><td><a href="#value-types">string</a></td></tr><tr><td>readonly</td><td>Permits read only from the PVC, prevents additions or modifications to its content</td><td><a href="#value-types">boolean</a></td></tr><tr><td>ReadwriteOnce</td><td>Requesting claim that can be mounted in read/write mode to exactly 1 host. If none of the modes are specified, the default is readWriteOnce.</td><td><a href="#value-types">boolean</a></td></tr><tr><td>size</td><td>Requested size for the PVC. Mandatory when existing PVC is false</td><td><a href="#value-types">string</a></td></tr><tr><td>storageClass</td><td>Storage class name to associate with the PVC. This parameter may be omitted if there is a single storage class in the system, or you are using the default storage class. Further details at <a href="https://kubernetes.io/docs/concepts/storage/storage-classes.">Kubernetes storage classes</a>.</td><td><a href="#value-types">string</a></td></tr><tr><td>readOnlyMany</td><td>Requesting claim that can be mounted in read-only mode to many hosts</td><td><a href="#value-types">boolean</a></td></tr><tr><td>readWriteMany</td><td>Requesting claim that can be mounted in read/write mode to many hosts</td><td><a href="#value-types">boolean</a></td></tr></tbody></table>

</details>

<details>

<summary>NFS Field Details</summary>

* **Description:** Specifies NFS volume to mount into the container running the workload
* **Value type:** [itemized](#value-types)
* **Example policy snippet:**

```yaml
defaults:
 storage:
   nfs:
     instances:
       - path: nfs-path
         readOnly: true
         server: nfs-server
         mountPath: nfs-mount
rules:
  storage:
    nfs:
      instances:
        canAdd: false
```

<table><thead><tr><th>nfs fields</th><th width="341">Description</th><th>Value type</th></tr></thead><tbody><tr><td>mountPath</td><td>The path that the NFS volume is mounted to when in use</td><td><a href="#value-types">string</a></td></tr><tr><td>path</td><td>Path that is exported by the NFS server</td><td><a href="#value-types">string</a></td></tr><tr><td>readOnly</td><td>Whether to force the NFS export to be mounted with read-only permissions</td><td><a href="#value-types">boolean</a></td></tr><tr><td>nfsServer</td><td>The hostname or IP address of the NFS server</td><td><a href="#value-types">string</a></td></tr></tbody></table>

</details>

<details>

<summary>S3 Field Details</summary>

* **Description:** Specifies S3 buckets to mount into the container running the workload
* **Value type:** [itemized](#value-types)
* **Example policy snippet:**

```yaml
defaults:
  storage:
    s3:
      instances:
        - bucket: bucket-opt-1
          path: /s3/path
          accessKeySecret: s3-access-key
          secretKeyOfAccessKeyId: s3-secret-id
          secretKeyOfSecretKey: s3-secret-key
      attributes:
        url: https://amazonaws.s3.com
```

<table><thead><tr><th>s3 fields</th><th width="341">Description</th><th>Value type</th></tr></thead><tbody><tr><td>Bucket</td><td>The name of the bucket</td><td><a href="#value-types">string</a></td></tr><tr><td>path</td><td>Local path within the workspace to which the S3 bucket is mapped</td><td><a href="#value-types">string</a></td></tr><tr><td>url</td><td>The URL of the S3 service provider. The default is the URL of the Amazon AWS S3 service</td><td><a href="#value-types">string</a></td></tr></tbody></table>

</details>

<details>

<summary>EmptyDirVolume Field Details</summary>

* **Description:** A list of emptyDir volumes to mount in the workload
* **Value type:** [itemized](#value-types)
* **Example policy snippet:**

```yaml
defaults:
  storage:
    emptyDirVolume:
      instances:
        - name: storage-instance-a
          path: /mnt/emptydir
          medium: ""  # Leave empty for disk-backed, or set to "Memory"
          sizeLimit: 1G
          exclude: false
```

<table><thead><tr><th>emptyDirVolume fields</th><th width="341">Description</th><th>Value type</th></tr></thead><tbody><tr><td>name</td><td>Unique name to identify the instance. Primarily used for policy locked rules.</td><td><a href="#value-types">string</a></td></tr><tr><td>path</td><td>Local path within the workload to which the emptyDir volume is mapped.</td><td><a href="#value-types">string</a></td></tr><tr><td>medium</td><td>The type of storage medium for the volume. Use <code>Memory</code> for memory-backed storage, or leave empty for disk-backed storage.</td><td><a href="#value-types">string</a></td></tr><tr><td>sizeLimit</td><td>The total amount of local storage or memory required for the emptyDir volume. Specify using Kubernetes quantity format (for example, <code>1G</code>, <code>500Mi</code>).</td><td><a href="#value-types">string</a></td></tr><tr><td>exclude</td><td>If set to true, excludes this volume from the workload.</td><td><a href="#value-types">boolean</a></td></tr></tbody></table>

</details>

### Serving Port Fields

<table><thead><tr><th width="133.30859375">Fields</th><th width="311.24609375">Description</th><th width="133.33984375">Value type</th><th width="187.6953125">Supported NVIDIA Run:ai workload type</th></tr></thead><tbody><tr><td>container / port</td><td>Specifies the port that the container running the inference service exposes</td><td><a href="#value-types">integer</a></td><td><ul><li>Inference</li><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>protocol</td><td><p>Specifies the protocol used by the port. Defaults to http.</p><p>Enum: "http", "grpc"</p></td><td><a href="#value-types">string</a></td><td><ul><li>Inference</li><li>Distributed inference (API only)</li></ul></td></tr><tr><td>authorizationType</td><td><p>Specifies the authorization type for serving port URL access. Defaults to public, which means no authorization is required. If set to authenticatedUsers, only authenticated NVIDIA Run:ai users are allowed to access the URL. If set to authorizedUsersOrGroups, only users or groups specified in authorizedUsers or authorizedGroups are allowed to access the URL. Supported from cluster version 2.19.</p><p>Enum: "public", "authenticatedUsers", "authorizedUsersOrGroups"</p></td><td><a href="#value-types">string</a></td><td><ul><li>Inference</li><li>Distributed inference (API only)</li></ul></td></tr><tr><td>authorizedUsers</td><td>Specifies the list of users that are allowed to access the URL. Note that authorizedUsers and authorizedGroups are mutually exclusive.</td><td><a href="#value-types">array</a></td><td><ul><li>Inference</li><li>Distributed inference (API only)</li></ul></td></tr><tr><td>authorizedGroups</td><td>Specifies the list of groups that are allowed to access the URL. Note that authorizedUsers and authorizedGroups are mutually exclusive.</td><td><a href="#value-types">array</a></td><td><ul><li>Inference</li><li>Distributed inference (API only)</li></ul></td></tr><tr><td>clusterLocalAccessOnly</td><td>Configures the serving port URL to be available only on the cluster-local network, and not externally. Defaults to false.</td><td><a href="#value-types">boolean</a></td><td>Inference</td></tr><tr><td>exposeExternally</td><td>Indicates whether the inference serving endpoint should be accessible outside the cluster. If set to true, the endpoint will be exposed externally. To enable external access, your administrator must configure the cluster as described in the <a href="/pages/Ody4Af15R7fOat8L9zlX#inference">inference requirements</a> section.</td><td><a href="#value-types">boolean</a></td><td><ul><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>exposedUrl</td><td>The custom URL to use for the serving port. If empty (default), an autogenerated URL will be used.</td><td><a href="#value-types">string</a></td><td><ul><li>Distributed inference (API only)</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>serviceType</td><td>The type of Kubernetes service to create for the inference deployment. Options include 'ClusterIP' (default), 'NodePort', 'LoadBalancer', and 'ExternalName'. Default: "ClusterIP"</td><td><a href="#value-types">string</a></td><td>NVIDIA NIM services (API only)</td></tr><tr><td>grpcPort</td><td>The GRPC port that the container running the inference service exposes.</td><td><a href="#value-types">integer</a></td><td>NVIDIA NIM services (API only)</td></tr><tr><td>metricsPort</td><td>The port where metrics are exposed, required only if it's different than the main port.</td><td><a href="#value-types">integer</a></td><td>NVIDIA NIM services (API only)</td></tr><tr><td>exposedProtocol</td><td>The protocol to use for the exposed URL. If grpcPort is set, this defaults to grpc. Otherwise, it defaults to http. Enum: "http" "grpc"</td><td><a href="#value-types">string</a></td><td>NVIDIA NIM services (API only)</td></tr></tbody></table>

### Autoscaling Fields

<table><thead><tr><th width="133.05859375">Fields</th><th width="311.234375">Description</th><th width="132.8515625">Value type</th><th width="187.83984375">Supported NVIDIA Run:ai workload type</th></tr></thead><tbody><tr><td>metricThresholdPercentage</td><td>Specifies the percentage of metric threshold value to use for autoscaling. Defaults to 70. Applicable only with the 'throughput' and 'concurrency' metrics.</td><td><a href="#value-types">number</a></td><td>Inference</td></tr><tr><td>minReplicas</td><td>Specifies the minimum number of replicas for autoscaling. Defaults to 1. Use 0 to allow scale-to-zero.</td><td><a href="#value-types">integer</a></td><td><ul><li>Inference</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>maxReplicas</td><td>Specifies the maximum number of replicas for autoscaling. Defaults to minReplicas, or to 1 if minReplicas is set to 0.</td><td><a href="#value-types">integer</a></td><td><ul><li>Inference</li><li>NVIDIA NIM services (API only)</li></ul></td></tr><tr><td>initialReplicas</td><td>Specifies the number of replicas to run when initializing the workload for the first time. Defaults to minReplicas, or to 1 if minReplicas is set to 0.</td><td><a href="#value-types">integer</a></td><td>Inference</td></tr><tr><td>activationReplicas</td><td>Specifies the number of replicas to run when scaling-up from zero. Defaults to minReplicas, or to 1 if minReplicas is set to 0.</td><td><a href="#value-types">integer</a></td><td>Inference</td></tr><tr><td>concurrencyHardLimit</td><td>Specifies the maximum number of requests allowed to flow to a single replica at any time. 0 means no limit.</td><td><a href="#value-types">integer</a></td><td>Inference</td></tr><tr><td>scaleToZeroRetentionSeconds</td><td>Specifies the minimum amount of time (in seconds) that the last replica will remain active after a scale-to-zero decision. Defaults to 0. Available only if minReplicas is set to 0.</td><td><a href="#value-types">integer</a></td><td>Inference</td></tr><tr><td>scaleDownDelaySeconds</td><td>Specifies the minimum amount of time (in seconds) that a replica will remain active after a scale-down decision</td><td><a href="#value-types">integer</a></td><td>Inference</td></tr><tr><td>scaleWindowSeconds</td><td>The time window for autoscaling decisions, in seconds. Defaults to 300 seconds.</td><td><a href="#value-types">integer</a></td><td>NVIDIA NIM services (API only)</td></tr><tr><td>metric</td><td>Specifies the metric to use for autoscaling. Mandatory if minReplicas &#x3C; maxReplicas, except for the special case where minReplicas is set to 0 and maxReplicas is set to 1, as in this case autoscaling decisions are made according to network activity rather than metrics. Use one of the built-in metrics of 'throughput', 'concurrency' or 'latency', or any other available custom metric. Only the 'throughput' and 'concurrency' metrics support scale-to-zero.</td><td><a href="#value-types">string</a></td><td>Inference</td></tr><tr><td>metricThreshold</td><td>Specifies the threshold to use with the specified metric for autoscaling. Mandatory if metric is specified.</td><td><a href="#value-types">integer</a></td><td><ul><li>Inference</li><li>NVIDIA NIM services (API only)</li></ul></td></tr></tbody></table>

### Serving Configuration Fields

<table><thead><tr><th width="133.41796875">Fields</th><th width="310.875">Description</th><th width="133.421875">Value type</th><th width="187.91015625">Supported NVIDIA Run:ai workload type</th></tr></thead><tbody><tr><td>initializationTimeoutSeconds</td><td>Specifies the maximum time (in seconds) allowed for a workload to initialize and become ready. If the workload does not start within this time, it will be moved to failed state.</td><td><a href="#value-types">integer</a></td><td>Inference</td></tr><tr><td>requestTimeoutSeconds<br></td><td>Specifies the maximum time (in seconds) allowed to process an end-user request. If no response is returned within this time, the request will be ignored.</td><td><a href="#value-types">integer</a></td><td>Inference</td></tr></tbody></table>

## Value Types

Each field has a specific value type. The following value types are supported.

<table><thead><tr><th width="132.99609375">Value type</th><th width="311.35546875">Description</th><th width="164.80078125">Supported rule type</th><th width="187">Defaults</th></tr></thead><tbody><tr><td>Boolean</td><td>A binary value that can be either True or False</td><td><ul><li><a href="#rule-types">canEdit</a></li><li><a href="#rule-types">required</a></li></ul></td><td>true/false</td></tr><tr><td>String</td><td>A sequence of characters used to represent text. It can include letters, numbers, symbols, and spaces</td><td><ul><li><a href="#rule-types">canEdit</a></li><li><a href="#rule-types">required</a></li><li><a href="#rule-types">options</a></li></ul></td><td>abc</td></tr><tr><td>Itemized</td><td>An ordered collection of items (objects), which can be of different types (all items in the list are of the same type). For further information see the chapter below the table.</td><td><ul><li><a href="#rule-types">canAdd</a></li><li><a href="#rule-types">locked</a></li></ul></td><td>See below</td></tr><tr><td>Integer</td><td>An Integer is a whole number without a fractional component.</td><td><ul><li><a href="#rule-types">canEdit</a></li><li><a href="#rule-types">required</a></li><li><a href="#rule-types">min</a></li><li><a href="#rule-types">max</a></li><li><a href="#rule-types">step</a></li><li><a href="#rule-types">defaultFrom</a></li></ul></td><td>100</td></tr><tr><td>Number</td><td>Capable of having non-integer values</td><td><ul><li><a href="#rule-types">canEdit</a></li><li><a href="#rule-types">required</a></li><li><a href="#rule-types">min</a></li><li><a href="#rule-types">defaultFrom</a></li></ul></td><td>10.3</td></tr><tr><td>Quantity</td><td>Holds a string composed of a number and a unit representing a quantity</td><td><ul><li><a href="#rule-types">canEdit</a></li><li><a href="#rule-types">required</a></li><li><a href="#rule-types">min</a></li><li><a href="#rule-types">max</a></li><li><a href="#rule-types">defaultFrom</a></li></ul></td><td>5M</td></tr><tr><td>Array</td><td>Set of values that are treated as one, as opposed to Itemized in which each item can be referenced separately.</td><td><ul><li><a href="#rule-types">canEdit</a></li><li><a href="#rule-types">required</a></li></ul></td><td><ul><li>node-a</li><li>node-b</li><li>node-c</li></ul></td></tr></tbody></table>

## Itemized

Workload fields of type itemized have multiple instances, however in comparison to objects, each can be referenced by a key field. The key field is defined for each field.

Consider the following workload spec:

```yaml
spec:
  image: ubuntu
  compute:
    extendedResources:
      - resource: added/cpu
        quantity: 10
      - resource: added/memory
        quantity: 20M
```

In this example, extendedResources have two instances, each has two attributes: resource (the key attribute) and quantity.

In policy, the defaults and rules for itemized fields have two sub sections:

* Instances: default items to be added to the policy or rules which apply to an instance as a whole.
* Attributes: defaults for attributes within an item or rules which apply to attributes within each item.

Consider the following example:

```yaml
defaults:
  compute:
    extendedResources:
      instances: 
        - resource: default/cpu
          quantity: 5
        - resource: default/memory
          quantity: 4M
      attributes:
        quantity: 3
rules:
  compute:
    extendedResources:
      instances:
        locked: 
          - default/cpu
      attributes:
        quantity: 
          required: true
```

Assume the following workload submission is requested:

```yaml
spec:
  image: ubuntu
  compute:
    extendedResources:
      - resource: default/memory
        exclude: true
      - resource: added/cpu
      - resource: added/memory
        quantity: 5M
```

The effective policy for the above mentioned workload has the following extendedResources instances:

<table><thead><tr><th width="155.73046875">Resource</th><th width="174.46875">Source of the instance</th><th width="153.49609375">Quantity</th><th width="263.421875">Source of the attribute quantity</th></tr></thead><tbody><tr><td>default/cpu</td><td>Policy defaults</td><td>5</td><td>The default of this instance in the policy defaults section</td></tr><tr><td>added/cpu</td><td>Submission request</td><td>3</td><td>The default of the quantity attribute from the attributes section</td></tr><tr><td>added/memory</td><td>Submission request</td><td>5M</td><td>Submission request</td></tr></tbody></table>

{% hint style="info" %}
**Note**

The default/memory is not populated to the workload, this is because it has been excluded from the workload using “exclude: true”.
{% endhint %}

A workload submission request cannot exclude the default/cpu resource, as this key is included in the locked rules under the instances section. {#a-workload-submission-request-cannot-exclude-the-default/cpu-resource,-as-this-key-is-included-in-the-locked-rules-under-the-instances-section.}

## Rule Types

<table><thead><tr><th width="156.01953125">Rule types</th><th width="414.84375">Description</th><th width="169.328125">Supported value types</th></tr></thead><tbody><tr><td>canAdd</td><td>Whether the submission request can add items to an itemized field other than those listed in the policy defaults for this field.</td><td><a href="#value-types">itemized</a></td></tr><tr><td>locked</td><td>Set of items that the workload is unable to modify or exclude. In this example, a workload policy default is given to HOME and USER, that the submission request cannot modify or exclude from the workload.</td><td><a href="#value-types">itemized</a></td></tr><tr><td>blocked</td><td>Blocks the field to prevent workloads from specifying any values for it.</td><td><ul><li><a href="#value-types">string</a></li><li><a href="#value-types">boolean</a></li><li><a href="#value-types">integer</a></li><li><a href="#value-types">number</a></li><li><a href="#value-types">quantity</a></li><li><a href="#value-types">array</a></li></ul></td></tr><tr><td>canEdit</td><td>Whether the submission request can modify the policy default for this field. In this example, it is assumed that the policy has default for imagePullPolicy. As canEdit is set to false, submission requests are not able to alter this default.</td><td><ul><li><a href="#value-types">string</a></li><li><a href="#value-types">boolean</a></li><li><a href="#value-types">integer</a></li><li><a href="#value-types">number</a></li><li><a href="#value-types">quantity</a></li><li><a href="#value-types">array</a></li></ul></td></tr><tr><td>required</td><td>When set to true, the workload must have a value for this field. The value can be obtained from policy defaults. If no value specified in the policy defaults, a value must be specified for this field in the submission request.</td><td><ul><li><a href="#value-types">string</a></li><li><a href="#value-types">boolean</a></li><li><a href="#value-types">integer</a></li><li><a href="#value-types">number</a></li><li><a href="#value-types">quantity</a></li><li><a href="#value-types">array</a></li></ul></td></tr><tr><td>min</td><td>The minimal value for the field</td><td><ul><li><a href="#value-types">integer</a></li><li><a href="#value-types">number</a></li><li><a href="#value-types">quantity</a></li></ul></td></tr><tr><td>max</td><td>The maximal value for the field</td><td><ul><li><a href="#value-types">integer</a></li><li><a href="#value-types">number</a></li><li><a href="#value-types">quantity</a></li></ul></td></tr><tr><td>step</td><td>The allowed gap between values for this field. In this example the allowed values are: 1, 3, 5, 7</td><td><ul><li><a href="#value-types">integer</a></li><li><a href="#value-types">number</a></li></ul></td></tr><tr><td>options</td><td>Set of allowed values for this field</td><td><a href="#value-types">string</a></td></tr><tr><td>defaultFrom</td><td>Set a default value for a field that will be calculated based on the value of another field</td><td><ul><li><a href="#value-types">integer</a></li><li><a href="#value-types">number</a></li><li><a href="#value-types">quantity</a></li></ul></td></tr></tbody></table>

### Rule Type Examples

<details>

<summary>canAdd</summary>

```yaml
storage:
  hostPath:
     instances:
       canAdd: false
```

</details>

<details>

<summary>locked</summary>

```yaml
storage:
  hostPath:
    Instances:
      locked:
        - HOME
        - USER
```

</details>

<details>

<summary>canEdit</summary>

```yaml
imagePullPolicy:
    canEdit: false
```

</details>

<details>

<summary>required</summary>

```yaml
image:
    required: true
```

</details>

<details>

<summary>min</summary>

```yaml
compute:
  gpuDevicesRequest:
    min: 3
```

</details>

<details>

<summary>max</summary>

```yaml
compute:
  gpuMemoryRequest:
     max: 2G
```

</details>

<details>

<summary>step</summary>

```yaml
compute:
  cpuCoreRequest:
    min: 1
    max: 7
    Step: 2
```

</details>

<details>

<summary>options</summary>

```yaml
image:
  options:
    - value: image-1
    - value: image-2
```

</details>

<details>

<summary>defaultFrom</summary>

```yaml
cpuCoreRequest:
  defaultFrom:
    field: compute.cpuCoreLimit
    factor: 0.5
```

</details>

## Policy Spec Sections

For each field of a specific policy, you can specify both rules and defaults. A policy spec consists of the following sections:

* Rules
* Defaults
* Imposed Assets

### Rules

Rules set up constraints on workload policy fields. For example, consider the following policy:

```yaml
rules:
  compute:
    gpuDevicesRequest: 
      max: 8
  security:
    runAsUid: 
      min: 500
```

Such a policy restricts the maximum value for gpuDeviceRequests to 8, and the minimal value for runAsUid, provided in the security section to 500.

### Defaults

The defaults section is used for providing defaults for various workload fields. For example, consider the following policy:

```yaml
defaults:
  imagePullPolicy: Always
  security:
    runAsNonRoot: true
    runAsUid: 500
```

Assume a submission request with the following values:

* Image: ubuntu
* runAsUid: 501

The effective workload that runs has the following set of values:

| Field                 | Value  | Source             |
| --------------------- | ------ | ------------------ |
| Image                 | Ubuntu | Submission request |
| ImagePullPolicy       | Always | Policy defaults    |
| security.runAsNonRoot | true   | Policy defaults    |
| security.runAsUid     | 501    | Submission request |

{% hint style="info" %}
**Note**

It is possible to specify a rule for each field, which states if a submission request is allowed to change the policy default for that given field, for example:

```yaml
defaults:
imagePullPolicy: Always
security:
    runAsNonRoot: true
    runAsUid: 500
 rules:
 security:
    runAsUid:
    canEdit: false
```

If this policy is applied, the submission request above fails, as it attempts to change the value of secuirty.runAsUid from 500 (the policy default) to 501 (the value provided in the submission request), which is forbidden due to canEdit rule set to false for this field.
{% endhint %}

### Imposed Assets

Default instances of a storage field can be provided using a datasource containing the details of this storage instance. To add such instances in the policy, specify those asset IDs in the imposedAssets section of the policy.

```yaml
defaults: null
rules: null
imposedAssets:
  - f12c965b-44e9-4ff6-8b43-01d8f9e630cc
```

Assets with references to credential assets (for example: private S3, containing reference to an AccessKey asset) cannot be used as imposedAssets.


# Supported Workload Types Policies

NVIDIA Run:ai workload policies can target any supported workload type registered in the platform (such as PyTorchJob, Deployment, and others). When you create a policy for a supported workload type, you can define its rules using the **Rule builder**, an interactive interface that maps the fields of the selected workload type, or by writing the policy YAML directly. For the list of workload types you can target, see [Supported workload types](/saas/workloads-in-nvidia-run-ai/workload-types/supported-workload-types).

{% hint style="info" %}
**Note**

* Workload policies are enabled by default. If you do not see it in the menu, contact your administrator to enable it under **General settings** → Workloads → Policies.
* Supported workload type policies can only be scoped to a project. Cluster, department, and system scopes are not available for these policies.
* For policies targeting native workload types (Workspace, Training, and Inference), see [NVIDIA Run:ai native workload policies](/saas/platform-management/policies/native-workload-policies).
* To apply policies to a new [workload type](/saas/workloads-in-nvidia-run-ai/workload-types/extending-workload-support) your organization has registered, you must first map it to the policy engine using the mapping APIs. See [Workload types policy mapping](/saas/platform-management/policies/supported-workload-type-policies/workload-type-policy-mapping).
  {% endhint %}

## Workload Policies Table

The Workload policies table can be found under **Policies** in the NVIDIA Run:ai platform.

The Workload policies table provides a list of all the policies defined in the platform, and allows you to manage them.

<figure><img src="/files/Xd3J56mPoTGBOklP31fT" alt=""><figcaption></figcaption></figure>

The Workload policies table consists of the following columns:

| Column              | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Policy              | The policy name which is a combination of the policy scope and the policy type                                                                                                                                                                                                                                                                                                                                                                                                 |
| Type                | The policy type is per NVIDIA Run:ai workload type. This allows administrators to set different policies for each [workload type](/saas/workloads-in-nvidia-run-ai/workload-types).                                                                                                                                                                                                                                                                                            |
| Policy architecture | Indicates whether the policy targets a **Standard** workload or a **Distributed** workload.                                                                                                                                                                                                                                                                                                                                                                                    |
| Status              | <p>The current state of the policy. Possible values:</p><ul><li><strong>Ready</strong> - the policy is active and enforcing workloads.</li><li><strong>Degraded</strong> - one or more elements referenced in the policy no longer exist in the workload type mapping (for example, because the mapping was changed after the policy was created). The policy remains in place but may not enforce correctly until the mapping is restored or the policy is updated.</li></ul> |
| Scope               | The scope the policy affects. Click the name of the scope to view the organizational tree diagram. You can only view the parts of the organizational tree for which you have permission to view.                                                                                                                                                                                                                                                                               |
| Created by          | The user who created the policy                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| Creation time       | The timestamp for when the policy was created                                                                                                                                                                                                                                                                                                                                                                                                                                  |
| Last updated        | The last time the policy was updated                                                                                                                                                                                                                                                                                                                                                                                                                                           |

### Customizing the Table View

* Filter - Click ADD FILTER, select the column to filter by, and enter the filter values
* Search - Click SEARCH and type the value to search by
* Sort - Click each column header to sort by
* Column selection - Click COLUMNS and select the columns to display in the table
* Refresh - Click REFRESH to update the table with the latest data

## Policy Structure

A rule has two parts: **target** (which field, spec selector, and item selector it applies to) and **behavior** (what enforcements it applies). The Rule builder walks you through these in sequence.

**Target:**

* **Field** - The configuration attribute you want to control, such as CPU limit, privilege escalation, or environment variables. NVIDIA Run:ai pre-maps all policy-controllable fields for each registered workload type, so the available fields depend on the workload type selected for the policy.
* **Spec selector** - Identifies which part or parts of the workload spec the rule applies to. Many workload types have more than one spec subtree; for example, a workload may have multiple containers, each with its own configuration. The spec selector narrows the rule to the relevant subtrees, such as "all containers" or "a specific container by name." Some selectors require additional input - once selected, the form shows the fields you need to fill in (for example, choosing `appContainerByName` reveals a `containerName` field). The available spec selectors depend on the field you selected.
* **Item selector** - Applies only to fields that hold a list of items, such as a list of environment variables or network ports. The item selector identifies which item within the list the rule targets, for example "all environment variables sourced from a ConfigMap" or "a port with a specific port number." Some selectors require additional input - once selected, the form shows the fields you need to fill in. The item selector step only appears when the chosen field is a list-type field.

**Behavior:**

* **Enforcement** - Defines how the rule constrains the field during workload submission. You can add multiple enforcements to a single rule. See [Enforcement types](#enforcement-types) for the full list.

### Enforcement Types

Each enforcement defines how the selected field is controlled during workload submission. The available enforcement types depend on the data type of the selected field - not all enforcement types are valid for all fields. The following enforcement types are available:

| Enforcement type | Description                                                                                                                                                                                                                        |
| ---------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Set              | Automatically applies a specific value to the field (supports multiple values for object fields).                                                                                                                                  |
| Default          | Sets an initial value that users can modify.                                                                                                                                                                                       |
| Required         | Mandates user input before saving or proceeding.                                                                                                                                                                                   |
| Blocked          | Disables the field to prevent data entry.                                                                                                                                                                                          |
| Options          | Limits input to a predefined choice list.                                                                                                                                                                                          |
| Range            | Restricts numeric values by minimum, maximum, and step increment.                                                                                                                                                                  |
| Quantity         | Enforces a Kubernetes resource quantity to fall within the inclusive range \[min, max], optionally restricted to increments of step starting from min. All bounds must be valid Kubernetes quantities such as 100m, 1Gi, or 500Mi. |

## Adding a Policy

To create a new supported workload type policy:

1. Click **+NEW POLICY**
2. In the **Scope** section, select a **project**
3. In the **Workload type** section, select a [supported workload type](/saas/workloads-in-nvidia-run-ai/workload-types/supported-workload-types)
4. Under **Set how the policy should be configured**, select:
   * **Rule builder** - Configure rules interactively using an interface that maps the fields of the selected workload type. See [Adding rules](#adding-rules).
   * **YAML editor** - Write or paste the policy YAML directly. See [Using the YAML editor](#using-the-yaml-editor).
5. Click **SAVE POLICY**

### Adding Rules

The **Rule builder** lets you add rules that set the limits and suggested defaults shown during workload submission.

To add a rule:

1. In the **Rules** section, click **+ RULE**
2. Under **Select a field from the spec**, choose a **Field**. The available fields are mapped from the spec of the selected workload type.
3. Under **Set the spec selector and its parameters**, choose a **Spec selector**. If the selector requires additional input, fill in the fields that appear below it.
4. When applicable, under **Set the item selector and its parameters**, choose an **Item selector**. If the selector requires additional input, fill in the fields that appear below it. The item selector step only appears for fields that require it.
5. Under **Add enforcements and set their values**, click **+ ENFORCEMENT**, select an **Enforcement type**, and set its values. You can add multiple enforcements. The enforcement section becomes available once the required fields, spec selector, and item selector (if applicable) are configured. See [Enforcement types](#enforcement-types) for more details.
6. Click **SAVE RULE**
7. Repeat for additional rules, then click **SAVE POLICY**

{% hint style="info" %}
**Note**

All fields in a rule are mandatory. If a rule has unresolved issues when you save it, the form prompts you to review and fix them, and the rule cannot be saved until they are resolved.
{% endhint %}

### Using the YAML Editor

The YAML editor accepts a `rules` array that defines the same constraints available in the Rule builder. Each rule targets a field using the same four components described in [Policy structure](#policy-structure): Field, Spec selector, optional Item selector, and Enforcement.

Rule structure:

```yaml
rules:
  - field: <field-name>
    specSelector: <spec-selector>
    itemSelector: <item-selector>    # optional; only for list-type fields
    selectorParams:                   # optional; required by some selectors
      <param>: <value>
    enforcements:
      - <enforcement-type>: <value>
```

To discover the fields and spec selectors available for a specific workload type, use the following APIs:

* [`POST /api/v1/policies/mappings/minimal`](https://run-ai-docs.nvidia.com/api/policies/policy#post-api-v1-policies-mappings-minimal) - returns minimal information about the mapping for the workload type, including the list of fields and the spec selectors for each field
* [`POST /api/v1/policies/mappings/field-info`](https://run-ai-docs.nvidia.com/api/policies/policy#post-api-v1-policies-mappings-field-info) - returns detailed field information including item selectors and formatter information

See [Enforcement types](#enforcement-types) for the full list of enforcement types and their value formats. For worked examples, see [Policy YAML examples](/saas/platform-management/policies/supported-workload-type-policies/policy-yaml-examples-supported).

## Editing a Policy

1. Select the policy you want to edit
2. Click **EDIT**
3. Update the policy and click **APPLY**
4. Click **SAVE POLICY**

## Viewing a Policy

To view a policy:

1. Select the policy you want to view.
2. Click **VIEW POLICY**
3. In the **Rules** section, view the rules defined for the policy. Each rule shows:
   * **Field**\
     The configuration attribute the rule targets
   * **Spec selector**\
     The part of the workload spec the rule applies to
   * **Item selector**\
     The specific item within a list-type field the rule targets. Shown only when the field is a list-type field.
   * **Rule**\
     The enforcements applied to the field

## Deleting a Policy

1. Select the policy you want to delete
2. Click **DELETE**
3. On the dialog, click **DELETE** to confirm the deletion

## Using API

Go to the [Policies](https://run-ai-docs.nvidia.com/api/policies/policy) API reference to view the available actions.


# Workload Types Policy Mapping

NVIDIA Run:ai provides policy mappings for many supported workload types. For those types, you can create workload policies directly. See [Supported workload types policies](/saas/platform-management/policies/supported-workload-type-policies).

This document describes how to register mapping resources for a workload type that is not yet mapped, or for a new workload type you are adding to the platform using the [Workload Types API](https://run-ai-docs.nvidia.com/api/workloads/workload-properties#post-api-v1-workload-types).

{% hint style="info" %}
**Note**

* Workload type policy mapping is configured via API only.
* All mapping APIs are currently experimental.
* To register mapping resources, follow the instructions in this document or [contact NVIDIA Run:ai support](https://www.nvidia.com/en-eu/support/enterprise/#contact-us).
* To check which workload types are mapped, see [Manage mappings](#managing-mappings).
  {% endhint %}

## How It Works

{% hint style="info" %}
**Note**

For workload types that use a standard Kubernetes pod template, NVIDIA Run:ai provides system-provided formatters and field groups (`podSpec`, `podSpecContainer`, `podSpecAndContainer`, `metadata`) that cover standard Kubernetes fields out of the box. In this case you only need to create the spec mapping and can skip the formatter and field group steps.
{% endhint %}

Three building blocks work together to enable policies for a new workload type. The dependency order is: formatters are referenced by field groups, which are referenced by spec mappings.

| Resource         | Purpose                                                                                                                                              | Endpoint                        |
| ---------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------- |
| **Formatter**    | Defines how to read and write a specific field value in the CRD spec using JQ expressions                                                            | `/api/v1/policies/formatters`   |
| **Field group**  | Bundles policy-controllable fields together; each field references a formatter (or a direct spec path) and lists the item selectors that apply to it | `/api/v1/policies/field-groups` |
| **Spec mapping** | Connects your workload type (GVK) to spec selectors (JQ queries that locate subtrees of the CRD spec), each linked to one or more field groups       | `/api/v1/policies/mappings`     |

## Before You Begin

* Your new workload type must be registered using the [Workload Types API](https://run-ai-docs.nvidia.com/api/workloads/workload-properties#post-api-v1-workload-types). See [Extending workload support with Karta](/saas/workloads-in-nvidia-run-ai/workload-types/extending-workload-support).
* Identify the **group**, **version**, and **kind** (GVK) of your workload type as defined in its CRD.
* You have the `policies:create` permission at the **tenant** scope. Department or project administrators with policy creation permissions scoped to a department, cluster, or project cannot create mappings.

## Creating Formatters

A formatter defines how to read and write a single policy-controllable field using JQ expressions. **In most cases, you do not need to create a custom formatter.** NVIDIA Run:ai provides two built-in formatters that cover the most common field structures:

* **`KeyValue`** - for simple key-value fields, such as `image` or `privileged`.
* **`KeyValueMap`** - for key-value map fields, such as `labels` or `options`.

To reference these in a field group, see [Creating Field Groups](#creating-field-groups).

Create a custom formatter only when the built-in formatters do not cover your field's structure, for example when reading or writing a field requires custom JQ logic.

Send a `POST` request to `/api/v1/policies/formatters`.

**Request body:**

```json
{
  "name": "<formatter-name>",
  "description": "<optional>",
  "fields": {
    "<field-name>": {
      "getExpr": "<jq-expression>",
      "setExpr": "<jq-expression>",
      "argsInfo": {
        "<arg-name>": "<type>"
      },
      "argsEnum": {
        "<arg-name>": ["<value>"]
      },
      "arrayKey": {
        "argName": "<arg-name>"
      }
    }
  }
}
```

| Field                          | Required | Description                                                                                                                                                                                                                        |
| ------------------------------ | -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `name`                         | Yes      | Unique formatter name (alphanumeric with optional internal hyphens)                                                                                                                                                                |
| `fields.<field-name>.getExpr`  | No       | JQ expression that reads the field value from the CRD spec. Required if you want the field to support `set` or `default` enforcements.                                                                                             |
| `fields.<field-name>.setExpr`  | No       | JQ expression that writes the field value into the CRD spec. Required if you want the field to support `set` or `default` enforcements.                                                                                            |
| `fields.<field-name>.argsInfo` | No       | Argument names and their types (`string`, `integer`, `boolean`, `number`, `quantity`, `ipaddress`). Argument names must match the names referenced in `getExpr` and `setExpr` using standard JQ notation: `$args.named.<argName>`. |
| `fields.<field-name>.argsEnum` | No       | Allowed values for string-typed arguments                                                                                                                                                                                          |
| `fields.<field-name>.arrayKey` | No       | The unique key argument per item, for itemized (array) fields                                                                                                                                                                      |

**Example: formatter for a ConfigMap volume field**

This field requires a custom formatter because writing it involves constructing a nested Kubernetes volume object from multiple arguments. A `KeyValue` or `specRef` approach cannot handle this structure.

```json
{
  "name": "podSpecVolumes",
  "description": "Formatter for pod spec volume fields",
  "fields": {
    "configMapVolumes": {
      "setExpr": "{ \"volumes\": [{ \"name\": $ARGS.named.name, \"configMap\": { \"name\": $ARGS.named.configMapName, \"defaultMode\": $ARGS.named.defaultMode, \"items\": $ARGS.named.items, \"optional\": $ARGS.named.optional } }] }",
      "arrayKey": {
        "argName": "name"
      },
      "argsInfo": {
        "name": "string",
        "configMapName": "string",
        "defaultMode": "integer",
        "items": "object",
        "optional": "boolean"
      }
    }
  }
}
```

A successful response returns the created formatter with its `id` and `scope: tenant`.

To retrieve available system-provided formatters for reference: `GET /api/v1/policies/formatters`

{% hint style="info" %}
**Note**

System-provided formatters cannot be modified. You can only update formatters that you have created.
{% endhint %}

## Creating Field Groups

A field group bundles policy-controllable fields. Each field either references a formatter by name or provides a direct path (`specRef`) into the spec for simple scalar fields. Field groups also define item selectors (JQ queries used to filter specific items within array fields).

Send a `POST` request to `/api/v1/policies/field-groups`.

**Request body:**

```json
{
  "name": "<field-group-name>",
  "description": "<optional>",
  "fields": {
    "<field-name>": {
      "formatter": "<formatter-name>",
      "description": "<optional>",
      "type": "<field-type>",
      "itemSelectors": ["<item-selector-name>"]
    }
  },
  "itemSelectors": {
    "<item-selector-name>": {
      "description": "<optional>",
      "query": "<jq-expression>",
      "argsInfo": {
        "<arg-name>": "<type>"
      }
    }
  }
}
```

| Field                               | Required | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| ----------------------------------- | -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `name`                              | Yes      | Unique field group name                                                                                                                                                                                                                                                                                                                                                                                                                                                         |
| `fields.<field-name>.formatter`     | No\*     | Name of the formatter that reads and writes this field. Use this or `specRef`. Must be either a built-in formatter (`KeyValue` or `KeyValueMap`) or the name of a formatter you created using `POST /api/v1/policies/formatters`.                                                                                                                                                                                                                                               |
| `fields.<field-name>.specRef`       | No\*     | Dot-separated path to the field relative to the spec element targeted by the spec selector. By default, the system assumes the field is found directly under that spec element and shares the same name. Use `specRef` when the field is nested deeper or has a different name in the spec. For example, if the field is named `privileged` in the policy but sits under `securityContext` in the spec, set `specRef` to `securityContext.privileged`. Use this or `formatter`. |
| `fields.<field-name>.type`          | No       | Field data type. One of: `string`, `boolean`, `integer`, `number`, `quantity`, `ipaddress`, `strings`, `integers`, `numbers`, `namesArray`, `object`                                                                                                                                                                                                                                                                                                                            |
| `fields.<field-name>.itemSelectors` | No       | Item selectors that apply to this field. Only needed when the field is stored in the spec as an array of items.                                                                                                                                                                                                                                                                                                                                                                 |
| `itemSelectors.<name>.query`        | No       | JQ expression that filters specific items within an array field                                                                                                                                                                                                                                                                                                                                                                                                                 |

**Example: field group for a custom job workload type**

```json
{
  "name": "acmeJobSpec",
  "description": "Policy-controllable fields for AcmeJob",
  "fields": {
    "parallelism": {
      "formatter": "KeyValue",
      "description": "Maximum number of pods running in parallel",
      "type": "integer"
    },
    "privileged": {
      "specRef": "securityContext.privileged",
      "description": "Whether the container runs in privileged mode",
      "type": "boolean"
    }
  }
}
```

A successful response returns the created field group with its `id` and `scope: tenant`.

To retrieve available system-provided field groups for reference: `GET /api/v1/policies/field-groups`

{% hint style="info" %}
**Note**

Unlike formatters, you can modify system-provided field groups by adding fields or item selectors to them. You cannot modify system-provided formatters; you can only create new formatters at the tenant scope.
{% endhint %}

## Creating the Spec Mapping

A spec mapping connects your workload type (GVK) to spec selectors. Each spec selector is a JQ query that locates a subtree of the CRD spec and links it to one or more field groups, exposing those fields for policy rules at that location.

Send a `POST` request to `/api/v1/policies/mappings`.

**Request body:**

```json
{
  "workloadType": {
    "group": "<crd-group>",
    "version": "<crd-version>",
    "kind": "<crd-kind>"
  },
  "description": "<optional>",
  "specSelectors": {
    "<selector-name>": {
      "description": "<optional>",
      "query": "<jq-query>",
      "fieldGroups": ["<field-group-name>"],
      "dependency": {
        "specSelector": "<parent-spec-selector-name>",
        "path": "<dot-separated-path>"
      }
    }
  }
}
```

| Field                                          | Required | Description                                                                                                                                                                                                                                                  |
| ---------------------------------------------- | -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `workloadType.group`                           | No       | The API group of the CRD, for example `acme.io`. Omit for core Kubernetes types that have no group, such as `Pod` (`v1`).                                                                                                                                    |
| `workloadType.version`                         | Yes      | The API version, for example `v1`                                                                                                                                                                                                                            |
| `workloadType.kind`                            | Yes      | The resource kind, for example `AcmeJob`                                                                                                                                                                                                                     |
| `specSelectors.<name>.query`                   | Yes      | JQ query that selects a subtree of the CRD spec                                                                                                                                                                                                              |
| `specSelectors.<name>.fieldGroups`             | Yes      | Field groups whose fields are exposed at this spec location                                                                                                                                                                                                  |
| `specSelectors.<name>.dependency.specSelector` | No       | The name of a parent spec selector that this selector depends on. Together with `path`, this forms a dependency chain the engine traces back to ensure all intermediate spec nodes exist before applying `set` or `default` rules.                           |
| `specSelectors.<name>.dependency.path`         | No       | Dot-separated path that must exist under the parent spec selector's target before this selector's `set` or `default` rules are applied. For example, if the parent selector targets `spec.template`, a `path` of `spec` ensures `spec.template.spec` exists. |

**Example: spec mapping with system and custom field groups**

```json
{
  "workloadType": {
    "group": "acme.io",
    "version": "v1",
    "kind": "AcmeJob"
  },
  "description": "AcmeJob workload type policy mapping",
  "specSelectors": {
    "acmeJobSpec": {
      "description": "the job (top) spec",
      "query": ".spec",
      "fieldGroups": ["acmeJobSpec"],
      "dependency": {
        "path": "spec"
      }
    },
    "acmePodSpec": {
      "description": "the pod spec",
      "query": ".spec.template.spec",
      "fieldGroups": ["podSpec", "podSpecAndContainer"],
      "dependency": {
        "specSelector": "acmeJobSpec",
        "path": "template.spec"
      }
    },
    "allContainers": {
      "description": "all app and init containers",
      "query": "(\n  .spec.template.spec.containers[]?,\n  .spec.template.spec.initContainers[]?\n)\n| select(.)",
      "fieldGroups": ["podSpecContainer", "podSpecAndContainer"]
    }
  }
}
```

A successful response returns the created mapping with its `id` and `scope: tenant`.

{% hint style="info" %}
**Note**

As a best practice, include an `allContainers` spec selector in any workload mapping. This ensures container-level fields (such as image, resource limits, and environment variables) are exposed for policy rules across all app and init containers.
{% endhint %}

## Managing Formatters

| Operation           | Method   | Endpoint                                      |
| ------------------- | -------- | --------------------------------------------- |
| List all formatters | `GET`    | `/api/v1/policies/formatters`                 |
| Get a formatter     | `GET`    | `/api/v1/policies/formatters/{FormatterName}` |
| Update a formatter  | `PATCH`  | `/api/v1/policies/formatters/{FormatterName}` |
| Delete a formatter  | `DELETE` | `/api/v1/policies/formatters/{FormatterName}` |

## Managing Field Groups

| Operation             | Method   | Endpoint                                         |
| --------------------- | -------- | ------------------------------------------------ |
| List all field groups | `GET`    | `/api/v1/policies/field-groups`                  |
| Get a field group     | `GET`    | `/api/v1/policies/field-groups/{FieldGroupName}` |
| Update a field group  | `PATCH`  | `/api/v1/policies/field-groups/{FieldGroupName}` |
| Delete a field group  | `DELETE` | `/api/v1/policies/field-groups/{FieldGroupName}` |

## Managing Mappings

| Operation               | Method   | Endpoint                                                                            |
| ----------------------- | -------- | ----------------------------------------------------------------------------------- |
| List all mappings       | `GET`    | `/api/v1/policies/mappings`                                                         |
| Filter by workload type | `GET`    | `/api/v1/policies/mappings?filterBy=group==<group>,kind==<kind>,version==<version>` |
| Update a mapping        | `PATCH`  | `/api/v1/policies/mappings`                                                         |
| Delete a mapping        | `DELETE` | `/api/v1/policies/mappings`                                                         |

To delete a mapping, provide the GVK in the request body:

```json
{
  "workloadType": {
    "group": "acme.io",
    "version": "v1",
    "kind": "AcmeJob"
  }
}
```

## Next Steps

Once the spec mapping is registered, create a policy for your new workload type the same way as any other supported workload type. See [Supported workload types policies](/saas/platform-management/policies/supported-workload-type-policies#adding-a-policy).


# Policy YAML Examples

This page provides YAML examples for supported workload type policies. These policies use a flat `rules` array where each rule specifies the field to control, the spec selector that identifies which part of the workload spec the rule applies to, and one or more enforcements.

For a description of all rule components, see [Supported workload type policies](/saas/platform-management/policies/supported-workload-type-policies).

## Rule Structure

```yaml
rules:
  - field: <field-name>
    specSelector: <spec-selector>
    itemSelector: <item-selector>    # optional; only for list-type fields
    selectorParams:                   # optional; required by some selectors
      <param>: <value>
    enforcements:
      - <enforcement-type>: <value>
```

## Deployment: Container Security Baseline

Requires an image on all containers, blocks privilege escalation, and caps CPU and memory limits.

```yaml
rules:
  - field: image
    specSelector: allContainers
    enforcements:
      - required: true
  - field: allowPrivilegeEscalation
    specSelector: allContainers
    enforcements:
      - set:
          value: false
  - field: resourceLimits
    specSelector: allContainers
    enforcements:
      - set:
          key: cpu
          value: "2"
      - set:
          key: memory
          value: "4Gi"
```

## PyTorchJob: Distributed Training Guardrails

Sets resource limits across all containers, pins the master replica count to 1, restricts the worker replica range, and limits retries on the job.

```yaml
rules:
  - field: resourceLimits
    specSelector: allContainers
    enforcements:
      - set:
          key: cpu
          value: "4"
      - set:
          key: memory
          value: "8Gi"
  - field: replicas
    specSelector: masterReplicaSpec
    enforcements:
      - set:
          value: 1
  - field: replicas
    specSelector: workerReplicaSpec
    enforcements:
      - range:
          min: 1
          max: 8
  - field: runPolicy-BackoffLimit
    specSelector: workloadSpec
    enforcements:
      - set:
          value: 3
```


# Scheduling Rules

Scheduling rules are restrictions applied to workloads. These restrictions apply to either the resources (nodes) on which workloads can run or the duration of the run time. Scheduling rules are set for [Projects](/saas/platform-management/aiinitiatives/organization/projects) or [Departments](/saas/platform-management/aiinitiatives/organization/departments) and apply to specific workload types. Once scheduling rules are set for a project or department, all matching workloads associated with the project have the restrictions applied to them, as defined, when the workload was submitted. New scheduling rules added to a project are not applied over previously created workloads associated with that project.

## Workload Duration (Time Limit)

This rule limits the duration of a workload run time. Workload run time is calculated as the total time in which the workload was in status Running. You can apply a single rule per workload type - Preemptible Workspaces, Non-preemptible Workspaces, and Training.

## Idle GPU Time Limit

This rule limits the total GPU time of a workload. Workload idle time is counted from the first time the workload is in status Running and the GPU was idle. Idleness is calculated by employing the `runai_gpu_idle_seconds_per_workload` metric. This metric determines the total duration of zero GPU utilization within each 30-second interval. If the GPU remains idle throughout the 30-second window, 30 seconds are added to the idleness sum; otherwise, the idleness count is reset. You can apply a single rule per workload type - “Preemptible” Workspaces, “Non-preemptible” Workspaces, and Training.

{% hint style="info" %}
**Note**

To make Idle GPU timeout effective, it must be set to a shorter duration than the workload duration of the same workload type.
{% endhint %}

## Node Type (Affinity)

Node type is used to select a group of nodes, typically with specific characteristics such as a hardware feature, storage type, fast networking interconnection, etc. The [Scheduler](/saas/platform-management/runai-scheduler/scheduling/how-the-scheduler-works) uses node type as an indication of which nodes should be used for your workloads, within this project.

Node type is a label in the form of `run.ai/type` and a value (e.g. run.ai/type = dgx200) that the administrator uses to tag a set of nodes. Adding the node type to the project’s scheduling rules mandates the user to submit workloads with a node type label/value pairs from this list, according to the workload type - Workspace or Training. The Scheduler then schedules workloads using a node selector, targeting nodes tagged with the NVIDIA Run:ai node type label/value pair. Node pools and a node type can be used in conjunction. For example, specifying a node pool and a smaller group of nodes from that node pool that includes a fast SSD memory or other unique characteristics.

### Labelling Nodes for Node Types Grouping

The administrator should use a node label with the key of `run.ai/type` and any coupled value

To assign a label to nodes you want to group, set the ‘node type (affinity)’ on each relevant node:

1. Obtain the list of nodes and their current labels by copying the following to your terminal:

```bash
kubectl get nodes --show-labels
```

2. Annotate a specific node with a new label by copying the following to your terminal:

```bash
kubectl label node <node-name> run.ai/type=<value>
```


# Monitor Performance and Health


# Before You Start

NVIDIA Run:ai provides [metrics and telemetry](/saas/platform-management/monitor-performance/metrics) for both physical cluster entities such as clusters, nodes, and node pools and application organization entities such as departments and projects. Metrics represent over-time data while telemetry represents current analytics data. This data is essential for monitoring and analyzing the performance and health of your platform.

## Consuming Metrics and Telemetry Data

Users can consume the data based on their permissions:

1. **API** - Access the data programmatically through the [NVIDIA Run:ai API](https://run-ai-docs.nvidia.com/api/).
2. **CLI** - Use the NVIDIA Run:ai [Command Line Interface](/saas/reference/cli) to query and manage the data.
3. **UI** - Visualize the data through the NVIDIA Run:ai user interface.

### **API**

* **Metrics API** - Access over-time detailed analytics data programmatically.
* **Telemetry API** - Access current analytics data programmatically.

Refer to [metrics and telemetry](/saas/platform-management/monitor-performance/metrics) to see the full list of supported metrics and telemetry APIs.

### **CLI**

Use the `list` and `describe` commands to fetch and manage the data. See [CLI reference](/saas/reference/cli/runai) for more details.

<figure><img src="/files/ZRNkJ3EaYtx7OxCQaZYd" alt=""><figcaption><p>Describe a specific workload telemetry</p></figcaption></figure>

<figure><img src="/files/fA2D3B4o9XhB9yAXT3et" alt=""><figcaption><p>List projects and view their telemetry and metrics</p></figcaption></figure>

### **UI Views**

Refer to [metrics and telemetry](/saas/platform-management/monitor-performance/metrics) to see the full list of supported metrics and telemetry.

* **Overview dashboard** - Provides a high-level summary of the cluster's health and performance, including key metrics such as GPU utilization, memory usage, and node status. Allows administrators to quickly identify any potential issues or areas for optimization. Offers advanced analytics capabilities for analyzing GPU usage patterns and identifying trends. Helps administrators optimize resource allocation and improve cluster efficiency.
* **Quota management** - Enables administrators to monitor and manage GPU quotas across the cluster. Includes features for setting and adjusting quotas, tracking usage, and receiving alerts when quotas are exceeded.
* **Workload visualizations** - Provides detailed insights into the resource usage and utilization of each GPU in the cluster. Includes metrics such as GPU memory utilization, core utilization, and power consumption. Allows administrators to identify GPUs that are under-utilized and overloaded.
* **Node and node pool visualizations** - Similar to workload visualizations, but focused on the resource usage and utilization of each GPU within a specific node or node pool. Helps administrators identify potential issues or bottlenecks at the node level.
* **Advanced NVIDIA metrics** - Provides access to a range of advanced NVIDIA metrics, such as GPU temperature, fan speed, and voltage. Enables administrators to monitor the health and performance of GPUs in greater detail. This data is available at the node and workload level. To enable these metrics, contact NVIDIA Run:ai customer support.


# Monitor Workloads by Category

A workload category represents the role or purpose of a workload such as training, building, or deploying models. Each workload type is automatically assigned a default category to ensure consistent classification across the platform.

Categories appear in the Overview dashboard, allowing administrators to filter, group, and monitor workloads based on their function. Administrators can modify the default category mapping for a workload type using the NVIDIA Run:ai API.

## Default Category Mapping

NVIDIA Run:ai defines the following default mappings of workload types to categories. To retrieve the default category per workload type, refer to the [List workload types](https://run-ai-docs.nvidia.com/api/workloads/workload-properties#get-api-v1-workload-types) API.

{% hint style="info" %}
**Note**

* For more information on workload support, see [Introduction to workloads](/saas/workloads-in-nvidia-run-ai/introduction-to-workloads).
* To see the default priority assigned to each of the workload types listed below, refer to [Workload priority control](/saas/platform-management/runai-scheduler/scheduling/workload-priority-control).
  {% endhint %}

### NVIDIA Run:ai Native Workloads

<table><thead><tr><th>Workload Type</th><th data-type="checkbox">Build</th><th data-type="checkbox">Train</th><th data-type="checkbox">Deploy</th></tr></thead><tbody><tr><td>Workspaces</td><td>true</td><td>false</td><td>false</td></tr><tr><td>Standard training</td><td>false</td><td>true</td><td>false</td></tr><tr><td>Distributed training</td><td>false</td><td>true</td><td>false</td></tr><tr><td>Custom inference</td><td>false</td><td>false</td><td>true</td></tr><tr><td>NVIDIA NIM inference</td><td>false</td><td>false</td><td>true</td></tr><tr><td>Hugging Face inference</td><td>false</td><td>false</td><td>true</td></tr><tr><td>Distributed inference</td><td>false</td><td>false</td><td>true</td></tr></tbody></table>

### Supported Workload Types

<table><thead><tr><th>Workload Type</th><th data-type="checkbox">Build</th><th data-type="checkbox">Train</th><th data-type="checkbox">Deploy</th></tr></thead><tbody><tr><td>AMLJob</td><td>false</td><td>true</td><td>false</td></tr><tr><td>CronJob</td><td>false</td><td>false</td><td>true</td></tr><tr><td>Deployment</td><td>false</td><td>false</td><td>true</td></tr><tr><td>DevWorkspace</td><td>true</td><td>false</td><td>false</td></tr><tr><td>DynamoGraphDeployment</td><td>false</td><td>false</td><td>true</td></tr><tr><td>InferenceService (KServe)</td><td>false</td><td>false</td><td>true</td></tr><tr><td>JAXJob</td><td>false</td><td>true</td><td>false</td></tr><tr><td>Job</td><td>true</td><td>false</td><td>false</td></tr><tr><td>JobSet</td><td>false</td><td>true</td><td>false</td></tr><tr><td>LeaderWorkerSet (LWS)</td><td>false</td><td>false</td><td>true</td></tr><tr><td>MPIJob</td><td>false</td><td>true</td><td>false</td></tr><tr><td>NIMCache</td><td>false</td><td>false</td><td>true</td></tr><tr><td>NIMServices</td><td>false</td><td>false</td><td>true</td></tr><tr><td>NodeSet (Slinky)</td><td>false</td><td>false</td><td>true</td></tr><tr><td>Notebook</td><td>true</td><td>false</td><td>false</td></tr><tr><td>PipelineRun</td><td>false</td><td>false</td><td>true</td></tr><tr><td>Pod</td><td>false</td><td>false</td><td>true</td></tr><tr><td>PyTorchJob</td><td>false</td><td>true</td><td>false</td></tr><tr><td>RayCluster</td><td>false</td><td>true</td><td>false</td></tr><tr><td>RayJob</td><td>false</td><td>true</td><td>false</td></tr><tr><td>RayService</td><td>false</td><td>false</td><td>true</td></tr><tr><td>ReplicaSet</td><td>false</td><td>false</td><td>true</td></tr><tr><td>ScheduledWorkflow</td><td>false</td><td>false</td><td>true</td></tr><tr><td>SeldonDeployment</td><td>false</td><td>false</td><td>true</td></tr><tr><td>Service</td><td>false</td><td>false</td><td>true</td></tr><tr><td>SPOTRequest</td><td>false</td><td>false</td><td>false</td></tr><tr><td>StatefulSet</td><td>false</td><td>false</td><td>true</td></tr><tr><td>TaskRun</td><td>true</td><td>false</td><td>false</td></tr><tr><td>TFJob</td><td>false</td><td>true</td><td>false</td></tr><tr><td>VirtualMachineInstance</td><td>false</td><td>true</td><td>false</td></tr><tr><td>Workflow</td><td>false</td><td>false</td><td>true</td></tr><tr><td>XGBoostJob</td><td>false</td><td>true</td><td>false</td></tr></tbody></table>

## Update the Default Category Mapping

Administrators can change the default category assigned to a workload type by updating the category mapping using the [NVIDIA Run:ai API](https://run-ai-docs.nvidia.com/api/). To update the category mapping:

1. Retrieve the list of workload types and their IDs using `GET /api/v1/workload-types`.
2. Identify the `workloadTypeId` of the workload type you want to modify.
3. Retrieve the list of available categories and their IDs using `GET /api/v1/workload-categories`.
4. Send a request to update the workload type with the new category using\
   `PUT /api/v1/workload-types/{workloadTypeId}` and include the `categoryId` in the request body.

## Using API

Go to the [Workload properties](https://run-ai-docs.nvidia.com/api/workloads/workload-properties#get-api-v1-workload-categories) API reference to view the available actions.


# GPU Profiling Metrics

This guide describes how to enable advanced GPU profiling metrics from **NVIDIA Data Center GPU Manager (DCGM)**. These metrics provide deep visibility into GPU performance, including SM utilization, memory bandwidth, tensor core activity, and compute pipeline behavior - extending beyond standard GPU utilization metrics. For more details on metric definitions, see [NVIDIA profiling metrics](https://docs.nvidia.com/datacenter/dcgm/latest/user-guide/feature-overview.html#metrics).

## Available Metrics

Once enabled, NVIDIA Run:ai exposes the following GPU profiling metrics:

* SM activity - SM active cycles and occupancy
* Memory bandwidth - DRAM active cycles
* Compute pipelines - FP16/FP32/FP64 and tensor core activity
* Graphics engine - GR engine utilization
* PCIe/NVLink - Data transfer rates

NVIDIA Run:ai automatically aggregates these metrics at multiple levels:

* Per GPU device
* Per pod
* Per workload
* Per node

## Configuring the DCGM Exporter

The **DCGM Exporter** is responsible for exposing GPU performance metrics to Prometheus. To configure it for advanced metrics:

1. Create the metrics configuration file and save it as `dcgm-metrics.csv`:

   ```csv
   # DCGM FIELD, Prometheus metric type, help message

   # Clocks
   DCGM_FI_DEV_SM_CLOCK,  gauge, SM clock frequency (in MHz).
   DCGM_FI_DEV_MEM_CLOCK, gauge, Memory clock frequency (in MHz).

   # Temperature
   DCGM_FI_DEV_MEMORY_TEMP, gauge, Memory temperature (in C).
   DCGM_FI_DEV_GPU_TEMP,    gauge, GPU temperature (in C).

   # Power
   DCGM_FI_DEV_POWER_USAGE,              gauge, Power draw (in W).
   DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION, counter, Total energy consumption since boot (in mJ).

   # PCIE
   DCGM_FI_DEV_PCIE_REPLAY_COUNTER, counter, Total number of PCIe retries.

   # Utilization
   DCGM_FI_DEV_GPU_UTIL,      gauge, GPU utilization (in %).
   DCGM_FI_DEV_MEM_COPY_UTIL, gauge, Memory utilization (in %).
   DCGM_FI_DEV_ENC_UTIL,      gauge, Encoder utilization (in %).
   DCGM_FI_DEV_DEC_UTIL ,     gauge, Decoder utilization (in %).

   # Errors
   DCGM_FI_DEV_XID_ERRORS, gauge, Value of the last XID error encountered.

   # Memory
   DCGM_FI_DEV_FB_FREE, gauge, Framebuffer memory free (in MiB).
   DCGM_FI_DEV_FB_USED, gauge, Framebuffer memory used (in MiB).

   # NVLink
   DCGM_FI_DEV_NVLINK_BANDWIDTH_TOTAL, counter, Total number of NVLink bandwidth counters for all lanes.
   DCGM_FI_DEV_NVLINK_BANDWIDTH_L0,    counter, The number of bytes of active NVLink rx or tx data including both header and payload.

   # vGPU
   DCGM_FI_DEV_VGPU_LICENSE_STATUS, gauge, vGPU License status

   # Remapped rows
   DCGM_FI_DEV_UNCORRECTABLE_REMAPPED_ROWS, counter, Number of remapped rows for uncorrectable errors
   DCGM_FI_DEV_CORRECTABLE_REMAPPED_ROWS,   counter, Number of remapped rows for correctable errors
   DCGM_FI_DEV_ROW_REMAP_FAILURE,           gauge,   Whether remapping of rows has failed

   # Labels
   DCGM_FI_DRIVER_VERSION, label, Driver Version

   # DCP Profiling Metrics (Advanced)
   DCGM_FI_PROF_GR_ENGINE_ACTIVE,   gauge, Ratio of time the graphics engine is active (in %).
   DCGM_FI_PROF_SM_ACTIVE,          gauge, The ratio of cycles an SM has at least 1 warp assigned (in %).
   DCGM_FI_PROF_SM_OCCUPANCY,       gauge, The ratio of number of warps resident on an SM (in %).
   DCGM_FI_PROF_PIPE_TENSOR_ACTIVE, gauge, Ratio of cycles the tensor (HMMA) pipe is active (in %).
   DCGM_FI_PROF_DRAM_ACTIVE,        gauge, Ratio of cycles the device memory interface is active sending or receiving data (in %).
   DCGM_FI_PROF_PIPE_FP64_ACTIVE,   gauge, Ratio of cycles the fp64 pipes are active (in %).
   DCGM_FI_PROF_PIPE_FP32_ACTIVE,   gauge, Ratio of cycles the fp32 pipes are active (in %).
   DCGM_FI_PROF_PIPE_FP16_ACTIVE,   gauge, Ratio of cycles the fp16 pipes are active (in %).
   DCGM_FI_PROF_PCIE_TX_BYTES,      gauge, The rate of data transmitted over the PCIe bus - including both protocol headers and data payloads - in bytes per second.
   DCGM_FI_PROF_PCIE_RX_BYTES,      gauge, The rate of data received over the PCIe bus - including both protocol headers and data payloads - in bytes per second.
   DCGM_FI_PROF_NVLINK_TX_BYTES,    gauge, The number of bytes of active NvLink tx (transmit) data including both header and payload.
   DCGM_FI_PROF_NVLINK_RX_BYTES,    gauge, The number of bytes of active NvLink rx (read) data including both header and payload
   ```
2. Create the following Helm values file and save it as `extended-dcgm-metrics-values.yaml`:

   ```yaml
   dcgmExporter:
     config:
       name: metrics-config
     env:
       - name: DCGM_EXPORTER_COLLECTORS
         value: /etc/dcgm-exporter/dcgm-metrics.csv
   ```
3. Run the following to create the ConfigMap and upgrade the GPU operator:

   ```bash
   # Get GPU Operator version
   GPU_OPERATOR_VERSION=$(helm ls -A | grep gpu-operator | awk '{ print $10 }')

   # Create ConfigMap with metrics configuration
   kubectl create configmap metrics-config -n gpu-operator --from-file=dcgm-metrics.csv

   # Upgrade GPU Operator with new configuration
   helm upgrade -i gpu-operator nvidia/gpu-operator \
     -n gpu-operator \
     --version $GPU_OPERATOR_VERSION \
     --reuse-values \
     -f extended-dcgm-metrics-values.yaml
   ```

## Enabling NVIDIA Run:ai Metric Aggregation

Enable NVIDIA Run:ai to create enriched metrics from the DCGM profiling data. This configures Prometheus recording rules that aggregate raw DCGM metrics per pod, workload, and node. See [Advanced cluster configurations](/saas/infrastructure-setup/advanced-setup/cluster-config#prometheus) for more details.

* **Using Helm** - Set the following value in your `values.yaml` file under `clusterConfig` and upgrade the chart:

  ```bash
  clusterConfig:  
    prometheus:
      spec: # PrometheusSpec
        config:
          advancedMetricsEnabled: true
  ```
* **Using runaiconfig at runtime** - Use the following `kubectl` patch command:

  ```bash
  kubectl patch runaiconfig runai -n runai \
    --type=merge \
    -p '{
      "spec": {
        "prometheus": {
          "config": {
            "advancedMetricsEnabled": true
          }
        }
      }
    }'
  ```

## Enabling GPU Profiling Metrics Settings

GPU profiling metrics are disabled by default. To enable:

1. Go to **General settings** and navigate to Analytics
2. Enable **GPU profiling metrics**
3. Metrics become visible under the **Workloads** and **Nodes** pages once active

## Verification

{% tabs %}
{% tab title="UI" %}
Workloads:

1. Navigate to **Workload manager** → Workloads
2. Click a row in the Workloads table and then click the SHOW DETAILS button at the upper-right side of the action bar. The details pane appears, presenting the **Metrics** tab.
3. In the **Type** dropdown, verify the **GPU profiling** option is available

Nodes:

1. Navigate to **Resources** → Nodes
2. Click a row in the Nodes table and then click the SHOW DETAILS button at the upper-right side of the action bar. The details pane appears, presenting the **Metrics** tab.
3. In the **Type** dropdown, verify the **GPU profiling** option is available
   {% endtab %}

{% tab title="CLI v2" %}

```bash
# Check DCGM exporter pods are running
kubectl get pods -n gpu-operator -l app=nvidia-dcgm-exporter

# Check for advanced metrics in Prometheus (if accessible)

# Look for metrics like: DCGM_FI_PROF_SM_ACTIVE, DCGM_FI_PROF_DRAM_ACTIVE

# And Run:ai enriched metrics like: runai_gpu_sm_active_per_pod_per_gpu
```

{% endtab %}

{% tab title="API" %}

* Workloads - To view the GPU profiling metrics per pod, refer to the [Pods](https://run-ai-docs.nvidia.com/api/workloads/pods) API
* Nodes - To view the GPU profiling metrics for nodes, refer to the [Nodes](https://run-ai-docs.nvidia.com/api/organizations/nodes) API
  {% endtab %}
  {% endtabs %}


# Metrics and Telemetry

Metrics are numeric measurements recorded **over time** that are emitted from the NVIDIA Run:ai cluster and telemetry is a numeric measurement recorded in real-time when emitted from the NVIDIA Run:ai cluster.

## Scopes

NVIDIA Run:ai provides control-plane API which supports and aggregates analytics at various levels.

<table><thead><tr><th width="298.04296875">Level</th><th>Description</th></tr></thead><tbody><tr><td>Cluster</td><td>A cluster is a set of nodes pools and nodes. With Cluster metrics, metrics are aggregated at the Cluster level. In the NVIDIA Run:ai user interface, metrics are available in the Overview dashboard.</td></tr><tr><td>Node</td><td>Data is aggregated at the node level.</td></tr><tr><td>Node pool</td><td>Data is aggregated at the node pool level.</td></tr><tr><td>Workload</td><td>Data is aggregated at the workload level. In some workloads, e.g. with distributed workloads, these metrics aggregate data from all worker pods.</td></tr><tr><td>Pod</td><td>The basic unit of execution.</td></tr><tr><td>Project</td><td>The basic organizational unit. Projects are the tool to implement resource allocation policies as well as the segregation between different initiatives.</td></tr><tr><td>Department</td><td>Departments are a grouping of projects.</td></tr></tbody></table>

## Supported Metrics

| Metric name in API               | Applicable API endpoint                                                            | Metric name in UI per grid                                           | Applicable UI grid                                                        |
| -------------------------------- | ---------------------------------------------------------------------------------- | -------------------------------------------------------------------- | ------------------------------------------------------------------------- |
| `ALLOCATED_GPU`                  | <ul><li>Clusters</li><li>Node pools</li></ul>                                      | <ul><li>GPU devices (allocated)</li><li>Allocated GPUs</li></ul>     | <ul><li>Overview dashboard</li><li>Node pools</li></ul>                   |
| `AVG_WORKLOAD_WAIT_TIME`         | <ul><li>Clusters</li><li>Node pools</li></ul>                                      |                                                                      |                                                                           |
| `CPU_LIMIT_CORES`                | Workloads                                                                          | CPU limit                                                            | Workloads                                                                 |
| `CPU_MEMORY_LIMIT_BYTES`         | Workloads                                                                          | CPU memory limit                                                     | Workloads                                                                 |
| `CPU_MEMORY_REQUEST_BYTES`       | Workloads                                                                          | CPU memory request                                                   | Workloads                                                                 |
| `CPU_MEMORY_USAGE_BYTES`         | <ul><li>Workloads</li><li>Pods</li></ul>                                           | CPU memory usage                                                     | Workloads                                                                 |
| `CPU_MEMORY_UTILIZATION`         | <ul><li>Clusters</li><li>Node pools</li><li>Nodes</li></ul>                        | CPU memory utilization                                               | <ul><li>Overview dashboard</li><li>Node pools</li><li>Nodes</li></ul>     |
| `CPU_REQUEST_CORES`              | Workloads                                                                          | CPU request                                                          | Workloads                                                                 |
| `CPU_USAGE_CORES`                | <ul><li>Nodes</li><li>Workloads</li><li>Pods</li></ul>                             | CPU usage                                                            | Workloads                                                                 |
| `CPU_UTILIZATION`                | <ul><li>Clusters</li><li>Node pools</li><li>Nodes</li></ul>                        | <ul><li>CPU compute utilization</li><li>CPU utilization</li></ul>    | <ul><li>Overview dashboard and Node pools</li><li>Nodes</li></ul>         |
| `GPU_ALLOCATION`                 | <ul><li>Workloads</li><li>Projects</li><li>Departments</li></ul>                   | GPU devices (allocated)                                              | Overview dashboard                                                        |
| `GPU_MEMORY_REQUEST_BYTES`       | Workloads                                                                          | GPU memory request                                                   | Workloads                                                                 |
| `GPU_MEMORY_USAGE_BYTES`         | <ul><li>Workloads</li><li>Pods</li><li>Nodes</li></ul>                             | GPU memory usage                                                     | Workloads                                                                 |
| `GPU_MEMORY_USAGE_BYTES_PER_GPU` | <ul><li>Nodes</li><li>Pods</li></ul>                                               | GPU memory usage per GPU                                             | Workloads per pod                                                         |
| `GPU_MEMORY_UTILIZATION`         | <ul><li>Clusters</li><li>Node pools</li></ul>                                      | GPU memory utilization                                               | <ul><li>Overview dashboard</li><li>Node pools</li></ul>                   |
| `GPU_MEMORY_UTILIZATION_PER_GPU` | Nodes                                                                              | GPU memory utilization per GPU                                       | Nodes                                                                     |
| `GPU_QUOTA`                      | <ul><li>Clusters</li><li>Node pools</li><li>Projects</li><li>Departments</li></ul> | Quota                                                                | Quota management                                                          |
| `GPU_UTILIZATION`                | <ul><li>Clusters</li><li>Node pools</li><li>Workloads</li><li>Pods</li></ul>       | GPU compute utilization                                              | <ul><li>Overview dashboard</li><li>Node pools</li><li>Workloads</li></ul> |
| `GPU_UTILIZATION_PER_GPU`        | <ul><li>Nodes</li><li>Pods</li></ul>                                               | GPU utilization per GPU                                              | Nodes                                                                     |
| `TOTAL_GPU`                      | <ul><li>Clusters</li><li>Node pools</li></ul>                                      | <ul><li>GPU devices total</li><li>Total GPUs</li></ul>               | <ul><li>Overview dashboard</li><li>Node pools</li></ul>                   |
| `TOTAL_GPU_NODES`                | <ul><li>Clusters</li><li>Node pools</li></ul>                                      |                                                                      |                                                                           |
| `GPU_UTILIZATION_DISTRIBUTION`   | <ul><li>Clusters</li><li>Node pools</li></ul>                                      | GPU utilization distribution                                         | Node pools                                                                |
| `UNALLOCATED_GPU`                | <ul><li>Clusters</li><li>Node pools</li></ul>                                      | <ul><li>GPU devices (unallocated)</li><li>Unallocated GPUs</li></ul> | <ul><li>Overview dashboard</li><li>Node pools</li></ul>                   |
| `CPU_QUOTA_MILLICORES`           | <ul><li>Projects</li><li>Departments</li></ul>                                     |                                                                      |                                                                           |
| `CPU_MEMORY_QUOTA_MB`            | <ul><li>Projects</li><li>Departments</li></ul>                                     |                                                                      |                                                                           |
| `CPU_ALLOCATION_MILLICORES`      | <ul><li>Projects</li><li>Departments</li></ul>                                     |                                                                      |                                                                           |
| `CPU_MEMORY_ALLOCATION_MB`       | <ul><li>Projects</li><li>Departments</li></ul>                                     |                                                                      |                                                                           |
| `POD_COUNT`                      | Workloads                                                                          |                                                                      |                                                                           |
| `RUNNING_POD_COUNT`              | Workloads                                                                          |                                                                      |                                                                           |
| `NVLINK_BANDWIDTH_TOTAL`         | <ul><li>Nodes</li><li>Pods</li></ul>                                               |                                                                      | <ul><li>Nodes</li><li>Workloads per pod</li></ul>                         |

### GPU Profiling

NVIDIA provides extended metrics as shown [here](https://docs.nvidia.com/datacenter/dcgm/latest/user-guide/feature-overview.html#profiling-metrics).

{% hint style="info" %}
**Note**

GPU profiling metrics are disabled by default. If unavailable, your administrator must enable it under **General settings** → Analytics → GPU profiling metrics. Before enabling, the administrator must configure GPU profiling through the DCGM Exporter and NVIDIA Run:ai Prometheus integration. For configuration steps, see [GPU profiling metrics](/saas/platform-management/monitor-performance/gpu-profiling-metrics).
{% endhint %}

<table><thead><tr><th width="177">Metric name in API</th><th width="177">Applicable API endpoint</th><th width="177">Metric name in UI</th><th>Applicable UI table</th></tr></thead><tbody><tr><td><code>GPU_FP16_ENGINE_ACTIVITY_PER_GPU</code></td><td><ul><li>Pods</li><li>Nodes</li></ul></td><td>GPU FP16 engine activity</td><td><ul><li>Workloads</li><li>Nodes</li></ul></td></tr><tr><td><code>GPU_FP32_ENGINE_ACTIVITY_PER_GPU</code></td><td><ul><li>Pods</li><li>Nodes</li></ul></td><td>GPU FP32 engine activity</td><td><ul><li>Workloads</li><li>Nodes</li></ul></td></tr><tr><td><code>GPU_FP64_ENGINE_ACTIVITY_PER_GPU</code></td><td><ul><li>Pods</li><li>Nodes</li></ul></td><td>GPU FP64 engine activity</td><td><ul><li>Workloads</li><li>Nodes</li></ul></td></tr><tr><td><code>GPU_GRAPHICS_ENGINE_ACTIVITY_PER_GPU</code></td><td><ul><li>Pods</li><li>Nodes</li></ul></td><td>Graphics engine activity</td><td><ul><li>Workloads</li><li>Nodes</li></ul></td></tr><tr><td><code>GPU_MEMORY_BANDWIDTH_UTILIZATION_PER_GPU</code></td><td><ul><li>Pods</li><li>Nodes</li></ul></td><td>Memory bandwidth utilization</td><td><ul><li>Workloads</li><li>Nodes</li></ul></td></tr><tr><td><code>GPU_NVLINK_RECEIVED_BANDWIDTH_PER_GPU</code></td><td><ul><li>Pods</li><li>Nodes</li></ul></td><td>NVLink received bandwidth</td><td><ul><li>Workloads</li><li>Nodes</li></ul></td></tr><tr><td><code>GPU_NVLINK_TRANSMITTED_BANDWIDTH_PER_GPU</code></td><td><ul><li>Pods</li><li>Nodes</li></ul></td><td>NVLink transmitted bandwidth</td><td><ul><li>Workloads</li><li>Nodes</li></ul></td></tr><tr><td><code>GPU_PCIE_RECEIVED_BANDWIDTH_PER_GPU</code></td><td><ul><li>Pods</li><li>Nodes</li></ul></td><td>PCIe received bandwidth</td><td><ul><li>Workloads</li><li>Nodes</li></ul></td></tr><tr><td><code>GPU_PCIE_TRANSMITTED_BANDWIDTH_PER_GPU</code></td><td><ul><li>Pods</li><li>Nodes</li></ul></td><td>PCIe transmitted bandwidth</td><td><ul><li>Workloads</li><li>Nodes</li></ul></td></tr><tr><td><code>GPU_SM_ACTIVITY_PER_GPU</code></td><td><ul><li>Pods</li><li>Nodes</li></ul></td><td>GPU SM activity</td><td><ul><li>Workloads</li><li>Nodes</li></ul></td></tr><tr><td><code>GPU_SM_OCCUPANCY_PER_GPU</code></td><td><ul><li>Pods</li><li>Nodes</li></ul></td><td>GPU SM occupancy</td><td><ul><li>Workloads</li><li>Nodes</li></ul></td></tr><tr><td><code>GPU_TENSOR_ACTIVITY_PER_GPU</code></td><td><ul><li>Pods</li><li>Nodes</li></ul></td><td>GPU tensor activity</td><td><ul><li>Workloads</li><li>Nodes</li></ul></td></tr><tr><td><code>GPU_OOMKILL_SWAP_OUT_OF_RAM_COUNT_PER_GPU</code></td><td>Nodes</td><td>OOMKill swap out of RAM count</td><td>Nodes</td></tr><tr><td><code>GPU_OOMKILL_BURST_COUNT_PER_GPU</code></td><td>Nodes</td><td>OOMKill burst count</td><td>Nodes</td></tr><tr><td><code>GPU_OOMKILL_IDLE_COUNT_PER_GPU</code></td><td>Nodes</td><td>OOMKill idle count</td><td>Nodes</td></tr><tr><td><code>GPU_SWAP_MEMORY_BYTES_PER_GPU</code></td><td>Pods</td><td>GPU swap memory</td><td>Workloads</td></tr></tbody></table>

### NVIDIA NIM

NVIDIA NIM metrics provide workload-level observability, including key runtime and performance data such as request throughput, latency, and token usage for LLMs. See [NIM observability metrics via API](https://run-ai-docs.nvidia.com/api/api-guides/nim-observability-metrics-via-api) for more details.

<table><thead><tr><th width="177.42578125">Metric name in API</th><th width="177">Applicable API endpoint</th><th width="176.8828125">Metric name in UI</th><th>Applicable UI table</th></tr></thead><tbody><tr><td><code>NIM_NUM_REQUESTS_RUNNING</code></td><td><ul><li>Pods</li><li>Workloads</li></ul></td><td>Request concurrency by status</td><td>Workloads</td></tr><tr><td><code>NIM_NUM_REQUESTS_WAITING</code></td><td><ul><li>Pods</li><li>Workloads</li></ul></td><td>Request concurrency by status</td><td>Workloads</td></tr><tr><td><code>NIM_NUM_REQUEST_MAX</code></td><td><ul><li>Pods</li><li>Workloads</li></ul></td><td>Request concurrency by status</td><td>Workloads</td></tr><tr><td><code>NIM_REQUEST_SUCCESS_TOTAL</code></td><td><ul><li>Pods</li><li>Workloads</li></ul></td><td>Request count by status</td><td>Workloads</td></tr><tr><td><code>NIM_REQUEST_FAILURE_TOTAL</code></td><td><ul><li>Pods</li><li>Workloads</li></ul></td><td>Request count by status</td><td>Workloads</td></tr><tr><td><code>NIM_GPU_CACHE_USAGE_PERC</code></td><td><ul><li>Pods</li><li>Workloads</li></ul></td><td>GPU KV cache utilization</td><td>Workloads</td></tr><tr><td><code>NIM_TIME_TO_FIRST_TOKEN_SECONDS</code></td><td><ul><li>Pods</li><li>Workloads</li></ul></td><td>Time to first token (TTFT)</td><td>Workloads</td></tr><tr><td><code>NIM_E2E_REQUEST_LATENCY_SECONDS</code></td><td><ul><li>Pods</li><li>Workloads</li></ul></td><td>End to end request latency</td><td>Workloads</td></tr><tr><td><code>NIM_TIME_TO_FIRST_TOKEN_SECONDS_PERCENTILES</code></td><td><ul><li>Pods</li><li>Workloads</li></ul></td><td>Time to first token (TTFT) by percentiles</td><td>Workloads</td></tr><tr><td><code>NIM_E2E_REQUEST_LATENCY_SECONDS_PERCENTILES</code></td><td><ul><li>Pods</li><li>Workloads</li></ul></td><td>End to end request latency by percentiles</td><td>Workloads</td></tr></tbody></table>

## Supported Telemetry

| Metric                              | Applicable API endpoint                                          | Metric name in UI              | Applicable UI table                                          |
| ----------------------------------- | ---------------------------------------------------------------- | ------------------------------ | ------------------------------------------------------------ |
| `WORKLOADS_COUNT`                   | Workloads                                                        |                                |                                                              |
| `ALLOCATED_GPUS`                    | Nodes                                                            | Allocated GPUs                 | Nodes                                                        |
| `GPU_allocation`                    | <ul><li>Workloads</li><li>Projects</li><li>Departments</li></ul> |                                |                                                              |
| `READY_GPU_NODES`                   | Nodes                                                            | Ready / Total GPU nodes        | Overview dashboard                                           |
| `READY_GPUS`                        | Nodes                                                            | Ready / Total GPU devices      | Overview dashboard                                           |
| `TOTAL_GPU_NODES`                   | Nodes                                                            | Ready / Total GPU nodes        | Overview dashboard                                           |
| `TOTAL_GPUS`                        | Nodes                                                            | Ready / Total GPU devices      | Overview dashboard                                           |
| `IDLE_ALLOCATED_GPUS`               | Nodes                                                            | Idle allocated GPU devices     | Overview dashboard                                           |
| `FREE_GPUS`                         | Nodes                                                            | Free GPU devices               | Nodes                                                        |
| `TOTAL_CPU_CORES`                   | Nodes                                                            | CPU (Cores)                    | Nodes                                                        |
| `USED_CPU_CORES`                    | Nodes                                                            |                                |                                                              |
| `ALLOCATED_CPU_CORES`               | <ul><li>Nodes</li><li>Projects</li><li>Departments</li></ul>     | <p>Allocated CPU cores<br></p> | Nodes                                                        |
| `TOTAL_GPU_MEMORY_BYTES`            | Nodes                                                            | GPU memory                     | Nodes                                                        |
| `USED_GPU_MEMORY_BYTES`             | Nodes                                                            | Used GPU memory                | Nodes                                                        |
| `TOTAL_CPU_MEMORY_BYTES`            | Nodes                                                            | CPU memory                     | Nodes                                                        |
| `USED_CPU_MEMORY_BYTES`             | Nodes                                                            | Used CPU memory                | Nodes                                                        |
| `ALLOCATED_CPU_MEMORY_BYTES`        | <ul><li>Nodes</li><li>Projects</li><li>Departments</li></ul>     | Allocated CPU memory           | <ul><li>Nodes</li><li>Projects</li><li>Departments</li></ul> |
| `GPU_QUOTA`                         | <ul><li>Projects</li><li>Departments</li></ul>                   | GPU quota                      | <ul><li>Projects</li><li>Departments</li></ul>               |
| `CPU_QUOTA`                         | <ul><li>Projects</li><li>Departments</li></ul>                   |                                |                                                              |
| `MEMORY_QUOTA`                      | <ul><li>Projects</li><li>Departments</li></ul>                   |                                |                                                              |
| `GPU_ALLOCATION_NON_PREEMPTIBLE`    | <ul><li>Projects</li><li>Departments</li></ul>                   |                                |                                                              |
| `CPU_ALLOCATION_NON_PREEMPTIBLE`    | <ul><li>Projects</li><li>Departments</li></ul>                   |                                |                                                              |
| `MEMORY_ALLOCATION_NON_PREEMPTIBLE` | <ul><li>Projects</li><li>Departments</li></ul>                   |                                |                                                              |




---

[Next Page](/llms-full.txt/1)

