AI Potluck
Back to Gap Map Infrastructure / ML orchestration

NVIDIA Run:ai

NVIDIA
closed / Overall score: n/a

Enterprise platform for AI workloads and GPU orchestration. It pools GPU capacity across public clouds, private clouds, hybrid environments and on-premise data centres, allocates it dynamically including fractionally, and places AI workloads against quotas set per team.

NVIDIA describes the open KAI Scheduler as 'based on NVIDIA Run:ai'; KAI now lives in its own GitHub organization and is scored separately on this map, so the open component is not credited to this record. The product page is dynamic and returned a different digest on two fetches an hour apart, so a re-fetch will read as drift. Verified 2026-09-15 via the NVIDIA Run:ai product page.

Openness

1 high confidence
1.0
license
Proprietary
service
proprietary(NVIDIA commercial platform)
source
closed(the platform is unpublished

A closed commercial platform. NVIDIA open-sourced KAI Scheduler, which its own page describes as 'based on NVIDIA Run:ai', but that is a separate published project with its own record here rather than the platform's source - the same shape as google-cloud-run beside gVisor. The open component is recorded on its own and does not lift this product's source dimension.

  • https://www.nvidia.com/en-us/software/run-ai/ recorded 2026-09-15

    NVIDIA's Run:ai product page: 'The enterprise platform for AI workloads and GPU orchestration', with dynamic resource allocation and GPU pooling across public clouds, private clouds, hybrid environments and on-premise data centres, and the note that the open-source KAI Scheduler is 'based on NVIDIA Run:ai'. Establishes a closed platform with a separately published component.

Adoption

not assessed

No usage figure is published for this service specifically, and it publishes no countable artifact - no package, repository or registry entry - so no level is assigned rather than one being inferred from the parent platform.

Capability

5 medium confidence
5.0

Band 5 on NVIDIA's own scheduler documentation, which states that the Scheduler 'always uses the pod group associated with each pod and handles the pod as part of a larger group of pods' with 'gang scheduling, minimum replicas or workers... applied consistently across the entire pod group', over 'distributed workloads using multiple pods, each running on a node'. That is a multi-machine set reserved as one allocation by the product's own scheduler. It was drafted at 4 from the product marketing page, which describes pooling and fractional allocation and attributes gang scheduling to Grove with KAI; reading the scheduler documentation rather than the marketing page is what corrected it, and treating a marketing page's silence as a capability ceiling is the error.

  • https://www.nvidia.com/en-us/software/run-ai/ recorded 2026-09-15

    NVIDIA's Run:ai page describing dynamic scheduling and orchestration, GPU pooling and fractional inference allocation, and attributing gang-scheduled deployments to Grove with KAI Scheduler.

  • https://run-ai-docs.nvidia.com/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles recorded 2026-09-16

    NVIDIA Run:ai's own Scheduler documentation: the Scheduler 'always uses the pod group associated with each pod and handles the pod as part of a larger group of pods, while taking into consideration the common characteristics of the pod group, such as gang scheduling, minimum replicas or workers, workload priority class, workload preemptibility policy, topology information', applied 'consistently across the entire pod group'; workloads range to 'distributed workloads using multiple pods, each running on a node'. Establishes gang scheduling by the platform's own scheduler rather than by a separate component.

Verified 2026-09-15