NVIDIA NIM requirements on your own server: GPUs, drivers, support matrix and licence
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- NIM for LLMs 2.0.13 needs an NVIDIA driver 580 or later, Docker 24.0 or later and the NVIDIA Container Toolkit 1.14.0 or later on AMD64 or ARM64 Linux, with Ubuntu 22.04 LTS or later recommended (NVIDIA prerequisites, 6 October 2026)
- The support matrix of 6 October 2026 lists the H200 NVL, the RTX PRO 6000 Blackwell Server Edition and the L40S as verified GPUs for the Llama 3.1 8B and Llama 3.3 70B, gpt-oss and Nemotron 3 Nano and Super NIMs; the L4 is in none of these lists
- NIM picks a profile (precision and tensor parallelism) from the GPU and its memory;
list-model-profilesshows which profiles are runnable, low on memory or incompatible on your host - Models are cached in /opt/nim/.cache, and an air-gapped host runs from a cache or model store prepared on a connected machine, with no NGC or Hugging Face key set
- NIM Certified Production Branch requires an NVIDIA AI Enterprise subscription, licensed per GPU, and NVIDIA’s NIM FAQ requires one for any production use, while the offerings page lists the Feature Branch at no charge; each H200 NVL includes five years
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
NVIDIA NIM requirements in brief
NVIDIA NIM for large language models runs on a Linux host with an NVIDIA driver from branch 580 or later, Docker 24.0 or later and the NVIDIA Container Toolkit 1.14.0 or later, on a GPU that has a matching profile for the model. NVIDIA’s support matrix of 6 October 2026 names the H200 NVL, the RTX PRO 6000 Blackwell Server Edition and the L40S among the verified GPUs for the Llama, gpt-oss and Nemotron 3 Nano and Super NIMs. For production use, NVIDIA’s NIM FAQ requires an NVIDIA AI Enterprise licence, counted per GPU, and each H200 NVL includes a five-year subscription; the licence section below covers the Feature Branch that NVIDIA offers at no charge.
This guide covers NIM for LLMs at release 2.0.13, as NVIDIA’s documentation stood when we read it on 10 October 2026. Embedding and other NIM families keep their own support matrices. Counting, activation and renewal of AI Enterprise are covered in our NVIDIA AI Enterprise licensing guide.
What a NIM container contains
A NIM is a container image that combines a serving engine, an OpenAI-compatible API on port 8000 and a set of model profiles. NVIDIA’s legal notice for NIM for LLMs states that it “packages vLLM directly as its inference backend”, and the release notes of 2.0.13 name vLLM 0.28.0. A profile is one combination of precision, such as BF16, FP8, NVFP4 or MXFP4, and tensor parallelism (TP), the number of GPUs one copy of the model is split over.
There are two kinds of image. A model-specific NIM, such as nim/meta/llama-3. on NVIDIA’s registry nvcr.io, carries the profiles NVIDIA built for that model and downloads the weights on first start. The model-free NIM serves a model you point it to with NIM_MODEL_PATH; NVIDIA lists NGC, Hugging Face, Amazon S3, Google Cloud Storage, ModelScope and local storage as sources.
Host requirements: driver, CUDA and container runtime
NVIDIA’s prerequisites page, last updated on 6 October 2026, sets these minimums and gives no figures for system RAM, CPU cores or disk.
| REQUIREMENT | NVIDIA STATES | OUR READING |
|---|---|---|
| CPU architecture | AMD64, ARM64 | x86 servers and Arm hosts both qualify |
| Operating system | Ubuntu 22.04 LTS or later recommended | other distributions “have not been officially validated” |
| GPU driver | 580 or later | R580 and R595 both qualify |
| CUDA SDK | 12.9 or later | the image brings its own CUDA runtime |
| Docker | 24.0 or later | Kubernetes has its own guides |
| NVIDIA Container Toolkit | 1.14.0 or later | needed on every GPU host |
| GPU memory | per model | Llama 3.1 8B Instruct: at least 24 GB |
| NVIDIA access | Developer Program or AI Enterprise | NGC Personal API key for Production Branch models |
NVIDIA NIM for LLMs prerequisites and quickstart pages, both updated 6 October 2026 and read on 10 October 2026; the right-hand column is our reading.
R580, a long-term support branch to June 2028, and R595, a production branch to March 2027, both meet it, as our guide to NVIDIA driver branches and CUDA versions explains. The 2.0.13 release notes update the cuda-compat-13-0 package inside the image to 580.159.03 and add: “The associated host-driver vulnerabilities require a host GPU driver update.” The fixed versions in NVIDIA’s security bulletin of 30 September 2026 are 580.178.04 and 595.91.07. On RTX PRO Blackwell cards the driver must use the open kernel modules.
The NGC Personal API key is needed only to download Production Branch models or NIMs released before 2.0.10, and the page warns that “Legacy API keys are not supported by NIM LLM and VLM.”
Supported GPUs in the NIM support matrix
The support matrix “lists supported models, deployment profiles, and verified hardware SKUs for NIM LLM and VLM”. Each model has its own list of verified GPUs. The page does not define what verification covers; we read a missing card as untested by NVIDIA for that model, not as blocked by the container.
| MODEL | PROFILES | H200 NVL | RTX PRO 6000 SE | L40S |
|---|---|---|---|---|
| Llama 3.1 8B Instruct | BF16, FP8, NVFP4; TP1 | verified | verified | verified |
| Llama 3.3 70B Instruct | BF16, FP8, NVFP4; TP1 to 8 | verified | verified | verified |
| gpt-oss-120b | MXFP4; TP1 to 8 | verified | verified | verified |
| gpt-oss-20b | MXFP4; TP1 to 8 | verified | verified | verified |
| Nemotron 3 Nano | BF16, FP8, NVFP4; TP1 to 8 | verified | verified | verified |
| Nemotron 3 Super 120B | BF16, FP8, NVFP4; TP1 to 8 | verified | verified | verified |
| Llama 3.3 Nemotron Super 49B | BF16, FP8, NVFP4; TP1 to 8 | verified | verified | verified |
| Nemotron 3 Ultra 550B | BF16 TP8; NVFP4 TP2, 4, 8 | not listed | not listed | not listed |
| GLM-5.2 | FP8, NVFP4; TP4, TP8 | not listed | not listed | not listed |
| DeepSeek V4 Pro 0813 | FP8; TP8 | not listed | not listed | not listed |
NVIDIA NIM for LLMs support matrix, last updated 6 October 2026, read on 10 October 2026. Most profiles also exist with LoRA adapters. SE: Server Edition.
Every list in the table that names the RTX PRO 6000 Server Edition also names the RTX PRO 4500 Blackwell Server Edition. The GB10, the chip in DGX Spark, is verified for Llama 3.1 8B, gpt-oss-20b, Nemotron 3 Nano and Llama 3.3 Nemotron Super 49B. None of the lists above names the L4, the RTX PRO 5000 or the Workstation and Max-Q editions of the RTX PRO 6000.
The three largest models in the table are verified on SXM and Grace-based GPUs such as the H200 (SXM), B200 and GB200, and Ultra’s list also names the H100 NVL; none of the three lists names a card we supply. NVIDIA’s note for Super says that lower-TP profiles need substantially more memory per device, so “some verified GPUs support only TP4 or TP8 profiles”. Memory per model and per user is covered in our Nemotron hardware requirements and our guide to how much VRAM an LLM needs.
Release 2.0.13 makes Llama 3.1 70B in BF16 at TP2 “supported on NVIDIA RTX PRO 6000 Blackwell Server Edition”. It lists the BF16 LoRA profile of Nemotron 3 Super at TP2 as unsupported on the H200 NVL, and several Llama 3.3 70B LoRA profiles as unsupported on the L40S for lack of GPU memory.
We build AI servers to order with the RTX PRO 6000 Server Edition, the L40S or the H200 NVL. Send us the NIMs you plan to run and your peak number of users, and we propose cards that the matrix lists for those models.
Model profiles and GPU memory
If no profile is set, “NIM automatically selects the best compatible profile from the manifest based on your hardware”, using the GPU model, its memory and the parallelism settings. Running the image with list-model-profiles shows the result before a deployment. It sorts profiles into “Compatible with system and runnable”, “Compatible with system but low memory” and “Incompatible with system”, the last meaning that the model weights alone exceed the GPU memory. NIM_MODEL_PROFILE pins one profile by its ID or description, so several hosts run the same profile.
Precision also depends on the architecture. For Nemotron 3.5 Lightning the matrix gives a minimum per GPU and an architecture for every profile. NVFP4 at TP1 needs 30 GB and “Blackwell or newer (SM 10.0+)”, which the RTX PRO 6000 meets with compute capability 12.0. The H200 NVL is a Hopper card, so it takes the W4A16 profile (32 GB at TP1) or BF16 (66 GB at TP1) instead. On a 48 GB L40S the W4A16 profile fits one card, while BF16 needs TP2 at 35 GB per GPU.
NVIDIA notes that these minimums cover the weights and runtime allocations, without headroom for a large KV cache. The memory left after the weights decides how many conversations one copy of the model serves, so a card that meets a minimum may still hold too few users.
Model cache, disk space and air-gapped installation
The NIM stores model files in /opt/nim/.cache inside the container, set by NIM_CACHE_PATH, and the quickstart mounts a host directory there, which the prerequisites set to ~/.cache/nim. NVIDIA’s model download page gives no disk figure. The cache holds the weights of every profile downloaded, and download-to-cache --all fetches all of them, so plan local NVMe for at least the checkpoints of the profiles you will run. For Nemotron 3 Super, the NVFP4 checkpoint alone is 80.3 GB on Hugging Face.
The air-gap guide, updated on 6 October 2026, splits the work into a connected phase and an offline phase.
- On a machine with internet access, run
list-model-profilesand note the profile ID for your GPU and TP size. - Run
download-to-cache -pwith that ID, orcreate-model-store -pwith an output directory set by-m. - Copy the cache or the model store to the offline host, by archive, rsync or physical media.
- Start the NIM there with the cache mounted and
NIM_MODEL_PROFILEset to the profile ID, or withNIM_MODEL_PATHpointing to the model store, and withoutNGC_API_KEYorHF_TOKEN; NVIDIA states “In the air-gapped phase, do not set NGC_API_KEY or HF_TOKEN.”
On Kubernetes, “Every image referenced by the downloaded Helm chart must be available from a registry that the air-gapped cluster can reach”, including the GPU Operator’s images. AI Enterprise needs no licence server for containers on bare metal, so an offline host has no licence check to reach.
NIM licence: NIM, NIM Certified and NVIDIA AI Enterprise
NVIDIA’s offerings page, updated on 6 October 2026, describes NIM and two branches of NIM Certified. “NIM is available at no charge under the applicable license terms”; it is published within about 72 hours of the upstream model, “validated to be functional on a small set of NVIDIA GPUs” and “not covered by NVIDIA Enterprise support.” The page states that “NIM Certified Feature Branch is available at no charge under the applicable license terms” and lists it for “Production deployments that prioritize newer capabilities and can adopt frequent software updates”, with enterprise support only under an AI Enterprise subscription. For the other branch it states: “NIM Certified Production Branch requires an active NVIDIA AI Enterprise subscription.”
NVIDIA’s other documents set narrower limits. Its NIM General FAQ, last modified on 6 August 2026, states: “Using NIM in production requires an NVIDIA AI Enterprise license.” It defines production as any use other than development, testing, research or evaluation, and limits self-hosted Developer Program use to those purposes on up to 16 GPUs. The NIM legal page refers to NVIDIA’s Product-Specific Terms for AI Products, last modified on 15 April 2026, which state: “Software offered as part of the developer program is not for use, distribution or deployment in production.” Their section on NIMs for workstations allows designated NIMs without a subscription on a PC or workstation with NVIDIA RTX GPUs that is “not used in a commercial kiosk, server or other system used to service multiple users”.
As of 10 October 2026, these pages do not say in one place whether a Feature Branch NIM may serve staff in production without a subscription, so have the quote state the offering and the licence it rests on. The Production Branch and NVIDIA support need AI Enterprise in every case, counted per GPU. The H200 NVL includes a five-year subscription, activated with the card’s serial number; the RTX PRO 6000 Server Edition, the L40S and the L4 include none, and DGX Spark has a separate AI Enterprise product. Whether a given deployment counts as production under these terms is a legal assessment for your legal department.
We supply NVIDIA AI Enterprise licences on one EU contract and invoice with the GPUs they run on. Tell us how many GPUs will serve NIM in production and which cards you already own, and the quote shows what each card includes.
NIM in virtual machines and on Kubernetes
NVIDIA’s vGPU deployment page for NIM asks you to “Install and license NVIDIA vGPU software on the hypervisor host” and uses C-series compute profiles in its examples. It also notes that “the usable GPU frame buffer inside the VM is less than the configured vGPU profile size”, so leave a margin when you choose the vGPU profile for a NIM. Compute (C-series) vGPU profiles are licensed only through AI Enterprise, per GPU.
On Kubernetes, NVIDIA documents NIM with Helm, KServe, OpenShift, Run:ai and the NIM Operator, whose NIMCache resource downloads models to network storage. Our article on a private LLM platform on Kubernetes shows where NIM sits beside vLLM and KServe.
What we supply
We supply the cards this guide discusses, the RTX PRO 6000 Server Edition, the RTX PRO 4500 Server Edition, the L40S and the H200 NVL, as cards or in AI servers built to order, and we check the rack, power and airflow before we quote. On request the servers arrive with the operating system, drivers, CUDA and a container runtime installed. NVIDIA AI Enterprise licences for NIM, listed on our AI Enterprise and vGPU licensing page, come on the same invoice as the hardware, which carries manufacturer warranty. The models, RAG and MLOps on top are our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
What are the NVIDIA NIM requirements?
Which GPUs does NVIDIA NIM support?
Does NIM run on the RTX PRO 6000?
Does NIM run on the H200 NVL?
Do I need an NVIDIA AI Enterprise licence for NIM?
Can NVIDIA NIM run on-premise without internet access?
Send us the NIM models you plan to run, the GPUs or servers you already have, the hypervisor if there is one and whether the NIM will serve staff in production. We reply within one business day with a configuration, the AI Enterprise licence count and a quote in writing.
Talk to an expertWe reply within one business day