BLOG · GUIDE ·

NVIDIA NIM requirements on your own server: GPUs, drivers, support matrix and licence

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • NIM for LLMs 2.0.13 needs an NVIDIA driver 580 or later, Docker 24.0 or later and the NVIDIA Container Toolkit 1.14.0 or later on AMD64 or ARM64 Linux, with Ubuntu 22.04 LTS or later recommended (NVIDIA prerequisites, 6 October 2026)
  • The support matrix of 6 October 2026 lists the H200 NVL, the RTX PRO 6000 Blackwell Server Edition and the L40S as verified GPUs for the Llama 3.1 8B and Llama 3.3 70B, gpt-oss and Nemotron 3 Nano and Super NIMs; the L4 is in none of these lists
  • NIM picks a profile (precision and tensor parallelism) from the GPU and its memory; list-model-profiles shows which profiles are runnable, low on memory or incompatible on your host
  • Models are cached in /opt/nim/.cache, and an air-gapped host runs from a cache or model store prepared on a connected machine, with no NGC or Hugging Face key set
  • NIM Certified Production Branch requires an NVIDIA AI Enterprise subscription, licensed per GPU, and NVIDIA’s NIM FAQ requires one for any production use, while the offerings page lists the Feature Branch at no charge; each H200 NVL includes five years

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

NVIDIA NIM requirements in brief

NVIDIA NIM for large language models runs on a Linux host with an NVIDIA driver from branch 580 or later, Docker 24.0 or later and the NVIDIA Container Toolkit 1.14.0 or later, on a GPU that has a matching profile for the model. NVIDIA’s support matrix of 6 October 2026 names the H200 NVL, the RTX PRO 6000 Blackwell Server Edition and the L40S among the verified GPUs for the Llama, gpt-oss and Nemotron 3 Nano and Super NIMs. For production use, NVIDIA’s NIM FAQ requires an NVIDIA AI Enterprise licence, counted per GPU, and each H200 NVL includes a five-year subscription; the licence section below covers the Feature Branch that NVIDIA offers at no charge.

This guide covers NIM for LLMs at release 2.0.13, as NVIDIA’s documentation stood when we read it on 10 October 2026. Embedding and other NIM families keep their own support matrices. Counting, activation and renewal of AI Enterprise are covered in our NVIDIA AI Enterprise licensing guide.

What a NIM container contains

A NIM is a container image that combines a serving engine, an OpenAI-compatible API on port 8000 and a set of model profiles. NVIDIA’s legal notice for NIM for LLMs states that it “packages vLLM directly as its inference backend”, and the release notes of 2.0.13 name vLLM 0.28.0. A profile is one combination of precision, such as BF16, FP8, NVFP4 or MXFP4, and tensor parallelism (TP), the number of GPUs one copy of the model is split over.

There are two kinds of image. A model-specific NIM, such as nim/meta/llama-3.1-8b-instruct:2.0.13 on NVIDIA’s registry nvcr.io, carries the profiles NVIDIA built for that model and downloads the weights on first start. The model-free NIM serves a model you point it to with NIM_MODEL_PATH; NVIDIA lists NGC, Hugging Face, Amazon S3, Google Cloud Storage, ModelScope and local storage as sources.

Host requirements: driver, CUDA and container runtime

NVIDIA’s prerequisites page, last updated on 6 October 2026, sets these minimums and gives no figures for system RAM, CPU cores or disk.

REQUIREMENTNVIDIA STATESOUR READING
CPU architectureAMD64, ARM64x86 servers and Arm hosts both qualify
Operating systemUbuntu 22.04 LTS or later recommendedother distributions “have not been officially validated”
GPU driver580 or laterR580 and R595 both qualify
CUDA SDK12.9 or laterthe image brings its own CUDA runtime
Docker24.0 or laterKubernetes has its own guides
NVIDIA Container Toolkit1.14.0 or laterneeded on every GPU host
GPU memoryper modelLlama 3.1 8B Instruct: at least 24 GB
NVIDIA accessDeveloper Program or AI EnterpriseNGC Personal API key for Production Branch models

NVIDIA NIM for LLMs prerequisites and quickstart pages, both updated 6 October 2026 and read on 10 October 2026; the right-hand column is our reading.

R580, a long-term support branch to June 2028, and R595, a production branch to March 2027, both meet it, as our guide to NVIDIA driver branches and CUDA versions explains. The 2.0.13 release notes update the cuda-compat-13-0 package inside the image to 580.159.03 and add: “The associated host-driver vulnerabilities require a host GPU driver update.” The fixed versions in NVIDIA’s security bulletin of 30 September 2026 are 580.178.04 and 595.91.07. On RTX PRO Blackwell cards the driver must use the open kernel modules.

The NGC Personal API key is needed only to download Production Branch models or NIMs released before 2.0.10, and the page warns that “Legacy API keys are not supported by NIM LLM and VLM.”

Supported GPUs in the NIM support matrix

The support matrix “lists supported models, deployment profiles, and verified hardware SKUs for NIM LLM and VLM”. Each model has its own list of verified GPUs. The page does not define what verification covers; we read a missing card as untested by NVIDIA for that model, not as blocked by the container.

MODELPROFILESH200 NVLRTX PRO 6000 SEL40S
Llama 3.1 8B InstructBF16, FP8, NVFP4; TP1verifiedverifiedverified
Llama 3.3 70B InstructBF16, FP8, NVFP4; TP1 to 8verifiedverifiedverified
gpt-oss-120bMXFP4; TP1 to 8verifiedverifiedverified
gpt-oss-20bMXFP4; TP1 to 8verifiedverifiedverified
Nemotron 3 NanoBF16, FP8, NVFP4; TP1 to 8verifiedverifiedverified
Nemotron 3 Super 120BBF16, FP8, NVFP4; TP1 to 8verifiedverifiedverified
Llama 3.3 Nemotron Super 49BBF16, FP8, NVFP4; TP1 to 8verifiedverifiedverified
Nemotron 3 Ultra 550BBF16 TP8; NVFP4 TP2, 4, 8not listednot listednot listed
GLM-5.2FP8, NVFP4; TP4, TP8not listednot listednot listed
DeepSeek V4 Pro 0813FP8; TP8not listednot listednot listed

NVIDIA NIM for LLMs support matrix, last updated 6 October 2026, read on 10 October 2026. Most profiles also exist with LoRA adapters. SE: Server Edition.

Every list in the table that names the RTX PRO 6000 Server Edition also names the RTX PRO 4500 Blackwell Server Edition. The GB10, the chip in DGX Spark, is verified for Llama 3.1 8B, gpt-oss-20b, Nemotron 3 Nano and Llama 3.3 Nemotron Super 49B. None of the lists above names the L4, the RTX PRO 5000 or the Workstation and Max-Q editions of the RTX PRO 6000.

The three largest models in the table are verified on SXM and Grace-based GPUs such as the H200 (SXM), B200 and GB200, and Ultra’s list also names the H100 NVL; none of the three lists names a card we supply. NVIDIA’s note for Super says that lower-TP profiles need substantially more memory per device, so “some verified GPUs support only TP4 or TP8 profiles”. Memory per model and per user is covered in our Nemotron hardware requirements and our guide to how much VRAM an LLM needs.

Release 2.0.13 makes Llama 3.1 70B in BF16 at TP2 “supported on NVIDIA RTX PRO 6000 Blackwell Server Edition”. It lists the BF16 LoRA profile of Nemotron 3 Super at TP2 as unsupported on the H200 NVL, and several Llama 3.3 70B LoRA profiles as unsupported on the L40S for lack of GPU memory.

We build AI servers to order with the RTX PRO 6000 Server Edition, the L40S or the H200 NVL. Send us the NIMs you plan to run and your peak number of users, and we propose cards that the matrix lists for those models.

Model profiles and GPU memory

If no profile is set, “NIM automatically selects the best compatible profile from the manifest based on your hardware”, using the GPU model, its memory and the parallelism settings. Running the image with list-model-profiles shows the result before a deployment. It sorts profiles into “Compatible with system and runnable”, “Compatible with system but low memory” and “Incompatible with system”, the last meaning that the model weights alone exceed the GPU memory. NIM_MODEL_PROFILE pins one profile by its ID or description, so several hosts run the same profile.

Precision also depends on the architecture. For Nemotron 3.5 Lightning the matrix gives a minimum per GPU and an architecture for every profile. NVFP4 at TP1 needs 30 GB and “Blackwell or newer (SM 10.0+)”, which the RTX PRO 6000 meets with compute capability 12.0. The H200 NVL is a Hopper card, so it takes the W4A16 profile (32 GB at TP1) or BF16 (66 GB at TP1) instead. On a 48 GB L40S the W4A16 profile fits one card, while BF16 needs TP2 at 35 GB per GPU.

NVIDIA notes that these minimums cover the weights and runtime allocations, without headroom for a large KV cache. The memory left after the weights decides how many conversations one copy of the model serves, so a card that meets a minimum may still hold too few users.

Model cache, disk space and air-gapped installation

The NIM stores model files in /opt/nim/.cache inside the container, set by NIM_CACHE_PATH, and the quickstart mounts a host directory there, which the prerequisites set to ~/.cache/nim. NVIDIA’s model download page gives no disk figure. The cache holds the weights of every profile downloaded, and download-to-cache --all fetches all of them, so plan local NVMe for at least the checkpoints of the profiles you will run. For Nemotron 3 Super, the NVFP4 checkpoint alone is 80.3 GB on Hugging Face.

The air-gap guide, updated on 6 October 2026, splits the work into a connected phase and an offline phase.

  1. On a machine with internet access, run list-model-profiles and note the profile ID for your GPU and TP size.
  2. Run download-to-cache -p with that ID, or create-model-store -p with an output directory set by -m.
  3. Copy the cache or the model store to the offline host, by archive, rsync or physical media.
  4. Start the NIM there with the cache mounted and NIM_MODEL_PROFILE set to the profile ID, or with NIM_MODEL_PATH pointing to the model store, and without NGC_API_KEY or HF_TOKEN; NVIDIA states “In the air-gapped phase, do not set NGC_API_KEY or HF_TOKEN.”

On Kubernetes, “Every image referenced by the downloaded Helm chart must be available from a registry that the air-gapped cluster can reach”, including the GPU Operator’s images. AI Enterprise needs no licence server for containers on bare metal, so an offline host has no licence check to reach.

NIM licence: NIM, NIM Certified and NVIDIA AI Enterprise

NVIDIA’s offerings page, updated on 6 October 2026, describes NIM and two branches of NIM Certified. “NIM is available at no charge under the applicable license terms”; it is published within about 72 hours of the upstream model, “validated to be functional on a small set of NVIDIA GPUs” and “not covered by NVIDIA Enterprise support.” The page states that “NIM Certified Feature Branch is available at no charge under the applicable license terms” and lists it for “Production deployments that prioritize newer capabilities and can adopt frequent software updates”, with enterprise support only under an AI Enterprise subscription. For the other branch it states: “NIM Certified Production Branch requires an active NVIDIA AI Enterprise subscription.”

NVIDIA’s other documents set narrower limits. Its NIM General FAQ, last modified on 6 August 2026, states: “Using NIM in production requires an NVIDIA AI Enterprise license.” It defines production as any use other than development, testing, research or evaluation, and limits self-hosted Developer Program use to those purposes on up to 16 GPUs. The NIM legal page refers to NVIDIA’s Product-Specific Terms for AI Products, last modified on 15 April 2026, which state: “Software offered as part of the developer program is not for use, distribution or deployment in production.” Their section on NIMs for workstations allows designated NIMs without a subscription on a PC or workstation with NVIDIA RTX GPUs that is “not used in a commercial kiosk, server or other system used to service multiple users”.

As of 10 October 2026, these pages do not say in one place whether a Feature Branch NIM may serve staff in production without a subscription, so have the quote state the offering and the licence it rests on. The Production Branch and NVIDIA support need AI Enterprise in every case, counted per GPU. The H200 NVL includes a five-year subscription, activated with the card’s serial number; the RTX PRO 6000 Server Edition, the L40S and the L4 include none, and DGX Spark has a separate AI Enterprise product. Whether a given deployment counts as production under these terms is a legal assessment for your legal department.

We supply NVIDIA AI Enterprise licences on one EU contract and invoice with the GPUs they run on. Tell us how many GPUs will serve NIM in production and which cards you already own, and the quote shows what each card includes.

NIM in virtual machines and on Kubernetes

NVIDIA’s vGPU deployment page for NIM asks you to “Install and license NVIDIA vGPU software on the hypervisor host” and uses C-series compute profiles in its examples. It also notes that “the usable GPU frame buffer inside the VM is less than the configured vGPU profile size”, so leave a margin when you choose the vGPU profile for a NIM. Compute (C-series) vGPU profiles are licensed only through AI Enterprise, per GPU.

On Kubernetes, NVIDIA documents NIM with Helm, KServe, OpenShift, Run:ai and the NIM Operator, whose NIMCache resource downloads models to network storage. Our article on a private LLM platform on Kubernetes shows where NIM sits beside vLLM and KServe.

What we supply

We supply the cards this guide discusses, the RTX PRO 6000 Server Edition, the RTX PRO 4500 Server Edition, the L40S and the H200 NVL, as cards or in AI servers built to order, and we check the rack, power and airflow before we quote. On request the servers arrive with the operating system, drivers, CUDA and a container runtime installed. NVIDIA AI Enterprise licences for NIM, listed on our AI Enterprise and vGPU licensing page, come on the same invoice as the hardware, which carries manufacturer warranty. The models, RAG and MLOps on top are our Private AI/ML service, with engineering by our partner Vixen.UNO.

FAQ

What are the NVIDIA NIM requirements?
NIM for LLMs 2.0.13 needs an NVIDIA driver 580 or later, Docker 24.0 or later, the NVIDIA Container Toolkit 1.14.0 or later and an AMD64 or ARM64 Linux host, with Ubuntu 22.04 LTS or later recommended. The GPU must have enough memory for a profile of the model, for example at least 24 GB for Llama 3.1 8B Instruct. Access comes through the NVIDIA Developer Program or an AI Enterprise licence.
Which GPUs does NVIDIA NIM support?
Each model in NVIDIA’s support matrix has its own list of verified GPUs. As of 6 October 2026, the H200 NVL, the RTX PRO 6000 Blackwell Server Edition, the RTX PRO 4500 Blackwell Server Edition and the L40S are verified for the Llama 3.1 8B and Llama 3.3 70B, gpt-oss and Nemotron 3 Nano and Super NIMs. The largest models, such as Nemotron 3 Ultra and GLM-5.2, are not verified on any of these four cards.
Does NIM run on the RTX PRO 6000?
The RTX PRO 6000 Blackwell Server Edition is a verified GPU for the Llama, gpt-oss and Nemotron 3 Nano and Super NIMs in the matrix of 6 October 2026, alongside the RTX PRO 4500 Server Edition. The Workstation and Max-Q editions do not appear in those lists.
Does NIM run on the H200 NVL?
Yes, the H200 NVL is a verified GPU for the Llama, gpt-oss and Nemotron 3 Nano and Super NIMs in the matrix of 6 October 2026. NIM selects the profile from the card’s 141 GB and its Hopper architecture, and profiles that require Blackwell, such as the NVFP4 profile of Nemotron 3.5 Lightning, do not apply to it. Each H200 NVL includes a five-year NVIDIA AI Enterprise subscription.
Do I need an NVIDIA AI Enterprise licence for NIM?
For production, NVIDIA’s NIM FAQ of 6 August 2026 says yes, and NIM Certified Production Branch requires an active AI Enterprise subscription, counted per GPU, while NVIDIA’s offerings page of 6 October 2026 lists NIM and the Feature Branch at no charge under the applicable licence terms. Have the quote state which offering and licence a production server rests on. Development and testing through the Developer Program need no subscription, and NVIDIA offers a 90-day evaluation licence.
Can NVIDIA NIM run on-premise without internet access?
Yes. NVIDIA’s air-gap guide prepares a model cache or model store on a connected machine with download-to-cache or create-model-store, then runs the NIM offline without NGC_API_KEY or HF_TOKEN set. On Kubernetes, every image of the Helm chart has to be mirrored to a registry the cluster can reach.

Send us the NIM models you plan to run, the GPUs or servers you already have, the hypervisor if there is one and whether the NIM will serve staff in production. We reply within one business day with a configuration, the AI Enterprise licence count and a quote in writing.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna