BLOG · GUIDE ·

Nemotron hardware requirements: Nemotron 3 Nano, Super and Ultra on-premise, with NIM

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • By our memory estimate, Nemotron 3 Super (120B, 12B active) runs for one user on one RTX PRO 6000 or one DGX Spark (128 GB) in its 80.3 GB NVFP4 checkpoint, and on one H200 NVL or two RTX PRO 6000 in its 128.4 GB FP8 version
  • Only 8 of Super’s 88 layers are attention layers, so a 32K conversation costs about 0.28 GiB of FP8 cache and Mamba state, and one RTX PRO 6000 holds 29 such conversations beside the NVFP4 weights
  • Nemotron 3 Nano 30B-A3B (32.7 GB in FP8), Nemotron 3.5 Lightning (21.6 GB in NVFP4) and Nano 4B fit one card or one Spark with room for 100 conversations at 32K
  • Nemotron 3 Ultra (550B, 55B active) ships in NVFP4 at 352.3 GB, with no FP8 checkpoint on Hugging Face, and needs four H200 NVL or eight RTX PRO 6000 by our estimate
  • NVIDIA’s NIM support matrix of 6 October 2026 lists the H200 NVL and the RTX PRO 6000 Server Edition as verified GPUs for Nano and Super but not for Ultra; Nano, Super and Nano 4B use the NVIDIA Nemotron Open Model License, Ultra and 3.5 Lightning OpenMDW-1.1

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

Nemotron hardware requirements in brief

Nemotron 3 Super, NVIDIA’s 120B open model of March 2026 with 12B active parameters, runs for one user on one RTX PRO 6000 or one DGX Spark in its 80.3 GB NVFP4 checkpoint, and on one H200 NVL or two RTX PRO 6000 in its 128.4 GB FP8 version. Nemotron 3 Nano 30B-A3B, Nemotron 3.5 Lightning and Nano 4B fit one card or one Spark with room for 100 conversations. Nemotron 3 Ultra, with 550B parameters, ships as a 352.3 GB NVFP4 checkpoint and needs four H200 NVL or eight RTX PRO 6000.

These are our memory estimates for conversations of 32,768 tokens, from NVIDIA’s Hugging Face files as read on 9 October 2026. Most Nemotron 3 layers keep no KV cache, so the weights usually decide the GPU count. Our LLM hardware requirements by model applies the same method to other model families.

Nemotron models on Hugging Face in October 2026

MODELPARAMETERSCHECKPOINTSCONTEXTLICENCE
Nemotron 3 Nano 4B3.97B, no expertsBF16 7.9 GB; FP8 5.3 GB262KNemotron Open Model
Nemotron 3 Nano 30B-A3B30B, 3.5B activeBF16 63.2 GB; FP8 32.7 GB; NVFP4 19.3 GB1MNemotron Open Model
Nemotron 3.5 Lightning30B, 3B activeBF16 65.8 GB; NVFP4 21.6 GB1MOpenMDW-1.1
Nemotron 3 Super120B, 12B activeBF16 247.2 GB; FP8 128.4 GB; NVFP4 80.3 GB1MNemotron Open Model
Nemotron 3 Ultra550B, 55B activeNVFP4 352.3 GB; BF161MOpenMDW-1.1
Llama 3.3 Nemotron Super 49B50B, denseFP8 52.0 GB; NVFP4 31.1 GB128KNVIDIA Open Model, Llama 3.3

Model cards, file lists and config.json files on huggingface.co/nvidia, read on 9 October 2026; sizes are the sums of the safetensors files. Release dates on the cards: Nano 30B-A3B 15 December 2025, Super 11 March 2026, Nano 4B 16 March 2026, Ultra 4 June 2026, 3.5 Lightning 11 August 2026, Super 49B v1.5 25 July 2025.

Nemotron 3 Nano, Super and Ultra interleave Mamba-2 layers with mixture-of-experts (MoE) layers and add a few attention layers. Super’s config.json lists 88 layers: 40 Mamba-2, 40 MoE and 8 attention, with 512 experts of which 22 are active per token. Its card calls the design “LatentMoE” and adds a multi-token prediction (MTP) layer. Ultra has 106 layers, 12 of them attention, and Nano 30B-A3B and 3.5 Lightning have 52 layers, 6 of them attention, with 6 of 128 experts active. Nano 4B has 42 layers, four with attention, and no experts.

Llama-3.3-Nemotron-Super-49B-v1.5 is derived from Llama 3.3 70B Instruct by Neural Architecture Search (NAS), which skips attention in some blocks; by our count of its config.json and weight index, 49 of its 80 layers keep attention.

The RTX PRO 6000 and the DGX Spark compute FP4 directly, while the H200 NVL computes FP8 but not FP4. We found no FP8 checkpoint of Ultra on NVIDIA’s Hugging Face pages, and its NIM documentation states that FP8 profiles were not released “because NVFP4 profiles work on all tested Hopper-architecture GPUs (for example, H100 and H200)”.

KV cache and Mamba state per Nemotron conversation

Only attention layers keep a KV cache that grows with the context: per token, attention layers × KV heads × head dimension × 2 (key and value) × bytes. For Super that is 8 × 2 × 128 × 2 × 1 byte in FP8, 4 KiB per token, against 160 KiB for Llama 3.3 70B with an FP8 cache. Each Mamba-2 layer keeps a fixed state per conversation instead: heads × head dimension × state size, 128 × 64 × 128 values for Super, or 4 MiB in the float32 that vLLM’s recipe sets with --mamba-ssm-cache-dtype float32, plus a small convolution buffer. The general method is in our guide to how much VRAM an LLM needs.

MODELATTENTION LAYERSKV PER TOKENMAMBA STATE32K; 1M
Nano 4B4 of 42, 8 KV heads8 KiB80 MiB0.33 GiB; 2.1 GiB at 262K
Nano 30B-A3B, 3.5 Lightning6 of 52, 2 KV heads3 KiB47 MiB0.14 GiB; 3.0 GiB
Nemotron 3 Super8 of 88, 2 KV heads4 KiB162 MiB0.28 GiB; 4.2 GiB
Nemotron 3 Ultra12 of 106, 2 KV heads6 KiB193 MiB0.38 GiB; 6.2 GiB
Super 49B v1.549 of 80, 8 KV heads196 KiB, 16-bitnone6.1 GiB; 128K maximum

Our arithmetic from the config.json files, read on 9 October 2026. KV cache in FP8, as NVIDIA’s vLLM commands set with --kv-cache-dtype fp8; Mamba state in float32 as an upper bound, Ultra’s in float16 as its card sets. Super 49B in 16-bit, vLLM’s default, as its card’s command sets no cache type.

Our budget follows the VRAM guide: 90 per cent of the memory the driver reports, less 3 GiB per card, which leaves 83.0 GiB per RTX PRO 6000 and 123.4 GiB per H200 NVL. For one DGX Spark (128 GB) we take 102 GB. With tensor parallelism we count the whole cache and state of a Nemotron 3 model on every card, an upper bound. Super 49B’s cache splits over 2, 4 or 8 cards by its eight KV heads, since “TP splits the KV cache by those heads first”, as vLLM’s blog of 7 August 2026 states.

GPUs for 1, 20 and 100 users of each Nemotron model

MODEL, FORMATONE DGX SPARKRTX PRO 6000H200 NVL
Nemotron 3 Nano 4B, BF16yes / yes / yes1 / 1 / 11 / 1 / 1
Nemotron 3 Nano 30B, FP8yes / yes / yes1 / 1 / 11 / 1 / 1
3.5 Lightning, NVFP4yes / yes / yes1 / 1 / 11 / 1 / 1, W4A16
Nemotron 3 Super, NVFP4yes / yes / no1 / 1 / 2see FP8
Nemotron 3 Super, FP8no / no / no2 / 2 / 41 / 2 / 2
Nemotron 3 Super, BF16no / no / no4 / 4 / 82 / 2 / 4
Nemotron 3 Ultra, NVFP4no / no / no8 / 8 / 84 / 4 / 4, weight-only
Super 49B v1.5, FP8yes / no / no1 / 4 / 81 / 2 / 8

Our estimates, not measurements: cards for 1, 20 and 100 concurrent conversations of 32,768 tokens in 1, 2, 4 or 8 cards with tensor parallelism, cache as in the table above. DGX Spark shows a memory fit only. W4A16 is NIM’s 4-bit weight profile for 3.5 Lightning, listed for Ampere or newer.

Nano 30B-A3B in FP8 leaves room for 377 conversations on one RTX PRO 6000, so the card is chosen for speed, not memory. Super 49B v1.5 behaves like a dense transformer: at 6.1 GiB per 32K conversation, one RTX PRO 6000 holds five beside the FP8 weights, so 20 users need four cards, or two H200 NVL, and eight hold 100 with no margin. An FP8 cache, set with --kv-cache-dtype fp8, halves the cache per conversation.

We build inference servers with 2 to 8 GPUs per node, sized by model size and concurrent users. Send us the Nemotron variant, its format, your context length and peak requests through the form below.

Nemotron 3 Super on one RTX PRO 6000, a DGX Spark or an H200 NVL

vLLM’s recipe for Nemotron 3 Super, updated on 8 October 2026, states: “The NVFP4 variant runs on a single RTX Pro 6000 at TP=1.” It requires vLLM 0.24.0 or later. Its command sets --gpu-memory-utilization 0.92 “to avoid an OOM at startup”, together with --kv-cache-dtype fp8, --mamba-ssm-cache-dtype float32, --max-num-seqs 8 and MTP speculative decoding with three tokens. By our rule, 8.2 GiB remain beside the 74.8 GiB of weights, room for 29 conversations of 32K; for 20 users, --max-num-seqs has to be raised above 8. We count one Mamba state per conversation, though the recipe’s prefix caching may keep more.

NVIDIA’s NVFP4 card names “1× B200 OR 1× DGX Spark” as the minimum GPU requirement, and its Spark command sets --max-model-len 1000000 with --max-num-seqs 4. By our estimate one Spark holds four conversations of 1M tokens at 4.2 GiB each. Its memory bandwidth is 273 GB/s, against 1,597 GB/s on the RTX PRO 6000 Server Edition, so speed limits a team before memory does.

The FP8 checkpoint’s card gives “2× H100-80GB” as the minimum GPU requirement. One H200 NVL holds the 119.5 GiB of FP8 weights with 3.8 GiB to spare, room for 13 conversations of 32K, so 20 users need a second card, joined by an NVLink bridge.

The RTX PRO 6000 has no NVLink, so the FP8 checkpoint on two cards splits every layer over PCIe, as our guide to one model on several GPUs explains. The NVFP4 checkpoint needs no split, and four copies on four cards hold 116 conversations, while two cards with tensor parallelism hold 160 by our estimate.

Nemotron 3 Ultra on four H200 NVL or eight RTX PRO 6000

Ultra’s card lists “4xGB200, 4xB200, 4x GB300, 4x B300, 8xH100” as the minimum GPU requirement. Its vLLM command for four B200 adds expert parallelism to --tensor-parallel-size 4 and sets --mamba-ssm-cache-dtype float16.

By our estimate four H200 NVL hold the 328.1 GiB of NVFP4 weights with 41.3 GiB per card to spare, room for 109 conversations of 32K. A four-way NVLink bridge joins them into one domain, as our guide to H200 NVL card counts for large models describes, and the Hopper card runs the 4-bit weights weight-only. Four RTX PRO 6000 hold the weights but leave only 1.0 GiB per card for the cache, which is not a usable configuration even for one user. Eight are the minimum and leave 42.0 GiB per card for 111 conversations, split over PCIe. Ultra’s NIM entry lists neither card among its verified GPUs, and its card gives eight H100 (640 GB) as the Hopper minimum, against 564 GB on four H200 NVL, so that layout rests on our memory estimate alone.

We build servers with four or eight H200 NVL and their NVLink bridges, and we check the rack, power and airflow before we quote. Describe the Nemotron model, your rack position and its power feed in the form below.

Nemotron NIM: support matrix and NVIDIA AI Enterprise

The current support matrix of NIM for LLMs, last updated on 6 October 2026 and showing no release number, lists NIMs for Nemotron 3 Nano, Super, Ultra, 3.5 Lightning and Super 49B v1.5. Super has BF16, FP8 and NVFP4 profiles at tensor parallelism of 1, 2, 4 or 8, and “some verified GPUs support only TP4 or TP8 profiles”. The verified GPUs for Nano and Super 49B include the H200 NVL, the RTX PRO 6000 Blackwell Server Edition and the GB10, the chip in DGX Spark; Super’s list includes the first two but not the GB10. For 3.5 Lightning it gives a minimum of 30 GB per GPU for NVFP4 at TP1 and 66 GB for BF16, without “additional headroom for large KV caches”.

Super’s card states that “The NIM container is governed by the NVIDIA Software License Agreement”, and production use of NIM needs an NVIDIA AI Enterprise licence per GPU, as our AI Enterprise licensing guide explains. NVIDIA states that the “H200 NVL comes with a five-year NVIDIA AI Enterprise subscription”, activated by serial number, while the RTX PRO 6000 Server Edition includes none and DGX Spark has its own AI Enterprise product. Without NIM, vLLM serves the same checkpoints under its Apache 2.0 licence.

Nemotron licences: Nemotron Open Model License and OpenMDW

Nano 4B, Nano 30B-A3B and Super are published under the NVIDIA Nemotron Open Model License, last modified on 15 December 2025, which states: “Works are commercially usable.” The licence ends for a Work if you start patent or copyright litigation claiming that it infringes. Ultra and 3.5 Lightning come under the OpenMDW License Agreement, version 1.1, which requires a copy of the agreement and the copyright notices on distribution and also ends on such litigation, unless it answers a suit first brought against you. Ultra’s card states: “This model is ready for commercial and non-commercial use.” Super 49B v1.5 names the NVIDIA Open Model License and the Llama 3.3 Community License Agreement, and its card carries “Built with Llama”. Whether a clause applies to your company is a legal assessment for your legal department.

What we supply

We supply the DGX Spark Founders Edition, the RTX PRO 6000 in its Workstation, Max-Q and Server editions, and the H200 NVL with its NVLink bridges, as cards or in AI servers built to order. NVIDIA AI Enterprise licences for NIM, listed on our software and licensing page, come on one EU contract and invoice with the hardware and its manufacturer warranty. The models, RAG and MLOps on top are our Private AI/ML service, with engineering by our partner Vixen.UNO.

FAQ

What are the hardware requirements for Nemotron 3 Super?
Nemotron 3 Super has 120B parameters with 12B active and ships as an 80.3 GB NVFP4 checkpoint, which vLLM’s recipe runs on one RTX PRO 6000 and NVIDIA’s card on one DGX Spark. The 128.4 GB FP8 version fits one H200 NVL or two RTX PRO 6000 for one user, and 20 users at 32K need two H200 NVL by our memory estimate. NVIDIA’s FP8 card names 2× H100-80GB as the minimum GPU requirement.
How much VRAM does Nemotron need?
The weights take 19.3 GB for Nano 30B-A3B in NVFP4, 32.7 GB in FP8, 80.3 GB for Super in NVFP4, 128.4 GB in FP8 and 352.3 GB for Ultra in NVFP4, as the Hugging Face file lists show. Because only a few layers keep a KV cache, a conversation of 32K tokens adds just 0.14 to 0.38 GiB of FP8 cache and Mamba state by our estimate. Plan on 90 per cent of the card’s memory, less 3 GiB, for weights plus cache.
Can I run Nemotron locally on a DGX Spark?
One DGX Spark (128 GB) holds Nano 4B, Nano 30B-A3B, 3.5 Lightning and Super in NVFP4, which NVIDIA’s card names with “1× B200 OR 1× DGX Spark” as the minimum. Super leaves about 20 GiB for the cache on a Spark, room for four conversations of 1M tokens or about 70 of 32K by our estimate. The Spark’s 273 GB/s of memory bandwidth limits speed before its memory does, and Ultra does not fit.
Which GPUs does the Nemotron NIM support?
NVIDIA’s NIM support matrix of 6 October 2026 lists the H200 NVL and the RTX PRO 6000 Blackwell Server Edition among the verified GPUs for Nemotron 3 Nano and Super, and the GB10 of DGX Spark for Nano but not for Super. The Ultra NIM lists B200, B300, GB200, GB300, H100 and H200, has no FP8 profiles, and its smallest profile uses two high-memory Blackwell GPUs, the others four or eight. Production use of a NIM needs an NVIDIA AI Enterprise licence per GPU, which the H200 NVL includes for five years.
What GPU does Llama Nemotron Super 49B need?
Llama-3.3-Nemotron-Super-49B-v1.5 is 52.0 GB in FP8 and 31.1 GB in NVFP4, so one RTX PRO 6000, H200 NVL or DGX Spark holds it for a single user. As a dense transformer with 49 attention layers it needs about 6.1 GiB of 16-bit cache per 32K conversation, so 20 users need four RTX PRO 6000 or two H200 NVL by our estimate. An FP8 cache halves that figure.
How much VRAM does Nemotron Nano need?
Nemotron 3 Nano 30B-A3B is 63.2 GB in BF16, 32.7 GB in FP8 and 19.3 GB in NVFP4, and Nano 4B is 7.9 GB in BF16. Each 32K conversation adds about 0.14 GiB on Nano 30B-A3B and 0.33 GiB on Nano 4B, so one RTX PRO 6000, one H200 NVL or one DGX Spark serves 100 such conversations by memory. NVIDIA’s NIM matrix lists the GB10 of DGX Spark, the RTX PRO 6000 Server Edition and the H200 NVL as verified GPUs for Nano.

Send us the Nemotron model and format, the context length you will configure, your peak number of concurrent requests and whether you plan to run NIM or vLLM. We reply within one business day with a configuration and a quote, with the rack, power and airflow checked before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna