BLOG · GUIDE ·

H200 NVFP4 and MXFP4: running FP4 models on H200 NVL, L40S and other GPUs without FP4 Tensor Cores

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • The H200 NVL (Hopper, compute capability 9.0) and the L40S, L4 and RTX Ada cards (8.9) have FP8 Tensor Cores but no FP4 ones; NVIDIA’s H200 and L40S specification tables list FP8 and no FP4
  • vLLM still loads NVFP4 and MXFP4 checkpoints on these GPUs: without a native FP4 GEMM kernel it falls back to weight-only W4A16 execution through Marlin and logs a warning, and SGLang does the same for NVFP4 on SM80 to SM90
  • TensorRT-LLM 1.3.0rc29 lists NVFP4 and MXFP4 only for Blackwell; its gpt-oss guide documents one Hopper path, MXFP4 MoE weights with BF16 activations on the H200 through the Triton backend
  • Weight-only FP4 keeps the memory saving: on one H200 NVL, Llama 3.3 70B leaves room for about 66 concurrent 8K sessions with NVIDIA’s NVFP4 checkpoint and 44 with its FP8 one, by our estimate, but the maths runs in 16-bit
  • On an H200 NVL or L40S fleet, the FP8 checkpoint is the default where it fits, INT4 AWQ or GPTQ is the documented 4-bit route, and gpt-oss runs in its MXFP4 release on Hopper

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

FP4 support on the H200 NVL, L40S and RTX Ada cards

The H200 NVL does not compute in FP4, and neither do the L40S, the L4 or the RTX Ada workstation cards. Their Tensor Cores go down to FP8 for floating-point formats, and FP4 Tensor Cores arrived with Blackwell. NVFP4 and MXFP4 models still run on these GPUs where the serving engine has weight-only kernels: vLLM and SGLang keep the weights in 4-bit in GPU memory and unpack them for 16-bit arithmetic. The memory saving remains, while the compute gain of FP4 does not apply.

Our guide to FP8, NVFP4, MXFP4, INT4 and GGUF explains the formats: NVFP4 stores 4-bit values with an FP8 scale for every 16 of them, MXFP4 with a power-of-two scale for every 32. This article is for teams with H200 NVL or L40S servers and a model published in FP4. It sets out what each engine documents for these cards as of October 2026 and which checkpoint to deploy.

Which precisions the H200 NVL, L40S and RTX Ada cards compute

NVIDIA’s specification table for the H200 NVL has Tensor Core rows for FP64, TF32, BFLOAT16, FP16, FP8 and INT8, and none for FP4. The L40S table lists TF32, BFLOAT16, FP16, FP8, INT8 and INT4, also without FP4, and NVIDIA describes the card’s fourth-generation Tensor Cores “with support for FP8”. NVIDIA’s compute capability list puts the H200 at 9.0 and the L4, L40S, RTX 6000 Ada, RTX 5000 Ada, RTX 4500 Ada and RTX 4000 Ada at 8.9. The RTX PRO Blackwell cards are listed at 12.0 and DGX Spark at 12.1, and these compute FP4 directly.

GPUCAPABILITYTENSOR FORMATSFP4 CHECKPOINTS
H200 NVL (Hopper)9.0FP8, INT8, BF16, FP16, TF32, FP64weight-only W4A16 in vLLM and SGLang; gpt-oss MXFP4 in TensorRT-LLM with BF16 activations
L40S (Ada)8.9FP8, INT8, INT4, BF16, FP16, TF32weight-only W4A16 in vLLM, and in SGLang by its SM80 to SM90 range; not in TensorRT-LLM
L4 and RTX Ada cards8.9FP8 and the 16-bit formats, as all 8.9 GPUsas the L40S
RTX PRO Blackwell12.0FP4 (NVFP4, MXFP4), FP8native FP4 in TensorRT-LLM (sm120) and SGLang (SM120)

NVIDIA H200 and L40S product pages and CUDA compute capability list (which names the H200, not the H200 NVL), read on 10 October 2026; vLLM quantisation docs (14 September 2026) and Model Optimizer page (2 October 2026); SGLang quantisation docs; TensorRT-LLM quantisation matrix 1.3.0rc29.

How vLLM runs NVFP4 and MXFP4 on Hopper and Ada

For an NVFP4 checkpoint, vLLM selects the matrix-multiplication kernel at load time from the backends the platform offers, CUTLASS, FlashInfer, Marlin and others. Its page on NVIDIA Model Optimizer checkpoints, dated 2 October 2026, states: “On GPUs without a supported native FP4 GEMM kernel, vLLM falls back to weight-only (W4A16) execution via Marlin”. It adds that vLLM logs a warning and that this “may reduce throughput for compute-heavy workloads”. On an H200 NVL or an L40S, this is the path an NVFP4 model takes by default. The option --linear-backend overrides the automatic choice and replaces the deprecated VLLM_NVFP4_GEMM_BACKEND variable.

The Marlin FP4 code in vLLM’s source requires compute capability 7.5 or higher and unpacks NVFP4 in blocks of 16 and MXFP4 in blocks of 32. Its warning reads, in part: “Weight-only FP4 compression will be used leveraging the Marlin kernel.” vLLM’s quantisation table marks Marlin for GPTQ, AWQ, FP8 and FP4 as supported on Turing, Ampere, Ada and Hopper.

W4A16 means 4-bit weights and 16-bit activations. The weights are converted to 16-bit inside each multiplication, so the Tensor Cores do the same work as for a BF16 model. For the KV cache, TensorRT-LLM lists an FP8 format for Ada and Hopper, and vLLM sets one with --kv-cache-dtype fp8.

TensorRT-LLM and SGLang with FP4 checkpoints

TensorRT-LLM’s quantisation matrix, documentation 1.3.0rc29 updated on 26 September 2026, lists for Hopper FP8 per tensor, per block and per row, an FP8 KV cache, and INT4 AWQ and GPTQ as W4A16 and W4A8. Ada gets FP8 per tensor, the FP8 KV cache and the same four INT4 modes. NVFP4 and MXFP4 appear only in the Blackwell rows. NVIDIA’s NVFP4 checkpoint of Llama 3.3 70B gives “NVIDIA Blackwell” as its supported microarchitecture and TensorRT-LLM as its runtime engine.

The one Hopper exception is gpt-oss. TensorRT-LLM’s gpt-oss deployment guide states: “For Hopper, the default MoE backend is TRITON.” Its support table pairs MXFP4 MoE weights with BF16 activations on the H200, so the 4-bit weights are computed at 16-bit precision there as well.

SGLang’s quantisation page lists the ModelOpt FP4 method for NVIDIA as “SM80-SM90 via Marlin; SM100+ native FP4” and describes it as a “Marlin W4A16 fallback on Ampere/Hopper”. Its automatic backend choice is stated as “On SM80-SM90, auto selects Marlin for NVFP4.” The page names Ampere and Hopper; Ada’s 8.9 lies inside that range. Our comparison of vLLM, SGLang, TensorRT-LLM and Ollama covers the engines beyond quantisation.

Memory and speed of weight-only FP4 on an H200 NVL

Weight-only FP4 reduces the memory the weights occupy and the bytes read for each generated token. It does not reduce the arithmetic. NVIDIA’s Model Optimizer guide explains when that matters: “Typically, in the context of small-batch inference scenarios (batch size ≤ 4), the inference is often ‘memory-bound’.” For such loads it names weight-only methods as the larger gain. For batches of 16 and more it recommends quantising activations too and writes: “We suggest prioritizing using FP8 first, as FP8 causes very little accuracy degradation and gives strong performance.” A company assistant with hundreds of users reaches batches of 16 and more in its busy hours.

Take Llama 3.3 70B on an H200 NVL. NVIDIA’s FP8 checkpoint is 72.7 GB (67.7 GiB) and its NVFP4 checkpoint 42.7 GB (39.8 GiB), by the file sizes on Hugging Face. Our sizing rule gives the KV cache 0.9 × the 140.4 GiB the driver reports, minus 3 GiB, minus the weights; an FP8 cache needs 1.25 GiB per session at 8,192 tokens. The FP8 checkpoint leaves 55.7 GiB, room for 44 concurrent sessions. The NVFP4 checkpoint leaves 83.6 GiB, room for 66. On eight H200 NVL in two servers, one copy of the model per card, that is 352 against 528 sessions held in memory.

Those are memory ceilings. Whether the NVFP4 copy serves more users at an acceptable response time depends on how Marlin’s 16-bit path compares with FP8 Tensor Core arithmetic at your batch size, and as of October 2026 we found no published measurement of the two on an H200 by NVIDIA or the vLLM project. NVIDIA’s model cards give MMLU scores of 81.1 for the FP4 checkpoint, 83.2 for FP8 and 83.3 for BF16, and the FP8 card lists Hopper and Lovelace among its supported architectures, the NVFP4 card only Blackwell. A load test of both checkpoints in your engine version, at your peak concurrency, settles it. Our RTX PRO 6000 and H200 NVL comparison shows the case where native FP4 on a Blackwell card is the better fit.

We supply the H200 NVL and build GPU servers with it to order. Send us the model, its checkpoint and your peak concurrency, and we reply within one business day with a configuration and quote.

gpt-oss MXFP4 on Hopper and Ada

OpenAI releases gpt-oss with its MoE weights in MXFP4, as it does the gpt-oss-safeguard models built on it. OpenAI’s model card states: “The models were post-trained with MXFP4 quantization of the MoE weights, making gpt-oss-120b run on a single 80GB GPU”, and names the H100 as an example. Here the MXFP4 checkpoint is the release, so on Hopper it is the checkpoint to run. Our gpt-oss-120b hardware guide sizes its 65.3 GB of weights on one H200 NVL.

vLLM’s gpt-oss recipe, updated on 22 September 2026, lists the H100, H200 and B200 among NVIDIA GPUs “with ongoing work for Ampere/Ada/RTX 5090”. For Hopper it says to run the Blackwell command without the FP8 KV cache setting and without the FlashInfer MoE flags. On a single A100 it names “Marlin MXFP4 MoE” as the default. The older guide in vLLM’s documentation keeps a commented-out VLLM_MXFP4_USE_MARLIN=1, a workaround for a performance regression of vLLM v0.12.0 on Hopper at small concurrency, which it says was fixed upstream.

On the L40S, gpt-oss-120b needs more than one card, since 65.3 GB of weights exceed 48 GB. vLLM’s quantisation table lists Marlin FP4 for Ada, while the gpt-oss recipe still calls Ada ongoing work. Treat gpt-oss on L40S servers as a configuration to test, not a documented recipe. Our L40S benchmarks article collects what has been measured on the card.

Which checkpoint to run on an H200 NVL or L40S fleet

RELEASE FORMATEXAMPLEH200 NVLL40S
FP8Qwen3-32B-FP8, Llama 3.3 70B FP8native FP8, the defaultnative FP8; a 70B model over two or four cards
NVFP4 beside FP8NVIDIA Llama 3.3 70B, Llama 4 ScoutFP8; NVFP4 only weight-only, for memoryFP8 if it fits, otherwise INT4 AWQ
FP4 expertsDeepSeek-V4-Flash previewvLLM recipe lists the H200no Ada target in vLLM’s recipe
MXFP4gpt-oss-120b, gpt-oss-20bMXFP4 in vLLM or TensorRT-LLMvLLM: Ada is ongoing work
INT4 (AWQ, GPTQ)Qwen3-32B-AWQ, Kimi K2 ThinkingW4A16 via Marlin; AWQ W4A8 in TensorRT-LLMas the H200 NVL

vLLM recipes for Llama 4 Scout (24 September 2026), gpt-oss (22 September 2026) and DeepSeek-V4-Flash (29 September 2026); TensorRT-LLM quantisation matrix 1.3.0rc29; model cards of NVIDIA’s Llama 3.3 70B FP8 and NVFP4 and of Kimi K2 Thinking on Hugging Face, read on 10 October 2026.

vLLM’s Llama 4 Scout recipe, written for “NVIDIA Blackwell & Hopper Hardware”, makes the choice explicit. Its launch script selects the FP4 checkpoint only on compute capability 10.0 and the FP8 checkpoint on every other GPU, so an H200 NVL serves Scout in FP8. That checkpoint is 111.6 GB, or 103.9 GiB, and fits one card with about 19 GiB left for the cache by the rule above.

Some models are published with their expert weights in a 4-bit format. vLLM’s DeepSeek-V4-Flash recipe, updated on 29 September 2026, states for the preview checkpoint that “MoE expert weights are stored in FP4” while the other parameters stay in FP8, and its H200 commands serve that checkpoint. It does not say which kernel computes the FP4 experts on Hopper. Kimi K2 Thinking uses integers instead: its model card describes “INT4 weight-only quantization to the MoE components” with quantisation-aware training, saved in the compressed-tensors format. That is the same W4A16 arithmetic as the FP4 fallback, on weights trained for 4-bit storage.

On an L40S, the choice is usually between FP8 and INT4 AWQ. Take Qwen3-32B on a server with eight L40S, one copy per card. Qwen’s FP8 checkpoint is 34.3 GB (32.0 GiB) and its AWQ checkpoint 19.3 GB (18.0 GiB). The L40S has GDDR6 with ECC, and NVIDIA’s CUDA C++ Best Practices Guide states that on GDDR memory with ECC enabled “the available DRAM is reduced by 6.25%”. We therefore start from about 44.8 GiB per card with ECC enabled and 47.8 GiB with ECC off, both our estimates, and apply the same rule. At 1 GiB of FP8 cache per 8K session, the FP8 copy holds about 5 sessions per card and the AWQ copy about 19 with ECC enabled, or 40 against 152 across the server, by our estimate. With ECC off the figures are 8 and 22 per card, and nvidia-smi -q shows the ECC mode of the delivered card. AWQ runs as W4A16 in vLLM and TensorRT-LLM on Ada, without an FP4 fallback.

We build L40S and H200 NVL servers to order and check the rack, power and airflow before we quote. Describe the models and checkpoints you plan to serve in the form below, with your peak number of users.

How to check which kernel an FP4 model uses

  1. Read the checkpoint’s quantisation metadata before you download it: hf_quant_config.json for NVIDIA Model Optimizer checkpoints, or the quantisation section of config.json.
  2. Start vLLM with the checkpoint on one card and search the log for the Marlin weight-only warning, which confirms the fallback path.
  3. Start the FP8 or AWQ checkpoint of the same model with the same engine version, context length and FP8 KV cache.
  4. Load both at your expected peak concurrency and record time to first token and tokens per second per user.
  5. Run your own evaluation set on both, since the accuracy cost differs between models and formats.

What we supply

We supply the H200 NVL, the L40S and the L4, the RTX Ada cards (RTX 6000 Ada, RTX 5880 Ada, RTX 5000 Ada, RTX 4500 Ada and RTX 4000 Ada) and the RTX PRO Blackwell cards, whose FP4 Tensor Cores run NVFP4 and MXFP4 directly. All professional NVIDIA GPUs come with manufacturer warranty on one EU contract and invoice. We build AI servers to order around them, with configuration and quote within one business day, and NVIDIA AI Enterprise licences on the same invoice.

FAQ

Does the H200 support FP4?
The H200 does not support FP4 in hardware. The H200 and H200 NVL are Hopper GPUs with Tensor Cores for FP8, INT8 and 16-bit formats, and NVIDIA’s specification table lists no FP4. FP4 Tensor Cores came with Blackwell, so FP4 checkpoints run on an H200 only through weight-only kernels that compute in 16-bit.
Can I run an NVFP4 model on an H200 NVL?
An NVFP4 model runs on an H200 NVL in vLLM and SGLang, which fall back to weight-only W4A16 execution through Marlin kernels when no native FP4 kernel exists. The weights stay in 4-bit in GPU memory, but the arithmetic runs in 16-bit, and vLLM’s documentation warns that this may reduce throughput for compute-heavy workloads. TensorRT-LLM lists NVFP4 only for Blackwell.
Does the L40S support NVFP4?
The L40S does not support NVFP4 natively, because it is an Ada GPU at compute capability 8.9 with FP8 but no FP4 Tensor Cores. vLLM and SGLang load NVFP4 checkpoints on it as weight-only W4A16 through Marlin, and TensorRT-LLM does not list NVFP4 for Ada. For 4-bit weights on an L40S, INT4 AWQ or GPTQ checkpoints have documented W4A16 and W4A8 paths in TensorRT-LLM.
Does MXFP4 work on Hopper?
MXFP4 works on Hopper for gpt-oss. OpenAI says the MXFP4 release of gpt-oss-120b runs on a single 80 GB GPU such as the H100, vLLM’s recipe lists the H100 and H200, and TensorRT-LLM runs it on the H200 with MXFP4 MoE weights and BF16 activations through its Triton backend. The 4-bit weights are computed at 16-bit precision, since Hopper has no FP4 Tensor Cores.
Can I run an NVFP4 model on an Ada GPU?
An NVFP4 model runs on an Ada GPU as weight-only W4A16 in vLLM, whose quantisation table lists its Marlin FP4 kernels for Turing, Ampere, Ada and Hopper, and in SGLang on SM80 to SM90. The memory saving remains while the maths runs in 16-bit. NVIDIA’s NVFP4 model card for Llama 3.3 70B lists Blackwell as the supported architecture.
Should I use NVFP4 or FP8 on an H200 NVL?
FP8 is the default on Hopper, since the H200 NVL computes it natively and NVIDIA’s Model Optimizer guide suggests FP8 first for larger batches. An NVFP4 checkpoint run weight-only frees memory for more concurrent sessions, about 66 against 44 for Llama 3.3 70B at 8K on one card by our estimate, but at 16-bit arithmetic and an MMLU score of 81.1 against 83.2 on NVIDIA’s model cards. A load test of both at your peak concurrency decides.

Send us the models and checkpoints you plan to serve, the context length, the peak number of concurrent users and the H200 NVL, L40S or other cards you run today. We reply within one business day with the checkpoint format that suits those cards, a configuration and a written quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna