H200 NVL and H100 NVL, supplied in Europe
Identical compute silicon. All the difference lives in memory, and whether that difference is worth paying for depends entirely on how memory-bound your workload is.
The premium buys memory, not compute. On some workloads that is 3.4×. On others it is exactly zero. We would rather you knew which before ordering.
| Card | Memory | Bandwidth | Notes |
|---|---|---|---|
| H200 NVL | 141 GB HBM3e | 4.8 TB/s | Six HBM stacks. Where the model no longer fits in 94 GB, this is the whole argument. |
| H100 NVL | 94 GB HBM3 | 3.9 TB/s | Still wider than H100 SXM at 3.35 TB/s. Often enough, and often cheaper per useful token. |
| NVLink bridges | – | – | The biggest generational difference between NVL pairs and single cards. |
| NVIDIA AI Enterprise | – | – | Licence term supplied with the cards, so it does not surface as a surprise later. |
Both cards and the bridges are available to order. Lead time is confirmed with every quote.
Cards alone are a fine order
You do not need a project or an engineering contract to buy from us. One H200 NVL to give a server you already own the memory it lacks, or a set of four bridged together, any quantity: we quote exactly what you ask for, at a competitive price, invoiced in the EU.
One counterparty in Vienna. No cross-border purchase to arrange, no import to clear on your side.
Where a card needs a licence to do the job you bought it for, it is in the quote, not discovered three months later.
Our partner Vixen.UNO designs, deploys and supports it. Under the same contract, or not at all – your call.
Manufacturer warranty handled through us, with support alongside the team who runs the system.
After TensorRT-LLM optimisations, NVIDIA’s own MLPerf report states that Llama 2 70B on H200 is limited by compute, not by memory bandwidth. If your workload is in that shape, the extra 47 GB buys you nothing at all.
Read the full analysis →The loudest figure in H200 materials is 110× on HPC. The footnote reveals it is four GPUs against a pair of Sapphire Rapids 8480 CPUs, a comparison with processors, not with H100. In another chart H100 ran batch 8 and H200 ran batch 32.
Read the full analysis →A 70B in FP16 with a 128K context is 140 GB of weights plus 40 GB of cache: 180 GB for one person. The counter-intuitive part: a 7B with multi-head attention spends 512 KB per token where a 70B with grouped-query attention spends 320.
Read the full analysis →H200 or H100, which do I need?
Why is the H200 NVL 141 GB and not a round number?
Do I need the NVLink bridges?
Is the AI Enterprise licence included?
Can you also supply the server?
141 vs 94 GB, 4.8 vs 3.9 TB/s: where the gain reaches 3.4× and where it is exactly zero.
Read the article →The working formulas for weights, KV cache and the overhead nobody puts on a slide.
Read the article →What to check before the servers arrive: the four things that stop a delivery.
Read the article →We reply within one business day