H200 NVL · H100 NVL · SUPPLY IN THE EU

H200 NVL and H100 NVL, supplied in Europe

Identical compute silicon. All the difference lives in memory, and whether that difference is worth paying for depends entirely on how memory-bound your workload is.

AVAILABLE TO ORDER
Both cards, and what sits between them

The premium buys memory, not compute. On some workloads that is 3.4×. On others it is exactly zero. We would rather you knew which before ordering.

CardMemoryBandwidthNotes
H200 NVL141 GB HBM3e4.8 TB/sSix HBM stacks. Where the model no longer fits in 94 GB, this is the whole argument.
H100 NVL94 GB HBM33.9 TB/sStill wider than H100 SXM at 3.35 TB/s. Often enough, and often cheaper per useful token.
NVLink bridgesThe biggest generational difference between NVL pairs and single cards.
NVIDIA AI EnterpriseLicence term supplied with the cards, so it does not surface as a surprise later.

Both cards and the bridges are available to order. Lead time is confirmed with every quote.

JUST THE CARDS?

Cards alone are a fine order

You do not need a project or an engineering contract to buy from us. One H200 NVL to give a server you already own the memory it lacks, or a set of four bridged together, any quantity: we quote exactly what you ask for, at a competitive price, invoiced in the EU.

WHAT COMES WITH IT
Not just a box on a pallet
EU contract and invoice

One counterparty in Vienna. No cross-border purchase to arrange, no import to clear on your side.

Licences sized with the hardware

Where a card needs a licence to do the job you bought it for, it is in the quote, not discovered three months later.

Engineering, if you want it

Our partner Vixen.UNO designs, deploys and supports it. Under the same contract, or not at all – your call.

Warranty and support

Manufacturer warranty handled through us, with support alongside the team who runs the system.

HOW TO CHOOSE
Which one is right for you
The premium is zero on compute-bound work

After TensorRT-LLM optimisations, NVIDIA’s own MLPerf report states that Llama 2 70B on H200 is limited by compute, not by memory bandwidth. If your workload is in that shape, the extra 47 GB buys you nothing at all.

Read the full analysis →
How to read the vendor numbers

The loudest figure in H200 materials is 110× on HPC. The footnote reveals it is four GPUs against a pair of Sapphire Rapids 8480 CPUs, a comparison with processors, not with H100. In another chart H100 ran batch 8 and H200 ran batch 32.

Read the full analysis →
Work out your memory budget first

A 70B in FP16 with a 128K context is 140 GB of weights plus 40 GB of cache: 180 GB for one person. The counter-intuitive part: a 7B with multi-head attention spends 512 KB per token where a 70B with grouped-query attention spends 320.

Read the full analysis →
FAQ
H200 NVL, in short
H200 or H100, which do I need?
Describe the model, the context length and the number of concurrent users. If the model does not fit in 94 GB, the answer is H200. If it does, the honest answer is usually H100 NVL.
Why is the H200 NVL 141 GB and not a round number?
It is 141 GB because of how the six HBM3e stacks are configured. Why the H100 NVL is 94 rather than 96, NVIDIA explains nowhere; we do not have a better answer than that.
Do I need the NVLink bridges?
If you are running one model across a pair of cards, yes. That is where the generational difference actually shows. For independent workloads on separate cards, no.
Is the AI Enterprise licence included?
It is quoted with the cards so the term and the cost are visible up front rather than discovered at deployment.
Can you also supply the server?
Yes, and if you already have one, we would rather check its power and airflow before you order than after.
Check price and availability

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry – see our privacy policy.

request@eurokommerz.at  ·  +43 1 585 1405 50  ·  Jordangasse 7, 1010 Vienna