BLOG · COMPARISON ·

GPU memory bandwidth comparison: HBM3e, GDDR7, GDDR6 and LPDDR5X across the cards we supply

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • For one user without speculative decoding, token generation is capped at memory bandwidth divided by the bytes read per token, so bandwidth ranks GPUs for chat speed before TFLOPS do
  • H200 NVL: 141 GB of HBM3e at 4.8 TB/s; RTX PRO 6000: 96 GB of GDDR7 at 1,792 GB/s, or 1,597 GB/s on the Server Edition; L40S: 48 GB of GDDR6 at 864 GB/s
  • RTX PRO 5000 1,344 GB/s, RTX PRO 4500 896 GB/s (Server Edition 800 GB/s), RTX PRO 4000 672 GB/s, RTX PRO 2000 288 GB/s; DGX Spark 273 GB/s of LPDDR5x
  • GDDR7 runs at 28 Gbps per pin on the RTX PRO 6000 against 20 Gbps for the GDDR6 of the RTX 6000 Ada; per tier, Blackwell raises bandwidth by 29 to 133 per cent
  • Llama 3.3 70B in FP8 reads about 70.6 GB per token, a ceiling of about 68 tokens per second on one H200 NVL and 23 on one RTX PRO 6000 Server Edition

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

How GPU memory bandwidth limits tokens per second

Each token a GPU generates requires reading the model’s weights from GPU memory once, plus the stored keys and values of the conversation. For one user, tokens per second therefore cannot exceed memory bandwidth divided by the bytes read per token, unless speculative decoding checks several tokens in one pass. NVIDIA’s blog on LLM inference optimisation of 17 November 2023 calls token generation “a memory-bound operation” and states that “The speed at which the data (weights, keys, values, activations) is transferred to the GPU from memory dominates the latency”.

That makes memory bandwidth the first figure to compare when the job is chat, coding assistants or agents, where users wait for each token. As of October 2026 NVIDIA lists the H200 NVL with 4.8 TB/s of HBM3e. The RTX PRO 6000 Blackwell has 1,792 GB/s of GDDR7 in the Workstation and Max-Q editions and 1,597 GB/s in the Server Edition. The L40S has 864 GB/s of GDDR6, the L4 300 GB/s and DGX Spark 273 GB/s of LPDDR5x.

Tensor Core throughput matters more for long prompts, large batches and training, which our GPU TFLOPS comparison by precision covers card by card.

GPU memory bandwidth comparison: every card we supply

The table gives the figures as NVIDIA publishes them; where NVIDIA states no bus width, the cell says so.

CARDMEMORYBUS WIDTHBANDWIDTHPER-PIN DATA RATE
H200 NVL141 GB HBM3enot stated4,800 GB/snot stated
RTX PRO 6000 Workstation96 GB GDDR7512-bit1,792 GB/s28 Gbps
RTX PRO 6000 Max-Q96 GB GDDR7512-bit1,792 GB/s28 Gbps
RTX PRO 6000 Server Edition96 GB GDDR7512-bit1,597 GB/sabout 25 Gbps
RTX PRO 5500 (preliminary)84 GB GDDR7not stated1,398 GB/snot stated
RTX PRO 5000 (48 or 72 GB)48 or 72 GB GDDR7384-bit1,344 GB/s28 Gbps
RTX PRO 450032 GB GDDR7256-bit896 GB/s28 Gbps
RTX PRO 4500 Server Edition32 GB GDDR7256-bit800 GB/s25 Gbps
RTX PRO 400024 GB GDDR7192-bit672 GB/s28 Gbps
RTX PRO 4000 SFF24 GB GDDR7192-bit432 GB/s18 Gbps
RTX PRO 200016 GB GDDR7128-bit288 GB/s18 Gbps
RTX 6000 Ada48 GB GDDR6384-bit960 GB/s20 Gbps
RTX 5880 Ada48 GB GDDR6384-bit960 GB/s20 Gbps
RTX 5000 Ada32 GB GDDR6256-bit576 GB/s18 Gbps
RTX 4500 Ada24 GB GDDR6192-bit432 GB/s18 Gbps
RTX 4000 Ada20 GB GDDR6160-bit360 GB/s18 Gbps
RTX 4000 SFF Ada20 GB GDDR6160-bit280 GB/s14 Gbps
RTX 2000 Ada16 GB GDDR6128-bit224 GB/s14 Gbps
L40S48 GB GDDR6not stated864 GB/snot stated
L424 GB GDDR6not stated300 GB/snot stated
DGX Spark (128 GB)128 GB LPDDR5x, unified256-bit273 GB/s8,533 MT/s

NVIDIA product pages for H200 (preliminary specifications), RTX PRO 6000 and 4500 Server Edition, RTX PRO 5500 (preliminary), L40S and L4, read 10 October 2026; NVIDIA vPC sizing guide (L4 memory type, updated 19 August 2026); NVIDIA datasheets of the RTX PRO 6000 Workstation (5349469, Jun 2026), Max-Q (5349650, Jun 2026), RTX PRO 5000 (5349550, Jul 2026), 4500 Workstation Edition (5108623, Apr 2026), 4000 (5323450, Jun 2026), 4000 SFF (5322500, Jun 2026) and 2000 (5322351, Jun 2026) and of the RTX Ada cards (2023 and 2024); DGX Spark user guide (updated 10 September 2026), which gives “LPDDR5X 8533”. Per-pin rates of 28 and 20 Gbps for the RTX PRO 6000 and RTX 6000 Ada from NVIDIA’s RTX Blackwell PRO architecture paper; the others are our arithmetic, bandwidth times 8 divided by bus width.

The H200 NVL has 3.0 times the bandwidth of the RTX PRO 6000 Server Edition and 5.6 times that of the L40S. By our arithmetic the Server Edition runs its 512-bit bus at about 25 Gbps per pin, where the Workstation and Max-Q cards run the same bus at 28 Gbps, so it has 1,597 instead of 1,792 GB/s. The RTX PRO 4500 Server Edition shows the same pattern; NVIDIA states no data rate for either Server Edition.

NVIDIA’s page lists the RTX PRO 5500 as “Coming Soon”, with “Preliminary product specifications, subject to change”. For DGX Spark (128 GB) the user guide gives “16 channels (256 bit) LPDDR5X 8533”; our DGX Spark benchmarks show what its 273 GB/s means in measured tokens per second.

GDDR7 vs HBM3e, GDDR6 and LPDDR5X: how fast each memory type is

GDDR7 is the memory of every RTX PRO Blackwell card. NVIDIA’s RTX Blackwell PRO architecture paper states that the RTX PRO 6000 Workstation and Max-Q editions “both ship with 96GB of 28 Gbps GDDR7 memory”, and gives 20 Gbps for the GDDR6 of the RTX 6000 Ada. By our arithmetic the RTX PRO 5000, 4500 and 4000 also run at 28 Gbps per pin, and the 70 W RTX PRO 4000 SFF and RTX PRO 2000 at 18 Gbps.

Bandwidth is the per-pin rate times the bus width. The RTX PRO 2000 keeps the 128-bit bus of the RTX 2000 Ada and gains 29 per cent from the faster memory alone, 288 against 224 GB/s. The RTX PRO 5000 combines a 384-bit bus with 28 Gbps and reaches 1,344 GB/s, 2.3 times the 576 GB/s of the RTX 5000 Ada on a 256-bit bus. At the top of the line, 512 bits at 28 Gbps give the RTX PRO 6000 its 1,792 GB/s, 1.87 times the RTX 6000 Ada.

For the HBM3e of the H200 NVL, NVIDIA publishes no bus width or per-pin rate, so the table marks those cells “not stated”. For sizing, the figure that counts is the 4.8 TB/s, 2.7 times the 1,792 GB/s of the RTX PRO 6000 and five times the RTX 6000 Ada.

The LPDDR5X of DGX Spark gives 273 GB/s from a 256-bit interface at 8,533 MT/s, slightly below the RTX PRO 2000, but with 128 GB of capacity shared by CPU and GPU.

From bandwidth to tokens per second: the ceiling for one user

For a dense model and one user, the ceiling in tokens per second is the bandwidth in GB/s divided by the GB read per token. The bytes read per token are the weights used for that token plus the KV cache of the conversation so far. For a mixture-of-experts model only the active experts and the shared layers count, which is why such models generate faster than their total size suggests.

Weight bytes follow from the parameter count and the precision, as our guide to how much VRAM an LLM needs explains: FP8 stores one byte per weight, NVFP4 a little more than half a byte. Halving the bytes per weight roughly doubles the ceiling on the same card. The KV cache grows with context length, so long conversations lower the ceiling as they go on.

The result is an upper bound, because kernels, scheduling and synchronisation take time as well, so measured speeds land below it. With speculative decoding the engine drafts several tokens and verifies them in one pass over the weights, so measured speed per user can exceed this figure. With several users in one batch, each decode step reads the weights once for all of them but every user’s cache separately, so the speed per user falls while the total output of the card rises, until the Tensor Cores rather than memory set the pace.

Prefill, the reading of the prompt, is in NVIDIA’s words “a matrix-matrix operation that’s highly parallelized”, so long RAG prompts depend on the Tensor Core rate more than on memory, as do training and fine-tuning.

Worked example: Llama 3.3 70B in FP8 on three cards

NVIDIA’s FP8 checkpoint of Llama 3.3 70B is 72.7 GB on disk. Its input embedding table, 128,256 tokens by 8,192 values in 16-bit precision, is about 2.1 GB, and only one row of it is read per token. By our arithmetic that leaves about 70.6 GB of weights read for each generated token. The model’s config.json gives 80 layers, 8 KV heads and a head dimension of 128, and its quantisation config sets an FP8 KV cache, so each token of context adds 160 KiB, about 5.4 GB for a conversation of 32,768 tokens.

The model fits on one H200 NVL or one RTX PRO 6000 Server Edition; on L40S cards it needs two, split by tensor parallelism. By the sizing rule of our VRAM guide, the 96 GB card has cache room beside the weights for about three conversations of 32K, the H200 NVL for about eleven.

CONFIGURATIONBANDWIDTHSHORT PROMPTAT 32K CONTEXT
1 × H200 NVL4,800 GB/sabout 68 tok/sabout 63 tok/s
1 × RTX PRO 6000 Server1,597 GB/sabout 23 tok/sabout 21 tok/s
2 × L40S, tensor parallel1,728 GB/sabout 24 tok/sabout 23 tok/s
4 × RTX PRO 6000 Server6,388 GB/sabout 90 tok/sabout 84 tok/s

Our arithmetic: bandwidth from NVIDIA’s product pages divided by 70.6 GB of weights, plus 5.4 GB of FP8 KV cache at 32K; weights and cache from NVIDIA’s Llama-3.3-70B-Instruct-FP8 files on Hugging Face, read 10 October 2026. Ceilings for one user, not measured speeds; multi-card rows assume tensor parallelism and ignore the time spent exchanging data between cards.

The H200 NVL ceiling is three times that of the RTX PRO 6000 Server Edition because its bandwidth is three times higher; the Tensor Core rates do not enter the calculation. With eight users at 32K context in one batch, the same arithmetic gives each user about 42 tokens per second on the H200 NVL. One RTX PRO 6000 Server Edition cannot hold the 43 GB of cache for those eight, so that load needs the model split across two or more cards. Our article on TTFT and tokens per second targets shows how to set the per-user figure a chat service has to meet before choosing the card.

We build AI servers to order with the H200 NVL, the RTX PRO 6000 Server Edition and the L40S. Tell us the model, its precision and your number of concurrent users, and we reply within one business day with a configuration and quote.

Several GPUs in one server: when bandwidth adds up

An eight-card server has 38.4 TB/s of memory bandwidth with H200 NVL cards and 12.8 TB/s with RTX PRO 6000 Server Edition cards, by our arithmetic. Whether one user sees that total depends on how the model is placed on the cards.

With replicas, each card holds a full copy of the model and serves its own users, so total output grows card by card while each user still gets the ceiling of one card. With pipeline parallelism, each card holds a range of layers and a token passes through the cards one after another; each card reads only its share of the weights, but the reads happen in sequence, so the ceiling per user stays close to that of one card with the whole model. Only tensor parallelism splits every layer across the cards, so that all of them read their part of each token’s weights at the same time; with four cards the ceiling per user rises towards four times that of one card.

Tensor parallelism pays for this with traffic between the cards on every layer. NVIDIA’s H200 page lists 900 GB/s per GPU over the 2- or 4-way NVLink bridge against 128 GB/s for PCIe Gen5; the RTX PRO 6000 and the L40S exchange this data over PCIe. That link bandwidth sets how much of the ceiling a split model keeps. Our comparison of one model on several GPUs over PCIe or NVLink gives the traffic per token and the parallelism to choose per card.

Memory capacity sets the other limit: a replica of the 70B model on each 96 GB card leaves little room for the KV cache of long conversations, while a split across two or four cards frees memory on each of them.

We check the rack, power and airflow before we quote a four- or eight-card server. Describe the rack, the models and the user count in the form below, and we reply with a configuration and quote.

What we supply

We supply every card in this comparison, from the H200 NVL, L40S and L4 to the RTX PRO Blackwell and RTX Ada ranges, and the DGX Spark Founders Edition. Our professional GPU page lists the RTX PRO and data-centre cards with their bandwidth, and our AI servers built to order take two to eight of the server cards, with configuration and quote within one business day. Servers are assembled and burn-in tested, with manufacturer warranty on every component. NVIDIA AI Enterprise and vGPU licences come on the same invoice, on one EU contract.

FAQ

What is the memory bandwidth of the H200 NVL?
NVIDIA lists the H200 NVL with 141 GB of HBM3e at 4.8 TB/s, the same bandwidth as the H200 SXM. That is 3.0 times the 1,597 GB/s of the RTX PRO 6000 Server Edition and 5.6 times the 864 GB/s of the L40S.
What is the RTX PRO 4000 Blackwell memory bandwidth?
The RTX PRO 4000 Blackwell has 24 GB of GDDR7 on a 192-bit bus at 672 GB/s, according to NVIDIA’s datasheet of June 2026. The RTX PRO 4000 SFF has the same 24 GB and bus at 432 GB/s, and the RTX 4000 Ada it replaces has 360 GB/s.
What is the RTX PRO 6000 memory bandwidth?
The RTX PRO 6000 Blackwell Workstation and Max-Q editions have 96 GB of GDDR7 on a 512-bit bus at 1,792 GB/s. The Server Edition has the same memory and bus at 1,597 GB/s, which NVIDIA lists on its product page.
How fast is GDDR7 compared with HBM3e?
NVIDIA states 28 Gbps per pin for the GDDR7 of the RTX PRO 6000, which gives 1,792 GB/s on its 512-bit bus. The HBM3e of the H200 NVL reaches 4.8 TB/s, 2.7 times as much; NVIDIA does not publish the bus width or per-pin rate of that memory.
How does memory bandwidth relate to tokens per second?
For one user, each generated token reads the model weights and the KV cache once, so without speculative decoding tokens per second cannot exceed bandwidth divided by the bytes read per token. For Llama 3.3 70B in FP8, about 70.6 GB per token, that gives ceilings of about 68 tokens per second on an H200 NVL and 23 on an RTX PRO 6000 Server Edition; measured speeds are lower.
How does DGX Spark memory bandwidth compare with RTX PRO cards?
DGX Spark (128 GB) has 273 GB/s of LPDDR5x unified memory on a 256-bit interface, a little below the 288 GB/s of the RTX PRO 2000 and far below the 1,792 GB/s of the RTX PRO 6000. Its 128 GB, shared by CPU and GPU, loads models that do not fit on any single RTX PRO card below the 6000.

Send us the model, its precision, the number of concurrent users, the context length you plan for and the rack the servers go into. We reply within one business day with the cards and card count we would use, and a configuration and quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna