GPU memory bandwidth comparison: HBM3e, GDDR7, GDDR6 and LPDDR5X across the cards we supply
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- For one user without speculative decoding, token generation is capped at memory bandwidth divided by the bytes read per token, so bandwidth ranks GPUs for chat speed before TFLOPS do
- H200 NVL: 141 GB of HBM3e at 4.8 TB/s; RTX PRO 6000: 96 GB of GDDR7 at 1,792 GB/s, or 1,597 GB/s on the Server Edition; L40S: 48 GB of GDDR6 at 864 GB/s
- RTX PRO 5000 1,344 GB/s, RTX PRO 4500 896 GB/s (Server Edition 800 GB/s), RTX PRO 4000 672 GB/s, RTX PRO 2000 288 GB/s; DGX Spark 273 GB/s of LPDDR5x
- GDDR7 runs at 28 Gbps per pin on the RTX PRO 6000 against 20 Gbps for the GDDR6 of the RTX 6000 Ada; per tier, Blackwell raises bandwidth by 29 to 133 per cent
- Llama 3.3 70B in FP8 reads about 70.6 GB per token, a ceiling of about 68 tokens per second on one H200 NVL and 23 on one RTX PRO 6000 Server Edition
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
How GPU memory bandwidth limits tokens per second
Each token a GPU generates requires reading the model’s weights from GPU memory once, plus the stored keys and values of the conversation. For one user, tokens per second therefore cannot exceed memory bandwidth divided by the bytes read per token, unless speculative decoding checks several tokens in one pass. NVIDIA’s blog on LLM inference optimisation of 17 November 2023 calls token generation “a memory-bound operation” and states that “The speed at which the data (weights, keys, values, activations) is transferred to the GPU from memory dominates the latency”.
That makes memory bandwidth the first figure to compare when the job is chat, coding assistants or agents, where users wait for each token. As of October 2026 NVIDIA lists the H200 NVL with 4.8 TB/s of HBM3e. The RTX PRO 6000 Blackwell has 1,792 GB/s of GDDR7 in the Workstation and Max-Q editions and 1,597 GB/s in the Server Edition. The L40S has 864 GB/s of GDDR6, the L4 300 GB/s and DGX Spark 273 GB/s of LPDDR5x.
Tensor Core throughput matters more for long prompts, large batches and training, which our GPU TFLOPS comparison by precision covers card by card.
GPU memory bandwidth comparison: every card we supply
The table gives the figures as NVIDIA publishes them; where NVIDIA states no bus width, the cell says so.
| CARD | MEMORY | BUS WIDTH | BANDWIDTH | PER-PIN DATA RATE |
|---|---|---|---|---|
| H200 NVL | 141 GB HBM3e | not stated | 4,800 GB/s | not stated |
| RTX PRO 6000 Workstation | 96 GB GDDR7 | 512-bit | 1,792 GB/s | 28 Gbps |
| RTX PRO 6000 Max-Q | 96 GB GDDR7 | 512-bit | 1,792 GB/s | 28 Gbps |
| RTX PRO 6000 Server Edition | 96 GB GDDR7 | 512-bit | 1,597 GB/s | about 25 Gbps |
| RTX PRO 5500 (preliminary) | 84 GB GDDR7 | not stated | 1,398 GB/s | not stated |
| RTX PRO 5000 (48 or 72 GB) | 48 or 72 GB GDDR7 | 384-bit | 1,344 GB/s | 28 Gbps |
| RTX PRO 4500 | 32 GB GDDR7 | 256-bit | 896 GB/s | 28 Gbps |
| RTX PRO 4500 Server Edition | 32 GB GDDR7 | 256-bit | 800 GB/s | 25 Gbps |
| RTX PRO 4000 | 24 GB GDDR7 | 192-bit | 672 GB/s | 28 Gbps |
| RTX PRO 4000 SFF | 24 GB GDDR7 | 192-bit | 432 GB/s | 18 Gbps |
| RTX PRO 2000 | 16 GB GDDR7 | 128-bit | 288 GB/s | 18 Gbps |
| RTX 6000 Ada | 48 GB GDDR6 | 384-bit | 960 GB/s | 20 Gbps |
| RTX 5880 Ada | 48 GB GDDR6 | 384-bit | 960 GB/s | 20 Gbps |
| RTX 5000 Ada | 32 GB GDDR6 | 256-bit | 576 GB/s | 18 Gbps |
| RTX 4500 Ada | 24 GB GDDR6 | 192-bit | 432 GB/s | 18 Gbps |
| RTX 4000 Ada | 20 GB GDDR6 | 160-bit | 360 GB/s | 18 Gbps |
| RTX 4000 SFF Ada | 20 GB GDDR6 | 160-bit | 280 GB/s | 14 Gbps |
| RTX 2000 Ada | 16 GB GDDR6 | 128-bit | 224 GB/s | 14 Gbps |
| L40S | 48 GB GDDR6 | not stated | 864 GB/s | not stated |
| L4 | 24 GB GDDR6 | not stated | 300 GB/s | not stated |
| DGX Spark (128 GB) | 128 GB LPDDR5x, unified | 256-bit | 273 GB/s | 8,533 MT/s |
NVIDIA product pages for H200 (preliminary specifications), RTX PRO 6000 and 4500 Server Edition, RTX PRO 5500 (preliminary), L40S and L4, read 10 October 2026; NVIDIA vPC sizing guide (L4 memory type, updated 19 August 2026); NVIDIA datasheets of the RTX PRO 6000 Workstation (5349469, Jun 2026), Max-Q (5349650, Jun 2026), RTX PRO 5000 (5349550, Jul 2026), 4500 Workstation Edition (5108623, Apr 2026), 4000 (5323450, Jun 2026), 4000 SFF (5322500, Jun 2026) and 2000 (5322351, Jun 2026) and of the RTX Ada cards (2023 and 2024); DGX Spark user guide (updated 10 September 2026), which gives “LPDDR5X 8533”. Per-pin rates of 28 and 20 Gbps for the RTX PRO 6000 and RTX 6000 Ada from NVIDIA’s RTX Blackwell PRO architecture paper; the others are our arithmetic, bandwidth times 8 divided by bus width.
The H200 NVL has 3.0 times the bandwidth of the RTX PRO 6000 Server Edition and 5.6 times that of the L40S. By our arithmetic the Server Edition runs its 512-bit bus at about 25 Gbps per pin, where the Workstation and Max-Q cards run the same bus at 28 Gbps, so it has 1,597 instead of 1,792 GB/s. The RTX PRO 4500 Server Edition shows the same pattern; NVIDIA states no data rate for either Server Edition.
NVIDIA’s page lists the RTX PRO 5500 as “Coming Soon”, with “Preliminary product specifications, subject to change”. For DGX Spark (128 GB) the user guide gives “16 channels (256 bit) LPDDR5X 8533”; our DGX Spark benchmarks show what its 273 GB/s means in measured tokens per second.
GDDR7 vs HBM3e, GDDR6 and LPDDR5X: how fast each memory type is
GDDR7 is the memory of every RTX PRO Blackwell card. NVIDIA’s RTX Blackwell PRO architecture paper states that the RTX PRO 6000 Workstation and Max-Q editions “both ship with 96GB of 28 Gbps GDDR7 memory”, and gives 20 Gbps for the GDDR6 of the RTX 6000 Ada. By our arithmetic the RTX PRO 5000, 4500 and 4000 also run at 28 Gbps per pin, and the 70 W RTX PRO 4000 SFF and RTX PRO 2000 at 18 Gbps.
Bandwidth is the per-pin rate times the bus width. The RTX PRO 2000 keeps the 128-bit bus of the RTX 2000 Ada and gains 29 per cent from the faster memory alone, 288 against 224 GB/s. The RTX PRO 5000 combines a 384-bit bus with 28 Gbps and reaches 1,344 GB/s, 2.3 times the 576 GB/s of the RTX 5000 Ada on a 256-bit bus. At the top of the line, 512 bits at 28 Gbps give the RTX PRO 6000 its 1,792 GB/s, 1.87 times the RTX 6000 Ada.
For the HBM3e of the H200 NVL, NVIDIA publishes no bus width or per-pin rate, so the table marks those cells “not stated”. For sizing, the figure that counts is the 4.8 TB/s, 2.7 times the 1,792 GB/s of the RTX PRO 6000 and five times the RTX 6000 Ada.
The LPDDR5X of DGX Spark gives 273 GB/s from a 256-bit interface at 8,533 MT/s, slightly below the RTX PRO 2000, but with 128 GB of capacity shared by CPU and GPU.
From bandwidth to tokens per second: the ceiling for one user
For a dense model and one user, the ceiling in tokens per second is the bandwidth in GB/s divided by the GB read per token. The bytes read per token are the weights used for that token plus the KV cache of the conversation so far. For a mixture-of-experts model only the active experts and the shared layers count, which is why such models generate faster than their total size suggests.
Weight bytes follow from the parameter count and the precision, as our guide to how much VRAM an LLM needs explains: FP8 stores one byte per weight, NVFP4 a little more than half a byte. Halving the bytes per weight roughly doubles the ceiling on the same card. The KV cache grows with context length, so long conversations lower the ceiling as they go on.
The result is an upper bound, because kernels, scheduling and synchronisation take time as well, so measured speeds land below it. With speculative decoding the engine drafts several tokens and verifies them in one pass over the weights, so measured speed per user can exceed this figure. With several users in one batch, each decode step reads the weights once for all of them but every user’s cache separately, so the speed per user falls while the total output of the card rises, until the Tensor Cores rather than memory set the pace.
Prefill, the reading of the prompt, is in NVIDIA’s words “a matrix-matrix operation that’s highly parallelized”, so long RAG prompts depend on the Tensor Core rate more than on memory, as do training and fine-tuning.
Worked example: Llama 3.3 70B in FP8 on three cards
NVIDIA’s FP8 checkpoint of Llama 3.3 70B is 72.7 GB on disk. Its input embedding table, 128,256 tokens by 8,192 values in 16-bit precision, is about 2.1 GB, and only one row of it is read per token. By our arithmetic that leaves about 70.6 GB of weights read for each generated token. The model’s config.json gives 80 layers, 8 KV heads and a head dimension of 128, and its quantisation config sets an FP8 KV cache, so each token of context adds 160 KiB, about 5.4 GB for a conversation of 32,768 tokens.
The model fits on one H200 NVL or one RTX PRO 6000 Server Edition; on L40S cards it needs two, split by tensor parallelism. By the sizing rule of our VRAM guide, the 96 GB card has cache room beside the weights for about three conversations of 32K, the H200 NVL for about eleven.
| CONFIGURATION | BANDWIDTH | SHORT PROMPT | AT 32K CONTEXT |
|---|---|---|---|
| 1 × H200 NVL | 4,800 GB/s | about 68 tok/s | about 63 tok/s |
| 1 × RTX PRO 6000 Server | 1,597 GB/s | about 23 tok/s | about 21 tok/s |
| 2 × L40S, tensor parallel | 1,728 GB/s | about 24 tok/s | about 23 tok/s |
| 4 × RTX PRO 6000 Server | 6,388 GB/s | about 90 tok/s | about 84 tok/s |
Our arithmetic: bandwidth from NVIDIA’s product pages divided by 70.6 GB of weights, plus 5.4 GB of FP8 KV cache at 32K; weights and cache from NVIDIA’s Llama-3.3-70B-Instruct-FP8 files on Hugging Face, read 10 October 2026. Ceilings for one user, not measured speeds; multi-card rows assume tensor parallelism and ignore the time spent exchanging data between cards.
The H200 NVL ceiling is three times that of the RTX PRO 6000 Server Edition because its bandwidth is three times higher; the Tensor Core rates do not enter the calculation. With eight users at 32K context in one batch, the same arithmetic gives each user about 42 tokens per second on the H200 NVL. One RTX PRO 6000 Server Edition cannot hold the 43 GB of cache for those eight, so that load needs the model split across two or more cards. Our article on TTFT and tokens per second targets shows how to set the per-user figure a chat service has to meet before choosing the card.
We build AI servers to order with the H200 NVL, the RTX PRO 6000 Server Edition and the L40S. Tell us the model, its precision and your number of concurrent users, and we reply within one business day with a configuration and quote.
Several GPUs in one server: when bandwidth adds up
An eight-card server has 38.4 TB/s of memory bandwidth with H200 NVL cards and 12.8 TB/s with RTX PRO 6000 Server Edition cards, by our arithmetic. Whether one user sees that total depends on how the model is placed on the cards.
With replicas, each card holds a full copy of the model and serves its own users, so total output grows card by card while each user still gets the ceiling of one card. With pipeline parallelism, each card holds a range of layers and a token passes through the cards one after another; each card reads only its share of the weights, but the reads happen in sequence, so the ceiling per user stays close to that of one card with the whole model. Only tensor parallelism splits every layer across the cards, so that all of them read their part of each token’s weights at the same time; with four cards the ceiling per user rises towards four times that of one card.
Tensor parallelism pays for this with traffic between the cards on every layer. NVIDIA’s H200 page lists 900 GB/s per GPU over the 2- or 4-way NVLink bridge against 128 GB/s for PCIe Gen5; the RTX PRO 6000 and the L40S exchange this data over PCIe. That link bandwidth sets how much of the ceiling a split model keeps. Our comparison of one model on several GPUs over PCIe or NVLink gives the traffic per token and the parallelism to choose per card.
Memory capacity sets the other limit: a replica of the 70B model on each 96 GB card leaves little room for the KV cache of long conversations, while a split across two or four cards frees memory on each of them.
We check the rack, power and airflow before we quote a four- or eight-card server. Describe the rack, the models and the user count in the form below, and we reply with a configuration and quote.
What we supply
We supply every card in this comparison, from the H200 NVL, L40S and L4 to the RTX PRO Blackwell and RTX Ada ranges, and the DGX Spark Founders Edition. Our professional GPU page lists the RTX PRO and data-centre cards with their bandwidth, and our AI servers built to order take two to eight of the server cards, with configuration and quote within one business day. Servers are assembled and burn-in tested, with manufacturer warranty on every component. NVIDIA AI Enterprise and vGPU licences come on the same invoice, on one EU contract.
FAQ
What is the memory bandwidth of the H200 NVL?
What is the RTX PRO 4000 Blackwell memory bandwidth?
What is the RTX PRO 6000 memory bandwidth?
How fast is GDDR7 compared with HBM3e?
How does memory bandwidth relate to tokens per second?
How does DGX Spark memory bandwidth compare with RTX PRO cards?
Send us the model, its precision, the number of concurrent users, the context length you plan for and the rack the servers go into. We reply within one business day with the cards and card count we would use, and a configuration and quote.
Talk to an expertWe reply within one business day