GPU TFLOPS comparison: FP4, FP8 and BF16 per card, dense vs sparse, and the 4 PFLOPS FP8 GPU server
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- NVIDIA prints most Tensor Core figures with 2:4 structured sparsity; the dense rate that ordinary model weights use is half, so the H200 NVL’s 1,671 BF16 and 3,341 FP8 TFLOPS are 835.5 and 1,670.5 dense
- The RTX PRO 6000 Server Edition page lists 4 PFLOPS of FP4, 2 of FP8 and 1 of BF16 without stating sparsity; the figures match the Workstation Edition’s sparse rates, rounded, whose dense FP8 rate is 1,007.6 TFLOPS
- A server’s FP8 or FP4 PFLOPS is the per-card figure times the card count: two RTX PRO 6000 Server Edition cards make 4 PFLOPS of FP8 as listed; if those figures are sparse, a dense 4 PFLOPS takes four, by our arithmetic
- For DGX Spark NVIDIA publishes one compute figure, up to 1 PFLOP of FP4 with sparsity, and no BF16, FP8 or FP32 figure; its GB10 chip has compute capability 12.1
- NVIDIA’s CUDA list gives the H200 compute capability 9.0, every RTX PRO Blackwell card 12.0 and the RTX 6000 Ada, L40S and L4 8.9; only the Blackwell cards list FP4, and the H200 NVL lists 30 TFLOPS of FP64 where the RTX PRO 6000 Workstation and RTX 6000 Ada chips run it at 1/64 of FP32
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
GPU TFLOPS comparison: sparse and dense figures
NVIDIA states most Tensor Core figures “with sparsity”, a rate that counts on weights pruned to a 2:4 pattern; the dense rate, which ordinary model weights use, is half of it. The H200 NVL’s 1,671 BF16 TFLOPS and 3,341 FP8 TFLOPS are therefore 835.5 and 1,670.5 TFLOPS dense. A server’s FP8 or FP4 PFLOPS figure is the per-card figure times the number of cards. Two RTX PRO 6000 Server Edition cards at 2 PFLOPS of FP8 each make a 4 PFLOPS FP8 GPU server on paper. NVIDIA does not say whether those 2 PFLOPS are sparse; if they are, the dense figure is about half.
The tables below give FP4, FP8, BF16, FP32 and FP64 for each card we supply and for DGX Spark, as NVIDIA’s datasheets, product pages and architecture whitepapers state them, read on 10 October 2026. Each figure is marked sparse or dense where NVIDIA marks it, and the compute capability comes from NVIDIA’s CUDA GPU list.
Tensor Core TFLOPS per card: FP4, FP8 and BF16
| CARD | FP4 TENSOR | FP8 TENSOR | BF16 TENSOR | BASIS STATED |
|---|---|---|---|---|
| H200 NVL | no FP4 | 3,341 | 1,671 | sparse |
| RTX PRO 6000 Server Edition | 4 PFLOPS | 2 PFLOPS | 1 PFLOP | not stated |
| RTX PRO 6000 Workstation | 4,030.4 / 2,015.2 | 2,015.2 / 1,007.6 | 1,007.6 / 503.8 | sparse / dense |
| RTX PRO 5000 (48 and 72 GB) | 2,064 | not listed | not listed | sparse |
| RTX PRO 4500 | 1,617 | not listed | not listed | sparse |
| RTX PRO 4500 Server Edition | 1.6 PFLOPS | 811 | 406 | not stated |
| RTX PRO 4000 | 1,178 | not listed | not listed | sparse |
| RTX PRO 2000 | 545 | not listed | not listed | sparse |
| RTX 6000 Ada | no FP4 | 1,457 / 728.5 | 728 / 364 | sparse / dense |
| RTX 5880 Ada | no FP4 | 1,108.4 | not listed | sparse |
| L40S | no FP4 | 1,466 / 733 | 733 / 362.05 | sparse / dense |
| L4 | no FP4 | 485 | 242 | sparse |
| DGX Spark (GB10) | up to 1 PFLOP | not published | not published | sparse |
TFLOPS unless marked. NVIDIA product pages of the H200 (marked “Preliminary specifications”), L40S, L4, RTX 6000 Ada and both Server Editions; RTX PRO datasheets 5349550 (5000), 5108623 (4500), 5323450 (4000) and 5322351 (2000), whose FP4 figure is “AI TOPS”; RTX Blackwell PRO architecture whitepaper v1.1, Table 4; RTX 5880 Ada datasheet 3082478; DGX Spark hardware overview, updated 10 September 2026; all read 10 October 2026.
The RTX PRO Blackwell workstation datasheets give one Tensor figure, in AI TOPS, and footnote it as “Effective FP4 TOPS with sparsity”; the RTX PRO 4000 sheet says “Theoretical FP4 TOPS using the sparsity feature”. NVIDIA publishes no FP8 or BF16 figure for the RTX PRO 5000, 4500, 4000 and 2000, and Table 4 of its whitepaper covers only the RTX PRO 6000 Workstation and Max-Q. For their GB202 chip, each halving of the bits doubles the rate: the Workstation Edition does 503.8 dense TFLOPS in BF16, 1,007.6 in FP8 and 2,015.2 in FP4, and its datasheet rounds the sparse FP4 figure to 4,000 AI TOPS. The RTX PRO 4500 Server Edition page shows the same steps, 1.6 PFLOPS, 811 and 406 TFLOPS. If the smaller cards keep this 4 to 2 to 1 step, dense FP8 is a quarter of the AI TOPS figure: about 516 TFLOPS on the RTX PRO 5000, 404 on the RTX PRO 4500, 295 on the RTX PRO 4000 and 136 on the RTX PRO 2000. These are our arithmetic, not NVIDIA figures.
The Ada cards print FP8 as their headline. The RTX 6000 Ada product page footnotes its 1,457 AI TOPS as “Theoretical FP8 TOPS using the sparsity feature”. The RTX 5880 Ada product page shows 1,108.4 Tensor TFLOPS with only a boost clock footnote, and NVIDIA’s datasheet marks the same figure as effective FP8 with sparsity, 554.2 dense by halving. For the L4, NVIDIA adds “Specifications are one-half lower without sparsity”, so its dense FP8 rate is 242.5 TFLOPS.
Sparse Tensor Core figures and when they apply
NVIDIA’s structured sparsity needs weights in a fixed pattern, described in its blog of 20 July 2021 as “In each contiguous block of four values, two values must be zero.” The Sparse Tensor Cores then skip the zeros and “can complete the same effective calculation in half the time.” NVIDIA describes the route to such weights as a training workflow: prune a dense network to the 2:4 pattern, then “Repeat the original training procedure.” Checkpoints downloaded in BF16, FP8 or NVFP4 are dense unless their publisher pruned them this way, so the dense column is the one to compare. Our guide to FP8, NVFP4, MXFP4 and INT4 explains the formats and which Tensor Cores compute them.
Two product pages state no basis. The RTX PRO 6000 Server Edition page lists 4 PFLOPS of FP4, 2 PFLOPS of FP8 and 1 PFLOP of BF16 without a footnote, and its datasheet (4682150, December 2025) gives only “4 PFLOPS” of peak FP4, with no mention of sparsity. These figures match the Workstation Edition’s sparse rates of 4,030.4, 2,015.2 and 1,007.6 TFLOPS, rounded, but NVIDIA does not say which basis they use. If they are sparse, the Server Edition’s dense FP8 rate is about 1 PFLOPS by our arithmetic, and the server table below marks every value that rests on this assumption. The RTX PRO 4500 Server Edition page lists 1.6 PFLOPS of FP4 without a basis either; the workstation RTX PRO 4500, at the same 51 TFLOPS of FP32, has 1,617 AI TOPS that NVIDIA marks as sparse.
FP32, FP64 and compute capability per card
FP32 here is the single-precision rate of the CUDA cores, without Tensor Cores, which simulation, rendering and many scientific codes use. FP64 decides cards for solvers that need double precision.
| CARD | FP32 TFLOPS | FP64 TFLOPS | COMPUTE CAPABILITY |
|---|---|---|---|
| H200 NVL | 60 | 30, 60 on Tensor Cores | 9.0 |
| RTX PRO 6000 Server Edition | 120 | not listed | 12.0 |
| RTX PRO 6000 Workstation | 125 (whitepaper 126.0) | 1/64 of FP32 | 12.0 |
| RTX PRO 5000 (48 and 72 GB) | 65 | not listed | 12.0 |
| RTX PRO 4500 | 51 | not listed | 12.0 |
| RTX PRO 4500 Server Edition | 51 | not listed | 12.0 |
| RTX PRO 4000 | 37 | not listed | 12.0 |
| RTX PRO 2000 | 17 | not listed | 12.0 |
| RTX 6000 Ada | 91.1 | 1/64 of FP32 | 8.9 |
| RTX 5880 Ada | 69.3 | not listed | not on the list |
| L40S | 91.6 | not listed | 8.9 |
| L4 | 30.3 | not listed | 8.9 |
| DGX Spark (GB10) | not published | not published | 12.1 |
Product pages and datasheets as in the first table; FP64 ratios from NVIDIA’s RTX Blackwell PRO (v1.1) and Ada (v2.02) architecture whitepapers; compute capability from developer.nvidia.com/cuda-gpus, read 10 October 2026, which names the H200 without a separate NVL entry.
Both whitepapers state that “The FP64 TFLOP rate is 1/64th the TFLOP rate of FP32 operations”, for the GB202 chip of the RTX PRO 6000 Workstation and Max-Q and the AD102 chip of the RTX 6000 Ada. The whitepaper does not cover the Server Edition; if it keeps the GB202 ratio, its 120 TFLOPS of FP32 give about 1.9 TFLOPS of FP64 by our arithmetic, against 30 TFLOPS on the H200 NVL. NVIDIA’s CUDA list gives compute capability 9.0 for the H200, 12.0 for every RTX PRO Blackwell card including both Server Editions, 12.1 for GB10 and 8.9 for the RTX 6000 Ada, L40S and L4. FP4 rates appear only for the Blackwell cards, and the whitepaper lists the RTX 6000 Ada’s FP4 rate as “N/A”.
DGX Spark FLOPS: 1 PFLOP of FP4 and no BF16 figure
NVIDIA’s DGX Spark hardware overview, updated on 10 September 2026, gives “up to 1 PFLOP (petaFLOP) at FP4 precision with sparsity” and “Up to 1,000 TOPS (trillion operations per second) inference”, with 6,144 CUDA cores and fifth-generation Tensor Cores. The product page lists “Up to 1 PFLOP FP4” and no BF16, FP8 or FP32 figure. Halving the sparse figure gives about 500 TFLOPS of dense FP4 by our arithmetic. We do not derive a BF16 figure for the Spark, because NVIDIA publishes no ratio between precisions for GB10. Token generation on the Spark follows its 273 GB/s of memory bandwidth rather than these FLOPS, as our DGX Spark benchmarks show with published measurements.
TFLOPS for prompts and training, bandwidth for token generation
NVIDIA’s blog on LLM inference optimisation, of 17 November 2023, describes the prefill phase, which reads the prompt, as “a matrix-matrix operation that’s highly parallelized” that “effectively saturates GPU utilization”. For generating tokens it states “this is a memory-bound operation.” Tensor Core TFLOPS therefore set the pace for long prompts such as RAG requests with retrieved passages, for large batches of concurrent requests and for training and fine-tuning steps. Memory bandwidth sets the tokens per second each user sees. The H200 NVL has 1.66 times the dense FP8 rate of the RTX PRO 6000 Workstation Edition, 1,670.5 against 1,007.6 TFLOPS, but 2.7 times its bandwidth, 4.8 TB/s against 1,792 GB/s. Our GPU memory bandwidth comparison gives the bandwidth of every card, and our H200 NVL and RTX PRO 6000 comparison for fine-tuning applies the tensor rates to training runs.
What 4 PFLOPS of FP8 means for a 4- or 8-card server
A quoted “4 PFLOPS FP8” is a per-card figure multiplied by a card count, and the precision and basis travel with it. NVIDIA’s H200 page shows “4 PetaFLOPS” of FP8 in its quick specs; its full table gives 3,958 TFLOPS with sparsity for the H200 SXM and 3,341 for the H200 NVL, so 4 PFLOPS sits between one and two H200 NVL. On NVIDIA’s listed figures, two RTX PRO 6000 Server Edition cards reach 4 PFLOPS of FP8; if those figures are sparse, a dense 4 PFLOPS takes four.
| SERVER | FP8 AS LISTED | FP8 DENSE | BF16 DENSE | FP4 AS LISTED |
|---|---|---|---|---|
| 4 × RTX PRO 6000 Server | 8 | about 4 if sparse | about 2 if sparse | 16 |
| 8 × RTX PRO 6000 Server | 16 | about 8 if sparse | about 4 if sparse | 32 |
| 4 × H200 NVL | 13.4 | 6.7 | 3.3 | no FP4 |
| 8 × H200 NVL | 26.7 | 13.4 | 6.7 | no FP4 |
| 8 × L40S | 11.7 | 5.9 | 2.9 | no FP4 |
| 8 × RTX PRO 4500 Server | 6.5 | about 3.2 if sparse | about 1.6 if sparse | 12.8 |
PFLOPS per server, our arithmetic from the per-card figures in the first table (card count × figure; dense = half of sparse). “If sparse” marks Server Edition values that assume NVIDIA’s unmarked figures are sparse, which NVIDIA does not state. Peak rates, not measured throughput.
In dense FP8, eight H200 NVL reach 13.4 PFLOPS, against about 8 for eight RTX PRO 6000 Server Edition cards if their listed figures are sparse. With FP4 weights the RTX PRO server lists 32 PFLOPS, while NVIDIA lists no FP4 rate for the H200 NVL, and vLLM runs FP4 checkpoints there with 16-bit arithmetic, as our formats guide explains. These totals are peaks at boost clock that assume every card works on its own share. A model split across cards also waits on the links between them, and our RTX PRO 6000 benchmarks collect published measurements of eight-card servers.
We build AI servers to order with four or eight RTX PRO 6000 Server Edition or H200 NVL cards, and we check the rack, power and airflow before we quote. Send us the model, its precision and your user count through the form below.
Checking TFLOPS in a datasheet or a quote
Four questions make two TFLOPS figures comparable, and they apply to a single card, a server or a cluster quote.
- Which precision: FP4, FP8, BF16 or FP32. “AI TOPS” means FP4 on RTX PRO Blackwell datasheets and FP8 on the RTX 6000 Ada.
- Sparse or dense: halve a figure marked “with sparsity” before comparing it with a dense one. Where a Tensor figure has no basis, ask the vendor, and until it answers compare it with both rates of the other card.
- Per card or per server: divide a server total by its card count, and check whether the count includes cards held in reserve.
- Which clock and power: NVIDIA’s datasheets give “Peak rates based on GPU Boost Clock”, and the 300 W Max-Q Workstation Edition is rated at 877.9 dense FP8 TFLOPS against 1,007.6 for the 600 W Workstation Edition.
If a quote lists PFLOPS without a basis, write to us with the card model and count, and we reply with a configuration and quote within one business day.
What we supply
We supply every card in the two per-card tables, DGX Spark as the Founders Edition and the RTX PRO 6000 Max-Q as well, with manufacturer warranty, on one EU contract and invoice. We build AI servers to order with these cards, assembled and burn-in tested, and size them from the model, its precision and the number of users. NVIDIA AI Enterprise and vGPU licences come on the same invoice. Our professional GPU line-up lists memory, bandwidth and power per card.
FAQ
How many BF16 TFLOPS does the H200 NVL have?
How many TFLOPS does the RTX PRO 6000 have?
How many BF16 FLOPS does DGX Spark have?
What is the difference between dense and sparse TFLOPS?
What does 4 PFLOPS of FP8 mean for a GPU server?
What is the compute capability of the H200?
Send us the model, its precision, how much of the load is long prompts or training, the number of users and the rack position’s power feed. We reply within one business day with a configuration and a quote, with the rack, power and airflow checked before we quote.
Talk to an expertWe reply within one business day