BLOG · COMPARISON ·

GPU TFLOPS comparison: FP4, FP8 and BF16 per card, dense vs sparse, and the 4 PFLOPS FP8 GPU server

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • NVIDIA prints most Tensor Core figures with 2:4 structured sparsity; the dense rate that ordinary model weights use is half, so the H200 NVL’s 1,671 BF16 and 3,341 FP8 TFLOPS are 835.5 and 1,670.5 dense
  • The RTX PRO 6000 Server Edition page lists 4 PFLOPS of FP4, 2 of FP8 and 1 of BF16 without stating sparsity; the figures match the Workstation Edition’s sparse rates, rounded, whose dense FP8 rate is 1,007.6 TFLOPS
  • A server’s FP8 or FP4 PFLOPS is the per-card figure times the card count: two RTX PRO 6000 Server Edition cards make 4 PFLOPS of FP8 as listed; if those figures are sparse, a dense 4 PFLOPS takes four, by our arithmetic
  • For DGX Spark NVIDIA publishes one compute figure, up to 1 PFLOP of FP4 with sparsity, and no BF16, FP8 or FP32 figure; its GB10 chip has compute capability 12.1
  • NVIDIA’s CUDA list gives the H200 compute capability 9.0, every RTX PRO Blackwell card 12.0 and the RTX 6000 Ada, L40S and L4 8.9; only the Blackwell cards list FP4, and the H200 NVL lists 30 TFLOPS of FP64 where the RTX PRO 6000 Workstation and RTX 6000 Ada chips run it at 1/64 of FP32

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

GPU TFLOPS comparison: sparse and dense figures

NVIDIA states most Tensor Core figures “with sparsity”, a rate that counts on weights pruned to a 2:4 pattern; the dense rate, which ordinary model weights use, is half of it. The H200 NVL’s 1,671 BF16 TFLOPS and 3,341 FP8 TFLOPS are therefore 835.5 and 1,670.5 TFLOPS dense. A server’s FP8 or FP4 PFLOPS figure is the per-card figure times the number of cards. Two RTX PRO 6000 Server Edition cards at 2 PFLOPS of FP8 each make a 4 PFLOPS FP8 GPU server on paper. NVIDIA does not say whether those 2 PFLOPS are sparse; if they are, the dense figure is about half.

The tables below give FP4, FP8, BF16, FP32 and FP64 for each card we supply and for DGX Spark, as NVIDIA’s datasheets, product pages and architecture whitepapers state them, read on 10 October 2026. Each figure is marked sparse or dense where NVIDIA marks it, and the compute capability comes from NVIDIA’s CUDA GPU list.

Tensor Core TFLOPS per card: FP4, FP8 and BF16

CARDFP4 TENSORFP8 TENSORBF16 TENSORBASIS STATED
H200 NVLno FP43,3411,671sparse
RTX PRO 6000 Server Edition4 PFLOPS2 PFLOPS1 PFLOPnot stated
RTX PRO 6000 Workstation4,030.4 / 2,015.22,015.2 / 1,007.61,007.6 / 503.8sparse / dense
RTX PRO 5000 (48 and 72 GB)2,064not listednot listedsparse
RTX PRO 45001,617not listednot listedsparse
RTX PRO 4500 Server Edition1.6 PFLOPS811406not stated
RTX PRO 40001,178not listednot listedsparse
RTX PRO 2000545not listednot listedsparse
RTX 6000 Adano FP41,457 / 728.5728 / 364sparse / dense
RTX 5880 Adano FP41,108.4not listedsparse
L40Sno FP41,466 / 733733 / 362.05sparse / dense
L4no FP4485242sparse
DGX Spark (GB10)up to 1 PFLOPnot publishednot publishedsparse

TFLOPS unless marked. NVIDIA product pages of the H200 (marked “Preliminary specifications”), L40S, L4, RTX 6000 Ada and both Server Editions; RTX PRO datasheets 5349550 (5000), 5108623 (4500), 5323450 (4000) and 5322351 (2000), whose FP4 figure is “AI TOPS”; RTX Blackwell PRO architecture whitepaper v1.1, Table 4; RTX 5880 Ada datasheet 3082478; DGX Spark hardware overview, updated 10 September 2026; all read 10 October 2026.

The RTX PRO Blackwell workstation datasheets give one Tensor figure, in AI TOPS, and footnote it as “Effective FP4 TOPS with sparsity”; the RTX PRO 4000 sheet says “Theoretical FP4 TOPS using the sparsity feature”. NVIDIA publishes no FP8 or BF16 figure for the RTX PRO 5000, 4500, 4000 and 2000, and Table 4 of its whitepaper covers only the RTX PRO 6000 Workstation and Max-Q. For their GB202 chip, each halving of the bits doubles the rate: the Workstation Edition does 503.8 dense TFLOPS in BF16, 1,007.6 in FP8 and 2,015.2 in FP4, and its datasheet rounds the sparse FP4 figure to 4,000 AI TOPS. The RTX PRO 4500 Server Edition page shows the same steps, 1.6 PFLOPS, 811 and 406 TFLOPS. If the smaller cards keep this 4 to 2 to 1 step, dense FP8 is a quarter of the AI TOPS figure: about 516 TFLOPS on the RTX PRO 5000, 404 on the RTX PRO 4500, 295 on the RTX PRO 4000 and 136 on the RTX PRO 2000. These are our arithmetic, not NVIDIA figures.

The Ada cards print FP8 as their headline. The RTX 6000 Ada product page footnotes its 1,457 AI TOPS as “Theoretical FP8 TOPS using the sparsity feature”. The RTX 5880 Ada product page shows 1,108.4 Tensor TFLOPS with only a boost clock footnote, and NVIDIA’s datasheet marks the same figure as effective FP8 with sparsity, 554.2 dense by halving. For the L4, NVIDIA adds “Specifications are one-half lower without sparsity”, so its dense FP8 rate is 242.5 TFLOPS.

Sparse Tensor Core figures and when they apply

NVIDIA’s structured sparsity needs weights in a fixed pattern, described in its blog of 20 July 2021 as “In each contiguous block of four values, two values must be zero.” The Sparse Tensor Cores then skip the zeros and “can complete the same effective calculation in half the time.” NVIDIA describes the route to such weights as a training workflow: prune a dense network to the 2:4 pattern, then “Repeat the original training procedure.” Checkpoints downloaded in BF16, FP8 or NVFP4 are dense unless their publisher pruned them this way, so the dense column is the one to compare. Our guide to FP8, NVFP4, MXFP4 and INT4 explains the formats and which Tensor Cores compute them.

Two product pages state no basis. The RTX PRO 6000 Server Edition page lists 4 PFLOPS of FP4, 2 PFLOPS of FP8 and 1 PFLOP of BF16 without a footnote, and its datasheet (4682150, December 2025) gives only “4 PFLOPS” of peak FP4, with no mention of sparsity. These figures match the Workstation Edition’s sparse rates of 4,030.4, 2,015.2 and 1,007.6 TFLOPS, rounded, but NVIDIA does not say which basis they use. If they are sparse, the Server Edition’s dense FP8 rate is about 1 PFLOPS by our arithmetic, and the server table below marks every value that rests on this assumption. The RTX PRO 4500 Server Edition page lists 1.6 PFLOPS of FP4 without a basis either; the workstation RTX PRO 4500, at the same 51 TFLOPS of FP32, has 1,617 AI TOPS that NVIDIA marks as sparse.

FP32, FP64 and compute capability per card

FP32 here is the single-precision rate of the CUDA cores, without Tensor Cores, which simulation, rendering and many scientific codes use. FP64 decides cards for solvers that need double precision.

CARDFP32 TFLOPSFP64 TFLOPSCOMPUTE CAPABILITY
H200 NVL6030, 60 on Tensor Cores9.0
RTX PRO 6000 Server Edition120not listed12.0
RTX PRO 6000 Workstation125 (whitepaper 126.0)1/64 of FP3212.0
RTX PRO 5000 (48 and 72 GB)65not listed12.0
RTX PRO 450051not listed12.0
RTX PRO 4500 Server Edition51not listed12.0
RTX PRO 400037not listed12.0
RTX PRO 200017not listed12.0
RTX 6000 Ada91.11/64 of FP328.9
RTX 5880 Ada69.3not listednot on the list
L40S91.6not listed8.9
L430.3not listed8.9
DGX Spark (GB10)not publishednot published12.1

Product pages and datasheets as in the first table; FP64 ratios from NVIDIA’s RTX Blackwell PRO (v1.1) and Ada (v2.02) architecture whitepapers; compute capability from developer.nvidia.com/cuda-gpus, read 10 October 2026, which names the H200 without a separate NVL entry.

Both whitepapers state that “The FP64 TFLOP rate is 1/64th the TFLOP rate of FP32 operations”, for the GB202 chip of the RTX PRO 6000 Workstation and Max-Q and the AD102 chip of the RTX 6000 Ada. The whitepaper does not cover the Server Edition; if it keeps the GB202 ratio, its 120 TFLOPS of FP32 give about 1.9 TFLOPS of FP64 by our arithmetic, against 30 TFLOPS on the H200 NVL. NVIDIA’s CUDA list gives compute capability 9.0 for the H200, 12.0 for every RTX PRO Blackwell card including both Server Editions, 12.1 for GB10 and 8.9 for the RTX 6000 Ada, L40S and L4. FP4 rates appear only for the Blackwell cards, and the whitepaper lists the RTX 6000 Ada’s FP4 rate as “N/A”.

DGX Spark FLOPS: 1 PFLOP of FP4 and no BF16 figure

NVIDIA’s DGX Spark hardware overview, updated on 10 September 2026, gives “up to 1 PFLOP (petaFLOP) at FP4 precision with sparsity” and “Up to 1,000 TOPS (trillion operations per second) inference”, with 6,144 CUDA cores and fifth-generation Tensor Cores. The product page lists “Up to 1 PFLOP FP4” and no BF16, FP8 or FP32 figure. Halving the sparse figure gives about 500 TFLOPS of dense FP4 by our arithmetic. We do not derive a BF16 figure for the Spark, because NVIDIA publishes no ratio between precisions for GB10. Token generation on the Spark follows its 273 GB/s of memory bandwidth rather than these FLOPS, as our DGX Spark benchmarks show with published measurements.

TFLOPS for prompts and training, bandwidth for token generation

NVIDIA’s blog on LLM inference optimisation, of 17 November 2023, describes the prefill phase, which reads the prompt, as “a matrix-matrix operation that’s highly parallelized” that “effectively saturates GPU utilization”. For generating tokens it states “this is a memory-bound operation.” Tensor Core TFLOPS therefore set the pace for long prompts such as RAG requests with retrieved passages, for large batches of concurrent requests and for training and fine-tuning steps. Memory bandwidth sets the tokens per second each user sees. The H200 NVL has 1.66 times the dense FP8 rate of the RTX PRO 6000 Workstation Edition, 1,670.5 against 1,007.6 TFLOPS, but 2.7 times its bandwidth, 4.8 TB/s against 1,792 GB/s. Our GPU memory bandwidth comparison gives the bandwidth of every card, and our H200 NVL and RTX PRO 6000 comparison for fine-tuning applies the tensor rates to training runs.

What 4 PFLOPS of FP8 means for a 4- or 8-card server

A quoted “4 PFLOPS FP8” is a per-card figure multiplied by a card count, and the precision and basis travel with it. NVIDIA’s H200 page shows “4 PetaFLOPS” of FP8 in its quick specs; its full table gives 3,958 TFLOPS with sparsity for the H200 SXM and 3,341 for the H200 NVL, so 4 PFLOPS sits between one and two H200 NVL. On NVIDIA’s listed figures, two RTX PRO 6000 Server Edition cards reach 4 PFLOPS of FP8; if those figures are sparse, a dense 4 PFLOPS takes four.

SERVERFP8 AS LISTEDFP8 DENSEBF16 DENSEFP4 AS LISTED
4 × RTX PRO 6000 Server8about 4 if sparseabout 2 if sparse16
8 × RTX PRO 6000 Server16about 8 if sparseabout 4 if sparse32
4 × H200 NVL13.46.73.3no FP4
8 × H200 NVL26.713.46.7no FP4
8 × L40S11.75.92.9no FP4
8 × RTX PRO 4500 Server6.5about 3.2 if sparseabout 1.6 if sparse12.8

PFLOPS per server, our arithmetic from the per-card figures in the first table (card count × figure; dense = half of sparse). “If sparse” marks Server Edition values that assume NVIDIA’s unmarked figures are sparse, which NVIDIA does not state. Peak rates, not measured throughput.

In dense FP8, eight H200 NVL reach 13.4 PFLOPS, against about 8 for eight RTX PRO 6000 Server Edition cards if their listed figures are sparse. With FP4 weights the RTX PRO server lists 32 PFLOPS, while NVIDIA lists no FP4 rate for the H200 NVL, and vLLM runs FP4 checkpoints there with 16-bit arithmetic, as our formats guide explains. These totals are peaks at boost clock that assume every card works on its own share. A model split across cards also waits on the links between them, and our RTX PRO 6000 benchmarks collect published measurements of eight-card servers.

We build AI servers to order with four or eight RTX PRO 6000 Server Edition or H200 NVL cards, and we check the rack, power and airflow before we quote. Send us the model, its precision and your user count through the form below.

Checking TFLOPS in a datasheet or a quote

Four questions make two TFLOPS figures comparable, and they apply to a single card, a server or a cluster quote.

  1. Which precision: FP4, FP8, BF16 or FP32. “AI TOPS” means FP4 on RTX PRO Blackwell datasheets and FP8 on the RTX 6000 Ada.
  2. Sparse or dense: halve a figure marked “with sparsity” before comparing it with a dense one. Where a Tensor figure has no basis, ask the vendor, and until it answers compare it with both rates of the other card.
  3. Per card or per server: divide a server total by its card count, and check whether the count includes cards held in reserve.
  4. Which clock and power: NVIDIA’s datasheets give “Peak rates based on GPU Boost Clock”, and the 300 W Max-Q Workstation Edition is rated at 877.9 dense FP8 TFLOPS against 1,007.6 for the 600 W Workstation Edition.

If a quote lists PFLOPS without a basis, write to us with the card model and count, and we reply with a configuration and quote within one business day.

What we supply

We supply every card in the two per-card tables, DGX Spark as the Founders Edition and the RTX PRO 6000 Max-Q as well, with manufacturer warranty, on one EU contract and invoice. We build AI servers to order with these cards, assembled and burn-in tested, and size them from the model, its precision and the number of users. NVIDIA AI Enterprise and vGPU licences come on the same invoice. Our professional GPU line-up lists memory, bandwidth and power per card.

FAQ

How many BF16 TFLOPS does the H200 NVL have?
NVIDIA lists the H200 NVL at 1,671 BF16 Tensor Core TFLOPS, marked “With sparsity”, which is 835.5 TFLOPS dense by halving. Its FP8 figure is 3,341 TFLOPS with sparsity, its FP32 rate 60 TFLOPS and its FP64 rate 30 TFLOPS, or 60 on the Tensor Cores.
How many TFLOPS does the RTX PRO 6000 have?
NVIDIA’s Server Edition page lists 120 TFLOPS of FP32, 4 PFLOPS of FP4, 2 PFLOPS of FP8 and 1 PFLOP of BF16 without stating whether the Tensor figures are sparse; they match the Workstation Edition’s sparse rates, rounded. The architecture whitepaper gives the Workstation Edition 126.0 TFLOPS of FP32, 125 on its datasheet, and 503.8 BF16, 1,007.6 FP8 and 2,015.2 FP4 TFLOPS dense, twice that with sparsity.
How many BF16 FLOPS does DGX Spark have?
NVIDIA publishes no BF16 figure for DGX Spark. Its only compute figure is up to 1 PFLOP of FP4 with sparsity, about 500 TFLOPS dense by halving, from a GB10 chip with 6,144 CUDA cores and compute capability 12.1. Its token generation speed follows its 273 GB/s of memory bandwidth more than its FLOPS.
What is the difference between dense and sparse TFLOPS?
Sparse figures assume weights pruned so that two of every four values are zero, which NVIDIA’s Sparse Tensor Cores process in half the time. Ordinary model checkpoints are dense, so the dense rate, half the sparse one, is the figure to compare. The L40S, for example, is listed at 733 dense and 1,466 sparse FP8 TFLOPS.
What does 4 PFLOPS of FP8 mean for a GPU server?
It is the per-card FP8 figure times the number of cards, and NVIDIA’s H200 page shows 4 PFLOPS of FP8 for one H200 SXM, a sparse figure. Two RTX PRO 6000 Server Edition cards at 2 PFLOPS each reach 4 PFLOPS as NVIDIA lists them, a figure without a stated basis. Eight H200 NVL list 26.7 PFLOPS of sparse FP8, or 13.4 dense.
What is the compute capability of the H200?
NVIDIA’s CUDA GPU list gives the H200 compute capability 9.0, without a separate entry for the NVL edition. The RTX PRO Blackwell cards have 12.0, the GB10 chip of DGX Spark 12.1, and the RTX 6000 Ada, L40S and L4 8.9. NVIDIA lists FP4 rates only for the Blackwell cards, so FP4 Tensor Core arithmetic needs 12.0 or 12.1 among the cards here.

Send us the model, its precision, how much of the load is long prompts or training, the number of users and the rack position’s power feed. We reply within one business day with a configuration and a quote, with the rack, power and airflow checked before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna