BLOG · HARDWARE REVIEW ·

NVIDIA RTX PRO 5000 Blackwell, 48 GB and 72 GB: real benchmarks and limits

IN BRIEF
  • The RTX PRO 5000 uses the RTX PRO 6000’s GB202 chip and has 14,080 CUDA cores, 1,344 GB/s and 300 W; its 48 GB and 72 GB versions differ in memory, MIG instance size and vGPU support
  • In December 2025 Puget Systems measured it 50 per cent ahead of the RTX 5000 Ada in Blender, 35 per cent in V-Ray GPU and 58 per cent in Octane, and in August 2026 it came 19 per cent behind the RTX PRO 6000 Workstation Edition in Topaz Video
  • Puget found the 48 GB and 72 GB versions within its margin of error in Topaz Video 1.6.1 in August 2026 and less than 1 per cent apart in Starlight Mini in September 2026
  • Published language model figures are thin: Phoronix and Puget give their token rates only in charts, the text of NVIDIA’s post on the 72 GB card names no baseline card, model or precision, and we found no MLPerf® Inference entry in the results we could read
  • By our arithmetic, 1,344 GB/s caps one stream of Llama 3.3 70B in NVFP4 at about 33 tokens per second without speculative decoding, three quarters of the ceiling of the 1,792 GB/s RTX PRO 6000 workstation cards

Where the card sits

The RTX PRO 5000 Blackwell is a cut-down GB202, the chip of the RTX PRO 6000, with 14,080 CUDA cores against the 6000’s 24,064. NVIDIA announced it in March 2025; a 72 GB version appeared on NVIDIA’s site in October 2025 and became generally available in December 2025. Both versions share every speed figure in NVIDIA’s datasheet; they differ in memory, MIG instance size and vGPU support.

SPECIFICATIONRTX PRO 4500RTX PRO 5000RTX PRO 6000 MAX-QRTX PRO 6000, 600 W
CUDA cores10,49614,08024,06424,064
Memory32 GB GDDR748 or 72 GB GDDR796 GB GDDR796 GB GDDR7
Bandwidth896 GB/s1,344 GB/s1,792 GB/s1,792 GB/s
AI TOPS (FP4, sparse)1,6172,0643,5114,000
Board power200 W300 W300 W600 W
MIG2 × 16 GB, unconfirmed2 × 24 GB or 2 × 36 GBup to 4 × 24 GBup to 4 × 24 GB
vGPUno72 GB only, Red Hat KVM 9.6nono

NVIDIA datasheets and product pages (2025 to 2026), MIG user guide of 11 September 2026, vGPU support matrix of 22 September 2026. The RTX PRO 5000’s bus is 384-bit in NVIDIA’s October 2025 datasheet and 512-bit in the June 2026 edition; NVIDIA’s December 2025 blog on the 72 GB card quotes 2,142 TOPS. The MIG guide lists an RTX PRO 4500 without naming its edition; the workstation datasheet omits MIG. The 4500 and 600 W 6000 columns are the workstation cards; the Server Editions run vGPU, the 4500 (800 GB/s) from vGPU 20.0 and the 6000 (1,597 GB/s) from 19.0.

Against the 600 W RTX PRO 6000, the 5000 has three quarters of the bandwidth and about half the tensor peak; against the RTX PRO 4500, 1.5 times the bandwidth. Bandwidth sets the pace of text generation; rendering and prompt reading depend mainly on compute. Our family comparison and the RTX PRO 4500 article cover the neighbours.

What has been measured, and by whom

The independent record is short. Puget Systems tested the 48 GB card in its content creation and engineering roundups of 18 December 2025, and both memory sizes in two Topaz Video tests in August and September 2026. Phoronix included the 48 GB card in its Linux review of 21 May 2026 (Ubuntu 26.04 LTS, NVIDIA driver 595.58.03) and publishes its results as charts. We found no desktop review by StorageReview, ServeTheHome or Level1Techs.

Two traps. NVIDIA’s laptop range has an RTX PRO 5000 Blackwell too, with 10,496 CUDA cores, 24 GB and 896 GB/s at 95 to 175 W; StorageReview’s RTX PRO 5000 reviews are of such laptops and say nothing about the desktop card. And several websites publish tokens-per-second tables for the desktop card that leave out the engine version, the concurrency or the test system. We leave those out: they cannot be compared fairly.

Rendering, video and CAD: Puget’s numbers

Puget’s December 2025 tests ran on a Ryzen 9 9950X3D under Windows 11 with NVIDIA driver 573.92, its 2026 Topaz tests on a Threadripper PRO 9965WX with driver 596.72.

PUGET TESTRTX PRO 5000 RESULTCOMPARED WITH
Blender50% fasterRTX 5000 Ada
V-Ray GPU35% faster; 11% ahead of the RTX 6000 AdaRTX 5000 Ada
Octane58% fasterRTX 5000 Ada
Topaz Video AI25% faster; 8% ahead of the RTX 6000 AdaRTX 5000 Ada
Unreal Engine18% fasterRTX 5000 Ada
Topaz Video 1.6.148 GB card 19% slowerRTX PRO 6000 Workstation Edition
Topaz Video 1.6.172 GB card 7% slowerRTX PRO 6000 Max-Q
Starlight Miniunder 11 s per frame, both versionsunder 8 s on the RTX PRO 6000 Workstation Edition
Starlight Precise 2.5just over 6 s per frameunder 5 s on the RTX PRO 6000 Workstation Edition

Puget Systems: 2025 Professional GPU Content Creation Roundup, 18 December 2025 (first five rows, 48 GB card); Topaz Video 1.6.1 analysis, 24 August 2026 (overall score); Starlight test, 17 September 2026.

For a buyer, memory size does not change speed when the work fits: Puget found a “negligible difference in performance between the two variants, small enough to fall within our margin of error” in Topaz Video 1.6.1, and less than 1 per cent in Starlight Mini. The card also lands closer to the RTX PRO 6000 Workstation Edition than its core count suggests: in Topaz Video it is 19 per cent behind with 59 per cent of the cores and half the board power.

CAD gains least. Puget’s engineering roundup found that “even the graphics tests in Inventor are relatively insensitive to GPUs once a certain threshold is reached”, that Revit “is primarily CPU-dependent”, and that every AMD card it tested beat the fastest NVIDIA card by at least 18 per cent in the SOLIDWORKS drawing test.

Language models: the published record is thin

Phoronix ran llama.cpp on its RTX PRO cards, the 48 GB 5000 among them, measuring generation of 128 tokens and prompt processing of 512 and 2,048 tokens. The models were Q8_0 files of Llama-3.1-Tulu-3-8B, Mistral-7B-Instruct-v0.3, DeepSeek-R1-Distill-Llama-8B and gpt-oss-20b on llama.cpp’s GPU and Vulkan back ends, plus GLM-4.7-Flash in IQ4_XS on the GPU back end and granite-3.0-3b-a800m-instruct in Q8_0 on Vulkan. The values are in chart images, the concurrency is not stated, and the text says only that the card “was outpacing the prior-gen flagship of the RTX 6000 Ada Generation”. Puget Systems’ own MLPerf® Client v1.0 runs put the Blackwell 6000 cards and the 5000 at the top of the time-to-first-token chart, “barely beating out the 6000 Ada”, and the Blackwell token rate on average 50 per cent above Ada, the RTX PRO 2000 excepted; the text names neither the model nor the execution path, and these results are unverified, not reviewed by MLCommons Association.

Result not verified by MLCommons Association. The MLPerf name and logo are registered and unregistered trademarks of MLCommons Association in the United States and other countries. All rights reserved. Unauthorized use strictly prohibited. See www.mlcommons.org for more information.

MLCommons’ MLPerf Client page has no result tables, and we found no system with the card in the MLPerf Inference results we could read. NVIDIA’s blog for the 72 GB card claims twice the text generation and 3.5 times the image generation performance of “prior-generation NVIDIA hardware” in “industry-standard benchmarks for generative AI”; the text of the post names no baseline card, model or precision. NVIDIA’s TensorRT-LLM performance overview of 21 September 2026 covers the RTX PRO 6000 Server Edition, not the 5000. So we found no tokens-per-second figure for this card that states the engine and its version, the precision, the concurrency and the test system together. What remains is the ceiling below and a load test with your own model; the technical assessment in our engineering partner Vixen.UNO’s AI/ML Integration service covers model and GPU selection and a pilot plan with metrics.

What 1,344 GB/s allows

Generating a token for one user reads every active weight once, so plain single-stream decoding cannot beat the bandwidth divided by the bytes read per token. That ceiling is an upper bound, not a forecast: real engines stay below it, and the KV cache adds reads as the context grows. Speculative decoding, in which one read of the weights checks several drafted tokens, is the exception that can pass it.

MODEL AND FORMATWEIGHTSREAD PER TOKENCEILING, ONE STREAMFITS ON
gpt-oss-20b, MXFP412.8 GiBabout 3.7 GBabout 360 tok/s48 and 72 GB
Qwen3-32B, NVFP419.3 GiBabout 19.1 GBabout 70 tok/s48 and 72 GB
Llama 3.3 70B, NVFP440 GiB40.6 GBabout 33 tok/s72 GB; 48 GB with a short context
gpt-oss-120b, MXFP460.8 GiBabout 5.0 GBabout 270 tok/s72 GB only

Our arithmetic: 1,344 GB/s divided by the bytes read per token (linear layers and output layer; the embedding table is only looked up). Weights in GiB, from OpenAI’s and NVIDIA’s checkpoints; bytes per token in GB, as bandwidth is quoted in GB/s. Fits: 90 per cent of the card for weights and cache. The 1,792 GB/s of the RTX PRO 6000 workstation cards raises each ceiling by a third.

The mixture-of-experts models read only the experts each token uses, about 5 GB per token in all for gpt-oss-120b; but the fewer bytes a token reads, the more the fixed cost of each step weighs, and the further real engines fall below the line.

With several users, a dense model’s weights are read once per step for the whole batch, so aggregate throughput grows with concurrency until compute or cache space runs out. Cache is where the 72 GB card pulls ahead: beside a 4-bit 70B model it leaves about 25 GiB after a 10 per cent reserve, the 48 GB card about 3 GiB; our 48 GB or 72 GB comparison has the full table. Reading the prompt is compute-bound, and there the gap to the 600 W RTX PRO 6000 widens from a quarter to about a half: 2,064 AI TOPS against 4,000, or by our arithmetic about 516 TFLOPS of dense FP8 against 1,000. Long RAG and coding-agent prompts feel that most.

Power, cooling and several cards

The card has a total board power of 300 W and one 16-pin PCIe CEM5 power connector; NVIDIA’s quick start guide lists an adapter from two 8-pin PCIe cables in the box. It is a full-height, dual-slot board of 4.4 × 10.5 inches; NVIDIA calls the cooler active, and its board partner Leadtek specifies a blower, which exhausts through the slot bracket. We found no published power, temperature or noise figures for the desktop card in text; Phoronix logged GPU power but shows it only in charts.

Four cards draw up to 1.2 kW before the processor. Without NVLink they talk over PCIe 5.0 x16, which suits one model or one team per card better than one model split across cards, whose tensor-parallel all-reduces then cross PCIe. The AI servers we build to order include workstations and short-depth servers on Threadripper PRO or Xeon W platforms with RTX PRO Blackwell Workstation cards, with a rack, power and airflow check before we quote and burn-in testing after assembly.

MIG, vGPU and software

NVIDIA’s MIG guide of 11 September 2026 lists the 48 GB card with two isolated 1g.24gb instances or one 2g.48gb, each also as a +gfx variant for graphics; the 72 GB card is not in the guide yet and its datasheet gives two instances of 36 GB, so once MIG mode is enabled, read its profile names with nvidia-smi mig -lgip. MIG needs Linux, driver 575.51.03 or later, vBIOS 98.02.73.00.00 or later on the 48 GB card, and the display mode switched to compute with DisplayModeSelector 1.72 or later, 1.76 or later for the 72 GB card according to NVIDIA staff. The card then drives no display, and NVIDIA says the system must be qualified for that mode.

vGPU covers only the 72 GB card, from vGPU 20.2 of August 2026, and as of September 2026 NVIDIA’s support matrices list it only for Red Hat Enterprise Linux with KVM 9.6; for VMware vSphere they name, of the RTX PRO Blackwell cards, only the RTX PRO 6000 and 4500 Server Editions. The card has compute capability 12.0: NVIDIA’s datasheet lists CUDA 12.8, and PyTorch added Blackwell support in release 2.7 with CUDA 12.8 wheels, so check containers built on older CUDA releases first.

Where it fits, and where it does not

The measured record supports the RTX PRO 5000 as a rendering and video card, well ahead of the RTX 5000 Ada and within 19 per cent of the RTX PRO 6000 Workstation Edition in Topaz Video at half the power. For language models its case rests on memory and arithmetic: 48 GB for models up to 32B in FP8 or a 4-bit 70B model for one or two users, 72 GB for gpt-oss-120b or a 4-bit 70B model serving a small team.

It is the wrong card for a 70B model in FP8: beside its 68 GiB of weights, 72 GB leaves too little room for the runtime and for the cache that useful concurrency needs. The new 84 GB RTX PRO 5500 holds such a model with little cache, the 96 GB RTX PRO 6000 with more. It is also the wrong card for vGPU outside Red Hat KVM 9.6, for a CAD seat alone, and for a model too large for one card, which without NVLink must be split over PCIe. At the same 300 W, the RTX PRO 6000 Max-Q brings 96 GB, 1,792 GB/s and 3,511 AI TOPS, and the 72 GB card scored 7 per cent below it in Puget’s Topaz test; whether that step pays depends on the model and the budget.

What we supply

Eurokommerz supplies the RTX PRO 5000 Blackwell in both memory sizes EU-wide, with manufacturer warranty and EU contract and invoicing, as a single card or in a configured workstation; see the catalogue entry. For the platform on top, our engineering partner Vixen.UNO delivers AI/ML Integration: private LLMs on your premises with vLLM, Ollama or NVIDIA AI Enterprise, RAG assistants that respect each user’s access rights, and a pilot on one process with clear metrics, scaling only what has proved its value. The contract stays with Eurokommerz.

FAQ

How many tokens per second does the RTX PRO 5000 generate?
We found no published figure for the desktop card that states the engine and its version, the precision, the concurrency and the test system; Phoronix and Puget Systems give their language model token rates only in charts. By our arithmetic, its 1,344 GB/s caps one stream at about 33 tokens per second for Llama 3.3 70B in NVFP4 and about 70 for Qwen3-32B in NVFP4 without speculative decoding, and real engines stay below that ceiling.
Is the 72 GB RTX PRO 5000 faster than the 48 GB version?
No. Puget Systems found the two within its margin of error in Topaz Video 1.6.1 and less than 1 per cent apart in Topaz’s Starlight Mini model. The 72 GB card helps only when a model, its cache or a scene needs the memory.
How much faster is the RTX PRO 5000 than the RTX 5000 Ada?
In Puget Systems’ December 2025 roundup it was 58 per cent faster in Octane, 50 per cent in Blender, 35 per cent in V-Ray GPU, 25 per cent in Topaz Video AI and 18 per cent in Unreal Engine. In V-Ray GPU it also beat the RTX 6000 Ada by 11 per cent.
How does the RTX PRO 5000 compare with the RTX PRO 6000?
In Puget’s Topaz Video 1.6.1 test of August 2026 the 48 GB card was 19 per cent slower than the 600 W RTX PRO 6000 Workstation Edition and the 72 GB card 7 per cent slower than the 300 W Max-Q. On paper it has three quarters of the bandwidth of those two cards, about half the FP4 tensor peak of the Workstation Edition, and 48 or 72 GB against 96 GB.
Does the RTX PRO 5000 support MIG and vGPU?
MIG splits either version into two instances, 24 GB each on the 48 GB card and, according to its datasheet, 36 GB each on the 72 GB card, which NVIDIA’s MIG guide does not list yet. It needs Linux, driver 575.51.03 or later, a minimum vBIOS (98.02.73.00.00 on the 48 GB card) and the display mode switched to compute. vGPU covers only the 72 GB card, from vGPU 20.2 and only on Red Hat Enterprise Linux with KVM 9.6, according to NVIDIA’s support matrix of 22 September 2026.
Do laptop RTX PRO 5000 reviews apply to the desktop card?
No. NVIDIA’s laptop RTX PRO 5000 Blackwell has 10,496 CUDA cores, 24 GB and 896 GB/s at 95 to 175 W, while the desktop card has 14,080 cores, 48 or 72 GB and 1,344 GB/s at 300 W. Results from one say nothing about the other.

Tell us the models and precision you plan to run and how many people will use them at once, or your renderer and your largest scene. We will tell you whether the 48 GB or the 72 GB RTX PRO 5000 fits, or another card, and configure the workstation around it. We reply within one business day.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry – see our privacy policy.

request@eurokommerz.at  ·  +43 1 585 1405 50  ·  Jordangasse 7, 1010 Vienna