NVIDIA RTX PRO 5000 Blackwell, 48 GB and 72 GB: real benchmarks and limits
- The RTX PRO 5000 uses the RTX PRO 6000’s GB202 chip and has 14,080 CUDA cores, 1,344 GB/s and 300 W; its 48 GB and 72 GB versions differ in memory, MIG instance size and vGPU support
- In December 2025 Puget Systems measured it 50 per cent ahead of the RTX 5000 Ada in Blender, 35 per cent in V-Ray GPU and 58 per cent in Octane, and in August 2026 it came 19 per cent behind the RTX PRO 6000 Workstation Edition in Topaz Video
- Puget found the 48 GB and 72 GB versions within its margin of error in Topaz Video 1.6.1 in August 2026 and less than 1 per cent apart in Starlight Mini in September 2026
- Published language model figures are thin: Phoronix and Puget give their token rates only in charts, the text of NVIDIA’s post on the 72 GB card names no baseline card, model or precision, and we found no MLPerf® Inference entry in the results we could read
- By our arithmetic, 1,344 GB/s caps one stream of Llama 3.3 70B in NVFP4 at about 33 tokens per second without speculative decoding, three quarters of the ceiling of the 1,792 GB/s RTX PRO 6000 workstation cards
Where the card sits
The RTX PRO 5000 Blackwell is a cut-down GB202, the chip of the RTX PRO 6000, with 14,080 CUDA cores against the 6000’s 24,064. NVIDIA announced it in March 2025; a 72 GB version appeared on NVIDIA’s site in October 2025 and became generally available in December 2025. Both versions share every speed figure in NVIDIA’s datasheet; they differ in memory, MIG instance size and vGPU support.
| SPECIFICATION | RTX PRO 4500 | RTX PRO 5000 | RTX PRO 6000 MAX-Q | RTX PRO 6000, 600 W |
|---|---|---|---|---|
| CUDA cores | 10,496 | 14,080 | 24,064 | 24,064 |
| Memory | 32 GB GDDR7 | 48 or 72 GB GDDR7 | 96 GB GDDR7 | 96 GB GDDR7 |
| Bandwidth | 896 GB/s | 1,344 GB/s | 1,792 GB/s | 1,792 GB/s |
| AI TOPS (FP4, sparse) | 1,617 | 2,064 | 3,511 | 4,000 |
| Board power | 200 W | 300 W | 300 W | 600 W |
| MIG | 2 × 16 GB, unconfirmed | 2 × 24 GB or 2 × 36 GB | up to 4 × 24 GB | up to 4 × 24 GB |
| vGPU | no | 72 GB only, Red Hat KVM 9.6 | no | no |
NVIDIA datasheets and product pages (2025 to 2026), MIG user guide of 11 September 2026, vGPU support matrix of 22 September 2026. The RTX PRO 5000’s bus is 384-bit in NVIDIA’s October 2025 datasheet and 512-bit in the June 2026 edition; NVIDIA’s December 2025 blog on the 72 GB card quotes 2,142 TOPS. The MIG guide lists an RTX PRO 4500 without naming its edition; the workstation datasheet omits MIG. The 4500 and 600 W 6000 columns are the workstation cards; the Server Editions run vGPU, the 4500 (800 GB/s) from vGPU 20.0 and the 6000 (1,597 GB/s) from 19.0.
Against the 600 W RTX PRO 6000, the 5000 has three quarters of the bandwidth and about half the tensor peak; against the RTX PRO 4500, 1.5 times the bandwidth. Bandwidth sets the pace of text generation; rendering and prompt reading depend mainly on compute. Our family comparison and the RTX PRO 4500 article cover the neighbours.
What has been measured, and by whom
The independent record is short. Puget Systems tested the 48 GB card in its content creation and engineering roundups of 18 December 2025, and both memory sizes in two Topaz Video tests in August and September 2026. Phoronix included the 48 GB card in its Linux review of 21 May 2026 (Ubuntu 26.04 LTS, NVIDIA driver 595.58.03) and publishes its results as charts. We found no desktop review by StorageReview, ServeTheHome or Level1Techs.
Two traps. NVIDIA’s laptop range has an RTX PRO 5000 Blackwell too, with 10,496 CUDA cores, 24 GB and 896 GB/s at 95 to 175 W; StorageReview’s RTX PRO 5000 reviews are of such laptops and say nothing about the desktop card. And several websites publish tokens-per-second tables for the desktop card that leave out the engine version, the concurrency or the test system. We leave those out: they cannot be compared fairly.
Rendering, video and CAD: Puget’s numbers
Puget’s December 2025 tests ran on a Ryzen 9 9950X3D under Windows 11 with NVIDIA driver 573.92, its 2026 Topaz tests on a Threadripper PRO 9965WX with driver 596.72.
| PUGET TEST | RTX PRO 5000 RESULT | COMPARED WITH |
|---|---|---|
| Blender | 50% faster | RTX 5000 Ada |
| V-Ray GPU | 35% faster; 11% ahead of the RTX 6000 Ada | RTX 5000 Ada |
| Octane | 58% faster | RTX 5000 Ada |
| Topaz Video AI | 25% faster; 8% ahead of the RTX 6000 Ada | RTX 5000 Ada |
| Unreal Engine | 18% faster | RTX 5000 Ada |
| Topaz Video 1.6.1 | 48 GB card 19% slower | RTX PRO 6000 Workstation Edition |
| Topaz Video 1.6.1 | 72 GB card 7% slower | RTX PRO 6000 Max-Q |
| Starlight Mini | under 11 s per frame, both versions | under 8 s on the RTX PRO 6000 Workstation Edition |
| Starlight Precise 2.5 | just over 6 s per frame | under 5 s on the RTX PRO 6000 Workstation Edition |
Puget Systems: 2025 Professional GPU Content Creation Roundup, 18 December 2025 (first five rows, 48 GB card); Topaz Video 1.6.1 analysis, 24 August 2026 (overall score); Starlight test, 17 September 2026.
For a buyer, memory size does not change speed when the work fits: Puget found a “negligible difference in performance between the two variants, small enough to fall within our margin of error” in Topaz Video 1.6.1, and less than 1 per cent in Starlight Mini. The card also lands closer to the RTX PRO 6000 Workstation Edition than its core count suggests: in Topaz Video it is 19 per cent behind with 59 per cent of the cores and half the board power.
CAD gains least. Puget’s engineering roundup found that “even the graphics tests in Inventor are relatively insensitive to GPUs once a certain threshold is reached”, that Revit “is primarily CPU-dependent”, and that every AMD card it tested beat the fastest NVIDIA card by at least 18 per cent in the SOLIDWORKS drawing test.
Language models: the published record is thin
Phoronix ran llama.cpp on its RTX PRO cards, the 48 GB 5000 among them, measuring generation of 128 tokens and prompt processing of 512 and 2,048 tokens. The models were Q8_0 files of Llama-3.1-Tulu-3-8B, Mistral-7B-Instruct-v0.3, DeepSeek-R1-Distill-Llama-8B and gpt-oss-20b on llama.cpp’s GPU and Vulkan back ends, plus GLM-4.7-Flash in IQ4_XS on the GPU back end and granite-3.0-3b-a800m-instruct in Q8_0 on Vulkan. The values are in chart images, the concurrency is not stated, and the text says only that the card “was outpacing the prior-gen flagship of the RTX 6000 Ada Generation”. Puget Systems’ own MLPerf® Client v1.0 runs put the Blackwell 6000 cards and the 5000 at the top of the time-to-first-token chart, “barely beating out the 6000 Ada”, and the Blackwell token rate on average 50 per cent above Ada, the RTX PRO 2000 excepted; the text names neither the model nor the execution path, and these results are unverified, not reviewed by MLCommons Association.
Result not verified by MLCommons Association. The MLPerf name and logo are registered and unregistered trademarks of MLCommons Association in the United States and other countries. All rights reserved. Unauthorized use strictly prohibited. See www.mlcommons.org for more information.
MLCommons’ MLPerf Client page has no result tables, and we found no system with the card in the MLPerf Inference results we could read. NVIDIA’s blog for the 72 GB card claims twice the text generation and 3.5 times the image generation performance of “prior-generation NVIDIA hardware” in “industry-standard benchmarks for generative AI”; the text of the post names no baseline card, model or precision. NVIDIA’s TensorRT-LLM performance overview of 21 September 2026 covers the RTX PRO 6000 Server Edition, not the 5000. So we found no tokens-per-second figure for this card that states the engine and its version, the precision, the concurrency and the test system together. What remains is the ceiling below and a load test with your own model; the technical assessment in our engineering partner Vixen.UNO’s AI/ML Integration service covers model and GPU selection and a pilot plan with metrics.
What 1,344 GB/s allows
Generating a token for one user reads every active weight once, so plain single-stream decoding cannot beat the bandwidth divided by the bytes read per token. That ceiling is an upper bound, not a forecast: real engines stay below it, and the KV cache adds reads as the context grows. Speculative decoding, in which one read of the weights checks several drafted tokens, is the exception that can pass it.
| MODEL AND FORMAT | WEIGHTS | READ PER TOKEN | CEILING, ONE STREAM | FITS ON |
|---|---|---|---|---|
| gpt-oss-20b, MXFP4 | 12.8 GiB | about 3.7 GB | about 360 tok/s | 48 and 72 GB |
| Qwen3-32B, NVFP4 | 19.3 GiB | about 19.1 GB | about 70 tok/s | 48 and 72 GB |
| Llama 3.3 70B, NVFP4 | 40 GiB | 40.6 GB | about 33 tok/s | 72 GB; 48 GB with a short context |
| gpt-oss-120b, MXFP4 | 60.8 GiB | about 5.0 GB | about 270 tok/s | 72 GB only |
Our arithmetic: 1,344 GB/s divided by the bytes read per token (linear layers and output layer; the embedding table is only looked up). Weights in GiB, from OpenAI’s and NVIDIA’s checkpoints; bytes per token in GB, as bandwidth is quoted in GB/s. Fits: 90 per cent of the card for weights and cache. The 1,792 GB/s of the RTX PRO 6000 workstation cards raises each ceiling by a third.
The mixture-of-experts models read only the experts each token uses, about 5 GB per token in all for gpt-oss-120b; but the fewer bytes a token reads, the more the fixed cost of each step weighs, and the further real engines fall below the line.
With several users, a dense model’s weights are read once per step for the whole batch, so aggregate throughput grows with concurrency until compute or cache space runs out. Cache is where the 72 GB card pulls ahead: beside a 4-bit 70B model it leaves about 25 GiB after a 10 per cent reserve, the 48 GB card about 3 GiB; our 48 GB or 72 GB comparison has the full table. Reading the prompt is compute-bound, and there the gap to the 600 W RTX PRO 6000 widens from a quarter to about a half: 2,064 AI TOPS against 4,000, or by our arithmetic about 516 TFLOPS of dense FP8 against 1,000. Long RAG and coding-agent prompts feel that most.
Power, cooling and several cards
The card has a total board power of 300 W and one 16-pin PCIe CEM5 power connector; NVIDIA’s quick start guide lists an adapter from two 8-pin PCIe cables in the box. It is a full-height, dual-slot board of 4.4 × 10.5 inches; NVIDIA calls the cooler active, and its board partner Leadtek specifies a blower, which exhausts through the slot bracket. We found no published power, temperature or noise figures for the desktop card in text; Phoronix logged GPU power but shows it only in charts.
Four cards draw up to 1.2 kW before the processor. Without NVLink they talk over PCIe 5.0 x16, which suits one model or one team per card better than one model split across cards, whose tensor-parallel all-reduces then cross PCIe. The AI servers we build to order include workstations and short-depth servers on Threadripper PRO or Xeon W platforms with RTX PRO Blackwell Workstation cards, with a rack, power and airflow check before we quote and burn-in testing after assembly.
MIG, vGPU and software
NVIDIA’s MIG guide of 11 September 2026 lists the 48 GB card with two isolated 1g.24gb instances or one 2g.48gb, each also as a +gfx variant for graphics; the 72 GB card is not in the guide yet and its datasheet gives two instances of 36 GB, so once MIG mode is enabled, read its profile names with nvidia-smi mig -lgip. MIG needs Linux, driver 575.51.03 or later, vBIOS 98.02.73.00.00 or later on the 48 GB card, and the display mode switched to compute with DisplayModeSelector 1.72 or later, 1.76 or later for the 72 GB card according to NVIDIA staff. The card then drives no display, and NVIDIA says the system must be qualified for that mode.
vGPU covers only the 72 GB card, from vGPU 20.2 of August 2026, and as of September 2026 NVIDIA’s support matrices list it only for Red Hat Enterprise Linux with KVM 9.6; for VMware vSphere they name, of the RTX PRO Blackwell cards, only the RTX PRO 6000 and 4500 Server Editions. The card has compute capability 12.0: NVIDIA’s datasheet lists CUDA 12.8, and PyTorch added Blackwell support in release 2.7 with CUDA 12.8 wheels, so check containers built on older CUDA releases first.
Where it fits, and where it does not
The measured record supports the RTX PRO 5000 as a rendering and video card, well ahead of the RTX 5000 Ada and within 19 per cent of the RTX PRO 6000 Workstation Edition in Topaz Video at half the power. For language models its case rests on memory and arithmetic: 48 GB for models up to 32B in FP8 or a 4-bit 70B model for one or two users, 72 GB for gpt-oss-120b or a 4-bit 70B model serving a small team.
It is the wrong card for a 70B model in FP8: beside its 68 GiB of weights, 72 GB leaves too little room for the runtime and for the cache that useful concurrency needs. The new 84 GB RTX PRO 5500 holds such a model with little cache, the 96 GB RTX PRO 6000 with more. It is also the wrong card for vGPU outside Red Hat KVM 9.6, for a CAD seat alone, and for a model too large for one card, which without NVLink must be split over PCIe. At the same 300 W, the RTX PRO 6000 Max-Q brings 96 GB, 1,792 GB/s and 3,511 AI TOPS, and the 72 GB card scored 7 per cent below it in Puget’s Topaz test; whether that step pays depends on the model and the budget.
What we supply
Eurokommerz supplies the RTX PRO 5000 Blackwell in both memory sizes EU-wide, with manufacturer warranty and EU contract and invoicing, as a single card or in a configured workstation; see the catalogue entry. For the platform on top, our engineering partner Vixen.UNO delivers AI/ML Integration: private LLMs on your premises with vLLM, Ollama or NVIDIA AI Enterprise, RAG assistants that respect each user’s access rights, and a pilot on one process with clear metrics, scaling only what has proved its value. The contract stays with Eurokommerz.
FAQ
How many tokens per second does the RTX PRO 5000 generate?
Is the 72 GB RTX PRO 5000 faster than the 48 GB version?
How much faster is the RTX PRO 5000 than the RTX 5000 Ada?
How does the RTX PRO 5000 compare with the RTX PRO 6000?
Does the RTX PRO 5000 support MIG and vGPU?
Do laptop RTX PRO 5000 reviews apply to the desktop card?
Tell us the models and precision you plan to run and how many people will use them at once, or your renderer and your largest scene. We will tell you whether the 48 GB or the 72 GB RTX PRO 5000 fits, or another card, and configure the workstation around it. We reply within one business day.
Talk to an expertWe reply within one business day