BLOG · COMPARISON ·

RTX PRO 5000 Blackwell vs RTX 5000 Ada: what 48 or 72 GB changes against 32 GB

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • The RTX PRO 5000 Blackwell has 48 or 72 GB of GDDR7 at 1,344 GB/s and 300 W; the RTX 5000 Ada has 32 GB of GDDR6 at 576 GB/s and 250 W, in the same 4.4 by 10.5 inch dual-slot format
  • FP32 (65 against 65.3 TFLOPS) and sparse FP8 Tensor throughput are close on paper; the Blackwell card adds FP4 Tensor Cores, MIG with up to two instances, PCIe 5.0 and DisplayPort 2.1b
  • Puget Systems measured the RTX PRO 5000 35 to 58 per cent ahead of the RTX 5000 Ada in V-Ray GPU, Blender and Octane in roundups of 18 December 2025, and found small generational gains for NVIDIA’s high-end cards in SOLIDWORKS
  • By our sizing rule, a 4-bit Qwen3-32B leaves room for about 7 concurrent 8K conversations on 32 GB, 22 on 48 GB and 43 on 72 GB, and only the 72 GB card serves a 4-bit 70B model to a team
  • The RTX 5000 Ada still fits where 32 GB is enough, where an Ada fleet shares one driver and image, and for vGPU on VMware vSphere, which NVIDIA supports for it and for neither RTX PRO 5000 version

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

RTX PRO 5000 vs RTX 5000 Ada: the short answer

The RTX PRO 5000 Blackwell has 48 or 72 GB of GDDR7 at 1,344 GB/s, against 32 GB of GDDR6 at 576 GB/s on the RTX 5000 Ada, at a board power of 300 W instead of 250 W in the same dual-slot format. It adds FP4 Tensor Cores and MIG with up to two isolated instances, and it moves to PCIe 5.0 and DisplayPort 2.1b. Its FP32 rate and its FP8 Tensor rate are close to the Ada card’s on paper, so the gains come from memory size, bandwidth and the newer architecture rather than from a larger compute peak.

For local language models the memory decides which models one card holds, and the bandwidth, 2.3 times higher, sets the upper limit on the speed of each stream. The RTX 5000 Ada keeps its place where 32 GB covers the work, where an Ada fleet runs one driver and container base, and for virtual workstations on VMware vSphere. Our overview of every RTX Ada and RTX PRO Blackwell tier places this pair in the whole range.

Specifications side by side

SPECIFICATIONRTX 5000 ADARTX PRO 5000
ArchitectureAda Lovelace, compute capability 8.9Blackwell, compute capability 12.0
CUDA cores12,80014,080
Tensor Cores400, fourth generation, FP8fifth generation, adds FP4
Memory32 GB GDDR6 with ECC48 or 72 GB GDDR7 with ECC
Memory interface256-bit384-bit
Bandwidth576 GB/s1,344 GB/s
FP3265.3 TFLOPS65 TFLOPS
RT Core performance151.0 TFLOPS196 TFLOPS
Tensor peak, sparse1,044.4 TFLOPS in FP82,064 TOPS in FP4
Board power250 W300 W, one 16-pin CEM5 connector
Host interfacePCIe 4.0 x16PCIe 5.0 x16
Display outputs4× DisplayPort 1.4a4× DisplayPort 2.1b
Video engines2 encoders, 2 decoders3 encoders, 3 decoders
Form factor4.4 × 10.5 inches, dual slot, active4.4 × 10.5 inches, dual slot, active
MIGnoup to 2 × 24 GB or 2 × 36 GB
vGPUfrom vGPU 16.1; vSphere and KVM72 GB only, from vGPU 20.2; KVM 9.6

NVIDIA RTX 5000 Ada datasheet (document 2898211, August 2023) and product page; NVIDIA RTX PRO 5000 Blackwell datasheet (document 5349550, July 2026), which gives the 384-bit interface and 1,344 GB/s for both versions, and product page; NVIDIA CUDA GPUs list, MIG user guide (11 September 2026), vGPU supported-GPU list (2 October 2026) and support matrices (29 September 2026); all read on 10 October 2026. Peak rates at the boost clock.

Both cards measure 4.4 by 10.5 inches and take two slots, so the Blackwell card goes into the same slot positions. Whether a given workstation takes it depends on 50 W more per card and the 16-pin CEM5 connector, which the maker confirms per model. In a PCIe 4.0 host the card runs at PCIe 4.0 speed, which matters little once a model sits in GPU memory.

FP8, FP4 and compute on paper

NVIDIA quotes the Ada card’s Tensor peak as 1,044.4 TFLOPS, “Effective FP8 teraFLOPS (TFLOPS) using sparsity”, and the Blackwell card’s as 2,064 AI TOPS, “Effective FP4 TOPS with sparsity”. The two headline figures use different precisions and cannot be compared directly. NVIDIA’s product page for the RTX PRO 6000 Server Edition lists FP8 at half its FP4 rate, 2 against 4 PFLOPS. On that ratio the RTX PRO 5000 reaches about 1,032 TFLOPS of sparse FP8 by our arithmetic, about 1 per cent below the Ada card’s 1,044.4, and FP32 is 65 against 65.3 TFLOPS.

The new precision on the Blackwell card is FP4. NVIDIA’s TensorRT-LLM support matrix (version 1.3.0rc29, updated 26 September 2026) marks NVFP4 and MXFP4 as supported on Blackwell sm120 and not on Ada Lovelace, which it lists with FP8 and with 4-bit AWQ weights. An Ada card runs 4-bit models as integer weights computed at 16-bit or FP8 precision.

Token generation reads every active weight once per token, so the bandwidth ratio of 2.3 carries over to the single-stream ceiling. For one stream of Qwen3-32B in Qwen’s 4-bit AWQ checkpoint of 19.3 GB, the RTX PRO 5000 therefore has a ceiling 2.3 times that of the RTX 5000 Ada. Both are arithmetic limits, not measurements, and measured rates stay below them.

Measured differences in rendering, video and CAD

Puget Systems compared the 48 GB RTX PRO 5000 with “its Ada counterpart” in its professional GPU content creation roundup of 18 December 2025, run with NVIDIA driver 573.92 on a Ryzen 9 9950X3D platform. It measured the Blackwell card 50 per cent ahead in Blender, 35 per cent in V-Ray GPU and 58 per cent in Octane. In Topaz Video AI the lead was 25 per cent, in After Effects 20 per cent and in Unreal Engine 18 per cent.

In SOLIDWORKS the step is smaller. Puget’s engineering roundup of the same date found “only minor generational improvements for the NVIDIA cards at the high end” and gave no figure for this pair. By that result, a SOLIDWORKS seat gains little from the swap unless its assemblies need more than 32 GB. Our RTX PRO 5000 benchmark review collects the published results for the Blackwell card, language model tests included.

Local LLMs and VLMs in 32, 48 and 72 GB

The sizing rule is the one in our guide to how much VRAM an LLM needs. We take 90 per cent of the memory the driver reports, subtract 3 GiB for activations and the runtime and then the weights, and the rest holds the KV cache. For the driver-reported memory we use 31.9, 47.8 and 71.7 GiB, in proportion to the 95.6 GiB of a 96 GB card, which leaves 25.7, 40.0 and 61.5 GiB for weights and cache.

MODEL AND FORMATWEIGHTSRTX 5000 ADA, 32 GBRTX PRO 5000, 48 GBRTX PRO 5000, 72 GB
Qwen3-14B, FP815.2 GiB163974
Qwen3-32B, 4-bit AWQ18.0 GiB72243
Qwen3-32B, FP832.0 GiBdoes not fit829
Llama 3.3 70B, NVFP439.8 GiBdoes not fitloads, under one17
gpt-oss-120b, MXFP460.8 GiBdoes not fitdoes not fitloads, under 1 GiB spare

Concurrent conversations at a full 8,192-token context with an FP8 KV cache: 0.625 GiB each for Qwen3-14B, 1 GiB for Qwen3-32B and 1.25 GiB for Llama 3.3 70B, from layers, KV heads and head size in each config.json. Weights are the checkpoint files on Hugging Face (Qwen’s FP8 and AWQ releases, NVIDIA’s Llama 3.3 70B FP4, OpenAI’s gpt-oss-120b), read on 10 October 2026. NVFP4 and MXFP4 run natively only on the Blackwell card. With ECC enabled, NVIDIA’s CUDA C++ Best Practices Guide states that the available DRAM on GDDR memory is reduced by 6.25 per cent, so the counts are lower. Our estimates, not measurements.

A 32 GB card serves a 14B model in FP8 to a team and a 32B model in 4-bit to a handful of parallel conversations. With 48 GB a 32B model runs in FP8, the precision of Qwen’s own FP8 release. With 72 GB a 4-bit 70B model serves a team from one card. The choice between the two Blackwell versions has its own comparison of the 48 GB and 72 GB RTX PRO 5000.

Vision-language models follow the same arithmetic, with the vision encoder added to the weights and each page image entering the context as tokens, so document pipelines need more cache room than chat. A GPU renderer or solver keeps the scene or mesh in card memory, or falls back to a slower out-of-core mode where it has one, so a dataset above 32 GB is a strong reason to move to the Blackwell card.

We supply both cards, the RTX PRO 5000 in its 48 and 72 GB versions. Tell us the models, their precision and the context length you plan, with the number of people using them at once, and we reply with the card that fits.

MIG, vGPU and the hypervisor

MIG exists only on the Blackwell card. Its datasheet gives up to two instances of 24 GB on the 48 GB version and up to two of 36 GB on the 72 GB version. NVIDIA’s MIG user guide of 11 September 2026 lists the 48 GB card with a maximum of two instances and does not list the RTX 5000 Ada. On a workstation card MIG has driver, vBIOS and display-mode requirements, set out in the comparison of the two versions linked above.

For vGPU the order is reversed. NVIDIA’s list of GPUs supported by vGPU (2 October 2026) gives the RTX 5000 Ada full support from vGPU release 16.1, and the support matrices of 29 September 2026 list it for VMware vSphere 8.0, VCF 9.0 and 9.1 and for eight Red Hat Enterprise Linux KVM releases from 8.10 to 10.2. Of the RTX PRO 5000, only the 72 GB Workstation Edition is on the list, from release 20.2, and the matrices name it for Red Hat Enterprise Linux KVM 9.6 alone. The vSphere matrix lists neither version.

A fleet that serves virtual workstations from vSphere hosts therefore keeps the RTX 5000 Ada, or moves to a Blackwell card that the vSphere matrix lists, the RTX PRO 6000 Server Edition or the RTX PRO 4500 Server Edition, both server cards.

Drivers, CUDA and an existing Ada fleet

NVIDIA lists the RTX 5000 Ada at compute capability 8.9 and the RTX PRO 5000 at 12.0. The Blackwell datasheet names CUDA 12.8 among its compute APIs. NVIDIA wrote on 17 July 2024 that for Blackwell “you must use the open-source GPU kernel modules”, and it recommends the same modules for Ada. Adding Blackwell cards to an Ada fleet therefore means a driver branch that covers both generations, on Linux the open kernel modules on the Blackwell hosts, and container images built for both compute capabilities.

NVIDIA’s Blackwell compatibility guide, updated 13 September 2026, states that application binaries which “do not include PTX (only include cubins), need to be rebuilt to run on the Blackwell GPUs”. A fleet that runs one image on every machine and needs nothing beyond 32 GB has a reason to add more Ada cards rather than start a second base. Our guide to mixing Ada and Blackwell GPUs covers hosts, node pools and the workloads that stay on Ada.

Which card for 4 to 8 workstations or a 4-card server

WORKLOADCARDWHY
SOLIDWORKS seatsRTX 5000 Adasmall generational gains in Puget’s SOLIDWORKS tests; same driver and image as the Ada fleet
VDI on vSphereRTX 5000 Adaon NVIDIA’s vSphere support matrix, where no RTX PRO 5000 version is
GPU rendering and videoRTX PRO 5000, 48 GB35 to 58 per cent ahead in V-Ray GPU, Blender and Octane in Puget’s tests
Datasets of 32 to 48 GBRTX PRO 5000, 48 GBfits a scene, mesh or model the 32 GB card cannot hold
LLMs of 14B to 32BRTX PRO 5000, 48 GB32B in FP8, or 4-bit 32B with about 22 conversations per card
4-bit 70B or gpt-oss-120bRTX PRO 5000, 72 GBabout 17 conversations on a 4-bit 70B; the only one that loads gpt-oss-120b
Two users or models per cardRTX PRO 5000MIG with two isolated instances

Our reading of the specifications, Puget Systems’ roundups of 18 December 2025, NVIDIA’s vGPU support matrices of 29 September 2026 and the sizing table above.

A 4-card server, in a chassis whose maker lists these actively cooled cards, shows the scale. Four RTX 5000 Ada give 128 GB at 1,000 W of board power, and four RTX PRO 5000 give 192 or 288 GB at 1,200 W. Serving Qwen3-32B in 4-bit as one copy per card, the sizing table gives about 28 concurrent 8K conversations on the Ada cards, 88 on the 48 GB cards and 172 on the 72 GB cards. With the 72 GB cards the same server can run a 4-bit 70B model instead, one copy per card, for about 68. Across eight workstations, the change from Ada to Blackwell adds 400 W of board power, 50 W per seat.

We supply either card and build GPU workstations and servers with the RTX PRO 5000 to order, with the rack, power and airflow checked before we quote. Send us the machine models or rack position and the largest model or scene.

What we supply

We supply the RTX 5000 Ada and the RTX PRO 5000 Blackwell in its 48 GB and 72 GB versions, with manufacturer warranty on one EU contract and invoice, as cards for your workstations and servers, and the RTX PRO 5000 also in GPU workstations and AI servers built to order. For models that outgrow 72 GB on one card, the 84 GB RTX PRO 5500 and the 96 GB RTX PRO 6000 come from the same range of professional NVIDIA GPUs. NVIDIA AI Enterprise and vGPU licences come on the same invoice as the hardware. We reply with a configuration and quote within one business day.

FAQ

RTX PRO 5000 vs RTX 5000 Ada: which is faster?
In Puget Systems’ roundup of 18 December 2025 the 48 GB RTX PRO 5000 was 50 per cent faster than the RTX 5000 Ada in Blender, 35 per cent in V-Ray GPU and 58 per cent in Octane. For language models its 1,344 GB/s against 576 GB/s raises the single-stream ceiling 2.3 times, while FP32 and sparse FP8 compute are close on paper.
Is the RTX PRO 5000 Blackwell the successor of the RTX 5000 Ada?
It is the Blackwell card in the same tier, and Puget Systems tests it against the RTX 5000 Ada as its Ada counterpart. It keeps the 4.4 by 10.5 inch dual-slot format and moves from 32 GB of GDDR6 at 250 W to 48 or 72 GB of GDDR7 at 300 W.
Is 32 GB on the RTX 5000 Ada enough for a local LLM?
For 14B models in FP8 and 32B models in 4-bit it is: by our sizing rule a 4-bit Qwen3-32B leaves room for about 7 concurrent conversations of 8K tokens with an FP8 cache. A 32B model in FP8 or any 70B model does not fit, and those need the 48 or 72 GB RTX PRO 5000.
Does the RTX 5000 Ada support FP4 or MIG?
No. NVIDIA’s TensorRT-LLM support matrix lists FP8 and 4-bit AWQ weights for Ada Lovelace but not NVFP4 or MXFP4, and NVIDIA’s MIG user guide does not list the card. The RTX PRO 5000 has FP4 Tensor Cores and splits into up to two MIG instances.
RTX 5000 Ada upgrade: can an RTX PRO 5000 go into the same workstation?
Both are 4.4 by 10.5 inch dual-slot cards, but the Blackwell card draws 300 W instead of 250 W through a 16-pin CEM5 connector, so check the power supply and the maker’s list for that model. It needs CUDA 12.8 or later, on Linux NVIDIA’s open kernel modules, and binaries that include PTX or are rebuilt for Blackwell.
Which RTX 5000 card supports vGPU on VMware vSphere?
The RTX 5000 Ada does. NVIDIA’s support matrices of 29 September 2026 list it for vSphere 8.0 and VCF 9.0 and 9.1. Of the RTX PRO 5000, only the 72 GB version supports vGPU, from release 20.2 and only on Red Hat Enterprise Linux KVM 9.6.

Send us the applications, the models with their precision and context length, the number of people using them at once and the workstation models or the server’s rack position. We reply within one business day with the card and memory size that fit, a configuration and a quote, with the rack, power and airflow checked before we quote a server.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna