RTX PRO 5500 Blackwell: where the new 84 GB card fits between the RTX PRO 5000 and the 6000
- NVIDIA put the RTX PRO 5500 Blackwell Workstation Edition on its website in September 2026 without a press release, with preliminary specifications and the status “Coming Soon”
- NVIDIA publishes 84 GB of GDDR7 with ECC, 1,398 GB/s, up to 600 W, PCIe 5.0 x16, three encoders and three decoders, and MIG in two instances of up to 42 GB
- Board partners list 21,760 CUDA cores, the count of the GeForce RTX 5090, one 16-pin connector and a dual-slot board; NVIDIA itself has published no core count, TOPS figure or bus width
- Bandwidth is only 4 per cent above the RTX PRO 5000, so on a model that fits both cards token rates should be close; the gain is 84 GB of capacity and, going by the partners’ core count, 55 per cent more cores
- On paper the RTX PRO 6000 Max-Q offers more memory and 28 per cent more bandwidth at a 300 W board power, and the 5500 is not on NVIDIA’s vGPU list
A quiet launch
NVIDIA did not announce the RTX PRO 5500 Blackwell at an event or in a press release. The product page went up quietly in early September 2026, and the first press reports followed on 14 September. It is marked “Preliminary product specifications, subject to change”, and where a buying link would be it shows “Coming Soon” with a “Notify Me” sign-up. NVIDIA describes the card as “designed for rack-mounted workstation deployment with a choice of air- or liquid-cooled thermal solutions”. When we compared the RTX PRO family on 14 September the card was a footnote; this is what NVIDIA and its board partners have published about it.
What NVIDIA publishes, and what partners add
| SPECIFICATION | NVIDIA PRODUCT PAGE | BOARD PARTNER LISTINGS |
|---|---|---|
| Memory | 84 GB GDDR7 with ECC | same |
| Bandwidth | 1,398 GB/s | same |
| Maximum power | up to 600 W | same |
| Interface | PCIe 5.0 x16 | same |
| Display outputs | up to 4× DisplayPort 2.1b | same |
| Video engines | 3 NVENC, 3 NVDEC | 3 NVENC, 3 NVDEC, 1 JPEG |
| MIG | up to 2 instances of up to 42 GB | 2 × 42 GB or 1 × 84 GB |
| Cooling | active, air-cooled; liquid-cooled version in the “RXM form factor” | air- or liquid-cooled |
| CUDA cores | not published | 21,760 |
| Memory bus | not published | 448-bit or 416-bit: the two listings disagree |
| Power connector | not published | one 16-pin |
| Board | not published | air-cooled version: 4.4 × 11.1 inches, full height, full length, dual slot |
| AI TOPS, FP32 TFLOPS | not published | not published |
| vGPU | not on NVIDIA’s vGPU list | not listed |
NVIDIA RTX PRO 5500 product page (September 2026); product listings of two NVIDIA board partners; NVIDIA vGPU supported-GPU list and MIG user guide, checked September 2026.
Three gaps matter. The core count comes from partners, not from NVIDIA, although 21,760 is the same figure the GeForce RTX 5090 carries. The bus width is open: 1,398 GB/s works out at about 25 Gbps on a 448-bit bus or about 27 Gbps on a 416-bit one, and only the 448-bit bus reaches 84 GB with identical 3 GB memory chips, but NVIDIA has not said which it is. And NVIDIA does not define “RXM” on the product page or in any NVIDIA document we found; press reports read it as a rack-module format for the liquid-cooled version, which NVIDIA has not confirmed. The card is not yet in NVIDIA’s MIG user guide either, so the MIG profile names are still to come.
Where it sits in the family
| SPECIFICATION | RTX PRO 5000 | RTX PRO 5500 | RTX PRO 6000 MAX-Q | RTX PRO 6000 WORKSTATION |
|---|---|---|---|---|
| Memory | 48 or 72 GB | 84 GB | 96 GB | 96 GB |
| Bandwidth | 1,344 GB/s | 1,398 GB/s | 1,792 GB/s | 1,792 GB/s |
| CUDA cores | 14,080 | 21,760 (partner data) | 24,064 | 24,064 |
| Board power | 300 W | up to 600 W | 300 W | 600 W |
| MIG | 2 × 24 or 2 × 36 GB | 2 × 42 GB | 4 × 24 GB | 4 × 24 GB |
| vGPU | 72 GB version, on Red Hat KVM 9.6 only | no | no | no |
NVIDIA datasheets (April 2025 to June 2026), product pages, MIG user guide and vGPU documentation, checked September 2026; the RTX PRO 5500 core count is from board-partner listings. The RTX PRO 5000 details are in our 48 or 72 GB comparison.
Against the RTX PRO 5000 the 5500 has 16.7 per cent more memory than the 72 GB version, twice the power ceiling, 4.0 per cent more bandwidth and, on the partners’ core count, 54.5 per cent more cores. Against the RTX PRO 6000 Workstation Edition it has 87.5 per cent of the memory, 78.0 per cent of the bandwidth and, on the same partner figure, 90.4 per cent of the cores, in the same 600 W envelope.
What that means for language models
Generating a token reads the active weights once, so on a model that fits both cards the 5500 should generate at almost exactly the rate of a 5000, since its bandwidth is only 4 per cent higher, and the RTX PRO 6000 should be about 28 per cent faster than the 5500 and a third faster than the 5000. Prompt processing, batch throughput, image generation and fine-tuning depend on compute, and there the partners’ core count, 90 per cent of a 6000’s, suggests the 5500 lands much closer to the 6000 than to the 5000, although NVIDIA has published no clocks or TOPS to confirm it. That is our reading of the specifications, not a measurement: no independent test of the card exists yet.
Capacity is where 84 GB earns its place. Using the method of our 70B sizing example, Llama 3.3 70B is 68 GiB in FP8 and 40 GiB in NVFP4, one 8,192-token conversation needs 1.25 GiB of FP8 cache, and about a tenth of the card goes to the runtime.
| LLAMA 3.3 70B | RTX PRO 5000, 72 GB | RTX PRO 5500, 84 GB | RTX PRO 6000, 96 GB |
|---|---|---|---|
| FP8 weights, 68 GiB | does not fit in practice | about 7.6 GiB of cache: six 8k conversations | about 18 GiB: fourteen 8k conversations |
| NVFP4 weights, 40 GiB | about 25 GiB: nineteen 8k conversations | about 36 GiB: twenty-eight | about 46 GiB: thirty-seven |
Our arithmetic: card memory less 10 per cent for the runtime, less the weights; FP8 KV cache at 160 KiB per token; an FP16 cache halves the conversation counts. Production servers keep more in reserve.
The FP8 row is the one that separates the 5500 from the 5000: a 70B model at 8-bit precision fits, for a handful of users. The same logic applies to gpt-oss-120b at 60.8 GiB, which leaves about 4 GiB on a 72 GB card and about 15 GiB on the 5500. In MIG mode each 42 GB half takes a 32B model in FP8, or gpt-oss-20b with long contexts, but not a 70B model in 4-bit with any useful context.
Rendering, video and the rack
The 5500 carries three NVENC and three NVDEC engines, like the 5000, where the 6000 has four of each, and it encodes and decodes H.264 and HEVC in 4:2:2 like the other RTX PRO Blackwell cards. NVIDIA’s page repeats the family claims, Tensor cores “up to 3x” and RT cores “up to 2x” the previous generation, without test conditions; rendering benchmarks will follow once independent testers have cards.
The positioning is the unusual part. NVIDIA places the card in rack-mounted workstations, where IT teams “centralize GPUs into managed racks, share them across users”: in practice the remote-workstation model, with the GPU in a rack and the user at a desk elsewhere. That puts it next to the RTX PRO 6000 Server Edition, which is passive, runs at 1,597 GB/s and supports vGPU. The 5500 is actively cooled, has display outputs, supports MIG and is not on NVIDIA’s vGPU list. At up to 600 W, and with the single 16-pin connector the partner listings show, it needs the same cabling and power supply headroom as a 6000 Workstation Edition.
Who should wait for it, and who should not
If the model fits in 72 GB, the RTX PRO 5000 72 GB gives almost the same token rate at a 300 W board power, and NVIDIA has published its full datasheet, while the 5500’s specifications are still marked preliminary.
If you need 96 GB or the most bandwidth, buy an RTX PRO 6000. For several cards in one machine, the Max-Q is the stronger choice on paper: more memory, 28 per cent more bandwidth and half the power ceiling of the 5500. NVIDIA has published neither clocks nor TOPS for the 5500, so how much extra compute its 600 W buys is the open question its datasheet will answer.
If you need 84 GB on one card, for a 70B model in FP8 or large render scenes, the 5500 is the card to evaluate once NVIDIA publishes the full specifications, including the compute figures that will show how close it comes to a 6000.
If you need vGPU, look at the RTX PRO 6000 Server Edition, or at the 5000 72 GB on Red Hat KVM 9.6.
What we supply
Eurokommerz supplies the RTX PRO Blackwell family EU-wide with manufacturer warranty and can include the RTX PRO 5500 in the same offer as the RTX PRO 5000 and 6000, with availability confirmed in writing, so the three can be compared side by side. Send us the workload, and we will tell you which of them it needs.
FAQ
When will the RTX PRO 5500 be available?
How many CUDA cores does the RTX PRO 5500 have?
Is the RTX PRO 5500 faster than the RTX PRO 5000 for language models?
Does the RTX PRO 5500 support MIG and vGPU?
What is the RXM form factor?
RTX PRO 5500 or RTX PRO 6000 Max-Q?
Send us the model or the scene, the number of users and where the machine will live. We will compare the RTX PRO 5000, 5500 and 6000 for that workload on one offer. We reply within one business day.
Talk to an expertWe reply within one business day