RTX PRO 4500 Server vs RTX PRO 6000 Server: eight smaller cards or four larger ones per server
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- Eight RTX PRO 4500 Server Edition cards give a server 256 GB of GPU memory in eight 32 GB pools at 1,320 W of board power; four RTX PRO 6000 Server Edition cards give 384 GB in four 96 GB pools at up to 2,400 W, and neither card has NVLink
- Summed memory bandwidth is almost the same, 6,400 GB/s against 6,388 GB/s, and both servers reach 16 MIG instances, of 16 GB on the smaller card and 24 GB on the larger one
- Models up to about 20 GiB of weights run on one card of either kind, while Llama 3.3 70B in FP8 fits one RTX PRO 6000 but needs four RTX PRO 4500 split over PCIe
- Eight smaller cards suit many small models, embeddings and video analytics, with 24 video decoders against 16; four larger cards suit one or two 70B-class models with one copy per card
- Where NVIDIA AI Enterprise is used, it is licensed per GPU, so eight cards need eight licences and four need four; neither card includes a subscription, and vWS and vPC licences are counted per concurrent user
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
RTX PRO 4500 Server vs RTX PRO 6000 Server: the short answer
Eight RTX PRO 4500 Blackwell Server Edition cards give a server 256 GB of GPU memory in eight pools of 32 GB. Four RTX PRO 6000 Blackwell Server Edition cards give 384 GB in four pools of 96 GB. The choice is set by the largest model that has to fit on one card, the number of separate workloads, the MIG slice sizes, the power at the rack position and the chassis the server maker lists for each card.
Many small models, embeddings and video streams favour eight smaller cards; one or two 70B-class models favour four larger ones. Lenovo’s product guides list “NVLink support” as “No” for both cards, so a model that does not fit one card is split over PCIe in either server. Card-level detail is in our guide to the RTX PRO 4500 Blackwell Server Edition.
The two configurations side by side
| ITEM | PER RTX PRO 4500 SE | 8 × RTX PRO 4500 SE | PER RTX PRO 6000 SE | 4 × RTX PRO 6000 SE |
|---|---|---|---|---|
| GPU memory | 32 GB GDDR7 | 256 GB in 8 pools | 96 GB GDDR7 | 384 GB in 4 pools |
| Memory bandwidth | 800 GB/s | 6,400 GB/s summed | 1,597 GB/s | 6,388 GB/s summed |
| CUDA cores | 10,496 | 83,968 | 24,064 | 96,256 |
| FP4 Tensor Core | 1.6 PFLOPS | 12.8 PFLOPS | 4 PFLOPS | 16 PFLOPS |
| FP8 Tensor Core | 811 TFLOPS | 6,488 TFLOPS | 2 PFLOPS | 8 PFLOPS |
| NVENC and NVDEC | 3 and 3 | 24 and 24 | 4 and 4 | 16 and 16 |
| MIG instances, maximum | 2 × 16 GB | 16 × 16 GB | 4 × 24 GB | 16 × 24 GB |
| Board power | 165 W | 1,320 W | 300 to 600 W | 1,200 to 2,400 W |
| Slot width, cooling | single slot, passive | 8 single-wide slots | dual slot, passive | 4 double-wide slots |
| AI Enterprise licences | 1 per GPU | 8 | 1 per GPU | 4 |
SE: Server Edition. NVIDIA product pages of both cards; NVIDIA MIG user guide (11 September 2026); NVIDIA product brief SP-12355-001_v02 for the 300 to 600 W range; NVIDIA AI Enterprise licensing guide (2 September 2026); all read on 10 October 2026. Server totals are our arithmetic.
The four larger cards have 50 per cent more memory and, by NVIDIA’s figures, about a quarter more FP4 and FP8 throughput; the eight smaller cards have 24 decoders against 16 and draw 1,320 W against up to 2,400 W.
Per card, a model that fits both has twice the memory bandwidth on the larger one. We found no NVIDIA tokens-per-second figure for the RTX PRO 4500 Server Edition; its product page claims “over 5x the performance of the previous-generation L4 GPU” for small and medium-sized models, without test conditions.
Largest model per card: what fits in 32 GB and in 96 GB
By our sizing rule, 0.9 × the card’s memory, less 3 GiB for activations and the runtime, holds the weights and the cache. That gives about 25.7 GiB on the 32 GB card with ECC off and 83.0 GiB on the RTX PRO 6000, whose driver reports 95.6 GiB. NVIDIA’s CUDA C++ Best Practices Guide states that on GDDR memory with ECC enabled “the available DRAM is reduced by 6.25%”, so with ECC on the 32 GB card keeps about 23.9 GiB, and nvidia-smi -q shows the mode. A copy split over several cards pools its cache, since vLLM’s blog of 7 August 2026 states that for grouped-query attention “TP splits the KV cache by those heads first”, and these models have 8 KV heads.
| MODEL, PRECISION | WEIGHTS | ON RTX PRO 4500 SE | ON RTX PRO 6000 SE | COPIES PER SERVER |
|---|---|---|---|---|
| gpt-oss-20b, MXFP4 | 12.8 GiB | 1 card, more than 130 | 1 card, more than 700 | 8 / 4 |
| Qwen3-14B, FP8 | 15.2 GiB | 1 card, about 16 | 1 card, about 108 | 8 / 4 |
| Qwen3-32B, NVFP4 | 19.3 GiB | 1 card, about 6 | 1 card, about 63 | 8 / 4 |
| Qwen3-32B, FP8 | 32.0 GiB | 2 cards, about 19 | 1 card, about 51 | 4 / 4 |
| Llama 3.3 70B, NVFP4 | 39.8 GiB | 2 cards, about 9 | 1 card, about 34 | 4 / 4 |
| Llama 3.3 70B, FP8 | 67.7 GiB | 4 cards, about 28 | 1 card, about 12 | 2 / 4 |
Our estimates, not measurements: concurrent conversations of 8,192 tokens at full length per copy, FP8 KV cache, 25.7 GiB per RTX PRO 4500 (ECC off) and 83.0 GiB per RTX PRO 6000; copies for eight RTX PRO 4500 / four RTX PRO 6000. Weights are the OpenAI, Qwen and NVIDIA checkpoints on Hugging Face as in our sizing guides; KV heads and layers from each config.json.
From Qwen3-32B in FP8 upwards, the smaller card needs two or four cards per copy, and tensor parallelism exchanges partial results between them at every layer, here over PCIe. Four RTX PRO 6000 keep one Llama 3.3 70B copy on each card with no traffic between cards. Eight RTX PRO 4500 hold two copies of four cards each and, by memory, more conversations in total, about 56 against 48, because the weights are stored twice instead of four times. A model above 96 GB needs several RTX PRO 6000 as well; our allocation guide for eight RTX PRO 6000 plans that server card by card.
Many small models and embeddings for 500 to 2,000 users
With the example values of our company-size guide, 500 staff produce about 20 requests in flight at the busiest moment and 2,000 staff about 80; the same values give about 40 for 1,000 staff. A 14B-class model on one RTX PRO 4500 holds about 13 full-length conversations of 8K with ECC on, so copies on four cards hold about 52, enough for 1,000 staff, and seven cards reach the 80 of 2,000 staff (three and five cards with ECC off). On one RTX PRO 6000 the same model has room for about 108, more than the largest example peak.
That surplus is the case for smaller cards when a server carries many models of modest size: an assistant, a coding model, embedding and reranking for RAG, speech-to-text and models per department, each with its own users and update cycle. Eight cards give eight places for an engine with its own memory and restart; with four, the rest share a card through MIG or several engine instances, as our guide to serving several models on one GPU server explains. A failed card takes one eighth of the GPU memory out of service on one server and one quarter on the other.
We build AI servers with 2 to 8 GPUs per node, sized by model size and concurrent users. Send us the models you plan and the users of each through the form below, and we reply with a configuration and quote within one business day.
MIG on both cards: two 16 GB or four 24 GB instances
NVIDIA’s MIG user guide, updated on 11 September 2026, lists two sizes for the RTX PRO 4500 Blackwell: 1g.16gb, up to two per card, each with half the memory and half the SMs, and 2g.32gb for the whole card. For the RTX PRO 6000 it lists 1g.24gb up to four times, 2g.48gb up to two times and 4g.96gb once. Both servers therefore reach at most 16 instances, of 16 GB on eight smaller cards and of 24 GB on four larger ones.
By the rule above, a 16 GB half holds NVIDIA’s 8.8 GiB FP8 checkpoint of Qwen3-8B with cache for about four conversations of 8K (two with ECC on), and Qwen3-14B in FP8 at 15.2 GiB does not fit with any cache. A 1g.24gb slice of the RTX PRO 6000 takes the same 14B model with cache for about five conversations. A 1g.16gb instance has one video decoder and one encoder, as does a 1g.24gb instance. Two plain halves leave the third decoder of an RTX PRO 4500 unused, so eight cards split into 16 halves offer 16 decoders, not 24; the 1g.16gb+me.all profile gives one half all three.
Video analytics and VDI on eight or four cards
For video analytics the decoders count. NVIDIA lists three NVENC and three NVDEC engines per RTX PRO 4500 Server Edition and four of each per RTX PRO 6000 Server Edition, so eight smaller cards carry 24 decoders and four larger cards 16. Streams per decoder depend on codec, resolution and frame rate, and are measured with your cameras. If the detection models fit 32 GB, the eight-card server gives more decoders at 1,320 W of board power against up to 2,400 W.
For VDI, both cards are on NVIDIA’s list of GPUs supported by vGPU software, updated on 2 October 2026: the RTX PRO 4500 Server Edition from release 20.0 and the RTX PRO 6000 Server Edition from release 19.0. Each vGPU profile has a fixed frame buffer size, so 32 GB and 96 GB cards divide into different numbers of desktops of a given size. vWS, vPC and vApps are licensed per concurrent user, so the number of cards does not change that count.
Power, slots and chassis: the Lenovo SR650a V4 as an example
NVIDIA specifies the RTX PRO 4500 Server Edition as a passive, single-slot, full-height, full-length card at 165 W with a PCIe 5.0 x16 link. The air-cooled RTX PRO 6000 Server Edition is a dual-slot card, and NVIDIA’s product brief SP-12355-001_v02 gives it a 600 W mode and a 450 W mode, each with a 300 W minimum. Four cards draw 2,400 W at 600 W or 1,800 W at 450 W. Capped at 300 W they draw 1,200 W, less than eight RTX PRO 4500, with lower throughput per card. NVIDIA’s product page does not state the power setting behind its FP4 and FP8 figures.
Lenovo’s 2U ThinkSystem SR650a V4 is listed for both cards. Its product guide, updated on 5 October 2026, lists “Support for up to 4x double-wide GPUs or 8x single-wide GPUs, installed in the front slots.” Lenovo’s guide for the RTX PRO 6000 Server Edition states that the card “can be configured power-capped to 450W to allow 4x GPUs to be installed in the SR650a V4”; the 600 W version is limited to two there. Lenovo’s guide for the RTX PRO 4500 Server Edition, updated on 28 July 2026, records that “The SR650a V4 now supports the GPU”. We could not read that guide’s maximum per server, so eight cards have to be confirmed in the maker’s configurator.
Either GPU total alone stays below the about 3.7 kW of a 16 A single-phase feed at 230 V; processors, memory and fans come on top, so plan the feed for the whole server. Whether one server or two serves the same cards better is covered in our comparison of one 8-GPU server or two 4-GPU servers.
NVIDIA AI Enterprise and vGPU licences per GPU
NVIDIA’s licensing guide, updated on 2 September 2026, states that NVIDIA AI Enterprise “is licensed on a per-GPU basis” and that a licence “is required for every GPU installed on the server or workstation” that hosts its software. Where it is used, eight cards therefore need eight licences and four cards four. The guide names the H100 PCIe and NVL, the A800 40GB Active and the H200 NVL as including a subscription, and neither RTX PRO Server Edition card. The guide does not mention MIG, so we count one licence per physical card however it is partitioned.
CUDA, vLLM and Triton run without the licence, as our guide to NVIDIA AI Enterprise licensing explains, so a bare-metal vLLM server needs none in either configuration.
Which configuration for which workload
| WORKLOAD | CONFIGURATION | WHY |
|---|---|---|
| Many models up to 20 GiB | 8 × RTX PRO 4500 SE | eight cards to place engines; a failure costs 32 GB |
| Embeddings and rerankers | either, with MIG | 16 instances on both servers |
| Department 14B models | 4 × RTX PRO 6000 SE with MIG | a 24 GB slice holds a 14B model in FP8 |
| Video analytics | 8 × RTX PRO 4500 SE | 24 decoders at 1,320 W of board power |
| VDI with vWS or vPC | either | both on NVIDIA’s vGPU list; profile sizes differ |
| One or two 70B-class models | 4 × RTX PRO 6000 SE | one copy per card, no traffic between cards |
| Power-limited rack position | either, by power cap | 1,320 W, or four RTX PRO 6000 capped at 300 W for 1,200 W |
Our reading of the sections above; sources as in the first two tables.
Describe your models, camera streams or desktops and the rack position in the form below, and we size both configurations side by side.
What we supply
We build AI servers to order with eight RTX PRO 4500 Server Edition or four RTX PRO 6000 Server Edition cards, assembled and burn-in tested, with manufacturer warranty on every component and delivery anywhere in the EU, on one EU contract and invoice. We check the rack, power and airflow before we quote, and return a configuration and quote within one business day. We also supply both cards on their own through our professional GPU range, alongside the L40S, L4 and H200 NVL. NVIDIA AI Enterprise and vGPU licences come on the same invoice where the setup needs them.
FAQ
RTX PRO 4500 Server Edition vs RTX PRO 6000 Server Edition: what is the difference?
Many small GPUs vs fewer large GPUs: which is better for inference?
What can a server with 8x RTX PRO 4500 Server Edition run?
How do I plan GPU density for an inference server?
How many NVIDIA AI Enterprise licences does a server with eight GPUs need?
Which GPU server suits many small models?
Send us the models you plan with their precision, the users or camera streams per workload, the desktops for vGPU and the power feed at the rack position. We reply within one business day with a configuration and quote for eight RTX PRO 4500 or four RTX PRO 6000 Server Edition cards, with the rack, power and airflow checked before we quote.
Talk to an expertWe reply within one business day