Four RTX PRO 6000 Max-Q in one workstation: 384 GB under a desk, and what it takes
- NVIDIA built the Max-Q edition for dense workstations: the full 96 GB and 1,792 GB/s at 300 W, “up to four GPUs in a single system”, 384 GB in total
- Towers with four Max-Q cards in their official specifications include the HP Z8 Fury G6i, the Lenovo ThinkStation PX and the Dell Precision 7960; Dell’s Pro Precision 9 T6 takes five. In these makers’ specifications the 600 W Workstation Edition stops at one or two per tower
- Four cards at x16 need 64 PCIe 5.0 lanes: Threadripper PRO 9000 WX and the larger Xeon 600 models offer 128, Xeon W-3500 offers 112
- Four cards and a 350 W processor come to about 1.55 kW before memory, drives and fans; plan for 1.8 to 2 kW at the wall, all of it heat in the room
- There is no NVLink, so large models are split over PCIe; MIG turns the four cards into up to sixteen isolated 24 GB instances for a team
Why the Max-Q exists
The RTX PRO 6000 Blackwell comes in three editions with the same 24,064 CUDA cores and 96 GB of GDDR7, as our edition comparison explains. The Workstation Edition runs at 600 W with a double flow-through cooler, the Server Edition is passive for rack chassis, and the Max-Q keeps the full memory and the full 1,792 GB/s at 300 W. NVIDIA’s product page states the purpose plainly: “it enables up to four GPUs in a single system”, with “up to 384 GB of combined memory across a four-GPU configuration”.
What the lower power costs is clock speed: 2,280 MHz boost instead of 2,617, 110 TFLOPS of FP32 instead of 125, 3,511 AI TOPS instead of 4,000. Puget Systems measured the Max-Q 5 to 13 per cent behind the Workstation Edition in Blender and about 14 per cent behind in DaVinci Resolve and Unreal Engine. NVIDIA describes the cooler only as “active”; Puget describes it as a blower that takes air from inside the case and pushes it out through the back, which is what lets four cards sit side by side. A double flow-through card releases its heat into the case instead, which, with its 600 W draw, is the likeliest reason the makers in the table below list only one or two per machine.
Which towers take four
| TOWER | PROCESSOR | MAX-Q CARDS | 600 W CARDS | POWER SUPPLY |
|---|---|---|---|---|
| HP Z8 Fury G6i | Intel Xeon 600, up to 86 cores | 4 | not listed | 1,700 W at 200 V or more; two supplies up to 2,700 W combined |
| Lenovo ThinkStation PX | two Intel Xeon Scalable processors | 4 | not listed | 1,850 W; up to 2,350 W with two supplies in team mode |
| Dell Precision 7960 | Intel Xeon W-3400 or W-3500 | 4 | not stated | 1,400 W or 2,200 W; the 2,200 W unit gives its full rating at 180 to 264 V |
| Dell Pro Precision 9 T6 | Intel Xeon 600 | 5 | 2 | 2,400 W at 200 to 240 V |
| Lenovo ThinkStation P8 | AMD Threadripper PRO 9000 or 7000 WX | 3 | 1 | 1,000 or 1,400 W |
Makers’ product pages, configurators and specification documents, September 2026; the Pro Precision 9 T6 input-voltage figure from StorageReview’s launch report. HP’s previous Z8 Fury G5 also took four Max-Q cards or one Workstation Edition.
Two things stand out. The towers that take four Max-Q cards do not list four 600 W cards: the most any of these makers lists is two, in Dell’s Pro Precision 9 T6. And every one of these power supplies delivers its full rating on a 230 V European socket, while the lower figures in the makers’ tables apply to 100 to 127 V networks.
By the power budget below, the 1,400 W Dell unit cannot carry four cards at full load and a single HP supply would run at its limit, so a four-card order should name the larger or the dual supply.
Lanes: the processor decides
Four cards at full width need 64 PCIe lanes before the NVMe drives and the network card take theirs. AMD’s Threadripper PRO 9000 WX processors provide 128 PCIe 5.0 lanes, Intel’s Xeon 600 for workstations also 128 on its larger models such as the 658X and 698X, and the Xeon W-3500 series 112. All of them carry four cards at x16.
The slot wiring matters as much as the processor. StorageReview tested two Max-Q cards in HP’s Z8 Fury G6i, which runs every GPU slot at PCIe 5.0 x16, and in a Threadripper-based Dell Precision 7875 where the second card ran at PCIe 4.0 x16. Serving Llama 3.1 8B in FP8 with vLLM at a batch of 256, the HP reached 23,004 tokens per second and the Dell 16,833; the review credits both the slots and the processor’s handling of vLLM’s scheduling. Read the slot table of a tower before its processor list.
Power: a 2 kW appliance
Four cards at 300 W make 1,200 W. A 350 W processor, the rated power of both a Threadripper PRO 9995WX and a Xeon 698X, whose turbo limit is 420 W, brings the total to 1,550 W before memory, drives, fans and the losses in the power supply. At the wall that is roughly 1.8 to 2 kW under full load. This is our estimate from the rated figures: we have found no independent test of a four-Max-Q tower with a measured wall figure.
On a 230 V circuit, 2 kW is about 9 A. A 16 A circuit, 3,680 W, carries the machine with room to spare, but not two of them together with the rest of an office. Size an uninterruptible power supply with headroom above 2 kW, and check that the building’s circuit is not already shared with a kitchen or a printer room.
Heat and noise
Practically all of that power ends up as heat in the room: 2 kW is about 6,800 BTU per hour, as much as a 2 kW fan heater running all day. A small office without air conditioning will warm up noticeably, and the machine’s own inlet temperature rises with it. Plan the room as well as the box.
We have found no independent noise measurement of a four-card Max-Q tower either. Ask for the configured system’s acoustic data before it goes next to someone’s desk, and put the rear of the machine, where the blowers exhaust, towards open space.
What 384 GB runs
| MODEL | PRECISION | WEIGHTS | CARDS NEEDED |
|---|---|---|---|
| gpt-oss-120b | MXFP4 | 60.8 GiB | one per instance: four independent servers |
| Llama 3.3 70B | FP8 | 68 GiB | one, with an FP8 cache for about fourteen 8k conversations; two for more |
| Qwen3-235B-A22B | NVFP4 | 134 GB | two |
| Qwen3-235B-A22B | FP8 | 236 GB | four |
| GLM-4.5 | FP8 | 361 GB | four on paper only: about 46 GiB left for the runtime and cache, below the model card’s minimum of eight H100 or four H200 |
| Llama 3.1 405B | FP8 | about 487 GB | does not fit |
| DeepSeek-V3 or R1 | FP8 | about 689 GB | does not fit |
Weights from the model cards and published checkpoints on Hugging Face, in GB or GiB as published; each card holds 95.6 GiB (about 102.6 GB) as the driver reports it. The Llama 3.1 405B FP8 figure is our sum from the checkpoint’s parameter types. The Qwen3-235B FP8 model card’s own serving examples use four GPUs; we found no published benchmark of it on four RTX PRO 6000 cards.
The useful configurations are the first four rows: one large mixture-of-experts model across all four cards, two or four smaller models side by side, or a 70B model with room for a team. For published multi-card numbers, StorageReview ran four RTX PRO 6000 Server Edition cards, which have about 11 per cent less bandwidth than the Max-Q, with vLLM splitting the model across all four: gpt-oss-120b in NVFP4 reached 105.8 tokens per second per user with four users, and Llama 2 70B reached 32.9 tokens per second for a single user.
Splitting one model over PCIe
The RTX PRO 6000 has no NVLink, so a model split across cards exchanges data over PCIe. vLLM’s documentation recommends tensor parallelism when a model is too large for one GPU but fits one node, and in a note on uneven GPU splits adds: “if the GPUs on the node do not have NVLINK interconnect (e.g. L40S), leverage pipeline parallelism instead of tensor parallelism for higher throughput and lower communication overhead”. Test both on your own model.
Direct transfers between the cards help. Google reports up to 168 per cent more throughput for tensor-parallel serving on RTX PRO 6000 Server Edition machines with a PCIe peer-to-peer data path, compared with instances without one. NVIDIA’s NCCL documentation explains one reason they fail: on bare-metal Linux, “CUDA and the NVIDIA driver stack do not support IOMMU-enabled PCIe peer-to-peer memory transfer”. Check the IOMMU and ACS settings in the BIOS before blaming the cards.
Sixteen slices for a team: MIG
Each Max-Q splits into up to four isolated 24 GB instances, so the tower can offer sixteen independent GPUs to a team of developers. MIG runs on Linux with driver 575.51.03 or later and a Max-Q vBIOS of 98.02.6A.00.00 or later, and each card has to be switched from graphics to compute mode with NVIDIA’s DisplayModeSelector first. On a card that drives the monitor, that switch turns the display outputs off, so a MIG workstation either runs headless or keeps a small separate card for the screen. Profiles with the “+gfx” suffix keep graphics support inside an instance. vGPU for virtual machines is not available on the Max-Q; it is a Server Edition feature, as our MIG and vGPU guide explains.
Tower or server
Four Server Edition cards in a rack server offer the same 384 GB with passive cooling, a configurable power limit of up to 600 W, vGPU support and none of the office problems, but they need a server room with its airflow and noise. Our article on how many GPUs fit in one server covers that side. The tower is the right answer for one team without a server room; the server is the right answer once the machine is shared across departments or runs around the clock for customers.
What we supply
Eurokommerz supplies the RTX PRO 6000 Blackwell Max-Q EU-wide with manufacturer warranty, as single cards or in a configured four-card workstation or server with the power supply, platform and cooling sized for the load. Send us the models and the number of users, and we will return the configuration and its power budget.
FAQ
Can four RTX PRO 6000 Workstation Edition cards go in one tower?
How much power does a four-card Max-Q workstation need?
Do four cards work as one 384 GB GPU?
Which processor does a four-GPU workstation need?
Can several people share the workstation?
How much slower is a Max-Q than a Workstation Edition?
Tell us the models you want to run, how many people will use the machine and where it will stand. We will size the cards, the platform and the power, and say whether a tower or a server fits better. We reply within one business day.
Talk to an expertWe reply within one business day