BLOG · GUIDE ·

Four RTX PRO 6000 Max-Q in one workstation: 384 GB under a desk, and what it takes

IN BRIEF
  • NVIDIA built the Max-Q edition for dense workstations: the full 96 GB and 1,792 GB/s at 300 W, “up to four GPUs in a single system”, 384 GB in total
  • Towers with four Max-Q cards in their official specifications include the HP Z8 Fury G6i, the Lenovo ThinkStation PX and the Dell Precision 7960; Dell’s Pro Precision 9 T6 takes five. In these makers’ specifications the 600 W Workstation Edition stops at one or two per tower
  • Four cards at x16 need 64 PCIe 5.0 lanes: Threadripper PRO 9000 WX and the larger Xeon 600 models offer 128, Xeon W-3500 offers 112
  • Four cards and a 350 W processor come to about 1.55 kW before memory, drives and fans; plan for 1.8 to 2 kW at the wall, all of it heat in the room
  • There is no NVLink, so large models are split over PCIe; MIG turns the four cards into up to sixteen isolated 24 GB instances for a team

Why the Max-Q exists

The RTX PRO 6000 Blackwell comes in three editions with the same 24,064 CUDA cores and 96 GB of GDDR7, as our edition comparison explains. The Workstation Edition runs at 600 W with a double flow-through cooler, the Server Edition is passive for rack chassis, and the Max-Q keeps the full memory and the full 1,792 GB/s at 300 W. NVIDIA’s product page states the purpose plainly: “it enables up to four GPUs in a single system”, with “up to 384 GB of combined memory across a four-GPU configuration”.

What the lower power costs is clock speed: 2,280 MHz boost instead of 2,617, 110 TFLOPS of FP32 instead of 125, 3,511 AI TOPS instead of 4,000. Puget Systems measured the Max-Q 5 to 13 per cent behind the Workstation Edition in Blender and about 14 per cent behind in DaVinci Resolve and Unreal Engine. NVIDIA describes the cooler only as “active”; Puget describes it as a blower that takes air from inside the case and pushes it out through the back, which is what lets four cards sit side by side. A double flow-through card releases its heat into the case instead, which, with its 600 W draw, is the likeliest reason the makers in the table below list only one or two per machine.

Which towers take four

TOWERPROCESSORMAX-Q CARDS600 W CARDSPOWER SUPPLY
HP Z8 Fury G6iIntel Xeon 600, up to 86 cores4not listed1,700 W at 200 V or more; two supplies up to 2,700 W combined
Lenovo ThinkStation PXtwo Intel Xeon Scalable processors4not listed1,850 W; up to 2,350 W with two supplies in team mode
Dell Precision 7960Intel Xeon W-3400 or W-35004not stated1,400 W or 2,200 W; the 2,200 W unit gives its full rating at 180 to 264 V
Dell Pro Precision 9 T6Intel Xeon 600522,400 W at 200 to 240 V
Lenovo ThinkStation P8AMD Threadripper PRO 9000 or 7000 WX311,000 or 1,400 W

Makers’ product pages, configurators and specification documents, September 2026; the Pro Precision 9 T6 input-voltage figure from StorageReview’s launch report. HP’s previous Z8 Fury G5 also took four Max-Q cards or one Workstation Edition.

Two things stand out. The towers that take four Max-Q cards do not list four 600 W cards: the most any of these makers lists is two, in Dell’s Pro Precision 9 T6. And every one of these power supplies delivers its full rating on a 230 V European socket, while the lower figures in the makers’ tables apply to 100 to 127 V networks.

By the power budget below, the 1,400 W Dell unit cannot carry four cards at full load and a single HP supply would run at its limit, so a four-card order should name the larger or the dual supply.

Lanes: the processor decides

Four cards at full width need 64 PCIe lanes before the NVMe drives and the network card take theirs. AMD’s Threadripper PRO 9000 WX processors provide 128 PCIe 5.0 lanes, Intel’s Xeon 600 for workstations also 128 on its larger models such as the 658X and 698X, and the Xeon W-3500 series 112. All of them carry four cards at x16.

The slot wiring matters as much as the processor. StorageReview tested two Max-Q cards in HP’s Z8 Fury G6i, which runs every GPU slot at PCIe 5.0 x16, and in a Threadripper-based Dell Precision 7875 where the second card ran at PCIe 4.0 x16. Serving Llama 3.1 8B in FP8 with vLLM at a batch of 256, the HP reached 23,004 tokens per second and the Dell 16,833; the review credits both the slots and the processor’s handling of vLLM’s scheduling. Read the slot table of a tower before its processor list.

Power: a 2 kW appliance

Four cards at 300 W make 1,200 W. A 350 W processor, the rated power of both a Threadripper PRO 9995WX and a Xeon 698X, whose turbo limit is 420 W, brings the total to 1,550 W before memory, drives, fans and the losses in the power supply. At the wall that is roughly 1.8 to 2 kW under full load. This is our estimate from the rated figures: we have found no independent test of a four-Max-Q tower with a measured wall figure.

On a 230 V circuit, 2 kW is about 9 A. A 16 A circuit, 3,680 W, carries the machine with room to spare, but not two of them together with the rest of an office. Size an uninterruptible power supply with headroom above 2 kW, and check that the building’s circuit is not already shared with a kitchen or a printer room.

Heat and noise

Practically all of that power ends up as heat in the room: 2 kW is about 6,800 BTU per hour, as much as a 2 kW fan heater running all day. A small office without air conditioning will warm up noticeably, and the machine’s own inlet temperature rises with it. Plan the room as well as the box.

We have found no independent noise measurement of a four-card Max-Q tower either. Ask for the configured system’s acoustic data before it goes next to someone’s desk, and put the rear of the machine, where the blowers exhaust, towards open space.

What 384 GB runs

MODELPRECISIONWEIGHTSCARDS NEEDED
gpt-oss-120bMXFP460.8 GiBone per instance: four independent servers
Llama 3.3 70BFP868 GiBone, with an FP8 cache for about fourteen 8k conversations; two for more
Qwen3-235B-A22BNVFP4134 GBtwo
Qwen3-235B-A22BFP8236 GBfour
GLM-4.5FP8361 GBfour on paper only: about 46 GiB left for the runtime and cache, below the model card’s minimum of eight H100 or four H200
Llama 3.1 405BFP8about 487 GBdoes not fit
DeepSeek-V3 or R1FP8about 689 GBdoes not fit

Weights from the model cards and published checkpoints on Hugging Face, in GB or GiB as published; each card holds 95.6 GiB (about 102.6 GB) as the driver reports it. The Llama 3.1 405B FP8 figure is our sum from the checkpoint’s parameter types. The Qwen3-235B FP8 model card’s own serving examples use four GPUs; we found no published benchmark of it on four RTX PRO 6000 cards.

The useful configurations are the first four rows: one large mixture-of-experts model across all four cards, two or four smaller models side by side, or a 70B model with room for a team. For published multi-card numbers, StorageReview ran four RTX PRO 6000 Server Edition cards, which have about 11 per cent less bandwidth than the Max-Q, with vLLM splitting the model across all four: gpt-oss-120b in NVFP4 reached 105.8 tokens per second per user with four users, and Llama 2 70B reached 32.9 tokens per second for a single user.

Splitting one model over PCIe

The RTX PRO 6000 has no NVLink, so a model split across cards exchanges data over PCIe. vLLM’s documentation recommends tensor parallelism when a model is too large for one GPU but fits one node, and in a note on uneven GPU splits adds: “if the GPUs on the node do not have NVLINK interconnect (e.g. L40S), leverage pipeline parallelism instead of tensor parallelism for higher throughput and lower communication overhead”. Test both on your own model.

Direct transfers between the cards help. Google reports up to 168 per cent more throughput for tensor-parallel serving on RTX PRO 6000 Server Edition machines with a PCIe peer-to-peer data path, compared with instances without one. NVIDIA’s NCCL documentation explains one reason they fail: on bare-metal Linux, “CUDA and the NVIDIA driver stack do not support IOMMU-enabled PCIe peer-to-peer memory transfer”. Check the IOMMU and ACS settings in the BIOS before blaming the cards.

Sixteen slices for a team: MIG

Each Max-Q splits into up to four isolated 24 GB instances, so the tower can offer sixteen independent GPUs to a team of developers. MIG runs on Linux with driver 575.51.03 or later and a Max-Q vBIOS of 98.02.6A.00.00 or later, and each card has to be switched from graphics to compute mode with NVIDIA’s DisplayModeSelector first. On a card that drives the monitor, that switch turns the display outputs off, so a MIG workstation either runs headless or keeps a small separate card for the screen. Profiles with the “+gfx” suffix keep graphics support inside an instance. vGPU for virtual machines is not available on the Max-Q; it is a Server Edition feature, as our MIG and vGPU guide explains.

Tower or server

Four Server Edition cards in a rack server offer the same 384 GB with passive cooling, a configurable power limit of up to 600 W, vGPU support and none of the office problems, but they need a server room with its airflow and noise. Our article on how many GPUs fit in one server covers that side. The tower is the right answer for one team without a server room; the server is the right answer once the machine is shared across departments or runs around the clock for customers.

What we supply

Eurokommerz supplies the RTX PRO 6000 Blackwell Max-Q EU-wide with manufacturer warranty, as single cards or in a configured four-card workstation or server with the power supply, platform and cooling sized for the load. Send us the models and the number of users, and we will return the configuration and its power budget.

FAQ

Can four RTX PRO 6000 Workstation Edition cards go in one tower?
The makers’ specifications list at most one 600 W Workstation Edition per tower in HP’s previous Z8 Fury G5 and the Lenovo ThinkStation P8, and two in Dell’s Pro Precision 9 T6. Four cards in one tower is what the 300 W Max-Q is for.
How much power does a four-card Max-Q workstation need?
About 1,550 W for four cards and a 350 W processor, and roughly 1.8 to 2 kW at the wall under full load by our estimate. A 230 V 16 A circuit carries it.
Do four cards work as one 384 GB GPU?
No. The RTX PRO 6000 has no NVLink, so software splits a large model across the cards with tensor or pipeline parallelism over PCIe.
Which processor does a four-GPU workstation need?
One with at least 64 PCIe lanes for the cards plus the rest of the system: Threadripper PRO 9000 WX and the larger Xeon 600 models offer 128 PCIe 5.0 lanes, Xeon W-3500 offers 112. Check that every GPU slot is wired at x16.
Can several people share the workstation?
Yes, with MIG: each Max-Q splits into up to four isolated 24 GB instances, sixteen in total, on Linux with driver 575.51.03 or later, a current vBIOS and the display mode set to compute.
How much slower is a Max-Q than a Workstation Edition?
Puget Systems measured 5 to 14 per cent lower performance in content-creation tests. Memory size and bandwidth are the same: 96 GB and 1,792 GB/s.

Tell us the models you want to run, how many people will use the machine and where it will stand. We will size the cards, the platform and the power, and say whether a tower or a server fits better. We reply within one business day.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry – see our privacy policy.

request@eurokommerz.at  ·  +43 1 585 1405 50  ·  Jordangasse 7, 1010 Vienna