Eight RTX PRO 6000 in one server: what an 8-GPU server runs for a company of 1,000 staff
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- Eight RTX PRO 6000 Server Edition cards give 768 GB of GPU memory as eight separate 96 GB pools: the card has no NVLink and connects over PCIe 5.0 x16, so the server is planned card by card
- In our example for 1,000 staff, four cards run gpt-oss-120b as one copy each for chat and RAG, one runs a coding model, one is split by MIG into four 24 GB instances, one serves a document model and one stays in reserve
- With the example values of our sizing guide, 1,000 staff produce about 40 requests in flight at the peak by our estimate; four copies of gpt-oss-120b hold 76 conversations at 32K, and 57 with one card out
- NVIDIA’s RTX PRO AI Factory reference architecture builds on an eight-card node it calls 2-8-5-200: 2 CPUs, 8 GPUs and 5 NICs at 200 Gbps each, with at least 128 GB of system memory per GPU
- Eight cards at up to 600 W draw 4.8 kW before processors and fans; at 2,000 staff the example peak doubles to 80, and two servers with five assistant copies each keep it if one fails
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
What eight RTX PRO 6000 in one server can run
An 8x RTX PRO 6000 server runs every LLM that fits in 96 GB as one or more separate copies, and larger models across several cards over PCIe. Its eight RTX PRO 6000 Blackwell Server Edition cards have 768 GB of GPU memory, but as eight separate pools of 96 GB, because the card has no NVLink and the cards exchange data only over PCIe 5.0 x16. The server is therefore planned card by card: every model that fits one card gets one or more cards of its own, and only a model too large for 96 GB is split across several.
For a company of 1,000 staff, our example gives four cards to the main chat and RAG model, one copy per card. One card runs a coding model, one is split by MIG into four 24 GB instances for embeddings, reranking, speech-to-text and a small task model, one serves a document vision-language model and batch jobs, and one stays in reserve as a test bed. The user numbers apply the example values of our guide to sizing a private ChatGPT server by company size, and the memory figures are our estimates, not measurements. One server has no failover for a host fault, which is why the same guide places two servers from 500 employees; the section on 2,000 staff covers the second server.
The RTX PRO Server and NVIDIA’s 2-8-5-200 node
NVIDIA announced RTX PRO Servers, built by its system partners, on 11 August 2025 in designs “capable of supporting two, four or eight NVIDIA RTX PRO 6000 Blackwell GPUs”, with eight-card configurations in 4U. Its RTX PRO AI Factory reference architecture, last updated on 18 May 2026, builds on the eight-card version and calls the node a “2-8-5-200 infrastructure configuration (2 CPUs, 8 GPUs, 5 NICs at 200 Gbps each)”. The five adapters are one BlueField-3 DPU for north-south traffic and four BlueField-3 SuperNICs for traffic between nodes. NVIDIA writes that the design “Efficiently supports inference for small and medium model sizes” and is “ideal for multi-user, single tenant workloads”, which describes one company serving its own staff.
| ITEM | PER CARD | EIGHT-CARD SERVER |
|---|---|---|
| GPU memory | 96 GB GDDR7 with ECC | 768 GB in eight pools |
| Memory bandwidth | 1,597 GB/s | 1,597 GB/s per pool |
| Maximum power | up to 600 W, configurable | up to 4.8 kW of GPU power |
| MIG | up to 4 instances of 24 GB | up to 32 instances |
| Host link | PCIe 5.0 x16, no NVLink | one Gen5 x16 link per GPU at least |
| System memory | set per server | at least 128 GB per GPU, 1 TB |
| Network | set per server | 1 BlueField-3 DPU, 4 SuperNICs |
NVIDIA RTX PRO 6000 Blackwell Server Edition product page and datasheet; Lenovo product guide LP2263 (updated 28 July 2026) for NVLink and host interface; NVIDIA RTX PRO AI Factory reference architecture (18 May 2026) for links, system memory and network; read on 10 October 2026.
Example allocation for a company of 1,000 staff
Our sizing guide uses four example values: 40 per cent of employees active in the busiest hour, 6 requests per user and hour, 30 seconds per request and a peak factor of 2. The guide works them through for 500 and 2,000 employees. Applied to 1,000, they give 400 busy-hour users, 2,400 requests an hour, about 20 requests in flight on average and 40 at the example peak, by our estimate. With gpt-oss-120b at a declared 32K context and a 16-bit KV cache, the same guide estimates 19 conversations per RTX PRO 6000, from 22.2 GiB of cache beside 60.8 GiB of weights.
| CARD | WORKLOAD | EXAMPLE MODEL | MEMORY USED |
|---|---|---|---|
| Cards 1 to 4 | chat and RAG assistant, one copy per card | gpt-oss-120b | 60.8 GiB weights, 22.2 GiB cache, 19 conversations at 32K per card |
| Card 5 | coding assistant and agents | Qwen3-Coder-30B-A3B, FP8 | 31.2 GB weights; 9 sessions of 64K, about 19 with an FP8 cache |
| Card 6, MIG 4 × 1g.24gb | embeddings, reranking, speech-to-text, task model | Qwen3 0.6B pair, Whisper large-v3, gpt-oss-20b | 1.2 GB, 1.2 GB, about 10 GB and 13.8 GB, one model per instance |
| Card 7 | document model by day, batch jobs at night | Qwen3.8-27B, FP8 | 30.9 GB weights, 25 conversations at 32K |
| Card 8 | reserve and test bed | new model versions, engine upgrades | empty, or a fifth gpt-oss-120b copy |
Example allocation, not a measurement. Memory figures are our estimates from our guides to private ChatGPT sizing, coding assistants, Qwen, embeddings and speech-to-text: 0.9 × the 95.6 GiB the driver reports, less 3 GiB per card (the coding guide keeps the 3 GiB); Whisper’s 10 GB is OpenAI’s figure; MIG profile from NVIDIA’s MIG user guide (11 September 2026).
Four copies of the assistant hold 76 conversations at 32K against the example peak of 40, and 57 if one card fails or is taken out for a driver problem. A load balancer spreads the requests over the four copies. The coding model has a card of its own because agents send prompts of tens of thousands of tokens on every step, and on a shared card that cache would compete with the assistant’s conversations. Qwen3.8-27B reads scanned pages and images, so card 7 answers questions about documents during the day and runs queued extraction jobs, such as invoices or contracts, overnight.
Card 8 is the place to try a new model version or a new vLLM release on the production hardware before it replaces a running copy. If the peak grows, it becomes a fifth assistant copy, and five copies hold 95 conversations. How several engines are started, pinned to cards and exposed under one endpoint is covered in our guide to serving several models on one GPU server.
We build servers with up to eight RTX PRO 6000 Server Edition cards, sized by the models and concurrent users. Tell us your use cases and the number of staff, and we propose a configuration with the card allocation.
Which department uses which cards
The same plan reads differently for the people who approve it. Department heads want to know which cards carry their workload and what makes that share grow.
| DEPARTMENT USE | CARDS | EXAMPLE MODEL | WHAT SETS THE SHARE |
|---|---|---|---|
| All staff, assistant and RAG | 1 to 4, retrieval on 6 | gpt-oss-120b | peak requests in flight |
| Software development | 5 | Qwen3-Coder-30B-A3B | prompt length of agent steps |
| Service desk, meetings | one MIG instance on 6 | Whisper large-v3 | live streams, hours of audio |
| Finance, procurement, legal | 7 | Qwen3.8-27B | pages per day, batch window |
| IT and AI team | 8 | candidates under test | releases to validate |
Example departments and models; shares follow the allocation table above. Model figures from our sizing guides as read on 10 October 2026.
Keeping each department’s workload on its own cards has a second effect. When the developers’ demand doubles, card 5 becomes two cards and the assistant copies stay as they are. A department that starts later, for example a translation service, takes the reserve card first and moves to a second server once the reserve is needed for failover again.
MIG on the shared card
NVIDIA’s MIG user guide, updated on 11 September 2026, offers three sizes on the RTX PRO 6000: 1g.24gb up to four times, 2g.48gb up to two times and 4g.96gb once. Each instance has its own memory and fault isolation, so a fault in the speech-to-text engine does not stop the embedding model. A plain 1g.24gb instance has one video decoder and one encoder, and the 1g.24gb-me variant leaves out all media engines for pure compute. There is no size between 24 and 48 GB, so a model of 30 GB takes a 2g.48gb instance.
MIG suits the small models of a RAG pipeline and the task model that writes chat titles and tags. It adds nothing for gpt-oss-120b, whose 60.8 GiB need the whole card. The Server Edition leaves the factory in display-off mode, so it needs no display-mode switch before MIG. On Hopper and later GPUs, Blackwell included, MIG mode and its instances do not survive a reboot or a driver reload and have to be recreated at start-up, as our MIG setup guide for RTX PRO Blackwell shows step by step.
PCIe topology, system memory and power
The reference architecture places the cards so that “8 GPUs are balanced across CPU sockets and root ports”, with at least one Gen5 x16 link per GPU and PCIe switches only where needed. In the example, every model fits one card, so no data moves between cards while the server serves requests. Topology matters once a model is split, for example a model above 96 GB with tensor parallelism. Its cards then belong on the shortest path that nvidia-smi topo -m shows. Where the GPUs of a node have no NVLink, vLLM’s documentation advises pipeline parallelism instead of tensor parallelism “for higher throughput and lower communication overhead”, and our article on one model across several GPUs over PCIe and NVLink covers the traffic each method sends.
The same document asks for at least 128 GB of system memory and 7 physical CPU cores per GPU, 1 TB and 56 cores for eight cards, and at least one 1 TB NVMe drive per CPU socket for inference servers. NVIDIA’s configuration guide for NVIDIA-Certified Systems sets a higher memory floor, twice the total GPU memory, which is 1,536 GB for eight cards. Model checkpoints, the RAG index and logs come on top of that minimum. Pin each engine to the processor socket its card is attached to, which the topology output also shows.
NVIDIA gives the card’s maximum power as “Up to 600W (configurable)”. Eight cards draw 4.8 kW before processors, memory and fans, which calls for three-phase feeds or several single-phase circuits at the rack position. Lenovo’s product guide documents a cap at 450 W “to support increased density”, which it uses to fit four cards in its SR650a V4; at 450 W, eight cards draw 3.6 kW. Check the server maker’s rules before planning a capped configuration.
What changes at 2,000 staff
With the same example values, 2,000 employees produce 800 busy-hour users and 80 requests in flight at the peak. Our sizing guide puts five RTX PRO 6000 copies of gpt-oss-120b on each of two servers, which hold 190 conversations together and 95 with one server down. On two eight-card servers, cards 1 to 5 run the assistant, card 6 the coding model, card 7 the MIG instances and card 8 the document model. Each server runs every model, so the second server takes over the reserve role that card 8 had at 1,000 staff.
At 1,000 staff, one server remains a single point of failure for a host fault, a firmware update or a driver change, even with a spare card inside. Our sizing guide places two servers from 500 employees for that reason. A company that cannot stop the assistant during a maintenance window plans the second server from the start, either a second eight-card server or two four-card servers, as compared in one 8-GPU server or two 4-GPU servers.
We supply AI servers with 2 to 8 GPUs per node, so one request can cover both layouts. Describe your headcount, sites and maintenance windows in the form below, and we propose a configuration for each layout.
What we supply
We build AI servers to order with up to eight RTX PRO 6000 Server Edition cards, assembled and burn-in tested, with manufacturer warranty on every component and delivery anywhere in the EU, on one EU contract and invoice. We check the rack, power and airflow before we quote and return a configuration and quote within one business day. The cards are also available on their own through our GPU supply, alongside the H200 NVL, L40S and L4, and NVIDIA AI Enterprise licences come on the same invoice where the setup needs them. Models, RAG and the platform on top are our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
What can an 8x RTX PRO 6000 server run?
Do eight RTX PRO 6000 act as one 768 GB GPU?
How many users can an 8-GPU RTX PRO 6000 server serve?
How much power does a server with eight RTX PRO 6000 need?
Can MIG split the cards of an 8-GPU RTX PRO 6000 server?
When does a company need a second 8-GPU server?
Send us your headcount, the use cases by department, the models you are considering and the rack position’s power feed. We reply within one business day with a configuration, including the card allocation, and a quote, with the rack, power and airflow checked before we quote.
Talk to an expertWe reply within one business day