BLOG · GUIDE ·

Eight RTX PRO 6000 in one server: what an 8-GPU server runs for a company of 1,000 staff

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • Eight RTX PRO 6000 Server Edition cards give 768 GB of GPU memory as eight separate 96 GB pools: the card has no NVLink and connects over PCIe 5.0 x16, so the server is planned card by card
  • In our example for 1,000 staff, four cards run gpt-oss-120b as one copy each for chat and RAG, one runs a coding model, one is split by MIG into four 24 GB instances, one serves a document model and one stays in reserve
  • With the example values of our sizing guide, 1,000 staff produce about 40 requests in flight at the peak by our estimate; four copies of gpt-oss-120b hold 76 conversations at 32K, and 57 with one card out
  • NVIDIA’s RTX PRO AI Factory reference architecture builds on an eight-card node it calls 2-8-5-200: 2 CPUs, 8 GPUs and 5 NICs at 200 Gbps each, with at least 128 GB of system memory per GPU
  • Eight cards at up to 600 W draw 4.8 kW before processors and fans; at 2,000 staff the example peak doubles to 80, and two servers with five assistant copies each keep it if one fails

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

What eight RTX PRO 6000 in one server can run

An 8x RTX PRO 6000 server runs every LLM that fits in 96 GB as one or more separate copies, and larger models across several cards over PCIe. Its eight RTX PRO 6000 Blackwell Server Edition cards have 768 GB of GPU memory, but as eight separate pools of 96 GB, because the card has no NVLink and the cards exchange data only over PCIe 5.0 x16. The server is therefore planned card by card: every model that fits one card gets one or more cards of its own, and only a model too large for 96 GB is split across several.

For a company of 1,000 staff, our example gives four cards to the main chat and RAG model, one copy per card. One card runs a coding model, one is split by MIG into four 24 GB instances for embeddings, reranking, speech-to-text and a small task model, one serves a document vision-language model and batch jobs, and one stays in reserve as a test bed. The user numbers apply the example values of our guide to sizing a private ChatGPT server by company size, and the memory figures are our estimates, not measurements. One server has no failover for a host fault, which is why the same guide places two servers from 500 employees; the section on 2,000 staff covers the second server.

The RTX PRO Server and NVIDIA’s 2-8-5-200 node

NVIDIA announced RTX PRO Servers, built by its system partners, on 11 August 2025 in designs “capable of supporting two, four or eight NVIDIA RTX PRO 6000 Blackwell GPUs”, with eight-card configurations in 4U. Its RTX PRO AI Factory reference architecture, last updated on 18 May 2026, builds on the eight-card version and calls the node a “2-8-5-200 infrastructure configuration (2 CPUs, 8 GPUs, 5 NICs at 200 Gbps each)”. The five adapters are one BlueField-3 DPU for north-south traffic and four BlueField-3 SuperNICs for traffic between nodes. NVIDIA writes that the design “Efficiently supports inference for small and medium model sizes” and is “ideal for multi-user, single tenant workloads”, which describes one company serving its own staff.

ITEMPER CARDEIGHT-CARD SERVER
GPU memory96 GB GDDR7 with ECC768 GB in eight pools
Memory bandwidth1,597 GB/s1,597 GB/s per pool
Maximum powerup to 600 W, configurableup to 4.8 kW of GPU power
MIGup to 4 instances of 24 GBup to 32 instances
Host linkPCIe 5.0 x16, no NVLinkone Gen5 x16 link per GPU at least
System memoryset per serverat least 128 GB per GPU, 1 TB
Networkset per server1 BlueField-3 DPU, 4 SuperNICs

NVIDIA RTX PRO 6000 Blackwell Server Edition product page and datasheet; Lenovo product guide LP2263 (updated 28 July 2026) for NVLink and host interface; NVIDIA RTX PRO AI Factory reference architecture (18 May 2026) for links, system memory and network; read on 10 October 2026.

Example allocation for a company of 1,000 staff

Our sizing guide uses four example values: 40 per cent of employees active in the busiest hour, 6 requests per user and hour, 30 seconds per request and a peak factor of 2. The guide works them through for 500 and 2,000 employees. Applied to 1,000, they give 400 busy-hour users, 2,400 requests an hour, about 20 requests in flight on average and 40 at the example peak, by our estimate. With gpt-oss-120b at a declared 32K context and a 16-bit KV cache, the same guide estimates 19 conversations per RTX PRO 6000, from 22.2 GiB of cache beside 60.8 GiB of weights.

CARDWORKLOADEXAMPLE MODELMEMORY USED
Cards 1 to 4chat and RAG assistant, one copy per cardgpt-oss-120b60.8 GiB weights, 22.2 GiB cache, 19 conversations at 32K per card
Card 5coding assistant and agentsQwen3-Coder-30B-A3B, FP831.2 GB weights; 9 sessions of 64K, about 19 with an FP8 cache
Card 6, MIG 4 × 1g.24gbembeddings, reranking, speech-to-text, task modelQwen3 0.6B pair, Whisper large-v3, gpt-oss-20b1.2 GB, 1.2 GB, about 10 GB and 13.8 GB, one model per instance
Card 7document model by day, batch jobs at nightQwen3.8-27B, FP830.9 GB weights, 25 conversations at 32K
Card 8reserve and test bednew model versions, engine upgradesempty, or a fifth gpt-oss-120b copy

Example allocation, not a measurement. Memory figures are our estimates from our guides to private ChatGPT sizing, coding assistants, Qwen, embeddings and speech-to-text: 0.9 × the 95.6 GiB the driver reports, less 3 GiB per card (the coding guide keeps the 3 GiB); Whisper’s 10 GB is OpenAI’s figure; MIG profile from NVIDIA’s MIG user guide (11 September 2026).

Four copies of the assistant hold 76 conversations at 32K against the example peak of 40, and 57 if one card fails or is taken out for a driver problem. A load balancer spreads the requests over the four copies. The coding model has a card of its own because agents send prompts of tens of thousands of tokens on every step, and on a shared card that cache would compete with the assistant’s conversations. Qwen3.8-27B reads scanned pages and images, so card 7 answers questions about documents during the day and runs queued extraction jobs, such as invoices or contracts, overnight.

Card 8 is the place to try a new model version or a new vLLM release on the production hardware before it replaces a running copy. If the peak grows, it becomes a fifth assistant copy, and five copies hold 95 conversations. How several engines are started, pinned to cards and exposed under one endpoint is covered in our guide to serving several models on one GPU server.

We build servers with up to eight RTX PRO 6000 Server Edition cards, sized by the models and concurrent users. Tell us your use cases and the number of staff, and we propose a configuration with the card allocation.

Which department uses which cards

The same plan reads differently for the people who approve it. Department heads want to know which cards carry their workload and what makes that share grow.

DEPARTMENT USECARDSEXAMPLE MODELWHAT SETS THE SHARE
All staff, assistant and RAG1 to 4, retrieval on 6gpt-oss-120bpeak requests in flight
Software development5Qwen3-Coder-30B-A3Bprompt length of agent steps
Service desk, meetingsone MIG instance on 6Whisper large-v3live streams, hours of audio
Finance, procurement, legal7Qwen3.8-27Bpages per day, batch window
IT and AI team8candidates under testreleases to validate

Example departments and models; shares follow the allocation table above. Model figures from our sizing guides as read on 10 October 2026.

Keeping each department’s workload on its own cards has a second effect. When the developers’ demand doubles, card 5 becomes two cards and the assistant copies stay as they are. A department that starts later, for example a translation service, takes the reserve card first and moves to a second server once the reserve is needed for failover again.

MIG on the shared card

NVIDIA’s MIG user guide, updated on 11 September 2026, offers three sizes on the RTX PRO 6000: 1g.24gb up to four times, 2g.48gb up to two times and 4g.96gb once. Each instance has its own memory and fault isolation, so a fault in the speech-to-text engine does not stop the embedding model. A plain 1g.24gb instance has one video decoder and one encoder, and the 1g.24gb-me variant leaves out all media engines for pure compute. There is no size between 24 and 48 GB, so a model of 30 GB takes a 2g.48gb instance.

MIG suits the small models of a RAG pipeline and the task model that writes chat titles and tags. It adds nothing for gpt-oss-120b, whose 60.8 GiB need the whole card. The Server Edition leaves the factory in display-off mode, so it needs no display-mode switch before MIG. On Hopper and later GPUs, Blackwell included, MIG mode and its instances do not survive a reboot or a driver reload and have to be recreated at start-up, as our MIG setup guide for RTX PRO Blackwell shows step by step.

PCIe topology, system memory and power

The reference architecture places the cards so that “8 GPUs are balanced across CPU sockets and root ports”, with at least one Gen5 x16 link per GPU and PCIe switches only where needed. In the example, every model fits one card, so no data moves between cards while the server serves requests. Topology matters once a model is split, for example a model above 96 GB with tensor parallelism. Its cards then belong on the shortest path that nvidia-smi topo -m shows. Where the GPUs of a node have no NVLink, vLLM’s documentation advises pipeline parallelism instead of tensor parallelism “for higher throughput and lower communication overhead”, and our article on one model across several GPUs over PCIe and NVLink covers the traffic each method sends.

The same document asks for at least 128 GB of system memory and 7 physical CPU cores per GPU, 1 TB and 56 cores for eight cards, and at least one 1 TB NVMe drive per CPU socket for inference servers. NVIDIA’s configuration guide for NVIDIA-Certified Systems sets a higher memory floor, twice the total GPU memory, which is 1,536 GB for eight cards. Model checkpoints, the RAG index and logs come on top of that minimum. Pin each engine to the processor socket its card is attached to, which the topology output also shows.

NVIDIA gives the card’s maximum power as “Up to 600W (configurable)”. Eight cards draw 4.8 kW before processors, memory and fans, which calls for three-phase feeds or several single-phase circuits at the rack position. Lenovo’s product guide documents a cap at 450 W “to support increased density”, which it uses to fit four cards in its SR650a V4; at 450 W, eight cards draw 3.6 kW. Check the server maker’s rules before planning a capped configuration.

What changes at 2,000 staff

With the same example values, 2,000 employees produce 800 busy-hour users and 80 requests in flight at the peak. Our sizing guide puts five RTX PRO 6000 copies of gpt-oss-120b on each of two servers, which hold 190 conversations together and 95 with one server down. On two eight-card servers, cards 1 to 5 run the assistant, card 6 the coding model, card 7 the MIG instances and card 8 the document model. Each server runs every model, so the second server takes over the reserve role that card 8 had at 1,000 staff.

At 1,000 staff, one server remains a single point of failure for a host fault, a firmware update or a driver change, even with a spare card inside. Our sizing guide places two servers from 500 employees for that reason. A company that cannot stop the assistant during a maintenance window plans the second server from the start, either a second eight-card server or two four-card servers, as compared in one 8-GPU server or two 4-GPU servers.

We supply AI servers with 2 to 8 GPUs per node, so one request can cover both layouts. Describe your headcount, sites and maintenance windows in the form below, and we propose a configuration for each layout.

What we supply

We build AI servers to order with up to eight RTX PRO 6000 Server Edition cards, assembled and burn-in tested, with manufacturer warranty on every component and delivery anywhere in the EU, on one EU contract and invoice. We check the rack, power and airflow before we quote and return a configuration and quote within one business day. The cards are also available on their own through our GPU supply, alongside the H200 NVL, L40S and L4, and NVIDIA AI Enterprise licences come on the same invoice where the setup needs them. Models, RAG and the platform on top are our Private AI/ML service, with engineering by our partner Vixen.UNO.

FAQ

What can an 8x RTX PRO 6000 server run?
It runs every model that fits one 96 GB card as one or more copies, and larger models split across cards over PCIe, as long as weights and cache fit the eight cards together, about 664 GiB by our sizing rule. In our example for 1,000 staff, four cards run gpt-oss-120b for chat and RAG, and one each runs a coding model, MIG instances for retrieval and speech-to-text, a document model and a reserve. NVIDIA’s reference architecture describes the eight-card node as suited to inference for small and medium model sizes.
Do eight RTX PRO 6000 act as one 768 GB GPU?
No. Each card keeps its own 96 GB, and the RTX PRO 6000 has no NVLink, so cards that share a model exchange data over PCIe 5.0 x16. Models that fit one card are therefore run as separate copies, one per card, with no traffic between the cards.
How many users can an 8-GPU RTX PRO 6000 server serve?
That depends on the requests in flight at the peak, not on the headcount. With the example values of our sizing guide, 1,000 employees produce about 40 requests in flight, and four cards running gpt-oss-120b hold 76 conversations at a declared 32K context, both by our estimate. The other four cards remain for coding, retrieval, documents and a reserve.
How much power does a server with eight RTX PRO 6000 need?
NVIDIA gives each card a maximum of up to 600 W, configurable, so eight cards draw 4.8 kW before processors, memory and fans. Lenovo documents a 450 W cap for higher density, which brings the cards to 3.6 kW. Plan three-phase feeds or several circuits at the rack position before ordering.
Can MIG split the cards of an 8-GPU RTX PRO 6000 server?
Yes. NVIDIA’s MIG user guide lists up to four 1g.24gb instances, two 2g.48gb or one 4g.96gb per RTX PRO 6000, each with its own memory and fault isolation. MIG suits embedding models, rerankers, speech-to-text and small task models, while a model such as gpt-oss-120b needs a whole card.
When does a company need a second 8-GPU server?
A second server is needed when the service must keep running through a host failure or a maintenance window, and when the peak outgrows one server. With our example values, 2,000 employees produce 80 requests in flight, and two servers with five gpt-oss-120b copies each still hold 95 conversations if one fails. Two four-card servers are an alternative where one server per site or a smaller failure domain is the goal.

Send us your headcount, the use cases by department, the models you are considering and the rack position’s power feed. We reply within one business day with a configuration, including the card allocation, and a quote, with the rack, power and airflow checked before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna