BLOG · GUIDE ·

Private AI platform architecture for 2,000 users: which servers do what

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • A private AI platform for 2,000 employees is a set of server roles (GPU inference servers, retrieval GPUs for embedding and reranking, a vector database and document store, a gateway with identity, observability, a model store and the Kubernetes control plane), not one large server
  • With the example values of our company-size guide, 2,000 employees peak at 80 requests in flight; two GPU servers with five RTX PRO 6000 Server Edition or two H200 NVL each carry that peak alone, so that one can fail
  • NVIDIA’s RAG sizing guide (22 August 2026) adds one reranker GPU, half an embedder GPU and a quarter GPU for the vector index to eight LLM GPUs; in our example, with every request at the peak reranking 40 passages, each inference server gets an L40S for retrieval
  • The CPU roles fit virtual machines on an existing cluster: three Kubernetes control plane nodes, because an etcd cluster of three tolerates one failure and a cluster of two tolerates none
  • The architecture stays the same from a 200-employee pilot on one two-card server through 1,000 employees on two servers with three RTX PRO 6000 or one H200 NVL each to 2,000 employees and a disaster recovery copy

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

The server roles of a private AI platform

The architecture of a private AI platform for 2,000 employees is a set of server roles rather than one large server. GPU inference servers run the language models, a smaller GPU or MIG instances run the embedding model and the reranker, and CPU servers or virtual machines carry the vector database, the document store, the gateway with identity, observability, the model store and the Kubernetes control plane. In the worked example below, two GPU servers each hold the whole peak with five RTX PRO 6000 Server Edition or two H200 NVL per server, and most CPU roles run in two or three replicas.

The GPU count of this enterprise LLM infrastructure comes from requests in flight, not from headcount, and this article takes it from our guide to sizing a private ChatGPT server by company size rather than repeating the arithmetic.

What each role does and what it runs on

GPU inference servers run the serving engine, such as vLLM, with one copy of the model per card when the model fits one card. NVIDIA’s Enterprise RAG Retrieval Scaling and Sizing Guide, updated on 22 August 2026, says “The NIM LLM is the main component affecting performance” and scales the other services in ratios to it.

Retrieval models are small and are sized by their rate, not their memory. The guide’s baseline pairs one LLM GPU, running a 49B Nemotron model per card, with one GPU for the reranking service and half a GPU for the embedding service, “with Run:ai or MIG”. At eight times that scale, eight LLM GPUs still share one reranker GPU, half an embedder GPU and the quarter GPU of the index node used for ingestion, 9.75 GPUs on two nodes in all. The baseline is a chat load of 128 input and 128 output tokens at a concurrency of 20 per LLM GPU, so longer RAG prompts call for a test of these ratios.

The vector database and the document store need CPU cores, ECC memory and NVMe rather than GPUs. In the same baseline, the Milvus query nodes run on CPU, listed as two nodes of 16 vCPU and 40 GiB each and scaled at one node per 4 million vectors up to 40 million, and the RAG server that assembles prompts from the retrieved passages has 16 vCPU and 64 GiB. The document store keeps the source files and the parsed chunk text.

The gateway is the only way in for users and applications. It checks each user against the directory service through single sign-on, applies keys and limits per team and writes the query log, as our article on the LLM gateway with keys, budgets and logging explains. Observability collects the serving engine’s metrics and the GPUs’ health; NVIDIA’s observability guide for its reference architectures, updated on 14 July 2026, uses Kube Prometheus and Grafana, with the DCGM exporter among its data sources. The model store keeps every checkpoint in use and the previous one, centrally and on NVMe in each GPU server, so that a restarted server loads from local disk. gpt-oss-120b is a 65.3 GB checkpoint, which needs about 21 seconds at the full line rate of 25 Gbit/s, before protocol overhead.

Worked example: 2,000 employees on two GPU servers

Our company-size guide uses example values of 40 per cent busy-hour users, 6 requests an hour each, 30 seconds per request and a peak factor of 2. For 2,000 employees they give a peak of 80 requests in flight. With gpt-oss-120b at a declared 32K context and a 16-bit KV cache, one RTX PRO 6000 holds about 19 conversations and one H200 NVL about 55, by that guide’s estimate. Two servers that each carry the peak alone need five RTX PRO 6000 or two H200 NVL per server, and keep 95 or 110 conversations if one fails.

ROLEHARDWAREIN THE EXAMPLEAVAILABILITY
LLM inferenceGPU server with 5 × RTX PRO 6000 Server Edition or 2 × H200 NVL2 serverseach holds the peak of 80 alone
Embedding and reranking1 × L40S in each inference server, or 24 GB MIG instances on a sixth RTX PRO 60002 cardsone per server, so the surviving server has both
Vector databaseCPU nodes or VMs with ECC memory and NVMe3 nodescollections replicated on at least two nodes
Document storeexisting file or object storage1 replicated volumesnapshots and backup like any file service
Gateway and identityVMs2 instancesbehind a load balancer, single sign-on
Kubernetes control planeVMs or small servers3 nodesetcd keeps its quorum with one node down
Observability and query logVMs2outside the GPU servers, so they report a GPU server failure
Model storeNVMe in each GPU server plus a central copy2 local, 1 centrala restarted server loads locally

GPU counts from our company-size and high-availability guides (gpt-oss-120b at 32K, 16-bit KV cache, peak of 80 by example values); etcd failure tolerance from the etcd v3.6 FAQ; the other counts are the example’s design choices.

Each server must hold every model of the service, so the retrieval card sits in both, and after a failure one card carries the whole retrieval load. Our embedding and reranker guide estimates from NVIDIA’s figures that, at 40 passages per query, an L4 reranks about 1.3 queries per second with NVIDIA’s 1B reranker and an L40S about 4.3. This example assumes that every request at the peak is a RAG query with 40 passages. Then 80 requests in flight at 30 seconds each arrive at about 2.7 per second, so the example plans an L40S per server. With fewer passages per query, or retrieval on only part of the requests, the load falls in proportion, and the L4 that our company-size guide places for the 0.6B Qwen3 models may carry it. Test the card with your own reranker and passage count.

With three servers, the two left after a failure hold the peak with three RTX PRO 6000 or one H200 NVL each, nine cards instead of ten or three instead of four. Our guide to high availability across two GPU nodes compares these layouts.

We build AI servers to order with 2 to 8 GPUs per node, sized by model size and concurrent users. Send us your headcount, models and RAG corpus size through the form below, and we return a configuration and quote within one business day.

What NVIDIA’s Enterprise Reference Architectures describe

NVIDIA describes its Enterprise Reference Architectures as “validated, repeatable designs” that “bring together compute, networking, storage, and software”. Its RTX PRO AI Factory design, last updated on 18 May 2026, is “based on a 2-8-5-200 infrastructure configuration (2 CPUs, 8 GPUs, 5 NICs at 200 Gbps each)”. Its examples have 16 nodes with 128 GPUs and 32 nodes with 256.

Two of its choices carry over to a platform of two or three GPU servers. The first is separate control plane nodes, since “control plane nodes are needed to run the software that manages the cluster and provides access for users”. Its example has three for the Kubernetes control plane besides those for cluster management and Slurm, each with two 32-core processors and at least 256 GB of memory. The second is separate networks, named in the design as compute east-west, converged north-south and out-of-band management networks.

The east-west fabric and the 200 Gbps NICs serve jobs that span nodes, while a platform that runs one copy of the model per card sends no request across servers. For two or three GPU servers, virtual machines carry the control plane roles.

Bare-metal Kubernetes or virtual machines for the CPU roles

In the example, the GPU servers are bare-metal Kubernetes worker nodes. The serving engine sees whole cards and MIG instances without a hypervisor layer, and the questions of passthrough, vGPU profiles and their licences do not arise. Kubernetes keeps the CPU workloads off these servers with taints, which in its documentation “allow a node to repel a set of pods”; only pods with a matching toleration, such as the serving engine and the retrieval models, can be placed there. The software layers on these nodes are covered in our guide to a private LLM platform on Kubernetes.

The CPU roles fit virtual machines on an existing virtualisation cluster. Place the replicas of each role on different hosts with anti-affinity rules. A control plane in virtual machines then depends on the availability of that cluster, and etcd needs a majority of its members. The etcd v3.6 FAQ states that “An etcd cluster needs a majority of nodes, a quorum, to agree on updates to the cluster state”, and its table gives three members a failure tolerance of one and two members none.

Bare-metal CPU servers for these roles make sense where no suitable virtualisation cluster exists, or where the platform must be administered separately from it.

Network zones for an on-premise AI platform

ZONEWHAT SITS THEREWHO MAY CONNECT
User accessbrowsers, client applications, integrationsthe gateway only, over HTTPS
Gatewaygateway, chat front end, RAG server, identity connectionusers; the directory service
InferenceGPU servers, model endpoints, metrics portsgateway, RAG server and monitoring only
Datavector database, document store, connectors, query loggateway, RAG server, ingestion jobs, restricted administrators
ManagementBMCs, Kubernetes API, monitoring, model storeadministrators; GPU servers to the model store

Zones are our example; NVIDIA’s RTX PRO AI Factory design names compute east-west, converged north-south and out-of-band management networks.

User traffic is streamed text and needs little bandwidth. The larger flows are checkpoints copied from the model store to the GPU servers and documents moved into the data zone during ingestion, so these two paths get the fast links. The model endpoints accept requests from the gateway and the RAG server only, which keeps keys, limits and the query log in force for every request. The query log holds prompts and answers, so it sits in the data zone with restricted access and a set retention.

From a 200-employee pilot to 2,000 employees

With the same example values, a pilot with 200 employees peaks at about 8 requests in flight, 1,000 employees at about 40 and 2,000 at 80. The roles stay the same, and each phase adds cards and copies.

PHASEPEAK IN FLIGHTGPU SERVERSWHAT IS ADDED
Pilot, 200 employeesabout 81 server, 2 × RTX PRO 6000 Server Edition: one runs the model, one in MIG for retrievalgateway with query log, one vector database node, monitoring
1,000 employeesabout 402 servers, each 3 × RTX PRO 6000 (57 held) or 1 × H200 NVL (55), plus a retrieval cardsecond server for N+1, three control plane nodes, replicated vector database
2,000 employees802 servers, each 5 × RTX PRO 6000 or 2 × H200 NVL plus an L40S, or 3 serversmore cards per server, a disaster recovery copy at a second site

Peaks by the example values of our company-size guide, as above; conversations per card for gpt-oss-120b at 32K with a 16-bit KV cache from the same guide.

The pilot server matches the Private AI starter on our AI servers page and holds about 19 conversations against a peak of 8, without failover. At 1,000 employees, two cards per server would hold 38 conversations, below the peak of 40, which is why that phase takes three. An eight-GPU chassis bought for the second phase takes the cards of the third without a new server. Before each step, replace the example values with figures from the gateway logs of the phase before.

Our Private AI/ML service starts with a pilot on one process with clear metrics. Tell us which department would go first and which document sources its assistant should use.

Where the disaster recovery copy sits

N+1 in one server room protects against the loss of a server, and a second site protects against the loss of the room. The disaster recovery copy is a further placement of the roles at a second site, with the vector index, the documents, the gateway configuration and the query log replicated, and the models copied from the model store. Its GPU capacity depends on which AI services must survive a site loss, and our article on disaster recovery for a private AI platform covers that decision.

What we supply

We supply the GPU servers of this architecture as AI servers built to order with 2 to 8 GPUs per node: the RTX PRO 6000 Server Edition, the H200 NVL with NVLink bridges and the L40S and L4 for retrieval, burn-in tested and with manufacturer warranty. NVIDIA AI Enterprise and vGPU licences come on the same invoice, under one EU contract. We check the rack, power and airflow before we quote. The platform on top, private LLMs, RAG, a Kubernetes-based platform and logging of queries and answers, is our Private AI/ML service, with engineering by our partner Vixen.UNO.

FAQ

What does a private AI platform architecture consist of?
It consists of server roles: GPU inference servers for the language models, a small GPU or MIG instances for the embedding model and the reranker, a vector database and document store, a gateway with identity and the query log, observability, a model store and the Kubernetes control plane. Only the inference and retrieval roles need GPUs, and the others run on CPU servers or virtual machines. Each role is sized and made redundant on its own.
How many GPU servers does an LLM platform for 2,000 users need?
With the example values of our company-size guide, 2,000 employees produce a peak of about 80 requests in flight. With gpt-oss-120b at 32K and a 16-bit KV cache, two servers with five RTX PRO 6000 or two H200 NVL each carry that peak alone, so that one can fail, and three servers need three RTX PRO 6000 or one H200 NVL each. Replace the example values with figures from a pilot before ordering.
Is there an on-premise AI platform reference architecture from NVIDIA?
NVIDIA publishes Enterprise Reference Architectures, which it calls validated, repeatable designs bringing together compute, networking, storage and software. Its RTX PRO AI Factory design builds on a 2-8-5-200 node with 2 CPUs, 8 GPUs and 5 NICs at 200 Gbps, with examples of 16 and 32 nodes. For two or three GPU servers its separate control plane nodes and separate networks carry over, while the multi-node fabric does not.
Should an AI platform run on bare-metal Kubernetes or in virtual machines?
A common split is bare-metal Kubernetes on the GPU servers, so the serving engine sees whole cards and MIG instances without passthrough or vGPU, and virtual machines for the CPU roles on an existing virtualisation cluster. The Kubernetes control plane needs three members, because an etcd cluster of three tolerates one failure and a cluster of two tolerates none. Replicas of each role belong on different hosts.
How should enterprise AI infrastructure grow from a pilot?
Keep the same roles and add cards and copies per phase. With our example values, a pilot of 200 employees peaks at about 8 requests in flight and fits one server with two RTX PRO 6000, 1,000 employees need two servers with three RTX PRO 6000 or one H200 NVL each, and 2,000 employees five RTX PRO 6000 or two H200 NVL per server. A disaster recovery copy at a second site follows once the platform carries the whole company.
Which network zones does an on-premise LLM platform need?
A workable split has a user access zone that reaches only the gateway, a gateway zone, an inference zone with the GPU servers, a data zone with the vector database, documents and the query log, and a management zone for BMCs, the Kubernetes API and monitoring. The model endpoints accept requests only from the gateway and the RAG server, so keys, limits and logging apply to every request. The large flows are checkpoint copies and document ingestion, not user traffic.

Send us your headcount, the use cases and models, the RAG corpus size, the virtualisation cluster the CPU roles could use and the rack positions available. We reply within one business day with a configuration and quote for the GPU servers, with the rack, power and airflow checked before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna