Private ChatGPT server for 100, 500 or 2,000 employees: sizing the GPUs by company size
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- A private ChatGPT server is sized by the peak number of requests in flight, not by headcount: employees active in the busiest hour, times their requests per hour, times the time each request takes, gives the average in flight
- We found no published ratio of active or concurrent users to employees, so the shares in this article are example values that a pilot replaces with figures from its own gateway logs
- With gpt-oss-120b at a declared 32K context and a 16-bit KV cache, one RTX PRO 6000 holds about 19 conversations and one H200 NVL about 55 by our estimate; at 8K the same cards hold about 79 and 222
- With the example values, 100 employees need one copy of the model on one RTX PRO 6000, 500 employees two copies and 2,000 employees five RTX PRO 6000 copies or two H200 NVL copies; from 500 employees upwards each of two servers holds the whole peak, so that one can fail
- A RAG assistant adds an embedding model, a reranker and a vector index; the 0.6B Qwen3 embedding and reranker models fit a small card such as the L4 or a 24 GB MIG instance of an RTX PRO 6000
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
How to size a private ChatGPT server for your company
A private ChatGPT server, an on-premise LLM for the whole company, is sized by the peak number of requests in flight, and you get there from the number of employees in three steps. Estimate how many employees use the assistant in the busiest hour, convert their requests into requests in flight with the time each one takes, and check that the GPUs hold the model weights plus a KV cache for every request at that peak. With gpt-oss-120b at a declared 32K context, one RTX PRO 6000 holds about 19 conversations and one H200 NVL about 55, by the estimate in our LLM hardware requirements by model.
With the example values below, that gives one copy of the model on one card for 100 employees, two copies for 500 and five RTX PRO 6000 copies or two H200 NVL copies for 2,000. From 500 employees two servers each hold the whole peak, so that one can fail, and every size adds capacity for the RAG models. The usage figures are examples to replace with your own, and the memory figures are our estimates, not measurements.
From employees to busy-hour users to requests in flight
The first step is the share of employees who use the assistant in the busiest hour. We found no published ratio between headcount and active or concurrent users, and none of the NVIDIA sizing documents we read assumes one. NVIDIA’s blog on LLM inference cost, of June 2025, makes the related point about concurrency: “Note that this isn’t the same as the number of concurrent users, because not all will have an active request at once.” It sizes from peak requests per second and latency targets instead.
The second step converts users into requests in flight. Little’s law from queueing theory states that the average number in a system equals the arrival rate times the average time each item spends in it. For an assistant, that is requests per second at the busiest hour times the end-to-end latency of a request, from submission to the last token. A peak factor then lifts the hourly average to the busiest minutes.
We use these example values throughout this article. They are not market statistics, and a pilot replaces each one with a figure from its gateway logs.
- Busy-hour users: 40 per cent of employees (example value).
- Requests per busy-hour user: 6 per hour (example value).
- Time in flight: 30 seconds per request, including reasoning tokens (example value).
- Peak factor: 2 times the hourly average (example value).
For 500 employees, 200 users send 1,200 requests an hour, or one every 3 seconds. At 30 seconds each, 10 requests are in flight on average and 20 at the example peak. The time in flight depends on the GPU, the engine and the load, so measure it on the GPU type production will use, as our guide on what a private AI pilot should measure explains. A slower test GPU shows more requests in flight for the same traffic.
Memory per conversation and conversations per card
The third step is memory. Each request in flight keeps a KV cache on the GPU, and the cards must hold it next to the weights. Our sizing rule, the same as in our guide to how many users one RTX PRO 6000 serves, takes 90 per cent of the memory the driver reports, minus about 3 GiB per card, minus the weights. The driver reports 95.6 GiB on a 96 GB RTX PRO 6000 and 140.4 GiB on a 141 GB H200 NVL.
gpt-oss-120b, with 117B parameters of which 5.1B are active, is a 65.3 GB checkpoint with its expert weights in MXFP4, about 60.8 GiB. OpenAI’s model card places it in use cases “that fit into a single 80GB GPU”. Its config.json gives 36 layers, of which 18 attend over a 128-token window and 18 over the whole context, with 8 key/value heads of dimension 64. A token therefore costs 36 KiB of 16-bit cache in the full-attention layers, and a conversation at 32K costs 1.125 GiB.
| GPU AND CONTEXT | LEFT FOR CACHE | PER CONVERSATION | CONVERSATIONS |
|---|---|---|---|
| RTX PRO 6000, 32K | 22.2 GiB | 1.125 GiB | 19 |
| H200 NVL, 32K | 62.5 GiB | 1.125 GiB | 55 |
| RTX PRO 6000, 8K | 22.2 GiB | 0.28 GiB | 79 |
| H200 NVL, 8K | 62.5 GiB | 0.28 GiB | 222 |
| RTX PRO 6000, 128K | 22.2 GiB | 4.5 GiB | 4 |
Our estimates for gpt-oss-120b with a 16-bit KV cache, one copy per card: 0.9 × driver-visible memory, less 3 GiB, less 60.8 GiB of weights; layers and heads from the model’s config.json on Hugging Face, read on 9 October 2026.
These are conversations filled to the declared length at the same moment, and shorter ones leave room for more. They limit memory only, and response time under load needs its own test. NVIDIA’s benchmarking blog notes that a busy system serves more requests through batching, but “this comes at the cost of increased latency”. When the cache runs out, vLLM’s tuning guide says “there are times when KV cache space is insufficient to handle all batched requests”, and it preempts requests and recomputes them later. Our gpt-oss-120b hardware guide covers the model’s variants.
Worked sizes for 100, 500 and 2,000 employees
The model fits one card, so more capacity comes from copies, one vLLM instance per card behind a load balancer, rather than from splitting the model. Each size below counts the copies the example peak needs, and from 500 employees upwards places them on two servers that each hold the whole peak, so that the service survives the loss of one.
| EMPLOYEES | BUSY-HOUR USERS | PEAK IN FLIGHT | CONFIGURATION | 32K SESSIONS HELD |
|---|---|---|---|---|
| 100 | 40 | 4 | one server, 2 × RTX PRO 6000 Server Edition: one runs gpt-oss-120b, one in MIG for RAG models | 19 |
| 500 | 200 | 20 | two servers, each 2 × RTX PRO 6000 Server Edition and 1 × L4 | 76, or 38 with one server down |
| 2,000 | 800 | 80 | two servers, each 2 × H200 NVL and 1 × L4; or each 5 × RTX PRO 6000 and 1 × L4 | 220 or 190; 110 or 95 with one down |
Users and peaks from the example values above (40 per cent, 6 requests an hour, 30 seconds, peak factor 2); sessions by our memory estimate for gpt-oss-120b at 32K from the table above.
For 100 employees the peak of 4 is far below the 19 conversations one card holds, so one card carries the model and the second card is split into 24 GB MIG instances for the embedding model, the reranker and a small task model. Failover for this size needs a second server of the same build. While the model is still being chosen, one DGX Spark (128 GB) holds gpt-oss-120b with about 30 conversations at 32K in memory, but its 273 GB/s of memory bandwidth makes it a pilot and development machine rather than the production server.
For 500 employees, 20 requests at the peak need two copies, and two servers with two cards each hold four. If one server stops, the other still holds 38 conversations. For 2,000 employees, 80 requests need five RTX PRO 6000 copies or two H200 NVL copies. Two servers with two H200 NVL each keep 110 conversations if one fails, and two servers with five RTX PRO 6000 each keep 95. With four RTX PRO 6000 per server, the remaining server would hold 76, below the example peak of 80.
If the evaluation set calls for a larger model, such as DeepSeek-V4-Flash under the MIT licence, our hub places one copy on two RTX PRO 6000 with tensor parallelism. Its configuration has one key/value head, and vLLM’s blog of 7 August 2026 notes that “once TP exceeds the number of KV heads, the cache starts duplicating across GPUs”. Each card of a copy therefore keeps the whole cache in the about 5.3 GiB left beside its half of the 167 GB -0731 weights, room for about 43 conversations at 0.12 GiB each in the FP8 cache that vLLM’s recipe sets, by our estimate. Two copies cover the example peak of 80, and two servers with four cards each keep it if one fails.
We build inference servers with 2 to 8 GPUs per node, sized by model size and concurrent users. Send us your headcount, the use cases and the models you are considering through the form below, and we size the server from them.
Embedding models, rerankers and index storage for RAG
An assistant that answers from company documents adds an embedding model, usually a reranker and a vector index, each with its own memory. The Qwen3 Embedding and Reranker series, under Apache 2.0, comes in 0.6B, 4B and 8B sizes with a 32K sequence length. Its 0.6B embedding model returns vectors of 1,024 dimensions and the 8B model up to 4,096. In BF16 the weights take about 1.2 GB for a 0.6B model and about 16 GB for an 8B model, by parameters times two bytes before activations.
The two 0.6B models fit any of the small cards we supply, from the 16 GB RTX PRO 2000 to the 24 GB L4 and RTX PRO 4000. The 8B pair takes about 32 GB of weights, which points to the 48 GB L40S, or to two 24 GB MIG instances on an RTX PRO 6000. NVIDIA’s MIG guide, updated on 11 September 2026, lists up to four instances on every RTX PRO 6000 edition and seven on the H200 NVL. MIG helps with these small models and with a task model for chat titles and tags, but not with gpt-oss-120b, whose 60.8 GiB need the whole card. Our private ChatGPT alternative guide explains why the front end’s background tasks should run on a model of their own.
The index needs fast storage sized from the corpus. At four bytes per value, one million chunks with 1,024-dimension vectors take about 4.1 GB before index structures and the chunk text, and 4,096 dimensions take four times as much. Plan NVMe space for the index, the source documents, every model checkpoint you keep and the query log.
What changes the size
| FACTOR | EFFECT ON THE SERVER | WHAT TO DECIDE |
|---|---|---|
| Declared context | cache per conversation grows with it: 19 at 32K, 79 at 8K, 4 at 128K on one RTX PRO 6000 | the context requests need, not the model’s maximum |
| Reasoning effort | longer answers raise the time in flight and fill more context | the level per use case |
| RAG prompts | retrieved passages lengthen every prompt | chunks per answer and their size |
| Front-end tasks | titles, tags and autocomplete add requests | a separate task model, or off |
| Peaks and rollout | a launch day or a company-wide training session can raise the peak factor | the peak to serve without queueing |
| Failover | one card or one server is a single point of failure | copies and hosts beyond the peak |
Context figures from the table above; reasoning levels from OpenAI’s gpt-oss-120b model card; front-end tasks from our private ChatGPT alternative guide.
Declared context has the largest effect on memory, and reasoning has a large effect on time. OpenAI’s card describes gpt-oss’s reasoning levels as “Low: Fast responses for general dialogue” up to “High: Deep and detailed analysis”, set in the system prompt. Reasoning tokens are generated like any answer, so a high level keeps each request in flight longer and raises the concurrency at the same request rate. Measure the time in flight at the level each use case will run.
Checking the estimate in a pilot
Every result above depends on the example values, so replace them before ordering hardware. Run a pilot with one department, log every request at the gateway with its start and end time, and take the busiest hour’s arrival rate and the distribution of prompt and answer lengths. Then load-test the candidate GPU at the expected peak with the pilot’s own prompts and check time to first token and per-user speed against targets you set before the test. Scale the arrival rate to the full headcount, not the concurrency the pilot showed.
If you want the pilot planned and measured for you, our Private AI/ML service starts with a pilot on one process with clear metrics. Tell us which process it should cover and how many employees would use the assistant.
What we supply
We supply the RTX PRO 6000 in its Workstation, Max-Q and Server editions, the H200 NVL with NVLink bridges, the DGX Spark Founders Edition and the small cards for RAG models, the L4, L40S and RTX PRO 2000, 4000 and 4500. They come as cards or in AI servers built to order with 2 to 8 GPUs per node, burn-in tested, with manufacturer warranty, on one EU contract and invoice. We check the rack, power and airflow before we quote and return a configuration and quote within one business day. Models, RAG and the platform on top are our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
How many GPUs does an LLM server for 500 users need?
How do I calculate concurrent users for an on-premise LLM?
How many users can one GPU serve with gpt-oss-120b?
What hardware does a self-hosted ChatGPT for a business need?
Is one GPU enough for a company of 100 employees?
Does MIG help when sizing an LLM server?
Send us your headcount, the use cases, the models you are considering, the context length you plan to declare and any pilot figures for requests in flight. We reply within one business day with a configuration and a quote, and we check the rack, power and airflow before we quote.
Talk to an expertWe reply within one business day