GPU capacity planning for an LLM platform: when to add the next server and how to decide it
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- Plan LLM capacity from busy-hour measurements, not daily averages: peak requests in flight, KV cache usage, queue time and 95th-percentile time to first token, set against thresholds you write down, plus an adoption forecast by quarter
- For planning, a server’s capacity is the lower of two figures: the conversations its KV cache holds at the declared context, and the concurrency at which a load test still meets the latency target
- With N+1, count capacity with one server down: for a peak of 80 requests in flight of gpt-oss-120b at 32K, two servers with five RTX PRO 6000 each keep 95 sessions, so the peak uses 84 per cent of them, above our example limit of 80 per cent, while four servers with two cards each keep 114 (our estimates)
- Model decisions move the load in steps: by our earlier estimates, 8,000-token reasoning raises the cards for 500 employees from 2 to between 5 and 16 RTX PRO 6000, and a declared 128K context cuts gpt-oss-120b sessions per RTX PRO 6000 from 19 to 4
- Kubernetes autoscaling deploys more pods, not more GPUs; on-premise, the next server is a purchase, so decide it at a quarterly review once the forecast crosses your threshold, well before an alert fires
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
LLM capacity planning: when to add the next GPU server
In GPU capacity planning for an LLM platform, add the next server when the measured busy-hour load, projected forward by your adoption forecast, crosses a headroom threshold that you set for the capacity left with one server down. The inputs are the platform’s metrics at the busiest hour of each working day (requests and tokens, KV cache usage, queue time and the 95th percentile of time to first token) and a forecast of users and use cases per quarter. Model decisions such as a larger model, a longer declared context, reasoning or agents change the load in steps, often by more than a new department does. On-premise capacity is a purchase, so the decision belongs in a quarterly review and should be taken early.
The worked example in this article is a platform for 2,000 employees running gpt-oss-120b at a declared 32K context with a 16-bit KV cache. It uses the example values of our guide to private ChatGPT servers by company size: 40 per cent of employees active in the busiest hour, 6 requests an hour each, 30 seconds per request and a peak factor of 2. With these values the peak of requests in flight is 4 per cent of the employees who have access, 80 for all 2,000 by our arithmetic.
Inputs: busy-hour metrics and an adoption forecast
vLLM’s documentation states that “These metrics are exposed via the /metrics endpoint on the vLLM OpenAI compatible API server.” Capacity planning uses the gauges vllm:num_requests_running, vllm:num_requests_waiting and vllm:kv_cache_usage_perc, the histograms for time to first token and request queue time, and the token counters vllm:prompt_tokens and vllm:generation_tokens, which Prometheus shows with the suffix _total. What each metric measures, and the alerts we suggest on it, are in our guide to LLM serving metrics in Prometheus. Keep one row per working day with the busy-hour peak of running plus waiting requests, the highest KV cache usage, the share of requests that waited, the 95th percentile of time to first token at the gateway and the tokens generated, for at least a year.
DCGM-Exporter, in NVIDIA’s documentation updated on 27 May 2026, “exposes GPU metrics at an HTTP endpoint (/metrics) for monitoring solutions such as Prometheus”, on port 9400 by default. Its power draw and temperatures are per GPU. The whole server’s draw, processors and fans included, comes from its BMC or the rack PDU, and both show how much of the rack feed and cooling the next server can still use. Memory in use is no load signal under vLLM, whose --gpu-memory-utilization sets “The fraction of GPU memory to be used for the model executor”, 0.92 by default. The engine claims its share at startup, and the cache inside that share fills and empties with the traffic. GPU utilisation is, in nvidia-smi’s documentation, the “Percent of time over the past sample period during which one or more kernels was executing on the GPU”, whatever share of the GPU those kernels use.
The forecast comes from the business. For each use case, write down per quarter the employees who get access, the share you expect at the busiest hour, requests per hour, tokens per request, the model and the declared context.
Capacity per server: memory and latency
For planning, a server’s capacity is the lower of two figures. One is the number of conversations its KV cache holds at the declared context beside the weights, which vLLM prints at startup and which our estimates put at 19 per RTX PRO 6000 and 55 per H200 NVL for the example model. The other is the number of requests in flight at which a load test with your own prompts still meets the targets for time to first token and tokens per second per user.
NVIDIA’s NIM benchmarking guide, updated on 20 July 2026, describes why the second figure is often lower: “As the number of concurrent requests increases, total system TPS increases while TPS per user decreases as latency increases.” Total throughput rises until the server “saturates the available GPU compute resources”, and the guide adds that “Beyond that point, TPS can decrease.” Run the load test on the GPU type production uses, and again after each engine upgrade, model change or change of context length.
Planning thresholds and actions
Planning thresholds sit below alert thresholds. An alert tells the operations team that users wait now; a planning threshold tells you that the next server should be decided while the current ones still cope. Read each value at the busy hour, on most working days of a month, not as an average across the day.
| METRIC, BUSY HOUR | EXAMPLE THRESHOLD | ACTION |
|---|---|---|
| Peak requests in flight | above 80 per cent of capacity with one server down | start the decision on the next step |
| KV cache usage | above 0.8 on most working days | check context and cache type, then add capacity |
| Requests waiting | queue on more than 5 busy hours a month | check scheduler limits, then add capacity |
| p95 time to first token | above target, under 95 per cent of requests within it, on more than 5 busy hours a month | compare queue and prefill time; add capacity if queue time grew |
| Preemptions | rising from month to month | more KV cache: FP8 cache after an evaluation, or cards |
| Generated tokens per hour | growth above 20 per cent a quarter | update the forecast and pull the review forward |
| Server power draw | above 80 per cent of the rack feed | plan the feed before the next server |
Thresholds are our example values, not vLLM’s or NVIDIA’s; set your own and write them down. Metric names from vLLM’s metrics documentation (latest, read on 10 October 2026); server power from the BMC or rack PDU, GPU power from DCGM-Exporter.
The alert examples in the monitoring guide fire on a KV cache above 0.9 for 15 minutes, while the planning value here is 0.8 at the busy hour. A queue with room left in the cache points to a scheduler setting, as the monitoring guide explains.
Headroom policy and N+1 capacity
A headroom policy states how much of the capacity the forecast peak may use. Our example policy allows 80 per cent of the capacity left with the largest server down, a margin for a peak factor that is only an estimate and for latency that rises before the cache is full. N+1 means that the remaining servers carry the peak while one is down for a failure or an update, and our guide to LLM high availability with two GPU nodes shows the health checks and failover behind it.
Each server must hold every model of the service on its own, including the embedding model and reranker of a RAG assistant. For the example of 2,000 employees, two servers with five RTX PRO 6000 each keep 95 sessions with one server down. That covers the peak of 80, but 80 is 84 per cent of 95, above the example limit. Two servers with two H200 NVL each keep 110 and stay within it.
Growth path from one to two to four servers
The table follows the example rollout from one department to all 2,000 employees, with one copy of the model per card behind a load balancer and servers of the same build; the alternatives at the last stage keep two servers.
| STAGE | PEAK IN FLIGHT | LAYOUT | ONE SERVER DOWN | WITHIN 80 PER CENT |
|---|---|---|---|---|
| First department, 250 | 10 | 1 server, 2 RTX PRO 6000 | no failover | not N+1 |
| Rollout to 500 | 20 | 2 servers × 2 RTX PRO 6000 | 38 | yes, limit 30 |
| Rollout to 1,000 | 40 | 3 servers × 2 RTX PRO 6000 | 76 | yes, limit 60 |
| All 2,000 | 80 | 4 servers × 2 RTX PRO 6000 | 114 | yes, limit 91 |
| All 2,000, alternative | 80 | 2 servers × 5 RTX PRO 6000 | 95 | no, limit 76 |
| All 2,000, alternative | 80 | 2 servers × 2 H200 NVL | 110 | yes, limit 88 |
Our estimates: peaks from the example values of our company-size guide; sessions of gpt-oss-120b at 32K with a 16-bit KV cache, 19 per RTX PRO 6000 and 55 per H200 NVL, one copy per card; 80 per cent is our example limit; RAG models need room on every server in addition.
More servers of two cards each lose a smaller share of capacity in a failure, a quarter at the last stage, and each purchase is a small step. They also need more rack positions, power feeds, switch ports and nodes to patch. Fewer, larger servers keep the node count low, but each step is larger, and with two servers half of all cards are reserve at the peak. Growing a server from two cards to four or eight works only if it was ordered for the final count, as our guide to planning a GPU server for growth explains, so decide the path at the first order.
We build AI servers with 2 to 8 GPUs per node and supply each added server on one EU contract and invoice. Send us your busy-hour figures and rollout forecast through the form below, and we size the next step from them.
Step changes from model decisions, not only from users
Users add load gradually, while a model decision changes it on the day it goes live. Each such change belongs in the forecast with its planned date.
| GROWTH TRIGGER | CAPACITY CHANGE | SOURCE |
|---|---|---|
| More employees, same use | peak in flight grows in proportion, 20 for 500 to 80 for 2,000 | company-size example values |
| Declared context 32K to 128K | gpt-oss-120b sessions per RTX PRO 6000 from 19 to 4, if requests fill it | company-size guide |
| 8,000-token reasoning | time in flight from 30 to 444 s; 500 employees from 2 to between 5 and 16 RTX PRO 6000, before N+1 | reasoning guide |
| Agents instead of chat | 102,400 input tokens per 8-call task against 3,500 for a chat turn, load shifts to prefill; time per step sets the cards | agents guide, assumed values |
| Larger model | fewer copies per card, or one copy split across cards | model hardware guides |
| RAG added | embedding model, reranker and longer prompts on every server | company-size guide |
Our earlier estimates and illustrative examples, linked in the text; none is a measurement, and a pilot or load test with your prompts replaces each one.
Our guide to reasoning model hardware derives the reasoning row from time in flight, and our guide to GPU sizing for AI agents shows why time per step, more than memory, sets the hardware for agents. Before such a change goes live, repeat the load test and recount the capacity per server.
Quarterly capacity review and the purchase decision
Kubernetes describes its Horizontal Pod Autoscaler in these words: “Horizontal scaling means that the response to increased load is to deploy more Pods.” On a fixed set of GPU servers, a new vLLM pod needs a GPU, or MIG instance, of the right type that no other pod holds, so autoscaling can shift GPUs between models but adds none. The next server is a purchase decision, and a quarterly review is a workable rhythm for it.
- Export the busy-hour rows of the quarter and compare them with the thresholds in the first table.
- Recount capacity with the largest server down, from the last load test and the current engine version.
- Update the forecast with the business: new departments, new use cases and planned changes of model, context or reasoning level.
- Project the peak quarter by quarter for the next year and mark the first quarter that crosses the headroom rule.
- If a quarter crosses it, choose the step (cards in servers ordered for them, another server, or a different card) and start the order together with the rack, power and network checks.
- Record the thresholds, the decision and its reason, so the next review starts from them.
Take the decision at the review where the forecast first crosses the rule, because the new server also needs a rack position, a power feed, network ports, installation and a load test before it carries traffic.
We give you a configuration and quote within one business day and check the rack, power and airflow before we quote. Describe the step your last review pointed to in the form below.
What we supply
We build AI servers to order with 2 to 8 GPUs per node, sized by model size and concurrent users, assembled and burn-in tested, with manufacturer warranty on every component and on one EU contract and invoice. The cards include the RTX PRO 6000 Server Edition, the H200 NVL with NVLink bridges, the L40S and the L4 from our range of professional NVIDIA GPUs, and NVIDIA AI Enterprise and vGPU licences come on the same invoice. For an existing server we check the platform, power and cooling before you add cards. What runs on top, the private LLMs, RAG and MLOps, is our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
What is GPU capacity planning for LLMs?
When should I add another GPU server for LLM inference?
How do you plan LLM inference capacity?
How much headroom should an LLM platform keep?
Is GPU utilisation a good metric for LLM capacity planning?
Does Kubernetes autoscaling scale LLM inference on GPU servers?
Send us your busy-hour figures (peak requests in flight, KV cache usage, time to first token), the model and declared context, the number of servers and cards today and your rollout forecast by quarter. We reply within one business day with a configuration and a quote for the next step, with the rack, power and airflow checked before we quote.
Talk to an expertWe reply within one business day