BLOG · GUIDE ·

GPU capacity planning for an LLM platform: when to add the next server and how to decide it

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • Plan LLM capacity from busy-hour measurements, not daily averages: peak requests in flight, KV cache usage, queue time and 95th-percentile time to first token, set against thresholds you write down, plus an adoption forecast by quarter
  • For planning, a server’s capacity is the lower of two figures: the conversations its KV cache holds at the declared context, and the concurrency at which a load test still meets the latency target
  • With N+1, count capacity with one server down: for a peak of 80 requests in flight of gpt-oss-120b at 32K, two servers with five RTX PRO 6000 each keep 95 sessions, so the peak uses 84 per cent of them, above our example limit of 80 per cent, while four servers with two cards each keep 114 (our estimates)
  • Model decisions move the load in steps: by our earlier estimates, 8,000-token reasoning raises the cards for 500 employees from 2 to between 5 and 16 RTX PRO 6000, and a declared 128K context cuts gpt-oss-120b sessions per RTX PRO 6000 from 19 to 4
  • Kubernetes autoscaling deploys more pods, not more GPUs; on-premise, the next server is a purchase, so decide it at a quarterly review once the forecast crosses your threshold, well before an alert fires

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

LLM capacity planning: when to add the next GPU server

In GPU capacity planning for an LLM platform, add the next server when the measured busy-hour load, projected forward by your adoption forecast, crosses a headroom threshold that you set for the capacity left with one server down. The inputs are the platform’s metrics at the busiest hour of each working day (requests and tokens, KV cache usage, queue time and the 95th percentile of time to first token) and a forecast of users and use cases per quarter. Model decisions such as a larger model, a longer declared context, reasoning or agents change the load in steps, often by more than a new department does. On-premise capacity is a purchase, so the decision belongs in a quarterly review and should be taken early.

The worked example in this article is a platform for 2,000 employees running gpt-oss-120b at a declared 32K context with a 16-bit KV cache. It uses the example values of our guide to private ChatGPT servers by company size: 40 per cent of employees active in the busiest hour, 6 requests an hour each, 30 seconds per request and a peak factor of 2. With these values the peak of requests in flight is 4 per cent of the employees who have access, 80 for all 2,000 by our arithmetic.

Inputs: busy-hour metrics and an adoption forecast

vLLM’s documentation states that “These metrics are exposed via the /metrics endpoint on the vLLM OpenAI compatible API server.” Capacity planning uses the gauges vllm:num_requests_running, vllm:num_requests_waiting and vllm:kv_cache_usage_perc, the histograms for time to first token and request queue time, and the token counters vllm:prompt_tokens and vllm:generation_tokens, which Prometheus shows with the suffix _total. What each metric measures, and the alerts we suggest on it, are in our guide to LLM serving metrics in Prometheus. Keep one row per working day with the busy-hour peak of running plus waiting requests, the highest KV cache usage, the share of requests that waited, the 95th percentile of time to first token at the gateway and the tokens generated, for at least a year.

DCGM-Exporter, in NVIDIA’s documentation updated on 27 May 2026, “exposes GPU metrics at an HTTP endpoint (/metrics) for monitoring solutions such as Prometheus”, on port 9400 by default. Its power draw and temperatures are per GPU. The whole server’s draw, processors and fans included, comes from its BMC or the rack PDU, and both show how much of the rack feed and cooling the next server can still use. Memory in use is no load signal under vLLM, whose --gpu-memory-utilization sets “The fraction of GPU memory to be used for the model executor”, 0.92 by default. The engine claims its share at startup, and the cache inside that share fills and empties with the traffic. GPU utilisation is, in nvidia-smi’s documentation, the “Percent of time over the past sample period during which one or more kernels was executing on the GPU”, whatever share of the GPU those kernels use.

The forecast comes from the business. For each use case, write down per quarter the employees who get access, the share you expect at the busiest hour, requests per hour, tokens per request, the model and the declared context.

Capacity per server: memory and latency

For planning, a server’s capacity is the lower of two figures. One is the number of conversations its KV cache holds at the declared context beside the weights, which vLLM prints at startup and which our estimates put at 19 per RTX PRO 6000 and 55 per H200 NVL for the example model. The other is the number of requests in flight at which a load test with your own prompts still meets the targets for time to first token and tokens per second per user.

NVIDIA’s NIM benchmarking guide, updated on 20 July 2026, describes why the second figure is often lower: “As the number of concurrent requests increases, total system TPS increases while TPS per user decreases as latency increases.” Total throughput rises until the server “saturates the available GPU compute resources”, and the guide adds that “Beyond that point, TPS can decrease.” Run the load test on the GPU type production uses, and again after each engine upgrade, model change or change of context length.

Planning thresholds and actions

Planning thresholds sit below alert thresholds. An alert tells the operations team that users wait now; a planning threshold tells you that the next server should be decided while the current ones still cope. Read each value at the busy hour, on most working days of a month, not as an average across the day.

METRIC, BUSY HOUREXAMPLE THRESHOLDACTION
Peak requests in flightabove 80 per cent of capacity with one server downstart the decision on the next step
KV cache usageabove 0.8 on most working dayscheck context and cache type, then add capacity
Requests waitingqueue on more than 5 busy hours a monthcheck scheduler limits, then add capacity
p95 time to first tokenabove target, under 95 per cent of requests within it, on more than 5 busy hours a monthcompare queue and prefill time; add capacity if queue time grew
Preemptionsrising from month to monthmore KV cache: FP8 cache after an evaluation, or cards
Generated tokens per hourgrowth above 20 per cent a quarterupdate the forecast and pull the review forward
Server power drawabove 80 per cent of the rack feedplan the feed before the next server

Thresholds are our example values, not vLLM’s or NVIDIA’s; set your own and write them down. Metric names from vLLM’s metrics documentation (latest, read on 10 October 2026); server power from the BMC or rack PDU, GPU power from DCGM-Exporter.

The alert examples in the monitoring guide fire on a KV cache above 0.9 for 15 minutes, while the planning value here is 0.8 at the busy hour. A queue with room left in the cache points to a scheduler setting, as the monitoring guide explains.

Headroom policy and N+1 capacity

A headroom policy states how much of the capacity the forecast peak may use. Our example policy allows 80 per cent of the capacity left with the largest server down, a margin for a peak factor that is only an estimate and for latency that rises before the cache is full. N+1 means that the remaining servers carry the peak while one is down for a failure or an update, and our guide to LLM high availability with two GPU nodes shows the health checks and failover behind it.

Each server must hold every model of the service on its own, including the embedding model and reranker of a RAG assistant. For the example of 2,000 employees, two servers with five RTX PRO 6000 each keep 95 sessions with one server down. That covers the peak of 80, but 80 is 84 per cent of 95, above the example limit. Two servers with two H200 NVL each keep 110 and stay within it.

Growth path from one to two to four servers

The table follows the example rollout from one department to all 2,000 employees, with one copy of the model per card behind a load balancer and servers of the same build; the alternatives at the last stage keep two servers.

STAGEPEAK IN FLIGHTLAYOUTONE SERVER DOWNWITHIN 80 PER CENT
First department, 250101 server, 2 RTX PRO 6000no failovernot N+1
Rollout to 500202 servers × 2 RTX PRO 600038yes, limit 30
Rollout to 1,000403 servers × 2 RTX PRO 600076yes, limit 60
All 2,000804 servers × 2 RTX PRO 6000114yes, limit 91
All 2,000, alternative802 servers × 5 RTX PRO 600095no, limit 76
All 2,000, alternative802 servers × 2 H200 NVL110yes, limit 88

Our estimates: peaks from the example values of our company-size guide; sessions of gpt-oss-120b at 32K with a 16-bit KV cache, 19 per RTX PRO 6000 and 55 per H200 NVL, one copy per card; 80 per cent is our example limit; RAG models need room on every server in addition.

More servers of two cards each lose a smaller share of capacity in a failure, a quarter at the last stage, and each purchase is a small step. They also need more rack positions, power feeds, switch ports and nodes to patch. Fewer, larger servers keep the node count low, but each step is larger, and with two servers half of all cards are reserve at the peak. Growing a server from two cards to four or eight works only if it was ordered for the final count, as our guide to planning a GPU server for growth explains, so decide the path at the first order.

We build AI servers with 2 to 8 GPUs per node and supply each added server on one EU contract and invoice. Send us your busy-hour figures and rollout forecast through the form below, and we size the next step from them.

Step changes from model decisions, not only from users

Users add load gradually, while a model decision changes it on the day it goes live. Each such change belongs in the forecast with its planned date.

GROWTH TRIGGERCAPACITY CHANGESOURCE
More employees, same usepeak in flight grows in proportion, 20 for 500 to 80 for 2,000company-size example values
Declared context 32K to 128Kgpt-oss-120b sessions per RTX PRO 6000 from 19 to 4, if requests fill itcompany-size guide
8,000-token reasoningtime in flight from 30 to 444 s; 500 employees from 2 to between 5 and 16 RTX PRO 6000, before N+1reasoning guide
Agents instead of chat102,400 input tokens per 8-call task against 3,500 for a chat turn, load shifts to prefill; time per step sets the cardsagents guide, assumed values
Larger modelfewer copies per card, or one copy split across cardsmodel hardware guides
RAG addedembedding model, reranker and longer prompts on every servercompany-size guide

Our earlier estimates and illustrative examples, linked in the text; none is a measurement, and a pilot or load test with your prompts replaces each one.

Our guide to reasoning model hardware derives the reasoning row from time in flight, and our guide to GPU sizing for AI agents shows why time per step, more than memory, sets the hardware for agents. Before such a change goes live, repeat the load test and recount the capacity per server.

Quarterly capacity review and the purchase decision

Kubernetes describes its Horizontal Pod Autoscaler in these words: “Horizontal scaling means that the response to increased load is to deploy more Pods.” On a fixed set of GPU servers, a new vLLM pod needs a GPU, or MIG instance, of the right type that no other pod holds, so autoscaling can shift GPUs between models but adds none. The next server is a purchase decision, and a quarterly review is a workable rhythm for it.

  1. Export the busy-hour rows of the quarter and compare them with the thresholds in the first table.
  2. Recount capacity with the largest server down, from the last load test and the current engine version.
  3. Update the forecast with the business: new departments, new use cases and planned changes of model, context or reasoning level.
  4. Project the peak quarter by quarter for the next year and mark the first quarter that crosses the headroom rule.
  5. If a quarter crosses it, choose the step (cards in servers ordered for them, another server, or a different card) and start the order together with the rack, power and network checks.
  6. Record the thresholds, the decision and its reason, so the next review starts from them.

Take the decision at the review where the forecast first crosses the rule, because the new server also needs a rack position, a power feed, network ports, installation and a load test before it carries traffic.

We give you a configuration and quote within one business day and check the rack, power and airflow before we quote. Describe the step your last review pointed to in the form below.

What we supply

We build AI servers to order with 2 to 8 GPUs per node, sized by model size and concurrent users, assembled and burn-in tested, with manufacturer warranty on every component and on one EU contract and invoice. The cards include the RTX PRO 6000 Server Edition, the H200 NVL with NVLink bridges, the L40S and the L4 from our range of professional NVIDIA GPUs, and NVIDIA AI Enterprise and vGPU licences come on the same invoice. For an existing server we check the platform, power and cooling before you add cards. What runs on top, the private LLMs, RAG and MLOps, is our Private AI/ML service, with engineering by our partner Vixen.UNO.

FAQ

What is GPU capacity planning for LLMs?
It is the regular comparison of measured busy-hour load on an LLM platform with the capacity of its GPU servers, projected forward with a forecast of users and use cases. The load comes from serving metrics such as requests in flight, KV cache usage, queue time and time to first token. The result is a decision, made in advance, on when to add cards or the next server.
When should I add another GPU server for LLM inference?
Add it when the forecast peak of requests in flight crosses the headroom you set for the capacity left with one server down, for example 80 per cent. Take the decision at a quarterly review when the forecast first crosses that line, not when an alert fires, because a new server also needs a rack position, power, network ports and a load test before it carries traffic.
How do you plan LLM inference capacity?
Record the busy-hour peak of running and waiting requests, KV cache usage, queue time, the 95th percentile of time to first token and tokens generated on each working day. Set them against a capacity per server, the lower of the conversations its KV cache holds and the concurrency at which a load test meets your latency target. Then project the peak with a forecast that includes planned model, context and reasoning changes.
How much headroom should an LLM platform keep?
There is no published standard, so set a figure and write it down; our example policy lets the forecast peak use 80 per cent of the capacity left with the largest server down. For 80 requests in flight of gpt-oss-120b at 32K, two servers with five RTX PRO 6000 each keep 95 sessions with one down, so the peak uses 84 per cent, above that rule, while two servers with two H200 NVL each keep 110 by our estimates.
Is GPU utilisation a good metric for LLM capacity planning?
It is a weak one, because nvidia-smi defines GPU utilisation as the share of time during which one or more kernels run, whatever share of the GPU they use. Memory in use does not help either, since vLLM reserves its share of GPU memory at startup, 0.92 by default. Requests in flight, KV cache usage, queue time and time to first token from the serving engine show the load.
Does Kubernetes autoscaling scale LLM inference on GPU servers?
It scales only within the GPUs you have. The Horizontal Pod Autoscaler deploys more pods, and on a fixed set of GPU servers each new vLLM pod needs a GPU or MIG instance of the right type that no other pod holds, so it can shift GPUs between models but adds none. More capacity on-premise means adding cards or servers, the purchase that a quarterly capacity review prepares.

Send us your busy-hour figures (peak requests in flight, KV cache usage, time to first token), the model and declared context, the number of servers and cards today and your rollout forecast by quarter. We reply within one business day with a configuration and a quote for the next step, with the rack, power and airflow checked before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna