LLM high availability on-premise: two GPU nodes, failover and maintenance windows
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- An on-premise LLM service survives the loss of a node when each GPU node runs every model of the service on its own and carries the whole peak, and a load balancer or gateway sends requests only to replicas that pass their health check
- Sizing is N+1 at the peak: for the 80 requests in flight of our 2,000-employee example with gpt-oss-120b at 32K, two nodes need 5 RTX PRO 6000 or 2 H200 NVL each, three nodes 3 RTX PRO 6000 or 1 H200 NVL each
- vLLM 0.31.0 answers /health with 200 while its engine runs and 503 once the engine is dead; NIM for LLMs separates /v1/health/live from /v1/health/ready, which returns 200 only when the model is loaded
- A failover loses the requests in flight on the failed node and its prefix cache; NGINX by default does not retry a POST already sent upstream and cannot resume a response once part of it has reached the client
- On GPU nodes without a spare card, the default rolling update of a two-replica Deployment stalls, because a 25 per cent surge rounds up to one extra pod; set maxSurge to 0, update drivers node by node and run etcd on three members
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
How an on-premise LLM stays available when a node fails
LLM high availability on-premise means that the service stays up when one GPU server fails. That holds when every node can run each model of the service on its own and still carry the whole peak, and a load balancer or gateway sends requests only to replicas that pass a health check. With two redundant nodes that means N+1 sizing, in which either server alone holds the peak number of requests in flight. Drivers and models are then updated one node at a time in a maintenance window. A failover still costs the requests running on the failed node.
The sizing arithmetic comes from our guide to private ChatGPT servers by company size, and the platform layers from our guide to a private LLM platform on Kubernetes.
Sizing two GPU nodes for N+1 at the peak
Each node must hold every model of the service: a RAG assistant also needs its embedding model, its reranker and the vector index on the surviving node. In our company-size guide, gpt-oss-120b at a declared 32K context with a 16-bit KV cache leaves room for about 19 conversations on one RTX PRO 6000 and about 55 on one H200 NVL, and its example of 2,000 employees produces a peak of 80 requests in flight.
| LAYOUT | CARDS, PEAK 80 | ONE NODE DOWN | DURING MAINTENANCE | ETCD MEMBERS |
|---|---|---|---|---|
| One node | 5 × RTX PRO 6000 or 2 × H200 NVL | service down | service down for the window | 1, tolerates no failure |
| Two nodes | 5 × RTX PRO 6000 or 2 × H200 NVL per node, 10 or 4 in all | 95 or 110 sessions held | the other node holds the peak, with no failover reserve | 2, still tolerate no failure: add a third control-plane node |
| Three nodes | 3 × RTX PRO 6000 or 1 × H200 NVL per node, 9 or 3 in all | 114 or 110 sessions held | 114 or 110 held; a second failure leaves 57 or 55 | 3, tolerate one failure |
Sessions are conversations of gpt-oss-120b at 32K with a 16-bit KV cache, 19 per RTX PRO 6000 and 55 per H200 NVL, our estimates from the company-size guide; the peak of 80 rests on its example values. etcd failure tolerance from the etcd v3.6 FAQ.
With two nodes, one node’s GPUs are failover reserve at the busiest hour. With three, each node holds half the peak, which takes 9 RTX PRO 6000 instead of 10, or 3 H200 NVL instead of 4. During maintenance on one node, both layouts run without failover reserve, so schedule the window outside the peak.
We build AI servers to order with 2 to 8 GPUs per node, sized by model size and concurrent users, and supply both nodes on one EU contract and invoice. Send us your models and peak requests in flight through the form below.
Health checks and load balancing in front of vLLM
vLLM’s online serving documentation lists /health and /metrics among its endpoints. In vLLM 0.31.0, released on 5 October 2026, /health returns 200 when the engine passes its health check and 503 once the engine is dead. By default the API server does not outlive its engine: vLLM’s VLLM_, off by default, keeps the server “alive even after the underlying AsyncLLMEngine errors”. Point the load balancer at /health rather than at the TCP port, and let a liveness probe or the service manager restart the process. NVIDIA NIM for LLMs splits the check in two, with /v1/health/live returning 200 “when the container is running” and /v1/health/ready returning 200 “when the model is loaded and inference is available”.
A health check shows whether the engine runs, not whether it keeps up; queue, KV cache and latency come from the metrics in our guide to LLM serving monitoring.
Without Kubernetes, a reverse proxy on a separate host spreads the requests. Open-source NGINX checks passively: by default one failed attempt (max_fails) within 10 seconds (fail_timeout) marks a server unavailable for 10 seconds. Connection errors, timeouts and invalid headers count as failed attempts, but an HTTP 503 counts only if proxy_next_upstream lists http_503. Active checks with its health_check directive are, in NGINX’s words, “available as part of our commercial subscription”, and their default URI is /, so set it to /health. Chat requests are POST requests, which NGINX does not pass to the next server once they have been sent to an upstream server, unless the non_idempotent option is set. Its documentation adds that “passing a request to the next server is only possible if nothing has been sent to a client yet”. Its proxy_read_timeout of 60 s applies between two reads, so a non-streamed answer that takes longer than a minute ends in a timeout unless you raise it.
The proxy is itself a single point of failure; keepalived, at 2.4.3 of July 2026, runs two proxies on one virtual IP address with VRRP.
What a failover loses: requests in flight and cached prefixes
When a node fails, the requests running on it are lost, and a streamed answer stops with an error. The conversation survives, because a request to the OpenAI-compatible chat API carries the messages of the conversation, and the chat front end keeps them in its own database. The user regenerates the answer, and the next request goes to the surviving node.
That node starts without the lost node’s cached prefixes. vLLM’s automatic prefix caching “caches the KV cache of existing queries” for reuse by a new query with the same prefix, in the GPU memory of one server. After a failover, each moved conversation is processed again from its first token, so time to first token rises until the cache has refilled, at the moment the survivor takes the whole load.
The chat front end and its database, the gateway and the vector database each need a second instance too. Keep the model weights on local NVMe in each node, with the revision pinned and a checksum manifest, so that a failed file server cannot stop a replica from starting. What each of these holds and how to restore it is in our guide to backing up an AI server.
Failures and changes: what happens and the design answer
| EVENT | WHAT HAPPENS | DESIGN ANSWER |
|---|---|---|
| Whole node lost | requests in flight on it end in an error; the balancer removes it after its check fails | each node carries the peak alone; separate power feeds and switch ports |
| GPU off the bus, Xid 79 | the vLLM engine on that card fails | the health check removes the replica; NVIDIA’s triage guide says to drain the node |
| Engine dead | /health answers 503, and by default the API server then exits | health checks and liveness probes on /health, not on the port |
| Restart after a crash | the replica serves again only after its weights have loaded | a startup probe that covers the load time; weights on local NVMe |
| Load balancer host lost | no traffic reaches either healthy node | two proxies sharing one virtual IP address with VRRP |
| Driver or vLLM update | the node’s replica stops for the update | one node at a time, in a window, after checking the other node’s capacity |
| Two etcd members, one lost | etcd has no majority and the cluster cannot store changes | three control-plane members, which can be small servers without GPUs |
vLLM 0.31.0 source code (/health, environment variables); NVIDIA’s Xid catalogue and triage guide as summarised in our DCGM guide; Kubernetes documentation v1.37; etcd v3.6 FAQ; keepalived project page.
Alert per node, because a pool with one failed node answers normally until the next peak. Compare the peak of running plus waiting requests, summed over both nodes, with what one node holds alone: when the sum approaches that figure, the failover reserve is gone. Our guide to GPU server monitoring with DCGM lists the XIDs that call for a drain.
Kubernetes: two replicas, probes and rolling updates without a spare GPU
On Kubernetes, give the model’s Deployment two replicas and a required pod anti-affinity with the topology key kubernetes.io/hostname, which in the Kubernetes documentation’s example “tells the scheduler to avoid placing multiple replicas” on one node. The Kubernetes example in vLLM’s latest documentation probes /health on port 8000 for liveness and readiness, each with an initial delay of 60 seconds. A pod whose readiness probe fails “will not receive traffic from any services”, and a failing liveness probe makes the kubelet restart the container. vLLM’s page warns about a failureThreshold “too low for the time needed to start up the server”; add a startup probe whose failureThreshold × periodSeconds covers the slowest model load you have timed.
Rolling updates need care on nodes whose GPUs are all in use. A Deployment defaults to 25 per cent for maxSurge and maxUnavailable, with maxSurge rounded up and maxUnavailable rounded down, so with two replicas Kubernetes creates one extra pod first and removes none. That pod needs an unused GPU and, with the anti-affinity above, a third node; without them it stays Pending. The old replicas keep serving while the rollout stalls. Setting maxSurge to 0 and maxUnavailable to 1 replaces one pod at a time on the GPU it releases. In its Standard deployment mode, KServe 0.20.0, the current release as of October 2026, calls the setting with maxSurge at 0 its ResourceAware mode; its Availability mode, with maxUnavailable at 0, needs the spare GPU.
NVIDIA’s Helm chart for NIM for LLMs deploys a StatefulSet by default. Kubernetes updates it one pod at a time, deleting each old pod before its replacement starts and waiting until that is “Running and Ready”, so no spare GPU is needed. For the chart’s Deployment mode, its documentation suggests Recreate when the cluster cannot spare an extra GPU pod, but Recreate stops all pods before new ones start, which takes a two-replica service down, so keep the StatefulSet default.
On a drain or an update, the kubelet sends SIGTERM and, after terminationGracePeriodSeconds, 30 seconds by default, SIGKILL. In vLLM 0.31.0, --shutdown-timeout defaults to 0, which its help text explains as “0 = abort, >0 = wait”. The documentation says no more about requests in flight, so set it to cover a long answer, set the grace period above it and test a drain under load. A PodDisruptionBudget with minAvailable 1 lets kubectl drain evict one replica at a time, but it does not limit a Deployment’s rolling update, and involuntary disruptions, in Kubernetes’ words, “cannot be prevented by PDBs”. Two GPU nodes alone do not make a highly available cluster: etcd’s FAQ gives a two-member cluster a failure tolerance of 0 and recommends “an odd number of members”, so run three control-plane members.
Driver and model updates node by node
The NVIDIA GPU Operator, at 26.7.1 the newest release in its notes as of October 2026, updates containerised drivers through an upgrade controller that is enabled by default. With no upgrade policy set it upgrades one node at a time, cordons it, deletes the pods that hold GPUs before it reloads the driver, and leaves draining off, in NVIDIA’s words “By default, drain is disabled”. Drivers installed on the host are outside its scope, and our guide to NVIDIA driver branches and CUDA versions covers which branch to choose. By hand, the order is the same.
- Check in the monitoring that the other node alone held the peak of recent working days.
- Take the node out of the load balancer and wait until
vllm:num_requests_runningreaches 0; on Kubernetes, cordon and drain it. - Update the driver, the images or vLLM, reboot where the driver needs it, and confirm in nvidia-smi and DCGM that every GPU is back.
- Run your fixed evaluation set against the node directly, before it receives traffic again.
- Return the node to the pool, watch it through the next peak hour, then repeat on the second node.
A new model version follows the same order; keep the previous weights on each node’s local disk, so that a rollback needs no download.
Our Private AI/ML service, with engineering by our partner Vixen.UNO, builds a Kubernetes-based platform with KServe and deploys models with vLLM, with ongoing support under an agreed SLA. Describe your set-up and maintenance windows in the form below.
What we supply
We supply both GPU nodes as AI servers built to order with 2 to 8 GPUs per node, with RTX PRO 6000 Server Edition or H200 NVL cards for the language model and the L4 or L40S for embedding and reranking models. They are load-tested before shipment, with manufacturer warranty on every component, on one EU contract and invoice. We check the rack, power and airflow before we quote, and the configuration and quote follow within one business day. The platform on top, with a query log and a team trained to run it, is our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
How do you make an on-premise LLM highly available?
Does vLLM support high availability?
How many GPU nodes do I need for LLM failover?
What happens to running chats when an LLM server fails?
How do I update an LLM without downtime?
Can a two-node Kubernetes cluster be highly available?
Send us the models you serve, the peak number of requests in flight, how long the service may be unavailable and whether it runs on Kubernetes. We reply within one business day with a configuration and quote for both nodes. The first call about the platform is free of charge.
Talk to an expertWe reply within one business day