BLOG · GUIDE ·

LLM high availability on-premise: two GPU nodes, failover and maintenance windows

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • An on-premise LLM service survives the loss of a node when each GPU node runs every model of the service on its own and carries the whole peak, and a load balancer or gateway sends requests only to replicas that pass their health check
  • Sizing is N+1 at the peak: for the 80 requests in flight of our 2,000-employee example with gpt-oss-120b at 32K, two nodes need 5 RTX PRO 6000 or 2 H200 NVL each, three nodes 3 RTX PRO 6000 or 1 H200 NVL each
  • vLLM 0.31.0 answers /health with 200 while its engine runs and 503 once the engine is dead; NIM for LLMs separates /v1/health/live from /v1/health/ready, which returns 200 only when the model is loaded
  • A failover loses the requests in flight on the failed node and its prefix cache; NGINX by default does not retry a POST already sent upstream and cannot resume a response once part of it has reached the client
  • On GPU nodes without a spare card, the default rolling update of a two-replica Deployment stalls, because a 25 per cent surge rounds up to one extra pod; set maxSurge to 0, update drivers node by node and run etcd on three members

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

How an on-premise LLM stays available when a node fails

LLM high availability on-premise means that the service stays up when one GPU server fails. That holds when every node can run each model of the service on its own and still carry the whole peak, and a load balancer or gateway sends requests only to replicas that pass a health check. With two redundant nodes that means N+1 sizing, in which either server alone holds the peak number of requests in flight. Drivers and models are then updated one node at a time in a maintenance window. A failover still costs the requests running on the failed node.

The sizing arithmetic comes from our guide to private ChatGPT servers by company size, and the platform layers from our guide to a private LLM platform on Kubernetes.

Sizing two GPU nodes for N+1 at the peak

Each node must hold every model of the service: a RAG assistant also needs its embedding model, its reranker and the vector index on the surviving node. In our company-size guide, gpt-oss-120b at a declared 32K context with a 16-bit KV cache leaves room for about 19 conversations on one RTX PRO 6000 and about 55 on one H200 NVL, and its example of 2,000 employees produces a peak of 80 requests in flight.

LAYOUTCARDS, PEAK 80ONE NODE DOWNDURING MAINTENANCEETCD MEMBERS
One node5 × RTX PRO 6000 or 2 × H200 NVLservice downservice down for the window1, tolerates no failure
Two nodes5 × RTX PRO 6000 or 2 × H200 NVL per node, 10 or 4 in all95 or 110 sessions heldthe other node holds the peak, with no failover reserve2, still tolerate no failure: add a third control-plane node
Three nodes3 × RTX PRO 6000 or 1 × H200 NVL per node, 9 or 3 in all114 or 110 sessions held114 or 110 held; a second failure leaves 57 or 553, tolerate one failure

Sessions are conversations of gpt-oss-120b at 32K with a 16-bit KV cache, 19 per RTX PRO 6000 and 55 per H200 NVL, our estimates from the company-size guide; the peak of 80 rests on its example values. etcd failure tolerance from the etcd v3.6 FAQ.

With two nodes, one node’s GPUs are failover reserve at the busiest hour. With three, each node holds half the peak, which takes 9 RTX PRO 6000 instead of 10, or 3 H200 NVL instead of 4. During maintenance on one node, both layouts run without failover reserve, so schedule the window outside the peak.

We build AI servers to order with 2 to 8 GPUs per node, sized by model size and concurrent users, and supply both nodes on one EU contract and invoice. Send us your models and peak requests in flight through the form below.

Health checks and load balancing in front of vLLM

vLLM’s online serving documentation lists /health and /metrics among its endpoints. In vLLM 0.31.0, released on 5 October 2026, /health returns 200 when the engine passes its health check and 503 once the engine is dead. By default the API server does not outlive its engine: vLLM’s VLLM_KEEP_ALIVE_ON_ENGINE_DEATH, off by default, keeps the server “alive even after the underlying AsyncLLMEngine errors”. Point the load balancer at /health rather than at the TCP port, and let a liveness probe or the service manager restart the process. NVIDIA NIM for LLMs splits the check in two, with /v1/health/live returning 200 “when the container is running” and /v1/health/ready returning 200 “when the model is loaded and inference is available”.

A health check shows whether the engine runs, not whether it keeps up; queue, KV cache and latency come from the metrics in our guide to LLM serving monitoring.

Without Kubernetes, a reverse proxy on a separate host spreads the requests. Open-source NGINX checks passively: by default one failed attempt (max_fails) within 10 seconds (fail_timeout) marks a server unavailable for 10 seconds. Connection errors, timeouts and invalid headers count as failed attempts, but an HTTP 503 counts only if proxy_next_upstream lists http_503. Active checks with its health_check directive are, in NGINX’s words, “available as part of our commercial subscription”, and their default URI is /, so set it to /health. Chat requests are POST requests, which NGINX does not pass to the next server once they have been sent to an upstream server, unless the non_idempotent option is set. Its documentation adds that “passing a request to the next server is only possible if nothing has been sent to a client yet”. Its proxy_read_timeout of 60 s applies between two reads, so a non-streamed answer that takes longer than a minute ends in a timeout unless you raise it.

The proxy is itself a single point of failure; keepalived, at 2.4.3 of July 2026, runs two proxies on one virtual IP address with VRRP.

What a failover loses: requests in flight and cached prefixes

When a node fails, the requests running on it are lost, and a streamed answer stops with an error. The conversation survives, because a request to the OpenAI-compatible chat API carries the messages of the conversation, and the chat front end keeps them in its own database. The user regenerates the answer, and the next request goes to the surviving node.

That node starts without the lost node’s cached prefixes. vLLM’s automatic prefix caching “caches the KV cache of existing queries” for reuse by a new query with the same prefix, in the GPU memory of one server. After a failover, each moved conversation is processed again from its first token, so time to first token rises until the cache has refilled, at the moment the survivor takes the whole load.

The chat front end and its database, the gateway and the vector database each need a second instance too. Keep the model weights on local NVMe in each node, with the revision pinned and a checksum manifest, so that a failed file server cannot stop a replica from starting. What each of these holds and how to restore it is in our guide to backing up an AI server.

Failures and changes: what happens and the design answer

EVENTWHAT HAPPENSDESIGN ANSWER
Whole node lostrequests in flight on it end in an error; the balancer removes it after its check failseach node carries the peak alone; separate power feeds and switch ports
GPU off the bus, Xid 79the vLLM engine on that card failsthe health check removes the replica; NVIDIA’s triage guide says to drain the node
Engine dead/health answers 503, and by default the API server then exitshealth checks and liveness probes on /health, not on the port
Restart after a crashthe replica serves again only after its weights have loadeda startup probe that covers the load time; weights on local NVMe
Load balancer host lostno traffic reaches either healthy nodetwo proxies sharing one virtual IP address with VRRP
Driver or vLLM updatethe node’s replica stops for the updateone node at a time, in a window, after checking the other node’s capacity
Two etcd members, one lostetcd has no majority and the cluster cannot store changesthree control-plane members, which can be small servers without GPUs

vLLM 0.31.0 source code (/health, environment variables); NVIDIA’s Xid catalogue and triage guide as summarised in our DCGM guide; Kubernetes documentation v1.37; etcd v3.6 FAQ; keepalived project page.

Alert per node, because a pool with one failed node answers normally until the next peak. Compare the peak of running plus waiting requests, summed over both nodes, with what one node holds alone: when the sum approaches that figure, the failover reserve is gone. Our guide to GPU server monitoring with DCGM lists the XIDs that call for a drain.

Kubernetes: two replicas, probes and rolling updates without a spare GPU

On Kubernetes, give the model’s Deployment two replicas and a required pod anti-affinity with the topology key kubernetes.io/hostname, which in the Kubernetes documentation’s example “tells the scheduler to avoid placing multiple replicas” on one node. The Kubernetes example in vLLM’s latest documentation probes /health on port 8000 for liveness and readiness, each with an initial delay of 60 seconds. A pod whose readiness probe fails “will not receive traffic from any services”, and a failing liveness probe makes the kubelet restart the container. vLLM’s page warns about a failureThreshold “too low for the time needed to start up the server”; add a startup probe whose failureThreshold × periodSeconds covers the slowest model load you have timed.

Rolling updates need care on nodes whose GPUs are all in use. A Deployment defaults to 25 per cent for maxSurge and maxUnavailable, with maxSurge rounded up and maxUnavailable rounded down, so with two replicas Kubernetes creates one extra pod first and removes none. That pod needs an unused GPU and, with the anti-affinity above, a third node; without them it stays Pending. The old replicas keep serving while the rollout stalls. Setting maxSurge to 0 and maxUnavailable to 1 replaces one pod at a time on the GPU it releases. In its Standard deployment mode, KServe 0.20.0, the current release as of October 2026, calls the setting with maxSurge at 0 its ResourceAware mode; its Availability mode, with maxUnavailable at 0, needs the spare GPU.

NVIDIA’s Helm chart for NIM for LLMs deploys a StatefulSet by default. Kubernetes updates it one pod at a time, deleting each old pod before its replacement starts and waiting until that is “Running and Ready”, so no spare GPU is needed. For the chart’s Deployment mode, its documentation suggests Recreate when the cluster cannot spare an extra GPU pod, but Recreate stops all pods before new ones start, which takes a two-replica service down, so keep the StatefulSet default.

On a drain or an update, the kubelet sends SIGTERM and, after terminationGracePeriodSeconds, 30 seconds by default, SIGKILL. In vLLM 0.31.0, --shutdown-timeout defaults to 0, which its help text explains as “0 = abort, >0 = wait”. The documentation says no more about requests in flight, so set it to cover a long answer, set the grace period above it and test a drain under load. A PodDisruptionBudget with minAvailable 1 lets kubectl drain evict one replica at a time, but it does not limit a Deployment’s rolling update, and involuntary disruptions, in Kubernetes’ words, “cannot be prevented by PDBs”. Two GPU nodes alone do not make a highly available cluster: etcd’s FAQ gives a two-member cluster a failure tolerance of 0 and recommends “an odd number of members”, so run three control-plane members.

Driver and model updates node by node

The NVIDIA GPU Operator, at 26.7.1 the newest release in its notes as of October 2026, updates containerised drivers through an upgrade controller that is enabled by default. With no upgrade policy set it upgrades one node at a time, cordons it, deletes the pods that hold GPUs before it reloads the driver, and leaves draining off, in NVIDIA’s words “By default, drain is disabled”. Drivers installed on the host are outside its scope, and our guide to NVIDIA driver branches and CUDA versions covers which branch to choose. By hand, the order is the same.

  1. Check in the monitoring that the other node alone held the peak of recent working days.
  2. Take the node out of the load balancer and wait until vllm:num_requests_running reaches 0; on Kubernetes, cordon and drain it.
  3. Update the driver, the images or vLLM, reboot where the driver needs it, and confirm in nvidia-smi and DCGM that every GPU is back.
  4. Run your fixed evaluation set against the node directly, before it receives traffic again.
  5. Return the node to the pool, watch it through the next peak hour, then repeat on the second node.

A new model version follows the same order; keep the previous weights on each node’s local disk, so that a rollback needs no download.

Our Private AI/ML service, with engineering by our partner Vixen.UNO, builds a Kubernetes-based platform with KServe and deploys models with vLLM, with ongoing support under an agreed SLA. Describe your set-up and maintenance windows in the form below.

What we supply

We supply both GPU nodes as AI servers built to order with 2 to 8 GPUs per node, with RTX PRO 6000 Server Edition or H200 NVL cards for the language model and the L4 or L40S for embedding and reranking models. They are load-tested before shipment, with manufacturer warranty on every component, on one EU contract and invoice. We check the rack, power and airflow before we quote, and the configuration and quote follow within one business day. The platform on top, with a query log and a team trained to run it, is our Private AI/ML service, with engineering by our partner Vixen.UNO.

FAQ

How do you make an on-premise LLM highly available?
Run each model of the service on at least two GPU nodes, each able to carry the whole peak of requests in flight on its own, and put a load balancer or a Kubernetes Service in front that sends requests only to replicas passing a health check on vLLM’s /health endpoint. Update drivers and models one node at a time in a maintenance window. The embedding model, reranker, vector database, gateway and chat front end need a second instance too.
Does vLLM support high availability?
vLLM provides what a load balancer needs: in vLLM 0.31.0, /health returns 200 while the engine runs and 503 once it is dead, and /metrics exports Prometheus metrics for queue, KV cache and latency. Failover itself comes from running one replica per node behind a load balancer or a Kubernetes Service that stops routing to a replica whose check fails.
How many GPU nodes do I need for LLM failover?
At least two, each sized for the full peak, or three, each sized for half of it. With our example of 80 requests in flight for gpt-oss-120b at 32K, two nodes need 5 RTX PRO 6000 or 2 H200 NVL each, while three nodes need 3 RTX PRO 6000 or 1 H200 NVL each. A Kubernetes control plane needs three etcd members in either case.
What happens to running chats when an LLM server fails?
Answers being generated on the failed server stop with an error, and a proxy such as NGINX cannot move a response that has already started to another server. The conversation itself survives in the chat front end, which sends the whole history with the next request, so the surviving server answers after processing that history again without the lost prefix cache.
How do I update an LLM without downtime?
Update one node at a time while the other carries the peak: take it out of the load balancer, wait until its running requests reach zero, update, test it with a fixed evaluation set and return it. On Kubernetes without a spare GPU, set maxSurge to 0 and maxUnavailable to 1, because the default 25 per cent surge creates an extra pod that waits for an unused GPU and stalls the rollout.
Can a two-node Kubernetes cluster be highly available?
Not on its own, because etcd needs a majority of its members and a two-member cluster tolerates no failure, according to etcd’s FAQ. Run three control-plane members, which can be small servers or virtual machines without GPUs, and keep the two GPU nodes as workers with one model replica each.

Send us the models you serve, the peak number of requests in flight, how long the service may be unavailable and whether it runs on Kubernetes. We reply within one business day with a configuration and quote for both nodes. The first call about the platform is free of charge.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna