BLOG · GUIDE ·

Disaster recovery for a private AI platform: second site, GPU capacity and failover

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • Decide per AI service, from its business impact analysis, whether it must survive the loss of a site, then what waits at the second site: full GPU capacity, reduced capacity for an agreed minimum service level, or a rebuild from backups
  • Replicate what exists only on the platform (adapters, vector index, document store state, chat history, gateway configuration, API keys and query logs) and rebuild the inference nodes from pinned images and weights kept ready at the second site
  • For a 2,000-employee example with a peak of 80 requests in flight on gpt-oss-120b at 32K, full capacity at the second site is one server with 5 RTX PRO 6000 or 2 H200 NVL; half the peak takes 3 RTX PRO 6000 or 1 H200 NVL
  • vSphere 9.0 lists snapshots as unavailable for DirectPath I/O VMs, so passthrough inference VMs are rebuilt from templates rather than replicated through snapshots, and vGPU for Compute VMs at the second site must reach an NVIDIA licence service
  • A DR test for AI times the switch and also runs the fixed evaluation set, retrieval queries per access level and existing API keys against the recovered stack, with its own pass mark for a smaller model

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

Disaster recovery for an AI platform: what to decide first

Disaster recovery for a private AI platform starts with one decision per AI service: whether it must keep running when the main site is lost, and at what minimum level. That answer sets what waits at the second site. The options are full GPU capacity, reduced capacity with fewer cards, a smaller or quantised model or fewer concurrent users, or no GPUs at all and a rebuild from backups once servers are available again. In every option the data that exists only on the platform is replicated or backed up to the second site, and inference nodes are rebuilt from pinned sources.

This guide covers the loss of a whole site. Surviving the loss of one GPU server at the same site is high availability, covered in our guide to LLM high availability with two GPU nodes, and the backup of weights, adapters and vector stores is in our guide to backing up an AI server.

Recovery tiers for AI services

AI services enter the business impact analysis like any other system, with their process owners. A code assistant for developers, an internal chat assistant, a RAG assistant for customer service and a nightly document extraction job differ in how fast the damage grows when they stop.

ENISA’s technical implementation guidance for NIS2 (version 1.0, June 2025) adds a measure that suits AI services well: the service delivery objective, “the minimum level of performance that needs to be reached” by business functions while they run in the alternate processing mode. For an AI service the minimum level is stated in requests in flight, context length and model quality. Agree it with the process owner before sizing the second site.

A RAG assistant needs its embedding model, reranker, vector index, document connectors, gateway and identity provider at the second site. NIST SP 800-34 Rev. 1 states that “the RTO must normally be shorter than the MTD”, so the time to start every one of these components, in order, has to fit inside the RTO of the AI service.

What to replicate and what to rebuild

The rule for an AI platform is to replicate state and to rebuild compute. Inference nodes hold nothing unique when weights are pinned and configuration lives in git; the databases around them hold data that exists nowhere else.

COMPONENTREPLICATE OR REBUILDMETHOD
Base model weightskeep ready at the second sitecopy each pinned revision when it changes, with a SHA-256 manifest
Adapters, fine-tuned weightsreplicatecopy each version with its base revision and evaluation set
Vector indexreplicate, rebuild as fallbackpgvector: PostgreSQL WAL replication; other stores: scheduled snapshots copied over
Documents, connector statereplicatedatabase replication, or VM replication of the CPU tier
Chat history, user settingsreplicatedatabase replication of the chat front end
Gateway configuration, keysreplicateconfiguration in git; keys in a replicated secrets store
Query and answer logsreplicatelog shipping under the same retention rules
Inference nodes and VMsrebuildimages by digest, deployment manifests from git
Kubernetes resourcesrebuild or restoreGitOps from git; Velero for the rest
Licence service (vGPU)second instancea DLS instance reachable from the second site

Our classification. pgvector README (FAQ “Is replication supported?”), Velero v1.18 documentation, Argo CD documentation and NVIDIA License System user guide, read on 10 October 2026.

For pgvector the route is the database’s own replication: its README answers “Yes, pgvector uses the write-ahead log (WAL), which allows for replication and point-in-time recovery.” Stores without WAL-based replication are copied as snapshots or backups on a schedule, and the newest copy sets the recovery point of the index. Keep the embedding model at the second site at the same revision as at the main site. Vectors from a different embedding model do not match the stored ones, and switching models means re-embedding the whole corpus.

Keep the base weights at the second site before you need them, since a download on the day depends on the hub, on access to a gated model and on the revision still existing. gpt-oss-120b is a 65.3 GB checkpoint, and with several models and versions the model store can reach terabytes, so copy each new revision once, when it is approved, not on every backup run.

GPU capacity at the second site

The worked example comes from our company-size guide for private ChatGPT servers: 2,000 employees produce a peak of 80 requests in flight on gpt-oss-120b at a declared 32K context with a 16-bit KV cache, and one RTX PRO 6000 holds about 19 such conversations, one H200 NVL about 55. The main site in our high-availability guide runs two servers, each with 5 RTX PRO 6000 or 2 H200 NVL, so that either one holds the peak.

DR OPTIONGPUS WAITINGEXAMPLE: PEAK 80RECOVERY CLASS
Full capacityone server sized for the whole peak5 × RTX PRO 6000 or 2 × H200 NVL, plus 1 × L4 or L40Sthe service as before, without failover reserve at the second site
Reduced capacityfewer cards for the agreed minimum level40 in flight: 3 × RTX PRO 6000 or 1 × H200 NVL, plus 1 × L4 or L40Susers served, with queues at the busiest hour
Smaller or quantised modelcards for the smaller modelsized from the smaller model’s weights and cacheanswers of a different quality, tested in advance
Rebuild from backupsnonenonethe service is down until GPU servers are installed; its data is preserved

Sessions at 32K with a 16-bit cache, 19 per RTX PRO 6000 and 55 per H200 NVL, and the L4 for embedding and reranking, from our company-size guide; the L40S where every request reranks 40 passages, as in our platform architecture guide; the minimum level of 40 is an example value. The recovery classes are descriptive, not targets of our service.

Two RTX PRO 6000 hold about 38 conversations, below the example minimum of 40, which is why the reduced option takes three. A smaller model also differs in context length, system prompt behaviour and tool calls, so route to it under a name the process owners know, or vLLM’s --served-model-name can expose it under the existing name, which its documentation describes as “The model name(s) used in the API”. Two sites can also share the everyday load, each holding the agreed minimum level alone, at the cost of running the platform and its updates in both places.

We build AI servers to order with 2 to 8 GPUs per node, for the main site and the second site, on one EU contract and invoice. Tell us the minimum service level for each AI service and where the second site is.

GPU virtual machines: passthrough, vGPU and replication

Where inference runs in vSphere VMs, the GPU mode matters. Broadcom’s vSphere 9.0 documentation, updated on 24 August 2026, lists snapshots among the features that are “unavailable for virtual machines configured with DirectPath”, so tools that replicate a VM through its snapshots cannot copy a running passthrough VM. vSphere Replication works differently: Broadcom’s VCF 9.1 documentation, updated on 10 September 2026, describes an agent that “sends changed blocks in the virtual machine disks from the source site to the target site”. We found no statement on passthrough devices in that documentation, so prove the route with one test replication and recovery before relying on it.

The simpler design keeps the inference VMs out of replication. They are rebuilt at the second site from a template, with the weights on local disks and the configuration from git. Dynamic DirectPath I/O helps here: per Broadcom’s KB 312208, a VM can list valid vendor and device combinations and let ESXi choose a device “based on what is available at the time of VM power-on”, so a template does not depend on one card’s address. A vGPU VM needs a host whose GPU offers its vGPU type. The matching rules for both modes are in our guide to GPUs in vSphere: passthrough or vGPU.

Licences come along as a dependency. VMs with vGPU for Compute check out a licence from the NVIDIA License System, whose on-premises form, in NVIDIA’s words, “is hosted on-premises at a location that is accessible from your private network”. The second site needs a route to such an instance, or an instance of its own with licences from your entitlement allocated to it on the NVIDIA Licensing Portal. NVIDIA adds that after the failure of one instance in a cluster of two, “the remaining instance becomes a single point of failure.”

Kubernetes at the recovery site: GitOps and Velero

On Kubernetes, the second site runs its own cluster rather than a stretched one, because a control plane spread over two sites depends on the link between them. The cluster is built from the same definitions as the main one. Argo CD describes its principle as “Application definitions, configurations, and environments should be declarative and version controlled”, and lists the “Ability to manage and deploy to multiple clusters”. With the model deployments, gateway and front end in git, the second cluster starts the same versions, and the git server itself belongs in the replicated tier.

Velero covers what is not in git, and its v1.18 documentation lists “Migrate cluster resources to other clusters” among its uses. Its migration guide points the Velero instances of both clusters “to the same cloud object storage location”, so that storage must sit outside the main site. A volume snapshot on the main site’s storage is lost with the site, and for moving volume data between clusters Velero suggests “the file system backup or the snapshot data mover”. It also states that “Velero doesn’t support restoring into a cluster with a lower Kubernetes version than where the backup was taken”, so keep both clusters on the same Kubernetes version and upgrade the second site first.

DNS and gateway failover

Users and applications reach the platform through one name, the gateway’s, so the failover of an AI service is mostly a change of where that name points. RFC 1035 defines the TTL of a DNS record as the time it “may be cached before the source of the information should again be consulted”. Set the TTL of the gateway’s record to the delay you accept, well before any disaster, because a lower TTL set on the day reaches clients only after the old one has expired. Serve the zone from name servers at both sites, or the change cannot be published once the main site is down.

The gateway at the second site needs the same certificate name, the same API keys and the same routing rules as the main one, and the identity provider behind it must work at that site.

DR tests that check answer quality

A DR test for an AI service measures the time to a working service against the RTO, like any disaster recovery test, and adds checks that show the AI service came back with the right behaviour.

  1. Start the platform at the second site on a network with no route to production, and record when each component answers.
  2. Check the weights at the second site against the SHA-256 manifest of the main site.
  3. Run the fixed evaluation set through the gateway, against the pass mark agreed for the model that runs there; a smaller model gets its own mark.
  4. Run the fixed retrieval queries, one per access level, and compare the top results with the baseline.
  5. Call the API with existing keys from one application of each kind, and check that the query log records the calls.
  6. Note the age of the restore points used for the vector index, chat history and logs against their RPO.

Repeat the test after each change of model, embedding model, serving engine or GPU type at either site.

Our disaster recovery service replicates virtual machines with Veeam to a recovery site in Baltneta’s Tier-3 data centres in Lithuania and runs scheduled failover tests in an isolated environment. Describe the CPU tier of your AI platform in the form below.

What we supply

We build AI servers to order for both sites, with 2 to 8 GPUs per node: RTX PRO 6000 Server Edition or H200 NVL cards for the language models and the L4 or L40S for embedding and reranking, burn-in tested, with manufacturer warranty, on one EU contract and invoice. NVIDIA AI Enterprise and vGPU licences come on the same invoice, and we check the rack, power and airflow before we quote. Where the second site should be a data centre, our Private AI/ML service lists an EU data centre among its deployment options, Tier-3 in Lithuania, where data stays in the EU on dedicated hardware; the option is chosen at the assessment stage, and engineering is by our partner Vixen.UNO. For the virtual machines around the models, our disaster recovery service provides Veeam replication from a 15-minute interval, scheduled test restores and RPO and RTO fixed in the SLA.

FAQ

How do you do disaster recovery for an AI platform?
Decide per AI service from the business impact analysis whether it must survive the loss of a site and at what minimum level, then size the GPU capacity at the second site for that level. Replicate the state that exists only on the platform, such as adapters, the vector index, chat history, gateway configuration and logs, and rebuild the inference nodes from pinned images, weights kept ready at that site and configuration from git. Test the failover with the fixed evaluation set as well as the clock.
How do you plan LLM disaster recovery?
Plan the language model as stateless compute: keep the pinned weights and container images at the second site and start the serving engine there from the same configuration. The parts that cannot be downloaded again, such as fine-tuned adapters, conversations and query logs, are replicated or backed up to that site. Decide in advance whether the second site runs the same model at full capacity, at reduced capacity or a smaller model with its own accepted answers.
How many GPUs does an AI DR site need?
As many as the agreed minimum service level requires, sized like the main site. In our example of 2,000 employees with a peak of 80 requests in flight on gpt-oss-120b at 32K, full capacity is one server with 5 RTX PRO 6000 or 2 H200 NVL, while half the peak takes 3 RTX PRO 6000 or 1 H200 NVL. Services the business can do without for longer may have no GPUs waiting and be rebuilt once servers are installed.
How do you protect a RAG system against a site failure?
Replicate the vector index, the document store and the connector state to the second site, and keep the same embedding model at the same revision there, because vectors from another model do not match the stored ones. pgvector replicates with PostgreSQL’s write-ahead log, while other stores ship snapshots or backups on a schedule. After a failover, run fixed retrieval queries for each access level to check results and permissions.
Can GPU virtual machines be replicated for disaster recovery?
A running VM with a passthrough GPU cannot be replicated through snapshots, because vSphere 9.0 lists snapshots as unavailable for VMs configured with DirectPath I/O. Broadcom describes vSphere Replication as an agent that sends changed disk blocks, but its documentation says nothing about passthrough devices, so test it first. The simpler design rebuilds inference VMs at the second site from a template and replicates only the VMs that hold data.
How should AI virtual machine backups be handled?
Back up the data inside them rather than relying on VM snapshots alone: the vector store with its own tools, adapters and evaluation sets as files, and configuration in git. A VM with a passthrough GPU needs a guest agent while it runs, since vSphere offers no snapshot for it. Inference VMs whose weights and images are pinned can be rebuilt from a template instead of restored.

Send us the AI services you run, their models and peak requests in flight, the minimum service level each needs after a site loss and where your second site is. We reply within one business day with a configuration and quote for the GPU servers there, after we check the rack, power and airflow.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna