Failover vs failback: the DR failover steps on the day, from declaration to DNS, VPN and failback
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- Failover starts the replicas at the recovery site and switches the network so that users reach them; failback returns production to the repaired primary site with the changes made at the recovery site, and Veeam treats failover as an intermediate state that ends with undo failover, failback or permanent failover
- A named role declares the disaster against criteria written per tier; VMware Live Site Recovery then starts virtual machines by recovery priority 1 to 5, and Veeam failover plans set the order and delays so that a DNS server runs before the VMs that depend on it
- Recovered servers get new addresses through Veeam re-IP rules, which cover one guest operating system family only, or VMware IP customisation, or keep them on a stretched network or behind Veeam Cloud Connect network extension appliances, where re-IP is not supported and static addresses are required
- Resolvers may cache a DNS record for its TTL (RFC 1035), so lower the TTLs of records that would change well before a disaster; public addresses at the recovery site usually differ, so map them in advance and prepare VPN peers, partner allow-lists, certificates and licence servers for the recovery site
- Failback brings the primary systems to the state of the recovery site, so in our reading writes left on the primary after the last replication run are overwritten unless extracted first; Veeam Cloud Connect fails back on the tenant side only, VM by VM after a full site failover, and Live Site Recovery reprotects before and after the planned migration back
Eurokommerz × Vixen.UNO: Cloud Disaster Recovery Talk to an expert →
Failover vs failback: what happens on the day
Failover moves production to the recovery site: the replicas start there in a set order, and the network is switched so that users reach them. Failback returns production to the primary site once it is repaired, carrying back the changes made at the recovery site in a planned window. Veeam Backup & Replication 13 calls failover “an intermediate step that needs to be finalized” by undoing it, failing back or failing over permanently. A planned failover makes the same move before a known event while the primary still runs, after a full synchronisation and a shutdown of the source VMs. VMware Live Site Recovery, part of Protection and Recovery from VCF 9.1, runs a recovery plan as a planned migration, its term for this, or as a disaster recovery, and reprotects before it fails back.
| PHASE | ACTIONS | CHECKS |
|---|---|---|
| Declare | the named role applies the criteria, records the time, sets scope and restore point | provider, team and users informed |
| Isolate the primary | power off or disconnect what still runs there | no VM runs at both sites |
| Fail over by tier | directory service and DNS, then databases, application servers, front ends | each tier passes an application check |
| Switch the network | re-IP or network extension, DNS, public IP mappings, VPN peers, firewall rules | names resolve to recovery addresses inside and outside |
| Bring users back | VPN profiles, sign-in and MFA, a message to users | test transactions pass; owners sign off |
| Run at the recovery site | backups, monitoring, a log of every change | backup jobs succeed |
| Fail back | transfer the changes, switch in a planned window, restore the replication direction | divergent data recovered or written off |
Our checklist; terms as in Veeam’s version 13 guides and Broadcom’s Live Site Recovery 9.0 and Protection and Recovery 9.1 documentation (read 6 October 2026).
Each row belongs in the runbook of each system, as our guide to the disaster recovery runbook shows.
Declaring the disaster: who decides and on what criteria
The plan names the role that declares a disaster, with deputies reachable out of hours, and criteria per tier. We suggest declaring when the primary site will not be back within the tier’s RTO, is damaged or unsafe, or holds data that can no longer be trusted. For the entities in its Article 1, Implementing Regulation (EU) 2024/2690 lists “conditions for plan activation and deactivation” among the plan’s contents (Annex, point 4.1.2(d)); for others, which law applies is for their legal department to assess.
Because the RTO runs from the start of the outage, detection and decision have to fit into the RTO minus the failover time, which is the first hour for a one-hour failover and a two-hour RTO. Waiting is the better choice only for an outage certain to end within the RTO, because a failover costs the writes made after the last replication run and a failback later; our guide to RPO and RTO covers the targets per tier.
Broadcom separates the privilege to run a Live Site Recovery plan from the privilege to test one, because a run “makes changes at both sites that require significant time and effort to reverse”. With Veeam Cloud Connect, the tenant starts its cloud failover plan itself or, if its backup server is lost too, asks the provider to start it.
The declaration also fixes the restore point. After a fire or a hardware failure it is the latest one, after a faulty update one from before it, and after ransomware one from before the intrusion. Veeam can fail over “to the latest state of a replica or to any of its restore points”, but its version 12 Quick Start Guide limits VM replicas to 28 restore points, about 7 hours at a 15-minute interval; older clean points come from backups.
Failover order by tier and dependency
Systems start in the order their dependencies require, the directory service and DNS first. VMware Live Site Recovery starts virtual machines by recovery priority, from 1, the highest, to 5, with every virtual machine of a new plan at 3, and waits for the VMware Tools heartbeat of all virtual machines of one priority before starting the next; dependencies work only within a priority. Veeam failover plans set the order and a delay before the next VM, which “helps ensure that some VMs, such as a DNS server, are already running at the time the dependent VMs start”; a cloud failover plan starts at most 10 VMs at once.
A heartbeat shows only that the guest operating system runs, so check each tier before the next: the database accepts connections, the application server signs in to it, and a test user completes a transaction before the front ends open.
Re-IP, stretched networks or network extension
Recovered servers get new addresses from the recovery site’s subnets (re-IP), stay on production subnets stretched across both sites or moved there by routing, or keep their addresses behind appliances extending the network to a provider.
Veeam’s re-IP rules “map IPs in the production site to IPs in the disaster recovery (DR) site” at failover, for VMs of one guest operating system family only; Linux servers need another method, such as a script in the guest. VMware Live Site Recovery customises IP settings per virtual machine, in bulk with the DR IP Customizer tool or through subnet-level rules, and needs VMware Tools or VMware’s Operating System Specific Packages in the guest. Each new address then has to reach DNS, firewall rules, hosts files and application settings.
A stretched layer 2 network keeps the addresses, but its default gateway has to work at the recovery site without the primary, and a loop or broadcast storm reaches both sites. Routing can move a subnet only as a whole, which suits a full site failover (our reading).
Veeam Cloud Connect keeps production addresses too, since Veeam “does not support re-IP rules for VM replicas on the cloud host”, and because Cloud Connect replication “does not support DHCP”, replicated VMs need static addresses. Network extension appliances join the subnets through a VPN after a partial site failover and act as the replicas’ gateway after a full site failover, as our comparison of Veeam replication, backup copy and Cloud Connect explains.
DNS TTL, public IP addresses and VPN
RFC 1035 defines the TTL as the interval for which a record “may be cached before the source of the information should again be consulted”. After a change, clients can use the old address for as long as the TTL allows. Lower the TTL of records that would change at failover well before any disaster, because answers already cached keep the old TTL; we suggest a few minutes.
Public addresses usually stay with the primary site’s internet provider. In Veeam Cloud Connect the provider supplies public addresses through the hardware plan, and for a full site failover the tenant maps each public address and port to a replica’s internal address and port in the cloud failover plan. Public DNS records for the website, mail, the VPN gateway and partner interfaces then change, and partners that allow-list your addresses need the new ones in advance.
Configure site-to-site VPN peers, keys and routes at the recovery site before the day, and give remote-access clients the recovery gateway as a second server or a DNS name with a short TTL.
Our disaster recovery service replicates virtual machines via Veeam Cloud Connect to a recovery site in Baltneta’s Tier-3 data centres in Lithuania. Send us the subnets, public addresses and VPN peers your critical systems use today.
Firewall rules, certificates and licence servers
The recovery site’s firewalls have to allow the flows production needs, and after a re-IP every rule written for an address or subnet needs a counterpart there.
A certificate covers only the names it lists, and under RFC 9525 an automated client that finds no match “SHOULD terminate the communication attempt with a bad certificate error”, so a service reached under a new name needs that name on its certificate. Certificates on reverse proxies and load balancers outside the replicas need installing there too, and an internal certificate authority that publishes its revocation list only at the primary site can make clients that check revocation reject valid certificates.
Software that checks its licence against a licence server, a MAC address or a host ID can refuse to start at the recovery site, and the licence server may itself be tied to its address: NVIDIA’s License System, which licenses vGPU software, requires the platform of its Delegated License Service to have “a fixed (unchanging) IP address”.
User access and running at the recovery site
Users reach the recovery site through VPN gateways and published applications prepared there, and sign in against the directory service of the first tier; multi-factor authentication, RADIUS and any identity provider must be reachable from there too. Tell users what changes through a channel that does not depend on the primary site, since mail may be down.
From the first hour, back up the recovered systems, because the replica holds the only copy of the data written there. Log every change at the recovery site, from firewall rules to DNS records, because failback reverses or keeps each one. Rehearse this network side in partial failover tests, which isolated tests leave out, as our guide to disaster recovery testing describes.
The failback process: delta synchronisation and data divergence
Veeam’s failback transfers “all changes that took place while the VM replica was running to the original VM”, or to a new location if the source host is lost, and the changes are “only transferred but not published” until you test the original VM and commit failback. With Cloud Connect, failback “is available on the tenant side only”, which means a backup server lost with the site must be rebuilt and connected to the provider first, and after a full site failover the tenant fails back each VM of the cloud failover plan separately, in a written order.
In VMware Live Site Recovery, a protected site that comes back online after a disaster recovery creates what Broadcom calls “a split-brain scenario”, with production VMs at both sites, until the plan runs again as a planned migration, which powers them off there and completes the recovery; a reprotect, a planned migration back and a second reprotect then make up the failback, as our article on VMware Live Site Recovery after Broadcom explains.
Writes that reached the primary systems after the last replication run, until the primary was isolated, are the divergence, narrowed when a Live Site Recovery disaster recovery “first attempts a storage synchronization” and succeeds. Failback brings the primary disks to the recovery site’s state, and Broadcom’s reprotect “forces synchronization of the storage from the new protected site to the new recovery site”, so in our reading those writes are lost unless extracted first, from a copy of the old VM or its database logs. Undo failover works the other way round, returning to the original VM and discarding “all changes made to the VM replica while it was running”.
- Keep the repaired primary site’s VMs powered off or disconnected until the switch.
- Agree with the data owners which writes left on the primary are extracted and which are written off.
- Switch in an agreed maintenance window with a rollback plan, since the systems stop while the last changes move.
- Point DNS records, public address mappings, VPN peers and firewall rules back at the primary site.
- Let the owners test the original systems, then commit failback, or undo it and switch the network back to the replica.
- Restore replication in its original direction and check that backups protect the primary systems.
Permanent failover keeps the replica as production, which Veeam calls acceptable when original and replica “are located in the same site and are nearly equal in terms of resources”. With Cloud Connect it follows a full site failover and removes the replica’s restore points; “other jobs are not modified automatically”, so update backup jobs yourself.
Our disaster recovery service switches over by a rehearsed scenario on the day of a disaster and fails back once the primary is restored. Tell us which systems would run from the recovery site first and who may declare a disaster.
What we do
Under our disaster recovery service, our engineering partner Vixen.UNO defines with you the critical systems, the target RPO and RTO per tier and the disaster scenarios, and the output is a continuity plan. Virtual machines replicate from a 15-minute interval via Veeam Cloud Connect to Baltneta’s Tier-3 data centres in Lithuania (ISO 27001, PCI DSS), managed from your Veeam console or entirely on our side, with scheduled failover tests in an isolated environment and a report after each. On the day, switchover follows the rehearsed scenario and failback follows once the primary is restored, with support, RPO and RTO fixed in the SLA.
FAQ
What is the difference between failover and failback?
What are the steps of a DR failover?
What is a planned failover?
How does failback work after disaster recovery?
Do servers keep their IP addresses after a failover?
What DNS TTL should be set for disaster recovery?
Send us your critical systems by tier, how they replicate today, the subnets and public IP addresses they use and who may declare a disaster. We reply within one business day with a date for a first call, where we work through your critical systems, current backup and target RPO and RTO, and you leave with two or three possible DR scenarios. The first call is free of charge.
Talk to an expertWe reply within one business day