BLOG · GUIDE ·

Failover vs failback: the DR failover steps on the day, from declaration to DNS, VPN and failback

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • Failover starts the replicas at the recovery site and switches the network so that users reach them; failback returns production to the repaired primary site with the changes made at the recovery site, and Veeam treats failover as an intermediate state that ends with undo failover, failback or permanent failover
  • A named role declares the disaster against criteria written per tier; VMware Live Site Recovery then starts virtual machines by recovery priority 1 to 5, and Veeam failover plans set the order and delays so that a DNS server runs before the VMs that depend on it
  • Recovered servers get new addresses through Veeam re-IP rules, which cover one guest operating system family only, or VMware IP customisation, or keep them on a stretched network or behind Veeam Cloud Connect network extension appliances, where re-IP is not supported and static addresses are required
  • Resolvers may cache a DNS record for its TTL (RFC 1035), so lower the TTLs of records that would change well before a disaster; public addresses at the recovery site usually differ, so map them in advance and prepare VPN peers, partner allow-lists, certificates and licence servers for the recovery site
  • Failback brings the primary systems to the state of the recovery site, so in our reading writes left on the primary after the last replication run are overwritten unless extracted first; Veeam Cloud Connect fails back on the tenant side only, VM by VM after a full site failover, and Live Site Recovery reprotects before and after the planned migration back

Eurokommerz × Vixen.UNO: Cloud Disaster Recovery  Talk to an expert →

Failover vs failback: what happens on the day

Failover moves production to the recovery site: the replicas start there in a set order, and the network is switched so that users reach them. Failback returns production to the primary site once it is repaired, carrying back the changes made at the recovery site in a planned window. Veeam Backup & Replication 13 calls failover “an intermediate step that needs to be finalized” by undoing it, failing back or failing over permanently. A planned failover makes the same move before a known event while the primary still runs, after a full synchronisation and a shutdown of the source VMs. VMware Live Site Recovery, part of Protection and Recovery from VCF 9.1, runs a recovery plan as a planned migration, its term for this, or as a disaster recovery, and reprotects before it fails back.

PHASEACTIONSCHECKS
Declarethe named role applies the criteria, records the time, sets scope and restore pointprovider, team and users informed
Isolate the primarypower off or disconnect what still runs thereno VM runs at both sites
Fail over by tierdirectory service and DNS, then databases, application servers, front endseach tier passes an application check
Switch the networkre-IP or network extension, DNS, public IP mappings, VPN peers, firewall rulesnames resolve to recovery addresses inside and outside
Bring users backVPN profiles, sign-in and MFA, a message to userstest transactions pass; owners sign off
Run at the recovery sitebackups, monitoring, a log of every changebackup jobs succeed
Fail backtransfer the changes, switch in a planned window, restore the replication directiondivergent data recovered or written off

Our checklist; terms as in Veeam’s version 13 guides and Broadcom’s Live Site Recovery 9.0 and Protection and Recovery 9.1 documentation (read 6 October 2026).

Each row belongs in the runbook of each system, as our guide to the disaster recovery runbook shows.

Declaring the disaster: who decides and on what criteria

The plan names the role that declares a disaster, with deputies reachable out of hours, and criteria per tier. We suggest declaring when the primary site will not be back within the tier’s RTO, is damaged or unsafe, or holds data that can no longer be trusted. For the entities in its Article 1, Implementing Regulation (EU) 2024/2690 lists “conditions for plan activation and deactivation” among the plan’s contents (Annex, point 4.1.2(d)); for others, which law applies is for their legal department to assess.

Because the RTO runs from the start of the outage, detection and decision have to fit into the RTO minus the failover time, which is the first hour for a one-hour failover and a two-hour RTO. Waiting is the better choice only for an outage certain to end within the RTO, because a failover costs the writes made after the last replication run and a failback later; our guide to RPO and RTO covers the targets per tier.

Broadcom separates the privilege to run a Live Site Recovery plan from the privilege to test one, because a run “makes changes at both sites that require significant time and effort to reverse”. With Veeam Cloud Connect, the tenant starts its cloud failover plan itself or, if its backup server is lost too, asks the provider to start it.

The declaration also fixes the restore point. After a fire or a hardware failure it is the latest one, after a faulty update one from before it, and after ransomware one from before the intrusion. Veeam can fail over “to the latest state of a replica or to any of its restore points”, but its version 12 Quick Start Guide limits VM replicas to 28 restore points, about 7 hours at a 15-minute interval; older clean points come from backups.

Failover order by tier and dependency

Systems start in the order their dependencies require, the directory service and DNS first. VMware Live Site Recovery starts virtual machines by recovery priority, from 1, the highest, to 5, with every virtual machine of a new plan at 3, and waits for the VMware Tools heartbeat of all virtual machines of one priority before starting the next; dependencies work only within a priority. Veeam failover plans set the order and a delay before the next VM, which “helps ensure that some VMs, such as a DNS server, are already running at the time the dependent VMs start”; a cloud failover plan starts at most 10 VMs at once.

A heartbeat shows only that the guest operating system runs, so check each tier before the next: the database accepts connections, the application server signs in to it, and a test user completes a transaction before the front ends open.

Re-IP, stretched networks or network extension

Recovered servers get new addresses from the recovery site’s subnets (re-IP), stay on production subnets stretched across both sites or moved there by routing, or keep their addresses behind appliances extending the network to a provider.

Veeam’s re-IP rules “map IPs in the production site to IPs in the disaster recovery (DR) site” at failover, for VMs of one guest operating system family only; Linux servers need another method, such as a script in the guest. VMware Live Site Recovery customises IP settings per virtual machine, in bulk with the DR IP Customizer tool or through subnet-level rules, and needs VMware Tools or VMware’s Operating System Specific Packages in the guest. Each new address then has to reach DNS, firewall rules, hosts files and application settings.

A stretched layer 2 network keeps the addresses, but its default gateway has to work at the recovery site without the primary, and a loop or broadcast storm reaches both sites. Routing can move a subnet only as a whole, which suits a full site failover (our reading).

Veeam Cloud Connect keeps production addresses too, since Veeam “does not support re-IP rules for VM replicas on the cloud host”, and because Cloud Connect replication “does not support DHCP”, replicated VMs need static addresses. Network extension appliances join the subnets through a VPN after a partial site failover and act as the replicas’ gateway after a full site failover, as our comparison of Veeam replication, backup copy and Cloud Connect explains.

DNS TTL, public IP addresses and VPN

RFC 1035 defines the TTL as the interval for which a record “may be cached before the source of the information should again be consulted”. After a change, clients can use the old address for as long as the TTL allows. Lower the TTL of records that would change at failover well before any disaster, because answers already cached keep the old TTL; we suggest a few minutes.

Public addresses usually stay with the primary site’s internet provider. In Veeam Cloud Connect the provider supplies public addresses through the hardware plan, and for a full site failover the tenant maps each public address and port to a replica’s internal address and port in the cloud failover plan. Public DNS records for the website, mail, the VPN gateway and partner interfaces then change, and partners that allow-list your addresses need the new ones in advance.

Configure site-to-site VPN peers, keys and routes at the recovery site before the day, and give remote-access clients the recovery gateway as a second server or a DNS name with a short TTL.

Our disaster recovery service replicates virtual machines via Veeam Cloud Connect to a recovery site in Baltneta’s Tier-3 data centres in Lithuania. Send us the subnets, public addresses and VPN peers your critical systems use today.

Firewall rules, certificates and licence servers

The recovery site’s firewalls have to allow the flows production needs, and after a re-IP every rule written for an address or subnet needs a counterpart there.

A certificate covers only the names it lists, and under RFC 9525 an automated client that finds no match “SHOULD terminate the communication attempt with a bad certificate error”, so a service reached under a new name needs that name on its certificate. Certificates on reverse proxies and load balancers outside the replicas need installing there too, and an internal certificate authority that publishes its revocation list only at the primary site can make clients that check revocation reject valid certificates.

Software that checks its licence against a licence server, a MAC address or a host ID can refuse to start at the recovery site, and the licence server may itself be tied to its address: NVIDIA’s License System, which licenses vGPU software, requires the platform of its Delegated License Service to have “a fixed (unchanging) IP address”.

User access and running at the recovery site

Users reach the recovery site through VPN gateways and published applications prepared there, and sign in against the directory service of the first tier; multi-factor authentication, RADIUS and any identity provider must be reachable from there too. Tell users what changes through a channel that does not depend on the primary site, since mail may be down.

From the first hour, back up the recovered systems, because the replica holds the only copy of the data written there. Log every change at the recovery site, from firewall rules to DNS records, because failback reverses or keeps each one. Rehearse this network side in partial failover tests, which isolated tests leave out, as our guide to disaster recovery testing describes.

The failback process: delta synchronisation and data divergence

Veeam’s failback transfers “all changes that took place while the VM replica was running to the original VM”, or to a new location if the source host is lost, and the changes are “only transferred but not published” until you test the original VM and commit failback. With Cloud Connect, failback “is available on the tenant side only”, which means a backup server lost with the site must be rebuilt and connected to the provider first, and after a full site failover the tenant fails back each VM of the cloud failover plan separately, in a written order.

In VMware Live Site Recovery, a protected site that comes back online after a disaster recovery creates what Broadcom calls “a split-brain scenario”, with production VMs at both sites, until the plan runs again as a planned migration, which powers them off there and completes the recovery; a reprotect, a planned migration back and a second reprotect then make up the failback, as our article on VMware Live Site Recovery after Broadcom explains.

Writes that reached the primary systems after the last replication run, until the primary was isolated, are the divergence, narrowed when a Live Site Recovery disaster recovery “first attempts a storage synchronization” and succeeds. Failback brings the primary disks to the recovery site’s state, and Broadcom’s reprotect “forces synchronization of the storage from the new protected site to the new recovery site”, so in our reading those writes are lost unless extracted first, from a copy of the old VM or its database logs. Undo failover works the other way round, returning to the original VM and discarding “all changes made to the VM replica while it was running”.

  1. Keep the repaired primary site’s VMs powered off or disconnected until the switch.
  2. Agree with the data owners which writes left on the primary are extracted and which are written off.
  3. Switch in an agreed maintenance window with a rollback plan, since the systems stop while the last changes move.
  4. Point DNS records, public address mappings, VPN peers and firewall rules back at the primary site.
  5. Let the owners test the original systems, then commit failback, or undo it and switch the network back to the replica.
  6. Restore replication in its original direction and check that backups protect the primary systems.

Permanent failover keeps the replica as production, which Veeam calls acceptable when original and replica “are located in the same site and are nearly equal in terms of resources”. With Cloud Connect it follows a full site failover and removes the replica’s restore points; “other jobs are not modified automatically”, so update backup jobs yourself.

Our disaster recovery service switches over by a rehearsed scenario on the day of a disaster and fails back once the primary is restored. Tell us which systems would run from the recovery site first and who may declare a disaster.

What we do

Under our disaster recovery service, our engineering partner Vixen.UNO defines with you the critical systems, the target RPO and RTO per tier and the disaster scenarios, and the output is a continuity plan. Virtual machines replicate from a 15-minute interval via Veeam Cloud Connect to Baltneta’s Tier-3 data centres in Lithuania (ISO 27001, PCI DSS), managed from your Veeam console or entirely on our side, with scheduled failover tests in an isolated environment and a report after each. On the day, switchover follows the rehearsed scenario and failback follows once the primary is restored, with support, RPO and RTO fixed in the SLA.

FAQ

What is the difference between failover and failback?
Failover moves production to the recovery site by starting the replicas there and switching the network so that users reach them. Failback returns production to the repaired primary site and carries back the changes made while the replicas ran. Veeam treats failover as an intermediate state that ends with undo failover, failback or permanent failover.
What are the steps of a DR failover?
Declare the disaster against written criteria, isolate what still runs at the primary site and choose the restore point. Start the systems by tier, directory service and DNS first, then switch DNS records, public addresses, VPN peers and firewall rules. Bring users back with test sign-ins and owner sign-off, and back up the recovered systems while they run at the recovery site.
What is a planned failover?
A planned failover moves production to the recovery site before a known event, such as data centre maintenance or a migration, while the primary site still runs. Veeam runs the replication job to synchronise the replica fully, shuts down the source VM and fails over to the replica, and its Cloud Connect guide offers this for cloud replicas within a partial site failover. VMware Live Site Recovery calls it a planned migration, which attempts a graceful shutdown of the protected virtual machines, performs a final synchronisation before it powers them on at the recovery site and stops if errors occur.
How does failback work after disaster recovery?
Once the primary site is repaired, the changes made at the recovery site are transferred back and production switches over in a planned maintenance window. In Veeam you test the original VM with the transferred changes before you commit failback, and with Veeam Cloud Connect the tenant runs it, VM by VM after a full site failover. In VMware Live Site Recovery a planned migration completes the recovery, a reprotect reverses replication, and a planned migration back with a second reprotect completes the failback.
Do servers keep their IP addresses after a failover?
That depends on the network design at the recovery site. With re-IP, Veeam re-IP rules or VMware IP customisation give recovered servers addresses from the recovery subnets, and DNS, firewall rules and configurations must follow; on a stretched layer 2 network, on subnets moved to the recovery site by routing or behind Veeam Cloud Connect network extension appliances, servers keep their addresses. Veeam does not support re-IP rules for replicas on a Cloud Connect cloud host and requires static IP addresses there.
What DNS TTL should be set for disaster recovery?
A resolver may cache a record for its TTL, so after a change clients can keep using the old address for up to that long. Lower the TTL of records that would change at failover, such as public names, the VPN gateway and re-addressed servers, to a few minutes well before any disaster, because answers cached under the old TTL stay valid until it runs out. Records that keep their addresses through network extension or a stretched network do not need to change.

Send us your critical systems by tier, how they replicate today, the subnets and public IP addresses they use and who may declare a disaster. We reply within one business day with a date for a first call, where we work through your critical systems, current backup and target RPO and RTO, and you leave with two or three possible DR scenarios. The first call is free of charge.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry – see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna