Disaster recovery testing: from tabletop exercises to isolated and full failover tests
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- Disaster recovery tests range from tabletop exercises and drills of single functions through automated restore or replica checks and isolated failover tests to partial and full failovers of live service
- NIST SP 800-34 Rev. 1 (Table 3-6) pairs low-impact systems with a tabletop exercise, moderate-impact ones with a functional exercise that includes system recovery, and high-impact ones with a full-scale functional exercise plus an alternate-site test
- An isolated failover test starts replicas at the recovery site on networks with no path to production, as Veeam’s SureReplica does in a virtual lab and VMware Live Site Recovery does in a recovery plan test, whose auto-created isolated network does not span hosts
- Measure the time from the declared disaster to a working, owner-verified service per tier against its RTO, the age of the restore point used against the RPO, every failed or improvised step, and whether anything reached production
- The report says what came up, how fast and what to fix; Implementing Regulation (EU) 2024/2690 requires the entities it covers to document test results, take corrective action where needed and feed lessons learnt into the plans
Eurokommerz × Vixen.UNO: Cloud Disaster Recovery Talk to an expert →
Disaster recovery testing and the DR test types
Disaster recovery testing shows whether your systems come back at the recovery site within the recovery time objective (RTO) and with no more data loss than the recovery point objective (RPO), and whether the team and the plan get them there. From discussion to live service, the test types are a tabletop exercise that talks a scenario through, a drill of one function, automated restore or replica checks, an isolated failover test that brings up a whole tier at the recovery site without touching production, and partial and full failover tests that move live service.
Table 3-6 of NIST SP 800-34 Rev. 1, “ISCP TT&E Activities”, lists sample activities by the impact level of the system: a tabletop exercise for low-impact systems, a functional exercise for moderate-impact ones, described as a “simulation of a disruption with a system recovery component”, and a full-scale functional exercise plus a test at the alternate processing site for high-impact ones. Some DR plans use five other names: checklist review, structured walk-through, simulation, parallel test and full interruption test. The first two review the plan on paper or in a meeting, a simulation is close to a functional exercise, and the last two bring the recovery site into operation. Setting the targets for each tier is covered in our guide to RPO and RTO.
| TEST TYPE | WHAT IT PROVES | EFFORT | HOW OFTEN |
|---|---|---|---|
| Tabletop exercise | roles, decisions, contacts and gaps in the plan | a few hours; no systems touched | twice a year and after plan changes |
| Drill | one function, such as the call-out, one restore or a DNS change | an hour or two | quarterly, rotating the function |
| Restore or replica check | a backup or replica boots and passes its tests | automated once set up | after each job, or weekly |
| Isolated failover test | a tier starts in order at the recovery site within a measured time | a day; capacity at the recovery site | twice a year for the critical tier, yearly for the rest |
| Partial failover test | one application runs live from the recovery site, then fails back | a maintenance window and a rollback plan | yearly for critical applications |
| Full failover test | the business runs from the recovery site and returns | an outage window and every team | after the other tests pass, if the business accepts the risk |
Tabletop exercise after NIST SP 800-84 and SP 800-34 Rev. 1 (Table 3-6), drill after FEMA’s HSEEP doctrine (January 2020); the other rows, effort and frequency are our suggestion.
Tabletop exercises and DR drills
A tabletop exercise takes the people named in the plan through a scenario without touching a system. A facilitator releases the scenario in stages, for example a failed storage array, then the news that the newest clean restore point is a day old, and the participants say what they would do and whom they would call. CISA’s Tabletop Exercise Packages, with “template exercise objectives, scenarios, and discussion questions”, include ransomware scenarios.
The word drill has two senses, so check which one a policy or contract means. FEMA’s HSEEP doctrine (January 2020) defines a drill as “an operations-based exercise often employed to validate a single operation or function”; in DR that can be reaching the on-call team at night or restoring the directory service in a lab. In AWS Elastic Disaster Recovery, “A Recovery Drill is a non-disruptive test that performs all the same steps as an actual recovery”, and AWS recommends one “on at least a quarterly basis”.
Isolated failover tests in Veeam and VMware Live Site Recovery
An isolated failover test starts the replicas at the recovery site on networks with no path to production, so recovered servers keep their production addresses and names without clashing with the originals. NIST SP 800-34 Rev. 1 says testing “should be conducted in as close to an operating environment as possible”, and a test on the recovery hosts, from the replicas you would use on the day, comes close without disrupting operations.
Veeam Backup & Replication 13 calls the isolated environment a virtual lab. Veeam’s Help Center describes SureReplica as booting a regular or CDP replica “from the necessary restore point in the isolated environment”, running tests, powering it off and creating “a report on the VM replica state”. An application group first starts the machines that the verified machine depends on, typically domain controllers and DNS servers. Masquerade IP addresses give access from production through the lab’s proxy appliance, and custom verification scripts can check the application beyond predefined tests such as heartbeat and ping.
In VMware Live Site Recovery, part of Protection and Recovery from VCF 9.1, a recovery plan test starts the protected VMs at the recovery site in plan order. With the auto-created isolated network and no site-level mappings, Broadcom’s Live Site Recovery 9.0 documentation (updated 14 July 2026) puts them on “temporary networks that are not connected to any physical network”, and states that “An isolated test network does not span hosts”. Since VMs that interact must be recovered “to the same test network”, an application spread over several hosts needs an assigned test network spanning them, such as a VLAN without a route to production; that is our reading. The recovery plan history records “the start and end times for the whole plan and for each step”, for tests as well as runs.
Check the isolation before and after each test. Broadcom’s KB 444382 describes test VMs whose DNS PTR records reached production because the recovered domain controller had a network path to the production domain controllers, and says the issue “frequently occurs when using ‘User Defined’ network mappings instead of auto-generated ‘Isolated’ networks”. Veeam’s best practice guide warns that a multi-host virtual lab whose VLAN ID is in use in production is not isolated.
Our disaster recovery service runs scheduled failover tests in an isolated environment, with no impact on production systems. Tell us which systems you would test first and how you restore them today.
Partial and full failover tests
An isolated test leaves out the production network side: DNS changes reaching clients, public IP addresses, VPN tunnels, firewall rules, licence servers, users on their own devices and full load. A partial failover test covers these for one application or tier, which moves to the recovery site in a maintenance window, serves users and fails back; a full failover test does the same for the whole estate. In VMware Live Site Recovery this is a planned migration, which “shuts down virtual machines in reverse priority order” at the protected site and synchronises the replicated datastores before the switch, so it measures recovery time, not data loss after a sudden failure. For Veeam replicas it is a planned failover, a manually started switch from the primary VM to its replica.
In the five-name vocabulary, a parallel test runs the recovered systems alongside production, often on copies of the same input for comparison, while users stay on production; a full interruption test shuts production down, runs from the recovery site only and turns into an outage if the recovery fails.
Treat failback as part of the test, since the systems return with the changes made at the recovery site, and run partial and full tests only after isolated tests have passed, with a rollback plan. The steps on the day are in our guide to failover and failback steps.
How to run a DR test step by step
- Agree the scope, the scenario, the abort criteria and what counts as a working service with the system owners, for example a test order posted in the ERP system rather than a VM that answers a ping.
- Name the roles: the on-shift team executes the runbook, its author only watches so that gaps in the written steps show up, a data collector keeps the timeline, and a person with authority declares the simulated disaster and its end.
- Check that the test network has no route to production, and block its outbound traffic so that recovered applications cannot send mail or call partner interfaces.
- Declare the disaster, start the clock and let the team work from the runbook, with a timestamp for every phase.
- Let the application owners confirm their services with test transactions, and record the sign-off time per tier.
- Clean up, confirm that production DNS and directory data did not change, and hold the debrief the same day.
- Write the report, schedule the corrective actions and set the date of the retest.
NIST SP 800-84 adds two roles for functional exercises: controllers “who monitor, manage, and control exercise activity”, and simulators who play parties not taking part, such as a hosting provider. The per-system runbook is described in our guide to the disaster recovery runbook.
What to measure in a DR test
Give human steps such as decisions, sign-ins and checks their own timestamps, because the tool logs cover the machine steps.
| MEASURE | HOW TO CAPTURE IT | COMPARE WITH |
|---|---|---|
| Time to working service | timestamps per tier from the declared disaster to sign-off | the RTO of the tier |
| Data loss | time of the restore point used, or of the last committed transaction | the RPO |
| Start order | when each service answered; waits for dependencies | the order in the runbook |
| Failed or improvised steps | steps that failed, needed a workaround or a person not in the runbook | the runbook, step by step |
| Isolation | production DNS, directory and mail logs; outbound traffic from the test network | no change in production |
| Capacity | CPU, memory and storage latency during the checks | the sizing of the recovery site |
Our checklist; capture methods from Broadcom’s recovery plan history, Veeam’s SureReplica report and NIST SP 800-84.
In an isolated test, data loss is the gap between the declared disaster and the restore point the test starts from. With replication every 15 minutes it should stay within that interval plus the duration of one replication run; a larger gap means the schedule or the bandwidth does not support the RPO.
The DR test report and corrective actions
NIST SP 800-84 closes a test with a hotwash, “an informal test debrief with participants”, and an after action report that “determines how well the tested systems or components functioned”, with observations and recommendations from which action items are assigned. For a DR test the report says what came up, how fast and what to fix: scope and scenario, the timeline per tier against the RTO, the restore points against the RPO, failed or improvised steps, and corrective actions with an owner, a due date and a retest date.
For the entities in its Article 1, among them cloud computing, data centre and managed service providers, Implementing Regulation (EU) 2024/2690 makes documenting and acting on test results an obligation. Point 4.1.4 requires that “the plans incorporate lessons learnt from such tests”, and point 4.2.6, on recovery tests of backups and redundancies, adds that entities “shall document the results of the tests and, where needed, take corrective action”; points 3.5.5 and 4.3.4 cover tests of incident response procedures and of the crisis management plan. Other organisations in scope of NIS2 follow their national law and can use the regulation as a reference; which text applies is for the company’s legal department to assess. The related points are in our article on NIS2 backup and business continuity requirements.
After each scheduled failover test in our disaster recovery service you receive a report on what came up, how fast and what to fix. Send us the date and result of your last DR test through the form below.
How often to test and what triggers an extra test
NIST SP 800-84 asks for test, training and exercise events “periodically; following organizational changes, updates to an IT plan, or the issuance of new TT&E guidance; or as otherwise needed”, and SP 800-34 Rev. 1 leaves the frequency to the organisation. For DR, an extra test is due after changes to the recovery path: new storage, hypervisor or replication software versions at either site, a redesigned network, new dependencies of a critical application, another replication product or provider, or a new on-call team. How the EU texts word the interval is discussed in our article on why backup is not disaster recovery.
What we do
Under our disaster recovery service, our engineering partner Vixen.UNO defines with you the critical systems, the target RPO and RTO for each tier and the disaster scenarios. Virtual machines replicate from a 15-minute interval via Veeam Cloud Connect to a recovery site in Baltneta’s Tier-3 data centres in Lithuania. Scheduled failover tests run in an isolated environment with no impact on production systems, each with a report on what came up, how fast and what to fix. The standard targets are an RPO from 15 minutes and an RTO of 1 to 2 hours for critical systems, your targets per tier are fixed in the SLA, and the first call is free of charge.
FAQ
What is disaster recovery testing?
What are the types of DR tests?
How do you test a disaster recovery plan without affecting production?
What is a DR drill?
How often should a disaster recovery plan be tested?
What should a DR test report contain?
Send us the systems you would fail over first, how you restore them today and the date and result of your last DR test. We reply within one business day with a date for a first call, where we work through your critical systems, current backup and target RPO and RTO, and you leave with two or three possible DR scenarios. The first call is free of charge.
Talk to an expertWe reply within one business day