Disaster recovery runbook: how to write the steps for one system, with a template and an example
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- A disaster recovery runbook is the per-system procedure under the DR plan: purpose, owner and deputies, recovery objectives, preconditions, dependencies, invocation, steps with expected results and time budgets, verification and sign-off, rollback, failback, contacts and a version and test history
- NIST SP 800-34 Rev. 1 gives a system contingency plan three phases, activation and notification, recovery and reconstitution; for the entities it covers, Implementing Regulation (EU) 2024/2690 point 4.1.2(f) lists recovery plans for specific operations, including recovery objectives, among the plan contents
- Each step is written for someone other than its author, since NIST notes that a disruption can leave some personnel unavailable and AWS describes runbooks as processes with consistent outcomes no matter who uses them
- The worked example recovers a three-tier application on vSphere after the loss of the primary site, with an RTO of 4 hours: identity, DNS and time first, then the database, application and web tiers and user access; with illustrative step times, the owner signs off after 2 hours 55 minutes
- Recovery plans in VMware Live Site Recovery and Veeam Recovery Orchestrator automate the VM start order, checks and prompts, while invocation, external network changes and the business sign-off stay in the runbook, which every change request should be checked against
Eurokommerz × Vixen.UNO: Cloud Disaster Recovery Talk to an expert →
What a disaster recovery runbook is and how it differs from the DR plan
A disaster recovery runbook is the written procedure that brings one system back at the recovery site: who runs it, what must already be running, each step with its expected result and time budget, how the result is checked and signed off, and how to roll back or fail back. The DR plan above it sets when recovery starts, in which order systems return and who takes the decisions; the runbook says how, for one system, so that someone other than its author can follow it.
AWS’s Well-Architected Framework defines a runbook as “a documented process to achieve a specific outcome”. NIST SP 800-34 Rev. 1, still the current revision in October 2026, calls the per-system document an information system contingency plan, and its sample formats define three phases: activation and notification, recovery, and reconstitution, which includes activities “to test and validate system capability and functionality”.
For the entities listed in its Article 1, Implementing Regulation (EU) 2024/2690 point 4.1.2 lists the contents of the business continuity and disaster recovery plan, set out in our article on why backup is not disaster recovery. In our reading, per-system runbooks are the “recovery plans for specific operations, including recovery objectives” of its item (f). Other organisations in scope of NIS2 follow their national law, and which text applies is for the company’s legal department to assess.
DR runbook template: the sections and what goes in each
The fields of AWS’s runbook template map onto the twelve sections below, for example Special Permissions to preconditions and Escalation POC to contacts.
| SECTION | WHAT IT HOLDS | UPDATE WHEN |
|---|---|---|
| Purpose and scope | the system, its VMs, the outcome, what is out of scope | the application changes |
| Owner and deputies | runbook owner, named deputy, application owner who signs off | people change roles |
| Recovery objectives | RTO and RPO of the system’s tier | the impact analysis is reviewed |
| Preconditions | recovery site capacity, restore point age, accounts and tools that work without production | replication or accounts change |
| Dependencies | what must run first: identity, DNS, time, the database | an interface is added |
| Invocation | who may start it, on which plan decision, what is logged | the DR plan changes |
| Steps | one action each, with path or command, expected result, who, time budget | any version, address or configuration change |
| Verification and sign-off | checks per tier, a test transaction, who signs off | functions change |
| Rollback | abort criteria, how each stage is undone | the replication method changes |
| Failback | method, who decides, how changed data returns | the replication method changes |
| Contacts | team, deputies, vendors, provider, out-of-band channel | each review, and more often |
| Version and test history | changes with date, reason, author; tests with date, type, times, findings | every change and every test |
Our template, with fields from AWS’s runbook template (Well-Architected Framework, OPS07-BP03) and roles and review rules from NIST SP 800-34 Rev. 1, sections 3.4.6 and 3.6.
Invocation corresponds to NIST’s activation and notification phase, steps and rollback to recovery, and verification and failback to reconstitution. The objectives come from the business impact analysis, and the runbook repeats them so the operator can compare the clock with the target.
How to write runbook steps that someone else can follow
NIST SP 800-34 Rev. 1 asks the plan coordinator to consider “that a disruption could render some personnel unavailable to respond”, in which case the plan may need “personnel from another geographic area of the organization” or contractors and vendors. Write every step for that reader, a colleague from another site, a provider’s administrator or a deputy who has never run it; AWS describes runbooks as processes “that provide consistent outcomes no matter who uses them”.
Each step holds one action, given as the console path or command on the installed version, and the result that shows it worked. Where the result can differ, the step says what comes next: retry, use an older restore point or call the escalation contact it names. A time budget per step shows early when the recovery is running late.
The runbook names accounts and the vault holding their passwords, never the passwords, and the accounts must work while production is down: for the recovery site’s vCenter, an account in its own Single Sign-On domain, not one from the production directory service. NIST wants a copy of the plan stored “at the alternate site and with the backup media”; a runbook kept only in a wiki at the primary site, or a password vault that runs only there, is unavailable when that site is down.
Worked example: a three-tier application on vSphere
The example is an order-processing application on vSphere with two web servers behind a load balancer, two application servers and a database server. Two domain controllers provide the directory service, the internal DNS zone and the time source. The RTO is 4 hours and the RPO 30 minutes; the VMs replicate every 15 minutes to a recovery site that carries the same subnets, where a domain controller already runs and replicates with production.
| STEP | ACTION | EXPECTED RESULT | WHO | BUDGET (CUMULATIVE) |
|---|---|---|---|---|
| Invocation | log the declaration time and plan decision; call the deputy | owner and deputy on the call | runbook owner | 10 min (0:10) |
| Recovery site check | sign in with a break-glass account; check restore point times and capacity; confirm that the primary VMs are off or disconnected | newest point at most 30 minutes before the outage; no VM runs at both sites | VMware administrator | 15 min (0:25) |
| Identity, DNS and time | resolve the application’s names, test a sign-in, check the clock | names resolve; sign-in works; clocks agree | directory administrator | 20 min (0:45) |
| Database tier | fail over the database VM; note the last committed transaction | log recovery done; data loss within the RPO | database administrator | 30 min (1:15) |
| Application tier | fail over both application servers | health check passes; no database errors | application administrator | 20 min (1:35) |
| Web tier | fail over both web servers and the load balancer | HTTPS answers with a valid certificate | application administrator | 20 min (1:55) |
| User access | switch routing or DNS records as the DR plan specifies | a test client outside IT reaches the application | network administrator | 30 min (2:25) |
| Verification and sign-off | post a test order; check mail relay, partner interface, scheduled jobs | owner signs off; time recorded | application owner | 30 min (2:55) |
| Protection and hand-over | start backups of the recovered VMs; inform users | first backup running | VMware administrator | 15 min (3:10) |
Our example; the times are illustrative, not measurements, so use those recorded in your last test.
Identity, DNS and time come first because every later check depends on them. Kerberos, which the directory service uses for sign-ins, requires clocks that RFC 4120 calls “loosely synchronized”, with a tolerance “typically on the order of 5 minutes”, so Kerberos sign-ins from a server whose clock has drifted further fail. How users reach the recovery site is covered in our guide to failover and failback steps.
With these illustrative times, the owner signs off after 2 hours 55 minutes against an RTO of 4 hours. Time before the declaration counts against the RTO too, so the remaining 65 minutes have to cover detection and the decision.
The example covers the loss of the primary site. After ransomware, the replicas and the recovery site’s domain controller may carry the attacker’s changes, so a separate runbook restores into an isolated environment from points that predate the intrusion. It follows the order in our ransomware recovery guide: clean hosts and a new backup server first, then DNS and domain controllers, then vCenter, then the application tiers.
In our disaster recovery service, switchover on the day of a disaster follows a rehearsed scenario, with failback after the primary is restored. Describe the application you would write up first and its dependencies in the form below.
Verification, sign-off, rollback and failback
Checks run at three levels: the VM is up, the service answers, and the application works for a user. Veeam Recovery Orchestrator’s Verify DNS Port step “verifies the port used to connect to the recovered VM with the Domain Naming Service role”, port 53 by default, which shows that the port answers, not that the records are current. The runbook adds name lookups, a test order and interface calls, and the owner’s sign-off ends the recovery for the RTO.
The rollback section gives abort criteria per stage and the way back: an older restore point if the newest database copy does not start, a restore from backup if no replica point is clean, and a stop decision for whoever declared the disaster. Undoing a Veeam failover will “discard all changes made to the VM replica while it was running”, so the runbook names who may approve it.
In our reading, failback covers point 4.1.2(h) of the regulation, “restoring and resuming activities from temporary measures”. Veeam’s failback transfers “all changes that took place while the VM replica was running to the original VM”. In Live Site Recovery, a planned migration completes the recovery once the primary site is back, a reprotect reverses replication, and failback is a planned migration back followed by a second reprotect. The runbook names the method, who sets the date and how the reverse synchronisation is checked.
What Live Site Recovery and Veeam Recovery Orchestrator automate
In VMware Live Site Recovery, part of Protection and Recovery from VCF 9.1, the recovery priority “determines the shutdown and power-on order of virtual machines”, and dependencies “are only valid if the virtual machines have the same priority”, so the example gives its database a higher priority than its application servers. Custom steps “run commands or present messages to the user during a recovery”, and because a message waits for a response, the database check can become a pause in the plan. Broadcom’s VCF 9.1 FAQ of 3 September 2026 states that “Disaster recovery runbook orchestration requires standalone SRM or ACC entitlement”, meaning a Site Recovery Manager licence or Advanced Cyber Compliance. Broadcom also lets you export a plan’s steps “to keep a hard copy backup of your plans”.
Veeam Recovery Orchestrator 13 (user guide of 26 August 2026) builds plans from steps such as Check VM Heartbeat, Verify Domain Controller Port and Custom Script, and works with a Veeam Data Platform Premium licence or an Advanced licence with Orchestrator licences. By default it runs a daily readiness check of every enabled plan, confirming for example that “Replica VMs are detected and ready for failover” and “Required credentials are provided”. Its Plan Definition Report lists a plan’s steps and parameters, and Veeam says it “can be used to obtain a sign-off from application owners who need to verify plan configuration”. With either tool, invocation, changes by carriers, DNS providers or partners, business checks, the sign-off and messages to users stay in the runbook.
Keeping DR runbooks current with change management
AWS lists “Letting runbooks drift out of sync with system changes and automation” among the anti-patterns of OPS07-BP03, and NIST SP 800-34 Rev. 1 has plans reviewed and updated “as part of the organization’s change management process”. For the entities it covers, Implementing Regulation (EU) 2024/2690 has changes assessed “in view of the potential impact before being implemented” (point 6.4.2), and plans reviewed and, where appropriate, updated after significant incidents or significant changes to operations or risks (point 4.1.4).
- Add a field “recovery steps affected: yes or no” to every change request for a system with a runbook, checked by the runbook owner.
- Count as affected new or changed IP addresses, VLANs, DNS names, firewall rules, accounts and certificates, upgrades of vCenter, ESX, the database or the replication software, and new interfaces.
- Update the steps and contacts in the same change, and log it in the version history with date, reason and author.
- Run the changed steps in an isolated test before the change is closed.
- After each test or incident, record the time per step and the findings, and fix the runbook.
- Replace the copy at the recovery site with the new version.
Test types and the test report are covered in our article on disaster recovery testing.
Each scheduled failover test in our disaster recovery service ends with a report on what came up, how fast and what to fix. Tell us when your runbooks were last followed by someone other than their author.
What we do
Under our disaster recovery service, our engineering partner Vixen.UNO defines with you the critical systems, the target RPO and RTO for each tier and the disaster scenarios; the output is a continuity plan, not just copies. Virtual machines replicate from a 15-minute interval via Veeam Cloud Connect to a recovery site in Baltneta’s Tier-3 data centres in Lithuania, with scheduled failover tests in an isolated environment and a report after each. Switchover on the day of a disaster follows a rehearsed scenario; RPO and RTO are fixed in the SLA, and the first call is free of charge.
FAQ
What is a disaster recovery runbook?
What is the difference between a DR plan and a DR runbook?
What should a DR runbook template include?
How do you write a DR procedure for one system?
Is there a template for an IT disaster recovery plan?
How often should a disaster recovery runbook be updated?
Send us the application you would write a runbook for first, its VMs and dependencies, how it is protected today and the RTO and RPO it needs. We reply within one business day with a date for a first call, where we work through your critical systems, current backup and target RPO and RTO, and you leave with two or three possible DR scenarios. The first call is free of charge.
Talk to an expertWe reply within one business day