BLOG · GUIDE ·

Disaster recovery runbook: how to write the steps for one system, with a template and an example

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • A disaster recovery runbook is the per-system procedure under the DR plan: purpose, owner and deputies, recovery objectives, preconditions, dependencies, invocation, steps with expected results and time budgets, verification and sign-off, rollback, failback, contacts and a version and test history
  • NIST SP 800-34 Rev. 1 gives a system contingency plan three phases, activation and notification, recovery and reconstitution; for the entities it covers, Implementing Regulation (EU) 2024/2690 point 4.1.2(f) lists recovery plans for specific operations, including recovery objectives, among the plan contents
  • Each step is written for someone other than its author, since NIST notes that a disruption can leave some personnel unavailable and AWS describes runbooks as processes with consistent outcomes no matter who uses them
  • The worked example recovers a three-tier application on vSphere after the loss of the primary site, with an RTO of 4 hours: identity, DNS and time first, then the database, application and web tiers and user access; with illustrative step times, the owner signs off after 2 hours 55 minutes
  • Recovery plans in VMware Live Site Recovery and Veeam Recovery Orchestrator automate the VM start order, checks and prompts, while invocation, external network changes and the business sign-off stay in the runbook, which every change request should be checked against

Eurokommerz × Vixen.UNO: Cloud Disaster Recovery  Talk to an expert →

What a disaster recovery runbook is and how it differs from the DR plan

A disaster recovery runbook is the written procedure that brings one system back at the recovery site: who runs it, what must already be running, each step with its expected result and time budget, how the result is checked and signed off, and how to roll back or fail back. The DR plan above it sets when recovery starts, in which order systems return and who takes the decisions; the runbook says how, for one system, so that someone other than its author can follow it.

AWS’s Well-Architected Framework defines a runbook as “a documented process to achieve a specific outcome”. NIST SP 800-34 Rev. 1, still the current revision in October 2026, calls the per-system document an information system contingency plan, and its sample formats define three phases: activation and notification, recovery, and reconstitution, which includes activities “to test and validate system capability and functionality”.

For the entities listed in its Article 1, Implementing Regulation (EU) 2024/2690 point 4.1.2 lists the contents of the business continuity and disaster recovery plan, set out in our article on why backup is not disaster recovery. In our reading, per-system runbooks are the “recovery plans for specific operations, including recovery objectives” of its item (f). Other organisations in scope of NIS2 follow their national law, and which text applies is for the company’s legal department to assess.

DR runbook template: the sections and what goes in each

The fields of AWS’s runbook template map onto the twelve sections below, for example Special Permissions to preconditions and Escalation POC to contacts.

SECTIONWHAT IT HOLDSUPDATE WHEN
Purpose and scopethe system, its VMs, the outcome, what is out of scopethe application changes
Owner and deputiesrunbook owner, named deputy, application owner who signs offpeople change roles
Recovery objectivesRTO and RPO of the system’s tierthe impact analysis is reviewed
Preconditionsrecovery site capacity, restore point age, accounts and tools that work without productionreplication or accounts change
Dependencieswhat must run first: identity, DNS, time, the databasean interface is added
Invocationwho may start it, on which plan decision, what is loggedthe DR plan changes
Stepsone action each, with path or command, expected result, who, time budgetany version, address or configuration change
Verification and sign-offchecks per tier, a test transaction, who signs offfunctions change
Rollbackabort criteria, how each stage is undonethe replication method changes
Failbackmethod, who decides, how changed data returnsthe replication method changes
Contactsteam, deputies, vendors, provider, out-of-band channeleach review, and more often
Version and test historychanges with date, reason, author; tests with date, type, times, findingsevery change and every test

Our template, with fields from AWS’s runbook template (Well-Architected Framework, OPS07-BP03) and roles and review rules from NIST SP 800-34 Rev. 1, sections 3.4.6 and 3.6.

Invocation corresponds to NIST’s activation and notification phase, steps and rollback to recovery, and verification and failback to reconstitution. The objectives come from the business impact analysis, and the runbook repeats them so the operator can compare the clock with the target.

How to write runbook steps that someone else can follow

NIST SP 800-34 Rev. 1 asks the plan coordinator to consider “that a disruption could render some personnel unavailable to respond”, in which case the plan may need “personnel from another geographic area of the organization” or contractors and vendors. Write every step for that reader, a colleague from another site, a provider’s administrator or a deputy who has never run it; AWS describes runbooks as processes “that provide consistent outcomes no matter who uses them”.

Each step holds one action, given as the console path or command on the installed version, and the result that shows it worked. Where the result can differ, the step says what comes next: retry, use an older restore point or call the escalation contact it names. A time budget per step shows early when the recovery is running late.

The runbook names accounts and the vault holding their passwords, never the passwords, and the accounts must work while production is down: for the recovery site’s vCenter, an account in its own Single Sign-On domain, not one from the production directory service. NIST wants a copy of the plan stored “at the alternate site and with the backup media”; a runbook kept only in a wiki at the primary site, or a password vault that runs only there, is unavailable when that site is down.

Worked example: a three-tier application on vSphere

The example is an order-processing application on vSphere with two web servers behind a load balancer, two application servers and a database server. Two domain controllers provide the directory service, the internal DNS zone and the time source. The RTO is 4 hours and the RPO 30 minutes; the VMs replicate every 15 minutes to a recovery site that carries the same subnets, where a domain controller already runs and replicates with production.

STEPACTIONEXPECTED RESULTWHOBUDGET (CUMULATIVE)
Invocationlog the declaration time and plan decision; call the deputyowner and deputy on the callrunbook owner10 min (0:10)
Recovery site checksign in with a break-glass account; check restore point times and capacity; confirm that the primary VMs are off or disconnectednewest point at most 30 minutes before the outage; no VM runs at both sitesVMware administrator15 min (0:25)
Identity, DNS and timeresolve the application’s names, test a sign-in, check the clocknames resolve; sign-in works; clocks agreedirectory administrator20 min (0:45)
Database tierfail over the database VM; note the last committed transactionlog recovery done; data loss within the RPOdatabase administrator30 min (1:15)
Application tierfail over both application servershealth check passes; no database errorsapplication administrator20 min (1:35)
Web tierfail over both web servers and the load balancerHTTPS answers with a valid certificateapplication administrator20 min (1:55)
User accessswitch routing or DNS records as the DR plan specifiesa test client outside IT reaches the applicationnetwork administrator30 min (2:25)
Verification and sign-offpost a test order; check mail relay, partner interface, scheduled jobsowner signs off; time recordedapplication owner30 min (2:55)
Protection and hand-overstart backups of the recovered VMs; inform usersfirst backup runningVMware administrator15 min (3:10)

Our example; the times are illustrative, not measurements, so use those recorded in your last test.

Identity, DNS and time come first because every later check depends on them. Kerberos, which the directory service uses for sign-ins, requires clocks that RFC 4120 calls “loosely synchronized”, with a tolerance “typically on the order of 5 minutes”, so Kerberos sign-ins from a server whose clock has drifted further fail. How users reach the recovery site is covered in our guide to failover and failback steps.

With these illustrative times, the owner signs off after 2 hours 55 minutes against an RTO of 4 hours. Time before the declaration counts against the RTO too, so the remaining 65 minutes have to cover detection and the decision.

The example covers the loss of the primary site. After ransomware, the replicas and the recovery site’s domain controller may carry the attacker’s changes, so a separate runbook restores into an isolated environment from points that predate the intrusion. It follows the order in our ransomware recovery guide: clean hosts and a new backup server first, then DNS and domain controllers, then vCenter, then the application tiers.

In our disaster recovery service, switchover on the day of a disaster follows a rehearsed scenario, with failback after the primary is restored. Describe the application you would write up first and its dependencies in the form below.

Verification, sign-off, rollback and failback

Checks run at three levels: the VM is up, the service answers, and the application works for a user. Veeam Recovery Orchestrator’s Verify DNS Port step “verifies the port used to connect to the recovered VM with the Domain Naming Service role”, port 53 by default, which shows that the port answers, not that the records are current. The runbook adds name lookups, a test order and interface calls, and the owner’s sign-off ends the recovery for the RTO.

The rollback section gives abort criteria per stage and the way back: an older restore point if the newest database copy does not start, a restore from backup if no replica point is clean, and a stop decision for whoever declared the disaster. Undoing a Veeam failover will “discard all changes made to the VM replica while it was running”, so the runbook names who may approve it.

In our reading, failback covers point 4.1.2(h) of the regulation, “restoring and resuming activities from temporary measures”. Veeam’s failback transfers “all changes that took place while the VM replica was running to the original VM”. In Live Site Recovery, a planned migration completes the recovery once the primary site is back, a reprotect reverses replication, and failback is a planned migration back followed by a second reprotect. The runbook names the method, who sets the date and how the reverse synchronisation is checked.

What Live Site Recovery and Veeam Recovery Orchestrator automate

In VMware Live Site Recovery, part of Protection and Recovery from VCF 9.1, the recovery priority “determines the shutdown and power-on order of virtual machines”, and dependencies “are only valid if the virtual machines have the same priority”, so the example gives its database a higher priority than its application servers. Custom steps “run commands or present messages to the user during a recovery”, and because a message waits for a response, the database check can become a pause in the plan. Broadcom’s VCF 9.1 FAQ of 3 September 2026 states that “Disaster recovery runbook orchestration requires standalone SRM or ACC entitlement”, meaning a Site Recovery Manager licence or Advanced Cyber Compliance. Broadcom also lets you export a plan’s steps “to keep a hard copy backup of your plans”.

Veeam Recovery Orchestrator 13 (user guide of 26 August 2026) builds plans from steps such as Check VM Heartbeat, Verify Domain Controller Port and Custom Script, and works with a Veeam Data Platform Premium licence or an Advanced licence with Orchestrator licences. By default it runs a daily readiness check of every enabled plan, confirming for example that “Replica VMs are detected and ready for failover” and “Required credentials are provided”. Its Plan Definition Report lists a plan’s steps and parameters, and Veeam says it “can be used to obtain a sign-off from application owners who need to verify plan configuration”. With either tool, invocation, changes by carriers, DNS providers or partners, business checks, the sign-off and messages to users stay in the runbook.

Keeping DR runbooks current with change management

AWS lists “Letting runbooks drift out of sync with system changes and automation” among the anti-patterns of OPS07-BP03, and NIST SP 800-34 Rev. 1 has plans reviewed and updated “as part of the organization’s change management process”. For the entities it covers, Implementing Regulation (EU) 2024/2690 has changes assessed “in view of the potential impact before being implemented” (point 6.4.2), and plans reviewed and, where appropriate, updated after significant incidents or significant changes to operations or risks (point 4.1.4).

  1. Add a field “recovery steps affected: yes or no” to every change request for a system with a runbook, checked by the runbook owner.
  2. Count as affected new or changed IP addresses, VLANs, DNS names, firewall rules, accounts and certificates, upgrades of vCenter, ESX, the database or the replication software, and new interfaces.
  3. Update the steps and contacts in the same change, and log it in the version history with date, reason and author.
  4. Run the changed steps in an isolated test before the change is closed.
  5. After each test or incident, record the time per step and the findings, and fix the runbook.
  6. Replace the copy at the recovery site with the new version.

Test types and the test report are covered in our article on disaster recovery testing.

Each scheduled failover test in our disaster recovery service ends with a report on what came up, how fast and what to fix. Tell us when your runbooks were last followed by someone other than their author.

What we do

Under our disaster recovery service, our engineering partner Vixen.UNO defines with you the critical systems, the target RPO and RTO for each tier and the disaster scenarios; the output is a continuity plan, not just copies. Virtual machines replicate from a 15-minute interval via Veeam Cloud Connect to a recovery site in Baltneta’s Tier-3 data centres in Lithuania, with scheduled failover tests in an isolated environment and a report after each. Switchover on the day of a disaster follows a rehearsed scenario; RPO and RTO are fixed in the SLA, and the first call is free of charge.

FAQ

What is a disaster recovery runbook?
A disaster recovery runbook is the step-by-step procedure that brings one system back at the recovery site, written so that someone other than its author can follow it. It lists the preconditions and dependencies, each step with its expected result and time budget, how the recovered service is verified and signed off, and how to roll back or fail back. The DR plan above it decides when recovery starts and in which order systems return.
What is the difference between a DR plan and a DR runbook?
The DR plan covers recovery across the organisation: when it is activated and by whom, the order in which systems return and the resources they need. A runbook covers one system, with the exact steps, checks and timings to bring it back and to return it to the primary site later. NIST SP 800-34 Rev. 1 calls the per-system document an information system contingency plan, with activation and notification, recovery and reconstitution phases.
What should a DR runbook template include?
A DR runbook template should include purpose and scope, owner and deputies, recovery objectives, preconditions, dependencies, invocation, steps with expected results, owners and time budgets, verification and sign-off, rollback, failback, contacts and a version and test history. AWS’s runbook template adds fields such as special permissions, the author, the last update and an escalation contact. Credentials are never written into the runbook, which names the account and a vault that holds them and can be opened from the recovery site.
How do you write a DR procedure for one system?
Start from the system’s recovery objectives and dependencies, then write one action per step with the console path or command, the result that shows it worked, what to do if it does not, who does it and how long it may take. Have someone other than the author follow it in an isolated test, and rewrite every step that needed help. Keep a current copy at the recovery site, where it can be read when the primary site is down.
Is there a template for an IT disaster recovery plan?
NIST SP 800-34 Rev. 1 publishes sample contingency plan templates for low, moderate and high impact systems, as an appendix and as separate files on NIST’s website, together with a business impact analysis template. For the entities it covers, Implementing Regulation (EU) 2024/2690 point 4.1.2 lists what the business continuity and disaster recovery plan includes, from purpose and roles to the order of recovery and the return from temporary measures. Per-system runbooks then sit under that plan, one for each critical system.
How often should a disaster recovery runbook be updated?
Update it whenever the system, its dependencies or its recovery path changes, which is why the check belongs in change management, and after every test or incident that showed a gap. NIST SP 800-34 Rev. 1 asks for reviews at an organisation-defined frequency or whenever significant changes occur, with contact lists reviewed more often. Implementing Regulation (EU) 2024/2690 requires the entities it covers to test, review and, where appropriate, update their plans at planned intervals and after significant incidents or significant changes to operations or risks.

Send us the application you would write a runbook for first, its VMs and dependencies, how it is protected today and the RTO and RPO it needs. We reply within one business day with a date for a first call, where we work through your critical systems, current backup and target RPO and RTO, and you leave with two or three possible DR scenarios. The first call is free of charge.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry – see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna