DR site sizing: how much compute, storage, network and licences a recovery site needs
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- A DR site needs compute only for the tiers of systems that must run during the outage, as the business impact analysis sets them, and it is usually planned with less headroom than production
- Replica storage holds each VM’s disks, the snapshot restore points the replication job keeps and everything written while systems run there: Veeam writes all changes made in the failover state to the delta file of the chosen restore point
- In our worked example of 600 VMs in three tiers, the recovery site runs 300 VMs on 9 hosts against 16 in production, while it holds about 140 TB of data against 180 TB used in production
- Broadcom’s VCF programme document of June 2026 requires every core of a server where the software is installed to be licensed, with at least 16 per processor, and allows recovery and testing in evaluation mode for up to 15 days per year in the aggregate, without rights to more cores than those paid for
- Under Veeam’s licensing policy of June 2026 a protected VM consumes one licence instance, and target hosts of replication jobs need no licence under socket licensing
Eurokommerz × Vixen.UNO: Cloud Disaster Recovery Talk to an expert →
How much capacity a DR site needs
A DR site, whether your own recovery data centre or DRaaS, needs enough compute for the systems that must run during the outage, not for the whole estate, and enough storage for the replicas, the restore points each replication job keeps and the data written while systems run there. Development, test and archive systems usually wait until the primary site is back, so the compute at the recovery site can be half or less of production. Storage shrinks less, because every replicated disk sits there in full and grows while you work at the recovery site.
Network, licences, always-on services and room for isolated tests complete the list. NIST SP 800-34 Rev. 1 states that “the facility must be able to support system operations as defined in the contingency plan”, so the plan and the business impact analysis behind it decide the size, not the production cluster.
What each type of site keeps ready is covered in our comparison of hot, warm and cold DR sites. This guide assumes a site where VMs wait as replicas, at your own second data centre or at a DRaaS provider.
Inputs for disaster recovery capacity planning
Collect these inputs before anyone draws a host count, and record the period they were measured in.
| SIZING INPUT | WHERE IT COMES FROM | WHAT IT SIZES |
|---|---|---|
| Tier and RTO per system | business impact analysis | which VMs start at once, later or never |
| CPU and RAM demand | vCenter or monitoring data, month end included | host count at the recovery site |
| Configured RAM | VM inventory | RAM floor if memory is not overcommitted |
| Used and provisioned disk | datastore reports | base capacity of the replicas |
| Daily change rate | backup and replication job statistics | restore point deltas and the link |
| Restore points per job | replication job settings | snapshot space per replica |
| Days at the recovery site | BIA and contract | writes held in delta files |
| Always-on services | DR design | capacity used before any failover |
| Users and partners | BIA and access design | internet, VPN and public addresses |
| Test schedule | DR test plan | capacity for isolated tests |
Our checklist. Inputs from the business impact analysis follow NIST SP 800-34 Rev. 1, section 3.2; snapshot and delta file behaviour from Veeam’s Cloud Connect Guide (build 13.1.1.18) and Broadcom KB 318825.
The tier list is the input that removes the most capacity. Our guide to the business impact analysis for IT shows how process owners set the maximum tolerable downtime and how it becomes an RTO per system. A system with an RTO of 3 days that can be restored from a backup copy needs no CPU at the recovery site on the first day.
Compute: hosts for the tiers that run during the outage
Size CPU from the measured demand of the VMs that start rather than from their configured vCPU. vCenter keeps performance data per VM, by default at 5-minute intervals for the past day and 2-hour intervals for the past month, so take sustained demand from a month that includes a month end and short peaks from finer data. CPU can be overcommitted further at the recovery site than in production, because a few days of slower batch jobs are acceptable during a disaster. Broadcom describes memory as overcommitted when “the combined working memory footprint of all virtual machines” exceeds the host memory; ESX then reclaims memory through ballooning, memory sharing, compression and swapping; a VM whose memory is compressed or swapped to disk runs slower. For the top tier, plan the configured RAM of every VM that starts, with no overcommitment.
The recovery cluster also needs a reserve for a host failure. In Broadcom’s words, “vSphere HA uses admission control to ensure that sufficient resources are reserved for virtual machine recovery” when a host fails. Without that reserve, one host failure during a disaster stops part of the top tier a second time. Plan at least one host of reserve in the cluster that runs the top tier.
Some services run at the recovery site all the time, independent of any failover: domain controllers and DNS servers of the directory service, a vCenter, the backup or replication server of that site and a jump host for administrators. They use capacity every day and belong in the plan as a fixed amount.
Storage: replicas, restore points and writes during failover
A replica holds every disk of the VM. With thin-provisioned replica disks the base capacity is close to the used capacity of the source; with thick disks it is the provisioned size. On top come the restore points. In Veeam Cloud Connect, failover “reverts the VM replica to the necessary snapshot in the replica chain”, so each restore point is a vSphere snapshot whose delta file holds the blocks changed in its interval. The space for a replica’s restore points is therefore roughly the data changed over the time span they cover. Our comparison of Veeam replication and backup copy jobs explains how far back that chain reaches.
A storage item that is easy to miss comes from the failover itself. Veeam’s guide to partial site failover states that “All changes made to the VM replica while it is running in the Failover state are written to the delta file” of the restore point you rolled back to. Every write of production during the outage lands in delta files at the recovery site, until the failover is made permanent or failed back. Broadcom’s KB 318825 on snapshots states that “The snapshot file continues to grow in size the longer it is kept” and warns that this can make the storage location run out of space. Size that growth as the daily change rate of the started tiers multiplied by the days you plan to run there, an upper bound, since rewritten blocks are held once, and keep free space on the datastores above it.
If tier 3 is restored from backup copies held at the recovery site, the repository and the datastore those VMs are restored into need capacity of their own.
Network capacity for replication and user access
The replication link is sized from the data the busiest replication run transfers after compression, divided by the interval; our guide to replication bandwidth for disaster recovery covers that calculation. The second network item is the access path after a failover. Users, partners and interfaces reach the recovery site over the internet or a VPN that must not end at the primary site, and they need public addresses, firewall rules and enough internet capacity for the staff who work against the top tiers.
Count the networks the replicas connect to as well. A Veeam Cloud Connect hardware plan includes a “specified number of networks to which tenant VM replicas can connect”, and during full site failover Veeam starts a network extension appliance that acts as a gateway between the replica network and external networks.
Licences for hosts and VMs at the recovery site
Hosts you own at the recovery site run under the same licence terms as production hosts unless a vendor document says otherwise. Broadcom’s VMware Cloud Foundation Specific Program Documentation of June 2026 states that “Each Core on the Server where Software is installed must be licensed”, with a minimum of 16 core licences per processor. A recovery cluster with fewer or smaller hosts therefore needs fewer core licences, though never fewer than 16 per processor. The same document lets the customer perform recovery and testing activities “in the aggregate for up to 15 days per year” using the software’s embedded evaluation mode, and adds that this does not grant rights to more cores than those paid for. How that applies to your standby hosts is a question to put to Broadcom or your reseller in writing.
If you orchestrate failover with Broadcom’s Protection and Recovery (VMware Live Site Recovery before 9.1), its 9.1 documentation bases the licence on the number of protected VMs and states that “The same license is installed on both the protected and the recovery site.” Veeam’s licensing policy (version 19, June 2026) counts a protected workload as one with at least one restore point created in the past 31 days, and each protected VM consumes one instance. For socket licences, which Veeam no longer sells as new perpetual licences, it states that “Target hosts (for replication and migration jobs) do not need to be licensed.” Guest operating systems and applications follow their own vendors’ terms for standby and failover use.
Capacity for isolated failover tests
Failover tests start replicas on the recovery hosts, in a network isolated from production. Broadcom’s Protection and Recovery 9.1 documentation describes recovery to “a quarantined test network” and says testing “has no lasting effects on either the protected site or the recovery site”. A test of the full top tier needs the same CPU and RAM as a failover of that tier. If the site is sized only for the failover itself, test one tier at a time. Test runs also write to the replicas’ storage while they last, so plan a few hours of the tier’s change rate for them.
Our disaster recovery service runs scheduled failover tests in an isolated environment, with no impact on production systems and a report after each test. Tell us which tier you would test first and how much capacity your recovery site has today.
Worked example: a DR site for 600 VMs in three tiers
Take an invented company with 600 VMs on 16 production hosts, each with 64 cores and 1 TB of RAM. The BIA puts 90 VMs in tier 1, replicated every 15 minutes, and 210 VMs in tier 2, replicated hourly. The 300 VMs of tier 3 have backup copies only and are restored after the primary site returns.
| ITEM | TIER 1 | TIER 2 | TIER 3 | RECOVERY SITE |
|---|---|---|---|---|
| VMs | 90 | 210 | 300 | 300 started, plus always-on services |
| Configured vCPU | 540 | 840 | 1,020 | 1,380 |
| Configured RAM | 2.7 TB | 4.2 TB | 5.1 TB | 6.9 TB plus 0.2 TB always-on |
| Used storage | 45 TB | 65 TB | 70 TB | 110 TB of replicas |
| Daily change rate | 3 percent | 2 percent | not replicated | 2.65 TB per day |
| Restore points kept | 28 every 15 minutes | 24 hourly | none | about 2 TB of deltas |
Invented example; the figures, the change rates and the arithmetic in the text are ours, not targets of our service.
The host count follows from RAM, since 7.1 TB on hosts of 1 TB filled to 90 percent needs 8 hosts, and one more as the HA reserve gives 9 hosts with 576 cores. With one host down, 1,380 vCPU run on 512 cores, 2.7 vCPU per core against 2.3 in production, which the tier 2 batch jobs can live with for a few days.
The storage plan depends on how long the business runs at the recovery site. The BIA and the contract assume up to 10 days at the recovery site, so 2.65 TB of writes per day adds about 27 TB of delta files. With 110 TB of replicas, about 2 TB of restore point deltas and 1 TB for test runs, the site holds about 140 TB of data, and with 20 percent of each datastore kept free it needs about 175 TB of usable capacity. The repository for tier 3’s backup copies is not part of this figure. The recovery site thus has 9 hosts against 16, while its 140 TB of data compare with 180 TB used in production.
At a DRaaS provider using Veeam Cloud Connect, the same result becomes a hardware plan with a CPU limit and a RAM limit for all replicated VMs of the tenant, a storage quota on a datastore and a number of networks. Whether that capacity is held for you alone is a contract question, covered in our DRaaS checklist before you sign.
In our Backup and DRaaS service, capacity availability, support terms and target RPO and RTO are fixed in the SLA; the technical assessment covers the site sizing. Send us your VM counts per tier, with configured RAM, used storage and replication schedules.
What we do
Our disaster recovery service runs your recovery site in Baltneta’s Tier-3 data centres in Lithuania (ISO 27001, PCI DSS), geographically separate from your primary infrastructure, with EU data residency. In the DR strategy design we define together with you the critical systems, the target RPO and RTO for each tier and the disaster scenarios you protect against, and the paid technical assessment covers the site sizing, with its price fixed before work begins. Our engineering partner Vixen.UNO sets up virtual-machine replication from a 15-minute interval via Veeam Cloud Connect and runs scheduled failover tests in an isolated environment, with a report after each one. The service’s standard targets are an RPO from 15 minutes and an RTO of 1 to 2 hours for critical systems; your targets per tier are defined during the assessment and fixed in the SLA, together with support.
FAQ
How do you size a disaster recovery site?
Does a DR site need to be 1:1 with production?
How much storage do VM replicas need at the DR site?
How do you size DRaaS capacity?
Do hosts at a DR site need VMware licences?
Does Veeam need a licence for the DR site?
Send us your VM inventory by tier, with configured vCPU, RAM and used storage, your replication schedules and how long the business would run at the recovery site. We reply within one business day to arrange the first call, in which we work through your critical systems, current backup and target RPO and RTO, and you leave with 2 to 3 possible DR scenarios. The first call is free of charge.
Talk to an expertWe reply within one business day