BLOG · GUIDE ·

DR site sizing: how much compute, storage, network and licences a recovery site needs

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • A DR site needs compute only for the tiers of systems that must run during the outage, as the business impact analysis sets them, and it is usually planned with less headroom than production
  • Replica storage holds each VM’s disks, the snapshot restore points the replication job keeps and everything written while systems run there: Veeam writes all changes made in the failover state to the delta file of the chosen restore point
  • In our worked example of 600 VMs in three tiers, the recovery site runs 300 VMs on 9 hosts against 16 in production, while it holds about 140 TB of data against 180 TB used in production
  • Broadcom’s VCF programme document of June 2026 requires every core of a server where the software is installed to be licensed, with at least 16 per processor, and allows recovery and testing in evaluation mode for up to 15 days per year in the aggregate, without rights to more cores than those paid for
  • Under Veeam’s licensing policy of June 2026 a protected VM consumes one licence instance, and target hosts of replication jobs need no licence under socket licensing

Eurokommerz × Vixen.UNO: Cloud Disaster Recovery  Talk to an expert →

How much capacity a DR site needs

A DR site, whether your own recovery data centre or DRaaS, needs enough compute for the systems that must run during the outage, not for the whole estate, and enough storage for the replicas, the restore points each replication job keeps and the data written while systems run there. Development, test and archive systems usually wait until the primary site is back, so the compute at the recovery site can be half or less of production. Storage shrinks less, because every replicated disk sits there in full and grows while you work at the recovery site.

Network, licences, always-on services and room for isolated tests complete the list. NIST SP 800-34 Rev. 1 states that “the facility must be able to support system operations as defined in the contingency plan”, so the plan and the business impact analysis behind it decide the size, not the production cluster.

What each type of site keeps ready is covered in our comparison of hot, warm and cold DR sites. This guide assumes a site where VMs wait as replicas, at your own second data centre or at a DRaaS provider.

Inputs for disaster recovery capacity planning

Collect these inputs before anyone draws a host count, and record the period they were measured in.

SIZING INPUTWHERE IT COMES FROMWHAT IT SIZES
Tier and RTO per systembusiness impact analysiswhich VMs start at once, later or never
CPU and RAM demandvCenter or monitoring data, month end includedhost count at the recovery site
Configured RAMVM inventoryRAM floor if memory is not overcommitted
Used and provisioned diskdatastore reportsbase capacity of the replicas
Daily change ratebackup and replication job statisticsrestore point deltas and the link
Restore points per jobreplication job settingssnapshot space per replica
Days at the recovery siteBIA and contractwrites held in delta files
Always-on servicesDR designcapacity used before any failover
Users and partnersBIA and access designinternet, VPN and public addresses
Test scheduleDR test plancapacity for isolated tests

Our checklist. Inputs from the business impact analysis follow NIST SP 800-34 Rev. 1, section 3.2; snapshot and delta file behaviour from Veeam’s Cloud Connect Guide (build 13.1.1.18) and Broadcom KB 318825.

The tier list is the input that removes the most capacity. Our guide to the business impact analysis for IT shows how process owners set the maximum tolerable downtime and how it becomes an RTO per system. A system with an RTO of 3 days that can be restored from a backup copy needs no CPU at the recovery site on the first day.

Compute: hosts for the tiers that run during the outage

Size CPU from the measured demand of the VMs that start rather than from their configured vCPU. vCenter keeps performance data per VM, by default at 5-minute intervals for the past day and 2-hour intervals for the past month, so take sustained demand from a month that includes a month end and short peaks from finer data. CPU can be overcommitted further at the recovery site than in production, because a few days of slower batch jobs are acceptable during a disaster. Broadcom describes memory as overcommitted when “the combined working memory footprint of all virtual machines” exceeds the host memory; ESX then reclaims memory through ballooning, memory sharing, compression and swapping; a VM whose memory is compressed or swapped to disk runs slower. For the top tier, plan the configured RAM of every VM that starts, with no overcommitment.

The recovery cluster also needs a reserve for a host failure. In Broadcom’s words, “vSphere HA uses admission control to ensure that sufficient resources are reserved for virtual machine recovery” when a host fails. Without that reserve, one host failure during a disaster stops part of the top tier a second time. Plan at least one host of reserve in the cluster that runs the top tier.

Some services run at the recovery site all the time, independent of any failover: domain controllers and DNS servers of the directory service, a vCenter, the backup or replication server of that site and a jump host for administrators. They use capacity every day and belong in the plan as a fixed amount.

Storage: replicas, restore points and writes during failover

A replica holds every disk of the VM. With thin-provisioned replica disks the base capacity is close to the used capacity of the source; with thick disks it is the provisioned size. On top come the restore points. In Veeam Cloud Connect, failover “reverts the VM replica to the necessary snapshot in the replica chain”, so each restore point is a vSphere snapshot whose delta file holds the blocks changed in its interval. The space for a replica’s restore points is therefore roughly the data changed over the time span they cover. Our comparison of Veeam replication and backup copy jobs explains how far back that chain reaches.

A storage item that is easy to miss comes from the failover itself. Veeam’s guide to partial site failover states that “All changes made to the VM replica while it is running in the Failover state are written to the delta file” of the restore point you rolled back to. Every write of production during the outage lands in delta files at the recovery site, until the failover is made permanent or failed back. Broadcom’s KB 318825 on snapshots states that “The snapshot file continues to grow in size the longer it is kept” and warns that this can make the storage location run out of space. Size that growth as the daily change rate of the started tiers multiplied by the days you plan to run there, an upper bound, since rewritten blocks are held once, and keep free space on the datastores above it.

If tier 3 is restored from backup copies held at the recovery site, the repository and the datastore those VMs are restored into need capacity of their own.

Network capacity for replication and user access

The replication link is sized from the data the busiest replication run transfers after compression, divided by the interval; our guide to replication bandwidth for disaster recovery covers that calculation. The second network item is the access path after a failover. Users, partners and interfaces reach the recovery site over the internet or a VPN that must not end at the primary site, and they need public addresses, firewall rules and enough internet capacity for the staff who work against the top tiers.

Count the networks the replicas connect to as well. A Veeam Cloud Connect hardware plan includes a “specified number of networks to which tenant VM replicas can connect”, and during full site failover Veeam starts a network extension appliance that acts as a gateway between the replica network and external networks.

Licences for hosts and VMs at the recovery site

Hosts you own at the recovery site run under the same licence terms as production hosts unless a vendor document says otherwise. Broadcom’s VMware Cloud Foundation Specific Program Documentation of June 2026 states that “Each Core on the Server where Software is installed must be licensed”, with a minimum of 16 core licences per processor. A recovery cluster with fewer or smaller hosts therefore needs fewer core licences, though never fewer than 16 per processor. The same document lets the customer perform recovery and testing activities “in the aggregate for up to 15 days per year” using the software’s embedded evaluation mode, and adds that this does not grant rights to more cores than those paid for. How that applies to your standby hosts is a question to put to Broadcom or your reseller in writing.

If you orchestrate failover with Broadcom’s Protection and Recovery (VMware Live Site Recovery before 9.1), its 9.1 documentation bases the licence on the number of protected VMs and states that “The same license is installed on both the protected and the recovery site.” Veeam’s licensing policy (version 19, June 2026) counts a protected workload as one with at least one restore point created in the past 31 days, and each protected VM consumes one instance. For socket licences, which Veeam no longer sells as new perpetual licences, it states that “Target hosts (for replication and migration jobs) do not need to be licensed.” Guest operating systems and applications follow their own vendors’ terms for standby and failover use.

Capacity for isolated failover tests

Failover tests start replicas on the recovery hosts, in a network isolated from production. Broadcom’s Protection and Recovery 9.1 documentation describes recovery to “a quarantined test network” and says testing “has no lasting effects on either the protected site or the recovery site”. A test of the full top tier needs the same CPU and RAM as a failover of that tier. If the site is sized only for the failover itself, test one tier at a time. Test runs also write to the replicas’ storage while they last, so plan a few hours of the tier’s change rate for them.

Our disaster recovery service runs scheduled failover tests in an isolated environment, with no impact on production systems and a report after each test. Tell us which tier you would test first and how much capacity your recovery site has today.

Worked example: a DR site for 600 VMs in three tiers

Take an invented company with 600 VMs on 16 production hosts, each with 64 cores and 1 TB of RAM. The BIA puts 90 VMs in tier 1, replicated every 15 minutes, and 210 VMs in tier 2, replicated hourly. The 300 VMs of tier 3 have backup copies only and are restored after the primary site returns.

ITEMTIER 1TIER 2TIER 3RECOVERY SITE
VMs90210300300 started, plus always-on services
Configured vCPU5408401,0201,380
Configured RAM2.7 TB4.2 TB5.1 TB6.9 TB plus 0.2 TB always-on
Used storage45 TB65 TB70 TB110 TB of replicas
Daily change rate3 percent2 percentnot replicated2.65 TB per day
Restore points kept28 every 15 minutes24 hourlynoneabout 2 TB of deltas

Invented example; the figures, the change rates and the arithmetic in the text are ours, not targets of our service.

The host count follows from RAM, since 7.1 TB on hosts of 1 TB filled to 90 percent needs 8 hosts, and one more as the HA reserve gives 9 hosts with 576 cores. With one host down, 1,380 vCPU run on 512 cores, 2.7 vCPU per core against 2.3 in production, which the tier 2 batch jobs can live with for a few days.

The storage plan depends on how long the business runs at the recovery site. The BIA and the contract assume up to 10 days at the recovery site, so 2.65 TB of writes per day adds about 27 TB of delta files. With 110 TB of replicas, about 2 TB of restore point deltas and 1 TB for test runs, the site holds about 140 TB of data, and with 20 percent of each datastore kept free it needs about 175 TB of usable capacity. The repository for tier 3’s backup copies is not part of this figure. The recovery site thus has 9 hosts against 16, while its 140 TB of data compare with 180 TB used in production.

At a DRaaS provider using Veeam Cloud Connect, the same result becomes a hardware plan with a CPU limit and a RAM limit for all replicated VMs of the tenant, a storage quota on a datastore and a number of networks. Whether that capacity is held for you alone is a contract question, covered in our DRaaS checklist before you sign.

In our Backup and DRaaS service, capacity availability, support terms and target RPO and RTO are fixed in the SLA; the technical assessment covers the site sizing. Send us your VM counts per tier, with configured RAM, used storage and replication schedules.

What we do

Our disaster recovery service runs your recovery site in Baltneta’s Tier-3 data centres in Lithuania (ISO 27001, PCI DSS), geographically separate from your primary infrastructure, with EU data residency. In the DR strategy design we define together with you the critical systems, the target RPO and RTO for each tier and the disaster scenarios you protect against, and the paid technical assessment covers the site sizing, with its price fixed before work begins. Our engineering partner Vixen.UNO sets up virtual-machine replication from a 15-minute interval via Veeam Cloud Connect and runs scheduled failover tests in an isolated environment, with a report after each one. The service’s standard targets are an RPO from 15 minutes and an RTO of 1 to 2 hours for critical systems; your targets per tier are defined during the assessment and fixed in the SLA, together with support.

FAQ

How do you size a disaster recovery site?
Start from the business impact analysis: list the systems that must run during the outage, then size compute from their measured CPU and RAM demand, storage from their used capacity, restore points and the writes expected while they run there, and the network from change rate and user access. Add licences, always-on services such as domain controllers and capacity for isolated failover tests. Systems that can wait until the primary site returns need storage for their backups but no compute.
Does a DR site need to be 1:1 with production?
A DR site rarely needs to match production one to one. Development, test and archive systems usually stay off during a disaster, and the remaining tiers can run with less CPU headroom for a limited time, so compute can be half or less of production, as in our worked example of 9 hosts against 16. Storage shrinks less, because each replicated disk sits at the recovery site in full and grows with every write made there during the failover.
How much storage do VM replicas need at the DR site?
Each replica needs the capacity of its disks, close to the used capacity with thin provisioning, plus the snapshot deltas of the restore points the replication job keeps. Veeam writes all changes made to a replica in the failover state to the delta file of the chosen restore point, so add the daily change rate of the started systems multiplied by the days you plan to run there, and keep free space on the datastores.
How do you size DRaaS capacity?
The same way as a site of your own: compute for the tiers that start, storage for replicas, restore points and failover writes, and the networks the replicas connect to. In Veeam Cloud Connect these become a hardware plan with a CPU limit, a RAM limit, a storage quota and a number of networks. Whether that capacity is reserved for you is decided by the contract.
Do hosts at a DR site need VMware licences?
Broadcom’s VCF programme document of June 2026 requires every core of a server where the software is installed to be licensed, with at least 16 core licences per processor, and names no exception for standby hosts. It allows recovery and testing in the embedded evaluation mode for up to 15 days per year in the aggregate, without rights to additional cores. How this applies to your standby hosts is a question to put to Broadcom or your reseller in writing.
Does Veeam need a licence for the DR site?
Veeam’s licensing policy of June 2026 counts protected workloads, VMs with at least one restore point created in the past 31 days, and each protected VM consumes one instance. Under socket licensing, target hosts of replication and migration jobs need no licence. Guest operating systems and applications follow their own vendors’ terms for standby and failover use.

Send us your VM inventory by tier, with configured vCPU, RAM and used storage, your replication schedules and how long the business would run at the recovery site. We reply within one business day to arrange the first call, in which we work through your critical systems, current backup and target RPO and RTO, and you leave with 2 to 3 possible DR scenarios. The first call is free of charge.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna