Isolated recovery environment: how to design a clean room for ransomware recovery
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- An isolated recovery environment (IRE), or cyber recovery clean room, is compute, storage and network the attacker never controlled, reached through one controlled entry point, where restore points are scanned and systems rebuilt before they return to production
- Broadcom’s best practices for the IRE in VMware Live Recovery Cloud (17 June 2026) say it is used for nothing else and that its VMs never get direct access to production; the outbound internet access EDR tools and patching need runs only through an NSX Tier-1 gateway
- Size it for tier-0 and tier-1 systems only: in a worked example of 1,000 VMs on 30 hosts, about 50 VMs run in the clean room at once, on four hosts including spare and forensic capacity
- Administer it with accounts outside the production directory: Veeam’s best practice guide puts backup components in a management workgroup or a management domain in a separate forest, with two-factor authentication
- Veeam’s clean room guidance says a backup repository must never be accessed by production and the clean room at the same time; some S3-compatible object storage allows the clean room limited read-only access instead
Eurokommerz × Vixen.UNO: Cyber Resilience Talk to an expert →
What an isolated recovery environment is
An isolated recovery environment (IRE), also called a cyber recovery clean room, is compute, storage and network that the attacker never controlled, reached through one controlled entry point. Restore points are scanned there and systems are rebuilt before they return to production. Broadcom’s documentation for VMware Live Recovery Cloud, updated on 1 September 2026, describes it as “a clean and secure network environment used specifically for recovery from ransomware attacks”. Commvault’s documentation for version 11.46 calls the same thing a cleanroom site.
The design has to settle where the IRE runs, how large it is, how it connects and who administers it. It also fixes how restore points get in, which tools wait inside and how a recovered system leaves. The order of work on the day of an attack is in our guide to the first 72 hours of ransomware recovery. This article covers the design that has to exist before that day.
| COMPONENT | DESIGN CHOICE | REASON |
|---|---|---|
| Location | spare hosts on their own segment, a second site, a DR site or cloud | the hosts the attacker used stay with the investigators |
| Size | tier-0 and tier-1 systems only | bulk restores go to production once it is clean |
| Network | no route to production, one entry point, outbound traffic through one filtered gateway | infected VMs cannot reach production, and production users cannot reach them |
| Administration | its own management network and a jump host inside the IRE | admin sessions do not cross the compromised network |
| Identity | admin accounts outside the production directory, with MFA | stolen production credentials do not open the IRE |
| Backup access | repository read by one side at a time, or read-only where the object storage allows it | Veeam does not support simultaneous access |
| Tools | malware scanning, EDR, forensic storage, golden images | each restore point is checked before use |
| Promotion | validated, staged, then moved to rebuilt production | only checked systems leave the clean room |
Design items after Broadcom’s Best Practices for the IRE (17 June 2026), Veeam’s clean room best practices, NIST SP 1800-11 and CISA’s #StopRansomware Guide; the middle column is our design summary.
Where to build it: location options
CISA’s #StopRansomware Guide lists “Retain backup hardware to rebuild systems if rebuilding the primary system is not preferred.” A Veeam blog post, updated on 5 August 2026, says “Clean rooms can be set up on any digital infrastructure, either on-premises or in the cloud.” Broadcom’s cloud service provisions an on-demand IRE on a recovery SDDC, which its documentation presents as an alternative to a physical, on-premises IRE. For a customer-owned IRE, Broadcom’s VCF blog of 14 May 2026 states that VCF 9.1 with the VMware Advanced Cyber Compliance add-on “allows you to build a fully customer-owned and managed clean room”. Commvault documents its cleanroom as Commvault-managed or self-managed, with cloud sites on AWS and one other public cloud, or fully on-premises.
Broadcom’s best practices state that the IRE is not used for any other purpose, and name test and development and burst capacity as examples. A DR site that must also carry a failover of production is therefore a weak fit unless its capacity covers both roles at once.
| LOCATION | WHAT IT NEEDS | STRENGTHS | LIMITS |
|---|---|---|---|
| Spare hosts, own segment | hosts, storage and switch ports outside production management | close to the repositories, fast restores | same building and power as the attacked site |
| Second own site | hosts and network at the second site | separate failure domain, under your control | hardware waits idle between tests |
| DR site | capacity beyond what failover needs, separate networks | hosts and copies already off site | failover may need the same hosts; replicas may hold the intruder |
| Public cloud on demand | the vendor’s service, data transfer from your copies | no idle hardware | restores cross a WAN link; data leaves your sites |
| Vendor-managed clean room | a contract with the backup vendor | the vendor operates it | depends on the vendor’s regions and service terms |
Options from Broadcom TechDocs (1 September 2026), Broadcom’s VCF blog (14 May 2026), Commvault documentation 11.46 and CISA’s #StopRansomware Guide; strengths and limits are our reading.
Our cyber resilience assessment reviews infrastructure, access and backups and ends with a risk map and a prioritised action plan. Tell us where your backup copies and spare hosts sit today and how many tier-0 and tier-1 systems you have.
Sizing the clean room for tier-0 and tier-1 systems
The IRE runs the systems that recovery depends on and the systems the business needs first. Everything else is restored into rebuilt production once the investigation allows it. Tier 0 covers DNS and domain controllers, a new backup server and its proxies, the jump host, the certificate authority and the scanning and EDR consoles. Tier 1 is the short list of business services with the tightest recovery time.
As a worked example, take a company with 1,500 staff and 1,000 VMs on 30 hosts, about 33 VMs per host. Its tier 0 has about 10 VMs and its tier 1, an ERP system with its database servers, file services and a production planning application, about 40. These 50 VMs fit on two hosts at the production ratio. A third host keeps recovery going when one fails, and a fourth carries scanning load and forensic analysis, which gives four hosts.
Storage holds the restored disks of these 50 VMs, room for a second restore point of each, because the first one checked may turn out to be infected, and space for forensic images. Broadcom’s IRE documentation lists access to snapshots “to enable rapid iteration of recovery points” for the same reason. The link to the repositories sets an upper limit on restore speed, and the repository’s read rate may set a lower one. At 10 Gbit/s the link moves at most about 4.5 TB an hour, so 20 TB of tier-0 and tier-1 disks need more than four hours of transfer alone.
Network design: no route to production, one entry point
Broadcom’s best practices for the IRE in VMware Live Recovery Cloud state that “VMs running within the IRE should never have direct access to the production environment” and that “No user machines at the production site should be able to directly access recovered VMs inside the IRE.” Broadcom recommends outbound internet access from the IRE, because most EDR and next-generation antivirus tools need cloud access for threat analysis and patches may come from public repositories, and states that it “should be allowed only through an NSX Tier-1 gateway and not routed elsewhere”. An IRE built without that service needs one filtered gateway at its edge for the same purpose. No inbound access from the internet is configured unless a specific requirement exists. Network isolation levels are set per VM, and Broadcom recommends using them when searching for hidden attack points. It also recommends keeping changes during recovery to a minimum: VMs keep their IP addresses, since changing them delays analysis, and sit behind NAT, so no address conflicts arise.
Recovered VMs can keep their production addresses because the IRE has no route to production and, in Broadcom’s design, its segments sit behind NAT. NIST SP 1800-11 recommends segregating the capabilities of its reference design, backup and corruption testing among them, on their own subnetwork, separate from the production network. CISA asks that a VLAN created for recovery receives only clean systems. In Veeam’s best practices for a Recovery Orchestrator clean room, network equipment between production and the clean room controls the isolation and allows connections for a limited time to transfer backup copies. Dedicated interfaces or switch ports are connected and disconnected on a schedule.
The IRE’s hosts are managed on their own network, by their own vCenter or standalone, never registered in the production vCenter. Administrators reach it through a jump host inside the IRE, from workstations that are not part of the production domain. Our guide to network segmentation and microsegmentation covers the production zones the IRE stays apart from.
Identity and administration inside the IRE
An IRE holds two identities. The first administers the IRE itself and exists before the incident. Veeam’s best practice guide says that for the most secure deployment backup components belong in a management workgroup or a management domain in a separate forest, with two-factor authentication. The aim is an infrastructure that “does not rely on the environment it is meant to protect”. CISA asks for administrator accounts separate from user accounts, and NIST SP 1800-11 asks to disable unnecessary predefined accounts and change default passwords on every component. Keep the IRE’s break-glass credentials offline and test them with each exercise.
The second identity is the restored production directory, which the recovered systems need. Broadcom recommends recovering a copy of production DNS in the clean room before other workloads, and a temporary copy of a domain controller for systems that use domain credentials. Commvault describes restoring and validating domain controllers in an isolated environment before dependent applications come back. Never use the restored directory to administer the IRE. How the directory itself is recovered is covered in our guide to domain controller recovery after ransomware.
Our cyber resilience service covers network segmentation, zero-trust access, multi-factor authentication and privileged access management. Describe how administrators reach your backup servers today in the form below.
Getting restore points into the clean room and checking them
In Veeam Recovery Orchestrator’s clean room scenario, backups are copied periodically from the production repository to a repository in the clean room. Recovery there works only from vSphere backups, not from Veeam agent backups. Veeam’s best practices warn that a repository “should never be simultaneously accessed by both the production and clean room”, because “this is not supported and may lead to data corruption”. The same guide adds that some S3-compatible object storage allows access from both sides by giving the clean room’s backup server limited read-only access to the bucket. How the copies themselves resist deletion is in our article on immutable backup.
Every restore point is checked before use. NIST SP 1800-11 determines the last known good state from logs and corruption testing. The Veeam blog post asks to “Always scan for malware” and notes that data can be validated against original hashes or signatures. If the clean room uses Veeam Threat Hunter, Veeam states that an internet connection is required for activation and for the latest signatures, one more reason for the filtered outbound gateway.
Keep the rebuild material inside the IRE too. CISA recommends maintaining golden images of critical systems, and keeping infrastructure-as-code templates under version control with offline backups. Forensic images of affected systems go to separate storage in the IRE, so investigators work on copies while restores continue.
The clean room workflow: scan, restore, validate, promote
In Broadcom’s VCF blog of 14 May 2026, VCF 9.1 with the Advanced Cyber Compliance add-on first places protected VMs in a validation state inside the IRE, where snapshots are scanned and cleaned with EDR sensors. Verified VMs move to a staging state to create a clean replica, then to a recovered state in the secondary site, before reprotection and failback. A comparable sequence works on other platforms, for each system:
- Take the candidate restore point from the investigation’s timeline and scan it with the backup software before restoring.
- Restore it into a quarantine segment with per-VM isolation, network adapters disconnected or on their own segment.
- Inspect it with EDR, compare files with known hashes, remove accounts and services the attacker added and patch the vulnerability used.
- Let the system owner validate the service with a test transaction against the tier-0 systems inside the IRE.
- Promote it to the staging network or rebuilt production after sign-off, and record which restore point was used.
Testing the IRE before an incident
Commvault’s documentation suggests using the cleanroom to simulate recovery scenarios and validate cyber recovery plans. Test the parts that daily operation never uses. Check that the jump host and break-glass accounts still work and that the scanning tools update through the filtered gateway. Make sure a domain controller and the backup server restore and start, and that no DNS record or directory change reaches production. Time each step against the recovery time of tier 1. Our guide to disaster recovery testing sets out isolated test methods and what to measure.
What we do
Under Cyber Resilience, our engineering partner Vixen.UNO segments the network and sets up zero-trust access with multi-factor authentication and privileged access management. It also provides Veeam-based backup with a recovery site in Baltneta’s Tier-3 data centres in Lithuania, regular test restores and an incident response plan with roles, actions and deadlines. Our disaster recovery service adds Veeam replication from a 15-minute interval and scheduled failover tests in an isolated environment, with a report after each test on what came up, how fast and what to fix. Eurokommerz holds the contract, the first call is free of charge and the price of the technical assessment is fixed before work begins.
FAQ
What is an isolated recovery environment?
How do you design a cyber recovery clean room?
Where should an isolated recovery environment run?
How big should a ransomware recovery environment be?
What is a cyber recovery vault?
Can the DR site be used as a clean room?
Send us the number of VMs and hosts, your tier-0 and tier-1 systems, where your backup copies are stored and how administrators sign in to the backup servers and hypervisor management today. We reply within one business day to arrange a first call, from which you leave with 2 to 3 possible solution scenarios. The first call is free of charge.
Talk to an expertWe reply within one business day