Emergency site selection is one of the decisions that, in an emergency, either works quietly and without incident – or fails very publicly. On paper, „recovery site“ sounds like a second data center or a cloud location. In practice, however, it is about more: whether your critical business processes can be restarted within the required time under real disruption scenarios (fire, power outage, ransomware, failure of a network provider, supply chain issue, natural event, staff unavailability) without creating new compliance and security risks.
This article provides an assessable matrix for recovery sites – with criteria, weightings, minimum requirements („knock-out“ criteria) and a compliance check suitable to serve as evidence in audits. The target audience is IT management, security, compliance, risk management and executive management with an IT remit: roles that are responsible for decisions and later must explain why a site was chosen (or rejected).
Terms and target metrics: What exactly must the recovery site deliver?
Before you compare sites, you must define the target metric – otherwise you are comparing apples and oranges.
- RTO (Recovery Time Objective): the maximum tolerable time to restore operations. Example: „ERP must be available again within 8 hours.“ RTO is a guiding metric for architecture, automation, capacity provisioning and operational processes.
- RPO (Recovery Point Objective): the maximum tolerable data loss measured in time. Example: „maximum 15 minutes of data loss.“ RPO determines replication methods, backup intervals and data consistency.
- Workload classes: not every system requires the same recovery site. Typical classes are „Tier 1 (business critical)“, „Tier 2 (important)“, „Tier 3 (lower priority)“ – each defined with RTO/RPO, dependencies and data classification.
- Recovery site type:
- Hot Site: operation possible almost immediately, high cost, low RTO/RPO.
- Warm Site: basic components in place, activation requires lead time, medium cost.
- Cold Site: infrastructure/space available, systems must be built up first, inexpensive but high RTO.
Important: RTO/RPO are not IT wishlist figures, but must be derived from a Business Impact Analysis (BIA) (impact analysis on processes, revenue, regulatory obligations, reputation). In audits it is increasingly expected that RTO/RPO are traceably derived, approved and regularly reviewed.
Site selection is more than geography: common misconceptions
„Being far away is automatically better“
Geographic distance reduces shared risks (e.g. regional power outage). At the same time it often increases latency, dependence on carrier routes and operational complexity (e.g. separate admin teams, different supply chains). What matters is not „kilometres“ but the separation of risk domains: energy supply, network providers, water supply, hazard zones, political risks, supply chains, personnel availability.
„Cloud is automatically compliant“
Cloud regions can be operationally very robust, but they do not absolve you of responsibility. For compliance, data residency, access controls, logging, key management (e.g. HSM/KMS), subcontractor chains and exit capability are relevant. The recovery site must not only „run“, but be demonstrably controlled.
“Backup is sufficient as a recovery site”
Backups are the last line of defense, but not a recovery architecture. Without tested RESTore procedures, sufficient compute capacity, network paths, DNS/certificate handling and IAM (Identity & Access Management), the backup remains a data store – not business continuity. In ransomware scenarios there is an additional requirement: backups must be protected from tampering (e.g. immutable, separate admin accounts, air-gap approaches).
Evaluation matrix for recovery sites: structure, weighting, knock-out criteria
A practical matrix combines:
- Knock-out criteria (Must): If not met, the site is rejected regardless of points.
- Scoring criteria (Should/Can): Rated 0–5 or 0–10, weighted by relevance.
- Evidence: For each criterion it must be clear which evidence is accepted (contract, audit report, architecture diagram, test protocol, process description).
Example: scale and weighting
A 0–5 scale has proven effective (0 = not present, 3 = fulfilled, 5 = exceeds), together with weighting per category (e.g. 25% Operations/Resilience, 25% Security/Compliance, 20% Network/Connectivity, 15% Data/Platform, 15% Cost/Contract). Weighting should depend on risk appetite and process classes. For Tier‑1 workloads, “Cost” is rarely weighted higher than “recovery assurance”.
Knock-out criteria (Must) – realistic for many organizations
- Data residency & legal jurisdiction: Data categories must be permitted to be processed at the site (data protection, sector-specific law, customer contracts).
- RTO/RPO fundamentally achievable: Plausible based on architecture and capacity, not merely asserted.
- Separate administration domains: Ability to operate the recovery environment with separate identities/keys (important against ransomware and insider risks).
- Provable physical and logical security controls: Access control, segmentation, logging, patch/vulnerability management, incident processes.
- Contractual emergency services: Activation rights, priorities in crisis, test windows, clear SLAs/OLAs, exit rules.
Category 1: risk and site profile (geography, infrastructure, cascade effects)
For emergency site selection you should consider the site profile as a “risk package”.
Assessment questions
- Shared risk: Do primary and recovery sites share the same power or carrier dependencies (same substation network, same fiber route, same access structures)?
- Hazard zones: Is the site located in flood, earthquake, industrial, or high-risk zones? Not only historically, but according to current maps and local developments.
Audit perspective: The decisive factor is not that you eliminate every natural hazard, but that you have consciously chosen and documented: risk assumptions, countermeasures and residual risk.
Category 2: Technical recovery capability (RTO/RPO, capacity, automation)
This determines whether the recovery site merely „exists“ or is operationally viable.
Capacity model instead of gut feeling
A reliable site assessment requires a capacity model: Which workloads should run in an emergency, with which performance requirements (CPU/RAM/IOPS), for how long and with which dependencies (directory services, PKI, DNS, Time/NTP, message brokers, interfaces)? A common weakness: only the application is considered, not the ecosystem (identity, monitoring, logging, secrets, batch jobs, integration partners).
Assess RESTart mechanisms
- Replication (synchronous/asynchronous): Synchronous replication reduces RPO but requires low latency and stable links; asynchronous is more robust but allows a data gap.
- Backup/RESTore: RESTore times must be measured (not estimated). This includes database recovery, index rebuilds, object storage RESToration and application configuration.
- Infrastructure as Code (IaC): Automated provisioning reduces errors under stress. Governance matters: IaC must be versioned, approved and tested.
- Runbooks: Step sequences for failover and failback, with responsibilities, approvals and communication channels.
Copyable evidence block: Minimal DR test evidence
For audits, a standardized test protocol helps. Example as a template:
DR test protocol (short form)
1) Scope
- Workloads / process class:
- Target values: RTO ____, RPO ____
- Test type: Tabletop / partial test / full test / unannounced
2) Prerequisites
- Approvals (IT, business unit, Compliance):
- Change window:
- Communication channels / contacts:
3) Execution (timestamps)
- Start time of incident simulation:
- Failover trigger:
- Identity/Access active:
- Data state verified (RPO measurement point):
- Application verification (smoke tests):
- Integration partners validated:
- Monitoring/Logging active:
- Business acceptance:
4) Result
- Achieved RTO:
- Achieved RPO:
- Deviations / causes:
- Immediate measures:
- Action plan with owner and deadline:
5) Evidence attachments
- Architecture state (diagram version):
- Ticket IDs / change records:
- Measurements / monitoring screenshots:
- Approvals / acceptances:Category 3: Network, connectivity and „Failover der Realität“
Many recovery concepts do not fail on compute, but on network details: DNS, routing, certificates, IP address spaces, firewall rules, partner connections, MPLS/VPN/Direct-Connect variants.
Evaluation criteria
- Independent physical paths: Redundancy is only effective if it does not traverse the same physical route corridor (keyword „Diverse Routing“).
- Failover mechanics: DNS-TTL, Anycast, BGP failover or load balancer switching – each with clearly defined operational responsibility.
- Partner and third-party connections: Can critical interfaces (e.g. payment service providers, logistics, EDI) switch over to the Recovery-Site? Are there contractual testing options?
- Segmentation: Separation of admin, data and application networks, including emergency operation (e.g. RESTricted access, jump hosts, MFA).
Governance note: Document which network changes are permissible in an emergency without CAB (Change Advisory Board), which „Break-Glass“ approvals apply and how post-documentation is performed.
Category 4: Data, protection requirements and data residency
Recovery-Site selection is always also a data architecture decision. Typical conflicts arise between fast recovery and strict data classification.
Checkpoints for data and platform
- Data classification: Which data may go where? (personal, confidential, export-controlled, trade secrets). The Recovery-Site must permit processing of these classes – including admin access and support access.
- Encryption: At-REST and in-transit. Crucial is who controls the keys (customer keys vs. provider keys) and how key rotation and emergency access are governed.
- Consistency models: Databases, message queues, filesystems. RPO is meaningless if applications are in inconsistent states after recovery (e.g. double bookings, orphaned orders).
- Retention & WORM/Immutability: For certain data immutability (Write Once Read Many) may be relevant. Verify whether the Recovery-Site supports this operating mode.
Category 5: Security controls against ransomware and cross-cutting risks
A Recovery-Site that runs in the same security context as the primary environment will be quickly pulled along in the event of an identity or admin compromise. Good site selection therefore also means: plan for isolation.
Specific criteria
- Separate IAM / separate admin accounts: At minimum separate roles and strong MFA, ideally a separate directory domain or a clearly segmented identity zone.
- Break-Glass process: Emergency access with documented approval, strong logging and subsequent review/control.
- Logging & Forensics: Security logs must be centrally available even in emergency operation (SIEM integration, immutable/tamper-evident logs, time synchronization via NTP).
- Vulnerability & Patch Management: Recovery site must not „sit idle for years“ and start unpatched in an incident. At least monthly updates, plus proof of the baseline.
Copyable policy block: Break-Glass (short policy)
Break-Glass access (policy - short form)
Purpose:
- Enables recovery operations in case regular admin access is unavailable/compromised.
Rules:
- Separate emergency accounts, not for day-to-day use.
- MFA mandatory, hardware tokens preferred.
- Activation only after two-person authorization (IT management + Security/Compliance).
- Every use generates an incident ticket and is documented retroactively within 24 hours.
- Sessions are logged (command logging / session recording, where available).
- Emergency accounts are tested quarterly and passwords/secrets are rotated.
Evidence:
- Account list, roles, last tests, activation logs, review records.Category 6: Compliance check (audit-ready): Which evidence you should request early
The compliance part is often considered too late – by then contracts are signed, technical directions are set and evidence is missing. For an audit-ready choice of a recovery location it is essential to anchor requirements in procurement, contractual documentation and operations.
Regulatory reference points (not exhaustive)
- ISO 22301 (Business Continuity Management): requires, among other things, BIA, risk analysis, recovery strategies, exercises, documented procedures and continuous improvement.
- DORA (Digital Operational Resilience Act): relevant for many financial market participants; emphasizes resilience, testability, third-party risk and governance.
- BAIT/VAIT/KAIT (supervisory requirements in Germany, depending on sector): typical focus on emergency concepts, outsourcing, information security and evidence.
- KRITIS/NIS2 context: depending on applicability, additional requirements for security measures, reporting obligations and resilience.
- GDPR: processing, data processing agreements, TOM annex (technical and organizational measures), third-country transfers, deletion concepts.
Important: You do not have to „comply“ with every standard, but you must know which ones apply to your organization — and how your recovery site supports those requirements.
Compliance checklist for recovery sites (evidence-oriented)
- Contract & outsourcing
- Service description for emergency operation (activation, prioritization, capacity commitments).
- Rights to tests/exercises (frequency, scope, cost arrangements).
- Subcontractor transparency and consent, including chain of locations.
- Exit strategy: data return, deletion, migration, deadlines, support.
- Data protection
- Roles (controller/processor), DPA, TOM annex.
- Data residency and third-country transfer rules, access from third countries.
- Deletion and retention concept also in emergency operation (e.g., test data, log data).
- Security & operations
- Evidence of access control, monitoring, incident management, change management.
- Logging and retention (audit logs, admin activities).
- Regular DR tests with documented results and action plans.
Category 7: Day-to-day operations: Who does what in a real emergency?
Recovery sites are procured technically and then organizationally forgotten. At the latest during an audit or an incident this becomes apparent. The decisive factor is operational capability under stress: roles, authorities, communication and decision logic.
Roles and responsibilities (RACI logic)
- IT leadership: decision „Failover yes/no“, resource prioritization, escalation to executive management.
- Service owners / application owners: validation of smoke tests, dependencies, functional acceptance.
- Security: approval of emergency access, monitoring, indicators of compromise (to avoid starting an „infected“ environment).
- Compliance/risk management: documentation and proof obligations, communication to regulators/oversight bodies depending on context.
- Provider/contractor management: activation of contractual services, coordination of subcontractors.
Practical rule: If your recovery site can only achieve failover with „the two senior admins“, that is a risk. Plan for shift coverage, deputies and knowledge transfer. This is a criterion in site evaluation because it determines the operational consequences.
Category 8: Cost and contractual reality: What you need to represent in the TCO model
Costs are often reduced to infrastructure („second site = double hardware“). In reality, cost blocks arise in operations, tests, licenses, connectivity and governance.
Typical TCO components
- Capacity reservation: reserved compute, storage, database clusters, where applicable license costs in standby.
- Connectivity: redundant links, cross-connects, additional firewalls, DDoS protection.
- Test operations: DR exercises, maintenance windows, business unit effort, documentation.
- Security controls: SIEM integration, EDR, PAM (Privileged Access Management), session recording.
- Exit & portability: data export, migration tools, duplicated skills/training.
Evaluation matrix tip: Costs should not only be included as an absolute figure, but also as cost risk (e.g. unclear pricing logic in an emergency, „Best Effort“ clauses, fees per test, missing capacity guarantees).
How to implement the evaluation matrix in practice (implementation model in 6 steps)
- Define scope: process classes, workloads, data classes, dependencies, RTO/RPO.
- Set mandatory criteria: knock-out list including compliance.
- Catalog criteria: categories (risk/site, technology, network, data, security, operations, contract, costs).
- Decide weighting: by IT, security, compliance and business unit representation (per tier class).
- Request evidence: accepted evidence per criterion; missing evidence counts as „not met“.
- Review & decision: result, residual risk, action plan, timing for re-assessment (e.g. annually or upon site/provider changes).
Copyable matrix block: criteria list (short form to start)
Assessment matrix Recovery Site – Quick scoring (0–5 points, weight per category)
A) Location / Risk profile
- Independence of power/carriers
- Hazard zones / Shared Risks
- Accessibility / personnel availability
B) RTO/RPO capability
- Replication/backup procedures verifiable
- Capacity and scaling under emergency conditions
- Automation (IaC/Runbooks)
- Regular tests (frequency, scope, results)
C) Network & connectivity
- Diversity of routing paths
- Failover mechanics (DNS/BGP/load balancer)
- Partner connections switchable
- Segmentation & admin access
D) Data & data protection
- Data residency / legal jurisdiction
- Encryption & key ownership
- Retention/deletion also during emergency operations
E) Security
- Separate admin domains / break-glass
- Logging/forensics/SIEM
- Patch/vulnerability management
F) Governance/contract/costs
- Activation rights, SLAs/OLAs, testing rights
- Transparency of subcontractors/supply chain
- Exit strategy and data return
- Cost model (incl. tests, emergency operations)Audit perspective: What auditors typically want to see
Whether the audit is internal or external: good results come when you document decisions as a traceable chain. Typical audit points:
- Rationale: BIA and risk analysis lead to RTO/RPO and site strategy.
- Implementation: Architecture and operations concepts show how the objectives are achieved.
- Effectiveness: Tests/exercises demonstrate that it is not merely theoretical.
- Improvement: Deviations lead to measures with an owner and a deadline.
- Third parties: Outsourcing control and exit are regulated, not just „somewhere in the contract“.
If you want to dig deeper internally: a structured audit questionnaire helps collect evidence early and close gaps before the next audit or exercise.
Conclusion: A sound emergency site selection is a decision with a burden of proof
A Recovery Site is not a symbol of resilience, but a dependable operating mode. The best emergency site selection arises when technology, operations and compliance are evaluated together: with mandatory criteria, traceable scoring, clear evidence and regular tests. This makes the decision viable not only in crisis mode but also defensible in budget cycles and audits — including the consciously accepted residual risks.
The decisive step is usually not selecting „the perfect site“, but the consistent implementation logic: fully capture dependencies, plan isolation against cross-site attacks, contractually secure testing rights and translate results into governance. Then the Recovery Site is not only in place, but operational.
Recovery Site and Disaster Recovery are also important for this topic. This article places these aspects in context and shows what matters in day-to-day operations.