HLD Group
Disaster recovery plan
Recovery from region loss, ransomware, and major outages.
Last updated: 24 July 2026
Version 1.0 · Review cycle: 365 days · View all frameworks
1. Purpose
This Disaster Recovery Policy defines how HLD Group restores technology services following a major disruption. It translates the recovery objectives set by the Business Continuity Policy into technical recovery capability, and works together with the Backup and Recovery Policy and the Incident Response Policy.
2. Scope
This policy applies to the production technology estate and the personnel, runbooks, and third-party dependencies required to recover it, including cloud infrastructure, data stores, networking, identity, and the tooling needed to rebuild services.
3. Definitions
- Disaster — a disruption causing, or likely to cause, an extended loss of a critical service beyond normal incident handling
- Failover — switching operations to standby infrastructure
- Failback — returning operations to primary infrastructure after recovery
- Warm standby — partially running standby capacity that can be brought to full service quickly
- Cold standby — provisioned but not running capacity requiring build-up before use
- Runbook — the documented, tested procedure for a specific recovery scenario
4. Policy statement
HLD Group maintains, documents, and tests the capability to recover critical technology services within their defined recovery time and recovery point objectives following a disaster. Recovery procedures name decision-makers, are executable under stress, and are validated through regular exercises rather than assumed to work.
5. Disaster scenarios
Recovery plans address a defined set of credible scenarios, each with a runbook naming decision-makers and technical steps.
- Loss of a cloud region or availability zone
- Ransomware or destructive compromise of production systems
- Failure or outage of a critical vendor or dependency
- Loss of the primary office or of physical access to it
- Large-scale data corruption or accidental destruction
- Loss of key personnel with rare operational knowledge
6. Recovery strategy and procedures
Recovery follows a documented order: assess and declare, communicate, execute technical recovery, validate service and data integrity, and then restore normal operations. Tier 1 services recover to warm or multi-region standby; lower tiers may use cold standby or rebuild-from-code approaches consistent with their recovery objectives.
- Declaration authority and activation criteria are defined and known to the on-call team
- Infrastructure is reconstitutable from version-controlled definitions and tested backups
- Recovery from a clean, known-good point is preferred where compromise is suspected
- Data integrity is validated before a recovered service is returned to customers
7. Communication during a disaster
Communication is a first-class part of recovery. Customers are notified within contractual timeframes through agreed channels, and internal status updates are issued to the response team on a regular cadence — at least every 60 minutes during an active disaster — using channels independent of the affected systems.
8. Testing and exercising
- Scenario-based recovery exercises for Tier 1 services at least every 18 months
- Validation of RTO and RPO achievement against the objectives set in the BIA
- Independent observation and documented findings tracked to closure
- Runbooks updated after every exercise and every real event
9. Post-recovery review
Following any disaster or major recovery exercise, a root cause analysis is completed within 10 business days. Findings, corrective actions, and improvements to strategy, tooling, and runbooks are tracked to closure and reported to leadership.
10. Framework alignment
- ISO/IEC 27031 (ICT readiness for business continuity)
- ISO 22301:2019 (business continuity management systems)
- ISO/IEC 27001:2022 Annex A control 5.30 (ICT readiness for business continuity)
- NIST SP 800-53 Rev. 5 control family CP (Contingency Planning)
- SOC 2 Trust Services Criteria availability category (A1.2, A1.3)
11. Roles, exceptions, and review
Disaster recovery is owned by the infrastructure and operations function under the CISO, with executive declaration authority. Exceptions require documented CISO approval with compensating controls and an expiry date. This policy is reviewed at least annually and after any disaster or major exercise.
Related frameworks
For contractual attestations or audit packs, contact [email protected].