Disaster Recovery & Business Continuity — Executive Review Deck
A rebuilt light-theme presentation for the Disaster Recovery & Business Continuity course: business framing, RPO/RTO, strategy choices, data protection, runbooks, testing, failback, and real case studies.
Slides
Disaster recovery in one line
High-level opening slide that reframes DR as the business capability that makes catastrophic failure survivable.
How the course fits together
Maps the three tutorial parts into one visual narrative: strategy, testing, and real incidents.
What actually counts as a disaster
Expands disaster beyond physical events to corruption, deletion, ransomware, control-plane loss, and regional failure.
HA, DR, and business continuity
Separates prevention, restoration, and full-business continuity responsibilities.
RPO and RTO
Visual explanation of the two core recovery numbers with a time window and memory anchors.
Why recovery speed gets expensive fast
Shows the non-linear cost jump as you move toward tighter RTO and RPO.
The four strategy spectrum
Positions backup-and-restore, pilot light, warm standby, and active-active on one visual rail.
Strategy 1 — backup and restore
Explains the coldest, cheapest option with rebuild-then-restore flow and fit criteria.
Strategy 2 — pilot light
Shows a light DR region with live data core but minimal application capacity.
Strategy 3 — warm standby
Shows a smaller but complete standby environment that can scale up quickly.
Strategy 4 — active-active
Shows fully live multi-region operation and the distributed-systems cost it brings.
How senior teams choose
Decision flow from business impact to service tiering to data design to testing cadence.
Reference Kubernetes DR architecture
Shows a realistic platform stack with Kubernetes, database, object storage, GitOps, backups, and observability.
Data protection is the center
Separates replication, backups, snapshots, and PITR so the audience sees why Terraform alone is insufficient.
Testing maturity ladder
Visual ladder from tabletop to full failover and drills, with what each level proves.
Runbook anatomy
The operational checklist of trigger, authority, commands, validation, communication, and failback.
Dependencies and recovery order
Shows why restore sequence and communication tooling are part of the service, not side concerns.
People and failback are failure domains
Adds knowledge concentration, approval design, and failback hazards to the technical plan.
Case study — GitLab 2017
Accidental deletion plus broken backup assumptions became the canonical restore-testing lesson.
Case study — AWS S3 us-east-1
A region event that exposed hidden dependencies, control-plane risk, and tooling guardrail gaps.
Case studies — Meta and Knight Capital
Shows that DR and business continuity include network control-plane loss and software deployment disasters.
Readiness checklist
Final review board with the minimum conditions for a trustworthy DR capability.