Course map
3 tutorial parts
02 / 22

Three parts. One operating story.

The tutorial starts with strategy, proves it through testing and runbooks, then grounds it in public incidents.

A strong DR deck should move from why to how to evidence.

Part 1

Strategy

Disaster scope, HA vs DR, RPO/RTO, strategy spectrum, and selection logic.

Part 2

Proof

Testing ladders, runbooks, recovery order, communication, and real execution discipline.

Part 3

Reality

GitLab, AWS S3, Meta, and Knight Capital show what bad assumptions look like in public.

Narrative arc

1. Define toleranceHow long down, how much loss.
2. Choose architectureCold, light, warm, or active-active.
3. Protect dataBackups, replication, PITR, immutability.
4. Rehearse recoveryRunbooks, drills, measured restoration.
5. Learn from failureCase studies feed the next design loop.