SRE Fundamentals — Review Deck
A fast-review, presenter-friendly companion to the 4-part SRE Fundamentals course: what SRE is, how SLOs work, how error budgets guide change, and how toil and postmortems close the loop.
Slides
SRE in one sentence
Cover slide for the course and the main framing line for SRE as a practical operating model.
What SRE actually changes
Shows the shift from manual operations to reliability managed through software, measurement, and policy.
SRE vs DevOps vs Platform
A comparison slide that separates philosophy, operating model, and platform-product responsibilities.
The reliability contract
Defines SLI, SLO, and SLA together and shows how measurement becomes an operating contract.
Why the nines matter
Connects reliability targets to downtime allowance and the non-linear cost of chasing extra nines.
Error budgets govern change
Turns reliability into a release decision system instead of a political argument between teams.
Toil is operational drag
Defines toil with the practical checklist SREs use to decide what should be automated away.
Blameless postmortems
Shows how incidents become systematic learning instead of person-focused blame and temporary fixes.
The full SRE loop
Connects measurement, budgets, engineering work, incidents, and learning into one operating cycle.
Fast recall: core terms
A compact review slide with the short phrases and distinctions that make the rest of the deck easier to retain.
Why Google created SRE
Places SRE in its original context: software engineers operating large production systems with explicit incentives.
The core principles behind the role
A slide on the principles that hold the course together: embracing risk, measuring users, automating toil, and learning fast.
What a good SLI looks like
Summarizes the design checklist for useful service indicators instead of noisy vanity metrics.
Where to measure the service
Compares client, edge, service, and dependency measurement points and what each one misses.
SLO windows and the math of nines
Connects 28-day windows, reliability targets, and why each extra nine sharply reduces room for error.
Burn rate tells you how fast you are losing the budget
Adds the fast operational metric behind budget-based alerting and policy decisions.
A simple error budget policy
Shows the minimal healthy structure for a policy that changes engineering behavior when reliability slips.
How teams actually reduce toil
Turns the definition of toil into a repeatable loop for prioritizing and removing operational drag.
From postmortem to deeper causes
Uses the five-whys style of thinking to move past symptoms and find the system conditions that made the incident likely.
Finish with the memory anchors
Final close with the short lines and distinctions that should survive after the full course or presentation.