AWS/2017-02-28/~4h
A typo took out us-east-1
Jan 2017Dec 2017
S3 in us-east-1 down for hours; thousands of dependent sites and services degraded — including AWS's own status dashboard.
Lesson
Tooling should refuse to remove capacity below safe thresholds. And your status page must never depend on the thing that's down.
Trigger
While debugging a billing-system slowdown, an engineer ran a playbook command with a mistyped parameter, removing far more capacity than intended.
Mechanism
Two core S3 subsystems (indexing and placement) required a full restart. They hadn't been fully restarted in years and took hours to come back; everything depending on S3 in the region degraded with them.
Interview lens
Blast-radius control: guardrails on destructive ops, plus out-of-band status reporting.
Recurring patterns