systemsfailed.devfailure archive
← All incidents
Incident report
AWS/2017-02-28/~4h

A typo took out us-east-1

Jan 2017Dec 2017

S3 in us-east-1 down for hours; thousands of dependent sites and services degraded — including AWS's own status dashboard.

Lesson

Tooling should refuse to remove capacity below safe thresholds. And your status page must never depend on the thing that's down.

While debugging a billing-system slowdown, an engineer ran a playbook command with a mistyped parameter, removing far more capacity than intended.

Two core S3 subsystems (indexing and placement) required a full restart. They hadn't been fully restarted in years and took hours to come back; everything depending on S3 in the region degraded with them.

Blast-radius control: guardrails on destructive ops, plus out-of-band status reporting.

Share this failurePost on XDownload card