systemsfailed.devfailure archive
← All incidents
Incident report
Slack/2021-01-04/~5h

The first-Monday thundering herd

Jan 2021Dec 2021

Widespread connection failures on the first work-Monday of the year, exactly when everyone logged on at once.

Lesson

Demand can outrun autoscaling. Pre-scale for known surges and make health checks resilient to transient loss so they don't amplify the failure.

A return-from-holiday traffic spike hit while Slack's AWS Transit Gateways were still scaling up.

Network saturation caused packet loss; that tripped health checks, which pulled servers from rotation, concentrating load on fewer machines — a downward spiral. Autoscaling lagged partly because the provisioning service itself was starved.

Health-check design under load, pre-warming for predictable spikes, and avoiding retry amplification.

Share this failurePost on XDownload card