systemsfailed.devfailure archive
← All incidents
Incident report
Slack/2021-01-04/~5h

The first-Monday thundering herd

Jan 2021Dec 2021

Widespread connection failures on the first work-Monday of the year, exactly when everyone logged on at once.

Lesson

Demand can outrun autoscaling. Pre-scale for known surges and make health checks resilient to transient loss so they don't amplify the failure.

A return-from-holiday traffic spike hit while Slack's AWS Transit Gateways were still scaling up.

Network saturation caused packet loss; that tripped health checks, which pulled servers from rotation, concentrating load on fewer machines — a downward spiral. Autoscaling lagged partly because the provisioning service itself was starved.

Prepare capacity for predictable spikes. Make health checks tolerate brief network failures so they do not remove healthy servers.

Share this failurePost on XDownload card