Slack/2021-01-04/~5h
The first-Monday thundering herd
Jan 2021Dec 2021
Widespread connection failures on the first work-Monday of the year, exactly when everyone logged on at once.
Lesson
Demand can outrun autoscaling. Pre-scale for known surges and make health checks resilient to transient loss so they don't amplify the failure.
Trigger
A return-from-holiday traffic spike hit while Slack's AWS Transit Gateways were still scaling up.
Mechanism
Network saturation caused packet loss; that tripped health checks, which pulled servers from rotation, concentrating load on fewer machines — a downward spiral. Autoscaling lagged partly because the provisioning service itself was starved.
Interview lens
Health-check design under load, pre-warming for predictable spikes, and avoiding retry amplification.
Recurring patterns