<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>systemsfailed.dev</title>
    <link>https://www.systemsfailed.dev/</link>
    <description>Real engineering postmortems, organised by how they failed.</description>
    <language>en</language>
    <atom:link href="https://www.systemsfailed.dev/feed.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Anthropic — Claude requests routed to the wrong context servers</title>
      <link>https://www.systemsfailed.dev/incident/anthropic-2025-context-window-routing</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/anthropic-2025-context-window-routing</guid>
      <pubDate>Sat, 08 Aug 2026 00:00:00 GMT</pubDate>
      <description>Degraded Sonnet 4 responses initially affected 0.8% of requests, reached 16% in the worst hour, and touched about 30% of active Claude Code users at least once.</description>
    </item>
    <item>
      <title>Cruise — The collision logic that chose to keep moving</title>
      <link>https://www.systemsfailed.dev/incident/cruise-2023-post-collision-pullover</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/cruise-2023-post-collision-pullover</guid>
      <pubDate>Sat, 08 Aug 2026 00:00:00 GMT</pubDate>
      <description>A driverless Cruise vehicle pulled a pedestrian forward after contact; Cruise paused its driverless fleet 24 days later and recalled the collision-detection software installed on 950 automated-driving systems.</description>
    </item>
    <item>
      <title>GitHub — A feature flag blocked every Copilot request</title>
      <link>https://www.systemsfailed.dev/incident/github-2025-copilot-rate-limiter</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/github-2025-copilot-rate-limiter</guid>
      <pubDate>Sat, 08 Aug 2026 00:00:00 GMT</pubDate>
      <description>Most GitHub Copilot features were degraded for 25 minutes, from 17:55 to 18:20 UTC, as the global rate limiter returned HTTP 403 responses for every request affected by the invalid configuration.</description>
    </item>
    <item>
      <title>Google — AI Overviews grounded answers in satire and trolls</title>
      <link>https://www.systemsfailed.dev/incident/google-2024-ai-overviews-data-voids</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/google-2024-ai-overviews-data-voids</guid>
      <pubDate>Sat, 08 Aug 2026 00:00:00 GMT</pubDate>
      <description>Some US Search users received odd, inaccurate, or unhelpful AI Overviews, including answers grounded in satire, sarcastic forum posts, and misleading user-generated advice, even though policy-violating responses were rare overall.</description>
    </item>
    <item>
      <title>OpenAI — The reward signal that made GPT-4o sycophantic</title>
      <link>https://www.systemsfailed.dev/incident/openai-2025-gpt4o-sycophancy</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/openai-2025-gpt4o-sycophancy</guid>
      <pubDate>Sat, 08 Aug 2026 00:00:00 GMT</pubDate>
      <description>ChatGPT became markedly over-agreeable, sometimes validating doubts, fueling anger, reinforcing negative emotions, or encouraging impulsive actions until OpenAI restored the previous GPT-4o version.</description>
    </item>
    <item>
      <title>AWS — Lambda Frontend Scaling Bug Breaks US-EAST-1</title>
      <link>https://www.systemsfailed.dev/incident/aws-2023-lambda-frontend-scaling</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/aws-2023-lambda-frontend-scaling</guid>
      <pubDate>Sat, 25 Jul 2026 00:00:00 GMT</pubDate>
      <description>A latent bug in Lambda&apos;s Frontend scaling logic degraded function invocations in one cell of US-EAST-1 for nearly four hours, causing elevated errors and latency across STS, Management Console, EKS cluster provisioning, Connect, and EventBridge.</description>
    </item>
    <item>
      <title>GitHub — Database Migration Cascades Into Site-Wide DNS Outage</title>
      <link>https://www.systemsfailed.dev/incident/github-2024-dns-database-migration</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/github-2024-dns-database-migration</guid>
      <pubDate>Sat, 25 Jul 2026 00:00:00 GMT</pubDate>
      <description>A single-site DNS outage degraded IDE code completions for 4% of Copilot users, delayed 25% of Actions workflows by over 5 minutes, and caused a complete code search outage for roughly 4 hours.</description>
    </item>
    <item>
      <title>GitHub — Expired internal SSL cert breaks Actions runner connectivity</title>
      <link>https://www.systemsfailed.dev/incident/github-2026-actions-runner-cert-expiry</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/github-2026-actions-runner-cert-expiry</guid>
      <pubDate>Sat, 25 Jul 2026 00:00:00 GMT</pubDate>
      <description>Self-hosted and larger GitHub Actions runners worldwide could not connect for nearly 5 hours, delaying or failing workflow jobs, while the resulting reconnection surge briefly degraded API Requests, Issues, and Pages with added latency and elevated error rates.</description>
    </item>
    <item>
      <title>GitHub — Public endpoints overwhelmed by abusive traffic</title>
      <link>https://www.systemsfailed.dev/incident/github-2026-abusive-traffic-504</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/github-2026-abusive-traffic-504</guid>
      <pubDate>Tue, 14 Jul 2026 00:00:00 GMT</pubDate>
      <description>Signed-out users experienced 504 errors (peaking at 34% error rate) for pull requests, issues, and releases across GitHub.com for two hours; authenticated users were unaffected.</description>
    </item>
    <item>
      <title>AWS — Kinesis cell manager mistakes healthy hosts for dead</title>
      <link>https://www.systemsfailed.dev/incident/aws-2024-kinesis-cell-management-shard-storm</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/aws-2024-kinesis-cell-management-shard-storm</guid>
      <pubDate>Fri, 10 Jul 2026 00:00:00 GMT</pubDate>
      <description>For nearly seven hours in US-EAST-1, an internal Kinesis Data Streams cell degraded, raising latencies and error rates across CloudWatch Logs, S3 event delivery, Data Firehose, ECS, Lambda, Redshift, and Glue, with log backlogs taking until the next day to clear.</description>
    </item>
    <item>
      <title>Basecamp — Primary key overflow on events table forces read-only mode</title>
      <link>https://www.systemsfailed.dev/incident/basecamp-2018-events-table-integer-overflow</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/basecamp-2018-events-table-integer-overflow</guid>
      <pubDate>Sat, 04 Jul 2026 00:00:00 GMT</pubDate>
      <description>Basecamp 3 was fully read-only for almost five hours, blocking all writes across the product — no new messages, todos, files, or edits — while reads continued to work.</description>
    </item>
    <item>
      <title>CircleCI — Database contention triggers day-long build queue collapse</title>
      <link>https://www.systemsfailed.dev/incident/circleci-2015-database-overload-build-queue-collapse</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/circleci-2015-database-overload-build-queue-collapse</guid>
      <pubDate>Sat, 04 Jul 2026 00:00:00 GMT</pubDate>
      <description>Linux builds were backed up or completely halted for roughly 17 hours across two days, affecting nearly all CircleCI customers&apos; CI/CD pipelines during peak usage.</description>
    </item>
    <item>
      <title>CircleCI — Backwards-incompatible schema change stalls all job execution</title>
      <link>https://www.systemsfailed.dev/incident/circleci-2021-backwards-incompatible-schema-change</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/circleci-2021-backwards-incompatible-schema-change</guid>
      <pubDate>Sat, 04 Jul 2026 00:00:00 GMT</pubDate>
      <description>All job executors were blocked from running customer jobs for roughly 53 minutes, with elevated queue times continuing for machine executors for over an hour longer, affecting all CircleCI customers.</description>
    </item>
    <item>
      <title>PagerDuty — Dual AWS region outage takes down notification dispatch</title>
      <link>https://www.systemsfailed.dev/incident/pagerduty-2013-dual-region-peering-outage</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/pagerduty-2013-dual-region-peering-outage</guid>
      <pubDate>Sat, 04 Jul 2026 00:00:00 GMT</pubDate>
      <description>Notification dispatch was completely down for 18 minutes after two of PagerDuty&apos;s three datacenters failed simultaneously; the events-ingestion API stayed up throughout.</description>
    </item>
    <item>
      <title>AWS — Config change killed Seoul region DNS for 84 minutes</title>
      <link>https://www.systemsfailed.dev/incident/aws-2018-seoul-dns-resolver</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/aws-2018-seoul-dns-resolver</guid>
      <pubDate>Sat, 27 Jun 2026 00:00:00 GMT</pubDate>
      <description>EC2 instances in the AWS AP-NORTHEAST-2 (Seoul) region were unable to resolve DNS for 84 minutes, affecting all customers running workloads in that region.</description>
    </item>
    <item>
      <title>Cloudflare — Dashboard bug overwhelmed auth API for 75 minutes</title>
      <link>https://www.systemsfailed.dev/incident/cloudflare-2025-tenant-api-thundering-herd</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/cloudflare-2025-tenant-api-thundering-herd</guid>
      <pubDate>Sat, 27 Jun 2026 00:00:00 GMT</pubDate>
      <description>Cloudflare&apos;s dashboard was fully unavailable and its APIs experienced two separate periods of severe degradation for roughly 75 minutes, affecting all customers attempting to make configuration changes.</description>
    </item>
    <item>
      <title>Datadog — Service discovery collapse took down all systems for 10 hours</title>
      <link>https://www.systemsfailed.dev/incident/datadog-2020-service-discovery-thundering-herd</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/datadog-2020-service-discovery-thundering-herd</guid>
      <pubDate>Sat, 27 Jun 2026 00:00:00 GMT</pubDate>
      <description>The US region of Datadog was degraded for roughly 10 hours on September 24–25, 2020, with the web tier hitting 60–90% error rates, leaving customers unable to reliably access dashboards, alerts, APM, or infrastructure monitoring.</description>
    </item>
    <item>
      <title>GitHub — Puppet bug corrupted DNS and cascaded into fileserver exhaustion</title>
      <link>https://www.systemsfailed.dev/incident/github-2014-dns-puppet-cascade</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/github-2014-dns-puppet-cascade</guid>
      <pubDate>Sat, 27 Jun 2026 00:00:00 GMT</pubDate>
      <description>GitHub.com was fully unavailable for 42 minutes on January 8, 2014, with an additional 95 minutes of degraded access for a subset of repositories while fileservers were triaged and restored.</description>
    </item>
    <item>
      <title>GitHub — ProxySQL file descriptor limit silently capped by system process manager</title>
      <link>https://www.systemsfailed.dev/incident/github-2020-proxysql-file-descriptor-exhaustion</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/github-2020-proxysql-file-descriptor-exhaustion</guid>
      <pubDate>Sat, 27 Jun 2026 00:00:00 GMT</pubDate>
      <description>GitHub.com experienced degraded service for a combined 8 hours 14 minutes across four events in late February 2020, with stalled writes on the mysql1 cluster affecting all authentication and core services that depended on it.</description>
    </item>
    <item>
      <title>incident.io — Database upgrade made incident IDs skip 32 numbers</title>
      <link>https://www.systemsfailed.dev/incident/incident-io-2021-postgres-sequence-skip</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/incident-io-2021-postgres-sequence-skip</guid>
      <pubDate>Sat, 27 Jun 2026 00:00:00 GMT</pubDate>
      <description>Multiple customer organizations saw their incident identifiers jump by up to 32 (e.g. from #INC-20 to #INC-52), eroding trust in the numbering system, though no data was lost or misattributed.</description>
    </item>
    <item>
      <title>incident.io — Unnecessary Slack transactions starved database connection pool</title>
      <link>https://www.systemsfailed.dev/incident/incident-io-2023-connection-pool-exhaustion</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/incident-io-2023-connection-pool-exhaustion</guid>
      <pubDate>Sat, 27 Jun 2026 00:00:00 GMT</pubDate>
      <description>incident.io&apos;s API experienced intermittent 20-second timeouts over two weeks in early 2023, causing customers to see request failures whenever connection pool slots were fully occupied.</description>
    </item>
    <item>
      <title>incident.io — GKE Dataplane V2 CPU starvation from concurrent TCP storms</title>
      <link>https://www.systemsfailed.dev/incident/incident-io-2023-gke-dataplane-tcp-storms</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/incident-io-2023-gke-dataplane-tcp-storms</guid>
      <pubDate>Sat, 27 Jun 2026 00:00:00 GMT</pubDate>
      <description>incident.io experienced intermittent Postgres and Memcache connection timeouts across production nodes for approximately one week following their migration to GKE, affecting API response reliability for all customers.</description>
    </item>
    <item>
      <title>incident.io — AWS us-east-1 outage cascaded through four hidden third-party dependencies</title>
      <link>https://www.systemsfailed.dev/incident/incident-io-2025-aws-outage-cascade</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/incident-io-2025-aws-outage-cascade</guid>
      <pubDate>Sat, 27 Jun 2026 00:00:00 GMT</pubDate>
      <description>incident.io&apos;s on-call notifications were delayed for over two hours, SAML authentication was intermittently unavailable, Scribe&apos;s AI note-taker was offline for ~5 hours, and the deployment pipeline was blocked — all driven by AWS us-east-1 dependency propagation through their telecom provider, auth provider, transcription provider, and Docker Hub.</description>
    </item>
    <item>
      <title>GitHub — Config change silenced database health checks for 36 minutes</title>
      <link>https://www.systemsfailed.dev/incident/github-2024-database-config-outage</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/github-2024-database-config-outage</guid>
      <pubDate>Fri, 26 Jun 2026 00:00:00 GMT</pubDate>
      <description>All GitHub.com services were fully inaccessible for all users worldwide for 36 minutes on the evening of August 14, 2024.</description>
    </item>
    <item>
      <title>Basecamp — Extortion DDoS took Basecamp down for 45 minutes</title>
      <link>https://www.systemsfailed.dev/incident/basecamp-2014-ddos-extortion</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/basecamp-2014-ddos-extortion</guid>
      <pubDate>Thu, 25 Jun 2026 00:00:00 GMT</pubDate>
      <description>Basecamp and its sibling services were fully unreachable for 45 minutes and intermittently degraded for nearly two hours, affecting all customers on a Monday morning.</description>
    </item>
    <item>
      <title>Allegro — A marketing campaign DDoSed itself</title>
      <link>https://www.systemsfailed.dev/incident/allegro-2018-scaling</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/allegro-2018-scaling</guid>
      <pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate>
      <description>Poland&apos;s largest e-commerce site went down at the moment a campaign was driving peak traffic.</description>
    </item>
    <item>
      <title>AWS — The autoscaler that fought the network</title>
      <link>https://www.systemsfailed.dev/incident/aws-2021-autoscale</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/aws-2021-autoscale</guid>
      <pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate>
      <description>A major us-east-1 disruption affecting a long list of AWS services and the customers built on them.</description>
    </item>
    <item>
      <title>AWS — A typo took out us-east-1</title>
      <link>https://www.systemsfailed.dev/incident/aws-s3-2017-typo</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/aws-s3-2017-typo</guid>
      <pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate>
      <description>S3 in us-east-1 down for hours; thousands of dependent sites and services degraded — including AWS&apos;s own status dashboard.</description>
    </item>
    <item>
      <title>Chef — Health checks too impatient to live</title>
      <link>https://www.systemsfailed.dev/incident/chef-2014-healthcheck</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/chef-2014-healthcheck</guid>
      <pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate>
      <description>The Supermarket community site crashed two hours after launch with intermittent unresponsiveness.</description>
    </item>
    <item>
      <title>CircleCI — When GitHub came back, the queue fell over</title>
      <link>https://www.systemsfailed.dev/incident/circleci-github-recovery</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/circleci-github-recovery</guid>
      <pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate>
      <description>Build throughput collapsed to roughly one transaction per minute; the backlog couldn&apos;t drain.</description>
    </item>
    <item>
      <title>Cloudflare — The regex that ate every CPU</title>
      <link>https://www.systemsfailed.dev/incident/cloudflare-2019-regex</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/cloudflare-2019-regex</guid>
      <pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate>
      <description>Global 502s across ~10% of the web — Discord, Shopify, and Cloudflare&apos;s own dashboard went dark.</description>
    </item>
    <item>
      <title>Cloudflare — A backbone change crushed 19 data centers</title>
      <link>https://www.systemsfailed.dev/incident/cloudflare-2022-backbone</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/cloudflare-2022-backbone</guid>
      <pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate>
      <description>19 of Cloudflare&apos;s busiest data centers dropped offline at once, hitting a large share of global traffic.</description>
    </item>
    <item>
      <title>Discord — A cloud live-migration tipped over Redis</title>
      <link>https://www.systemsfailed.dev/incident/discord-redis-failover</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/discord-redis-failover</guid>
      <pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate>
      <description>A real-time outage that forced a full service restart, reconnecting millions of clients over ~20 minutes.</description>
    </item>
    <item>
      <title>Fastly — One customer&apos;s config, the whole CDN down</title>
      <link>https://www.systemsfailed.dev/incident/fastly-2021-config</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/fastly-2021-config</guid>
      <pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate>
      <description>85% of Fastly&apos;s network returned errors within a minute — gov.uk, Reddit, Amazon and major news sites went down together.</description>
    </item>
    <item>
      <title>GitHub — 43 seconds that cost 24 hours</title>
      <link>https://www.systemsfailed.dev/incident/github-2018-splitbrain</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/github-2018-splitbrain</guid>
      <pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate>
      <description>A day of stale, read-only metadata; webhooks and Pages builds frozen; manual write reconciliation afterward.</description>
    </item>
    <item>
      <title>GitLab — rm -rf on the wrong database</title>
      <link>https://www.systemsfailed.dev/incident/gitlab-2017-rmrf</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/gitlab-2017-rmrf</guid>
      <pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate>
      <description>~300GB of production database gone; six hours of issues, merge requests and users lost for good.</description>
    </item>
    <item>
      <title>Knight Capital — $440M in 45 minutes</title>
      <link>https://www.systemsfailed.dev/incident/knight-2012-powerpeg</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/knight-2012-powerpeg</guid>
      <pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate>
      <description>A runaway trading algorithm lost roughly $460M in 45 minutes and effectively ended the firm.</description>
    </item>
    <item>
      <title>Meta — The day Facebook deleted itself from the internet</title>
      <link>https://www.systemsfailed.dev/incident/meta-2021-bgp</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/meta-2021-bgp</guid>
      <pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate>
      <description>Facebook, Instagram, WhatsApp gone globally for ~6 hours. Internal tools, badge access and conference systems were down too, slowing recovery.</description>
    </item>
    <item>
      <title>Roblox — 73 hours down via one coordination layer</title>
      <link>https://www.systemsfailed.dev/incident/roblox-2021-consul</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/roblox-2021-consul</guid>
      <pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate>
      <description>A near three-day full platform outage over Halloween weekend.</description>
    </item>
    <item>
      <title>Slack — The first-Monday thundering herd</title>
      <link>https://www.systemsfailed.dev/incident/slack-2021-herd</link>
      <guid isPermaLink="true">https://www.systemsfailed.dev/incident/slack-2021-herd</guid>
      <pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate>
      <description>Widespread connection failures on the first work-Monday of the year, exactly when everyone logged on at once.</description>
    </item>
  </channel>
</rss>
