A configuration error in AWS's billing system caused estimated bills to spike into the billions and trillions for over 24 hours. Although internal anomaly detection systems identified the issue, they failed to automatically halt bill generation or trigger engineer paging. The incident was only resolved after customer escalations alerted the company 4.5 hours later, during which time budget and cost anomaly alerts were disabled platform-wide.
- Internal anomaly detection failed to auto-remediate or page engineers despite clear billing spikes.
- Customer escalations were the primary driver for incident resolution, not automated systems.
- Budget and cost anomaly alerts were disabled platform-wide during the mitigation window.
- The billing configuration error persisted for over 24 hours before full resolution.