Buildkite experienced a major service disruption caused by an autoscaling feedback loop that overwhelmed their Redis cluster with connections. This resulted in widespread impact across the Web UI, Agent API, REST API, and job queue — affecting job scheduling, artifact uploads, and generating elevated error rates — in two windows: 18:43–19:21 UTC and 20:02–20:28 UTC. Full recovery was confirmed by 20:28 UTC, with a post-incident review to follow later that week.
Trusted by 1,000+ teams
Stop finding out about outages from your users. Monitor 6,320+ cloud services and get alerted the second something breaks.
We have seen full recovery for customers since 20:28 UTC.
We experienced an autoscaling feedback loop which increased the number of connections to our redis cluster above its ability to respond. This had widespread impact for all of our customers with Web UI, Agent API, REST API and job queue impact between 18:43-19:21 UTC, and again between 20:02-20:28 UTC. A full post incident review will be available later this week.
We are seeing improvements across the affected services, and are seeing services return to normal functionality. We are continuing to monitor and are determining the root cause.
We are continuing to investigate elevated error rates across multiple services, and are working to determine the cause.
We are continuing to investigate this issue. We are seeing impact on the Agent API which will affect job scheduling, artifact uploads, and an increase in 5xx responses from the Agent API endpoints.
We've spotted that something has gone wrong. We're currently investigating the issue, and will provide an update soon.
With IsDown, you can monitor all your critical services' official status pages from one centralized dashboard and receive instant alerts the moment an outage is detected. Say goodbye to constantly checking multiple sites for updates and stay ahead of outages with IsDown.
Start free trialNo credit card required · Cancel anytime · 6320 services available
Integrations with