Use cases
Software Products E-commerce MSPs Schools Development & Marketing DevOps Agencies Help Desk
Company
Internet Status Blog Pricing Log in Get started free

Outage in Blacksmith

Delays in job adoption

Resolved Minor
July 21, 2026 - Started 4 days ago - Lasted about 9 hours

Incident Report

Summary AI Generated

Blacksmith experienced a 9-hour control plane degradation triggered by a burst of delayed GitHub webhooks, which saturated the primary Redis instance and created a feedback loop that caused widespread caching, sticky disk, and job adoption failures across all regions. Jobs ran significantly slower than baseline, compounding queue buildup and delaying new job starts. The incident was resolved by redistributing load across multiple Redis instances, fixing Bazel caching, and draining regional job queues, with all regions returning to normal job adoption latencies by the end of the incident.

We are continuing to investigate degradation in our control plane services. This is causing a large % of caching and stickydisk requests to fail, which in turn cause customer jobs to run much slower than their baseline. The slower runs are exacerbating queueing of new jobs that are coming in to the system. We are actively investigating and rolling out fixes to unblock customer jobs.

Trusted by 1,000+ teams

The Status Page Aggregator with Early Outage Detection

Stop finding out about outages from your users. Monitor 6,320+ cloud services and get alerted the second something breaks.

IsDown status aggregator dashboard
Latest Updates ( sorted recent to last )
IDENTIFIED 3 days ago - at 07/21/2026 05:49PM

We are continuing to investigate degradation in our control plane services. This is causing a large % of caching and stickydisk requests to fail, which in turn cause customer jobs to run much slower than their baseline. The slower runs are exacerbating queueing of new jobs that are coming in to the system. We are actively investigating and rolling out fixes to unblock customer jobs.

INVESTIGATING 3 days ago - at 07/21/2026 06:39PM

We have rolled out fixes to our workload to bring database query latency back to baseline. The degradation has now moved to other parts of our stack. We are continuing to investigate the root cause here. Customers can still expect to see degraded cache interactions. We will provide an update within the next 30 minutes.

INVESTIGATING 3 days ago - at 07/21/2026 07:03PM

We have applied a change to our backend systems and metrics are showing partial improvement, though not yet back to baseline. Customers can still expect degraded cache interactions while we continue to investigate the root cause. We will provide an update within the next 30 minutes.

MONITORING 3 days ago - at 07/21/2026 07:38PM

Our primary Redis instance, which backs our control plane, hit a saturation point, leading to a feedback loop of load. The initial cause was a burst of deliveries of delayed webhooks from GitHub, and our reconciliation systems added further load to the control plane, making things worse. We have improved load balancing by spreading this workload across multiple Redis instances, and our control plane has fully recovered. The remaining impact is a backlog of queued jobs that we are actively draining, so some customers may still see delayed job starts until the queue clears. We will continue monitoring and update with our findings in the next 30 minutes.

MONITORING 3 days ago - at 07/21/2026 07:56PM

Cache and sticky disk operations are healthy again. Bazel caching is not yet operational but being actively investigated by our team. The remaining impact is a backlog of queued jobs that we are actively draining, so some customers may still see delayed job starts until the queue drains.

MONITORING 3 days ago - at 07/21/2026 09:25PM

Job adoption in us-west and eu-west have returned to normal. Customers in eu-central may still see delayed job starts while we clear the remaining queue backlog. We have identified the cause of the Bazel caching issue and are now implementing a fix. We will provide another update within the next hour.

MONITORING 3 days ago - at 07/21/2026 10:30PM

We are still watching over the job queue draining in one of our regions (eu-central). Customers in this region will see a delay for their jobs to be picked up.

MONITORING 3 days ago - at 07/21/2026 10:59PM

We are seeing job adoption return to normal latencies across all regions.

MONITORING 3 days ago - at 07/21/2026 11:19PM

We are currently scanning for any missed jobs during this degradation period and ensuring that they are being dispatched.

IDENTIFIED 4 days ago - at 07/21/2026 02:25PM

We have applied a mitigation for the Caching and Monitors degradation. We are seeing some delays in webhook processing on our control plane which we are investigating.

IDENTIFIED 4 days ago - at 07/21/2026 03:44PM

We are continuing to investigate the delays in webhook processing on our control plane. Customer's may also see some degradation with caching and sticky disk components.

The Status Page Aggregator with Early Outage Detection

With IsDown, you can monitor all your critical services' official status pages from one centralized dashboard and receive instant alerts the moment an outage is detected. Say goodbye to constantly checking multiple sites for updates and stay ahead of outages with IsDown.

Start free trial

No credit card required · Cancel anytime · 6320 services available

Integrations with Slack Microsoft Teams Google Chat Datadog PagerDuty Zapier Discord Webhook