Use cases
Software Products E-commerce MSPs Schools Development & Marketing DevOps Agencies Help Desk
Company
Internet Status Blog Pricing Log in Get started free

Outage in Blacksmith

Storage degradation in us-west

Resolved Minor
August 13, 2026 - Started 21 days ago - Lasted 1 day

Incident Report

Summary AI Generated

A thermal event (cooling failure) at Blacksmith's us-west datacenter provider caused storage degradation, taking down caching services including Sticky Disks, Actions Cache, Docker Container Cache, and Bazel Build Caching, while also triggering job queue delays across all regions as traffic was rebalanced away from us-west. Over the 30+ hour incident, the provider restored power and cooling in phases, with compute capacity recovering ahead of storage, and a concurrent GitHub webhooks outage compounded job pickup delays. All services were fully restored by the end of the incident, though some builds may run slower initially while caches rehydrate from the partial data loss during recovery.

Our datacenter provider has begun deploying mitigations at the affected us-west facility, and GitHub has resolved its separate webhooks incident, though some knock-on delays may persist while affected jobs requeue. The us-west region remains heavily affected and with capacity reduced while we rebalance traffic, jobs in all regions may take longer than usual to start.
Components affected
Blacksmith Github → Webhooks

Trusted by 1,000+ teams

The Status Page Aggregator with Early Outage Detection

Stop finding out about outages from your users. Monitor 6,320+ cloud services and get alerted the second something breaks.

IsDown status aggregator dashboard
Latest Updates ( sorted recent to last )
IDENTIFIED 21 days ago - at 08/13/2026 04:32PM

Our datacenter provider has begun deploying mitigations at the affected us-west facility, and GitHub has resolved its separate webhooks incident, though some knock-on delays may persist while affected jobs requeue. The us-west region remains heavily affected and with capacity reduced while we rebalance traffic, jobs in all regions may take longer than usual to start.

IDENTIFIED 21 days ago - at 08/13/2026 06:26PM

Our provider has begun powering servers back on at the us-west facility, and a portion of our us-west capacity is back online. We have started moving some customers back to us-west to spread load across regions. Jobs in all regions may still take longer than usual to start.

IDENTIFIED 21 days ago - at 08/13/2026 01:59PM

Conditions at our provider's us-west facility have not yet improved, and we are continuing to work around the issue to keep customer impact minimal. Such mitigations may include rebalancing customers to unaffected regions.

IDENTIFIED 21 days ago - at 08/13/2026 02:40PM

Conditions at our provider's us-west facility have not yet improved. We are continuing to issue mitigations to ensure that customer impact is minimal.

IDENTIFIED 21 days ago - at 08/13/2026 07:15PM

Our provider is continuing to bring us-west servers back online, and we are gradually moving customers back to us-west as capacity returns. The backlog has not yet cleared, so jobs in all regions may still take longer than usual to start. We will provide another update within the next hour.

IDENTIFIED 21 days ago - at 08/13/2026 08:44PM

Queues in us-west have come down substantially as restored capacity comes online and we continue moving customers back. Cache operations, including sticky disks, remain unavailable in us-west, and jobs in us-east may still take significantly longer than usual to start while we work through the remaining backlog. We will provide another update within the next hour.

IDENTIFIED 21 days ago - at 08/13/2026 05:00PM

Our provider has partially restored cooling at the us-west facility and expects to reach safe operating temperatures within the next few hours, at which point servers will be powered back on in phases. In the meantime jobs in all regions may still take significantly longer than usual to start. We will provide another update within the next hour.

IDENTIFIED 21 days ago - at 08/13/2026 12:06PM

Caching performance in US West is currently degraded. Our upstream provider has lost cooling power in their data center, leading to failure of some storage systems in US West. As a result jobs may take longer to complete. We are taking action to mitigate the issue.

INVESTIGATING 21 days ago - at 08/13/2026 11:40AM

Our cloud provider is experiencing a thermal event in their us-west datacenter that is affecting caching in the region. Sticky Disk, Actions Cache, Docker Container Cache, and Bazel Build Caching are impacted: cache operations may be slow or unavailable which may result in some builds running slower than usual.

IDENTIFIED 21 days ago - at 08/13/2026 07:59PM

More us-west capacity has come back online and queues in the region are steadily draining as we move customers back. Cache operations, including sticky disks, remain unavailable in us-west while our provider works to restore the storage portion of the facility, and because we shifted us-west traffic to other regions earlier today, jobs elsewhere, particularly in us-east, may still take longer than usual to start. We will provide another update within the next hour.

IDENTIFIED 21 days ago - at 08/13/2026 12:36PM

We are currently applying mitigations for affected customers in the region.

IDENTIFIED 21 days ago - at 08/13/2026 05:42PM

Our provider is continuing to restore cooling at the us-west facility; temperatures have not yet reached safe levels for servers to be powered back on, and we are staged to begin restoring runner capacity as soon as they are. The us-west region remains in a major outage, and jobs in all regions may still take longer than usual to start. We will provide another update within the next hour.

IDENTIFIED 21 days ago - at 08/13/2026 01:27PM

Conditions at our provider's us-west facility have not yet improved. As we rebalance customers away from us-west, other regions may see slightly longer pickup times than usual as a side effect

IDENTIFIED 21 days ago - at 08/13/2026 03:11PM

Conditions in our us-west region have not yet improved. We're actively working with our datacenter provider on the issue, and will provide updates as they occur. We are also manually rebalancing traffic out of the us-west region to aid with recovery.Github have also declared an incident affecting webhooks - u003chttps://www.githubstatus.com/incidents/k8vbzwqjkxznu003e which may cause some job adoption delays

IDENTIFIED 21 days ago - at 08/13/2026 01:02PM

Conditions at our provider's us-west facility have not yet improved, and we are continuing to work around the issue to keep customer impact minimal. Such mitigations may include rebalancing customers to unaffected regions.

IDENTIFIED 21 days ago - at 08/13/2026 09:48PM

Queue times in us-west have returned to near-normal levels, while us-east, eu-west, and eu-central are still working through their remaining backlogs. Cache operations, including sticky disks, remain unavailable in us-west while our provider works to restore the storage systems there. We will provide another update within the next hour.

IDENTIFIED 21 days ago - at 08/13/2026 10:54PM

Queue times are improving across all regions as capacity comes back online, but jobs may still take longer than usual to start. Cache operations in us-west, including sticky disks, remain unavailable while our provider restores the storage systems, and jobs on our largest runner sizes may see the longest delays. We will provide another update within the next hour.

IDENTIFIED 21 days ago - at 08/14/2026 12:00AM

Runner capacity in us-west continues to recover and queues in all regions are draining. Our provider has restored the access we need to begin bringing our us-west storage systems back online. Cache operations, including sticky disks, remain unavailable while that work completes, and jobs may still take longer than usual to start.

MONITORING 21 days ago - at 08/14/2026 01:15AM

Queue times for all runner sizes have returned to normal in all regions, and we are monitoring closely while eu-west clears the last of its backlog. Cache operations, including sticky disks, remain unavailable in us-west while we bring the restored storage hardware back online. We will provide another update within the next hour.

MONITORING 20 days ago - at 08/14/2026 03:06AM

Our us-west compute provider has restored power and cooling in their datacenter and the majority of our capacity has returned. Github Actions caching has been reenabled in the region but we are still working to restore Sticky Disks, Docker Container caching, and Incremental Docker Builders. We will post an update once these components are restored.

MONITORING 20 days ago - at 08/14/2026 10:57AM

We are continuing to work with our compute provider to restore the storage cluster for Sticky Disks, Docker Container Caching, and Incremental Docker Builders in us-west. This requires hands-on recovery work by our provider's team and may take a few more hours to fully restore. Workflows using the affected features in us-west may see slow or stalled Docker builds in the meantime.

MONITORING 21 days ago - at 08/14/2026 02:08AM

Job queues have fully recovered, and runners in all regions are operating normally for all runner sizes. Cache operations, including sticky disks, remain unavailable in us-west while we bring the restored storage hardware back online. We will provide another update within the next hour.

MONITORING 20 days ago - at 08/14/2026 03:26AM

We are noticing some issues with the storage cluster backing the Github Actions cache after it was restored. We are working on restoring this alongside the storage cluster backing Sticky Disks, Incremental Docker Builders, and Docker Container Caching.

MONITORING 20 days ago - at 08/14/2026 03:02PM

The storage cluster for Sticky Disks, Docker Container Caching, and Incremental Docker Builders in us-west has been restored, and disks are mounting and operating normally. As part of the recovery, a small amount of recently written cache data may need to be rebuilt, so some Docker builds may run slower over their first few runs while caches rehydrate. We are monitoring closely.

MONITORING 20 days ago - at 08/14/2026 04:09AM

We are still working with our compute provider to recover the storage clusters.

MONITORING 20 days ago - at 08/14/2026 06:06AM

We have restored the Github Actions Cache in us-west. We are still working with our compute provider to restore our remaining storage cluster for Sticky Disks, Docker Container Caching, and Incremental Docker Builders.

MONITORING 20 days ago - at 08/14/2026 04:16PM

All services have been restored, including sticky disks, Docker container caching, and incremental Docker builders in us-west, and job queues are operating normally in all regions. We are continuing to monitor the stability of the recovered storage cluster as it ramps back up with traffic, and some builds may run slower on their first runs while recently written cache data rebuilds. We will post a final update once we have confirmed stability.

The Status Page Aggregator with Early Outage Detection

With IsDown, you can monitor all your critical services' official status pages from one centralized dashboard and receive instant alerts the moment an outage is detected. Say goodbye to constantly checking multiple sites for updates and stay ahead of outages with IsDown.

Start free trial

No credit card required · Cancel anytime · 6320 services available

Integrations with Slack Microsoft Teams Google Chat Datadog PagerDuty Zapier Discord Webhook