A major infrastructure failure affecting multiple racks and nodes in DigitalOcean's MKC1 region caused an 8.5-hour outage impacting GPU workloads, Kubernetes (DOKS) worker nodes, Serverless Inference, and briefly disrupting CPU Droplets, managed databases, load balancers, block storage, and Spaces. Most services were restored within approximately 8 hours, with the regional control plane and non-GPU resources confirmed healthy. GPU Droplets remained partially impacted beyond incident resolution, with affected GPU customers receiving direct follow-up via Slack and email.
Trusted by 1,000+ teams
Stop finding out about outages from your users. Monitor 6,320+ cloud services and get alerted the second something breaks.
Our Engineering team has confirmed the regional control plane is healthy and CPU Droplets, Managed databases, Load balancers, Block storage, Kubernetes (DOKS) and Spaces are operating normally.
GPU Droplets continue to be impacted and our teams are working to restore all nodes. We will communicate with GPU Droplet customers separately via Slack and email with more information and regular updates.
Thank you for your patience throughout this incident. If you continue to experience any issues, please open a support ticket from within your account.
At this time, the regional control plane is fully healthy. CPU Droplets, Managed databases, Load balancers, Block storage and Spaces are operating normally. Droplet create, resize and other management operations are working. Kubernetes (DOKS) control planes are reachable.
Our teams are continuing to work with the facility to restore all equipment. GPU Droplets in MKC1 remain offline or unreachable. DOKS GPU worker nodes may remain NotReady, and GPU-backed inference endpoints in this region may be unavailable.
We will monitor the regional control plane for a short time and then resolve this incident. GPU customers will receive personalized updates with more information in lieu of this status page.
If you have questions about your affected resources, contact support and reference this incident.
Our Engineering team continues to work on the issue affecting the MKC1 region. We are actively working to restore connectivity and bring the impacted nodes back online.
We will provide another update as soon as we have more information.
We have identified the root cause of the issue in the MKC1 region affecting multiple racks and nodes. Our Engineering team is actively implementing remediation steps to restore connectivity and bring the impacted nodes back online.
During this time, customers may continue to experience disruptions to GPU workloads and Serverless Inference. Kubernetes (DOKS) worker nodes may also remain in a NotReady state, and customers may be unable to reach Kubernetes API endpoints or perform cluster-management operations in affected clusters. We will provide another update as soon as we have more information.
Our Engineering team is currently investigating an issue in the MKC1 region affecting multiple racks and nodes.
During this time, customers may experience disruptions to GPU workloads, and Kubernetes (DOKS) worker nodes may enter a NotReady state.
Our team is actively working to restore connectivity and bring impacted nodes back online. If you continue to experience issues, please open a support ticket from within your account.
With IsDown, you can monitor all your critical services' official status pages from one centralized dashboard and receive instant alerts the moment an outage is detected. Say goodbye to constantly checking multiple sites for updates and stay ahead of outages with IsDown.
Start free trialNo credit card required · Cancel anytime · 6320 services available
Integrations with