Use cases
Software Products E-commerce MSPs Schools Development & Marketing DevOps Agencies Help Desk
Company
Internet Status Blog Pricing Log in Get started free

Outage in Dagster Cloud

Brief Hybrid Agent Disruption on 2026-08-29

Resolved Minor
August 29, 2026 - Started about 1 month ago - Lasted 2 days
Official incident page

Incident Report

Impact: Dagster+ hybrid agents connecting to the US region via AWS PrivateLink experienced connection loss. Duration: ~30 minutes (17:35 – 18:05 UTC / 13:35 – 14:05 ET on 2026-08-29) Summary: As part of a planned infrastructure migration, we rotated the network load balancers behind our PrivateLink endpoint service. This involved disconnecting the previous set of consumer VPC endpoints so that connections would be reopened on the new infrastructure. During the ~30 minute window between endpoint disconnection and full re-establishment, hybrid agents connecting via PrivateLink were unable to reach the Dagster+ control plane and appeared as unhealthy in our health dashboard. Once new endpoints came online, agents automatically reconnected and returned to a healthy state without customer action. What was affected: Hybrid agents connecting to Dagster+ US region via AWS PrivateLink Job runs, code location updates, and metadata operations dependent on that connection were paused (queued locally) until the connection recovered What was NOT affected: The Dagster+ UI and API remained fully available Non-PrivateLink hybrid agents (using public network connectivity) Dagster+ Serverless (EU and US) Dagster+ EU hybrid customers Timeline (EDT): 12:35 — Existing PrivateLink consumer endpoints intentionally disconnected as part of the migration 13:05 — Full recovery to pre-migration baseline Root Cause: Though this was a planned migration, provisioning the new endpoints took significantly longer than expected due to transient errors, leading to noticeable customer impact. This possibility should have been accounted for and communicated to customers in advance.

Trusted by 1,000+ teams

The Status Page Aggregator with Early Outage Detection

Stop finding out about outages from your users. Monitor 6,320+ cloud services and get alerted the second something breaks.

IsDown status aggregator dashboard
Latest Updates ( sorted recent to last )
POSTMORTEM about 1 month ago - at 08/31/2026 06:15PM

**What happened**

During a planned network migration on 2026-08-29 17:35 UTC, we disconnected existing PrivateLink consumer VPC endpoints as part of the cutover. Hybrid customer agents connecting via PrivateLink lost their connections and were not able to reconnect until the new endpoints were provisioned. Full recovery at 18:05 UTC, ~30 min.

**What went well**

* No data loss; agents automatically re-attached with no customer action required.

**What could have gone better**

* Provisioning new endpoints took longer than expected.
* We did not proactively communicate to affected customers ahead of the maintenance window.

**Follow-ups**

* We will update our standard procedures for planned maintenance to ensure communication of operations that may cause delays in customer pipelines.

RESOLVED about 1 month ago - at 08/31/2026 06:11PM

Impact: Dagster+ hybrid agents connecting to the US region via AWS PrivateLink experienced connection loss.
Duration: ~30 minutes (17:35 – 18:05 UTC / 13:35 – 14:05 ET on 2026-08-29)

Summary:

As part of a planned infrastructure migration, we rotated the network load balancers behind our PrivateLink endpoint service. This involved disconnecting the previous set of consumer VPC endpoints so that connections would be reopened on the new infrastructure. During the ~30 minute window between endpoint disconnection and full re-establishment, hybrid agents connecting via PrivateLink were unable to reach the Dagster+ control plane and appeared as unhealthy in our health dashboard.

Once new endpoints came online, agents automatically reconnected and returned to a healthy state without customer action.

What was affected:

Hybrid agents connecting to Dagster+ US region via AWS PrivateLink
Job runs, code location updates, and metadata operations dependent on that connection were paused (queued locally) until the connection recovered

What was NOT affected:

The Dagster+ UI and API remained fully available
Non-PrivateLink hybrid agents (using public network connectivity)
Dagster+ Serverless (EU and US)
Dagster+ EU hybrid customers

Timeline (EDT):
12:35 — Existing PrivateLink consumer endpoints intentionally disconnected as part of the migration
13:05 — Full recovery to pre-migration baseline

Root Cause:

Though this was a planned migration, provisioning the new endpoints took significantly longer than expected due to transient errors, leading to noticeable customer impact. This possibility should have been accounted for and communicated to customers in advance.

Latest Dagster Cloud outages

Elevated API latency - over 1 year ago
Impact of GCP incident - over 1 year ago

The Status Page Aggregator with Early Outage Detection

With IsDown, you can monitor all your critical services' official status pages from one centralized dashboard and receive instant alerts the moment an outage is detected. Say goodbye to constantly checking multiple sites for updates and stay ahead of outages with IsDown.

Start free trial

No credit card required · Cancel anytime · 6320 services available

Integrations with Slack Microsoft Teams Google Chat Datadog PagerDuty Zapier Discord Webhook