Braze's EU02 cluster experienced approximately 5 hours of Currents data processing delays affecting a small subset of customers, caused by multiple Currents connectors failing to gracefully restart after routine maintenance while falsely appearing healthy in monitoring systems. The connectors built significant backlogs during this period, though no data was lost and the majority of traffic continued flowing normally throughout the incident. Resolution was achieved by performing graceful restarts and spinning up additional resources to clear the backlog, with all Currents events fully delivered to their destinations by 15:18 ET.
Trusted by 1,000+ teams
Stop finding out about outages from your users. Monitor 6,320+ cloud services and get alerted the second something breaks.
This incident is resolved. As of 15:18 ET / 19:18 UTC, we have fully processed the backlog of jobs, and our EU02 Currents connectors are all again processing within normal latency ranges. No data was lost, and all Currents events have been sent to their destinations.
We continue to make progress clearing the remaining backlog on EU02, improving our performance, and decreasing overall latency. We expect to be caught up within the next hour, although a small number of events may take slightly longer to fully process. The remaining impact is limited to less than half a percent of our overall traffic.
We'll provide another update once we're fully caught up.
We continue to investigate degraded performance in EU02. We have identified the root cause of the underlying latency, which stemmed from multiple Currents connectors failing to gracefully restart after routine maintenance, but appearing in a healthy state in our monitoring system.
Those specific servers build significant backlogs while in this state. We have completed a graceful restart, and have spun up additional resources to process through the backlogged items more rapidly.
At this point ~95% of data is flowing normally, while up to ~5% of our connectors, impacting a subset of customers, continue to experience latency.
We see throughput significantly improved due to config changes, and will provide an update in 1 hour or when we have an updated timeline.
We are currently investigating degraded performance in our EU02 cluster, which is causing delays in Currents processing for a small subset of customers. No data has been lost, and the majority of Currents data continues to flow normally. Our team is actively working toward a resolution and will provide updates as they become available.
With IsDown, you can monitor all your critical services' official status pages from one centralized dashboard and receive instant alerts the moment an outage is detected. Say goodbye to constantly checking multiple sites for updates and stay ahead of outages with IsDown.
Start free trialNo credit card required · Cancel anytime · 6320 services available
Integrations with