NinjaOne experienced a ~10-hour service availability disruption in the US caused by an AWS infrastructure outage that severed agent connections to backend services. As agents attempted to reconnect, a backoff/retry mechanism slowed recovery, and the device status processing service was overwhelmed by the volume of reconnection events, requiring capacity scaling and frontend web system restarts to resolve. The incident was fully resolved after mitigation efforts restored correct agent status reporting.
Trusted by 1,000+ teams
Stop finding out about outages from your users. Monitor 6,320+ cloud services and get alerted the second something breaks.
This incident has been resolved.
Mitigation efforts have shown positive results and we are seeing agents reporting status correctly at this time. We are continuing to monitor the incident.
To help mitigate some agents still showing offline we are restarting our frontend web systems. Some users may see a slight increase in response time during this process.
During the AWS Outage, Ninja Agents attempted and failed to reconnect to our backend services. This automatically triggers a backoff/retry mechanism in our agents that is designed to prevent traffic flooding or thundering herd behavior during significant outages. Agents are slowly reconnecting, but will randomize their reconnection attempts until successful. Additionally the service that was responsible for processing device status was struggling with the volume of agent status changes. We have scaled that capacity out and are seeing healthy metrics. Over the last hour we have seen our connected devices increase by over 150k devices, but recovery will likely be slow while the agents attempt to gracefully reestablish service
We are continuing to monitor for any further issues.
We have identified a broader connectivity issue within our cloud provider's infrastructure that is causing intermittent service disruptions for some customers. Recovery efforts are ongoing and service stability has improved, though some impact may persist while restoration work continues.
We have identified a broader connectivity issue within our cloud provider's infrastructure that is causing intermittent service disruptions for some customers, and we are actively monitoring recovery efforts.
We are currently investigating this issue.
With IsDown, you can monitor all your critical services' official status pages from one centralized dashboard and receive instant alerts the moment an outage is detected. Say goodbye to constantly checking multiple sites for updates and stay ahead of outages with IsDown.
Start free trialNo credit card required · Cancel anytime · 6320 services available
Integrations with