Use cases
Software Products E-commerce MSPs Schools Development & Marketing DevOps Agencies Help Desk
Company
Internet Status Blog Pricing Log in Get started free

Outage in LiveKit

Elevated timeouts on LiveKit Inference (google/gemma-4-31b-it)

Resolved Minor
August 04, 2026 - Started 1 day ago - Lasted about 2 hours
Official incident page

Incident Report

We are monitoring elevated timeouts that affected LiveKit Inference between 14:05 and approximately 15:00 UTC today. During that window, roughly 3% of LLM requests using the google/gemma-4-31b-it model timed out. Impacted customers would have seen LLM request timeouts in their agent sessions (for example, APITimeoutError in agent logs). Other models were not affected. We have identified the cause. During a period of increased traffic, response times for this model slowed, and a defect in our health-monitoring logic incorrectly marked a backup deployment as unhealthy and removed it from rotation, preventing requests from failing over as designed. Error rates returned to baseline by approximately 15:00 UTC, and a fix for the health-monitoring defect is in progress. We are continuing to monitor before marking this resolved. No action is needed on your part. If you continue to see timeouts, please reach out to support. We apologize for the disruption.

Trusted by 1,000+ teams

The Status Page Aggregator with Early Outage Detection

Stop finding out about outages from your users. Monitor 6,320+ cloud services and get alerted the second something breaks.

IsDown status aggregator dashboard
Latest Updates ( sorted recent to last )
RESOLVED 1 day ago - at 08/04/2026 10:11PM

This incident has been resolved.

MONITORING 1 day ago - at 08/04/2026 08:45PM

We are monitoring elevated timeouts that affected LiveKit Inference between 14:05 and approximately 15:00 UTC today. During that window, roughly 3% of LLM requests using the google/gemma-4-31b-it model timed out. Impacted customers would have seen LLM request timeouts in their agent sessions (for example, APITimeoutError in agent logs). Other models were not affected.

We have identified the cause. During a period of increased traffic, response times for this model slowed, and a defect in our health-monitoring logic incorrectly marked a backup deployment as unhealthy and removed it from rotation, preventing requests from failing over as designed. Error rates returned to baseline by approximately 15:00 UTC, and a fix for the health-monitoring defect is in progress. We are continuing to monitor before marking this resolved.

No action is needed on your part. If you continue to see timeouts, please reach out to support. We apologize for the disruption.

The Status Page Aggregator with Early Outage Detection

With IsDown, you can monitor all your critical services' official status pages from one centralized dashboard and receive instant alerts the moment an outage is detected. Say goodbye to constantly checking multiple sites for updates and stay ahead of outages with IsDown.

Start free trial

No credit card required · Cancel anytime · 6320 services available

Integrations with Slack Microsoft Teams Google Chat Datadog PagerDuty Zapier Discord Webhook