Summary
All affected models have been scaling correctly for more than 30 minutes at this point, and we see no residual prediction queues. Thank you for your patience!
Impact
minor
Timeline
[monitoring] Scaling decisions were delayed by nearly 1 hour after the controller responsible for emitting queue metrics failed to schedule on a soft-failed node. We have since cordoned and drained the node, and the controller is emitting queue metrics again.
via statuspage[resolved] All affected models have been scaling correctly for more than 30 minutes at this point, and we see no residual prediction queues. Thank you for your patience!
via statuspageLessons Learned
⚠Replicate has experienced 43 incidents in the past year. This frequency suggests systemic reliability challenges that may warrant additional monitoring.
📊Incidents related to capacity have occurred 71 times across all providers in the past year. This is one of the most common failure categories in cloud infrastructure.
💡This incident is categorized as: Capacity Issue. Consider implementing preventive measures specific to this failure category.
Similar Incidents
Delays in commit processing
GitHub · Sep 1, 2026
Customer support ticketing system not receiving new tickets and ticket replies
ElevenLabs · Aug 31, 2026
Workers Builds are Degraded
Cloudflare · Aug 27, 2026
A100 hardware partial outage
Replicate · Aug 26, 2026
Incident with GraphQL API Requests
GitHub · Aug 11, 2026