Delayed scaling due to node failure

lowReplicateAug 10, 2026 21:10
capacity
Capacity Issue

Summary

All affected models have been scaling correctly for more than 30 minutes at this point, and we see no residual prediction queues. Thank you for your patience!

Impact

minor

Timeline

Aug 10, 2026 21:05

[monitoring] Scaling decisions were delayed by nearly 1 hour after the controller responsible for emitting queue metrics failed to schedule on a soft-failed node. We have since cordoned and drained the node, and the controller is emitting queue metrics again.

via statuspage
+40m
Aug 10, 2026 21:45

[resolved] All affected models have been scaling correctly for more than 30 minutes at this point, and we see no residual prediction queues. Thank you for your patience!

via statuspage

Lessons Learned

Replicate has experienced 43 incidents in the past year. This frequency suggests systemic reliability challenges that may warrant additional monitoring.

📊Incidents related to capacity have occurred 71 times across all providers in the past year. This is one of the most common failure categories in cloud infrastructure.

💡This incident is categorized as: Capacity Issue. Consider implementing preventive measures specific to this failure category.