Summary
We've been below capacity for A100 hardware for the past 30 minutes and all queues have been drained. Thanks for your patience ๐
Impact
major
Timeline
[investigating] We're seeing most A100 hardware automatically cordoned due to possible hardware failure.
via statuspage[monitoring] We are seeing pods start up again now and they are working through the backlog.
via statuspage[resolved] We've been below capacity for A100 hardware for the past 30 minutes and all queues have been drained. Thanks for your patience ๐
via statuspageLessons Learned
โ Replicate has experienced 43 incidents in the past year. This frequency suggests systemic reliability challenges that may warrant additional monitoring.
๐Incidents related to api, capacity have occurred 996 times across all providers in the past year. This is one of the most common failure categories in cloud infrastructure.
๐กThis incident is categorized as: Capacity Issue, API Issue. Consider implementing preventive measures specific to this failure category.
Similar Incidents
Elevated number of R2 503 errors in Western North America region
Cloudflare ยท Sep 3, 2026
ChatGPT Work Mode High Error Rates
OpenAI ยท Sep 3, 2026
Elevated errors for Claude Sonnet 5
Anthropic ยท Sep 2, 2026
Elevated errors creating new accounts
OpenAI ยท Sep 2, 2026
Durable Objects increased errors in Western North America
Cloudflare ยท Sep 2, 2026