A100 hardware partial outage

highReplicateAug 26, 2026 21:20
apicapacity
Capacity IssueAPI Issue

Summary

We've been below capacity for A100 hardware for the past 30 minutes and all queues have been drained. Thanks for your patience ๐Ÿ’–

Impact

major

Timeline

Aug 26, 2026 21:17

[investigating] We're seeing most A100 hardware automatically cordoned due to possible hardware failure.

via statuspage
+1h 50m
Aug 26, 2026 23:07

[monitoring] We are seeing pods start up again now and they are working through the backlog.

via statuspage
+1h 18m
Aug 27, 2026 00:25

[resolved] We've been below capacity for A100 hardware for the past 30 minutes and all queues have been drained. Thanks for your patience ๐Ÿ’–

via statuspage

Lessons Learned

โš Replicate has experienced 43 incidents in the past year. This frequency suggests systemic reliability challenges that may warrant additional monitoring.

๐Ÿ“ŠIncidents related to api, capacity have occurred 996 times across all providers in the past year. This is one of the most common failure categories in cloud infrastructure.

๐Ÿ’กThis incident is categorized as: Capacity Issue, API Issue. Consider implementing preventive measures specific to this failure category.