Degraded scale-out due to failed setups pulling from huggingface
Summary
We're looking stable now, and with caching back on. Thanks again for your patience!
Impact
minor
Timeline
[identified] Models that pull from huggingface during setup have not been succeeding, which is resulting in scale-out delays and queue backups.
via statuspage[identified] We have confirmed that turning off our caching layer for huggingface downloads allows models to successfully set up, so we are making that change for all models while we continue to troubleshoot.
via statuspage[monitoring] All affected hardware types have recovered and we are continuing to monitor and troubleshoot. Thanks for your patience!
via statuspage[resolved] We're looking stable now, and with caching back on. Thanks again for your patience!
via statuspageLessons Learned
⚠Replicate has experienced 39 incidents in the past year. This frequency suggests systemic reliability challenges that may warrant additional monitoring.
📊Incidents related to api, capacity have occurred 865 times across all providers in the past year. This is one of the most common failure categories in cloud infrastructure.
💡This incident is categorized as: Capacity Issue, API Issue. Consider implementing preventive measures specific to this failure category.
Similar Incidents
Degraded availability GPT 5.6 Luna
GitHub · Aug 1, 2026
Increased HTTP 5XX Errors in IAD
Cloudflare · Jul 31, 2026
Increased HTTP Errors in London
Cloudflare · Jul 31, 2026
Cloudflare Analytics API Availability Reduced
Cloudflare · Jul 31, 2026
Enterprise & Education Chat Errors
OpenAI · Jul 31, 2026