New2026-07-15

New: SLA Calculator & RSS Feed

Calculate maximum allowed downtime for any SLA uptime target with our new SLA Calculator. Plus, subscribe to incident updates via our new RSS feed.

Try SLA Calculator
New2026-03-16

New: Documentation page launched

We have published a comprehensive docs page covering website usage, API reference, webhook integrations, error handling, and embed codes.

View docs
Resolved2026-03-15

Support email restored

Our support email at support@incidenthub-bay.com has been fully restored. Email alert delivery is operating normally.

AI Infrastructure Reliability

Is your AI API down? Know before your users do.

IncidentHub-Bay monitors OpenAI, Anthropic, Google AI, Mistral, and every major cloud provider in real time. Get instant alerts, compare reliability scores, and make smarter infrastructure decisions.

AI & cloud incidents

1457

Providers tracked

21

Critical incidents

89

Top cause tag

api

Updated every 5 minutes · Data from provider status pages

How it works

Three steps to stay ahead of AI and cloud outages.

1

Track

We monitor status pages for OpenAI, Anthropic, Google AI, AWS, and 50+ providers every 5 minutes.

2

Alert

Get notified via Slack, webhooks, or Google Chat the moment a provider reports an issue.

3

Analyze

Compare AI API reliability scores, review outage history, and decide when to add fallback providers.

AI API Reliability Dashboard

Which LLM provider is most reliable this month? Compare uptime, incident frequency, and mean resolution time across OpenAI, Anthropic, Google AI, and more.

View dashboard

Plans for every team size

Start free with 3 tracked AI services. Upgrade as your monitoring needs grow.

Compare plans →

Free

For individual developers exploring AI APIs

Free forever
$0/month

Track a small watchlist of AI and cloud services, test webhook alerts, and browse public reliability data.

3 tracked services

  • 3 AI/cloud service watchlist slots
  • Webhook and Google Chat alerts
  • Public reliability rankings and outage history
  • Blog and annual reports preview
Start free

Pro

For on-call engineers building on AI APIs

Most popular
$29/month

Monitor more AI providers, get reliability comparisons, and route alerts to the channels your team already uses.

10 tracked services

  • 10 AI/cloud service watchlist slots
  • Slack, Google Chat, PagerDuty, and webhook alerts
  • AI reliability comparison dashboard
  • Incident analytics and downtime statistics
Go Pro

Teams

For platform and infrastructure teams

Full platform
$79/month

Shared monitoring across your AI stack, API access for internal tooling, and executive-ready reliability reports.

25 tracked services

  • 25 AI/cloud service watchlist slots
  • Shared alert destinations and team workflows
  • Full API access for dashboards and internal tools
  • Monthly reliability reports for vendor reviews
Start Teams

Enterprise

For organizations with custom SLA and compliance needs

Custom
$299/month

Unlimited monitoring, dedicated support, SSO, SLA guarantees, and custom integrations for your infrastructure team.

Unlimited services

  • Unlimited AI/cloud service tracking
  • Custom SLA and uptime guarantees
  • SSO / SAML integration
  • Dedicated support and onboarding
Contact sales

Outage trackers by provider

Deep-dive into outage history, status timelines, and reliability data for AI and cloud services.

OpenAI Outage History

openai down

OpenAI and ChatGPT incident history, API uptime data, and recovery timelines for teams building on GPT models.

Anthropic / Claude Outages

claude api down

Anthropic Claude API and Claude Chat outage tracking, reliability scores, and incident history.

Google AI / Vertex Outages

google ai outage

Google Cloud and Vertex AI / Gemini incident history, outage timelines, and reliability data.

Mistral AI Outages

mistral ai status

Mistral AI platform status, API outage history, and reliability tracking for teams using Mistral models.

Cohere Outages

cohere api status

Cohere API outage history and reliability data for enterprise RAG and embedding workloads.

Replicate Outages

replicate status

Replicate platform status, open-source model hosting uptime, and incident tracking.

DeepSeek Outages

deepseek down

DeepSeek API and Chat outage history, incident timelines, and reliability tracking.

Groq Outages

groq api status

Groq inference API status, outage history, and uptime data for LPU-powered AI workloads.

Stability AI Outages

stability ai status

Stability AI platform status, API outage history, and reliability data for image generation workloads.

ElevenLabs Outages

elevenlabs down

ElevenLabs speech synthesis API status, outage history, and reliability tracking.

AI21 Labs Outages

ai21 api status

AI21 Labs Jamba API status, outage history, and reliability data for enterprise AI workloads.

AWS Outage History

aws outage history

AWS outages by year, timeline, and root-cause context for engineers tracking historical downtime.

Cloudflare Outages

cloudflare outage

Live and historical Cloudflare downtime tracking for teams who depend on edge infrastructure.

GitHub Outages

github down

GitHub and GitHub Actions outage tracking for developers who need answers during deploy failures.

AI API Reliability Compared

ai api reliability

Side-by-side reliability comparison of OpenAI, Anthropic, Google AI, Mistral, and other LLM providers.

Biggest Cloud Outages

biggest cloud outages

A ranked list of the longest and most impactful AI and cloud outages in our dataset.

Cloud Downtime Statistics

cloud downtime statistics

Incidents per year, average downtime, and provider-level patterns drawn from our full incident dataset.

Total Incidents

1457

Providers Tracked

21

Critical Incidents

89

Top Cause

api

Recent Incidents

View all →

Delays in credit purchases

Anthropic·Sep 1, 2026

We have identified and resolved an issue where users who reached a balance of zero usage credits saw delays in the availability of newly purchased credits, resulting in some requests to the Claude API receiving 'credit balance is too low' errors erroneously. This affected credits purchased from 05:10am PT / 12:10 UTC to 2:35pm PT / 21:35 UTC. At this time, newly purchased credits should be available as expected and we are working to resolve any remaining impact.

api
low

Pipelines Creation Issue

Cloudflare·Sep 1, 2026

A fix has been implemented and we are monitoring the results.

api
low

Network Route Leak in Palmas, Brazil

Cloudflare·Sep 1, 2026·19m

This incident has been resolved.

networkrouting
high

Cloudflare Access one-time PIN emails blocked by certain email security gateways, including Proofpoint

Cloudflare·Access·Sep 1, 2026

Some customers using Proofpoint may experience delayed Cloudflare notifications and login codes. Proofpoint is temporarily deferring messages with an SMTP 421 4.7.0 response. Cloudflare will continue retrying delivery.

authenticationsecurity
low

Increased deployment failures

Vercel·Sep 1, 2026·59m

This incident has been resolved. Builds have recovered.

computedeploymentrouting
high

Degraded performance on platform.claude.com and Claude for Microsoft Office 365

Anthropic·Sep 1, 2026·1h 2m

This issue has been resolved.

api
low

Degraded performance on claude.ai

Anthropic·Sep 1, 2026·20m

This issue has been resolved.

api
low

Multiple Endpoint Disruptions

Cohere·Sep 1, 2026

🎉 The issue has been resolved and we are back to normal.

api
low

Delays in commit processing

GitHub·Sep 1, 2026·1h 1m

This incident has been resolved. Thank you for your patience and understanding as we addressed this issue. A detailed root cause analysis will be shared as soon as it is available.

capacityapi
low

We are seeing issues with 400 Bad Request on Requests to D1

Cloudflare·Sep 1, 2026·3h 27m

This incident has been resolved.

apiauthentication
low

From the blog

AI reliability insights, outage analysis, and infrastructure decision-making.

All posts →