Is Replicate down right now?
HTTP endpoint monitoring - checked every 5 minutes · How we classify outages
Weekly AI pricing & uptime digest
Price drops, new model releases, and incident summaries - every Monday. Free.
Replicate is currently operational - 249ms HTTP response. Last checked . 90-day uptime: 99.9%.
Stay informed
90-day uptime
99.9%
12 incidents (90d)
Response now
249ms
HTTP p50 (7d)
463ms
median ping response
HTTP p95 (7d)
1290ms
tail ping response
Response time (7 days)
p50 463ms·p95 1290msⓘ HTTP response times from Tickerr's monitoring server to Replicate's status or homepage endpoint - not API inference latency or time-to-first-token (TTFT). For API benchmark data, see third-party sources like Artificial Analysis.
90-day uptime
Incident history
Scaling decisions were delayed by nearly 1 hour after the controller responsible for emitting queue metrics failed to schedule on a soft-failed node. We have since cordoned and drained the node, and t…
We are aware of a significant degradation with the service. We have identified the root cause and are taking steps to mitigate. Please hold tight for further information.
The API for creating predictions for A100s is currently degraded as a piece of backing infrastructure failed. We are working on getting it back online
Models that pull from huggingface during setup have not been succeeding, which is resulting in scale-out delays and queue backups.
We are over provisioned on H100s currently which is causing long queue times for some models running on H100s
We are aware of 504s being returned by models that reach out to HuggingFace during setup. We believe this is likely related to the disruption visible at [https://status.huggingface.co/](https://status…
Long queue times, especially on BFL models
We are seeing high contention on H100 hardware which is resulting in delays on predictions and scale-out.
We recently received a sharp increase in demand for H100 capacity which is resulting in delayed scale-out and queue backup.
We're seeing long setup times and high contention for models on some L40S and H200 clusters.
Long queue times for black-forest-labs/flux-2-klein-4b resulting in canceled predictions
Related pages
About Replicate status
Replicate is a platform for running machine learning models via API, including image, video, and language models. When Replicate is down, model predictions fail. Replicate uses a pay-per-second pricing model - billing stops during confirmed outages.
You can also check the official Replicate status page.