Is Groq down right now?
Authenticated API inference - 2 models monitored · How we classify outages
Weekly AI pricing & uptime digest
Price drops, new model releases, and incident summaries - every Monday. Free.
Groq is currently operational - but API inference failing on one or more models in the last 45 minutes - 305ms HTTP response. Last checked . 90-day uptime: 99.9%. Groq API: 0/2 models up - check inference section below.
Stay informed
HTTP uptime (90d)
99.9%
20 incidents (90d)
HTTP response now
305ms
HTTP p50 (7d)
412ms
median ping response
HTTP p95 (7d)
1794ms
tail ping response
API Inference Monitoring
Live · every 5 minBest TTFT (p50)
—
time to first token
Best throughput
—
output tokens/sec (24h avg)
Min success rate
0%
worst model (24h)
P50 = typical speed. P95 = worst case 95% of the time. Measured by Tickerr's independent inference checks. Requires ≥10 checks to display.
TTFT over 24 hours
ⓘ Authenticated streaming API calls via native fetch. TTFT = milliseconds from request start to first streamed token chunk. Throughput = output tokens ÷ generation time. Checks run from Vercel us-east-1. Independent of the provider's official status page.
Is Groq slow right now?
Groq is currently experiencing degraded performance. All 2 monitored models are experiencing elevated latency.Last checked 2m ago.
HTTP endpoint response: 305ms (7-day p50: 412ms). This measures basic reachability, not model inference speed.
TTFT (time-to-first-token) is measured via authenticated streaming API calls every 5 minutes. Current values are compared against 7-day and 30-day rolling baselines. A minimum of 5 successful checks is required before classification. Methodology
Groq speed FAQ
Is Groq slow right now?
Groq is currently experiencing degraded performance. All 2 monitored models are experiencing elevated latency.
Why is Groq so slow today?
Groq is experiencing elevated latency compared to its normal baseline. This can happen during high-demand periods, model updates, or infrastructure issues. Tickerr monitors time-to-first-token (TTFT) every 5 minutes across all Groq models to detect slowdowns automatically.
Is Groq down or just responding slowly?
Groq is not fully down, but it is experiencing degraded performance. Some features or models may be slower or less reliable than usual.
How does Tickerr determine whether Groq is slow?
Tickerr sends authenticated API requests to Groq every 5 minutes and measures time-to-first-token (TTFT). Current TTFT is compared against a rolling 7-day and 30-day baseline. A model is classified as slow when its current p50 TTFT exceeds 2x its baseline, and severely slow at 3.5x. At least 5 successful checks are required before any classification is made.
What is normal response time for Groq?
Tickerr is still collecting baseline data for Groq. Normal response times will be available after sufficient monitoring history has accumulated.
Can one Groq model be slow while others are working normally?
Yes. Groq runs multiple models (e.g. llama-3.3-70b-versatile, meta-llama/llama-4-scout-17b-16e-instruct), and each can have independent performance characteristics. Tickerr monitors each model separately. A slowdown on one model doesn't necessarily affect others.
Experimental capability checks
PUBLIC PILOTTickerr sends deterministic test prompts to Groq every hour and evaluates the output using code-based pass/fail checks. These results are informational onlyduring the pilot period and do not affect the tool's status verdict.
How Tickerr tests this
Every hour, Tickerr sends three deterministic test prompts to each monitored model via the provider's API. Each response is evaluated with a code-based pass/fail check — no subjective judgement.
- Instruction following: asks the model to reply with an exact word. Graded by exact string match.
- JSON schema: asks for a JSON object matching a fixed schema. Graded by parse + field validation.
- Tool call: asks the model to call a function with specific arguments. Graded by parsing the tool-call response.
Infrastructure errors (timeouts, rate limits, auth failures) are tracked separately and excluded from capability pass rates. Models that don't support tool calls are marked “not supported” rather than counted as failures. Full methodology
HTTP endpoint response time (7 days)
p50 412ms·p95 1794msⓘ HTTP response times to Groq's status endpoint - measures infrastructure availability, not API inference speed. For TTFT and model-level API status, see the Groq API Status section above.
90-day uptime
Incident history
Independent monitoring detected consecutive API failures for llama-3.3-70b-versatile. Tickerr measures time-to-first-token (TTFT) via live streaming API calls every 5 minutes.
Independent monitoring detected consecutive API failures for llama-3.3-70b-versatile. Tickerr measures time-to-first-token (TTFT) via live streaming API calls every 5 minutes.
Independent monitoring detected consecutive API failures for llama-3.3-70b-versatile. Tickerr measures time-to-first-token (TTFT) via live streaming API calls every 5 minutes.
Independent monitoring detected consecutive API failures for llama-3.3-70b-versatile. Tickerr measures time-to-first-token (TTFT) via live streaming API calls every 5 minutes.
Independent monitoring detected consecutive API failures for llama-3.3-70b-versatile. Tickerr measures time-to-first-token (TTFT) via live streaming API calls every 5 minutes.
Independent monitoring detected consecutive API failures for llama-3.3-70b-versatile. Tickerr measures time-to-first-token (TTFT) via live streaming API calls every 5 minutes.
Independent monitoring detected consecutive API failures for llama-3.3-70b-versatile. Tickerr measures time-to-first-token (TTFT) via live streaming API calls every 5 minutes.
Independent monitoring detected consecutive API failures for llama-3.3-70b-versatile. Tickerr measures time-to-first-token (TTFT) via live streaming API calls every 5 minutes.
Independent monitoring detected consecutive API failures for llama-3.3-70b-versatile. Tickerr measures time-to-first-token (TTFT) via live streaming API calls every 5 minutes.
Independent monitoring detected consecutive API failures for llama-3.3-70b-versatile. Tickerr measures time-to-first-token (TTFT) via live streaming API calls every 5 minutes.
Independent monitoring detected consecutive API failures for llama-3.3-70b-versatile. Tickerr measures time-to-first-token (TTFT) via live streaming API calls every 5 minutes.
Independent monitoring detected consecutive API failures for llama-3.3-70b-versatile. Tickerr measures time-to-first-token (TTFT) via live streaming API calls every 5 minutes.
Independent monitoring detected consecutive API failures for llama-3.3-70b-versatile. Tickerr measures time-to-first-token (TTFT) via live streaming API calls every 5 minutes.
Independent monitoring detected consecutive API failures for llama-3.3-70b-versatile. Tickerr measures time-to-first-token (TTFT) via live streaming API calls every 5 minutes.
Independent monitoring detected consecutive API failures for llama-3.3-70b-versatile. Tickerr measures time-to-first-token (TTFT) via live streaming API calls every 5 minutes.
Independent monitoring detected consecutive API failures for llama-3.3-70b-versatile. Tickerr measures time-to-first-token (TTFT) via live streaming API calls every 5 minutes.
Independent monitoring detected consecutive API failures for llama-3.3-70b-versatile. Tickerr measures time-to-first-token (TTFT) via live streaming API calls every 5 minutes.
Independent monitoring detected consecutive API failures for llama-3.3-70b-versatile. Tickerr measures time-to-first-token (TTFT) via live streaming API calls every 5 minutes.
Independent monitoring detected consecutive API failures for llama-3.3-70b-versatile. Tickerr measures time-to-first-token (TTFT) via live streaming API calls every 5 minutes.
Independent monitoring detected consecutive API failures for llama-3.3-70b-versatile. Tickerr measures time-to-first-token (TTFT) via live streaming API calls every 5 minutes.
Related pages
Groq API not working? Common error codes
If Groq's API is returning errors, the table below explains what each code means and how to fix it. If errors are widespread, check the live status above - a service incident will appear there within minutes.
| Error | What it means & what to do |
|---|---|
| HTTP 429 | Rate limit hit - check x-ratelimit-remaining headers; free tier is 30 RPM |
| HTTP 503 | Service overloaded - Groq queues fill fast under heavy load; retry |
Note: Tickerr monitors Groq's status endpoint, not individual API calls. An HTTP 429 or 500 in your app may be specific to your account tier - check the rate limits page for plan-specific thresholds.
About Groq status
Groq is an AI inference provider offering extremely fast LLM inference using custom LPU hardware. Groq downtime is rare due to their hardware architecture but can affect API users building latency-sensitive apps. Groq offers free and paid tiers with separate rate limits.