Accuracy — how often our own number held
How often our own number held — measured out of sample
| Model | p50 held | p90 held |
|---|---|---|
| OpenAI · gpt-5.6-luna | 80.3% | 99.7% |
| OpenAI · gpt-5-nano | 74.6% | 98.7% |
| OpenAI · gpt-5.6-sol | 56.5% | 96.3% |
| Anthropic · claude-haiku-4-5 | 60.5% | 95.3% |
| Google · gemini-3.7-flash | 46.5% | 80.6% |
| Google · gemini-2.5-flash | building (n=0) | |
| Google · gemini-2.5-flash-lite | building (n=0) | |
| OpenAI · gpt-4o-mini | building (n=0) | |
Source: /v1/calibration
Speed — median queue time by model
Median queue time by model
Minutes, not hours — and a large spread between models at the same provider. Fastest median at the top; each bar carries the n behind it and is coloured by how much we trust it. Batch endpoints bill at half the synchronous rate, so this wait is what you pay in time for that discount.
Source: /v1/coverage
The distribution behind the median — per model
Distribution, hour-of-day, weekday and size — per model
An average hides the tail, and the tail is the risk. Pick a model to see the whole distribution behind its median, when in the day its queue is actually fastest, its weekday pattern, and — the question the forums have had open since 2024 — whether bigger batches wait longer.
Distribution — share finished vs. elapsed time
Hour of day — when the queue is fastest (UTC)
Weekday pattern
Size effect — do bigger batches wait longer?
Reliability — incident timeline
Incident timeline
Severe and degraded windows we have measured, newest last. A provider's own status page rarely admits a batch slowdown; ours is drawn from what we actually watched happen.
Source: /v1/outages
Edge latency by region
Edge latency by region
batchwatch answers from Cloudflare's edge, so calling the API does not add a round trip to a distant origin. Time-to-first-byte from a real API call, measured from every region we can reach.
Source: /v1/latency
Coverage and provenance
Coverage and provenance
Right now every one of these 2,856 measurements is our own probe — 0 come from outside contributors. We publish that on purpose: a dashboard that hides its own single-source problem reads as marketing. One source today; here is how to become the second — every job you run through batchwatch adds a measurement nobody had before.
| Model | Measurements (n) | Outside contributors |
|---|---|---|
| gpt-5.6-luna OpenAI | 561 | 0 |
| claude-haiku-4-5 Anthropic | 531 | 0 |
| gpt-5.6-sol OpenAI | 520 | 0 |
| gpt-5-nano OpenAI | 508 | 0 |
| gemini-3.7-flash Google | 445 | 0 |
| gemini-2.5-flash Google | 238 | 0 |
| gemini-2.5-flash-lite Google | 53 | 0 |
Source: /v1/coverage