batchwatch

batchwatch / the data

Everything we measure, drawn — and how often our own number held

Eight public endpoints, no key required. Median queue time by model, the full distribution behind each median, when in the day the queue is fastest, the incident timeline, edge latency — and, given top billing because it is the point of the page, how often the wait we published actually held, measured out of sample. Every chart links to the endpoint that produced it, so any bar is one click from the raw data. Where the data is too thin to draw an honest chart, we say so rather than draw one.

Accuracy — how often our own number held

How often our own number held — measured out of sample

94.1%
of jobs finished inside the p90 we published · 1,494 jobs scored out of sample · 30-day window · live tier
Modelp50 heldp90 held
OpenAI · gpt-5.6-luna80.3%99.7%
OpenAI · gpt-5-nano74.6%98.7%
OpenAI · gpt-5.6-sol56.5%96.3%
Anthropic · claude-haiku-4-560.5%95.3%
Google · gemini-3.7-flash46.5%80.6%
Google · gemini-2.5-flashbuilding (n=0)
Google · gemini-2.5-flash-litebuilding (n=0)
OpenAI · gpt-4o-minibuilding (n=0)

Source: /v1/calibration

Speed — median queue time by model

Median queue time by model

Minutes, not hours — and a large spread between models at the same provider. Fastest median at the top; each bar carries the n behind it and is coloured by how much we trust it. Batch endpoints bill at half the synchronous rate, so this wait is what you pay in time for that discount.

0s 3 min 5 min gpt-5.6-sol 78s n=520 gemini-2.5-flash-lite 81s n=53 gpt-5-nano 88s n=508 gemini-2.5-flash 2 min n=238 claude-haiku-4-5 2 min n=531 gemini-3.7-flash 3 min n=445 gpt-5.6-luna 5 min n=561

Source: /v1/coverage

The distribution behind the median — per model

Distribution, hour-of-day, weekday and size — per model

An average hides the tail, and the tail is the risk. Pick a model to see the whole distribution behind its median, when in the day its queue is actually fastest, its weekday pattern, and — the question the forums have had open since 2024 — whether bigger batches wait longer.

Distribution — share finished vs. elapsed time

0s 6 hours 11 hours 0% 25% 50% 75% 90% 100% share of jobs finished →

Hour of day — when the queue is fastest (UTC)

Loading live…

Weekday pattern

Loading live…

Size effect — do bigger batches wait longer?

Loading live…

Source: /v1/curve · /v1/hours · /v1/patterns

Reliability — incident timeline

Incident timeline

Severe and degraded windows we have measured, newest last. A provider's own status page rarely admits a batch slowdown; ours is drawn from what we actually watched happen.

Loading live…
Loads live from /v1/outages with JavaScript on.

Source: /v1/outages

Edge latency by region

Edge latency by region

batchwatch answers from Cloudflare's edge, so calling the API does not add a round trip to a distant origin. Time-to-first-byte from a real API call, measured from every region we can reach.

Loading live…
The full edge-latency chart is at /latency. It loads live here from /v1/latency with JavaScript on.

Source: /v1/latency

Coverage and provenance

Coverage and provenance

Right now every one of these 2,856 measurements is our own probe — 0 come from outside contributors. We publish that on purpose: a dashboard that hides its own single-source problem reads as marketing. One source today; here is how to become the second — every job you run through batchwatch adds a measurement nobody had before.

ModelMeasurements (n)Outside contributors
gpt-5.6-luna OpenAI 561 0
claude-haiku-4-5 Anthropic 531 0
gpt-5.6-sol OpenAI 520 0
gpt-5-nano OpenAI 508 0
gemini-3.7-flash Google 445 0
gemini-2.5-flash Google 238 0
gemini-2.5-flash-lite Google 53 0

Source: /v1/coverage