What we measured
How predictable is it?
1.4x spread — p90 over p50
How much longer the slow tail runs than the typical job. Nearer 1× is more predictable.
Half of gemini-3.7-flash batch jobs at Google finished within 2 min, and nine in ten within 3 min. That is a spread of 1.4x: the tail sits close to the middle, so what you measure once is close to what you get again. It is the tightest distribution we measure across every model on the site — the most predictable batch queue we have found.
Spread is p90 divided by p50, both measured here over the last 30 days. No provider publishes it; it is the number that tells you whether a median is safe to plan on.
What this means for a deadline
A 1.4x spread is the whole difference between two ways of planning. When the tail sits close to the median, sizing a deadline against the worst case you are likely to see costs you almost nothing over sizing against the typical case — the two numbers are close together. On a model whose tail runs many times its median, the tail IS the plan: the median is a number you cannot safely book against, and the honest planning figure is far out on the right.
That is why the spread matters more than the median alone. This page gives you both, measured, so you can book Google gemini-3.7-flash against 3 min and know how often you will beat it.
Every measured job, slowest last
Share of jobs finished, against elapsed time. 2497 completions, 30-day window.
The curve is drawn from the pooled measurements, while the headline percentiles above are the median of each contributor's own figures - so the two need not line up exactly. That is deliberate: no single source can move a headline number by sending more data.
What you can plan for
3 min
Upper end of the measured interval
Plan for 3 min. That is the upper end of a 90% confidence interval around the p90, computed from the measurements themselves. It is deliberately not the median: being wrong toward synchronous costs twice the money, being wrong toward batch can miss a deadline, and those two mistakes are not equally expensive.
The slowest job we actually observed took 7 min. That is one job, not a bound.
How it compares
Every model we have measured enough of to publish a percentile, most predictable first. Spread is p90 over p50.
| Model | Median (p50) | p90 | Spread |
|---|---|---|---|
| Google gemini-3.7-flash | 2 min | 3 min | 1.4x |
| Anthropic claude-haiku-4-5 | 2 min | 7 min | 2.7x |
| OpenAI gpt-5.6-luna | 35s | 16 min | 28.0x |
| OpenAI gpt-5.6-sol | 30s | 17 min | 33.4x |
| OpenAI gpt-5-nano | 74s | 2 hours | 85.6x |
Each row is measured over the last 30 days. A model we have not yet measured enough of to stand behind a percentile is left out of this table entirely — never filled in with an estimate.
What we have not measured
The spread above is measured over the batch jobs contributors ran in the last 30 days, at the sizes they happened to run. We do not vary job size ourselves, so a much larger or much smaller batch than the ones behind this figure may sit differently — we publish a percentile only for what we have actually timed, and say so plainly rather than extrapolating past it. That is the same discipline that lets you trust the numbers we do print.
How much this answer is worth
Confidence: Low
Score 0.99 · basis: exact_model
Based on a single contributor, a figure pooled across contributors (not yet a per-contributor median).
Pooled across every measurement we have. As more contributors each build their own track record, this figure sharpens into a per-contributor median.
Where the numbers come from
Each measurement is one completed batch job. The duration is end minus start, both timestamped by our server - a client can never submit a duration, and a job we stopped waiting for is recorded as abandoned, never as a fast completion.
| Newest measurement | 2026-10-09T20:25:51.000Z |
|---|---|
| Oldest in window | 2026-09-09T20:46:48.000Z |
| Figures computed | 2026-10-09T20:43:41.000Z |
| Public delay | 15 minutes |
| Percentile method | pooled (1 voting source) |
Timestamps are absolute and in UTC on purpose: this page may be cached, and a relative age would quietly go stale in the cache while an absolute one stays true.
Questions this page answers
- How long does gemini-3.7-flash batch take at Google?
- Measured over the last 30 days: half of gemini-3.7-flash batch jobs at Google finished within 2 min and nine in ten within 3 min, across 2497 completed jobs from 0 sources. Google itself publishes only a 24 hours completion window and nothing tighter.
- How current is this gemini-3.7-flash figure?
- The newest gemini-3.7-flash batch job behind this page finished at 2026-10-09T20:25:51.000Z. Public figures are precomputed and delayed by 15 minutes; contributors see them at five minutes, and paid access is computed live.
The provider-wide version of this question has its own measured answer: How long does the Google (Gemini) batch API take?
Get this as JSON
Same numbers, same delay, no key required.
curl "https://batchwatch.dev/v1/curve?provider=google&model=gemini-3.7-flash"
Send five measurements in seven days and the delay is gone: you read every figure live, and the decision endpoints open (/v1/should-i-batch, /v1/estimate-batchtime). Start on the front page.