Nine routes answer questions about queue time. Five are open to anyone (/v1/wait, /v1/recommend, /v1/curve, /v1/coverage, /v1/accuracy); four are gated behind a contribution or the free trial (/v1/should-i-batch, /v1/estimate-batchtime, /v1/conditions, /v1/distribution).
All of them are GET. All of them are affected by the tier delay except /v1/coverage. See tiers.md.
Read interpreting.md before you act on any number here.
GET /v1/wait
"What is this queue doing?" The simplest route, open without a key.
Parameters
| Name | Type | Required | Default | Notes |
|---|---|---|---|---|
provider | string | no | openai | Not validated against a list on this route. An unknown provider simply has no measurements. |
model | string | yes | — | Missing gives 400 {"error":"model is required"}. |
mode | string | no | batch | batch or sync. Anything else matches nothing. |
input_tokens | number | no | — | Your job's input-token size. Size-matched for entitled callers, acknowledged-but-not-applied for everyone else (see below). Must be a positive whole number; a bad value is a 400 for every tier. |
Behaviour
input_tokensis size-matched for entitled callers, acknowledged (never silently ignored) for the rest. What varies is the entitlement, not the route: the URL and the parameter are stable, and only your status changes the answer. Whichever tier you are, the response carries aninput_tokensobject saying whether size was applied.
Entitled caller (a tier a contribution or payment unlocks — contributor, contributor_verified, paid; the same entitlement /v1/estimate-batchtime uses) with enough measurements near the size → the number is matched to your job size, computed by the same banding and robustness voting /v1/estimate-batchtime uses, so the two routes agree exactly for one model + token count. basis becomes size_matched, a size_band is disclosed, and the ack reads:
``json "input_tokens": { "value": 5000, "used": true, "why": "This estimate was matched to your job size, using measurements of comparable jobs on this model (the same size banding and robustness voting the size-matched estimate uses).", "size_band": { "from": 2500, "to": 10000, "factor": 2, "n": 12 }, "size_effect": { "detected": true, "ratio_vs_other_sizes": 1.4, "note": "…" } } ``
Entitled caller, but no measurements near the size → an honest fallback: the model-level number is kept unchanged and the ack says size could not be matched yet (used: false). We never fabricate a size-adjusted figure where there is no data to justify it. (Every prober job is currently ~8 tokens, so this is what an entitled caller sees today until the size ladder produces data.)
``json "input_tokens": { "value": 40, "used": false, "why": "You are entitled to a size-matched estimate, but there are not yet enough measurements near your job size on this model to compute one honestly. This is the model's overall queue time; it becomes size-matched automatically once comparable jobs have been measured." } ``
Unentitled caller (free / anonymous) → the parameter is accepted, not applied, and the ack names the (free) way to unlock it, with the exact next call. You change status, not route:
``json "input_tokens": { "value": 5000, "used": false, "why": "This number is for the model as a whole; it was NOT matched to your job size. A size-matched answer is an entitlement, and one way to earn it is free: contribute five measurements in the last seven days (POST /v1/calls, then PATCH it when the job finishes), or verify an email to lift the free tier. …" } ``
(Omit input_tokens and the field is absent; an existing caller sees no change.)
- Validation.
input_tokensmust be a positive whole number of tokens, for every tier. A non-positive, non-integer, or non-numeric value (0,-5,12.5,1e3,abc, empty) is rejected with400 {"error":"input_tokens must be a positive whole number of tokens (…).","parameter":"input_tokens"}— it is never silently absorbed or echoed back as if valid. - Same answer as
/v1/estimate-batchtime. For an entitled caller,/v1/wait'sp50_sequals/v1/estimate-batchtime'sestimate_batchtime_s(atrisk=p50) and the twoplan_for_sagree, for the same model andinput_tokens— one banding path, two routes.
- The headline percentiles come from the last hour when there are at least 5 completions in it; otherwise from the 30-day picture, and
based_on.notesays so. plan_for_sis the upper end of the confidence interval for p90, or the provider's published window when the tail cannot be bounded. It is not the median.plan_basisisinterval_upperorsla_prior.- With no measurements at all,
verdictisno_measurements_yetandplan_for_sis the provider's published window (24h for every provider insrc/confidence.js). oldest_running_sis only present forlive(paid) callers.vs_normalis the recent p50 divided by the 30-day p50, rounded to two decimals.
Example
Captured from https://batchwatch.dev, 2026-08-25 14:05 UTC, no key:
curl 'https://batchwatch.dev/v1/wait?provider=openai&model=gpt-5-nano'
{
"provider": "openai",
"model": "gpt-5-nano",
"mode": "batch",
"live": false,
"delayed_by_s": 900,
"based_on": {
"n": 1,
"window": "30d",
"note": "Too few completions in the last hour; showing the 30-day picture."
},
"coverage_30d": { "n": 1, "contributors": 1 },
"confidence": {
"level": "very_low",
"score": 0.13,
"basis": "exact_model",
"n": 1,
"contributors": 1,
"newest_age_s": 1497,
"why": "Limited by only 1 measurements, a single contributor, an upper bound we cannot pin down."
},
"robustness": {
"method": "pooled",
"contributors_voting": 0,
"note": "Pooled across all measurements: fewer than 3 contributors have enough data to vote, so a single contributor could move this number. Confidence is capped accordingly."
},
"basis": "exact_model",
"freshness": {
"last_measurement_at": "2026-08-25T13:39:43.000Z",
"last_measurement_age_s": 1497,
"last_measurement_age": "25 min",
"oldest_measurement_at": "2026-08-25T13:39:43.000Z",
"stale": false
},
"measurement": "observed completions",
"not": "a prediction for your job",
"plan_for_s": 86400,
"plan_basis": "sla_prior",
"p50_s": 1144,
"p90_s": 1144,
"p95_s": 1144,
"vs_normal": 1,
"hint": "Contribute measurements for live data and the full API."
}
That answer is honest about being thin: one measurement, one contributor, so plan_for_s falls back to OpenAI's published 24-hour window even though the one observed job took 1144s.
GET /v1/recommend
"Which rung should I actually run this on?" Every other route answers a question about one tier — how long is the batch queue, what is the p90. This one decides: given your job and your deadline, it names the rung (batch or flex) the measurements say is safe, and tells you why. Open without a key.
It exists because the intuitive choice is wrong in a way you would never guess. gpt-5-nano has the faster median (82s vs Claude Haiku's 132s) and a p90 that is 67× worse (1h33m vs 3½ minutes). Pick on "which is typically faster" and you pick the model that blows your deadline. The recommendation plans against the percentile you care about — the tail, by default — so it does not.
Parameters
| Name | Type | Required | Default | Notes |
|---|---|---|---|---|
provider | string | no | openai | An unknown provider simply has no measurements. |
model | string | yes | — | Missing gives 400 {"error":"model is required"}. |
deadline | number (seconds) | no | — | Your hard deadline. Without it we rank the rungs by measured latency but cannot call one feasible. A bad value is a 400. |
percentile | 50|90|95 (or p90) | no | 90 | The percentile the deadline is judged against. The default is the tail, because that is where deadlines live. |
n_requests | number | no | — | How many requests the job is. Drives the flex rate-limit throttle and the crossover. |
rpm / tpm | number | no | — | Your own requests- and tokens-per-minute limits. Used to throttle flex at volume; echoed back, and unverifiable by us. |
input_tokens | number | no | — | Tokens per request; only used with tpm. |
Behaviour — the decision rule
The rule is docs/flex-strategy.md §8, and this route is its executable form:
- Feasible? A rung's measured p-percentile completion time ≤ your deadline. Flex is additionally throttled by your rate limit (
n/RPM,n·tokens/TPM) — volume enters flex twice and batch barely at all, because a batch job is one file upload that bypasses those limits. A rung we have not measured is reportedunmeasured, never dropped and never guessed feasible. - Cheapest among the feasible rungs.
- On a price tie (flex vs batch cost the same after the discount — the normal case): most headroom at p wins; then batch at high
n, for its rate-limit and recovery advantages. - If nothing is feasible: we say so and name the fastest measured rung — never a fall-through to "cheapest".
- The crossover: the request count
nat which the answer flips from flex to batch under your rate limit, so you can see where your own volume moves the decision.
Real data only
The batch rung's percentiles are read through the same estimator /v1/wait quotes, and the flex rung's through the same rollup vote /v1/coverage publishes — so this route can never decide on a number a reader sees a different value for elsewhere. A rung whose deadline-percentile is not a real, measured, established number is reported as unmeasured with a reason, and the recommendation is made only over the rungs that are measured.
The flex rung is MEASURED on the models the flex prober covers. It went live with FLEX phase 01/02 (#374/#375), so a flex rung today carries a real percentile, its own n and its own confidence — openai/gpt-5-nano answers with unmeasured: false and its own n. A model the prober does not cover is still reported unmeasured with a reason, never dropped.
A measured flex rung carries measurement_basis, and it is worth reading before you act on one:
"measurement_basis": {
"measured_at_input_tokens": 8,
"floor_clause": "...",
"note": "..."
}
The two prose fields are quoted here only as "..." on purpose. This file is compiled into the site and served at /docs/query, and the real strings are generated from FLEX_INPUT_TOKENS in src/flex_basis.js — pasting them here would create a second copy that goes stale the day the probe changes, on a page answer engines cache. Call the endpoint to see them; a test asserts the doc and the code agree.
A flex call is synchronous, so its latency is the model's own work — prefill and decode — and it grows with your prompt. A batch figure is dominated by queue time, the wait to be scheduled, so the same extra work is a much smaller part of it. (Not work-free: a batch measurement is queue wait plus generation.) Our probes on both tiers send the same fixed small prompt, which is what makes the same-price comparison like for like; it also means a flex percentile is the floor a larger prompt starts from, while the batch one is the wait you can plan against. The field appears on the flex entries in rungs[] and on recommendation / fastest_measured when either names flex, and the why sentence states it in words. Batch rungs do not carry it — the basis describes a synchronous figure.
How well-measured is the advice — confidence and window per rung
Every measured rung — in the rungs array, and on the recommendation and fastest_measured it may name — carries two extra fields, so a recommendation is never a bare number:
confidence— the same grade/v1/waitand/v1/coveragepublish for the model (level,score,n,contributors, and a plain-Englishwhy). It is read straight from the estimator, so the confidence you see here and on/v1/waitfor one model can never disagree.window— the tail you should plan for:{ "plan_for_s", "basis" }. Same figure/v1/waitcallsplan_for_s.basisisinterval_upperwhen it is the measured upper bound of the confidence interval, orsla_priorwhen the tail cannot be bounded from the data yet and we fall back to the provider's own published completion window (labelled so you never mistake it for something we observed).
An unmeasured rung carries neither — it appears in unmeasured_rungs with only a reason, never a fabricated confidence or window (#30/#251).
Example — the tail decides
GET /v1/recommend?model=gpt-5-nano&deadline=600&percentile=90
{
"provider": "openai",
"model": "gpt-5-nano",
"inputs": { "deadline_s": 600, "percentile": 90, "n_requests": null },
"percentile": 90,
"deadline_s": 600,
"verdict": "infeasible",
"recommendation": null,
"fastest_measured": { "rung": "batch", "model": "gpt-5-nano", "measured_p_s": 5554 },
"unmeasured_rungs": [ { "rung": "flex", "model": "gpt-5-nano", "reason": "no_measurements" } ],
"why": "No measured rung makes your 600s deadline at p90. …"
}
At p90, gpt-5-nano cannot make a ten-minute deadline — its tail is 1h33m — and the route says so rather than recommending the model whose median looks fast. Ask about a model whose measured tail clears the deadline and you get "verdict": "feasible" with the rung named in recommendation — carrying its confidence and window so you can see how well-measured the advice is:
{
"verdict": "feasible",
"recommendation": {
"rung": "batch", "provider": "anthropic", "model": "claude-haiku-4-5",
"measured_p_s": 207, "effective_completion_s": 207, "headroom_s": 3393,
"confidence": { "level": "high", "n": 545, "contributors": 4, "why": "…" },
"window": { "plan_for_s": 260, "basis": "interval_upper" }
}
}
GET /v1/accuracy
"How often was your answer right?" Open without a key, deliberately: it is the one number that can be checked against us rather than taken on trust.
Every other route on this page tells you what the queue did. This one tells you how often our advice about it held.
Parameters
| Name | Type | Required | Default | Notes |
|---|---|---|---|---|
provider | string | no | all | Filters to one provider. |
model | string | no | all | Filters to one model. |
Window is fixed at 30 days.
Where the numbers come from
When a client calls should_batch() it caches the verdict it acted on, the deadline it asked against, and the p90 we quoted. When it later PATCHes the completion for that same job, it attaches those three values. The server judges them against its own measured duration — the difference between two server timestamps — and stores the outcome.
The client supplies only the figures we gave it. It never supplies the duration. That split is the whole integrity guarantee: a caller cannot make our hit rate look better by reporting a flattering runtime.
Only the authenticated PATCH completion path feeds this. The bulk history import does not, because its durations are client timestamps. Imported history must not be able to grade us.
Response
{
"provider": null, "model": null, "window": "30d",
"insufficient": true,
"n": 0, "sources": 0, "confidence": "insufficient",
"verdict_hit_rate": null, "p90_calibration": null, "p90_n": 0,
"by_verdict": {},
"required": { "n": 20, "sources": 3 }
}
| Field | Meaning |
|---|---|
n | Verifiable judgements in the window. A verdict we could not check (missing deadline, non-verifiable verdict) is not counted. |
sources | Distinct contributing keys. |
verdict_hit_rate | Share of judgements that turned out correct — null until the gate is met. |
p90_calibration | Share of jobs that finished within the p90 we quoted. Should sit near 0.90 if the estimator is honest. |
by_verdict | Per-verdict {n, correct}. |
required | The gate: what is still missing before a rate is published. |
The gate, and why there is one
Below 20 judgements from 3 distinct sources, insufficient is true and both rates are null. n, sources and confidence are still reported, so a page can honestly render "building (n=4)".
A hit rate computed from four measurements is not a weak number, it is a different kind of claim altogether — and publishing it would be exactly the fabrication this whole service exists to avoid. See interpreting.md.
run_sync counts toward the gate: "we said batch would miss" is verifiable too, and it is correct precisely when the job overran the deadline.
GET /v1/curve
The empirical distribution: "what share of jobs had finished after X seconds?" Open without a key, because the model picker on the front page calls it on every click.
Parameters
| Name | Type | Required | Default |
|---|---|---|---|
provider | string | no | openai |
model | string | yes | — |
mode | string | no | batch |
Behaviour
- Window is fixed at 30 days.
- The only hard limit is two measurements: a curve needs two points to connect. Below that,
drawableisfalsewith an emptypointsarray and awhystring. - Between two and eight the curve is drawn, but the response carries
thin: trueand awhythat says the number of observations. Drawing and claiming are separate: the points are shown so you can count them, but the page will not quote a ninetieth percentile off three jobs, because on three jobs that figure is literally the slowest of them. - The current thresholds are reported by
GET /v1/probeunderthresholds, so a client never has to hardcode them. They can be changed per deployment. pointsis an array of[duration_s, percent_completed]pairs, at most 60 of them, sampled evenly in rank (not in time) so the tail survives; the slowest measurement is always the last point, at100.p50_s/p90_s/p95_sare the robust per-contributor figures where three or more contributors have earned a vote (see interpreting.md). The curve itself is drawn from the contributor-capped pool.- Delayed callers may be served a precomputed rollup; when they are, the response carries
precomputed_at.
A percentile can be null, and that is an answer
Since the censoring-aware estimator shipped (task #107), a percentile field may be null with robustness.method set to "unidentified". Clients must handle this. A field that used to always be a number no longer is.
It means: too much of the evidence is still running (or was abandoned) for that percentile to be identified, so any number would be invented. It does not mean "no data" — method: "none" means that, and n will be 0.
Worked example, measured on a seeded dataset of 60 completed jobs at ~150 s plus 12 that had been queuing for two hours:
| before #107 | after #107 | |
|---|---|---|
p50_s | 152 | 152 |
p90_s | 164 | null |
robustness.method | pooled | unidentified |
The old figure of 164 s was produced by discarding the twelve running jobs. It said nine in ten jobs finish inside three minutes while twelve of them had been waiting two hours — the error is roughly 48×, and it points the dangerous way: a caller plans for three minutes and misses the deadline.
Note p50_s is unchanged. The estimator answers what is identified and declines only what is not; it does not refuse across the board out of caution.
What to do with a null: read robustness.note, which names the longest wait actually observed. The true figure is at least that. Treat it as a lower bound, not as an outage, and prefer the synchronous path if your deadline is near it.
Two counting fields go with this:
ncounts fully observed measurements only — a job we are still waiting on is not a measurement of anything.n_censoredcounts the right-censored rows behind the estimate: still running, abandoned, or expired. They shape the curve (that is the point) but they are deliberately not folded inton, so the two cannot be silently confused.
Example — below the hard limit (one measurement)
Captured from https://batchwatch.dev, 2026-08-25 14:05 UTC, no key:
curl 'https://batchwatch.dev/v1/curve?provider=openai&model=gpt-5-nano'
{
"provider": "openai",
"model": "gpt-5-nano",
"mode": "batch",
"live": false,
"delayed_by_s": 900,
"n": 1,
"contributors": 1,
"window": "30d",
"precomputed_at": "2026-08-25T14:00:44.000Z",
"freshness": {
"last_measurement_at": "2026-08-25T13:39:43.000Z",
"last_measurement_age_s": 1501,
"last_measurement_age": "25 min",
"stale": false
},
"measurement": "observed completions",
"not": "a prediction for your job",
"drawable": false,
"points": [],
"why": "1 measurement - a curve needs at least two points to connect.",
"confidence": {
"level": "very_low",
"score": 0.13,
"basis": "exact_model",
"n": 1,
"contributors": 1,
"newest_age_s": 1501,
"why": "Limited by only 1 measurements, a single contributor, an upper bound we cannot pin down."
}
}
Example — drawable
Captured from a local wrangler dev instance seeded with 24 synthetic measurements from 3 keys. The durations below were made up to show the shape; they are not measured queue times.
{
"provider": "openai",
"model": "gpt-5-mini-docs",
"mode": "batch",
"live": false,
"delayed_by_s": 900,
"n": 24,
"contributors": 3,
"window": "30d",
"freshness": {
"last_measurement_at": "2026-08-24T14:27:04.000Z",
"last_measurement_age_s": 85739,
"last_measurement_age": "24 hours",
"stale": false
},
"measurement": "observed completions",
"not": "a prediction for your job",
"drawable": true,
"points": [[790,4.2],[820,8.3],[860,12.5],[880,16.7],[910,20.8],[940,25],
[980,29.2],[1010,33.3],[1060,37.5],[1120,41.7],[1150,45.8],
[1180,50],[1250,54.2],[1290,58.3],[1330,62.5],[1390,66.7],
[1420,70.8],[1480,75],[1580,79.2],[1610,83.3],[1720,87.5],
[2300,91.7],[2400,95.8],[2600,100]],
"p50_s": 1150,
"p90_s": 2400,
"p95_s": 2400,
"robustness": { "method": "per_contributor" },
"max_s": 2600,
"confidence": {
"level": "very_low",
"score": 0.21,
"basis": "exact_model",
"n": 24,
"contributors": 3,
"newest_age_s": 85739,
"why": "Limited by an upper bound we cannot pin down, nothing measured for 24 hours."
}
}
(The API prints one points pair per line; the array is folded here for readability. Every value is verbatim.)
GET /v1/coverage
What has been measured at all, in the last 30 days. Open, never delayed, no parameters.
Behaviour
contributors counts distinct independent third-party sources — our own prober and our own backfill/import are excluded, because a measurement's provenance (probe / first_party / third_party / unknown) is recorded at ingest and only third_party is a contributor (#276). crowdsourced is true only when at least two such independent sources exist, so it can never read true over data that is all ours. Everything key-less counts as one source. probes still counts our own prober's completions, disclosed alongside, and probe_only is true when every call we hold for that series came from our own prober — full disclosure of how much outside corroboration a figure has, published on the row rather than left for you to infer.
n is the sample the percentiles rest on — the measurements behind the median and the tail on the very same row, and the same n /data.csv publishes for that series. Read n and p50_s together and you are reading one population, which is what makes the pair worth quoting: the sample size belongs to the figure standing beside it.
completed_calls is everything we hold for that series in the window — every completed, non-excluded call, including the extremes the estimator quarantines, the rows past the tier cut, and any without a usable duration. It is disclosed in full, under its own name, so the distance between the two numbers is visible rather than averaged away. That distance is a signal in its own right: it says we hold rows and do not let all of them into a published figure. n is never larger than completed_calls.
requested_service_tier is the pricing tier the row is for — standard (batch and full-price synchronous) or flex (the half-price synchronous tier). Each tier is its own series with its own latency, so a model measured on both appears as two rows and their numbers are never pooled.
answerable is the same verdict /v1/wait reaches on the same rows: it is true only when there is a median we would publish — a p50_s the estimator can identify. When a model's jobs are mostly still running or were abandoned the tail is not identified, so /v1/wait refuses and answerable here is false with p50_s: null. The two routes read the SAME estimator vote and cannot disagree; a null percentile is a gap, never a zero. completed_calls, contributors, probes and crowdsourced stay truthful regardless — a model we will not quote a percentile for still shows you everything we have on it.
p90_s and p95_s are the tail, read from the SAME estimator vote as p50_s (#301). p90_s is byte-for-byte the p90 llms.txt publishes for the same model in the same window — both read the one window-wide rollup vote, so those two surfaces cannot drift. (/v1/hours also reports a p90, but a different one: it is a nearest-rank percentile over a single hour-of-day bucket, a different method over a different population, so it is not — and is not claimed to be — the same number.) The tail is null on an honest gap in two cases: when the model is unanswerable, all three of p50_s/p90_s/p95_s are null together; and when the median is identified but the tail specifically is not (the estimator votes per-quantile), p50_s is a number while p90_s/p95_s are null. What never happens is a numeric p90_s beside a null p50_s — an unidentified median gates the whole row. A null percentile is always a gap, never a lowered threshold or a fabricated number. Only {p50,p90,p95} are computed; p75/p99 are not published here — inventing them from these would be a fabricated number, so they wait on an estimator change.
confidence here is the level string only, computed the same way as elsewhere.
measurement_basis appears on a flex row that publishes a percentile, and says what that latency was measured on (#405):
"measurement_basis": {
"measured_at_input_tokens": 8,
"floor_clause": "...",
"note": "..."
}
It is here because this endpoint returns a model's flex/sync row and its batch row in the same array — and a consumer will read them side by side, which is what the endpoint is for. A flex call is synchronous, so its latency is the model's own work and it grows with your prompt; a batch figure is dominated by queue time, the wait to be scheduled, so the same extra work is a much smaller part of it. Our probes on both tiers send the same fixed small prompt, so the same-price comparison is like for like — and the flex number is the floor a larger prompt starts from, while the batch number is the wait you can plan against. The refusal rate is unaffected by prompt size: admission control happens before inference, so a small probe measures it exactly as well as a large one would.
The field is absent on standard/batch rows (the basis describes a synchronous figure) and on a flex row whose percentiles are all null (there is no published number for it to qualify). Additive: no existing field changed shape.
Example
Illustrative of the contract. The first model has enough completions to identify a median and is answerable; note n (48) below completed_calls (51) — three rows are held but not carried by the published percentile, which is exactly the disclosure the two fields exist to make. The second model has a single measurement, so the tail is not identified: answerable is false and p50_s is null, exactly as /v1/wait would refuse it. Both rows here are the standard tier; a flex series for the same model would appear as its own row with "requested_service_tier": "flex".
curl https://batchwatch.dev/v1/coverage
{
"window": "30d",
"models": [
{
"provider": "anthropic",
"model": "claude-haiku-4-5",
"mode": "batch",
"requested_service_tier": "standard",
"n": 48,
"completed_calls": 51,
"contributors": 3,
"probes": 12,
"crowdsourced": true,
"probe_only": false,
"p50_s": 146,
"p90_s": 512,
"p95_s": 690,
"latest_at": 1787664703,
"last_measured_age_s": 1970,
"answerable": true,
"confidence": "medium"
},
{
"provider": "openai",
"model": "gpt-5-nano",
"mode": "batch",
"requested_service_tier": "standard",
"n": 1,
"completed_calls": 1,
"contributors": 1,
"probes": 1,
"crowdsourced": false,
"probe_only": true,
"p50_s": null,
"p90_s": null,
"p95_s": null,
"latest_at": 1787665183,
"last_measured_age_s": 1490,
"answerable": false,
"confidence": "very_low"
}
]
}
latest_at is a unix timestamp in seconds — the only timestamp in the API that is not also given as ISO-8601.
GET /v1/should-i-batch
Gated. "Batch or sync, given my deadline?" The route does not decide for you: you send your own limit and it answers against it.
Parameters
| Name | Type | Required | Default | Notes |
|---|---|---|---|---|
provider | string | no | openai | |
model | string | yes | — | |
risk | string | no | p90 | p50, p90 or p95. This is a decision threshold in disguise — see "What risk means" below. Not validated here: an unknown value silently falls back to p90 (RISK_Q[risk] ?? 0.9). Contrast /v1/estimate-batchtime, which rejects it. |
max_wait | duration | no | — | 900, 15m, 2h, 1d. A value that does not parse is treated as absent. |
input_tokens | number | no | — | Omit it and no cost figures are produced. |
output_tokens | number | no | — | The best cost basis: you know it. |
max_tokens | number | no | — | Used as a ceiling when output_tokens is absent; the saving shown is then the most you could save. |
n_shards | number | no | — | Presence adds the fanout block. The number of shards in one run. |
sync_budget | duration | no | 0 | Time to reserve for re-running stragglers synchronously. Only read when n_shards is present. |
Duration syntax: an integer or decimal, optionally followed by s, m, h or d. No suffix means seconds.
Fan-out: one run of N shards
A real job is rarely one request. You shard an eval across twenty calls and you need all twenty in before the deadline — and that is not the question "will one make it?".
The arithmetic is unforgiving. At p90 per shard, P(all 20 on time) is not 0.9. It is 0.9^20 = 0.12. Send n_shards and you get it computed against the measured distribution for that model:
curl 'https://batchwatch.dev/v1/should-i-batch?provider=openai&model=gpt-5.6-sol&max_wait=900&n_shards=20&sync_budget=60' -H "authorization: Bearer $BATCHWATCH_KEY"
{
"fanout": {
"verdict": "stage_release_window",
"n_shards": 20,
"n": 81,
"deadline_s": 900,
"p_ontime": 0.9167,
"p_correct": 1,
"p_shard_ok": 0.9167,
"p_all_ontime": 0.1755,
"p_all_ok": 0.1755,
"independence": "assumed_conservative",
"release_window": {
"cutoff_s": 840,
"sync_budget_s": 60,
"expected_outstanding_frac": 0.0833,
"expected_sync_fallback_shards": 2
},
"note": "Batch the bulk now: about 18 of your 20 shards land by 840s. Re-run the remaining ~2 synchronously at 840s and you hit your 900s deadline with 60s to spare."
}
}
release_window is the useful part. Rather than answering "batch or don't", it gives you a staged plan: batch everything, cut off at cutoff_s, and re-run whatever is still outstanding synchronously inside the budget you reserved.
Two things to know about the numbers
Independence is assumed, and that is conservative. Shards share the provider's queue, so they are correlated in reality. p^N assumes they are not, which overstates the risk that some shard misses. It is a safe upper bound on risk, never an optimistic one. The response says so in independence_note.
p_correct counts failures, not censoring. Jobs that came back failed or expired count against it. Jobs marked abandoned do not — those are runs where the measuring client stopped waiting, so the true duration is unknown and longer. Counting them as provider failures would blame the provider for our own impatience.
Without n_shards, the response is byte-for-byte what it was before the block existed.
Verdicts
verdict | Meaning |
|---|---|
run_batch | The planning number fits inside your max_wait. |
run_sync | It does not — or the median fits but the upper bound does not. If a measured quiet hour would fit, the note says so, but scheduling against it is suspended (see below). |
no_deadline_given | You sent no max_wait. Carries required_patience_s. |
insufficient_data | Only reachable when the handler is called without evidence, which the HTTP route never does. |
Withdrawn (ALG-4, #132): the
batch_atverdict and itsbatch_at_hour_utc/batch_wait_at_target_sfields are gone as of API2026-v2. Per-hour buckets are too thin (~6 rows/hour) and hour-of-day is not yet identifiable, so a "batch at 02:00" verdict cannot be made honestly today (docs/ALGORITHM.md §L9.4). When the queue is too slow now the verdict isrun_sync; if a quiet hour looks cheaper thenotereports it, without shipping a scheduling verdict we cannot stand behind. The/v1/hoursclock still serves the whole daily profile as measurements to read.
The decision is made on planning_wait_s (the interval's upper end, or the provider's window), never on observed_wait_s. Both are in the response, so the gap is visible.
The decision-shaped fields (ALG-5, #133)
A percentile alone makes the caller do the reasoning. Every quote now also carries the three figures the decision actually rests on:
| Field | Meaning |
|---|---|
decision_threshold | tau, what your risk asserts (p50→0.50, p90→0.90, p95→0.95). Always present.* |
meets_deadline_probability | p_hat = P(D ≤ your max_wait) = 1 − S(W), read off the same survival curve as the quoted wait. null with no deadline, no measurements, or an unidentified tail. |
expected_lateness_s | E[(D − W)+], the mean seconds you would be late. null unless the tail is identified — an unbounded tail cannot yield a finite figure, and we never invent one (#30). |
Both curve-derived figures are read off the same censoring-aware Kaplan–Meier the quoted wait comes from, so the probability and the wait cannot tell different stories about your deadline.
What risk means
risk is your loss ratio in disguise: tau* = 1 − L_sync/M, where L_sync is the certain saving from batching and M is your total cost of a missed deadline.
risk | asserts | tau* (decision_threshold) |
|---|---|---|
p50 | a missed deadline costs 2× the batch saving | 0.50 |
p90 | a missed deadline costs 10× the batch saving | 0.90 |
p95 | a missed deadline costs 20× the batch saving | 0.95 |
risk remains the only knob — no new parameter — but now you can see which loss ratio your choice picks (decision_threshold) and the probability behind the verdict (meets_deadline_probability).
Cost
Prices live in src/stats.js (PRICING), keyed provider/model. At the time of writing only nine OpenAI models have rates. Batch is exactly half of sync on both input and output.
If output tokens are unknown, no cost is produced: cost_basis.known is false and the raw rates are handed back instead. Absence gives absence — the code deliberately does not assume zero output tokens. If the model is not in PRICING at all, the whole cost_basis block is omitted.
Example — no deadline given, thin data
Captured from https://batchwatch.dev, 2026-08-25 14:05 UTC, no key. This call consumed one of the twenty free trial calls.
curl 'https://batchwatch.dev/v1/should-i-batch?provider=openai&model=gpt-5-nano&input_tokens=1000'
{
"provider": "openai",
"model": "gpt-5-nano",
"risk": "p90",
"batch": { "p50_s": 1144, "p90_s": 1144, "n": 1 },
"sync": null,
"cost_basis": {
"known": false,
"why": "Output tokens are decided by the model, not by you, and nobody has measured enough real jobs on this model for us to estimate them. We will not guess on your behalf.",
"rates_usd_per_mtok": { "input": 0.05, "output": 0.4 },
"note": "Batch is exactly half of sync on both rates, so your saving is half of whatever you would have paid. Send output_tokens if you have a figure, or max_tokens for an upper bound."
},
"confidence": {
"level": "very_low",
"score": 0.13,
"basis": "exact_model",
"n": 1,
"contributors": 1,
"newest_age_s": 1502,
"why": "Limited by only 1 measurements, a single contributor, an upper bound we cannot pin down."
},
"planning_basis": "sla_prior",
"robustness": {
"method": "pooled",
"contributors_voting": 0,
"note": "Pooled across all measurements: fewer than 3 contributors have enough data to vote, so a single contributor could move this number. Confidence is capped accordingly."
},
"observed_wait_s": 1144,
"planning_wait_s": 86400,
"decision_threshold": 0.9,
"meets_deadline_probability": null,
"expected_lateness_s": null,
"verdict": "no_deadline_given",
"required_patience_s": 86400,
"note": "Batch needs 1144s of patience at p90. Send max_wait and we answer against your limit.",
"contributors": 1,
"basis": "exact_model",
"freshness": {
"last_measurement_at": "2026-08-25T13:39:43.000Z",
"last_measurement_age_s": 1502,
"last_measurement_age": "25 min",
"stale": false
},
"trial": {
"calls_used": 4,
"calls_total": 20,
"calls_left": 16,
"note": "Free trial - no contribution needed yet."
}
}
Example — with a deadline
Captured from a local wrangler dev instance with 24 synthetic measurements from 3 keys. Synthetic durations, real response shape. Note the verdict: the median fits inside the 45-minute limit, but the tail cannot be bounded, so the answer is run_sync.
{
"provider": "openai",
"model": "gpt-5-mini-docs",
"risk": "p90",
"batch": { "p50_s": 1150, "p90_s": 2400, "n": 24 },
"sync": null,
"confidence": {
"level": "very_low",
"score": 0.21,
"basis": "exact_model",
"n": 24,
"contributors": 3,
"newest_age_s": 85633,
"why": "Limited by an upper bound we cannot pin down, nothing measured for 24 hours."
},
"planning_basis": "sla_prior",
"robustness": {
"method": "per_contributor",
"contributors_voting": 3,
"note": "Median of 3 established contributors' own figures. A contributor earns a vote by measuring on at least 3 separate days, so neither more data nor more accounts can move this number."
},
"observed_wait_s": 2400,
"planning_wait_s": 86400,
"decision_threshold": 0.9,
"meets_deadline_probability": null,
"expected_lateness_s": null,
"your_limit_s": 2700,
"batch_wait_s": 2400,
"verdict": "run_sync",
"note": "The median says 2400s, inside your 2700s limit - but on 24 measurements we cannot rule out 86400s. Limited by an upper bound we cannot pin down, nothing measured for 24 hours. Run sync until the data is thicker.",
"contributors": 3,
"basis": "exact_model",
"freshness": {
"last_measurement_at": "2026-08-24T14:27:04.000Z",
"last_measurement_age_s": 85633,
"last_measurement_age": "24 hours",
"stale": false
},
"trial": { "calls_used": 1, "calls_total": 20, "calls_left": 19, "note": "Free trial - no contribution needed yet." }
}
There is no cost_basis in that response because gpt-5-mini-docs is not in the price list.
GET /v1/estimate-batchtime
Gated. "How long will my job take?" — answered by finding the measurements that resemble your job in size and reporting what those did.
Parameters
| Name | Type | Required | Default | Notes |
|---|---|---|---|---|
provider | string | no | openai | |
model | string | yes | — | |
risk | string | no | p50 | p50, p90, p95. Validated: anything else is 422. Note the default differs from /v1/should-i-batch. |
input_tokens | number | no | — | Must be greater than 0 to be used; otherwise treated as absent. |
Size matching
The band is multiplicative, not additive: the route tries input_tokens / f to input_tokens * f for f in 2, 4, 8, and takes the narrowest band with at least 10 measurements. basis becomes size_matched when a band was found, all_sizes when it was not, sla_prior when there is nothing at all.
size_effect compares the band with jobs of other sizes. detected: true means the ratio fell below 0.85 or rose above 1.18. detected: null means only jobs of one rough size have ever been measured, so the question cannot be answered at all.
Example
Captured from https://batchwatch.dev, 2026-08-25 14:05 UTC, no key. Consumed a trial call.
curl 'https://batchwatch.dev/v1/estimate-batchtime?provider=openai&model=gpt-5-nano&input_tokens=50000&risk=p50'
{
"provider": "openai",
"model": "gpt-5-nano",
"risk": "p50",
"live": false,
"delayed_by_s": 900,
"input_tokens": 50000,
"estimate_batchtime_s": 1144,
"confidence": {
"level": "very_low",
"score": 0.13,
"basis": "exact_model",
"n": 1,
"contributors": 1,
"newest_age_s": 1522,
"why": "Limited by only 1 measurements, a single contributor, an upper bound we cannot pin down."
},
"basis": "all_sizes",
"plan_for_s": 86400,
"plan_basis": "sla_prior",
"interval_s": { "lo": 1144, "hi": null, "bounded": false },
"n": 1,
"contributors": 1,
"robustness": {
"method": "pooled",
"contributors_voting": 0,
"note": "Pooled across all measurements: fewer than 3 contributors have enough data to vote, so a single contributor could move this number. Confidence is capped accordingly."
},
"freshness": {
"last_measurement_at": "2026-08-25T13:39:43.000Z",
"last_measurement_age_s": 1522,
"last_measurement_age": "25 min",
"stale": false
},
"measurement": "observed completions of comparable jobs",
"not": "a model of how your job will behave",
"size_effect": {
"detected": false,
"note": "Not enough measurements near 50000 input tokens, so this is the model's overall queue time, not a size-matched one."
},
"trial": { "calls_used": 5, "calls_total": 20, "calls_left": 15, "note": "Free trial - no contribution needed yet." }
}
Example — rejected risk value
Captured from https://batchwatch.dev, 2026-08-25 14:05 UTC:
curl 'https://batchwatch.dev/v1/estimate-batchtime?provider=openai&model=gpt-5-nano&risk=p99'
{ "error": "risk must be p50, p90 or p95" }
Status 422. Verified in the same run that this did not consume a free trial call: the counter stood at 7 before it, the 422 came back, and the next successful gated call reported calls_used: 8.
GET /v1/conditions
Gated. "Is the queue slow right now?" — the last hour against the 30-day baseline.
Parameters
| Name | Type | Required | Default |
|---|---|---|---|
provider | string | no | openai |
model | string | yes | — |
Behaviour
The hour window is 3600 + delay seconds long and is cut at now - delay, so a delayed caller does not get a live signal for free. verdict is much_slower at ratio 3 or above, slower_than_usual at 1.5 or above, normal below that, and unknown when either side of the ratio is missing.
confidence on this route is the coarse string from confidenceOf() — insufficient, low, medium, high — not the object used elsewhere. It is insufficient below MIN_N (20) measurements or MIN_KEYS (3) contributors, and the route still answers.
Example
Captured from https://batchwatch.dev, 2026-08-25 14:05 UTC, no key. Consumed a trial call.
curl 'https://batchwatch.dev/v1/conditions?provider=openai&model=gpt-5-nano'
{
"provider": "openai",
"model": "gpt-5-nano",
"window": "1h",
"live": false,
"delayed_by_s": 900,
"completed": { "n": 1, "p50_s": 1144, "p90_s": 1144 },
"baseline_30d": { "n": 1, "p50_s": 1144, "p90_s": 1144 },
"ratio_p50": 1,
"confidence_note": "Thin data - treat this as a hint, not a reading.",
"verdict": "normal",
"running": { "oldest_s": null },
"confidence": "insufficient",
"contributors": 1,
"trial": { "calls_used": 6, "calls_total": 20, "calls_left": 14, "note": "Free trial - no contribution needed yet." }
}
GET /v1/distribution
Gated. p50 through p99 plus an hourly profile. The only route with a hard threshold: below 20 measurements from 3 distinct contributors it refuses, and explains why in the response.
Parameters
| Name | Type | Required | Default | Notes |
|---|---|---|---|---|
provider | string | no | openai | |
model | string | yes | — | |
mode | string | no | batch | |
window | duration | no | 30d | Same duration syntax as max_wait. An unparseable value falls back to 30 days. |
Thresholds come from the MIN_N and MIN_KEYS environment variables (currently 20 and 3 in wrangler.toml).
Example — refused
Captured from https://batchwatch.dev, 2026-08-25 14:05 UTC, no key. The refusal is a 200, so it consumed a trial call.
curl 'https://batchwatch.dev/v1/distribution?provider=openai&model=gpt-5-nano'
{
"provider": "openai",
"model": "gpt-5-nano",
"mode": "batch",
"verdict": "insufficient_data",
"live": false,
"delayed_by_s": 900,
"n": 1,
"contributors": 1,
"confidence": "insufficient",
"required": { "n": 20, "contributors": 3 },
"why": "This is the one route with a hard threshold. A full distribution with an hourly profile drawn on a handful of jobs would let you read one account's working pattern out of it. The other routes answer at any n, with the thinness stated.",
"instead": "/v1/wait?provider=openai&model=gpt-5-nano answers now, with graded confidence.",
"trial": { "calls_used": 7, "calls_total": 20, "calls_left": 13, "note": "Free trial - no contribution needed yet." }
}
Example — answered
Captured from a local wrangler dev instance with 24 synthetic measurements from 3 keys, called with a contributor key. Synthetic durations, real shape.
{
"provider": "openai",
"model": "gpt-5-mini-docs",
"mode": "batch",
"window_s": 2592000,
"live": false,
"delayed_by_s": 300,
"n": 24,
"p50_s": 1215,
"p75_s": 1505,
"p90_s": 2126,
"p95_s": 2385,
"p99_s": 2554,
"max_s": 2600,
"contributors": 3,
"confidence": "low",
"by_hour_utc": [
{ "hour_utc": 12, "n": 9, "p50_s": 1610, "p90_s": 2440, "p95_s": 2520 },
{ "hour_utc": 13, "n": 12, "p50_s": 1090, "p90_s": 1286, "p95_s": 1308 }
],
"quota": {
"tier": "contributor",
"calls_used": 4,
"calls_limit": 10000,
"calls_left": 9996,
"window": "7 days"
}
}
by_hour_utc only lists hours with at least 5 measurements. The percentiles are computed over the full set — the tail is not trimmed. (There used to be a dropped_outliers field counting measurements discarded for exceeding 8× the median; that median×8 cap was removed because it under-reported p90 by up to 40× when a real fat tail was mistaken for outliers, so the field is gone.)
Note that the percentiles on this route come from distribution() directly — they are not the per-contributor robust figures used by /v1/wait and /v1/curve. That is why p50_s here (1215) differs from p50_s on /v1/curve (1150) over the same 24 measurements. Both numbers are correct; they answer slightly different questions. See interpreting.md.
Bulk dataset download — GET /data.csv, GET /data.json
The routes above are query endpoints: you ask about one (provider, model) and get one answer. To take the whole measured dataset in one file — every provider, every model, both modes — download it:
GET /data.csv— the complete dataset as CSV, streamed, with aContent-Disposition: attachmentheader so a browser saves it as a file. The first line is a#-prefixed comment carrying the window, the licence and the generation time; the second line is the column header; every remaining line is one measurement row.GET /data.json— the same dataset as a single JSON document:
``json { "dataset": "batchwatch batch API queue-time measurements", "window": "30d", "generated_at": "2026-08-31T12:00:00Z", "license": "https://creativecommons.org/licenses/by/4.0/", "attribution": "batchwatch — CC BY 4.0", "source_url": "https://batchwatch.dev/data.csv", "schema": { "provider": "Provider slug, e.g. \"openai\".", "…": "…" }, "columns": ["provider", "provider_name", "model", "mode", "window", "n", "contributors", "p50_s", "p90_s", "p95_s", "max_s", "newest_at", "oldest_at", "computed_at", "requested_service_tier"], "count": 42, "rows": [ { "provider": "openai", "model": "gpt-4o", "mode": "batch", "n": 412, "p50_s": 146, "p90_s": 900, "…": "…" } ] } ``
Completeness
Both files are the complete published dataset — there is no cap, no truncation, no sampling, and no freshness filter. Every model we publish at the public 15-minute delay tier, in both batch and sync mode, with at least one completed measurement (n > 0), is a row. A model that has been seen but has no completed measurement in the window has nothing to export and is legitimately absent — that is the same n > 0 gate every other public surface applies, not a cap. The figures are the same rollup-derived percentiles that /v1/coverage and the per-model pages serve, so a number in the export can never disagree with the same number on a model page.
Columns
| Column | Meaning |
|---|---|
provider | Provider slug, e.g. openai. |
provider_name | Human-readable provider name. |
model | Model id, e.g. gpt-4o. |
mode | batch or sync. |
window | Rolling window the figures cover (always 30d). |
n | Completed jobs behind the figures (sample size). |
contributors | Distinct independent third-party sources. |
p50_s / p90_s / p95_s | Median / 90th / 95th percentile queue time, seconds. |
max_s | Slowest completion observed, seconds. |
newest_at / oldest_at | Unix seconds of the newest / oldest measurement in the window. |
computed_at | Unix seconds the rollup was computed. |
requested_service_tier | Pricing tier the figures are for: standard (batch and full-price sync) or flex (the half-price synchronous tier). Its own series, never pooled with standard. |
*A row is one series, and a series is (provider, model, mode, requested_service_tier).* The flex tier is measured separately end to end — its own probes, its own rollup, its own percentiles — because a half-price synchronous queue behaves nothing like the standard one, and pooling them would publish a median that describes neither. So a model we measure on both tiers has two rows, and requested_service_tier is what tells them apart. Group on all four columns.
Columns are only ever appended, never reordered or removed, so a pinned consumer keeps reading. requested_service_tier is the most recent addition and is the last column, exactly as that promise requires. Both files are public and cacheable (the same key-invariant cache as /v1/coverage); no API key is read, and no per-caller data is carried.
Licence
The measurement dataset is licensed CC BY 4.0 (operator decision, 2026-08-31) — the published aggregates are open data you may reuse, including commercially, for the price of a credit and a link back. The download carries the same canonical licence URL the site's schema.org/Dataset structured data and the in-band /v1/* licence block declare, plus a ready-to-paste attribution string (batchwatch — CC BY 4.0). See /license for the full statement. This open grant covers the published aggregates only; live, undelayed access under a paid plan is governed by the Terms, not by CC BY.