batchwatch

Nobody publishes how long each LLM price tier actually takes. We measure it, and we give it away.

batchwatch is a supplier of information. OpenAI, Anthropic and Google sell the same model output at more than one price — the asynchronous batch tier, a synchronous flex tier where they offer one, and the full-price standard call. What separates those tiers is latency and refusal, and not one of the three providers publishes either. “Within 24 hours” is a ceiling: the widest the wait can ever be, and silent on what it is today. So we run real jobs against all three, around the clock, and publish the whole distribution — median, p90, p95, each with the sample size behind it. That is the entire offer, and the numbers exist nowhere else. They are measured rather than estimated, they belong to no vendor, and they are free.

The median and the p90 rank models differently. Every public benchmark ranks on the median, and the median is not the figure a deadline cares about. Both are measured, both are in the table below, and both are in /v1/coverage with the sample size behind each.

Named precisely: the half-price tiers are the provider’s asynchronous batch API — the /v1/batches endpoints, where you submit now and the results come back within a published deadline — and, where a provider offers one, a synchronous flex tier at the same discount. That is a different product from in-server micro-batching such as vLLM’s, which coalesces requests in milliseconds and has no queue to measure. Every percentile we publish here is measured on the async batch queues at all three providers — the wait that comes with the discount, and the one number nobody else publishes at all.

We only ever receive timing and token counts — never your prompts, completions or any content. The client sends a fixed allowlist of fields and nothing else, by construction. See the send path on GitHub, the field-by-field list, or where the data lives.

One field, no email, no card — and the key comes back with the single command that sends your first measurement.

openai · gpt-5.6-sol · 9.7k in / 4.5k out
How long can you wait for this job?
…

Three prices for the same output, and no published latency for any of them

Every major provider will sell you the same tokens at half price if you can wait — 50% off, every input and output token, on the asynchronous batch tier and on the synchronous flex tier where they offer one. The discount is real and it is theirs to give. What none of them will tell you is the thing that decides whether you can use it: how long that tier actually takes today, and how often it refuses the work outright.

That is where it stays, because the only promise the provider gives you is “within 24 hours”. No engineer can build a feature on that and no product owner can plan around it, so the safe move is to pay double and keep the latency you understand. The recoverable half sits there, paying full price, and no pricing page will ever surface it.

The 50% and the 24 hours are the providers’ own published numbers. That the discount goes mostly untaken is our reading of the market, not a measurement — and saying which is which is the habit this whole site is built on.

A ceiling is not an estimate. The queue has good days and bad days, and the 24-hour number is the same on both. That is the whole gap: not that batch is slow, but that nobody tells you which day you are having.

That is the gap we fill, and filling it is the whole of what we do. We measure the tiers and we publish the measurements. Median, p90, p95, by model, by hour, with the sample size beside every figure and the method written down. We do not run your traffic, we do not take a cut of your bill, and we do not decide anything on your behalf — you read the distribution for your model and make the call yourself. The whole dataset is free at a 15-minute delay, published under CC BY 4.0, and the people who send us measurements read it live.

Batch price50%of synchronous, every token
Provider's promise24 hthe only number they give you
Half finish within30s openai · gpt-5.6-sol · n=1887 · updated every 10 minutes
Nine in ten within17 min openai · gpt-5.6-sol · n=1887 · updated every 10 minutes

See the full distribution behind that median — See the tail, every model on the data page.

How accurate are we? — we are the only ones who publish it

Our p90 held
92.6%
1,496 jobs scored out of sample · 30-day window · live tier

Measured, not promised — and checked out of sample. For every batch job we compute what we would have quoted using only measurements that finished before that job started, then compare it to what actually happened. A prediction never sees its own outcome, so the score is honest by construction. Everyone else in this category says “usually within a few hours.” We publish how often our own number was right — and you can check it in one call.

Every model we score, out of sample — the whole instrument, not just the headline.

Modelp50 heldp90 held
OpenAI · gpt-5.6-luna53.8%95.7%
Anthropic · claude-haiku-4-531.1%95.3%
OpenAI · gpt-5.6-sol54.3%94.7%
OpenAI · gpt-5-nano48.2%90.3%
Google · gemini-3.7-flash48.5%87.3%

This is the live tier — the coverage a paying caller’s real-time data scores at. Reproduce it yourself in one call: GET /v1/calibration, public to any key holder. The response carries n, the confidence interval and a not_measured block, so the figure is checkable end to end.

See per-model calibration — every model we score, out of sample, on the data page.

What the discount alone cannot tell you

What that is worth

What batching is worth on your spend

Measured, not assumed — and net: we subtract what you already batch and what the missed deadlines cost, so the number is one you can take into a meeting.

Not every job can move. Interactive features, anything a user waits on, stays synchronous. But the batchable half — nightly enrichment, evals, backfills, classification, summarisation, report generation — is usually the larger half, and it is paying double today.

Spend is what the whole workload would cost at synchronous prices — the baseline everything else is measured against, so that moving work to batch changes the bill instead of changing the question.

The two shares are sliders because they are your numbers, not ours. We have not measured your workload and will not pretend to. The 50% discount is the providers’ published rate; everything else above is arithmetic on what you typed.

A missed deadline costs more than it saves. A job that has to be re-run synchronously pays full price and has already paid for the batch attempt, so every one of them cancels out a job that made it. At a miss rate of 50% the whole thing is a wash — which is exactly the number this site exists to keep you away from.

Who this is for

●
Product ownersYour LLM bill is up and to the right, and the only levers anyone offers you cut scope, quality or headcount. Here is a different input: the measured latency of every tier your work could run on. Take the distribution into the meeting and decide with a number instead of an anecdote.
●
EngineersYou could move it to a cheaper tier — but you cannot tell in advance whether it will clear before your deadline, so you leave it on sync to be safe. One call gives you the measured distribution for that model and tier before you submit. What you do with it is your code’s decision, not ours — we supply the number, you write the fallback.
●
Platform teamsEvery team picks a tier by habit and folklore, and nobody owns the decision, so it is always wrong somewhere. Give them a shared source of truth instead: one measured feed every team reads, so the argument is about policy rather than about whose anecdote is newer.
●
Data & ML teamsNightly enrichment, backfills and classification are the textbook batch workload, and they run overnight anyway — yet they sit on synchronous pricing because moving them once felt risky. There is no deadline to miss at 3am; this is the cheapest half of your spend to reclaim.
●
Anyone with an eval suiteEvals are the perfect batch workload — large, deadline-free, embarrassingly parallel — and they almost always run synchronously out of habit, at full price, for no reason anyone can name. Move them once and every run after is half off.

“Why not just batch it myself?”

You can, and the easy part you should: submitting to a batch endpoint is a few lines. The hard part is not submitting — it is knowing whether it is safe to, and when. That judgement is the product, and it is the one part you cannot build from your own account.

1
The 50% is already yours — the confidence to capture it is not.The provider hands you the discount for free. What no pricing page hands you is which jobs will actually finish in time, so the deadline-bound half of your work stays on synchronous at full price. That half is exactly what we unlock — and it is the delta, not the discount, that decides whether batching pays.
2
You see your own account, after the fact. We see the queue filling before it reaches you.batchwatch is a leading indicator across accounts: measured queue times over the whole fleet, so a slowdown shows up in our data before it lands on your jobs. One account’s history cannot compute that — by definition.
3
Your own history is n = 1, and n = 1 hides the tail.A perfect personal record right up to the job that breaks it is the whole point of the nineteenth-job story below: one account never sees the slow one coming. The measured, cross-account distribution does.
4
Building it yourself means probing every provider and model, around the clock — and still seeing only your slice.The 50 lines of client are the easy part and we publish them open source. The measured, cross-account, continuous dataset behind the decision is the part that takes running probes every ten minutes across the fleet — and that is the product.

Read the send path before you install it, and self-batch every obvious job today — we would rather you did. Reach for us on the one that has a deadline, because that is the job where measurement, not a guess, is the difference between the discount and a negative bill.

Twenty shards is not one job, twenty times

A real workload is rarely one request. You shard an eval across twenty calls and you need all twenty in before the deadline. That is a different question, and the arithmetic is unforgiving: at 90% per shard, the odds that all twenty land are not 90%. They are 0.920 = 12%.

ModelOne shard on timeAll shards on time Batch now, sync the rest
Filled from the same measured curves as the table above. If you are reading this line, JavaScript is off — GET /v1/should-i-batch?…&n_shards=20 answers the same question without it.

What we cover

Providers

Coverage follows what people measure. Everything here is a real batch endpoint on a real provider — nothing is simulated. The status column is read from the running deployment, not written by hand.

ProviderEndpointDiscount Their promiseStatus
The status of each provider is read live from /v1/probe and /v1/coverage, so there is only one copy of it. If you are reading this line, JavaScript is off — GET /v1/probe answers the same question without it.

“Measured” means jobs have been submitted and timed — by our own prober, or by a contributor, and the row says which. “Configured” means the prober is running against it and the first result is on its way, and “accepted” means the route and the schema take that provider today — one measurement turns it into a row with numbers in it. Every column here is measured. An empty one means unmeasured, not zero — which is the reason to read coverage here rather than off a marketing page.

Any provider with a batch endpoint can be added — the schema is not OpenAI-shaped. What decides the order is where the measurements come from.

To be precise about which “batch” this is: we time the provider’s asynchronous batch API — the half-price, submit-and-wait tier with a published completion window. That is a different thing from in-server micro-batching (the sub-second request coalescing inside an inference server, as in vLLM’s continuous batching), which finishes in milliseconds and carries no queue to measure. The wait we publish is the one on the async tier, which is the one no one else measures.

How fast is batch, really

Share of jobs finished, by elapsed time

This is the chart the provider does not publish.

The distribution is not flat — it is front-loaded with a long, thin tail. That shape is why “up to 24 hours” is technically true and practically useless. We measure the tail, so you can route around it (one job in this dataset took eight hours) — and tell you the moment you are standing in it.

Find your model, and see who is fastest right now

Three pages built from the same measurements as everything above.

The whole of what we hand you

GET /v1/should-i-batch
      ?model=gpt-5.6-sol
      &input_tokens=9720   ← you know these
      &max_wait=15m     ← your deadline
      &risk=p90         ← how safe


{
  "verdict": "run_batch",
  "decision_threshold": 0.90,           ← what risk=p90 asserts
  "meets_deadline_probability": 0.94,   ← P(you make it)
  "expected_lateness_s": 0,             ← E[(D - deadline)+]

  "your_limit_s": 900,
  "batch_wait_s": 372,      ← observed p90
  "planning_wait_s": 680,   ← upper bound

  "batch": { "p50_s": 210, "p90_s": 372, "n": 1180 },
  "sync":  { "p50_s":  41, "p90_s":   88, "n":  204 },

  "saving_usd": 0.0647,
  "cost_basis": { "estimate": true, … },

  "confidence": { "level": "medium",
                  "why": "…", "n": 1180 }
}

That is the entire idea

No deadline given? You get required_patience_s instead — the number you would have to accept. Compare it to your own.

1
You set the deadline, not usWe cannot know whether 39 minutes ruins your day or means nothing at all. Send max_wait and we answer against your limit — we never guess it.
2
verdictrun_batch · run_sync · no_deadline_given · insufficient_data
3
Never a bare numberOn thin data we do not go quiet — we widen. batch_wait_s is what was observed, planning_wait_s is the bound we would actually route on, and they differ only when the data is thin. insufficient_data is reserved for a model that has never been measured — it is a coverage answer, not a shrug.
4
Two lines in your codeAsk before you submit. Route accordingly. That is the whole integration.
5
The client fails openAn outage here must never stop your job, so the failure is handled in the client, not promised by the server: should_batch() returns your default when it cannot reach us. Each library has a test against a dead port and a hung socket.

Just the number

GET /v1/wait?model=gpt-5.6-sol

{
  "model": "gpt-5.6-sol",
  "provider": "openai",

  "p50_s": 372,
  "p90_s": 2460,
  "oldest_running_s": 28800,  ← live only

  "based_on": { "n": 14, "window": "1h" },
  "coverage_30d": { "n": 1180, … },
  "vs_normal": 2.31,
  "confidence": { "level": "medium", … },
  "freshness": { "last_measurement_age": "4 min", … },

  "live": true,
  "delayed_by_s": 0,

  "measurement": "observed completions",
  "not": "a prediction for your job"
}

No key needed. Without one you get the same fields with "live": false and a delay in delayed_by_s — with one exception: oldest_running_s is a reading of the queue right now, so it is only in the answer when the answer is live. A delayed caller gets everything else.

One parameter, one answer

The simplest thing we can offer. Start here.

Not every caller wants the decision logic. Sometimes you just need to know whether the queue is 40 seconds deep or 8 hours deep, and you will decide what that means yourself.

Read the last two fields. This is what jobs finishing right now actually took. It is not a forecast for the job you are about to submit — nobody can give you that honestly, and we have the data to prove it.

oldest_running_s is the one people miss. Completed jobs describe the past. The oldest job still waiting is the earliest sign a queue has stalled.

Try it without signing up. Anyone gets this endpoint at 15 minutes delayed, starting with 20 free calls. That is enough to see the data is real and to link to it. Contribute your own measurements and the delay goes away entirely — the full quota, at live, free.

Pick the emptiest hour

Submit when the queue is shortest

The queue is not the same depth all day, and we measure the shape of it hour by hour. Line an overnight or next-morning batch up with the quietest hour and it clears faster — same work, same price, sooner done.

See it by hour of day on the data page, beside every other per-model chart.

What we receive

Timing and token counts. Nothing else.

The client library builds the submission from a fixed allowlist — what the job was and when it ran, nothing else — and drops everything outside it on the way out. It never touches the payload. Nothing larger is ever sent. A measurement is two small calls: the box on the left opens it, and a second, smaller one closes it with the id and the end time — that is how the duration gets measured against our clock instead of yours. There is no third request, and neither of them carries your payload.

Both calls are required. An integration that only opens measurements and never closes them contributes nothing — the percentiles read finished jobs, so an open row is invisible to them. The client library does both for you; if you are writing your own, see /docs/ingest.

{
  "mode": "batch",
  "provider": "openai",
  "model": "gpt-5.6-sol",
  "requests": 1,
  "input_tokens": 9720,
  "output_tokens": 4519,
  "started_at": "…16:24:27Z",
  "ended_at":   "…00:24:12Z",
  "status": "completed"
}
  • ×Your prompts — not sent, not hashed, not sampled
  • ×Your completions or any model output
  • ×System prompts, tool definitions, function schemas
  • ×Your API keys — the client never reads your provider credentials
  • ×End-user identifiers, request IDs, or anything that maps back to a person
  • ✓Country code, derived from your IP. We store the code, never the address

The client is open source and short enough to read in one sitting. Read the send path before you install it — that is the point of publishing it: clients/ on GitHub.

Getting it back out, and taking it away

Three calls, all authenticated with the key itself. There is no account to close and nobody to write to. It also names the one thing we cannot do: a contribution you sent without a token cannot be deleted — with no key, nothing ties that measurement to you to find and remove it.

GET    /v1/calls/mine        everything you sent, in full
DELETE /v1/calls/mine        exclude it, immediately
POST   /v1/keys/current/rotate   new token, old one revoked
DELETE /v1/keys/current      revoke the key

And this website

That was the client library. This is the site you are reading, which is a separate question with a separate answer — and one we would rather state than leave you to infer from a cookie banner.

What we measure, in full: which page you are on, clicks on Get an API key, whether the spend calculator and the model selector were used, whether you scrolled as far as the API section, and whether a key was created. That list is the whole of it. If we add anything, the banner asks again.

Declining costs you nothing. Every number, chart and endpoint on this site behaves identically either way — there is no reduced version. Change your mind whenever you like: . Withdrawing deletes the cookies again.

If your browser sends Global Privacy Control, we take that as a no and never ask.

Getting in

You already have this data. You just throw it away.

Every job you run on any tier is a measurement: when you submitted, when it landed, which model and tier, how many tokens. Nobody publishes that — not the providers, not any monitoring service. It is expensive to sample from the outside, because one data point costs one real call.

So the deal is simple. Send your measurements, get everyone else's. There is no other way this dataset can exist.

— models measured live — Contribute your measurements

The prober measures every ten minutes, live and continuously.

One step: get a key

No email, no password, no confirmation step, nothing to install. A key does not by itself carry any weight in the statistics — that is earned by measuring — so there is nothing to verify. You get it back with the one command that sends your first measurement, already filled in.

We only ever receive timing and token counts — never your prompts, completions or any content. The client sends a fixed allowlist of fields and nothing else, by construction: what we receive, the send path on GitHub, or where the data lives.

The label is your own note — “prod-pipeline”, “my laptop”. Nobody else sees it; it is how you tell your own keys apart when one has to be revoked.

You see the token once. We store a hash of it, not the token, so there is no “show it again” and no support request that can recover it. Copy it before you close the tab. Losing one costs nothing — ask for another — but the one you lost stays lost.

There is a small daily cap on new keys per IP address. It is there to stop a runaway loop, not to ration you: a key carries no weight in the statistics by itself, so hoarding them buys nothing. If you hit it, the page says how many you made and what the cap is — that is a limit doing its job, not the site breaking. One key per service is plenty; they are not per-machine.

You can undo all of it. If the key leaks, rotate it (POST /v1/keys/current/rotate) — you get a new token, the old one stops working, and everything you had measured moves across, so you do not start over. If you change your mind about the data, DELETE /v1/calls/mine takes your measurements out of the published figures immediately, including the public per-model pages. All three, in full.

Then wire it into your pipeline

One curl proves the account works. This is what makes it keep happening without anyone remembering to do it.

Python tests green
if bw.should_batch(
        "gpt-5.6-sol", max_wait="15m"):
    job = client.batches.create(…)

with bw.track("gpt-5.6-sol",
              input_tokens=9720) as t:
    r = wait_for(job)
    t.done(output_tokens=…)
TypeScript tests green
if (await bw.shouldBatch(
      "gpt-5.6-sol", { maxWait: "15m" })) {
  const job = await openai.batches.create(…);
}

const t = bw.track("gpt-5.6-sol",
                   { inputTokens: 9720 });
t.done({ outputTokens: … });
.NET tests green
if (await bw.ShouldBatchAsync(
      "gpt-5.6-sol", …))
{
    var job = await openai
        .Batches.CreateAsync(…);
}

using var t = bw.Track("gpt-5.6-sol",
                       inputTokens: 9720);
t.Done(outputTokens: …);

Nine languages, and not one external dependency. Python, TypeScript, Go, .NET, Java, PHP, Ruby, Rust and C++ — standard library only, every one of them. HTTP is hand-rolled where it had to be (std::net in Rust, java.net.http, PHP stream contexts, a raw BSD socket in C++) rather than pulling a client and a JSON package into your lock file, and each ships its own suite of 25–28 tests run against real loopback servers rather than stubs.

Every one of them is two lines to adopt, MIT, and short enough to read before you install it. Pick yours — all nine, with the install command for each. Two you can run right now:

pip install "git+https://github.com/batchwatch/client#subdirectory=python"
go get github.com/batchwatch/client/go@latest

All nine do the same three things: fail open, so a batchwatch outage never stops your job; never send content, only the fields listed under what we receive; and spool to disk, so a measurement taken while the network was bad is replayed later instead of lost. Each of those three has a test with a positive control, so a client that quietly sent nothing at all would fail them rather than pass.

Read the source on GitHub → All nine client libraries

Your coding agent can ask batchwatch directly — an MCP server, ready to install. The agent that writes the code is the one choosing batch or synchronous. Now it asks instead of guessing. Claude, Cursor, Cline or anything else that speaks MCP gets six tools — should_i_batch, estimate_batchtime, conditions, distribution, coverage, wait — each a thin proxy onto the route of the same name. It makes no decision and invents no numbers: every answer is the API's own JSON with its n, its confidence and its why intact, and a route that cannot answer returns an error rather than a figure.

It runs on the same tiers as the rest of the API. The server is a faithful proxy, so it carries the agent's own key and earns the agent's own tier — keyless calls read the public, 15-minute data, and a key buys freshness and live figures here exactly as it does everywhere else. One access model, one place to reason about it, and the live edge is worth paying for whether a human or an agent is asking. It installs with one command and no build step, and it is on its way to the MCP registries agents reach for by default — the Anthropic MCP registry, Smithery and mcp.so — so it becomes the tool an agent picks up on its own.

What it costs: nothing

Every queue time on this page is free to read, and always will be.

No account, no email, no card: the whole dashboard, every model page, the comparison tables, the outage feed and the CSV export. The read API answers the same measurements at 15 minutes delayed, and a new key starts with 20 free calls. These are measurements, not estimates — every percentile here came from a real call that was actually submitted to the provider’s own API on the tier it names, and actually finished. Nobody else publishes these numbers, at any price.

Contribute your own timings and the delay disappears. Send us the batch jobs you are running anyway and you read the queue as it is, not as it was 15 minutes ago. That is the whole trade, and it is the right one: the people whose measurements build the dataset should not be reading a delayed copy of it. How to contribute is below, and it is three steps.

No key

You are reading the site and want to see whether the answer is worth anything.

Free No account, no email, no card.
  • ✓The whole dashboard, every chart on this page
  • ✓20 free calls to the read API, then /v1/wait at 15 minutes delayed
  • ✓Submitting measurements is open to everyone — no key needed to give
  • —No weekly quota to draw on once the 20 are gone

Key

You want the trial to follow you between machines, and your own data back out again.

Free Self-service. No email, no password, no confirmation step.
  • ✓Everything above
  • ✓The 20 free calls count on the key, not on your IP, so they survive a new network
  • ✓Import finished jobs in bulk with POST /v1/calls/complete
  • ✓Export everything you sent, and delete it, from the key itself

Contributor

You run batch jobs anyway, and the timings are already sitting in your logs.

Free Paid in measurements: 5 in the last 7 days.
  • ✓Everything above
  • ✓live figures, no delay — the queue as it is right now, which is what you can route production traffic on
  • ✓5,000 calls a week, on a rolling seven days — not a calendar week, so usage frees up gradually
  • ✓Measure on 3 separate days and your figures start counting towards the published percentiles
Send your first measurement

5 in 7 days, not one a day. Twice a week is enough.

Verified

You are contributing and you want more room, without being asked for anything you would not want to give.

Free The 5 measurements plus a confirmed email — both, not either.
  • ✓Everything above
  • ✓10,000 calls a week, double the contributor quota, same rolling window
  • ✓We can reach you when a queue you depend on breaks, or when your key is about to stop working
  • ✓The email is an unlock on top of contributing, not a replacement — the better tier applies while you keep measuring
Get a key, then verify an email

Verify from the key with POST /v1/verify/start, or from /docs/keys. The address is the only personal data we ask for anywhere.

What decides which of these you are on: a key is something you ask for, and contribution is measured in your data. You cannot set your own tier — not as a hurdle, but because a dataset where accounts can promote themselves is worth nothing to read, and being worth reading is the entire product.

The dashboard never moves behind a plan. Free forever, for everyone, without an account. Charging for the page would mean charging the people whose measurements drew it.

How to contribute, exactly

Three steps, and the second one is a single request. If you already run batch jobs, the numbers are in your logs.

  1. Get a key. POST /v1/keys — one command, no email, no password, no confirmation step. The full recipe is at /docs/keys.
  2. Send finished jobs. POST /v1/calls/complete takes a batch job you have already run — model, token counts, submitted-at and finished-at — in a single request. Jobs still running can be timed live instead with POST /v1/calls and PATCH /v1/calls/{id}. Both are in /docs/ingest, and the client libraries do it in two lines in nine languages.
  3. Keep it current. 5 measurements inside 7 days holds live access — roughly twice a week, not once a day. We are asking you to stay current, not to change how you work.

Nothing you send identifies a prompt. The client builds the body from a fixed allowlist — provider, model, mode, endpoint, request count, token counts, timestamps — and there is no field for text at all. Everything you send comes back out again with GET /v1/calls/mine, and DELETE /v1/calls/mine removes it, both from the key itself.

Volume alone cannot move a published number. A figure counts towards the percentiles once it is backed by measurements on 3 separate days, so no single account can buy influence by flooding us — which is exactly why the published numbers are worth reading.

The published aggregate is licensed CC BY. Reuse it, including commercially, for the price of a credit and a link back — see the licence.

What each tier gets

Access is earned in recent measurements, not in signups.

YouDashboardSubmit APILive API
Anyone, no accountfull — 20 free calls, then /v1/wait, 15 min old
Account, no data yetfull yes 20 free calls, then /v1/wait, 15 min old
Contributing — 5+ measurements in the last 7 days full yes live, no delay, 5,000/week
Contributing & email verified — free full yes live, no delay, 10,000/week
Stopped contributingfull yes back to 15 min — any unspent free calls are still there

20 free calls, no signup, no card. Point curl at it and see what it says about your model before you write a line of integration code. Asking you to instrument your pipeline first, on the promise that the answer might be worth it, is the wrong way round.

Calls that error don't count. They are for finding out whether the answer is useful, not for finding out how the query string is spelled. Every response carries "trial": {"calls_left": 17}, so the wall is never a surprise.

5 in 7 days, not one a day. If you run batches twice a week you are still a contributor — we are asking you to stay current, not to change how you work.

Alerting is live, and it is open to everyone. Subscribe a webhook or a Slack hook to POST /v1/subscriptions and we tell you when a queue breaks — not when a job is slow, but when a model deviates from its own baseline, confirmed by three independent contributors and held for twenty minutes. Webhooks are HMAC-signed, and the secret is stored encrypted, never in the clear. /docs/alerts. Prefer to poll? /v1/outages and the Atom feed at /v1/outages.atom carry the same events.

These terms will change, and we will tell you first. A product this young that promised its terms were final would be making that promise at the moment it knows least, and we would rather commit to something we can keep: 30 days’ notice before any change that reduces what contributors get, and the open dataset stays open.

Your data, our aggregate — in plain words

This is the summary. The full terms say the same thing at greater length, and the data processing agreement is the version to hand to whoever owns the infrastructure you are measuring. Nothing here is hidden in them.

You keep
  • Your raw submissions. Export them any time, in full.
  • The right to delete everything you sent with a token — DELETE /v1/calls/mine, gone from the aggregate on the next rebuild. A keyless contribution has no token tying it to you, so it cannot be deleted.
  • The right to stop, with no effect on data you already got.
We take
  • A terminable licence to use your measurements in aggregate — including commercially. You can end it going forward at any time; aggregates already published stand, because a released statistic cannot be recalled.
  • The right to publish aggregate statistics openly, and to sell access to them.
  • Never the right to publish, sell or expose an individual submission, or anything traceable to you.

We are stating this up front rather than burying it, because you are handing over data from infrastructure you may not personally own, and you should be able to justify that to whoever does.

The evidence behind the answer

Queue conditions right now

Each model against its own 30-day baseline.

Loading measurements…

By weekday

Wait against each model's own normal, by UTC weekday.

If a backfill queued Friday evening behaves differently from the same job on Tuesday morning, it shows up here. On All models each bar is the wait against that model’s own normal, pooled — so a day looks slow only if it is slow across the board, never because it happened to sample a slower model. Bars appear per weekday once that day has enough measurements — a gap is a gap, not a zero.

Wait against job size

Wait against each model's own normal, by input-token band.

The question this answers: does splitting a big job into small ones actually make it start sooner? If the bands are level, it does not, and the effort is wasted. On All models each band is normalised to each model’s own normal, so the shape is the queue’s and not our probe mix. Jobs submitted without a token count are left out rather than counted as zero.

Queue outages

The record no provider publishes.

StartLengthModel× baselineCurrent status
Loading…

A queue outage is a model running far above its own 30-day baseline (confirmed by a majority of established contributors), or a run of our own probes all failing on it. Each row shows whether it is still ongoing or has since resolved, and how far above baseline it ran.

Built for the nineteenth job

The measurement that decided what this product returns.

19 jobs · same account
   · same model · same size

18 jobs   median   3 min
job 19             480 min

A perfect personal history said 3 minutes. The nineteenth job took 480 — off by a factor of 13. measured

That job is the entire reason to measure. A timestamp — "your job finishes at 14:32" — is right until the one time it matters, and by then it has already cost you the deadline. So you get the thing that survives job nineteen: a decision, and the distribution it was made from. What nine jobs in ten do, what the slowest one did, how much of that your deadline can absorb.

Route on the distribution and the tail stops being an incident. Job nineteen is just a slow job that your fallback already covered — which is worth considerably more than a precise number that was wrong.

Coverage

How much data sits behind each answer. Every row is counted, not estimated. We draw what we have and say what it is worth: a curve is drawn as soon as there are points to connect, and a thin one is served with a thinness flag rather than a smooth line’s authority. The one hard threshold left is the full distribution, and it is there so that an hourly profile over a handful of jobs cannot be read as one account’s working day — that route states its own numbers when it declines. Contributor counts are shown so you can see when a number is one account’s experience rather than a crowd.

Provider · modelJobs, 30dContributors MedianStatus
Loading…
🔒 The full distribution (/v1/distribution) is the one route behind the gate, for the privacy reason above. Send your first job to open it. The hourly clock and the outage history are not gated at all — /v1/hours and /v1/outages answer without a key, and this page reads them that way.

Sending a job shares only timing and token counts — never your prompts, completions or any content. The client builds a fixed allowlist of fields and nothing else, by construction. See the field-by-field list, the send path on GitHub, or where the data lives.

API

The API answers from Cloudflare's edge — and we do not ask you to take that on faith. We measure batchwatch.dev from seven countries and publish every run, the slow ones included: see the response times.

Submit — open to everyone

No account needed to contribute. That's deliberate.

POST /v1/calls
{ "mode": "batch",
  "provider": "openai",
  "model": "gpt-5.6-sol",
  "requests": 1,
  "input_tokens": 9720,
  "started_at": "2026-08-24T16:24:27Z" }

→ 201 { "id": "c_7f3a…" }

PATCH /v1/calls/c_7f3a…
{ "status": "completed",
  "output_tokens": 4519,
  "ended_at": "2026-08-25T00:24:12Z" }

→ 200 { "duration_s": 28785,
        "percentile": 99.4 }

The response tells you where your job landed in the distribution immediately. No key, no wait, no gate: your own percentile comes back on the very first call you make.

Read — 20 free calls, then contributors

Sync calls count too: mode: "sync".

open, no key at all:
  GET /v1/wait          how long is the queue?
  GET /v1/curve         the distribution, plotted
  GET /v1/hours         the 24-hour clock
  GET /v1/patterns      weekday and job size
  GET /v1/coverage      what we can answer
  GET /v1/outages       incidents (+ .atom)
  GET /v1/probe         our own measuring, and its budget
  GET /v1/plans         this deployment's limits

free calls, then contributors:
  GET /v1/should-i-batch    should I use it at all?
      &n_shards=20          …for a whole N-shard run
  GET /v1/estimate-batchtime  jobs the size of mine
  GET /v1/conditions        is it me or them?
  GET /v1/distribution      p50 … p99, filterable

your own key, your own rows:
  POST   /v1/subscriptions  tell me when a queue breaks
  GET    /v1/subscriptions  the ones you own
  DELETE /v1/subscriptions/1
  GET    /v1/savings        what your batching actually saved

Every answer that puts a number on the queue carries
its own uncertainty, never the number alone:
  "confidence": { "level": "high", "why": …, "n": … }
  "freshness": { "last_measurement_age": "4 min" }

  "no_coverage": false   on wait, curve, should-i-batch
                         and estimate-batchtime
  "trial": { "calls_left": 17 }   on the gated four,
                         while you are still trialling

Sync measurements are what make should-i-batch possible. Without both sides there is no trade-off to compute.

Alerts, when a queue actually breaks. Not a slow job: a sustained deviation from a model’s own baseline, seen by at least three independent contributors, held for twenty minutes. Webhook (HMAC-signed) or Slack — /docs/alerts. A failing subscription is never switched off silently; a channel that disables itself cannot be told from an alert that never fired.

The full reference is a real page, not this box: /docs — every route, every error code, and what the numbers mean.