batchwatch is a supplier of information. OpenAI, Anthropic and Google sell the same model output at more than one price — the asynchronous batch tier, a synchronous flex tier where they offer one, and the full-price standard call. What separates those tiers is latency and refusal, and not one of the three providers publishes either. “Within 24 hours” is a ceiling: the widest the wait can ever be, and silent on what it is today. So we run real jobs against all three, around the clock, and publish the whole distribution — median, p90, p95, each with the sample size behind it. That is the entire offer, and the numbers exist nowhere else. They are measured rather than estimated, they belong to no vendor, and they are free.
The median and the p90 rank models differently. Every public
benchmark ranks on the median, and the median is not the figure a
deadline cares about. Both are measured, both are in the table below,
and both are in /v1/coverage
with the sample size behind each.
Named precisely: the half-price tiers are the provider’s
asynchronous batch API — the /v1/batches
endpoints, where you submit now and the results come back within a
published deadline — and, where a provider offers one, a
synchronous flex tier at the same discount. That is a different
product from in-server micro-batching such as vLLM’s, which
coalesces requests in milliseconds and has no queue to measure.
Every percentile we publish here is measured on the async batch queues at
all three providers — the wait that comes with the discount, and
the one number nobody else publishes at all.
We only ever receive timing and token counts — never your prompts, completions or any content. The client sends a fixed allowlist of fields and nothing else, by construction. See the send path on GitHub, the field-by-field list, or where the data lives.
One field, no email, no card — and the key comes back with the single command that sends your first measurement.
Every major provider will sell you the same tokens at half price if you can wait — 50% off, every input and output token, on the asynchronous batch tier and on the synchronous flex tier where they offer one. The discount is real and it is theirs to give. What none of them will tell you is the thing that decides whether you can use it: how long that tier actually takes today, and how often it refuses the work outright.
That is where it stays, because the only promise the provider gives you is “within 24 hours”. No engineer can build a feature on that and no product owner can plan around it, so the safe move is to pay double and keep the latency you understand. The recoverable half sits there, paying full price, and no pricing page will ever surface it.
The 50% and the 24 hours are the providers’ own published numbers. That the discount goes mostly untaken is our reading of the market, not a measurement — and saying which is which is the habit this whole site is built on.
A ceiling is not an estimate. The queue has good days and bad days, and the 24-hour number is the same on both. That is the whole gap: not that batch is slow, but that nobody tells you which day you are having.
That is the gap we fill, and filling it is the whole of what we do. We measure the tiers and we publish the measurements. Median, p90, p95, by model, by hour, with the sample size beside every figure and the method written down. We do not run your traffic, we do not take a cut of your bill, and we do not decide anything on your behalf — you read the distribution for your model and make the call yourself. The whole dataset is free at a 15-minute delay, published under CC BY 4.0, and the people who send us measurements read it live.
See the full distribution behind that median — See the tail, every model on the data page.
Measured, not promised — and checked out of sample. For every batch job we compute what we would have quoted using only measurements that finished before that job started, then compare it to what actually happened. A prediction never sees its own outcome, so the score is honest by construction. Everyone else in this category says “usually within a few hours.” We publish how often our own number was right — and you can check it in one call.
Every model we score, out of sample — the whole instrument, not just the headline.
| Model | p50 held | p90 held |
|---|---|---|
| OpenAI · gpt-5.6-luna | 53.8% | 95.7% |
| Anthropic · claude-haiku-4-5 | 31.1% | 95.3% |
| OpenAI · gpt-5.6-sol | 54.3% | 94.7% |
| OpenAI · gpt-5-nano | 48.2% | 90.3% |
| Google · gemini-3.7-flash | 48.5% | 87.3% |
This is the live tier — the coverage a paying caller’s
real-time data scores at. Reproduce it yourself in one call:
GET /v1/calibration, public to any key holder. The response
carries n, the confidence interval and a not_measured
block, so the figure is checkable end to end.
See per-model calibration — every model we score, out of sample, on the data page.
Measured, not assumed — and net: we subtract what you already batch and what the missed deadlines cost, so the number is one you can take into a meeting.
Not every job can move. Interactive features, anything a user waits on, stays synchronous. But the batchable half — nightly enrichment, evals, backfills, classification, summarisation, report generation — is usually the larger half, and it is paying double today.
Spend is what the whole workload would cost at synchronous prices — the baseline everything else is measured against, so that moving work to batch changes the bill instead of changing the question.
The two shares are sliders because they are your numbers, not ours. We have not measured your workload and will not pretend to. The 50% discount is the providers’ published rate; everything else above is arithmetic on what you typed.
A missed deadline costs more than it saves. A job that has to be re-run synchronously pays full price and has already paid for the batch attempt, so every one of them cancels out a job that made it. At a miss rate of 50% the whole thing is a wash — which is exactly the number this site exists to keep you away from.
You can, and the easy part you should: submitting to a batch endpoint is a few lines. The hard part is not submitting — it is knowing whether it is safe to, and when. That judgement is the product, and it is the one part you cannot build from your own account.
Read the send path before you install it, and self-batch every obvious job today — we would rather you did. Reach for us on the one that has a deadline, because that is the job where measurement, not a guess, is the difference between the discount and a negative bill.
A real workload is rarely one request. You shard an eval across twenty calls and you need all twenty in before the deadline. That is a different question, and the arithmetic is unforgiving: at 90% per shard, the odds that all twenty land are not 90%. They are 0.920 = 12%.
| Model | One shard on time | All shards on time | Batch now, sync the rest |
|---|---|---|---|
Filled from the same measured curves
as the table above. If you are reading this line, JavaScript is off
— GET /v1/should-i-batch?…&n_shards=20
answers the same question without it. | |||
Coverage follows what people measure. Everything here is a real batch endpoint on a real provider — nothing is simulated. The status column is read from the running deployment, not written by hand.
| Provider | Endpoint | Discount | Their promise | Status |
|---|---|---|---|---|
The status of each provider is read
live from /v1/probe and /v1/coverage, so
there is only one copy of it. If you are reading this line,
JavaScript is off — GET /v1/probe answers the same
question without it. | ||||
“Measured” means jobs have been submitted and timed — by our own prober, or by a contributor, and the row says which. “Configured” means the prober is running against it and the first result is on its way, and “accepted” means the route and the schema take that provider today — one measurement turns it into a row with numbers in it. Every column here is measured. An empty one means unmeasured, not zero — which is the reason to read coverage here rather than off a marketing page.
Any provider with a batch endpoint can be added — the schema is not OpenAI-shaped. What decides the order is where the measurements come from.
To be precise about which “batch” this is: we time the provider’s asynchronous batch API — the half-price, submit-and-wait tier with a published completion window. That is a different thing from in-server micro-batching (the sub-second request coalescing inside an inference server, as in vLLM’s continuous batching), which finishes in milliseconds and carries no queue to measure. The wait we publish is the one on the async tier, which is the one no one else measures.
This is the chart the provider does not publish.
The distribution is not flat — it is front-loaded with a long, thin tail. That shape is why “up to 24 hours” is technically true and practically useless. We measure the tail, so you can route around it (one job in this dataset took eight hours) — and tell you the moment you are standing in it.
One page per model, with the distribution behind the number and what our confidence in it is.
Three pages built from the same measurements as everything above.
GET /v1/should-i-batch
?model=gpt-5.6-sol
&input_tokens=9720 ← you know these
&max_wait=15m ← your deadline
&risk=p90 ← how safe
{
"verdict": "run_batch",
"decision_threshold": 0.90, ← what risk=p90 asserts
"meets_deadline_probability": 0.94, ← P(you make it)
"expected_lateness_s": 0, ← E[(D - deadline)+]
"your_limit_s": 900,
"batch_wait_s": 372, ← observed p90
"planning_wait_s": 680, ← upper bound
"batch": { "p50_s": 210, "p90_s": 372, "n": 1180 },
"sync": { "p50_s": 41, "p90_s": 88, "n": 204 },
"saving_usd": 0.0647,
"cost_basis": { "estimate": true, … },
"confidence": { "level": "medium",
"why": "…", "n": 1180 }
}
No deadline given? You get
required_patience_s instead — the number you would
have to accept. Compare it to your own.
max_wait and we
answer against your limit — we never guess it.run_batch · run_sync ·
no_deadline_given ·
insufficient_databatch_wait_s is what was observed,
planning_wait_s is the bound we would actually route on,
and they differ only when the data is thin.
insufficient_data is reserved for a model that has never
been measured — it is a coverage answer, not a shrug.
should_batch() returns your
default when it cannot reach us. Each library has a test against a
dead port and a hung socket.GET /v1/wait?model=gpt-5.6-sol
{
"model": "gpt-5.6-sol",
"provider": "openai",
"p50_s": 372,
"p90_s": 2460,
"oldest_running_s": 28800, ← live only
"based_on": { "n": 14, "window": "1h" },
"coverage_30d": { "n": 1180, … },
"vs_normal": 2.31,
"confidence": { "level": "medium", … },
"freshness": { "last_measurement_age": "4 min", … },
"live": true,
"delayed_by_s": 0,
"measurement": "observed completions",
"not": "a prediction for your job"
}
No key needed. Without one you get the same fields with
"live": false and a delay in
delayed_by_s — with one exception:
oldest_running_s is a reading of the queue
right now, so it is only in the answer when the answer is
live. A delayed caller gets everything else.
The simplest thing we can offer. Start here.
Not every caller wants the decision logic. Sometimes you just need to know whether the queue is 40 seconds deep or 8 hours deep, and you will decide what that means yourself.
Read the last two fields. This is what jobs finishing right now actually took. It is not a forecast for the job you are about to submit — nobody can give you that honestly, and we have the data to prove it.
oldest_running_s is the one people miss. Completed jobs
describe the past. The oldest job still waiting is the earliest sign a
queue has stalled.
Try it without signing up. Anyone gets this endpoint at 15 minutes delayed, starting with 20 free calls. That is enough to see the data is real and to link to it. Contribute your own measurements and the delay goes away entirely — the full quota, at live, free.
The queue is not the same depth all day, and we measure the shape of it hour by hour. Line an overnight or next-morning batch up with the quietest hour and it clears faster — same work, same price, sooner done.
See it by hour of day on the data page, beside every other per-model chart.
The client library builds the submission from a fixed allowlist — what the job was and when it ran, nothing else — and drops everything outside it on the way out. It never touches the payload. Nothing larger is ever sent. A measurement is two small calls: the box on the left opens it, and a second, smaller one closes it with the id and the end time — that is how the duration gets measured against our clock instead of yours. There is no third request, and neither of them carries your payload.
Both calls are required. An integration that only opens measurements and never closes them contributes nothing — the percentiles read finished jobs, so an open row is invisible to them. The client library does both for you; if you are writing your own, see /docs/ingest.
{
"mode": "batch",
"provider": "openai",
"model": "gpt-5.6-sol",
"requests": 1,
"input_tokens": 9720,
"output_tokens": 4519,
"started_at": "…16:24:27Z",
"ended_at": "…00:24:12Z",
"status": "completed"
}
The client is open source and short enough to read in one sitting. Read the send path before you install it — that is the point of publishing it: clients/ on GitHub.
Three calls, all authenticated with the key itself. There is no account to close and nobody to write to. It also names the one thing we cannot do: a contribution you sent without a token cannot be deleted — with no key, nothing ties that measurement to you to find and remove it.
GET /v1/calls/mine everything you sent, in full DELETE /v1/calls/mine exclude it, immediately POST /v1/keys/current/rotate new token, old one revoked DELETE /v1/keys/current revoke the key
That was the client library. This is the site you are reading, which is a separate question with a separate answer — and one we would rather state than leave you to infer from a cookie banner.
What we measure, in full: which page you are on, clicks on Get an API key, whether the spend calculator and the model selector were used, whether you scrolled as far as the API section, and whether a key was created. That list is the whole of it. If we add anything, the banner asks again.
Declining costs you nothing. Every number, chart and endpoint on this site behaves identically either way — there is no reduced version. Change your mind whenever you like: . Withdrawing deletes the cookies again.
If your browser sends Global Privacy Control, we take that as a no and never ask.
Every job you run on any tier is a measurement: when you submitted, when it landed, which model and tier, how many tokens. Nobody publishes that — not the providers, not any monitoring service. It is expensive to sample from the outside, because one data point costs one real call.
So the deal is simple. Send your measurements, get everyone else's. There is no other way this dataset can exist.
The prober measures every ten minutes, live and continuously.
No email, no password, no confirmation step, nothing to install. A key does not by itself carry any weight in the statistics — that is earned by measuring — so there is nothing to verify. You get it back with the one command that sends your first measurement, already filled in.
We only ever receive timing and token counts — never your prompts, completions or any content. The client sends a fixed allowlist of fields and nothing else, by construction: what we receive, the send path on GitHub, or where the data lives.
The label is your own note — “prod-pipeline”, “my laptop”. Nobody else sees it; it is how you tell your own keys apart when one has to be revoked.
You see the token once. We store a hash of it, not the token, so there is no “show it again” and no support request that can recover it. Copy it before you close the tab. Losing one costs nothing — ask for another — but the one you lost stays lost.
There is a small daily cap on new keys per IP address. It is there to stop a runaway loop, not to ration you: a key carries no weight in the statistics by itself, so hoarding them buys nothing. If you hit it, the page says how many you made and what the cap is — that is a limit doing its job, not the site breaking. One key per service is plenty; they are not per-machine.
You can undo all of it. If the key leaks, rotate it
(POST /v1/keys/current/rotate) — you get a new
token, the old one stops working, and everything you had measured
moves across, so you do not start over. If you change your mind
about the data, DELETE /v1/calls/mine takes your
measurements out of the published figures immediately, including
the public per-model pages.
All three, in
full.
One curl proves the account works. This is what makes it keep happening without anyone remembering to do it.
if bw.should_batch(
"gpt-5.6-sol", max_wait="15m"):
job = client.batches.create(…)
with bw.track("gpt-5.6-sol",
input_tokens=9720) as t:
r = wait_for(job)
t.done(output_tokens=…)
if (await bw.shouldBatch(
"gpt-5.6-sol", { maxWait: "15m" })) {
const job = await openai.batches.create(…);
}
const t = bw.track("gpt-5.6-sol",
{ inputTokens: 9720 });
t.done({ outputTokens: … });
if (await bw.ShouldBatchAsync(
"gpt-5.6-sol", …))
{
var job = await openai
.Batches.CreateAsync(…);
}
using var t = bw.Track("gpt-5.6-sol",
inputTokens: 9720);
t.Done(outputTokens: …);
Nine languages, and not one external dependency. Python,
TypeScript, Go, .NET, Java, PHP, Ruby, Rust and C++ — standard
library only, every one of them. HTTP is hand-rolled where it had to
be (std::net in Rust, java.net.http, PHP
stream contexts, a raw BSD socket in C++) rather than pulling a client
and a JSON package into your lock file, and each ships its own suite
of 25–28 tests run against real loopback servers rather than
stubs.
Every one of them is two lines to adopt, MIT, and short enough to read before you install it. Pick yours — all nine, with the install command for each. Two you can run right now:
pip install "git+https://github.com/batchwatch/client#subdirectory=python" go get github.com/batchwatch/client/go@latest
All nine do the same three things: fail open, so a batchwatch outage never stops your job; never send content, only the fields listed under what we receive; and spool to disk, so a measurement taken while the network was bad is replayed later instead of lost. Each of those three has a test with a positive control, so a client that quietly sent nothing at all would fail them rather than pass.
Read the source on GitHub → All nine client libraries
Your coding agent can ask batchwatch directly — an MCP
server, ready to install. The agent that writes the code is the
one choosing batch or synchronous. Now it asks instead of guessing.
Claude, Cursor, Cline or anything else that speaks MCP gets six tools
— should_i_batch, estimate_batchtime,
conditions, distribution,
coverage, wait — each a thin proxy
onto the route of the same name. It makes no decision and invents no
numbers: every answer is the API's own JSON with its n,
its confidence and its why intact, and a
route that cannot answer returns an error rather than a figure.
It runs on the same tiers as the rest of the API. The server is a faithful proxy, so it carries the agent's own key and earns the agent's own tier — keyless calls read the public, 15-minute data, and a key buys freshness and live figures here exactly as it does everywhere else. One access model, one place to reason about it, and the live edge is worth paying for whether a human or an agent is asking. It installs with one command and no build step, and it is on its way to the MCP registries agents reach for by default — the Anthropic MCP registry, Smithery and mcp.so — so it becomes the tool an agent picks up on its own.
Every queue time on this page is free to read, and always will be.
No account, no email, no card: the whole dashboard, every model page, the comparison tables, the outage feed and the CSV export. The read API answers the same measurements at 15 minutes delayed, and a new key starts with 20 free calls. These are measurements, not estimates — every percentile here came from a real call that was actually submitted to the provider’s own API on the tier it names, and actually finished. Nobody else publishes these numbers, at any price.
Contribute your own timings and the delay disappears. Send us the batch jobs you are running anyway and you read the queue as it is, not as it was 15 minutes ago. That is the whole trade, and it is the right one: the people whose measurements build the dataset should not be reading a delayed copy of it. How to contribute is below, and it is three steps.
You are reading the site and want to see whether the answer is worth anything.
Free No account, no email, no card./v1/wait at 15 minutes
delayedYou want the trial to follow you between machines, and your own data back out again.
Free Self-service. No email, no password, no confirmation step.POST /v1/calls/completeYou run batch jobs anyway, and the timings are already sitting in your logs.
Free Paid in measurements: 5 in the last 7 days.5 in 7 days, not one a day. Twice a week is enough.
You are contributing and you want more room, without being asked for anything you would not want to give.
Free The 5 measurements plus a confirmed email — both, not either.Verify from the key with
POST /v1/verify/start, or from
/docs/keys. The address is the only
personal data we ask for anywhere.
What decides which of these you are on: a key is something you ask for, and contribution is measured in your data. You cannot set your own tier — not as a hurdle, but because a dataset where accounts can promote themselves is worth nothing to read, and being worth reading is the entire product.
The dashboard never moves behind a plan. Free forever, for everyone, without an account. Charging for the page would mean charging the people whose measurements drew it.
Three steps, and the second one is a single request. If you already run batch jobs, the numbers are in your logs.
POST /v1/keys — one command, no
email, no password, no confirmation step. The full recipe is at
/docs/keys.POST /v1/calls/complete takes
a batch job you have already run — model, token counts,
submitted-at and finished-at — in a single request. Jobs still
running can be timed live instead with POST /v1/calls and
PATCH /v1/calls/{id}. Both are in
/docs/ingest,
and the client
libraries do it in two lines in nine languages.
Nothing you send identifies a prompt. The client builds the body
from a fixed allowlist — provider, model, mode, endpoint, request
count, token counts, timestamps — and there is no field
for text at all. Everything you send comes back out again with
GET /v1/calls/mine, and
DELETE /v1/calls/mine removes it, both from the key itself.
Volume alone cannot move a published number. A figure counts towards the percentiles once it is backed by measurements on 3 separate days, so no single account can buy influence by flooding us — which is exactly why the published numbers are worth reading.
The published aggregate is licensed CC BY. Reuse it, including commercially, for the price of a credit and a link back — see the licence.
Access is earned in recent measurements, not in signups.
| You | Dashboard | Submit API | Live API |
|---|---|---|---|
| Anyone, no account | full | — | 20 free calls, then /v1/wait, 15 min old |
| Account, no data yet | full | yes | 20 free calls, then /v1/wait, 15 min old |
| Contributing — 5+ measurements in the last 7 days | full | yes | live, no delay, 5,000/week |
| Contributing & email verified — free | full | yes | live, no delay, 10,000/week |
| Stopped contributing | full | yes | back to 15 min — any unspent free calls are still there |
20 free calls, no signup, no card. Point curl at it and see what it says about your model before you write a line of integration code. Asking you to instrument your pipeline first, on the promise that the answer might be worth it, is the wrong way round.
Calls that error don't count. They are for finding out whether the
answer is useful, not for finding out how the query string is spelled.
Every response carries "trial": {"calls_left": 17}, so the
wall is never a surprise.
5 in 7 days, not one a day. If you run batches twice a week you are still a contributor — we are asking you to stay current, not to change how you work.
Alerting is live, and it is open to everyone. Subscribe a webhook
or a Slack hook to POST /v1/subscriptions and we tell you
when a queue breaks — not when a job is slow, but when a model
deviates from its own baseline, confirmed by three independent
contributors and held for twenty minutes. Webhooks are HMAC-signed, and
the secret is stored encrypted, never in the clear.
/docs/alerts.
Prefer to poll? /v1/outages and the Atom feed at
/v1/outages.atom carry the same events.
These terms will change, and we will tell you first. A product this young that promised its terms were final would be making that promise at the moment it knows least, and we would rather commit to something we can keep: 30 days’ notice before any change that reduces what contributors get, and the open dataset stays open.
This is the summary. The full terms say the same thing at greater length, and the data processing agreement is the version to hand to whoever owns the infrastructure you are measuring. Nothing here is hidden in them.
DELETE /v1/calls/mine, gone from the aggregate on the
next rebuild. A keyless contribution has no token tying it to you, so
it cannot be deleted.We are stating this up front rather than burying it, because you are handing over data from infrastructure you may not personally own, and you should be able to justify that to whoever does.
Each model against its own 30-day baseline.
Loading measurements…
Wait against each model's own normal, by UTC weekday.
If a backfill queued Friday evening behaves differently from the same job on Tuesday morning, it shows up here. On All models each bar is the wait against that model’s own normal, pooled — so a day looks slow only if it is slow across the board, never because it happened to sample a slower model. Bars appear per weekday once that day has enough measurements — a gap is a gap, not a zero.
Wait against each model's own normal, by input-token band.
The question this answers: does splitting a big job into small ones actually make it start sooner? If the bands are level, it does not, and the effort is wasted. On All models each band is normalised to each model’s own normal, so the shape is the queue’s and not our probe mix. Jobs submitted without a token count are left out rather than counted as zero.
The record no provider publishes.
| Start | Length | Model | × baseline | Current status |
|---|---|---|---|---|
| Loading… | ||||
A queue outage is a model running far above its own 30-day baseline (confirmed by a majority of established contributors), or a run of our own probes all failing on it. Each row shows whether it is still ongoing or has since resolved, and how far above baseline it ran.
The measurement that decided what this product returns.
19 jobs · same account · same model · same size 18 jobs median 3 min job 19 480 min
A perfect personal history said 3 minutes. The nineteenth job took 480 — off by a factor of 13. measured
That job is the entire reason to measure. A timestamp — "your job finishes at 14:32" — is right until the one time it matters, and by then it has already cost you the deadline. So you get the thing that survives job nineteen: a decision, and the distribution it was made from. What nine jobs in ten do, what the slowest one did, how much of that your deadline can absorb.
Route on the distribution and the tail stops being an incident. Job nineteen is just a slow job that your fallback already covered — which is worth considerably more than a precise number that was wrong.
How much data sits behind each answer. Every row is counted, not estimated. We draw what we have and say what it is worth: a curve is drawn as soon as there are points to connect, and a thin one is served with a thinness flag rather than a smooth line’s authority. The one hard threshold left is the full distribution, and it is there so that an hourly profile over a handful of jobs cannot be read as one account’s working day — that route states its own numbers when it declines. Contributor counts are shown so you can see when a number is one account’s experience rather than a crowd.
| Provider · model | Jobs, 30d | Contributors | Median | Status |
|---|---|---|---|---|
| Loading… | ||||
/v1/distribution) is the one route behind the gate, for
the privacy reason above.
Send your first job to
open it. The hourly clock and the outage history are not gated at all
— /v1/hours and /v1/outages answer
without a key, and this page reads them that way.
Sending a job shares only timing and token counts — never your prompts, completions or any content. The client builds a fixed allowlist of fields and nothing else, by construction. See the field-by-field list, the send path on GitHub, or where the data lives.
The API answers from Cloudflare's edge — and we do not ask you to take that on faith. We measure batchwatch.dev from seven countries and publish every run, the slow ones included: see the response times.
No account needed to contribute. That's deliberate.
POST /v1/calls
{ "mode": "batch",
"provider": "openai",
"model": "gpt-5.6-sol",
"requests": 1,
"input_tokens": 9720,
"started_at": "2026-08-24T16:24:27Z" }
→ 201 { "id": "c_7f3a…" }
PATCH /v1/calls/c_7f3a…
{ "status": "completed",
"output_tokens": 4519,
"ended_at": "2026-08-25T00:24:12Z" }
→ 200 { "duration_s": 28785,
"percentile": 99.4 }
The response tells you where your job landed in the distribution immediately. No key, no wait, no gate: your own percentile comes back on the very first call you make.
Sync calls count too: mode: "sync".
open, no key at all:
GET /v1/wait how long is the queue?
GET /v1/curve the distribution, plotted
GET /v1/hours the 24-hour clock
GET /v1/patterns weekday and job size
GET /v1/coverage what we can answer
GET /v1/outages incidents (+ .atom)
GET /v1/probe our own measuring, and its budget
GET /v1/plans this deployment's limits
free calls, then contributors:
GET /v1/should-i-batch should I use it at all?
&n_shards=20 …for a whole N-shard run
GET /v1/estimate-batchtime jobs the size of mine
GET /v1/conditions is it me or them?
GET /v1/distribution p50 … p99, filterable
your own key, your own rows:
POST /v1/subscriptions tell me when a queue breaks
GET /v1/subscriptions the ones you own
DELETE /v1/subscriptions/1
GET /v1/savings what your batching actually saved
Every answer that puts a number on the queue carries
its own uncertainty, never the number alone:
"confidence": { "level": "high", "why": …, "n": … }
"freshness": { "last_measurement_age": "4 min" }
"no_coverage": false on wait, curve, should-i-batch
and estimate-batchtime
"trial": { "calls_left": 17 } on the gated four,
while you are still trialling
Sync measurements are what make should-i-batch possible.
Without both sides there is no trade-off to compute.
Alerts, when a queue actually breaks. Not a slow job: a sustained deviation from a model’s own baseline, seen by at least three independent contributors, held for twenty minutes. Webhook (HMAC-signed) or Slack — /docs/alerts. A failing subscription is never switched off silently; a channel that disables itself cannot be told from an alert that never fired.
The full reference is a real page, not this box: /docs — every route, every error code, and what the numbers mean.