Two things are sold separately, and the split is the whole design:
- Contributing buys access to the data. You helped build it.
- Money buys everything that is not data — freshness, volume, live operation.
Source: src/tiers.js, src/trial.js, src/robust.js.
The five tiers
Contributing is what earns live data. A contributor reads every figure with no delay at all; confirming an email doubles the weekly volume on top of that. The email makes an otherwise anonymous contributor contactable — it is an opt-in unlock on top of contributing, never a replacement for it.
| Tier | Who | Delay | Weekly calls to gated routes | live |
|---|---|---|---|---|
anonymous | No key | 900s | 0 (uses the free trial) | false |
free | Has a key, but is not contributing right now | 900s | 0 (uses the free trial) | false |
contributor | 5+ completed measurements in the last 7 days | 0s | 5000 | true |
contributor_verified | A contributor who has confirmed an email | 0s | 10000 | true |
paid | Legacy, no longer offered | 0s | 1000000 | true |
Contributing earns live data outright: five completed measurements in the last seven days and the 15-minute delay is gone. A confirmed email doubles the weekly volume on top. Live access is not for sale at any price — it is earned by measuring, which is the only way this dataset can exist.
Verified against production 2026-08-25: an unauthenticated call to /v1/wait returned "live": false, "delayed_by_s": 900.
The numbers are configuration, not policy that has been decided. They can be overridden per deployment through environment variables:
| Variable | Affects |
|---|---|
PUBLIC_DELAY_S | anonymous and free delay |
CONTRIB_DELAY_S | contributor delay |
CONTRIB_WEEKLY_CALLS | contributor weekly calls |
CONTRIB_VERIFIED_DELAY_S | contributor_verified delay |
CONTRIB_VERIFIED_WEEKLY_CALLS | contributor_verified weekly calls |
PAID_WEEKLY_CALLS | paid weekly calls |
EMAIL_VERIFY_TTL_S | How long a verification link is valid |
TRIAL_CALLS | Free trial size (0 disables the trial) |
MIN_N, MIN_KEYS | /v1/distribution and /v1/conditions thresholds |
KEYS_PER_IP_PER_DAY | Key creation limit |
wrangler.toml currently sets PUBLIC_DELAY_S=900, MIN_N=20, MIN_KEYS=3, TRIAL_CALLS=20, PROBE_INTERVAL_S=600. The others fall back to the defaults in src/tiers.js.
How the tier is decided
tierOf(keyRow, contrib, emailVerified):
- No key at all →
anonymous. - The key's stored
tiercolumn ispaid→paid. The tier is retired: batchwatch stopped selling on 2026-09-02 and nothing grants it any more. It stays defined, and first in the order, so a key that already carries it keeps working exactly as before rather than losing access because the business model changed. It is no longer better than contributing — a contributor gets the same live data; only the volume headroom differs. - Not contributing →
free. - Contributing and email confirmed →
contributor_verified. - Contributing, no confirmed email →
contributor.
Only paid and the confirmed email are stored on the account (paid in the tier column, the address in verified_email / email_verified_at). contributor is derived from the data, deliberately: a field on the account could be set by the account.
The email is an unlock on top of contributing, not a replacement. A verified account that has stopped measuring falls back to free, exactly as an unverified one would — emailVerified only lifts a caller who is also contrib.ok.
Confirming an email
POST /v1/verify/start with your API key and an email. We send a link (valid for EMAIL_VERIFY_TTL_S, one hour by default); opening it confirms the address and moves you to contributor_verified while you keep contributing. You can remove the address any time with DELETE /v1/verify/email, which drops you back to contributor without losing any data.
We store the minimum — one confirmed address per key — and never put an email address in a log. The address exists so we can reach you when something changes — a queue you depend on breaks, or your key is about to stop working. This is the one place batchwatch keeps an email, and it is entirely opt-in.
What the delay actually does
delayed_by_s is not a cosmetic label. It is a cut-off timestamp: a delayed caller is served measurements up to now - delay and nothing newer.
- On
/v1/conditions, the "last hour" window is3600 + delayseconds long and ends atnow - delay. An anonymous caller therefore never sees a live signal for free — which is the point, since that route answers "is the queue slow right now?". - On
/v1/wait,/v1/curve,/v1/should-i-batchand/v1/estimate-batchtime, the delay also decides whether a precomputed rollup may be used. A delayed caller has already accepted a figure that is minutes old, so a cron-computed rollup is exactly that. A live caller gets it recomputed. live: falsein a response is there so a machine can see that it is routing on stale data instead of doing so in good faith. The delay hurts most in precisely the quarter-hour when a queue collapses.
/v1/outages and /v1/outages.atom are never delayed for anyone, and /v1/coverage has no delay logic at all.
Caching the advertised endpoints — key-aware, by design (#348/#364)
The endpoints Allowed to crawlers in robots.txt (CRAWLABLE_V1_PATHS in src/seo.js) are fetched directly by Google/Bing/the named AI crawlers. To spare the Worker and D1 a hit on every crawl, the identical-for-everyone ones are cached at the Cloudflare edge for 300 seconds. But the edge keys its cache on the URL, not on the API key, so it may only cache a response whose bytes do not depend on the caller's tier — otherwise one caller's response would be served to another.
The split (established from the code — the "which endpoints vary by key" question this document owns):
| Advertised endpoint | Varies by key? | Cache-Control |
|---|---|---|
/v1/coverage | no — hard-codes the 900s public rollup for every caller | public, max-age=300, s-maxage=300 |
/v1/latency | no — takes no key, applies no delay | public, max-age=300, s-maxage=300 |
/v1/outages | no — never delayed for any tier | public, max-age=300, s-maxage=300 |
/v1/outages.atom | no — same rows as /v1/outages | public, max-age=300 |
/v1/curve | yes — delayed rollup for anonymous, live path for paid | private, no-store |
/v1/hours | yes — rows filtered on now − tier.delayS | private, no-store |
/v1/patterns | yes — rows filtered on now − tier.delayS | private, no-store |
A key-varying endpoint is therefore never shareable at the edge: a paid caller's live figures can never be cached and handed to an anonymous crawler, and an anonymous crawler's delayed figures can never be served to a paying caller. The policy and its Step-0 classification live in src/cache_v1.js; the guard test/cache_key_aware_348.test.js drives the real Worker and reds if any key-varying response ever ships a shareable Cache-Control. /v1/calibration is not advertised (it is a per-request live compute behind the tier gate), so it is not cached and not crawlable.
Becoming a contributor
Five completed measurements in a rolling seven days (CONTRIB_MIN = 5, CONTRIB_WINDOW_S = 7 days in src/index.js). The count is over rows where key_id is yours, status = 'completed', ended_server is set, and started_server is inside the window.
Five in seven days rather than one a day: the latter would punish someone who runs batch twice a week but contributes loyally every time.
Contribution is checked on every request. There is nothing to activate, and it lapses on its own when you stop measuring.
Earning a vote — a separate, higher bar
Being a contributor gets you access. Having your numbers count towards the published percentiles is a different threshold, and it cannot be bought with volume or with new keys (src/robust.js):
| Requirement | Value |
|---|---|
| Has a key (not anonymous) | required |
| Measurements from that key | at least 5 (VOTE_MIN_N) |
| Distinct days measured on | at least 3 (VOTE_MIN_DAYS) |
| Voters needed before per-contributor mode engages | 3 (VOTE_MIN_SOURCES) |
Sources that have not earned a vote — fresh keys and everything anonymous — share one vote between them. Ten new accounts are worth exactly as much as one.
Time cannot be rushed, which is the point: an attacker now has to keep several accounts running over several days with plausible-looking data before any of them counts, and they are visible in the dataset the whole time. The documentation in src/robust.js is explicit that this makes an attack expensive and slow rather than impossible.
The quota
Only the four gated routes count against it, and only for a contributor or paid caller. It is a rolling seven days: the sum of the last seven daily counters, so usage frees up gradually rather than resetting on a boundary.
A successful gated response carries a quota block and an x-batchwatch-quota-left header. For a plain contributor the limit is 5000; a contributor_verified caller sees calls_limit: 10000:
"quota": {
"tier": "contributor",
"calls_used": 2,
"calls_limit": 5000,
"calls_left": 4998,
"window": "7 days"
}
Over the limit gives 429 — see errors.md.
The free trial
Twenty gated calls before you have to contribute anything. It exists to break a chicken-and-egg problem: without it the order is "write the ingest code, run it for a week, then find out whether the answer was worth anything".
How the twenty are counted
- Identity. If you send a key, the trial is counted on the key, so it follows you across networks. Otherwise it is counted on the client IP — never in the clear, but as
SHA-256(TRIAL_SALT + "|" + ip). The IP comes only fromcf-connecting-ip, which Cloudflare writes at the edge and which a caller cannot forge.X-Forwarded-Foris deliberately not used: Cloudflare appends to it, so a caller-supplied value would come first. - No salt, no trial. If the deployment has no
TRIAL_SALT, key-less callers get no trial at all rather than having their IPs stored as effectively-plaintext hashes. The402then saysreason: "no_salt_configured". - Order of consumption. Contribution is checked first. A loyal contributor never spends free calls, so they are still there if they take a break.
- Served first, counted after — and only if a response actually came out. A typo in the query string must not cost a free call. Verified against production: a
422did not move the counter. - The counter is atomic; the gate is not.
used = used + 1cannot be lost, but two concurrent calls can both see 19 remaining and both pass. That is a deliberate choice: locking would mean a write lock on every read, and the error points the harmless way.
Every successful trial call carries a trial block and an x-batchwatch-trial-left header:
"trial": {
"calls_used": 4,
"calls_total": 20,
"calls_left": 16,
"note": "Free trial - no contribution needed yet."
}
The note changes to a warning at 5 or fewer left, and to "That was your last free call..." at zero.
What the trial is not
It is not a security boundary. Someone who wants to can change IP and get twenty more, and src/trial.js says so in as many words. The data is aggregated percentiles, not secrets, and the real defence is that sustained use requires contributing.
What stays open regardless
/v1/wait, /v1/curve, /v1/coverage, /v1/status, /v1/probe, /v1/outages, /v1/outages.atom and /health never touch the trial counter. Neither does contributing — the ingest routes are open on purpose, because without contributions the dataset does not exist.
The MCP server rides these same tiers
The MCP server (mcp/) is a faithful proxy: each tool calls the read route of the same name and carries the agent's BATCHWATCH_TOKEN as the Authorization header, exactly as any client would. So the tiers above apply to it unchanged — there is no separate MCP tier and no MCP fast lane.
- No token → the agent is
anonymous: the same 15-minute delay and the same free trial as any keyless caller. - A token → the agent earns that key's tier —
contributor,contributor_verifiedorpaid— and the same freshness, quota andlivebehaviour that key gets at the HTTP API.
This is deliberate. Selling the live edge to an agent that routes real traffic is the same value proposition as selling it to a human, so it is priced the same way, through the same gate — one place to reason about access. Wiring is in ../../mcp/README.md.