batchwatch

batchwatch › Operational and public routes

Operational and public routes

Health, self-check, prober state, outage feeds, the forced rollup, and the HTML pages the same worker serves.

Health, self-check, prober state, outage feeds, the forced rollup, and the HTML pages the same worker serves.


GET /health

Liveness only. No key, no method check — POST /health also returns 200 (verified against production).

curl https://batchwatch.dev/health

Captured from https://batchwatch.dev, 2026-08-25 14:04 UTC:

{
  "ok": true,
  "ts": 1787666672
}

ts is unix seconds. A 200 here proves the worker is running and nothing else. That is what /v1/status is for.


GET /v1/status

The self-check. Open on purpose: a status page you can only see when logged in is marketing.

It measures the ways this service can fail while still answering 200 to everything:

CheckStatus meanings
cronfail if the newest rollup is older than 1800s (cron runs every 10 minutes); warn if no rollup has ever been computed.
ingestwarn if nothing has been measured for 86400s, or ever.
probewarn if a probe job has been outstanding past its own provider's window (so polling or abandonment has stopped).
probe_successwarn if an actively-probed provider has not completed a single probe recently (a provider failing every probe fast — #186). One flat three-hour backstop for every provider.
probe_stalenesswarn if a probed provider/model's pipeline has stopped turning over, measured against that model's own completion times rather than a shared constant (#382). Two faults: silence (nothing completed for far longer than this model's work takes) and stranded (its oldest outstanding job is past a wider ceiling, while other jobs may still be finishing). Tuned per model rather than flat — much tighter than probe_success for the fast models, and deliberately wider for OpenAI's long batch window, where a three-hour rule would be a false alarm. by_model carries one row per probed provider/model with the numbers behind its verdict — measured, slowest_completion_s, ceiling_s, stranded_ceiling_s, silence_s, oldest_outstanding_s, reason and the n they came from.
contributorswarn if there are fewer than two independent outside sources. Our own prober and our own backfill/import (provenance probe/first_party) are not counted — a dataset that is all ours is not crowdsourced (#276).
probe_pricingwarn if a model is configured for probing but has no price, so it is not probed at all (#379). A probe we cannot cost writes no spend row and would bill real money against a ceiling that cannot see it, so the prober fails closed. unpriced names each provider/model with the missing PRICING key, total is the true count, truncated says whether the list was capped (at 20 — this route is polled by uptime monitors, so it is bounded), and reason is carried once for the whole check rather than per entry. Fix: add the row to PRICING in src/stats.js in a pull request, from the provider's own pricing page — never a guess (#251, #264).
billing_reconcilewarn if the daily Stripe→tier reconcile has stopped, or is correcting drift (a rising count means the billing webhook is failing — #298).
page_reachabilityfail if a fresh synthetic probe of the real public site (/, /pricing, /docs) is down — a non-200, a 200 body under its byte floor, or a 200 missing its content marker (#307). warn if the probe is stale (the probe cron stopped) or has never run. This is the check that catches an edge-503 homepage while the data pipeline is green — the #320 gap. See /v1/uptime.
flex_proberwarn if the flex prober has gone dark — no flex measurement in 1800s, three times its own 600s cadence (FLEX_INTERVAL_S), so one missed run is noise and three is a prober that stopped (#374). It runs on its own independent runner and therefore gets its own check: its silence is never absorbed into everyone else's health. Until the first flex measurement lands the check reads ok — an expected silence, not a fault — and goes live with the first row. age_s carries the age of the newest flex measurement.

The overall status is the worst of the checks. The HTTP status code is 503 when status is fail, and 200 otherwise — including when it is warn. page_reachability is measured over the real public URL by the cron and read here from the last stored probe; /v1/status never fetches the site on your request (that per-request compute is exactly what caused the #304 outage).

Why probe_staleness exists next to probe (#382)

The probe check publishes oldest_age_s per provider, and that number cannot tell two very different states apart:

  1. the provider is genuinely still running the job — normal, an OpenAI batch window is 24 hours and we wait up to 26; and
  2. the poller never reaches that row, so it is never polled at all.

In both states oldest_age_s rises exactly 1:1 with the wall clock. Measured on production 2026-08-31 14:24–14:39 UTC: openai went 26063 → 26971 and google 1854 → 2762 over 908 seconds of wall clock — both exactly +908 — while probe jobs from all three providers were completing normally every tick. So a backlog age tracking the clock is not by itself evidence of a fault.

What is decidable is the provider's own completion time. When the pipeline is healthy a provider's silence is bounded: it submits on a fixed cadence and its jobs finish in a time we have thousands of measurements of. When the pipeline has stopped, that same quantity grows without bound. Two hypotheses, one measurement, opposite behaviour.

The window has to be per model, because a single constant cannot fit them. Straight off the live check at 18:05 UTC on the day it shipped — every column measured, none assumed:

Modelslowest completion (7d)silence ceilingstranded ceilingoldest outstanding
openai/gpt-5-nano49659s26.5h (capped)26.5h (capped)28580s
openai/gpt-5.6-luna40627s22.9h26.5h (capped)28580s
openai/gpt-5.6-sol29050s16.5h26.5h (capped)772s
google/gemini-3.7-flash14130s8.2h16.4h15171s
anthropic/claude-haiku-4-5660s42m84m171s

Anthropic is watched at forty-two minutes and OpenAI at a day, off the same rule, because that is what each model actually does. probe_success's single three-hour ceiling is far too tight for gpt-5-nano, which has genuinely run 13.8 hours, and far too loose for claude-haiku-4-5, which has never taken more than eleven minutes — three hours of silence there is ~45 missed completions read as green. The spread is 75× within OpenAI alone, which is why the ceiling is per model: one shared with nano would leave the others unwatched.

A note on where those figures come from, because it caught us out. The bulk export also publishes a max_s, and it is not the same number — it comes from the rollup table, over the rollup's window and exclusions, and read the same day it was up to 20× smaller (gemini 709s against the 14130s above). Neither bounds the other. The ceilings here are computed from the check's own query over completed probe rows, and the durations in it are the provider's reported end times, not the moment we happened to poll — so a slow poll cannot inflate a ceiling and blunt the check.

So each model's ceiling is multiple × (submit cadence + the slowest completion actually measured in the last 7 days), then capped at that provider's abandon threshold. The cap is what makes the check strictly an addition: it can fire earlier than probe does, never later. silence uses a multiple of 2 and stranded a multiple of 4 — "the oldest of N concurrent jobs" is an extreme order statistic with a much fatter tail than "the gap between consecutive completions", and the estimator is censored (a job that outruns its window is abandoned and never enters the maximum), so the measured ceiling is biased low and the wider rule carries the headroom.

A model with too few measured completions is still watched, at its provider's abandon threshold, and its row reports measured: false — we decline to publish a tighter ceiling we have not earned, but the check never goes quiet. That distinction matters: an earlier draft returned no ceiling at all in that case, which meant a poller dead long enough to empty the measurement window would have silenced the very alarm meant to catch it.

Example

Captured from https://batchwatch.dev, 2026-08-25 14:04 UTC — HTTP 200 with status: "warn". This capture predates six of the current checks — probe_success, probe_staleness, probe_pricing, billing_reconcile, page_reachability and flex_prober — so its checks array shows only four of the current set; the shape of each entry is unchanged, and the table above is the current, complete list. (Live today, /v1/status emits ten checks: cron, ingest, probe, probe_success, probe_staleness, contributors, probe_pricing, billing_reconcile, page_reachability, flex_prober.) (A fresh capture is not fabricated here — #30: we quote a real past response, not an invented one.)

curl https://batchwatch.dev/v1/status
{
  "status": "warn",
  "checked_at": "2026-08-25T14:04:33.000Z",
  "checks": [
    { "name": "cron", "status": "ok", "age_s": 229, "note": null },
    { "name": "ingest", "status": "ok", "age_s": 719, "note": null },
    { "name": "probe", "status": "ok", "outstanding": 1, "oldest_age_s": 2634, "note": null },
    {
      "name": "contributors",
      "status": "warn",
      "contributors": 1,
      "probe_share": 1,
      "note": "Fewer than two independent outside sources. The numbers are true but they are not crowdsourced, and coverage says so."
    }
  ],
  "note": "One or more subsystems can fail without any request returning an error. That is what these checks exist to catch."
}

The field names inside checks are English (name, age_s), matching every other route. They come straight from src/selftest.js. (They were Danish (navn, alder_s) until the de-Danish rename; the JSON above shows the current contract, so the field names differ from the raw 2026-08-25 capture while the measured values do not.)

The 503 case was not observed. Producing it would have required a live subsystem to stop, which is not something to arrange for a documentation example. From handleStatus in src/index.js, the response body is identical in shape with "status": "fail" and at least one check at fail.

Two checks can return fail, and they mean opposite things:

So read the failing check's name before deciding what a 503 means — the two are not interchangeable.


GET /v1/uptime

The page-reachability history behind the page_reachability check (#307). Open, like /v1/status: our own site's reachability is not a visitor's data.

The cron probes the real public URL over HTTP (/, /pricing, /docs) every few minutes and stores each result — status, byte count, latency, and whether the content marker was present. This endpoint is a cheap read of those stored rows; it does not fetch the site on your request (the #304 lesson). It returns the newest probe per route plus a recent window.

curl https://batchwatch.dev/v1/uptime
{
  "schema": "batchwatch.uptime.v1",
  "generated_at": "2026-08-30T16:40:00.000Z",
  "status": "operational",
  "routes": [
    { "route": "/", "url": "https://batchwatch.dev/", "up": true, "http_status": 200, "bytes": 41234, "latency_ms": 38, "probed_at": "2026-08-30T16:38:00.000Z", "age_s": 120, "reason": null },
    { "route": "/pricing", "url": "https://batchwatch.dev/pricing", "up": true, "http_status": 200, "bytes": 18922, "latency_ms": 31, "probed_at": "2026-08-30T16:38:00.000Z", "age_s": 120, "reason": null },
    { "route": "/docs", "url": "https://batchwatch.dev/docs", "up": true, "http_status": 200, "bytes": 9004, "latency_ms": 29, "probed_at": "2026-08-30T16:38:00.000Z", "age_s": 120, "reason": null }
  ],
  "recent": [ ... ],
  "note": "Every probed route returned a real page (200, above its byte floor, with its content marker) on its newest probe."
}

status is a measured verdict: operational only when every route's newest probe was up, degraded the moment one is down, and no_data when nothing has been probed yet — never a green by default (#30). A down route carries the measured reason (http 503, body 12 bytes < 2000 floor, marker "batchwatch" absent, or the fetch error). The illustrative values above are shaped, not a live capture — the live bytes/latency_ms are recorded by the first real cron probe.

Alerting. When a route fails N consecutive probes (default 3, ~10–15 min), an alert is sent through the same operator subscription path provider outages use (the batchwatch self-alert; see alerts.md), once per outage episode, so a human hears about a down homepage without anyone running a curl.


GET /v1/cron-trace

A per-tick breadcrumb ledger for the scheduled() cron (#315), so a mid-chain cron death is diagnosable over plain HTTP without wrangler or the Cloudflare dashboard. Open (not trial-gated) and PII-free: it is our own cron's health — tick names, statuses, wall-ms, and whitelisted scalar counts — never a visitor's data. A cheap indexed read of stored rows; it does not run the cron on your request.

The cron writes a __start__ breadcrumb first, one row after each tick returns, and an __end__ breadcrumb last (see scheduled-cron-wiring.md). A run that fired but died in a tick has __start__ + the completed ticks and no __end__ — so complete:false with last_completed_tick names the tick after which the invocation died (the killer is the next tick in the fixed cron order).

curl https://batchwatch.dev/v1/cron-trace          # newest ~5 fires
curl 'https://batchwatch.dev/v1/cron-trace?runs=10' # widen the window (≤20)
{
  "now": 1788153708,
  "newest": {
    "run_id": 1788153707,
    "started": true,
    "ended": true,
    "complete": true,
    "tick_count": 13,
    "last_completed_tick": "calibration",
    "degraded_ticks": [],
    "breadcrumbs": [
      { "seq": 0, "tick": "__start__", "status": "ok", "ms": null, "detail": null, "at": 1788153707 },
      { "seq": 1, "tick": "rollup", "status": "ok", "ms": 45, "detail": "{\"rollups\":8,\"attempted\":8,\"models\":8,\"timed_out\":0,\"ms\":17,\"failed\":0}", "at": 1788153707 },
      { "seq": 8, "tick": "provider_status", "status": "written", "ms": 452, "detail": "{\"status\":\"written\",\"rows\":34}", "at": 1788153708 },
      { "seq": 14, "tick": "__end__", "status": "ok", "ms": null, "detail": null, "at": 1788153708 }
    ]
  },
  "runs": [ ... ]
}

complete is a measured verdict, not a default: a run is complete only when it has both a __start__ and an __end__. degraded_ticks surfaces any tick that timed out or threw (tick_timeout/tick_threw/threw) without killing the whole invocation — a slow/erroring tick is a lead even on an otherwise-complete run. The breadcrumbs above are elided for brevity; a live fire carries all thirteen ticks in order between the two sentinels.


GET /v1/probe

The project's own measurement jobs. Open: a product that asks for trust in a dataset should be able to show how the dataset is made.

Captured from https://batchwatch.dev, 2026-08-25 14:05 UTC (the recent array is truncated here to three of its six entries):

curl https://batchwatch.dev/v1/probe
{
  "configured": [
    "openai/gpt-5.6-luna",
    "openai/gpt-5.6-sol",
    "openai/gpt-5-nano",
    "anthropic/claude-haiku-4-5",
    "google/gemini-3.7-flash"
  ],
  "not_configured": [],
  "interval_s": 3600,
  "outstanding": 1,
  "recent": [
    {
      "provider": "google",
      "model": "gemini-3.7-flash",
      "status": "completed",
      "started_at": "2026-08-25T13:50:39.000Z",
      "duration_s": 115
    },
    {
      "provider": "anthropic",
      "model": "claude-haiku-4-5",
      "status": "completed",
      "started_at": "2026-08-25T13:30:39.000Z",
      "duration_s": 64
    },
    {
      "provider": "openai",
      "model": "gpt-5.6-luna",
      "status": null,
      "started_at": "2026-08-25T13:20:39.000Z",
      "duration_s": null
    }
  ],
  "note": "These are our own measurement jobs - real batch calls against the real queue, not simulations. They count in the statistics and are marked source=probe so coverage can tell you when a dataset is really just us."
}

configured lists provider/model pairs, not providers, so a provider with three models appears three times. not_configured lists providers with no usable credentials. recent is the last 20 probe rows by start time; status: null with duration_s: null means the job is still outstanding.

outstanding counts the whole table, not just the twenty rows in recent. It used to be filtered out of that LIMIT 20 slice, so it could never report more than twenty and under-reported precisely when the backlog was growing — measured on production 2026-08-31 14:48 UTC it published 7 while /v1/status counted 25 outstanding over the same table at the same moment. Both surfaces now read the same predicate, so they cannot disagree again (#382).

unpriced (added by #379, so it postdates the capture above) reports the models that are configured but have no PRICING row and are therefore refused, not probed. A probe we cannot cost writes no spend row, so it would bill real money against a ceiling that cannot see it; the prober fails closed rather than measure blind.

It is an object, not a bare list, and it is bounded — 200 unpriced names would otherwise make this field 42.8 KB:

fieldmeaning
modelsup to 20 entries of provider / model / price_key — the PRICING key to add
totalhow many there really are, capped list or not (#251: a gap is a gap)
truncatedwhether models was shortened
reasonwhy, once for the whole response rather than repeated per entry; null when there is nothing to explain

The same fact is a named degraded condition on /v1/status (probe_pricing) and rides on the probe tick's /v1/cron-trace breadcrumb as unpriced, capped to ~80 characters there because that column degrades the whole row when it overflows. Design of record: probe-price-gate-379.md.

Probe rows are marked source=probe and do count in the statistics; that is what probes and crowdsourced in /v1/coverage exist to disclose.


GET /v1/outages

Outage state. Never delayed, for any tier — a feed that is fifteen minutes late is not a feed, and what is given away here is a boolean about a provider rather than the distribution that is the product.

HEAD mirrors GET here (same status and content-type, empty body), so a monitor that leads with a HEAD liveness check sees the resource as alive. Same for /v1/outages.atom and /v1/outages.rss.

status is one of:

ValueMeaning
degradedAt least one outage is open.
operationalNo open outage, and at least one provider/model can actually be watched.
insufficient_sourcesNothing can be watched: no model has 3 established contributors. This is not a claim that the providers are healthy.

A model is watchable when at least OUTAGE.MIN_SOURCES (3, same as the voting threshold) independent third-party keys have contributed in the rollup window. The contributors figure in coverage counts distinct third-party keys (#276 — our own prober and backfill/import are excluded, so three of our own keys cannot make a model look watchable by the crowd) and is an upper bound on voters: it does not apply the five-measurements-over-three-days rule that earns a vote. The response says so itself.

Detection uses hysteresis: an outage opens when the fresh median is 3× the baseline, and closes at 1.5×, with an absolute floor of 900s so that a queue going from 20s to 70s is not called an outage.

Captured from https://batchwatch.dev, 2026-08-25 14:04 UTC:

curl https://batchwatch.dev/v1/outages
{
  "schema": "batchwatch.status.v1",
  "generated_at": "2026-08-25T14:04:39.000Z",
  "status": "insufficient_sources",
  "watching": 0,
  "not_watchable": 0,
  "open": [],
  "recent": [],
  "coverage": [],
  "note": "Nothing here is a claim that the providers are healthy. We do not yet have enough independent contributors on any model to tell an outage from one account having a bad day.",
  "live": true,
  "delayed_by_s": 0,
  "delay_note": "Outage state is never delayed, unlike the percentile API. A feed that is fifteen minutes late is not a feed - and what is given away here is a boolean about a provider, not the distribution that is the product.",
  "coverage_note": "contributors counts distinct INDEPENDENT third-party keys (our own prober and backfill are excluded) and is an upper bound: it does not apply the five-measurements-over-three-days rule that earns a vote."
}

coverage is empty above because the query behind it requires key_id IS NOT NULL, and production held no completed measurement from a keyed contributor in the window at capture time. (The dataset was not empty — /v1/coverage listed five models — so the inference is that the prober's rows carry no key_id on this deployment. That was inferred from the two responses, not read out of the database.) open and recent were both empty, so the shape of an outage entry was not observed. From outageToJson in src/outage.js, each entry carries id, provider, model, mode, severity, started_at, ended_at, duration_s, peak_ratio, baseline_p50_s, ended_reason and ongoing. recent is capped at the 20 most recent closed outages.


GET /v1/outages.atom

The same state as an Atom feed, for anything that reads feeds. Content type application/atom+xml; charset=utf-8, and — unlike the JSON API — it is cacheable: cache-control: public, max-age=300.

Captured from https://batchwatch.dev, 2026-08-25 14:05 UTC, verbatim:

<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <id>tag:batchwatch.dev,2026:status</id>
  <title>batchwatch - batch queue status</title>
  <subtitle>Measured queue time on LLM batch APIs. We report a slowdown only when a majority of established contributors see it.</subtitle>
  <updated>2026-08-25T14:05:41.000Z</updated>
  <link rel="self" href="https://batchwatch.dev/v1/outages.atom"/>
  <link href="https://batchwatch.dev"/>

</feed>

The feed contains both open and closed outages as entries. There were none at capture time, so an entry was not observed.

Note: /v1/status.atom does not exist (verified: 404), even though it is the default selfUrl inside src/outage.js.


POST /v1/rollup/refresh

Forces a recomputation of the precomputed rollups instead of waiting for the ten-minute cron. Requires a key (401 {"error":"api key required"}), because it costs exactly what the rollups exist to avoid. Any valid key will do — no tier check.

Captured from a local wrangler dev instance:

{
  "models": 2,
  "rollups": 4,
  "models_total": 2,
  "deferred": 0,
  "ms": 48
}

models_total is how many provider/model/mode combinations were eligible; models is how many were reached before the time budget ran out, and deferred is the remainder, which the next run picks up. One rollup is written per delay tier, which is why rollups is a multiple of models.


HTML and crawler routes

The same worker serves the site. These are GET/HEAD only and are matched before authentication, so a crawler hitting a thousand pages costs no key lookups.

PathContent typeVerified
/, /index.htmltext/html200
/m/{provider}/{model}text/html200 for /m/openai/gpt-5-nano; 404 for /m/openai/not-a-model
/p/{provider}text/html200 for /p/openai; 404 for /p/notaprovider
/robots.txttext/plain200
/sitemap.xmlapplication/xml200

A model page exists only when there is a usable rollup for it; otherwise the worker returns a 404 page rather than an empty one.

Model names containing a slash (meta-llama/Llama-3) are supported: the model segment is everything after the provider, and encodeURIComponent writes the slash as %2F.

Captured from https://batchwatch.dev, 2026-08-25 14:05 UTC:

curl https://batchwatch.dev/robots.txt
User-agent: *
Allow: /
Disallow: /v1/
Disallow: /health

Sitemap: https://batchwatch.dev/sitemap.xml

The sitemap lists the front page, one entry per provider and one per model, with lastmod and priority.


Known discrepancies

Found while writing this reference, listed so nobody has to find them twice. None of these have been changed — they are reported, not fixed.

  1. docs/KRAVSPEC.md names six routes that do not exist: /v1/batches, /v1/batches/complete, /v1/forecast, /v1/tradeoff, /v1/anomalies, /v1/status.atom. All six returned 404 from production on 2026-08-25, in the same run in which /v1/coverage returned 200.
  1. DELETE is missing from the CORS preflight response. OPTIONS returns access-control-allow-methods: GET,POST,PATCH,OPTIONS, but DELETE /v1/calls/mine and DELETE /v1/keys/current both exist. Browser calls to those two routes will fail preflight. Verified against production.
  1. A revoked key gets the "no key given" message. After DELETE /v1/keys/current, reusing the same token on GET /v1/keys/current returns 401 {"error":"no key given"} — but a key was given. The caller is sent looking for a missing header rather than told the key is revoked. Verified on a local instance.
  1. README.md says the Anthropic and Google probers were "written, never run (no key yet)", but production /v1/probe on 2026-08-25 listed anthropic/claude-haiku-4-5 and google/gemini-3.7-flash as configured, with completed jobs 64s and 115s old. The README table is out of date.
  1. risk is validated on one route and not the other. /v1/estimate-batchtime rejects an unknown risk with 422; /v1/should-i-batch silently falls back to p90. The defaults also differ (p50 vs p90).