batchwatch

batchwatch › Contributing measurements

Contributing measurements

There are two ways to get a measurement into the dataset, and they are not equally trusted.

There are two ways to get a measurement into the dataset, and they are not equally trusted.

The live path — POST /v1/calls when your batch job starts, PATCH /v1/calls/{id} when it finishes — is open to everyone, key or no key. The duration is measured against the server's clock, so a fabricated eight-hour wait costs eight real hours.

The import path — POST /v1/calls/complete — takes your timestamps, so nothing can be checked against the server's clock. It therefore requires a key.

A key is not needed to contribute, but it is needed to count: measurements without a key are all lumped together as one anonymous source in the robust percentile calculation, and only a key can earn a vote. See interpreting.md.

Note on the examples below. These are captured from a local wrangler dev instance running this exact code, not from https://batchwatch.dev. Demonstrating the write routes against production would have meant putting measurements into the live dataset that were not real measurements, and the dataset is the entire product. The request and response shapes are real; the durations are not measurements of anything.


Idempotency-Key — retry without counting twice

Optional header on POST /v1/calls and POST /v1/calls/complete. Send no header and nothing changes — same behaviour as before, and not a single extra database lookup.

curl -X POST https://batchwatch.dev/v1/calls \
     -H "authorization: Bearer $BW_KEY" \
     -H "idempotency-key: $(uuidgen)" \
     -H 'content-type: application/json' \
     -d '{"provider":"openai","model":"gpt-5-nano","mode":"batch","requests":1200,"input_tokens":840000}'

Why you want it

A network timeout is ambiguous for you, not for us: you cannot tell whether the write landed. Either you drop the measurement or you send it again — and sending it again puts the same job in the dataset twice, where it counts twice in the percentiles. Neither of us can spot it afterwards: two identical jobs on one model is a perfectly legal thing to happen.

It is not even random noise. The calls that time out are disproportionately the slow ones, so duplicates pull the tail.

The rules

SituationResponse
Same key, same request, within 24hThe first response, replayed verbatim. Nothing new is written. x-batchwatch-idempotency-replayed: true.
Same key, different request409, reason: "body_differs". Nothing is written.
Same key, identical request still in flight409, reason: "in_progress". Retry in a moment.
Same key after 24hTreated as a new request.

The "different request" case is deliberately loud. It is not a retry — it is two different things sent under one key, and we cannot tell which you meant. Staying quiet and writing it anyway would put your bug in our dataset where nobody can see it.

The comparison is over the HTTP method, the route and the raw body bytes. Two bodies that mean the same thing but differ as text conflict. Normalising would mean we decide when two of your calls are "really" the same, and then we are the ones losing a measurement for you.

Scope, and what it means if you have no key

Keys are scoped per caller, so two customers can both pick "1" and never see each other's response.

Other details


POST /v1/calls

Open a measurement. No key required; send one if you want the measurement to count as yours.

Body (JSON)

FieldTypeRequiredDefaultNotes
providerstringyes—One of openai, anthropic, google, mistral, azure, other. Anything else is 422.
modelstringyes—Max 100 characters.
modestringnobatchbatch or sync. Anything else is 422.
endpointstringnonullFree text, stored as-is.
requestsnumbernonullMust be non-negative.
input_tokensnumbernonullMust be non-negative.
output_tokensnumbernonullValidated here but only stored on finish.
ttfb_msnumbernonullValidated here but only stored on finish.
started_atISO-8601 string or unix secondsno—Advisory only: used to compute clock skew, never as the start time.
sourcestringnouserThe literal string probe marks it as a project probe; anything else becomes user.
service_tierstringnonullThe pricing tier, orthogonal to mode: standard or flex. flex requires mode: "sync" and a flex-capable provider (openai, google); both are 422 otherwise. Omit it and the row keeps writing null, exactly as every pre-flex row reads.

The region is taken from Cloudflare's cf-ipcountry header, not from the body.

Response

201 with the assigned id. The id format is c_ followed by 20 hex characters, and only that format is accepted by the PATCH route.

curl -X POST http://localhost/v1/calls \
  -H "authorization: Bearer $BW_KEY" -H 'content-type: application/json' \
  -d '{"mode":"batch","provider":"openai","model":"gpt-5-nano",
       "endpoint":"/v1/chat/completions","requests":1200,
       "input_tokens":840000,"started_at":"2026-08-25T14:10:00Z"}'
{
  "id": "c_0ab78694c0034ebf9b10",
  "recorded_at": 1787666964
}

Clock skew warning

If started_at is more than 300 seconds away from the server's clock, the response carries a warning. The measurement is still accepted; the server's own timestamps are what get used.

{
  "id": "c_0df9295f56b24a189520",
  "recorded_at": 1787666979,
  "warning": "Your clock is -3600s off ours. We use our own timestamps, so your measurement is still valid — but check your clock."
}

Validation failure

422, with every problem listed at once:

curl -X POST http://localhost/v1/calls -H 'content-type: application/json' \
     -d '{"provider":"acme"}'
{
  "error": "validation failed",
  "details": [
    "provider must be one of: openai, anthropic, google, mistral, azure, other",
    "model is required (string, max 100 chars)"
  ]
}

PATCH /v1/calls/{id}

Close a measurement. The path must match ^/v1/calls/(c_[a-z0-9]+)$; anything else falls through to 404 unknown route.

Body (JSON)

FieldTypeRequiredDefaultNotes
statusstringnocompletedcompleted, failed, expired, cancelled, abandoned, timeout. Anything else is 422.
fail_reasonstringnonullWhy a non-completed call ended that way. Free text, control characters stripped, clipped at 300 characters. Never affects a duration — it only classifies.
ended_atISO-8601 string or unix secondsnoarrival timeAccepted only if it lies between the recorded start and now. Otherwise the server's arrival time is used instead, silently, and the deviation is stored.
output_tokensnumbernonull
ttfb_msnumbernonull

abandoned means "we stopped waiting" — the job did not expire, and the true duration is unknown but longer than what was recorded. Only completed rows are used in percentiles.

abandoned and timeout are not synonyms. abandoned is a batch job we stopped waiting for. timeout is a synchronous call the provider accepted and that outran the caller's own deadline. Both are right-censored; they are kept apart because the tiers they describe are.

Saying WHY, and why that matters most on flex

fail_reason is what turns one undifferentiated failed bucket into the facts inside it. On the flex tier that distinction is the most valuable thing we collect:

Outcomestatusfail_reasonWhat it is
servedcompleted(absent)the call ran. This one has a duration.
refusedfailedflex_resource_unavailablethe provider declined in admission control, before inference. Instant, free, and a direct reading of how busy the tier is.
timed outtimeouttimeoutaccepted, then outran your deadline. The true wait is unknown but longer.
erroredfailedyour own reason, verbatimtransport, auth, a 500.

A refusal is not an error, and reporting it as one loses the measurement entirely: our aggregation counts a refusal as exactly status = 'failed' AND fail_reason = 'flex_resource_unavailable', and anything else that failed as an error. Use the string verbatim.

Contributing a flex measurement is therefore two extra fields on the calls you already report: mode: "sync" with service_tier: "flex" on the open, and status + fail_reason on the close. Every client library does it for you with a named method per outcome, so there is nothing to look up.

Response

{
  "id": "c_0ab78694c0034ebf9b10",
  "duration_s": 0,
  "status": "completed"
}

(duration_s is 0 here because the example opened and closed the measurement within the same second. The write still succeeds — this path is fail-open — but a completed row with a zero or negative duration is invalid data, not a fast job, so it is marked excluded on write and never enters a percentile.)

When the status is completed and at least 10 completed measurements exist for that provider/model/mode over the last 30 days, the response also carries percentile (where this job landed, 0–100, one decimal) and sometimes a note: "That was unusually slow for this model." at p95 or above, or "Faster than most." at p20 or below.

Errors

StatusBodyWhen
400{"error":"body must be JSON"}Body did not parse.
404{"error":"unknown call id"}No row with that id.
403{"error":"this call belongs to another key"}The row has a different key_id than the authenticating key.
409{"error":"already finished"}The row already has an end time.
422{"error":"unknown status: ..."}Status not in the list above.

Verified 409 and 404 against a local instance:

{ "error": "already finished" }
{ "error": "unknown call id" }

POST /v1/calls/complete

Bulk import of finished jobs, using your own timestamps. Requires a key.

Without one, the route explains itself rather than just refusing:

curl -X POST https://batchwatch.dev/v1/calls/complete -d '{}'

Captured from https://batchwatch.dev, 2026-08-25 14:05 UTC — 401:

{
  "error": "api key required",
  "why": "This route takes your own timestamps, so it cannot be verified against our clock. Anonymous history import would let anyone assert any duration instantly.",
  "instead": "POST /v1/calls when the job starts and PATCH it when it finishes - that path is open to everyone, because the duration is measured against our clock and cannot be faked."
}

Body

Either a single object or an array of them. Maximum 500 records per request (413 above that). An empty array is 400.

Each record takes the same fields as POST /v1/calls, plus:

FieldTypeRequiredNotes
started_atISO-8601 or unix secondsyes
ended_atISO-8601 or unix secondsyesMust be strictly after started_at. A zero- or negative-duration record (ended_at at or before started_at) is rejected: a job that starts and ends in the same second never went through a queue, and a zero-second batch "wait" is invalid data, not a fast job.
statusstringnoDefaults to completed. Same list as PATCH, timeout included.
fail_reasonstringnoSame as PATCH. A spooled refusal replays through this route, so the reason travels with it — otherwise a network blip would quietly downgrade a refusal to an anonymous failure.
output_tokensnumberno

Here both client timestamps are stored and used as the server timestamps — that is exactly why the route needs a key.

Test-mode keys. A key whose label is exactly batchwatch-e2e-test is a designated test-mode key: its rows are still written (the call succeeds) but are marked excluded on ingest, so they never enter a published percentile. This exists so our own end-to-end test suite — which runs against production every release — cannot move the public numbers. It applies to POST /v1/calls and POST /v1/calls/complete alike. There is no reason to use it for real measurements: excluded rows count for nothing.

Provenance (who a measurement came from). Every row records a provenance at ingest, derived from the key: probe (our own prober), first_party (ours — a key we have marked first-party, or a test-mode key), third_party (a genuine outside contributor), or unknown (a historical row written before provenance existed; never a backdated guess). This is what the published contributor count and the crowdsourced flag are computed over: only distinct third_party sources count, so a dataset that is entirely ours (prober + our own backfill) is never reported as crowdsourced. It is derived from the key and stored on the row, so a later change to the key cannot rewrite the history. You do not set it — it is decided for you.

Response

Partial success is normal: valid records are stored, invalid ones are reported by index. Status is 201 when anything was accepted, 422 when everything was rejected.

{
  "accepted": 1,
  "rejected": 1,
  "ids": ["c_684f85eb12f34bf3b365"],
  "errors": [
    { "index": 1, "errors": ["duration must be positive (ended_at after started_at)"] }
  ]
}

Partial completion — how a split batch is measured

A batch of 20,000 requests does not come back as one thing. Some land, some fail per-request, and some are still outstanding when the 24-hour expiry hits. The measurement schema must not collapse those into "completed", and it does not have to — status already carries the three readings, and the rule below is what the client SDK's job.split() and every hand-rolled importer should follow so the dataset stays honest.

The job as a wholestatus to reportWhy
Every request landedcompletedThe clean case. Counts in percentiles.
Some requests failed per-request, the batch itself completedcompletedThe batch held its wait — the per-request failures are the caller's payload problem, not a queue-time signal. Report the batch's real duration.
Outstanding requests at the 24h cutoffexpiredThe queue was too slow to finish in the window. Not completed (that pollutes p90 with a censored duration) and not failed (that throws away the "the queue was slow" signal, which is the whole product).

Only completed rows are used in percentiles (see PATCH), so an expired job is recorded and visible but does not drag the median toward a duration the batch never actually reached. This is the same reasoning as abandoned ("we stopped waiting"): a censored wait is real information, but it is not a completed-duration measurement and must never be counted as one.

The client never invents a fourth status. expired, failed and completed are already accepted by both write routes above; the SDK maps a partial batch onto them and nothing new is added to the schema.


GET /v1/calls/mine

Everything the service holds that came from your key. Requires a key (401 {"error":"api key required"} without one — verified against production).

Parameters

NameTypeDefaultNotes
afternumber (unix seconds)0Returns rows with started_server strictly greater than this.
limitnumber500Clamped to 1–1000.

Pagination is keyed on started_server, not on an offset, so rows arriving during a walk cannot make you skip anything. Follow the next field; it is null on the last page.

Response

Captured from a local wrangler dev instance, limit=2:

{
  "label": "prod-pipeline",
  "count": 2,
  "next": "/v1/calls/mine?after=1787666964&limit=2",
  "calls": [
    {
      "id": "c_684f85eb12f34bf3b365",
      "mode": "batch",
      "provider": "anthropic",
      "model": "claude-haiku-4-5",
      "endpoint": null,
      "requests": 40,
      "input_tokens": 120000,
      "output_tokens": 38000,
      "started_client": 1787216400,
      "ended_client": 1787218860,
      "started_server": 1787216400,
      "ended_server": 1787218860,
      "status": "completed",
      "ttfb_ms": null,
      "region": null,
      "source": "user",
      "provenance": "third_party",
      "excluded": false,
      "clock_skew_s": -448119,
      "provider_job_id": null,
      "started_at": "2026-08-20T09:00:00.000Z",
      "ended_at": "2026-08-20T09:41:00.000Z",
      "duration_s": 2460
    }
  ],
  "note": "Everything we hold that came from this key. Server timestamps are the ones used in statistics; client timestamps are advisory and kept for diagnostics."
}

(One of the two returned rows is shown; the second is elided.)

provenance is who the measurement came from, decided at ingest from your key: third_party for an ordinary key like this one, first_party if the key is one of ours, probe for our prober, unknown for a row written before provenance existed. It is what the public contributor count and crowdsourced flag are computed over — see interpreting.md.

Rows come back ordered by started_server ascending. There is no route to anyone else's rows.


DELETE /v1/calls/mine

Take your measurements out of every aggregate. Requires a key.

The rows are flagged excluded = 1, not deleted. They leave every percentile immediately — there is no rebuild to wait for, because every query filters on the flag. Keeping the rows means a mistaken request can be undone by the operator.

Captured from a local wrangler dev instance:

{
  "excluded": 8,
  "precomputed_rows_dropped": 1,
  "note": "8 measurements are out of every aggregate as of now, and the 1 precomputed row built on them have been dropped. That covers the public model pages too, not just the live queries.",
  "still_stored": "The rows are flagged, not deleted, so a mistaken request can be undone. Write to us for physical deletion.",
  "reversible_by": "the operator, on request from this key"
}

Calling it again when there is nothing left to exclude:

{
  "excluded": 0,
  "note": "Nothing to exclude - this key has no measurements counting today.",
  "still_stored": "The rows are flagged, not deleted, so a mistaken request can be undone. Write to us for physical deletion.",
  "reversible_by": "the operator, on request from this key"
}

Note the difference from revoking a key: revoking stops the key from authenticating and leaves the measurements in the dataset. This route is the one that removes their influence.