Version 1.0 · 31 August 2026
1. The short version
- We receive metadata about your batch job — which model, how many tokens, when it started, when it finished, whether it succeeded. Nothing from inside the request.
- We never receive prompts, completions, system prompts, tool definitions or your provider API keys. Not sent, not hashed, not sampled.
- There are no passwords and no accounts in the usual sense. There are exactly two circumstances in which we hold an email address, and both are yours to switch off.
- An email you chose to verify, via
POST /v1/verify/start. It lifts your key to the verified-contributor tier and lets us contact you, andDELETE /v1/verify/emailremoves it. If you never verify, we hold no email for you. - A recovery address, if you bought a subscription and left key recovery on. Because we store a hash of your token and never the token itself, a lost paid key used to be unrecoverable — you had paid for access you could not reach. So for paying customers only, and only while you leave the box ticked at checkout, we keep the email address on your Stripe receipt, purely so we can issue you a replacement key. It is used for that and nothing else, and
DELETE /v1/keys/recovery-emailremoves it. If you have not paid, this one does not apply to you at all. Nothing else about you is stored as an account. - API keys are stored as a SHA-256 hash, never in clear text. We cannot recover your token, not even for you.
- The free trial is counted against a salted hash of your IP address, never the address itself — and if the salt is missing there is no trial rather than a hash that only pretends to be anonymous.
- We keep a country code derived from your IP. Not the address.
- Measurements are not deleted when a key is revoked. That is a deliberate product decision, and §7 explains exactly how it interacts with your right to erasure.
2. Everything we store, field by field
This section is the schema, in English. If a field is not listed, we do not store it.
call — one row per measured job
| Field | Contents | Source |
|---|---|---|
id | opaque row id, c_ + 20 hex chars | us |
key_id | which API key submitted it, or NULL if anonymous | us |
mode | batch or sync | you |
provider | openai, anthropic, google, mistral, azure, other | you |
model | the model name, verbatim | you |
endpoint | e.g. /v1/chat/completions | you |
requests | number of requests in the batch | you |
input_tokens, output_tokens | token counts | you |
started_client, ended_client | your timestamps — advisory only | you |
started_server, ended_server | our timestamps — authoritative | us |
status | completed, failed, expired, cancelled, abandoned, timeout | you |
fail_reason | why a non-completed job ended that way — e.g. flex_resource_unavailable for a provider refusal. Free text, control characters stripped, capped at 300 characters | you |
ttfb_ms | time to first token, synchronous calls only | you |
region | two-letter country code only, from Cloudflare's CF-IPCountry header | Cloudflare edge |
source | user or probe (our own measurement jobs) | us |
excluded | quality flag; excluded rows are removed from every published figure | us |
clock_skew_s | difference between your clock and ours, for diagnostics | us |
provider_job_id | the provider's own job id — only ever set on our own prober's rows, never on yours | us |
requested_service_tier | the pricing tier you asked for: standard or flex, or empty | you |
served_service_tier | the tier the provider reported serving, verbatim, or empty when it reported none | provider |
planned_output_tokens | the output cap a probe set out to request — our own prober's rows only | us |
provenance | probe, first_party, third_party or unknown — who the row came from, derived from the key at ingest | us |
quarantined | quality flag, like excluded | us |
We do not store the IP address on this row. Only the country code derived from it, at the edge, by Cloudflare.
Three fields carry free text you control, and they are the only place personal data could reach us by accident:
modelandendpointare strings you send.modelis capped at 100 characters and is expected to be a model name.endpointis stored verbatim.fail_reasonsays why a job did not complete. Control characters are stripped and it is capped at 300 characters, but the text itself is stored as you sent it. Send the provider's error code, never raw request or response text. Our own client libraries never fill this field from an exception for exactly that reason: a caller-side exception sends the fixed stringclient_exception, never the exception's message.- Do not put customer identifiers, usernames, ticket numbers, internal hostnames or anything else personal into these fields. They are stored and they are not scrubbed.
api_key
| Field | Contents |
|---|---|
id | integer |
token_hash | SHA-256 of the bearer token. The token itself is never stored |
label | your own note, e.g. prod-pipeline. Control characters stripped, trimmed to 60 characters, never shown to anyone else |
created_at, revoked_at | timestamps |
tier | free or paid |
verified_email | the email you chose to verify, or NULL. Set only when you opt in via POST /v1/verify/start and confirm; you can remove it with DELETE /v1/verify/email |
email_verified_at | timestamp of that confirmation, or NULL |
recovery_email | the email on your Stripe receipt, or NULL. Written only when a payment completes on a paid key and you left key recovery ticked at checkout. Used for one thing: issuing you a replacement key if you lose this one. Never used for marketing, never shared. Remove it with DELETE /v1/keys/recovery-email |
recovery_email_at | when that address was stored, or NULL |
recovery_sent_at | when we last sent a recovery link for this key, or NULL. A cooldown, so that nobody who knows your address can use the recovery route to flood your inbox. Overwritten each time; not kept as history |
Retention of recovery_email: for as long as the key exists, because that is exactly as long as it can be useful. It moves with the key when the key is rotated or recovered, and it is deleted from the old key at that moment. It is also deleted the moment you ask, without affecting your subscription.
Do not put a personal name or an email address in label. It is free text and we do not scrub it. The two emails we hold on purpose are verified_email (only if you opt in) and recovery_email (only if you pay and leave recovery on).
email_verification — pending opt-in verifications
| Field | Contents |
|---|---|
key_id | which key asked to verify an email |
email | the address awaiting confirmation |
issued_at | when the verification link was issued |
One row per key, overwritten by a fresh request; it holds a pending (not yet confirmed) address until confirmation moves it to api_key.verified_email and the row is cleared.
trial — the free-trial counter
| Field | Contents |
|---|---|
identity | either key:<id> or ip:<salted SHA-256 hash> — never a raw IP address |
used | how many free calls have been consumed |
first_seen, last_seen | timestamps |
A row is written only while the trial is running — at most 20 times per identity, ever — and never for a caller who is already contributing.
usage — the rolling quota counter
| Field | Contents |
|---|---|
identity | key:<id> for quota, or newkey:<salted SHA-256 hash> for the key-creation limit |
day | unix day, floor(ts / 86400) |
calls | count for that day |
Note the second form: the daily key-creation limit is also counted against a salted hash of your IP address, under the prefix newkey:. Same salt, same construction, same reasoning as the trial counter.
crawler_hit — how often an AI crawler fetched a page
This is the only store derived from a request rather than from a batch-job submission, and it is deliberately aggregate: one row per crawler, per path, per day.
| Field | Contents |
|---|---|
agent | the declared AI-crawler name, e.g. GPTBot, ClaudeBot, PerplexityBot — the matched token, not the raw User-Agent header |
path | the request path, without query string or fragment |
day | unix day, floor(ts / 86400) |
hits | how many times that crawler fetched that path that day |
Nothing here identifies a person. A row is only ever written when the User-Agent matches one of a fixed list of AI crawlers (robots, not visitors); a human visitor, a browser, or an ordinary bot writes no row. We store no IP address, no raw User-Agent, and no per-request row — only the running count per agent, per path, per day. It exists so we can see which AI crawlers index which pages (the leading indicator of AI-search visibility), and it holds only counts and crawler names.
Your browser
The dashboard stores exactly one item in localStorage, under the key bw.spend: the monthly spend figure and batchable-share percentage you typed into the savings calculator. It stays on your device and is never sent to us. There is no other client-side storage. We set no cookies.
3. How the IP hash works, and why the salt is mandatory
identity = "ip:" + SHA256(TRIAL_SALT + "|" + client_ip)
The client IP is read only from Cloudflare's CF-Connecting-IP header, which Cloudflare writes at the edge and which a caller cannot forge. X-Forwarded-For is deliberately not read: Cloudflare appends to that header, so a caller's own value would come first, and reading it would let anyone reset their trial with a header flag.
TRIAL_SALT is a Cloudflare secret. It is not in the repository, not in wrangler.toml, and not in any deployment config — deliberately, and with a comment explaining that wrangler deploy will happily overwrite a secret with a [vars] entry of the same name, which is how this went wrong once before.
Why the salt is not optional. An unsalted SHA-256 of an IPv4 address is effectively clear text: there are only about four billion of them, so the whole space can be enumerated on a laptop in an afternoon. A hash you can reverse by brute force is not a protective measure — it is the address with extra steps.
What happens without the salt. Keyless callers get no free trial at all. The service does not silently fall back to storing hashes that pretend to be anonymous. This is a deliberate fail-closed choice in src/trial.js, and it is the correct thing to fail loudly on.
4. Is a salted IP hash personal data?
Our position: yes. We treat it as pseudonymised personal data, not anonymous data, and it is in scope of the GDPR.
The reasoning:
- An IP address is personal data (Breyer, C-582/14), at least where the controller has means reasonably likely to be used to identify the subscriber.
- Hashing does not change that on its own. The mapping is deterministic: we hold the salt, so given a candidate IP address we can compute the hash and confirm a match. That is exactly the "singling out and verification" capability that the EDPB and the former Article 29 Working Party (WP216 on anonymisation) treat as leaving data pseudonymised rather than anonymised.
- Recital 26 asks whether identification is reasonably likely. Because we hold the salt, for us it is trivially likely. The salt reduces the risk for anyone who steals the database without it — which is real and worth having — but it does not put the data outside the Regulation for us.
Two honest consequences of taking that position:
- The hashes are subject to your rights — access, erasure, objection. §7 says what we can actually do about that.
- Destroying the salt would genuinely anonymise the existing hashes, at least as regards re-identification from an IP address. If the trial mechanism is ever retired, rotating or destroying
TRIAL_SALTis a real and cheap privacy improvement, not a gesture. It would also irretrievably reset every trial counter, which is the intended effect.
The country code (region) is a two-letter code. On its own it does not identify anyone. Combined with a key_id it is one more attribute of a key holder, and it is treated as part of that record.
5. Legal basis, per category
| Data | Purpose | Legal basis |
|---|---|---|
call rows submitted with a key | Building and publishing the aggregate; giving you your percentile back; contributor status | Art. 6(1)(b) performance of the arrangement you entered into by submitting, and Art. 6(1)(f) legitimate interest in operating and publishing a measurement dataset |
call rows submitted anonymously | Same | Art. 6(1)(f). In practice these rows are not attributable to any person by us at all — see §7 |
api_key.token_hash, label, tier | Authenticating you; deciding your tier | Art. 6(1)(b) |
trial.identity (ip: hash) | Enforcing the 20-call free trial | Art. 6(1)(f) — we have a legitimate interest in a trial that cannot be reset for free, and hashing with a secret salt is the least intrusive way we found to do it. The alternative was storing the address |
usage.identity (newkey: hash) | The 5-keys-per-IP-per-day noise limit | Art. 6(1)(f) — preventing runaway key creation |
usage.identity (key: form) | Enforcing the rolling weekly quota | Art. 6(1)(b) |
region country code | Understanding regional queue behaviour | Art. 6(1)(f) |
api_key.recovery_email, recovery_email_at | Re-issuing a paid key to the person who bought it, when the one-time display was lost at checkout | Art. 6(1)(b) — a paid key we cannot re-issue is a product we cannot deliver, so holding the buyer's address is part of performing the contract they entered into. Paid keys only, captured at checkout, never for free or contributor keys |
api_key.recovery_sent_at | Rate-limiting recovery emails per address | Art. 6(1)(f) — the recover route is necessarily unauthenticated (the caller has lost their only credential), so without a cooldown anyone knowing a customer's address could use it to flood their inbox |
usage.identity (recover: hash) | The daily per-caller ceiling on recovery requests | Art. 6(1)(f) — preventing an unauthenticated route from being used to send mail at volume. Same salted-hash construction as trial.identity; the address itself is never written |
On the legitimate-interest balancing (Art. 6(1)(f)): the data is metadata about machine behaviour, not about a person's activity; we receive no content; the IP is hashed with a secret salt and the address itself is never written; the retention is bounded by the purpose; and there is no profiling of individuals. A formal Legitimate Interests Assessment should still be written up before launch — this paragraph is a sketch of one, not a substitute for it.
No cookies are set, so ePrivacy Art. 5(3) consent is not engaged for the API. The one localStorage item (bw.spend) holds values you typed into a calculator on your own screen and is read only to put them back in the fields; we consider that strictly necessary for functionality you explicitly requested.
6. Retention
Stated plainly: the code implements no automatic deletion of anything. There is no TTL, no scheduled purge and no DELETE statement anywhere in the service. This was verified by reading every query in src/.
What that means in practice:
- Published figures are computed over rolling windows — 30 days for distributions, 1 hour for current conditions, 7 days for contribution and quota. Rows outside those windows stop influencing any answer.
- But they are retained, because the historical record is the asset.
trialandusagerows are small and bounded (at most 20 writes per trial identity, ever) and are likewise retained.
Before launch, a retention period must be set and implemented — at minimum for trial.identity and usage.identity, which are the rows that contain hashed personal data and which serve no purpose once their window has passed. A trial row whose last_seen is a year old is doing nothing except existing. This is an open item, not a policy.
7. Your rights, and what we can honestly do
You have the rights in Chapter III of the GDPR: access, rectification, erasure, restriction, objection and portability. Write to hello@batchwatch.dev. There is no self-service route for any of this except key revocation; requests are handled manually by the operator.
To act on a request we need to be able to tie the data to you. In practice that means authenticating with the API key in question — that is the only identifier that links rows to a submitter.
What we can delete or exclude
| Data | Can we? | How |
|---|---|---|
| Your API key | Yes, immediately, yourself. DELETE /v1/keys/current | Sets revoked_at; the key stops authenticating |
The api_key row entirely | Yes, on request | Manual database operation |
Your trial row (the IP hash) | Yes, on request | Manual database operation |
Your usage rows | Yes, on request | Manual database operation |
Your call rows | Yes, immediately, yourself. DELETE /v1/calls/mine sets excluded = 1 on every row from your key | See the correction below — it does not yet reach every published surface |
Physical deletion of call rows | Yes, on request | Manual database operation |
Correction, 25 August 2026. An earlier version of this section said that setting excluded = 1 removes the rows from every published figure "on the very next request", and that this "was verified across all of them". That was checked by reading the queries, not by running them, and it is wrong. It was re-tested against a running worker on 25 August 2026, and the result is:
- Every query that reads the
calltable directly does filterexcluded = 0./v1/coverageand/v1/patternswent empty immediately. That part was right. - But the precomputed
rolluptable is not re-filtered. It stores figures computed before the exclusion./v1/curvekept serving all 28 test measurements, unchanged, after the purge. - And the public model page kept serving them permanently. Once every row for a model is excluded, that model drops out of the rollup-refresh candidate list entirely, so its stale row is never recomputed. The page at
/m/<provider>/<model>reads that row without any freshness check and returned HTTP 200 with the full curve after a forced refresh reportedmodels: 0. On the live site an HTTP cache (s-maxage=1800, stale-while-revalidate=86400) sits on top of that.
So today, DELETE /v1/calls/mine removes your measurements from the API but not from the indexed public pages. Until that is fixed, an erasure request must be handled manually by the operator, who must also delete the corresponding rollup rows. The fix is described in docs/privacy-wiring.md §3 and is an open item below.
Note separately that exclusion is suppression, not erasure. The row still exists. For an Art. 17 erasure request, exclusion is not enough on its own; the row has to actually go, and that has no API route — it is a manual operation by the operator against the D1 database.
What we cannot do, and why
- We cannot delete anonymous submissions on request. If you submitted without a key,
key_idis NULL and there is nothing in the row that links it to you. We are not able to identify which rows are yours, and we will not guess. Under Art. 11 GDPR, where we cannot identify the data subject, the erasure right does not oblige us to acquire additional information purely to be able to comply. If you may want your submissions removable later, submit them with a key.
- We cannot recall answers already given. A percentile that was computed and served to a caller last week cannot be un-served. Removing your rows changes future answers, not past ones.
- We cannot retract aggregates already published or exported. Once a figure is on the dashboard, in someone's screenshot, or in a third party's cache, it is out. The figures are aggregates across contributors, so this is not a disclosure of your submissions — but it is a limit on what deletion achieves.
- We cannot recover your API token. We only have the hash. This is a feature.
The conflict you should know about
There are, today, two inconsistent statements about deletion, and this document will not pretend otherwise:
DELETE /v1/keys/currentreturns: "Measurements you already sent stay in the dataset — they were true when they were taken, and a record that can be deleted backwards is not a record." That is the product philosophy, and it is a good one for a measurement dataset.- The front page promises: "The right to delete everything you sent. We remove it from the aggregate on the next rebuild."
These do not say the same thing. Our position, pending a decision: key revocation alone does not delete measurements, and the API's wording is correct about that. A separate, explicit erasure request naming your key will be honoured — we will delete or exclude those call rows.
The API response text and the front-page wording need to be reconciled before launch. This is flagged as an open item, not resolved here.
8. Aggregation and disclosure risk
Published figures are aggregates. We never publish an individual submission.
However, an aggregate over a very thin bucket can leak. The old design intent — stated in wrangler.toml — was that a bucket should not be published below MIN_N measurements (20) from MIN_KEYS distinct keys (3), "otherwise you could read one competitor's usage out of a thin bucket". That threshold is enforced on /v1/distribution and nowhere else.
As of 25 August 2026 this was re-examined by measurement, not by reading, and the honest position is set out below. The full test, route by route, is in docs/privacy-wiring.md §2; the rule we intend to adopt is in src/privacy.js.
What a reader can actually learn today
A test dataset of 28 measurements from a single key, submitted through the ordinary ingest route, was queried through every public surface. With no API key at all, a reader could obtain:
- A weekday profile from
/v1/patterns: medians for Monday through Friday and nothing on Saturday or Sunday. That is a working calendar. - Every individual measurement from
/v1/curve. The curve is downsampled to at most 60 points, so below 60 measurements the "points" are the rows themselves, one per job, sorted. - A public, indexed HTML page at
/m/<provider>/<model>carrying the same curve, and asitemap.xmlentry actively submitting that URL to search engines.
/v1/wait returned percentiles with contributors: 1 and confidence: very_low, which we consider acceptable — a queue median is a property of the provider, not of the contributor.
The part that is not about thresholds at all
model is free text that the submitter sends, capped only at 100 characters. A fine-tune identifier such as ft:gpt-4o-mini:acme-corp::9xYz names its owner, and we build a public page per model and announce it in sitemap.xml. No contributor threshold protects against that: the disclosure is in the key of the aggregate, not in the value.
What we are changing
The rule we are adopting, and the reasoning, are in src/privacy.js:
- Existence, counts and percentiles may be published from a single source, with the source count shown.
- Anything showing shape (curves, individual durations, maxima) or timing (weekday, hour of day, job size) requires at least two contributors other than our own prober.
- A dataset consisting entirely of our own prober's jobs may be published in full. There is no third party to protect, and we do not hide behind a threshold to avoid saying so.
- Model names we cannot recognise as public are kept out of
sitemap.xmland out of enumerations, and get no indexed page. They are still answered for by the API if you already know the name.
Until that is wired into the routes, the position stated in the previous version of this section still holds and you should act on it: while a model has a single contributor, that contributor's queue-time distribution, weekday pattern and job-size mix are effectively public. If your batch timings are commercially sensitive, that is a reason to weigh before contributing — and a reason not to put a private model name in the model field.
9. Who else processes the data
Cloudflare, Inc. is our sole sub-processor. The service is a Cloudflare Worker with a Cloudflare D1 database, on Cloudflare's network. Cloudflare processes the request itself (including your IP address, at the edge, before it reaches our code) and stores the database.
The D1 region is not pinned in our configuration. wrangler.toml declares the database binding and id but no location_hint; the primary region was chosen by Cloudflare when the database was created. This must be established and stated here before publication — see DATA-PROCESSING.md §5. Do not assert an EU region in this document until it has been checked in the Cloudflare dashboard.
Cloudflare is a US company. Transfers are covered by Cloudflare's own Data Processing Addendum and Standard Contractual Clauses, which must be executed and referenced before launch.
Our prober submits tiny real jobs (8 tokens in, 5 out) to LLM providers using our own accounts. No user data of any kind is sent to those providers. We never see, hold or transmit your provider credentials.
9a. Analytics — only if you say yes
We use Google Analytics 4 (property G-SLNPNGC0EB) on the public website, and only after you have actively allowed it. Nothing is loaded before that:
- On your first visit a banner asks. Until you choose, no analytics script is fetched, no analytics cookie is set, and no request is made to Google.
- Decline and nothing loads, on that visit or any later one. Declining is a real choice, not a delay — it is remembered in your browser.
- Allow and the Google Analytics script is loaded from
googletagmanager.com, which is a third-party request and can see your IP address as any HTTP request does. IP anonymisation is on, and we do not send advertising signals. - You can change your mind at any time from the "analytics" link on the page; withdrawing consent stops any further collection.
We use no advertising, no tag manager, and no other third-party script. The API (/v1/…) loads and reports nothing to anyone — analytics exist on the marketing pages only, never on the measurement routes.
This section previously read "We do not use analytics, advertising, tag managers, or any third-party script." That was written before analytics were added and was not true of the running site. It was corrected on 2026-08-26. If you relied on the earlier wording, this paragraph is what the site actually does, and it did so on consent only.
10. Logging
The only application logging in the service is a one-line summary of each prober run: how many jobs were polled, finished, abandoned and submitted, plus any error strings. It contains no user data.
We also keep a per-day count of how often each AI crawler fetched each page (crawler_hit, §2). That is a store of counts and crawler names, not a request log: it records no IP address, no raw User-Agent, and no per-request row, and it is only ever written for a recognised crawler — never for a human visitor.
Cloudflare's own edge logging and analytics are Cloudflare's processing and are governed by their documentation and DPA. Assume request-level metadata, including IP addresses, exists there for Cloudflare's retention period.
11. Automated decision-making
The service returns a verdict (run_batch, run_sync) about a queue. It makes no decision about a person, and produces no legal or similarly significant effect on anyone within the meaning of Art. 22.
12. Children
The service is not directed at children and is not usable in any meaningful way without an LLM provider account.
13. Complaints
You may complain to your supervisory authority. In Denmark that is Datatilsynet, Carl Jacobsens Vej 35, 2500 Valby, dt@datatilsynet.dk.
14. Changes to this policy
Material changes will be announced on the dashboard and in this repository, with the same 30 days' notice we commit to for terms changes that reduce what contributors get.
localStorage item.
- Whether contributors should have to warrant that they are permitted to measure on the provider account they submit from. See
docs/compliance.md§7.