Docs/Privacy and scrubbing

How scrubbing works

How the scrubber classifies each column, replaces identifiers with pseudonyms, coarsens dates and amounts, cleans free text and strips secrets.

The scrubber runs on Bun with no runtime dependencies. On its own it reads CSV, JSON or JSON Lines files, and the SDK calls the same code on every run, with each object type (invoices, contacts, episode steps) profiled as its own table. Each run works in five steps: profile the columns, transform each cell by its column's role, check re-identification risk, check for leaks, and write the report.

Column roles

The scrubber reads each column's header and values and assigns one of five roles.

Role Examples What happens
Direct identifier Names, emails, phones, addresses, bank account numbers, IBANs, sort codes, tax IDs, card numbers, IP addresses, URLs and domains Replaced with a pseudonym
Quasi-identifier Dates, amounts, postcodes, company names, job titles, locations, and categories with rare values Generalized, and used in the re-identification check
Free text Notes, comments, descriptions, message bodies Redacted, or dropped
Safe Record IDs, common categories, numeric measurements Record IDs are re-keyed. Categories get a light scan for secrets, emails and known identifiers. Numbers pass through.
Drop Password, token, API key and secret columns Removed from the output

The header decides first. A column named customer_email is an email column whatever it contains. When the header says "date" or "amount" but the values disagree, or when the header matches no rule, the values decide: a column where most values look like email addresses is treated as emails.

Profiling is heuristic, so a column with an odd header and mixed values can be classified wrongly. Read the role table in the report for every new source and fix mistakes in the config:

json
{
  "columns": {
    "account_code": "safe",
    "owner": { "role": "direct", "kind": "name" },
    "lead_source": { "role": "safe", "kind": "category" }
  }
}

Pseudonyms

Direct identifiers and company names become tokens such as PERSON_3f9a1c2b7d4e or ORG_8b21e0c4a9f1. Each token is the first 12 hex characters of an HMAC-SHA256 over the normalized value, keyed with a secret salt. Normalizing first means ann@example.com and ANN@example.com get the same token, and phone numbers match regardless of spacing.

The same value always gets the same token, in every file of a run and in every run that uses the same salt. Joins keep working: an invoice and a CRM contact that refer to the same company still point at the same ORG_ token. Record IDs such as invoice, sample and contact numbers are re-keyed the same way with an ID_ prefix.

Keep the salt fixed across runs so pseudonyms in March's batch match those in February's. The SDK creates and stores one for you and refuses to upload without it; see Configuration. When you run the scrubber on its own, it reads the salt from SCRUB_SALT, or another variable named by saltEnv, and without one it picks a random salt whose pseudonyms will not match any later run. The salt is never written to any output.

Dates

Dates are generalized to the month by default, so 2024-03-15 becomes 2024-03. You can choose the year instead, keep dates as they are, or shift them.

In shift mode, every date belonging to one entity (one customer, say) moves by the same number of days, up to a maximum you set. The offset comes from the HMAC of that entity's value, so the intervals between one customer's events stay exact while the calendar dates change. Shift mode needs dates.entityColumn.

Ambiguous dates such as 01/02/2024 are read day-first or month-first according to dates.order. Use dmy for UK exports and mdy for US exports. A value with a day above 12 is read correctly either way.

Amounts

Amounts keep two significant figures by default, so 12,345.67 becomes 12000. You can instead round to a step, such as the nearest 100, bucket values into ranges like 1000-5000, or keep them as they are.

Postcodes and rare categories

UK postcodes are cut to the outward code (LS6 2AB becomes LS6) and US ZIP codes to their first three digits. In quasi-identifier category columns, such as job titles or cities, any value seen fewer than rareK times (five by default) becomes OTHER.

Free text

Free text is the hardest part, so it gets three passes in this order.

  1. Secrets. Private keys, credentials inside URLs, JWTs, AWS, GitHub, Stripe, Slack, Google and LLM API keys, bearer tokens, and password= style assignments become [SECRET].
  2. Known identifiers. Every value already seen in an identifier column, in any file of the run, is replaced with its token. A customer's name inside a notes field becomes that customer's ORG_ or PERSON_ token. Parts of known person names, such as a first name on its own, become [NAME].
  3. Patterns. Emails, URLs, IBANs, card numbers that pass the Luhn check, US, UK and EIN tax IDs, sort codes, phone numbers, IP addresses, UK postcodes, names after an honorific or greeting (Dr Chen, Dear Ann), dates of birth, and number-like IDs of five or more digits become placeholders such as [EMAIL] or [NAME].

The rules miss some things, such as a first name of someone who appears in no other column ("spoke to Dave in finance"). Set "freeText": "drop" to remove free-text columns entirely, which is the safer choice for any column nobody has read. See Limits.

Optional model stages

Two extra stages can run on top of the rules. Both are off unless you turn them on, and if one is turned on without its settings, the run records "not configured" in the report and sends nothing.

Column reviewer

The reviewer gives a second opinion on each column's role, using either Jev or any OpenAI-compatible endpoint, including a server on your own machine. It receives only the file and column names, the role the rules chose and why, counts, and three shape-masked samples per column, where every letter becomes a or A and every digit becomes 9. Priya Raman is sent as Aaaaa Aaaaa. It never receives a raw value.

The reviewer only suggests. The rules still decide every transform, and the output is identical with or without it. Every column where the reviewer disagrees with the rules is listed under Review needed in the report, for a person to settle with a columns override.

Entity detection in free text

An entity detector catches names and places the rules miss, such as the "Dave" in "spoke to Dave in finance". The supported adapter calls a Presidio Analyzer that you run yourself, for example in a container on the same machine. Each entity it finds is replaced with a placeholder such as [PERSON] or [LOCATION] after the rule pass.

The detector sees free-text cells from the rows being released, with secrets already stripped, so it does receive personal text. Run it locally, never as a hosted service. It lowers the free-text risk but does not remove it, because models miss names too.

Secrets in other columns

If a secret turns up in any column other than free text, the scrubber blanks that cell. Every secret it finds is listed in the report by row, column and type, never by value. Removing a secret from the output does not revoke it, so rotate anything the report lists.

Leak check

After transforming everything, the scrubber scans every output cell for any identifier value it saw in the input. If one survived, the run fails with exit code 2 and the report says do not release.

Reports

Each run writes the report three ways: report.json for machines (the SDK uploads this one with each batch), report.md for reading in a terminal or a pull request, and report.html, a self-contained page with no scripts or external assets.

Configuration reference

All keys are optional.

Key Default Meaning
columns none Role overrides by column header. A bare role string, or { "role", "kind" }.
kThreshold 5 Records whose quasi-identifier combination is shared by fewer than this many are flagged.
rareK same as kThreshold Category values seen fewer than this many times become OTHER.
dates.mode month month, year, shift or keep.
dates.order dmy How to read ambiguous dates: dmy or mdy.
dates.entityColumn none Column that keys the per-entity shift. Required for shift.
dates.maxShiftDays 30 Largest shift in days, in either direction.
amounts.mode sigfig sigfig, round, bucket or keep.
amounts.precision 2 Significant figures for sigfig, or the step for round.
amounts.edges 0, 100, 1000, 10000, 100000, 1000000 Bucket boundaries for bucket.
freeText redact redact or drop.
saltEnv SCRUB_SALT Environment variable that holds the salt.
quasiIdentifiers the generalized quasi-identifier columns Columns used in the re-identification check.