Docs/Privacy and scrubbing

Limits

What the scrubber does not catch, what k-anonymity does not measure, and what we check by hand instead.

The scrubber reduces the risk that a record can be traced to a person or a company. It does not make data anonymous in the legal sense under GDPR, UK GDPR or CCPA, and no tool can. A person reviews the output and the report before anything goes to a buyer.

Column profiling is heuristic

Roles come from header names and value patterns. A column with an odd header and mixed values can be classified wrongly. We read the role table for every new source and fix mistakes with overrides in the config.

Free text is the weakest point

The rules do not catch:

  • names of people who appear in no identifier column and have no honorific or greeting, such as "spoke to Dave in finance"
  • addresses written differently from the address column
  • partial company names, such as "Halden" for "Halden Chemicals Ltd"
  • identifiers in text that is not in English
  • secrets in formats the scrubber does not know, because it has no entropy-based detection

The optional entity detector catches many of the names the rules miss, but models miss names too. For any free-text column nobody has read, we drop it with "freeText": "drop" or read a sample by hand before it ships.

Pseudonyms are linkable by design

Every record about one person or company keeps the same token, so joins keep working and a profile of that token can also be built from the records. Anyone with the salt can confirm a guess by hashing a candidate value, so the salt stays with you. Tokens are 48 bits long, so collisions are negligible below millions of distinct values.

k-anonymity has a narrow scope

The re-identification check runs per file, over the chosen quasi-identifier columns only. It does not account for:

  • attributes outside the quasi-identifier set
  • joining several delivered files on their shared tokens
  • outside data a buyer may already hold

Small samples have small groups, so their k is lower. Company columns are pseudonymized and left out of the check, but a company can still be singled out from its amounts, dates and volumes.

Shifted dates keep patterns

In date shift mode, intervals within one entity are preserved on purpose. Seasonality and exact gaps between events can still identify that entity.

Format details

Dates come out in ISO form (2024-03 or 2024-03-15). Nested JSON values are treated as text and written back as JSON strings. Masked examples in the report reveal one character.

Removing a secret does not revoke it

The scrubber blanks secrets and lists them in the report. The credential is still valid in the system it came from until you rotate it.

The scrubber does not check rights, consent, contracts, export control or sector rules. We handle those in the rights review, with your counsel where the data is regulated.

Trained models cannot forget

If you withdraw a dataset, we stop selling it. Data a buyer already used to train a model cannot be pulled back out of that model.