The scrubber reduces the risk that a record can be traced to a person or a company. It does not make data anonymous in the legal sense under GDPR, UK GDPR or CCPA, and no tool can. A person reviews the output and the report before anything goes to a buyer.
Column profiling is heuristic
Roles come from header names and value patterns. A column with an odd header and mixed values can be classified wrongly. We read the role table for every new source and fix mistakes with overrides in the config.
Free text is the weakest point
The rules do not catch:
- names of people who appear in no identifier column and have no honorific or greeting, such as "spoke to Dave in finance"
- addresses written differently from the address column
- partial company names, such as "Halden" for "Halden Chemicals Ltd"
- identifiers in text that is not in English
- secrets in formats the scrubber does not know, because it has no entropy-based detection
The optional entity detector catches many of the names the rules miss, but models miss names too. For any free-text column nobody has read, we drop it with "freeText": "drop" or read a sample by hand before it ships.
Pseudonyms are linkable by design
Every record about one person or company keeps the same token, so joins keep working and a profile of that token can also be built from the records. Anyone with the salt can confirm a guess by hashing a candidate value, so the salt stays with you. Tokens are 48 bits long, so collisions are negligible below millions of distinct values.
k-anonymity has a narrow scope
The re-identification check runs per file, over the chosen quasi-identifier columns only. It does not account for:
- attributes outside the quasi-identifier set
- joining several delivered files on their shared tokens
- outside data a buyer may already hold
Small samples have small groups, so their k is lower. Company columns are pseudonymized and left out of the check, but a company can still be singled out from its amounts, dates and volumes.
Shifted dates keep patterns
In date shift mode, intervals within one entity are preserved on purpose. Seasonality and exact gaps between events can still identify that entity.
Format details
Dates come out in ISO form (2024-03 or 2024-03-15). Nested JSON values are treated as text and written back as JSON strings. Masked examples in the report reveal one character.
Removing a secret does not revoke it
The scrubber blanks secrets and lists them in the report. The credential is still valid in the system it came from until you rotate it.
Legal questions are out of scope
The scrubber does not check rights, consent, contracts, export control or sector rules. We handle those in the rights review, with your counsel where the data is regulated.
Trained models cannot forget
If you withdraw a dataset, we stop selling it. Data a buyer already used to train a model cannot be pulled back out of that model.