Privacy and law

k-anonymity in plain English

A dataset is k-anonymous when every record looks identical to at least k-1 others on the fields an outsider could know. Here is how it works and where it fails.

·5 min read

A dataset is k-anonymous when every record shares its combination of identifying fields with at least k-1 other records, so no one can be narrowed down to fewer than k people. If k is 5, anyone trying to find a person in the data ends up with at least five equally likely matches.

The idea comes from Latanya Sweeney, who published the k-anonymity model in 2002. It is one of the simplest ways to test whether de-identified data can still point to someone. This page is general information, not legal advice.

The problem it solves

Removing names doesn't remove identity. Sweeney had earlier shown that 87% of the US population was likely unique on just ZIP code, gender and date of birth. None of those fields is a name, but together they work like one.

Fields like these are called quasi-identifiers. They are ordinary on their own and identifying in combination, and an outsider might already know them from a public source. k-anonymity asks one question about them: how many records share each combination?

A worked example

Here is a small, made-up table from a support system. Names are gone. The quasi-identifiers are age, city and job title. The last column is the sensitive value a buyer wants to learn from.

Age City Job title Ticket outcome
34 Manchester Head of Finance Refund issued
36 Manchester Finance Manager Refund refused
38 Manchester Finance Analyst Refund issued
52 Leeds Operations Director Escalated
57 Leeds Operations Manager Escalated
41 Bristol CTO Refund issued

Every row here is unique on its quasi-identifiers. Someone who knows a 41-year-old CTO in Bristol filed a ticket finds that person's row at once, so this table has k = 1.

Now generalize. Ages become 10-year bands, and job titles become broad roles:

Age City Role Ticket outcome
30-39 Manchester Finance Refund issued
30-39 Manchester Finance Refund refused
30-39 Manchester Finance Refund issued
50-59 Leeds Operations Escalated
50-59 Leeds Operations Escalated
40-49 Bristol Technology Refund issued

The Manchester group has three matching rows and the Leeds group has two. The Bristol row is still alone, so the whole table is only 1-anonymous. The fix is suppression: drop the Bristol row, or generalize city to region until it joins a group. With the Bristol row gone, the table is 2-anonymous, because the smallest group has two rows.

Those two moves, generalization and suppression, are how k-anonymity is reached in practice. Each one costs some detail, so the work is finding the least generalization that reaches the k you need.

Where k-anonymity fails

Look at the Leeds group again. Both rows say "Escalated". Anyone who knows a person in their fifties in operations in Leeds is in the data learns their ticket was escalated without knowing which row is theirs. The group hid the row and leaked the answer anyway.

This is the homogeneity problem. k-anonymity protects against finding a record. It says nothing about what the matching records reveal when they all agree.

A second gap is background knowledge. If an attacker already knows something that rules out some rows in a group, the effective group shrinks.

l-diversity fixes the homogeneity problem

l-diversity, proposed by Machanavajjhala and colleagues in 2007, adds a second rule: every group must contain at least l well-represented values of the sensitive field.

In the example, the Manchester group has two different outcomes, so it is 2-diverse. The Leeds group has one outcome, so it fails any l above 1. To fix it you would merge Leeds into a larger group with mixed outcomes, generalize further, or suppress those rows.

Other refinements exist. They all check the same thing: whether the people sharing a group also share too much else.

Choosing k

No law sets a single value of k. Higher k means more protection and less detail. The right number depends on how sensitive the data is, how many people or companies it covers, and what an outsider could plausibly know.

Business data needs care with small populations. A dataset of 40 client firms in one niche can't hide anyone well, because the firms themselves are few. In that case the safer route is to coarsen client attributes sharply or drop them and keep the behavior: what was invoiced, how it was paid, what happened next.

What k-anonymity doesn't prove

Passing a k-anonymity check is evidence, not a legal finding. The GDPR's Recital 26 asks whether identification is reasonably likely by any means, and a 2019 study estimated that 99.98% of Americans could be re-identified from 15 demographic attributes. Datasets with many columns are hard to make k-anonymous without losing most of their value, which is why free text and long attribute lists get dropped or reviewed.

Differential privacy is a different approach that adds calibrated noise to results. It suits statistics and aggregates better than the record-level data most AI buyers want.

How we use it

We treat k-anonymity as one check among several. Our scrub step removes direct identifiers, swaps clients and people for stable pseudonyms, generalizes amounts, dates and titles, and drops unreviewed free text. Then we test whether any record can be singled out from a combination of fields and remove any that can. The result goes into the scrub report you see before anything is offered to a buyer, and every buyer license bans re-identification. More on the full process is in de-identification vs anonymization and how to prepare a data export.

See what your records are worth

Generalized, tested data still sells, because labs care about the decisions and outcomes in it. Try the calculator, or read how licenses keep paying in recurring data revenue.

Frequently asked questions

What does the k in k-anonymity mean?

It is the minimum group size. In a 5-anonymous dataset, every combination of identifying fields appears in at least five records, so any one person hides among at least four others.

Is a k-anonymous dataset anonymous under the GDPR?

Not automatically. The GDPR asks whether anyone could reasonably identify a person by any means, and k-anonymity only checks the fields you chose as quasi-identifiers. It is strong evidence when the field choice and k are sensible, alongside other safeguards.

What is the difference between k-anonymity and l-diversity?

k-anonymity makes sure each record hides in a group. l-diversity also makes sure each group contains several different sensitive values, so the group can't give away the answer when every member shares it.

Does generalizing data make it worthless to AI labs?

Usually not. Labs want the sequence of actions and outcomes, which survives when ages become bands and amounts become ranges. The detail that gets removed is mostly the detail that identifies people.

Find out what your records are worth.

Value my data