Databounties

How to anonymise business data before selling it to AI companies

How to clean and anonymise company data for AI training: what to remove, techniques that work, free-text risks and re-identification testing.

By Databounties Editorial Team · Updated 6 October 2026 · 7 min read

A hand inspecting a redacted document with a magnifying glass, representing anonymisation

Key points

  • Anonymisation means individuals can no longer be identified by any means reasonably likely to be used, not just that names have been removed.
  • Direct identifiers (names, emails, account numbers) are easy to remove. Indirect identifiers and free text are where most risk lies.
  • Common techniques include removal, consistent tokenisation, generalisation, date shifting and aggregation.
  • Test for re-identification risk before sharing anything, and keep a written record of your method.

Anonymisation is what makes most company data sellable. Done well, it removes the privacy risk and keeps nearly all of the value. Done badly (names deleted, everything else left in), it leaves you exposed and buyers will spot it straight away. Here's how to do it properly.

What counts as "anonymous"

Under UK and EU GDPR, data is anonymous when individuals can no longer be identified, directly or indirectly, by any means reasonably likely to be used. That's a higher bar than removing names, and it's what takes the data outside GDPR. (Why that matters is covered in is it legal to sell company data?)

Step 1: Map where personal data lives

  • Direct identifiers: names, emails, phone numbers, addresses, account and card numbers, national IDs, employee IDs.
  • Indirect identifiers: job titles, exact dates, postcodes, rare values, small groups.
  • Free text: notes, descriptions, ticket bodies, email content, comments.
  • Attachments and metadata: PDFs, images, document authors, file paths.

Step 2: Choose the right technique for each field

TechniqueWhat it doesUse it for
RemovalDeletes the field entirelyFields with no training value (phone numbers, bank details)
Consistent tokenisationReplaces a value with the same random token everywhere, with no key keptCustomer, supplier and employee names, where links between records matter
GeneralisationReduces precisionFull postcode to area, exact age to a band, title to a role family
Date shiftingMoves dates by a consistent random offsetTimelines where intervals matter more than exact dates
AggregationGroups small populationsRare categories that would single someone out
Free-text redactionDetects and replaces identifiers in textNotes, tickets, emails, reports

Destroy any mapping key. If you can reverse the tokens, the data is only pseudonymised and still personal data.

Step 3: Deal with free text carefully

Free text is where anonymisation most often fails, and it's often the most valuable part of the dataset. Use automated entity detection to find names, organisations, locations, emails and numbers. Then sample the output and review it by hand, especially signatures, greetings and pasted email threads.

Step 4: Remove commercially sensitive information too

Privacy isn't the only concern. Decide what you don't want to share: client names, specific pricing, trade secrets, security details. These can be tokenised or removed just like personal data.

Step 5: Test for re-identification

  • Look for rare combinations of field values that could single someone out.
  • Try to re-identify a sample using only the data and public information (a "motivated intruder" test).
  • Check that free-text redaction has worked on a random sample.
  • Write down the method, the tests and the results. Buyers and regulators will ask for them.

Step 6: Back it up in the licence

Your licence should ban re-identification and linking with other data to identify individuals. See what to include in an AI data licensing agreement.

Sources

  1. 1.
    Opinion 05/2014 on Anonymisation Techniques (WP216)

    Article 29 Data Protection Working Party, 2014

  2. 2.
    De-Identification of Personal Information (NISTIR 8053)

    National Institute of Standards and Technology, 2015

  3. 3.
    Anonymisation guidance

    Information Commissioner's Office (ICO)

Frequently asked questions

Is removing names enough to anonymise data?

No. People can be identified from combinations of other fields (job title plus location plus date, for example) and from free-text notes. Genuine anonymisation deals with direct identifiers, indirect identifiers and free text.

Will anonymisation make my data less valuable?

Usually only slightly. AI labs want patterns, workflows and outcomes, not identities. Replacing names with consistent tokens keeps the relationships between records, which is what holds most of the value.

Can AI tools anonymise free text?

They help. Named-entity recognition and LLM-based redaction catch most names, addresses and identifiers in free text, but they aren't perfect. Combine automated redaction with sampling and manual review.

Written by

Databounties Editorial Team

Data licensing and privacy

The Databounties team works with companies selling data to AI labs, handling valuation, anonymisation and licensing. Our guides draw on that work and on primary sources such as ICO and EU guidance, which are cited in every article.