How to anonymise business data before selling it to AI companies
How to clean and anonymise company data for AI training: what to remove, techniques that work, free-text risks and re-identification testing.
By Databounties Editorial Team · Updated 6 October 2026 · 7 min read

Key points
- Anonymisation means individuals can no longer be identified by any means reasonably likely to be used, not just that names have been removed.
- Direct identifiers (names, emails, account numbers) are easy to remove. Indirect identifiers and free text are where most risk lies.
- Common techniques include removal, consistent tokenisation, generalisation, date shifting and aggregation.
- Test for re-identification risk before sharing anything, and keep a written record of your method.
Anonymisation is what makes most company data sellable. Done well, it removes the privacy risk and keeps nearly all of the value. Done badly (names deleted, everything else left in), it leaves you exposed and buyers will spot it straight away. Here's how to do it properly.
What counts as "anonymous"
Under UK and EU GDPR, data is anonymous when individuals can no longer be identified, directly or indirectly, by any means reasonably likely to be used. That's a higher bar than removing names, and it's what takes the data outside GDPR. (Why that matters is covered in is it legal to sell company data?)
Step 1: Map where personal data lives
- Direct identifiers: names, emails, phone numbers, addresses, account and card numbers, national IDs, employee IDs.
- Indirect identifiers: job titles, exact dates, postcodes, rare values, small groups.
- Free text: notes, descriptions, ticket bodies, email content, comments.
- Attachments and metadata: PDFs, images, document authors, file paths.
Step 2: Choose the right technique for each field
| Technique | What it does | Use it for |
|---|---|---|
| Removal | Deletes the field entirely | Fields with no training value (phone numbers, bank details) |
| Consistent tokenisation | Replaces a value with the same random token everywhere, with no key kept | Customer, supplier and employee names, where links between records matter |
| Generalisation | Reduces precision | Full postcode to area, exact age to a band, title to a role family |
| Date shifting | Moves dates by a consistent random offset | Timelines where intervals matter more than exact dates |
| Aggregation | Groups small populations | Rare categories that would single someone out |
| Free-text redaction | Detects and replaces identifiers in text | Notes, tickets, emails, reports |
Destroy any mapping key. If you can reverse the tokens, the data is only pseudonymised and still personal data.
Step 3: Deal with free text carefully
Free text is where anonymisation most often fails, and it's often the most valuable part of the dataset. Use automated entity detection to find names, organisations, locations, emails and numbers. Then sample the output and review it by hand, especially signatures, greetings and pasted email threads.
Step 4: Remove commercially sensitive information too
Privacy isn't the only concern. Decide what you don't want to share: client names, specific pricing, trade secrets, security details. These can be tokenised or removed just like personal data.
Step 5: Test for re-identification
- Look for rare combinations of field values that could single someone out.
- Try to re-identify a sample using only the data and public information (a "motivated intruder" test).
- Check that free-text redaction has worked on a random sample.
- Write down the method, the tests and the results. Buyers and regulators will ask for them.
Step 6: Back it up in the licence
Your licence should ban re-identification and linking with other data to identify individuals. See what to include in an AI data licensing agreement.
Sources
- 1.Opinion 05/2014 on Anonymisation Techniques (WP216)
Article 29 Data Protection Working Party, 2014
- 2.De-Identification of Personal Information (NISTIR 8053)
National Institute of Standards and Technology, 2015
- 3.Anonymisation guidance
Information Commissioner's Office (ICO)
Frequently asked questions
Is removing names enough to anonymise data?
No. People can be identified from combinations of other fields (job title plus location plus date, for example) and from free-text notes. Genuine anonymisation deals with direct identifiers, indirect identifiers and free text.
Will anonymisation make my data less valuable?
Usually only slightly. AI labs want patterns, workflows and outcomes, not identities. Replacing names with consistent tokens keeps the relationships between records, which is what holds most of the value.
Can AI tools anonymise free text?
They help. Named-entity recognition and LLM-based redaction catch most names, addresses and identifiers in free text, but they aren't perfect. Combine automated redaction with sampling and manual review.


