How to sell lab and chemistry data to AI companies
How labs and chemical companies can license LIMS, experimental, formulation and batch data to AI labs, why failed experiments sell, and how to protect IP.
By Databounties Editorial Team · Updated 6 October 2026 · 6 min read

Key points
- AI for scientific discovery is limited by a lack of real experimental data, especially negative results.
- Valuable data includes LIMS records, lab notebooks, assay and QC results, formulations, batch records and spectra.
- Failed and abandoned experiments rarely get published, which makes them unusually valuable.
- Licences can be scoped to protect core IP, for example by excluding active programmes or using field-limited terms.
Scientific AI is one of the fastest-growing areas of research, from materials discovery to process optimisation. Its biggest constraint is data: real experimental records with all their noise, failures and context. Contract labs, chemical manufacturers, formulators and life-science companies hold years of it, often sitting unused in a LIMS or on shared drives.
What buyers want
- LIMS records: samples, tests, results, specifications and out-of-spec events.
- Electronic lab notebooks: procedures, observations and conclusions.
- Assay and QC data, including stability studies.
- Formulations and recipes, linked to the properties they produced.
- Batch and process records: parameters, deviations, yields.
- Instrument data: spectra, chromatograms, microscopy.
- SOPs, method validations and investigation reports.
Why negative results are worth so much
Journals publish what worked. Your notebooks record everything: the routes that failed, the batches that went out of spec, the formulations that separated. That's the data AI needs to learn what doesn't work, and almost nobody else has it.
Protecting your IP
- Scope carefully: exclude active programmes and anything under patent prosecution.
- Start with legacy data: discontinued products and older projects often carry little competitive risk.
- Generalise sensitive details: parameter ranges instead of exact set points, coded compound identifiers.
- Limit use in the licence: training only, no use to build competing products, field-limited exclusivity.
Read more in AI data licensing agreements.
Client and regulatory considerations
Contract research organisations need to check client agreements, since results produced for a client usually belong to that client. Regulated data (GMP, clinical) may have further restrictions. Personal data is usually limited to analyst names and signatures, which should be removed.
Sources
- 1.Negative results are disappearing from most disciplines and countries
Fanelli, Scientometrics, 2012
- 2.Scaling deep learning for materials discovery
Merchant et al., Nature, 2023
- 3.Guidance on GxP data integrity
Medicines and Healthcare products Regulatory Agency (MHRA), 2018
Frequently asked questions
Why would AI companies want our failed experiments?
Published literature is heavily biased towards successes. Models trained only on successes don't learn what doesn't work. Records of negative results and abandoned routes are rare, which is exactly why they're valuable.
Will licensing our data expose our IP?
It doesn't have to. You decide what's in scope: you can exclude active programmes, use only older data, generalise sensitive parameters and restrict use to model training in the licence. Many sellers start with legacy or discontinued projects.
Does our data need to be perfectly structured?
No. Structured LIMS data is easier to work with, but notebooks, PDF reports and instrument outputs are valuable too. Structuring them is part of the preparation work.


