AI labs pay most for data they can't scrape from the web or generate with their own models: real records of decisions people made inside working businesses, linked to what happened next. Long history, steady fresh records and clear rights to license them raise the price further.
Labs already have the public internet. Epoch AI projects that, if trends continue, language models will fully use the stock of human-generated public text between 2026 and 2032. Text that models wrote themselves teaches them little new, because it is mostly what a model already produced. That leaves private records of real work as one of the few sources of new signal.
Decisions
A record that shows a choice is worth more than a record that shows a state. An invoice tells a lab what was billed. The audit trail around it shows that someone flagged it, chased the client twice and wrote off part of the balance. That sequence teaches a model how an accountant handles a late payer, which is the job labs are training AI to do.
Decisions show up in many systems:
- Field changes in a CRM as a deal moves through stages
- Adjustments and reclassifications in a general ledger
- Status transitions on a support ticket or Jira issue
- Revisions in an electronic lab notebook
- A claim moving from review to approval, denial or appeal
Software teams record decisions in plans and briefs: what the team chose to build, the options it turned down and why. See our guide on selling software work history.
Outcomes
Labs want to know how things turned out. Did the deal close? Was the invoice paid? Did the experiment produce the compound? Did the change ship, or was it rolled back? An outcome lets a model learn which actions worked.
Agent training puts a price on this directly. Labs buy tasks where a grader checks whether the model succeeded, and Epoch AI reports those tasks commonly cost $200 to $2,000 each. A history of real situations with known outcomes is raw material for the same kind of product. We package it as episodes: situation, human action and outcome, in order.
Failures
Published sources overrepresent success. Papers report the reaction that worked, and case studies describe the project that shipped. Records of failure are scarce, and they teach a model where the boundaries are.
Chemistry has the clearest evidence. In a 2016 Nature paper, researchers trained a model on failed "dark reactions" taken from lab notebooks. It outperformed traditional human strategies and predicted conditions for new products with an 89% success rate. A 2022 study in Angewandte Chemie found that removing negative results changed what reaction-prediction models concluded. Our guide on selling lab notebook data covers this in more depth.
History
Years of records show how practice changed, how rare cases were handled and what happened over the long run. A rare event might appear once a year, so a ten-year history has ten examples where a one-year history has one. Our calculator weights years of history for this reason, with diminishing returns after the first several years.
Freshness
Old data describes old practice. Tax rules change, software changes and new kinds of cases appear. A dataset that keeps getting new records stays useful to a lab long after a frozen archive goes stale. That is the basis of recurring data revenue: each month of new records is a new batch that a buyer with a refresh term pays for.
Rarity
Data a lab can get elsewhere sells for less. Standard SaaS fields, templates and boilerplate look the same at every company. Records from specialized work are harder to find: semiconductor yield histories, insurance claims files, contract chemistry notebooks, legal work product. The narrower and more expert the work, the fewer other sellers there are.
Structure and provenance
Labs need to know where data came from and what each field means. A clean dataset includes:
- The source system, its version and the date range
- A description of every field
- A record of every change made during de-identification
- A statement that the seller has the right to license it
Missing provenance slows diligence and can stop a deal. We attach it to every dataset we deliver.
Clear rights
Rights are the floor under every other factor. Data that a seller can't prove it may license is close to worthless, because a lab can't risk training on it. If you hold data on behalf of customers, read can I sell data I hold for my customers? before anything else.
What lowers value
- Mostly boilerplate or default fields
- Free text full of personal details that has to be removed
- Short history or long gaps
- Data copied from public sources the lab already has
- Model-written content, such as AI agent output
- Unclear or restricted rights
Check how your data scores
Our calculator turns your industry, systems, years of records, team size and rights into an estimated range in about a minute. Value your data, or read how one source becomes three products in episodes, records and environments.
Frequently asked questions
What kind of data do AI companies want most?
They want records of real work that they can't scrape or generate: decisions people made in business systems, linked to outcomes, over several years. Audit trails, revision histories and status changes are especially useful because they show the steps.
Is AI-generated content valuable to labs?
Generally not. Content written by a model is mostly what a frontier model already produced, so training on it teaches another model little. The human parts around it, such as the goal, the corrections and whether the result was accepted, are worth more.
Does more data always mean more value?
No. A smaller dataset with decisions, outcomes and a clean audit trail can be worth more than a large export of default fields. Volume multiplies value that is already there.
Why do failed experiments or lost deals matter?
They show a model where the boundaries are. Positive-only data teaches what worked without teaching what separates success from failure, which is why research on chemistry models found negative results changed their conclusions.