AI labs buy company data in three forms: records, which are de-identified rows from a business system; episodes, which show a situation, the action a person took and what happened next; and environments, which are working replicas of a system that an AI agent can practice in. The same company data can be packaged all three ways, and episodes and environments usually sell at task and environment prices, which are higher per unit of effort than a one-off dataset.
Why the form matters
A lab that trains on text learns what documents look like. A lab that trains agents to do work needs more: what a person saw, what they did about it, and whether it worked. Public text can't supply that, and Epoch AI projects that language models will fully use the stock of human-generated public text somewhere between 2026 and 2032 if current trends continue.
Company systems hold exactly the missing piece, because business software records decisions and their results as a side effect of normal work. The packaging decides which piece of that a lab can use.
Records
A record is one row from a business system after scrubbing: an invoice, a lab experiment, a support ticket, a claim. A set of records from one domain is a corpus.
Records show labs what real work in a domain looks like. A real chart of accounts across many client firms, real reaction conditions with real yields, or real ticket categories with resolution times are all hard to fake, and a lab can't generate them from public sources.
Records are the simplest product to prepare. They come from an export or a read-only connector, go through the scrub, and are delivered with field notes and a scrub report. They are priced as dataset licenses, with a price for the history and, where the license includes a refresh term, a price for each monthly batch.
Episodes
An episode puts a person's actions in order: the situation they faced, each step they took and the outcome. Episodes are built from audit trails and event histories, such as Salesforce field history, the Xero audit log, Jira status transitions, ELN revisions and MES logs.
Our episodes are delivered as one JSON line each. Here is an illustrative example from a ticketing system, with every identifier replaced by a pseudonym and every date generalized to the month:
{
"episode_id": "EP_9C21",
"source": "jira",
"actor": "PERSON_14",
"started_at": "2026-03",
"steps": [
{ "at": "2026-03", "state": { "status": "open", "priority": "high" }, "action": "assigned to PERSON_07", "outcome": "in progress" },
{ "at": "2026-03", "state": { "status": "in progress" }, "action": "linked PR_A3", "outcome": "tests passed" },
{ "at": "2026-04", "state": { "status": "in review" }, "action": "merged", "outcome": "deployed" }
],
"outcome": "resolved, reopened once"
}
Labs value episodes because they look like training tasks. A task has a starting state, a goal and a way to check whether the goal was reached. An episode from a real company supplies the starting state and a known outcome, which gives a lab a way to check an agent's answer.
Prices for tasks are higher than prices for rows. Epoch AI interviewed people building reinforcement learning environments and reported that multiple interviewees put task costs at $200 to $2,000 each, with $20,000 per task rare and reserved for especially complex software engineering work. Those are prices for tasks built by experts; an episode from your systems is raw material for such a task, and its price depends on how much work a lab still has to do to turn it into one.
Episodes are also where history pays off. A system with years of audit logs holds thousands of situations and outcomes, and only systems that recorded changes as they happened can produce them.
Environments
An RL environment is a working replica of a system that an AI agent can act inside: a copy of an accounting package, a CRM or a ticket tracker with realistic data in it. The agent tries a task, the environment changes in response, and a grader checks the result.
Our environments copy the structure of a real system and fill it with synthetic data drawn from the real data's distributions. The buyer gets the shape of a real company's tables, workflows and edge cases without receiving the real company's records.
Environments are the most expensive form. Epoch AI reports, citing SemiAnalysis, that website replicas cost around $20,000 each, and one interviewee put higher-quality replicas of complex products like Slack at around $300,000. The same article reports that environment contracts are often six to seven figures per quarter, and that two founders independently estimated exclusive deals at roughly four to five times the price of non-exclusive ones.
Environments raise a question we still treat as open: how faithful a synthetic replica can be before it leaks the real company's data. That makes environments the product we offer most cautiously, and you review any environment before a buyer sees it.
Side by side
| Records | Episodes | Environments | |
|---|---|---|---|
| What it is | De-identified rows from one domain | Situation, human action and outcome, in order | A working replica filled with synthetic data |
| Built from | Exports and connectors | Audit trails and event histories | Schema and data distributions from the records |
| What a lab learns | What real work in a domain looks like | How people make decisions and what follows | How to act inside a real tool |
| Typical price signal | Dataset license, history plus refresh | Per task, like RL tasks | Per environment |
| Needs from the seller | An export or connector | Audit logs or event history | A system worth copying, with enough records |
Which one your data supports
Every system that can export rows supports records. Episodes need history: field history, audit logs or status transitions with timestamps. Environments need a system whose structure is worth copying and enough records to learn realistic distributions from.
When you prepare an export, include audit logs and event histories wherever the system offers them. They cost nothing extra to export and they decide whether your data can become episodes.
For software teams, the strongest episodes link a plan to the tasks, code changes and outcome that followed. Selling software work history covers how that chain gets recorded.
How each one pays over time
All three forms can recur. Records and episodes refresh naturally: the SDK pulls new records and new audit events each month, scrubs them on your machine and uploads a new batch, and each license with a refresh term pays per accepted batch. An environment can be refreshed when the underlying system changes. Recurring data revenue explains the refresh model.
Find out which products your data supports
The valuation calculator gives a range based on your systems and years of history, and on a 30-minute call we tell you which of the three products your data could support. For prices across industries, see what AI labs pay for data.
Frequently asked questions
Do I have to choose one product?
No. One connection to a system can produce records, episodes and, later, an environment. Each is licensed separately, so the same data can earn from more than one buyer under non-exclusive terms.
Are episodes the same as AI agent transcripts?
No. Episodes come from audit trails of human work in business systems. Transcripts of AI coding agents are mostly model output, which teaches a lab little, and AI providers' terms generally forbid using their outputs to build competing models. We confirm with a lawyer before collecting any transcripts.
Does an environment contain my real data?
No. It copies the structure of your system and fills it with synthetic data drawn from the real distributions. How closely a replica can follow real data without leaking it is still an open question for us, so you review any environment before it is offered.
Why are episodes worth more than records?
An episode has a starting state and a known outcome, which is what a lab needs to build a task it can grade. Records show what work looks like, while episodes show decisions and their results. Prices reported for RL tasks run from about $200 to $2,000 per task according to Epoch AI's interviews.