Docs/For AI labs

Data formats

How records, episodes and environments are delivered, with the fields and conventions every batch follows.

Each dataset comes from one source system at one company and is sold as one product: records, episodes or an environment. Records and episodes arrive as monthly batches.

Batches

A batch is one scrubbed delivery of a dataset for one period, written as YYYY-MM. The first batch carries the history. Each batch after that holds the new records for its period, so a license with a refresh term receives a stream of monthly batches. Every batch comes with its record count, its schema and its scrub report.

Files are gzipped JSON Lines: one JSON object per line. The examples below are formatted across several lines for reading.

Records

A record is one row from the source system after scrubbing, such as an invoice, a deal, a ticket or a lab sample. Field names follow the source system, and each connector page lists them. Nested values are flattened into dotted names, such as CustomerRef.name. The _object field names the object type, such as invoice or invoice_line, since one batch can hold several.

json
{
  "_object": "invoice",
  "invoice_id": "ID_9c1e2a7f04b3",
  "customer": "ORG_8b21e0c4a9f1",
  "contact_email": "EMAIL_51d0c9e2a7b8",
  "invoice_date": "2024-03",
  "postcode": "LS6",
  "total": 1200,
  "status": "PAID",
  "description": "Monthly retainer for ORG_8b21e0c4a9f1, [NAME] approved"
}

The scrubber changes values in predictable ways:

Value In the delivered record
Names, emails, phones, addresses, account numbers, tax IDs, URLs, IPs Tokens such as PERSON_…, EMAIL_…, PHONE_…
Company names ORG_… tokens
Record IDs ID_… tokens
Dates Month (2024-03) by default; year, or shifted full dates in some datasets
Amounts Two significant figures by default; rounded or bucketed (1000-5000) in some datasets
Postcodes UK outward code (LS6) or ZIP3
Rare categories OTHER
Free text Placeholders such as [EMAIL], [NAME], [SECRET], or the field is absent

Tokens are consistent within a dataset. The same customer has the same ORG_ token in every record and every batch, so joins and per-entity histories work. The scrub report for each batch states the exact settings used.

Episodes

An episode is the ordered history of one thing in a system, such as a ticket, an issue, a pull request or an opportunity: the state it was in, what a person did, and what happened next. Episodes are built from the source's audit trail or change history.

json
{
  "episode_id": "EP_4a2f9e1b7c30",
  "source": "jira",
  "entity": "issue",
  "actor": "PERSON_3f9a1c2b7d4e",
  "started_at": "2026-09",
  "steps": [
    {
      "at": "2026-09",
      "actor": "PERSON_3f9a1c2b7d4e",
      "state": {
        "issue_type": "Bug",
        "project": "ENG"
      },
      "action": "created",
      "outcome": "created"
    },
    {
      "at": "2026-09",
      "actor": "PERSON_9d02b5e1c6a7",
      "state": {
        "issue_type": "Bug",
        "project": "ENG",
        "status": "To Do"
      },
      "action": "status: To Do -> In Progress",
      "outcome": "status = In Progress"
    },
    {
      "at": "2026-09",
      "actor": "PERSON_9d02b5e1c6a7",
      "state": {
        "issue_type": "Bug",
        "project": "ENG",
        "status": "In Progress"
      },
      "action": "commented: Fixed in the retry handler, [NAME] to verify",
      "outcome": "comment added"
    }
  ],
  "outcome": "status: In Progress"
}
Field Meaning
episode_id Pseudonymous ID of the episode, derived from the scrubbed ID of the thing it describes
source The connector it came from, such as jira or zendesk
entity What the episode is about: ticket, issue, pull_request, opportunity and so on
actor Token for the person who acted most often in the episode
started_at Time of the first step, at the dataset's date granularity
steps The steps in order
steps[].at When the step happened
steps[].actor Token for the person who took the step
steps[].state The tracked fields before the step, such as status, priority or assignee
steps[].action What happened: created, a field change written as field: old -> new, or a comment, review, merge or label written as commented: <text>
steps[].outcome The immediate result, such as status = In Review, approved, merged or comment added
outcome The final status, state, resolution or stage, with merged first for merged pull requests, or open if none is set

Every value in an episode goes through the same scrubber as records, with the same pseudonyms. People in actors, assignees and owners become PERSON_ tokens, dates are generalized, and comment text is redacted like any other free text.

An episode covers the whole history of its item up to the end of the batch's window. When the item changes again later, it appears in a later batch with the same episode_id and its longer history. Keep the latest version.

Each connector page lists whether it produces episodes and from which history.

Environments

An environment is a working replica of a system: its structure plus synthetic data drawn from the distributions in the real records, for agents to practice in. The SDK does not produce environments, and there is no standard delivery format for them yet. Ask us about a specific system.