Docs/SDK

Incremental runs and state

How the SDK remembers where each source left off, so every run sends only new and changed records.

The SDK keeps a small state file next to your config. For each source it records where the last successful run ended, and the next run starts there.

The run window

Each run pulls one window of time per source.

  • The first run has no start. It pulls the full history the connector can reach, up to the moment the run started.
  • Every later run starts where the last successful upload ended and runs up to the moment it started.

The start is inclusive and the end is exclusive, so a record updated exactly at the boundary goes out once. Each connector applies the window to its own timestamp, usually the last-updated time. Records edited after they were sent fall into a later window and go out again with their new values. Each connector page says which timestamp it uses.

The period

Each batch is labeled with a period, the month that holds the middle of the run window. A monthly run on October 1 that covers September is labeled 2026-09. The first run, which has no start, is labeled with the month it ran in.

The state file

The state lives in .datayield/state.json next to the config file, unless stateFile in the config says otherwise. For each source it holds:

Field Meaning
sourceId The source's ID from the API
lastUntil The end of the window of the last successful upload. The next run starts here.
lastRunAt When the last successful upload finished
lastBatches The batches that upload created, with product, period and record count

The state only moves forward after a successful upload, or after a run that found nothing new. If a run fails partway, stops at the leak check, or is a dry run, the next run pulls the same window again, so nothing is skipped.

The SDK writes the state to a temporary file and renames it into place, so a crash never leaves a half-written state file.

Running on more than one machine

The state file belongs to the machine and folder that ran the SDK. If you move the SDK to a new machine, copy datayield.config.json and .datayield/state.json with it, and use the same salt.

If the state file is missing, an upload run asks the API for the source's batches and starts from the newest one that was not rejected. A fresh CI runner therefore carries on where the last run left off instead of resending the full history. A dry run with no state file previews the full history.

Two machines should not run the same source. Each would keep its own state, and both would upload the same records.

Starting from a date

datayield run --since <date> overrides the start of the window for that run. Use it to resend a period after fixing a scrub setting, or to limit a first run to recent history.

Episodes and the window

For connectors that produce episodes, the window selects which items changed, and the SDK then reads each item's full history so its episode is complete. An item that changes again next month has its episode rebuilt and sent again with the new steps.