Docs/SDK

Scheduling

Run the SDK every month with cron on Linux, launchd on macOS, or a GitHub Actions workflow.

Once datayield preview looks right and your first datayield run has delivered your history, schedule it. Each scheduled run sends the new records since the last one, and nobody has to touch it.

bash
datayield schedule            # print the entry for this machine
datayield schedule --install  # install it
datayield schedule --github   # write a GitHub Actions workflow instead

Runs are monthly by default, at 06:00 on the 1st. Pass --weekly to run at 06:00 every Monday instead. Weekly runs still label each batch with the month in the middle of its window. Times are the machine's local time for cron and launchd, and UTC on GitHub Actions.

Linux: cron

datayield schedule prints a crontab line like this:

bash
0 6 1 * * cd <config folder> && <path to bun> <path to sdk>/src/cli.ts run --config <config folder>/datayield.config.json >> <config folder>/.datayield/run.log 2>&1 # datayield

The printed line has the real paths to Bun, the SDK and your config on that machine filled in. Add it with crontab -e, or run datayield schedule --install to add it for you. Installing again replaces the earlier datayield line for the same config instead of adding a second one.

Output from every run goes to .datayield/run.log next to the config.

macOS: launchd

On macOS, datayield schedule prints a launchd agent and the command to load it. --install writes it to ~/Library/LaunchAgents/ai.datayield.run.<id>.plist and loads it with launchctl. The agent runs at the same times as the cron line and logs to the same .datayield/run.log.

launchd runs agents only while you are logged in. For a machine that may be logged out on the 1st, use a server with cron or GitHub Actions.

GitHub Actions

bash
datayield schedule --github [--repo DIR] [--force]

--github writes .github/workflows/datayield.yml into the repository at --repo, or the current folder. It refuses to overwrite an existing workflow unless you pass --force. The workflow:

  • runs on the same cron schedule, and can also be started by hand from the Actions tab
  • installs Bun and runs the same SDK version you have, with bunx @datayield/sdk@<version> run, so it needs the published package
  • has read-only access to the repository contents
  • caches the .datayield state folder between runs

The command lists the repository secrets to add under Settings → Secrets and variables → Actions:

Secret Value
DATAYIELD_API_KEY An API key for this workflow. Create a separate key named for it, such as "GitHub Actions".
DATAYIELD_SALT The salt from your ~/.datayield/credentials. It must be the same salt, or pseudonyms will not match earlier batches.
One per env:NAME option The source credentials your config refers to, such as JIRA_TOKEN

CI cannot read your local credentials file. If a source's token is stored there, schedule --github names it. Run datayield connect again with -o key=env:NAME and add NAME as a repository secret.

If the cached state is missing, for example after the cache expires, the run asks the API for the newest batch and continues from there. See Incremental runs.

The runner must be able to reach your source systems. GitHub-hosted runners work for cloud systems such as Xero, HubSpot or Jira. For a database inside your network, use a self-hosted runner or cron on a server.

Checking that it ran

bash
datayield status

status shows each source's last successful run and its recent batches. The portal shows the same batches under Datasets. If a month is missing, read .datayield/run.log or the Actions run log, then see Troubleshooting.