Docs/Getting started

Quickstart for sellers

Install the SDK, connect one system, preview the scrubbed output and send your first batch.

You need an organization in the portal, an API key from Settings, and a machine with Bun 1.2 or later that can reach the system you want to connect. The SDK runs on that machine, and raw records never leave it.

1. Create an API key

In the portal, open Settings → API keys and create a key. Keys starting with dy_test_ are for trying things out and never create billable batches. Keys starting with dy_live_ deliver real batches. Start with a test key. See API keys.

2. Install and initialize

Install Bun and the SDK as described in Install, then run:

bash
datayield init

datayield init asks for your API key, checks it, and writes datayield.config.json in the current folder. It stores the key in ~/.datayield/credentials, readable only by your user, along with a new salt that keeps pseudonyms stable from one batch to the next. Neither is written to the config file. Back up the salt. See Install and Configuration.

3. Connect a source

bash
datayield connect hubspot

connect adds a source to the config, asks for any required option you did not pass with -o key=value, and checks the credentials straight away, so a bad token fails here instead of on the first scheduled run. Tokens go to the credentials file, not the config. Each connector page lists the token it needs and the scopes to grant; start with Connectors.

4. Preview

bash
datayield preview --out review

preview pulls records, scrubs them on your machine and prints the scrub report without uploading anything. --out review also writes the scrubbed records and the report to a review folder, so you can read exactly what would be sent. Read the column roles in the report. If a column is classified wrongly, fix it with an override in the config. See How scrubbing works.

5. Run

bash
datayield run

run does the same work as preview and then uploads the scrubbed batch with its report, one batch for records and one for episodes if the source produces them. The first run of a source sends its history. Later runs send only records created or changed since the last successful run. If the leak check finds an identifier in the output, the run stops with exit code 2 and uploads nothing.

6. Schedule

bash
datayield schedule --install

schedule prints a cron line on Linux or a launchd agent on macOS. Add --install to install it, or use --github to write a GitHub Actions workflow instead. Once it is in place, new batches go out every month without anyone touching them. See Scheduling.

Check what happened

bash
datayield status

status shows the last run of each source and every batch with its status. The same batches appear in the portal under Datasets.