Linking Soda Cloud as a source

Let AI connect your sources for you

Skip the manual setup — run this in your project and the wizard auto-detects your databases and APIs and connects them to PostHog.

Learn more
PostHog Wizard hedgehog

Alpha release

This source is currently in alpha. The interface and available tables may change.

The Soda Cloud connector syncs your data quality monitoring data – datasets, checks, and incidents – into PostHog, so you can analyze data quality trends alongside your product data.

Prerequisites

You need a Soda Cloud account with API access. To generate API keys:

  1. Log in to Soda Cloud.
  2. Click your avatar and select Profile.
  3. Go to API Keys and generate a new key.

This gives you an API key ID and API key secret. The synced data reflects the View dataset permissions of the API key owner – only datasets and checks visible to that user are included.

Adding a data source

  1. In PostHog, go to the Sources tab of the data pipeline section.
  2. Click + New source and click Link next to this source.
  3. Enter your credentials (see Configuration below) and click Next.
  4. Select the tables you want to sync, choose a sync method and frequency, then click Import.

Once the syncs are complete, you can start querying this data in PostHog.

When linking Soda Cloud, you'll need:

  • API key ID – the key ID from your Soda Cloud profile.
  • API key secret – the corresponding secret.
  • Region – select Europe or United States to match the region your Soda Cloud account is hosted on. Defaults to Europe.

Sync modes

Each table can be synced in one of several modes, depending on what the source supports:

  • Webhook (when available) – the source pushes changes to PostHog in real time. Fastest freshness, lowest ongoing cost, and the only mode that reliably captures updates and deletes.
  • Incremental – only new or updated rows are synced on each run, using a cursor field (such as an updated_at timestamp). Cheaper than a full refresh, but deletes aren't captured.
  • Append only – new rows are appended using a cursor field; existing rows are never updated. Ideal for immutable, append-only tables like event logs.
  • Full refresh – the whole table is reloaded on every sync. Use it when a table has no reliable cursor or when you need deletions reflected.

See sync methods for a full explanation of how each mode works and how to choose between them.

The datasets table supports incremental sync using the lastUpdated field, so only datasets modified since the last sync are fetched. Full refresh is also available.

The checks and incidents tables use full refresh only. Incremental sync isn't supported for these tables because creation-based filters would miss later changes to check results and incident statuses.

Configuration

OptionTypeRequired
API key IDtextYes
API key secretpasswordYes
RegionselectYes

Supported tables

TableDescriptionSync methodIncremental fieldPrimary key
datasets

Datasets visible to the API key owner, with check coverage and the latest quality status.

Incremental, Full refreshlastUpdated—
checks

Checks with their latest evaluation, associated datasets, and linked incidents.

Full refresh——
incidents

Data quality incidents with their status, severity, linked checks, and assigned lead.

Full refresh——

Checks include the current check result values. Separate historical check results and scan data are not synced.

Troubleshooting

  • If the connection fails with an authentication error, check that the API key ID, key secret, and region are correct. Keys are region-specific – a Europe key won't work against the United States endpoint.
  • If a table syncs fewer rows than expected, verify that the API key owner has View dataset permissions for the data you want to sync.

If your sync is failing or data looks wrong, see the Data warehouse troubleshooting guide. If that doesn't help, contact support – we're happy to help.

Still have questions?

Was this page useful?