Linking Scale AI as a source

Let AI connect your sources for you

Skip the manual setup — run this in your project and the wizard auto-detects your databases and APIs and connects them to PostHog.

Learn more
PostHog Wizard hedgehog

Alpha release

This source is currently in alpha. The interface and available tables may change.

The Scale AI connector syncs tasks, batches, projects, and more into the PostHog data warehouse, so you can analyze them alongside your product data.

Prerequisites

Credentials that can read the data you want to sync. PostHog only reads data, so read access is enough.

Adding a data source

  1. In PostHog, go to the Sources tab of the data pipeline section.
  2. Click + New source and click Link next to this source.
  3. Enter your credentials (see Configuration below) and click Next.
  4. Select the tables you want to sync, choose a sync method and frequency, then click Import.

Once the syncs are complete, you can start querying this data in PostHog.

Enter your Scale AI API key to sync your Scale data labeling data.

You can find your API key in the Scale dashboard under Settings → API Keys. Use a live-mode key - test-mode keys have fully isolated data. Only account Managers and Admins can access API keys.

You'll be asked for:

  • API key: for example live_....

Sync modes

Each table can be synced in one of several modes, depending on what the source supports:

  • Webhook (when available) – the source pushes changes to PostHog in real time. Fastest freshness, lowest ongoing cost, and the only mode that reliably captures updates and deletes.
  • Incremental – only new or updated rows are synced on each run, using a cursor field (such as an updated_at timestamp). Cheaper than a full refresh, but deletes aren't captured.
  • Append only – new rows are appended using a cursor field; existing rows are never updated. Ideal for immutable, append-only tables like event logs.
  • Full refresh – the whole table is reloaded on every sync. Use it when a table has no reliable cursor or when you need deletions reflected.

See sync methods for a full explanation of how each mode works and how to choose between them.

All Scale AI tables are full refresh. Each sync replaces the contents of the table.

Configuration

OptionTypeRequired
API keypasswordYes

Supported tables

TableDescriptionSync methodIncremental fieldPrimary key
tasks

A unit of data labeling work submitted to Scale, with its type, status, timestamps, and review results.

Incremental, Full refreshupdated_at, created_attask_id
batches

Incremental sync filters on created_at, so it catches newly created batches but not status changes to existing ones; use a full refresh to re-read batch status.

Incremental, Full refreshcreated_atname
projects

A container that defines the labeling task type and instructions shared by its tasks and batches.

Full refresh—name

Troubleshooting

  • If the connection fails with an authorization error, the API key is wrong, expired, or has been revoked. Create a new one, then reconnect the source.
  • If a table syncs no rows, the credential may not have access to that data. Check its permissions, then reconnect the source.

If your sync is failing or data looks wrong, see the Data warehouse troubleshooting guide. If that doesn't help, contact support – we're happy to help.

Still have questions?

Was this page useful?