Linking Dagster+ (Dagster Cloud) as a source
Let AI connect your sources for you
Skip the manual setup — run this in your project and the wizard auto-detects your databases and APIs and connects them to PostHog.

This source is currently in alpha. The interface and available tables may change.
The Dagster+ (Dagster Cloud) connector syncs runs, backfills, assets, and more into the PostHog data warehouse, so you can analyze them alongside your product data.
Prerequisites
Credentials that can read the data you want to sync. PostHog only reads data, so read access is enough.
Adding a data source
- In PostHog, go to the Sources tab of the data pipeline section.
- Click + New source and click Link next to this source.
- Enter your credentials (see Configuration below) and click Next.
- Select the tables you want to sync, choose a sync method and frequency, then click Import.
Once the syncs are complete, you can start querying this data in PostHog.
Connect your Dagster+ deployment to sync run history, backfills, and your asset catalog.
Create a user token under Organization settings → Tokens in Dagster+, then enter your organization name, deployment name, and the token below.
You'll be asked for:
- Organization: for example
your-org. - Deployment: for example
prod. - API token: for example
user:....
Sync modes
Each table can be synced in one of several modes, depending on what the source supports:
- Webhook (when available) – the source pushes changes to PostHog in real time. Fastest freshness, lowest ongoing cost, and the only mode that reliably captures updates and deletes.
- Incremental – only new or updated rows are synced on each run, using a cursor field (such as an
updated_attimestamp). Cheaper than a full refresh, but deletes aren't captured. - Append only – new rows are appended using a cursor field; existing rows are never updated. Ideal for immutable, append-only tables like event logs.
- Full refresh – the whole table is reloaded on every sync. Use it when a table has no reliable cursor or when you need deletions reflected.
See sync methods for a full explanation of how each mode works and how to choose between them.
Some Dagster+ (Dagster Cloud) tables sync incrementally, so later runs only fetch new or updated rows. The rest are full refresh.
Configuration
| Option | Type | Required |
|---|---|---|
Organization | text | Yes |
Deployment | text | Yes |
API token | password | Yes |
Supported tables
| Table | Description | Sync method | Incremental field | Primary key |
|---|---|---|---|---|
runs | Historical job/pipeline run records for the deployment, used to compute reliability, DORA, and SLA metrics. | Incremental, Full refresh | updateTime, creationTime | — |
backfills | Partition backfills launched in the deployment, including asset backfills. | Full refresh | — | — |
assets | Catalog of assets known to the deployment. | Full refresh | — | — |
schedules | Schedule definitions in each code location, with their cron expression and default status. | Full refresh | — | — |
sensors | Sensor definitions in each code location, with their type, tick interval, and default status. | Full refresh | — | — |
instigation_states | Live state of every schedule and sensor, including the cursor a sensor last stored. | Full refresh | — | — |
instigation_ticks | Tick history for every schedule and sensor — one row per evaluation, with its outcome and the runs it requested. Used for orchestration reliability analysis. | Incremental, Full refresh | timestamp | — |
asset_nodes | Asset definition metadata — group, owning jobs, dependencies, partitioning, and freshness policy. Resolves the bare asset keys in the assets table. | Full refresh | — | — |
asset_materializations | Materialization event history per asset — one row each time an asset was materialized, with the metadata the run attached. The fact table behind the asset catalog. | Incremental, Full refresh | timestamp | — |
asset_observations | Observation event history per asset — one row each time an observable asset was observed, with the metadata the observation reported. | Incremental, Full refresh | timestamp | — |
deployments | Every deployment in the organization, both the full deployments and the branch deployments open against them. The lookup resolving which Dagster+ deployment a run or asset belongs to. | Full refresh | — | — |
run_logs | Event log for each run, covering step starts, successes, failures, retries and log messages. Resolves what happened inside the runs table's rows. | Incremental, Full refresh | runUpdateTime | — |
users | Organization members and the roles granted to each, for attributing runs and backfills to people. Needs organization-level permissions. | Full refresh | — | — |
teams | Teams in the organization, their members, and the roles granted to each team. Needs organization-level permissions. | Full refresh | — | — |
custom_roles | Custom roles defined in the organization and the permissions each one carries. Resolves the custom role identifiers on the users and teams tables. Needs organization-level permissions. | Full refresh | — | — |
insights_job_metrics | Dagster+ Insights metrics per job over time, such as credits consumed, execution time, and materialization and retry counts. One row per metric, job and time bucket. | Incremental, Full refresh | timestamp | — |
insights_asset_metrics | Dagster+ Insights metrics per asset over time, such as credits consumed, execution time, and materialization and retry counts. One row per metric, asset and time bucket. | Incremental, Full refresh | timestamp | — |
insights_deployment_metrics | Dagster+ Insights metrics per deployment over time, such as credits consumed, execution time, and materialization and retry counts. One row per metric, deployment and time bucket. | Incremental, Full refresh | timestamp | — |
Troubleshooting
- If the connection fails with an authorization error, the API token is wrong, expired, or has been revoked. Create a new one, then reconnect the source.
- If a table syncs no rows, the credential may not have access to that data. Check its permissions, then reconnect the source.
If your sync is failing or data looks wrong, see the Data warehouse troubleshooting guide. If that doesn't help, contact support – we're happy to help.