Linking Unstructured as a source
Let AI connect your sources for you
Skip the manual setup — run this in your project and the wizard auto-detects your databases and APIs and connects them to PostHog.

This source is currently in alpha. The interface and available tables may change.
The Unstructured connector syncs metadata from your Unstructured platform account – workflows, job runs, and source/destination connectors – into the PostHog data warehouse so you can monitor document pipeline health alongside the rest of your data.
Unstructured is a document ETL platform that turns raw files (PDFs, Office documents, emails, and more) into RAG-ready structured data.
Prerequisites
- An Unstructured platform account (a free tier is available).
- An Unstructured API key. The key has account-wide read access, so no extra scopes are required.
Adding a data source
- In PostHog, go to the Sources tab of the data pipeline section.
- Click + New source and click Link next to this source.
- Enter your credentials (see Configuration below) and click Next.
- Select the tables you want to sync, choose a sync method and frequency, then click Import.
Once the syncs are complete, you can start querying this data in PostHog.
When linking Unstructured, you'll need:
- API key – sign in to the Unstructured platform dashboard, open API Keys, and generate a new key. Copy the key and paste it into the API key field in PostHog.
Leave API host blank unless Unstructured provisioned your account with a custom API URL, in which case enter that origin (for example https://platform.unstructuredapp.io).
Sync modes
Each table can be synced in one of several modes, depending on what the source supports:
- Webhook (when available) – the source pushes changes to PostHog in real time. Fastest freshness, lowest ongoing cost, and the only mode that reliably captures updates and deletes.
- Incremental – only new or updated rows are synced on each run, using a cursor field (such as an
updated_attimestamp). Cheaper than a full refresh, but deletes aren't captured. - Append only – new rows are appended using a cursor field; existing rows are never updated. Ideal for immutable, append-only tables like event logs.
- Full refresh – the whole table is reloaded on every sync. Use it when a table has no reliable cursor or when you need deletions reflected.
See sync methods for a full explanation of how each mode works and how to choose between them.
All Unstructured tables sync as full refresh only. The API doesn't expose a reliable server-side incremental filter, and workflows and connectors are mutable configuration objects whose fields (status, schedule, timestamps) change over time. A full refresh each sync is inexpensive for these low-volume inventory tables and always reflects the current state.
Configuration
| Option | Type | Required |
|---|---|---|
API key | password | Yes |
API host | text | No |
Supported tables
| Table | Description | Sync method | Incremental field | Primary key |
|---|---|---|---|---|
workflows | Workflows that define how Unstructured ingests, processes, and routes documents from sources to destinations. | Full refresh | — | — |
jobs | Individual runs of a workflow at a point in time, including status and runtime. | Full refresh | — | — |
sources | Source connectors that ingest files or data into Unstructured from an external location. | Full refresh | — | — |
destinations | Destination connectors that send Unstructured's processed data to an external location. | Full refresh | — | — |
Troubleshooting
- Invalid API key – regenerate the key in the Unstructured platform dashboard and reconnect. The key is account-wide, so a valid key can read every table.
- Custom API host – if your account uses a custom API URL, make sure the API host field matches the origin Unstructured gave you. The scheme is forced to HTTPS.
If your sync is failing or data looks wrong, see the Data warehouse troubleshooting guide. If that doesn't help, contact support – we're happy to help.