Linking SFTP as a source

Let AI connect your sources for you

Skip the manual setup — run this in your project and the wizard auto-detects your databases and APIs and connects them to PostHog.

Learn more
PostHog Wizard hedgehog

Alpha release

This source is currently in alpha. The interface and available tables may change.

The SFTP connector syncs your file storage data into the PostHog data warehouse, so you can analyze it alongside your product data.

Prerequisites

Credentials that can read the data you want to sync. PostHog only reads data, so read access is enough.

Adding a data source

  1. In PostHog, go to the Sources tab of the data pipeline section.
  2. Click + New source and click Link next to this source.
  3. Enter your credentials (see Configuration below) and click Next.
  4. Select the tables you want to sync, choose a sync method and frequency, then click Import.

Once the syncs are complete, you can start querying this data in PostHog.

Import CSV and JSON files from an SFTP server. PostHog lists the files in the folder you point it at, including subfolders, and creates one table per file, or one combined table if you prefer. Every sync reads the files in full, so each table matches what's on the server right now.

You'll be asked for:

  • Host: for example sftp.example.com.
  • Port: for example 22.
  • Username: for example posthog.
  • Authentication type: choose between Password and SSH private key.
  • Folder: for example /incoming.
  • File format: choose between Detect from file extension, CSV, JSON Lines and JSON.

Sync modes

Each table can be synced in one of several modes, depending on what the source supports:

  • Webhook (when available) – the source pushes changes to PostHog in real time. Fastest freshness, lowest ongoing cost, and the only mode that reliably captures updates and deletes.
  • Incremental – only new or updated rows are synced on each run, using a cursor field (such as an updated_at timestamp). Cheaper than a full refresh, but deletes aren't captured.
  • Append only – new rows are appended using a cursor field; existing rows are never updated. Ideal for immutable, append-only tables like event logs.
  • Full refresh – the whole table is reloaded on every sync. Use it when a table has no reliable cursor or when you need deletions reflected.

See sync methods for a full explanation of how each mode works and how to choose between them.

All SFTP tables are full refresh. Each sync replaces the contents of the table.

Configuration

OptionDescription
Host
Type: text
Required: True
Port
Type: number
Required: True
Username
Type: text
Required: True
Authentication type
Type: select
Required: True
Folder
Type: text
Required: True
File pattern (optional)
Type: text
Required: False

A regular expression matched against each file's path inside the folder. Leave it empty to import every file.

File format
Type: select
Required: True
CSV delimiter (optional)
Type: text
Required: False

Use \t for tab-separated files. Defaults to a comma.

Combine every file into one table?
Type: switch-group
Required: False

Turn this on when the folder holds files that share the same columns, such as a daily export. All matching files land in one table instead of one table per file.

Supported tables

The tables available from this source are discovered from your account when you connect it, so the exact list depends on your data. Once connected, you can pick which tables to sync from the sources tab.

Troubleshooting

  • If the connection fails with an authorization error, the password is wrong, expired, or has been revoked. Create a new one, then reconnect the source.
  • If a table syncs no rows, the credential may not have access to that data. Check its permissions, then reconnect the source.

If your sync is failing or data looks wrong, see the Data warehouse troubleshooting guide. If that doesn't help, contact support – we're happy to help.

Still have questions?

Was this page useful?