AWS Glue Data Catalog

AWS Glue Data Catalog

Let AI connect your sources for you

Skip the manual setup — run this in your project and the wizard auto-detects your databases and APIs and connects them to PostHog.

Learn more
PostHog Wizard hedgehog

Connect AWS Glue Data Catalog to PostHog to sync your data into the PostHog data warehouse for analysis and modeling.

Sync catalog metadata, job history, and crawlers from one AWS region. Grant glue:GetDatabases, glue:GetTables, and glue:GetPartitions for catalog tables. Grant glue:GetJobs, glue:GetJobRuns, and glue:GetCrawlers for jobs and crawlers. Lake Formation permissions can also restrict catalog access. Enter the access key ID and secret access key. Add a session token for temporary credentials.

Configuration

OptionDescription
AWS access key ID
Type: text
Required: True
AWS secret access key
Type: password
Required: True
AWS session token
Type: password
Required: False

Temporary credentials expire. Replace them before the next scheduled sync.

AWS region
Type: text
Required: True

Linking AWS Glue Data Catalog to PostHog

  1. Go to the Data pipeline page in PostHog
  2. Click New source and select AWS Glue Data Catalog
  3. Fill in the required configuration fields
  4. Click Next, select the tables you want to sync, and then press Import

Supported tables

TableDescriptionSync methodIncremental fieldPrimary key
databases

Databases in the account's default Data Catalog.

Full refresh——
tables

Table definitions, columns, and storage metadata for each database.

Full refresh——
partitions

Partition values and storage metadata for each table.

Full refresh——
jobs

Job definitions and execution settings.

Full refresh——
job_runs

Job execution history, status, duration, and resource usage.

Full refresh——
crawlers

Crawler settings, state, and the latest crawl result.

Full refresh——