Bot and traffic detection

Contents

PostHog classifies traffic by user agent and source IP address so you can tell humans apart from bots, crawlers, and automation directly in your queries. It categorizes each request – so you can welcome, measure, or exclude AI agents, search crawlers, and automation independently, rather than treating every bot the same. The classification runs in SQL, so it works anywhere HogQL does – the SQL editor, insights, trends, and Web Analytics breakdowns.

New feature

This is a brand new feature, and we're actively working on improving it – expanding the list of detected bots and refining classification over time. Function names and behavior may still change. If you have feedback or a bot we should detect, leave a comment on this page or open a pull request against the bot definitions or bot IP definitions.

This is different from the client-side bot blocking in the PostHog JavaScript SDK, which stops detected bots from sending events in the first place. The functions here classify traffic that has already been captured, so you can include, exclude, or break down by traffic type at query time without losing the underlying data.

Capturing server-side traffic with $http_log

Most bots don't run JavaScript, so the PostHog JavaScript SDK never fires a $pageview for them. To see this traffic, forward your HTTP access logs to PostHog as $http_log events – log entries from your web server, CDN, or edge network. They carry the raw user agent, so the functions below classify them the same way they classify $pageview and $screen events.

Set $raw_user_agent on each event, plus $host, $current_url, $pathname, and status_code for richer breakdowns. There are a few ways to send them:

  • Log drain source – in PostHog, go to Data pipeline > Sources and add the Vercel logs source. PostHog generates an endpoint, registers it with your project, and captures each log line as a $http_log event with the user agent populated.
  • Edge worker – intercept requests at your CDN or edge (for example, a Cloudflare Worker) and send a $http_log event to the capture API. This works on any plan and lets you control exactly what's logged.
  • Capture API directly – send $http_log events from any web server or reverse proxy:
JSON
{
"api_key": "<ph_project_token>",
"event": "$http_log",
"distinct_id": "server-log",
"properties": {
"$raw_user_agent": "Mozilla/5.0 (compatible; GPTBot/1.2; +https://openai.com/gptbot)",
"$current_url": "https://yoursite.com/blog/hello-world",
"$pathname": "/blog/hello-world",
"status_code": 200,
"method": "GET"
}
}

For the full payload reference and copy-paste implementations (Cloudflare Worker, server middleware), see sending HTTP log events.

Once $http_log events are flowing, the functions and virtual properties below classify them alongside your $pageview and $screen events.

Person processing and distinct IDs

Server requests don't carry a PostHog cookie, so for any $http_log ingestion you decide two things per event: what distinct_id to assign, and whether the event creates a person profile. Both are worth thinking about:

  • Person processing. Backend events are identified by default, which creates a person profile per distinct_id and bills on the person-profiles line. For high-cardinality log traffic that adds cost and slows person-joined queries. Capturing these as anonymous events ($process_person_profile set to false) avoids that, and bot detection still works since it reads event properties, not the profile. Use identified only if you need person-level stitching.
  • Distinct ID. When events are anonymous, the distinct ID no longer affects cost – it only affects unique-visitor counts. A derived, per-client ID – for example a hash of IP, host, and user agent – keeps one stable identity per client. Avoid a single shared ID, which counts all traffic as one visitor; a random per-request ID makes every request its own visitor.

How you set these depends on how you ingest. With the capture API or an edge worker you set $process_person_profile and the distinct_id directly on each event. The Vercel logs source exposes both as settings, and defaults to anonymous person processing with a fixed-salt distinct ID (one stable ID per client) – you can change either.

Classification functions

Each function takes a user agent string and returns its classification. PostHog stores the user agent as $raw_user_agent, which the JavaScript SDK, server-side capture, and $http_log events all set. Some SDKs send $user_agent instead, so the examples below pass coalesce(properties.$raw_user_agent, properties.$user_agent) to cover both. To skip passing it, use the virtual properties below, which read $raw_user_agent for you.

FunctionReturnsExample
isLikelyBot(user_agent)true if the user agent matches a bot or automation pattern, otherwise false. An empty user agent counts as a bot.true
getTrafficType(user_agent)One of AI Agent, Bot, Automation, or Regular.Bot
getTrafficCategory(user_agent)A subcategory such as ai_crawler, ai_search, search_crawler, seo_crawler, social_crawler, monitoring, http_client, or headless_browser. Returns regular for human traffic.search_crawler
getBotType(user_agent)The same subcategory as getTrafficCategory, but returns an empty string for human traffic – handy for filtering.search_crawler
getBotName(user_agent)The bot's name, such as Googlebot, ChatGPT, or curl. Empty for human traffic.Googlebot
getBotOperator(user_agent)The company or operator behind the bot, such as Google, OpenAI, or Anthropic. Empty for human traffic.Google

isLikelyBot is named "likely" because detection is based on user agent heuristics – it can't confirm with certainty that a request is automated.

Traffic types

getTrafficType sorts every request into one of four traffic types:

Traffic typeDescriptionExamples
RegularNormal human visitorsChrome, Safari, Firefox
AI AgentAI crawlers, search bots, and assistantsGPTBot, ClaudeBot, PerplexityBot
BotTraditional crawlers and monitoringGooglebot, Bingbot, AhrefsBot
AutomationHTTP clients, headless browsers, and requests with no user agentcurl, Puppeteer, HeadlessChrome

Bot categories

getTrafficCategory and getBotType return a finer-grained category within each traffic type:

CategoryTraffic typeDescriptionExamples
ai_crawlerAI AgentTraining data collectionGPTBot, ClaudeBot, Google-Extended
ai_searchAI AgentAI-powered search resultsOAI-SearchBot, Claude-SearchBot, Applebot
ai_assistantAI AgentReal-time user-facing AIChatGPT-User, Claude-User, Perplexity-User
search_crawlerBotTraditional search enginesGooglebot, Bingbot, Baidu
seo_crawlerBotSEO analysis toolsAhrefsBot, SemrushBot, Majestic
social_crawlerBotSocial media preview crawlersFacebook, Twitter, LinkedIn, Slack
monitoringBotUptime and health monitoringPingdom, UptimeRobot, Datadog
http_clientAutomationHTTP client librariescurl, Wget, Python requests, axios
headless_browserAutomationAutomated browsersHeadlessChrome, Puppeteer, Playwright
no_user_agentAutomationEmpty or missing user agent–

Example queries

Break down pageviews by traffic type:

SQL
SELECT
getTrafficType(coalesce(properties.$raw_user_agent, properties.$user_agent)) AS traffic_type,
count() AS events
FROM events
WHERE event = '$pageview'
GROUP BY traffic_type
ORDER BY events DESC

Count human pageviews by excluding bots:

SQL
SELECT count() AS human_pageviews
FROM events
WHERE event = '$pageview'
AND NOT isLikelyBot(coalesce(properties.$raw_user_agent, properties.$user_agent))

Find which bots hit your site most often:

SQL
SELECT
getBotName(coalesce(properties.$raw_user_agent, properties.$user_agent)) AS bot,
getBotOperator(coalesce(properties.$raw_user_agent, properties.$user_agent)) AS operator,
count() AS hits
FROM events
WHERE event = '$pageview'
AND isLikelyBot(coalesce(properties.$raw_user_agent, properties.$user_agent))
GROUP BY bot, operator
ORDER BY hits DESC

Virtual properties

For convenience, PostHog exposes the classification as virtual event properties. They read $raw_user_agent for you, so you don't have to pass it in. Events that only carry $user_agent are classified as having no user agent. They're available wherever you select event properties, including breakdowns:

PropertyName in the UIEquivalent to
$virt_is_botIs botisLikelyBot(...)
$virt_traffic_typeTraffic typegetTrafficType(...)
$virt_traffic_categoryTraffic categorygetTrafficCategory(...)
$virt_bot_nameBot namegetBotName(...)
$virt_bot_operatorBot operatorgetBotOperator(...)

The Property column is the name you use in SQL. The Name in the UI column is what you search for in the property picker when building an insight, dashboard, or breakdown.

For example, to break down traffic without writing out the function:

SQL
SELECT
properties.$virt_traffic_type AS traffic_type,
count() AS events
FROM events
WHERE event = '$pageview'
GROUP BY traffic_type
ORDER BY events DESC

Use in Product Analytics

You don't need to write SQL to break traffic down. As long as your events carry $raw_user_agent – set by the JavaScript SDK, server-side SDKs, and $http_log events – the virtual properties work as breakdowns and filters in any Product Analytics insight: trends, funnels, retention, paths, and more.

For example:

  • Exclude bots from an insight – add a filter where Is bot equals false.
  • Break down traffic by type – set the breakdown to Traffic type to split a trend into Regular, AI Agent, Bot, and Automation.
  • See which crawlers hit a page – filter to Is bot is true and break down by Bot name.

Because the classification reads the raw user agent, this works for any event that carries one – including server-side and $http_log traffic – using the standard insight builder.

How classification works

Classification uses two signals:

  • User agent patterns – the user agent is matched against a maintained list of known bot, crawler, and automation patterns. The list is open source and pull requests are welcome – you can find it in the PostHog repository.

  • IP address ranges – the source IP is checked against operator-published crawler IP ranges. Some crawlers – like ChatGPT-User browsing or Bing preview fetches – send real browser user agents with no bot token, so the source IP is the only reliable signal. PostHog currently checks ranges published by Google, OpenAI, Microsoft (Bing), Apple, Perplexity, and Ahrefs. The IP definitions are also open source.

If either signal matches, the request is classified as bot traffic. You can also add your own rules on top of both lists – see custom bot rules.

Custom bot rules

The built-in lists only cover bots that PostHog knows about. If a bot matters to you but isn't detected – an internal load test, a partner integration, a niche crawler, or a scraper that sends a normal browser user agent – add a custom bot rule in your project settings.

A custom rule counts as a bot everywhere Is bot is available, including insights, Web Analytics, and SQL. It updates all the virtual properties, not only $virt_is_bot, and applies at query time, so it also classifies events you've already captured.

The classification functions only see the arguments you pass them, so they apply rules on the user agent but skip rules on other properties – use the virtual properties to get every rule.

Add a rule

  1. Open the custom bots settings

    Go to Settings > Customization > Custom bots, or open them directly.

  2. Name the bot and pick a category

    Click Add bot, then enter a name, like Acme scraper, and pick a category. The name becomes the Bot name and Bot operator for matching events. The category sets the Traffic category and the Traffic type:

    CategoryTraffic type
    Custom (default), Search crawler, SEO crawler, Social crawler, MonitoringBot
    AI crawler, AI search, AI assistantAI Agent
    HTTP client, Headless browserAutomation

  3. Add conditions

    Each condition matches one event property. Click Add condition to combine several, then choose whether all or any of them need to match.

  4. Save

    Click Save. PostHog checks every pattern when you save and tells you if one can't be used.

You can also manage rules without the web app. Ask an AI agent to list, create, or delete them over the PostHog MCP, or call the web analytics API from your own scripts.

Properties and matchers

A condition can match on any of these event properties:

Property in the ruleEvent property
Raw user agent$raw_user_agent
IP address$ip
Library$lib
Host$host
Path name$pathname
Current URL$current_url
Browser$browser
OS$os
Browser language$browser_language
Screen width$screen_width
Screen height$screen_height
Country code$geoip_country_code
Referrer$referrer
Referring domain$referring_domain

Each condition uses one of these matchers:

  • contains – the property contains the text anywhere. Case-insensitive.
  • equals – the property is exactly the text. Case-sensitive.
  • matches regex – the property matches a regular expression. Case-sensitive unless the pattern starts with (?i).
  • is in range – the IP address is inside a CIDR range, like 192.0.2.0/24. Only available for IP address.

Regular expressions run in ClickHouse, which supports less than most regex engines. Lookaheads, lookbehinds, backreferences, and atomic groups aren't supported, and PostHog rejects a rule that uses them when you save it.

Which rule wins

Your rules are checked before PostHog's built-in lists, in the order they're listed, and the first matching rule decides the classification. Drag rules in the editor to reorder them. Because your rules come first, you can also rename or recategorize a bot PostHog already detects. For example, if your own synthetic checks run in headless Chrome, a rule that matches HeadlessChrome with the category Monitoring makes those events count as Bot traffic instead of Automation.

Examples

  • An internal load test – IP address is in range 198.51.100.0/24, in the HTTP client category.
  • A scraper that names itself – Raw user agent contains AcmeBot.
  • A scraper that pretends to be a browser – a combination no real visitor sends: Screen width equals 800 and Screen height equals 600 and OS equals Linux.
  • A partner integration – Library equals posthog-python and Path name matches regex ^/api/.

Limits

  • Up to 50 rules per project
  • Up to 10 conditions per rule, and 100 conditions across all rules
  • Patterns up to 200 characters, and bot names up to 100 characters
Custom rules tag events, they don't drop them

A custom rule only changes how events are classified, so you can still include or break down that traffic whenever you want. To stop bot events from being stored at all, use the Filter Bot Events transformation, which accepts its own custom user agent patterns and IP ranges.

Limitations

Detection is best-effort, so a few cases are worth keeping in mind:

  • No user agent – requests with an empty or missing user agent (server-to-server calls, misconfigured SDKs) can't be classified from the user agent alone. They're treated as Automation with the no_user_agent category, so isLikelyBot returns true for them.
  • Spoofing – some bots disguise themselves with regular browser user agents, and some legitimate tools use bot-like ones. IP-based detection catches known crawlers that use real browser user agents (such as ChatGPT-User and Bing preview fetches), but only covers operators that publish their crawler IP ranges. Bots from unpublished IP ranges with spoofed user agents can still evade detection, which is why the boolean function is named isLikelyBot. If you can identify one by its IP range or by a combination of properties, add a custom bot rule.

Still have questions?

Was this page useful?