> AI agents: this is one page from PostHog's docs. Full index of Markdown docs for LLMs: https://posthog.com/llms.txt # Most "AEO experts" are bluffing - [ ![](https://res.cloudinary.com/dmukukwp6/image/upload/c_scale,w_50/Natalia_s_Portrait_1_fd2c5fe102) Natalia Amorim](/community/profiles/35321.md) Sep 24, 2026 - [Marketing](/blog/marketing.md) In [a previous article](/blog/aeo-advice.md), I said "there has never been an easier time to be a charlatan when it comes to AEO performance". If you are here expecting a retraction of that statement, I have bad news for you: I stand by it more than ever. [Answer engine optimization (AEO) is a fickle discipline](/handbook/aeo-guide.md) riddled with questionable reporting methodologies, and it happens to be both remarkably easy to manipulate and hard to verify; that's a dangerous combination, and it makes it the hottest party trick in marketing right now. Unfortunately, the confidence people express about AEO performance is often entirely unearned. Hot take? Maybe, but between you and me, some of the success claims going around have gotten a little ridiculous lately, and I think it's time we call their bluff. ## Most AEO metrics are educated guesses in a trench coat Every AEO number you have ever seen, from visibility scores to citation rate to average rank[1](#fn-1), rests on the **unverifiable** assumption that the prompts being tracked and reported on are the prompts people actually type. This doesn't make AEO unimportant; at PostHog, it continues to be our fastest-growing channel driving thousands of sign-ups every month, and I'm not about to argue myself out of a job (not in this economy). But there is no real prompt volume data available to us mortals. None. No LLM provider publishes what people ask it, and nothing exists that plays the role that "search volume" plays in SEO. And yes, I know about clickstream data; while I can concede it's the best bad option available, that's still different from being good. Thankfully us marketers are resourceful creatures, so we do what we can with what's available to us: we proxy. We ground prompts in traditional search volume, leaning on the reasonable-enough theory that people ask LLMs roughly what they'd Google. We ask our own users what they typed. We do competitive research, social listening, the whole nine yards. And while these may be steps in the right direction, they are still educated guesses stacked on educated guesses, not a scientifically sound methodology with a path to accuracy. We're operating inside a gigantic blind spot, and I've been seeing a lot of tools and so-called experts talking about it as if it's been solved. A biblically accurate representation of the AEO industry Newsflash, it has not. Which is why some of these "who's winning on AEO" posts give me chills, and not in a good way. ## "\[Brand\] is demolishing at AEO 🚀" I get a debilitating wave of second-hand embarrassment seeing these on my feed. On the bright side, I've learned just how far back my eyes can roll. Before anyone comes for me: I am NOT saying brands claiming they have good AEO performance are necessarily lying or being dodgy about it. I know plenty of companies who are doing a spectacular job at capturing demand in the space, and they back it up with signup data, referral traffic, self-reported attribution, etc... you know, actual business outcomes. What makes me cringe is when the **entire** proof of someone's AEO success relies solely on a visibility score, calculated from a prompt set they chose... against competitors they selected… for only certain models, on certain days, at a sample size they didn't disclose. The infuriating bit is that these posts are technically true; anyone can produce a chart showing they're "winning" with numbers that are not necessarily fake (just lacking the necessary context). Don't believe me? Let me show you how it's done. ## How to manufacture a perfect AEO score in a few easy steps *A guide for the aspiring thought leader.* ### 1\. Start with a metric nobody can check. Do not report attributed signups, referral traffic, or revenue. Report on **brand visibility** or **citation rate** instead: numbers that sound like outcomes, sit at the very top of the funnel, and can't be verified by anyone. Being mentioned doesn't mean anyone clicked, clicking doesn't mean anyone signed up, and signing up doesn't mean anyone stuck around. But "our citations grew 60%" sounds like a Certifiable Business Result™ to [bring to your next QBR](/blog/are-you-actually-data-driven.md), and hopefully nobody will ask what this actually means or happened next. Every step below follows from this one. ### 2\. Invent the demand, then meet it. What makes all of this possible is that nobody can verify that a single prompt you track has ever been typed by an actual human being. Which means you can invent a question, track it, win it, and describe it as "what buyers are asking." Magic, am I right? So: track a huge set across the funnel – some generic, some specific, some branded, some not. Run it. See where you happen to win. Delete or ignore the rest. Need a boost? Put some branded prompts in the mix. You show up in every *"\[you\] vs \[competitor\]"* answer by definition, and because you're reporting on visibility rather than sentiment, you never have to admit which way it went. ### 3\. Pick competitors who are bad at the thing you're good at. Nobody audits your comparison set. Pick the two competitors who are weakest in your strongest category and simply... don't include the ones that beat you. If asked, say you chose "the most relevant players for these prompts." ### 4\. Cherry-pick the timeframe, the model, and the sample. This is free money, because these metrics are volatile; some of ours have swung multiple percentage points in a single quarter. There is almost always a slice that looks good. - **Timeframe:** when a line looks like a heart rate monitor, there is always a start date that makes it go up. Pick that one, and call it "the last 30 days" so it sounds like a neutral choice. - **Model:** you have at least four or five to choose from and you'll perform differently across all of them. At PostHog we tend to do well on Claude and Gemini, and struggle on Perplexity and ChatGPT. Report the winners only, and justify it with "we focused on these given their market share and ICP relevance." - **Sample:** the same prompt returns different answers on different runs. Run it a few times, keep the good one, and never disclose the run count. Nobody in your comments is going to ask "how many samples per prompt?" and if they do, you can always just block them. ### 5\. Flip flop between metrics that flatter you. Mentions up but citations down? Report mention rate and call it "visibility." Citations up but mentions flat? Report citation rate and call it "authority." Both down? Report share of voice against a competitor set you get to choose (refer back to step 3). There are enough metrics in this space that at least one of them is destined to be having a good week. ### 6\. Compare your number to someone else's number. Take your citation rate, measured with your tool, on your prompts. Put it next to a figure a competitor published using a different tool, a different prompt set, and a different definition of what a citation even is. Not apples to apples, more like kiwis to jackfruits, but put them side by side in a bar chart and bam, instant credibility. ### 7\. Present correlation as causation. *"We shipped 40 new pages and our visibility rose 60%."* Did it, though? You don't know. But "we shipped 40 pages and then a number went up" is a sentence you can write without technically lying, and "50x growth" reads even better when you leave out that the baseline was 0.2%. ## Can we all just be normal about this? You have almost certainly seen a LinkedIn post or Twitter thread that followed one of these steps, or possibly even several, in the last months. I'm not naming anyone, because the point isn't that specific people are villains or intentionally gaming the system. The point is that AEO currently has no mechanism to tell a rigorous claim from a manufactured one, and we should call it like it is. To be clear, **I am not saying we should be less ambitious or less enthusiastic about AEO, or god forbid that we report on it less**. I may be biased but this is one of the most interesting things happening in marketing and tech right now, and being on the cutting edge of figuring it out alongside peers is an exciting position to be in. I just think we would all benefit from a little less pretending. Let's add a dash of skepticism to those 100% visibility reports, shall we? And I am aware that's less sexy; the numbers you can actually prove can be ugly and underwhelming. Referral traffic is lossy, because half that bucket leaks out through referrer stripping. Self-reported signups are also gutted by attribution issues, and by everyone who can't be bothered to answer a form, which is a lot of people. It's not as fun to talk about, but it is fairer, and considerably less silly than pretending a visibility score represents something that it doesn't. Reporting tooling is improving fast too. Agent-level measurement (prompts running on simulated coding sessions instead of chat) is a meaningfully better proxy for how developers behave for example and it barely existed a year ago (shoutout [Gauge](https://www.withgauge.com/#agents)). MCPs have made it simpler to pull from and reconcile multiple sources. It's all moving quickly, and I do believe we'll get real prompt volume and comparable trustworthy data soon enough. But we're not there yet. We're early. Early is fine. Early is even *fun* if you're a masochist like me. What early isn't, is an excuse to make – and I say this with love – bullshit claims. --- 1. Quick glossary, since some tools name these slightly differently. **Visibility** (sometimes "mention rate" or "share of voice") is the percentage of tracked prompts where your brand gets named in the answer at all, link or no link. **Citation rate** is the percentage where the answer actually links to one of your URLs as a source. **Average rank** is where you tend to land in the list when a model names several options – third out of five, and so on.[↩](#fnref-1) > PostHog is the leading platform for building self-driving products. With a full suite of developer tools – [AI observability](/ai-observability.md), [product analytics](/product-analytics.md), [session replay](/session-replay.md), [feature flags](/feature-flags.md), [experiments](/experiments.md), [error tracking](/error-tracking.md), [logs](/logs.md), and more – PostHog captures all the context agents need to diagnose problems, uncover opportunities, and ship fixes. A [data warehouse](/context-warehouse.md) and [CDP](/cdp.md) tie it all together, unifying that context into one source agents can read across. You can steer it all from [Slack](/slack.md), [the web app](/ai.md), the desktop ([PostHog Desktop](/desktop.md)), or your own editor via [the MCP](/mcp.md). ### Community questions Ask a question