Updated
Updated What changed on
- Added the spread between published AI Overview prevalence estimates, which run from Pew's 18% on real browsing data to Ahrefs' 48% on its keyword panel.
- Added Pew Research's click measurement as the benchmark for what a citation is actually worth in traffic.
How do I measure if GEO is working?
The 30-second answer
Three metrics worth tracking: citation frequency (what % of your priority queries cite you), share-of-voice (your citations divided by all named-brand citations across the same queries), query coverage (how many distinct priority queries cite you at all). Sample weekly or fortnightly - daily is noise. Pair with GA4's new AI Assistant channel as a referrer-based leading indicator. Ignore total AI mentions, raw bot crawl counts, and composite vendor "AI visibility scores" with no methodology behind them.
Why the headline numbers disagree with each other
of Google searches trigger an AI Overview, depending entirely on whose panel you read.
How it was measured: Pew measured 18% on real-user browsing data. Ahrefs measured 48% of its keyword set in March 2026. Similarweb reported 43%, up from 15% a year earlier, via TechCrunch on 27 July 2026. The panels sample different query mixes, so the spread is a methodology artefact. Anyone quoting a single figure as the number is overstating what is known.
Sources disagree on this one. The range is the honest answer.
is how often people clicked a search result when an AI summary was present, against when it was not. Roughly half the click rate.
How it was measured: Browsing data from 900 US adults who agreed to share their activity, covering March 2025. US only, one month, and now over a year old. It remains the largest published real-user measurement rather than a keyword-panel estimate.
of visits to a page carrying an AI summary resulted in a click on a source cited inside that summary.
How it was measured: Same 900-person US panel, March 2025. This is the figure that matters most for anyone selling AI citation as a traffic channel.
The three metrics that matter
Measurement is the fourth stage of the wider playbook in my GEO field guide; the section below is the practical detail for that stage.
1. Citation frequency - of your 10-20 priority buyer queries, what percentage actually cite you in the AI engine answer? The headline number. Weekly or fortnightly sample.
2. Citation share-of-voice - your citations divided by all named-brand citations across the same queries. Tells you not just whether you appear, but whether you appear more than competitors. The metric that drives competitive insight.
3. Query coverage - how many of your priority queries cite you at least once across all monitored engines. Tells you breadth - are you a one-engine wonder or showing up across the stack?
The metrics worth ignoring
- Total AI mentions - includes retrieval-without-citation noise (ChatGPT pulls Reddit for context constantly without citing).
- Raw AI bot crawl counts - bots crawl widely, citations are the actual signal. Crawl volume tells you nothing about whether you are being quoted.
- Composite "AI visibility scores" from vendor dashboards without published methodology - usually inflated to look comparable to SEMrush scores.
- Daily readings - AI outputs vary per session, per prompt phrasing, per user context. Daily creates noise without insight.
Pair with GA4's AI Assistant channel
Google added AI Assistant as a default channel group in GA4 in May 2026. It buckets visits from chat.openai.com, perplexity.ai, claude.ai, gemini.google.com and copilot.microsoft.com. Useful as a leading indicator but it undercounts real influence because many AI surfaces strip the referrer header. Trend matters more than absolute count. More on the GA4 AI Assistant channel.
What a credible monthly report includes
- Citation frequency this month vs last (on the same 10-20 priority queries)
- Share-of-voice vs 3-5 named competitors
- Per-engine breakdown (ChatGPT, AI Overviews, Perplexity etc - each engine separately)
- Query coverage trend
- GA4 AI Assistant referral volume + trend
- Two-line commentary: what shifted, why we think it shifted, what we are doing next
Tools that produce this
Free baseline: the AI Search Visibility tool on this site. Paid trackers for ongoing monitoring: Profound (enterprise), AIClicks ($59/mo mid-market), Otterly.AI ($29/mo founder-led). Full comparison here. GA4 AI Assistant is free and automatic.
Two measurement methodologies worth naming
Reverse prompting - rather than guessing what buyers ask, take a known buyer journey or sales call and reverse-engineer the prompts that would produce that conversation. Then run those prompts. AnalyticaHouse's framework calls this out as the main alternative to keyword-list-style query construction. It produces query lists that map cleanly to actual buyer intent.
AI benchmarking - the same prompt run repeatedly across 4-6 engines (ChatGPT, AI Overviews, Perplexity, Claude, Gemini, Copilot), citations logged per engine per run. Reveals which engines cite you, which do not, and where the gap is. The honest version of "AI visibility score".
Sentiment as a 4th metric (when reputation matters)
For brands in regulated or reputation-sensitive sectors (financial services, law, healthcare, food, recruitment), citation count is not enough. Michael Brito's piece argues for sentiment analysis on the answer text itself: of the times your brand is mentioned, what tone is the AI describing you in? Positive, neutral, negative? Especially useful for catching reputational issues that competitors have surfaced into the AI graph (negative Reddit threads, regulatory action coverage, contested reviews).
What you cannot measure yet
Search Engine Land's 2025 piece and iPullRank's "Measurement Chasm" both list the genuinely unmeasured things as of mid-2026:
- Prompt volume - the AI equivalent of search volume does not exist. We do not know how often any one prompt is actually being asked.
- Why the model picked you - LLMs do not surface citation reasoning. We see the result, not the why.
- True share of brand discovery influence - if a buyer reads an AI answer with your brand named and then searches for you in Google a day later, traditional attribution credits Google, not AI. The actual influence is under-counted.
- Per-prompt reproducibility - same prompt, different outputs on different days/sessions. Sample size mitigates this; perfect tracking does not solve it.
These gaps argue for measurement honesty - report the trackable metrics confidently, flag the gaps explicitly, do not invent precision that does not exist.
More on GEO measurement.
What is the single best GEO metric?
Citation frequency on priority buyer queries. Sample weekly or fortnightly: of your top 10 buyer queries, how many actually mention you in the AI engine answer? Track the trend, not the absolute number.
Should I track AI Assistant in GA4?
Yes as a leading indicator, no as the only metric. GA4 AI Assistant is referrer-based and many AI surfaces strip the referrer, so the number undercounts. Useful for trend; pair with citation tracker data for the full picture.
How often should I report on GEO?
Monthly is the right cadence for stakeholder reporting. Weekly or fortnightly is right for the tracking team because AI outputs vary per session. Quarterly is too slow if you are actively running work.
What numbers should I ignore?
Total AI mentions (includes retrieval-without-citation noise). Raw AI bot crawl counts (bots crawl, citations are the signal). Composite vendor "AI visibility scores" without methodology transparency. Daily readings (noise).
Can I measure without a paid tool?
Yes for baseline. The free AI Search Visibility tool on this site scores any brand against the major engines. For ongoing tracking at scale, a paid tracker (Profound, AIClicks, Otterly) is more practical than running 200 manual queries a month.
What is reverse prompting?
The technique of building your tracked query list from real buyer journeys (sales calls, support tickets, qualifying questions) rather than guessing what buyers would type. AnalyticaHouse names it as the main alternative to keyword-list construction. Produces a query list that maps cleanly to actual buyer intent rather than to what marketers think buyers ask.
Should I track sentiment of AI answers about my brand?
Worth doing in reputation-sensitive sectors (financial services, law, healthcare, food, recruitment). Citation count tells you visibility; sentiment tells you whether the AI is recommending you, neutrally listing you or warning against you. Especially useful for catching reputational issues that competitors or contested coverage have pushed into the AI graph.
What can\'t I measure yet about GEO?
Four genuine gaps as of mid-2026: prompt volume (the AI equivalent of search volume does not exist), why the model picked your page over a competitor (LLMs do not surface citation reasoning), true share of brand-discovery influence (attribution credits whatever channel the user converted on, not the AI that triggered the awareness), and per-prompt reproducibility (same prompt different outputs across days/sessions). Sample size mitigates the last one but the others remain unsolved.
Jason Burns
Independent UK SEO, GEO and AI consultant. 17 years in search. Portfolio includes 3M, BlackRock, Unilever and E.ON. Owner of SEO Moves Ltd since 2014.
More about Jason →Want me to set up a baseline reading for your brand?
Send your 10 priority buyer queries and 3 competitor names. I run them across 4 AI engines, 3 times each, and send back the citation frequency, share-of-voice and query coverage numbers as a one-page baseline report.
Request the baseline