
How to Measure AI Search Visibility Without Pretending a Screenshot Is a KPI
Manual screenshots are useful evidence, but they are not a reporting system. Here is a reproducible way to track mentions, supporting links, competitors and real business outcomes.
- 01Separate Google search performance from manual AI-answer observations: Search Console cannot break AI Overview traffic out from ordinary Web reporting.
- 02Track a fixed, versioned query set by topic and buyer intent, then report mention rate, supporting-link rate, competitor share and consistency over time.
- 03Treat custom AI metrics as directional, and connect them to first-party outcomes such as qualified visits, leads and conversions.
The measurement problem
A screenshot can prove that a brand appeared in one answer at one time. It cannot tell a client whether visibility improved, whether competitors moved, whether the result repeated, or whether anyone clicked.
That is the gap in most AEO and GEO reporting. The agency runs a few prompts, takes the best screenshot and calls it visibility. The client cannot reproduce the method and cannot tell whether next month's answer is better.
A useful system needs two layers:
- First-party search performance, from Google Search Console and Analytics.
- A controlled AI-answer panel, used directionally to observe mentions and supporting links across repeated prompts.
Neither layer is complete alone.
What Google actually reports
Google says AI Overview and AI Mode traffic is included inside Search Console's ordinary Web search type. There is no documented separate AI Overview citation count or share-of-voice metric. The Search Analytics API exposes dimensions such as query, page, country, device, date and search appearance, but AI Overview is not documented as its own dimension.
That means Search Console can tell you:
- which queries produced impressions and clicks;
- which pages gained or lost visibility;
- how CTR and average position changed;
- whether performance differs by country or device;
- whether a page's trend improved over consistent date ranges.
It cannot tell you, on its own, that a particular impression came from an AI Overview source card rather than an ordinary web result.
Use Analytics for conversions, engagement and time on site. Without correctly configured analytics and conversion events, search visibility cannot be tied reliably to leads or AI-referral quality.
Build a fixed query set before you test
Do not choose prompts after seeing the answers. Define them first.
A small business benchmark can start with 20 to 30 queries across five groups:
| Group | What it measures | Example for a mattress business |
|---|---|---|
| Brand | Whether the engine understands the entity | what is Lullaflex? |
| Category | Inclusion in broad recommendations | best cooling mattress singapore |
| Problem | Visibility for customer pain | mattress for hot sleepers without aircon |
| Comparison | Decision-stage visibility | single vs super single for hdb room |
| Local/action | Ability to complete a task | bedroom layout planner singapore |
Google's AI features can use query fan-out, issuing related searches across subtopics. So measure both exact prompts and topic clusters. A page can surface through a related subquestion even if it does not repeat the user's exact wording.
Version the query set. If you add or remove prompts, record the change rather than quietly moving the goalposts.
Record the conditions, not just the answer
For every manual check, record:
- exact query;
- platform and feature: Google AI Overview, AI Mode, ChatGPT Search or Perplexity;
- date and local time;
- country and language;
- mobile or desktop;
- account state and browser profile;
- whether the AI feature appeared;
- whether the brand appeared in the answer;
- whether a link to the brand's domain appeared;
- exact source URL;
- named competitors;
- screenshot or exported answer.
Incognito is useful because it reduces some account history. It is not neutral. Google can still vary results by location, language, device, query and time. ChatGPT and Perplexity can also return different answers across runs.
Do not automate Google result scraping with Playwright. Google's spam policy explicitly covers automated rank-checking queries. Use Search Console for Google performance and keep Google AI-answer sampling manual unless you have an approved source. Use platform APIs only where their terms permit it.
Use metrics that describe what was observed
The metrics below are custom reporting metrics. They are not Google metrics, and they should be labelled that way.
Mention rate
The percentage of tested answers that name the brand.
answers naming brand ÷ valid answers tested
Supporting-link rate
The percentage of tested answers that link to the brand's domain. Keep this separate from mentions: an engine may name a brand without linking it, or link a page without naming the brand prominently.
answers linking to brand domain ÷ valid answers tested
Source-page distribution
Count which URLs appear. This shows whether visibility depends on one page or spans products, tools, service pages and guides.
Competitor share of observed mentions
For each topic cluster, count the brands named across the valid answers.
brand mentions ÷ all tracked-brand mentions
Call this observed share, not market share. The answer panel is a controlled sample, not the whole market.
Answer consistency
Run each prompt three times on the same day under the same recorded conditions. Report how often the brand or URL repeats. One of three is a weak observation. Three of three is more consistent, but still not permanent.
Accuracy and sentiment
Tag whether the answer is positive, neutral or negative, then record factual errors separately. Do not collapse sentiment and accuracy into one score. A positive answer with the wrong price is still a bad result.
A minimum viable protocol for a client
A 30-query, four-platform, three-run benchmark creates 360 answer checks. That is too much for an early-stage agency to run manually every week. It also creates false precision from a small, volatile sample.
Start smaller:
- 15 prompts, grouped into five topic clusters;
- three platforms: Google AI Overview or AI Mode where available, ChatGPT Search, Perplexity;
- one primary run per prompt each month;
- three repeats only for prompts where the client or a tracked competitor appears;
- one fixed country, language and device, plus a small mobile control set;
- a frozen baseline for 90 days.
This is 45 primary observations, with repeats focused where they add information. It is affordable to run and honest about variance.
What the client report should show
A monthly report should include:
- Search Console clicks, impressions, CTR and average position by query group and page.
- New and lost organic queries.
- AI-answer mention and supporting-link rates by topic cluster.
- Competitors gaining or losing observed share.
- URLs that surfaced, not just the brand name.
- Factual errors or negative descriptions found in answers.
- Qualified visits and conversions from Analytics when configured.
- A change log: what was published, updated or promoted, and when.
Do not claim that one change caused a movement when several changes launched together. Search effects can take days or months, demand changes, and rankings fluctuate. Record the rollout and use careful language: after the change, not because of the change, unless the design isolates the variable.
Build the reporting stack in the right order
A reliable reporting system should separate three different forms of evidence:
- Search Console: clicks, impressions, CTR, queries, pages and average position.
- Analytics: qualified visits, engagement, enquiries and conversions.
- A fixed AI-answer panel: observed mentions, supporting links, competitors and factual errors under recorded conditions.
Search Console totals should come from the property-level aggregate, not from summing grouped page rows. Different dimensions can aggregate search features differently, so a page-row sum is not always the true site total.
Start by validating Search Console and Analytics, then freeze a 15-prompt panel for 90 days. Report observed AI mentions and links separately from organic performance, and keep screenshots as evidence rather than treating them as the KPI.
Why screenshots still matter
Screenshots are useful evidence when they include the query and conditions. Our Lullaflex screenshot from 12 August 2026 shows the Bedroom Layout Planner beside a Google AI Overview for one incognito query. That observation belongs in a case study.
It does not become a percentage until it is part of a defined sample. It does not become traffic until a user clicks. It does not become ROI until the visit produces a useful outcome.
That sequence is the difference between proof and reporting theatre.
Frequently asked questions
Can Search Console tell me exactly how many times Google AI Overview cited my site?
No. Google currently reports AI-feature traffic inside the Web search type and does not document a separate citation-count or share-of-voice report.
Is incognito mode enough for an unbiased test?
No. It reduces some account history, but location, language, device, query wording, time and changing search systems still affect the answer.
Should we run every query every day?
Usually not. Daily manual checks create cost and noise before a site has enough visibility to learn from them. A monthly fixed panel with focused repeats is a better starting point.
Can we automate Google AI Overview checks?
Do not scrape Google Search for automated rank checking. Google's policy explicitly prohibits automated queries without permission. Use Search Console for Google performance and manual sampling for directional AI-answer observations.
What is the difference between a mention and a citation?
A mention names the brand. A supporting link sends users to the brand's page. Track them separately.