Measuring AI visibility without fooling yourself

The measurement layer. Metrics worth tracking, how to design a question panel, what run-to-run volatility does to your conclusions, and how to report honestly.

Part 4 of 4 · 17 min read · Updated August 25, 2026

The short answer

Measure AI visibility by running a fixed set of real buyer questions on a schedule across the engines your buyers use, then tracking mention rate, share of voice, citation mix, and uncontested questions over a window rather than per run. Single runs are unreliable: in our monitoring, 17.3% of consecutive runs disagreed with the one before them even though nothing on the site had changed.

In this part
  1. The four numbers worth keeping
  2. Designing the question set
  3. One run is a reading, not a result
  4. A cadence that matches the decision
  5. Connecting visibility to something the business recognizes
  6. A monthly report that holds up
  7. What good looks like after a quarter
  8. Where PromptScout fits in this part
  9. Where to go from here

The four numbers worth keeping

AI visibility dashboards produce dozens of metrics, most of which move without meaning anything. Four survive contact with a real decision.

Metric What it is What it tells you to do
Mention rate Share of answers, across your fixed question set, that name your brand Whether you are present at all, and whether that is trending
Share of voice Of the answers that name any brand, the share that name yours Whether you are gaining or losing ground against named rivals
Citation and source mix Which domains and which kinds of source the engines actually used Where to invest off-site, and which of your pages are doing the work
Uncontested questions Answers that named nobody at all The cheapest openings in the category, before anyone owns them

Mention rate is the headline. Share of voice is the competitive read, and you need both: mention rate can fall while your position improves, if the whole category got less brand-heavy that month. Citation mix tells you why the other two moved. Uncontested questions are your opportunity backlog.

Worked example. You track 30 questions on ChatGPT. Over four weekly runs that is 120 answers. You are named in 42 of them, so mention rate is 35%. Of those 120, 90 named at least one brand, so share of voice is 42 ÷ 90 = 47%. The 30 answers that named nobody are your backlog. Next month, mention rate falling to 32% while share of voice rises to 51% means the category got less brand-heavy, not that you lost ground.

Track all four per engine, never averaged. Mention rate ran from 38% to 67% across five engines on the same questions in the same window; an average across that range describes nothing that exists.

Skip sentiment scoring. At the volumes most teams run it is noise. A model summarizing a category is neither praising nor criticizing you in a way a five-point scale captures, and the score moves between runs for the same reason everything else does. Read the answer text when a description is wrong, and fix the source behind it.

Metrics that sound useful and are not

  • "AI visibility score." Any single composite number built from weightings a vendor chose. It cannot be audited, it moves for unexplainable reasons, and nobody can act on it.
  • Total mentions. Grows when you add prompts. Going from 40 to 60 mentions means nothing if you also went from 20 to 35 questions. Use rates, not counts.
  • Estimated AI traffic. Modelled numbers presented with a confidence they do not have.
  • Position in the answer. Tempting, because it feels like rank. In practice it is unstable between runs and the ordering rarely means what a ranked list means.

Designing the question set

Everything downstream depends on the questions, and this is the part teams rush and regret two months later, when the trend line turns out to describe the panel rather than the market.

Three failure modes account for most bad panels:

  • Branded questions. A panel of questions with your name in them will show a flattering mention rate and tell you nothing, because you were the subject of the question.
  • Panels that are too large. Expensive to run, slow to react, and no more reliable than a well-chosen smaller set.
  • Panels that change. A question set that quietly gains and loses members every month is not a trend line. It is a series of unrelated snapshots.

Five question shapes to cover

Swap in your own category and the panel writes itself.

Shape Example What it tests
Category What are the best project management tools for a small agency? Whether you are considered part of the category at all
Comparison Asana vs Monday for a ten-person team How you are framed against the alternative buyers already know
Constraint Project management software under $10 per user with time tracking Whether your specifics are findable, not just your positioning
Use case How do agencies bill clients for retainer hours? Whether you show up in the problem, before the category is named
Alternatives Alternatives to Trello for client work Whether you are the answer when someone is already leaving a rival

Aim for roughly a third category questions, a third comparison and alternatives, a third constraint and use case. Keep branded questions to a small slice and read them separately: they test description accuracy, not visibility.

Rules that keep a panel honest

  • Use the words buyers use, not your internal category name.
  • Fix the panel and leave it fixed. Add questions deliberately, note the date you added them, and never quietly swap one out.
  • Twenty to forty questions is enough for most categories. Precision comes from repetition, not from panel size.
  • Include the questions you are losing. The instinct is to track the ones where you already appear, because the chart looks better. Those have the least to teach you. A panel that includes the comparisons you currently lose is the one that produces work worth doing.
  • Write down why each question is in the panel. In six months somebody will ask, and "it seemed important" is not a defense.

One run is a reading, not a result

We took every case where the same question ran on the same engine twice in a row, and asked how often the second run disagreed with the first about whether the brand was mentioned. Nothing on any of the sites changed in between.

PromptScout monitoring data

How often the same question changes its mind

Share of consecutive run pairs where the same prompt on the same engine flipped between mentioning the brand and not.

  • AI Overviews26.0%

    275 flips across 1,058 consecutive pairs

  • ChatGPT19.0%

    194 flips across 1,020 consecutive pairs

  • Gemini16.7%

    170 flips across 1,020 consecutive pairs

  • Bing Copilot10.5%

    31 flips across 294 consecutive pairs

  • Perplexity9.3%

    99 flips across 1,064 consecutive pairs

of consecutive runs disagreed

Across all five engines, 17.3% of consecutive runs disagreed with the one before them, with nothing changed on the site in between. A single run is a reading, not a trend.

Source: PromptScout monitoring, May 28 – August 25, 2026. 5,436 completed answers across 135 tracked prompts. Aggregated across all monitored brands.

Across all five engines, 17.3% of consecutive runs disagreed. On Google AI Overviews it was 26%. That is your noise floor. Any single-run change smaller than it is indistinguishable from nothing happening.

What this rules out

  • Reporting a one-run change as a result.
  • Claiming that a page you shipped on Tuesday caused a mention on Wednesday.
  • Ranking two competitors against each other on a single day's answers.
  • Any alert that fires on a single missing mention. It will cry wolf roughly one time in six.

What it does not rule out

The thing that actually matters. A mention rate moving from 30% to 55% across four weeks of a fixed panel is well outside this noise floor. A competitor appearing in 20% more answers over a month has genuinely gained ground. Volatility is an argument for windows and cohorts, not for giving up.

Three rules follow:

  1. Compare windows, not runs. Four weeks against the previous four weeks, not Tuesday against Monday.
  2. Watch cohorts, not single questions. Group the questions a change was meant to affect and read them together.
  3. State the noise floor in the report itself. Once, near the top. It stops a two-point move being read as a result by somebody who was not in the room.

A cadence that matches the decision

Different decisions need different windows. Checking a daily chart to make a weekly decision is how teams end up chasing noise.

Window What you are looking for What you decide
Each run Something material moved: a new competitor, a lost mention on a key question, an unfamiliar source Whether anything needs inspecting today
Weekly Which gap is now best supported by evidence The one piece of work to do this week
Monthly Mention rate, share of voice, and source mix across the full window Whether the strategy is working, and what to report
After a change The specific cohort the change was meant to affect, over the following runs Whether to keep, extend, or abandon the approach

State the caveat on that last row out loud in your report: you are observing a correlation over a watch window, not proving causation. Engines update, competitors publish, and the index shifts underneath you. The honest phrasing costs nothing and protects the numbers that do hold up.

Connecting visibility to something the business recognizes

Mention rate is the right operating metric and the wrong boardroom metric. Two signals bridge the gap, and both are worth wiring up early.

Referral traffic from assistants

Assistant referrals arrive with identifiable referrers and can be segmented in analytics. Volumes are small by design, because the whole point of an answer is that it answers, so read the trend and the behavior of those visitors, not the raw count. AI-sourced sessions arrive later in the decision, which usually shows up as shorter paths to a demo or signup.

Crawler activity in your server logs

Which AI crawlers fetch which pages, and how often, tells you what the engines consider worth reading. A page that never gets fetched cannot be cited, and that is a diagnosis you can act on the same day. Rising crawl frequency on a section you just rewrote is the earliest available signal that the rewrite landed.

Setting expectations before someone else sets them

Do this in the first report, not the third. Pew found that people click a source cited inside an AI summary on roughly 1% of visits, and Cloudflare's crawl-to-refer measurements show AI platforms fetching far more pages than they return visitors for.

Report AI visibility as a traffic channel and it will look like it is failing. Report it as presence in the recommendation, with referral traffic as a secondary indicator, and you are describing what is actually happening. Same numbers, different framing, and the framing decides whether the work survives its second quarter.

A monthly report that holds up

Six things, in this order. This structure survives someone senior pushing back on it.

  1. State the panel. How many questions, which engines, over what window, unchanged since when.
  2. State the noise floor. One sentence. "Runs disagree with each other about one time in six, so we read windows rather than single runs."
  3. Lead with mention rate and share of voice for the window, per engine, against the previous window.
  4. Show the answers behind the biggest movements. Two or three verbatim excerpts. A reader who can check your work trusts the number more than one who cannot.
  5. Separate the two kinds of absence. Competitive losses and uncontested questions are different line items with different owners.
  6. Say what you changed and when, and describe the relationship as observed rather than proven.

Then add the section most reports omit: what you were wrong about last month. It costs a paragraph and buys more credibility than any chart.

What good looks like after a quarter

Roughly what a team doing this properly has after three months:

  • A fixed panel of 20 to 40 questions, unchanged since week two, with a documented reason for each.
  • A mention rate per engine with twelve weekly readings behind it, so the trend is legible through the noise.
  • A short list of pages that are reliably cited, and a shorter list of competitor pages that beat you repeatedly.
  • Two or three uncontested questions converted into questions you now win.
  • A clear read on which engines matter for your buyers, and permission to stop spending on the ones that do not.

What they do not have, and should not claim, is proof that any individual change caused any individual mention. That is not available at this sample size from any tool, and a vendor offering it is selling you a story.

Where PromptScout fits in this part

PromptScout is built around this cadence. A fixed prompt panel runs on a schedule across the engines. AI Visibility Changes surfaces what moved since the last run, so run-level inspection takes minutes rather than an afternoon. Weekly Tasks record the reason and the watch window alongside each piece of work. Reports read the window rather than the latest run, and use the observation-not-proof phrasing above.

For the cadence question on its own, how often to review AI search performance goes deeper on run, weekly, monthly, and post-change reviews. The weekly workflow documentation covers the operational loop.

Where to go from here

Four things worth doing this week, in order:

  1. Check robots.txt for a blocked search crawler.
  2. curl your three most important pages and look for the sentences you care about.
  3. Write down twenty questions your buyers actually ask.
  4. Run them once to see where you stand.

Everything else in this guide is refinement on top of those four steps.

Common questions

How often should I check AI visibility?
Inspect each completed run for material movement, choose work weekly, read the trend monthly, and add a check after any significant site or content change. Single runs vary too much to support a decision on their own.
Why do AI answers change when nothing on my site changed?
Generated answers are not deterministic, and retrieval shifts as the underlying index and the engines themselves change. In our monitoring, 17.3% of consecutive runs of the same question on the same engine disagreed about whether the brand was mentioned, rising to 26% on Google AI Overviews.
How many prompts should I track?
Twenty to forty real buyer questions is enough for most categories. Reliability comes from repeating a fixed panel over time rather than from tracking a larger set once, and a panel that changes between periods cannot produce a trend.
Can I prove that a content change caused an AI mention?
Not from a single run, and rarely with certainty at all. The defensible claim is that a specific cohort of questions moved over a watch window after a specific change, stated as an observation rather than as proof of cause.
What is a good mention rate?
There is no universal benchmark, because it depends entirely on how competitive your question set is and how many brands the category supports. Your own trend over a fixed panel is the only meaningful comparison. A mention rate that rose from 30% to 55% over six weeks tells you far more than any industry average.
Should I track sentiment in AI answers?
Usually not as a chart. At the volumes most teams run, sentiment scoring on generated answers is noise. Read the answer text when the description of your product is wrong, and fix the underlying source; that is more useful than a five-point scale that moves for no reason.

Sources cited in this part

Primary sources are published by the party that runs the system. Third-party studies are labeled as such, because vendor research in this field disagrees more than the headlines suggest.