AI

How to Measure Your Brand’s Visibility in AI Search Results

AI search visibility can be measured, but only partially and only if you understand what each data source actually shows. This guide covers the platform reporting available today, a workable KPI set, and a monthly framework that survives contact with a board meeting.

Measurement is where most AI search programmes fall apart. The work gets done, something probably improves, and nobody can demonstrate it. Meanwhile a tool vendor produces a dashboard with a confident number on it and the number turns out to be an estimate derived from sampling.

This is a solvable problem, but only if you accept from the outset that AI visibility measurement is partial. Some of it is properly instrumented, some is sampled, and some is genuinely unavailable. Building a framework that acknowledges the difference produces more useful reporting than pretending the gaps are not there.

Why Traditional Rankings Are Not Enough

Rank tracking answers a question that has become less complete. It tells you where a page sits in a list, but not whether the answer above that list already resolved the query, whether your brand was named in it, or whether a competitor was recommended instead.

Three specific blind spots matter.

Mentions without links. An AI answer that names your company without linking to your site is commercially valuable and completely invisible in analytics. This is arguably the most important AI visibility outcome and the least measurable one.

Position metrics that flatten AI features. In Search Console, each element of a results page counts as one position, so links appearing within an AI Overview register at the top position rather than at some deeper rank. This makes position data harder to interpret than it looks, and it is why a page can appear to be performing well on paper while its click behaviour says otherwise.

Answers you never appear in at all. Ranking data tells you about queries you already compete for. It says nothing about prompts where an assistant is recommending three suppliers in your category and you are not among them.

What You Can Actually Measure Today

Data sources fall into three tiers, and it is worth being explicit about which tier a number comes from whenever you report it.

Tier One: Platform Reported Data

The most reliable, and the most limited in coverage.

Google Search Console provides a generative AI performance report showing how content is performing in generative AI features across Search and Discover. This is first party data from the platform itself, which puts it in a different category from anything a third party tool can offer.

Bing Webmaster Tools added an AI performance report covering citations in Microsoft Copilot, AI generated summaries in Bing, and selected AI partner integrations. It reports total citations and average cited pages for a period, and breaks results down by grounding queries, described as the key phrases the AI used when retrieving the cited content, and by the pages cited. It reports frequency of citation rather than importance or position, and it provides no traffic or click through data. It also cannot be filtered by individual surface, so you cannot separate Copilot from Bing summaries.

Neither report tells you what was said about you. Both tell you that you were used.

Tier Two: Referral Data

Where an AI answer produces a click, that visit lands in your analytics like any other referral. OpenAI’s publisher documentation notes that publishers allowing OAI-SearchBot can track ChatGPT referral traffic through analytics platforms, and these referrals typically carry an identifiable source parameter. Perplexity and other assistants generally pass identifiable referrers too.

Set up a segment or channel grouping that captures these sources and treat it as its own acquisition channel. Volumes are usually modest. Behaviour is often notably different from organic search, since these visitors have already read a synthesised explanation before clicking.

Tier Three: Prompt Monitoring

This is sampling, not measurement, and it should always be labelled as such.

The method is straightforward. Define a fixed set of prompts a buyer might realistically ask, run them across the platforms that matter to you on a fixed schedule, and record whether your brand appears, whether it is cited, which competitors appear, and how you are characterised.

The critical methodological point is repetition. Generative systems are non deterministic, so the same prompt can return different sources on consecutive runs, as explained in our walkthrough of how AI citations work. A single check tells you almost nothing. A prompt run five times a month for six months tells you something real.

Tools exist to automate this at scale, and several are genuinely useful for workflow. Treat their output as directional. Google warns site owners to be wary of third party tools claiming to use internal Google metrics, on the straightforward basis that none have access to internal ranking or AI systems.

A Workable KPI Set

Metric What it tells you Source Reliability
Generative AI performance (clicks and impressions) How your content performs in Google’s AI features Search Console Platform reported
Total citations How often your pages are used in Copilot and Bing AI answers Bing Webmaster Tools Platform reported
Grounding queries The expanded queries that retrieved your content Bing Webmaster Tools Platform reported
AI referral sessions Traffic arriving from assistants Analytics Measured, undercounts mentions
AI referral conversion rate Commercial quality of that traffic Analytics Measured, small samples
Prompt appearance rate Share of tracked prompts where you appear at all Prompt monitoring Sampled
Citation rate Share of appearances that include a link to you Prompt monitoring Sampled
Competitive share of answers How often you appear relative to named competitors Prompt monitoring Sampled
Mention sentiment and framing How you are described when named Manual review Qualitative
AI crawler activity Whether AI search crawlers are reaching your pages Server logs Measured

That last one is underrated. Server logs will tell you unambiguously whether OAI-SearchBot and other AI crawlers are fetching your important pages and how often. It is the cheapest diagnostic available and the first place to look when visibility is flat.

Branded and Non Branded Prompts

Split your tracked prompts into two groups, because they answer different questions.

Branded prompts ask about you directly. “What does BrandingX do”, “is BrandingX any good”, “BrandingX pricing”. These measure whether the picture being formed about your company is accurate and fair. Errors here are urgent, because they are being repeated to people who already know your name.

Non branded prompts ask about your category. “Best branding agencies in Manchester for challenger brands”, “how much does a rebrand cost in the UK”. These measure whether you are entering consideration sets at all. This is the growth metric.

Most businesses discover their branded picture is broadly accurate and their non branded presence is close to zero. That gap is the work.

Building the Measurement Framework

Define twenty to forty prompts. Enough to be representative, few enough to review properly. Cover the buying journey: problem awareness, solution comparison, supplier selection, and your branded set.

Fix the platforms. Pick the three or four that matter for your audience rather than tracking everything. For most UK B2B businesses that means ChatGPT, Google AI Mode and AI Overviews, and Perplexity. Consumer businesses may weight differently.

Take a baseline. Run everything before you change anything. Without this, you will spend a year arguing about whether things improved.

Set the cadence. Monthly for prompt monitoring and platform reports. Weekly is noise. Quarterly loses the pattern.

Record the qualitative detail. Screenshot answers where you appear and where a competitor is preferred. Over six months this archive explains more than the numbers do.

An Example Monthly Reporting Framework

The figures below are illustrative only. They are not benchmarks, and no reliable public benchmarks exist for these metrics.

Section Contents Illustrative example
Platform data Search Console generative AI performance, Bing citation totals 142 total citations, up from 118
Referral performance Sessions, conversion rate, top landing pages by AI source 96 sessions, 4.2 per cent conversion
Prompt visibility Appearance rate and citation rate across tracked prompts Appeared in 11 of 32 prompts
Competitive position Share of answers versus three named competitors Third of four on non branded prompts
Framing review How the brand was described, with any inaccuracies flagged Two answers cited superseded pricing
Technical health AI crawler activity, indexing, snippet eligibility OAI-SearchBot fetched 340 URLs
Actions What changed last month, what changes next Pricing page updated, comparison content briefed

Keep the actions section. Reporting without a decision attached becomes a monthly ritual nobody reads.

Connecting Visibility to Commercial Outcomes

This is the honest weak point, and it is better to name it than to fabricate attribution.

Direct attribution works only where a click occurred. Everything else is influence you cannot trace. A prospect who read an AI answer naming your firm, then searched your brand a fortnight later, appears in your data as direct or branded organic traffic.

Two partial workarounds. First, watch branded search volume alongside AI visibility, since a rise in branded queries without a corresponding campaign is at least consistent with increased mention exposure. Second, add a question to your enquiry form asking how people came across you. Self reported attribution is imperfect, but “ChatGPT recommended you” is a data point you will not get anywhere else.

The Limitations Worth Stating in Every Report

  • Prompt monitoring samples a non deterministic system and cannot be treated as a census.
  • Mentions without links are invisible in analytics and are probably your largest untracked outcome.
  • Platform reporting covers Google and Microsoft properly and most other assistants barely at all.
  • Personalisation and location affect answers, so your results are not universal.
  • No third party tool has access to internal platform systems, whatever the interface implies.

Stating these once in a methodology note protects the credibility of everything else you report. A measurement framework that admits its edges is more persuasive than one that does not, particularly to a finance director who has seen confident marketing numbers before.

Frequently Asked Questions

What is a good AI visibility score?

There is no established benchmark, and any figure presented as an industry standard is invented or derived from a vendor’s own sample. Measure your own trajectory and your position relative to named competitors instead.

How often should I run prompt monitoring?

Monthly, with each prompt run several times per cycle. Because generation is non deterministic, frequency of sampling matters more than frequency of reporting.

Can I see which ChatGPT conversations mentioned my brand?

No. OpenAI does not provide this reporting to most publishers. Referral traffic and prompt sampling are the available proxies.

Are AI visibility tools worth paying for?

They can save considerable manual effort at scale, which is a legitimate reason to buy one. They do not have privileged access to platform systems, so judge them on workflow rather than on claimed accuracy.

Should I replace my SEO reporting with AI visibility reporting?

No. Add a section rather than swapping the report. Conventional search still accounts for the majority of measurable revenue for most UK businesses, as discussed in our comparison of GEO and SEO.


Published by BrandingX UK.


DS

Daniel Sullivan

Part-time blogger and full-time SEO leader at a leading web, app and software development company in Rickmansworth, UK, driving organic growth and digital visibility.