← Back to blog
Mara Ellison

What Is AI Visibility? Metrics, Measurement and What Good Looks Like

AI visibility is how often AI systems mention, cite, and recommend your brand. The five metrics that define it and how each one is measured.

MeasurementStrategyCitations
On this page

AI visibility is how often, how prominently, and how accurately AI systems name your brand when people ask questions in your category. It is a category of measurement rather than a single number, made of five metrics: mention rate, citation rate, share of voice, position, and sentiment. None of them come out of a rank tracker, which is why the term exists at all.

The rest of this page defines each metric, explains how it is measured and mis-measured, and is honest about the part most articles skip: there are almost no trustworthy industry benchmarks to compare yourself against.

Why the term exists

Search visibility had a settled meaning for twenty years. You ranked, or you did not, and the number was stable enough to put in a weekly report.

Answer engines broke three assumptions behind that. There is no ranked list to hold a position in, because the engine composes one response. There is no stable result, because identical prompts return different answers. And there is frequently no click, so the analytics that historically stood in for visibility record nothing at all.

What replaced it is a set of rates measured over repeated samples. That shift from a position to a rate is the whole conceptual move, and everything below follows from it. The wider discipline built on top of these metrics is covered in the AEO pillar.

The five metrics

Framework of five AI visibility metrics: mention rate, citation rate, share of voice, position and sentiment, each showing the question it answers and the work it points at
Five metrics, five different questions. A single blended score answers none of them.

Mention rate

The question: across repeated runs of a prompt, how often does the answer name your brand?

How it is measured: run the prompt N times per window, count the runs naming you, divide. A 60% mention rate means you appeared in three of five runs, and it is meaningless without that N attached.

How to read it: this is the headline number and the one most sensitive to sampling. It moves for real reasons (a new comparison page, a competitor's launch) and for structural reasons (model updates, retrieval variance), which is why trend lines over weeks beat any single reading.

The common mistake: measuring from a personal account. Assistants personalize on history, so your own logged-in ChatGPT reports your relationship with your brand, not the market's.

Citation rate

The question: how often does your domain appear among the sources an answer links?

How it is measured: count answers citing your domain, divide by answers where citations appeared at all. Which denominator you use matters enormously, because engines differ wildly in how often they cite anything.

How to read it: citation and mention are different events and come apart constantly. Being recommended without a citation means the engines believe you on other people's evidence. Being cited without a recommendation means your content is useful but your brand is not in the consideration set. Those two situations need opposite work, and what is a good AI citation rate goes through both.

The common mistake: tracking only your own domain. The cited-domain list for a prompt is the more valuable artifact, because it names the pages shaping your category's answers whether or not you appear on them.

Share of voice

The question: of all brand mentions across your tracked prompts, what proportion are yours?

How it is measured: count mentions of you and of each named competitor on the same runs, then express yours as a share. Same runs is the load-bearing phrase; comparing your Tuesday numbers to a competitor's Thursday numbers measures the weather.

How to read it: this is the metric that makes a mention rate mean something. A 40% mention rate is dominant in a fragmented category and dire in one where the leader sits at 90%. Our AI share of voice guide covers scoring competitors properly.

The common mistake: tracking your brand alone, then having no idea whether 40% is good.

Position

The question: when you are named, where in the answer do you land?

How it is measured: score each mention as first recommendation, mid-list, or afterthought. Some tools report an average position; the ordinal is usually more useful than the average.

How to read it: being the first brand named in a short answer is the closest thing to position one in this medium. Being fifth in a list of "other options" is technically a mention and practically invisible. Position also matters less on engines that answer in tight lists than on those that write conversationally.

The common mistake: treating position as the headline. On most engines the answer names three or four brands, so presence is the binary that matters and ordering is the tiebreak.

Sentiment and accuracy

The question: when the answer describes you, is it right, and is it favourable?

How it is measured: classify each mention as positive, neutral, or negative, and separately check factual claims about pricing, features, and positioning.

How to read it: a mention that says "X is cheaper but limited for teams over fifty" is a mention you would rather fix than celebrate. Accuracy problems are often more urgent than sentiment ones, because a confidently wrong price is repeated to every buyer who asks. Where the negativity comes from and how to work on it is covered in fixing negative brand sentiment in AI answers.

The common mistake: counting mentions without reading them, which turns a reputational problem into a green number on a dashboard.

Every rate needs its run count

Every metric above is a rate, and a rate is only as trustworthy as the sampling behind it.

We have first-party numbers on how unstable that sampling is. In our August 2026 answer study, 30 prompts were run five times each through ChatGPT with web search enabled. Only 68.8% of brands that appeared in at least one run appeared in all five. Not one prompt returned the same citation list twice, with a median overlap of 4.8% between same-prompt runs.

Read that as a measurement warning rather than an engine complaint. A brand that "lost" a mention between two single checks may have lost nothing, and a competitor that "appeared" may have been there all along. Anything derived from one run per prompt is a sample of one, however precise the percentage looks.

So the first question to ask of any AI visibility number, yours or a vendor's, is how many runs produced it. If the answer is not visible in the interface, do not put the number in a report. The mechanics of why answers move are in why AI answers change.

Visibility is not one thing per engine

The second measurement trap is aggregation. Engines behave differently enough that a blended figure hides the only information you can act on.

The same study ran 200 buyer prompts through ChatGPT and Perplexity. Perplexity attached citations to 100% of its answers, averaging 9.9 sources each, and cited 569 distinct domains. ChatGPT cited on 39% of answers, averaging 1.5 sources, across 138 domains. On 77% of prompts, the two engines cited zero overlapping domains.

Three consequences follow. Citation rate is not comparable across engines, so a drop in a blended citation number can just mean your prompt mix shifted. Winning one engine predicts almost nothing about another, so per-engine reporting is not a nice-to-have. And the work differs by engine: Perplexity leans heavily on communities and video, so presence there matters, while ChatGPT's citations skewed toward tech and business media.

This is also the case against the "AI presence score" that several tools now sell. A composite of five metrics across eight engines compresses genuinely different situations into the same number, and the compression discards exactly the detail that tells you what to do next. Track the components. Build a composite afterwards if an executive needs one line, and keep the components visible underneath.

Reading the metrics as a set

Individually each metric is ambiguous. In combination they are diagnostic, and the combination usually names the work.

What you seeMost likely readingWhere the work goes
High mentions, low citationsEngines believe you on other people's evidencePublish the answer-shaped pages that let engines cite you directly
Low mentions, high citationsYour content is useful, your brand is not in the consideration setThe comparison surface: "best X" and "alternatives to Y" pages
High mentions, low share of voiceYou are present but the category is crowded or a leader dominatesDifferentiation in the third-party sources, not more volume
Good rates on one engine, zero on anotherRetrieval or source-diet difference, not a brand problemAudit that engine separately, starting with crawler access
Mentions rising, sentiment fallingYou are being named as the cheap or limited optionThe source pages carrying that framing
Everything flat for two windows after a shipped fixEither the fix was upstream of nothing, or you are in the slow bucketCheck eligibility first, then accept the third-party timescale

The last row is the one worth sitting with, because it is the most common and the most misread. Structural fixes to your own pages can move within a window or two. Third-party presence moves in quarters. Memory-mode answers move with the next model. Flat numbers after three weeks of outreach are expected, not evidence of failure.

Worked numeric examples of the first two rows, with the arithmetic spelled out, are in the AI rank trackers guide.

How often to measure

Cadence should follow how fast the surface actually moves, not how often you would like a report.

Weekly is the default. It is fast enough to catch a competitor displacing you and slow enough that you are reading movement rather than sampling noise. Most programs never need more.

Daily earns its cost in three situations: a launch, the weeks after a model update when answers reshuffle, and active crisis work on sentiment or a factual error being repeated.

Google's surfaces reward a slower hand. AI Overviews and AI Mode move with the index rather than with sampling temperature, so weekly resolution is genuinely sufficient and daily checking mostly buys noise.

Re-measure after every ship, regardless of cadence. The point of a baseline is the comparison, and a fix you cannot connect to a number is a fix you will repeat blindly.

Why there are almost no real benchmarks

The question everyone asks next is what a good mention rate looks like. The honest answer is that nobody has a defensible cross-industry number, and you should be suspicious of anyone quoting one.

Three reasons. Category concentration varies enormously, so a 30% mention rate in a market with forty credible vendors is a stronger position than 60% in a market with three. Prompt sets are not comparable, since a set weighted toward branded questions produces flattering numbers and one weighted toward "best X" questions produces honest ones. And there is no shared methodology: two vendors running different run counts, engines, and detection rules will report different rates for the same brand on the same day.

Vendors are also poorly placed to publish benchmarks, ours included. We measure our own customers' workspaces, which is both a self-selected sample and their data rather than ours to aggregate and publish. Any tool quoting an industry benchmark is either using a public research corpus, which is worth asking to see, or telling you about its customer base rather than your market.

What to do instead: benchmark against yourself and your named competitors on the same runs. Your mention rate six weeks ago and your closest rival's rate this week are both real comparisons. An industry average is not.

Where AI visibility actually comes from

Three sources feeding an AI answer: your own pages, third-party pages about you such as reviews and communities, and model memory from training, with the relative weight of each
Three inputs, and the one most teams work on hardest is the smallest.

Visibility has three supply lines, and they respond to different work on different timescales.

Your own pages, reached through live retrieval. The fastest to change and the most directly controllable, which is why it absorbs most of the effort. It is also the smallest share of what engines cite: vendor-owned domains took 15.3% of ChatGPT's citations and 3.9% of Perplexity's across the 200-prompt study.

Third-party pages about you: reviews, comparison roundups, community threads, trade coverage. The majority of citations, and the slowest thing to influence, because it means outreach rather than publishing.

Model memory, the answers given with no retrieval at all, drawn from training data. Not directly addressable, and moved only by sustained third-party presence that ends up in a future corpus.

An honest AI visibility program allocates against that shape rather than against what is easiest to control. The playbook that does is how to do AEO.

Measuring it without buying anything

The manual protocol, in full:

  1. Write 20 buyer questions the way buyers type them. Finding buyer prompts covers where to mine them.
  2. Pick two engines your audience actually uses.
  3. Run each prompt five times in fresh sessions, signed out.
  4. Score each run: named or not, position, your domain cited or not, which competitors appeared, and whether the description is accurate.
  5. Compute mention rate, citation rate, and share of voice per prompt, with the run count next to each.
  6. Rank every cited domain by frequency. That table is the list of pages to go influence.

A couple of hours produces a baseline you assembled yourself, on your own prompts. Repeat weekly and you have trends, which is where the numbers start deciding things. The engine-specific walkthroughs go deeper: tracking brand mentions in AI search for the general method, and per-engine guides for ChatGPT, Perplexity, and Copilot, where Microsoft gives you citation data free.

What AI visibility is not

It is not traffic. Answers often resolve without a click, so visibility can rise while sessions stay flat. Connecting the two is a separate and harder problem, covered in AI search attribution.

It is not a ranking. There is no position three, and vendors describing their product as a rank tracker for AI are using a familiar word for an unfamiliar thing.

It is not a one-time audit. A single measurement is a sample, and the answers move underneath you.

It is not free of your SEO program. Retrieval runs on indexes built from the same signals your rankings depend on, which is why AEO vs SEO argues for one strategy measured in two places.

Where to go next

If you want the metric definitions turned into a running program, how to do AEO is the ordered playbook and tracking brand mentions is the measurement half in detail. If you are evaluating software, best AI visibility tools compares the category on documented pricing and method, and LLM visibility tools covers the checker end of the market.

Whatever you use, hold it to the standard this page has argued for throughout: rates over repeated runs, reported per engine, with the run count visible and competitors scored on the same answers. Tools that meet that bar disagree with each other far less than tools that do not. It is the bar our own prompt tracking was built around, run counts on screen included.