What Do AI Citations vs Mentions and Share of Voice Measure?

Ask three vendors to define AI citations vs mentions and you'll get three answers, all confident. A citation is a source URL the model surfaced as supporting material, countable and tied to an address. A mention is your brand appearing in the answer text, whether or not a source sits alongside it. Share of voice is a ratio of your appearances against a competitor set that somebody chose. Two of those are counts. The third is an editorial decision with a percentage sign on it.
So which number belongs in the board deck? Would a rival tool return the same figure? And if your share of voice doubled tomorrow, could you name the page that did it?
If the third question stalled you, that's the gap this covers.
Key takeaways
- Citations are the most diagnostic of the three, because they expose the source layer behind an answer.
- Mentions are exposure. Your brand appeared, and the count on its own says nothing about why.
- Mentions and citations aren't alternatives. An answer can name you and carry citations at the same time, to your site or somebody else's.
- Share of voice is a ratio, and it's only readable against a fixed prompt set and a fixed competitor set.
- Vendors use "citations" and "mentions" interchangeably in the same formula, and some contradict themselves on a single page, so figures from two tools aren't comparable.
- Track all three. Act on citations.
What does AI citations vs mentions compare?
It compares two things that sit at different depths of the same answer. A mention is what the model said. A citation is where the model pointed. They often occur together, and either can occur without the other, which is why teams tracking one number keep getting surprised by the other.
The wider family includes a fourth term, the AI visibility score, which vendors typically publish as a composite index. Here's the full set, side by side.
Two are counts, one is a ratio, and one is a composite that reduces the other three to a single figure your CFO can misread. The ratio is the one that ends up on slides.
What counts as an AI citation?
A citation is a source URL or reference the AI surfaces as supporting material for its answer, rendered as a link or a footnote. It isn't always clean attribution mapped to one specific claim, but it does arrive with an address, which makes it the diagnostic in the set. When a citation goes to a competitor's comparison page, you know exactly which page to answer.
It's also the metric with a lever attached. Kevin Indig's analysis of 18,012 verified citations, drawn from 1.2 million ChatGPT answers and reported by Search Engine Land, found that 44.2% of citations came from the first 30% of a page, 31.1% from the middle 40%, and 24.7% from the final 30%. That's ChatGPT specifically, not AI in general, and keeping the two apart is the habit this whole article is arguing for. The distribution strongly favours earlier sections, which makes front-loading your definitions a sensible bet. It's an observed association rather than a controlled test, so read it as a reason to put the answer near the top, not a promise that moving a paragraph upwards earns you a citation.
That's the tactical layer, and it's a separate job from measurement. We've covered the content signals that earn citations elsewhere. Here the point is narrower: citations are countable, they name a source, and they show you where the answer's evidence came from. No other metric on this list does all three.
What counts as a brand mention?
A brand mention is your company name appearing in generated answer text. That's the whole test: the brand is in the answer. It may sit alongside citations or alongside none at all, and most vendor share-of-voice methods count the mention either way.
The uncited case is the one that catches teams out. The model recommends you, the buyer reads your name, and the trail ends there.
That gap is wider than most teams assume. A practitioner scan of B2B SaaS brands across ChatGPT, Perplexity and Gemini, posted to r/b2bmarketing, found one bootstrapped tool where 88% of the citations in its own AI answers pointed at third-party pages and only 5% at its own domain. That's one brand, self-reported, and the person who ran it sells a tool in this category, so treat it as a signal rather than a study. The mechanism it describes is familiar enough: the model recommends you, and somebody else gets the click.
Which explains why mention counts and referral analytics never reconcile. Referral data counts people who clicked. Plenty of people reading an AI answer never click through at all, and a mention that produces no click leaves no trace in GA4. Exposure is real and your analytics can't see it.
What is AI share of voice?
AI share of voice is generally a relative ratio of your brand's appearances against some defined competitive universe, across a prompt set somebody chose. The important part is that vendors disagree about what belongs in that denominator. Trakkr draws a clean published line between absolute visibility, meaning how often you appear at all, and relative share, meaning how much of the category you hold, and says plainly that the relative figure is sensitive to which competitors you track.
The words doing the quiet work are "defined competitive universe". Many tools let you nominate the rivals you're measured against, so adding a competitor that earns mentions can drop your share without anything changing in the market. Remove one and it climbs. The number moved because a person edited a list. Waikay calls that closed denominator a structural flaw in rival tools and publishes the open version instead: your mentions over every brand the model names in any response, with no preset competitor list at all. The two also split on weighting. Trakkr argues for weighting by position, on the grounds that being named first matters; Waikay argues against it, because response order is unstable between runs. There is no settled formula here, which is rather the point.
None of that makes the metric useless. It makes it a positional reading rather than a performance reading, and positional readings need their configuration stated every time they're quoted. If you want the mechanics of running one properly, we've written up how to measure AI share of voice separately.
Why do dashboards disagree on numbers?
Dashboards disagree because they're not counting the same events, and almost nobody says so on the pricing page. Two of the three pages below also disagree with themselves.
Citations, mentions and responses are three different events, and the same brand can be handed very different percentages in the same week without any tool making an arithmetic error. Two of these pages change their own denominator inside a single scroll. The third rejects the method the rest of the category uses. Nobody is lying. They counted different things and printed the same word on the chart.
Three practical consequences. Don't compare a share-of-voice figure to one from another tool until you've reconciled what each one counts. Don't change tools mid-quarter without restating the baseline in the new tool's terms. And when a vendor quotes your share against competitors, ask which competitors, because the competitor set determines the answer more than your content does.
Why do mentions and citations diverge?
There's a hypothesis in the practitioner community worth knowing, with a caveat attached before the sentence rather than after it: this is a mechanism claim from experienced people, not published research, and we haven't yet tested it against our own Brand Radar data.
The claim is that some ungrounded brand mentions originate in model knowledge, while retrieval supplies the sources used to support or refine the answer. On that reading, a share of citations are attached after the fact: the model named you because of what it absorbed months ago, then went looking for something linkable to hang the sentence on.
Current evidence doesn't establish that as the universal mechanism, and the tidy version is almost certainly too tidy. Reporting on ChatGPT's retrieval stack describes separate discovery, reading and source-selection stages running before and during synthesis rather than bolted on afterwards, so a given answer may draw on model knowledge, retrieved sources, the conversation so far, or all three.
It's still worth holding loosely, because it's consistent with two things practitioners keep reporting: citations only loosely related to the sentence they support, and citations that resolve to dead links. And if it holds even partly, earning mentions and earning citations are partly different programs. Presence in third-party sources plausibly helps the first. Being the most retrievable answer at the moment of the query plausibly helps the second. Same goal, overlapping but not identical work.
The honest position, and the one worth adopting in front of a board, is that no current tool can measure every citation your brand receives across every user's prompts, sessions, engines and locations. You measure observable signals from a prompt set you invented, and the tools worth paying for are the ones that say so.
Which metric should you act on?
Act on citations, in most cases. They're the most diagnostic of the three, because they expose the source layer behind the answer.
Share of voice belongs in the quarterly review as a direction of travel. Mentions belong in the same place, as evidence that exposure is growing even when nobody clicks. Citations belong in the weekly content meeting, because a citation gap resolves into a named source and a decision about it. Sometimes that decision is to improve a page you own. Often it isn't: the cited source is a Reddit thread, a G2 listing, a Wikipedia entry or a journalist's comparison, and the work is earning a place there rather than editing your own site. If your reporting only carries the ratio, your team has a number to discuss and nothing to do.
What breaks each of these numbers?
Sampling breaks them, mostly. Rand Fishkin and Patrick O'Donnell ran 12 identical prompts across ChatGPT, Claude and Google's AI, close to 3,000 test runs with around 600 volunteers, and Search Engine Land covered the result: the odds of getting the same list of brands twice came in under 1 in 100, and closer to 1 in 1,000 for the same list in the same order. A single run isn't a reading. It's one roll of a very large die, and any tool reporting a share of voice from a thin sample is reporting variance.
Angela Skane, writing in Search Engine Land, offers a working rule of thumb: treat something as a pattern when it appears in at least three of every four outputs, holds across two different models, and survives repeated runs of the same prompt. She's careful to say there's no statistical basis for that number and that a lower bar would serve as well. Take it as a discipline for reading your own data rather than a standard anyone can hold you to.
Three other failure modes are worth naming. Blended engines can hide how often ChatGPT, Perplexity and Gemini disagree, so a single averaged score conceals the divergence you most need to see. Sampling depth varies, so ask a vendor how many prompts it runs and how often rather than assuming the dashboard reflects the whole set. And polling frequency varies enormously between tools, which means two vendors watching the same brand in the same week can honestly disagree about what happened.
How should you evaluate a tool?
Ask five questions, drawn from what buyers in this category argue about rather than what vendors put on comparison pages:
- Does it show the source URLs behind each answer, or only that your brand appeared?
- Does it report engines separately instead of blending them into one average?
- Can you tell whether your page was cited, a rival's page was cited, or the answer rested on Reddit or a listicle?
- Does a prompt gap turn into an action, meaning a page to repair or a source to pursue?
- After you change something, does the tool show the source mix changing, not only the score?
The last one separates measurement from theater. A score that moves without the source mix moving tells you the model had a different day. If the underlying evidence base shifted, something you did worked. Teams new to this discipline may want the plain-English map of GEO, SEO and AEO before they start buying software, because a fair amount of what's sold as AI visibility tooling is a rank tracker with a new logo.
FAQs
Is an AI citation the same as a backlink?
No. A backlink is a permanent link on someone else's page that you can audit any time. An AI citation is a source the model surfaced in one answer, in one session, for one user. Run the same prompt an hour later and it may cite something else entirely. Backlinks accumulate. Citations are events.
Can you have mentions without citations?
Yes. A model can mention your brand without citing a source, and it can also mention your brand while citing a third-party source rather than your own site. Depending on the system and the query, the mention may draw on model knowledge, retrieved sources, or both. Either way, a rising mention count doesn't on its own mean your content is being read.
What is a good AI share of voice?
There's no threshold that travels between tools, because the denominators differ. The benchmark with meaning is your own figure in your own tool, tracked against the same prompt set and the same competitor list over time. A vendor quoting you an industry average is quoting an average of numbers that were never calculated the same way.
Which metric belongs in a board report?
Share of voice as a trend line, with the competitor set named in the footnote, plus one citation-level example that shows the work. The ratio gives leadership the position. The citation example gives them the evidence that somebody is doing something about it. Reporting the ratio alone invites the question you can't answer, which is what changed and why.
Start with citations for a fortnight. Pull the prompts your buyers would realistically type, run them across at least two engines, and record which domains the answers rest on. You'll find a short list of third-party pages doing the persuading on your behalf, and many teams find their own site barely appears.
That list is the work. If you'd rather have it built for you, our GEO playbook for B2B walks through the same process end to end, or you can talk to us about running it as a diagnostic.
