top of page

AI Optimisation Metrics: Mentions, Citations, Rank, Sentiment and Share of Voice

Aug 18
10 min read

This guide was produced by AI, TELL ME!, a Berlin-based AEO and GEO agency with its own AI search monitoring platform.

Summary

Most brands can quote their Google ranking to the decimal point and have no clear read on what their AI optimisation produces once someone asks ChatGPT the same question, which is the gap this piece closes. It works through the five metrics that describe how a brand shows up in AI-generated answers: mentions, citations, rank, sentiment and share of voice, and what each one measures underneath the label. It shows how to calculate each metric with a working formula rather than a vague description, what a strong score looks like against a weak one, and where teams misread the numbers because they are still thinking in Google-rank terms. By the end it is clear why treating these five as one dashboard, not five separate reports, is the only way the numbers carry any weight.

Abstract illustration of five AI visibility metrics — mentions, citations, rank, sentiment and share of voice — converging into one central AI optimisation score.

A brand's Google position and its presence in an AI answer are no longer the same question, and most reporting still treats them as though they are. Someone on the marketing team pulls a rank tracker, sees the domain sitting in position two or three, and assumes visibility is handled. Meanwhile ChatGPT, Claude, Gemini or Perplexity answers the exact buyer question without mentioning the brand at all, or mentions a competitor first, or cites a review site instead of the brand's own page. None of that shows up on a rank tracker, because a rank tracker was built for a different kind of retrieval system.

AI optimisation, in the sense that matters for reporting, is the practice of tracking a different set of numbers: whether a brand gets mentioned, whether it gets cited as a source, where it lands within an answer, how it is described when it does appear and how its visibility compares with the brands it competes against. Each of those is a separate metric with its own method of calculation, its own failure modes and its own blind spots, and conflating them is how a lot of otherwise careful marketing teams end up reporting a number that means very little.

Why AI Optimisation Needs Its Own Metrics

Traditional SEO reporting assumes a deterministic system: a domain either ranks in position four for a keyword or it does not, and that position is stable enough to check weekly. LLMs (large language models) do not work that way. The same prompt asked twice in the same session can produce a different set of brands, in a different order, with different framing, because the model is generating a response rather than retrieving a fixed list. That single fact is why AI optimisation cannot borrow a rank tracker's logic wholesale and needs its own set of metrics built around probability and pattern rather than a fixed position.

Adoption numbers explain why this now matters at reporting level rather than as a future concern. In the United States, 49% of adults report using an AI chatbot, with 42% specifically using one to search for information, according to Pew Research Center. In the European Union, roughly a third of people aged 16 to 74 had used a generative AI tool in the past three months as of the most recent Eurostat figures, with usage running well above that average in Denmark, Estonia and Finland. Buyers are already asking these questions before a brand's own analytics ever register them, which makes the five metrics below closer to a pipeline report than a marketing vanity metric.

The Five AI Visibility Metrics AI Optimisation Depends On

Five metrics cover what happens when an LLM answers a buyer's question: whether the brand is mentioned at all, whether it is cited with a source link, where it lands within the answer, how it is described and how its visibility compares against competitors answering the same prompts. These five sit inside the broader question of what AEO monitoring measures as a discipline, but each one deserves its own formula and benchmark rather than a single-line definition, which is what the sections below set out to give.

Mentions

A mention is any point where a brand's name appears in an LLM's response to a relevant prompt, cited or not. It is the coarsest of the five metrics and the easiest to track, which also makes it the easiest to over-read. A brand mentioned in ten out of fifty relevant prompts has a mention rate of 20%, calculated as mentions divided by total prompts tested across a fixed prompt set.

Mention rate answers one question only: does the brand exist in the model's frame of reference for a given category. It says nothing about whether that mention helped or hurt, or whether a buyer reading the answer would come away more likely to choose the brand. A mention rate above 40% across a well-built prompt set is generally strong for an established category player, while anything under 10% usually means the brand has not built enough third-party presence for the model to associate it with the category at all.

Citations

A citation is narrower than a mention: it is a case where the model names a specific source, usually with a link, and that source is the brand's own domain rather than a review site, forum or news outlet describing it. Citation rate is citations divided by total mentions, not total prompts, because a brand can be mentioned without ever being cited as the source.

This distinction matters because a high mention rate paired with a low citation rate usually means the brand is known to the model through secondary sources such as review platforms rather than through its own content being trusted enough to link back to, and the fix for that gap sits mostly in content and site structure, covered in more detail in the guide to getting a website cited by ChatGPT.

Rank

Rank measures where within an answer a brand appears, on the logic that a brand named first in a list of five carries more weight with the reader than one named last, even though both count equally in a raw mention count. Average rank is calculated by recording the brand's position in every response where it appears and averaging those positions across the prompt set, so a brand that appears first in three answers and third in two more has an average rank of 1.8.

Rank inversions are common and often surprising. The 5W AI Visibility Index, which tracked 780 buyer prompts across 325 brands in 13 industries, found that the largest airline by capacity ranked fifth in AI citation share while a smaller competitor led, and that the world's biggest luggage maker trailed a much newer rival in the same category. Market share and AI rank measure different things, and a brand with a strong average rank but a thin mention rate is a different problem from a brand with a high mention rate and a weak average rank, which is why the two numbers need to sit next to each other rather than stand in for one another. AI, TELL ME!'s own reporting tracks this nuance separately as First Mention Rate, alongside average rank, since being named first shapes buyer recall more than being named third in the same answer.

Sentiment

Sentiment records how a brand is described when it appears, not just whether it appears. An LLM might mention a brand as the market leader, as a budget option, as a legacy tool worth avoiding, or in neutral, factual terms, and those four outcomes have very different commercial value even though a basic mention tracker would count them identically. Sentiment is usually scored on a simple three-point scale, positive, neutral or negative, applied per mention and then rolled up into a percentage split across the prompt set.

A brand can have a healthy mention rate and a healthy citation rate while still losing on sentiment, typically because it gets grouped with lower-tier competitors or described using outdated positioning that no longer matches its product. Sentiment in AI answers is the metric most often skipped in a basic AI optimisation setup, because it requires reading the text of each response rather than detecting whether a brand name appears, and it is usually the one that explains a visibility number that looks fine on paper but does not translate into pipeline.

Share of Voice

Share of voice compares a brand's visibility against a defined competitor set answering the same prompts, rather than measuring the brand in isolation. It is calculated as the brand's mentions divided by total mentions across all tracked competitors for the same prompt set, expressed as a percentage. A brand with a 20% AI share of voice across a competitor set of five rivals, six entities in total including the brand itself, is matching an exactly even split, while anything higher signals dominance in the model's eyes and anything lower means competitors are being named more often for the same prompts.

Share of voice is the metric that turns the other four into a competitive picture rather than a standalone report. A brand can improve its own mention rate every month and still lose share of voice if competitors are improving faster, a distinction that a single-brand dashboard cannot show and a comparative one always will.

The Five Metrics at a Glance

Metric

What it measures

How it's calculated

Strong benchmark

Common blind spot

Mentions

Whether the brand appears at all

Mentions ÷ total prompts tested

Above 40% for an established player

Says nothing about position or tone

Citations

Whether the brand is named as the source

Citations ÷ total mentions

Above 30% of mentions carrying a citation

High mentions with low citations signals second-hand visibility

Rank

Where the brand lands within an answer

Average position across all appearances

Average rank under 2.5

A good average can hide a thin mention rate

Sentiment

How the brand is described when it appears

Share positive, neutral and negative per mention

Majority positive or neutral, negative under 10%

Requires reading text, not just detecting names

Share of voice

Visibility relative to competitors

Brand mentions ÷ total competitor-set mentions

At or above an even split for the competitor count

Can mask decline if rivals are improving faster

Reading the Metrics Together

None of these five numbers means much when read alone, and a high mention rate with a weak share of voice usually means the whole category is well covered by LLMs and the brand is keeping pace rather than leading it. A strong average rank paired with a low citation rate suggests the model trusts secondary sources describing the brand more than the brand's own site, a content and structure problem rather than a monitoring one that overlaps closely with the technical and content layers covered in the guide to AI optimisation for websites, while a healthy citation rate paired with negative sentiment is arguably the worst combination of the five, because the brand is visible and technically well cited, and the model is still steering buyers away from it in the framing.

The practical fix is to read the five as a single scorecard rather than five separate charts, checking for combinations that flag a specific kind of problem instead of treating each metric as pass or fail on its own. That is roughly the logic behind rolling all five into one composite figure, which is what a GEO Score is built to do: a single number a team can track over time, backed by the five underlying metrics that explain why it moved.

Turning AI Search Metrics into a Reporting Habit

Tracking these five metrics once amounts to an audit, while tracking them on a schedule is what makes them useful, because LLM output shifts from one model update to the next and a single snapshot goes stale within weeks. A workable reporting habit usually follows the same handful of steps regardless of company size.

  1. Fix a prompt set of 20 to 40 buyer queries covering category, comparison and problem-first phrasing, and keep it stable enough that month-on-month numbers stay comparable.

  2. Define a competitor set of three to six brands that buyers would realistically consider alongside the brand being tracked.

  3. Run the prompt set across every LLM the target audience uses, since a single-model report understates total exposure across the audience.

  4. Score each response for mentions, citations, rank and sentiment, then calculate share of voice against the competitor set for the same run.

  5. Compare results month over month rather than reacting to any single response, since individual answers vary even when the underlying pattern is stable.

Brands that skip straight to fixing website content without first establishing this baseline tend to make changes they cannot evaluate, since they have nothing to compare the after state against. The methodology for building and testing the prompt set itself is covered in more detail in the guide to checking whether ChatGPT recommends a brand, and the five metrics above are what that testing process should be scored against once the prompts are running.

FAQ

What is the difference between a mention and a citation in AI search?

A mention is any appearance of a brand's name in an LLM's answer, with or without a source. A citation is a mention where the model names the specific source, typically the brand's own domain, usually with a link a user could follow. A brand can have a high mention rate driven mostly by review sites and forums describing it while its own citation rate stays low, which is a distinct problem from not being mentioned at all.

How is AI share of voice calculated?

AI share of voice is a brand's mentions divided by the total mentions recorded across a defined competitor set answering the same prompts, expressed as a percentage. It requires tracking mentions for every brand in the competitor set on every run, not just the brand being reported on, since the number is only meaningful relative to rivals.

Can a brand have positive mentions but still lose on AI optimisation metrics?

Yes. A brand can be mentioned frequently and even cited regularly while still being described in outdated or unfavourable terms, or while consistently ranking behind competitors within the same answers. Mentions and citations measure presence, while sentiment and rank measure the quality of that presence, and a brand needs all four moving in the same direction for the visibility to convert into anything commercially useful.

How often should these AI visibility metrics be tracked?

Monthly is a reasonable baseline for most categories, since LLM outputs shift gradually enough that weekly tracking rarely shows meaningful movement. Categories with heavy competitive activity or frequent model updates benefit from tracking every two weeks, mainly to catch sudden drops in mention rate or sentiment before they compound.

Do these LLM visibility metrics differ across ChatGPT, Claude, Gemini and Perplexity?

Yes, often substantially. The LLMs a brand's audience uses draw on different source mixes and weight citations differently, so a brand can show a strong share of voice on one model and a weak one on another using the exact same prompt set. Tracking a single model and reporting it as overall AI optimisation performance is one of the more common reporting mistakes teams make.

Building this scorecard from scratch, prompt set and all, is the work AI, TELL ME!'s GEO Monitoring is built to do, tracking mentions, citations, rank, sentiment and share of voice across the LLMs a brand's buyers use and rolling the result into a single GEO Score that moves as the underlying numbers do. Anyone wanting to see where a brand currently stands on these five metrics before changing anything can run a free AEO Monitoring check at AI, TELL ME!.

If you want to see how AI currently describes your  clinic, device or treatment  and where run a free AEO Monitoring check at aipleasetellme.com.



bottom of page