Quick answer
An AI mention occurs when a brand's name appears in an observed answer. An AI citation occurs, in this protocol, when a URL or domain belonging to the brand is explicitly shown as a source. AI share of voice is the share of mention events a brand earns among a set of competitors, over a fixed panel of prompts, platforms, and repetitions. These measures are neither a universal ranking, a real audience, nor a revenue attribution.
To produce a defensible number, publish the exact prompt, the platform, the date, the market, the number of runs, the normalization rules, and the denominator. Repeat every prompt, since a generative answer varies. Keep ChatGPT, Gemini, AI Overviews, AI Mode, and Perplexity separate before any summary. Finally, keep the raw outputs: if someone else can't recompute the rate from the export, the score is a proprietary indicator, not an auditable metric.
Key takeaways
- A brand can be mentioned without being cited, and cited without being named in the text.
- The denominator must be made of eligible runs, with failures explained.
- Prompt coverage measures the diversity of questions touched; the mention rate measures the frequency of positive runs.
- Share of voice depends entirely on the chosen competitor and prompt panel.
- Don't merge a third-party panel, a Search Console impression, a click, and a conversion.
The measurement contract
Before the first run, write a dictionary. The definitions below are this page's own; a provider can adopt different ones, but it must state them.
| Measure | Counted event | Denominator | Doesn't prove |
|---|---|---|---|
| Brand mention | Name or validated variant appears in the answer | Eligible runs | Recommendation, link, positive sentiment, or a read |
| Domain citation | Owned URL or domain appears as an explicit source | Eligible runs | That every sentence comes from the source or that a click happened |
| Prompt coverage | Unique prompt with at least one mention across its repeats | Eligible unique prompts | Frequency of the prompt in the real population |
| Mention-to-citation conversion | Run containing both a mention and a citation | Runs with a mention, per convention | Causality between the name and the choice of source |
| AI share of voice | Brand's mention events within the competitive panel | All mention events of tracked brands | Market share, audience, or real preference |
| Generative impression | Impression counted by the proprietary platform | Impressions per its own rules | Textual mention or observed citation |
| AI referral session | Visit attributed to a recognized referrer | Sessions per analytics | All exposures with no click or commercial causality |
A "response" and a "run" are synonyms here for a recorded execution of a prompt on a surface. A run is eligible if the platform returned a complete, usable answer under the rule set before the test. Timeouts, refusals, login pages, and technical errors stay in a failure log; they aren't silently dropped.
The core formulas
Mention rate
Mention rate = eligible runs with a mention / eligible runs × 100
A run mentioning the brand twice counts once for this rate, unless the protocol explicitly measures the number of occurrences. Avoid letting one long answer dominate the result.
Citation rate
Citation rate = eligible runs with at least one domain citation / eligible runs × 100
Decide whether several URLs from the same domain count as one run event or several citations. Publish both columns if both questions matter: cited_runs and cited_URLs.
Prompt coverage
Coverage = unique prompts with ≥ 1 mention / eligible unique prompts × 100
This metric keeps a small number of very favorable questions from masking an absence across the rest of the journey. It's still sensitive to the number of repeats: five chances make a prompt more likely to earn at least one mention than a single run.
Conditional mention-to-citation rate
Mention-to-citation rate = runs with mention AND citation / runs with mention × 100
If you instead compute cited runs / mentioned runs, a citation with no textual mention can produce an odd ratio. The convention above requires both events in the numerator. Also keep a citation_without_mention category.
AI share of voice
Brand M's share of voice =
M's mention events /
sum of mention events across every brand in the panel
× 100
Several brands can coexist in one answer, in which case each generates an event. The denominator is therefore not the number of runs. Another definition might attribute a single point per answer, split the point, or weight order. None is universal: pick one, name it, and don't change it mid-series.
Normalization decisions to make before testing
| Situation | Possible convention | Risk if left implicit |
|---|---|---|
| Generic name or homonym | Require product context or domain | Brand false positives |
| Subsidiary and parent company | Count separately, then publish a secondary aggregate | Inflated or lost presence |
| Misspelling | Closed list of validated variants | Unstable detection |
| Local and international domain | Normalize to one entity, keep the raw URL | Mixing markets |
| Subdomains | State which owned subdomains are included | Omitted citations or third-party domains included |
| Redirect and tracking URL | Resolve to the destination, keep both | Double counting |
| Third-party profile citation | "Third-party presence" category, not owned domain | Confusing the brand with an owned asset |
| Answer with no source list | Mention possible, citation not observable | Inferred citation with no evidence |
| Refusal or timeout | Exclude from the rate, count in the failure rate | Artificial improvement through selective suppression |
| Personalized answer | Separate stratum or documented fresh session | Non-reproducible result |
Sentiment can be annotated in addition, but it requires a separate rubric. "Alternative to X" isn't automatically negative; "not recommended for this case" can be useful and accurate. Have two humans check a sample and publish inter-annotator agreement before automating.
Reproducible protocol in eight steps
1. Define the decision
Choose what the measurement should inform: comparing themes, monitoring a campaign, diagnosing sources, or tracking competitors. A metric with no decision behind it produces an expensive chart with no action.
2. Build the prompt panel
Sample the journey stages: definition, problem, method, comparison, risk, cost, alternative, and action. Deduplicate near-identical rewordings. Note the origin: research data, customer support, sales, community, or hypothesis. An internal panel doesn't automatically represent overall demand.
3. Freeze the entities
List the brand, products, owned domains, and competitors before observing answers. Adding the most visible competitor after the fact changes the denominator and invalidates the historical comparison.
4. Fix the surfaces
Document application, mode, account, subscription, country, language, date, displayed model, and memory. Don't present "Google" as a single surface if you're testing AI Overviews and AI Mode separately.
5. Repeat and spread out
Run several executions per prompt and spread them over time when the goal is a trend. Five repeats are an example protocol, not a statistical truth. The stronger the variance, the more the interval and the sample size need scrutiny.
6. Keep the raw observations
Record identifier, prompt, answer, sources, timestamp, engine, parameters, status, and error. Keep the URLs before and after normalization. A single screenshot is hard to recompute; a single score is impossible to audit.
7. Annotate, then recompute
Apply the brand taxonomy and citation rules. Recompute the rates locally from the CSV. Inspect every ambiguous case, or at least a documented random sample.
8. Publish with uncertainty and breaks
Show numerator, denominator, failure rate, platforms, and period. If the panel, model, or rule changes, place a series break. A percentage to one decimal place doesn't imply real precision.
Worked example: 360 runs
A team tests 24 prompts on three platforms, five times each: 24 × 3 × 5 = 360 eligible runs. The example is fictional and shows the calculations, not the performance of SEOryon or a real brand.
| Platform | Runs | Runs with mention | Runs with citation | Mention rate | Citation rate |
|---|---|---|---|---|---|
| ChatGPT | 120 | 18 | 4 | 15.0% | 3.3% |
| Perplexity | 120 | 32 | 22 | 26.7% | 18.3% |
| Gemini | 120 | 22 | 4 | 18.3% | 3.3% |
| Descriptive total | 360 | 72 | 30 | 20.0% | 8.3% |
The total gives 72 mentions out of 360 runs, or 20%. The 30 citations give 8.33%. Assume the 30 cited runs also contain a mention: the conditional rate is 30 / 72 = 41.7%. If 14 of the 24 prompts got at least one mention on one platform, the descriptive coverage is 14 / 24 = 58.3%; a more rigorous analysis would also publish coverage per platform.
For share of voice, the competitive panel counts 72 events for the brand, 90 for competitor A, and 48 for competitor B:
Share of voice = 72 / (72 + 90 + 48) = 34.3%
That 34.3% doesn't mean 34.3% of users prefer the brand. It only means it represents 34.3% of mention events among these three brands, on this panel, over this period. Adding a fourth competitor or transactional prompts can immediately change the result.
If six additional runs failed, publish "360 eligible out of 366 attempted, 1.64% failure rate." Dropping them from the log would lose operational information; including them as answers with no mention would answer a different question. The convention must be explicit.
Connecting the panel to proprietary data
A prompt panel observes synthetic answers chosen by the analyst. Proprietary reports observe other events. Google announced in June 2026 a Search Console report on generative performance, initially for a subset of properties, with impressions, pages, countries, devices, and trends (Google's announcement and scope). These impressions don't necessarily reveal the prompt, the textual mention, or the citation as defined here.
Bing launched AI Performance in public preview in February 2026, including total citations, cited pages, and sampled anchor phrases. Bing warns that the figures don't represent placement, authority, or rank (Bing documentation). Keep Bing metrics under their own name instead of adding them to a tracker's runs.
Your analytics finally measures recognized sessions and conversions per its own configuration. An answer can influence a later brand search with no referrer. Treat that influence as a hypothesis to study with experiments, surveys, or time series, not as a certain conversion.
What the data proves and doesn't prove
Semrush's 2026 AI Visibility Index also distinguishes mentions and citations and publishes results per platform. Per its methodology and press release, the analysis covers 126 million deduplicated US prompts, from January to April 2026, drawn from Datos clickstream and AI Overviews keyword data, across 22 sectors and on ChatGPT, Google AI Mode, AI Overviews, and Gemini (Semrush methodology, Semrush press release).
This scale shows that segmentation by platform and sector is possible. It doesn't make the proprietary score universal: the full details of deduplication, prompt execution, model versions, repeats, and uncertainty aren't all public. This is a commercial provider. Use this study to understand a market method, not to automatically calibrate your local rate.
Google describes query fan-out and retrieving several results for its AI search experiences, while reiterating that SEO fundamentals still apply (Google's guide). That explains why an answer isn't necessarily a mirror of the visible query's top 10. It doesn't reveal every sub-query, nor a formula that guarantees a citation.
Common mistakes and stopping conditions
- Dividing the number of name occurrences by the number of prompts, when each prompt has several runs.
- Counting a Wikipedia page about the brand as an owned domain.
- Detecting a homonym with no context check.
- Mixing a cited URL, a cited name, and an unseen grounding source.
- Changing the competitive panel after seeing the results.
- Comparing a single run in January to five runs in July.
- Aggregating platforms without showing their separate rows.
- Presenting a panel score as audience share or market share.
Stop publishing if the denominator is unknown, if more observations are missing than a predefined threshold, if the raw outputs aren't kept, or if a protocol change was hidden. The decision can be "insufficient data."
Reusable asset: SEOryon's AI visibility protocol
The assets/protocole-visibilite-ia.csv file serves as a row-by-row log. Keep at least: prompt_id, prompt_version, theme, platform, surface, market, language, run, timestamp, status, answer, raw_sources, detected_brands, normalized_domains, annotator, and rule_version.
Add a summary sheet that computes this page's formulas, never the reverse. The aggregate must be reconstructible from the rows. Version the file before any change to a brand, a prompt, or a rule.
Where to go next
- Go back to the 45 definitions of the SEO, GEO, and AI search glossary to frame the neighboring units.
- Build the panel with the full protocol for measuring AI visibility.
- Compare the features and exports of an AI citation tracking tool.
- Interpret the gap between AI citations and Google positions without jumping to causality.
How SEOryon fits in
SEOryon can be evaluated or used within this framework with no favorable definition. The measurement must expose observations, denominators, and limitations, and public claims must stay tied to the panel in question. No "SEOryon score" should replace raw rates or promise future visibility.
Measurable exercise: build an auditable mini-benchmark
- Choose ten prompts representing five journey stages.
- Test two platforms, five runs per prompt: 100 attempted runs.
- Define two brands and their domains before execution.
- Annotate mentions, citations, failures, and ambiguous cases.
- Compute rates per platform, coverage, and share of voice.
- Hand the CSV to a second person for recomputation.
Deliverable: 100 attempted rows, a failure log, and a summary with numerators. Success criterion: zero recomputation gap, 100% of ambiguous cases decided, platforms kept separate, and no unsupported audience or revenue conclusion.
FAQ
Does a mention with no link have value?
It can signal a presence in the answer, but its commercial value isn't automatic. Track sentiment, brand search, traffic, and conversions separately, then test the relationships without claiming an unobserved causality.
How many repeats does a prompt need?
There's no universal number. Start with several runs, measure the variance, and increase the sample when the decision or the instability justifies it. Always publish the chosen number.
Can you add up the results across all platforms?
A descriptive total is possible if the units are identical, but it can hide large gaps and overweight one surface. Show each platform first and make any weighting explicit.
Is a citation necessarily clickable?
No. Interfaces and definitions vary. The protocol must specify whether the source is a visible link, a displayed domain, a numbered reference, or grounding data reported by the platform.
How do you track a change after a prompt update?
Keep the old panel as a fixed cohort and launch the new one in parallel for a bridging period. Only chart a continuous series over the prompts and rules shared by both.
References
- Semrush: AI Visibility Index methodology2. Semrush: 2026 AI Visibility Index press release3. Google Search Central: Generative AI performance reporting5. Bing Webmaster Blog: AI Performance public preview6. Google Search Central: AI features optimization guide
Method and update note
Page reviewed 22 July 2026, translated and edited from the French original (reviewed 16 July 2026). The formulas are declared analytical conventions, not standards imposed by platforms. The 360-run example is fictional. Check the Google, Bing, and Semrush methodology reports each quarter; place a series break on any change to prompts, platforms, models, normalization rules, or eligibility definitions.