Quick answer
An AI citation tracking tool doesn't observe "a brand's ranking in artificial intelligence." It typically runs a defined panel of prompts on certain platforms, then detects brand mentions, cited links, and sometimes their position, sentiment, or competitors present. Its result therefore depends on the choice of prompts, market, account, model, date, and number of repetitions.
To choose a tracker, demand the raw observations: exact prompt, response, engine, timestamp, cited URL, and normalization rule. Repeat the same protocol and measure the variance. A score with no denominator or export isn't comparable. Finally, combine this third-party panel with the proprietary data available in Google Search Console, Bing Webmaster Tools, and your own analytics: mention, citation, impression, session, and revenue are different stages.
Key takeaways
- A mention names the brand; a citation uses or displays a source. The two aren't interchangeable.
- The prompt panel only measures the questions tested, not the whole of real conversations.
- A single run is too fragile to measure a trend; document repetitions and variance.
- Google's and Bing's proprietary reports don't measure the same thing as third-party trackers.
- The best tool is the one that exposes its data, its definitions, and its errors, not the one showing the highest score.
The four layers of measurement
A GEO (generative engine optimization) tracker can blend four layers. Separate them before comparing:
- Execution: the tool sends a prompt to a platform under a given market, language, and set of parameters.
- Extraction: it detects brand names, domains, URLs, source order, and associated text.
- Aggregation: it computes rates, scores, trends, and competitive comparisons.
- Attribution: it may reconcile these observations with sessions and conversions received on the site.
An error at the first layer contaminates the rest. A clumsily translated prompt, a truncated response, or a poorly normalized URL can change the final score without the brand's real visibility having changed at all.
What each data source knows and doesn't know
| Source | Direct observation | What it doesn't prove | Essential check |
|---|---|---|---|
| Prompt tracker | Response obtained for the panel's prompts and runs | Real frequency of the prompts in the population; exposure of every user | Exportable prompts, repetitions, market, account, model, and timestamp |
| Google Search Console, generative report | Impressions, pages, countries, devices, and trends on AI Overviews and AI Mode for eligible properties | Brand mention in the text, sentiment, causal conversion, or every platform | Report availability, property/page aggregation, canonicals, thresholds, and row limits |
| Bing Webmaster Tools AI Performance | Citations, cited pages, and sampled grounding phrases on covered surfaces | Placement, authority, or "ranking" in AI | Preview scope, sampling, and evolving surfaces |
| Web analytics | Sessions attributed to a referrer and on-site actions | Influence without a click, a response seen, the exact prompt, or a citation that preceded the visit | Channel rules, dark/direct traffic, consent, conversion, and deduplication |
| User survey | Reported use, preference, or influence | Observed behavior and incrementality | Population, wording, period, geography, and weighting |
Google announced a dedicated Search Console report for generative performance in June 2026, initially available for a subset of sites. It shows impressions, pages, countries, and devices, and the data stays included in the overall reports (Google announcement and scope). An impression is neither a verified citation nor a click.
Bing launched AI Performance in public preview in February 2026, with total citations, average cited pages, sampled anchor phrases, and URL trends. Bing states that these figures indicate neither placement, nor authority, nor rank (Bing documentation). This caveat should stop anyone from merging Bing counts with a ChatGPT score as if they shared a unit.
The buying grid for a citation tracker
| Criterion | Question for the vendor | Observable test | Warning sign |
|---|---|---|---|
| Coverage | Which platforms, surfaces, countries, languages, and account types? | Run the same prompt on every advertised surface | A logo shown with no detail on the mode covered |
| Prompts | Who chooses them, and can you export the exact text? | Import, edit, version, and export twenty prompts | Prompt hidden or automatically rewritten with no trace |
| Repetitions | How many runs per data point, and when? | Repeat five times, then inspect each response | A single run presented as a stable measurement |
| Model and context | Are version, account, memory, personalization, and location logged? | Compare a fresh session with a personalized one | Unknown parameters but a very precise score |
| Extraction | How are brands, domains, subdomains, and redirects normalized? | Test variants, homonyms, and a redirected URL | Brand false positive or merged domain |
| Sources | Is the full response and are the URLs retained? | Open an old result and reconstruct the calculation | Only an aggregated chart |
| Variance | Does the tool show stability or a range? | Compare positive runs per prompt | A trend with no sample size |
| Export/API | Can you retrieve observations and definitions? | CSV/API export, then local recomputation | Proprietary score that can't be reproduced |
| Costs | Does the price depend on prompts, runs, engines, markets, or seats? | Simulate the quarterly volume and a spike | "Prompt" billed with no detail on repetitions |
| Governance | Are roles, clients, retention, and deletion controllable? | Revoke a client and export their history | Client data visible in a shared space |
Seven-step reproducibility protocol
1. Define the prompt population
Build a panel covering discovery, problem, comparison, selection, usage, and risk. Don't just pick queries where the brand is already known. For each prompt, note audience, market, language, intent, and business reason.
2. Freeze a version
Keep the exact text and an identifier. If the prompt changes, create a new version; don't rewrite history. A translation is a different prompt.
3. Define the run
Document platform, surface, whether an account is connected, location, language, date, time, visible model, memory, and available parameters. Engines evolve; missing information should be recorded as "not shown," not guessed.
4. Repeat
Choose a number of repetitions compatible with your budget and the observed variability. Five runs is a pedagogical starting point, not a statistical standard. If a prompt goes from zero to five mentions depending on the run, increase the sample or conclude it's unstable.
5. Keep the responses
Record text, links, order, citations, date, and error. Respect the platforms' terms and data protection obligations. A technical error still counts in the denominator of attempts but can be excluded from eligible runs, provided the rule is published.
6. Recompute locally
Calculate the mention and citation rate from the export. If your results don't match the dashboard, ask for the aggregation rule, weightings, and exclusions.
7. Connect to the business, carefully
Add referred sessions, conversions, brand search, and sales feedback, but don't claim a citation caused a sale without an incrementality setup. A recommendation seen without a click can still influence; it's just as hard to attribute.
Minimal formulas, no magic score
Mention rate = eligible runs mentioning the brand / eligible runs
Citation rate = eligible runs citing at least one URL of the domain / eligible runs
Prompt coverage = prompts with at least one positive run / prompts tested
Prompt stability = positive runs / repetitions of that prompt
Always publish the numerator and the denominator. A 50% rate on two runs doesn't carry the same weight as 50% on 2,000 runs. Keep platforms separate before any total: an engine that cites abundantly can artificially dominate a combined score.
Worked example: 20 prompts, three platforms
A SaaS brand tests 20 prompts on three platforms, five times each: 300 planned runs. Six errors are excluded under an announced rule, leaving 294 eligible runs. The tracker detects 42 runs with a mention and 18 with at least one citation of the domain.
- Mention rate:
42 / 294 = 14.3%. - Citation rate:
18 / 294 = 6.1%. - Mention-to-citation conversion:
18 / 42 = 42.9%, to be read only among mentioned runs. - Eleven of the twenty prompts have at least one mention: coverage
11 / 20 = 55%.
The total masks a difference: 13 of the 18 citations come from a source-oriented platform, while a conversational platform often mentions the brand with no link. The decision isn't therefore "improve the overall score," but to separate two tasks: make the pages more useful as sources on the first platform and clarify the brand entity on the second.
The figures are fictional and illustrate the formulas. They are not a SEOryon benchmark.
What the data proves and doesn't prove
Semrush's AI Visibility Index, announced in June 2026, separates mentions and citations and analyzes 126 million US prompts across ChatGPT, Google AI Mode, AI Overviews, and Gemini, in 22 industries. The methodology states deduplicated prompts sourced from Datos and AIO data, and presents results by platform (Semrush methodology). It doesn't fully detail runs, model versions, repetitions, coverage, or uncertainty; the score remains proprietary. This corpus shows the value of separating platforms, not that the index can replace your own commercial panel.
Google explains that its generative features may use query fan-out and that results and links vary (Google guide updated July 10, 2026). This description concerns Google Search, doesn't reveal the weightings, and doesn't prove that a given change will trigger a citation.
Limits and stopping conditions
- Platforms can personalize based on account, history, language, location, and availability.
- The model or retrieval system can change between two waves.
- A vendor can silently change prompts, normalization, or weightings.
- The real volume of user prompts is rarely known; a panel isn't a market measurement.
- A citation can be positive, critical, outdated, or out of context.
- A cited URL can redirect, be canonicalized elsewhere, or belong to a subdomain that isn't included.
- Responses with no link aren't necessarily "unsourced"; the system may not display all its influences.
- Direct traffic or a brand search can have been influenced with no observable referrer.
Stop the purchase if you can't export the prompts and observations, if the score changes with no methodology log, if the advertised coverage isn't reproducible, or if client data isn't separated.
Reusable asset: run log
The file ai-visibility-protocol.csv provides the minimal columns: UTC date, market, language, platform, visible version, prompt, run, mention, citation, URL, position, controlled sentiment, competitor, sessions, and conversions. The example row must be replaced; it isn't a SEOryon result.
Where SEOryon fits in
SEOryon should be tested like any other tracker: exact prompts, multiple runs, raw export, known rules, and comparison with proprietary data. The protocol published here lets you run this test independently of the product. Only keep a SEOryon measurement if you can reproduce its calculation and document its scope; appearing on this page attests to neither superiority nor universal coverage.
Measurable exercise
Choose ten prompts, two platforms, and five repetitions, i.e. 100 planned runs. Log parameters and responses, calculate mentions, citations, coverage, and stability, then redo the calculation from the vendor's export. Success: a gap of less than one percentage point or a fully explained difference, zero hidden prompts, all errors classified, and at least one specific operational decision per platform.
Where to go next
- Go back to the guide for choosing SEO and AI visibility software.
- First build a measurement of AI visibility with repetitions and a denominator.
- Standardize outputs with the definitions of mention, citation, and AI share of voice.
- Then check whether citations diverge from classic Google rankings.
FAQ
Are citation and mention the same thing?
No. A response can name a brand without showing its domain, or cite an article without explicitly naming the brand in the text. Measure both.
How many prompts should you track?
There's no universal number. Cover the decision stages and markets that matter, then sample enough repetitions to see the variance. 30 defensible prompts beat 3,000 opaque ones.
Can you compare two vendors' scores?
Only if prompts, platforms, markets, dates, repetitions, normalization, and formula are equivalent. Otherwise, compare the raw observations or run a joint test.
Does Google Search Console replace a GEO tracker?
It brings proprietary data on Google's generative surfaces for eligible properties. It doesn't cover every platform, doesn't replace a prompt panel, and doesn't measure the whole business journey.
Does a rise in citations guarantee traffic?
No. It increases an exposure measured within the panel. The click depends on the surface, the intent, the response, and the interface; measure sessions and conversions separately.
References
- Google: Introducing Search Generative AI performance reports- Microsoft Bing: AI Performance in Bing Webmaster Tools- Google: Guide to optimizing for generative AI features- Semrush: AI Visibility Index methodology- Semrush: announcement of the 126-million-prompt index- Google: AI features and your website
Method note
Reviewed 16 July 2026, translated and edited 22 July 2026. Vendor interfaces, models, surfaces, reports, and formulas change quickly. Revalidate coverage and archive the aggregation rules at every measurement wave; don't rewrite historical results with a new formula without versioning them.