AI Overview citations are volatile enough that a one-time screenshot is weak evidence. Ahrefs found that 45.5% of cited URLs changed between consecutive responses, while entity overlap was more stable at about 54%. The practical response is to repeat a frozen prompt panel, report inclusion frequency and source churn, then act only when the apparent change survives those repeats.

That answer is deliberately narrower than a claim that citations are random or that a page has permanently won or lost visibility. The observation concerns cited URLs in consecutive responses in one third-party study. It does not turn a small panel into a market estimate, prove why a source changed, or establish a universal ranking effect. For publishers, that restraint is useful: it tells you what to measure before assigning an expensive content, technical, or procurement response to a noisy signal.

What the evidence supports, and what it does not

The Ahrefs finding is simple to repeat but easy to overread. In the study, 45.5% of cited URLs changed between consecutive observations. At the same time, answer entities overlapped more consistently, at about 54%. Exact grounding sources can therefore rotate even when the answer is still about much the same set of entities, and citation presence should not be read as a proxy for classic ranking position either, since only about 38% of AI Overview citations rank in Google's top 10.

Question Evidence boundary Operational reading
Do cited URLs change? Ahrefs observed 45.5% URL change between consecutive responses. One captured citation position is a noisy observation.
Can the answer remain conceptually similar? Entity overlap was about 54%. Track answer entities and sources as separate signals.
Can a prompt panel stand for all searchers? No. Prompts, locale, account context, surface and time can affect output. Freeze inputs or log each difference.
Does a citation change prove a content problem? No causal test is supplied. Diagnose across repeated windows before changing a page.

Diagram of a repeat-sampling protocol for AI Overview citation volatility

The distinction between an entity and a URL is central. A response can continue to mention the same brand, product, condition, place, or concept while citing a different page. Conversely, the same domain can appear through a different URL after an editorial refresh, a canonical change, or a source-selection shift. Conflating these observations makes the resulting dashboard look more certain than it is.

Google's guidance on AI features and your website is relevant as a scope check, not as proof of a citation tactic. It frames AI-feature eligibility alongside established Search guidance. The Search Essentials remain a useful reference when a team is tempted to treat an unstable citation as evidence that it should bypass basic crawlability, policy, or quality checks.

Freeze the observation before you interpret it

A repeat is meaningful only when it is comparable to the first observation. For every prompt, make a small capture record before running it. The record need not be elaborate, but it must prevent a later analyst from mistaking a changed setup for a changed result.

Field Record it this way Why it matters
Prompt text Store the exact string, including qualifiers. A rewritten question can retrieve a different answer.
Locale Record country and language. Local intent and language can alter sources.
Surface or model Name the AI-search surface observed. Different surfaces are not interchangeable samples.
Account context Note signed-in, signed-out, and known personalization state. Context may influence the response.
Capture time Use a timestamp and observation window. A date-less screenshot cannot be compared reliably.
Raw output Save the response and every cited URL. A count without evidence cannot be audited.

This is panel discipline, not bureaucratic overhead. If a team first captures a prompt in France, then reruns a translated version in Germany while signed in, it has changed several inputs at once. Any difference may be real, but it cannot honestly be called source churn for the original observation. A frozen panel creates a modest, defensible unit of analysis: the result of this exact prompt, in this documented context, at this point in time.

The Ahrefs analysis of AI Overview changes is the dated source for the 45.5% figure. Keep its observation date and methodology attached to the number. A future change in its parser, sample, or surface may be a methodology break rather than evidence that volatility itself rose or fell.

The minimum repeat-sampling protocol

Run each business-critical prompt at least three times per observation window. Three is a floor for detecting obvious instability, not a magic sample size. Use the same documented setup for each run, capture the raw answer and citations, then repeat the whole window before declaring a gain or loss. Two or more windows are the minimum practical guardrail unless a verified technical failure requires immediate action.

For each prompt, calculate three related measures:

  1. URL inclusion frequency: the number of runs citing a specific URL divided by the number of comparable runs.
  2. Domain inclusion frequency: the number of runs citing any URL on a domain divided by comparable runs.
  3. Source churn: the degree to which the cited-source set differs from one run to the next. A Jaccard comparison is a transparent option.

For source sets A and B, Jaccard similarity is |A intersection B| / |A union B|. Define churn as 1 minus similarity. If run one cites {a.com/x, b.com/y, c.com/z} and run two cites {a.com/x, b.com/new, d.com/q}, the intersection contains one URL and the union contains five. Similarity is 1/5, or 20%; churn is 80%.

That calculation does not say that 80% of the answer changed, nor that any website became worse. It records the difference between two documented sets. A domain-level version can be useful alongside it: in the example, the shared domain a.com may tell a different story from exact URL overlap. Report both rather than choosing the friendlier number.

A worked decision example

Consider a six-run panel for one frozen query across two weekly windows. Your URL appears in two of the first three runs and in one of the next three. Its URL inclusion frequency is 3/6, or 50%. Your domain appears in four runs because a second page is cited once, so domain inclusion is 4/6, or about 66.7%.

The raw counts matter more than the percentages alone. A move from 2/3 to 1/3 looks dramatic in a chart, but it represents one changed observation in each three-run window. If the page was technically available, the prompt and context were frozen, and a second window produces the same pattern, a content or canonical review may be justified. If the next window returns 2/3, the earlier dip was not enough evidence for a confident diagnosis.

Use a simple decision table so the response matches the evidence.

Pattern after comparable repeats Most defensible reading Next action
One missing citation, with later reappearance Normal panel variation remains plausible. Record it; do not rush a rewrite.
URL falls, domain remains present URL selection changed, not necessarily topical presence. Check canonical and page roles before changing content.
Domain is absent in two or more windows A signal worth investigation, not a causal conclusion. Review coverage and technical availability, then test a controlled action.
Verified rendering, indexing, or canonical failure The measurement issue has a known technical cause. Fix the verified failure and retain the pre-fix evidence.

Failure tests that prevent false alarms

Before treating absence as a visibility loss, test whether the observation itself failed. Confirm that the cited page still resolves as expected, the canonical decision has not changed, and the capture did not silently switch locale, account state, or prompt wording. These checks do not establish an AI-citation cause. They remove a few avoidable alternatives.

Do not use a dashboard threshold as a substitute for judgment. A low inclusion rate can result from a genuinely mixed source panel, a narrow prompt, a context difference, or a change outside the site's control. A high rate is also not a promise of future appearance. The responsible report names the numerator, denominator, dates, and setup before it labels a trend.

For a query-specific operating checklist, keep the following with the evidence packet:

  • Exact prompt and prompt version
  • Locale, language, surface, account state, and timestamps
  • Raw responses, cited URLs, extracted domains, and entity notes
  • URL and domain inclusion numerators and denominators
  • Pairwise source-churn calculation and the sets used
  • Technical checks performed, including any verified failure
  • The decision taken, or an explicit decision to take no action

The last item matters. If repeated evidence does not support an intervention, no action is a valid result. It preserves budget and avoids turning normal output variation into serial, unmeasured rewrites.

Where SEOryon fits

For this AI Overview citation volatility update, SEOryon belongs in the operating layer rather than the headline. Use it to connect an observed signal to current site coverage, choose refresh or new content deliberately, publish through the chosen approval mode, and measure the result later. It should not be used to manufacture certainty from one capture.

SEOryon's role here is limited to research, a canonical decision, controlled content action, publishing, and later measurement. This article does not extend that verified scope into undocumented product, integration, security, or performance claims. Evaluate SEOryon on your own site with one controlled topic cluster before expanding automation, or compare it first against other AI visibility tools that turn citation gaps into concrete page fixes. That remains the appropriate procurement caveat because the protocol is an observation method, not proof that a tool will produce a particular outcome.

Frequently asked questions

How often should a publisher repeat an AI Overview query?

Run a business-critical frozen query at least three times in an observation window, then compare at least two windows before calling a gain or loss. This is a practical noise check, not a claim that three runs represent all users or all queries.

Should URL and domain citations be reported together?

Yes. Exact URLs can rotate while the same domain remains present through another page. Reporting both inclusion frequencies prevents a URL-level change from being misread as a complete loss of topical presence.

Does repeated absence prove that a page is less useful?

No. Repeated absence is a signal to investigate after setup and technical checks are documented. The cited evidence does not identify a causal mechanism, and it does not guarantee that a content change will restore a citation.

Sources and evidence notes

Official documentation supports product and Search guidance. The 45.5% URL-change finding comes from a third-party observational study; it should retain its sample, method, and observation date rather than becoming a universal causal claim.