Quick answer
The SEOryon Observatory isn't looking for "the true number" in AI search: it makes measurements comparable that don't cover the same thing. A panel study tracks people, a keyword study observes SERPs, Google describes its product, and a visibility tool samples prompts. The usable output is therefore a register where every number keeps its sample, period, geography, unit of analysis, and limitations. Without those five elements, an impressive percentage isn't a basis for a decision.
Key takeaways
- "Zero-click," a CTR drop, an AI citation, a brand mention, and a referral visit are five different metrics.
- Two apparently contradictory results can both be valid within their own population and period.
- Official documentation proves a surface's stated behavior, not the causal effect of an optimization.
- Relative figures must always be translated into percentage points and expected volumes.
- The right decision comes from a reproducible protocol applied to your market, not from a vendor's average.
The Observatory's evidence model
We classify every claim into one of four categories. This classification avoids giving the same weight to an official rule and to a practitioner's opinion.
| Level | What it contains | What it allows you to write | Example |
|---|---|---|---|
| Official behavior | Engine documentation, a metric definition, a policy | "Google states that…" | Google describes query fan-out and the eligibility conditions |
| Observed behavior | A panel, analytics data, or a measured experiment | "In this sample and this period…" | Pew observes the clicks of 900 American adults |
| Estimate or model | A projection, estimated traffic, a calculated counterfactual | "The model estimates… under these assumptions" | Ahrefs estimates a counterfactual CTR with no AI Overview |
| Operating hypothesis | A plausible explanation to test | "We will test whether…" | Does a more precise answer increase citations? |
A study can occupy several levels. Google's documentation on AI features proves that Google can launch several related searches through query fan-out. It doesn't prove that adding ten subheadings to a page increases its probability of being cited. That second proposition is a hypothesis, unless it's evaluated in a controlled protocol.
Comparative register of the main studies
The table below keeps the unit of analysis next to the result. That's the minimum needed before putting two figures side by side.
| Source | Population, period, and territory | Unit and method | Useful result | Limitation that changes the interpretation |
|---|---|---|---|---|
| SparkToro / Similarweb, zero-click 2026 | Mobile and desktop web panel in the United States, January to April 2026 | Journeys after a Google search; weighting estimated at two thirds mobile | 68.01% of observed searches end with no click | The Google app is excluded, mobile ad clicks are estimated, and the panel size isn't published in the article |
| Pew Research Center | 900 American adults, 68,879 Google searches in March 2025 | Measured navigation; SERPs recreated from 7 to 17 April | 8% clicks on a classic result when an AI summary is present, versus 15% without | The SERP was collected after the visit and could have changed |
| Ahrefs, AI Overviews CTR | 300,000 keywords, December 2023 and 2025 data, desktop | Two cohorts and aggregated GSC data; modeled counterfactual | Position 1 CTR observed at 1.57% versus 3.73% estimated without an AIO, roughly 58% lower in relative terms | Correlation and a model, not a randomized experiment; desktop only |
| Seer Interactive | 53 brands, 5.47 million queries and 2.43 billion impressions, January 2025 to February 2026 | Full cohort, intents classified; recent AIO status applied retroactively | AIO CTR from 1.3% in December 2025 to 2.4% in February 2026 | A client sample; a static AIO label; no seasonal correction |
| Semrush / Datos, AIO study | More than 10 million keywords; January to November 2025 | Keyword database, 200,000 terms for zero-click, before and after cohort | In the cohort of terms that gained an AIO, zero-click falls from 33.75% to 31.53% | A volume threshold above 100 and proprietary classifications; other SERP changes not controlled |
| Ahrefs, citations and the top 10 | 863,000 SERPs and 4 million cited URLs, study published in March 2026 | The same URL in the AIO and in the classic result for the same query | 37.1% of citations rank in the top 10; 36.7% aren't in the top 100 | Parsing changed since the 2025 study; a snapshot of a volatile system |
| Adobe Digital Insights | More than one trillion visits to US retail sites; January to March 2026, with a survey of more than 5,000 Americans | Adobe Analytics transactions and a separate survey | In March, AI visits convert 42% better than non-AI traffic within this scope | US commerce only; the detailed definition of sources and the absolute rates aren't provided in the article |
| Semrush, Reddit and AI search | 217,000 prompts and 248,000 Reddit URLs; data updated in October 2025 | Public outputs from Google AI Mode, Perplexity, and ChatGPT Search | Reddit appears in 13%, 4%, and 9% of answers depending on the platform; more than half of the citations are Q&A threads | Prompts come from the vendor and the platforms are in a dated state; no causality with votes |
This register isn't there to crown a winning study. It's there to help you pick the one whose population most resembles the decision under review. To forecast organic clicks for a portfolio of queries, a cohort of queries plus GSC is closer to the problem than a survey. To understand what people do after a search, a browsing panel is more relevant than a keyword database.
Procedure: auditing a figure in seven steps
- Write the decision before the figure. For example: "should we lower the click forecast for informational queries that trigger an AIO?"
- Identify the denominator. Searches, impressions, sessions, users, AI answers, and citations aren't interchangeable.
- Record the population. Note the language, country, device, sector, recruitment source, and exclusions.
- Fix the period. Interfaces and models change fast; a value with no month of observation ages badly.
- Qualify the study design. Observation, cohort comparison, before and after, declarative survey, model, or experiment.
- Calculate the absolute effect. Convert the relative percentage into clicks, leads, or euros within your own volumes.
- Record the decisive limitation. Not a generic sentence: the limitation capable of reversing your decision.
Worked example: turning "58% lower" into a scenario
An agency manages 120,000 monthly impressions on a group of queries where position 1 historically earned a 4% CTR, meaning 4,800 clicks. Mechanically applying a 58% relative drop gives:
scenario CTR = 4% × (1 - 0.58) = 1.68%
scenario clicks = 120,000 × 1.68% = 2,016
The modeled loss is 2,784 clicks, but that isn't a sufficient forecast. The Ahrefs study covered informational queries, on desktop, in December 2025, with a counterfactual. At minimum the agency should produce three scenarios (relative drops of 25%, 40%, and 58%), then segment by intent and check which terms actually trigger an AIO. If only 45% of impressions are exposed, the high scenario becomes 54,000 × 1.68% + 66,000 × 4% = 3,547 clicks, not 2,016.
How to reconcile contradictory results
Ahrefs, Pew, Seer, and Semrush aren't asking the same question. Pew compares visits with and without an AI summary in a human panel. Ahrefs estimates the CTR gap for the first position across keyword cohorts. Seer tracks brands and observes a recovery in early 2026. Semrush/Datos tracks the same set of terms before and after an AIO appears and finds a slight drop in zero-click.
The contradiction shrinks once you align eight columns: period, territory, device, intent, query selection, definition of exposure, definition of a click, and study design. What remains after alignment is genuine uncertainty. It should become a scenario range, not be hidden behind an average.
What the data proves and doesn't prove
The sources converge on one point: generated answers change how clicks are distributed and make it necessary to measure more than position. They also show strong heterogeneity by intent, platform, and moment.
They don't prove that an AI citation causes a sale, that all AI traffic converts better, that every AIO removes 58% of clicks, or that Reddit is favored by a particular algorithmic signal. Nor do they let you attribute an overall traffic drop to a single feature. An update, demand, the season, a ranking change, or the device mix can all act at the same time.
Common mistakes and stopping conditions
- Adding up followed citations across different platforms without normalizing the number of prompts.
- Comparing a rate per answer with a rate per cited URL.
- Presenting a rise from 1% to 2% as "+1%" instead of +1 point and +100% relative.
- Reusing a US value for a French market with no local test.
- Treating a 2028 projection as an observation.
- Stopping the analysis if the source publishes neither denominator nor period: the figure can be mentioned as a claim, not as a benchmark.
Reusable asset: the evidence register
Download or copy the SEOryon evidence register. For each claim, fill in: claim_id, exact wording, URL, publisher, publication date, data period, territory, sample, unit, method, result, limitation, verification date, and the page using it. Add two decision fields: applicable_yes_no and reason.
The validation rule is simple: a figure only reaches publication if the period, population, method, and limitation fields are filled in. Documentation with no statistic can pass with the note "official source: stated behavior."
How SEOryon fits in
SEOryon can centralize observations from Search Console, tracked pages, and visibility tests when those functions are enabled in the product. The relevant role here is keeping dated series and comparable segments, not manufacturing a universal score. Third party data must keep its provenance and its limitations. Before any commercial claim about a feature, check that it's actually public on seoryon.com and available in the plan concerned.
Measurable exercise
Pick a claim used in a client report, for example "AIOs cut our traffic." Deliver a one page sheet with: the target decision, the denominator, the period, the country and device, three alternative causes, two contradictory sources, and a test on your own data. The exercise succeeds if a second analyst can reproduce the calculation and identify precisely what would make you abandon the hypothesis.
FAQ
What's the best figure for measuring AI search?
There is no single figure. Use impressions and clicks for acquisition, mentions and citations per prompt for visibility, then conversions and revenue for value. Keep the denominators separate.
Can you compare two GEO tracking tools?
Yes, provided you use the same prompt set, the same language, the same region, the same model, the same day, and several repeats. Then compare reproducibility and coverage, not just the raw number of citations.
Is a vendor study unusable?
No. It can be very useful if its sample and method are published. Flag its commercial interest and avoid extending its results beyond the observed population.
How often should this observatory be updated?
Official documentation and measurement interfaces should be reviewed monthly. Benchmarks can be reviewed quarterly, with an immediate alert whenever a source changes its method.
Recommended path through the Observatory
Start with what zero-click measures in 2026, then compare the AI Overviews CTR studies. For source selection, continue with AI citations and Google rankings. Then connect visibility to AI traffic conversion and examine Reddit, forums, and UGC. The neighboring frameworks are the AI search guide and the link based authority method.
References
- Google: AI features and your website- Google: Generative AI optimization guide- Pew Research Center: click behavior with an AI summary- SparkToro / Similarweb: zero-click 2026- Ahrefs: AI Overviews and CTR- Seer Interactive: 2026 CTR update- Semrush / Datos: AI Overviews study- Adobe Digital Insights: AI traffic in US retail
Method and update note
Reviewed 16 July 2026, translated and edited 22 July 2026. Figures are reproduced with their population and study design as published. Product links, parsing methods, and platform scopes can change; every revision must keep the old value and explain any break in the series.