fix(geo): I6 — every stat was real and attached to the wrong claim

Audited each statistic in agents/resources/ against primary sources after
the VSI fiction (I2) showed WebSearch launders SEO-blog consensus.

The failure mode is not invention — it is plausible recombination, which is
what a model half-remembering a search result produces:

- "Adding statistics increases AI visibility by up to 40% (Aggarwal et al.)"
  — paper real (KDD 2024), number real, SCOPE WRONG: 40% is the aggregate
  over the whole method set, domain-dependent. No per-technique figure
  exists.
- "Pages not updated quarterly are 3x more likely to lose AI citations
  (LLMRefs)" — LLMrefs' actual 3x says brand mentions correlate ~3x more
  strongly with AI visibility than backlinks. DIFFERENT SUBJECT. No source
  supports a quarterly decay multiplier.
- "QAPage cited 58% more often than Article" — uncited. Nearest real number:
  AccuraCast 2025, `Person` schema at 58.9% PREVALENCE among cited sources —
  wrong type, and its FAQPage figure (1.8%) points the opposite way to the
  claim it propped up. This one drove Tier 1 ranking.
- "62% of searches involve voice" — uncited; 62% circulates as smart-speaker
  ADOPTION. Same family as the "50% by 2020" myth ComScore denied (origin: a
  2014 Andrew Ng interview).

Corrected my own framing too: I claimed three times these stats "drive axis
weights". They do not — the weight tables carry no citations. They drive
Tier/priority recommendations and, worse, geo-analyzer's "Cite sources" rule
pushed them into CLIENT reports as research-backed.

Fixes: recommendations kept on mechanism, fabricated numbers removed with
the incident documented inline so they are not re-added. Unverified stats
(48% AI Overviews, 2.5B queries/day, Gartner -25%) labelled [UNVERIFIED]
rather than asserted or deleted — I did not check them.

Structural, not just exhortation: resources/README.md now mandates
`<claim> — <source, year, venue|vendor> — measured: <what the source ACTUALLY
measured> — <link>`. `measured:` is the field that catches this — all four
errors survive a source name; none survives stating the real measurement
next to the claim. WebSearch demoted from verification to crawler/tool-name
lookup only.

Verified: make test 35 GREEN / 0 RED.
This commit is contained in:
Bastien Chanot
2026-07-16 20:32:55 +02:00
parent e70e1d6c71
commit 9da1dec9e6
5 changed files with 123 additions and 19 deletions
+16 -6
View File
@@ -13,10 +13,13 @@ Apple Intelligence**. Google classical search is handled by the
## Context — why GEO is its own discipline in 2026
- AI Overviews trigger on ~48% of Google searches (April 2026).
- ChatGPT processes 2.5B queries/day.
- Gartner projects commercial organic search traffic to fall 25% by
end-2026 as discovery shifts to AI engines.
- `[UNVERIFIED — 2026-07-16]` AI Overviews trigger on ~48% of Google
searches (April 2026); ChatGPT processes 2.5B queries/day; Gartner
projects commercial organic search traffic to fall 25% by end-2026 as
discovery shifts to AI engines. Framing only — **never quote these to a
client** until each carries `source + measured: + link` per
`resources/README.md`. GEO is worth doing on mechanism; it does not need
these numbers to be true.
- Classical SEO ≠ GEO. Some signals overlap (headings, Schema.org)
but the optimization levers differ: entity clarity, definition
architecture, citable stats, crawler permissions.
@@ -936,8 +939,15 @@ PROCHAINE ETAPE : <highest-priority>
unknown fields → `[À COMPLÉTER]`, never a value copied from a sibling
on-site source.
- **Remove deprecated schemas rather than keep broken ones.**
- **Cite sources.** When emitting stats in the report, link
`content-shape-for-ai.md` research citations.
- **Cite sources, and only citable ones.** A stat reaches the client only
if it carries `source + measured: + link` per `resources/README.md`.
Anything marked `[UNVERIFIED]` is framing for you, never a line in the
report. Quote the source's ACTUAL measurement, never a widened or
re-subjected version of it — the 2026-07-16 audit found every stat in
that directory real but attached to the wrong claim, and this rule is
what pushed them into client deliverables as research-backed.
A recommendation that only stands up with a number you cannot source was
never standing up: make it on mechanism, or drop it.
### Process
- **Every user action lists automation options.** Mandatory from
+46 -1
View File
@@ -17,7 +17,52 @@ Loaded on demand — keep each file focused and current.
These files capture state as of 2026-04. Crawler lists, Schema.org
deprecations, and tool landscape shift fast. Agents MUST cross-check
via WebSearch on each run when FULL depth is selected.
crawler lists and tool names via WebSearch on each run when FULL depth is
selected.
## Citation standard (mandatory for every statistic)
**WebSearch is NOT verification for a number.** It ranks SEO blogs, and SEO
blogs cross-cite each other into a consensus that looks like corroboration.
Two 2026-07-16 audits of this directory show how it fails:
- A "VSI (Visual Stability Index) — new 2026 Core Web Vital" lived in
`seo-analyzer.md`. Ten blogs asserted it; several claimed CrUX already
collected it. It is absent from the CrUX API metric list and from
web.dev. WebSearch returned the echo, not the truth.
- Every stat in this directory was real **and attached to the wrong
subject**: the GEO paper's 40% (all methods) pinned on one technique;
LLMrefs' 3x (brand mentions vs backlinks) pinned on freshness decay;
AccuraCast's 58.9% (Person schema prevalence) pinned on QAPage lift, with
its meaning inverted; a smart-speaker adoption figure sold as voice-search
share.
The failure mode is not invention — it is **plausible recombination**, which
is exactly what a model half-remembering a search result produces. So the
format has to make an unsourced number conspicuous:
```
<claim> — <source, year, venue|vendor> — measured: <what the source ACTUALLY
measured> — <link>
```
`measured:` is the field that catches it. All four errors above survive a
source name; none survives having to state the source's real measurement
next to the claim.
Rules:
1. **Primary source or no number.** Peer-reviewed paper, the vendor's own
published study, or an official API/doc. `developer.chrome.com/docs/crux`
is decisive for metrics: what CrUX cannot return, we cannot score.
2. **Name the tier.** Peer review ≠ vendor marketing. LLMrefs, AccuraCast,
Ahrefs publish useful data and sell products — say "vendor".
3. **Never widen scope.** An aggregate result is not a per-technique result.
4. **No number beats a wrong number.** A recommendation that only stands up
with a fabricated statistic was never standing up. Delete the stat, keep
the recommendation if it survives on mechanism.
5. **Unverified ⇒ labelled.** `[UNVERIFIED — <date>]` inline. Never quote an
unverified number to a client: `geo-analyzer.md` ("Cite sources") sends
these into client reports as research-backed.
## Loading pattern
+11 -3
View File
@@ -4,9 +4,17 @@ Tools that track whether your brand appears in AI-generated answers
across ChatGPT, Perplexity, Gemini, Copilot, Claude, and Google AI
Overviews.
Context: Google AI Overviews trigger on ~48% of searches; ChatGPT
processes 2.5B queries/day; Gartner projects commercial organic
search traffic will drop 25% by 2026. Monitoring is no longer optional.
Context `[UNVERIFIED — 2026-07-16]`: Google AI Overviews trigger on ~48% of
searches; ChatGPT processes 2.5B queries/day; Gartner projects commercial
organic search traffic will drop 25% by 2026.
> Not checked against primary sources in the 2026-07-16 audit that corrected
> the rest of this directory — flagged rather than asserted or deleted, per
> the citation standard in `README.md` (rule 5). The Gartner projection at
> least names its source; the other two float. Treat all three as
> motivation, not evidence: **do NOT quote them to a client** until each
> carries `source + measured: + link`. Their only job here is to explain why
> this file exists, and that argument does not need numbers.
## Commercial tools
+26 -5
View File
@@ -61,9 +61,18 @@ query. A one-sentence self-contained answer has the highest density.
### 4. Citations and statistics (strongest measured lever)
Adding peer-cited statistics with clear sources increases AI visibility
**by up to 40%** (Aggarwal et al., 2024 "GEO: Generative Engine
Optimization").
Aggarwal et al., 2024 ("GEO: Generative Engine Optimization", KDD 2024)
report that their optimisation methods **collectively** boost visibility
**by up to 40%** in generative-engine responses, and state the effect
**varies across domains**. Citations/statistics/quotations are among those
methods.
> **Attribute this correctly.** Until 2026-07-16 this section read "Adding
> peer-cited statistics with clear sources increases AI visibility by up to
> 40%" — pinning the paper's *aggregate* result on this *one* technique. The
> paper publishes no separate figure per technique. When quoting it to a
> client: "up to 40%, across the method set, domain-dependent" — never "+40%
> if you add stats".
Pattern: embed specific numbers with attribution.
@@ -100,8 +109,20 @@ Comparison tables are even stronger. Structure:
### 6. Freshness signals
Pages not updated at least quarterly are **3x more likely to lose AI
citations** (LLMRefs 2026 study).
Freshness is a real retrieval input: RAG systems fetch live and read
timestamps, so a page updated this quarter carries a stronger recency
signal than the same page last touched years ago. LLMrefs (a **vendor**,
not peer review) reports cited content running **~25.7% fresher** than
organic top-10 across ~17M citations. Substantive updates only — bumping a
date string is not freshness.
> **The "3x" that lived here was grafted from another claim.** Until
> 2026-07-16 this read "Pages not updated at least quarterly are 3x more
> likely to lose AI citations (LLMRefs 2026 study)". LLMrefs' actual "3x"
> says **brand mentions correlate ~3x more strongly with AI visibility than
> backlinks** — a different subject entirely. No source supports a quarterly
> decay multiplier. Recommend quarterly refresh on its merits; do not price
> it with a borrowed number.
What to maintain:
- Visible "Last updated: YYYY-MM-DD" at the top of content pages
+24 -4
View File
@@ -21,8 +21,20 @@ existing instances. They no longer produce rich results.
### QAPage — single Q&A format
Pages cited 58% more often by ChatGPT vs basic Article schema.
Use when the page is built around ONE primary question.
Use when the page is built around ONE primary question. Emitting the type
that matches the content shape beats wrapping everything in a generic
`Article`.
> **No lift figure here — the one that lived here was wrong.** Until
> 2026-07-16 this read "Pages cited 58% more often by ChatGPT vs basic
> Article schema", uncited. Nothing supports it. The nearest real number is
> AccuraCast 2025 (~2,000 prompts across ChatGPT / AI Overviews /
> Perplexity, ~9,000 cited sources): **`Person` schema appeared in 58.9%**
> of cited sources — a *prevalence* count for a *different type* — while
> **`FAQPage` appeared in 1.8%**, which points the opposite way to the claim
> it was propping up. Q&A shape is still worth doing on genuinely
> single-question pages; it is not worth a fabricated number. Do NOT quote a
> QAPage lift % to a client — there isn't one.
```json
{
@@ -81,8 +93,16 @@ visible content.
### Speakable — voice + AI extraction marker
62% of searches in 2026 involve voice. Speakable flags the passage
best suited for voice readout and AI summary.
Speakable flags the passage best suited for voice readout and AI summary.
> **No voice-share figure — the one that lived here was a conflation.**
> Until 2026-07-16 this read "62% of searches in 2026 involve voice",
> uncited. No primary source carries it; 62% circulates as a *smart-speaker
> adoption* number, not a share of searches. It is the same family as the
> "50% of searches will be voice by 2020" myth — attributed to ComScore,
> who **denied it**; the real origin is a 2014 Andrew Ng interview. Speakable
> is cheap and harmless, so keep recommending it on TL;DR / summary blocks —
> but justify it by extraction shape, never by a voice-share statistic.
```json
{