From 9da1dec9e6a360f71bf7de34690f7769f58eb520 Mon Sep 17 00:00:00 2001 From: Bastien Chanot Date: Thu, 16 Jul 2026 20:32:55 +0200 Subject: [PATCH] =?UTF-8?q?fix(geo):=20I6=20=E2=80=94=20every=20stat=20was?= =?UTF-8?q?=20real=20and=20attached=20to=20the=20wrong=20claim?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Audited each statistic in agents/resources/ against primary sources after the VSI fiction (I2) showed WebSearch launders SEO-blog consensus. The failure mode is not invention — it is plausible recombination, which is what a model half-remembering a search result produces: - "Adding statistics increases AI visibility by up to 40% (Aggarwal et al.)" — paper real (KDD 2024), number real, SCOPE WRONG: 40% is the aggregate over the whole method set, domain-dependent. No per-technique figure exists. - "Pages not updated quarterly are 3x more likely to lose AI citations (LLMRefs)" — LLMrefs' actual 3x says brand mentions correlate ~3x more strongly with AI visibility than backlinks. DIFFERENT SUBJECT. No source supports a quarterly decay multiplier. - "QAPage cited 58% more often than Article" — uncited. Nearest real number: AccuraCast 2025, `Person` schema at 58.9% PREVALENCE among cited sources — wrong type, and its FAQPage figure (1.8%) points the opposite way to the claim it propped up. This one drove Tier 1 ranking. - "62% of searches involve voice" — uncited; 62% circulates as smart-speaker ADOPTION. Same family as the "50% by 2020" myth ComScore denied (origin: a 2014 Andrew Ng interview). Corrected my own framing too: I claimed three times these stats "drive axis weights". They do not — the weight tables carry no citations. They drive Tier/priority recommendations and, worse, geo-analyzer's "Cite sources" rule pushed them into CLIENT reports as research-backed. Fixes: recommendations kept on mechanism, fabricated numbers removed with the incident documented inline so they are not re-added. Unverified stats (48% AI Overviews, 2.5B queries/day, Gartner -25%) labelled [UNVERIFIED] rather than asserted or deleted — I did not check them. Structural, not just exhortation: resources/README.md now mandates ` — — measured: — `. `measured:` is the field that catches this — all four errors survive a source name; none survives stating the real measurement next to the claim. WebSearch demoted from verification to crawler/tool-name lookup only. Verified: make test 35 GREEN / 0 RED. --- agents/geo-analyzer.md | 22 ++++++++--- agents/resources/README.md | 47 +++++++++++++++++++++++- agents/resources/ai-visibility-tools.md | 14 +++++-- agents/resources/content-shape-for-ai.md | 31 +++++++++++++--- agents/resources/geo-schemas.md | 28 ++++++++++++-- 5 files changed, 123 insertions(+), 19 deletions(-) diff --git a/agents/geo-analyzer.md b/agents/geo-analyzer.md index c3e9e3b..b3034db 100644 --- a/agents/geo-analyzer.md +++ b/agents/geo-analyzer.md @@ -13,10 +13,13 @@ Apple Intelligence**. Google classical search is handled by the ## Context — why GEO is its own discipline in 2026 -- AI Overviews trigger on ~48% of Google searches (April 2026). -- ChatGPT processes 2.5B queries/day. -- Gartner projects commercial organic search traffic to fall 25% by - end-2026 as discovery shifts to AI engines. +- `[UNVERIFIED — 2026-07-16]` AI Overviews trigger on ~48% of Google + searches (April 2026); ChatGPT processes 2.5B queries/day; Gartner + projects commercial organic search traffic to fall 25% by end-2026 as + discovery shifts to AI engines. Framing only — **never quote these to a + client** until each carries `source + measured: + link` per + `resources/README.md`. GEO is worth doing on mechanism; it does not need + these numbers to be true. - Classical SEO ≠ GEO. Some signals overlap (headings, Schema.org) but the optimization levers differ: entity clarity, definition architecture, citable stats, crawler permissions. @@ -936,8 +939,15 @@ PROCHAINE ETAPE : unknown fields → `[À COMPLÉTER]`, never a value copied from a sibling on-site source. - **Remove deprecated schemas rather than keep broken ones.** -- **Cite sources.** When emitting stats in the report, link - `content-shape-for-ai.md` research citations. +- **Cite sources, and only citable ones.** A stat reaches the client only + if it carries `source + measured: + link` per `resources/README.md`. + Anything marked `[UNVERIFIED]` is framing for you, never a line in the + report. Quote the source's ACTUAL measurement, never a widened or + re-subjected version of it — the 2026-07-16 audit found every stat in + that directory real but attached to the wrong claim, and this rule is + what pushed them into client deliverables as research-backed. + A recommendation that only stands up with a number you cannot source was + never standing up: make it on mechanism, or drop it. ### Process - **Every user action lists automation options.** Mandatory from diff --git a/agents/resources/README.md b/agents/resources/README.md index e685885..6283ea0 100644 --- a/agents/resources/README.md +++ b/agents/resources/README.md @@ -17,7 +17,52 @@ Loaded on demand — keep each file focused and current. These files capture state as of 2026-04. Crawler lists, Schema.org deprecations, and tool landscape shift fast. Agents MUST cross-check -via WebSearch on each run when FULL depth is selected. +crawler lists and tool names via WebSearch on each run when FULL depth is +selected. + +## Citation standard (mandatory for every statistic) + +**WebSearch is NOT verification for a number.** It ranks SEO blogs, and SEO +blogs cross-cite each other into a consensus that looks like corroboration. +Two 2026-07-16 audits of this directory show how it fails: + +- A "VSI (Visual Stability Index) — new 2026 Core Web Vital" lived in + `seo-analyzer.md`. Ten blogs asserted it; several claimed CrUX already + collected it. It is absent from the CrUX API metric list and from + web.dev. WebSearch returned the echo, not the truth. +- Every stat in this directory was real **and attached to the wrong + subject**: the GEO paper's 40% (all methods) pinned on one technique; + LLMrefs' 3x (brand mentions vs backlinks) pinned on freshness decay; + AccuraCast's 58.9% (Person schema prevalence) pinned on QAPage lift, with + its meaning inverted; a smart-speaker adoption figure sold as voice-search + share. + +The failure mode is not invention — it is **plausible recombination**, which +is exactly what a model half-remembering a search result produces. So the +format has to make an unsourced number conspicuous: + +``` + — — measured: — +``` + +`measured:` is the field that catches it. All four errors above survive a +source name; none survives having to state the source's real measurement +next to the claim. + +Rules: +1. **Primary source or no number.** Peer-reviewed paper, the vendor's own + published study, or an official API/doc. `developer.chrome.com/docs/crux` + is decisive for metrics: what CrUX cannot return, we cannot score. +2. **Name the tier.** Peer review ≠ vendor marketing. LLMrefs, AccuraCast, + Ahrefs publish useful data and sell products — say "vendor". +3. **Never widen scope.** An aggregate result is not a per-technique result. +4. **No number beats a wrong number.** A recommendation that only stands up + with a fabricated statistic was never standing up. Delete the stat, keep + the recommendation if it survives on mechanism. +5. **Unverified ⇒ labelled.** `[UNVERIFIED — ]` inline. Never quote an + unverified number to a client: `geo-analyzer.md` ("Cite sources") sends + these into client reports as research-backed. ## Loading pattern diff --git a/agents/resources/ai-visibility-tools.md b/agents/resources/ai-visibility-tools.md index 5f54efe..7c62e80 100644 --- a/agents/resources/ai-visibility-tools.md +++ b/agents/resources/ai-visibility-tools.md @@ -4,9 +4,17 @@ Tools that track whether your brand appears in AI-generated answers across ChatGPT, Perplexity, Gemini, Copilot, Claude, and Google AI Overviews. -Context: Google AI Overviews trigger on ~48% of searches; ChatGPT -processes 2.5B queries/day; Gartner projects commercial organic -search traffic will drop 25% by 2026. Monitoring is no longer optional. +Context `[UNVERIFIED — 2026-07-16]`: Google AI Overviews trigger on ~48% of +searches; ChatGPT processes 2.5B queries/day; Gartner projects commercial +organic search traffic will drop 25% by 2026. + +> Not checked against primary sources in the 2026-07-16 audit that corrected +> the rest of this directory — flagged rather than asserted or deleted, per +> the citation standard in `README.md` (rule 5). The Gartner projection at +> least names its source; the other two float. Treat all three as +> motivation, not evidence: **do NOT quote them to a client** until each +> carries `source + measured: + link`. Their only job here is to explain why +> this file exists, and that argument does not need numbers. ## Commercial tools diff --git a/agents/resources/content-shape-for-ai.md b/agents/resources/content-shape-for-ai.md index 8c0f6a3..592eccc 100644 --- a/agents/resources/content-shape-for-ai.md +++ b/agents/resources/content-shape-for-ai.md @@ -61,9 +61,18 @@ query. A one-sentence self-contained answer has the highest density. ### 4. Citations and statistics (strongest measured lever) -Adding peer-cited statistics with clear sources increases AI visibility -**by up to 40%** (Aggarwal et al., 2024 "GEO: Generative Engine -Optimization"). +Aggarwal et al., 2024 ("GEO: Generative Engine Optimization", KDD 2024) +report that their optimisation methods **collectively** boost visibility +**by up to 40%** in generative-engine responses, and state the effect +**varies across domains**. Citations/statistics/quotations are among those +methods. + +> **Attribute this correctly.** Until 2026-07-16 this section read "Adding +> peer-cited statistics with clear sources increases AI visibility by up to +> 40%" — pinning the paper's *aggregate* result on this *one* technique. The +> paper publishes no separate figure per technique. When quoting it to a +> client: "up to 40%, across the method set, domain-dependent" — never "+40% +> if you add stats". Pattern: embed specific numbers with attribution. @@ -100,8 +109,20 @@ Comparison tables are even stronger. Structure: ### 6. Freshness signals -Pages not updated at least quarterly are **3x more likely to lose AI -citations** (LLMRefs 2026 study). +Freshness is a real retrieval input: RAG systems fetch live and read +timestamps, so a page updated this quarter carries a stronger recency +signal than the same page last touched years ago. LLMrefs (a **vendor**, +not peer review) reports cited content running **~25.7% fresher** than +organic top-10 across ~17M citations. Substantive updates only — bumping a +date string is not freshness. + +> **The "3x" that lived here was grafted from another claim.** Until +> 2026-07-16 this read "Pages not updated at least quarterly are 3x more +> likely to lose AI citations (LLMRefs 2026 study)". LLMrefs' actual "3x" +> says **brand mentions correlate ~3x more strongly with AI visibility than +> backlinks** — a different subject entirely. No source supports a quarterly +> decay multiplier. Recommend quarterly refresh on its merits; do not price +> it with a borrowed number. What to maintain: - Visible "Last updated: YYYY-MM-DD" at the top of content pages diff --git a/agents/resources/geo-schemas.md b/agents/resources/geo-schemas.md index da2f746..9d0eba9 100644 --- a/agents/resources/geo-schemas.md +++ b/agents/resources/geo-schemas.md @@ -21,8 +21,20 @@ existing instances. They no longer produce rich results. ### QAPage — single Q&A format -Pages cited 58% more often by ChatGPT vs basic Article schema. -Use when the page is built around ONE primary question. +Use when the page is built around ONE primary question. Emitting the type +that matches the content shape beats wrapping everything in a generic +`Article`. + +> **No lift figure here — the one that lived here was wrong.** Until +> 2026-07-16 this read "Pages cited 58% more often by ChatGPT vs basic +> Article schema", uncited. Nothing supports it. The nearest real number is +> AccuraCast 2025 (~2,000 prompts across ChatGPT / AI Overviews / +> Perplexity, ~9,000 cited sources): **`Person` schema appeared in 58.9%** +> of cited sources — a *prevalence* count for a *different type* — while +> **`FAQPage` appeared in 1.8%**, which points the opposite way to the claim +> it was propping up. Q&A shape is still worth doing on genuinely +> single-question pages; it is not worth a fabricated number. Do NOT quote a +> QAPage lift % to a client — there isn't one. ```json { @@ -81,8 +93,16 @@ visible content. ### Speakable — voice + AI extraction marker -62% of searches in 2026 involve voice. Speakable flags the passage -best suited for voice readout and AI summary. +Speakable flags the passage best suited for voice readout and AI summary. + +> **No voice-share figure — the one that lived here was a conflation.** +> Until 2026-07-16 this read "62% of searches in 2026 involve voice", +> uncited. No primary source carries it; 62% circulates as a *smart-speaker +> adoption* number, not a share of searches. It is the same family as the +> "50% of searches will be voice by 2020" myth — attributed to ComScore, +> who **denied it**; the real origin is a 2014 Andrew Ng interview. Speakable +> is cheap and harmless, so keep recommending it on TL;DR / summary blocks — +> but justify it by extraction shape, never by a voice-share statistic. ```json {