forked from bchanot/claude
fix(geo): I6 — every stat was real and attached to the wrong claim
Audited each statistic in agents/resources/ against primary sources after the VSI fiction (I2) showed WebSearch launders SEO-blog consensus. The failure mode is not invention — it is plausible recombination, which is what a model half-remembering a search result produces: - "Adding statistics increases AI visibility by up to 40% (Aggarwal et al.)" — paper real (KDD 2024), number real, SCOPE WRONG: 40% is the aggregate over the whole method set, domain-dependent. No per-technique figure exists. - "Pages not updated quarterly are 3x more likely to lose AI citations (LLMRefs)" — LLMrefs' actual 3x says brand mentions correlate ~3x more strongly with AI visibility than backlinks. DIFFERENT SUBJECT. No source supports a quarterly decay multiplier. - "QAPage cited 58% more often than Article" — uncited. Nearest real number: AccuraCast 2025, `Person` schema at 58.9% PREVALENCE among cited sources — wrong type, and its FAQPage figure (1.8%) points the opposite way to the claim it propped up. This one drove Tier 1 ranking. - "62% of searches involve voice" — uncited; 62% circulates as smart-speaker ADOPTION. Same family as the "50% by 2020" myth ComScore denied (origin: a 2014 Andrew Ng interview). Corrected my own framing too: I claimed three times these stats "drive axis weights". They do not — the weight tables carry no citations. They drive Tier/priority recommendations and, worse, geo-analyzer's "Cite sources" rule pushed them into CLIENT reports as research-backed. Fixes: recommendations kept on mechanism, fabricated numbers removed with the incident documented inline so they are not re-added. Unverified stats (48% AI Overviews, 2.5B queries/day, Gartner -25%) labelled [UNVERIFIED] rather than asserted or deleted — I did not check them. Structural, not just exhortation: resources/README.md now mandates `<claim> — <source, year, venue|vendor> — measured: <what the source ACTUALLY measured> — <link>`. `measured:` is the field that catches this — all four errors survive a source name; none survives stating the real measurement next to the claim. WebSearch demoted from verification to crawler/tool-name lookup only. Verified: make test 35 GREEN / 0 RED.
This commit is contained in:
@@ -21,8 +21,20 @@ existing instances. They no longer produce rich results.
|
||||
|
||||
### QAPage — single Q&A format
|
||||
|
||||
Pages cited 58% more often by ChatGPT vs basic Article schema.
|
||||
Use when the page is built around ONE primary question.
|
||||
Use when the page is built around ONE primary question. Emitting the type
|
||||
that matches the content shape beats wrapping everything in a generic
|
||||
`Article`.
|
||||
|
||||
> **No lift figure here — the one that lived here was wrong.** Until
|
||||
> 2026-07-16 this read "Pages cited 58% more often by ChatGPT vs basic
|
||||
> Article schema", uncited. Nothing supports it. The nearest real number is
|
||||
> AccuraCast 2025 (~2,000 prompts across ChatGPT / AI Overviews /
|
||||
> Perplexity, ~9,000 cited sources): **`Person` schema appeared in 58.9%**
|
||||
> of cited sources — a *prevalence* count for a *different type* — while
|
||||
> **`FAQPage` appeared in 1.8%**, which points the opposite way to the claim
|
||||
> it was propping up. Q&A shape is still worth doing on genuinely
|
||||
> single-question pages; it is not worth a fabricated number. Do NOT quote a
|
||||
> QAPage lift % to a client — there isn't one.
|
||||
|
||||
```json
|
||||
{
|
||||
@@ -81,8 +93,16 @@ visible content.
|
||||
|
||||
### Speakable — voice + AI extraction marker
|
||||
|
||||
62% of searches in 2026 involve voice. Speakable flags the passage
|
||||
best suited for voice readout and AI summary.
|
||||
Speakable flags the passage best suited for voice readout and AI summary.
|
||||
|
||||
> **No voice-share figure — the one that lived here was a conflation.**
|
||||
> Until 2026-07-16 this read "62% of searches in 2026 involve voice",
|
||||
> uncited. No primary source carries it; 62% circulates as a *smart-speaker
|
||||
> adoption* number, not a share of searches. It is the same family as the
|
||||
> "50% of searches will be voice by 2020" myth — attributed to ComScore,
|
||||
> who **denied it**; the real origin is a 2014 Andrew Ng interview. Speakable
|
||||
> is cheap and harmless, so keep recommending it on TL;DR / summary blocks —
|
||||
> but justify it by extraction shape, never by a voice-share statistic.
|
||||
|
||||
```json
|
||||
{
|
||||
|
||||
Reference in New Issue
Block a user