Audited each statistic in agents/resources/ against primary sources after the VSI fiction (I2) showed WebSearch launders SEO-blog consensus. The failure mode is not invention — it is plausible recombination, which is what a model half-remembering a search result produces: - "Adding statistics increases AI visibility by up to 40% (Aggarwal et al.)" — paper real (KDD 2024), number real, SCOPE WRONG: 40% is the aggregate over the whole method set, domain-dependent. No per-technique figure exists. - "Pages not updated quarterly are 3x more likely to lose AI citations (LLMRefs)" — LLMrefs' actual 3x says brand mentions correlate ~3x more strongly with AI visibility than backlinks. DIFFERENT SUBJECT. No source supports a quarterly decay multiplier. - "QAPage cited 58% more often than Article" — uncited. Nearest real number: AccuraCast 2025, `Person` schema at 58.9% PREVALENCE among cited sources — wrong type, and its FAQPage figure (1.8%) points the opposite way to the claim it propped up. This one drove Tier 1 ranking. - "62% of searches involve voice" — uncited; 62% circulates as smart-speaker ADOPTION. Same family as the "50% by 2020" myth ComScore denied (origin: a 2014 Andrew Ng interview). Corrected my own framing too: I claimed three times these stats "drive axis weights". They do not — the weight tables carry no citations. They drive Tier/priority recommendations and, worse, geo-analyzer's "Cite sources" rule pushed them into CLIENT reports as research-backed. Fixes: recommendations kept on mechanism, fabricated numbers removed with the incident documented inline so they are not re-added. Unverified stats (48% AI Overviews, 2.5B queries/day, Gartner -25%) labelled [UNVERIFIED] rather than asserted or deleted — I did not check them. Structural, not just exhortation: resources/README.md now mandates `<claim> — <source, year, venue|vendor> — measured: <what the source ACTUALLY measured> — <link>`. `measured:` is the field that catches this — all four errors survive a source name; none survives stating the real measurement next to the claim. WebSearch demoted from verification to crawler/tool-name lookup only. Verified: make test 35 GREEN / 0 RED.
3.5 KiB
SEO/GEO shared resources
Knowledge base shared by seo-analyzer and geo-analyzer agents.
Loaded on demand — keep each file focused and current.
| File | Owner agents | Topic |
|---|---|---|
ai-crawlers-2026.md |
seo + geo | User-agent strings, categories (training vs search), robots.txt strategy |
llms-txt-template.md |
geo | /llms.txt + /llms-full.txt structure, generation patterns |
geo-schemas.md |
geo | Schema.org types for AI extraction (QAPage, Speakable, Person, Article) + deprecated list |
entity-seo.md |
geo | Wikidata QID, sameAs network, Knowledge Graph wiring |
content-shape-for-ai.md |
geo | Definition Lead, TL;DR, Q→A, stats, citations — content patterns LLMs cite |
ai-visibility-tools.md |
geo | Monitoring tools (OtterlyAI, Peec, Trendos, ZipTie, HubSpot AEO, SE Ranking) |
automation-catalog.md |
seo + geo | For every user-action in SEO.md §11 — what tool can automate it |
Update policy
These files capture state as of 2026-04. Crawler lists, Schema.org deprecations, and tool landscape shift fast. Agents MUST cross-check crawler lists and tool names via WebSearch on each run when FULL depth is selected.
Citation standard (mandatory for every statistic)
WebSearch is NOT verification for a number. It ranks SEO blogs, and SEO blogs cross-cite each other into a consensus that looks like corroboration. Two 2026-07-16 audits of this directory show how it fails:
- A "VSI (Visual Stability Index) — new 2026 Core Web Vital" lived in
seo-analyzer.md. Ten blogs asserted it; several claimed CrUX already collected it. It is absent from the CrUX API metric list and from web.dev. WebSearch returned the echo, not the truth. - Every stat in this directory was real and attached to the wrong subject: the GEO paper's 40% (all methods) pinned on one technique; LLMrefs' 3x (brand mentions vs backlinks) pinned on freshness decay; AccuraCast's 58.9% (Person schema prevalence) pinned on QAPage lift, with its meaning inverted; a smart-speaker adoption figure sold as voice-search share.
The failure mode is not invention — it is plausible recombination, which is exactly what a model half-remembering a search result produces. So the format has to make an unsourced number conspicuous:
<claim> — <source, year, venue|vendor> — measured: <what the source ACTUALLY
measured> — <link>
measured: is the field that catches it. All four errors above survive a
source name; none survives having to state the source's real measurement
next to the claim.
Rules:
- Primary source or no number. Peer-reviewed paper, the vendor's own
published study, or an official API/doc.
developer.chrome.com/docs/cruxis decisive for metrics: what CrUX cannot return, we cannot score. - Name the tier. Peer review ≠ vendor marketing. LLMrefs, AccuraCast, Ahrefs publish useful data and sell products — say "vendor".
- Never widen scope. An aggregate result is not a per-technique result.
- No number beats a wrong number. A recommendation that only stands up with a fabricated statistic was never standing up. Delete the stat, keep the recommendation if it survives on mechanism.
- Unverified ⇒ labelled.
[UNVERIFIED — <date>]inline. Never quote an unverified number to a client:geo-analyzer.md("Cite sources") sends these into client reports as research-backed.
Loading pattern
Agents reference resources like this:
Load: ~/.claude/agents/resources/ai-crawlers-2026.md
Do not inline these contents into agent prompts — read them at step time.