feat(seo/geo): split into parallel seo + geo agents with shared resources

Refactor the monolithic seo-analyzer into two specialist agents
orchestrated in parallel by the /seo skill, plus a standalone /geo
skill for AI-only audits.

Changes
- agents/seo-analyzer.md: refocused on classical engines (Google, Bing,
  DuckDuckGo). Adds Core Web Vitals 2.0 (LCP/INP/CLS + VSI), CSP + full
  security headers, hreflang audit, video SEO (transcripts), accessibility
  as ranking signal, image/video sitemaps.
- agents/geo-analyzer.md: new agent for AI engines (ChatGPT, Claude,
  Perplexity, Gemini, Google AI Overviews, Copilot). Covers AI crawler
  policy, llms.txt/llms-full.txt, Schema.org for AI extraction (QAPage,
  Speakable, Person+Article, Organization graph), entity SEO (Wikidata,
  sameAs, Knowledge Panel), content shape (Definition Lead, TL;DR,
  Q->A, citable stats, freshness), AI visibility testing.
- agents/resources/: shared knowledge base referenced by both agents —
  ai-crawlers-2026.md (25+ bots, training vs retrieval categories,
  permissive/restrictive templates), llms-txt-template.md, geo-schemas.md
  (incl. deprecated list: ClaimReview, CourseInfo, etc. removed June 2025),
  entity-seo.md, content-shape-for-ai.md, ai-visibility-tools.md,
  automation-catalog.md.
- skills/seo/SKILL.md: becomes parallel dispatcher. Collects context
  once (depth + business), spawns both agents in a single message for
  concurrent execution, merges envelopes into unified SEO.md. Includes
  authoritative file-ownership matrix to prevent parallel-edit races.
- skills/geo/SKILL.md: new standalone wrapper for GEO-only audits.

Scoring
- Combined score: GLOBAL = 0.80 * SEO + 0.20 * GEO (local B2C),
  0.75 * SEO + 0.25 * GEO (SaaS/national/content).
- GEO axis weight raised from 5% (old) to first-class dimension.

Policy
- AI crawlers: permissive default (maximise AI citations). Restrictive
  template available for premium/regulated content.
- Every user action in SEO.md section 11 must cite automation options
  from automation-catalog.md.

Tools
- WebFetch + WebSearch added to allowed-tools of both skills and
  both agents (needed for live CWV via PageSpeed API, AI visibility
  testing, Wikidata/Knowledge Panel lookups, competitor analysis).

Research basis (2026 state of the art validated via WebSearch):
- Core Web Vitals 2.0 (VSI signal, Google core update March 2026)
- AI Overviews trigger on ~48% of Google searches
- ClaimReview + 6 other schema types deprecated June 2025
- Definition Lead Architecture (CMU KDD 2024, +impression score)
- Citations + stats add up to 40% AI visibility (Aggarwal 2024)
- Wikidata grounds every major LLM (ChatGPT, Claude, Gemini, Perplexity)

Backup
- agents/seo-analyzer.md.bak kept for rollback reference.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
bastien
2026-04-21 16:16:30 +02:00
co-authored by Claude Opus 4.7
parent 53d04db480
commit 95347d2e47
13 changed files with 4003 additions and 499 deletions
+30
View File
@@ -0,0 +1,30 @@
# SEO/GEO shared resources
Knowledge base shared by `seo-analyzer` and `geo-analyzer` agents.
Loaded on demand — keep each file focused and current.
| File | Owner agents | Topic |
|---|---|---|
| `ai-crawlers-2026.md` | seo + geo | User-agent strings, categories (training vs search), robots.txt strategy |
| `llms-txt-template.md` | geo | `/llms.txt` + `/llms-full.txt` structure, generation patterns |
| `geo-schemas.md` | geo | Schema.org types for AI extraction (QAPage, Speakable, Person, Article) + deprecated list |
| `entity-seo.md` | geo | Wikidata QID, sameAs network, Knowledge Graph wiring |
| `content-shape-for-ai.md` | geo | Definition Lead, TL;DR, Q→A, stats, citations — content patterns LLMs cite |
| `ai-visibility-tools.md` | geo | Monitoring tools (OtterlyAI, Peec, Trendos, ZipTie, HubSpot AEO, SE Ranking) |
| `automation-catalog.md` | seo + geo | For every user-action in SEO.md §11 — what tool can automate it |
## Update policy
These files capture state as of 2026-04. Crawler lists, Schema.org
deprecations, and tool landscape shift fast. Agents MUST cross-check
via WebSearch on each run when FULL depth is selected.
## Loading pattern
Agents reference resources like this:
```
Load: ~/.claude/agents/resources/ai-crawlers-2026.md
```
Do not inline these contents into agent prompts — read them at step time.
+209
View File
@@ -0,0 +1,209 @@
# AI crawlers — 2026 reference
State as of 2026-04. Cross-check via WebSearch on FULL audits — new
bots and renames ship monthly.
## The two categories that matter
The blanket "block AI" strategy of 2024 is obsolete. Bots now split
into two roles, and treating them the same loses traffic.
### Training bots — scrape content to train future models
No direct user traffic. No citation back. Content vanishes into weights.
| User-agent | Company | Notes |
|---|---|---|
| `GPTBot` | OpenAI | Training for GPT models |
| `Google-Extended` | Google | Opt-out for Gemini training |
| `CCBot` | Common Crawl | Feeds many LLMs (open dataset) |
| `anthropic-ai` | Anthropic | Legacy training bot (being phased out) |
| `ClaudeBot` | Anthropic | Current training bot |
| `Bytespider` | ByteDance / TikTok | Aggressive scraper, frequent complaints |
| `Meta-ExternalAgent` | Meta | Training for Llama family |
| `Meta-ExternalFetcher` | Meta | Per-request fetch |
| `Applebot-Extended` | Apple | Opt-out for Apple Intelligence training |
| `Amazonbot` | Amazon | Alexa + internal LLMs |
| `cohere-ai` | Cohere | Training |
| `Diffbot` | Diffbot | Knowledge Graph construction |
| `omgilibot` | Webz.io | Data resale |
| `img2dataset` | Various | Image dataset builders |
| `Timpibot` | Timpi | Search-index + training hybrid |
### Search / retrieval bots — fetch content to cite in live answers
User asked a question → bot fetches → cites your URL → traffic returns.
| User-agent | Company | Notes |
|---|---|---|
| `OAI-SearchBot` | OpenAI | Powers ChatGPT Search |
| `ChatGPT-User` | OpenAI | On-demand fetch when user asks ChatGPT about a URL |
| `Claude-SearchBot` | Anthropic | Powers Claude web search |
| `Claude-User` | Anthropic | On-demand fetch inside Claude |
| `Claude-Web` | Anthropic | Legacy retrieval bot |
| `PerplexityBot` | Perplexity | Index builder |
| `Perplexity-User` | Perplexity | On-demand fetch |
| `GoogleOther` | Google | Various Google retrieval use cases |
| `FacebookBot` | Meta | Meta AI search |
| `DuckAssistBot` | DuckDuckGo | DuckAssist answers |
| `YouBot` | You.com | You.com retrieval |
| `MistralAI-User` | Mistral | On-demand fetch |
## Recommended default strategy — PERMISSIVE
Rationale: the user's stated goal is to maximise AI visibility. The
future-of-search brief favours being cited over being protected.
```
# robots.txt — PERMISSIVE default (allow everything, block problem bots)
# --- Training bots: allow (contributes to brand visibility long-term) ---
User-agent: GPTBot
Allow: /
User-agent: Google-Extended
Allow: /
User-agent: ClaudeBot
Allow: /
User-agent: Applebot-Extended
Allow: /
User-agent: Meta-ExternalAgent
Allow: /
User-agent: CCBot
Allow: /
# --- Search / retrieval bots: always allow (direct traffic) ---
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: Claude-User
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: Perplexity-User
Allow: /
# --- Block only known-abusive bots (aggressive scraping, no return value) ---
User-agent: Bytespider
Disallow: /
User-agent: omgilibot
Disallow: /
User-agent: img2dataset
Disallow: /
# --- Default: allow the rest ---
User-agent: *
Allow: /
Sitemap: https://example.com/sitemap.xml
```
## Alternative — RESTRICTIVE (for premium content, paywalled, regulated)
```
# robots.txt — RESTRICTIVE (block training, allow retrieval)
# Block all training bots
User-agent: GPTBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: anthropic-ai
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: Bytespider
Disallow: /
User-agent: Meta-ExternalAgent
Disallow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: Amazonbot
Disallow: /
User-agent: cohere-ai
Disallow: /
User-agent: Diffbot
Disallow: /
User-agent: Timpibot
Disallow: /
# Allow search/retrieval (keeps citations flowing)
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: Claude-User
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: Perplexity-User
Allow: /
User-agent: *
Allow: /
Sitemap: https://example.com/sitemap.xml
```
## Common mistakes
- **Only blocking `ClaudeBot`** — does not block `Claude-SearchBot` or `Claude-User`. Same for other families.
- **Using `GPTBot` to block ChatGPT Search** — wrong. `OAI-SearchBot` and `ChatGPT-User` are the search bots.
- **Blocking `CCBot`** — has knock-on effects across dozens of downstream LLMs that train on Common Crawl.
- **Using wildcards** (e.g. `User-agent: *AI*`) — robots.txt wildcards are not universally supported.
- **Relying on meta robots** — `<meta name="robots">` is less respected than robots.txt by AI crawlers. Use both.
## Verification
Each bot should return 200 for allowed, 403 for blocked, via simulated requests:
```bash
DOMAIN="example.com"
for UA in "GPTBot" "ClaudeBot" "PerplexityBot" "OAI-SearchBot" "ChatGPT-User" "Google-Extended"; do
CODE=$(curl -sI -A "$UA" -o /dev/null -w "%{http_code}" "https://$DOMAIN/")
echo "$UA: $CODE"
done
```
This hits the page, not robots.txt directly — but if the origin respects
robots.txt via CDN/WAF rules, you'll see the difference.
## Sources to refresh this doc
- https://platform.openai.com/docs/bots
- https://darkvisitors.com/agents (community-maintained)
- https://github.com/ai-robots-txt/ai.robots.txt
- Anthropic docs: https://docs.anthropic.com/
- Cloudflare AI crawlers dashboard (if account available)
+99
View File
@@ -0,0 +1,99 @@
# AI visibility monitoring tools — 2026
Tools that track whether your brand appears in AI-generated answers
across ChatGPT, Perplexity, Gemini, Copilot, Claude, and Google AI
Overviews.
Context: Google AI Overviews trigger on ~48% of searches; ChatGPT
processes 2.5B queries/day; Gartner projects commercial organic
search traffic will drop 25% by 2026. Monitoring is no longer optional.
## Commercial tools
| Tool | Platforms covered | Strong points | Weak points |
|---|---|---|---|
| **OtterlyAI** (otterly.ai) | ChatGPT, Perplexity, Gemini, AI Overviews, Copilot | Mature, 20k+ users, Gartner-recognised | Pricing mid-to-high |
| **Peec AI** (peec.ai) | ChatGPT, Perplexity, Gemini, AI Overviews | Good SaaS-brand focus, sentiment analysis | Narrower platform scope |
| **Profound** (tryprofound.com) | ChatGPT, Perplexity, Gemini, Copilot | Enterprise-grade, full-response capture | Enterprise pricing |
| **ZipTie** (ziptie.dev) | ChatGPT, Perplexity, AI Overviews | Competitive benchmarking, source attribution | Smaller team, newer |
| **HubSpot AEO** (hubspot.com/products/aeo) | ChatGPT, Gemini, Perplexity | Integrates with HubSpot ecosystem | Best if already HubSpot user |
| **Trendos** (trendos by Tesonet) | ChatGPT, Gemini, AI Search, Perplexity, DeepSeek | Added DeepSeek coverage, 2026 launch | Unproven longevity |
| **SE Ranking AI Tracker** (seranking.com) | ChatGPT, Perplexity, Gemini, AI Mode, AI Overviews | Bundled with classical SEO suite | Less specialised |
| **LLMrefs** (llmrefs.com) | ChatGPT, Perplexity, Gemini, Claude | GEO focus, research-backed | Newer, less tested |
## Free / manual methods (zero budget)
For clients/projects with no monitoring budget, a manual process works
at lower frequency. Recommended cadence: monthly for established
brands, weekly during optimization sprints.
### Query list construction
Build a list of 20-40 queries covering:
1. **Branded queries** — "what is [brand]", "is [brand] good", "[brand] reviews"
2. **Generic category queries** — "best [category] in [location]", "how to [problem]"
3. **Comparison queries** — "[brand] vs [competitor]", "alternatives to [brand]"
4. **Problem queries** — the actual questions the target persona asks
### Manual check workflow
For each query, run across:
- **ChatGPT** (web version with search enabled, chatgpt.com)
- **Perplexity** (perplexity.ai)
- **Google AI Overviews** (google.com — appears for ~48% of searches)
- **Claude** (claude.ai with web search)
- **Gemini** (gemini.google.com)
- **Copilot** (copilot.microsoft.com)
- **Brave Search AI** (search.brave.com)
- **DuckAssist** (duckduckgo.com)
Record for each:
- Mentioned? (yes/no)
- Cited with link? (yes/no + which page)
- Position in answer? (1st mention / buried / listed)
- Sentiment? (positive / neutral / negative / misleading)
### Spreadsheet template
| Date | Query | ChatGPT | Perplexity | Google AIO | Claude | Gemini | Copilot |
|---|---|---|---|---|---|---|---|
| 2026-04-21 | best plombier Évry | Mentioned, ranked 3, cited | Not mentioned | Top 3, no cite | — | — | — |
## KPIs to track
From GEO research and industry consensus (GenOptima, HubSpot 2026):
| Metric | Definition | Benchmark |
|---|---|---|
| **Mention Rate** | % of AI answers that mention brand name | Varies; track trend, not absolute |
| **Citation Rate** | % of AI answers with a clickable link to domain | Target 20%+ for established brands |
| **Position** | When cited, is brand 1st mention vs buried? | First mention = best |
| **Sentiment** | Tone of brand mention (positive/neutral/negative) | Track for negative drift |
| **Source Diversity** | Which of your pages get cited? | Aim for 5+ distinct pages/domain |
| **Competitor Share** | % of category queries where competitor cited vs brand | Track gap |
## Integration into SEO.md
In `SEO.md §11 — Actions utilisateur requises`:
> ### Monitor AI visibility monthly
>
> **Automatisation possible avec:** OtterlyAI, Peec AI, ZipTie, HubSpot
> AEO, SE Ranking AI Tracker. Budget: 50-500 EUR/mois selon le tool.
>
> **Alternative manuelle gratuite:** template spreadsheet + 20 queries
> testées mensuellement sur ChatGPT, Perplexity, Google AI Overviews.
> Temps: ~1h/mois.
## Methodology caveats
- AI engines are **non-deterministic**. Same query twice can return
different answers. Always take 3 samples and track the median.
- **Personalisation** affects results. Test in logged-out / private
mode for reproducibility.
- **Geographic bias** — ChatGPT's answers about local businesses vary
by IP. Test from the target market's geography.
- **Freshness lag** — content updates take days to weeks to propagate
into AI answers. Don't expect instant reflection of changes.
+215
View File
@@ -0,0 +1,215 @@
# Automation catalog — for SEO.md §11 user actions
For every action that requires the human, this catalog lists tools
that can partially or fully automate it. Both agents cite this file
when emitting user actions into `SEO.md §11`.
**Format rule in SEO.md §11**: every entry MUST include:
```
- **<Action>** — <what to do>
**Automatisation possible avec:** <tool 1>, <tool 2>, <tool 3>
**Budget:** <free / XX EUR/mois / one-time XX EUR>
**Effort manuel:** <time estimate>
```
## Local SEO actions
### Google Business Profile — claim / create / optimize
- **Google Business Profile API** (free, requires Google Cloud project)
→ post updates, reply to reviews, sync hours automatically
- **Yext** (enterprise, 500-5000 EUR/mois) → syncs GMB across
200+ directories
- **BrightLocal** (30-80 USD/mois) → GMB management + rank tracking
- **Moz Local** (14-33 USD/mois/location) → listing management
- **Uberall** (enterprise) → multi-location listing sync
- **LocalFalcon** (30-60 USD/mois) → GMB rank visualisation
- **PlePer** (~25 EUR/mois) → GMB post scheduling
- Manual workflow: 30 min/week via https://business.google.com
### Review management — collect, reply, aggregate
- **Trustpilot / Google Reviews API** (via GBP API) → read/reply programmatically
- **Birdeye** (290+ USD/mois) → review aggregation + auto-reply
- **Podium** (enterprise) → SMS-based review requests
- **NiceJob** (90 USD/mois) → review request automation
- **Grade.us** (110 USD/mois) → multi-platform aggregation
- Manual: monitor GMB + reply within 48h (legal: L121-1 Code conso FR)
### Directory citations — PagesJaunes, Yelp, Mappy, Bing Places, Apple
- **Yext** → 200+ directories incl. French
- **BrightLocal / Moz Local** → coverage varies, check French support
- **Uberall** → strong in European markets
- **Rio SEO** (enterprise) → big brands
- Manual: one-time 4-8h to register on top 10 directories
## AI visibility actions
### Monitor brand in AI engines
See `ai-visibility-tools.md`. Summary:
- **OtterlyAI, Peec AI, Trendos, ZipTie, HubSpot AEO, SE Ranking** — commercial
- **Manual spreadsheet + 20 queries/mois** — free, ~1h/mois
### Submit to AI indexes directly
- **Bing Webmaster Tools** → submits to Bing + Copilot + ChatGPT Search (which uses Bing index)
- **IndexNow protocol** (indexnow.org) → proactive ping to Bing/Yandex
- **Google Search Console + URL Inspection** → request indexing (no ChatGPT index direct submit exists in 2026)
### Maintain llms.txt / llms-full.txt
- **llms-txt-action** (GitHub Action) → rebuild on deploy
- **Mintlify / Fern / ReadMe** → auto-generated for supported docs hosts
- **Custom cron + script** → pull from CMS, regenerate weekly
- Manual: monthly review if content changes rarely
## Entity / Knowledge Graph actions
### Create or optimize Wikidata entry
- **Kalicube** (commercial, custom pricing) → specialised Knowledge Panel + Wikidata
- **InLinks** (40-350 USD/mois) → entity optimization + Schema.org graph
- **WordLift** (30-300 USD/mois) → WordPress plugin with Wikidata linking
- **Entity.ai** → entity signal auditing
- Manual: https://www.wikidata.org, 2-4h initial + sources required
### Claim / optimize Google Knowledge Panel
- **Kalicube** — best specialisation
- **Manual via Google Search** — click "Claim this Knowledge Panel" (requires verification)
- Cannot be forced; appears when entity signals are strong enough
### LinkedIn / Crunchbase / industry directory entities
- **Yext** → includes Crunchbase sync in enterprise tiers
- **Manual** → LinkedIn Company page free, Crunchbase free profile claim
- **Brandify** (enterprise) → multi-directory entity management
## Content production actions
### Create city/service landing pages (30/70 rule)
- **Surfer SEO** (89-219 USD/mois) → content optimization with AI
- **Frase.io** (45-115 USD/mois) → SERP-driven briefs
- **Clearscope** (170+ USD/mois) → keyword + semantic briefs
- **Manual + AI writer** → use Claude/ChatGPT with explicit 30/70 instruction
Agent note: Batch D in `seo-analyzer` triage can handle the
CREATION of these pages if confirmation granted — city pages are
typically batch D (structural change, user approval needed).
### Produce blog content on schedule
- **Frase / Surfer / Clearscope** (see above)
- **MarketMuse** (enterprise) → content planning
- **Jasper AI / Copy.ai** → AI drafting (quality review mandatory)
- Human editor remains the bottleneck — AI drafts need domain expert review
### Refresh existing content quarterly
- **ContentKing** (now part of Conductor) → change detection
- **SEOClarity** (enterprise) → content decay tracking
- **Manual** — spreadsheet of top 50 pages + quarterly review cycle
## Technical SEO actions
### Generate sitemaps
- **Framework plugin** — `@astrojs/sitemap`, `next-sitemap`, `@nuxtjs/sitemap`, `rails-sitemap-generator`, etc.
- **Yoast / RankMath** (WordPress) → auto-generate
- **Screaming Frog** (200 GBP/an) → crawler-based generation
- Manual: only as last resort, hand-maintained sitemaps go stale fast
### Implement Schema.org at scale
- **Yoast / RankMath / SEOPress** (WordPress) → Article/Organization/LocalBusiness auto-graph
- **Schema App** (enterprise) → multi-CMS
- **Merkle Schema Markup Generator** (free) → one-off generation
- **Manual + `geo-schemas.md` templates** — for frameworks without plugins
### Optimize Core Web Vitals
- **PageSpeed Insights API** (free) → measure + monitor
- **WebPageTest** (free tier + paid) → detailed waterfalls
- **Cloudflare Speed** (free tier with Cloudflare) → CDN-level optimizations
- **Nitropack** (35-175 USD/mois) → WordPress speed automation
- **Vercel Speed Insights** (free for Vercel projects)
- Manual: Lighthouse + manual fixes guided by its recommendations
### Security headers (CSP, HSTS, X-Frame-Options, Referrer-Policy)
- **securityheaders.com** (free audit)
- **Cloudflare Page Rules** → header injection
- **Vercel `next.config.js` headers** → declarative
- **`.htaccess`** → Apache hosts
- Manual: one-time config, ~1-2h setup
## Social presence actions
### Create / maintain social profiles (Facebook, Instagram, LinkedIn, TikTok, YouTube)
- **Buffer** (6-120 USD/mois) → multi-platform scheduling
- **Hootsuite** (99-249 USD/mois) → full social suite
- **Later** (16-80 USD/mois) → visual content scheduling
- **Metricool** (18-50 USD/mois) → analytics + scheduling
- Manual: 30-60 min/semaine for basic maintenance
### Monitor brand mentions on social / forums / Reddit
- **Brand24** (99-299 USD/mois)
- **Mention** (41-149 USD/mois)
- **Google Alerts** (free, basic)
- **Reddit search + saved queries** — free, manual
- **BuzzSumo** (199+ USD/mois) → trend + mention discovery
## Legal compliance actions (FR)
### Install cookie consent management (CMP)
- **Axeptia / Axeptio** (free to 100 EUR/mois) → French-focused CMP
- **Cookiebot** (11-96 USD/mois) → international CMP, CNIL-compliant
- **OneTrust** (enterprise) → enterprise compliance
- **tarteaucitron.js** (free, open source) → CNIL-compliant, self-hosted
- **Didomi** (enterprise) → strong French legal context
### Generate legal pages (mentions légales, politique de confidentialité, CGV)
- **Legalstart / Captain Contrat** (one-time 50-200 EUR) → FR templates
- **Genius Legal** → template generators
- **Legalbuddy** → questionnaire-driven legal pages
- Agent fallback: Batch B in `seo-analyzer` creates templates with
`[À COMPLÉTER]` placeholders for SIREN, capital, etc.
## Reporting format in SEO.md §11
Example entry generated by the agents:
```markdown
### Créer / réclamer la fiche Google Business Profile
**Action:** Vérifier que la fiche GMB existe, est réclamée, et les
informations sont cohérentes avec le site (NAP).
**Lien direct:** https://business.google.com
**Automatisation possible avec:**
- Google Business Profile API (gratuit, technique)
- BrightLocal (30-80 USD/mois, gestion + rank tracking)
- Yext (500+ EUR/mois, multi-directories)
- LocalFalcon (30-60 USD/mois, rank visualisation)
**Effort manuel:** 30 min initial + 30 min/semaine maintenance
**Impact SEO local:** critique (base du SEO local)
```
## Maintenance of this catalog
Tool landscape shifts fast. Cross-check quarterly:
- Have tool URLs changed?
- Has pricing moved tier?
- Have new tools emerged (especially in AI visibility monitoring)?
- Are deprecated tools still listed?
Use WebSearch on FULL audits to validate before emitting in SEO.md.
+250
View File
@@ -0,0 +1,250 @@
# Content shape for LLM extraction
How to write pages so AI engines quote, cite, and recommend them.
Based on peer-reviewed GEO research (CMU KDD 2024, Aggarwal et al.)
and tracked citation patterns across ChatGPT, Perplexity, Claude,
Gemini, Google AI Overviews (2025-2026).
## The six patterns that measurably increase AI citations
### 1. Definition Lead Architecture
Open the page (or first paragraph after each major heading) with:
> **[Entity] is a [category] that [differentiator].**
Research backing: CMU GEO framework (KDD 2024) — pages with explicit
definitional openings score significantly higher in LLM retrieval
impression scores.
**Good**: "Astro is a static site generator that ships zero JavaScript by default, producing HTML at build time that search engines and AI crawlers can index without running a browser."
**Bad**: "In today's fast-paced digital landscape, choosing the right framework can feel overwhelming. At Acme, we know how important it is to..."
### 2. TL;DR / Answer Box above the fold
Insert an explicit summary block at the top of long content. AI engines
preferentially quote from these blocks because the content is
pre-summarised.
```html
<aside class="tldr">
<strong>TL;DR</strong> —
Next.js 15 removes the pages/ directory entirely in favour of App
Router. Migration requires rewriting route handlers, layouts, and
data fetching. Estimated effort: 2-5 days for a medium project.
</aside>
```
CSS: no class requirement, but mark it semantically (e.g. `aria-label="summary"`
or Speakable schema targeting this selector).
### 3. Question-then-direct-answer structure
Each H2/H3 heading phrased as a likely user query. First sentence
after the heading: a single-sentence direct answer. Supporting detail
follows.
**Pattern**:
```
## How much does a Qualibat RGE certification cost in France?
A Qualibat RGE certification costs between 500 and 1500 EUR for the
initial audit, plus an annual fee of 200-400 EUR. The cost varies by
trade category and company size.
[Detailed breakdown follows...]
```
Why it works: LLMs grade passages by answer-density relative to the
query. A one-sentence self-contained answer has the highest density.
### 4. Citations and statistics (strongest measured lever)
Adding peer-cited statistics with clear sources increases AI visibility
**by up to 40%** (Aggarwal et al., 2024 "GEO: Generative Engine
Optimization").
Pattern: embed specific numbers with attribution.
**Good**: "According to the ADEME 2024 energy report, French households spent an average of 2,137 EUR on heating in 2023 — a 12% increase from 2021."
**Bad**: "Heating costs have increased a lot recently."
Source attribution matters: link the citation to the original source
(`<a href>`), ideally with `rel="cite"`. AI engines use link graphs
to validate factual claims.
### 5. Structured lists and comparison tables
LLMs quote list items and table rows more readily than prose of the
same content. Convert what you can:
**Before** (prose):
"The best frameworks for public sites are Astro for static content,
Next.js for dynamic server-rendered apps, and Nuxt for Vue-based
projects."
**After** (list):
"Best frameworks for public sites by use case:
- **Astro** — static content (blog, docs, portfolio)
- **Next.js** — dynamic SSR with React
- **Nuxt** — dynamic SSR with Vue"
Comparison tables are even stronger. Structure:
| Framework | Rendering | Best for | JS by default |
|---|---|---|---|
| Astro | SSG + islands | Public content | 0 KB |
| Next.js | SSG + SSR | Hybrid apps | Large |
### 6. Freshness signals
Pages not updated at least quarterly are **3x more likely to lose AI
citations** (LLMRefs 2026 study).
What to maintain:
- Visible "Last updated: YYYY-MM-DD" at the top of content pages
- `dateModified` in Article/BlogPosting JSON-LD (ISO 8601)
- HTTP header `Last-Modified` in sync with content change
- Changelog on evergreen reference pages
Do NOT fake dates — AI engines and Google increasingly validate
freshness against actual content diffs.
## Anti-patterns — what to avoid
### Pronoun-heavy writing
LLMs resolve pronouns by context window, which costs them confidence.
Prefer explicit entity names.
**Bad**: "It was founded in 2015. Its founders wanted to solve a problem. They saw that..."
**Good**: "Acme Corp was founded in 2015. Acme's founders, Jane Doe and John Smith, wanted to solve..."
### Marketing fluff before facts
AI engines typically truncate retrieval windows. Fluff at the top
wastes the budget. Put factual claims FIRST.
**Bad** (first 200 chars wasted): "In today's fast-moving digital landscape, businesses are constantly looking for ways to stay competitive..."
**Good** (first 200 chars dense): "Our API processes 50M requests/day at p99 latency of 47ms across 8 regions, with a 99.99% SLA. Pricing starts at 99 EUR/month for the 10K requests tier."
### Claims without sources
Any numerical or comparative claim without a linked source degrades
trust. AI engines can detect the pattern "number without citation" and
weight those passages lower.
### Cookie-cutter content across pages (especially city pages)
The 30/70 rule: when creating per-city or per-service variants,
at most 30% of the content should be templated. 70% must be
unique per page (local landmarks, specific testimonials, unique
stats, real photos).
Generic city pages get filtered out as "doorway pages" by both
classical search and AI engines.
## Page templates by type
### Service page (local business)
```
<h1>[Service] in [City] — [Business Name]</h1>
<div class="tldr">
<strong>En résumé :</strong> [Business] offers [service] in [city + surrounding].
[Key differentiator — price, response time, certifications]. Open [hours].
Call [phone] or request a quote online.
</div>
<h2>What is [service]?</h2>
<p>[Service] is a [category] that [differentiator]. In [city], demand
is driven by [local factor — housing stock, climate, regulations].</p>
<h2>How much does [service] cost in [city]?</h2>
<p>[Specific price range] for a typical [job type], based on [n]
projects completed in [year]. Factors affecting cost: [list].</p>
<h2>Why choose [Business] for [service]?</h2>
<ul>
<li>[Certification 1] — [what it means]</li>
<li>[Certification 2]</li>
<li>[N+ years] experience on [specific housing stock]</li>
</ul>
<h2>FAQ</h2>
[QAPage or FAQPage schema + visible Q&A]
```
### Blog post / guide
```
<h1>[Clear, question-style or noun-phrase headline]</h1>
<p class="byline">By [Author Name] — Updated [Date]</p>
<div class="tldr">
[3-5 sentence summary. Include the key number, the key conclusion,
and any nuance.]
</div>
<h2>[Question 1]</h2>
<p>[One-sentence answer.] [Supporting detail with cited statistics.]</p>
<h2>[Question 2]</h2>
...
<h2>Sources</h2>
<ul>
<li><a href="...">Source 1 — author, year</a></li>
<li><a href="...">Source 2 — author, year</a></li>
</ul>
```
### Homepage / landing
```
<h1>[Entity] is a [category] that [differentiator].</h1>
<!-- The H1 IS the Definition Lead. Yes, really. -->
<p class="hero-subtitle">
[Elaboration on the H1. Include one concrete stat or proof point.]
</p>
[Primary CTA]
<section>
<h2>What [Entity] does</h2>
<p>[Functional description, one paragraph.]</p>
</section>
<section>
<h2>Who uses [Entity]</h2>
<ul><li>[Use case 1]</li><li>[Use case 2]</li>...</ul>
</section>
<section>
<h2>How it works</h2>
<!-- HowTo schema + visible steps -->
</section>
<section>
<h2>Frequently asked</h2>
<!-- FAQPage schema + visible Q&A -->
</section>
```
## Self-audit — is this page AI-friendly?
- [ ] First sentence: `[Entity] is a [category] that [differentiator]` ?
- [ ] TL;DR or summary block above the fold ?
- [ ] Every H2/H3 phrased as a likely user question ?
- [ ] First sentence under each heading: direct answer ?
- [ ] At least 2-3 specific numerical claims with linked sources ?
- [ ] Visible "Last updated" date + matching `dateModified` in JSON-LD ?
- [ ] Lists or tables instead of dense prose where possible ?
- [ ] Entity names used explicitly, not pronouns ?
- [ ] If it's a city/service variant: ≥70% unique content ?
+163
View File
@@ -0,0 +1,163 @@
# Entity SEO — Wikidata, Knowledge Graph, sameAs
Why this matters: every major AI engine (ChatGPT, Claude, Gemini,
Perplexity, Apple Intelligence) grounds factual claims against
Wikidata. A business without a clean entity footprint is effectively
invisible to AI grounding pipelines, regardless of on-site SEO.
## The entity identity stack
Think of your entity as having five layers, from strongest to weakest
identity signal:
1. **Wikidata QID** — globally unique, machine-readable identifier.
2. **Wikipedia article** — human-readable notability signal.
3. **Google Knowledge Panel** — surfaced directly in Google results.
4. **Authoritative third-party IDs** — Crunchbase, Bloomberg, SIRENE (FR), Companies House (UK), OpenCorporates.
5. **Social + directory profiles** — LinkedIn, Facebook, PagesJaunes, industry directories.
Each layer reinforces the ones below. Wikidata is the most leveraged
because it's structured, open, and explicitly consumed by LLMs.
## Audit checklist
### Does the entity have a Wikidata QID?
Search: https://www.wikidata.org/wiki/Special:Search — by name + city.
If found:
- Record QID (format `Q` + number, e.g. `Q12345678`)
- Verify: official website property (P856) points to the current domain
- Verify: VAT (P3608), SIRET (P3893), category (P31) are correct
If NOT found:
- For businesses meeting Wikidata notability: creation is possible
(requires verifiable third-party sources)
- For non-notable businesses: skip Wikidata, focus on other identity layers
- Flag in SEO.md §11 as user action (Wikidata requires human judgement
+ source citations)
### Does the entity have a Wikipedia article?
- Search by exact business name. If found and matches: record URL.
- If not found: flag as long-term goal (long-term — notability bar is high).
### Is there a Google Knowledge Panel?
Search Google: exact business name. Look for the right-side panel.
- Present + claimed → verify info is correct
- Present + unclaimed → user action: claim via https://www.google.com/business/
- Absent → Knowledge Panels are generated automatically when entity
signals are strong enough (GMB + Wikidata + consistent citations)
### Is `sameAs` complete in on-site JSON-LD?
The `sameAs` property is how you declare "these external URLs represent
the same entity as this page". It's the single most impactful entity
signal after Wikidata.
Minimum recommended `sameAs` for a local business:
```json
"sameAs": [
"https://www.wikidata.org/wiki/Q123456789", // if exists
"https://www.linkedin.com/company/name",
"https://www.facebook.com/businessname",
"https://www.instagram.com/businessname",
"https://www.pagesjaunes.fr/pros/12345", // FR
"https://fr.wikipedia.org/wiki/Nom_Entreprise" // if exists
]
```
For a SaaS / international brand, add:
```json
"https://www.crunchbase.com/organization/name",
"https://github.com/organization",
"https://www.g2.com/products/name",
"https://www.producthunt.com/products/name"
```
For a Person (author, founder):
```json
"sameAs": [
"https://www.wikidata.org/wiki/Q987654321",
"https://www.linkedin.com/in/name",
"https://twitter.com/name",
"https://github.com/name",
"https://scholar.google.com/citations?user=XYZ", // academics
"https://orcid.org/0000-0000-0000-0000" // academics
]
```
### Is `@id` used consistently?
Across all JSON-LD blocks on the site, the same entity MUST use the
same `@id`. Pattern: `https://example.com/#org` for the organization,
`https://example.com/about#author-{slug}` for people.
Split across multiple pages? Use `@id` with fragment identifiers to
tie them back to one canonical entity node.
## The Wikidata playbook for businesses
Not every business qualifies for Wikidata. Criteria (simplified):
- Multiple independent third-party sources (press articles, books,
academic papers) covering the entity.
- Some form of public notability (not just "we exist").
If qualified, the creation workflow:
1. Create Wikidata account.
2. Use "Create a new item" → name, label, description.
3. Add statements with sources:
- `instance of (P31)` → `enterprise (Q6881511)` or more specific
- `country (P17)` → `France (Q142)`
- `headquarters location (P159)` → city QID
- `official website (P856)` → domain URL
- `inception (P571)` → founding date
- `industry (P452)` → industry QID
- `SIRET (P3893)` → SIRET number (FR)
- `VAT number (P3608)` → VAT ID
4. Each statement must cite a reference (URL of press article,
official registry, etc.).
5. Wait for community review. Items without sources get merged or deleted.
This is labor-intensive and failure-prone for non-notable entities.
Do NOT invent sources. Better to skip Wikidata than create a deletable item.
## Automation options (for SEO.md §11)
- **Kalicube** — paid service specialised in Knowledge Panel + Wikidata
optimization for businesses and executives.
- **Entity.ai** / **InLinks** — tools that help structure entity
signals on-site + track Knowledge Panel status.
- **WordLift** — WordPress/plugin with Wikidata linking + Schema.org
graph generation.
- **Yext Knowledge Graph** — enterprise platform syncing entity data
across 200+ directories.
- **BrightLocal / Moz Local / Uberall** — focus on local citations
+ directory sync (not Wikidata-specific).
For Wikidata specifically: no full-automation tool is reliable because
it requires sourced statements. Human curation is the bottleneck.
## Common mistakes
- **Fake Wikidata entries** — flagged and deleted by community, damages
reputation.
- **`sameAs` pointing to dead profiles** — validate each URL resolves.
- **Inconsistent entity names across platforms** ("Dupont Plomberie"
vs "Plomberie Dupont" vs "DUPONT PLOMBERIE SAS") — pick one, apply
everywhere.
- **Missing VAT/SIREN on Organization schema** — easy credibility
signal, often forgotten.
- **Treating @id as a URL that must resolve** — `@id` is an identifier,
not a mandatory-resolvable URL (though resolvable is better).
## Verification tools
- https://www.wikidata.org/wiki/Special:Search — find QID
- https://tools.wmflabs.org/reasonator/ — human-readable Wikidata view
- https://kalicube.com — commercial Knowledge Panel audit
- https://www.google.com/search?q=%22business+name%22 — check Knowledge Panel
- Schema validator (see `geo-schemas.md`) — check `@id` + `sameAs` integrity
+343
View File
@@ -0,0 +1,343 @@
# Schema.org for GEO — types that matter in 2026
All examples use JSON-LD (the only format Google recommends in 2026).
Place inside `<script type="application/ld+json">` in `<head>` or
before `</body>`.
## DEPRECATED — do not emit
Google deprecated these in June 2025. Stop emitting them and remove
existing instances. They no longer produce rich results.
- `ClaimReview` (was a fact-check signal)
- `CourseInfo`
- `EstimatedSalary`
- `LearningVideo`
- `SpecialAnnouncement`
- `VehicleListing`
- `Book` actions (ReadAction, BuyAction on Book)
## TIER 1 — highest GEO impact
### QAPage — single Q&A format
Pages cited 58% more often by ChatGPT vs basic Article schema.
Use when the page is built around ONE primary question.
```json
{
"@context": "https://schema.org",
"@type": "QAPage",
"mainEntity": {
"@type": "Question",
"name": "What is the best framework for a public website in 2026?",
"text": "Should I use React SPA, Next.js, or Astro for a public-facing website in 2026?",
"answerCount": 1,
"acceptedAnswer": {
"@type": "Answer",
"text": "For public-facing websites, Astro is the 2026 default because it ships static HTML by default, preserves SEO/GEO signals, and allows React/Vue/Svelte islands only where interactivity is needed. React SPAs are only appropriate for authenticated, non-indexed surfaces.",
"dateCreated": "2026-04-21",
"upvoteCount": 0,
"author": {
"@type": "Person",
"name": "Author Name",
"url": "https://example.com/about"
}
}
}
}
```
### FAQPage — multiple Q&A
Only valid when the page visibly contains all listed questions and
answers. Google will penalise pages with FAQ schema that doesn't match
visible content.
```json
{
"@context": "https://schema.org",
"@type": "FAQPage",
"mainEntity": [
{
"@type": "Question",
"name": "How long does shipping take?",
"acceptedAnswer": {
"@type": "Answer",
"text": "Standard shipping takes 2 to 5 business days in France."
}
},
{
"@type": "Question",
"name": "Do you offer refunds?",
"acceptedAnswer": {
"@type": "Answer",
"text": "Yes — refunds are available within 30 days of purchase."
}
}
]
}
```
### Speakable — voice + AI extraction marker
62% of searches in 2026 involve voice. Speakable flags the passage
best suited for voice readout and AI summary.
```json
{
"@context": "https://schema.org",
"@type": "Article",
"headline": "Article headline",
"speakable": {
"@type": "SpeakableSpecification",
"cssSelector": [".article-summary", ".tldr"]
}
}
```
Or via xpath for non-CSS-targetable content:
```json
"speakable": {
"@type": "SpeakableSpecification",
"xpath": ["/html/head/title", "//div[@class='tldr']"]
}
```
### Article + Person — E-E-A-T backbone
The single most important pattern for non-local content. Couples
content to a real author with verifiable credentials.
```json
{
"@context": "https://schema.org",
"@type": "Article",
"headline": "Exact title of the article",
"description": "One-sentence summary matching meta description.",
"image": ["https://example.com/images/hero-1x1.jpg", "https://example.com/images/hero-4x3.jpg", "https://example.com/images/hero-16x9.jpg"],
"datePublished": "2026-04-15T09:00:00+02:00",
"dateModified": "2026-04-21T14:30:00+02:00",
"author": {
"@type": "Person",
"@id": "https://example.com/about#author-jane",
"name": "Jane Doe",
"url": "https://example.com/authors/jane-doe",
"image": "https://example.com/images/jane-doe.jpg",
"jobTitle": "Senior Plumber",
"description": "Master plumber with 15 years of experience in Paris region.",
"knowsAbout": ["plumbing", "boiler repair", "leak detection"],
"alumniOf": "Lycée Professionnel Diderot",
"award": ["Qualibat RGE certification", "Artisan de l'année 2024 Essonne"],
"worksFor": {
"@type": "Organization",
"@id": "https://example.com/#org"
},
"sameAs": [
"https://www.linkedin.com/in/jane-doe-plomberie",
"https://twitter.com/janedoeplumbing",
"https://www.wikidata.org/wiki/Q123456789"
]
},
"publisher": {
"@type": "Organization",
"@id": "https://example.com/#org",
"name": "Business Name",
"logo": {
"@type": "ImageObject",
"url": "https://example.com/logo.png"
}
},
"mainEntityOfPage": {
"@type": "WebPage",
"@id": "https://example.com/article-slug"
}
}
```
## TIER 2 — solid GEO contribution
### HowTo — procedural content
```json
{
"@context": "https://schema.org",
"@type": "HowTo",
"name": "How to reset a Chaffoteaux Talia Green boiler",
"description": "Step-by-step reset procedure for the Talia Green combi boiler.",
"totalTime": "PT5M",
"estimatedCost": {"@type": "MonetaryAmount", "currency": "EUR", "value": "0"},
"tool": [{"@type": "HowToTool", "name": "None"}],
"step": [
{
"@type": "HowToStep",
"name": "Locate the reset button",
"text": "The reset button is on the front panel, marked with a flame icon.",
"url": "https://example.com/guides/reset#step1",
"image": "https://example.com/img/step1.jpg"
},
{
"@type": "HowToStep",
"name": "Press and hold for 3 seconds",
"text": "Press the reset button until the red light turns off.",
"url": "https://example.com/guides/reset#step2"
}
]
}
```
### BreadcrumbList — navigation context for AI
Gives AI the hierarchical position of the page. Nearly universal to
add, low cost.
```json
{
"@context": "https://schema.org",
"@type": "BreadcrumbList",
"itemListElement": [
{"@type": "ListItem", "position": 1, "name": "Accueil", "item": "https://example.com/"},
{"@type": "ListItem", "position": 2, "name": "Services", "item": "https://example.com/services"},
{"@type": "ListItem", "position": 3, "name": "Dépannage chaudière", "item": "https://example.com/services/depannage-chaudiere"}
]
}
```
### LocalBusiness — local services (required for local SEO)
Must be consistent with GMB. Any divergence is a NAP inconsistency.
```json
{
"@context": "https://schema.org",
"@type": "Plumber",
"@id": "https://example.com/#business",
"name": "Plomberie Dupont",
"image": "https://example.com/img/shopfront.jpg",
"url": "https://example.com",
"telephone": "+33123456789",
"priceRange": "€€",
"address": {
"@type": "PostalAddress",
"streetAddress": "12 rue des Lilas",
"addressLocality": "Évry-Courcouronnes",
"postalCode": "91000",
"addressRegion": "Île-de-France",
"addressCountry": "FR"
},
"geo": {
"@type": "GeoCoordinates",
"latitude": 48.62939,
"longitude": 2.44199
},
"openingHoursSpecification": [
{
"@type": "OpeningHoursSpecification",
"dayOfWeek": ["Monday", "Tuesday", "Wednesday", "Thursday", "Friday"],
"opens": "08:00",
"closes": "18:00"
}
],
"areaServed": [
{"@type": "City", "name": "Évry-Courcouronnes"},
{"@type": "City", "name": "Corbeil-Essonnes"},
{"@type": "AdministrativeArea", "name": "Essonne"}
],
"sameAs": [
"https://www.facebook.com/plomberiedupont",
"https://www.instagram.com/plomberiedupont",
"https://www.pagesjaunes.fr/pros/12345",
"https://www.wikidata.org/wiki/Q999999999"
]
}
```
Use the most specific subclass of `LocalBusiness` available (`Plumber`,
`Dentist`, `Restaurant`, `AutoRepair`, etc.) — list at
https://schema.org/LocalBusiness under "More specific Types".
### Organization — company-level entity
Separate from `LocalBusiness` when brand > single location.
```json
{
"@context": "https://schema.org",
"@type": "Organization",
"@id": "https://example.com/#org",
"name": "Company Name",
"legalName": "Company Name SAS",
"url": "https://example.com",
"logo": "https://example.com/logo.png",
"foundingDate": "2015-03-01",
"founders": [{"@type": "Person", "name": "Founder Name"}],
"numberOfEmployees": {"@type": "QuantitativeValue", "value": "12"},
"vatID": "FR12345678901",
"iso6523Code": "0199:123456789",
"sameAs": [
"https://www.wikidata.org/wiki/Q123456",
"https://www.linkedin.com/company/companyname",
"https://www.crunchbase.com/organization/companyname"
],
"contactPoint": {
"@type": "ContactPoint",
"telephone": "+33123456789",
"contactType": "customer service",
"availableLanguage": ["fr", "en"]
}
}
```
### Dataset — factual reference content
Use for data-heavy pages (statistics, research, public-data reports).
```json
{
"@context": "https://schema.org",
"@type": "Dataset",
"name": "French boiler energy consumption by model, 2020-2025",
"description": "Average annual kWh consumption for 47 boiler models installed in France.",
"license": "https://creativecommons.org/licenses/by/4.0/",
"creator": {"@type": "Organization", "@id": "https://example.com/#org"},
"distribution": {
"@type": "DataDownload",
"encodingFormat": "text/csv",
"contentUrl": "https://example.com/data/boilers-2020-2025.csv"
}
}
```
## TIER 3 — niche but high-leverage when applicable
- **`Product`** — e-commerce (required for Merchant Center)
- **`Recipe`** — food sites
- **`Event`** — event listings
- **`JobPosting`** — job boards
- **`Review` / `AggregateRating`** — only when backed by verifiable public reviews (fraud risk otherwise)
- **`VideoObject`** — any embedded video (transcripts are critical for AI)
- **`DefinedTerm` / `DefinedTermSet`** — glossary pages, taxonomy (great for entity disambiguation)
- **`Course` / `EducationalOccupationalCredential`** — training/cert providers
- **`MedicalBusiness`, `PhysiologicalFeature`, `Drug`** — health (YMYL, demand extra rigour)
## Graph linking — @id patterns
Use `@id` to build a single graph across multiple JSON-LD blocks:
```json
{"@context":"https://schema.org","@graph":[
{"@type":"Organization","@id":"https://example.com/#org","name":"..."},
{"@type":"WebSite","@id":"https://example.com/#website","publisher":{"@id":"https://example.com/#org"}},
{"@type":"WebPage","@id":"https://example.com/page#webpage","isPartOf":{"@id":"https://example.com/#website"}},
{"@type":"Article","mainEntityOfPage":{"@id":"https://example.com/page#webpage"},"author":{"@id":"https://example.com/about#author-jane"}}
]}
```
This is the pattern Yoast, RankMath, and modern headless-CMS plugins
output. It lets AI engines traverse entities without duplicating them.
## Validation
- https://validator.schema.org — strict Schema.org validator
- https://search.google.com/test/rich-results — Google Rich Results Test
- https://developers.google.com/search/docs/appearance/structured-data — type-by-type Google docs
+153
View File
@@ -0,0 +1,153 @@
# llms.txt / llms-full.txt — template and strategy
## Status as of 2026-04
**Honest assessment**: llms.txt is a proposed standard by Jeremy Howard
(Answer.AI, Sept 2024). No major AI crawler has publicly confirmed they
extract content via `/llms.txt`. A Search Engine Land study (2025) found
8 of 9 sites saw no measurable traffic change after adoption.
**Why include it anyway**:
- Low cost (small static file).
- Real value for developer-facing sites — AI coding assistants (Cursor,
Continue, Claude Code, GitHub Copilot Chat) DO read it for doc retrieval.
- Signals intent to AI ecosystem. Early mover advantage if adoption grows.
- Reduces RAG token consumption when third parties ingest your content.
**Do not promise ranking gains.** Frame as "no-regret hedge", not "quick win".
## Where it goes
- `/llms.txt` — root of domain. Index of your content in markdown.
- `/llms-full.txt` — root of domain. Full text of your most important pages
concatenated. Optional but recommended for docs/blog/knowledge base.
Both MUST be reachable over HTTPS, content-type `text/plain` or
`text/markdown`, and NOT blocked in robots.txt.
## Canonical structure
```markdown
# <Site or Project Name>
> <One-sentence elevator pitch. This is the single line AI systems extract
> as your site summary. Be concrete. Include entity + category + differentiator.>
<Optional free-form paragraph providing more context. Keep under 400 chars.>
## Docs
- [Getting started](https://example.com/docs/getting-started): What it does, how to install.
- [API reference](https://example.com/docs/api): All endpoints with examples.
- [Tutorials](https://example.com/docs/tutorials): Step-by-step walkthroughs.
## Examples
- [Quickstart example](https://example.com/examples/quickstart.md): Minimal working demo.
## Optional
- [Changelog](https://example.com/changelog.md): Version history.
- [Blog](https://example.com/blog/index.md): In-depth articles.
```
## Structure rules (Jeremy Howard spec)
1. First line: `# <Name>` (H1 with project/site name).
2. Second non-comment line: `> summary` (blockquote, one sentence).
3. Optional paragraphs of free-form context after the blockquote.
4. H2 sections grouping links: `## Docs`, `## Examples`, `## Optional`, etc.
5. Each link: `[Title](URL): description.` — description under 120 chars.
6. Any link pointing to a `.md` version of the page is preferred.
7. Total file: target under 8 KB. If larger, split into `llms-full.txt`.
## llms-full.txt
Concatenation of the full text (stripped of nav/footer/ads) of your most
important pages. Separator between pages:
```
---
URL: https://example.com/docs/getting-started
Title: Getting Started
---
<full markdown content of that page>
---
URL: https://example.com/docs/api
Title: API Reference
---
<full markdown content of that page>
```
Target under 500 KB. If your corpus is larger, trim to highest-value pages
(most-linked, most-traffic, most-updated).
## Generation patterns
### Static sites (Astro, Hugo, Jekyll, 11ty, Next.js SSG)
Best practice: generate both files at build time from the same source as
your regular pages. Examples:
**Astro**: add a `src/pages/llms.txt.ts` endpoint:
```typescript
import { getCollection } from 'astro:content';
export async function GET() {
const docs = await getCollection('docs');
const body = [
'# My Project',
'',
'> One-sentence pitch.',
'',
'## Docs',
...docs.map(d => `- [${d.data.title}](https://example.com/docs/${d.slug}): ${d.data.description}`),
].join('\n');
return new Response(body, { headers: { 'Content-Type': 'text/plain' } });
}
```
**Next.js App Router**: `app/llms.txt/route.ts`:
```typescript
export async function GET() {
// similar — pull from your CMS/MDX/db
return new Response(body, { headers: { 'Content-Type': 'text/plain' } });
}
```
**Hugo**: custom output format `llms` → `llms.txt` template in layouts.
### CMS (WordPress, Drupal, Ghost)
Use a plugin OR a cron job that regenerates files weekly. Flag stale
files (older than site content) in audits.
### Static HTML / PHP
Hand-maintained file. Flag in audits if older than 90 days.
## Automation tools (for SEO.md §11 "automatisation possible")
- **`llms-txt-action`** (GitHub Action) — generates on each deploy
- **Mintlify** — auto-generates for Mintlify-hosted docs
- **Fern** — auto-generates for Fern-generated API docs
- **`llmstxt-hub`** — community directory of examples
- Custom script + cron — works for any static content source
## What NOT to put in llms.txt
- Login walls / private content
- Pricing tables (change frequently → stale risk)
- Testimonials (authenticity risk if AI quotes them)
- Marketing fluff without factual anchors
## Validation checklist
- [ ] File reachable at `/llms.txt` over HTTPS
- [ ] Content-type `text/plain` or `text/markdown`
- [ ] H1 + blockquote present as first two non-comment lines
- [ ] All linked URLs resolve (200)
- [ ] No broken markdown (valid CommonMark)
- [ ] Mentioned in `/sitemap.xml`? Optional, debated
- [ ] NOT blocked in `/robots.txt`