forked from bchanot/claude
fix(seo,geo): C1c — split COVERAGE source/live; the 30/70 rule needed the
opposite sample I5 made COVERAGE mandatory and told the agent to "sample by risk, one per template". Grounding that against a real Astro site shows the rule is half wrong, and that one number was hiding two. 86 URLs collapse into 8 families; 75 of them (87%) come from 3 dynamic [dept] templates. So the same 12 sampled pages are simultaneously 14% LIVE coverage and ~100% SOURCE coverage. Reporting only 14% understates the audit; reporting only 100% oversells it. Both lines now, in both agents — and they bound different findings, so they must not be averaged: SOURCE bounds CODE (one template renders its whole family: a missing canonical in [dept]/index.astro breaks all 25 identically). LIVE bounds CONTENT (title wording, thin pages, duplication — written per page, so a template says nothing about its 25 instances). The sharper half: "one per template" is CORRECT for code and WRONG for the 30/70 rule, which this same spec mandates at :954 (city pages: 30% shared, 70% unique). You cannot tell whether 25 dept pages are 70% unique by reading one of them. The spec was mandating a check its own sampling made structurally impossible. Sampling is now keyed to the finding class: 1 per family for code, >=3 from the LARGEST family for duplication, spread for per-page content. Families come from the first path segment of the sitemap URLs (C1b) — a good enough proxy for "same template" that needs no framework routing knowledge, verified against the real distribution. geo gets the same split, cut differently: JSON-LD lives in a shared layout so SOURCE bounds Schema.org, while Definition Lead / TL;DR are written per page so LIVE bounds Content Shape. Site-wide axes (crawler policy, llms.txt) stay unbounded — single files, fully read. Verified: full suite green. Caught and fixed one self-inflicted contradiction before commit — geo's prose demanded both coverages while its output block still had a single line.
This commit is contained in:
+13
-1
@@ -653,7 +653,9 @@ Score each axis. Use concrete findings from STEP 2-9.
|
||||
|
||||
```
|
||||
GEO SCORING (<depth>)
|
||||
COVERAGE : <N> of <M> sitemap URLs (<P>%) | <N> pages, total UNKNOWN
|
||||
COVERAGE SOURCE : <N> of <M> page templates (<P>%) — bounds Schema.org
|
||||
COVERAGE LIVE : <N> of <M> sitemap URLs (<P>%) — bounds Content Shape
|
||||
| UNKNOWN (no sitemap / fetch degraded)
|
||||
AI Crawlers Policy : XX/20 <justification>
|
||||
llms.txt : XX/20 <justification>
|
||||
Schema.org for AI : XX/20 <justification>
|
||||
@@ -670,6 +672,16 @@ Schema.org. Site-wide axes (AI Crawlers Policy, llms.txt) are unaffected:
|
||||
robots.txt and llms.txt are single files, fully read. Say which is which
|
||||
rather than letting one ratio discredit the whole report.
|
||||
|
||||
**Same source/live split as seo-analyzer STEP 9 (C1c), and it cuts your axes
|
||||
differently.** A JSON-LD block lives in a shared layout, so one sampled page
|
||||
per URL family proves the SCHEMA for the whole family — SOURCE coverage is
|
||||
what bounds it. Content Shape does NOT work that way: Definition Lead, TL;DR
|
||||
and heading wording are written per page, so a template says nothing about
|
||||
its 25 instances. Bound Schema.org by SOURCE, Content Shape by LIVE, and
|
||||
never quote the flattering one alone. Get the URL families from
|
||||
`fetch.sh sitemap` (first path segment); if `/seo` already ran it, reuse the
|
||||
count rather than re-fetching.
|
||||
|
||||
Per user instruction: **GEO weight in combined SEO+GEO report = 20% for
|
||||
local, 25% for national/SaaS/content.**
|
||||
|
||||
|
||||
Reference in New Issue
Block a user