fix(seo,geo): C1c — split COVERAGE source/live; the 30/70 rule needed the
opposite sample I5 made COVERAGE mandatory and told the agent to "sample by risk, one per template". Grounding that against a real Astro site shows the rule is half wrong, and that one number was hiding two. 86 URLs collapse into 8 families; 75 of them (87%) come from 3 dynamic [dept] templates. So the same 12 sampled pages are simultaneously 14% LIVE coverage and ~100% SOURCE coverage. Reporting only 14% understates the audit; reporting only 100% oversells it. Both lines now, in both agents — and they bound different findings, so they must not be averaged: SOURCE bounds CODE (one template renders its whole family: a missing canonical in [dept]/index.astro breaks all 25 identically). LIVE bounds CONTENT (title wording, thin pages, duplication — written per page, so a template says nothing about its 25 instances). The sharper half: "one per template" is CORRECT for code and WRONG for the 30/70 rule, which this same spec mandates at :954 (city pages: 30% shared, 70% unique). You cannot tell whether 25 dept pages are 70% unique by reading one of them. The spec was mandating a check its own sampling made structurally impossible. Sampling is now keyed to the finding class: 1 per family for code, >=3 from the LARGEST family for duplication, spread for per-page content. Families come from the first path segment of the sitemap URLs (C1b) — a good enough proxy for "same template" that needs no framework routing knowledge, verified against the real distribution. geo gets the same split, cut differently: JSON-LD lives in a shared layout so SOURCE bounds Schema.org, while Definition Lead / TL;DR are written per page so LIVE bounds Content Shape. Site-wide axes (crawler policy, llms.txt) stay unbounded — single files, fully read. Verified: full suite green. Caught and fixed one self-inflicted contradiction before commit — geo's prose demanded both coverages while its output block still had a single line.
This commit is contained in:
+13
-1
@@ -653,7 +653,9 @@ Score each axis. Use concrete findings from STEP 2-9.
|
|||||||
|
|
||||||
```
|
```
|
||||||
GEO SCORING (<depth>)
|
GEO SCORING (<depth>)
|
||||||
COVERAGE : <N> of <M> sitemap URLs (<P>%) | <N> pages, total UNKNOWN
|
COVERAGE SOURCE : <N> of <M> page templates (<P>%) — bounds Schema.org
|
||||||
|
COVERAGE LIVE : <N> of <M> sitemap URLs (<P>%) — bounds Content Shape
|
||||||
|
| UNKNOWN (no sitemap / fetch degraded)
|
||||||
AI Crawlers Policy : XX/20 <justification>
|
AI Crawlers Policy : XX/20 <justification>
|
||||||
llms.txt : XX/20 <justification>
|
llms.txt : XX/20 <justification>
|
||||||
Schema.org for AI : XX/20 <justification>
|
Schema.org for AI : XX/20 <justification>
|
||||||
@@ -670,6 +672,16 @@ Schema.org. Site-wide axes (AI Crawlers Policy, llms.txt) are unaffected:
|
|||||||
robots.txt and llms.txt are single files, fully read. Say which is which
|
robots.txt and llms.txt are single files, fully read. Say which is which
|
||||||
rather than letting one ratio discredit the whole report.
|
rather than letting one ratio discredit the whole report.
|
||||||
|
|
||||||
|
**Same source/live split as seo-analyzer STEP 9 (C1c), and it cuts your axes
|
||||||
|
differently.** A JSON-LD block lives in a shared layout, so one sampled page
|
||||||
|
per URL family proves the SCHEMA for the whole family — SOURCE coverage is
|
||||||
|
what bounds it. Content Shape does NOT work that way: Definition Lead, TL;DR
|
||||||
|
and heading wording are written per page, so a template says nothing about
|
||||||
|
its 25 instances. Bound Schema.org by SOURCE, Content Shape by LIVE, and
|
||||||
|
never quote the flattering one alone. Get the URL families from
|
||||||
|
`fetch.sh sitemap` (first path segment); if `/seo` already ran it, reuse the
|
||||||
|
count rather than re-fetching.
|
||||||
|
|
||||||
Per user instruction: **GEO weight in combined SEO+GEO report = 20% for
|
Per user instruction: **GEO weight in combined SEO+GEO report = 20% for
|
||||||
local, 25% for national/SaaS/content.**
|
local, 25% for national/SaaS/content.**
|
||||||
|
|
||||||
|
|||||||
+43
-12
@@ -479,11 +479,25 @@ point of use (same contract as the sameAs check in geo-analyzer).
|
|||||||
|
|
||||||
### Meta tags per page (sample 5-15 key pages)
|
### Meta tags per page (sample 5-15 key pages)
|
||||||
|
|
||||||
Sample by risk, not convenience: homepage + top templates (one per page
|
**Group the sitemap URLs into families first** — first path segment is a
|
||||||
type: service, city, blog, product, legal) + any page GSC flags as a
|
good enough proxy for "same template", and it needs no framework routing
|
||||||
position 4-10 quick win. Same template audited twice buys nothing; an
|
knowledge. Measured on a real Astro site: 86 URLs collapse into 8 families,
|
||||||
un-sampled template is an un-audited template — name the templates you
|
and 75 of them (87%) come from just 3 dynamic `[dept]` templates.
|
||||||
skipped.
|
|
||||||
|
**Sample by finding class, because the classes need opposite samples:**
|
||||||
|
|
||||||
|
| Looking for | Sample | Why |
|
||||||
|
|---|---|---|
|
||||||
|
| Code defects (canonical, OG, `<img>` dims, hreflang) | **1 per family** | one template renders the whole family — a missing canonical in `[dept]/index.astro` breaks all 25 identically. 1 per family ≈ 100% SOURCE coverage for ~8 fetches. |
|
||||||
|
| **Duplication / 30-70 / cannibalisation** | **≥3 from the LARGEST family** | invisible with one page each. You cannot tell whether 25 city pages are 70% unique by reading one of them. |
|
||||||
|
| Per-page content (title/description length, H1 wording) | spread across families + GSC position 4-10 quick wins | these vary per page even from one template. |
|
||||||
|
|
||||||
|
"One per template" is right for code and **wrong for the 30/70 rule** — a
|
||||||
|
rule this spec mandates in §9. Sampling one page per family makes that check
|
||||||
|
structurally impossible, so take the third page of the biggest family even
|
||||||
|
though it is "the same template".
|
||||||
|
|
||||||
|
An un-sampled family is an un-audited family. Name the ones you skipped.
|
||||||
|
|
||||||
For each sampled page:
|
For each sampled page:
|
||||||
```
|
```
|
||||||
@@ -868,8 +882,9 @@ misroutes the client-handover gate and the user's effort.
|
|||||||
|
|
||||||
```
|
```
|
||||||
SEO SCORING (<depth>)
|
SEO SCORING (<depth>)
|
||||||
COVERAGE : <N> of <M> sitemap URLs (<P>%) — templates skipped: <list|none>
|
COVERAGE SOURCE: <N> of <M> page templates (<P>%) — skipped: <list|none>
|
||||||
| <N> pages, total UNKNOWN (no sitemap)
|
COVERAGE LIVE : <N> of <M> sitemap URLs (<P>%) — families: <fam N/M, …>
|
||||||
|
| UNKNOWN (no sitemap / fetch degraded)
|
||||||
Technical : XX/20 <justification>
|
Technical : XX/20 <justification>
|
||||||
On-page : XX/20 <justification>
|
On-page : XX/20 <justification>
|
||||||
SEO Local : XX/20 | N/A
|
SEO Local : XX/20 | N/A
|
||||||
@@ -881,11 +896,27 @@ Legal : XX/20 <justification>
|
|||||||
SEO GLOBAL (weighted): XX.X/20 (<depth>)
|
SEO GLOBAL (weighted): XX.X/20 (<depth>)
|
||||||
```
|
```
|
||||||
|
|
||||||
**COVERAGE is mandatory, never omitted, never rounded up.** It is the
|
**Both COVERAGE lines are mandatory, never omitted, never rounded up.** They
|
||||||
honesty bound on every page-level axis: On-page and the on-page share of
|
are the honesty bound on every page-level axis: On-page and the on-page share
|
||||||
Technical are extrapolations from the sample. If coverage < 25%, repeat it
|
of Technical are extrapolations from the sample, and `/client-handover` gates
|
||||||
in §0 as a major alert — a 17/20 drawn from 3% of a site is not a 17/20, and
|
on these numbers.
|
||||||
`/client-handover` gates on these numbers.
|
|
||||||
|
**Report both, because they bound different findings — do not average them
|
||||||
|
into one comforting number.**
|
||||||
|
- **SOURCE** bounds CODE findings. One template renders its whole family, so
|
||||||
|
1 page per family can legitimately reach 100% here. High SOURCE coverage is
|
||||||
|
a real claim: the code paths were seen.
|
||||||
|
- **LIVE** bounds CONTENT findings — title/description wording, thin pages,
|
||||||
|
30/70 duplication. It stays low by design and that is fine, as long as it
|
||||||
|
is printed. Measured on a real site: 12 of 86 URLs is 14% LIVE while the
|
||||||
|
same 12 pages are 100% SOURCE. Reporting only the 14% understates the audit;
|
||||||
|
reporting only the 100% oversells it. Both, or neither means anything.
|
||||||
|
- LIVE < 25% → repeat in §0. A 17/20 for content drawn from 3% of a site is
|
||||||
|
not a 17/20.
|
||||||
|
- SOURCE < 100% → name the skipped templates in §0. That is not a sampling
|
||||||
|
choice, it is code nobody read.
|
||||||
|
- Denominator UNKNOWN (no sitemap, or `sitemap` degraded) → print UNKNOWN.
|
||||||
|
Never let silence imply full coverage.
|
||||||
|
|
||||||
Per user instruction: this score represents **80% of the combined
|
Per user instruction: this score represents **80% of the combined
|
||||||
final score for local B2C (20% for GEO), or 75% for SaaS/national
|
final score for local B2C (20% for GEO), or 75% for SaaS/national
|
||||||
|
|||||||
Reference in New Issue
Block a user