opposite sample
I5 made COVERAGE mandatory and told the agent to "sample by risk, one per
template". Grounding that against a real Astro site shows the rule is half
wrong, and that one number was hiding two.
86 URLs collapse into 8 families; 75 of them (87%) come from 3 dynamic
[dept] templates. So the same 12 sampled pages are simultaneously 14% LIVE
coverage and ~100% SOURCE coverage. Reporting only 14% understates the audit;
reporting only 100% oversells it. Both lines now, in both agents — and they
bound different findings, so they must not be averaged:
SOURCE bounds CODE (one template renders its whole family: a missing
canonical in [dept]/index.astro breaks all 25 identically).
LIVE bounds CONTENT (title wording, thin pages, duplication — written per
page, so a template says nothing about its 25 instances).
The sharper half: "one per template" is CORRECT for code and WRONG for the
30/70 rule, which this same spec mandates at :954 (city pages: 30% shared,
70% unique). You cannot tell whether 25 dept pages are 70% unique by reading
one of them. The spec was mandating a check its own sampling made
structurally impossible. Sampling is now keyed to the finding class: 1 per
family for code, >=3 from the LARGEST family for duplication, spread for
per-page content.
Families come from the first path segment of the sitemap URLs (C1b) — a good
enough proxy for "same template" that needs no framework routing knowledge,
verified against the real distribution.
geo gets the same split, cut differently: JSON-LD lives in a shared layout so
SOURCE bounds Schema.org, while Definition Lead / TL;DR are written per page
so LIVE bounds Content Shape. Site-wide axes (crawler policy, llms.txt) stay
unbounded — single files, fully read.
Verified: full suite green. Caught and fixed one self-inflicted contradiction
before commit — geo's prose demanded both coverages while its output block
still had a single line.