forked from bchanot/claude
opposite sample I5 made COVERAGE mandatory and told the agent to "sample by risk, one per template". Grounding that against a real Astro site shows the rule is half wrong, and that one number was hiding two. 86 URLs collapse into 8 families; 75 of them (87%) come from 3 dynamic [dept] templates. So the same 12 sampled pages are simultaneously 14% LIVE coverage and ~100% SOURCE coverage. Reporting only 14% understates the audit; reporting only 100% oversells it. Both lines now, in both agents — and they bound different findings, so they must not be averaged: SOURCE bounds CODE (one template renders its whole family: a missing canonical in [dept]/index.astro breaks all 25 identically). LIVE bounds CONTENT (title wording, thin pages, duplication — written per page, so a template says nothing about its 25 instances). The sharper half: "one per template" is CORRECT for code and WRONG for the 30/70 rule, which this same spec mandates at :954 (city pages: 30% shared, 70% unique). You cannot tell whether 25 dept pages are 70% unique by reading one of them. The spec was mandating a check its own sampling made structurally impossible. Sampling is now keyed to the finding class: 1 per family for code, >=3 from the LARGEST family for duplication, spread for per-page content. Families come from the first path segment of the sitemap URLs (C1b) — a good enough proxy for "same template" that needs no framework routing knowledge, verified against the real distribution. geo gets the same split, cut differently: JSON-LD lives in a shared layout so SOURCE bounds Schema.org, while Definition Lead / TL;DR are written per page so LIVE bounds Content Shape. Site-wide axes (crawler policy, llms.txt) stay unbounded — single files, fully read. Verified: full suite green. Caught and fixed one self-inflicted contradiction before commit — geo's prose demanded both coverages while its output block still had a single line.