fix(seo-data,seo): backtest on a second, native site — two real bugs

Everything on this branch was grounded on ONE Astro repo. A native PHP site
(lavageangels356.fr) broke two things that looked fine there.

BUG 1 — sitemap counted images as pages. _locs matched
`el.tag.endswith("}loc")`, and <image:loc> from Google's image-sitemap
namespace ALSO ends with '}loc'. Astro's sitemap has no image extension, so
this was invisible. The native site's does: 24 <url> + 3 <image:loc> came back
as count=27. The COVERAGE denominator was 12.5% too high and img/logo.png was
about to be sampled and audited as a page.
Fixed with two locks: walk the DIRECT children of each <url>/<sitemap> instead
of root.iter() (which alone excludes <image:image><image:loc>), and test the
sitemaps.org namespace explicitly. Regression fixture carries the image
extension; the old endswith code returns 9 URLs against it, the new one 7 with
zero images.
Verified both sites: native 27 -> 24, zero images; Astro unchanged at 86.

BUG 2 — the C1c family heuristic was tuned to one URL layout. "First path
segment" works for NESTED city pages (/creation-site-internet/essonne-91/ →
25 pages, 1 family) and FAILS for FLAT ones (/lavage-auto-pomponne,
/lavage-auto-torcy → 8 pages, 8 singletons). Consequence: C1c's rule "sample
>=3 from the largest family" would have targeted /services (5) and missed the
8 city pages entirely — the exact doorway-page risk the 30/70 rule exists to
catch.
Family is now "shared parent path OR shared slug prefix (>=3 URLs sharing 2+
hyphen tokens)", with both real layouts as the worked examples, plus a
sanity-check: a sitemap yielding almost as many families as URLs has defeated
the heuristic, not proved the site has no templates. Fixed in seo-analyzer and
in the geo pointer that referenced it.

Backtest results on the native site for everything else: url-guard accepts the
domain; source-scope excludes only .git (no dist/build/out exists — the
exclusions are correctly no-ops, and cache/ holds only .htaccess+.gitignore so
it is rightly untouched); the sameAs check runs and finds zero (a real GEO gap
for that site, not a tool bug); links are present in the served HTML (PHP is
SSR), so C3 is feasible there.

Verified: seo-data 119 -> 122 pass, 0 fail; full suite green; py_compile clean.
This commit is contained in:
Bastien Chanot
2026-07-17 12:28:36 +02:00
parent 3a15643c2c
commit dca977bb27
5 changed files with 68 additions and 14 deletions
+28 -7
View File
@@ -57,19 +57,40 @@ def _refuse_dtd(raw):
if b"<!DOCTYPE" in head or b"<!ENTITY" in head:
raise UnsafeXML("DTD in sitemap")
SITEMAP_NS = "{http://www.sitemaps.org/schemas/sitemap/0.9}"
def _is_page_loc(tag):
"""A PAGE <loc>: sitemaps.org namespace, or namespace-less.
NOT <image:loc> or <video:loc>. Those live in Google's extension
namespaces and name an ASSET inside a <url>, not a page of its own. An
endswith('}loc') test matches them too — that shipped, and a real site
caught it: 24 <url> + 3 <image:loc> came back as a count of 27, so the
COVERAGE denominator was 12.5% too high and img/logo.png was about to be
sampled and audited as a page.
"""
return tag == SITEMAP_NS + "loc" or tag == "loc"
def _locs(raw):
"""(<loc> texts, is_sitemapindex). Namespace-agnostic: real sitemaps carry
the sitemaps.org xmlns and often xhtml too."""
"""(page <loc> texts, is_sitemapindex).
Walks the DIRECT children of each <url>/<sitemap> rather than root.iter():
that alone excludes <image:image><image:loc>, and the namespace test above
is the second lock. XML comments iterate as elements with no children, so
they fall through harmlessly.
"""
import xml.etree.ElementTree as ET # stdlib, lazy
_refuse_dtd(raw)
root = ET.fromstring(raw)
is_index = root.tag.endswith("sitemapindex")
out = []
for el in root.iter():
if el.tag.endswith("}loc") or el.tag == "loc":
text = (el.text or "").strip()
if text:
out.append(text)
for entry in root: # <url> | <sitemap>
for child in entry: # direct children only
if _is_page_loc(child.tag):
text = (child.text or "").strip()
if text:
out.append(text)
break # one <loc> per entry
return out, is_index
def _sane(u):