fix(seo-data,seo): backtest on a second, native site — two real bugs

Everything on this branch was grounded on ONE Astro repo. A native PHP site
(lavageangels356.fr) broke two things that looked fine there.

BUG 1 — sitemap counted images as pages. _locs matched
`el.tag.endswith("}loc")`, and <image:loc> from Google's image-sitemap
namespace ALSO ends with '}loc'. Astro's sitemap has no image extension, so
this was invisible. The native site's does: 24 <url> + 3 <image:loc> came back
as count=27. The COVERAGE denominator was 12.5% too high and img/logo.png was
about to be sampled and audited as a page.
Fixed with two locks: walk the DIRECT children of each <url>/<sitemap> instead
of root.iter() (which alone excludes <image:image><image:loc>), and test the
sitemaps.org namespace explicitly. Regression fixture carries the image
extension; the old endswith code returns 9 URLs against it, the new one 7 with
zero images.
Verified both sites: native 27 -> 24, zero images; Astro unchanged at 86.

BUG 2 — the C1c family heuristic was tuned to one URL layout. "First path
segment" works for NESTED city pages (/creation-site-internet/essonne-91/ →
25 pages, 1 family) and FAILS for FLAT ones (/lavage-auto-pomponne,
/lavage-auto-torcy → 8 pages, 8 singletons). Consequence: C1c's rule "sample
>=3 from the largest family" would have targeted /services (5) and missed the
8 city pages entirely — the exact doorway-page risk the 30/70 rule exists to
catch.
Family is now "shared parent path OR shared slug prefix (>=3 URLs sharing 2+
hyphen tokens)", with both real layouts as the worked examples, plus a
sanity-check: a sitemap yielding almost as many families as URLs has defeated
the heuristic, not proved the site has no templates. Fixed in seo-analyzer and
in the geo pointer that referenced it.

Backtest results on the native site for everything else: url-guard accepts the
domain; source-scope excludes only .git (no dist/build/out exists — the
exclusions are correctly no-ops, and cache/ holds only .htaccess+.gitignore so
it is rightly untouched); the sameAs check runs and finds zero (a real GEO gap
for that site, not a tool bug); links are present in the served HTML (PHP is
SSR), so C3 is feasible there.

Verified: seo-data 119 -> 122 pass, 0 fail; full suite green; py_compile clean.
This commit is contained in:
Bastien Chanot
2026-07-17 12:28:36 +02:00
parent 3a15643c2c
commit dca977bb27
5 changed files with 68 additions and 14 deletions
+4 -2
View File
@@ -679,8 +679,10 @@ what bounds it. Content Shape does NOT work that way: Definition Lead, TL;DR
and heading wording are written per page, so a template says nothing about and heading wording are written per page, so a template says nothing about
its 25 instances. Bound Schema.org by SOURCE, Content Shape by LIVE, and its 25 instances. Bound Schema.org by SOURCE, Content Shape by LIVE, and
never quote the flattering one alone. Get the URL families from never quote the flattering one alone. Get the URL families from
`fetch.sh sitemap` (first path segment); if `/seo` already ran it, reuse the `fetch.sh sitemap`, grouped as seo-analyzer STEP 5 describes — shared parent
count rather than re-fetching. path OR shared slug prefix, because both layouts are real: first-segment
alone reads 8 flat `/lavage-auto-<city>` pages as 8 singletons. If `/seo`
already ran it, reuse the count rather than re-fetching.
Per user instruction: **GEO weight in combined SEO+GEO report = 20% for Per user instruction: **GEO weight in combined SEO+GEO report = 20% for
local, 25% for national/SaaS/content.** local, 25% for national/SaaS/content.**
+19 -4
View File
@@ -508,10 +508,25 @@ point of use (same contract as the sameAs check in geo-analyzer).
### Meta tags per page (sample 5-15 key pages) ### Meta tags per page (sample 5-15 key pages)
**Group the sitemap URLs into families first** — first path segment is a **Group the sitemap URLs into families first** — a family is "pages one
good enough proxy for "same template", and it needs no framework routing template renders". You do not need framework routing knowledge to see them,
knowledge. Measured on a real Astro site: 86 URLs collapse into 8 families, but you DO need to look at the actual URL shape, because it varies:
and 75 of them (87%) come from just 3 dynamic `[dept]` templates.
| Layout | Example | Family signal |
|---|---|---|
| Nested | `/creation-site-internet/essonne-91/`, `/creation-site-internet/seine-et-marne-77/` | **shared parent path** → 25 pages, 1 family |
| **Flat** | `/lavage-auto-pomponne`, `/lavage-auto-torcy`, `/lavage-auto-chelles` | **shared slug prefix** → 8 pages, 1 family |
Both are real, measured on two live sites. First-path-segment alone handles
the nested case and **fails the flat one**: those 8 city pages read as 8
unrelated singletons, so the largest "family" becomes `/services` (5) and the
doorway-page risk — the exact thing the 30/70 rule exists to catch — is
invisible. Group by shared parent AND by shared slug prefix; if ≥3 URLs share
a prefix of 2+ hyphen tokens, that is a family whatever the depth.
Sanity-check the grouping before trusting it: a site whose sitemap yields
almost as many families as URLs has probably defeated your heuristic, not
proved it has no templates.
**Sample by finding class, because the classes need opposite samples:** **Sample by finding class, because the classes need opposite samples:**
+11 -1
View File
@@ -1,9 +1,19 @@
<?xml version="1.0" encoding="UTF-8"?> <?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9" xmlns:xhtml="http://www.w3.org/1999/xhtml"> <urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9"
xmlns:xhtml="http://www.w3.org/1999/xhtml"
xmlns:image="http://www.google.com/schemas/sitemap-image/1.1">
<!-- image:loc also ends with }loc — it must NOT be counted as a page -->
<url> <url>
<loc>https://ex.com/</loc> <loc>https://ex.com/</loc>
<changefreq>weekly</changefreq> <changefreq>weekly</changefreq>
<xhtml:link rel="alternate" hreflang="en" href="https://ex.com/en/" /> <xhtml:link rel="alternate" hreflang="en" href="https://ex.com/en/" />
<image:image>
<image:loc>https://ex.com/img/logo.png</image:loc>
<image:title>Logo</image:title>
</image:image>
<image:image>
<image:loc>https://ex.com/img/hero.jpeg</image:loc>
</image:image>
</url> </url>
<url><loc>https://ex.com/services</loc></url> <url><loc>https://ex.com/services</loc></url>
<url><loc>https://ex.com/blog</loc></url> <url><loc>https://ex.com/blog</loc></url>
+6
View File
@@ -111,6 +111,12 @@ hasnt "sitemap drops non-http" "$SM" 'ftp://'
hasnt "sitemap drops shell-meta" "$SM" 'bad"quote' hasnt "sitemap drops shell-meta" "$SM" 'bad"quote'
# namespace-agnostic: real sitemaps carry sitemaps.org xmlns (+ xhtml here) # namespace-agnostic: real sitemaps carry sitemaps.org xmlns (+ xhtml here)
has "sitemap reads namespaced" "$SM" '"https://ex.com/services"' has "sitemap reads namespaced" "$SM" '"https://ex.com/services"'
# REGRESSION: <image:loc> also ends with '}loc'. An endswith test counted image
# sitemap entries as pages — a real native site returned 27 for 24 <url>, and
# img/logo.png was about to be sampled and audited as a page.
hasnt "image:loc is not a page" "$SM" '/img/logo.png'
hasnt "image:loc jpeg not a page" "$SM" '/img/hero.jpeg'
has "image ns does not inflate count" "$SM" '"count": 4'
IDX="$(SEO_DATA_MOCK_DIR="$SD/fixtures-sitemap-index" python3 "$SD/sitemap.py" \ IDX="$(SEO_DATA_MOCK_DIR="$SD/fixtures-sitemap-index" python3 "$SD/sitemap.py" \
--url https://ex.com/sitemap.xml)" --url https://ex.com/sitemap.xml)"
+28 -7
View File
@@ -57,19 +57,40 @@ def _refuse_dtd(raw):
if b"<!DOCTYPE" in head or b"<!ENTITY" in head: if b"<!DOCTYPE" in head or b"<!ENTITY" in head:
raise UnsafeXML("DTD in sitemap") raise UnsafeXML("DTD in sitemap")
SITEMAP_NS = "{http://www.sitemaps.org/schemas/sitemap/0.9}"
def _is_page_loc(tag):
"""A PAGE <loc>: sitemaps.org namespace, or namespace-less.
NOT <image:loc> or <video:loc>. Those live in Google's extension
namespaces and name an ASSET inside a <url>, not a page of its own. An
endswith('}loc') test matches them too — that shipped, and a real site
caught it: 24 <url> + 3 <image:loc> came back as a count of 27, so the
COVERAGE denominator was 12.5% too high and img/logo.png was about to be
sampled and audited as a page.
"""
return tag == SITEMAP_NS + "loc" or tag == "loc"
def _locs(raw): def _locs(raw):
"""(<loc> texts, is_sitemapindex). Namespace-agnostic: real sitemaps carry """(page <loc> texts, is_sitemapindex).
the sitemaps.org xmlns and often xhtml too."""
Walks the DIRECT children of each <url>/<sitemap> rather than root.iter():
that alone excludes <image:image><image:loc>, and the namespace test above
is the second lock. XML comments iterate as elements with no children, so
they fall through harmlessly.
"""
import xml.etree.ElementTree as ET # stdlib, lazy import xml.etree.ElementTree as ET # stdlib, lazy
_refuse_dtd(raw) _refuse_dtd(raw)
root = ET.fromstring(raw) root = ET.fromstring(raw)
is_index = root.tag.endswith("sitemapindex") is_index = root.tag.endswith("sitemapindex")
out = [] out = []
for el in root.iter(): for entry in root: # <url> | <sitemap>
if el.tag.endswith("}loc") or el.tag == "loc": for child in entry: # direct children only
text = (el.text or "").strip() if _is_page_loc(child.tag):
if text: text = (child.text or "").strip()
out.append(text) if text:
out.append(text)
break # one <loc> per entry
return out, is_index return out, is_index
def _sane(u): def _sane(u):