forked from bchanot/claude
fix(seo-data,seo): backtest on a second, native site — two real bugs
Everything on this branch was grounded on ONE Astro repo. A native PHP site
(lavageangels356.fr) broke two things that looked fine there.
BUG 1 — sitemap counted images as pages. _locs matched
`el.tag.endswith("}loc")`, and <image:loc> from Google's image-sitemap
namespace ALSO ends with '}loc'. Astro's sitemap has no image extension, so
this was invisible. The native site's does: 24 <url> + 3 <image:loc> came back
as count=27. The COVERAGE denominator was 12.5% too high and img/logo.png was
about to be sampled and audited as a page.
Fixed with two locks: walk the DIRECT children of each <url>/<sitemap> instead
of root.iter() (which alone excludes <image:image><image:loc>), and test the
sitemaps.org namespace explicitly. Regression fixture carries the image
extension; the old endswith code returns 9 URLs against it, the new one 7 with
zero images.
Verified both sites: native 27 -> 24, zero images; Astro unchanged at 86.
BUG 2 — the C1c family heuristic was tuned to one URL layout. "First path
segment" works for NESTED city pages (/creation-site-internet/essonne-91/ →
25 pages, 1 family) and FAILS for FLAT ones (/lavage-auto-pomponne,
/lavage-auto-torcy → 8 pages, 8 singletons). Consequence: C1c's rule "sample
>=3 from the largest family" would have targeted /services (5) and missed the
8 city pages entirely — the exact doorway-page risk the 30/70 rule exists to
catch.
Family is now "shared parent path OR shared slug prefix (>=3 URLs sharing 2+
hyphen tokens)", with both real layouts as the worked examples, plus a
sanity-check: a sitemap yielding almost as many families as URLs has defeated
the heuristic, not proved the site has no templates. Fixed in seo-analyzer and
in the geo pointer that referenced it.
Backtest results on the native site for everything else: url-guard accepts the
domain; source-scope excludes only .git (no dist/build/out exists — the
exclusions are correctly no-ops, and cache/ holds only .htaccess+.gitignore so
it is rightly untouched); the sameAs check runs and finds zero (a real GEO gap
for that site, not a tool bug); links are present in the served HTML (PHP is
SSR), so C3 is feasible there.
Verified: seo-data 119 -> 122 pass, 0 fail; full suite green; py_compile clean.
This commit is contained in:
@@ -679,8 +679,10 @@ what bounds it. Content Shape does NOT work that way: Definition Lead, TL;DR
|
|||||||
and heading wording are written per page, so a template says nothing about
|
and heading wording are written per page, so a template says nothing about
|
||||||
its 25 instances. Bound Schema.org by SOURCE, Content Shape by LIVE, and
|
its 25 instances. Bound Schema.org by SOURCE, Content Shape by LIVE, and
|
||||||
never quote the flattering one alone. Get the URL families from
|
never quote the flattering one alone. Get the URL families from
|
||||||
`fetch.sh sitemap` (first path segment); if `/seo` already ran it, reuse the
|
`fetch.sh sitemap`, grouped as seo-analyzer STEP 5 describes — shared parent
|
||||||
count rather than re-fetching.
|
path OR shared slug prefix, because both layouts are real: first-segment
|
||||||
|
alone reads 8 flat `/lavage-auto-<city>` pages as 8 singletons. If `/seo`
|
||||||
|
already ran it, reuse the count rather than re-fetching.
|
||||||
|
|
||||||
Per user instruction: **GEO weight in combined SEO+GEO report = 20% for
|
Per user instruction: **GEO weight in combined SEO+GEO report = 20% for
|
||||||
local, 25% for national/SaaS/content.**
|
local, 25% for national/SaaS/content.**
|
||||||
|
|||||||
+19
-4
@@ -508,10 +508,25 @@ point of use (same contract as the sameAs check in geo-analyzer).
|
|||||||
|
|
||||||
### Meta tags per page (sample 5-15 key pages)
|
### Meta tags per page (sample 5-15 key pages)
|
||||||
|
|
||||||
**Group the sitemap URLs into families first** — first path segment is a
|
**Group the sitemap URLs into families first** — a family is "pages one
|
||||||
good enough proxy for "same template", and it needs no framework routing
|
template renders". You do not need framework routing knowledge to see them,
|
||||||
knowledge. Measured on a real Astro site: 86 URLs collapse into 8 families,
|
but you DO need to look at the actual URL shape, because it varies:
|
||||||
and 75 of them (87%) come from just 3 dynamic `[dept]` templates.
|
|
||||||
|
| Layout | Example | Family signal |
|
||||||
|
|---|---|---|
|
||||||
|
| Nested | `/creation-site-internet/essonne-91/`, `/creation-site-internet/seine-et-marne-77/` | **shared parent path** → 25 pages, 1 family |
|
||||||
|
| **Flat** | `/lavage-auto-pomponne`, `/lavage-auto-torcy`, `/lavage-auto-chelles` | **shared slug prefix** → 8 pages, 1 family |
|
||||||
|
|
||||||
|
Both are real, measured on two live sites. First-path-segment alone handles
|
||||||
|
the nested case and **fails the flat one**: those 8 city pages read as 8
|
||||||
|
unrelated singletons, so the largest "family" becomes `/services` (5) and the
|
||||||
|
doorway-page risk — the exact thing the 30/70 rule exists to catch — is
|
||||||
|
invisible. Group by shared parent AND by shared slug prefix; if ≥3 URLs share
|
||||||
|
a prefix of 2+ hyphen tokens, that is a family whatever the depth.
|
||||||
|
|
||||||
|
Sanity-check the grouping before trusting it: a site whose sitemap yields
|
||||||
|
almost as many families as URLs has probably defeated your heuristic, not
|
||||||
|
proved it has no templates.
|
||||||
|
|
||||||
**Sample by finding class, because the classes need opposite samples:**
|
**Sample by finding class, because the classes need opposite samples:**
|
||||||
|
|
||||||
|
|||||||
@@ -1,9 +1,19 @@
|
|||||||
<?xml version="1.0" encoding="UTF-8"?>
|
<?xml version="1.0" encoding="UTF-8"?>
|
||||||
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9" xmlns:xhtml="http://www.w3.org/1999/xhtml">
|
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9"
|
||||||
|
xmlns:xhtml="http://www.w3.org/1999/xhtml"
|
||||||
|
xmlns:image="http://www.google.com/schemas/sitemap-image/1.1">
|
||||||
|
<!-- image:loc also ends with }loc — it must NOT be counted as a page -->
|
||||||
<url>
|
<url>
|
||||||
<loc>https://ex.com/</loc>
|
<loc>https://ex.com/</loc>
|
||||||
<changefreq>weekly</changefreq>
|
<changefreq>weekly</changefreq>
|
||||||
<xhtml:link rel="alternate" hreflang="en" href="https://ex.com/en/" />
|
<xhtml:link rel="alternate" hreflang="en" href="https://ex.com/en/" />
|
||||||
|
<image:image>
|
||||||
|
<image:loc>https://ex.com/img/logo.png</image:loc>
|
||||||
|
<image:title>Logo</image:title>
|
||||||
|
</image:image>
|
||||||
|
<image:image>
|
||||||
|
<image:loc>https://ex.com/img/hero.jpeg</image:loc>
|
||||||
|
</image:image>
|
||||||
</url>
|
</url>
|
||||||
<url><loc>https://ex.com/services</loc></url>
|
<url><loc>https://ex.com/services</loc></url>
|
||||||
<url><loc>https://ex.com/blog</loc></url>
|
<url><loc>https://ex.com/blog</loc></url>
|
||||||
|
|||||||
@@ -111,6 +111,12 @@ hasnt "sitemap drops non-http" "$SM" 'ftp://'
|
|||||||
hasnt "sitemap drops shell-meta" "$SM" 'bad"quote'
|
hasnt "sitemap drops shell-meta" "$SM" 'bad"quote'
|
||||||
# namespace-agnostic: real sitemaps carry sitemaps.org xmlns (+ xhtml here)
|
# namespace-agnostic: real sitemaps carry sitemaps.org xmlns (+ xhtml here)
|
||||||
has "sitemap reads namespaced" "$SM" '"https://ex.com/services"'
|
has "sitemap reads namespaced" "$SM" '"https://ex.com/services"'
|
||||||
|
# REGRESSION: <image:loc> also ends with '}loc'. An endswith test counted image
|
||||||
|
# sitemap entries as pages — a real native site returned 27 for 24 <url>, and
|
||||||
|
# img/logo.png was about to be sampled and audited as a page.
|
||||||
|
hasnt "image:loc is not a page" "$SM" '/img/logo.png'
|
||||||
|
hasnt "image:loc jpeg not a page" "$SM" '/img/hero.jpeg'
|
||||||
|
has "image ns does not inflate count" "$SM" '"count": 4'
|
||||||
|
|
||||||
IDX="$(SEO_DATA_MOCK_DIR="$SD/fixtures-sitemap-index" python3 "$SD/sitemap.py" \
|
IDX="$(SEO_DATA_MOCK_DIR="$SD/fixtures-sitemap-index" python3 "$SD/sitemap.py" \
|
||||||
--url https://ex.com/sitemap.xml)"
|
--url https://ex.com/sitemap.xml)"
|
||||||
|
|||||||
+26
-5
@@ -57,19 +57,40 @@ def _refuse_dtd(raw):
|
|||||||
if b"<!DOCTYPE" in head or b"<!ENTITY" in head:
|
if b"<!DOCTYPE" in head or b"<!ENTITY" in head:
|
||||||
raise UnsafeXML("DTD in sitemap")
|
raise UnsafeXML("DTD in sitemap")
|
||||||
|
|
||||||
|
SITEMAP_NS = "{http://www.sitemaps.org/schemas/sitemap/0.9}"
|
||||||
|
|
||||||
|
def _is_page_loc(tag):
|
||||||
|
"""A PAGE <loc>: sitemaps.org namespace, or namespace-less.
|
||||||
|
|
||||||
|
NOT <image:loc> or <video:loc>. Those live in Google's extension
|
||||||
|
namespaces and name an ASSET inside a <url>, not a page of its own. An
|
||||||
|
endswith('}loc') test matches them too — that shipped, and a real site
|
||||||
|
caught it: 24 <url> + 3 <image:loc> came back as a count of 27, so the
|
||||||
|
COVERAGE denominator was 12.5% too high and img/logo.png was about to be
|
||||||
|
sampled and audited as a page.
|
||||||
|
"""
|
||||||
|
return tag == SITEMAP_NS + "loc" or tag == "loc"
|
||||||
|
|
||||||
def _locs(raw):
|
def _locs(raw):
|
||||||
"""(<loc> texts, is_sitemapindex). Namespace-agnostic: real sitemaps carry
|
"""(page <loc> texts, is_sitemapindex).
|
||||||
the sitemaps.org xmlns and often xhtml too."""
|
|
||||||
|
Walks the DIRECT children of each <url>/<sitemap> rather than root.iter():
|
||||||
|
that alone excludes <image:image><image:loc>, and the namespace test above
|
||||||
|
is the second lock. XML comments iterate as elements with no children, so
|
||||||
|
they fall through harmlessly.
|
||||||
|
"""
|
||||||
import xml.etree.ElementTree as ET # stdlib, lazy
|
import xml.etree.ElementTree as ET # stdlib, lazy
|
||||||
_refuse_dtd(raw)
|
_refuse_dtd(raw)
|
||||||
root = ET.fromstring(raw)
|
root = ET.fromstring(raw)
|
||||||
is_index = root.tag.endswith("sitemapindex")
|
is_index = root.tag.endswith("sitemapindex")
|
||||||
out = []
|
out = []
|
||||||
for el in root.iter():
|
for entry in root: # <url> | <sitemap>
|
||||||
if el.tag.endswith("}loc") or el.tag == "loc":
|
for child in entry: # direct children only
|
||||||
text = (el.text or "").strip()
|
if _is_page_loc(child.tag):
|
||||||
|
text = (child.text or "").strip()
|
||||||
if text:
|
if text:
|
||||||
out.append(text)
|
out.append(text)
|
||||||
|
break # one <loc> per entry
|
||||||
return out, is_index
|
return out, is_index
|
||||||
|
|
||||||
def _sane(u):
|
def _sane(u):
|
||||||
|
|||||||
Reference in New Issue
Block a user