forked from bchanot/claude
Merge feature/seo-data-cherry-picks into develop
This commit is contained in:
@@ -396,6 +396,8 @@ rules:
|
||||
- BDR-068 (close-auto-persist) MERGED to develop + pushed. Then cut + pushed **v1.1.0** (minor, that feature). Standard forward bump → sonnet release-executor ran BOTH spans (prep + finish+tag); lineage continued 1.0.0→1.1.0 not 5.x (validates [[BDR-067]]). origin: main=2f8dc6b, develop=21b1e21, tags v1.0.0 + v1.1.0. WATCH-ITEM: a stale local tag `v4.0.0` reappeared during the release — NOT from origin (origin never regained it; `push.followTags` off; its commit unreachable from develop/main). Inert (push targeted main/develop/v1.1.0 explicitly + deleted the local copy; origin verified clean). Mechanism unexplained — if `v4.0.0` resurfaces locally after a `gitflow` op, trace the release lib (gitflow.sh / release-executor) for stray tag re-creation.
|
||||
|
||||
## 2026-07-17
|
||||
- content_quality verb shipped via /feat (2nd cherry-pick, stacked on feature/seo-data-cherry-picks): deterministic filler/AI-slop signal (QRG list intact, no LLM), advisory-not-verdict wired into geo STEP 8. GATE 1 CONFORME 10/10 both verbs, seo-data 190→210. Two easy claude-seo picks DONE; url_safety (DNS-rebinding) still deferred pending threat-model. Branch carries 2 feat + 1 journal commit, UNMERGED (human gate).
|
||||
- Gap-revisit claude-seo after the 21-commit build: remaining cherry-pick value narrowed to 2 clean stdlib picks + url_safety (DNS-rebinding, deferred on threat-model). schema_gen verb shipped via /feat (honors [[BDR-070]] adapt-not-copy): generates JSON-LD (Reservation/OrderAction/DiscussionForumPosting/ProfilePage), the system only audited before. GATE 1 CONFORME 10/10, seo-data 167→190 pass. content_quality next (same /feat, stacked — shares fetch.sh/test/README).
|
||||
- seo/geo parity vs github.com/AgriciDaniel/claude-seo (11.5k★, MIT): full 20-point plan built from a 3-subagent inventory, then executed. Verdict cherry-pick-never-install ([[BDR-070]]). 21 commits: Phase 1 (I1-I8 integrity, markdown specs) MERGED to develop (02c7a6f, 8 commits); Phases 2-7 on bugfix/seo-geo-integrity UNMERGED (13 commits, human gate). `fetch.sh` 5→11 verbs (richresults via inspect, sitemap, rendercheck, linkgraph, cannibal, drift, score); seo-data test suite 85→167 pass, 0 fail. Dogfooded on 2 live sites (zenquality Astro + lavageangels356 native PHP) — the second caught 2 bugs Astro hid (image:loc counted as page, flat-URL family heuristic).
|
||||
- 4 features KILLED at measurement, not built: B1/B2 (Common Crawl edges = 17.3 GB, ref impl reads 2.9% and calls it a profile — [[BDR-071]]), B3 (GSC Links API doesn't exist), W2 (Bing OAuth swamp — [[BLK-017]]). 30/70 similarity refused (needs content extraction), Playwright refused (R2 [[BDR-072]]), defusedxml refused (DTD-reject keeps stdlib-only). The most trustworthy output was the code NOT written ([[EVAL-025]]).
|
||||
- BDR-070/071/072/073 + LRN-131/132/133 + BLK-017 + EVAL-025 capitalized; checked 14 TODO done (I1-I5,W1,W3,C1-C3,B3,R2,H1,H2), W2+R1 left unchecked (deferred/rejected). 2 learnings dropped as dup of [[LRN-074]] (grep/find gitignore + detector-proof). Red thread [[LRN-133]]: an omission must stay legible. Verification discipline [[LRN-131]]/[[LRN-132]]: WebSearch ≠ verification, subagent summary = claim not fact (7 disproven, 3 self-reproduced).
|
||||
|
||||
+23
-1
@@ -551,6 +551,13 @@ sample of a 300-page site says nothing about the other 294.
|
||||
pronouns?
|
||||
8. **Lists/tables vs prose** — structured where possible?
|
||||
9. **30/70 rule** (if city/service variants exist) — ≥70% unique?
|
||||
10. **Filler/AI-slop signal (deterministic)** — feed each sampled page's
|
||||
body text to `fetch.sh content_quality`. It is a DETERMINISTIC input
|
||||
that INFORMS checks 1-9 (word-list/density heuristics, no LLM call);
|
||||
it never replaces your read of them. A low `overall_quality` or a
|
||||
`filler`/`ai-patterns` flag is a candidate for human review, not an
|
||||
automatic finding — do not let the number become the verdict, and do
|
||||
not claim a page "is AI-written" from it.
|
||||
|
||||
### Sampling command
|
||||
|
||||
@@ -561,6 +568,11 @@ for f in index.html $(find . "${FEXCL[@]}" -maxdepth 3 \( -name "*.astro" -o -na
|
||||
echo "=== $f ==="
|
||||
grep -oE '<(h1|h2|h3)[^>]*>[^<]+</(h1|h2|h3)>|^#{1,3} .+' "$f" 2>/dev/null | head -20
|
||||
done
|
||||
|
||||
# Filler/AI-slop signal (Check 10) — strip markup to plain body text, then
|
||||
# score it. Advisory only: pair the number with your own read of Checks 1-9.
|
||||
sed -e 's/<[^>]*>//g' index.html | \
|
||||
bash ~/.claude/lib/seo-data/fetch.sh content_quality
|
||||
```
|
||||
|
||||
### Findings
|
||||
@@ -576,6 +588,9 @@ CITED STATISTICS : <avg per page>
|
||||
FRESHNESS VISIBLE : <n/N pages>
|
||||
PRONOUN-HEAVY : <n/N pages flagged>
|
||||
30/70 RULE : pass | fail | N/A
|
||||
FILLER/AI-SLOP SIGNAL : <avg overall_quality>/100, flags: <n/N pages flagged>
|
||||
(deterministic, advisory — informs checks 1-9, never
|
||||
a verdict, never scored on its own)
|
||||
PRIORITY ACTIONS : <top 5>
|
||||
```
|
||||
|
||||
@@ -813,7 +828,14 @@ to act without your audit context. Embed per item:
|
||||
- **Templates + context** — G2/G6 paste the expected JSON-LD from
|
||||
`geo-schemas.md` + business context (entity name, sameAs, @id canonical)
|
||||
+ framework note. G4 follows `llms-txt-template.md` exactly. G1 pastes
|
||||
the correct variant from `ai-crawlers-2026.md`.
|
||||
the correct variant from `ai-crawlers-2026.md`. When a G2 item needs a
|
||||
`Reservation`/`OrderAction`/`DiscussionForumPosting`/`ProfilePage` block,
|
||||
generate the skeleton via `fetch.sh schema_gen
|
||||
<reservation|order|discussion|profile> [flags]`
|
||||
(`~/.claude/lib/seo-data/fetch.sh`) and fill in the real values, rather
|
||||
than hand-writing that markup. The data-integrity rule still applies on
|
||||
top of it: `schema_gen` only generates STRUCTURE — unknown field values
|
||||
stay `[À COMPLÉTER]`, never invented to fill a flag the verb needs.
|
||||
- **PERMISSIVE default** on G1 unless the client flagged premium/regulated.
|
||||
|
||||
### Output shape
|
||||
|
||||
@@ -215,6 +215,74 @@ fetch.sh score --findings <path.json | ->
|
||||
• Malformed input is an error, never a silently wrong number — unlike the
|
||||
fetch verbs, a degrade here would mean bad input, not a network fact.
|
||||
|
||||
fetch.sh schema_gen <reservation|order|discussion|profile> [flags] [--script-tag]
|
||||
→ {"status":"ok","source":"schema_gen","type":"<@type>","jsonld":{…}}
|
||||
→ {"status":"error","reason":"bad_usage"} # a REQUIRED flag omitted
|
||||
→ {"status":"degraded","reason":"…"} # a required flag given, empty
|
||||
|
||||
fetch.sh schema_gen reservation --provider "Marea NYC" \
|
||||
--start 2026-06-04T19:30:00-04:00 --party-size 4
|
||||
fetch.sh schema_gen order --merchant "Acme Pizza" --order-url https://acme.example/order
|
||||
fetch.sh schema_gen discussion --headline "…" --author "Sara Park" \
|
||||
--url https://forum.example.com/t/123 --date 2026-05-12T14:00:00Z
|
||||
fetch.sh schema_gen profile --name "Daniel Agrici" --url https://agricidaniel.com/about \
|
||||
--same-as https://github.com/AgriciDaniel --knows-about "SEO" "Schema markup"
|
||||
|
||||
Adapted from claude-seo's `schema_generate.py` (MIT) into this contract.
|
||||
Our system only AUDITS existing markup elsewhere; this is the one verb
|
||||
that GENERATES it — deterministic JSON-LD skeletons for the four v2
|
||||
high-leverage Schema.org types, so geo-analyzer's G2 batch stops
|
||||
hand-writing markup by hand. It only generates STRUCTURE: unknown field
|
||||
VALUES are the caller's job, `[À COMPLÉTER]` for anything unconfirmed —
|
||||
this verb never invents a sameAs, an email, or a business name.
|
||||
• Stdlib only, no network, no auth — runs even without the venv.
|
||||
• `--script-tag` wraps the cleaned jsonld in
|
||||
`<script type="application/ld+json">…</script>` under a `script` key,
|
||||
still inside the `ok` envelope. It must be given AFTER the type
|
||||
(`schema_gen reservation … --script-tag`, not before) — argparse
|
||||
subcommand flags only parse after their subcommand.
|
||||
• Never emits a JSON `null`: fields left unset are omitted from the
|
||||
`jsonld` object entirely rather than serialised as `null`.
|
||||
• A REQUIRED flag omitted → `{"status":"error","reason":"bad_usage"}`,
|
||||
exit 2 (bad usage, like every other verb). A required flag GIVEN but
|
||||
empty (argparse cannot catch that) → `{"status":"degraded",...}`,
|
||||
exit 0 — fail-open, never a traceback.
|
||||
|
||||
fetch.sh content_quality [--file <path.txt>] < text_on_stdin
|
||||
→ {"status":"ok","source":"content_quality","filler_score":0,"ai_pattern_score":0,
|
||||
"information_density":1.0,"overall_quality":90,"flags":[],
|
||||
"matches":{"filler":[],"ai_patterns":[]}}
|
||||
→ {"status":"degraded","reason":"empty_input"|"<file error>"}
|
||||
|
||||
fetch.sh content_quality --file article.txt
|
||||
printf '%s' "$BODY_TEXT" | fetch.sh content_quality
|
||||
|
||||
Adapted from claude-seo's `content_quality.py` (MIT) into this contract.
|
||||
100% deterministic — regex/word-lists (QRG §4.6 filler phrases + a
|
||||
Wikipedia "AI Cleanup" catalogue of LLM-typical phrasings, CC BY-SA 4.0),
|
||||
no LLM call, no network. Reads the text to score from `--file <path>` or,
|
||||
when `--file` is `-` or omitted, from stdin — the same idiom `score.py`
|
||||
uses for `--findings`.
|
||||
• **ADVISORY, NOT A VERDICT.** The output never claims "this text is
|
||||
AI-written" — modern generative tools can pass every heuristic here,
|
||||
and human writers use some of these phrases too. `flags` are
|
||||
candidates for HUMAN REVIEW, never an automatic finding. geo-analyzer
|
||||
STEP 8 (Content Shape for AI) treats `overall_quality`/`flags` as ONE
|
||||
measured input that INFORMS the axis; the axis itself stays an LLM
|
||||
judgement (30/70, Definition Lead), never replaced by this score.
|
||||
• `filler_score`/`ai_pattern_score` (0-100, higher = worse) count
|
||||
phrase-list hits scaled per 1000 tokens; `information_density`
|
||||
(0.0-1.0) is entities + numbers per 100 tokens; `overall_quality`
|
||||
(0-100, higher is better) is the weighted composite (also folds in a
|
||||
bigram-repetition penalty even though that score isn't itself a
|
||||
top-level field). `flags` fires at fixed thresholds: `filler`,
|
||||
`ai-patterns`, `low-density`, `repetitive`.
|
||||
• Stdlib only (argparse/json/re/sys/collections/typing) — runs even
|
||||
without the venv. Empty/whitespace-only input degrades rather than
|
||||
returning a false zero-value "ok": an empty analysis is not a result.
|
||||
• This is filler/AI-pattern SHAPE, not fact-checking — a text can be
|
||||
dense and well-cited yet still wrong; that stays a human/LLM call.
|
||||
|
||||
fetch.sh drift --url https://ex.com/sitemap.xml [--max 500]
|
||||
→ {"status":"ok","baseline":true,"captured":"…","pages":24,"store":"…"}
|
||||
→ {"status":"ok","baseline":false,"since":"…","gone":[…],"new":[…],
|
||||
|
||||
@@ -0,0 +1,242 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Deterministic filler / AI-slop content-quality scorer. Stdlib only.
|
||||
|
||||
Adapted from claude-seo (github.com/AgriciDaniel/claude-seo, MIT),
|
||||
content_quality.py — rewritten to the lib/seo-data fail-open contract.
|
||||
|
||||
Scores a block of text against three regex/word-list heuristics: padding
|
||||
"filler" phrases (QRG §4.6), LLM-typical phrasings ("AI-pattern" list),
|
||||
and a measured information density (entities + numbers per token). 100%
|
||||
deterministic — no LLM call, no network.
|
||||
|
||||
ADVISORY, NOT A VERDICT. This never claims "this text is AI-written" —
|
||||
modern generative tools can pass every heuristic here, and human writers
|
||||
use some of these phrases too. A low overall_quality or a filler/
|
||||
ai-patterns flag is a candidate for human review, nothing more. In
|
||||
geo-analyzer's STEP 8 (Content Shape for AI) it is ONE measured input
|
||||
that INFORMS the axis, which stays an LLM judgement (30/70, Definition
|
||||
Lead) — never a replacement for it, and never auto-filed as a finding on
|
||||
its own.
|
||||
|
||||
Attribution: the AI-pattern list draws from the Wikipedia "AI Cleanup"
|
||||
project's catalogue of LLM-typical phrasings (CC BY-SA 4.0), the same
|
||||
list claude-seo cites.
|
||||
|
||||
Envelope (see `_cli`)::
|
||||
|
||||
{"status": "ok", "source": "content_quality",
|
||||
"filler_score": 0..100, # higher = more filler-like
|
||||
"ai_pattern_score": 0..100, # higher = more AI-pattern hits
|
||||
"information_density": 0.0..1.0,
|
||||
"overall_quality": 0..100, # composite, higher is better
|
||||
"flags": ["filler", "ai-patterns", "low-density", "repetitive"],
|
||||
"matches": {"filler": [...], "ai_patterns": [...]}}
|
||||
{"status": "degraded", "reason": "empty_input" | "<why>"}
|
||||
"""
|
||||
import argparse, json, re, sys
|
||||
from collections import Counter
|
||||
from typing import Iterable
|
||||
|
||||
# Padding / filler phrases QRG §4.6 flags as "little-to-no value". The
|
||||
# lists are the value of this module — kept intact from the source, not
|
||||
# trimmed.
|
||||
_FILLER_PHRASES = (
|
||||
"it's important to note that",
|
||||
"in this article, we'll explore",
|
||||
"in this article we will explore",
|
||||
"in today's fast-paced world",
|
||||
"in today's digital age",
|
||||
"in today's competitive landscape",
|
||||
"needless to say",
|
||||
"at the end of the day",
|
||||
"when it comes to",
|
||||
"when all is said and done",
|
||||
"in the realm of",
|
||||
"in the world of",
|
||||
"the bottom line is",
|
||||
"without further ado",
|
||||
"first and foremost",
|
||||
"last but not least",
|
||||
"for what it's worth",
|
||||
"it goes without saying",
|
||||
"as we all know",
|
||||
"the truth is that",
|
||||
"the fact of the matter is",
|
||||
"more often than not",
|
||||
"let's dive in",
|
||||
"let's dive into",
|
||||
"let's take a closer look",
|
||||
"let's take a deeper look",
|
||||
)
|
||||
|
||||
# LLM-typical phrasings (Wikipedia AI Cleanup catalogue, CC BY-SA 4.0;
|
||||
# also used by claude-seo, MIT). Conservative: only phrases that
|
||||
# disproportionately appear in LLM output. Adding to this list should
|
||||
# require corpus evidence, not intuition.
|
||||
_AI_PATTERNS = (
|
||||
"delve into",
|
||||
"delve deeper into",
|
||||
"in the ever-evolving",
|
||||
"ever-evolving landscape",
|
||||
"ever-changing landscape",
|
||||
"in the dynamic landscape",
|
||||
"navigating the",
|
||||
"navigate the complexities",
|
||||
"tapestry of",
|
||||
"rich tapestry",
|
||||
"intricate tapestry",
|
||||
"embark on a journey",
|
||||
"embarking on this",
|
||||
"a testament to",
|
||||
"a beacon of",
|
||||
"the cornerstone of",
|
||||
"a cornerstone of",
|
||||
"at the heart of",
|
||||
"at its core",
|
||||
"in essence,",
|
||||
"in conclusion,",
|
||||
"ultimately,",
|
||||
"moreover,",
|
||||
"furthermore,",
|
||||
"however, it's worth noting",
|
||||
"it's worth noting that",
|
||||
"by leveraging",
|
||||
"leverage the power of",
|
||||
"leveraging the power of",
|
||||
"harness the power of",
|
||||
"unlock the potential",
|
||||
"unlock the full potential",
|
||||
"the realm of possibilities",
|
||||
"open up a world of",
|
||||
"a world of possibilities",
|
||||
"elevate your",
|
||||
"transform your",
|
||||
"revolutionize the way",
|
||||
"game-changer",
|
||||
"game-changing",
|
||||
"cutting-edge",
|
||||
"state-of-the-art",
|
||||
"in summary,",
|
||||
"to summarize,",
|
||||
"to put it simply,",
|
||||
"in a nutshell,",
|
||||
)
|
||||
|
||||
_TOKEN_RE = re.compile(r"[A-Za-z][A-Za-z'\-]*")
|
||||
_NUMBER_RE = re.compile(r"\b\d+(?:[.,]\d+)?(?:%|st|nd|rd|th)?\b")
|
||||
# Capitalised multi-word names: rough proper-noun heuristic. Two or more
|
||||
# capitalised tokens in a row count as one entity.
|
||||
_ENTITY_RE = re.compile(r"\b(?:[A-Z][a-z]+(?:\s+[A-Z][a-z]+)+)\b")
|
||||
|
||||
|
||||
def _count_phrase_hits(text: str, patterns: Iterable[str]) -> list:
|
||||
"""Patterns that appear at least once in text (case-insensitive)."""
|
||||
lowered = text.lower()
|
||||
return [p for p in patterns if p in lowered]
|
||||
|
||||
|
||||
def _repetition_score(tokens):
|
||||
"""Bigram repetition: fraction of bigrams that recur more than once."""
|
||||
if len(tokens) < 4:
|
||||
return 0.0
|
||||
bigrams = [tokens[i] + " " + tokens[i + 1] for i in range(len(tokens) - 1)]
|
||||
counts = Counter(bigrams)
|
||||
repeated = sum(1 for v in counts.values() if v > 1)
|
||||
return repeated / max(1, len(counts))
|
||||
|
||||
|
||||
def analyse(text):
|
||||
"""Score text against the filler / AI-pattern / density / repetition
|
||||
heuristics. Advisory only — see module docstring."""
|
||||
tokens = [t.lower() for t in _TOKEN_RE.findall(text)]
|
||||
n_tokens = len(tokens)
|
||||
|
||||
filler_hits = _count_phrase_hits(text, _FILLER_PHRASES)
|
||||
ai_hits = _count_phrase_hits(text, _AI_PATTERNS)
|
||||
|
||||
# Density: entities + numbers per 100 tokens. A high-density article
|
||||
# (case studies, data journalism) lands at ~5+; generic filler <2.
|
||||
entities = len(_ENTITY_RE.findall(text))
|
||||
numbers = len(_NUMBER_RE.findall(text))
|
||||
density_per_100 = (entities + numbers) * 100.0 / max(1, n_tokens)
|
||||
information_density = min(1.0, density_per_100 / 10.0)
|
||||
|
||||
rep_score = int(round(_repetition_score(tokens) * 100))
|
||||
|
||||
# Scale to per-1000 tokens so the score is comparable across lengths.
|
||||
scale = max(1.0, n_tokens / 1000.0)
|
||||
filler_score = min(100, int(round(len(filler_hits) / scale * 25)))
|
||||
ai_pattern_score = min(100, int(round(len(ai_hits) / scale * 15)))
|
||||
|
||||
flags = []
|
||||
if filler_score >= 50:
|
||||
flags.append("filler")
|
||||
if ai_pattern_score >= 40:
|
||||
flags.append("ai-patterns")
|
||||
if information_density < 0.20:
|
||||
flags.append("low-density")
|
||||
if rep_score >= 30:
|
||||
flags.append("repetitive")
|
||||
|
||||
# Composite: invert penalty signals, weight by impact. Same weights
|
||||
# as the source — the length bonus caps at 1000 tokens.
|
||||
overall = (
|
||||
(100 - filler_score) * 0.25
|
||||
+ (100 - ai_pattern_score) * 0.25
|
||||
+ information_density * 100 * 0.25
|
||||
+ (100 - rep_score) * 0.15
|
||||
+ min(100, n_tokens / 10.0) * 0.10
|
||||
)
|
||||
|
||||
return {
|
||||
"filler_score": filler_score,
|
||||
"ai_pattern_score": ai_pattern_score,
|
||||
"information_density": round(information_density, 3),
|
||||
"overall_quality": int(round(overall)),
|
||||
"flags": flags,
|
||||
"matches": {"filler": filler_hits, "ai_patterns": ai_hits},
|
||||
}
|
||||
|
||||
|
||||
def _build_parser():
|
||||
p = argparse.ArgumentParser(
|
||||
description="Deterministic filler / AI-slop content-quality scorer."
|
||||
)
|
||||
p.add_argument("--store", default=None) # accepted+ignored (dispatch)
|
||||
p.add_argument(
|
||||
"--file", default="-",
|
||||
help="Path to a text file, or - for stdin (default -).",
|
||||
)
|
||||
return p
|
||||
|
||||
|
||||
def _read_input(path):
|
||||
"""Read the analysis target from stdin ('-'/omitted) or a plain file.
|
||||
Plain `open()` only — no pathlib, to stay stdlib-minimal per contract."""
|
||||
if path in (None, "-"):
|
||||
return sys.stdin.read()
|
||||
return open(path, encoding="utf-8", errors="replace").read()
|
||||
|
||||
|
||||
def _cli():
|
||||
try:
|
||||
args = _build_parser().parse_args()
|
||||
text = _read_input(args.file)
|
||||
if not text or not text.strip():
|
||||
print(json.dumps({"status": "degraded", "reason": "empty_input"}))
|
||||
return
|
||||
envelope = {"status": "ok", "source": "content_quality"}
|
||||
envelope.update(analyse(text))
|
||||
print(json.dumps(envelope, indent=2))
|
||||
except SystemExit as e:
|
||||
if e.code not in (0, None):
|
||||
print(json.dumps({"status": "error", "reason": "bad_usage"}))
|
||||
raise
|
||||
except Exception as e:
|
||||
# Fail-open: a missing --file, an unreadable/binary file, or any
|
||||
# other unexpected error degrades rather than crashing the caller.
|
||||
print(json.dumps({"status": "degraded", "reason": str(e)}))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
_cli()
|
||||
@@ -34,6 +34,10 @@ case "$cmd" in
|
||||
exec "$PY" "$HERE/sitemap.py" --store "$STORE" "$@" ;;
|
||||
score)
|
||||
exec "$PY" "$HERE/score.py" --store "$STORE" "$@" ;;
|
||||
schema_gen)
|
||||
exec "$PY" "$HERE/schema_gen.py" --store "$STORE" "$@" ;;
|
||||
content_quality)
|
||||
exec "$PY" "$HERE/content_quality.py" --store "$STORE" "$@" ;;
|
||||
drift)
|
||||
exec "$PY" "$HERE/drift.py" --store "$STORE" "$@" ;;
|
||||
rendercheck)
|
||||
@@ -52,6 +56,6 @@ case "$cmd" in
|
||||
fi
|
||||
echo '{"status":"error","reason":"usage: fetch.sh forget {--label <label>|--all} (label charset: A-Za-z0-9._-)"}'
|
||||
exit 2 ;;
|
||||
*) echo '{"status":"error","reason":"usage: fetch.sh {accounts|crux|queries|inspect|cannibal|sitemap|rendercheck|linkgraph|drift|score|forget} [flags]"}'
|
||||
*) echo '{"status":"error","reason":"usage: fetch.sh {accounts|crux|queries|inspect|cannibal|sitemap|rendercheck|linkgraph|drift|score|schema_gen|content_quality|forget} [flags]"}'
|
||||
exit 2 ;;
|
||||
esac
|
||||
|
||||
@@ -0,0 +1,301 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Deterministic JSON-LD generators for four Schema.org types. Stdlib only.
|
||||
|
||||
Adapted from claude-seo (github.com/AgriciDaniel/claude-seo, MIT),
|
||||
schema_generate.py — rewritten to the lib/seo-data fail-open contract.
|
||||
|
||||
Everywhere else in this repo we AUDIT existing markup (google_seo.py
|
||||
`inspect`, geo-analyzer's JSON-LD rules); this is the one verb that
|
||||
GENERATES it. Reservation + potentialAction matter now that AI Mode
|
||||
executes restaurant reservations; DiscussionForumPosting is a live SERP
|
||||
feature; ProfilePage with sameAs/knowsAbout is the cheapest entity-graph
|
||||
builder for AI citation correlation. geo-analyzer's G2 batch calls this
|
||||
instead of hand-writing the markup — it only generates STRUCTURE, unknown
|
||||
field VALUES stay the caller's `[À COMPLÉTER]` placeholder, never invented
|
||||
here.
|
||||
"""
|
||||
import argparse, json
|
||||
|
||||
|
||||
def reservation(provider, start, *, end=None, party_size=None,
|
||||
reservation_id=None, reservation_for_name=None,
|
||||
customer_name=None, customer_email=None,
|
||||
kind="FoodEstablishmentReservation"):
|
||||
"""Reservation JSON-LD block. Defaults to FoodEstablishment."""
|
||||
payload = {
|
||||
"@context": "https://schema.org",
|
||||
"@type": kind,
|
||||
"reservationStatus": "https://schema.org/ReservationConfirmed",
|
||||
"provider": {"@type": "Organization", "name": provider},
|
||||
"reservationFor": {
|
||||
"@type": "FoodEstablishment"
|
||||
if kind == "FoodEstablishmentReservation" else "Place",
|
||||
"name": reservation_for_name or provider,
|
||||
},
|
||||
"startTime": start,
|
||||
"endTime": end,
|
||||
"partySize": party_size,
|
||||
"reservationId": reservation_id,
|
||||
}
|
||||
if customer_name or customer_email:
|
||||
payload["underName"] = {"@type": "Person", "name": customer_name,
|
||||
"email": customer_email}
|
||||
return payload
|
||||
|
||||
|
||||
def order_action(merchant, *, order_url, name="Order online",
|
||||
accepted_payment_method=None, delivery_method=None):
|
||||
"""OrderAction potentialAction block. Attach to a Product/Service via
|
||||
{"@type": "Product", "potentialAction": <this dict>}."""
|
||||
payload = {
|
||||
"@context": "https://schema.org",
|
||||
"@type": "OrderAction",
|
||||
"name": name,
|
||||
"target": {
|
||||
"@type": "EntryPoint",
|
||||
"urlTemplate": order_url,
|
||||
"inLanguage": "en-US",
|
||||
"actionPlatform": [
|
||||
"https://schema.org/DesktopWebPlatform",
|
||||
"https://schema.org/MobileWebPlatform",
|
||||
],
|
||||
},
|
||||
"deliveryMethod": delivery_method or [
|
||||
"https://schema.org/OnSitePickup",
|
||||
"https://schema.org/ParcelService",
|
||||
],
|
||||
"priceSpecification": {
|
||||
"@type": "PriceSpecification",
|
||||
"eligibleTransactionVolume": {
|
||||
"@type": "PriceSpecification",
|
||||
"minPrice": 0,
|
||||
"priceCurrency": "USD",
|
||||
},
|
||||
},
|
||||
"merchant": {"@type": "Organization", "name": merchant},
|
||||
}
|
||||
if accepted_payment_method:
|
||||
payload["acceptedPaymentMethod"] = [
|
||||
{"@type": "PaymentMethod", "name": m}
|
||||
for m in accepted_payment_method
|
||||
]
|
||||
return payload
|
||||
|
||||
|
||||
def discussion(headline, author, *, url, date_published, text=None,
|
||||
date_modified=None, interaction_count=None,
|
||||
comment_count=None):
|
||||
"""DiscussionForumPosting JSON-LD block."""
|
||||
payload = {
|
||||
"@context": "https://schema.org",
|
||||
"@type": "DiscussionForumPosting",
|
||||
"headline": headline,
|
||||
"author": {"@type": "Person", "name": author},
|
||||
"datePublished": date_published,
|
||||
"dateModified": date_modified,
|
||||
"url": url,
|
||||
"mainEntityOfPage": {"@type": "WebPage", "@id": url},
|
||||
"text": text,
|
||||
"commentCount": comment_count,
|
||||
}
|
||||
if interaction_count:
|
||||
payload["interactionStatistic"] = [
|
||||
{"@type": "InteractionCounter",
|
||||
"interactionType": "https://schema.org/%s" % k,
|
||||
"userInteractionCount": v}
|
||||
for k, v in interaction_count.items()
|
||||
]
|
||||
return payload
|
||||
|
||||
|
||||
def profile(name, *, url, description=None, same_as=None, knows_about=None,
|
||||
works_for=None, image=None, job_title=None):
|
||||
"""ProfilePage JSON-LD block. sameAs + knowsAbout is the entity-graph
|
||||
helper for AI citation correlation — Wikipedia/GitHub/LinkedIn/ORCID
|
||||
URLs in sameAs disambiguate the person across knowledge graphs."""
|
||||
person = {
|
||||
"@type": "Person",
|
||||
"name": name,
|
||||
"url": url,
|
||||
"description": description,
|
||||
"sameAs": list(same_as) if same_as else None,
|
||||
"knowsAbout": list(knows_about) if knows_about else None,
|
||||
"worksFor": {"@type": "Organization", "name": works_for}
|
||||
if works_for else None,
|
||||
"image": image,
|
||||
"jobTitle": job_title,
|
||||
}
|
||||
return {"@context": "https://schema.org", "@type": "ProfilePage",
|
||||
"mainEntity": person, "url": url}
|
||||
|
||||
|
||||
def _strip_nones(value):
|
||||
"""Recursively drop dict keys AND list elements whose value is None —
|
||||
the emitted JSON-LD must never contain a null."""
|
||||
if isinstance(value, dict):
|
||||
return {k: _strip_nones(v) for k, v in value.items() if v is not None}
|
||||
if isinstance(value, list):
|
||||
return [_strip_nones(v) for v in value if v is not None]
|
||||
return value
|
||||
|
||||
|
||||
def _need(value, field):
|
||||
"""Raise on a schema-required field that is present but empty — the
|
||||
case argparse's `required=True` cannot catch (an empty string is a
|
||||
given flag, not a missing one)."""
|
||||
if value is None or not str(value).strip():
|
||||
raise ValueError("missing required field: %s" % field)
|
||||
return value
|
||||
|
||||
|
||||
def _generate(kind, args):
|
||||
"""Route to the matching generator, enforcing schema-required fields."""
|
||||
if kind == "reservation":
|
||||
return reservation(
|
||||
_need(args.provider, "provider"), _need(args.start, "start"),
|
||||
end=args.end, party_size=args.party_size,
|
||||
reservation_id=args.reservation_id,
|
||||
reservation_for_name=args.reservation_for_name,
|
||||
customer_name=args.customer_name,
|
||||
customer_email=args.customer_email, kind=args.reservation_kind,
|
||||
)
|
||||
if kind == "order":
|
||||
return order_action(
|
||||
_need(args.merchant, "merchant"),
|
||||
order_url=_need(args.order_url, "order_url"), name=args.name,
|
||||
accepted_payment_method=args.accepted_payment_method,
|
||||
delivery_method=args.delivery_method,
|
||||
)
|
||||
if kind == "discussion":
|
||||
interaction = {"LikeAction": args.likes} if args.likes else None
|
||||
return discussion(
|
||||
_need(args.headline, "headline"), _need(args.author, "author"),
|
||||
url=_need(args.url, "url"),
|
||||
date_published=_need(args.date_published, "date_published"),
|
||||
text=args.text, date_modified=args.date_modified,
|
||||
interaction_count=interaction, comment_count=args.comment_count,
|
||||
)
|
||||
if kind == "profile":
|
||||
return profile(
|
||||
_need(args.name, "name"), url=_need(args.url, "url"),
|
||||
description=args.description, same_as=args.same_as,
|
||||
knows_about=args.knows_about, works_for=args.works_for,
|
||||
image=args.image, job_title=args.job_title,
|
||||
)
|
||||
raise ValueError("unknown kind: %r" % kind) # pragma: no cover — argparse
|
||||
|
||||
|
||||
def _envelope(payload, script_tag):
|
||||
cleaned = _strip_nones(payload)
|
||||
out = {"status": "ok", "source": "schema_gen",
|
||||
"type": cleaned.get("@type"), "jsonld": cleaned}
|
||||
if script_tag:
|
||||
pretty = json.dumps(cleaned, indent=2, ensure_ascii=False)
|
||||
out["script"] = ('<script type="application/ld+json">\n%s\n</script>'
|
||||
% pretty)
|
||||
return out
|
||||
|
||||
|
||||
def _script_tag_parent():
|
||||
"""`--script-tag` as a shared parent parser, so it is valid on every
|
||||
subcommand — `fetch.sh schema_gen <type> [flags]` puts the type FIRST,
|
||||
and argparse only accepts a flag after a subcommand token if that flag
|
||||
was declared on the subparser, not the top-level one."""
|
||||
parent = argparse.ArgumentParser(add_help=False)
|
||||
parent.add_argument(
|
||||
"--script-tag", action="store_true",
|
||||
help="Wrap jsonld in <script type=application/ld+json>.",
|
||||
)
|
||||
return parent
|
||||
|
||||
|
||||
def _add_reservation_args(sub, parents):
|
||||
p = sub.add_parser("reservation", parents=parents,
|
||||
help="FoodEstablishmentReservation et al.")
|
||||
p.add_argument("--provider", required=True)
|
||||
p.add_argument("--start", required=True, help="ISO 8601 startTime.")
|
||||
p.add_argument("--end")
|
||||
p.add_argument("--party-size", type=int)
|
||||
p.add_argument("--reservation-id")
|
||||
p.add_argument("--reservation-for-name")
|
||||
p.add_argument("--customer-name")
|
||||
p.add_argument("--customer-email")
|
||||
p.add_argument(
|
||||
"--reservation-kind", dest="reservation_kind",
|
||||
default="FoodEstablishmentReservation",
|
||||
choices=(
|
||||
"FoodEstablishmentReservation", "LodgingReservation",
|
||||
"RentalCarReservation", "TaxiReservation", "EventReservation",
|
||||
"TrainReservation", "FlightReservation",
|
||||
),
|
||||
)
|
||||
|
||||
|
||||
def _add_order_args(sub, parents):
|
||||
p = sub.add_parser("order", parents=parents,
|
||||
help="OrderAction (potentialAction).")
|
||||
p.add_argument("--merchant", required=True)
|
||||
p.add_argument("--order-url", required=True)
|
||||
p.add_argument("--name", default="Order online")
|
||||
p.add_argument("--accepted-payment-method", nargs="*", default=None)
|
||||
p.add_argument("--delivery-method", nargs="*", default=None)
|
||||
|
||||
|
||||
def _add_discussion_args(sub, parents):
|
||||
p = sub.add_parser("discussion", parents=parents,
|
||||
help="DiscussionForumPosting.")
|
||||
p.add_argument("--headline", required=True)
|
||||
p.add_argument("--author", required=True)
|
||||
p.add_argument("--url", required=True)
|
||||
p.add_argument("--date", dest="date_published", required=True)
|
||||
p.add_argument("--text")
|
||||
p.add_argument("--date-modified")
|
||||
p.add_argument("--comment-count", type=int)
|
||||
p.add_argument("--likes", type=int, default=None,
|
||||
help="LikeAction count (interactionStatistic).")
|
||||
|
||||
|
||||
def _add_profile_args(sub, parents):
|
||||
p = sub.add_parser("profile", parents=parents,
|
||||
help="ProfilePage with sameAs / knowsAbout.")
|
||||
p.add_argument("--name", required=True)
|
||||
p.add_argument("--url", required=True)
|
||||
p.add_argument("--description")
|
||||
p.add_argument("--same-as", nargs="*", default=None)
|
||||
p.add_argument("--knows-about", nargs="*", default=None)
|
||||
p.add_argument("--works-for")
|
||||
p.add_argument("--image")
|
||||
p.add_argument("--job-title")
|
||||
|
||||
|
||||
def _build_parser():
|
||||
p = argparse.ArgumentParser(
|
||||
description="Schema.org JSON-LD generators (stdlib, deterministic)."
|
||||
)
|
||||
p.add_argument("--store", default=None) # accepted+ignored (dispatch)
|
||||
sub = p.add_subparsers(dest="kind", required=True)
|
||||
parents = [_script_tag_parent()]
|
||||
_add_reservation_args(sub, parents)
|
||||
_add_order_args(sub, parents)
|
||||
_add_discussion_args(sub, parents)
|
||||
_add_profile_args(sub, parents)
|
||||
return p
|
||||
|
||||
|
||||
def _cli():
|
||||
try:
|
||||
args = _build_parser().parse_args()
|
||||
payload = _generate(args.kind, args)
|
||||
print(json.dumps(_envelope(payload, args.script_tag), indent=2))
|
||||
except SystemExit as e:
|
||||
if e.code not in (0, None):
|
||||
print(json.dumps({"status": "error", "reason": "bad_usage"}))
|
||||
raise
|
||||
except Exception as e:
|
||||
# Fail-open: a missing required field or any other unexpected error
|
||||
# is a normal outcome here, never a traceback or empty stdout.
|
||||
print(json.dumps({"status": "degraded", "reason": str(e)}))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
_cli()
|
||||
@@ -233,6 +233,136 @@ NCHG="$(printf '%s' "$D2" | python3 -c 'import sys,json; print(len(json.load(sys
|
||||
|| no "reworded title is a change, not a regression" "got $NCHG"
|
||||
rm -rf "$DH"
|
||||
|
||||
echo "── schema_gen ──"
|
||||
SG() { python3 "$SD/schema_gen.py" "$@"; }
|
||||
RES="$(SG reservation --provider "Chez X" --start "2026-08-01T19:00")"
|
||||
has "reservation ok" "$RES" '"status": "ok"'
|
||||
has "reservation type surfaced" "$RES" '"type": "FoodEstablishmentReservation"'
|
||||
has "jsonld has @context" "$RES" '"@context": "https://schema.org"'
|
||||
has "reservation keeps provider" "$RES" 'Chez X'
|
||||
has "reservation keeps start" "$RES" '2026-08-01T19:00'
|
||||
PROF="$(SG profile --name "Jane Doe" --url https://ex.com/about)"
|
||||
has "profile ok" "$PROF" '"status": "ok"'
|
||||
has "profile type surfaced" "$PROF" '"type": "ProfilePage"'
|
||||
ORD="$(SG order --merchant "Acme" --order-url https://ex.com/order)"
|
||||
has "order ok" "$ORD" '"status": "ok"'
|
||||
has "order type surfaced" "$ORD" '"type": "OrderAction"'
|
||||
DISC="$(SG discussion --headline "Q" --author "Jo" --url https://ex.com/t/1 \
|
||||
--date 2026-05-01T00:00:00Z)"
|
||||
has "discussion ok" "$DISC" '"status": "ok"'
|
||||
has "discussion type surfaced" "$DISC" '"type": "DiscussionForumPosting"'
|
||||
# argparse required=True catches an OMITTED flag → bad usage, exit 2
|
||||
BADRES="$(SG reservation --start 2026-08-01T19:00 2>/dev/null)"; BADRC=$?
|
||||
hasnt "missing --provider is not ok" "$BADRES" '"status": "ok"'
|
||||
[ "$BADRC" = "2" ] && ok "missing --provider exit 2" \
|
||||
|| no "missing --provider exit 2" "got $BADRC"
|
||||
# a required field argparse ALLOWS through (flag given, value empty) must
|
||||
# still fail open — degraded, not a crash, exit 0
|
||||
EMPTYRES="$(SG reservation --provider "" --start 2026-08-01T19:00)"; EMPTYRC=$?
|
||||
hasnt "empty --provider is not ok" "$EMPTYRES" '"status": "ok"'
|
||||
has "empty --provider degrades" "$EMPTYRES" '"status": "degraded"'
|
||||
[ "$EMPTYRC" = "0" ] && ok "empty --provider exit 0" \
|
||||
|| no "empty --provider exit 0" "got $EMPTYRC"
|
||||
# --script-tag must work AFTER the type, matching `fetch.sh schema_gen
|
||||
# <type> [flags]` — the shape the dispatcher actually calls it with. The
|
||||
# envelope is JSON, so the `script` field's own quotes are backslash-escaped
|
||||
# in the raw stdout — decode it to check the LITERAL wrapper string.
|
||||
SCRIPT="$(SG profile --name "Jane Doe" --url https://ex.com/about --script-tag)"
|
||||
SCRIPT_TAG="$(printf '%s' "$SCRIPT" | \
|
||||
python3 -c 'import sys,json; print(json.load(sys.stdin)["script"])')"
|
||||
has "script-tag wraps output" "$SCRIPT_TAG" '<script type="application/ld+json">'
|
||||
# an omitted optional field must never surface as a JSON null
|
||||
hasnt "no null ever emitted" "$RES" 'null'
|
||||
# stdlib ONLY — no requests/httpx/bs4/any third-party import
|
||||
IMPORTS="$(grep -E '^(import|from) ' "$SD/schema_gen.py")"
|
||||
if printf '%s' "$IMPORTS" | grep -qiE 'requests|httpx|bs4'; then
|
||||
no "schema_gen stdlib only" "third-party import found: $IMPORTS"
|
||||
else
|
||||
ok "schema_gen stdlib only"
|
||||
fi
|
||||
# dispatch wiring: --store precedes the type (fetch.sh's own convention),
|
||||
# --script-tag comes after it (the caller's convention) — both must work
|
||||
# through the real fetch.sh entrypoint, not just the bare script
|
||||
FSG="$(SEO_DATA_ENV_FILE=/dev/null SEO_DATA_STORE=/nonexistent bash "$SD/fetch.sh" \
|
||||
schema_gen reservation --provider "Chez X" --start 2026-08-01T19:00 --script-tag)"
|
||||
has "fetch dispatches schema_gen" "$FSG" '"status": "ok"'
|
||||
FSG_TAG="$(printf '%s' "$FSG" | \
|
||||
python3 -c 'import sys,json; print(json.load(sys.stdin)["script"])')"
|
||||
has "fetch schema_gen script-tag" "$FSG_TAG" '<script type="application/ld+json">'
|
||||
|
||||
echo "── content_quality ──"
|
||||
CQ() { python3 "$SD/content_quality.py" "$@"; }
|
||||
# feed the phrase list's OWN entries so the match is exact, not paraphrased —
|
||||
# a detector proven only on the maintainer's paraphrase proves nothing
|
||||
FILLER_TXT="In today's fast-paced world, it's important to note that this \
|
||||
article will delve into the ever-evolving landscape of technology. Let's \
|
||||
dive in and navigate the complexities together, leveraging the power of \
|
||||
innovation to unlock the potential of your business. Ultimately, this \
|
||||
cutting-edge, state-of-the-art approach is a testament to progress. \
|
||||
Moreover, furthermore, in conclusion, transform your outcomes today."
|
||||
CLEAN_TXT="The 2024 ADEME report found French households spent 2,137 EUR \
|
||||
on heating, up 12% from 2021."
|
||||
FILLER_OUT="$(printf '%s' "$FILLER_TXT" | CQ)"
|
||||
CLEAN_OUT="$(printf '%s' "$CLEAN_TXT" | CQ)"
|
||||
has "filler text is ok" "$FILLER_OUT" '"status": "ok"'
|
||||
has "clean text is ok" "$CLEAN_OUT" '"status": "ok"'
|
||||
# flags is a JSON array — extract it in isolation so the check can't be
|
||||
# fooled by the always-present "matches": {"filler": [...]} key sharing
|
||||
# the same quoted word
|
||||
FILLER_FLAGS="$(printf '%s' "$FILLER_OUT" | \
|
||||
python3 -c 'import sys,json; print(",".join(json.load(sys.stdin)["flags"]))')"
|
||||
CLEAN_FLAGS="$(printf '%s' "$CLEAN_OUT" | \
|
||||
python3 -c 'import sys,json; print(",".join(json.load(sys.stdin)["flags"]))')"
|
||||
case "$FILLER_FLAGS" in
|
||||
*filler*|*ai-patterns*) ok "filler-heavy text is flagged" ;;
|
||||
*) no "filler-heavy text is flagged" "flags: $FILLER_FLAGS" ;;
|
||||
esac
|
||||
hasnt "clean text is not flagged filler" "$CLEAN_FLAGS" 'filler'
|
||||
hasnt "clean text is not flagged ai-patterns" "$CLEAN_FLAGS" 'ai-patterns'
|
||||
# proves BOTH directions: an always-flag or a never-flag detector is useless
|
||||
FILLER_Q="$(printf '%s' "$FILLER_OUT" | \
|
||||
python3 -c 'import sys,json; print(json.load(sys.stdin)["overall_quality"])')"
|
||||
CLEAN_Q="$(printf '%s' "$CLEAN_OUT" | \
|
||||
python3 -c 'import sys,json; print(json.load(sys.stdin)["overall_quality"])')"
|
||||
[ "$FILLER_Q" -lt 50 ] && ok "filler-heavy text scores LOW overall_quality" \
|
||||
|| no "filler-heavy text scores LOW overall_quality" "got $FILLER_Q"
|
||||
[ "$CLEAN_Q" -gt "$FILLER_Q" ] && ok "clean dense text scores higher" \
|
||||
|| no "clean dense text scores higher" "$CLEAN_Q vs $FILLER_Q"
|
||||
# empty / whitespace-only input never crashes and never claims a result
|
||||
EMPTY_OUT="$(printf '' | CQ)"
|
||||
has "empty input degrades" "$EMPTY_OUT" '"status": "degraded"'
|
||||
has "empty input reason" "$EMPTY_OUT" 'empty_input'
|
||||
WS_OUT="$(printf ' \n\t ' | CQ)"
|
||||
has "whitespace-only degrades" "$WS_OUT" '"status": "degraded"'
|
||||
# --file path works, no fixture committed — mktemp + rm
|
||||
CQTMP="$(mktemp)"; printf '%s' "$CLEAN_TXT" > "$CQTMP"
|
||||
FILE_OUT="$(CQ --file "$CQTMP")"
|
||||
has "file input is ok" "$FILE_OUT" '"status": "ok"'
|
||||
rm -f "$CQTMP"
|
||||
# a missing --file degrades, never a traceback
|
||||
MISSING_OUT="$(CQ --file /nonexistent/path/content-quality-test.txt)"
|
||||
has "missing --file degrades" "$MISSING_OUT" '"status": "degraded"'
|
||||
# stdlib ONLY — asserted, not assumed
|
||||
CQ_IMPORTS="$(grep -E '^(import|from) ' "$SD/content_quality.py")"
|
||||
if printf '%s' "$CQ_IMPORTS" | grep -qivE '^(import argparse, json, re, sys|from collections import counter|from typing import iterable)$'; then
|
||||
no "content_quality stdlib only" "unexpected import: $CQ_IMPORTS"
|
||||
else
|
||||
ok "content_quality stdlib only"
|
||||
fi
|
||||
# ADVISORY HONESTY (LRN-131/133): a heuristic signal, never a verdict
|
||||
hasnt "never claims ai-written" "$FILLER_OUT" 'ai-written'
|
||||
hasnt "never claims is AI verdict" "$FILLER_OUT" 'is AI'
|
||||
# dispatch wiring: --store precedes the verb (fetch.sh's own convention);
|
||||
# both stdin AND --file must work through the real entrypoint
|
||||
FCQ_STDIN="$(printf '%s' "$CLEAN_TXT" | \
|
||||
SEO_DATA_ENV_FILE=/dev/null SEO_DATA_STORE=/nonexistent bash "$SD/fetch.sh" content_quality)"
|
||||
has "fetch dispatches content_quality (stdin)" "$FCQ_STDIN" '"status": "ok"'
|
||||
CQTMP2="$(mktemp)"; printf '%s' "$CLEAN_TXT" > "$CQTMP2"
|
||||
FCQ_FILE="$(SEO_DATA_ENV_FILE=/dev/null SEO_DATA_STORE=/nonexistent bash "$SD/fetch.sh" \
|
||||
content_quality --file "$CQTMP2")"
|
||||
has "fetch dispatches content_quality (--file)" "$FCQ_FILE" '"status": "ok"'
|
||||
rm -f "$CQTMP2"
|
||||
|
||||
echo "── fetch.sh ──"
|
||||
FETCH="$SD/fetch.sh"
|
||||
# SEO_DATA_ENV_FILE=/dev/null: tests must NEVER source the real ~/.claude/.env —
|
||||
@@ -358,6 +488,8 @@ tf "analyzer calls fetch crux" "$REPO/agents/seo-analyzer.md" "fetch.sh crux"
|
||||
tf "analyzer calls fetch queries" "$REPO/agents/seo-analyzer.md" "fetch.sh queries"
|
||||
tf "analyzer gsc subsection" "$REPO/agents/seo-analyzer.md" "Performance GSC"
|
||||
tf "catalog gsc oauth entry" "$REPO/agents/resources/automation-catalog.md" "make seo-connect"
|
||||
tf "geo-analyzer wires schema_gen" "$REPO/agents/geo-analyzer.md" "fetch.sh schema_gen"
|
||||
tf "geo-analyzer wires content_quality" "$REPO/agents/geo-analyzer.md" "fetch.sh content_quality"
|
||||
|
||||
echo "── account-mgmt locks ──"
|
||||
tf "skill routes account verbs" "$REPO/skills/seo/SKILL.md" "forget --all"
|
||||
@@ -370,6 +502,9 @@ tf "readme documents fetch.sh" "$REPO/lib/seo-data/README.md" "fetch.sh"
|
||||
tf "readme documents seo-connect" "$REPO/lib/seo-data/README.md" "make seo-connect"
|
||||
tf "readme documents forget" "$REPO/lib/seo-data/README.md" "forget --all"
|
||||
tf "readme revocation note" "$REPO/lib/seo-data/README.md" "myaccount.google.com/permissions"
|
||||
tf "readme documents schema_gen" "$REPO/lib/seo-data/README.md" "schema_gen"
|
||||
tf "readme documents content_quality" "$REPO/lib/seo-data/README.md" "content_quality"
|
||||
tf "readme states advisory caveat" "$REPO/lib/seo-data/README.md" "ADVISORY, NOT A VERDICT"
|
||||
|
||||
echo ""
|
||||
echo "seo-data engine: $PASS pass, $FAIL fail"
|
||||
|
||||
Reference in New Issue
Block a user