Two or more project paths dispatch one general-purpose runner per repo,
all in a single message, instead of processing repos one by one. The
runner inherits the session model — no pin, it carries tour's reflection
(fix decisions, convergence) — and every agent inside keeps its defined
tier (security-auditor sonnet, Phase B opus, doc-syncer sonnet). A dead
or mute runner becomes an explicit RUNNER FAILED summary row; the gated
capitalize offer stays in the main loop, never in a runner.
Bounded LRN-083 derogation recorded in BDR-084: the per-project fix loop
moves into its runner, but nothing a runner decides touches shared state
— independent repos, per-repo chore branches, branches left unmerged for
human review exactly as inline. Mechanics proven before building: nested
probe, 3 sub-agent windows all overlapping, 9.1s vs ~18s sequential.
Census §12: 6 locks, flip-tested. Single-project path unchanged.
The include is authoritative, but feat/bugfix/ship-feature/init-project
each restate the verify loop inline — an orchestrator following the
restatement alone would have skipped the floor. Each now carries the
GATE 0 bullet ahead of GATE 1 (4 new structure locks, flip-tested).
The contract-interview weight table stops promising a hotfix oracle
nothing executes: hotfix runs no floor, the hotfixer runs the suite
itself. CHANGELOG extended with the wiring + the RED result.
BDR-083 records what was taken from unlazy and, more usefully, what was
refused and why. LRN-141: an external skill's machinery encodes its threat
model, not yours — take the invariants, refuse the machinery. LRN-142:
structure locks are fixed-string, so reflowing a doctrine paragraph reds
them; fix the doc, not the lock.
GATE 1 is an LLM dispatch and the verifier's mandatory PROOF: line is a line
the verifier writes — nothing structurally stops it being produced without
anything being executed. Nothing deterministic sat between the executor and
that dispatch.
An acceptance criterion can now carry an oracle: indented CHECK: (command),
EXPECT: (success-only marker), EVIDENCE: (slot). lib/gates.sh runs them
fail-closed — MET requires exit 0 AND the marker, so a nonzero process never
passes on its error text carrying the token — and writes the outcome back
into the contract, so the fresh verifier reads evidence as fact rather than
trusting the executor's report.
GATE 0 runs that floor before any verifier is dispatched; a red build sends
the executor back for free, on its own iteration budget. ABANDON: <id>
<reason> turns an impossible criterion into a visible handoff that blocks
CONFORME and routes to the human gate, via the new ABANDONED(n) verdict —
a distinct token because it routes distinctly, never a dev loop. feater and
bugfixer gain a four-pass completion discipline, scoped so a pass can never
widen the contract.
The runner's parse fails closed on partial oracles, duplicate ids,
unindented attributes and runnable criteria with no EVIDENCE: line, and
executes nothing at all when the ledger is malformed. status never executes
and never writes; run always re-executes, since trusting written evidence is
the failure being closed.
Adapted from the unlazy skill (Leonxlnx/unlazy, MIT). Its Stop hook,
approval store, .unlazy/ tree, depth-tree arithmetic and Node checker were
deliberately refused — BDR-083 records each reason.
64 assertions in lib/tests/gates.test.sh, non-execution proved by sentinel
with its own positive control asserted first.
skills-external/emil-design-eng/SKILL.md is curl'd from emilkowalski/skill
by install-plugins.sh when absent and re-fetched unconditionally by every
update-all.sh run, so tracking it produced a repo diff on each upstream edit
(latest: Radix vars dropped for Base UI). Same category as frontend-design/
and impeccable/, already ignored on that rationale — a fresh clone re-fetches
it, so no offline copy is needed and nothing was pinned here anyway.
design-motion-principles/ has the same overwrite-on-update behaviour but NOT
the same bootstrap: install-plugins.sh only warns instead of cloning it, so it
stays tracked until that gap is closed.
Opus 5 follows conservative-reporting clauses literally; 'a manufactured
concern is a failure' risked suppressing real low-confidence findings.
In-place reword: ungrounded stays noise, grounded-but-uncertain files as
[MINOR] with the uncertainty in WHY:. OUTPUT grammar byte-identical;
census row added.
3rd tightening pass (series LRN-1005/1007): bare "ux" matched inside
French prose (2 logged FPs, both FR — latest "changement ux vu").
\bui\b kept: zero logged FP, one logged true positive, now locked by a
must-fire test row. Flip-tested: quiet row fired pre-change.
BDR-065 delete-side now automated (lib/gitflow.sh _gitflow_purge_transient).
LRN-138: gitignore != delete for run-time artifacts read from disk — use
commit-during-run + auto-delete at the integration boundary. TODO checked,
journal line.
_gitflow_purge_transient removes docs/superpowers/{specs,plans} on the
feature/bugfix branch just before the directed merge, so develop's tip
lands clean while the feature commits stay reachable as the archive
(git show <sha>:...). Best-effort: never aborts a finish (no-op when
absent, skip on dirty paths, restore index+tree on commit failure).
Opt-out GITFLOW_PURGE_TRANSIENT=0; purge-transient CLI verb. Automates
the manual post-merge cleanup BDR-065 left as doctrine (slipped once,
655e364). Universal via the ~/.claude/lib symlink. gitflow-test T17 a-d;
shellcheck clean; make test exit 0.
- Fresh-install block: clone URL → github.com/bchanot/claude, bash
install.sh/doctor.sh → make install / make doctor (user pass)
- magic MCP example: placeholder key line instead of WRONG/RIGHT contrast
- new subsection: SEO data layer needs GOOGLE_OAUTH_CLIENT_ID/SECRET +
CRUX_API_KEY in ~/.claude/.env (GCP steps, make seo-connect, graceful
degradation) — mirrors .env.example
- docs/plans + docs/specs (deploy-skill 2026-06-27): predate the BDR-065
lifecycle codification, never swept
- docs/superpowers/{plans,specs} (model-routing 2026-07-15): 6-wave
chantier — final W6 merge closed it without the purge step
Git history at the feature commits is the archive (BDR-065).
- MANAGED_EXTERNALS (emil-design-eng, frontend-design,
design-motion-principles, impeccable) + MANAGED_MCPS (magic):
cmd_set now trims both when the profile does not list them —
design leftovers no longer survive a 'set backend'
- cmd_set refactored to 4 symmetric trim helpers; nothing outside
the MANAGED_* allowlists is ever auto-toggled (darwin-skill manual)
- enable_skill external: from-source fallback (ln -sf
skills-external/<name>), mirrors toggle-external.sh
- stale usage() NOTE + SKILL.md updated to the both-ways reality
- hermetic test: 16 checks, fixture repo + fake claude shim (gstack
on-demand, from-source, park/restore, magic add/remove, non-managed
untouched); shellcheck + full make test green
Fresh-opus whole-chantier ronde (17fbe51..HEAD, EVAL-023 style): axes
severed-wires / gate-regressions / fail-open PROVEN CLEAN (all 7 new
handoffs traced end-to-end both sides). Fixed: init 5b now dispatches the
FULL-AUDIT path (auto-mode gated a missing README as SIGNIFICANT →
[CREATE-AUTO] unconditional restored, sole greenfield README path);
census locks added for /geo ERROR CONTRACT + never-re-derive (deleting
the fail-closed handler would have stayed green), doc-audit
model=opus override x5 flows, SYNTH REPORT grammar; doc-syncer ex-STEP-8
prose repointed to the dispatcher gate; plugin-advisor anim rows read
the PROBE REPORT ANIM field (no Bash anymore); client-handover 9.7
cleans the transient draft. Census caught one more line-wrapped lock
before it shipped vacuous. 133 pass / 0 fail, make test exit 0.
seo-analyzer + geo-analyzer gain MODE: collect|judge|template around the
dispatcher (mode-based, zero body-text moves — seo-data fetch-wiring
locks survive; opus pin kept = fail-safe direction, a forgotten override
over-tiers but never downgrades judgment). Run-scoped gitignored
signals handoff (.audit/*-signals-<RUNID>.md + COLLECTION COMPLETE
sentinel), judge fails closed on absent/mismatched/unsealed signals.
/seo rewired to 3 phases (domains parallel per phase) + DISPATCHER ERROR
CONTRACT (mute/ERROR judge never carried into templating; retry once,
escalate); /geo same single-domain; legacy no-MODE single-shot kept on
the opus pin for /harden narrow-scope + /onboard report-only. Dropped
/geo's 'ask and I relay' fiction (dispatched agents cannot ask).
In-wave smokes PASSED disk-verified: collect signals+sentinel; judge
ERROR-verdict on wrong RUNID; real judge = honest N/A + deterministic
engine + full scoring grammar; template = complete envelope + verbatim
sentinel + zero re-derivation. Census §18 (125 pass — one vacuous
line-wrapped lock caught by the census itself and fixed), make test
exit 0.
handover-doc-writer: MODE synthesize (model="opus" call-site — STEP
9/10/12 → .audit/handover-draft-<RUNID>.md + DRAFT COMPLETE sentinel) /
MODE render (sonnet pin — STEP 13-16 from the draft, fail-closed on
absent/mismatched RUNID). Mode-based, not a file split: the §9
name+dispatch census locks survive untouched. client-handover-writer 9.6
dispatches twice with the FULL PACKAGE both times (LRN-126) + RUNID mint
+ post-run draft cleanup. In-wave smokes PASSED disk-verified: draft
written+sentinel+gitignored; render BLOCKED on wrong RUNID (no phantom
synthesis); render consumed draft + honored skip-write, report grammar
intact. Census §17 (111 pass), make test exit 0.
Plan challenged by 3 blind lenses + 1 confirmation pass (1 BLOCKER closed
by fable-dispatch spike, 6 MAJORs + 8 MINORs fixed by named changes, 0
deferred). TODO: seo-geo-integrity 'UNMERGED' note was stale (92301fe
already in develop) — corrected.
Reverses the BDR-066 rejected alternative (opus pins on audit agents):
session default is now Fable, so inherit burned Fable quota on every
dispatched audit/challenge. analyzer, plan-challenger, seo/geo/
validator-analyzer pinned model: opus; onboard's 6 general-purpose
audit dispatches carry model="opus"; tour Phase B repointed.
interviewer + client-handover-writer stay unpinned (inline-load only,
a pin there is inert). settings.json default: claude-fable-5[1m].
Census flipped: model-routing §3 + new §11 (61 pass), loops-light 35,
full make test green.
STEP 1.8 (Option B): skip purely cosmetic fixes (CSS/copy/typo), run the 3-lens
challenge only when the fix touches control flow/behaviour (off-by-one, wrong
operator, behaviour-changing config, execution-altering import); a BLOCKER means
it was never a hotfix -> escalate to /bugfix. 12th orchestrator wired.
- skills/hotfix/SKILL.md — STEP 1.8 guarded challenge
- lib/tests/plan-challenger.test.sh — lock hotfix into the census (43 assertions)
After a plan/reflection is elaborated and before it executes, three fresh blind
sub-agents (correctness / robustness / simplicity) attack it on the big model;
the main loop RE-THINKS every aspect a BLOCKER lands (a named plan change, or
[deferred]) and re-challenges once if the plan materially changed. Advisory into
each skill's existing human gate — the human stays the decider.
- lib/challenge-plan.md — reusable phase: fail-safe (never fail open),
severity-driven (any single-lens BLOCKER = must-address), RE-THINK loop
- agents/plan-challenger.md — challenger role (read-only, big-model per BDR-066)
- lib/tests/plan-challenger.test.sh — 41-assertion structure lock
- wired into 11 orchestrators: ship-feature/init-project/feat/bugfix (build-plan),
onboard/audit-delta/code-clean (proposals), seo/geo/harden/web-validate (fix-bundle)
Hardened by dogfooding: 3 blind challengers reviewed this feature's own v1 plan
and caught 4 BLOCKERs (fail-open, consensus-buries-lone-finding, wrong model
tier vs BDR-066, false on-disk-plan premise) — all fixed here.
Full removal per user request: the PreToolUse hook that blocked model
Edit/Write on quality-gate files (settings.json, gitflow.sh, .githooks,
doctor.sh, hooks, lib/tests, lint configs) plus its one-shot sentinel.
- delete hooks/config-protection.sh
- delete lib/tests/config-protection.test.sh
- deregister the hook from settings.json (rtk-rewrite PreToolUse kept)
- drop the README mention
Residual protection unchanged: gitflow pre-commit guard + Gitea branch
protection still block direct code commits to main/develop.
By-principle hardening. H1's url-guard validates the NAME; urlopen then resolved
AND connected — two DNS lookups with a window a hostile authority uses to answer
PUBLIC to validation and PRIVATE (169.254.169.254 metadata, 127.0.0.1, the LAN)
to the connect. A name-level guard cannot see that rebind.
safe_fetch collapses the two lookups into one: resolve ONCE, validate every IP
(ipaddress, dual-stack v4+v6), refuse if ANY is non-public (the multi-A vector),
connect to the exact validated IP with Host+SNI+cert for the real host — no
second resolution to poison. Redirects re-validate each hop (urlopen followed
them blind). One seam: sitemap._fetch, which linkgraph/render_check/drift all
call, so every network verb inherits it.
The load-bearing property (confirmed by the security review): classification is
on the OS-resolved address (sockaddr[0]), never the URL text — so octal/hex/
decimal literals, IPv4-mapped IPv6, NAT64, 6to4 are all defeated structurally,
not by enumeration.
Better than the source idea (claude-seo url_safety.py, MIT): dual-stack (theirs
IPv4-only), no global monkeypatch so thread-safe by construction (theirs locks a
patched getaddrinfo), stdlib-only (no requests). Proven end-to-end before
writing: pinned connect keeps SNI+cert for the real host.
NOT covered, stated not silent: shell `curl` in the agent specs (separate
process, unpinnable here). Smaller surface; `curl --resolve` is a separate change.
REVIEW-SURFACED (fresh security-auditor, adversarial, VERDICT PASS) — two real
holes it found while attacking the diff, both fixed here:
- billion-laughs REOPENED in C1b: _refuse_dtd scanned only raw[:4096], so a
>4KB leading comment pushed <!DOCTYPE past the window while ET parsed AND
EXPANDED the entities. Proven (&lol2; → "lollollollollol"), now a full-doc
case-insensitive scan. This is a genuine fix to already-merged C1b, not this
feature — fixed here rather than filed, per root-cause discipline.
- 192.88.99.0/24 (6to4-relay anycast) passed is_global as public — added to an
extra special-use deny list.
Verified: rebind-to-metadata refused BEFORE any connect (injected resolver),
multi-A public+private refused, classifier fuzzed dual-stack incl. CGNAT/6to4,
non-http scheme refused, both review fixes proven with no false positive; real
fetch still works (zenquality 86 loc, lavageangels 24) through the pinned path;
all 4 verbs work end-to-end via fetch.sh; seo-data 210 → 221 pass, 0 fail; full
suite green; shellcheck + py_compile clean.
Cherry-picked from claude-seo (github.com/AgriciDaniel/claude-seo, MIT)
content_quality.py, rewritten to the lib/seo-data contract per BDR-070. The
Content Shape axis was 100% LLM judgement; this gives it a measured input.
fetch.sh content_quality (stdin or --file) → {filler_score, ai_pattern_score,
information_density, overall_quality, flags[], matches{}}. 100% deterministic:
QRG §4.6 filler list (26 phrases) + AI-pattern list (46) kept intact, regex
matching, no LLM. Stdlib only (argparse/json/re/sys/collections/typing).
Advisory, NOT a verdict — the point of the wiring. It never claims a page "is
AI-written" (LRN-131/133); flags are candidates for human review. geo-analyzer
STEP 8 Check 10 makes it a deterministic input that INFORMS checks 1-9, never
replaces them, never scored on its own. A low number is not an automatic
finding.
Detection proven both directions (a detector that always- or never-flags is
useless): filler+slop text → flags [filler, low-density], overall 34-49; clean
dense factual text (dates/EUR/percentages) → no flags, overall 90. Empty input →
degraded/empty_input, never zeros-as-a-result.
Verified: GATE 1 verifier CONFORME 10/10 (both directions exercised live, lists
diffed intact vs source, advisory language confirmed); GATE 2 self-scan clean
(only sink is read-only open() for --file); seo-data 190 → 210 pass, 0 fail;
full suite green; shellcheck + py_compile clean.
Cherry-picked from claude-seo (github.com/AgriciDaniel/claude-seo, MIT)
schema_generate.py, rewritten to the lib/seo-data contract per BDR-070 — adapt,
never copy. The system audited JSON-LD but could not generate it; geo-analyzer's
G2 batch hand-wrote markup. Now it calls the verb.
fetch.sh schema_gen {reservation|order|discussion|profile} → fail-open envelope
{"status":"ok","source":"schema_gen","type":…,"jsonld":{…}}. Types: Reservation
(7 subtypes), OrderAction, DiscussionForumPosting, ProfilePage (sameAs/knowsAbout
for the entity graph). Stdlib only (import argparse, json — zero third-party),
_strip_nones so a null is never emitted, --script-tag wraps for direct paste.
Fail-open mirrors score.py's _cli exactly (the contract's named pattern): a
flag-omitted required field → argparse exit 2 + {"status":"error","bad_usage"};
a flag-present-but-empty field → {"status":"degraded","reason":"missing required
field: …"} exit 0. Never a traceback, never empty stdout (LRN-133: the
can't-generate case stays legible).
geo-analyzer G2 wiring preserves the data-integrity rule — the verb generates
STRUCTURE, unknown values stay [À COMPLÉTER], never invented.
Verified: GATE 1 verifier CONFORME 10/10 (exercised the fail-open edge cases +
diffed field surface against the source); GATE 2 self-scan clean (no
network/shell/secret/eval sink); seo-data 167 → 190 pass, 0 fail; full suite
green; shellcheck + py_compile clean.
/harden has a real scale (SKILL.md:435 — Critique -15, Haute -8, Moyenne -3,
Basse -1, clamp [0,100]). /seo had none: every axis was felt, so two runs over
identical code could disagree. That is a credibility problem on its own, and
/client-handover gates on 17/20 — a wobbling number makes the gate arbitrary.
H2 sharpened it: now that drift reports what actually changed, a score moving
on its own is visibly noise.
The split is the whole point. WHICH findings exist and how severe each is
stays the LLM's judgement — irreducible, and I am not pretending otherwise.
The arithmetic stops being judgement: same findings in, same score out. Same
principle as grouping cannibalisation rows in the engine rather than handing a
model 1000 rows to add up.
Reuses /harden's scale, /5 into /20, so the family speaks one vocabulary
instead of two.
Two things it makes real that were prose:
- **N/A is not a zero.** R2 (client-rendered on-page) and I1 (unauditable
off-page) both mandate excluding an axis and renormalising the rest. Both
left that arithmetic to the model. Now the engine does it and refuses to let
N/A behave like a zero — verified: all-20 axes with two N/A still yields
global 20.0, not a dragged-down mean.
- **Prevalence.** affected/sampled shift severity ONE step (>=50% escalates, a
single page de-escalates). A defect on 1 of 12 pages is not the defect on
12 of 12, and flattening the two is part of what made the old numbers move.
Malformed input is an error, never a silently wrong number — unlike the fetch
verbs, a degrade here would mean bad input, not a network fact. Unknown
severity and unknown profile both rejected, tested.
Verified: hand-checkable arithmetic (haute+moyenne = 100-11 = 89 → 17.8;
critique+haute = 77 → 15.4), identical global across repeated runs, weights
renormalised to sum 1.0 with two axes N/A. seo-data 155 -> 167 pass, 0 fail;
full suite green; shellcheck + py_compile clean.
seo-analyzer.md:1365 keeps history as "date + score + key changes" — prose the
LLM writes about its own previous prose. Lossy, unreproducible, and
machine-uncomparable, so "the redesign silently dropped 40 canonicals" is
invisible unless someone happens to notice.
drift snapshots title/description/canonical/robots/h1_count/jsonld_types per
URL and diffs them. Stdlib only, no auth.
The classification IS the feature: LOSING a signal is a regression, CHANGING
one is a change that may well be intended. The engine says which kind; the
agent judges. A reworded title is not an alert; an evaporated canonical is.
Runs over the WHOLE sitemap, never a sample — caught while designing: a drift
computed over a sample that changes between runs compares nothing.
NOT rank tracking. That is the common misread of this same feature elsewhere;
positions come from GSC `queries`. This is on-page regression detection.
Also caught in my own draft before testing: _capture reused
sm._mock("page.html"), the exact single-fixture flaw I had already fixed in
linkgraph — one fixture cannot express a multi-page snapshot, every URL would
read identical. Now pages.json, same convention.
Proved on a planted failure rather than a happy path — two clean sites would
look identical to a detector that always returns []:
v1 -> v2: canonical lost on /a, h1 + jsonld lost on /, title reworded,
/gone removed, /neuve added
→ 3 regressions, 1 change, gone/new both detected, title correctly NOT a
regression.
Store is ~/.claude/seo-data/drift/<host>.json, 0700, written via os.replace so
a crash never leaves a half-written baseline; a corrupt store degrades to
"first run" instead of killing the audit.
Verified: seo-data 144 -> 155 pass, 0 fail; full suite green.