Commit Graph
176 Commits
Author SHA1 Message Date
bastien 58c3a3e9b7 fix(effort): re-raise judgment dispatches, planning re-asserts, pairing caveat, dedupe audit script (final review I1-I3) 2026-09-28 20:48:27 +02:00
bastien 98ef991958 docs(effort): BDR-107 id, CHANGELOG entry, spec corrected for the rulings (vendored pins, exclusions, pairing rule) 2026-09-28 20:20:05 +02:00
bastien a117e7ed67 fix(effort): shifts are sent with the step's first tool call (harness pairing rule); challenge shifts under their heading; fence indentation 2026-09-28 19:55:47 +02:00
bastien 3c58160d0c feat(effort): wire phase shifts in the 13 orchestrators and the handover writer 2026-09-28 19:44:22 +02:00
bastien 9223fda99f feat(effort): pin effort on the 20 repo-authored agents (BDR-077 second axis) 2026-09-28 18:59:49 +02:00
bastien ddea411491 chore(config): superpowers citers by bare name, routing map, docs, settings
Every superpowers-prefixed skill call in ship-feature, init-project, tour,
deploy, audit-delta, plugin-advisor and lib/analyze-before-plan now names
the vendored skill directly. finishing-a-development-branch is described
as the upstream skill this config does not vendor (gitflow finish is the
integration path). CLAUDE.global.md Skill routing maps the four
non-vendored skills the vendored text still references. settings.json
loses the plugin key and its marketplace block; README, USAGE,
plugin-advisor and the profile skill describe superpowers as vendored
skills, always on, zero plugin cost. CHANGELOG entry with a known
residual.
2026-09-28 14:54:53 +02:00
bastien 4c86d6dc70 chore(config): security-guidance Stop review off, plugins off, routing and docs
settings.json: ENABLE_STOP_REVIEW=0 (the plugin's own switch: no more
Opus call on every turn that changes code, 0 findings in 6 days, 1
recorded false positive; the regex layer and the commit/push agentic
review stay on), brightdata-plugin@synced false (keyless-useless, its MCP
skill would hijack WebFetch/WebSearch), frontend-design official plugin
entry gone (uninstalled: byte-identical to the managed copy).

CLAUDE.global.md routes Ship/PR to ship-feature (gstack ship takes
origin/HEAD = main as base), drops ship/context-save from the gstack-off
list and 21st-ui-review from the design review line (trio is max-only).
deploy's table no longer points at land-and-deploy/setup-deploy.
plugin-advisor.md describes security-guidance's real mechanics. CHANGELOG
Unreleased entry with a Known residual section.
2026-09-28 11:49:01 +02:00
bastien 6617889b77 feat(verifier): floor-guard waivers outside test files need a CLARIFICATIONS ack
Security-gate MEDIUM: a self-service floor-guard: allow <reason> neutralised
the detector in the same commit. User chose strict: the tool prints WAIVED,
the contract authorizes, the verifier counts the rest as gaps. BDR-102
amendment.
2026-09-27 21:20:36 +02:00
bastien 2b25cb4704 feat(lib): floor-guard, diff-scoped detector of a weakened quality bar
SUPPRESS / SKIP / DELETED_TEST / ASSERT_DROP / STUB / THRESHOLD_DOWN over
git diff <base> (untracked files included), floor-guard: allow <reason>
waiver printed as WAIVED, rc 0/2/3. Mandatory verifier STEP 3, documented
under GATE 1 of verify-secure-loop.md. Suite: 6 kinds + WAIVED + CLEAN,
flip-tested. Adapted from agent-skills constraint-driven-development.
2026-09-27 20:17:36 +02:00
bastien 0d035fcab6 feat(profile): default profile = full; reset applies it, current is label-driven
No profile selected (.active-profile absent, empty or legacy "none") now
means the `full` profile is in force: DEFAULT_PROFILE declared once in
lib/profile.sh, resolved by active_profile(); the statusline reads the
constant and shows `full` instead of `?`; `gstack off` trims to it instead
of erroring. `reset` goes to the default profile (= `set full`: enables its
list, parks any non-listed gstack or managed item). `current` names the
active label and scores that profile only, saying `default — not applied
yet` until a set/apply/reset wrote the cache; the "none" sentinel and the
cross-profile best-guess scan are gone (they keyed on the parked-gstack
count, which says nothing under BDR-030's gstack-off default). Hermetic
suite lib/tests/profile-default.test.sh (29 checks) seeds gstack as OFF like
a real tree. Citers updated: profile SKILL, Makefile help, plugin-advisor
PROFILE line + reset paragraph, toggle-external header.
2026-09-25 16:28:21 +02:00
bastien 27f201d4aa feat(guardrails): refusal ends the attempt; doctrine-citers census; make test suite=
Root causes of the 2026-09-24 errors turned into mechanisms (BDR-100). hard_deny 'Routing around a guardrail': a refused command is never rerun through a wrapper, alias, heredoc, Makefile target, env file, other shell or other agent; the same clause in 14 agents and in the doctrine's sub-agent rule. make test suite=<file> runs one suite hermetically so the denied env-prefix form is never needed by hand. lib/tests/doctrine-citers.test.sh: every CLAUDE.md "Section" / § Label citation across skills, agents, lib, rules and hooks must resolve to a heading or bold label (flip-tested); its first run fixed rest-api-node.md. Doctrine 'After code changes' step 4: a changed rule, heading, label or threshold → grep every citer in the same commit.
2026-09-24 20:58:25 +02:00
bastien d82c06f572 refactor(doctrine): C2 coherence — 30 doctrine/skill tensions resolved, doctrine wins (BDR-099)
One ask policy; mandated executors exempt from the delegation rule; skill plan satisfies the planning rule; journal line exempt from the approval gate; chore = maintenance without new behaviour; small fix on develop = bugfix; BDR-068 written as the one auto-finish exception; deploy routes to /deploy. Skills and agents follow: hotfix types by base + skips the design gate on trivial; capitalize/close create missing registries; commit-change asks the branch type; doc/seo/web-validate/refactor branch through the aiguillage; tour reports BREAKING fixes as needs-decision and runs doc-syncer two-mode; client-handover applies audit bundles from its main loop behind one gate; init-project/onboard use the 200-file graphify signal and bootstrap memory; release-candidate gates the tag push only; push wording aligned with the BDR-095 hooks; stale pointers fixed (§ Language, .gsd/ROADMAP.md, handover script path, design-gate lists).
2026-09-24 20:25:40 +02:00
bastien c81b1731af feat(graphify): threshold signal from 200 tracked code files, the banner informs and the user decides
lib/graphify-gate.sh counts tracked code files (graphify's AST extension set, vendored trees excluded) and, from 200 with no graphify-out/graph.json, prints one banner-sized line; session-start shows it with the /graphify hint. Nothing is built, installed or updated: the rule is the user's (BDR-097), grounded in the LRN-162 measurements (AST build 2.3 s, 0 tokens, a query 2 to 3k tokens). Doctrine section and plugin-advisor thresholds follow the same rule; graphify claude install stays rejected. Test: 11 checks. GRAPHIFY_MIN_CODE_FILES overrides the threshold.
2026-09-24 12:12:16 +02:00
bastien 9da5d8d52c feat(guardrails): push every commit, static deny for destructive tools, brief carries no user authority
Layer C of the plan written after the 2026-09-21 wipe (BDR-095): a reviewer
sub-agent traced `lftp mirror --delete` against a local file:// tree, the
prose tiers named neither lftp nor a local trace, the brief had authorized
it, and four days of commits had never left the machine.

- gitflow: `start` pushes the branch with its upstream, merge targets are
  pushed after each merge, and `init`/`install-hook` write post-commit and
  post-merge hooks that push every commit as it lands (warn, never block;
  GITFLOW_NO_PUSH=1 for throwaway repos). T18 + T19 (installed == emitted).
- hooks/unpushed-guard.sh on SessionStart and Stop: branch ahead of its
  upstream, no upstream, or no origin. Non-blocking systemMessage.
- settings.json: static deny for transfer and mirror tools, rsync --delete,
  xargs rm, pipe-to-shell, chmod/chown -R, sudo/doas/pkexec, disk tools,
  chattr, docker volume drops/prune/--privileged/socket/-v /:, git history
  destruction, --no-verify and core.hooksPath; new hard_deny "destructive
  tool against a local path, brief carries no user authority"; soft_deny
  reworded + discarding uncommitted work; environment records the incident.
- CLAUDE.global.md "Destructive tools & data loss"; the four report-only
  agents trace by reading, never by running, whatever the brief says.
- lib/tests/guard-bash.test.sh: executable spec of the PreToolUse guard
  (214 cases). The hook itself is not shipped (BLK-022); the spec skips.
2026-09-22 07:43:12 +02:00
bastien 7c05f75eab feat(21st): replace the magic MCP with the @21st-dev CLI + skill pack
Upstream supersedes `@21st-dev/magic` with `@21st-dev/cli` (bin `21st`):
same endpoint, `21st login` in place of an API key, no MCP process loaded
into every session.

- install-plugins.sh Step 8.7: `npm i -g @21st-dev/cli` (pinned in
  plugins.lock.json), staged `21st skills install`, TTY-only login offer,
  pack disabled by default. update-all.sh 7.4 refreshes both.
- The documented `21st install-skill` cannot be used: the installer refuses
  to follow a symlink on the target path and `~/.claude/skills` is one. The
  install runs under a throwaway HOME and the result moves into
  skills-external/21st-* (gitignored), symlinked on demand.
- toggle-external.sh manages `21st` as a pack (names globbed from
  skills-external/21st-*, parked under plain names). `magic` is gone.
- The 5 design skills join design/web/web-full/full and MANAGED_EXTERNALS;
  21st-registry and 21st-design-sync stay parked. MANAGED_MCPS is now empty
  and profile.sh's dead magic branches are removed.
- Design gate: GATE-BLOCK gains `21st` (required-manual, magic's old slot)
  and `21st-ui-build`; PATH repair extended to the npm global bin.
- settings.json: the 4 mcp__magic__* ask entries go; the outward-facing
  21st verbs land in autoMode.soft_deny, the tier that holds under auto
  mode (LRN-153).
- Docs: README, CLAUDE.global.md, design-gate.md, profile SKILL.md,
  .env.example, .gitleaks.toml, link.sh. BDR-093, LRN-158.

Tests: profile-set-managed 17/17, make test green except 2 pre-existing
gitflow FAILs (gitleaks binary absent on this host), shellcheck clean.
2026-09-22 02:53:31 +00:00
Bastien Chanot 17370d7e4c feat(executors): NEED-DECISION and BLOCKED carry a CLASS tag 2026-09-16 22:14:13 +02:00
Bastien Chanot 7cc95952bd feat(interviewer): a visible or public choice is asked, never assumed 2026-09-16 22:13:50 +02:00
Bastien Chanot 6eac7fbca9 fix(hotfix): batch-skeptic residuals — RULES restore mandate file-scoped; hotfixer FILE(S) marks created files (new) 2026-08-26 22:28:31 +02:00
Bastien Chanot ad4985f410 fix(security-auditor,close): hotfix no-verifier carve-out documented; close enumerates STEP 5C + --no-push passthrough 2026-08-26 21:59:08 +02:00
Bastien Chanot c983f1ff94 fix(handover-writers): stale ch.4 refs post-NAP-renumbering (glossary/tone->6, cross-links/THRESHOLD->5); 14.5 verification deferred to post-write; anchor gate ordered into STEP 16 2026-08-26 21:55:13 +02:00
Bastien Chanot 27f17939d7 fix(plan-challenger): ERROR verdict added to the load-bearing OUTPUT grammar (STEP 1 emitted it, parser enum omitted it) 2026-08-26 21:54:35 +02:00
Bastien Chanot 9428b86880 optimize status-reporter r2: d5 — passive cost sourced from doctor.sh constants (skeptic's find), count-only fallback kept 2026-08-26 19:44:03 +02:00
Bastien Chanot a3f1624b15 optimize status-reporter: d5 — unproducible token field replaced by /plugin-check deferral; dead ROADMAP.md row rewritten for post-ADR-013 gsd layout 2026-08-26 19:35:46 +02:00
Bastien Chanot e9c6bf52fa optimize analyze-system: d1 triggers in skill description + d2 ordered TASKS mapped to OUTPUT sections in analyzer 2026-08-26 19:22:36 +02:00
Bastien Chanot 080d2d9f03 optimize plugin-probe+advisor: d8 — FRAMEWORK-DEPS exact dep@version (no preact false-hit, fallback fires), advisor derives frontend/fast-libs from it, PLAN echoed-or-unknown (no invention) 2026-08-26 19:16:09 +02:00
Bastien Chanot c0a2a8069d optimize refactor-system: d4 — no-tests STOP gate (agent) + user arbitration loop (dispatcher) + mid-run test-failure revert; code-cleaner inline carve-out 2026-08-26 19:05:29 +02:00
Bastien Chanot 562a42e936 optimize onboarder: d8 — BRIEF contract split REQUIRED/OPTIONAL, draft placeholders replace blanket STOP (fixes first-dispatch bounce vs /onboard STEP 2 minimal brief) 2026-08-26 15:12:24 +02:00
Bastien Chanot 9db213b0a1 optimize interviewer: d9 — DO NOT blacklist (no design, no invented values, no budget overrun) 2026-08-26 12:35:29 +02:00
Bastien Chanot e0923d6c4b optimize interviewer: d3 — failure-mode table (vague/idk/contradiction/partial/balloon) + 2-round budget 2026-08-26 12:34:15 +02:00
Bastien Chanot 63310467ca feat(gates): deterministic floor (GATE 0) under the fresh verifier
GATE 1 is an LLM dispatch and the verifier's mandatory PROOF: line is a line
the verifier writes — nothing structurally stops it being produced without
anything being executed. Nothing deterministic sat between the executor and
that dispatch.

An acceptance criterion can now carry an oracle: indented CHECK: (command),
EXPECT: (success-only marker), EVIDENCE: (slot). lib/gates.sh runs them
fail-closed — MET requires exit 0 AND the marker, so a nonzero process never
passes on its error text carrying the token — and writes the outcome back
into the contract, so the fresh verifier reads evidence as fact rather than
trusting the executor's report.

GATE 0 runs that floor before any verifier is dispatched; a red build sends
the executor back for free, on its own iteration budget. ABANDON: <id>
<reason> turns an impossible criterion into a visible handoff that blocks
CONFORME and routes to the human gate, via the new ABANDONED(n) verdict —
a distinct token because it routes distinctly, never a dev loop. feater and
bugfixer gain a four-pass completion discipline, scoped so a pass can never
widen the contract.

The runner's parse fails closed on partial oracles, duplicate ids,
unindented attributes and runnable criteria with no EVIDENCE: line, and
executes nothing at all when the ledger is malformed. status never executes
and never writes; run always re-executes, since trusting written evidence is
the failure being closed.

Adapted from the unlazy skill (Leonxlnx/unlazy, MIT). Its Stop hook,
approval store, .unlazy/ tree, depth-tree arithmetic and Node checker were
deliberately refused — BDR-083 records each reason.

64 assertions in lib/tests/gates.test.sh, non-execution proved by sentinel
with its own positive control asserted first.
2026-08-24 13:12:38 +02:00
Bastien Chanot c7646a9c8a feat(agents): de-prescribe geo-analyzer for Opus 5 (C1 P3)
Same invariant as adafa35. Census 71/0 green; contract surface
byte-identical; 1106→1107 lines (single-shot scoping line).

- MANDATORY/MUST caps on AI-index submission → plain content rule
- 'Print the plan before STEP 13' → single-shot-scoped (conf#1);
  tier-mapping kept in BOTH judge and template ranges (rob#1/#9,
  PERMISSIVE :873 named survivor kept)
- vestigial ':1106 Transparency §14' reworded to truth: change log is
  dispatcher's SEO.md §15 (folded into 'Dispatcher verifies')
- 'copy these patterns' → reporting-shape-to-match; FAQ '20-50'
  quantity → 'typically dozens'; 'EVERY finding' → outcome bar
- dedup after inspection: ZERO merges (PERMISSIVE ×3 cross-range;
  never-apply ×4 distinct obligations; content_quality pair =
  spec rule vs emitted-artifact caveat — annex counts corrected)
- FROZEN untouched: guard-first :273, NAP direction rule, cite-sources,
  WebSearch-freshness (already when-shaped, rob#6), all STEP headers
- deltas: MANDATORY 1→0, MUST 4→3, NEVER 9→8
2026-08-02 01:30:12 +02:00
Bastien Chanot adafa350da feat(agents): de-prescribe seo-analyzer for Opus 5 (C1 P2)
Choreography → when-guidance under the audience×mode-range invariant
(plan §4b Q3). Census 71/0 + model-routing 133/0 + seo-data 221/0 green
throughout; contract surface (§2a/2b) byte-identical; 1528→1503 lines.

- self-output verification: ':970 run it twice' → deterministic-engine
  integrity guard (conf#8); ':1217 do not proceed' → single-shot-scoped
  (conf#1); grouping sanity-check → when-guidance detector (rob#8)
- vestigial pre-BDR-061 'Transparency' line deleted; §15 ownership
  folded into 'Dispatcher verifies'
- dedup (inspection-corrected: most annex 'twins' are distinct
  obligations — kept): only true same-range dups removed ('Handoff to
  dispatcher' ≈ sentinel note; 'Landing page rule' block ≈ payload
  instance + RULES line)
- completeness checklist reshaped to routing map, rows verbatim (rob#3)
- caps softened: 'P0 rule' MUST/ALWAYS → plain content rules (CMS
  plugin-first folded with STEP 2 twin, corr#3); First-action/ordering
  emphasis dropped; C1a + sampling essays compressed (rules + LRN
  citations kept); WebSearch → drifting-externals when-guidance (rob#6)
- FROZEN untouched: guard-first orderings :287, denominator-before-
  sampling :550, R2 refuse, COVERAGE obligations, all STEP headers
- deltas: 'P0 rule' 2→0, ALWAYS 1→0, MUST 5→4; NEVER 9→9 (class-B
  named bans, kept by design LRN-105)
2026-08-02 01:30:02 +02:00
Bastien Chanot c3d3f4d465 feat(agents): plan-challenger — route grounded doubts to [MINOR]
Opus 5 follows conservative-reporting clauses literally; 'a manufactured
concern is a failure' risked suppressing real low-confidence findings.
In-place reword: ungrounded stays noise, grounded-but-uncertain files as
[MINOR] with the uncertainty in WHY:. OUTPUT grammar byte-identical;
census row added.
2026-07-30 12:58:29 +02:00
Bastien Chanot b7026e4bda feat(ctx7): coverage extension — fast-libs single source + reminder hook + executor briefs (BDR-078)
- lib/fast-libs.sh: detect/cache-status verbs, JS+Python manifests,
  7-day cache freshness, LC_ALL=C sort — replaces 3 hardcoded lists
  (ship-feature 0c, init-project 5c, onboard 3.5)
- hooks/ctx7-reminder.sh: once-per-session UserPromptSubmit nudge when
  the project carries fast-libs and .ctx7-cache/ is missing/stale
- find-docs: before-writing-code trigger + cache-first rule; dist is
  machine-owned (gitignored) so the durable patch lives in
  install-plugins.sh STEP ctx7 (idempotent, grep-guarded)
- feater/bugfixer briefs: fast-lib docs rule (fresh cache read, else
  2-topic ctx7 fetch, else NOTES cache miss + proceed)
- tests: lib/tests/fast-libs.test.sh (11 checks); shellcheck + full
  make test green (review-guards 5/0)
2026-07-20 10:45:06 +02:00
Bastien Chanot 444c79acb2 fix(routing): W6 ronde — 6 findings closed (README bootstrap path, 2 census gaps, 3 stale refs)
Fresh-opus whole-chantier ronde (17fbe51..HEAD, EVAL-023 style): axes
severed-wires / gate-regressions / fail-open PROVEN CLEAN (all 7 new
handoffs traced end-to-end both sides). Fixed: init 5b now dispatches the
FULL-AUDIT path (auto-mode gated a missing README as SIGNIFICANT →
[CREATE-AUTO] unconditional restored, sole greenfield README path);
census locks added for /geo ERROR CONTRACT + never-re-derive (deleting
the fail-closed handler would have stayed green), doc-audit
model=opus override x5 flows, SYNTH REPORT grammar; doc-syncer ex-STEP-8
prose repointed to the dispatcher gate; plugin-advisor anim rows read
the PROBE REPORT ANIM field (no Bash anymore); client-handover 9.7
cleans the transient draft. Census caught one more line-wrapped lock
before it shipped vacuous. 133 pass / 0 fail, make test exit 0.
2026-07-19 23:56:58 +02:00
Bastien Chanot 07253e093c feat(doctrine): W6 — prose sweep + BDR-077 + LRN-137 + plan execution notes
LRN-113 whole-surface sweep: client-handover x2 + commit-change prose
repointed to the two-mode reality; code-cleaner/status historic notes
kept (accurate). Memory: BDR-077 (full architecture), LRN-137
(mode-based re-tiering + fail-safe pin rule), journal. Plan carries
as-built EXECUTION NOTES.
2026-07-19 23:44:24 +02:00
Bastien Chanot 9e4ebb4cf4 feat(agents): W5 seo/geo 3-mode pipelines — collect sonnet / judge opus pin / template sonnet (BDR-077)
seo-analyzer + geo-analyzer gain MODE: collect|judge|template around the
dispatcher (mode-based, zero body-text moves — seo-data fetch-wiring
locks survive; opus pin kept = fail-safe direction, a forgotten override
over-tiers but never downgrades judgment). Run-scoped gitignored
signals handoff (.audit/*-signals-<RUNID>.md + COLLECTION COMPLETE
sentinel), judge fails closed on absent/mismatched/unsealed signals.
/seo rewired to 3 phases (domains parallel per phase) + DISPATCHER ERROR
CONTRACT (mute/ERROR judge never carried into templating; retry once,
escalate); /geo same single-domain; legacy no-MODE single-shot kept on
the opus pin for /harden narrow-scope + /onboard report-only. Dropped
/geo's 'ask and I relay' fiction (dispatched agents cannot ask).
In-wave smokes PASSED disk-verified: collect signals+sentinel; judge
ERROR-verdict on wrong RUNID; real judge = honest N/A + deterministic
engine + full scoring grammar; template = complete envelope + verbatim
sentinel + zero re-derivation. Census §18 (125 pass — one vacuous
line-wrapped lock caught by the census itself and fixed), make test
exit 0.
2026-07-19 23:23:10 +02:00
Bastien Chanot d2a10de08b feat(agents): W4 handover two-mode — synthesize opus / render sonnet, run-scoped draft handoff (BDR-077)
handover-doc-writer: MODE synthesize (model="opus" call-site — STEP
9/10/12 → .audit/handover-draft-<RUNID>.md + DRAFT COMPLETE sentinel) /
MODE render (sonnet pin — STEP 13-16 from the draft, fail-closed on
absent/mismatched RUNID). Mode-based, not a file split: the §9
name+dispatch census locks survive untouched. client-handover-writer 9.6
dispatches twice with the FULL PACKAGE both times (LRN-126) + RUNID mint
+ post-run draft cleanup. In-wave smokes PASSED disk-verified: draft
written+sentinel+gitignored; render BLOCKED on wrong RUNID (no phantom
synthesis); render consumed draft + honored skip-write, report grammar
intact. Census §17 (111 pass), make test exit 0.
2026-07-19 22:59:34 +02:00
Bastien Chanot 5e8bb0c22e feat(routing): W3 tier moves — validator-analyzer opus→sonnet, commit-changer propose=opus override (BDR-077)
validator-analyzer tiered down (deterministic validator-runner + fixed
deduction tables — no deep judgment; supersedes its BDR-076 opus pin).
commit-changer: MODE propose dispatched model="opus" (narrative
reconstruction + capitalize routing = judgment), MODE apply on the sonnet
pin. Typed-pin precedence smoke PASSED: sonnet-pinned verifier dispatched
model="haiku" ran on claude-haiku-4-5 — call-site wins, documented +
now behaviorally proven. Census §11 flip + §16 (106 pass), make test
exit 0.
2026-07-19 22:45:00 +02:00
Bastien Chanot 18075a38db feat(agents): W2/S2 doc pipeline two-mode + last inline conversions (BDR-077)
doc-syncer: ONE agent, TWO dispatch modes around the dispatcher's gate —
MODE: audit (model="opus" call-site override, READ-ONLY, drafts + PATCH
PLAN) / MODE: patch (sonnet pin, applies the APPROVED plan, shape oracle
w/ revert-on-fail, emits CHANGE SUMMARY + PATCHED_FILES). Deviation from
plan's 2-file split, per the challenge's own commit-changer mode
precedent: zero text duplication, zero lock moves. Fixes a LATENT DEFECT:
/doc dispatched an agent whose STEP 8 gate could never fire (dispatched
agents cannot ask) — the gate now lives in the dispatcher (DISPATCHER
PROTOCOL section). doc-commit.md consumes the patcher's CHANGE SUMMARY
(the in-thread context now crosses the dispatch boundary, LRN-126).
Consumers rewired: /doc (audit→gate→patch→commit), onboard (audit
report-only, opus), doc-commit steps in bugfix/hotfix/feat/ship-feature/
init-project(5b+10c); scaffolder loses PHASE 6 (README = init 5b's job);
scaffolder + onboarder now DISPATCHED in init-project/onboard (pins live,
was inline on session model). In-wave planted-drift smoke PASSED
end-to-end, disk-verified (audit caught npm-run-dev drift → [MINOR] plan
→ patch applied → summary crossed). Census §14-15 (103 pass), make test
exit 0. Typed plugin-probe dispatch resolution verified post-restart.
2026-07-19 22:28:09 +02:00
Bastien Chanot 74528a6910 feat(agents): W2/S1 plugin split — probe (sonnet) + advisor reasoner (opus) + plugin-gate include (BDR-077)
plugin-advisor keeps its name, becomes the opus REASONER: PHASE 1 bash
extracted to new plugin-probe (sonnet, facts-only PROBE REPORT), PHASE 4
apply + checkpoint hoisted to new lib/plugin-gate.md (main-loop include,
doc-commit.md x6 pattern). Fail-closed: advisor ERRORs on missing report.
4 consumers rewired (plugin-check, onboard, init-project, ship-feature).
In-wave planted-input smoke PASSED: probe report complete w/ fallbacks;
advisor consumed every planted field (monorepo per-package note, fast-libs
ctx7 reco) with zero re-detection; ERROR verdict on absent report.
Census §13 (81 pass). Note: new subagent_type registers next session —
resolution re-check before wave merge.
2026-07-19 20:57:10 +02:00
Bastien Chanot 3f7c754239 feat(routing): W1 no-inherit — fable skill-runners, opus review dispatches, dispatch-tier doctrine (BDR-077)
No dispatched agent inherits the session model anymore:
- client-handover-writer's 7 general-purpose skill-runner dispatch sites
  carry model: "fable" (+ normative rule; spike-verified alias — resolves
  claude-fable-5, enum-validated, loud failure, never silent fallback)
- ship-feature STEP 6 + init-project STEP 10 code-review dispatches carry
  model: "opus" (was: inherit — the leak the maps exposed)
- model-gate.md §4: dispatch-tier doctrine (typed = frontmatter pin,
  built-ins = explicit model= at every call site)
- census §12 (66 pass), make test green
2026-07-19 19:51:06 +02:00
Bastien Chanot 354ff2644f feat(agents): pin dispatched judgment agents to opus — Fable = inline reflection only (BDR-076)
Reverses the BDR-066 rejected alternative (opus pins on audit agents):
session default is now Fable, so inherit burned Fable quota on every
dispatched audit/challenge. analyzer, plan-challenger, seo/geo/
validator-analyzer pinned model: opus; onboard's 6 general-purpose
audit dispatches carry model="opus"; tour Phase B repointed.
interviewer + client-handover-writer stay unpinned (inline-load only,
a pin there is inert). settings.json default: claude-fable-5[1m].
Census flipped: model-routing §3 + new §11 (61 pass), loops-light 35,
full make test green.
2026-07-19 17:38:55 +02:00
Bastien Chanot 6bfc0543e5 feat(skills): add 3-way adversarial plan-challenge phase to reflection orchestrators
After a plan/reflection is elaborated and before it executes, three fresh blind
sub-agents (correctness / robustness / simplicity) attack it on the big model;
the main loop RE-THINKS every aspect a BLOCKER lands (a named plan change, or
[deferred]) and re-challenges once if the plan materially changed. Advisory into
each skill's existing human gate — the human stays the decider.

- lib/challenge-plan.md — reusable phase: fail-safe (never fail open),
  severity-driven (any single-lens BLOCKER = must-address), RE-THINK loop
- agents/plan-challenger.md — challenger role (read-only, big-model per BDR-066)
- lib/tests/plan-challenger.test.sh — 41-assertion structure lock
- wired into 11 orchestrators: ship-feature/init-project/feat/bugfix (build-plan),
  onboard/audit-delta/code-clean (proposals), seo/geo/harden/web-validate (fix-bundle)

Hardened by dogfooding: 3 blind challengers reviewed this feature's own v1 plan
and caught 4 BLOCKERs (fail-open, consensus-buries-lone-finding, wrong model
tier vs BDR-066, false on-disk-plan premise) — all fixed here.
2026-07-17 22:51:50 +02:00
Bastien Chanot b271e83fb6 feat(seo-data): content_quality verb — deterministic filler/AI-slop signal
Cherry-picked from claude-seo (github.com/AgriciDaniel/claude-seo, MIT)
content_quality.py, rewritten to the lib/seo-data contract per BDR-070. The
Content Shape axis was 100% LLM judgement; this gives it a measured input.

fetch.sh content_quality (stdin or --file) → {filler_score, ai_pattern_score,
information_density, overall_quality, flags[], matches{}}. 100% deterministic:
QRG §4.6 filler list (26 phrases) + AI-pattern list (46) kept intact, regex
matching, no LLM. Stdlib only (argparse/json/re/sys/collections/typing).

Advisory, NOT a verdict — the point of the wiring. It never claims a page "is
AI-written" (LRN-131/133); flags are candidates for human review. geo-analyzer
STEP 8 Check 10 makes it a deterministic input that INFORMS checks 1-9, never
replaces them, never scored on its own. A low number is not an automatic
finding.

Detection proven both directions (a detector that always- or never-flags is
useless): filler+slop text → flags [filler, low-density], overall 34-49; clean
dense factual text (dates/EUR/percentages) → no flags, overall 90. Empty input →
degraded/empty_input, never zeros-as-a-result.

Verified: GATE 1 verifier CONFORME 10/10 (both directions exercised live, lists
diffed intact vs source, advisory language confirmed); GATE 2 self-scan clean
(only sink is read-only open() for --file); seo-data 190 → 210 pass, 0 fail;
full suite green; shellcheck + py_compile clean.
2026-07-17 19:06:58 +02:00
Bastien Chanot cfdd89e73b feat(seo-data): schema_gen verb — generate JSON-LD, not just audit it
Cherry-picked from claude-seo (github.com/AgriciDaniel/claude-seo, MIT)
schema_generate.py, rewritten to the lib/seo-data contract per BDR-070 — adapt,
never copy. The system audited JSON-LD but could not generate it; geo-analyzer's
G2 batch hand-wrote markup. Now it calls the verb.

fetch.sh schema_gen {reservation|order|discussion|profile} → fail-open envelope
{"status":"ok","source":"schema_gen","type":…,"jsonld":{…}}. Types: Reservation
(7 subtypes), OrderAction, DiscussionForumPosting, ProfilePage (sameAs/knowsAbout
for the entity graph). Stdlib only (import argparse, json — zero third-party),
_strip_nones so a null is never emitted, --script-tag wraps for direct paste.

Fail-open mirrors score.py's _cli exactly (the contract's named pattern): a
flag-omitted required field → argparse exit 2 + {"status":"error","bad_usage"};
a flag-present-but-empty field → {"status":"degraded","reason":"missing required
field: …"} exit 0. Never a traceback, never empty stdout (LRN-133: the
can't-generate case stays legible).

geo-analyzer G2 wiring preserves the data-integrity rule — the verb generates
STRUCTURE, unknown values stay [À COMPLÉTER], never invented.

Verified: GATE 1 verifier CONFORME 10/10 (exercised the fail-open edge cases +
diffed field surface against the source); GATE 2 self-scan clean (no
network/shell/secret/eval sink); seo-data 167 → 190 pass, 0 fail; full suite
green; shellcheck + py_compile clean.
2026-07-17 14:30:31 +02:00
Bastien Chanot 4818c6116f feat(seo-data): I7 — compute the score instead of feeling it
/harden has a real scale (SKILL.md:435 — Critique -15, Haute -8, Moyenne -3,
Basse -1, clamp [0,100]). /seo had none: every axis was felt, so two runs over
identical code could disagree. That is a credibility problem on its own, and
/client-handover gates on 17/20 — a wobbling number makes the gate arbitrary.
H2 sharpened it: now that drift reports what actually changed, a score moving
on its own is visibly noise.

The split is the whole point. WHICH findings exist and how severe each is
stays the LLM's judgement — irreducible, and I am not pretending otherwise.
The arithmetic stops being judgement: same findings in, same score out. Same
principle as grouping cannibalisation rows in the engine rather than handing a
model 1000 rows to add up.

Reuses /harden's scale, /5 into /20, so the family speaks one vocabulary
instead of two.

Two things it makes real that were prose:
- **N/A is not a zero.** R2 (client-rendered on-page) and I1 (unauditable
  off-page) both mandate excluding an axis and renormalising the rest. Both
  left that arithmetic to the model. Now the engine does it and refuses to let
  N/A behave like a zero — verified: all-20 axes with two N/A still yields
  global 20.0, not a dragged-down mean.
- **Prevalence.** affected/sampled shift severity ONE step (>=50% escalates, a
  single page de-escalates). A defect on 1 of 12 pages is not the defect on
  12 of 12, and flattening the two is part of what made the old numbers move.

Malformed input is an error, never a silently wrong number — unlike the fetch
verbs, a degrade here would mean bad input, not a network fact. Unknown
severity and unknown profile both rejected, tested.

Verified: hand-checkable arithmetic (haute+moyenne = 100-11 = 89 → 17.8;
critique+haute = 77 → 15.4), identical global across repeated runs, weights
renormalised to sum 1.0 with two axes N/A. seo-data 155 -> 167 pass, 0 fail;
full suite green; shellcheck + py_compile clean.
2026-07-17 13:29:14 +02:00
Bastien Chanot d6b8edc8ea fix(seo): B1 KILLED — Common Crawl backlinks measured, not assumed
The plan said Common Crawl was the free backlink source and the 70/100 cap
was therefore mandatory. Measured before building, and both premises die.

HEAD against data.commoncrawl.org, live:
  cc-main-2026-feb-mar-apr-domain-edges.txt.gz    17.3 GB   gzipped
  cc-main-2026-feb-mar-apr-domain-ranks.txt.gz     2.3 GB
  cc-main-2026-feb-mar-apr-domain-vertices.txt.gz  879 MB

Finding one domain's inbound links means scanning the edges file end to end,
per audit. That is not slow, it is non-viable — and abusive toward a
nonprofit serving the data free.

Worse, the reference implementation everyone points at
(claude-seo scripts/commoncrawl_graph.py:169) does this:

    max_compressed_bytes = 500 * 1024 * 1024   # 500 MiB safety cap
    if total_downloaded > max_compressed_bytes: break

500 MiB of 17.3 GB is **2.9% of the edges file**, which is sorted by source
ID — so it reads an arbitrary slice of source domains and reports whatever
backlinks happened to be in it, as a backlink profile, capped at "70/100
health". Nothing in the output says 3%. That is a random sample wearing a
measurement's clothes: the exact failure class this branch exists to remove,
and I was one step from copying it.

B2 dies with B1: nothing left to cap.

CONSEQUENCE, and it is the point: I1's narrowed Off-page axis — brand
mentions only, backlinks + authority declared unauditable in §14 — is the
FINAL state, not a placeholder waiting for data. Corrected my own I1 text,
which pointed at Common Crawl as the "nearest free source": that sends a
future reader into a 17 GB dead end. The §14 line now records what was
measured and why no number beats a fabricated one.

Also corrects the B3 note, whose follow-on ("so Common Crawl is the only free
source") was wrong for the same reason. The only free viable backlink source
is Bing's GetUrlLinks — first-party only, never a competitor, and blocked on
the client's Bing account. That raises W2's value; it does not unblock it.

Verified: full suite green, seo-data 144 pass / 0 fail.
2026-07-17 13:17:18 +02:00
Bastien Chanot 20d3082542 feat(seo-data,seo,geo): R2 — refuse to score what JS paints; no Playwright
Arbitrated (user): honest refusal on SPA, no headless browser.

STEP 2 has recorded `RENDERING: SSR/SSG/SPA/hybrid` since forever and NOTHING
ever acted on it. (The inventory claimed a "SPA severely limited" §0 flag
compensated — it does not exist. Seventh subagent claim this branch has had to
disprove.) So on a client-rendered site the FULL audit curls an empty shell,
every meta/H1/JSON-LD check reports "missing", and the agent emits a page of
false findings — plus a bundle that would "fix" tags which already exist.

rendercheck reads the verdict from what the server SENT. package.json cannot
tell a React SPA from a Next.js SSR app; the served bytes can. Stdlib only.

The refusal is the point:
- client-rendered → On-page is N/A, excluded from the weighted global, NOT
  scored zero. A zero says "your on-page is bad"; N/A says "we could not see
  it". Only one is true, and /client-handover gates on this number.
- No bundle item may come from a live on-page check on such a site.
- The report still says what IS auditable (robots, sitemap, headers, config,
  CrUX field data — real users, hydration included — GSC, legal, images)
  rather than returning an empty verdict.
- geo refuses Content Shape the same way, and states the sharper fact: AI
  crawlers are WORSE at JS than Googlebot. GPTBot/PerplexityBot/ClaudeBot
  fetch HTML and largely do not execute it, so a client-rendered site is not
  merely unauditable by us — it is near-invisible to the engines this audit
  exists to serve. §0 alert + SSR/SSG as the top user action.

Script/style text is not page text: a React shell with a fat inline
window.__INITIAL_STATE__ measures 7 chars. Without that skip a 200 KB bundle
reads as a rich page — the detector would fail exactly where it matters.

Verified on both extremes, not just the happy path: zenquality 7650 chars/1
h1/9 jsonld and lavageangels356 13973/1/1 → server-rendered, no warning; a
Vite/React shell fixture → client-rendered, 7/0/0, warned.

seo-data 136 -> 144 pass, 0 fail; full suite green; shellcheck + py_compile
clean.
2026-07-17 12:44:43 +02:00
Bastien Chanot fe41986be9 feat(seo-data): C3 — internal link graph; orphans + click depth, measured
seo-analyzer.md:613 asks "Every important page reachable within 3 clicks?"
and :616 asks "Orphan pages (no inbound internal links)?". Neither ever had a
command — same shape as the sameAs check before W3. This is that command.

My earlier reservation ("costs a lot of network") was wrong and the
measurement killed it: 24 pages in 2.7s, 86 in 3.8s. Cheap enough to always
run on FULL.

EXHAUSTIVE OR NOTHING is the design constraint, not a nicety. Orphans cannot
be sampled: proving a page has no inbound link means having read every other
page. So when the crawl is capped or any page fails, orphans are WITHHELD —
`orphans_withheld: true` and no list. A false orphan ("page X has no inbound
links" when it does) sends a client fixing what is not broken; that is the
worst finding this tool could emit. The cap does not degrade the result, it
invalidates it.

SPA refusal: on a client-rendered site the links are not in the HTML and
every page reads as orphaned. That is catastrophic, so an empty graph returns
degraded/no_links_in_html instead of a full false-positive list. No JS
rendering by design — that is the R1/R2 arbitration, not something to smuggle
in here.

Verified against BOTH live sites and against a planted failure, because two
clean results are not evidence a detector detects:
- native PHP: 24 pages, 335 links, depth 2, 0 orphans
- Astro: 86 pages, 2015 links, depth 2, 0 orphans
- fixture with a planted orphan + a 4-click chain: both found. Filters proven
  on real shapes seen live — /css/main.css?v=1778157313, #anchors, mailto:,
  tel:, external hosts, .png. /b/ in markup vs /b in sitemap unify to one node
  rather than a phantom orphan pair.

Fixed a flaw in my own mock while writing that test: a single page.html
fixture cannot express a GRAPH (every node gets identical links), so the mock
is now pages.json = {url: html}.

Verified: seo-data 122 -> 136 pass, 0 fail; full suite green; shellcheck +
py_compile clean.
2026-07-17 12:32:42 +02:00