Commit Graph
178 Commits
Author SHA1 Message Date
bchanot 6104545e76 feat(skills): push state read from facts, never pushed by the skills
Run C2 of manual-push mode (BDR-111/BDR-112). The four flows that pushed
on their own, or claimed the branch was on origin, now read the truth
after the fact and hand the user the exact command:

- client-handover-writer: the "Push to origin now?" question and its
  push block are gone (the hooks had already pushed in auto-push mode;
  push-guard denies it in manual mode). A reusable PUSH STATE READ
  (branch, origin probe, `git rev-list --count origin/<br>..<br>`, the
  verb only to word the reason) runs after commit-change, at the top of
  the deploy pause, after "Deployed" and before each end report. The
  branch name is validated against an allowlist before it is placed in
  any command or hint (a hostile branch name is otherwise a shell
  injection). Pending → the user pushes BEFORE the deploy pause; the
  deploy brief says "after your push". `Push:` line in both reports.
- release-candidate STEP 6: two ahead counts + the verb; anything other
  than auto with both counts 0 prints one user command
  `! git push --atomic origin main develop v<X.Y.Z>` and stops; the tag
  gate stays for auto mode; `hold` notes --follow-tags; version regex.
- release-executor: push claims qualified (auto-push mode, best effort).
- tour: mode-agnostic rule; STEP 3 reads one `git -C <project>` fact per
  project (suffix-aware branch, --remotes=origin, origin probe) and the
  summary row says on origin / local only with the user command.
2026-10-07 14:02:51 +02:00
bastien 535022186c docs(plugin-advisor): never recommend the Higgsfield toggles from project signals 2026-09-30 18:17:32 +02:00
bastien 58c3a3e9b7 fix(effort): re-raise judgment dispatches, planning re-asserts, pairing caveat, dedupe audit script (final review I1-I3) 2026-09-28 20:48:27 +02:00
bastien 98ef991958 docs(effort): BDR-107 id, CHANGELOG entry, spec corrected for the rulings (vendored pins, exclusions, pairing rule) 2026-09-28 20:20:05 +02:00
bastien a117e7ed67 fix(effort): shifts are sent with the step's first tool call (harness pairing rule); challenge shifts under their heading; fence indentation 2026-09-28 19:55:47 +02:00
bastien 3c58160d0c feat(effort): wire phase shifts in the 13 orchestrators and the handover writer 2026-09-28 19:44:22 +02:00
bastien 9223fda99f feat(effort): pin effort on the 20 repo-authored agents (BDR-077 second axis) 2026-09-28 18:59:49 +02:00
bastien ddea411491 chore(config): superpowers citers by bare name, routing map, docs, settings
Every superpowers-prefixed skill call in ship-feature, init-project, tour,
deploy, audit-delta, plugin-advisor and lib/analyze-before-plan now names
the vendored skill directly. finishing-a-development-branch is described
as the upstream skill this config does not vendor (gitflow finish is the
integration path). CLAUDE.global.md Skill routing maps the four
non-vendored skills the vendored text still references. settings.json
loses the plugin key and its marketplace block; README, USAGE,
plugin-advisor and the profile skill describe superpowers as vendored
skills, always on, zero plugin cost. CHANGELOG entry with a known
residual.
2026-09-28 14:54:53 +02:00
bastien 4c86d6dc70 chore(config): security-guidance Stop review off, plugins off, routing and docs
settings.json: ENABLE_STOP_REVIEW=0 (the plugin's own switch: no more
Opus call on every turn that changes code, 0 findings in 6 days, 1
recorded false positive; the regex layer and the commit/push agentic
review stay on), brightdata-plugin@synced false (keyless-useless, its MCP
skill would hijack WebFetch/WebSearch), frontend-design official plugin
entry gone (uninstalled: byte-identical to the managed copy).

CLAUDE.global.md routes Ship/PR to ship-feature (gstack ship takes
origin/HEAD = main as base), drops ship/context-save from the gstack-off
list and 21st-ui-review from the design review line (trio is max-only).
deploy's table no longer points at land-and-deploy/setup-deploy.
plugin-advisor.md describes security-guidance's real mechanics. CHANGELOG
Unreleased entry with a Known residual section.
2026-09-28 11:49:01 +02:00
bastien 6617889b77 feat(verifier): floor-guard waivers outside test files need a CLARIFICATIONS ack
Security-gate MEDIUM: a self-service floor-guard: allow <reason> neutralised
the detector in the same commit. User chose strict: the tool prints WAIVED,
the contract authorizes, the verifier counts the rest as gaps. BDR-102
amendment.
2026-09-27 21:20:36 +02:00
bastien 2b25cb4704 feat(lib): floor-guard, diff-scoped detector of a weakened quality bar
SUPPRESS / SKIP / DELETED_TEST / ASSERT_DROP / STUB / THRESHOLD_DOWN over
git diff <base> (untracked files included), floor-guard: allow <reason>
waiver printed as WAIVED, rc 0/2/3. Mandatory verifier STEP 3, documented
under GATE 1 of verify-secure-loop.md. Suite: 6 kinds + WAIVED + CLEAN,
flip-tested. Adapted from agent-skills constraint-driven-development.
2026-09-27 20:17:36 +02:00
bastien 0d035fcab6 feat(profile): default profile = full; reset applies it, current is label-driven
No profile selected (.active-profile absent, empty or legacy "none") now
means the `full` profile is in force: DEFAULT_PROFILE declared once in
lib/profile.sh, resolved by active_profile(); the statusline reads the
constant and shows `full` instead of `?`; `gstack off` trims to it instead
of erroring. `reset` goes to the default profile (= `set full`: enables its
list, parks any non-listed gstack or managed item). `current` names the
active label and scores that profile only, saying `default — not applied
yet` until a set/apply/reset wrote the cache; the "none" sentinel and the
cross-profile best-guess scan are gone (they keyed on the parked-gstack
count, which says nothing under BDR-030's gstack-off default). Hermetic
suite lib/tests/profile-default.test.sh (29 checks) seeds gstack as OFF like
a real tree. Citers updated: profile SKILL, Makefile help, plugin-advisor
PROFILE line + reset paragraph, toggle-external header.
2026-09-25 16:28:21 +02:00
bastien 27f201d4aa feat(guardrails): refusal ends the attempt; doctrine-citers census; make test suite=
Root causes of the 2026-09-24 errors turned into mechanisms (BDR-100). hard_deny 'Routing around a guardrail': a refused command is never rerun through a wrapper, alias, heredoc, Makefile target, env file, other shell or other agent; the same clause in 14 agents and in the doctrine's sub-agent rule. make test suite=<file> runs one suite hermetically so the denied env-prefix form is never needed by hand. lib/tests/doctrine-citers.test.sh: every CLAUDE.md "Section" / § Label citation across skills, agents, lib, rules and hooks must resolve to a heading or bold label (flip-tested); its first run fixed rest-api-node.md. Doctrine 'After code changes' step 4: a changed rule, heading, label or threshold → grep every citer in the same commit.
2026-09-24 20:58:25 +02:00
bastien d82c06f572 refactor(doctrine): C2 coherence — 30 doctrine/skill tensions resolved, doctrine wins (BDR-099)
One ask policy; mandated executors exempt from the delegation rule; skill plan satisfies the planning rule; journal line exempt from the approval gate; chore = maintenance without new behaviour; small fix on develop = bugfix; BDR-068 written as the one auto-finish exception; deploy routes to /deploy. Skills and agents follow: hotfix types by base + skips the design gate on trivial; capitalize/close create missing registries; commit-change asks the branch type; doc/seo/web-validate/refactor branch through the aiguillage; tour reports BREAKING fixes as needs-decision and runs doc-syncer two-mode; client-handover applies audit bundles from its main loop behind one gate; init-project/onboard use the 200-file graphify signal and bootstrap memory; release-candidate gates the tag push only; push wording aligned with the BDR-095 hooks; stale pointers fixed (§ Language, .gsd/ROADMAP.md, handover script path, design-gate lists).
2026-09-24 20:25:40 +02:00
bastien c81b1731af feat(graphify): threshold signal from 200 tracked code files, the banner informs and the user decides
lib/graphify-gate.sh counts tracked code files (graphify's AST extension set, vendored trees excluded) and, from 200 with no graphify-out/graph.json, prints one banner-sized line; session-start shows it with the /graphify hint. Nothing is built, installed or updated: the rule is the user's (BDR-097), grounded in the LRN-162 measurements (AST build 2.3 s, 0 tokens, a query 2 to 3k tokens). Doctrine section and plugin-advisor thresholds follow the same rule; graphify claude install stays rejected. Test: 11 checks. GRAPHIFY_MIN_CODE_FILES overrides the threshold.
2026-09-24 12:12:16 +02:00
bastien 9da5d8d52c feat(guardrails): push every commit, static deny for destructive tools, brief carries no user authority
Layer C of the plan written after the 2026-09-21 wipe (BDR-095): a reviewer
sub-agent traced `lftp mirror --delete` against a local file:// tree, the
prose tiers named neither lftp nor a local trace, the brief had authorized
it, and four days of commits had never left the machine.

- gitflow: `start` pushes the branch with its upstream, merge targets are
  pushed after each merge, and `init`/`install-hook` write post-commit and
  post-merge hooks that push every commit as it lands (warn, never block;
  GITFLOW_NO_PUSH=1 for throwaway repos). T18 + T19 (installed == emitted).
- hooks/unpushed-guard.sh on SessionStart and Stop: branch ahead of its
  upstream, no upstream, or no origin. Non-blocking systemMessage.
- settings.json: static deny for transfer and mirror tools, rsync --delete,
  xargs rm, pipe-to-shell, chmod/chown -R, sudo/doas/pkexec, disk tools,
  chattr, docker volume drops/prune/--privileged/socket/-v /:, git history
  destruction, --no-verify and core.hooksPath; new hard_deny "destructive
  tool against a local path, brief carries no user authority"; soft_deny
  reworded + discarding uncommitted work; environment records the incident.
- CLAUDE.global.md "Destructive tools & data loss"; the four report-only
  agents trace by reading, never by running, whatever the brief says.
- lib/tests/guard-bash.test.sh: executable spec of the PreToolUse guard
  (214 cases). The hook itself is not shipped (BLK-022); the spec skips.
2026-09-22 07:43:12 +02:00
bastien 7c05f75eab feat(21st): replace the magic MCP with the @21st-dev CLI + skill pack
Upstream supersedes `@21st-dev/magic` with `@21st-dev/cli` (bin `21st`):
same endpoint, `21st login` in place of an API key, no MCP process loaded
into every session.

- install-plugins.sh Step 8.7: `npm i -g @21st-dev/cli` (pinned in
  plugins.lock.json), staged `21st skills install`, TTY-only login offer,
  pack disabled by default. update-all.sh 7.4 refreshes both.
- The documented `21st install-skill` cannot be used: the installer refuses
  to follow a symlink on the target path and `~/.claude/skills` is one. The
  install runs under a throwaway HOME and the result moves into
  skills-external/21st-* (gitignored), symlinked on demand.
- toggle-external.sh manages `21st` as a pack (names globbed from
  skills-external/21st-*, parked under plain names). `magic` is gone.
- The 5 design skills join design/web/web-full/full and MANAGED_EXTERNALS;
  21st-registry and 21st-design-sync stay parked. MANAGED_MCPS is now empty
  and profile.sh's dead magic branches are removed.
- Design gate: GATE-BLOCK gains `21st` (required-manual, magic's old slot)
  and `21st-ui-build`; PATH repair extended to the npm global bin.
- settings.json: the 4 mcp__magic__* ask entries go; the outward-facing
  21st verbs land in autoMode.soft_deny, the tier that holds under auto
  mode (LRN-153).
- Docs: README, CLAUDE.global.md, design-gate.md, profile SKILL.md,
  .env.example, .gitleaks.toml, link.sh. BDR-093, LRN-158.

Tests: profile-set-managed 17/17, make test green except 2 pre-existing
gitflow FAILs (gitleaks binary absent on this host), shellcheck clean.
2026-09-22 02:53:31 +00:00
Bastien Chanot 17370d7e4c feat(executors): NEED-DECISION and BLOCKED carry a CLASS tag 2026-09-16 22:14:13 +02:00
Bastien Chanot 7cc95952bd feat(interviewer): a visible or public choice is asked, never assumed 2026-09-16 22:13:50 +02:00
Bastien Chanot 6eac7fbca9 fix(hotfix): batch-skeptic residuals — RULES restore mandate file-scoped; hotfixer FILE(S) marks created files (new) 2026-08-26 22:28:31 +02:00
Bastien Chanot ad4985f410 fix(security-auditor,close): hotfix no-verifier carve-out documented; close enumerates STEP 5C + --no-push passthrough 2026-08-26 21:59:08 +02:00
Bastien Chanot c983f1ff94 fix(handover-writers): stale ch.4 refs post-NAP-renumbering (glossary/tone->6, cross-links/THRESHOLD->5); 14.5 verification deferred to post-write; anchor gate ordered into STEP 16 2026-08-26 21:55:13 +02:00
Bastien Chanot 27f17939d7 fix(plan-challenger): ERROR verdict added to the load-bearing OUTPUT grammar (STEP 1 emitted it, parser enum omitted it) 2026-08-26 21:54:35 +02:00
Bastien Chanot 9428b86880 optimize status-reporter r2: d5 — passive cost sourced from doctor.sh constants (skeptic's find), count-only fallback kept 2026-08-26 19:44:03 +02:00
Bastien Chanot a3f1624b15 optimize status-reporter: d5 — unproducible token field replaced by /plugin-check deferral; dead ROADMAP.md row rewritten for post-ADR-013 gsd layout 2026-08-26 19:35:46 +02:00
Bastien Chanot e9c6bf52fa optimize analyze-system: d1 triggers in skill description + d2 ordered TASKS mapped to OUTPUT sections in analyzer 2026-08-26 19:22:36 +02:00
Bastien Chanot 080d2d9f03 optimize plugin-probe+advisor: d8 — FRAMEWORK-DEPS exact dep@version (no preact false-hit, fallback fires), advisor derives frontend/fast-libs from it, PLAN echoed-or-unknown (no invention) 2026-08-26 19:16:09 +02:00
Bastien Chanot c0a2a8069d optimize refactor-system: d4 — no-tests STOP gate (agent) + user arbitration loop (dispatcher) + mid-run test-failure revert; code-cleaner inline carve-out 2026-08-26 19:05:29 +02:00
Bastien Chanot 562a42e936 optimize onboarder: d8 — BRIEF contract split REQUIRED/OPTIONAL, draft placeholders replace blanket STOP (fixes first-dispatch bounce vs /onboard STEP 2 minimal brief) 2026-08-26 15:12:24 +02:00
Bastien Chanot 9db213b0a1 optimize interviewer: d9 — DO NOT blacklist (no design, no invented values, no budget overrun) 2026-08-26 12:35:29 +02:00
Bastien Chanot e0923d6c4b optimize interviewer: d3 — failure-mode table (vague/idk/contradiction/partial/balloon) + 2-round budget 2026-08-26 12:34:15 +02:00
Bastien Chanot 63310467ca feat(gates): deterministic floor (GATE 0) under the fresh verifier
GATE 1 is an LLM dispatch and the verifier's mandatory PROOF: line is a line
the verifier writes — nothing structurally stops it being produced without
anything being executed. Nothing deterministic sat between the executor and
that dispatch.

An acceptance criterion can now carry an oracle: indented CHECK: (command),
EXPECT: (success-only marker), EVIDENCE: (slot). lib/gates.sh runs them
fail-closed — MET requires exit 0 AND the marker, so a nonzero process never
passes on its error text carrying the token — and writes the outcome back
into the contract, so the fresh verifier reads evidence as fact rather than
trusting the executor's report.

GATE 0 runs that floor before any verifier is dispatched; a red build sends
the executor back for free, on its own iteration budget. ABANDON: <id>
<reason> turns an impossible criterion into a visible handoff that blocks
CONFORME and routes to the human gate, via the new ABANDONED(n) verdict —
a distinct token because it routes distinctly, never a dev loop. feater and
bugfixer gain a four-pass completion discipline, scoped so a pass can never
widen the contract.

The runner's parse fails closed on partial oracles, duplicate ids,
unindented attributes and runnable criteria with no EVIDENCE: line, and
executes nothing at all when the ledger is malformed. status never executes
and never writes; run always re-executes, since trusting written evidence is
the failure being closed.

Adapted from the unlazy skill (Leonxlnx/unlazy, MIT). Its Stop hook,
approval store, .unlazy/ tree, depth-tree arithmetic and Node checker were
deliberately refused — BDR-083 records each reason.

64 assertions in lib/tests/gates.test.sh, non-execution proved by sentinel
with its own positive control asserted first.
2026-08-24 13:12:38 +02:00
Bastien Chanot c7646a9c8a feat(agents): de-prescribe geo-analyzer for Opus 5 (C1 P3)
Same invariant as adafa35. Census 71/0 green; contract surface
byte-identical; 1106→1107 lines (single-shot scoping line).

- MANDATORY/MUST caps on AI-index submission → plain content rule
- 'Print the plan before STEP 13' → single-shot-scoped (conf#1);
  tier-mapping kept in BOTH judge and template ranges (rob#1/#9,
  PERMISSIVE :873 named survivor kept)
- vestigial ':1106 Transparency §14' reworded to truth: change log is
  dispatcher's SEO.md §15 (folded into 'Dispatcher verifies')
- 'copy these patterns' → reporting-shape-to-match; FAQ '20-50'
  quantity → 'typically dozens'; 'EVERY finding' → outcome bar
- dedup after inspection: ZERO merges (PERMISSIVE ×3 cross-range;
  never-apply ×4 distinct obligations; content_quality pair =
  spec rule vs emitted-artifact caveat — annex counts corrected)
- FROZEN untouched: guard-first :273, NAP direction rule, cite-sources,
  WebSearch-freshness (already when-shaped, rob#6), all STEP headers
- deltas: MANDATORY 1→0, MUST 4→3, NEVER 9→8
2026-08-02 01:30:12 +02:00
Bastien Chanot adafa350da feat(agents): de-prescribe seo-analyzer for Opus 5 (C1 P2)
Choreography → when-guidance under the audience×mode-range invariant
(plan §4b Q3). Census 71/0 + model-routing 133/0 + seo-data 221/0 green
throughout; contract surface (§2a/2b) byte-identical; 1528→1503 lines.

- self-output verification: ':970 run it twice' → deterministic-engine
  integrity guard (conf#8); ':1217 do not proceed' → single-shot-scoped
  (conf#1); grouping sanity-check → when-guidance detector (rob#8)
- vestigial pre-BDR-061 'Transparency' line deleted; §15 ownership
  folded into 'Dispatcher verifies'
- dedup (inspection-corrected: most annex 'twins' are distinct
  obligations — kept): only true same-range dups removed ('Handoff to
  dispatcher' ≈ sentinel note; 'Landing page rule' block ≈ payload
  instance + RULES line)
- completeness checklist reshaped to routing map, rows verbatim (rob#3)
- caps softened: 'P0 rule' MUST/ALWAYS → plain content rules (CMS
  plugin-first folded with STEP 2 twin, corr#3); First-action/ordering
  emphasis dropped; C1a + sampling essays compressed (rules + LRN
  citations kept); WebSearch → drifting-externals when-guidance (rob#6)
- FROZEN untouched: guard-first orderings :287, denominator-before-
  sampling :550, R2 refuse, COVERAGE obligations, all STEP headers
- deltas: 'P0 rule' 2→0, ALWAYS 1→0, MUST 5→4; NEVER 9→9 (class-B
  named bans, kept by design LRN-105)
2026-08-02 01:30:02 +02:00
Bastien Chanot c3d3f4d465 feat(agents): plan-challenger — route grounded doubts to [MINOR]
Opus 5 follows conservative-reporting clauses literally; 'a manufactured
concern is a failure' risked suppressing real low-confidence findings.
In-place reword: ungrounded stays noise, grounded-but-uncertain files as
[MINOR] with the uncertainty in WHY:. OUTPUT grammar byte-identical;
census row added.
2026-07-30 12:58:29 +02:00
Bastien Chanot b7026e4bda feat(ctx7): coverage extension — fast-libs single source + reminder hook + executor briefs (BDR-078)
- lib/fast-libs.sh: detect/cache-status verbs, JS+Python manifests,
  7-day cache freshness, LC_ALL=C sort — replaces 3 hardcoded lists
  (ship-feature 0c, init-project 5c, onboard 3.5)
- hooks/ctx7-reminder.sh: once-per-session UserPromptSubmit nudge when
  the project carries fast-libs and .ctx7-cache/ is missing/stale
- find-docs: before-writing-code trigger + cache-first rule; dist is
  machine-owned (gitignored) so the durable patch lives in
  install-plugins.sh STEP ctx7 (idempotent, grep-guarded)
- feater/bugfixer briefs: fast-lib docs rule (fresh cache read, else
  2-topic ctx7 fetch, else NOTES cache miss + proceed)
- tests: lib/tests/fast-libs.test.sh (11 checks); shellcheck + full
  make test green (review-guards 5/0)
2026-07-20 10:45:06 +02:00
Bastien Chanot 444c79acb2 fix(routing): W6 ronde — 6 findings closed (README bootstrap path, 2 census gaps, 3 stale refs)
Fresh-opus whole-chantier ronde (17fbe51..HEAD, EVAL-023 style): axes
severed-wires / gate-regressions / fail-open PROVEN CLEAN (all 7 new
handoffs traced end-to-end both sides). Fixed: init 5b now dispatches the
FULL-AUDIT path (auto-mode gated a missing README as SIGNIFICANT →
[CREATE-AUTO] unconditional restored, sole greenfield README path);
census locks added for /geo ERROR CONTRACT + never-re-derive (deleting
the fail-closed handler would have stayed green), doc-audit
model=opus override x5 flows, SYNTH REPORT grammar; doc-syncer ex-STEP-8
prose repointed to the dispatcher gate; plugin-advisor anim rows read
the PROBE REPORT ANIM field (no Bash anymore); client-handover 9.7
cleans the transient draft. Census caught one more line-wrapped lock
before it shipped vacuous. 133 pass / 0 fail, make test exit 0.
2026-07-19 23:56:58 +02:00
Bastien Chanot 07253e093c feat(doctrine): W6 — prose sweep + BDR-077 + LRN-137 + plan execution notes
LRN-113 whole-surface sweep: client-handover x2 + commit-change prose
repointed to the two-mode reality; code-cleaner/status historic notes
kept (accurate). Memory: BDR-077 (full architecture), LRN-137
(mode-based re-tiering + fail-safe pin rule), journal. Plan carries
as-built EXECUTION NOTES.
2026-07-19 23:44:24 +02:00
Bastien Chanot 9e4ebb4cf4 feat(agents): W5 seo/geo 3-mode pipelines — collect sonnet / judge opus pin / template sonnet (BDR-077)
seo-analyzer + geo-analyzer gain MODE: collect|judge|template around the
dispatcher (mode-based, zero body-text moves — seo-data fetch-wiring
locks survive; opus pin kept = fail-safe direction, a forgotten override
over-tiers but never downgrades judgment). Run-scoped gitignored
signals handoff (.audit/*-signals-<RUNID>.md + COLLECTION COMPLETE
sentinel), judge fails closed on absent/mismatched/unsealed signals.
/seo rewired to 3 phases (domains parallel per phase) + DISPATCHER ERROR
CONTRACT (mute/ERROR judge never carried into templating; retry once,
escalate); /geo same single-domain; legacy no-MODE single-shot kept on
the opus pin for /harden narrow-scope + /onboard report-only. Dropped
/geo's 'ask and I relay' fiction (dispatched agents cannot ask).
In-wave smokes PASSED disk-verified: collect signals+sentinel; judge
ERROR-verdict on wrong RUNID; real judge = honest N/A + deterministic
engine + full scoring grammar; template = complete envelope + verbatim
sentinel + zero re-derivation. Census §18 (125 pass — one vacuous
line-wrapped lock caught by the census itself and fixed), make test
exit 0.
2026-07-19 23:23:10 +02:00
Bastien Chanot d2a10de08b feat(agents): W4 handover two-mode — synthesize opus / render sonnet, run-scoped draft handoff (BDR-077)
handover-doc-writer: MODE synthesize (model="opus" call-site — STEP
9/10/12 → .audit/handover-draft-<RUNID>.md + DRAFT COMPLETE sentinel) /
MODE render (sonnet pin — STEP 13-16 from the draft, fail-closed on
absent/mismatched RUNID). Mode-based, not a file split: the §9
name+dispatch census locks survive untouched. client-handover-writer 9.6
dispatches twice with the FULL PACKAGE both times (LRN-126) + RUNID mint
+ post-run draft cleanup. In-wave smokes PASSED disk-verified: draft
written+sentinel+gitignored; render BLOCKED on wrong RUNID (no phantom
synthesis); render consumed draft + honored skip-write, report grammar
intact. Census §17 (111 pass), make test exit 0.
2026-07-19 22:59:34 +02:00
Bastien Chanot 5e8bb0c22e feat(routing): W3 tier moves — validator-analyzer opus→sonnet, commit-changer propose=opus override (BDR-077)
validator-analyzer tiered down (deterministic validator-runner + fixed
deduction tables — no deep judgment; supersedes its BDR-076 opus pin).
commit-changer: MODE propose dispatched model="opus" (narrative
reconstruction + capitalize routing = judgment), MODE apply on the sonnet
pin. Typed-pin precedence smoke PASSED: sonnet-pinned verifier dispatched
model="haiku" ran on claude-haiku-4-5 — call-site wins, documented +
now behaviorally proven. Census §11 flip + §16 (106 pass), make test
exit 0.
2026-07-19 22:45:00 +02:00
Bastien Chanot 18075a38db feat(agents): W2/S2 doc pipeline two-mode + last inline conversions (BDR-077)
doc-syncer: ONE agent, TWO dispatch modes around the dispatcher's gate —
MODE: audit (model="opus" call-site override, READ-ONLY, drafts + PATCH
PLAN) / MODE: patch (sonnet pin, applies the APPROVED plan, shape oracle
w/ revert-on-fail, emits CHANGE SUMMARY + PATCHED_FILES). Deviation from
plan's 2-file split, per the challenge's own commit-changer mode
precedent: zero text duplication, zero lock moves. Fixes a LATENT DEFECT:
/doc dispatched an agent whose STEP 8 gate could never fire (dispatched
agents cannot ask) — the gate now lives in the dispatcher (DISPATCHER
PROTOCOL section). doc-commit.md consumes the patcher's CHANGE SUMMARY
(the in-thread context now crosses the dispatch boundary, LRN-126).
Consumers rewired: /doc (audit→gate→patch→commit), onboard (audit
report-only, opus), doc-commit steps in bugfix/hotfix/feat/ship-feature/
init-project(5b+10c); scaffolder loses PHASE 6 (README = init 5b's job);
scaffolder + onboarder now DISPATCHED in init-project/onboard (pins live,
was inline on session model). In-wave planted-drift smoke PASSED
end-to-end, disk-verified (audit caught npm-run-dev drift → [MINOR] plan
→ patch applied → summary crossed). Census §14-15 (103 pass), make test
exit 0. Typed plugin-probe dispatch resolution verified post-restart.
2026-07-19 22:28:09 +02:00
Bastien Chanot 74528a6910 feat(agents): W2/S1 plugin split — probe (sonnet) + advisor reasoner (opus) + plugin-gate include (BDR-077)
plugin-advisor keeps its name, becomes the opus REASONER: PHASE 1 bash
extracted to new plugin-probe (sonnet, facts-only PROBE REPORT), PHASE 4
apply + checkpoint hoisted to new lib/plugin-gate.md (main-loop include,
doc-commit.md x6 pattern). Fail-closed: advisor ERRORs on missing report.
4 consumers rewired (plugin-check, onboard, init-project, ship-feature).
In-wave planted-input smoke PASSED: probe report complete w/ fallbacks;
advisor consumed every planted field (monorepo per-package note, fast-libs
ctx7 reco) with zero re-detection; ERROR verdict on absent report.
Census §13 (81 pass). Note: new subagent_type registers next session —
resolution re-check before wave merge.
2026-07-19 20:57:10 +02:00
Bastien Chanot 3f7c754239 feat(routing): W1 no-inherit — fable skill-runners, opus review dispatches, dispatch-tier doctrine (BDR-077)
No dispatched agent inherits the session model anymore:
- client-handover-writer's 7 general-purpose skill-runner dispatch sites
  carry model: "fable" (+ normative rule; spike-verified alias — resolves
  claude-fable-5, enum-validated, loud failure, never silent fallback)
- ship-feature STEP 6 + init-project STEP 10 code-review dispatches carry
  model: "opus" (was: inherit — the leak the maps exposed)
- model-gate.md §4: dispatch-tier doctrine (typed = frontmatter pin,
  built-ins = explicit model= at every call site)
- census §12 (66 pass), make test green
2026-07-19 19:51:06 +02:00
Bastien Chanot 354ff2644f feat(agents): pin dispatched judgment agents to opus — Fable = inline reflection only (BDR-076)
Reverses the BDR-066 rejected alternative (opus pins on audit agents):
session default is now Fable, so inherit burned Fable quota on every
dispatched audit/challenge. analyzer, plan-challenger, seo/geo/
validator-analyzer pinned model: opus; onboard's 6 general-purpose
audit dispatches carry model="opus"; tour Phase B repointed.
interviewer + client-handover-writer stay unpinned (inline-load only,
a pin there is inert). settings.json default: claude-fable-5[1m].
Census flipped: model-routing §3 + new §11 (61 pass), loops-light 35,
full make test green.
2026-07-19 17:38:55 +02:00
Bastien Chanot 6bfc0543e5 feat(skills): add 3-way adversarial plan-challenge phase to reflection orchestrators
After a plan/reflection is elaborated and before it executes, three fresh blind
sub-agents (correctness / robustness / simplicity) attack it on the big model;
the main loop RE-THINKS every aspect a BLOCKER lands (a named plan change, or
[deferred]) and re-challenges once if the plan materially changed. Advisory into
each skill's existing human gate — the human stays the decider.

- lib/challenge-plan.md — reusable phase: fail-safe (never fail open),
  severity-driven (any single-lens BLOCKER = must-address), RE-THINK loop
- agents/plan-challenger.md — challenger role (read-only, big-model per BDR-066)
- lib/tests/plan-challenger.test.sh — 41-assertion structure lock
- wired into 11 orchestrators: ship-feature/init-project/feat/bugfix (build-plan),
  onboard/audit-delta/code-clean (proposals), seo/geo/harden/web-validate (fix-bundle)

Hardened by dogfooding: 3 blind challengers reviewed this feature's own v1 plan
and caught 4 BLOCKERs (fail-open, consensus-buries-lone-finding, wrong model
tier vs BDR-066, false on-disk-plan premise) — all fixed here.
2026-07-17 22:51:50 +02:00
Bastien Chanot b271e83fb6 feat(seo-data): content_quality verb — deterministic filler/AI-slop signal
Cherry-picked from claude-seo (github.com/AgriciDaniel/claude-seo, MIT)
content_quality.py, rewritten to the lib/seo-data contract per BDR-070. The
Content Shape axis was 100% LLM judgement; this gives it a measured input.

fetch.sh content_quality (stdin or --file) → {filler_score, ai_pattern_score,
information_density, overall_quality, flags[], matches{}}. 100% deterministic:
QRG §4.6 filler list (26 phrases) + AI-pattern list (46) kept intact, regex
matching, no LLM. Stdlib only (argparse/json/re/sys/collections/typing).

Advisory, NOT a verdict — the point of the wiring. It never claims a page "is
AI-written" (LRN-131/133); flags are candidates for human review. geo-analyzer
STEP 8 Check 10 makes it a deterministic input that INFORMS checks 1-9, never
replaces them, never scored on its own. A low number is not an automatic
finding.

Detection proven both directions (a detector that always- or never-flags is
useless): filler+slop text → flags [filler, low-density], overall 34-49; clean
dense factual text (dates/EUR/percentages) → no flags, overall 90. Empty input →
degraded/empty_input, never zeros-as-a-result.

Verified: GATE 1 verifier CONFORME 10/10 (both directions exercised live, lists
diffed intact vs source, advisory language confirmed); GATE 2 self-scan clean
(only sink is read-only open() for --file); seo-data 190 → 210 pass, 0 fail;
full suite green; shellcheck + py_compile clean.
2026-07-17 19:06:58 +02:00
Bastien Chanot cfdd89e73b feat(seo-data): schema_gen verb — generate JSON-LD, not just audit it
Cherry-picked from claude-seo (github.com/AgriciDaniel/claude-seo, MIT)
schema_generate.py, rewritten to the lib/seo-data contract per BDR-070 — adapt,
never copy. The system audited JSON-LD but could not generate it; geo-analyzer's
G2 batch hand-wrote markup. Now it calls the verb.

fetch.sh schema_gen {reservation|order|discussion|profile} → fail-open envelope
{"status":"ok","source":"schema_gen","type":…,"jsonld":{…}}. Types: Reservation
(7 subtypes), OrderAction, DiscussionForumPosting, ProfilePage (sameAs/knowsAbout
for the entity graph). Stdlib only (import argparse, json — zero third-party),
_strip_nones so a null is never emitted, --script-tag wraps for direct paste.

Fail-open mirrors score.py's _cli exactly (the contract's named pattern): a
flag-omitted required field → argparse exit 2 + {"status":"error","bad_usage"};
a flag-present-but-empty field → {"status":"degraded","reason":"missing required
field: …"} exit 0. Never a traceback, never empty stdout (LRN-133: the
can't-generate case stays legible).

geo-analyzer G2 wiring preserves the data-integrity rule — the verb generates
STRUCTURE, unknown values stay [À COMPLÉTER], never invented.

Verified: GATE 1 verifier CONFORME 10/10 (exercised the fail-open edge cases +
diffed field surface against the source); GATE 2 self-scan clean (no
network/shell/secret/eval sink); seo-data 167 → 190 pass, 0 fail; full suite
green; shellcheck + py_compile clean.
2026-07-17 14:30:31 +02:00
Bastien Chanot 4818c6116f feat(seo-data): I7 — compute the score instead of feeling it
/harden has a real scale (SKILL.md:435 — Critique -15, Haute -8, Moyenne -3,
Basse -1, clamp [0,100]). /seo had none: every axis was felt, so two runs over
identical code could disagree. That is a credibility problem on its own, and
/client-handover gates on 17/20 — a wobbling number makes the gate arbitrary.
H2 sharpened it: now that drift reports what actually changed, a score moving
on its own is visibly noise.

The split is the whole point. WHICH findings exist and how severe each is
stays the LLM's judgement — irreducible, and I am not pretending otherwise.
The arithmetic stops being judgement: same findings in, same score out. Same
principle as grouping cannibalisation rows in the engine rather than handing a
model 1000 rows to add up.

Reuses /harden's scale, /5 into /20, so the family speaks one vocabulary
instead of two.

Two things it makes real that were prose:
- **N/A is not a zero.** R2 (client-rendered on-page) and I1 (unauditable
  off-page) both mandate excluding an axis and renormalising the rest. Both
  left that arithmetic to the model. Now the engine does it and refuses to let
  N/A behave like a zero — verified: all-20 axes with two N/A still yields
  global 20.0, not a dragged-down mean.
- **Prevalence.** affected/sampled shift severity ONE step (>=50% escalates, a
  single page de-escalates). A defect on 1 of 12 pages is not the defect on
  12 of 12, and flattening the two is part of what made the old numbers move.

Malformed input is an error, never a silently wrong number — unlike the fetch
verbs, a degrade here would mean bad input, not a network fact. Unknown
severity and unknown profile both rejected, tested.

Verified: hand-checkable arithmetic (haute+moyenne = 100-11 = 89 → 17.8;
critique+haute = 77 → 15.4), identical global across repeated runs, weights
renormalised to sum 1.0 with two axes N/A. seo-data 155 -> 167 pass, 0 fail;
full suite green; shellcheck + py_compile clean.
2026-07-17 13:29:14 +02:00
Bastien Chanot d6b8edc8ea fix(seo): B1 KILLED — Common Crawl backlinks measured, not assumed
The plan said Common Crawl was the free backlink source and the 70/100 cap
was therefore mandatory. Measured before building, and both premises die.

HEAD against data.commoncrawl.org, live:
  cc-main-2026-feb-mar-apr-domain-edges.txt.gz    17.3 GB   gzipped
  cc-main-2026-feb-mar-apr-domain-ranks.txt.gz     2.3 GB
  cc-main-2026-feb-mar-apr-domain-vertices.txt.gz  879 MB

Finding one domain's inbound links means scanning the edges file end to end,
per audit. That is not slow, it is non-viable — and abusive toward a
nonprofit serving the data free.

Worse, the reference implementation everyone points at
(claude-seo scripts/commoncrawl_graph.py:169) does this:

    max_compressed_bytes = 500 * 1024 * 1024   # 500 MiB safety cap
    if total_downloaded > max_compressed_bytes: break

500 MiB of 17.3 GB is **2.9% of the edges file**, which is sorted by source
ID — so it reads an arbitrary slice of source domains and reports whatever
backlinks happened to be in it, as a backlink profile, capped at "70/100
health". Nothing in the output says 3%. That is a random sample wearing a
measurement's clothes: the exact failure class this branch exists to remove,
and I was one step from copying it.

B2 dies with B1: nothing left to cap.

CONSEQUENCE, and it is the point: I1's narrowed Off-page axis — brand
mentions only, backlinks + authority declared unauditable in §14 — is the
FINAL state, not a placeholder waiting for data. Corrected my own I1 text,
which pointed at Common Crawl as the "nearest free source": that sends a
future reader into a 17 GB dead end. The §14 line now records what was
measured and why no number beats a fabricated one.

Also corrects the B3 note, whose follow-on ("so Common Crawl is the only free
source") was wrong for the same reason. The only free viable backlink source
is Bing's GetUrlLinks — first-party only, never a competitor, and blocked on
the client's Bing account. That raises W2's value; it does not unblock it.

Verified: full suite green, seo-data 144 pass / 0 fail.
2026-07-17 13:17:18 +02:00