From 354ff2644fbc7a14171fe210f0a3c2c3a334c2b6 Mon Sep 17 00:00:00 2001 From: Bastien Chanot Date: Sun, 19 Jul 2026 17:38:55 +0200 Subject: [PATCH] =?UTF-8?q?feat(agents):=20pin=20dispatched=20judgment=20a?= =?UTF-8?q?gents=20to=20opus=20=E2=80=94=20Fable=20=3D=20inline=20reflecti?= =?UTF-8?q?on=20only=20(BDR-076)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Reverses the BDR-066 rejected alternative (opus pins on audit agents): session default is now Fable, so inherit burned Fable quota on every dispatched audit/challenge. analyzer, plan-challenger, seo/geo/ validator-analyzer pinned model: opus; onboard's 6 general-purpose audit dispatches carry model="opus"; tour Phase B repointed. interviewer + client-handover-writer stay unpinned (inline-load only, a pin there is inert). settings.json default: claude-fable-5[1m]. Census flipped: model-routing §3 + new §11 (61 pass), loops-light 35, full make test green. --- agents/analyzer.md | 1 + agents/geo-analyzer.md | 1 + agents/plan-challenger.md | 7 +++++-- agents/seo-analyzer.md | 1 + agents/validator-analyzer.md | 1 + lib/challenge-plan.md | 8 +++++--- lib/tests/model-routing.test.sh | 23 +++++++++++++++-------- settings.json | 2 +- skills/onboard/SKILL.md | 10 ++++++++-- skills/tour/SKILL.md | 6 +++--- 10 files changed, 41 insertions(+), 19 deletions(-) diff --git a/agents/analyzer.md b/agents/analyzer.md index 0ce6727..135d405 100644 --- a/agents/analyzer.md +++ b/agents/analyzer.md @@ -2,6 +2,7 @@ name: analyzer description: Analyze code, codebase, or problem before any modification. Produces a factual report without proposing solutions. Use proactively before any refactoring, design, or implementation. tools: Read, Grep, Glob, Bash +model: opus memory: project --- diff --git a/agents/geo-analyzer.md b/agents/geo-analyzer.md index 5d5cd77..c787427 100644 --- a/agents/geo-analyzer.md +++ b/agents/geo-analyzer.md @@ -2,6 +2,7 @@ name: geo-analyzer description: GEO audit agent for AI search engines — dispatched by /geo and /seo. Audits AI crawlers, llms.txt, entity signals, Schema.org; emits a fix bundle (dispatcher applies), scored report. Classical SEO → seo-analyzer agent. tools: Read, Edit, Write, Bash, Grep, Glob, WebFetch, WebSearch +model: opus --- # GEO — Generative Engine Optimization audit, fix & strategy diff --git a/agents/plan-challenger.md b/agents/plan-challenger.md index 38d361e..e5c0fda 100644 --- a/agents/plan-challenger.md +++ b/agents/plan-challenger.md @@ -2,6 +2,7 @@ name: plan-challenger description: Fresh independent plan challenger — reads a PLAN file from disk and adversarially attacks it through ONE assigned lens (correctness | robustness | simplicity), then renders structured findings + a verdict. Report-only, never fixes, never implements. Dispatched fresh; blind to the other lenses. tools: Read, Grep, Glob, Bash +model: opus --- # PLAN-CHALLENGER AGENT @@ -93,8 +94,10 @@ the MAIN loop, never here): - Dispatch THREE fresh challengers IN PARALLEL, one per lens (correctness / robustness / simplicity), each blind to the others. -- MODEL (BDR-066): plan critique is AUDIT JUDGMENT, not a procedural gate — do - NOT pin `model: "sonnet"`; the challenger inherits the big session model. +- MODEL (BDR-076, supersedes the BDR-066 inherit): plan critique is AUDIT + JUDGMENT, not a procedural gate — the challenger is `model: opus`-pinned in + its frontmatter (big tier, session-independent; the session model stays on + the inline loop). Never `model: "sonnet"` — a silent judgment downgrade. (Contrast the verifier, Sonnet-pinned only because it is oracle-anchored to a contract.) - FAIL-SAFE — never fail open: a malformed/empty verdict, a missing `PROOF`, or diff --git a/agents/seo-analyzer.md b/agents/seo-analyzer.md index 82cb309..6f6f384 100644 --- a/agents/seo-analyzer.md +++ b/agents/seo-analyzer.md @@ -2,6 +2,7 @@ name: seo-analyzer description: 'Classical SEO audit agent (Google, Bing) — dispatched from /seo. Live audit: Core Web Vitals, on-page, technical, local SEO, legal (FR). Emits a fix bundle (dispatcher applies) + scored report. AI/GEO → geo-analyzer agent.' tools: Read, Edit, Write, Bash, Grep, Glob, WebFetch, WebSearch +model: opus --- # SEO — Classical Search Engines audit, fix & strategy diff --git a/agents/validator-analyzer.md b/agents/validator-analyzer.md index c78b15b..081765b 100644 --- a/agents/validator-analyzer.md +++ b/agents/validator-analyzer.md @@ -2,6 +2,7 @@ name: validator-analyzer description: Web standards audit agent — W3C HTML validity (validator.nu), W3C CSS validity (jigsaw.w3.org), WCAG 2.1 accessibility (axe-core, pa11y, WAVE). Dispatched from /web-validate. Produces scored .claude/audits/VALIDATE.md report with concrete diffs for auto-fixable issues and user actions for judgment-required fixes. Complementary to /harden (security), /seo (indexability), /geo (AI extraction). tools: Read, Edit, Write, Bash, Grep, Glob, WebFetch +model: opus --- # Validator — W3C + WCAG audit diff --git a/lib/challenge-plan.md b/lib/challenge-plan.md index a69ecbd..91e72ef 100644 --- a/lib/challenge-plan.md +++ b/lib/challenge-plan.md @@ -42,9 +42,11 @@ Agent(subagent_type="plan-challenger", description="challenge:", prompt="" """) ``` -**MODEL (BDR-066):** plan critique is AUDIT JUDGMENT — do NOT pin -`model: "sonnet"`; the challengers inherit the big session model. (The executor -gates stay sonnet; the challenger does not.) +**MODEL (BDR-076, supersedes the BDR-066 inherit):** plan critique is AUDIT +JUDGMENT — the challengers are `model: opus`-pinned in their frontmatter: a big +tier, session-independent, off the session model. The session model (Fable) +keeps only this loop — synthesis, RE-THINK, gate. Never sonnet: that would +silently downgrade the judgment. (The executor gates stay sonnet.) **Lens framing by `KIND`** (the agent's three lenses, read against the artifact): - `build-plan` — will it WORK / will it BREAK / is it needlessly COMPLEX. diff --git a/lib/tests/model-routing.test.sh b/lib/tests/model-routing.test.sh index f0d08f5..8b1b724 100755 --- a/lib/tests/model-routing.test.sh +++ b/lib/tests/model-routing.test.sh @@ -1,5 +1,5 @@ #!/usr/bin/env bash -# lib/tests/model-routing.test.sh — census: gate wiring + pins + executor shape (BDR-066) +# lib/tests/model-routing.test.sh — census: gate wiring + pins + executor shape (BDR-066, BDR-076) set -u R="$(cd "$(dirname "$0")/../.." && pwd)" pass=0; fail=0 @@ -22,7 +22,7 @@ has "agents/feater.md" 'model: sonnet' has "agents/hotfixer.md" 'model: sonnet' has "agents/verifier.md" 'model: sonnet' has "agents/security-auditor.md" 'model: sonnet' -fm_lacks "agents/analyzer.md" 'model:' +has "agents/analyzer.md" 'model: opus' # 4) /feat executor shape has "skills/feat/SKILL.md" 'subagent_type="feater"' has "skills/feat/SKILL.md" 'verify-secure-loop.md' @@ -54,17 +54,24 @@ lacks "agents/handover-doc-writer.md" 'AskUserQuestion' lacks "agents/handover-doc-writer.md" 'Agent(' has "agents/client-handover-writer.md" 'subagent_type="handover-doc-writer"' # 10) post-merge edge fixes (ronde): F1 feater applier carve-out, F2 /refactor -# dispatch + pin, F3 /analyze gated (in loop 1), F4 interviewer un-pinned, -# F5 audit agents' ABSENT pin locked (a stray sonnet pin would silently -# downgrade a live audit even though the skill's gate passed) +# dispatch + pin, F3 /analyze gated (in loop 1) has "agents/feater.md" 'Applier path' has "skills/refactor/SKILL.md" 'subagent_type="refactorer"' has "agents/refactorer.md" 'model: sonnet' -fm_lacks "agents/seo-analyzer.md" 'model:' -fm_lacks "agents/geo-analyzer.md" 'model:' -fm_lacks "agents/validator-analyzer.md" 'model:' +# 11) BDR-076 — session model (Fable) = orchestration + inline reflection ONLY. +# Dispatched judgment agents pinned OPUS (big tier, session-independent; +# never sonnet — that would silently downgrade a live audit). Inline-load- +# only agents (interviewer, client-handover-writer) STAY unpinned: they run +# IN the main loop, a frontmatter pin there is inert and misleads. +has "agents/seo-analyzer.md" 'model: opus' +has "agents/geo-analyzer.md" 'model: opus' +has "agents/validator-analyzer.md" 'model: opus' +has "agents/plan-challenger.md" 'model: opus' fm_lacks "agents/client-handover-writer.md" 'model:' fm_lacks "agents/interviewer.md" 'model:' +has "skills/onboard/SKILL.md" 'model="opus"' +has "skills/tour/SKILL.md" 'model="opus"' +has "lib/challenge-plan.md" 'BDR-076' printf 'model-routing census: %d pass, %d fail\n' "$pass" "$fail" [ "$fail" -eq 0 ] diff --git a/settings.json b/settings.json index 43bdbf8..7a09992 100644 --- a/settings.json +++ b/settings.json @@ -250,7 +250,7 @@ "disableBypassPermissionsMode": "disable", "additionalDirectories": [] }, - "model": "opus[1m]", + "model": "claude-fable-5[1m]", "hooks": { "SessionStart": [ { diff --git a/skills/onboard/SKILL.md b/skills/onboard/SKILL.md index 367ab11..c6f3ed6 100644 --- a/skills/onboard/SKILL.md +++ b/skills/onboard/SKILL.md @@ -354,7 +354,7 @@ Lire le bloc `audit_stack:` du fichier `~/.claude/lib/project-archetypes/. ARCHETYPE: . @@ -406,6 +407,7 @@ bash $HOME/.claude/lib/toggle-external.sh list 2>/dev/null | grep -E "^gstack\s+ ``` Agent( subagent_type="general-purpose", + model="opus", description="Onboard — security audit fallback (archetype-adaptive)", prompt=""" READ-ONLY security audit. No file modifications. @@ -647,6 +649,7 @@ Si le skill ne supporte pas `--output`, capturer la sortie et écrire à la main ``` Agent( subagent_type="general-purpose", + model="opus", description="Onboard — static design review fallback", prompt=""" AUDIT-ONLY mode — NO edits. Static design review du code UI. @@ -690,6 +693,7 @@ Puis parser le JSON Lighthouse (scores perf/a11y/bp/seo/pwa + top opportunities) ``` Agent( subagent_type="general-purpose", + model="opus", description="Onboard — static perf audit", prompt=""" AUDIT-ONLY mode — NO edits. @@ -732,6 +736,7 @@ Parser axe-core résultats (violations, incomplete, inapplicable, passes) → `. ``` Agent( subagent_type="general-purpose", + model="opus", description="Onboard — static a11y audit", prompt=""" AUDIT-ONLY mode — NO edits. @@ -777,6 +782,7 @@ Spawn un subagent synthétiseur (isolé, chargé uniquement du contenu de `.onbo ``` Agent( subagent_type="general-purpose", + model="opus", description="Onboard — synthèse vers .claude/audits/", prompt=""" Lire tous les fichiers de /.onboard-audit/ : diff --git a/skills/tour/SKILL.md b/skills/tour/SKILL.md index 8e7ef8a..d7aefaa 100644 --- a/skills/tour/SKILL.md +++ b/skills/tour/SKILL.md @@ -111,9 +111,9 @@ honestly in the summary. Never loop past 3. ### Phase B — CLEAN -1. Dispatch a read-only cleanup audit (analyzer or general-purpose — - inherits the big session model; NOT the sonnet code-cleaner, which is - now a fix executor): dead code, unused imports/exports, +1. Dispatch a read-only cleanup audit (analyzer — opus-pinned, BDR-076 — + or general-purpose with `model="opus"`; NOT the sonnet code-cleaner, + which is now a fix executor): dead code, unused imports/exports, commented-out blocks, stale flags, norm violations. Findings as `id | file:line | finding | proposed fix`. 2. Apply **behavior-preserving** fixes only. A finding that would change