feat(model-router): wave 2-B — orchestrators declare phases, shifters and pins removed, frontmatter = off-state floor

The 15 Skill(effort-*) citers now call mcp__model-router__route per phase
(orchestrate at a dispatch span, reflect/plan for the skill's own level,
apply at the bookkeeping tail, escalate at the verify-secure caps); built-in
judgment dispatches carry an explicit effort= param. lib/effort-shift.md is
the route doctrine, lib/model-gate.md the mod rule (route answer = witness,
/route on as remedy). Deleted: skills/effort-*, lib/effort-pins.txt/.sh,
lib/model-check.sh, their tests, the installers' re-apply blocks. The mod
drops its Skill(effort-*) bridge. The tracked model:/effort: frontmatter
stays as the off-state floor, census-locked equal to the rows
(lib/tests/effort-routing.test.sh rewritten, 140 checks; analyzer → xhigh).

Contract .claude/tasks/contracts/2026-10-10-model-router-w2b-1045.md, plan
r4 § W2-B: GATE 0 MET, verifier ECARTS(7) then CONFORME 10/10, security
PASS, full make test green (design-tool-gate env red only).
This commit is contained in:
bchanot
2026-10-10 11:24:54 +02:00
parent 65dff0e768
commit 1f2d33b7a6
43 changed files with 318 additions and 811 deletions
+1 -1
View File
@@ -3,7 +3,7 @@ name: analyzer
description: Analyze code, codebase, or problem before any modification. Produces a factual report without proposing solutions. Use proactively before any refactoring, design, or implementation.
tools: Read, Grep, Glob, Bash
model: opus
effort: high
effort: xhigh
memory: project
---
+10 -10
View File
@@ -97,7 +97,7 @@ Parse `$ARGUMENTS` for optional flags:
---
EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op.
ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route).
## STEP 1 — PRE-FLIGHT
@@ -227,7 +227,7 @@ Store `DEPLOYED_URL` for STEP 7. If empty, ask user during STEP 6.
---
## STEP 3 — BASELINE AUDITS (parallel)
First: `Skill(effort-high)` (effort-shift: judgment dispatch; the fable skill-runners are built-ins and inherit the level in force; high is the entry level of the audits they run).
Every skill-runner dispatch below carries an explicit `effort="high"` (route: built-ins never inherit a main route; high is the entry level of the audits they run).
Goal: capture `SCORE_*_BEFORE` so the client doc shows the delta.
@@ -262,10 +262,10 @@ the gate.
**Model routing (BDR-077):** EVERY `general-purpose` skill-runner dispatch in
this pipeline (initial audits, fix-loop re-dispatches, commit-change,
web-validate) carries `model: "fable"` — the child hosts gated orchestration
on the pipeline's behalf; it must never inherit the session model.
web-validate) carries `model: "fable"` and `effort="high"` — the child hosts gated
orchestration on the pipeline's behalf; it must never inherit the session model.
For web projects, dispatch in **a single message with two parallel Agent calls** (each with `model: "fable"`):
For web projects, dispatch in **a single message with two parallel Agent calls** (each with `model: "fable"`, `effort="high"`):
| Audit (web) | Subagent | Prompt template |
|---------------|-------------------|-----------------|
@@ -424,7 +424,7 @@ for CSO) in the format
### Re-dispatch prompt template (SEO + GEO loop)
Send to `general-purpose` subagent (`model: "fable"`):
Send to `general-purpose` subagent (`model: "fable"`, `effort="high"`):
> Read `~/.claude/skills/seo/SKILL.md` and re-run it on this project in
> **conservative (audit-only) intervention mode**: re-score and leave both
@@ -455,7 +455,7 @@ Send to `general-purpose` subagent (`model: "fable"`):
### Re-dispatch prompt template (HARDEN loop)
Send to `general-purpose` subagent (`model: "fable"`):
Send to `general-purpose` subagent (`model: "fable"`, `effort="high"`):
> Read `~/.claude/skills/harden/SKILL.md` and re-run it with `--fix` ONLY
> up to the bundle: stop at `READY TO APPLY — awaiting dispatcher
@@ -468,7 +468,7 @@ Send to `general-purpose` subagent (`model: "fable"`):
### Re-dispatch prompt template (CSO loop — non-web only)
Send to `general-purpose` subagent (`model: "fable"`):
Send to `general-purpose` subagent (`model: "fable"`, `effort="high"`):
> Read `~/.claude/skills/cso/SKILL.md` and re-run it in **daily mode**.
> Previous score: **`<SCORE_CSO_PREVIOUS>`/20** — below threshold.
@@ -557,7 +557,7 @@ listed changes by hand before deploy." Continue to STEP 6.
If `PENDING_CHANGES` non-empty → invoke /commit-change skill via subagent:
> Dispatch `general-purpose` subagent (`model: "fable"`). Prompt:
> Dispatch `general-purpose` subagent (`model: "fable"`, `effort="high"`). Prompt:
>
> "Read `~/.claude/skills/commit-change/SKILL.md` and execute. All pending
> changes were produced by the client-handover ship pipeline during the
@@ -677,7 +677,7 @@ Skip if `VALIDATE_SKIPPED=true` or `PROJECT_TYPE != web` (in either case
ensure `VALIDATE_SKIPPED=true` is set so the gate logic in STEP 8 treats
VALIDATE as not-applicable rather than failed).
Dispatch `general-purpose` subagent (`model: "fable"`):
Dispatch `general-purpose` subagent (`model: "fable"`, `effort="high"`):
> Read `~/.claude/skills/web-validate/SKILL.md` and execute against the
> deployed URL: `<DEPLOYED_URL>`. Audit W3C HTML validity (validator.nu),
+1 -1
View File
@@ -10,7 +10,7 @@ effort: high
> MODEL (BDR-077): `MODE: propose` is dispatched with `model="opus"` (the
> call-site override — narrative reconstruction + capitalize routing are
> judgment); `MODE: apply` runs on the sonnet frontmatter pin (mechanical
> judgment); `MODE: apply` runs on its sonnet row (frontmatter = off-state floor; mechanical
> staging/committing of an approved plan).
Reconstruct the development narrative from a working directory. The goal
+2 -2
View File
@@ -60,8 +60,8 @@ audit, report, and patch.
Parse `$ARGUMENTS`:
- **`MODE: patch`** — the dispatcher approved a PATCH PLAN and re-dispatches
this agent to APPLY it. Jump to MODE: PATCH section. Runs on the sonnet
frontmatter pin.
this agent to APPLY it. Jump to MODE: PATCH section. Runs on its sonnet
row (frontmatter = off-state floor).
- **`MODE: audit`** (or no explicit MODE — audit is the default) — analysis
half, dispatched with `model: "opus"` (judgment tier; the call-site
override takes precedence over the sonnet pin). **READ-ONLY: Write and
+1 -1
View File
@@ -104,7 +104,7 @@ Mirror of seo-analyzer's pipeline contract. Parse the MODE line:
to the run-scoped, gitignored `.audit/geo-signals-<RUNID>.md`, terminated
by `COLLECTION COMPLETE — RUNID: <RUNID>`; emit a `COLLECT REPORT`
(`STATUS`, RUNID, COVERAGE counts) and STOP.
- **`MODE: judge`** — opus frontmatter pin. Fail-closed load of
- **`MODE: judge`** — opus row (frontmatter = off-state floor). Fail-closed load of
`.audit/geo-signals-<RUNID>.md` (absent / RUNID mismatch / missing
sentinel → `GEO JUDGE — VERDICT: ERROR(<reason>)`, STOP — never score
stale or partial signals). Then STEP 6-12 (schema, entity — including
+1 -1
View File
@@ -57,7 +57,7 @@ The parent dispatches this agent TWICE, with the FULL PACKAGE both times
(`STATUS: DONE | BLOCKED`, RUNID, phase-cluster count, per-chapter word
counts) and STOPS — STEP 13-16, the final MD, HTML and PDF are NEVER
this mode's job.
- **`MODE: render`** — runs on the sonnet frontmatter pin. FIRST loads the
- **`MODE: render`** — runs on its sonnet row (frontmatter = off-state floor). FIRST loads the
draft: absent file, RUNID mismatch, or missing `DRAFT COMPLETE` sentinel
→ `STATUS: BLOCKED` naming the cause (fail closed — never synthesize a
missing draft, never render a partial one). Then runs STEP 13 → 14 →
+5 -5
View File
@@ -106,11 +106,11 @@ the MAIN loop, never here):
- Dispatch THREE fresh challengers IN PARALLEL, one per lens
(correctness / robustness / simplicity), each blind to the others.
- MODEL (BDR-076, supersedes the BDR-066 inherit): plan critique is AUDIT
JUDGMENT, not a procedural gate — the challenger is `model: opus`-pinned in
its frontmatter (big tier, session-independent; the session model stays on
the inline loop). Never `model: "sonnet"` — a silent judgment downgrade.
(Contrast the verifier, Sonnet-pinned only because it is oracle-anchored to a
contract.)
JUDGMENT, not a procedural gate — the challenger is routed to the `judge`
row (opus) by the model-router, the `model: opus` frontmatter being the off-state floor (big
tier; the session model stays on the inline loop). Never `model: "sonnet"` —
a silent judgment downgrade. (Contrast the verifier on its `verify` row and
the executors on their `implement` row, both sonnet, oracle-anchored to a contract.)
- FAIL-SAFE — never fail open: a malformed/empty verdict, a missing `PROOF`, or
a dead challenger → retry ONCE fresh; a 2nd failure → escalate to the human and
NAME the lens. Never report "plan challenged" on a silently dropped lens (same
+1 -1
View File
@@ -38,7 +38,7 @@ The dispatcher (/seo) runs this agent as a 3-stage pipeline; /harden and
terminated by the line `COLLECTION COMPLETE — RUNID: <RUNID>`, then
emits a short `COLLECT REPORT` (`STATUS: DONE | BLOCKED`, RUNID,
COVERAGE counts) and STOPS. No scoring, no findings, no bundle.
- **`MODE: judge`** — runs on the opus frontmatter pin (audit judgment).
- **`MODE: judge`** — runs on its opus row (audit judgment; frontmatter = off-state floor).
FIRST loads `.audit/seo-signals-<RUNID>.md`: absent, RUNID mismatch, or
missing `COLLECTION COMPLETE` sentinel → emit
`SEO JUDGE — VERDICT: ERROR(<reason>)` and STOP (fail closed — NEVER