Files
claude/.claude/tasks/contracts/2026-10-10-model-router-w2b-1045.md
T

56 lines
9.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# CONTRACT — model-router-w2b
- date: 2026-10-10 | flow: feat | branch: feature/model-router-w2
- status: active
- wave: W2-B (repo migration), plan r4 `.claude/tasks/plans/2026-10-09-model-router-w2-1546.md` § D1-D4, § Row tables, § W2-B B0-B7. Follows W2-A (bb56f3e) and the A→B gate.
## REQUEST (verbatim — IMMUTABLE)
tu peux lancer la vague 2
(TODO W2 line: 15 skills `Skill(effort-*)` → `route` tool calls; remove `skills/effort-*`, `lib/effort-pins.txt/.sh`, install/update steps, `effort:` frontmatter on skills and agents; repo agents into the mod's `agents` table (verify the spawn/first-step ordering first); `lib/model-gate.md` + `lib/model-check.sh` → mod rule; census tests repointed; `lib/effort-shift.md` rewritten; docs)
## CLARIFICATIONS
Q: pass B 2026-10-09 (see the W2-A contract) / A: effort-* deleted; pins reworked into phase rows with the tracked frontmatter kept as census-locked off-state floor; orchestrators declare phases; slim model gate [gated 2026-10-09]
Q: A→B gate evidence (engine records 2026-10-10) / A: typed `/status` → main effort low then medium during the dispatch (row applied to a typed slash); analyzer spawn → opus/xhigh from step 0 while its frontmatter still said high (spawn bookkeeping precedes the first step; the row beats the frontmatter); status-reporter haiku from step 0. Probe 4 (typed skill while a background agent is live: marker vs fallback) NOT observed → recorded as a limit in `lib/effort-shift.md`, re-checked later in the records; not blocking [gated 2026-10-10]
Q: slim gate design / A: the include ALWAYS calls the route tool at entry (`mcp__model-router__route` phase = the skill's row) and reads the id in the answer; no self-check wording dependency; the tool is deferred for the model → `ToolSearch("select:mcp__model-router__route")` once per session when it is not loaded [gated 2026-10-10]
Q: FLOOR items (verifier 1): deletion of `lib/tests/effort-pins.test.sh` and `lib/tests/model-check.test.sh`, removal of the two `Skill(effort-low)` bridge kit tests, the from-scratch rewrite of `lib/tests/effort-routing.test.sh` (has/ok/ko helpers, its `# shellcheck disable=SC2015,SC2016` directive carried over from HEAD) / A: all authorized — they are the plan's B0 and B4 steps, accepted at pass B (shifters, pins and witness deleted with their tests) [gated 2026-10-10]
Q: scope addition (verifier 1, criterion 9a): `agents/plan-challenger.md:110-112` twins `lib/challenge-plan.md:46` ("`model: opus`-pinned in its frontmatter, session-independent") / A: added to FILE SCOPE, same rewording (row + off-state floor); the ~20 loose "sonnet pin"/"opus pin" shorthand sites stay (true: alias == row tier by census) [gated 2026-10-10, orchestrator — minor, user may veto]
## ACCEPTANCE CRITERIA
1. No `Skill(effort-` call and no `EFFORT SHIFTS:` header line remains in tracked skills, agents and lib (the mod and CHANGELOG excepted).
CHECK: ! git grep -nE 'Skill\(effort-|EFFORT SHIFTS:' -- skills agents lib ':!mods' >/dev/null && echo W2B-NO-SHIFTER-CITER
EXPECT: W2B-NO-SHIFTER-CITER
EVIDENCE: MET exit=0 marker-found :: W2B-NO-SHIFTER-CITER
2. Deleted: `skills/effort-{low,medium,high,xhigh,max}`, `lib/effort-pins.txt`, `lib/effort-pins.sh`, `lib/model-check.sh`, `lib/tests/effort-pins.test.sh`, `lib/tests/model-check.test.sh`; no tracked reference to `effort-pins`, `apply_effort_pins` or `model-check.sh` outside CHANGELOG.md, `.claude/`, README.md and USAGE.md (the two docs are the STEP 6 doc-sync targets [gated 2026-10-10]).
CHECK: for p in skills/effort-low skills/effort-medium skills/effort-high skills/effort-xhigh skills/effort-max lib/effort-pins.txt lib/effort-pins.sh lib/model-check.sh lib/tests/effort-pins.test.sh lib/tests/model-check.test.sh; do [ ! -e "$p" ] || { echo "still here: $p"; exit 1; }; done && ! git grep -nE 'effort-pins|apply_effort_pins|model-check\.sh' -- . ':!CHANGELOG.md' ':!.claude' ':!README.md' ':!USAGE.md' >/dev/null && echo W2B-DELETIONS-DONE
EXPECT: W2B-DELETIONS-DONE
EVIDENCE: MET exit=0 marker-found :: W2B-DELETIONS-DONE
3. Mod: the `Skill(effort-*)` bridge is gone (`EFFORT_SKILL`, `effortBridge` absent), the kit suite and the mods suite are green.
CHECK: ! grep -qE 'EFFORT_SKILL|effortBridge' mods/model-router/hooks/register.ts && cd mods/model-router && out="$(claude plugin test . 2>&1)" && printf '%s\n' "$out" | grep -qE '[0-9]+ pass' && ! printf '%s\n' "$out" | grep -qE '[1-9][0-9]* fail' && cd ../.. && make test suite=lib/tests/mods.test.sh 2>&1 | grep -q 'all suites green' && echo W2B-MOD-GREEN
EXPECT: W2B-MOD-GREEN
EVIDENCE: MET exit=0 marker-found :: W2B-MOD-GREEN
4. `lib/tests/effort-routing.test.sh` is rewritten as the wave-2 census (drift lock rows ↔ frontmatter parsed from `register.ts`, no `effort:`-less routed skill, D3 wiring markers, no shifter citers) and is green; `lib/tests/model-routing.test.sh` and `lib/tests/higgsfield.test.sh` are green.
CHECK: grep -q 'register.ts' lib/tests/effort-routing.test.sh && bash lib/tests/effort-routing.test.sh >/dev/null 2>&1 && bash lib/tests/model-routing.test.sh >/dev/null 2>&1 && bash lib/tests/higgsfield.test.sh >/dev/null 2>&1 && echo W2B-CENSUS-GREEN
EXPECT: W2B-CENSUS-GREEN
EVIDENCE: MET exit=0 marker-found :: W2B-CENSUS-GREEN
5. Wiring follows plan D3 at every former shifter site: dispatch span → `mcp__model-router__route` phase orchestrate; own level high → reflect, xhigh → plan; bookkeeping tail → apply; escalation (verify-secure-loop ×3, ship-feature error recovery) → escalate; judgment dispatches of built-ins (`general-purpose` `model="opus"` in ship-feature, init-project, onboard, tour; `model: "fable"` skill-runners in client-handover-writer) carry an explicit `effort=` param instead of a shift; the "turn reset" re-assert sites become a route call at the first step of the resumed turn. Judged by reading each site against the plan.
6. `lib/model-gate.md` is the mod rule, ≤ 30 lines: always call `mcp__model-router__route` at entry with the skill's row phase, STOP unless the answer names a fable or opus id (tool missing → ToolSearch once; "is off" → STOP with the remedy); keeps the dispatch-tier table with `model: "fable"` for skill-runners; no `model-check` reference.
CHECK: [ "$(wc -l < lib/model-gate.md)" -le 30 ] && grep -q 'mcp__model-router__route' lib/model-gate.md && grep -q 'model: "fable"' lib/model-gate.md && ! grep -q 'model-check' lib/model-gate.md && echo W2B-GATE-SLIM
EXPECT: W2B-GATE-SLIM
EVIDENCE: MET exit=0 marker-found :: W2B-GATE-SLIM
7. `lib/effort-shift.md` is rewritten (≤ 60 lines) as the route doctrine: the tool name and the ToolSearch note, the D3 wiring points, explicit `effort=` for built-in judgment dispatches, the run slot and `/route clear`, levers (`ultrathink`, `/route effort=max`; builtin `/effort` is not a lever inside a run), last rowed skill wins and an unrowed one changes nothing, the probe-4 limit, headless OK, `effort-audit.py`.
CHECK: [ "$(wc -l < lib/effort-shift.md)" -le 60 ] && grep -q 'mcp__model-router__route' lib/effort-shift.md && grep -q 'ToolSearch' lib/effort-shift.md && grep -q 'effort-audit.py' lib/effort-shift.md && ! grep -q 'Skill(effort' lib/effort-shift.md && echo W2B-DOCTRINE
EXPECT: W2B-DOCTRINE
EVIDENCE: MET exit=0 marker-found :: W2B-DOCTRINE
8. Frontmatter aligned to the rows: `agents/analyzer.md` carries `effort: xhigh`; every other tracked `effort:`/`model:` value equals its row (criterion 4 census proves it); `CLAUDE.global.md` Design-work lines no longer cite `lib/effort-pins.txt` or the lone-Skill-call rule.
CHECK: grep -q '^effort: xhigh' agents/analyzer.md && ! grep -q 'effort-pins' CLAUDE.global.md && ! grep -q 'lone Skill call' CLAUDE.global.md && echo W2B-FRONTMATTER
EXPECT: W2B-FRONTMATTER
EVIDENCE: MET exit=0 marker-found :: W2B-FRONTMATTER
9. Prose sites that stated pin semantics ("sonnet by frontmatter pin", "`model: opus`-pinned in their frontmatter, session-independent", STOP texts naming `/effort-max`) are reworded to the row/off-state-floor semantics and the D4 levers. Judged by reading.
10. The suites the diff touches are green (doctrine-citers, loops-light, plan-challenger, skill-routing-census, profile-census, portability-census, gitflow-test if present) and `shellcheck *.sh hooks/*.sh lib/*.sh` is clean; the full `make test` runs once before the merge (LRN-209), the orchestrator reports it.
CHECK: for t in doctrine-citers loops-light plan-challenger skill-routing-census profile-census portability-census; do [ -f "lib/tests/$t.test.sh" ] || continue; bash "lib/tests/$t.test.sh" >/dev/null 2>&1 || { echo "red: $t"; exit 1; }; done; shellcheck ./*.sh hooks/*.sh lib/*.sh && echo W2B-SUITE-GREEN
EXPECT: W2B-SUITE-GREEN
EVIDENCE: MET exit=0 marker-found :: W2B-SUITE-GREEN
## FILE SCOPE
mods/model-router/hooks/register.ts, mods/model-router/hooks/register.test.ts; lib/effort-shift.md, lib/model-gate.md, lib/verify-secure-loop.md, lib/challenge-plan.md; skills/{ship-feature,init-project,feat,bugfix,web-validate,seo,hotfix,geo,harden,code-clean,audit-delta,onboard,tour,client-handover,analyze,commit-change}/SKILL.md; agents/client-handover-writer.md, agents/analyzer.md, agents/plan-challenger.md (prose only), agents/{commit-changer,handover-doc-writer,geo-analyzer,doc-syncer,seo-analyzer}.md (prose only); deletions listed in criterion 2; install-plugins.sh, update-all.sh; lib/tests/{effort-routing,model-routing,higgsfield}.test.sh; CLAUDE.global.md (Design-work lines only).