Files
claude/.claude/tasks/contracts/2026-10-10-model-router-w2b-1045.md
T

9.3 KiB
Raw Blame History

CONTRACT — model-router-w2b

  • date: 2026-10-10 | flow: feat | branch: feature/model-router-w2
  • status: active
  • wave: W2-B (repo migration), plan r4 .claude/tasks/plans/2026-10-09-model-router-w2-1546.md § D1-D4, § Row tables, § W2-B B0-B7. Follows W2-A (bb56f3e) and the A→B gate.

REQUEST (verbatim — IMMUTABLE)

tu peux lancer la vague 2 (TODO W2 line: 15 skills Skill(effort-*) → route tool calls; remove skills/effort-*, lib/effort-pins.txt/.sh, install/update steps, effort: frontmatter on skills and agents; repo agents into the mod's agents table (verify the spawn/first-step ordering first); lib/model-gate.md + lib/model-check.sh → mod rule; census tests repointed; lib/effort-shift.md rewritten; docs)

CLARIFICATIONS

Q: pass B 2026-10-09 (see the W2-A contract) / A: effort-* deleted; pins reworked into phase rows with the tracked frontmatter kept as census-locked off-state floor; orchestrators declare phases; slim model gate [gated 2026-10-09] Q: A→B gate evidence (engine records 2026-10-10) / A: typed /status → main effort low then medium during the dispatch (row applied to a typed slash); analyzer spawn → opus/xhigh from step 0 while its frontmatter still said high (spawn bookkeeping precedes the first step; the row beats the frontmatter); status-reporter haiku from step 0. Probe 4 (typed skill while a background agent is live: marker vs fallback) NOT observed → recorded as a limit in lib/effort-shift.md, re-checked later in the records; not blocking [gated 2026-10-10] Q: slim gate design / A: the include ALWAYS calls the route tool at entry (mcp__model-router__route phase = the skill's row) and reads the id in the answer; no self-check wording dependency; the tool is deferred for the model → ToolSearch("select:mcp__model-router__route") once per session when it is not loaded [gated 2026-10-10]

Q: FLOOR items (verifier 1): deletion of lib/tests/effort-pins.test.sh and lib/tests/model-check.test.sh, removal of the two Skill(effort-low) bridge kit tests, the from-scratch rewrite of lib/tests/effort-routing.test.sh (has/ok/ko helpers, its # shellcheck disable=SC2015,SC2016 directive carried over from HEAD) / A: all authorized — they are the plan's B0 and B4 steps, accepted at pass B (shifters, pins and witness deleted with their tests) [gated 2026-10-10] Q: scope addition (verifier 1, criterion 9a): agents/plan-challenger.md:110-112 twins lib/challenge-plan.md:46 ("model: opus-pinned in its frontmatter, session-independent") / A: added to FILE SCOPE, same rewording (row + off-state floor); the ~20 loose "sonnet pin"/"opus pin" shorthand sites stay (true: alias == row tier by census) [gated 2026-10-10, orchestrator — minor, user may veto]

ACCEPTANCE CRITERIA

  1. No Skill(effort- call and no EFFORT SHIFTS: header line remains in tracked skills, agents and lib (the mod and CHANGELOG excepted). CHECK: ! git grep -nE 'Skill(effort-|EFFORT SHIFTS:' -- skills agents lib ':!mods' >/dev/null && echo W2B-NO-SHIFTER-CITER EXPECT: W2B-NO-SHIFTER-CITER EVIDENCE: MET exit=0 marker-found :: W2B-NO-SHIFTER-CITER
  2. Deleted: skills/effort-{low,medium,high,xhigh,max}, lib/effort-pins.txt, lib/effort-pins.sh, lib/model-check.sh, lib/tests/effort-pins.test.sh, lib/tests/model-check.test.sh; no tracked reference to effort-pins, apply_effort_pins or model-check.sh outside CHANGELOG.md, .claude/, README.md and USAGE.md (the two docs are the STEP 6 doc-sync targets [gated 2026-10-10]). CHECK: for p in skills/effort-low skills/effort-medium skills/effort-high skills/effort-xhigh skills/effort-max lib/effort-pins.txt lib/effort-pins.sh lib/model-check.sh lib/tests/effort-pins.test.sh lib/tests/model-check.test.sh; do [ ! -e "$p" ] || { echo "still here: $p"; exit 1; }; done && ! git grep -nE 'effort-pins|apply_effort_pins|model-check.sh' -- . ':!CHANGELOG.md' ':!.claude' ':!README.md' ':!USAGE.md' >/dev/null && echo W2B-DELETIONS-DONE EXPECT: W2B-DELETIONS-DONE EVIDENCE: MET exit=0 marker-found :: W2B-DELETIONS-DONE
  3. Mod: the Skill(effort-*) bridge is gone (EFFORT_SKILL, effortBridge absent), the kit suite and the mods suite are green. CHECK: ! grep -qE 'EFFORT_SKILL|effortBridge' mods/model-router/hooks/register.ts && cd mods/model-router && out="$(claude plugin test . 2>&1)" && printf '%s\n' "$out" | grep -qE '[0-9]+ pass' && ! printf '%s\n' "$out" | grep -qE '[1-9][0-9]* fail' && cd ../.. && make test suite=lib/tests/mods.test.sh 2>&1 | grep -q 'all suites green' && echo W2B-MOD-GREEN EXPECT: W2B-MOD-GREEN EVIDENCE: MET exit=0 marker-found :: W2B-MOD-GREEN
  4. lib/tests/effort-routing.test.sh is rewritten as the wave-2 census (drift lock rows ↔ frontmatter parsed from register.ts, no effort:-less routed skill, D3 wiring markers, no shifter citers) and is green; lib/tests/model-routing.test.sh and lib/tests/higgsfield.test.sh are green. CHECK: grep -q 'register.ts' lib/tests/effort-routing.test.sh && bash lib/tests/effort-routing.test.sh >/dev/null 2>&1 && bash lib/tests/model-routing.test.sh >/dev/null 2>&1 && bash lib/tests/higgsfield.test.sh >/dev/null 2>&1 && echo W2B-CENSUS-GREEN EXPECT: W2B-CENSUS-GREEN EVIDENCE: MET exit=0 marker-found :: W2B-CENSUS-GREEN
  5. Wiring follows plan D3 at every former shifter site: dispatch span → mcp__model-router__route phase orchestrate; own level high → reflect, xhigh → plan; bookkeeping tail → apply; escalation (verify-secure-loop ×3, ship-feature error recovery) → escalate; judgment dispatches of built-ins (general-purpose model="opus" in ship-feature, init-project, onboard, tour; model: "fable" skill-runners in client-handover-writer) carry an explicit effort= param instead of a shift; the "turn reset" re-assert sites become a route call at the first step of the resumed turn. Judged by reading each site against the plan.
  6. lib/model-gate.md is the mod rule, ≤ 30 lines: always call mcp__model-router__route at entry with the skill's row phase, STOP unless the answer names a fable or opus id (tool missing → ToolSearch once; "is off" → STOP with the remedy); keeps the dispatch-tier table with model: "fable" for skill-runners; no model-check reference. CHECK: [ "$(wc -l < lib/model-gate.md)" -le 30 ] && grep -q 'mcp__model-router__route' lib/model-gate.md && grep -q 'model: "fable"' lib/model-gate.md && ! grep -q 'model-check' lib/model-gate.md && echo W2B-GATE-SLIM EXPECT: W2B-GATE-SLIM EVIDENCE: MET exit=0 marker-found :: W2B-GATE-SLIM
  7. lib/effort-shift.md is rewritten (≤ 60 lines) as the route doctrine: the tool name and the ToolSearch note, the D3 wiring points, explicit effort= for built-in judgment dispatches, the run slot and /route clear, levers (ultrathink, /route effort=max; builtin /effort is not a lever inside a run), last rowed skill wins and an unrowed one changes nothing, the probe-4 limit, headless OK, effort-audit.py. CHECK: [ "$(wc -l < lib/effort-shift.md)" -le 60 ] && grep -q 'mcp__model-router__route' lib/effort-shift.md && grep -q 'ToolSearch' lib/effort-shift.md && grep -q 'effort-audit.py' lib/effort-shift.md && ! grep -q 'Skill(effort' lib/effort-shift.md && echo W2B-DOCTRINE EXPECT: W2B-DOCTRINE EVIDENCE: MET exit=0 marker-found :: W2B-DOCTRINE
  8. Frontmatter aligned to the rows: agents/analyzer.md carries effort: xhigh; every other tracked effort:/model: value equals its row (criterion 4 census proves it); CLAUDE.global.md Design-work lines no longer cite lib/effort-pins.txt or the lone-Skill-call rule. CHECK: grep -q '^effort: xhigh' agents/analyzer.md && ! grep -q 'effort-pins' CLAUDE.global.md && ! grep -q 'lone Skill call' CLAUDE.global.md && echo W2B-FRONTMATTER EXPECT: W2B-FRONTMATTER EVIDENCE: MET exit=0 marker-found :: W2B-FRONTMATTER
  9. Prose sites that stated pin semantics ("sonnet by frontmatter pin", "model: opus-pinned in their frontmatter, session-independent", STOP texts naming /effort-max) are reworded to the row/off-state-floor semantics and the D4 levers. Judged by reading.
  10. The suites the diff touches are green (doctrine-citers, loops-light, plan-challenger, skill-routing-census, profile-census, portability-census, gitflow-test if present) and shellcheck *.sh hooks/*.sh lib/*.sh is clean; the full make test runs once before the merge (LRN-209), the orchestrator reports it. CHECK: for t in doctrine-citers loops-light plan-challenger skill-routing-census profile-census portability-census; do [ -f "lib/tests/$t.test.sh" ] || continue; bash "lib/tests/$t.test.sh" >/dev/null 2>&1 || { echo "red: $t"; exit 1; }; done; shellcheck ./.sh hooks/.sh lib/*.sh && echo W2B-SUITE-GREEN EXPECT: W2B-SUITE-GREEN EVIDENCE: MET exit=0 marker-found :: W2B-SUITE-GREEN

FILE SCOPE

mods/model-router/hooks/register.ts, mods/model-router/hooks/register.test.ts; lib/effort-shift.md, lib/model-gate.md, lib/verify-secure-loop.md, lib/challenge-plan.md; skills/{ship-feature,init-project,feat,bugfix,web-validate,seo,hotfix,geo,harden,code-clean,audit-delta,onboard,tour,client-handover,analyze,commit-change}/SKILL.md; agents/client-handover-writer.md, agents/analyzer.md, agents/plan-challenger.md (prose only), agents/{commit-changer,handover-doc-writer,geo-analyzer,doc-syncer,seo-analyzer}.md (prose only); deletions listed in criterion 2; install-plugins.sh, update-all.sh; lib/tests/{effort-routing,model-routing,higgsfield}.test.sh; CLAUDE.global.md (Design-work lines only).