chore(memory): BDR-115 amendment 2 + LRN-210/211 + EVAL-042 + journal/TODO/contract w2b — feat model-router wave 2

This commit is contained in:
bchanot
2026-10-10 11:32:58 +02:00
parent 6f0df37f5b
commit c1767a5032
6 changed files with 78 additions and 3 deletions
@@ -0,0 +1,55 @@
# CONTRACT — model-router-w2b
- date: 2026-10-10 | flow: feat | branch: feature/model-router-w2
- status: active
- wave: W2-B (repo migration), plan r4 `.claude/tasks/plans/2026-10-09-model-router-w2-1546.md` § D1-D4, § Row tables, § W2-B B0-B7. Follows W2-A (bb56f3e) and the A→B gate.
## REQUEST (verbatim — IMMUTABLE)
tu peux lancer la vague 2
(TODO W2 line: 15 skills `Skill(effort-*)` → `route` tool calls; remove `skills/effort-*`, `lib/effort-pins.txt/.sh`, install/update steps, `effort:` frontmatter on skills and agents; repo agents into the mod's `agents` table (verify the spawn/first-step ordering first); `lib/model-gate.md` + `lib/model-check.sh` → mod rule; census tests repointed; `lib/effort-shift.md` rewritten; docs)
## CLARIFICATIONS
Q: pass B 2026-10-09 (see the W2-A contract) / A: effort-* deleted; pins reworked into phase rows with the tracked frontmatter kept as census-locked off-state floor; orchestrators declare phases; slim model gate [gated 2026-10-09]
Q: A→B gate evidence (engine records 2026-10-10) / A: typed `/status` → main effort low then medium during the dispatch (row applied to a typed slash); analyzer spawn → opus/xhigh from step 0 while its frontmatter still said high (spawn bookkeeping precedes the first step; the row beats the frontmatter); status-reporter haiku from step 0. Probe 4 (typed skill while a background agent is live: marker vs fallback) NOT observed → recorded as a limit in `lib/effort-shift.md`, re-checked later in the records; not blocking [gated 2026-10-10]
Q: slim gate design / A: the include ALWAYS calls the route tool at entry (`mcp__model-router__route` phase = the skill's row) and reads the id in the answer; no self-check wording dependency; the tool is deferred for the model → `ToolSearch("select:mcp__model-router__route")` once per session when it is not loaded [gated 2026-10-10]
Q: FLOOR items (verifier 1): deletion of `lib/tests/effort-pins.test.sh` and `lib/tests/model-check.test.sh`, removal of the two `Skill(effort-low)` bridge kit tests, the from-scratch rewrite of `lib/tests/effort-routing.test.sh` (has/ok/ko helpers, its `# shellcheck disable=SC2015,SC2016` directive carried over from HEAD) / A: all authorized — they are the plan's B0 and B4 steps, accepted at pass B (shifters, pins and witness deleted with their tests) [gated 2026-10-10]
Q: scope addition (verifier 1, criterion 9a): `agents/plan-challenger.md:110-112` twins `lib/challenge-plan.md:46` ("`model: opus`-pinned in its frontmatter, session-independent") / A: added to FILE SCOPE, same rewording (row + off-state floor); the ~20 loose "sonnet pin"/"opus pin" shorthand sites stay (true: alias == row tier by census) [gated 2026-10-10, orchestrator — minor, user may veto]
## ACCEPTANCE CRITERIA
1. No `Skill(effort-` call and no `EFFORT SHIFTS:` header line remains in tracked skills, agents and lib (the mod and CHANGELOG excepted).
CHECK: ! git grep -nE 'Skill\(effort-|EFFORT SHIFTS:' -- skills agents lib ':!mods' >/dev/null && echo W2B-NO-SHIFTER-CITER
EXPECT: W2B-NO-SHIFTER-CITER
EVIDENCE: MET exit=0 marker-found :: W2B-NO-SHIFTER-CITER
2. Deleted: `skills/effort-{low,medium,high,xhigh,max}`, `lib/effort-pins.txt`, `lib/effort-pins.sh`, `lib/model-check.sh`, `lib/tests/effort-pins.test.sh`, `lib/tests/model-check.test.sh`; no tracked reference to `effort-pins`, `apply_effort_pins` or `model-check.sh` outside CHANGELOG.md, `.claude/`, README.md and USAGE.md (the two docs are the STEP 6 doc-sync targets [gated 2026-10-10]).
CHECK: for p in skills/effort-low skills/effort-medium skills/effort-high skills/effort-xhigh skills/effort-max lib/effort-pins.txt lib/effort-pins.sh lib/model-check.sh lib/tests/effort-pins.test.sh lib/tests/model-check.test.sh; do [ ! -e "$p" ] || { echo "still here: $p"; exit 1; }; done && ! git grep -nE 'effort-pins|apply_effort_pins|model-check\.sh' -- . ':!CHANGELOG.md' ':!.claude' ':!README.md' ':!USAGE.md' >/dev/null && echo W2B-DELETIONS-DONE
EXPECT: W2B-DELETIONS-DONE
EVIDENCE: MET exit=0 marker-found :: W2B-DELETIONS-DONE
3. Mod: the `Skill(effort-*)` bridge is gone (`EFFORT_SKILL`, `effortBridge` absent), the kit suite and the mods suite are green.
CHECK: ! grep -qE 'EFFORT_SKILL|effortBridge' mods/model-router/hooks/register.ts && cd mods/model-router && out="$(claude plugin test . 2>&1)" && printf '%s\n' "$out" | grep -qE '[0-9]+ pass' && ! printf '%s\n' "$out" | grep -qE '[1-9][0-9]* fail' && cd ../.. && make test suite=lib/tests/mods.test.sh 2>&1 | grep -q 'all suites green' && echo W2B-MOD-GREEN
EXPECT: W2B-MOD-GREEN
EVIDENCE: MET exit=0 marker-found :: W2B-MOD-GREEN
4. `lib/tests/effort-routing.test.sh` is rewritten as the wave-2 census (drift lock rows ↔ frontmatter parsed from `register.ts`, no `effort:`-less routed skill, D3 wiring markers, no shifter citers) and is green; `lib/tests/model-routing.test.sh` and `lib/tests/higgsfield.test.sh` are green.
CHECK: grep -q 'register.ts' lib/tests/effort-routing.test.sh && bash lib/tests/effort-routing.test.sh >/dev/null 2>&1 && bash lib/tests/model-routing.test.sh >/dev/null 2>&1 && bash lib/tests/higgsfield.test.sh >/dev/null 2>&1 && echo W2B-CENSUS-GREEN
EXPECT: W2B-CENSUS-GREEN
EVIDENCE: MET exit=0 marker-found :: W2B-CENSUS-GREEN
5. Wiring follows plan D3 at every former shifter site: dispatch span → `mcp__model-router__route` phase orchestrate; own level high → reflect, xhigh → plan; bookkeeping tail → apply; escalation (verify-secure-loop ×3, ship-feature error recovery) → escalate; judgment dispatches of built-ins (`general-purpose` `model="opus"` in ship-feature, init-project, onboard, tour; `model: "fable"` skill-runners in client-handover-writer) carry an explicit `effort=` param instead of a shift; the "turn reset" re-assert sites become a route call at the first step of the resumed turn. Judged by reading each site against the plan.
6. `lib/model-gate.md` is the mod rule, ≤ 30 lines: always call `mcp__model-router__route` at entry with the skill's row phase, STOP unless the answer names a fable or opus id (tool missing → ToolSearch once; "is off" → STOP with the remedy); keeps the dispatch-tier table with `model: "fable"` for skill-runners; no `model-check` reference.
CHECK: [ "$(wc -l < lib/model-gate.md)" -le 30 ] && grep -q 'mcp__model-router__route' lib/model-gate.md && grep -q 'model: "fable"' lib/model-gate.md && ! grep -q 'model-check' lib/model-gate.md && echo W2B-GATE-SLIM
EXPECT: W2B-GATE-SLIM
EVIDENCE: MET exit=0 marker-found :: W2B-GATE-SLIM
7. `lib/effort-shift.md` is rewritten (≤ 60 lines) as the route doctrine: the tool name and the ToolSearch note, the D3 wiring points, explicit `effort=` for built-in judgment dispatches, the run slot and `/route clear`, levers (`ultrathink`, `/route effort=max`; builtin `/effort` is not a lever inside a run), last rowed skill wins and an unrowed one changes nothing, the probe-4 limit, headless OK, `effort-audit.py`.
CHECK: [ "$(wc -l < lib/effort-shift.md)" -le 60 ] && grep -q 'mcp__model-router__route' lib/effort-shift.md && grep -q 'ToolSearch' lib/effort-shift.md && grep -q 'effort-audit.py' lib/effort-shift.md && ! grep -q 'Skill(effort' lib/effort-shift.md && echo W2B-DOCTRINE
EXPECT: W2B-DOCTRINE
EVIDENCE: MET exit=0 marker-found :: W2B-DOCTRINE
8. Frontmatter aligned to the rows: `agents/analyzer.md` carries `effort: xhigh`; every other tracked `effort:`/`model:` value equals its row (criterion 4 census proves it); `CLAUDE.global.md` Design-work lines no longer cite `lib/effort-pins.txt` or the lone-Skill-call rule.
CHECK: grep -q '^effort: xhigh' agents/analyzer.md && ! grep -q 'effort-pins' CLAUDE.global.md && ! grep -q 'lone Skill call' CLAUDE.global.md && echo W2B-FRONTMATTER
EXPECT: W2B-FRONTMATTER
EVIDENCE: MET exit=0 marker-found :: W2B-FRONTMATTER
9. Prose sites that stated pin semantics ("sonnet by frontmatter pin", "`model: opus`-pinned in their frontmatter, session-independent", STOP texts naming `/effort-max`) are reworded to the row/off-state-floor semantics and the D4 levers. Judged by reading.
10. The suites the diff touches are green (doctrine-citers, loops-light, plan-challenger, skill-routing-census, profile-census, portability-census, gitflow-test if present) and `shellcheck *.sh hooks/*.sh lib/*.sh` is clean; the full `make test` runs once before the merge (LRN-209), the orchestrator reports it.
CHECK: for t in doctrine-citers loops-light plan-challenger skill-routing-census profile-census portability-census; do [ -f "lib/tests/$t.test.sh" ] || continue; bash "lib/tests/$t.test.sh" >/dev/null 2>&1 || { echo "red: $t"; exit 1; }; done; shellcheck ./*.sh hooks/*.sh lib/*.sh && echo W2B-SUITE-GREEN
EXPECT: W2B-SUITE-GREEN
EVIDENCE: MET exit=0 marker-found :: W2B-SUITE-GREEN
## FILE SCOPE
mods/model-router/hooks/register.ts, mods/model-router/hooks/register.test.ts; lib/effort-shift.md, lib/model-gate.md, lib/verify-secure-loop.md, lib/challenge-plan.md; skills/{ship-feature,init-project,feat,bugfix,web-validate,seo,hotfix,geo,harden,code-clean,audit-delta,onboard,tour,client-handover,analyze,commit-change}/SKILL.md; agents/client-handover-writer.md, agents/analyzer.md, agents/plan-challenger.md (prose only), agents/{commit-changer,handover-doc-writer,geo-analyzer,doc-syncer,seo-analyzer}.md (prose only); deletions listed in criterion 2; install-plugins.sh, update-all.sh; lib/tests/{effort-routing,model-routing,higgsfield}.test.sh; CLAUDE.global.md (Design-work lines only).