chore(memory): BDR-115 amendment 2 + LRN-210/211 + EVAL-042 + journal/TODO/contract w2b — feat model-router wave 2

This commit is contained in:
bchanot
2026-10-10 11:32:58 +02:00
parent 6f0df37f5b
commit c1767a5032
6 changed files with 78 additions and 3 deletions
+1 -1
View File
@@ -1402,4 +1402,4 @@ Branch feature/user-writing-web-rules, UNMERGED (human gate).
- **Gates**: plan r1 → r3 through 3 challengers + 1 confirmation (2 BLOCKER + 10 MAJOR closed, [[EVAL-040]]); feater DONE first pass; GATE 0 MET; verifier CONFORME 6/6; security PASS + hardening round (criteria 7-11). - **Gates**: plan r1 → r3 through 3 challengers + 1 confirmation (2 BLOCKER + 10 MAJOR closed, [[EVAL-040]]); feater DONE first pass; GATE 0 MET; verifier CONFORME 6/6; security PASS + hardening round (criteria 7-11).
- **Refs**: plan `.claude/tasks/plans/2026-10-08-model-router-mod.md`, contract `.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md`, [[BDR-107]], [[BDR-108]], [[BLK-029]], [[LRN-205]], [[LRN-206]]. - **Refs**: plan `.claude/tasks/plans/2026-10-08-model-router-mod.md`, contract `.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md`, [[BDR-107]], [[BDR-108]], [[BLK-029]], [[LRN-205]], [[LRN-206]].
- **Amendment (2026-10-09, user decisions 2026-10-08 evening)**: (a) LOAD supersedes the "Load:" line: tracked relative symlink `skills/model-router` → `../mods/model-router`, loaded in place as `model-router@skills-dir` wherever link.sh links `~/.claude/skills`; `CLAUDE_CODE_PLUGIN_DIRS` dropped (absolute path, settings `env` has no `$HOME` expansion, settings.json tracked), local marketplace dropped (`add` writes an absolute path into settings.json). Proven by fresh-process `claude plugin list --json`. (b) PRECEDENCE amended: `ultrathink` and a typed `/effort-<l>` are the main turn's DEFAULT and MINIMUM (floor slot `turnFloor`): sticky `/route` effort > turn route effort > floor > engine, then floored; per axis; mid-turn prompt floors the running turn and the next (`wait` ignored). Rationale: user "un choix explicite bat la phase déduite"; a pure floor made `/effort-low` a no-op (challenge finding). (c) Per-machine kill switch `"enabled": false` in the untracked `~/.claude/model-router.json` (survives `/clear`, a failed reload keeps the previous config); `enabledPlugins` would dirty the tracked settings.json on every machine. (d) Hardening: `/route` composer-only; agent loops effort-only (model fixed at spawn); config caps; typed slash attested at `prompt.submit`. Commits 346d6ae, 1ff608a, 6430ac6; contracts `2026-10-08-model-router-floor-1835`, `2026-10-08-model-router-wiring-1835`; residuals parked in TODO. - **Amendment (2026-10-09, user decisions 2026-10-08 evening)**: (a) LOAD supersedes the "Load:" line: tracked relative symlink `skills/model-router` → `../mods/model-router`, loaded in place as `model-router@skills-dir` wherever link.sh links `~/.claude/skills`; `CLAUDE_CODE_PLUGIN_DIRS` dropped (absolute path, settings `env` has no `$HOME` expansion, settings.json tracked), local marketplace dropped (`add` writes an absolute path into settings.json). Proven by fresh-process `claude plugin list --json`. (b) PRECEDENCE amended: `ultrathink` and a typed `/effort-<l>` are the main turn's DEFAULT and MINIMUM (floor slot `turnFloor`): sticky `/route` effort > turn route effort > floor > engine, then floored; per axis; mid-turn prompt floors the running turn and the next (`wait` ignored). Rationale: user "un choix explicite bat la phase déduite"; a pure floor made `/effort-low` a no-op (challenge finding). (c) Per-machine kill switch `"enabled": false` in the untracked `~/.claude/model-router.json` (survives `/clear`, a failed reload keeps the previous config); `enabledPlugins` would dirty the tracked settings.json on every machine. (d) Hardening: `/route` composer-only; agent loops effort-only (model fixed at spawn); config caps; typed slash attested at `prompt.submit`. Commits 346d6ae, 1ff608a, 6430ac6; contracts `2026-10-08-model-router-floor-1835`, `2026-10-08-model-router-wiring-1835`; residuals parked in TODO.
- **Amendment 2 (2026-10-10, wave 2 closed, commits bb56f3e + 1f2d33b)**: (a) rule 4 closed: rows for every repo skill (56) + agent (21) + Explore/Plan, PHASES by role (plan/reflect/orchestrate/escalate best · judge big · implement/write/verify/explore/apply work · mechanical cheap; `write` work/high + `apply` work/low added). User: pins "deleted or reworked", not copied → rows are the live source, tracked `model:`/`effort:` frontmatter KEPT as off-state floor, census-locked equal to rows (`lib/tests/effort-routing.test.sh`). Robustness BLOCKER closed: mod off → agents would inherit parent model. (b) Rule 5 amended: unrowed skill load changes nothing; best-tier skill row lives in `runMain` slot surviving turn end (precedence userMain > turnMain > runMain > floor > engine), dropped by `/route clear`, `/route off`, user `/model`, typed non-best skill; turn writers (route tool, prompt rules) never touch it. (c) Agents: model written at spawn WITHIN tier, upward only (never below frontmatter alias); explicit Agent params win; project-defined agent (agent.offer source projectSettings|localSettings) skipped. (d) Typed slash: name-bound marker at prompt.submit (composer|sdk|bridge) + pending slot + idle fallback (no live/spawning loop). (e) Shifters `effort-*`, `effort-pins.*`, `model-check.sh` DELETED; orchestrators call `mcp__model-router__route` per phase; gate witness = route answer id, remedy `/route on`; builtin `/effort` not a lever inside a run (levers `ultrathink`, `/route effort=max`). (f) SemVer: typed `/effort-*` removal = breaking → next release 3.0.0. Plan `.claude/tasks/plans/2026-10-09-model-router-w2-1546.md` r4, contracts `2026-10-09-model-router-w2a-1546`, `2026-10-10-model-router-w2b-1045`. Links [[LRN-210]], [[LRN-211]], [[EVAL-042]].
+7
View File
@@ -62,6 +62,7 @@ rules:
| EVAL-039 | 2026-09-30 | ship-feature run higgsfield-pack: plan dry-run in scratch → 0 executor failure on 7 tasks; challenge found 7 MAJOR I missed; floor-guard caught 2 shellcheck suppressions of mine; final review found README/code gap | keep | | EVAL-039 | 2026-09-30 | ship-feature run higgsfield-pack: plan dry-run in scratch → 0 executor failure on 7 tasks; challenge found 7 MAJOR I missed; floor-guard caught 2 shellcheck suppressions of mine; final review found README/code gap | keep |
| EVAL-040 | 2026-10-08 | model-router w1a plan: 3 challengers + 1 confirmation found 2 BLOCKER + 14 MAJOR on a plan judged closed; executor then passed every gate first time | keep the round, never dispatch a mod plan without it | | EVAL-040 | 2026-10-08 | model-router w1a plan: 3 challengers + 1 confirmation found 2 BLOCKER + 14 MAJOR on a plan judged closed; executor then passed every gate first time | keep the round, never dispatch a mod plan without it |
| EVAL-041 | 2026-10-09 | model-router W1-C plan: 3 lenses FATAL (4 BLOCKER + 20 MAJOR) then 2 confirmations each FATAL with a NEW BLOCKER in my own revision; executor DONE first pass, 3 short text/hardening rounds | one confirmation is not enough when a revision removes a whole mechanism; the plan carried the risk, the code almost none | | EVAL-041 | 2026-10-09 | model-router W1-C plan: 3 lenses FATAL (4 BLOCKER + 20 MAJOR) then 2 confirmations each FATAL with a NEW BLOCKER in my own revision; executor DONE first pass, 3 short text/hardening rounds | one confirmation is not enough when a revision removes a whole mechanism; the plan carried the risk, the code almost none |
| EVAL-042 | 2026-10-10 | model-router W2 plan r1 → r4: 3 lenses (1 BLOCKER), 2 confirmations (1 BLOCKER then 0); W2-A verifier 3× ECARTS on coverage clauses only, W2-B ECARTS(7) → CONFORME; gate A→B read from engine records | second confirmation paid again (BLOCKER on my own r2 slot); compound coverage criterion = endless ECARTS; engine jsonl replaces the live log |
--- ---
@@ -387,3 +388,9 @@ Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itse
- **Anomaly**: 6 BLOCKER + 29 MAJOR over four revisions, each confirmation found a flaw my own fix had introduced; the doctrine cap (one confirmation) would have shipped r2 with a 5-hour false outage. The executor never needed a re-dispatch for logic: all later rounds were text truthfulness and hardening. - **Anomaly**: 6 BLOCKER + 29 MAJOR over four revisions, each confirmation found a flaw my own fix had introduced; the doctrine cap (one confirmation) would have shipped r2 with a 5-hour false outage. The executor never needed a re-dispatch for logic: all later rounds were text truthfulness and hardening.
- **Action**: when a revision REMOVES or REPLACES a mechanism, re-challenge once more (state the deviation); keep plan sections additive with an explicit precedence line (r4 > r3 > r2) so executors and verifiers read one law; name superseded clauses of prior contracts in the Disposition. Links [[EVAL-040]], [[LRN-207]], [[BDR-115]]. - **Action**: when a revision REMOVES or REPLACES a mechanism, re-challenge once more (state the deviation); keep plan sections additive with an explicit precedence line (r4 > r3 > r2) so executors and verifiers read one law; name superseded clauses of prior contracts in the Disposition. Links [[EVAL-040]], [[LRN-207]], [[BDR-115]].
## EVAL-042 — model-router W2: challenge value held on the second confirmation; coverage-shaped criterion burned three verifiers
- **Date**: 2026-10-10
- **Output checked**: plan `.claude/tasks/plans/2026-10-09-model-router-w2-1546.md` r1 → r4; W2-A diff (register.ts, 88 kit tests); W2-B diff (43 files, 140-check census).
- **Method**: 3 blind lenses (simplicity CONCERNS(5), correctness CONCERNS(10), robustness FATAL(8) → BLOCKER: mod off = agents inherit parent model); confirmation 1 robustness FATAL(8) → BLOCKER introduced by r2 (route calls wiped the sticky slot); confirmation 2 correctness CONCERNS(3), no BLOCKER → r4. W2-A: feater DONE + 4 rounds (1 internal decision, 3 coverage), GATE 0 MET, verifier ECARTS(3)/(1)/(1) all coverage, user accepted at cap; security PASS. W2-B: feater DONE first pass, verifier ECARTS(7) (5 FLOOR items = planned deletions needing a CLARIFICATIONS line, 2 prose, 1 scope add) → CONFORME 10/10; security PASS; full `make test` once (env red only). Gate A→B: 3/4 probes answered from engine jsonl, probe 4 (typed skill with a live agent) unobserved, recorded as a limit.
- **Anomaly**: the user's rework answer ("delete or rework, not copy") changed the design mid-plan; r2's own fix carried a BLOCKER again (as in EVAL-041). Coverage criterion: 3 verifiers, 0 defects. FLOOR guard needs the test deletions named in CLARIFICATIONS, not only in criteria.
- **Action**: keep the "second confirmation after a mechanism change" rule; write coverage criteria one clause each (LRN-210); when a plan deletes tests, write the authorizing CLARIFICATIONS line BEFORE the first verifier. Links [[EVAL-041]], [[LRN-210]], [[LRN-211]], [[BDR-115]].
+3
View File
@@ -591,3 +591,6 @@ rules:
- /hotfix make-test-names-red-suites (user: '9 min pour un merge?'): measured from transcript timestamps, the merge took <20 s; 454 s went to make test run TWICE (full + per-suite sweep to name the red suite, because the aggregate rc is silent). Fix: Makefile test prints FAIL <suite> + summary, rc unchanged (GNU make returns 2 on a failed recipe; my first oracle expected 1). hotfixer DONE, GATE 0 MET 2/2, security PASS. Branch bugfix/make-test-names-red-suites UNMERGED (human signal). Method note: a mods/-only diff needs only the mods suite + doctrine census, the full run once before merge. - /hotfix make-test-names-red-suites (user: '9 min pour un merge?'): measured from transcript timestamps, the merge took <20 s; 454 s went to make test run TWICE (full + per-suite sweep to name the red suite, because the aggregate rc is silent). Fix: Makefile test prints FAIL <suite> + summary, rc unchanged (GNU make returns 2 on a failed recipe; my first oracle expected 1). hotfixer DONE, GATE 0 MET 2/2, security PASS. Branch bugfix/make-test-names-red-suites UNMERGED (human signal). Method note: a mods/-only diff needs only the mods suite + doctrine census, the full run once before merge.
- User go 'all, merge le, et je clear': LRN-207/208/209 + EVAL-041 written (ca96458); gitflow finish bugfix/make-test-names-red-suites → develop 5e0e5c0, branch removed. develop 28 commits ahead of origin, NOT pushed (manual mode, user publishes). Only settings.json dirty (user's own /model change). Session closes; next: wave 2 of model-router (TODO W2 line). - User go 'all, merge le, et je clear': LRN-207/208/209 + EVAL-041 written (ca96458); gitflow finish bugfix/make-test-names-red-suites → develop 5e0e5c0, branch removed. develop 28 commits ahead of origin, NOT pushed (manual mode, user publishes). Only settings.json dirty (user's own /model change). Session closes; next: wave 2 of model-router (TODO W2 line).
- model-router W2-A landed (bb56f3e, feature/model-router-w2). User go 'lance la vague 2' + 4 pass-B answers: shifters deleted, pins reworked into phase rows (not copied), phases declared, slim gate. Plan r1→r4: 3 lenses (1 BLOCKER: mod off → agents inherit parent model → frontmatter kept as off-state floor), confirmation 1 FATAL (BLOCKER: route calls wiped the sticky slot → separate runMain), confirmation 2 CONCERNS (deviation: 2 confirmations, stated). Feater DONE + 4 gap rounds (A5 in-agent decision, 6+3+3 coverage tests, every one mutation-proven). GATE 0 MET; verifier ECARTS(3)/(1)/(1) all on coverage of criterion 3 clauses, never code → diagnosis at max: compound criterion → user accepted at the cap. Security PASS (2 MEDIUM pre-existing). Kit 58→88. Next: user /reload-plugins + live probe, then W2-B. Lesson for the next contract: one coverage clause per criterion, not twelve. - model-router W2-A landed (bb56f3e, feature/model-router-w2). User go 'lance la vague 2' + 4 pass-B answers: shifters deleted, pins reworked into phase rows (not copied), phases declared, slim gate. Plan r1→r4: 3 lenses (1 BLOCKER: mod off → agents inherit parent model → frontmatter kept as off-state floor), confirmation 1 FATAL (BLOCKER: route calls wiped the sticky slot → separate runMain), confirmation 2 CONCERNS (deviation: 2 confirmations, stated). Feater DONE + 4 gap rounds (A5 in-agent decision, 6+3+3 coverage tests, every one mutation-proven). GATE 0 MET; verifier ECARTS(3)/(1)/(1) all on coverage of criterion 3 clauses, never code → diagnosis at max: compound criterion → user accepted at the cap. Security PASS (2 MEDIUM pre-existing). Kit 58→88. Next: user /reload-plugins + live probe, then W2-B. Lesson for the next contract: one coverage clause per criterion, not twelve.
## 2026-10-10
- model-router W2-B landed (1f2d33b) + gate A→B. User /reload-plugins + typed /status; I read the engine jsonl instead of the UI log: main low on typed /status, analyzer opus/xhigh at step 0 → ordering + row-over-frontmatter proven; probe 4 unobserved (user away), limit recorded. Route tool checked live (`mcp__model-router__route`, deferred → ToolSearch once; answer names the id); session itself routed through it (reflect → orchestrate). W2-B: contract 10 criteria, feater DONE first pass (43 files), verifier ECARTS(7): 5 FLOOR items (planned test deletions → CLARIFICATIONS line), `/route on` wording, false CLAUDE.global.md sentence, plan-challenger scope add → CONFORME 10/10; I folded 3 observations by hand (SDD sonnet implementers `effort="medium"`, run-slot droppers, 80-col). Security PASS (push-guard false positive on a grep pattern; it also bit my registry heredoc). Full make test green (env red only). Doc audit SIGNIFICANT → user: apply all P1-P8, SemVer BREAKING → 3.0.0, registries all → BDR-115 amendment 2, LRN-210/211, EVAL-042. Pending: doc commit, memory commit, user merge decision (gitflow finish, human signal), publish by hand.
+7
View File
@@ -1777,3 +1777,10 @@ Rule: when editing a doctrine file under structure locks, grep the test's lock s
- **Context**: 2026-10-09 the pre-merge check took 454 s = full `make test` (48 suites) + a per-suite re-run to NAME the red one (aggregate rc, permanent env red design-tool-gate). Three more full passes earlier that day for mods/-only diffs. Hotfix efdd491: the recipe prints `FAIL <suite>` + `all suites green` / `<n> suite(s) red: …`. My oracle expected rc 1; GNU make returns 2 when a recipe line fails. - **Context**: 2026-10-09 the pre-merge check took 454 s = full `make test` (48 suites) + a per-suite re-run to NAME the red one (aggregate rc, permanent env red design-tool-gate). Three more full passes earlier that day for mods/-only diffs. Hotfix efdd491: the recipe prints `FAIL <suite>` + `all suites green` / `<n> suite(s) red: …`. My oracle expected rc 1; GNU make returns 2 when a recipe line fails.
- **Apply**: a diff confined to one component runs that component's suite + the doctrine census; the full suite runs ONCE before `gitflow finish`; read the FAIL lines, never re-run per suite. Oracles on `make` test `[ $rc -ne 0 ]`, not `-eq 1`. Links [[LRN-173]]. - **Apply**: a diff confined to one component runs that component's suite + the doctrine census; the full suite runs ONCE before `gitflow finish`; read the FAIL lines, never re-run per suite. Oracles on `make` test `[ $rc -ne 0 ]`, not `-eq 1`. Links [[LRN-173]].
## LRN-210 — One coverage clause per contract criterion; a compound "covered by kit tests" criterion yields ECARTS forever with zero defects
- **Context**: W2-A contract criterion 3 bundled ~12 clauses under one "Covered by kit tests". Three fresh verifiers: ECARTS(3), (1), (1), each a NEW untested guard, code judged correct 3×. Cap hit, diagnosis at max, user accepted. 4 feater rounds added 30 tests (58 → 88), every one mutation-proven.
- **Apply**: split coverage criteria: one clause = one criterion with its own test name; or make the oracle a mutation script (`CHECK:` flips the guard, expects a red test). Verifier reads clauses literally: write only what one test can prove. Links [[LRN-209]], [[EVAL-042]].
## LRN-211 — Typed-slash routing: name-bound marker + idle fallback; run slot separate from turn routes; engine records are the live oracle
- **Context**: mod needs "user typed /feat" from `skill.prompt`, which carries no origin. `prompt.submit` sees the raw `/name` first (composer|sdk|bridge): store the NAME (not a boolean; a bare flag leaked to the next preload), pending slot when mid-turn, consume only on the matching `skill.prompt`; fallback = no live/spawning loop AND allowed origin. Sticky run route in the SAME slot as turn routes was wiped by the first `route()` call (confirmation BLOCKER) → separate `runMain`, best-tier rows only (work/cheap rows leak low effort across turns). Live facts read from `~/.claude/projects/<repo>/<session>.jsonl` (+ `subagents/agent-*.jsonl`): `effort` + `message.model` per step. Typed `/status` → main low; analyzer step 0 opus/xhigh with frontmatter high (spawn bookkeeping precedes step 0; row beats frontmatter).
- **Apply**: hook-side user-intent markers: bind to a name, add a pending slot for mid-turn, keep an ordering-independent fallback. Two lifetimes = two slots, never one slot with a source tag. Verify engine behaviour in the transcript jsonl, not in `$.ui.log`. Links [[BDR-115]], [[LRN-206]].
+5 -2
View File
@@ -18,9 +18,12 @@ migration of shifters/pins/model-gate in wave 2 after proof; names model-router
- [x] W1-C live verification part 1 (2026-10-09 after `/reload-plugins`, skills-dir copy loaded, 12 hooks): derived orchestrate on a real background Explore (main high → medium while it ran → high after its end); Explore on sonnet/medium; `route plan` → next request xhigh (engine record); `Skill(effort-low)` bridged → next request low (engine record); `/route show` resolves all 10 phases to full ids, `down: none`; steps arrive as bare `claude-fable-5-1` (no `[1m]`) so the suffix carry never fires on this session - [x] W1-C live verification part 1 (2026-10-09 after `/reload-plugins`, skills-dir copy loaded, 12 hooks): derived orchestrate on a real background Explore (main high → medium while it ran → high after its end); Explore on sonnet/medium; `route plan` → next request xhigh (engine record); `Skill(effort-low)` bridged → next request low (engine record); `/route show` resolves all 10 phases to full ids, `down: none`; steps arrive as bare `claude-fable-5-1` (no `[1m]`) so the suffix carry never fires on this session
- [ ] W1-C live verification part 2 (needs a real incident): StopFailure vs turn.complete order; `PostModelSwitch` `auto` semantics and its `from_model` after a router upgrade; whether a router rewrite raises `auto`; `[1m]` carry validity on opus (only on a session whose steps carry it) - [ ] W1-C live verification part 2 (needs a real incident): StopFailure vs turn.complete order; `PostModelSwitch` `auto` semantics and its `from_model` after a router upgrade; whether a router rewrite raises `auto`; `[1m]` carry validity on opus (only on a session whose steps carry it)
- [x] W2-A mod (2026-10-09, bb56f3e on feature/model-router-w2, contract `2026-10-09-model-router-w2a-1546`, plan r4 `2026-10-09-model-router-w2-1546`): user decisions — effort-* skills deleted (W2-B), pins REWORKED not deleted (rows = phases by role; frontmatter `model:`/`effort:` kept as census-locked off-state floor after a robustness BLOCKER), orchestrators declare phases, slim model gate. Phases `write` work/high + `apply` work/low; 56 skill rows, 21 agent rows; agents' model at spawn within tier, upward only, project-defined agents skipped (agent.offer); typed slash → name-bound marker (composer|sdk|bridge) + pending slot + idle fallback; best-tier rows in a `runMain` slot surviving turn end; unrowed skill leaves the route; route answer always names the id; `null` override rows. 3 lenses + 2 confirmations (1 BLOCKER each round closed), feater + 4 gap rounds, GATE 0 MET, verifier 3× ECARTS on coverage clauses only → user accepted at the cap, security PASS (2 MEDIUM pre-existing: first-load kill switch fails open, ReDoS size-bounded). Kit 58 → 88 tests. Doc-sync skipped for A (mod-only, docs at B). - [x] W2-A mod (2026-10-09, bb56f3e on feature/model-router-w2, contract `2026-10-09-model-router-w2a-1546`, plan r4 `2026-10-09-model-router-w2-1546`): user decisions — effort-* skills deleted (W2-B), pins REWORKED not deleted (rows = phases by role; frontmatter `model:`/`effort:` kept as census-locked off-state floor after a robustness BLOCKER), orchestrators declare phases, slim model gate. Phases `write` work/high + `apply` work/low; 56 skill rows, 21 agent rows; agents' model at spawn within tier, upward only, project-defined agents skipped (agent.offer); typed slash → name-bound marker (composer|sdk|bridge) + pending slot + idle fallback; best-tier rows in a `runMain` slot surviving turn end; unrowed skill leaves the route; route answer always names the id; `null` override rows. 3 lenses + 2 confirmations (1 BLOCKER each round closed), feater + 4 gap rounds, GATE 0 MET, verifier 3× ECARTS on coverage clauses only → user accepted at the cap, security PASS (2 MEDIUM pre-existing: first-load kill switch fails open, ReDoS size-bounded). Kit 58 → 88 tests. Doc-sync skipped for A (mod-only, docs at B).
- [ ] W2 gate A→B (user): `/reload-plugins`, then the live probe of plan § Gate (4 points: typed `/status` marker vs fallback in the verbose log; a real rowed spawn line + `step 0 agent` effort = spawn/first-step ordering; sonnet session typed `/feat` self-check + route answer; probe 1 again with a background agent alive). `typed-marker` never seen → W2-B blocked, A4 re-planned. - [x] W2 gate A→B (2026-10-10, engine records): typed `/status` → main low then medium during the dispatch; analyzer opus/xhigh from step 0 (frontmatter high) = spawn bookkeeping first, row beats frontmatter; status-reporter haiku step 0. Probe 3 dropped (gate always calls route). Probe 4 (typed skill while a background agent is live: marker vs fallback) NOT observed → limit in `lib/effort-shift.md`; re-check in the records whenever the user types a skill with an agent alive.
- [ ] W2-B repo migration (plan § W2-B B0-B7 + STEP 6/7): bridge removal, `lib/effort-shift.md` rewrite, 15 citers → `route`, `effort=` on opus general-purpose dispatches, slim `lib/model-gate.md`, delete `model-check.sh` + `effort-pins.*` + their tests + install/update blocks, delete `skills/effort-*`, analyzer `effort: xhigh`, census rewrite (drift lock rows ↔ frontmatter), docs + registries (BDR-115 amendment, LRN typed-slash/run slot, EVAL) - [x] W2-B repo migration (2026-10-10, 1f2d33b, contract `2026-10-10-model-router-w2b-1045`): 15 citers → `mcp__model-router__route` per phase, `effort=` on built-in judgment dispatches + SDD sonnet implementers (`effort="medium"`), `lib/effort-shift.md` + `lib/model-gate.md` rewritten (witness = route answer, `/route on`), deletions (effort-*, effort-pins.*, model-check.sh, tests, installer blocks), bridge removed from the mod, census rewritten (140 checks, drift lock rows ↔ frontmatter), analyzer xhigh, CLAUDE.global.md Design-work line. Verifier ECARTS(7) → CONFORME 10/10, security PASS, full make test green (design-tool-gate env red). Docs P1-P8 user-approved; SemVer: breaking → next release 3.0.0 (MIGRATION section at release).
- [ ] (was) W2 gate A→B (user): `/reload-plugins`, then the live probe of plan § Gate (4 points: typed `/status` marker vs fallback in the verbose log; a real rowed spawn line + `step 0 agent` effort = spawn/first-step ordering; sonnet session typed `/feat` self-check + route answer; probe 1 again with a background agent alive). `typed-marker` never seen → W2-B blocked, A4 re-planned.
- [x] (was) W2-B repo migration (plan § W2-B B0-B7 + STEP 6/7): bridge removal, `lib/effort-shift.md` rewrite, 15 citers → `route`, `effort=` on opus general-purpose dispatches, slim `lib/model-gate.md`, delete `model-check.sh` + `effort-pins.*` + their tests + install/update blocks, delete `skills/effort-*`, analyzer `effort: xhigh`, census rewrite (drift lock rows ↔ frontmatter), docs + registries (BDR-115 amendment, LRN typed-slash/run slot, EVAL)
- [ ] W2-A residuals (security, accepted): first-load failure of the override activates the router despite `enabled:false` (fix = treat a failed first load as off); ReDoS on a self-authored prompt pattern (size-bounded); error text in the local log; model alias keys unvalidated (PHASE_KEY would do); `offers` map uncapped; `__proto__`/`constructor` override keys untested. Known limits: `skillCalls`/`spawning` counters are global; offers keyed by name only; builtin `/effort` is not a lever inside a run. - [ ] W2-A residuals (security, accepted): first-load failure of the override activates the router despite `enabled:false` (fix = treat a failed first load as off); ReDoS on a self-authored prompt pattern (size-bounded); error text in the local log; model alias keys unvalidated (PHASE_KEY would do); `offers` map uncapped; `__proto__`/`constructor` override keys untested. Known limits: `skillCalls`/`spawning` counters are global; offers keyed by name only; builtin `/effort` is not a lever inside a run.
- [ ] W2 residuals: probe 4 unobserved (above); the push-guard hook denies a read-only grep (or a heredoc) whose TEXT contains the push verb next to `git` (security-auditor + orchestrator 2026-10-10, false positives, reworded); ~27 loose "sonnet pin"/"opus pin" shorthand sites kept (true by census); README config key list omits `tiers`/`fallback`/`cooldownMinutes`/`mainUpgrade`/`upgradeMaxTokens` (doc audit item 6, pre-existing); MIGRATION.md "Upgrading to 3.0.0" at release time; `skillCalls`/`spawning` counters global; offers keyed by name.
- [ ] W3 optional: step heuristics, haiku classifier, quota-aware downgrade, A/B - [ ] W3 optional: step heuristics, haiku classifier, quota-aware downgrade, A/B
## 2026-09-30 — Higgsfield pack: CLI + skills in the install process, off by default (feature/higgsfield-pack) ## 2026-09-30 — Higgsfield pack: CLI + skills in the install process, off by default (feature/higgsfield-pack)
@@ -0,0 +1,55 @@
# CONTRACT — model-router-w2b
- date: 2026-10-10 | flow: feat | branch: feature/model-router-w2
- status: active
- wave: W2-B (repo migration), plan r4 `.claude/tasks/plans/2026-10-09-model-router-w2-1546.md` § D1-D4, § Row tables, § W2-B B0-B7. Follows W2-A (bb56f3e) and the A→B gate.
## REQUEST (verbatim — IMMUTABLE)
tu peux lancer la vague 2
(TODO W2 line: 15 skills `Skill(effort-*)` → `route` tool calls; remove `skills/effort-*`, `lib/effort-pins.txt/.sh`, install/update steps, `effort:` frontmatter on skills and agents; repo agents into the mod's `agents` table (verify the spawn/first-step ordering first); `lib/model-gate.md` + `lib/model-check.sh` → mod rule; census tests repointed; `lib/effort-shift.md` rewritten; docs)
## CLARIFICATIONS
Q: pass B 2026-10-09 (see the W2-A contract) / A: effort-* deleted; pins reworked into phase rows with the tracked frontmatter kept as census-locked off-state floor; orchestrators declare phases; slim model gate [gated 2026-10-09]
Q: A→B gate evidence (engine records 2026-10-10) / A: typed `/status` → main effort low then medium during the dispatch (row applied to a typed slash); analyzer spawn → opus/xhigh from step 0 while its frontmatter still said high (spawn bookkeeping precedes the first step; the row beats the frontmatter); status-reporter haiku from step 0. Probe 4 (typed skill while a background agent is live: marker vs fallback) NOT observed → recorded as a limit in `lib/effort-shift.md`, re-checked later in the records; not blocking [gated 2026-10-10]
Q: slim gate design / A: the include ALWAYS calls the route tool at entry (`mcp__model-router__route` phase = the skill's row) and reads the id in the answer; no self-check wording dependency; the tool is deferred for the model → `ToolSearch("select:mcp__model-router__route")` once per session when it is not loaded [gated 2026-10-10]
Q: FLOOR items (verifier 1): deletion of `lib/tests/effort-pins.test.sh` and `lib/tests/model-check.test.sh`, removal of the two `Skill(effort-low)` bridge kit tests, the from-scratch rewrite of `lib/tests/effort-routing.test.sh` (has/ok/ko helpers, its `# shellcheck disable=SC2015,SC2016` directive carried over from HEAD) / A: all authorized — they are the plan's B0 and B4 steps, accepted at pass B (shifters, pins and witness deleted with their tests) [gated 2026-10-10]
Q: scope addition (verifier 1, criterion 9a): `agents/plan-challenger.md:110-112` twins `lib/challenge-plan.md:46` ("`model: opus`-pinned in its frontmatter, session-independent") / A: added to FILE SCOPE, same rewording (row + off-state floor); the ~20 loose "sonnet pin"/"opus pin" shorthand sites stay (true: alias == row tier by census) [gated 2026-10-10, orchestrator — minor, user may veto]
## ACCEPTANCE CRITERIA
1. No `Skill(effort-` call and no `EFFORT SHIFTS:` header line remains in tracked skills, agents and lib (the mod and CHANGELOG excepted).
CHECK: ! git grep -nE 'Skill\(effort-|EFFORT SHIFTS:' -- skills agents lib ':!mods' >/dev/null && echo W2B-NO-SHIFTER-CITER
EXPECT: W2B-NO-SHIFTER-CITER
EVIDENCE: MET exit=0 marker-found :: W2B-NO-SHIFTER-CITER
2. Deleted: `skills/effort-{low,medium,high,xhigh,max}`, `lib/effort-pins.txt`, `lib/effort-pins.sh`, `lib/model-check.sh`, `lib/tests/effort-pins.test.sh`, `lib/tests/model-check.test.sh`; no tracked reference to `effort-pins`, `apply_effort_pins` or `model-check.sh` outside CHANGELOG.md, `.claude/`, README.md and USAGE.md (the two docs are the STEP 6 doc-sync targets [gated 2026-10-10]).
CHECK: for p in skills/effort-low skills/effort-medium skills/effort-high skills/effort-xhigh skills/effort-max lib/effort-pins.txt lib/effort-pins.sh lib/model-check.sh lib/tests/effort-pins.test.sh lib/tests/model-check.test.sh; do [ ! -e "$p" ] || { echo "still here: $p"; exit 1; }; done && ! git grep -nE 'effort-pins|apply_effort_pins|model-check\.sh' -- . ':!CHANGELOG.md' ':!.claude' ':!README.md' ':!USAGE.md' >/dev/null && echo W2B-DELETIONS-DONE
EXPECT: W2B-DELETIONS-DONE
EVIDENCE: MET exit=0 marker-found :: W2B-DELETIONS-DONE
3. Mod: the `Skill(effort-*)` bridge is gone (`EFFORT_SKILL`, `effortBridge` absent), the kit suite and the mods suite are green.
CHECK: ! grep -qE 'EFFORT_SKILL|effortBridge' mods/model-router/hooks/register.ts && cd mods/model-router && out="$(claude plugin test . 2>&1)" && printf '%s\n' "$out" | grep -qE '[0-9]+ pass' && ! printf '%s\n' "$out" | grep -qE '[1-9][0-9]* fail' && cd ../.. && make test suite=lib/tests/mods.test.sh 2>&1 | grep -q 'all suites green' && echo W2B-MOD-GREEN
EXPECT: W2B-MOD-GREEN
EVIDENCE: MET exit=0 marker-found :: W2B-MOD-GREEN
4. `lib/tests/effort-routing.test.sh` is rewritten as the wave-2 census (drift lock rows ↔ frontmatter parsed from `register.ts`, no `effort:`-less routed skill, D3 wiring markers, no shifter citers) and is green; `lib/tests/model-routing.test.sh` and `lib/tests/higgsfield.test.sh` are green.
CHECK: grep -q 'register.ts' lib/tests/effort-routing.test.sh && bash lib/tests/effort-routing.test.sh >/dev/null 2>&1 && bash lib/tests/model-routing.test.sh >/dev/null 2>&1 && bash lib/tests/higgsfield.test.sh >/dev/null 2>&1 && echo W2B-CENSUS-GREEN
EXPECT: W2B-CENSUS-GREEN
EVIDENCE: MET exit=0 marker-found :: W2B-CENSUS-GREEN
5. Wiring follows plan D3 at every former shifter site: dispatch span → `mcp__model-router__route` phase orchestrate; own level high → reflect, xhigh → plan; bookkeeping tail → apply; escalation (verify-secure-loop ×3, ship-feature error recovery) → escalate; judgment dispatches of built-ins (`general-purpose` `model="opus"` in ship-feature, init-project, onboard, tour; `model: "fable"` skill-runners in client-handover-writer) carry an explicit `effort=` param instead of a shift; the "turn reset" re-assert sites become a route call at the first step of the resumed turn. Judged by reading each site against the plan.
6. `lib/model-gate.md` is the mod rule, ≤ 30 lines: always call `mcp__model-router__route` at entry with the skill's row phase, STOP unless the answer names a fable or opus id (tool missing → ToolSearch once; "is off" → STOP with the remedy); keeps the dispatch-tier table with `model: "fable"` for skill-runners; no `model-check` reference.
CHECK: [ "$(wc -l < lib/model-gate.md)" -le 30 ] && grep -q 'mcp__model-router__route' lib/model-gate.md && grep -q 'model: "fable"' lib/model-gate.md && ! grep -q 'model-check' lib/model-gate.md && echo W2B-GATE-SLIM
EXPECT: W2B-GATE-SLIM
EVIDENCE: MET exit=0 marker-found :: W2B-GATE-SLIM
7. `lib/effort-shift.md` is rewritten (≤ 60 lines) as the route doctrine: the tool name and the ToolSearch note, the D3 wiring points, explicit `effort=` for built-in judgment dispatches, the run slot and `/route clear`, levers (`ultrathink`, `/route effort=max`; builtin `/effort` is not a lever inside a run), last rowed skill wins and an unrowed one changes nothing, the probe-4 limit, headless OK, `effort-audit.py`.
CHECK: [ "$(wc -l < lib/effort-shift.md)" -le 60 ] && grep -q 'mcp__model-router__route' lib/effort-shift.md && grep -q 'ToolSearch' lib/effort-shift.md && grep -q 'effort-audit.py' lib/effort-shift.md && ! grep -q 'Skill(effort' lib/effort-shift.md && echo W2B-DOCTRINE
EXPECT: W2B-DOCTRINE
EVIDENCE: MET exit=0 marker-found :: W2B-DOCTRINE
8. Frontmatter aligned to the rows: `agents/analyzer.md` carries `effort: xhigh`; every other tracked `effort:`/`model:` value equals its row (criterion 4 census proves it); `CLAUDE.global.md` Design-work lines no longer cite `lib/effort-pins.txt` or the lone-Skill-call rule.
CHECK: grep -q '^effort: xhigh' agents/analyzer.md && ! grep -q 'effort-pins' CLAUDE.global.md && ! grep -q 'lone Skill call' CLAUDE.global.md && echo W2B-FRONTMATTER
EXPECT: W2B-FRONTMATTER
EVIDENCE: MET exit=0 marker-found :: W2B-FRONTMATTER
9. Prose sites that stated pin semantics ("sonnet by frontmatter pin", "`model: opus`-pinned in their frontmatter, session-independent", STOP texts naming `/effort-max`) are reworded to the row/off-state-floor semantics and the D4 levers. Judged by reading.
10. The suites the diff touches are green (doctrine-citers, loops-light, plan-challenger, skill-routing-census, profile-census, portability-census, gitflow-test if present) and `shellcheck *.sh hooks/*.sh lib/*.sh` is clean; the full `make test` runs once before the merge (LRN-209), the orchestrator reports it.
CHECK: for t in doctrine-citers loops-light plan-challenger skill-routing-census profile-census portability-census; do [ -f "lib/tests/$t.test.sh" ] || continue; bash "lib/tests/$t.test.sh" >/dev/null 2>&1 || { echo "red: $t"; exit 1; }; done; shellcheck ./*.sh hooks/*.sh lib/*.sh && echo W2B-SUITE-GREEN
EXPECT: W2B-SUITE-GREEN
EVIDENCE: MET exit=0 marker-found :: W2B-SUITE-GREEN
## FILE SCOPE
mods/model-router/hooks/register.ts, mods/model-router/hooks/register.test.ts; lib/effort-shift.md, lib/model-gate.md, lib/verify-secure-loop.md, lib/challenge-plan.md; skills/{ship-feature,init-project,feat,bugfix,web-validate,seo,hotfix,geo,harden,code-clean,audit-delta,onboard,tour,client-handover,analyze,commit-change}/SKILL.md; agents/client-handover-writer.md, agents/analyzer.md, agents/plan-challenger.md (prose only), agents/{commit-changer,handover-doc-writer,geo-analyzer,doc-syncer,seo-analyzer}.md (prose only); deletions listed in criterion 2; install-plugins.sh, update-all.sh; lib/tests/{effort-routing,model-routing,higgsfield}.test.sh; CLAUDE.global.md (Design-work lines only).