chore(memory): BDR-115 amendment 2 + LRN-210/211 + EVAL-042 + journal/TODO/contract w2b — feat model-router wave 2
This commit is contained in:
@@ -18,9 +18,12 @@ migration of shifters/pins/model-gate in wave 2 after proof; names model-router
|
||||
- [x] W1-C live verification part 1 (2026-10-09 after `/reload-plugins`, skills-dir copy loaded, 12 hooks): derived orchestrate on a real background Explore (main high → medium while it ran → high after its end); Explore on sonnet/medium; `route plan` → next request xhigh (engine record); `Skill(effort-low)` bridged → next request low (engine record); `/route show` resolves all 10 phases to full ids, `down: none`; steps arrive as bare `claude-fable-5-1` (no `[1m]`) so the suffix carry never fires on this session
|
||||
- [ ] W1-C live verification part 2 (needs a real incident): StopFailure vs turn.complete order; `PostModelSwitch` `auto` semantics and its `from_model` after a router upgrade; whether a router rewrite raises `auto`; `[1m]` carry validity on opus (only on a session whose steps carry it)
|
||||
- [x] W2-A mod (2026-10-09, bb56f3e on feature/model-router-w2, contract `2026-10-09-model-router-w2a-1546`, plan r4 `2026-10-09-model-router-w2-1546`): user decisions — effort-* skills deleted (W2-B), pins REWORKED not deleted (rows = phases by role; frontmatter `model:`/`effort:` kept as census-locked off-state floor after a robustness BLOCKER), orchestrators declare phases, slim model gate. Phases `write` work/high + `apply` work/low; 56 skill rows, 21 agent rows; agents' model at spawn within tier, upward only, project-defined agents skipped (agent.offer); typed slash → name-bound marker (composer|sdk|bridge) + pending slot + idle fallback; best-tier rows in a `runMain` slot surviving turn end; unrowed skill leaves the route; route answer always names the id; `null` override rows. 3 lenses + 2 confirmations (1 BLOCKER each round closed), feater + 4 gap rounds, GATE 0 MET, verifier 3× ECARTS on coverage clauses only → user accepted at the cap, security PASS (2 MEDIUM pre-existing: first-load kill switch fails open, ReDoS size-bounded). Kit 58 → 88 tests. Doc-sync skipped for A (mod-only, docs at B).
|
||||
- [ ] W2 gate A→B (user): `/reload-plugins`, then the live probe of plan § Gate (4 points: typed `/status` marker vs fallback in the verbose log; a real rowed spawn line + `step 0 agent` effort = spawn/first-step ordering; sonnet session typed `/feat` self-check + route answer; probe 1 again with a background agent alive). `typed-marker` never seen → W2-B blocked, A4 re-planned.
|
||||
- [ ] W2-B repo migration (plan § W2-B B0-B7 + STEP 6/7): bridge removal, `lib/effort-shift.md` rewrite, 15 citers → `route`, `effort=` on opus general-purpose dispatches, slim `lib/model-gate.md`, delete `model-check.sh` + `effort-pins.*` + their tests + install/update blocks, delete `skills/effort-*`, analyzer `effort: xhigh`, census rewrite (drift lock rows ↔ frontmatter), docs + registries (BDR-115 amendment, LRN typed-slash/run slot, EVAL)
|
||||
- [x] W2 gate A→B (2026-10-10, engine records): typed `/status` → main low then medium during the dispatch; analyzer opus/xhigh from step 0 (frontmatter high) = spawn bookkeeping first, row beats frontmatter; status-reporter haiku step 0. Probe 3 dropped (gate always calls route). Probe 4 (typed skill while a background agent is live: marker vs fallback) NOT observed → limit in `lib/effort-shift.md`; re-check in the records whenever the user types a skill with an agent alive.
|
||||
- [x] W2-B repo migration (2026-10-10, 1f2d33b, contract `2026-10-10-model-router-w2b-1045`): 15 citers → `mcp__model-router__route` per phase, `effort=` on built-in judgment dispatches + SDD sonnet implementers (`effort="medium"`), `lib/effort-shift.md` + `lib/model-gate.md` rewritten (witness = route answer, `/route on`), deletions (effort-*, effort-pins.*, model-check.sh, tests, installer blocks), bridge removed from the mod, census rewritten (140 checks, drift lock rows ↔ frontmatter), analyzer xhigh, CLAUDE.global.md Design-work line. Verifier ECARTS(7) → CONFORME 10/10, security PASS, full make test green (design-tool-gate env red). Docs P1-P8 user-approved; SemVer: breaking → next release 3.0.0 (MIGRATION section at release).
|
||||
- [ ] (was) W2 gate A→B (user): `/reload-plugins`, then the live probe of plan § Gate (4 points: typed `/status` marker vs fallback in the verbose log; a real rowed spawn line + `step 0 agent` effort = spawn/first-step ordering; sonnet session typed `/feat` self-check + route answer; probe 1 again with a background agent alive). `typed-marker` never seen → W2-B blocked, A4 re-planned.
|
||||
- [x] (was) W2-B repo migration (plan § W2-B B0-B7 + STEP 6/7): bridge removal, `lib/effort-shift.md` rewrite, 15 citers → `route`, `effort=` on opus general-purpose dispatches, slim `lib/model-gate.md`, delete `model-check.sh` + `effort-pins.*` + their tests + install/update blocks, delete `skills/effort-*`, analyzer `effort: xhigh`, census rewrite (drift lock rows ↔ frontmatter), docs + registries (BDR-115 amendment, LRN typed-slash/run slot, EVAL)
|
||||
- [ ] W2-A residuals (security, accepted): first-load failure of the override activates the router despite `enabled:false` (fix = treat a failed first load as off); ReDoS on a self-authored prompt pattern (size-bounded); error text in the local log; model alias keys unvalidated (PHASE_KEY would do); `offers` map uncapped; `__proto__`/`constructor` override keys untested. Known limits: `skillCalls`/`spawning` counters are global; offers keyed by name only; builtin `/effort` is not a lever inside a run.
|
||||
- [ ] W2 residuals: probe 4 unobserved (above); the push-guard hook denies a read-only grep (or a heredoc) whose TEXT contains the push verb next to `git` (security-auditor + orchestrator 2026-10-10, false positives, reworded); ~27 loose "sonnet pin"/"opus pin" shorthand sites kept (true by census); README config key list omits `tiers`/`fallback`/`cooldownMinutes`/`mainUpgrade`/`upgradeMaxTokens` (doc audit item 6, pre-existing); MIGRATION.md "Upgrading to 3.0.0" at release time; `skillCalls`/`spawning` counters global; offers keyed by name.
|
||||
- [ ] W3 optional: step heuristics, haiku classifier, quota-aware downgrade, A/B
|
||||
|
||||
## 2026-09-30 — Higgsfield pack: CLI + skills in the install process, off by default (feature/higgsfield-pack)
|
||||
|
||||
@@ -0,0 +1,55 @@
|
||||
# CONTRACT — model-router-w2b
|
||||
- date: 2026-10-10 | flow: feat | branch: feature/model-router-w2
|
||||
- status: active
|
||||
- wave: W2-B (repo migration), plan r4 `.claude/tasks/plans/2026-10-09-model-router-w2-1546.md` § D1-D4, § Row tables, § W2-B B0-B7. Follows W2-A (bb56f3e) and the A→B gate.
|
||||
|
||||
## REQUEST (verbatim — IMMUTABLE)
|
||||
tu peux lancer la vague 2
|
||||
(TODO W2 line: 15 skills `Skill(effort-*)` → `route` tool calls; remove `skills/effort-*`, `lib/effort-pins.txt/.sh`, install/update steps, `effort:` frontmatter on skills and agents; repo agents into the mod's `agents` table (verify the spawn/first-step ordering first); `lib/model-gate.md` + `lib/model-check.sh` → mod rule; census tests repointed; `lib/effort-shift.md` rewritten; docs)
|
||||
|
||||
## CLARIFICATIONS
|
||||
Q: pass B 2026-10-09 (see the W2-A contract) / A: effort-* deleted; pins reworked into phase rows with the tracked frontmatter kept as census-locked off-state floor; orchestrators declare phases; slim model gate [gated 2026-10-09]
|
||||
Q: A→B gate evidence (engine records 2026-10-10) / A: typed `/status` → main effort low then medium during the dispatch (row applied to a typed slash); analyzer spawn → opus/xhigh from step 0 while its frontmatter still said high (spawn bookkeeping precedes the first step; the row beats the frontmatter); status-reporter haiku from step 0. Probe 4 (typed skill while a background agent is live: marker vs fallback) NOT observed → recorded as a limit in `lib/effort-shift.md`, re-checked later in the records; not blocking [gated 2026-10-10]
|
||||
Q: slim gate design / A: the include ALWAYS calls the route tool at entry (`mcp__model-router__route` phase = the skill's row) and reads the id in the answer; no self-check wording dependency; the tool is deferred for the model → `ToolSearch("select:mcp__model-router__route")` once per session when it is not loaded [gated 2026-10-10]
|
||||
|
||||
Q: FLOOR items (verifier 1): deletion of `lib/tests/effort-pins.test.sh` and `lib/tests/model-check.test.sh`, removal of the two `Skill(effort-low)` bridge kit tests, the from-scratch rewrite of `lib/tests/effort-routing.test.sh` (has/ok/ko helpers, its `# shellcheck disable=SC2015,SC2016` directive carried over from HEAD) / A: all authorized — they are the plan's B0 and B4 steps, accepted at pass B (shifters, pins and witness deleted with their tests) [gated 2026-10-10]
|
||||
Q: scope addition (verifier 1, criterion 9a): `agents/plan-challenger.md:110-112` twins `lib/challenge-plan.md:46` ("`model: opus`-pinned in its frontmatter, session-independent") / A: added to FILE SCOPE, same rewording (row + off-state floor); the ~20 loose "sonnet pin"/"opus pin" shorthand sites stay (true: alias == row tier by census) [gated 2026-10-10, orchestrator — minor, user may veto]
|
||||
|
||||
## ACCEPTANCE CRITERIA
|
||||
1. No `Skill(effort-` call and no `EFFORT SHIFTS:` header line remains in tracked skills, agents and lib (the mod and CHANGELOG excepted).
|
||||
CHECK: ! git grep -nE 'Skill\(effort-|EFFORT SHIFTS:' -- skills agents lib ':!mods' >/dev/null && echo W2B-NO-SHIFTER-CITER
|
||||
EXPECT: W2B-NO-SHIFTER-CITER
|
||||
EVIDENCE: MET exit=0 marker-found :: W2B-NO-SHIFTER-CITER
|
||||
2. Deleted: `skills/effort-{low,medium,high,xhigh,max}`, `lib/effort-pins.txt`, `lib/effort-pins.sh`, `lib/model-check.sh`, `lib/tests/effort-pins.test.sh`, `lib/tests/model-check.test.sh`; no tracked reference to `effort-pins`, `apply_effort_pins` or `model-check.sh` outside CHANGELOG.md, `.claude/`, README.md and USAGE.md (the two docs are the STEP 6 doc-sync targets [gated 2026-10-10]).
|
||||
CHECK: for p in skills/effort-low skills/effort-medium skills/effort-high skills/effort-xhigh skills/effort-max lib/effort-pins.txt lib/effort-pins.sh lib/model-check.sh lib/tests/effort-pins.test.sh lib/tests/model-check.test.sh; do [ ! -e "$p" ] || { echo "still here: $p"; exit 1; }; done && ! git grep -nE 'effort-pins|apply_effort_pins|model-check\.sh' -- . ':!CHANGELOG.md' ':!.claude' ':!README.md' ':!USAGE.md' >/dev/null && echo W2B-DELETIONS-DONE
|
||||
EXPECT: W2B-DELETIONS-DONE
|
||||
EVIDENCE: MET exit=0 marker-found :: W2B-DELETIONS-DONE
|
||||
3. Mod: the `Skill(effort-*)` bridge is gone (`EFFORT_SKILL`, `effortBridge` absent), the kit suite and the mods suite are green.
|
||||
CHECK: ! grep -qE 'EFFORT_SKILL|effortBridge' mods/model-router/hooks/register.ts && cd mods/model-router && out="$(claude plugin test . 2>&1)" && printf '%s\n' "$out" | grep -qE '[0-9]+ pass' && ! printf '%s\n' "$out" | grep -qE '[1-9][0-9]* fail' && cd ../.. && make test suite=lib/tests/mods.test.sh 2>&1 | grep -q 'all suites green' && echo W2B-MOD-GREEN
|
||||
EXPECT: W2B-MOD-GREEN
|
||||
EVIDENCE: MET exit=0 marker-found :: W2B-MOD-GREEN
|
||||
4. `lib/tests/effort-routing.test.sh` is rewritten as the wave-2 census (drift lock rows ↔ frontmatter parsed from `register.ts`, no `effort:`-less routed skill, D3 wiring markers, no shifter citers) and is green; `lib/tests/model-routing.test.sh` and `lib/tests/higgsfield.test.sh` are green.
|
||||
CHECK: grep -q 'register.ts' lib/tests/effort-routing.test.sh && bash lib/tests/effort-routing.test.sh >/dev/null 2>&1 && bash lib/tests/model-routing.test.sh >/dev/null 2>&1 && bash lib/tests/higgsfield.test.sh >/dev/null 2>&1 && echo W2B-CENSUS-GREEN
|
||||
EXPECT: W2B-CENSUS-GREEN
|
||||
EVIDENCE: MET exit=0 marker-found :: W2B-CENSUS-GREEN
|
||||
5. Wiring follows plan D3 at every former shifter site: dispatch span → `mcp__model-router__route` phase orchestrate; own level high → reflect, xhigh → plan; bookkeeping tail → apply; escalation (verify-secure-loop ×3, ship-feature error recovery) → escalate; judgment dispatches of built-ins (`general-purpose` `model="opus"` in ship-feature, init-project, onboard, tour; `model: "fable"` skill-runners in client-handover-writer) carry an explicit `effort=` param instead of a shift; the "turn reset" re-assert sites become a route call at the first step of the resumed turn. Judged by reading each site against the plan.
|
||||
6. `lib/model-gate.md` is the mod rule, ≤ 30 lines: always call `mcp__model-router__route` at entry with the skill's row phase, STOP unless the answer names a fable or opus id (tool missing → ToolSearch once; "is off" → STOP with the remedy); keeps the dispatch-tier table with `model: "fable"` for skill-runners; no `model-check` reference.
|
||||
CHECK: [ "$(wc -l < lib/model-gate.md)" -le 30 ] && grep -q 'mcp__model-router__route' lib/model-gate.md && grep -q 'model: "fable"' lib/model-gate.md && ! grep -q 'model-check' lib/model-gate.md && echo W2B-GATE-SLIM
|
||||
EXPECT: W2B-GATE-SLIM
|
||||
EVIDENCE: MET exit=0 marker-found :: W2B-GATE-SLIM
|
||||
7. `lib/effort-shift.md` is rewritten (≤ 60 lines) as the route doctrine: the tool name and the ToolSearch note, the D3 wiring points, explicit `effort=` for built-in judgment dispatches, the run slot and `/route clear`, levers (`ultrathink`, `/route effort=max`; builtin `/effort` is not a lever inside a run), last rowed skill wins and an unrowed one changes nothing, the probe-4 limit, headless OK, `effort-audit.py`.
|
||||
CHECK: [ "$(wc -l < lib/effort-shift.md)" -le 60 ] && grep -q 'mcp__model-router__route' lib/effort-shift.md && grep -q 'ToolSearch' lib/effort-shift.md && grep -q 'effort-audit.py' lib/effort-shift.md && ! grep -q 'Skill(effort' lib/effort-shift.md && echo W2B-DOCTRINE
|
||||
EXPECT: W2B-DOCTRINE
|
||||
EVIDENCE: MET exit=0 marker-found :: W2B-DOCTRINE
|
||||
8. Frontmatter aligned to the rows: `agents/analyzer.md` carries `effort: xhigh`; every other tracked `effort:`/`model:` value equals its row (criterion 4 census proves it); `CLAUDE.global.md` Design-work lines no longer cite `lib/effort-pins.txt` or the lone-Skill-call rule.
|
||||
CHECK: grep -q '^effort: xhigh' agents/analyzer.md && ! grep -q 'effort-pins' CLAUDE.global.md && ! grep -q 'lone Skill call' CLAUDE.global.md && echo W2B-FRONTMATTER
|
||||
EXPECT: W2B-FRONTMATTER
|
||||
EVIDENCE: MET exit=0 marker-found :: W2B-FRONTMATTER
|
||||
9. Prose sites that stated pin semantics ("sonnet by frontmatter pin", "`model: opus`-pinned in their frontmatter, session-independent", STOP texts naming `/effort-max`) are reworded to the row/off-state-floor semantics and the D4 levers. Judged by reading.
|
||||
10. The suites the diff touches are green (doctrine-citers, loops-light, plan-challenger, skill-routing-census, profile-census, portability-census, gitflow-test if present) and `shellcheck *.sh hooks/*.sh lib/*.sh` is clean; the full `make test` runs once before the merge (LRN-209), the orchestrator reports it.
|
||||
CHECK: for t in doctrine-citers loops-light plan-challenger skill-routing-census profile-census portability-census; do [ -f "lib/tests/$t.test.sh" ] || continue; bash "lib/tests/$t.test.sh" >/dev/null 2>&1 || { echo "red: $t"; exit 1; }; done; shellcheck ./*.sh hooks/*.sh lib/*.sh && echo W2B-SUITE-GREEN
|
||||
EXPECT: W2B-SUITE-GREEN
|
||||
EVIDENCE: MET exit=0 marker-found :: W2B-SUITE-GREEN
|
||||
|
||||
## FILE SCOPE
|
||||
mods/model-router/hooks/register.ts, mods/model-router/hooks/register.test.ts; lib/effort-shift.md, lib/model-gate.md, lib/verify-secure-loop.md, lib/challenge-plan.md; skills/{ship-feature,init-project,feat,bugfix,web-validate,seo,hotfix,geo,harden,code-clean,audit-delta,onboard,tour,client-handover,analyze,commit-change}/SKILL.md; agents/client-handover-writer.md, agents/analyzer.md, agents/plan-challenger.md (prose only), agents/{commit-changer,handover-doc-writer,geo-analyzer,doc-syncer,seo-analyzer}.md (prose only); deletions listed in criterion 2; install-plugins.sh, update-all.sh; lib/tests/{effort-routing,model-routing,higgsfield}.test.sh; CLAUDE.global.md (Design-work lines only).
|
||||
Reference in New Issue
Block a user