# PLAN r4 — model-router wave 2 (migration), 2026-10-09 r1 → r2 after 3 lenses (simplicity CONCERNS(5), correctness CONCERNS(10), robustness FATAL(8)); r2 → r3 after the confirmation pass (robustness FATAL(8): 1 BLOCKER); r3 → r4 after a second confirmation (correctness CONCERNS(3), wording-level). Every BLOCKER/MAJOR is closed by a NAMED change (§ Challenge ledger). Mod = the router while on; the tracked frontmatter (`model:` + `effort:` on agents, `effort:` on skills) STAYS as the off-state floor, census-locked equal to the rows (r3). Two sub-runs on `feature/model-router-w2`: W2-A (mod) then, after a user `/reload-plugins` + live probe, W2-B (repo migration). Docs + registries = W2-B's own STEP 6/7 (feat pipeline tail); W2-A runs no doc-sync (deviation, stated: a mod-only diff has no public doc of its own until B lands). ## Decisions (pass B, user 2026-10-09) — amended by the challenge - D1 the five `effort-*` skills are DELETED in W2-B. The mod's `Skill(effort-*)` bridge is deleted in W2-B too (same commit as the citers, robustness 7: edits are live on the symlinked tree). The typed `/effort-` floor code goes in W2-A (A1). The `ultrathink` floor stays. - D2 pins reworked, not deleted (r3): rows are PHASES by ROLE, no bare levels; the mod routes agents on both axes while on (model written at spawn from the tier, only UPWARD in rank; effort per step; explicit Agent params win). The tracked frontmatter stays as the OFF-STATE FLOOR (mod off / `enabled:false` / unloaded / a hook failing open → the engine applies the frontmatter as today: never the parent model, never session effort on an xhigh gate agent; r2 robustness 1 + confirmation 8). The census locks every row equal to its frontmatter (tier head == `model:` alias, phase effort == `effort:`), so there is one declared value, two carriers. Deleted: the five shifters, `lib/effort-pins.*` (vendored levels become rows; off-state = session level for them, as before BDR-108), the witness script. - D3 the orchestrators declare PHASES: dispatch span → `route(phase= "orchestrate")`; own level high → `reflect`, xhigh → `plan`; bookkeeping tail → `route(phase="apply")` (work/low); escalation → `escalate`. Judgment dispatches of BUILT-INS (`general-purpose` with `model="opus"` or `"fable"`) carry an explicit `effort=` param (a main route never reaches a child; correctness 8). A best-tier skill row SURVIVES the end of the turn in its own slot (A6 `runMain`): a run spans prose gates; turn-scoped routes (`route` calls, prompt rules, the bridge) never touch it. - D4 `lib/model-gate.md` = the mod rule; the witness is the `route` tool's own answer (correctness 7): self-check big → silent; small → call `route(phase=)` and STOP unless the answer names a fable or opus id; tool absent or "is off" → STOP with the remedy. `model-check.sh` + its test deleted. The route answer ALWAYS names the id the next main step runs on (A3d). Relaunch levers in STOP texts: `ultrathink` in the relaunch prompt (turn floor) or `/route effort=max` (sticky, `/route clear` after). Builtin `/effort` is NOT a lever inside a run (rows and routes rank above the engine effort; r2's engine-effort detector dropped as unsafe, confirmation 4) — documented, named to the user. ## Phase table (A2a) — two rows added to `phases` | phase | tier | effort | role | | plan | best | xhigh | brainstorm, plan, architecture, audit verdict | | reflect | best | high | diagnosis, review, contract, day-to-day orchestrator entry | | orchestrate | best | medium | between dispatches | | escalate | best | max | stuck, cap reached | | judge | big | xhigh | dispatched challengers, analyzers, audits | | implement | work | medium | code from a closed plan | | write | work | **high** (was medium) | docs, commits, refactors, handover prose on sonnet at high (BDR-107 "high judgment on sonnet") | | verify | work | xhigh | verifier, security-auditor | | explore | work | medium | Explore | | **apply** (new) | work | low | low appliers (BDR-107): small fixes, release mechanics, probes, validators; bookkeeping tail of the main loop | | mechanical | cheap | low | listing, status, profile toggles | ## Row tables (A2b) skills (56): - plan: ship-feature init-project onboard tour audit-delta analyze code-clean client-handover brainstorming writing-plans requesting-code-review 21st-ui-review - reflect: feat hotfix bugfix refactor web-validate harden seo geo site-motion frontend-design emil-design-eng design-motion-principles 21st-ui-build scroll-world-storytelling build-threejs-scroll-worlds scroll-scrubbed-visual-sequence scroll-scrubbed-word-reveal scroll-progress-timeline subagent-driven-development writing-skills deprecation-and-migration 21st-ai 21st-ui-explore - implement: gitflow prune-memory pdf-translate ci-cd-and-automation observability-and-instrumentation test-driven-development - apply: commit-change release-candidate doc capitalize close reconcile deploy (work tier: never haiku on main even with the switch on, robustness 15) - mechanical: status profile plugin-check skills-perso using-git-worktrees 21st-cli-use 21st-registry 21st-design-sync agents (21 + 2 built-ins), model = frontmatter alias, unchanged everywhere; `effort:` frontmatter updated where the row differs (analyzer → xhigh): - implement (sonnet/medium): feater bugfixer code-cleaner scaffolder onboarder - write (sonnet/high): commit-changer doc-syncer handover-doc-writer refactorer - apply (sonnet/low): hotfixer release-executor plugin-probe validator-analyzer - verify (sonnet/xhigh): verifier security-auditor - judge (opus/xhigh): plan-challenger plugin-advisor seo-analyzer geo-analyzer analyzer (the ONE level delta: high → xhigh, BDR-108 rung "audit before validation"; named at the gate) - mechanical (haiku/low): status-reporter - built-ins: Explore explore, Plan judge - no row: interviewer, client-handover-writer (inline-load), impeccable-* (gitignored vendor output, untouched) Deltas named at the gate: analyzer effort; best-tier raise now also reaches refactor + the design/vendored reflect skills on a small session (simplicity 6: the raise is the point of the mod). ## W2-A — mod (contract 2026-10-09-model-router-w2a-1546) - [ ] A1 `register.ts`: delete the typed-floor code — `slashEffort`, `guardedSlash`, `Source` member `'slash'`, `floorSource`'s `typed /` branch, the State comments on the typed `/effort-` floor. KEEP `EFFORT_SKILL` + `effortBridge` (tool.call bridge, removed in W2-B) and `turnFloor`/`higherFloor`/`floorWord` (ultrathink). - [ ] A2 `DEFAULT_CONFIG`: phases per § Phase table (`write` high, `apply` new); `skills` + `agents` per § Row tables; the "Built-ins only" comment → "phase = role; one row per routed repo skill/agent (wave 2); a project-level agent of the same name shadows its row (A3)". - [ ] A3 spawn: (a) explicit model = `e.model !== undefined` at spawn (the engine does not pre-fill the frontmatter; no tool_use_id map); explicit effort as today (`explicitEffort`). (b) `spawnRoute`: `frozen` → none; row lookup for every provider; a row is SKIPPED when the agent's definition `source`, recorded per `subagentType` from the `agent.offer` event, is `projectSettings` or `localSettings` (a foreign repo's own `verifier.md`); `userSettings`, `built-in`, `plugin` or no record → the row applies (fail-open on routing, as W1). Known limit, stated in a comment: the record is keyed by name only (an offer fired inside a sub-agent with another cwd overwrites it). (c) `spawnTarget` for a rowed spawn without explicit model: resolve WITHIN the tier only, and write only an alias ranked ≥ the tier head in `fallback` (agents move UP, never below their frontmatter alias): `big` with opus down → fable; opus + fable down → no write + one log line (deduped per agent+tier per turn) "model-router: tier down, frontmatter model kept". (d) `routedText`/`mainAnswer` always name the id the next main step runs on (`model ()`), the sticky/floor note APPENDED, never substituted (the gate reads this answer). - [ ] A4 typed slash: `typedSlash: string | null` = the first token of a slash prompt at `prompt.submit` when `origin.kind` ∈ {composer, sdk, bridge} (allowlist; floor/default rules stay composer-only), stored only when it is a `cfg.skills` key; mid-turn → `pendingSlash`, promoted at `endMainTurn`. `prompt.submit` also records `st.promptAllowed` (origin in the allowlist) for the turn it opens (pending slot for a mid-turn prompt, like the marker). `skill.prompt` on main (`skillCalls === 0`): apply the row when `e.skill === typedSlash` (then null it) OR when `promptAllowed && loops.size === 0 && spawning === 0` (`spawning` = a counter held from spawn-hook entry to after `next`); otherwise `next(e)`. Verbose log names which path fired (`typed-marker` / `typed-fallback`). - [ ] A5 `onSkillLoad`: a skill WITHOUT a row leaves `turnMain` untouched (find-docs / gstack / plugin skills mid-run no longer clear the run's route; correctness 11); a rowed skill replaces it. Same rule INSIDE a sub-agent (gated 2026-10-09, feater NEED-DECISION): an unrowed skill leaves the agent loop's effort untouched; a rowed one writes it (test: `feater` loop at medium, `Skill(find-docs)` with that agentId → next step still medium). - [ ] A6 run slot: new `runMain: Routed | null`. A rowed skill load ON MAIN whose phase tier is `best` writes BOTH `turnMain` (source `skill`, as today) and `runMain`; a non-best row loaded by the model's `Skill` tool writes `turnMain` only (helper skills such as `using-git-worktrees` inside SDD never drop the run); a USER-TYPED rowed skill (marker path) replaces `runMain` with its row when best, drops it when not. Precedence `userMain > turnMain > runMain > floor > engine` (per axis, as today); `mainRoute` (statusline, spinner, planStep) includes it; `endMainTurn` leaves it; `pushOrchestrate` writes `turnMain` when empty exactly as today (derived orchestrate overrides the run default for the dispatch span, like a `route orchestrate` call); cleared by `/route clear` (text: "run slot dropped" when one held), `/route off`, a user `/model` switch (`PostModelSwitch` source command|picker|sdk); `route(clear=true)` from the model clears `turnMain` only and its answer says "run still holds". A `Skill` call inside a sub-agent never touches it. `/route show` prints `run ` when no turn route is in force. - [ ] A7 (dropped in r3: engine-effort detector, confirmation 4). - [ ] A8 `mergeTable` accepts `null` in the override's `skills`/`agents` to drop a default row (robustness 4). - [ ] A9 `register.test.ts`: delete the typed `/effort-*` floor tests; keep the bridge test; add — typed `/feat` (prompt.submit `/feat x`, origin composer, then skill.prompt feat) → `main: skill reflect`, `effort high`, `[tier best]`; origin `channel`: not armed AND the idle fallback refused (main untouched); skill.prompt `feat` with a live sub-agent and no marker → main untouched; same with no live loop and an allowed origin → routed; a mid-turn `/status` → pending, applied after turn.complete; `Skill(find-docs)` (no row) after `route plan` keeps plan — the existing test 'floor: survives a skill load' (register.test.ts:341-348) is REWRITTEN to assert orchestrate KEPT (A5 behavior, authorized here); run slot: `Skill(feat)`, then `route orchestrate`, then `turn.complete` → `/route show` back to `run reflect`; `Skill(using-git-worktrees)` (mechanical, model path) keeps `run reflect` under a mechanical turn route; a TYPED `/status` drops the run slot; `/route clear` drops it; `/route off` drops it; PostModelSwitch source `command` drops it; a `Skill(feat)` with an agentId leaves `runMain`; `Skill(effort-low)` bridge then `turn.complete` → session defaults (not sticky); `feater` spawn (provider `{plugin:'engine',tier:'core'}`, no model) → `started.model` = sonnet full id and the first `turn.step` of that agentId at medium; `plan-challenger` with opus down → fable id; opus AND fable down → no write, one log line; explicit `model`/`effort` win; `fork: true` untouched; a project-source `agent.offer` record → no write; route answer text names the id with a floor in force; override `agents: { verifier: null }` drops the row (bottom `fs`/`env` mocked like `session.model`; if the kit refuses, the case is dropped and said so). - Disposition: honors BDR-115 (one writer per axis while on, full ids, config tables, closure state; rule 4 closed, rule 5 amended by A5/A6), BDR-107/108 (roles + levels preserved except analyzer), BDR-076/077 (judge rows = opus, off-state floor kept), LRN-205, LRN-206, LRN-207 (breaker feeds the tier resolution). ## Gate between A and B — live probe (user runs `/reload-plugins`) Evidence into the W2-B contract: (1) typed `/status` → `/route show` main `skill mechanical`; (2) a real `plugin-probe` or `feater` dispatch with verbose on → spawn log line (provider shape, model id written) + `step 0 agent …` effort line = spawn/first-step ordering fact the TODO asks for; (3) on a SONNET session (`/model sonnet`, then back): typed `/feat` → the self-check wording of the system prompt + the `route` answer id (gate witness fact); (4) probe (1) again with a background Explore alive (marker path vs fallback path in the verbose log). The verbose log line (`typed-marker` / `typed-fallback`, `spawn … → `, `step 0 agent …`) is the witness for (1), (2), (4), not `/route show` after the turn (turn-scoped routes are gone by then). Decision rules: (2) step 0 BEFORE the spawn bookkeeping → step 0 runs on the frontmatter `effort:` (kept, D2), recorded as a known limit; (4) `typed-marker` never seen → W2-B blocked, A4 re-planned (the `prompt.submit` text fact does not hold). Only then W2-B. ## W2-B — repo migration (own contract) - [ ] B0 mod: delete `EFFORT_SKILL`, `effortBridge`, the Skill-hook branch, their test; `agents/`+`skills/` rows unchanged. - [ ] B1 `lib/effort-shift.md` rewritten (~40 lines): route doctrine, the wiring points in `route` terms (D3), judgment built-ins carry `effort=`, sticky skill route + `/route clear`, no pairing rule, headless OK (hooks run under -p), levers = `ultrathink` / `/route effort=max` (builtin `/effort` is not a lever inside a run), last ROWED skill loaded wins (an unrowed one changes nothing). - [ ] B2 citers `Skill(effort-*)` → `route`: ship-feature 9, init-project 5, feat 4, bugfix 4, web-validate 3, seo 3, hotfix 3, geo 3, verify-secure-loop 3, harden 2, code-clean 2, audit-delta 2, onboard 1, client-handover-writer 1; EFFORT SHIFTS header lines in the 13 orchestrators + tour:33; `effort=` added to every `general-purpose` judgment dispatch (`model="opus"`: ship-feature, init-project, onboard ×7, tour; `model: "fable"` skill-runners in client-handover-writer); STOP/remedy texts (challenge-plan.md:62, verify-secure-loop.md:109) → the D4 levers; the 21 prose sites stating "sonnet by frontmatter pin" / "`model: opus`-pinned, session-independent" (challenge-plan:46, feat:146, bugfix:160, hotfix:138, code-clean:175, seo:463, 6 agents…) reworded "routed by the model-router row (frontmatter = off-state floor)". - [ ] B3 `lib/model-gate.md` → D4 (≤ 25 lines, keeps `model: "fable"` for skill-runners and the dispatch-tier table); delete `lib/model-check.sh`, `lib/tests/model-check.test.sh`. - [ ] B4 delete `lib/effort-pins.txt`, `lib/effort-pins.sh`, `lib/tests/effort-pins.test.sh`; remove the re-apply blocks + comments in `install-plugins.sh` (941-943, 1008 comment, 1131-1136) and `update-all.sh` (472, 567-572); `lib/tests/higgsfield.test.sh` :324 + the `before-pins` check (:330-356) dropped. - [ ] B5 `skills/effort-*` deleted (5 dirs; `profile`/catalog lists grepped). Tracked frontmatter KEPT and ALIGNED: `agents/analyzer.md` `effort: xhigh`; the vendored tracked `design-motion-principles` keeps `effort: high`; impeccable-* (gitignored) untouched; every other `model:`/`effort:` value already equals its row. - [ ] B6 census: `lib/tests/effort-routing.test.sh` rewritten — every tracked `effort:` equals its row's phase effort and every agent `model:` alias equals its row's tier head (drift lock, both directions, parsed from `register.ts`; haiku rows exempt from the effort direction: no effort on haiku); enumeration = tracked `SKILL.md` under `skills/` + `skills-external/` (`git ls-files`), no-row list = graphify, model-router, find-docs, impeccable, the agents interviewer/client-handover-writer/impeccable-*; no `Skill(effort-` anywhere in skills/agents/lib; D3 wiring markers (orchestrate in the 11, apply tail in the 5, escalate ×3 in verify-secure-loop + ship-feature, `effort=` on every `model="opus"` general-purpose dispatch); every tracked skill/agent (minus the no-row list) has a row in `register.ts` with the EXPECTED PHASE (per-name asserts, not mere presence); `model-routing.test.sh` :96-97 model-gate locks updated (the 18 `model:` locks, loops-light.test.sh:74,79 and plan-challenger.test.sh:20 stay valid: frontmatter kept); `CLAUDE.global.md` Design work lines 299-301 → "rows in the mod, an unrowed member changes nothing". - [ ] B7 doctrine-citers census; full `make test` once before merge. - [ ] STEP 6/7 of the B run: doc-syncer audit → README/USAGE/ARCHITECTURE/ CHANGELOG; BDR-115 amendment (wave 2 closed, D1-D4, A5/A6), LRN (typed-slash marker + sticky skill route), EVAL on the run; TODO W2 lines checked. ## Challenge ledger (r1 → r2) - robustness 1 BLOCKER (mod off → agents inherit the parent model) → D2: `model:` frontmatter kept as off-state floor; A3c row overrides it while on. - correctness 1/2, robustness 2 (marker unbound, ordering unproven) → A4 name-bound marker + pending slot + loops-size fallback; live probe gate. - correctness 3 (show text) → A9 asserts `skill reflect`. - correctness 4 (56) → fixed everywhere. - correctness 5, simplicity 1/2, robustness 13 (level deltas) → `write` high + `apply` phase; one delta left (analyzer), named. - correctness 6 (raise lasts one turn) → A6 sticky skill route. - correctness 7, robustness 5, simplicity 8 (gate witness) → D4 route answer. - correctness 8, robustness 11, simplicity 7 (point 5) → explicit `effort=`. - correctness 9, robustness 6 (levers) → D4 texts (r2's A7 engine-effort detector dropped again in r3). - correctness 10, robustness 9/10 (census) → B6 per-phase asserts, suites listed; no `effort="high"` call-site edits (write = high). - correctness 11 (unrowed skill clears) → A5. - correctness 12, robustness 12, simplicity 11 (counts, impeccable) → B5 from `git ls-files`; impeccable untouched. - correctness 13 (dead code) → A1 list; contract criterion 5 grep widened. - correctness 14 (override test needs fs) → A9 bottom mocks or drop, stated. - correctness 15, robustness 4/8 (provider shape, collisions) → A3b source from `agent.offer` + A8 null rows; tests use `engine/core`. - robustness 3 (haiku via fallback chain) → A3c tier-only + log. - robustness 7 (live edits mid-migration) → bridge stays until B0; probe gate before B. - robustness 14 (tour:33, install comment) → B2/B4. - robustness 15 (deploy on haiku) → `apply` rows. - simplicity 3 (tail) → kept as `route(phase="apply")` (phases only; the switch-on cost is the user's opt-in). [kept, reasoned] - simplicity 4 (W2-C) → folded into W2-B STEP 6/7. - simplicity 12 (two parsers) → contract oracle = one-shot values; B6 = durable per-phase census. [kept, reasoned] ## Confirmation ledger (r2 → r3) - conf 1 BLOCKER (route calls wipe the sticky slot) → A6 separate `runMain` slot, turn writers never touch it. - conf 2 (low/haiku leaks across turns) → only best-tier rows enter `runMain`. - conf 3 (bridge sticky in the A→B window) → bridge writes `turnMain` only. - conf 4 (engine-effort detector) → A7 dropped; builtin `/effort` named as a non-lever to the user; levers = `ultrathink`, `/route effort=max`. - conf 5 (route answer without id) → A3d always names the id, note appended; probe (3) on a sonnet session. - conf 6 (fs/cwd shadow unreliable) → A3b definition source from `agent.offer`; fs dropped. - conf 7 (judge → sonnet in-tier) → A3c rank ≥ tier head, else no write + log. - conf 8 (off-state quality drop, step-0 ordering) → D2 frontmatter kept as floor on both axes, census-locked; probe decision rule written. - conf 9 (pre-fill premise) → explicit = `e.model !== undefined`, no map. - conf 10 (origin denylist) → allowlist composer|sdk|bridge. - conf 11 (preload before trackLoop) → spawning counter; log names the path. - conf 12 (kit has no fs) → fs dropped. conf 13 → deduped log. - conf 14 (skill rows shadow) → accepted, named: skill rows affect main only; a foreign skill of the same name gets the row for its turn. - conf 15 (artefact drift) → contract clarification amended; status-reporter listed under mechanical. ## Confirmation 2 ledger (r3 → r4) - conf2 1 (helper skill drops the run) → A6: model-loaded non-best rows write turnMain only; typed rowed skills manage runMain. - conf2 2 (test :347 goes red) → A9 authorizes its rewrite. - conf2 3 (fallback ignores origin) → A4 `promptAllowed` flag on the fallback. - conf2 4/5/6 (runMain wiring) → A6 names both slots, mainRoute, clearLoop text, pushOrchestrate, `/model` + `/route off` + sub-agent cases in A9. - conf2 7 (probe witnesses) → gate: verbose lines + decision rules. - conf2 8 (grep) → contract criterion 5 narrowed. - conf2 9 (source tokens) → A3b names projectSettings|localSettings + limit. - conf2 10/11 (census) → B6 enumeration, haiku exemption, cites fixed.