Files
claude/.claude/tasks/plans/2026-10-09-model-router-w2-1546.md
T

346 lines
22 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# PLAN r4 — model-router wave 2 (migration), 2026-10-09
r1 → r2 after 3 lenses (simplicity CONCERNS(5), correctness CONCERNS(10),
robustness FATAL(8)); r2 → r3 after the confirmation pass (robustness
FATAL(8): 1 BLOCKER); r3 → r4 after a second confirmation (correctness
CONCERNS(3), wording-level). Every BLOCKER/MAJOR is closed by a NAMED change
(§ Challenge ledger). Mod = the router while on; the tracked frontmatter
(`model:` + `effort:` on agents, `effort:` on skills) STAYS as the
off-state floor, census-locked equal to the rows (r3).
Two sub-runs on `feature/model-router-w2`: W2-A (mod) then, after a user
`/reload-plugins` + live probe, W2-B (repo migration). Docs + registries =
W2-B's own STEP 6/7 (feat pipeline tail); W2-A runs no doc-sync (deviation,
stated: a mod-only diff has no public doc of its own until B lands).
## Decisions (pass B, user 2026-10-09) — amended by the challenge
- D1 the five `effort-*` skills are DELETED in W2-B. The mod's
`Skill(effort-*)` bridge is deleted in W2-B too (same commit as the citers,
robustness 7: edits are live on the symlinked tree). The typed `/effort-`
floor code goes in W2-A (A1). The `ultrathink` floor stays.
- D2 pins reworked, not deleted (r3): rows are PHASES by ROLE, no bare
levels; the mod routes agents on both axes while on (model written at
spawn from the tier, only UPWARD in rank; effort per step; explicit
Agent params win). The tracked frontmatter stays as the OFF-STATE FLOOR
(mod off / `enabled:false` / unloaded / a hook failing open → the engine
applies the frontmatter as today: never the parent model, never session
effort on an xhigh gate agent; r2 robustness 1 + confirmation 8). The
census locks every row equal to its frontmatter (tier head == `model:`
alias, phase effort == `effort:`), so there is one declared value, two
carriers. Deleted: the five shifters, `lib/effort-pins.*` (vendored
levels become rows; off-state = session level for them, as before
BDR-108), the witness script.
- D3 the orchestrators declare PHASES: dispatch span → `route(phase=
"orchestrate")`; own level high → `reflect`, xhigh → `plan`; bookkeeping
tail → `route(phase="apply")` (work/low); escalation → `escalate`.
Judgment dispatches of BUILT-INS (`general-purpose` with `model="opus"`
or `"fable"`) carry an explicit `effort=` param (a main route never
reaches a child; correctness 8). A best-tier skill row SURVIVES the end
of the turn in its own slot (A6 `runMain`): a run spans prose gates;
turn-scoped routes (`route` calls, prompt rules, the bridge) never
touch it.
- D4 `lib/model-gate.md` = the mod rule; the witness is the `route` tool's
own answer (correctness 7): self-check big → silent; small → call
`route(phase=<entry phase>)` and STOP unless the answer names a fable or
opus id; tool absent or "is off" → STOP with the remedy. `model-check.sh`
+ its test deleted. The route answer ALWAYS names the id the next main
step runs on (A3d). Relaunch levers in STOP texts: `ultrathink` in the
relaunch prompt (turn floor) or `/route effort=max` (sticky, `/route
clear` after). Builtin `/effort` is NOT a lever inside a run (rows and
routes rank above the engine effort; r2's engine-effort detector dropped
as unsafe, confirmation 4) — documented, named to the user.
## Phase table (A2a) — two rows added to `phases`
| phase | tier | effort | role |
| plan | best | xhigh | brainstorm, plan, architecture, audit verdict |
| reflect | best | high | diagnosis, review, contract, day-to-day orchestrator entry |
| orchestrate | best | medium | between dispatches |
| escalate | best | max | stuck, cap reached |
| judge | big | xhigh | dispatched challengers, analyzers, audits |
| implement | work | medium | code from a closed plan |
| write | work | **high** (was medium) | docs, commits, refactors, handover prose on sonnet at high (BDR-107 "high judgment on sonnet") |
| verify | work | xhigh | verifier, security-auditor |
| explore | work | medium | Explore |
| **apply** (new) | work | low | low appliers (BDR-107): small fixes, release mechanics, probes, validators; bookkeeping tail of the main loop |
| mechanical | cheap | low | listing, status, profile toggles |
## Row tables (A2b)
skills (56):
- plan: ship-feature init-project onboard tour audit-delta analyze
code-clean client-handover brainstorming writing-plans
requesting-code-review 21st-ui-review
- reflect: feat hotfix bugfix refactor web-validate harden seo geo
site-motion frontend-design emil-design-eng design-motion-principles
21st-ui-build scroll-world-storytelling build-threejs-scroll-worlds
scroll-scrubbed-visual-sequence scroll-scrubbed-word-reveal
scroll-progress-timeline subagent-driven-development writing-skills
deprecation-and-migration 21st-ai 21st-ui-explore
- implement: gitflow prune-memory pdf-translate ci-cd-and-automation
observability-and-instrumentation test-driven-development
- apply: commit-change release-candidate doc capitalize close reconcile
deploy (work tier: never haiku on main even with the switch on,
robustness 15)
- mechanical: status profile plugin-check skills-perso using-git-worktrees
21st-cli-use 21st-registry 21st-design-sync
agents (21 + 2 built-ins), model = frontmatter alias, unchanged everywhere;
`effort:` frontmatter updated where the row differs (analyzer → xhigh):
- implement (sonnet/medium): feater bugfixer code-cleaner scaffolder onboarder
- write (sonnet/high): commit-changer doc-syncer handover-doc-writer refactorer
- apply (sonnet/low): hotfixer release-executor plugin-probe validator-analyzer
- verify (sonnet/xhigh): verifier security-auditor
- judge (opus/xhigh): plan-challenger plugin-advisor seo-analyzer
geo-analyzer analyzer (the ONE level delta: high → xhigh, BDR-108 rung
"audit before validation"; named at the gate)
- mechanical (haiku/low): status-reporter
- built-ins: Explore explore, Plan judge
- no row: interviewer, client-handover-writer (inline-load), impeccable-*
(gitignored vendor output, untouched)
Deltas named at the gate: analyzer effort; best-tier raise now also reaches
refactor + the design/vendored reflect skills on a small session
(simplicity 6: the raise is the point of the mod).
## W2-A — mod (contract 2026-10-09-model-router-w2a-1546)
- [ ] A1 `register.ts`: delete the typed-floor code — `slashEffort`,
`guardedSlash`, `Source` member `'slash'`, `floorSource`'s `typed /`
branch, the State comments on the typed `/effort-<l>` floor. KEEP
`EFFORT_SKILL` + `effortBridge` (tool.call bridge, removed in W2-B)
and `turnFloor`/`higherFloor`/`floorWord` (ultrathink).
- [ ] A2 `DEFAULT_CONFIG`: phases per § Phase table (`write` high, `apply`
new); `skills` + `agents` per § Row tables; the "Built-ins only"
comment → "phase = role; one row per routed repo skill/agent (wave 2);
a project-level agent of the same name shadows its row (A3)".
- [ ] A3 spawn: (a) explicit model = `e.model !== undefined` at spawn (the
engine does not pre-fill the frontmatter; no tool_use_id map);
explicit effort as today (`explicitEffort`). (b) `spawnRoute`:
`frozen` → none; row lookup for every provider; a row is SKIPPED when
the agent's definition `source`, recorded per `subagentType` from the
`agent.offer` event, is `projectSettings` or `localSettings` (a
foreign repo's own `verifier.md`); `userSettings`, `built-in`,
`plugin` or no record → the row applies (fail-open on routing, as
W1). Known limit, stated in a comment: the record is keyed by name
only (an offer fired inside a sub-agent with another cwd overwrites it). (c) `spawnTarget` for a rowed spawn
without explicit model: resolve WITHIN the tier only, and write only
an alias ranked ≥ the tier head in `fallback` (agents move UP, never
below their frontmatter alias): `big` with opus down → fable; opus +
fable down → no write + one log line (deduped per agent+tier per
turn) "model-router: <agent> tier <t> down, frontmatter model kept".
(d) `routedText`/`mainAnswer` always name the id the next main step
runs on (`model <id> (<why>)`), the sticky/floor note APPENDED, never
substituted (the gate reads this answer).
- [ ] A4 typed slash: `typedSlash: string | null` = the first token of a
slash prompt at `prompt.submit` when `origin.kind` ∈ {composer, sdk,
bridge} (allowlist; floor/default rules stay composer-only), stored
only when it is a `cfg.skills` key; mid-turn → `pendingSlash`,
promoted at `endMainTurn`. `prompt.submit` also records
`st.promptAllowed` (origin in the allowlist) for the turn it opens
(pending slot for a mid-turn prompt, like the marker). `skill.prompt`
on main (`skillCalls === 0`): apply the row when `e.skill ===
typedSlash` (then null it) OR when `promptAllowed && loops.size === 0
&& spawning === 0` (`spawning` = a counter held from spawn-hook entry
to after `next`); otherwise `next(e)`. Verbose log names which path
fired (`typed-marker` / `typed-fallback`).
- [ ] A5 `onSkillLoad`: a skill WITHOUT a row leaves `turnMain` untouched
(find-docs / gstack / plugin skills mid-run no longer clear the
run's route; correctness 11); a rowed skill replaces it. Same rule
INSIDE a sub-agent (gated 2026-10-09, feater NEED-DECISION): an
unrowed skill leaves the agent loop's effort untouched; a rowed one
writes it (test: `feater` loop at medium, `Skill(find-docs)` with
that agentId → next step still medium).
- [ ] A6 run slot: new `runMain: Routed | null`. A rowed skill load ON
MAIN whose phase tier is `best` writes BOTH `turnMain` (source
`skill`, as today) and `runMain`; a non-best row loaded by the
model's `Skill` tool writes `turnMain` only (helper skills such as
`using-git-worktrees` inside SDD never drop the run); a USER-TYPED
rowed skill (marker path) replaces `runMain` with its row when best,
drops it when not. Precedence `userMain > turnMain > runMain > floor
> engine` (per axis, as today); `mainRoute` (statusline, spinner,
planStep) includes it; `endMainTurn` leaves it; `pushOrchestrate`
writes `turnMain` when empty exactly as today (derived orchestrate
overrides the run default for the dispatch span, like a `route
orchestrate` call); cleared by `/route clear` (text: "run slot
dropped" when one held), `/route off`, a user `/model` switch
(`PostModelSwitch` source command|picker|sdk); `route(clear=true)`
from the model clears `turnMain` only and its answer says "run
<phase> still holds". A `Skill` call inside a sub-agent never touches
it. `/route show` prints `run <phase>` when no turn route is in force.
- [ ] A7 (dropped in r3: engine-effort detector, confirmation 4).
- [ ] A8 `mergeTable` accepts `null` in the override's `skills`/`agents`
to drop a default row (robustness 4).
- [ ] A9 `register.test.ts`: delete the typed `/effort-*` floor tests;
keep the bridge test; add — typed `/feat` (prompt.submit `/feat x`,
origin composer, then skill.prompt feat) → `main: skill reflect`,
`effort high`, `[tier best]`; origin `channel`: not armed AND the
idle fallback refused (main untouched); skill.prompt `feat` with a
live sub-agent and no marker → main untouched; same with no live loop
and an allowed origin → routed; a mid-turn `/status` → pending,
applied after turn.complete; `Skill(find-docs)` (no row) after
`route plan` keeps plan — the existing test 'floor: survives a skill
load' (register.test.ts:341-348) is REWRITTEN to assert orchestrate
KEPT (A5 behavior, authorized here); run slot: `Skill(feat)`, then
`route orchestrate`, then `turn.complete` → `/route show` back to
`run reflect`; `Skill(using-git-worktrees)` (mechanical, model path)
keeps `run reflect` under a mechanical turn route; a TYPED `/status`
drops the run slot; `/route clear` drops it; `/route off` drops it;
PostModelSwitch source `command` drops it; a `Skill(feat)` with an
agentId leaves `runMain`; `Skill(effort-low)` bridge then
`turn.complete` → session defaults (not sticky); `feater` spawn (provider
`{plugin:'engine',tier:'core'}`, no model) → `started.model` = sonnet
full id and the first `turn.step` of that agentId at medium;
`plan-challenger` with opus down → fable id; opus AND fable down →
no write, one log line; explicit `model`/`effort` win; `fork: true`
untouched; a project-source `agent.offer` record → no write; route
answer text names the id with a floor in force; override `agents: {
verifier: null }` drops the row (bottom `fs`/`env` mocked like
`session.model`; if the kit refuses, the case is dropped and said so).
- Disposition: honors BDR-115 (one writer per axis while on, full ids,
config tables, closure state; rule 4 closed, rule 5 amended by A5/A6),
BDR-107/108 (roles + levels preserved except analyzer), BDR-076/077
(judge rows = opus, off-state floor kept), LRN-205, LRN-206, LRN-207
(breaker feeds the tier resolution).
## Gate between A and B — live probe (user runs `/reload-plugins`)
Evidence into the W2-B contract: (1) typed `/status` → `/route show` main
`skill mechanical`; (2) a real `plugin-probe` or `feater` dispatch with
verbose on → spawn log line (provider shape, model id written) + `step 0
agent …` effort line = spawn/first-step ordering fact the TODO asks for;
(3) on a SONNET session (`/model sonnet`, then back): typed `/feat` → the
self-check wording of the system prompt + the `route` answer id (gate
witness fact); (4) probe (1) again with a background Explore alive (marker
path vs fallback path in the verbose log). The verbose log line (`typed-marker` /
`typed-fallback`, `spawn … → <id>`, `step 0 agent …`) is the witness for
(1), (2), (4), not `/route show` after the turn (turn-scoped routes are
gone by then). Decision rules: (2) step 0 BEFORE the spawn bookkeeping →
step 0 runs on the frontmatter `effort:` (kept, D2), recorded as a known
limit; (4) `typed-marker` never seen → W2-B blocked, A4 re-planned (the
`prompt.submit` text fact does not hold). Only then W2-B.
## W2-B — repo migration (own contract)
- [ ] B0 mod: delete `EFFORT_SKILL`, `effortBridge`, the Skill-hook branch,
their test; `agents/`+`skills/` rows unchanged.
- [ ] B1 `lib/effort-shift.md` rewritten (~40 lines): route doctrine, the
wiring points in `route` terms (D3), judgment built-ins carry
`effort=`, sticky skill route + `/route clear`, no pairing rule,
headless OK (hooks run under -p), levers = `ultrathink` / `/route
effort=max` (builtin `/effort` is not a lever inside a run), last ROWED
skill loaded wins
(an unrowed one changes nothing).
- [ ] B2 citers `Skill(effort-*)` → `route`: ship-feature 9, init-project
5, feat 4, bugfix 4, web-validate 3, seo 3, hotfix 3, geo 3,
verify-secure-loop 3, harden 2, code-clean 2, audit-delta 2, onboard
1, client-handover-writer 1; EFFORT SHIFTS header lines in the 13
orchestrators + tour:33; `effort=` added to every `general-purpose`
judgment dispatch (`model="opus"`: ship-feature, init-project,
onboard ×7, tour; `model: "fable"` skill-runners in
client-handover-writer); STOP/remedy texts (challenge-plan.md:62,
verify-secure-loop.md:109) → the D4 levers; the 21 prose sites
stating "sonnet by frontmatter pin" / "`model: opus`-pinned,
session-independent" (challenge-plan:46, feat:146, bugfix:160,
hotfix:138, code-clean:175, seo:463, 6 agents…) reworded "routed by
the model-router row (frontmatter = off-state floor)".
- [ ] B3 `lib/model-gate.md` → D4 (≤ 25 lines, keeps `model: "fable"` for
skill-runners and the dispatch-tier table); delete
`lib/model-check.sh`, `lib/tests/model-check.test.sh`.
- [ ] B4 delete `lib/effort-pins.txt`, `lib/effort-pins.sh`,
`lib/tests/effort-pins.test.sh`; remove the re-apply blocks + comments
in `install-plugins.sh` (941-943, 1008 comment, 1131-1136) and
`update-all.sh` (472, 567-572); `lib/tests/higgsfield.test.sh`
:324 + the `before-pins` check (:330-356) dropped.
- [ ] B5 `skills/effort-*` deleted (5 dirs; `profile`/catalog lists
grepped). Tracked frontmatter KEPT and ALIGNED: `agents/analyzer.md`
`effort: xhigh`; the vendored tracked `design-motion-principles`
keeps `effort: high`; impeccable-* (gitignored) untouched; every
other `model:`/`effort:` value already equals its row.
- [ ] B6 census: `lib/tests/effort-routing.test.sh` rewritten — every
tracked `effort:` equals its row's phase effort and every agent
`model:` alias equals its row's tier head (drift lock, both
directions, parsed from `register.ts`; haiku rows exempt from the
effort direction: no effort on haiku); enumeration = tracked
`SKILL.md` under `skills/` + `skills-external/` (`git ls-files`),
no-row list = graphify, model-router, find-docs, impeccable, the
agents interviewer/client-handover-writer/impeccable-*; no
`Skill(effort-` anywhere in
skills/agents/lib; D3 wiring markers (orchestrate in the 11, apply
tail in the 5, escalate ×3 in verify-secure-loop + ship-feature,
`effort=` on every `model="opus"` general-purpose dispatch); every
tracked skill/agent (minus the no-row list) has a row in
`register.ts` with the EXPECTED PHASE (per-name asserts, not mere
presence); `model-routing.test.sh` :96-97 model-gate locks
updated (the 18 `model:` locks, loops-light.test.sh:74,79 and
plan-challenger.test.sh:20 stay valid: frontmatter kept); `CLAUDE.global.md` Design
work lines 299-301 → "rows in the mod, an unrowed member changes
nothing".
- [ ] B7 doctrine-citers census; full `make test` once before merge.
- [ ] STEP 6/7 of the B run: doc-syncer audit → README/USAGE/ARCHITECTURE/
CHANGELOG; BDR-115 amendment (wave 2 closed, D1-D4, A5/A6), LRN
(typed-slash marker + sticky skill route), EVAL on the run; TODO W2
lines checked.
## Challenge ledger (r1 → r2)
- robustness 1 BLOCKER (mod off → agents inherit the parent model) →
D2: `model:` frontmatter kept as off-state floor; A3c row overrides it
while on.
- correctness 1/2, robustness 2 (marker unbound, ordering unproven) → A4
name-bound marker + pending slot + loops-size fallback; live probe gate.
- correctness 3 (show text) → A9 asserts `skill reflect`.
- correctness 4 (56) → fixed everywhere.
- correctness 5, simplicity 1/2, robustness 13 (level deltas) → `write`
high + `apply` phase; one delta left (analyzer), named.
- correctness 6 (raise lasts one turn) → A6 sticky skill route.
- correctness 7, robustness 5, simplicity 8 (gate witness) → D4 route answer.
- correctness 8, robustness 11, simplicity 7 (point 5) → explicit `effort=`.
- correctness 9, robustness 6 (levers) → D4 texts (r2's A7 engine-effort
detector dropped again in r3).
- correctness 10, robustness 9/10 (census) → B6 per-phase asserts, suites
listed; no `effort="high"` call-site edits (write = high).
- correctness 11 (unrowed skill clears) → A5.
- correctness 12, robustness 12, simplicity 11 (counts, impeccable) → B5
from `git ls-files`; impeccable untouched.
- correctness 13 (dead code) → A1 list; contract criterion 5 grep widened.
- correctness 14 (override test needs fs) → A9 bottom mocks or drop, stated.
- correctness 15, robustness 4/8 (provider shape, collisions) → A3b
source from `agent.offer` + A8 null rows; tests use `engine/core`.
- robustness 3 (haiku via fallback chain) → A3c tier-only + log.
- robustness 7 (live edits mid-migration) → bridge stays until B0; probe
gate before B.
- robustness 14 (tour:33, install comment) → B2/B4.
- robustness 15 (deploy on haiku) → `apply` rows.
- simplicity 3 (tail) → kept as `route(phase="apply")` (phases only; the
switch-on cost is the user's opt-in). [kept, reasoned]
- simplicity 4 (W2-C) → folded into W2-B STEP 6/7.
- simplicity 12 (two parsers) → contract oracle = one-shot values; B6 =
durable per-phase census. [kept, reasoned]
## Confirmation ledger (r2 → r3)
- conf 1 BLOCKER (route calls wipe the sticky slot) → A6 separate `runMain`
slot, turn writers never touch it.
- conf 2 (low/haiku leaks across turns) → only best-tier rows enter `runMain`.
- conf 3 (bridge sticky in the A→B window) → bridge writes `turnMain` only.
- conf 4 (engine-effort detector) → A7 dropped; builtin `/effort` named as
a non-lever to the user; levers = `ultrathink`, `/route effort=max`.
- conf 5 (route answer without id) → A3d always names the id, note appended;
probe (3) on a sonnet session.
- conf 6 (fs/cwd shadow unreliable) → A3b definition source from
`agent.offer`; fs dropped.
- conf 7 (judge → sonnet in-tier) → A3c rank ≥ tier head, else no write + log.
- conf 8 (off-state quality drop, step-0 ordering) → D2 frontmatter kept
as floor on both axes, census-locked; probe decision rule written.
- conf 9 (pre-fill premise) → explicit = `e.model !== undefined`, no map.
- conf 10 (origin denylist) → allowlist composer|sdk|bridge.
- conf 11 (preload before trackLoop) → spawning counter; log names the path.
- conf 12 (kit has no fs) → fs dropped. conf 13 → deduped log.
- conf 14 (skill rows shadow) → accepted, named: skill rows affect main
only; a foreign skill of the same name gets the row for its turn.
- conf 15 (artefact drift) → contract clarification amended; status-reporter
listed under mechanical.
## Confirmation 2 ledger (r3 → r4)
- conf2 1 (helper skill drops the run) → A6: model-loaded non-best rows
write turnMain only; typed rowed skills manage runMain.
- conf2 2 (test :347 goes red) → A9 authorizes its rewrite.
- conf2 3 (fallback ignores origin) → A4 `promptAllowed` flag on the fallback.
- conf2 4/5/6 (runMain wiring) → A6 names both slots, mainRoute, clearLoop
text, pushOrchestrate, `/model` + `/route off` + sub-agent cases in A9.
- conf2 7 (probe witnesses) → gate: verbose lines + decision rules.
- conf2 8 (grep) → contract criterion 5 narrowed.
- conf2 9 (source tokens) → A3b names projectSettings|localSettings + limit.
- conf2 10/11 (census) → B6 enumeration, haiku exemption, cites fixed.