chore(tasks): model-router w1a contract + plan r3

This commit is contained in:
bchanot
2026-10-08 16:26:57 +02:00
parent b721c94dcb
commit e8ca713d9e
2 changed files with 394 additions and 0 deletions
@@ -0,0 +1,347 @@
# PLAN — model-router wave 1-A: the mod (dispatch-ready) — REVISED r3
r3 changes (confirmation pass, 4 MAJOR + minors): explicit Agent params are
frozen on the Loop and never overridden by in-agent `route`/`Skill(effort-*)`;
`pendingPrompt` honours `e.wait`; tests assert on the `main:` line of `show`
only, with full typed inputs; haiku gets `effort: undefined`; `userMain`
beats every turn route including a typed `/effort-<l>` (the text says so);
`skill.prompt` acts only when no Skill call is in flight; route tool handles
`clear` first and answers "off" in place; `turn.step` catch is a generator;
windows keyed by resolved id; validate user entries BEFORE merging; truthful
answers; unpinned-skill reset accepted and documented.
Contract: .claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md
Wave plan + harness facts: .claude/tasks/plans/2026-10-08-model-router-mod.md
Base: the spike `~/.claude/dev-mods/385f7190-70f5-4bdd-b0d8-e4566cd412fd/model-router/hooks/register.ts`
(read it first; keep its proven hook shapes; drop its spike levers `via`,
`stepModel`, `agentsDefault`, the tool's `model`/`scope`/`mainModelSwitch`
params and the hardcoded PHASES).
API reference: `<spike>/.claude-plugin/types/claude-code/index.d.ts` (grep the
event or noun; `declare module 'claude-code/testing'` for the test kit) and
`<spike>/.claude-plugin/types/claude-code-tools/index.d.ts` (`Skill: {`,
`Agent: {`, and the Skill RESULT schema near line 5139).
r2 changes (challenge round, 3 lenses): agents table = built-ins only; one
writer per agent model (spawn), no per-step model rewrite unless an in-agent
`route` call changed it and the engine did not fall back; every write goes
to the CALLING loop; `agentsNext` / `scope` / `/route agents` / tool `model`
dropped; Skill bridge answers in the tool's output shape; unconditional
skill-load reset (prompt route kept); config validated at load, refs
resolved once; state in the `register` closure, cloned defaults; `/route off`
kill switch; queued `ultrathink` promoted to its own turn; AC1/AC5 amended.
## Files
- [ ] mods/model-router/.claude-plugin/plugin.json — `{ "name": "model-router", "version": "0.1.0", "description": "<one line>", "author": { "name": "bchanot" } }`
- [ ] mods/model-router/hooks/hooks.json — `{ "modules": ["./register.ts"] }`
- [ ] mods/model-router/hooks/register.ts — the hooks module (below)
- [ ] mods/model-router/hooks/register.test.ts — `claude plugin test` suite (below)
The engine lays `./tsconfig.json` and `.claude-plugin/types/` beside a loaded
mod; both are ignored by AC1 and gitignored in wave 1-B. Never commit them.
## Config (one shape, defaults in code, optional override on disk)
```ts
type Level = 'low' | 'medium' | 'high' | 'xhigh' | 'max'
type Route = { model?: string; effort?: Level } // model = alias OR full id
type Config = {
models: Record<string, string> // alias → full id
windows: Record<string, number> // alias → context window (tokens)
phases: Record<string, Route>
agents: Record<string, string> // built-in subagentType → phase name
skills: Record<string, string> // skill name → phase name
prompt: { pattern: string; phase: string }[] // regex source, flag i
mainModelSwitch: boolean; verbose: boolean; spinner: boolean
}
```
DEFAULT_CONFIG values:
- models: haiku→`claude-haiku-4-5-20251001`, sonnet→`claude-sonnet-5-5`,
opus→`claude-opus-5-5`, fable→`claude-fable-5-1`.
- windows: `claude-haiku-4-5-20251001`→200000 (keyed by FULL id; others
unknown: absent = no check).
- phases: plan {effort xhigh}, reflect {effort high}, orchestrate {effort medium},
escalate {effort max}, judge {opus, xhigh}, implement {sonnet, medium},
write {sonnet, medium}, verify {sonnet, xhigh}, explore {sonnet, medium},
mechanical {haiku, low}. A phase without `model` keeps the loop's model.
- agents: Explore→explore, Plan→judge. NOTHING else in wave 1: every repo
agent keeps its frontmatter pin (the engine applies it); the pins move into
this table in wave 2, in the same change that deletes the frontmatter.
- skills: {} (wave 2 fills it).
- prompt: [{ pattern: '\\bultrathink\\b', phase: 'escalate' }].
- mainModelSwitch false, verbose false, spinner true. (Pass B, user
2026-10-08: verbose default off but ON for now through the override file,
spinner on, user `/route` sticky, Explore → sonnet/medium.)
`loadConfig($, log)` → `Config` (never throws):
1. `home = await $.env.get('HOME')`; path `${home}/.claude/model-router.json`;
`$.fs.exists` then `$.fs.read`; `JSON.parse`. Any failure → the defaults.
2. `mergeConfig(D, u, log)`: FIXED merge, no recursion, VALIDATE EACH USER
ENTRY BEFORE IT REPLACES A DEFAULT (an invalid user `models.sonnet` is
dropped and the default kept, so the phases on `sonnet` stay valid):
for each table (`models`, `windows`, `phases`, `agents`, `skills`) take
`u.<table>` only when it is a plain object, then per key: valid → over
the default, invalid → `log(...)` and keep the default. Scalars
(`mainModelSwitch`, `verbose`, `spinner`) taken only when boolean.
`prompt` taken only when an array; each rule validated or dropped.
Type guards on `unknown`, no `any`.
3. Validity rules (`log(...)` ALWAYS, not only verbose: a config error must
be seen once):
- models value: string matching `/^claude-[a-z0-9.-]+$/`;
- windows: key a full id (same regex), value a positive integer;
- phase: plain object; `effort` absent or in LEVELS; `model` absent, a
`models` key or a full id; at least one of the two;
- agents / skills value: a phase name (checked after phases merged);
- prompt rule: `{ pattern: string, phase: <phase name> }` whose pattern
compiles (`new RegExp(p, 'i')` in try/catch).
Lookups use `Object.hasOwn`, never bare indexing on user keys.
4. Returns `structuredClone`d data: the defaults constant is never handed
out by reference.
`compileRules(cfg)` → `{ re: RegExp; phase: string }[]` once per load.
`resolveModel(cfg, name)`: `Object.hasOwn(cfg.models, name) ? cfg.models[name] : name`.
`isModelName(cfg, v)`: a `models` key or the full-id regex. `isLevel(v)`.
## State — ONE object built inside `register`, passed to every helper
```ts
type Source = 'user' | 'model' | 'skill' | 'prompt' | 'slash'
type Routed = { phase: string; route: Route; source: Source }
type Loop = {
effort?: Level; model?: string // routed by the table or an in-agent call
spawnModel: string; frozen: boolean // engine's model at spawn; fork/workflow
explicitModel: boolean; explicitEffort: boolean // Agent params given → axis frozen
}
type State = {
cfg: Config; rules: Rule[]
userMain: Routed | null // /route by the user; sticky until /route clear
turnMain: Routed | null // tool / skill / prompt / slash; dropped at turn end
pendingPrompt: Routed | null // prompt rule typed mid-turn, promoted next turn
loops: Map<string, Loop> // agentId → that loop's routing
explicitEffort: Map<string, Level> // Agent tool_use_id → explicit effort param
skillCalls: number // Skill tool calls in flight (hook 4 ± around next)
off: boolean // /route off: every hook passes through
lastMain: string // "model/effort" of the last main step (spinner)
}
```
`newState(cfg)` builds it; `register` calls it once; `session.start` reloads
`cfg` + `rules` into it; `session.end` rebuilds it (`/clear` fires no
`session.start`, so sticky routes must not survive a clear).
Effective main route `mainRoute(st)`: `st.userMain ?? st.turnMain`. ONE
order, transitive: user `/route` (sticky) > the latest turn route (tool,
skill, slash, prompt all share `turnMain`; last writer wins) > session. A
typed `/effort-<l>` while a sticky route is in force does not apply; its
text says so (hook 5).
Loop lookups: `loopOf(st, e.agentId)`; a missing entry is created on first
write as `{ spawnModel: '', frozen: false, explicitModel: false,
explicitEffort: false }`. In-agent writes (hooks 3 and 4) never set an axis
whose `explicit*` flag is true: an explicit Agent param wins for the whole
run.
## Hooks
Rule for failures: hooks that only observe or rewrite carry
`.catch(($, e, next) => next(e))`; `turn.step` streams, so its catch is the
generator form `async function* ($, e, next) { return yield* next(e) }` (a
plain function there is a type error). The four hooks that ANSWER without `next`
(command.run, the route tool, the Skill `effort-*` bridge, `skill.prompt`)
carry a `.catch` that answers in place: `{ text: 'route failed (<kind>)' }`,
`{ result: 'route failed (<kind>); nothing routed' }`, and for the two skill
hooks `next(e)` (the skill then loads normally — a safe fallback). State is
mutated only AFTER the input validated.
1. `session.start`: `st.cfg = await loadConfig(...)`, `st.rules = compileRules`;
`registerTool($, st)` (`$.tool.register({ name: 'route', description,
inputSchema })`: properties `phase` (enum = Object.keys(st.cfg.phases)),
`effort` (enum LEVELS), `clear` (boolean); no `required`). The description
tells the model: declare the phase before a span changes nature; acts on
the calling loop only; no model choice here. `$.command.register({ name:
'route', description, argumentHint: '[show|clear|off|on|reload|<phase>|
model=<alias|id> effort=<level>|switch on|off|verbose on|off]', immediate:
true })` in try/catch (log on failure, keep going). `$.ui.status(statusLine(st))`.
2. `command.run {command:'route'}` → `{ text: handleCommand($, st, e.args) }`:
`show`/empty → `show(st)`; `clear` → userMain = turnMain = pendingPrompt =
null; `off` / `on` → st.off; `reload` → loadConfig + compileRules +
`registerTool` again (the phase enum follows the config) + 'config
reloaded' + show; `switch on|off` → cfg.mainModelSwitch; `verbose on|off`;
otherwise `parseRoute(st.cfg, args)`: a phase name, or tokens `model=<m>`
/ `effort=<l>` / bare alias / bare level, each validated by `isModelName`
/ `isLevel` → `st.userMain = { phase, route, source: 'user' }`; any
unknown token → error text listing the phases and the levels. Never
calls `next`.
3. `tool.call {tool:'mcp__model-router__route'}` → `handleRouteTool`, in
this order: (i) `st.off` → `{ result: 'model-router is off (/route on to
resume); nothing routed' }`; (ii) `clear` → main: `turnMain = null`; agent:
unset the loop's `effort`/`model` → `{ result: 'route cleared for <loop>' }`;
(iii) validate: `phase` given and not a `phases` key → `{ deny: 'unknown
phase "<p>"; phases: …' }` (even with a valid `effort`); `effort` given and
not a level → deny naming the levels; neither given → deny. (iv) route =
`{ ...phases[phase], ...(effort ? { effort } : {}) }` (an explicit effort
overrides the phase's). Target = the CALLING loop: main → `turnMain = {
phase: phase ?? 'effort-' + effort, route, source: 'model' }`; agent →
`loop.effort = route.effort` unless `loop.explicitEffort`; `loop.model =
route.model` unless `loop.explicitModel` (applied at step only under the
fallback guard, never on a frozen loop). (v) Answer on the EFFECTIVE
outcome: main with a sticky `userMain` → `'recorded <phase> for this turn,
but a sticky /route <userMain.phase> is in force; it wins until /route
clear'`; otherwise `'routed <main|this agent> to <phase>: effort <level|
unchanged>, model <resolved id|unchanged>'`, and when a model is part of
the route on main while `mainModelSwitch` is off, say `model unchanged
(switch off)`. Verbose → log.
4. `tool.call {tool:'Skill'}`:
a. `e.skill` matches `/^effort-(low|medium|high|xhigh|max)$/` → the
CALLING loop: main → `turnMain = { phase: e.skill, route: { ...st.turnMain?.route, effort }, source: 'skill' }`
(effort merged over the current turn route, last loaded wins); agent →
`loopOf(...).effort = level` unless `loop.explicitEffort` (entry created
if missing: an untabled agent's shift must still land). Answer WITHOUT
`next`, in the Skill tool's output shape (claude-code-tools ~5139; a
string result is refused and the skill would load):
`{ result: { success: true, commandName: e.skill, status: 'inline' },
context: [<line>] }` where `<line>` states the effective outcome:
`'model-router: effort → <l> for this loop from the next request on; the
effort-<l> skill text was not loaded.'`, or when main has a sticky
`userMain`: `'model-router: effort-<l> recorded, but a sticky /route
<phase> is in force and wins until /route clear.'`, or when the agent
axis is explicit: `'model-router: this agent was dispatched with an
explicit effort; the shift does not apply.'`
b. Any other skill: `st.skillCalls += 1` before `next(e)`, `-= 1` after
(try/finally). Main → `turnMain = null` when its source is 'model',
'skill' or 'slash' (a 'prompt' route such as `ultrathink` stays unless
the skill has a table entry); `Object.hasOwn(cfg.skills, e.skill)` →
`turnMain = { phase, route, source: 'skill' }`. Agent → unset the loop's
non-explicit `effort`/`model`; table entry → `loop.effort = route.effort`
unless explicit (never model). Accepted change vs the legacy shifters:
loading an UNPINNED skill after a shift returns main to the harness
level (orchestrators already re-assert after a nested skill,
lib/effort-shift.md § Re-assert). Then `return next(e)`.
5. `skill.prompt {skill: /^effort-/}`: `st.skillCalls > 0` (reached through
the bridge's fallback or a Skill call) → `next(e)`. Otherwise it is a
user-typed `/effort-<l>` (or a preload, unsupported: treated the same):
level parse; `turnMain = { phase: e.skill, route: { effort }, source:
'slash' }`; return `{ text: <line> + '\n' + e.text }` (prepend, never
replace: args ride in the text) where `<line>` is `'Effort shifted to <l>
by model-router for this turn.'` or, with a sticky `userMain`, `'Effort
<l> recorded; the sticky /route <phase> wins until /route clear.'`.
Unknown suffix → `next(e)`.
6. `tool.call {tool:'Agent'}`: `isLevel(e.effort) && typeof e.tool_use_id ===
'string'` → `st.explicitEffort.set(e.tool_use_id, e.effort)`; always
`return next(e)` unchanged (params are never rewritten).
7. `agent.spawn`: `frozen = e.fork || e.workflow !== undefined`. Route:
`!frozen && e.provider.plugin === 'engine' && Object.hasOwn(cfg.agents, e.subagentType)`
→ `cfg.phases[cfg.agents[e.subagentType]]`, else none. Model rewrite ONLY
when route?.model is set AND `e.model === undefined` (an explicit param
wins): `next({ ...e, model: resolveModel(cfg, route.model) })`, else
`next(e)`. On a non-deny result with `agentId`: `loops.set(agentId, {
spawnModel: result.model, frozen, explicitModel: e.model !== undefined,
explicitEffort: given !== undefined, effort: given ? undefined : route?.effort })`
where `given = explicitEffort.get(e.tool_use_id)` (then deleted). `model`
is NOT stored at spawn: the engine already runs the agent on it. Verbose
log `spawn <type>: <e.model ?? '-'> → <result.model>`. Known limit,
documented in a comment: `provider.plugin === 'engine'` is the best
available test for a built-in at spawn; a user agent named `Explore` in a
foreign project would also match (wave 1 impact: sonnet/medium on it).
8. `turn.step` (async generator). `st.off` → log when verbose, `yield* next(e)`.
Agent loop (`e.agentId`): `loop = loops.get(...)`; `effort = loop?.effort ??
e.effort`; `model = loop?.model && !loop.frozen && e.model === loop.spawnModel
? resolveModel(cfg, loop.model) : e.model` (an engine fallback — `e.model`
differs from the spawn model — is never fought). Main: `set = mainRoute(st)`;
`effort = set?.route.effort ?? e.effort`; `model = set?.route.model &&
cfg.mainModelSwitch && windowOk ? resolved : e.model`, where `resolved =
resolveModel(cfg, set.route.model)` and `windowOk` = no
`cfg.windows[resolved]` entry (windows are keyed by FULL id; the defaults
key haiku's full id) or `(await $.session.usage()).context.tokens` is a
number below it; an absent `tokens` or a failed `usage()` → no switch,
logged once per turn. If the model actually sent starts with
`claude-haiku`, send `effort: undefined` (omit it entirely; haiku takes
none and a hook-set effort on it is unproven). Main → `lastMain =
'<model without claude->/<effort>'`, `$.ui.status(statusLine(st))`. Verbose →
log before (`step <i> <loop>: <from> → <to>`) and after (`answered by
<usage.model>`). `const r = yield* next(changed ? { ...e, model, effort } : e); return r`.
9. `prompt.submit`: `e.origin.kind !== 'composer'` → `next(e)`. First rule in
`st.rules` whose `re.test(e.text)` → `routed = { phase, route, source: 'prompt' }`;
`e.turnId !== undefined && e.wait` (typed mid-turn and asked to wait: it
belongs to the NEXT turn) → `st.pendingPrompt = routed`; otherwise
(idle, or delivered INTO the running turn) → `st.turnMain = routed`.
`return next(e)`.
10. `turn.complete`: `e.agentId` → `loops.delete(e.agentId)`. Main →
`turnMain = pendingPrompt; pendingPrompt = null; explicitEffort.clear();
lastMain = ''`; `$.ui.status(statusLine(st))`. `return next(e)`.
11. `ui.render {component:'Spinner'}`: `cfg.spinner && lastMain` →
`next({ ...e, props: { ...e.props, suffix: ' · ' + lastMain + '…' } })` else `next(e)`.
12. `session.end`: `Object.assign(st, newState(st.cfg))` (keeps the loaded
config, drops every route and map). `return next(e)`.
`statusLine(st)`: `'route: ' + (st.off ? 'off' : describe(mainRoute(st)) )`
where `describe` = `'<source> <phase>'` or `'session defaults'`, plus
`' · switch on'` when `cfg.mainModelSwitch`.
`show(st)`: main (effective, with its source), off/on, switch, verbose,
spinner, live loops count, phases as `name=<resolved id|session>/<effort>`
(resolved ids printed, so a wrong `models` entry is visible), config source
line (`defaults` or the override path).
## Tests (register.test.ts, `import { test, expect } from 'claude-code/testing'`)
Read the kit's declarations first (`declare module 'claude-code/testing'`):
`test(name, async ($, on) => …)`; events are fired as calls on `$` with the
event's FULL input (the kit's `$` is `EngineCall<E> = (e: Args<E>)`, and
AC2 type-checks the test file): `$.command.run({ command: 'route', args:
'show', origin: { kind: 'composer' }, presentation: <a valid value from the
types> })`, `$.prompt.submit({ text: 'ultrathink please', wait: false,
origin: { kind: 'composer' } })`, `$.agent.spawn({ tool_use_id: 't1',
prompt: 'x', description: 'd', subagentType: 'Explore', provider: { plugin:
'engine', tier: 'core' }, parentModel: 'claude-fable-5-1', background: false,
fork: false })`, `$.tool.call({ tool: 'Skill', skill: 'effort-low' })`. Read
each input type and fill every required field; never relax a test to dodge
a type. Establish from the kit whether `session.start` fires at load; if
not, fire `$.session.start(...)` first in every test. An event whose hook
calls `next` needs a BOTTOM hook registered by the test through its `on`
(the kit's bottom throws otherwise), e.g. `on('prompt.submit', ($, e) => ({
text: e.text }))`, `on('agent.spawn', ($, e) => ({ model: e.model, agentId:
'a1' }))`. A helper `mainLine(text)` returns the `main:` line of `show`;
EVERY assertion on a route reads that line only (the phases listing always
contains every id and level, so matching the whole text proves nothing).
Tests (one per contract item 3a-3g):
- 3a `Skill(effort-low)` via `$.tool.call`: resolves with `result.success ===
true` and `result.commandName === 'effort-low'`; a bottom `on('tool.call',
{ tool: 'Skill' })` registered by the test is NOT reached (flag); `mainLine`
contains `skill effort-low` and `low`.
- 3b route tool `{ phase: 'orchestrate' }` → `mainLine` contains `model
orchestrate` and `medium`.
- 3c `/route clear` → `mainLine` contains `session defaults`.
- 3d `/route bogus` → text contains `unknown` and every phase name.
- 3e `$.prompt.submit` with `ultrathink`, `wait: false`, no `turnId` →
`mainLine` contains `prompt escalate`.
- 3f `/route model=sonnet` → `mainLine` contains `claude-sonnet-5-5`; `/route
model=claude-x-9` → `mainLine` contains `claude-x-9`; `/route model=sonet` →
text contains `unknown`.
- 3g spawn path: `$.agent.spawn(...)` for `Explore` without `model`, bottom
hook captures `e.model === 'claude-sonnet-5-5'`; with `model: 'opus'` given
→ captured `e.model === 'opus'`.
No fs, network or process in tests: the defaults path is the one exercised.
## Edge cases
- `$.command.register` throws when `/route` is taken → log, keep the tool.
- `loadConfig` never throws out of `session.start`; invalid entries dropped with a log.
- `/clear` → `session.end` rebuilds the state; `/route reload` re-registers the tool.
- Remote agents raise no `turn.complete`; denied Agent calls never spawn: both
maps are bounded by `explicitEffort.clear()` at main turn end and `loops`
entries only for started agents (a leak of a few entries per turn is accepted).
- `loops.delete` at an agent's `turn.complete`: a resumed agent (SendMessage,
woken teammate) runs its later turns at the engine's effort. Accepted in
wave 1 (built-ins only); revisit with the pins in wave 2.
- Spawn-vs-first-step race: the loop entry is set after `next(e)` resolves,
so step 0 of a tabled built-in may run at the engine's effort. Accepted in
wave 1 (Explore/Plan only); wave 2 verifies the ordering before pins move.
- Unpinned agents (general-purpose, interviewer, client-handover-writer) no
longer inherit a shifted level: the bridge does not move the harness
level. Accepted: BDR-077 already requires explicit call-site params for
built-ins; the two inline-load agents run on main's own route.
- The word `any` must not appear as a TypeScript type in register.ts
(AC5 greps `: any`, `<any>`, `as any`).
## Disposition (STEP 0.6)
- honors BDR-066/076/077 (tiers): wave 1 touches only the two built-ins that
carry no pin (Explore sonnet/medium, Plan opus/xhigh); every repo agent
keeps its frontmatter as the single writer; explicit call-site params win.
- honors BDR-107/108: levels and aliases unchanged; aliases stay the config's
vocabulary, full ids are resolved by the mod (LRN-203, BLK-029).
- LRN-180/181 made moot: the bridge answers `Skill(effort-*)` itself, no
pairing rule. "Last loaded wins" holds among pinned skills and shifts; the
one behaviour change, accepted: loading an UNPINNED skill after a shift
returns main to the harness level (orchestrators re-assert after nested
skills already, lib/effort-shift.md § Re-assert).
- LRN-204: main-loop model switch behind `mainModelSwitch` (default false) and
a context-window guard.
- BDR-044 not contradicted: the mod routes model/effort, never skills.
- Deferred (minor, challenge r1): a ceiling on model-declared efforts; the
spawn-vs-step race.