From b73d1b127e3adbec21d158557681177627e7a05a Mon Sep 17 00:00:00 2001 From: bchanot Date: Thu, 8 Oct 2026 15:25:11 +0200 Subject: [PATCH 01/22] =?UTF-8?q?chore(memory):=20model-router=20wave=200?= =?UTF-8?q?=20=E2=80=94=20plan,=20LRN-203/204,=20BLK-029,=20journal?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- .claude/memory/blockers.md | 9 + .claude/memory/journal.md | 2 + .claude/memory/learnings.md | 9 + .claude/tasks/TODO.md | 9 + .../plans/2026-10-08-model-router-mod.md | 159 ++++++++++++++++++ 5 files changed, 188 insertions(+) create mode 100644 .claude/tasks/plans/2026-10-08-model-router-mod.md diff --git a/.claude/memory/blockers.md b/.claude/memory/blockers.md index cfe9f7f..cbf3a64 100644 --- a/.claude/memory/blockers.md +++ b/.claude/memory/blockers.md @@ -48,6 +48,7 @@ rules: | BLK-026 | 2026-10-06 | `make test` red on macOS: 13 suites, GNU-only idioms in suite + 7 libs (SIGPIPE under pipefail, `sed -i`, `wc` padding, `stat -c`, `realpath -m`, bare `timeout`, `grep -oP`) | resolved | | BLK-027 | 2026-10-06 | this machine never ran `make link`/`make plugin`: no global `core.hooksPath` → post-commit push never fired, branches landed ahead of upstream; 11 vendored skills + `~/.claude/.env` missing | resolved (link) / open (plugin) | | BLK-028 | 2026-10-06 | notify-attention on a VS Code client: bell + toast silent-degradation faults (merge of BLK-019 + BLK-020): terminalBell sound default off, ext hooks only terminals born after activation, Code muted in Windows mixer | resolved | +| BLK-029 | 2026-10-08 | Claude Code mods (2.1.294): alias → id resolution for a model set by a hook lags the Agent tool's (`sonnet` → `claude-sonnet-5`, 404); Agent tool schema refuses full ids | upstream | --- @@ -305,3 +306,11 @@ rules: - **Probe order (do FIRST, before server archaeology)**: fresh VS Code terminal, `printf '\a\a\033]777;notify;Test;hello\033\\'` → splits terminal path from client renderer; palette `Help: List Signal Sounds` → Terminal Bell preview bypasses terminal/BEL/hook/dtach/ext, isolates renderer audio in one step. - **Status**: resolved (BLK-019 2026-09-01, BLK-020 A+B 2026-09-02/03). Sources superseded by this entry; bodies kept for history. - **Reference**: `~/.claude/hooks/notify-attention.sh` header documents the setting; [[LRN-145]] terminalSequence-not-/dev/tty; silent-degradation class [[LRN-047]]; sources [[BLK-019]], [[BLK-020]]. + +## BLK-029 — Mods: model alias set by a hook resolves to a stale id (`sonnet` → `claude-sonnet-5`, 404) — 2026-10-08 +- **Friction**: model-router spike. Sub-agent routed by `agent.spawn` or `tool.call Agent` rewrite with alias `sonnet`/`haiku` → "model_not_found HTTP 404, model sent to the API: claude-sonnet-5". Same alias typed by the model in the Agent tool param → `claude-sonnet-5-5`, OK. +- **Real cause**: two alias tables in the CLI (2.1.294): the Agent tool's is current, the function-hooks path's is stale. Not an access issue (`/model` lists all four tiers; explicit haiku/sonnet dispatches answered). +- **Solution**: hooks write full ids (`claude-sonnet-5-5`, `claude-haiku-4-5-20251001`, `claude-opus-5-5`, `claude-fable-5-1`) from the mod's own table; Agent tool param rewrite stays alias-only (schema enum) → the mod sets the model at `agent.spawn`, not at the param. +- **Status**: upstream (report to anthropics/claude-code with the request id `req_011Cfppp7VFt8z2Pd3zJrpUi`); workaround in model-router. +- **Reference**: [[LRN-203]], plan `.claude/tasks/plans/2026-10-08-model-router-mod.md`. + diff --git a/.claude/memory/journal.md b/.claude/memory/journal.md index d59a38e..739c1be 100644 --- a/.claude/memory/journal.md +++ b/.claude/memory/journal.md @@ -577,3 +577,5 @@ rules: - Merge (user go "tu peux merge dans develop"): final full suite green (46 suites minus the declared env red) + Health Stack shellcheck clean on the branch tip → `gitflow finish feature manual-push-mode` → develop 669db06, pushed, branch removed local + origin. 19 commits (runs A, B, C1/C2, D1/D2/D3 + docs + memory). User answered: only pushes change; commits/branches/local merges untouched; invalid value now fail-closed everywhere. User plan: dotfiles installer prompts for `gitflow.autopush` (default false) — told them the gitconfig template also needs `core.hooksPath` (the install wiped it). Open: user probe `! git push --dry-run` under autopush=false; AC6 env red (design-tool-gate); post-run-D residuals in TODO. - User tested manual-push mode on their machine: works (no auto push under `false`, bang-prefixed dry-run passes). Prompt handed over for the dotfiles repo: gitconfig template gets `core.hooksPath = ~/.claude/githooks` + `[gitflow] autopush = @AUTOPUSH@` rendered from an install question (default false, true/false only, unrendered placeholder = render failure). BDR-112 amended. +## 2026-10-08 +- model-router mod, wave 0 spike (user ask: one mod routes model + effort per request, replaces effort-* shifters + pins). Plan `.claude/tasks/plans/2026-10-08-model-router-mod.md`, 4 decisions by AskUserQuestion (spike-first main-loop switch, `CLAUDE_CODE_PLUGIN_DIRS` load, migration wave 2, names model-router / route / /route), rule "pin = entry default, sub-tasks route finer". Spike in dev-mods, hot reload on: `turn.step` effort rewrite proven (transcript `effort` field is the oracle, not `CLAUDE_EFFORT`); sub-agent model at `agent.spawn` + effort per step by agentId proven; main-loop fable → sonnet-5-5/low for 3 steps then back: works, one cold-cache step per switch INTO a model, return free. Found: hook-side alias resolver stale (`sonnet` → `claude-sonnet-5`, 404; Agent tool enum resolves the same alias to 5-5) → mod writes full ids only. feature/model-router-mod open, nothing committed yet (plan + TODO + journal pending). diff --git a/.claude/memory/learnings.md b/.claude/memory/learnings.md index c3b2bb9..c46d9a0 100644 --- a/.claude/memory/learnings.md +++ b/.claude/memory/learnings.md @@ -1748,3 +1748,12 @@ Rule: when editing a doctrine file under structure locks, grep the test's lock s ## LRN-202 — Reading a stderr-then-stdout verb from a hook: `out=$(cmd 2>&1)`, last line = word, prefix line = reason; no temp file; lib path absolute before any cd - **Context**: unpushed-guard plan used `2>"${TMPDIR:-/tmp}/x.$$"` + cat + rm: fail-OPEN when TMPDIR is full (`|| mode=auto`), predictable path, symlink-followable on shared /tmp, leaked on kill; `mode=$(cmd 2>&1 >/dev/null)` captures ONLY stderr. push-guard's `mktemp` variant added `set -u` trap hazards. The verb writes its stderr line BEFORE its stdout word in one process, so `${out##*$'\n'}` is the word and the `gitflow.sh push-mode:` line is the reason (select by prefix, not `head -1`: a bash startup warning could precede it). Resolve the lib path to an absolute one BEFORE the hook's `cd "$cwd"` (a relative invocation otherwise resolves into the target repo). - **Apply**: hooks never touch temp files for a one-line capture; anything but the expected word is treated as the fail-closed state, never as the default. Links [[BDR-114]], [[LRN-196]], [[LRN-199]]. + +## LRN-203 — Mod hooks: model ALIAS set by a hook resolves through stale table (`sonnet` → `claude-sonnet-5`, 404); write full ids; oracle = transcript fields, not `CLAUDE_EFFORT` +- **Context**: model-router spike 2026-10-08, CLI 2.1.294. `agent.spawn` or `tool.call Agent` param rewrite with alias `sonnet` → API got `claude-sonnet-5`, HTTP 404 model_not_found. Same alias passed by the model in the Agent tool param → `claude-sonnet-5-5`, fine. Full id `claude-sonnet-5-5` from hook → fine, every step answered by 5-5. Agent tool schema enum refuses full ids, so full ids reach API only via hooks. Effort rewrite at `turn.step` proven by transcript record field `effort` (high → medium); `CLAUDE_EFFORT` env + `perTurnEffort` stay at turn setting, blind to per-request rewrite. +- **Apply**: any mod that sets a model carries its own alias → full-id table (one place to bump per tier release). Verify routing with transcript `message.model` + record `effort`, never env vars. Links [[BDR-108]] (aliases as pins: still right at the Agent-tool call site, wrong inside hooks), [[BLK-029]]. + +## LRN-204 — Main-loop model switch mid-turn works, costs one cold-cache step on the full context per switch INTO a model; return free (per-model cache, 1 h TTL) +- **Context**: spike 2026-10-08, fable → `claude-sonnet-5-5`/low for 3 steps on ~260k context, then back. Conversation intact (tools, results, thinking blocks from another model in history: no error). First sonnet step cache_read 0 (full 260k billed), next steps 237k cached; return to fable step read 263k cached. +- **Apply**: switch the main loop only for spans long enough to amortize one uncached read of the whole context (many mechanical steps), never per tool call; short mechanical work → small-context haiku sub-agent. Haiku 4.5 window 200k: a long main loop cannot go to haiku at all. Flag off by default in model-router. Links [[LRN-203]], [[BDR-107]]. + diff --git a/.claude/tasks/TODO.md b/.claude/tasks/TODO.md index c624d12..6eb8418 100644 --- a/.claude/tasks/TODO.md +++ b/.claude/tasks/TODO.md @@ -1,5 +1,14 @@ # TODO +## 2026-10-08 — model-router mod: one mod routes model + effort per request (feature/model-router-mod) +Plan `.claude/tasks/plans/2026-10-08-model-router-mod.md`. Decisions 2026-10-08: main-loop +model switch spike-first then flag off; load via `CLAUDE_CODE_PLUGIN_DIRS` + link.sh; +migration of shifters/pins/model-gate in wave 2 after proof; names model-router / route / /route. +- [x] W0 spike in dev-mods (hot reload): facts a-d established 2026-10-08 (plan file § Spike facts); e moved to W1.10 +- [ ] W1 core mod in `mods/model-router/` (config, route tool, /route, agents, skills, prompt rules, visibility, tests, install) +- [ ] W2 migration: 15 skills off `Skill(effort-*)`, remove shifters + effort-pins + model-gate, census repointed, docs +- [ ] W3 optional: step heuristics, haiku classifier, quota-aware downgrade, A/B + ## 2026-09-30 — Higgsfield pack: CLI + skills in the install process, off by default (feature/higgsfield-pack) Contract `.claude/tasks/contracts/2026-09-30-higgsfield-pack-1412.md`, spec + plan under `docs/superpowers/` (transient). Approved 2026-09-30: toggle pack off by default, two toggles, diff --git a/.claude/tasks/plans/2026-10-08-model-router-mod.md b/.claude/tasks/plans/2026-10-08-model-router-mod.md new file mode 100644 index 0000000..20f3832 --- /dev/null +++ b/.claude/tasks/plans/2026-10-08-model-router-mod.md @@ -0,0 +1,159 @@ +# PLAN — model-router mod (feature/model-router-mod) + +User ask 2026-10-08: one Claude Code mod that routes every request to the +model and effort its task deserves, main loop and sub-agents alike, declared +or automatic, configurable, loaded in every session (user scope). Replaces +the five `effort-*` shifter skills, the `effort:` / `model:` frontmatter +pins, `lib/effort-pins.txt` + `.sh` and `lib/model-gate.md` once proven. + +Decisions taken 2026-10-08 (user, AskUserQuestion): +- Main-loop MODEL switch: spike first, then behind a userConfig flag, off by + default, applied only on explicit declaration. Main-loop EFFORT always routed. +- Loading: `CLAUDE_CODE_PLUGIN_DIRS` in settings.json `env`, mod lives in the + repo under `mods/model-router/`, symlinked by link.sh. No marketplace. +- Migration: wave 2, after wave 1 is proven. Mod = single source of truth. +- Names: mod `model-router`, tool `route` (model sees `mcp__model-router__route`), + command `/route`, config `~/.claude/model-router.json`. + +## Routing rule (user, 2026-10-08): pin = entry default, sub-tasks route finer +The existing pins and `effort-*` shifters were built for this same goal with +the tools of their time; the mod replaces them (or nearly). A pin is not to +be contested one by one, but some were forced: a skill pinned to one level +does many different things inside one run. So: +- the entry pin of a skill or agent (today frontmatter / effort-pins.txt, + tomorrow the config table) = the DEFAULT route of the run, never a ceiling + or a floor; +- inside the run every sub-task routes to its own phase: declared by the + skill through the `route` tool (replaces `Skill(effort-*)` + pairing rule), + or derived by the mod (Agent dispatch → orchestrate, Skill load → its + entry phase, Read/Grep result → comprehension level, bookkeeping tail → + mechanical); +- an explicit per-call choice (Agent `model`/`effort` param, `/route`, + `ultrathink`) beats the derived phase for that span; +- wave 1 keeps the frontmatter pins as the entry defaults (the mod reads the + same values), wave 2 moves them into the config table and deletes the + frontmatter + shifters. Skills get their intra-run `route` calls in wave 2 + (the 15 `Skill(effort-*)` citers first). + +## Spike facts so far (2026-10-08) +- (a) main-loop effort rewrite at `turn.step` reaches the API: transcript + records flip `effort: high` → `medium` after a `route` call. `CLAUDE_EFFORT` + is NOT an oracle (turn-level setting); the transcript `effort` field is. +- (c) sub-agent steps are visible by `agentId`; an explicit Agent `model` + param resolves fine (sonnet → claude-sonnet-5-5, haiku → claude-haiku-4-5-20251001; + haiku steps carry `effort: undefined`, no effort on that model). +- ROOT CAUSE of the 404 (T1-T3, 2026-10-08): a model set by a hook as an + ALIAS is resolved by a stale table (`sonnet` -> `claude-sonnet-5`, 404); + the Agent tool's own enum resolves the same alias to `claude-sonnet-5-5`. + T1 param rewrite + alias: 404. T2 spawn rewrite + alias: 404. T3b spawn + rewrite + full id `claude-sonnet-5-5`: OK, every step answered by + claude-sonnet-5-5 at effort medium (turn.step by agentId also OK). + Rule for the mod: ALWAYS write full model ids from its own alias -> id + table in the config (one place to bump when a tier ships). The Agent tool + schema accepts only aliases, so full ids can only come from the hooks. + To report upstream: hook-side alias resolution lags the tool's. +- T4 main-loop model switch (fable -> claude-sonnet-5-5/low, switch on): WORKS. + Steps 11 and 12 answered by claude-sonnet-5-5 at effort low, transcript + records agree, thinking still produced (4.6 s), conversation intact (897 + messages, tools and results carried across). COST: the first step after a + switch read 0 cached tokens on a ~260k context (cache is per model), the + next step read 237k. Switching BACK to fable at step 14 read 263k cached + tokens: the fable cache survived three sonnet steps (per-model caches, + 1 h TTL), so the return is free. Every switch INTO another model pays one + cold-cache step on the full context. Consequence for the design: main-loop + model switches only for spans long enough to amortize (many mechanical + steps), never per tool call; short mechanical work goes to a haiku + sub-agent whose context is small. Flag stays off by default. +- Open: T3a (param rewrite + full id), fable/opus ids for the table, + switching back mid-turn, behavior with thinking blocks from another model + in history (no error seen), headless `-p` run. + +## Harness facts (types 2.1.292, CLI 2.1.294) +- `turn.step` (async generator) rewrites `model` and `effort` per request; + `e.agentId` set inside a sub-agent loop. Pinned: turn, index, messageCount. +- `agent.spawn` rewrites `model` (not effort); result carries `agentId`. +- `tool.call {tool:'Agent'}` sees and rewrites the call's `model` / `effort` + params; `tool.call {tool:'Skill'}` names the skill loading. +- `$.tool.register` / `$.command.register` (`immediate: true` runs mid-turn). +- `$.model.classify(text, labels)`, `$.session.usage().rateLimits`. +- Hooks run under `claude -p` too (closes the BDR-107 headless gap). +- Hook budget 10 s own code; `$` calls do not count. +- Honest limit: a hook cannot know what the NEXT request will decide to do. + Routing = declared phase (skill table, `route` tool, `/route`, prompt + rules) + conservative after-the-fact heuristics on the following step. + +## Phase table (proposal, config-driven) +| phase | model | effort | when | +|---|---|---|---| +| plan | session | xhigh | brainstorm, plan, architecture, challenge synthesis, audit verdict | +| reflect | session | high | diagnosis, reading to understand, contract, review | +| orchestrate | session | medium | between dispatches, reading a report | +| escalate | session | max | `ultrathink`, stuck loop, STOP relaunch | +| judge | opus | xhigh | dispatched challengers, analyzers, audits (BDR-076) | +| implement | sonnet | medium | code from a closed plan (feater, bugfixer, …) | +| write | sonnet | medium | docs, prose from decided content | +| verify | sonnet | xhigh | verifier, security-auditor | +| mechanical | haiku | low | cp/mv, git bookkeeping, status collection, listing | + +## Wave 0 — spike (dev-mods folder, hot reload, this session) +- [x] W0.1 minimal mod: `/route` command, `route` tool, `turn.step` logging + + rewrite, `agent.spawn` rewrite, `ultrathink` → max, spinner suffix +- [x] W0.2 `claude plugin validate` clean; type-check with the header tsconfig +- [x] W0.3 facts to establish, each with its evidence (usage.model, CLAUDE_EFFORT, + debug log): (a) effort rewrite on main loop takes effect; (b) model rewrite + on main loop mid-turn: works / breaks (thinking signatures, cache, tools); + (c) sub-agent model via spawn + effort via step by agentId; (d) `/route` + immediate mid-turn; (e) load via `CLAUDE_CODE_PLUGIN_DIRS` from settings env → moved to W1.10 +- [x] W0.4 record facts → journal + BDR draft; freeze wave 1 scope (facts in this file; registries pending user go) + +## Wave 1 — core (repo `mods/model-router/`) +- [ ] W1.1 config loader: `~/.claude/model-router.json` (phases, agents, + skills, prompt rules, defaults); schema check; `/route reload` +- [ ] W1.2 state: per-loop phase (main + agentId map), reset at `turn.start` + to the prompt-derived phase; explicit > table > heuristic +- [ ] W1.3 `route` tool + `/route [phase|show|reload|clear]` (immediate) +- [ ] W1.4 agents: `tool.call Agent` param rewrite + `agent.spawn` model + + `turn.step` effort by agentId; covers built-ins (Explore, Plan, general-purpose) +- [ ] W1.5 skills: `tool.call Skill` → phase from the skills table; any + skill load RESETS the main route to that skill's entry phase (table, + else the frontmatter `effort:` the harness just applied), so no route + declared earlier in the turn survives a skill change silently +- [ ] W1.5b single-writer bridge for the legacy shifters (user, 2026-10-08: + doublon + silent one-way conflict): `tool.call {tool:'Skill', skill: + /^effort-/}` answers WITHOUT `next` (skill text never loaded, no pairing + rule) and translates the level into a route on the calling loop; + `skill.prompt {skill:/^effort-/}` does the same for a user-typed + `/effort-max` and returns a one-line text. The 15 citers keep working + untouched until wave 2 rewrites them to `route`. Rule: every effort + change goes through the mod's state; frontmatter values are inputs. +- [ ] W1.6 prompt rules: `ultrathink` → escalate; keyword → phase (effort up + only; never a main-loop model change without declaration) +- [ ] W1.7 visibility: Spinner suffix `· /`, `$.ui.status`, + `$.ui.log` when verbose, `/route show` +- [ ] W1.8 userConfig: `mainLoopModelSwitch` (false), `verbose` (false), + `classifier` (false) +- [ ] W1.9 tests `*.test.ts` under `claude plugin test`; `claude plugin validate` +- [ ] W1.10 install: `mods/` symlink + `CLAUDE_CODE_PLUGIN_DIRS` in settings.json + env via link.sh; doctor line; README/USAGE/CHANGELOG +- [ ] W1.11 contract + GATE 0 + fresh verifier + security gate; `make test` + +## Wave 2 — migration (after wave 1 proven) +- [ ] W2.1 15 skills `Skill(effort-*)` → `route` tool calls (lib/effort-shift.md rewritten) +- [ ] W2.2 remove `skills/effort-*`, `lib/effort-pins.txt`, `lib/effort-pins.sh`, + install/update steps, `effort:` frontmatter on skills and agents +- [ ] W2.3 `lib/model-gate.md` + `lib/model-check.sh` → mod rule (reflect on a + small model → raise); census tests repointed to the config table +- [ ] W2.4 docs + CHANGELOG + registries (BDR, LRN, EVAL via effort-audit.py) + +## Wave 3 — optional +- [ ] W3.1 per-step heuristics (Read/Grep → +1 level next step; Agent return → orchestrate) +- [ ] W3.2 haiku classifier on `prompt.submit` (`$.model.classify`) +- [ ] W3.3 quota-aware downgrade from `$.session.usage().rateLimits` +- [ ] W3.4 A/B via `lib/effort-audit.py` + +## Risks +- Main-loop model switch mid-turn unproven (W0.3b decides). +- A mod bug cuts all routing at once: fail-open (`.catch` → `next(e)`), never deny. +- Two sources of truth during wave 1 (pins + mod): mod must agree with the + pins until wave 2 removes them. +- API early access: types change between releases; pin the CLI version in the README. From b721c94dcbae8167d3fc1ce36ccf2b506f2ff9e9 Mon Sep 17 00:00:00 2001 From: bchanot Date: Thu, 8 Oct 2026 16:26:57 +0200 Subject: [PATCH 02/22] =?UTF-8?q?feat(mods):=20model-router=20mod,=20wave?= =?UTF-8?q?=201-A=20=E2=80=94=20per-request=20model/effort=20routing?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Function-hooks plugin under mods/model-router: routes effort (and, behind a flag, the model) of every main-loop request, sets built-in sub-agents' model at spawn with full ids, answers Skill(effort-*) itself (single writer, no pairing rule), exposes the route tool and /route, validates the optional ~/.claude/model-router.json. 11 plugin tests, validate + tsc clean. Contract .claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md. --- mods/model-router/.claude-plugin/plugin.json | 8 + mods/model-router/hooks/hooks.json | 1 + mods/model-router/hooks/register.test.ts | 202 +++++ mods/model-router/hooks/register.ts | 834 +++++++++++++++++++ 4 files changed, 1045 insertions(+) create mode 100644 mods/model-router/.claude-plugin/plugin.json create mode 100644 mods/model-router/hooks/hooks.json create mode 100644 mods/model-router/hooks/register.test.ts create mode 100644 mods/model-router/hooks/register.ts diff --git a/mods/model-router/.claude-plugin/plugin.json b/mods/model-router/.claude-plugin/plugin.json new file mode 100644 index 0000000..548269f --- /dev/null +++ b/mods/model-router/.claude-plugin/plugin.json @@ -0,0 +1,8 @@ +{ + "name": "model-router", + "version": "0.1.0", + "description": "Routes every model request (main loop and sub-agents) to the model and effort its phase deserves, from a user config, a route tool, /route and skill or prompt rules.", + "author": { + "name": "bchanot" + } +} diff --git a/mods/model-router/hooks/hooks.json b/mods/model-router/hooks/hooks.json new file mode 100644 index 0000000..5fcba7c --- /dev/null +++ b/mods/model-router/hooks/hooks.json @@ -0,0 +1 @@ +{ "modules": ["./register.ts"] } diff --git a/mods/model-router/hooks/register.test.ts b/mods/model-router/hooks/register.test.ts new file mode 100644 index 0000000..5c0c634 --- /dev/null +++ b/mods/model-router/hooks/register.test.ts @@ -0,0 +1,202 @@ +import { test, expect } from 'claude-code/testing' +import type { Engine } from 'claude-code/testing' +import type { On } from 'claude-code' + +/** Fires session.start so the mod loads its config and registers /route. */ +async function boot($: Engine, on: On): Promise { + on('session.start', ($, e) => ({ cwd: e.cwd })) + await $.session.start({ cwd: '/tmp', surface: null, isInteractive: false }) +} + +/** Runs `/route ` as the user would type it. */ +async function route($: Engine, args: string): Promise { + const out = await $.command.run({ + command: 'route', + args, + origin: { kind: 'composer' }, + presentation: { isFullscreen: false, columns: 80 }, + }) + return out.text ?? '' +} + +/** The `main:` line of `show`: the phases listing holds every id and level. */ +function mainLine(text: string): string { + return text.split('\n').find(l => l.startsWith('main:')) ?? '' +} + +const spawnInput = (model?: string) => ({ + tool_use_id: 't1', + prompt: 'x', + description: 'd', + subagentType: 'Explore', + provider: { plugin: 'engine', tier: 'core' as const }, + parentModel: 'claude-fable-5-1', + background: false, + fork: false, + ...(model === undefined ? {} : { model }), +}) + +test('Skill(effort-low) is answered without next, route shows low', async ( + $, on) => { + let reached = false + on('tool.call', { tool: 'Skill' }, () => { + reached = true + return { result: { success: true, commandName: 'bottom' } } + }) + await boot($, on) + const out = await $.tool.call({ tool: 'Skill', skill: 'effort-low' }) + expect(out).toMatchObject({ + result: { success: true, commandName: 'effort-low' }, + }) + expect(reached).toBe(false) + const line = mainLine(await route($, 'show')) + expect(line).toContain('skill effort-low') + expect(line).toContain('effort low') +}) + +test('route tool with phase orchestrate sets medium on main', async ( + $, on) => { + await boot($, on) + const out = await $.tool.call({ + tool: 'mcp__model-router__route', + phase: 'orchestrate', + }) + expect(out).toHaveProperty('result') + const line = mainLine(await route($, 'show')) + expect(line).toContain('model orchestrate') + expect(line).toContain('effort medium') +}) + +test('/route clear drops the route', async ($, on) => { + await boot($, on) + await $.tool.call({ tool: 'mcp__model-router__route', phase: 'plan' }) + expect(mainLine(await route($, 'show'))).toContain('model plan') + await route($, 'clear') + expect(mainLine(await route($, 'show'))).toContain('session defaults') +}) + +test('/route bogus names the phases', async ($, on) => { + await boot($, on) + const text = await route($, 'bogus') + expect(text).toContain('unknown') + for (const phase of ['plan', 'judge', 'explore', 'mechanical']) { + expect(text).toContain(phase) + } +}) + +test('ultrathink in a prompt sets escalate on main', async ($, on) => { + on('prompt.submit', ($, e) => ({ text: e.text })) + await boot($, on) + await $.prompt.submit({ + text: 'ultrathink please', + wait: false, + origin: { kind: 'composer' }, + }) + expect(mainLine(await route($, 'show'))).toContain('prompt escalate') +}) + +test('/route model: alias resolved, id passed, typo refused', async ( + $, on) => { + await boot($, on) + const alias = await route($, 'model=sonnet') + expect(mainLine(alias)).toContain('claude-sonnet-5-5') + const id = await route($, 'model=claude-x-9') + expect(mainLine(id)).toContain('claude-x-9') + expect(await route($, 'model=sonet')).toContain('unknown') +}) + +test('Explore spawn gets the sonnet id; an explicit model wins', async ( + $, on) => { + const seen: (string | undefined)[] = [] + on('agent.spawn', ($, e) => { + seen.push(e.model) + return { model: e.model ?? e.parentModel, agentId: 'a1' } + }) + await boot($, on) + await $.agent.spawn(spawnInput()) + await $.agent.spawn(spawnInput('opus')) + expect(seen).toEqual(['claude-sonnet-5-5', 'opus']) +}) + +test('a sticky /route wins over a model-declared route', async ($, on) => { + await boot($, on) + await route($, 'judge') + const out = await $.tool.call({ + tool: 'mcp__model-router__route', + phase: 'mechanical', + }) + expect(JSON.stringify(out)).toContain('sticky') + expect(mainLine(await route($, 'show'))).toContain('user judge') +}) + +test('/route off makes the route tool a no-op', async ($, on) => { + await boot($, on) + await route($, 'off') + const out = await $.tool.call({ + tool: 'mcp__model-router__route', + phase: 'plan', + }) + expect(JSON.stringify(out)).toContain('off') + expect(mainLine(await route($, 'show'))).toContain('session defaults') +}) + +type Seen = { model: string; effort: unknown } + +/** A bottom turn.step hook that records what reaches the model. */ +function recordSteps(on: On, seen: Seen[]): void { + on('turn.step', async function* ($, e) { + seen.push({ model: e.model, effort: e.effort }) + return { + turnId: e.turnId, + index: e.index, + answer: '', + toolUses: [], + stopReason: 'end_turn' as const, + usage: null, + } + }) +} + +const stepInput = (agentId?: string) => ({ + turnId: 'u1', + index: 0, + model: 'claude-fable-5-1', + effort: 'low' as const, + messageCount: 1, + ...(agentId === undefined ? {} : { agentId }), +}) + +/** Streams one turn.step to its end; the hooks run as the chunks flow. */ +async function runStep( + $: Engine, + input: ReturnType, +): Promise { + const stream = $.turn.step(input) + for await (const _chunk of stream) { + // chunks are not under test + } + await stream.result +} + +test('the main step carries the routed effort, model untouched', async ( + $, on) => { + const seen: Seen[] = [] + recordSteps(on, seen) + await boot($, on) + await $.tool.call({ tool: 'mcp__model-router__route', phase: 'plan' }) + await runStep($, stepInput()) + expect(seen).toEqual([{ model: 'claude-fable-5-1', effort: 'xhigh' }]) +}) + +test('a tabled agent steps at its table effort', async ($, on) => { + const seen: Seen[] = [] + recordSteps(on, seen) + on('agent.spawn', ($, e) => ({ + model: e.model ?? e.parentModel, + agentId: 'a1', + })) + await boot($, on) + await $.agent.spawn(spawnInput()) + await runStep($, stepInput('a1')) + expect(seen).toEqual([{ model: 'claude-fable-5-1', effort: 'medium' }]) +}) diff --git a/mods/model-router/hooks/register.ts b/mods/model-router/hooks/register.ts new file mode 100644 index 0000000..9bf4ae4 --- /dev/null +++ b/mods/model-router/hooks/register.ts @@ -0,0 +1,834 @@ +// model-router: routes each model request (main loop and sub-agents) to the +// model and effort its phase deserves. Phases, agents, skills and prompt +// rules come from DEFAULT_CONFIG, overridable by ~/.claude/model-router.json. +// One writer per concern: the Agent tool's own params are never rewritten, +// the spawn hook sets an agent's model once, turn.step sets efforts. +import type { EngineInterface, On, Register, TurnStepInput } from 'claude-code' + +type Api = EngineInterface +type Level = 'low' | 'medium' | 'high' | 'xhigh' | 'max' +type Route = { model?: string; effort?: Level } // model: alias or full id +type PromptRule = { pattern: string; phase: string } +type Config = { + models: Record // alias -> full id + windows: Record // full id -> context window (tokens) + phases: Record + agents: Record // built-in subagentType -> phase + skills: Record // skill name -> phase + prompt: PromptRule[] + mainModelSwitch: boolean + verbose: boolean + spinner: boolean +} +type Rule = { re: RegExp; phase: string } +type Source = 'user' | 'model' | 'skill' | 'prompt' | 'slash' +type Routed = { phase: string; route: Route; source: Source } +type Loop = { + effort?: Level + model?: string // routed by the table or an in-agent call + spawnModel: string // the engine's model at spawn + frozen: boolean // fork or workflow agent: never re-modelled + explicitModel: boolean // Agent call gave a model: axis frozen + explicitEffort: boolean // Agent call gave an effort: axis frozen +} +type State = { + cfg: Config + rules: Rule[] + source: string // 'defaults' or the override path + userMain: Routed | null // /route by the user, sticky until /route clear + turnMain: Routed | null // tool, skill, slash, prompt; dropped at turn end + pendingPrompt: Routed | null // typed mid-turn, promoted next turn + loops: Map // agentId -> that loop's routing + explicitEffort: Map // Agent tool_use_id -> effort param + skillCalls: number // Skill tool calls in flight + off: boolean // /route off: every hook passes through + lastMain: string // "model/effort" of the last main step (spinner) + windowWarned: boolean // context-window warning already logged this turn +} +type Log = (text: string) => void +type StepIn = Readonly +type Plan = { model: string; effort: StepIn['effort'] } +type RouteInput = { + phase?: unknown + effort?: unknown + clear?: unknown + agentId?: string +} +type Picked = { phase: string; route: Route } + +const LEVELS: readonly Level[] = ['low', 'medium', 'high', 'xhigh', 'max'] +const MODEL_ID = /^claude-[a-z0-9.-]+$/ +const TOOL = 'mcp__model-router__route' +const EFFORT_SKILL = /^effort-(low|medium|high|xhigh|max)$/ +const OVERRIDE = '.claude/model-router.json' +const HAIKU = 'claude-haiku' + +const DEFAULT_CONFIG: Config = { + models: { + haiku: 'claude-haiku-4-5-20251001', + sonnet: 'claude-sonnet-5-5', + opus: 'claude-opus-5-5', + fable: 'claude-fable-5-1', + }, + windows: { 'claude-haiku-4-5-20251001': 200000 }, + phases: { + plan: { effort: 'xhigh' }, + reflect: { effort: 'high' }, + orchestrate: { effort: 'medium' }, + escalate: { effort: 'max' }, + judge: { model: 'opus', effort: 'xhigh' }, + implement: { model: 'sonnet', effort: 'medium' }, + write: { model: 'sonnet', effort: 'medium' }, + verify: { model: 'sonnet', effort: 'xhigh' }, + explore: { model: 'sonnet', effort: 'medium' }, + mechanical: { model: 'haiku', effort: 'low' }, + }, + // Built-ins only: repo agents keep their frontmatter pin (wave 2). + agents: { Explore: 'explore', Plan: 'judge' }, + skills: {}, + prompt: [{ pattern: '\\bultrathink\\b', phase: 'escalate' }], + mainModelSwitch: false, + verbose: false, + spinner: true, +} + +// ---- config ---------------------------------------------------------- + +const isRecord = (v: unknown): v is Record => + typeof v === 'object' && v !== null && !Array.isArray(v) +const isLevel = (v: unknown): v is Level => LEVELS.some(l => l === v) +const isFullId = (v: unknown): v is string => + typeof v === 'string' && MODEL_ID.test(v) +const hasKey = (table: object, key: string): boolean => + Object.hasOwn(table, key) + +function isModelName(models: Record, v: unknown): v is string { + return typeof v === 'string' && (hasKey(models, v) || isFullId(v)) +} + +function resolveModel(cfg: Config, name: string): string { + return hasKey(cfg.models, name) ? (cfg.models[name] ?? name) : name +} + +function phaseRoute(cfg: Config, phase: string): Route | undefined { + return hasKey(cfg.phases, phase) ? cfg.phases[phase] : undefined +} + +/** Merges a user table over a default one, dropping invalid entries. */ +function mergeTable( + base: Record, + user: unknown, + name: string, + accept: (key: string, value: unknown) => T | undefined, + log: Log, +): Record { + const out = { ...base } + if (!isRecord(user)) return out + for (const [key, value] of Object.entries(user)) { + const ok = key === '__proto__' ? undefined : accept(key, value) + if (ok === undefined) log(`model-router: config ${name}.${key} ignored`) + else out[key] = ok + } + return out +} + +const acceptModel = (_key: string, v: unknown): string | undefined => + isFullId(v) ? v : undefined + +const acceptWindow = (key: string, v: unknown): number | undefined => + isFullId(key) && typeof v === 'number' && Number.isInteger(v) && v > 0 + ? v + : undefined + +function acceptPhase(models: Record) { + return (_key: string, v: unknown): Route | undefined => { + if (!isRecord(v)) return undefined + const route: Route = {} + if (v.effort !== undefined) { + if (!isLevel(v.effort)) return undefined + route.effort = v.effort + } + if (v.model !== undefined) { + if (!isModelName(models, v.model)) return undefined + route.model = v.model + } + return route.effort || route.model ? route : undefined + } +} + +function acceptPhaseRef(phases: Record) { + return (_key: string, v: unknown): string | undefined => + typeof v === 'string' && hasKey(phases, v) ? v : undefined +} + +function compiles(pattern: string): boolean { + try { + new RegExp(pattern, 'i') + return true + } catch { + return false + } +} + +function acceptRule(phases: Record, v: unknown) { + if (!isRecord(v)) return undefined + const { pattern, phase } = v + if (typeof pattern !== 'string' || typeof phase !== 'string') return undefined + return hasKey(phases, phase) && compiles(pattern) + ? { pattern, phase } + : undefined +} + +/** A user prompt array replaces the default rules; bad rules are dropped. */ +function mergePrompt( + base: PromptRule[], + user: unknown, + phases: Record, + log: Log, +): PromptRule[] { + if (!Array.isArray(user)) return base + const rules: PromptRule[] = [] + for (const item of user as unknown[]) { + const rule = acceptRule(phases, item) + if (rule) rules.push(rule) + else log('model-router: config prompt rule ignored') + } + return rules +} + +const pickBool = (v: unknown, fallback: boolean): boolean => + typeof v === 'boolean' ? v : fallback + +/** Defaults overlaid with the user's entries, each validated first. */ +function mergeConfig(user: unknown, log: Log): Config { + const base = structuredClone(DEFAULT_CONFIG) + if (!isRecord(user)) return base + const models = mergeTable( + base.models, user.models, 'models', acceptModel, log) + const phases = mergeTable( + base.phases, user.phases, 'phases', acceptPhase(models), log) + const ref = acceptPhaseRef(phases) + return { + models, + windows: mergeTable( + base.windows, user.windows, 'windows', acceptWindow, log), + phases, + agents: mergeTable(base.agents, user.agents, 'agents', ref, log), + skills: mergeTable(base.skills, user.skills, 'skills', ref, log), + prompt: mergePrompt(base.prompt, user.prompt, phases, log), + mainModelSwitch: pickBool(user.mainModelSwitch, base.mainModelSwitch), + verbose: pickBool(user.verbose, base.verbose), + spinner: pickBool(user.spinner, base.spinner), + } +} + +async function readOverride( + $: Api, + log: Log, +): Promise<{ path: string; data: unknown } | undefined> { + try { + const home = await $.env.get('HOME') + if (!home) return undefined + const path = `${home}/${OVERRIDE}` + if (!(await $.fs.exists(path))) return undefined + return { path, data: JSON.parse(await $.fs.read(path)) } + } catch (err) { + log(`model-router: ${OVERRIDE} unreadable (${String(err)}); defaults`) + return undefined + } +} + +/** Never throws; a failed read or parse leaves the defaults. */ +async function loadConfig( + $: Api, + log: Log, +): Promise<{ cfg: Config; source: string }> { + const found = await readOverride($, log) + if (!found) return { cfg: mergeConfig(undefined, log), source: 'defaults' } + return { cfg: mergeConfig(found.data, log), source: found.path } +} + +function compileRules(cfg: Config): Rule[] { + return cfg.prompt.map(r => ({ + re: new RegExp(r.pattern, 'i'), + phase: r.phase, + })) +} + +// ---- state ----------------------------------------------------------- + +function newState(cfg: Config, source: string): State { + return { + cfg, + rules: compileRules(cfg), + source, + userMain: null, + turnMain: null, + pendingPrompt: null, + loops: new Map(), + explicitEffort: new Map(), + skillCalls: 0, + off: false, + lastMain: '', + windowWarned: false, + } +} + +/** The main loop's effective route: user /route > latest turn route. */ +const mainRoute = (st: State): Routed | null => st.userMain ?? st.turnMain + +function loopOf(st: State, agentId: string): Loop { + const known = st.loops.get(agentId) + if (known) return known + const fresh: Loop = { + spawnModel: '', + frozen: false, + explicitModel: false, + explicitEffort: false, + } + st.loops.set(agentId, fresh) + return fresh +} + +/** Writes a route on an agent loop, never on an axis given explicitly. */ +function writeLoop(loop: Loop, route: Route): void { + if (!loop.explicitEffort) loop.effort = route.effort + if (!loop.explicitModel) loop.model = route.model +} + +function clearRoutes(st: State): void { + st.userMain = null + st.turnMain = null + st.pendingPrompt = null +} + +// ---- text ------------------------------------------------------------ + +const modelText = (cfg: Config, model: string | undefined): string => + model === undefined ? '-' : resolveModel(cfg, model) + +function mainText(st: State): string { + const r = mainRoute(st) + if (!r) return 'main: session defaults' + const model = modelText(st.cfg, r.route.model) + return `main: ${r.source} ${r.phase} · model ${model} · effort ${ + r.route.effort ?? '-'}` +} + +function phasesText(cfg: Config): string { + const entries = Object.entries(cfg.phases).map(([name, r]) => + `${name}=${r.model ? resolveModel(cfg, r.model) : 'session'}/${ + r.effort ?? 'session'}`) + return `phases: ${entries.join(' ')}` +} + +function show(st: State): string { + const c = st.cfg + const flag = (b: boolean) => (b ? 'on' : 'off') + return [ + mainText(st), + `router: ${st.off ? 'off' : 'on'} · switch: ${flag(c.mainModelSwitch)} · ` + + `verbose: ${flag(c.verbose)} · spinner: ${flag(c.spinner)}`, + `live loops: ${st.loops.size}`, + phasesText(c), + `config: ${st.source}`, + ].join('\n') +} + +function statusLine(st: State): string { + const r = mainRoute(st) + const now = st.off ? 'off' : r ? `${r.source} ${r.phase}` : 'session defaults' + return `route: ${now}${st.cfg.mainModelSwitch ? ' · switch on' : ''}` +} + +const refresh = ($: Api, st: State): void => $.ui.status(statusLine(st)) + +function vlog($: Api, st: State, text: string): void { + if (st.cfg.verbose) $.ui.log(text) +} + +const unknownText = (cfg: Config, token: string): string => + `unknown token "${token}"; phases: ${Object.keys(cfg.phases).join(' ')}; ` + + `levels: ${LEVELS.join(' ')}; models: ${Object.keys(cfg.models).join(' ')} ` + + 'or a full claude-* id' + +// ---- /route command -------------------------------------------------- + +/** One token: `model=x`, `effort=y`, a bare alias, id or level. */ +function applyToken(cfg: Config, route: Route, token: string): boolean { + const eq = token.indexOf('=') + const key = eq < 0 ? '' : token.slice(0, eq) + const value = eq < 0 ? token : token.slice(eq + 1) + if (key !== '' && key !== 'model' && key !== 'effort') return false + if (key !== 'model' && isLevel(value)) route.effort = value + else if (key !== 'effort' && isModelName(cfg.models, value)) { + route.model = value + } else return false + return true +} + +function parseRoute(cfg: Config, args: string): Routed | string { + const words = args.trim().split(/\s+/) + const only = words.length === 1 ? words[0] : undefined + const named = only === undefined ? undefined : phaseRoute(cfg, only) + if (only !== undefined && named) { + return { phase: only, route: { ...named }, source: 'user' } + } + const route: Route = {} + for (const word of words) { + if (!applyToken(cfg, route, word)) return unknownText(cfg, word) + } + return { phase: 'custom', route, source: 'user' } +} + +function toggle(st: State, what: string, arg: string | undefined): string { + if (arg !== 'on' && arg !== 'off') return `usage: /route ${what} on|off` + if (what === 'switch') st.cfg.mainModelSwitch = arg === 'on' + else st.cfg.verbose = arg === 'on' + return show(st) +} + +function setUserRoute($: Api, st: State, args: string): string { + const parsed = parseRoute(st.cfg, args) + if (typeof parsed === 'string') return parsed + st.userMain = parsed + refresh($, st) + return show(st) +} + +/** (Re)loads the config into the state and re-registers the route tool. */ +async function reloadConfig($: Api, st: State): Promise { + const loaded = await loadConfig($, text => $.ui.log(text)) + st.cfg = loaded.cfg + st.rules = compileRules(loaded.cfg) + st.source = loaded.source + await registerTool($, st) +} + +async function handleCommand($: Api, st: State, args: string): Promise { + const [head = '', ...rest] = args.trim().split(/\s+/) + switch (head) { + case '': + case 'show': + return show(st) + case 'clear': + clearRoutes(st) + refresh($, st) + return 'route cleared\n' + show(st) + case 'on': + case 'off': + st.off = head === 'off' + refresh($, st) + return show(st) + case 'reload': + await reloadConfig($, st) + refresh($, st) + return 'config reloaded\n' + show(st) + case 'switch': + case 'verbose': + return toggle(st, head, rest[0]) + default: + return setUserRoute($, st, args) + } +} + +// ---- route tool ------------------------------------------------------ + +async function registerTool($: Api, st: State): Promise { + try { + await $.tool.register({ + name: 'route', + description: + 'Declare the phase of the work ahead so the next model requests ' + + 'run at the effort (and model) it deserves. Call it before a span ' + + 'of work changes nature (planning, orchestrating, mechanical ' + + 'work). It acts on the calling loop only; no model choice here. ' + + `Phases: ${Object.keys(st.cfg.phases).join(', ')}.`, + inputSchema: { + type: 'object', + properties: { + phase: { type: 'string', enum: Object.keys(st.cfg.phases) }, + effort: { type: 'string', enum: [...LEVELS] }, + clear: { type: 'boolean', description: 'drop this loop\'s route' }, + }, + }, + }) + } catch (err) { + $.ui.log(`model-router: route tool not registered: ${String(err)}`) + } +} + +async function registerCommand($: Api): Promise { + try { + await $.command.register({ + name: 'route', + description: 'model-router: show or set the model and effort route', + argumentHint: + '[show|clear|off|on|reload||model= effort=' + + '|switch on|off|verbose on|off]', + immediate: true, + }) + } catch (err) { + $.ui.log(`model-router: /route not registered: ${String(err)}`) + } +} + +function pickRoute(cfg: Config, phase: unknown, effort: unknown) { + const phases = Object.keys(cfg.phases).join(', ') + const named = typeof phase === 'string' ? phaseRoute(cfg, phase) : undefined + if (phase !== undefined && !named) { + return `unknown phase "${String(phase)}"; phases: ${phases}` + } + if (effort !== undefined && !isLevel(effort)) { + return `unknown effort "${String(effort)}"; levels: ${LEVELS.join(', ')}` + } + if (phase === undefined && effort === undefined) { + return `give a phase (${phases}), an effort, or clear` + } + const route: Route = { ...named, ...(isLevel(effort) ? { effort } : {}) } + const name = typeof phase === 'string' ? phase : `effort-${String(effort)}` + return { phase: name, route } +} + +function applyRoute(st: State, agentId: string | undefined, p: Picked): void { + if (agentId === undefined) { + st.turnMain = { phase: p.phase, route: p.route, source: 'model' } + } else { + writeLoop(loopOf(st, agentId), p.route) + } +} + +function clearLoop(st: State, agentId: string | undefined): string { + if (agentId === undefined) { + st.turnMain = null + return 'route cleared for main' + } + const loop = st.loops.get(agentId) + if (loop) writeLoop(loop, {}) + return 'route cleared for this agent' +} + +/** Truthful answer: states what the calling loop will actually do. */ +function routedText(st: State, agentId: string | undefined, p: Picked): string { + if (agentId === undefined && st.userMain) { + return `recorded ${p.phase} for this turn, but a sticky /route ` + + `${st.userMain.phase} is in force; it wins until /route clear` + } + const loop = agentId === undefined ? undefined : st.loops.get(agentId) + const effort = loop?.explicitEffort ? undefined : p.route.effort + const model = agentId === undefined && !st.cfg.mainModelSwitch + ? undefined + : loop?.explicitModel ? undefined : p.route.model + const modelNote = p.route.model && model === undefined ? ' (switch off)' : '' + return `routed ${agentId === undefined ? 'main' : 'this agent'} to ` + + `${p.phase}: effort ${effort ?? 'unchanged'}, model ` + + `${model === undefined ? 'unchanged' : resolveModel(st.cfg, model)}` + + modelNote +} + +function handleRouteTool(st: State, e: RouteInput) { + if (st.off) { + return { + result: 'model-router is off (/route on to resume); nothing routed', + } + } + if (e.clear === true) return { result: clearLoop(st, e.agentId) } + const picked = pickRoute(st.cfg, e.phase, e.effort) + if (typeof picked === 'string') return { deny: picked } + applyRoute(st, e.agentId, picked) + return { result: routedText(st, e.agentId, picked) } +} + +// ---- skills ---------------------------------------------------------- + +const skillResult = (skill: string, line: string) => ({ + result: { success: true, commandName: skill, status: 'inline' as const }, + context: [line], +}) + +/** Answers Skill(effort-) in place: one writer, the skill never loads. */ +function effortBridge(st: State, agentId: string | undefined, skill: string, + level: Level) { + if (agentId === undefined) { + const route = { ...st.turnMain?.route, effort: level } + st.turnMain = { phase: skill, route, source: 'skill' } + return skillResult(skill, st.userMain + ? `model-router: ${skill} recorded, but a sticky /route ` + + `${st.userMain.phase} is in force and wins until /route clear.` + : `model-router: effort → ${level} for this loop from the next ` + + `request on; the ${skill} skill text was not loaded.`) + } + const loop = loopOf(st, agentId) + if (loop.explicitEffort) { + return skillResult(skill, 'model-router: this agent was dispatched with ' + + 'an explicit effort; the shift does not apply.') + } + loop.effort = level + return skillResult(skill, `model-router: effort → ${level} for this ` + + `loop from the next request on; the ${skill} skill text was not loaded.`) +} + +/** A non-effort skill load: resets the loop's route, applies its table row. */ +function onSkillLoad(st: State, skill: string, agentId: string | undefined) { + const table = hasKey(st.cfg.skills, skill) ? st.cfg.skills[skill] : undefined + const route = table === undefined ? undefined : phaseRoute(st.cfg, table) + if (agentId === undefined) { + if (st.turnMain && st.turnMain.source !== 'prompt') st.turnMain = null + if (table !== undefined && route) { + st.turnMain = { phase: table, route, source: 'skill' } + } + return + } + const loop = route ? loopOf(st, agentId) : st.loops.get(agentId) + if (loop) writeLoop(loop, { effort: route?.effort }) +} + +/** A user-typed /effort-: prepends one line, args ride in the text. */ +function slashEffort(st: State, skill: string, text: string) { + const level = EFFORT_SKILL.exec(skill)?.[1] + if (!isLevel(level)) return undefined + st.turnMain = { phase: skill, route: { effort: level }, source: 'slash' } + const line = st.userMain + ? `Effort ${level} recorded; the sticky /route ${st.userMain.phase} ` + + 'wins until /route clear.' + : `Effort shifted to ${level} by model-router for this turn.` + return { text: line + '\n' + text } +} + +// ---- agents ---------------------------------------------------------- + +type SpawnIn = { + tool_use_id: string + subagentType: string + provider: { plugin: string } + model?: string + fork: boolean + workflow?: unknown +} + +/** + * The table row of a built-in agent. Known limit: provider.plugin === + * 'engine' is the best built-in test at spawn; a user agent named Explore + * in a foreign project also matches (wave 1: sonnet/medium on it). + */ +function spawnRoute(cfg: Config, e: SpawnIn, frozen: boolean) { + if (frozen || e.provider.plugin !== 'engine') return undefined + if (!hasKey(cfg.agents, e.subagentType)) return undefined + const phase = cfg.agents[e.subagentType] + return phase === undefined ? undefined : phaseRoute(cfg, phase) +} + +function trackLoop(st: State, e: SpawnIn, started: { + model: string + agentId?: string +}, route: Route | undefined, frozen: boolean): void { + const given = st.explicitEffort.get(e.tool_use_id) + st.explicitEffort.delete(e.tool_use_id) + if (started.agentId === undefined) return + st.loops.set(started.agentId, { + spawnModel: started.model, + frozen, + explicitModel: e.model !== undefined, + explicitEffort: given !== undefined, + effort: given ? undefined : route?.effort, + }) +} + +// ---- turn steps ------------------------------------------------------ + +function agentPlan(st: State, e: StepIn): Plan { + const loop = e.agentId === undefined ? undefined : st.loops.get(e.agentId) + const wanted = loop?.model + const reroute = loop && wanted !== undefined && !loop.frozen && + e.model === loop.spawnModel + return { + model: reroute ? resolveModel(st.cfg, wanted) : e.model, + effort: loop?.effort ?? e.effort, + } +} + +/** True when the context still fits the target model's known window. */ +async function windowOk($: Api, st: State, id: string): Promise { + const limit = hasKey(st.cfg.windows, id) ? st.cfg.windows[id] : undefined + if (limit === undefined) return true + let tokens: number | undefined + try { + tokens = (await $.session.usage()).context.tokens + } catch { + tokens = undefined + } + if (typeof tokens === 'number' && tokens < limit) return true + if (!st.windowWarned) { + st.windowWarned = true + $.ui.log(`model-router: no switch to ${id}: context not known to fit`) + } + return false +} + +async function mainPlan($: Api, st: State, e: StepIn): Promise { + const set = mainRoute(st) + const effort = set?.route.effort ?? e.effort + const wanted = set?.route.model + if (wanted === undefined || !st.cfg.mainModelSwitch) { + return { model: e.model, effort } + } + const id = resolveModel(st.cfg, wanted) + return { model: (await windowOk($, st, id)) ? id : e.model, effort } +} + +async function planStep($: Api, st: State, e: StepIn): Promise { + const plan = e.agentId === undefined + ? await mainPlan($, st, e) + : agentPlan(st, e) + // Haiku takes no effort: omit it rather than send a hook-set value. + return plan.model.startsWith(HAIKU) ? { ...plan, effort: undefined } : plan +} + +function withPlan(e: StepIn, plan: Plan): StepIn { + const { effort: _replaced, ...rest } = e + const base = { ...rest, model: plan.model } + return plan.effort === undefined ? base : { ...base, effort: plan.effort } +} + +function stepLog(e: StepIn, plan: Plan): string { + const id = e.agentId + const loop = id === undefined ? 'main' : `agent ${id.slice(0, 8)}` + return `step ${e.index} ${loop}: ${e.model}/${String(e.effort)} → ` + + `${plan.model}/${String(plan.effort)}` +} + +function noteMain($: Api, st: State, plan: Plan): void { + st.lastMain = `${plan.model.replace(/^claude-/, '')}/${plan.effort ?? '-'}` + refresh($, st) +} + +function endMainTurn($: Api, st: State): void { + st.turnMain = st.pendingPrompt + st.pendingPrompt = null + st.explicitEffort.clear() + st.lastMain = '' + st.windowWarned = false + refresh($, st) +} + +// ---- registration ---------------------------------------------------- + +function registerSession(on: On, st: State): void { + on('session.start', async ($, e, next) => { + await reloadConfig($, st) + await registerCommand($) + refresh($, st) + return next(e) + }).catch(($, e, next) => next(e)) + on('session.end', async ($, e, next) => { + Object.assign(st, newState(st.cfg, st.source)) + return next(e) + }).catch(($, e, next) => next(e)) + on('command.run', { command: 'route' }, async ($, e) => ({ + text: await handleCommand($, st, e.args), + })).catch(($, e, next) => ({ text: `route failed (${next.error.kind})` })) +} + +function registerRouteTool(on: On, st: State): void { + on('tool.call', { tool: TOOL }, async ($, e) => { + const out = handleRouteTool(st, e) + refresh($, st) + vlog($, st, `route ${e.agentId ?? 'main'}: ${JSON.stringify(out)}`) + return out + }).catch(($, e, next) => ({ + result: `route failed (${next.error.kind}); nothing routed`, + })) +} + +function registerSkills(on: On, st: State): void { + on('tool.call', { tool: 'Skill' }, async ($, e, next) => { + if (st.off) return next(e) + const level = EFFORT_SKILL.exec(e.skill)?.[1] + if (isLevel(level)) return effortBridge(st, e.agentId, e.skill, level) + st.skillCalls += 1 + try { + onSkillLoad(st, e.skill, e.agentId) + return await next(e) + } finally { + st.skillCalls -= 1 + } + }).catch(($, e, next) => next(e)) + on('skill.prompt', async ($, e, next) => { + if (st.off || st.skillCalls > 0) return next(e) + return slashEffort(st, e.skill, e.text) ?? next(e) + }).catch(($, e, next) => next(e)) +} + +function registerAgents(on: On, st: State): void { + on('tool.call', { tool: 'Agent' }, async ($, e, next) => { + if (!st.off && isLevel(e.effort) && typeof e.tool_use_id === 'string') { + st.explicitEffort.set(e.tool_use_id, e.effort) + } + return next(e) + }).catch(($, e, next) => next(e)) + on('agent.spawn', async ($, e, next) => { + if (st.off) return next(e) + const frozen = e.fork || e.workflow !== undefined + const route = spawnRoute(st.cfg, e, frozen) + // An explicit model param on the Agent call always wins. + const wanted = e.model === undefined ? route?.model : undefined + const started = await next(wanted === undefined + ? e + : { ...e, model: resolveModel(st.cfg, wanted) }) + if (started.deny !== undefined) return started + trackLoop(st, e, started, route, frozen) + vlog($, st, `spawn ${e.subagentType}: ${e.model ?? '-'} → ${started.model}`) + return started + }).catch(($, e, next) => next(e)) +} + +function registerTurns(on: On, st: State): void { + on('turn.step', async function* ($, e, next) { + if (st.off) return yield* next(e) + const plan = await planStep($, st, e) + const changed = plan.model !== e.model || plan.effort !== e.effort + if (e.agentId === undefined) noteMain($, st, plan) + vlog($, st, stepLog(e, plan)) + const result = yield* next(changed ? withPlan(e, plan) : e) + vlog($, st, `step ${e.index} answered by ${result.usage?.model ?? '?'}`) + return result + }).catch(async function* ($, e, next) { + return yield* next(e) + }) + on('turn.complete', async ($, e, next) => { + if (e.agentId !== undefined) st.loops.delete(e.agentId) + else endMainTurn($, st) + return next(e) + }).catch(($, e, next) => next(e)) +} + +function registerPrompt(on: On, st: State): void { + on('prompt.submit', async ($, e, next) => { + if (st.off || e.origin.kind !== 'composer') return next(e) + const rule = st.rules.find(r => r.re.test(e.text)) + const route = rule ? phaseRoute(st.cfg, rule.phase) : undefined + if (rule && route) { + const routed: Routed = { phase: rule.phase, route, source: 'prompt' } + // Typed mid-turn and asked to wait: it belongs to the NEXT turn. + if (e.turnId !== undefined && e.wait) st.pendingPrompt = routed + else st.turnMain = routed + refresh($, st) + } + return next(e) + }).catch(($, e, next) => next(e)) + on('ui.render', { component: 'Spinner' }, async ($, e, next) => { + if (st.off || !st.cfg.spinner || !st.lastMain) return next(e) + const suffix = ` · ${st.lastMain}…` + return next({ ...e, props: { ...e.props, suffix } }) + }).catch(($, e, next) => next(e)) +} + +export const register: Register = on => { + const st = newState(mergeConfig(undefined, () => undefined), 'defaults') + registerSession(on, st) + registerRouteTool(on, st) + registerSkills(on, st) + registerAgents(on, st) + registerTurns(on, st) + registerPrompt(on, st) +} From e8ca713d9e16ea9cb183a0fe2cc0c956a9bee9e5 Mon Sep 17 00:00:00 2001 From: bchanot Date: Thu, 8 Oct 2026 16:26:57 +0200 Subject: [PATCH 03/22] chore(tasks): model-router w1a contract + plan r3 --- .../2026-10-08-model-router-w1a-1533.md | 47 +++ .../plans/2026-10-08-model-router-w1a-1533.md | 347 ++++++++++++++++++ 2 files changed, 394 insertions(+) create mode 100644 .claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md create mode 100644 .claude/tasks/plans/2026-10-08-model-router-w1a-1533.md diff --git a/.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md b/.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md new file mode 100644 index 0000000..9c07039 --- /dev/null +++ b/.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md @@ -0,0 +1,47 @@ +# CONTRACT — model-router-w1a (wave 1-A: the mod itself) +- date: 2026-10-08 | flow: feat | branch: feature/model-router-mod +- status: active + +## REQUEST (verbatim — IMMUTABLE) +Skill args: "model-router mod, wave 1-A: the mod itself under mods/model-router/ (plugin.json, hooks.json, register.ts, config.json alias→id + phases/agents/skills/prompt tables, register.test.ts); spike code in ~/.claude/dev-mods/385f7190-70f5-4bdd-b0d8-e4566cd412fd/model-router/ is the base; plan .claude/tasks/plans/2026-10-08-model-router-mod.md W1.1-W1.9" +User (fr, same session): "go pour le registre et go sur la vague 1". Earlier framing (verbatim excerpts): "repartir correctement chaque tache au model qui lui correspond […] Il faut que l'effort aussi soit en consequence […] plus propre, plus unifier et plus automatique (meme dans la discussion courante ou d'un agent on puisse switch d'un model / effort a un autre. Et le mieux que ca soit configurable et qu'on puisse l'installer et qu'il soit actif sur toutes les session en userscope"; "Il faut un pin pour le global, mais toutes les sous taches fait pas le routage donne au model correspondant"; "si la route modifie deja les efforts, alors les skills pour changer les efforts devienne inutile mais vont quand meme etre trigger. Ca fait doublon, des token pour rien used, et peut etre meme des conflits non ?" + +## CLARIFICATIONS +Q: config.json as a 5th file? / A: no — defaults live in register.ts (`DEFAULT_CONFIG`), the optional user override is `~/.claude/model-router.json` (deep-merged); `claude plugin test` runs without fs, so the mod must work with no file at all. 4 files. [orchestrator — internal, derived from the test sandbox] +Q: userConfig (W1.8) / A: dropped for 1-A — `mainModelSwitch`, `verbose`, `spinner` are keys of the same config (one source), toggled live by `/route`. [orchestrator — internal] +Q: model ids / A: hooks always write FULL ids from `config.models` (alias → id); the Agent tool param is never rewritten (its schema accepts aliases only, and the hook-side alias resolver is stale, LRN-203 / BLK-029). [orchestrator — in-force learning] +Q: precedence / A: user `/route` (sticky until `/route clear`) > the latest turn-scoped route on main (model `route` tool, a skill load's table phase or a `Skill(effort-*)` shift, a typed `/effort-*`, a prompt rule: one slot, last writer wins; a non-effort skill load resets the slot except a prompt rule) > session settings. Explicit Agent-call `model`/`effort` params always win for that agent, for its whole run: an in-agent `route` call or `Skill(effort-*)` never touches an axis given explicitly. [orchestrator — derived from the user's "pin = entry default, sub-tasks route finer"; r3 after the confirmation challenge] +Q: verbose default / A: `verbose: false` in DEFAULT_CONFIG; for now the user wants it ON to watch the routing → after the build the orchestrator writes `~/.claude/model-router.json` with `{"verbose": true}` (user-home file, outside FILE SCOPE). [gated 2026-10-08] +Q: spinner suffix default / A: on (`spinner: true`). [gated 2026-10-08] +Q: `/route` typed by the user / A: sticky until `/route clear`, wins over model-declared routes. [gated 2026-10-08] +Q: Explore built-in / A: `explore` phase = sonnet / medium (supersedes the BDR-066 wave-3 inherit for Explore). [gated 2026-10-08] +Q: legacy `Skill(effort-*)` / A: answered by the mod without loading the skill (single writer, no pairing rule); `/effort-*` typed by the user → `skill.prompt` sets the same route and returns a one-line text. [user 2026-10-08: "doublon … conflits"] + +## ACCEPTANCE CRITERIA +1. `mods/model-router/` holds exactly `.claude-plugin/plugin.json`, `hooks/hooks.json`, `hooks/register.ts`, `hooks/register.test.ts` (the engine-laid `.claude-plugin/types/` folder and `./tsconfig.json` are ignored, never committed); `claude plugin validate` passes with no warning. + CHECK: cd mods/model-router && [ "$(find . -type f | grep -v '/.claude-plugin/types/' | grep -v '^./tsconfig.json$' | sort | tr '\n' ' ')" = "./.claude-plugin/plugin.json ./hooks/hooks.json ./hooks/register.test.ts ./hooks/register.ts " ] && out=$(claude plugin validate . 2>&1) && echo "$out" | grep -q 'Validation passed' && ! echo "$out" | grep -qi 'warning' && echo FILES-VALIDATE-OK + EXPECT: FILES-VALIDATE-OK + EVIDENCE: MET exit=0 marker-found :: FILES-VALIDATE-OK +2. Type-check clean against this build's declarations (the engine-laid copy beside the spike mod). + CHECK: T=/Users/b.chanot/.claude/dev-mods/385f7190-70f5-4bdd-b0d8-e4566cd412fd/model-router/.claude-plugin/types; W=$(mktemp -d) && printf '{"compilerOptions":{"target":"es2023","lib":["es2023"],"types":[],"module":"esnext","moduleResolution":"bundler","strict":true,"noUncheckedIndexedAccess":true,"noEmit":true,"skipLibCheck":true,"jsx":"react","jsxFactory":"h","jsxFragmentFactory":"Fragment"},"include":["%s/claude-code/index.d.ts","%s/claude-code-tools/index.d.ts","%s/hooks"]}' "$T" "$T" "$PWD/mods/model-router" > "$W/tsconfig.json" && (cd "$W" && npx --yes -p typescript@5 tsc -p tsconfig.json) && echo TSC-OK + EXPECT: TSC-OK + EVIDENCE: MET exit=0 marker-found :: TSC-OK +3. `claude plugin test mods/model-router` passes; the suite covers: (a) `Skill(effort-low)` via `$.tool.call` is answered without `next` in the Skill tool's output shape (`success`, `commandName`) and the route shows `low` on main; (b) the `route` tool with `phase: "orchestrate"` sets `medium` on main and `/route show` prints it; (c) `/route clear` drops it; (d) `/route bogus` returns an error text naming the phases; (e) a `prompt.submit` text holding `ultrathink` sets `escalate` on main; (f) `/route model=sonnet` shows `claude-sonnet-5-5`, a full id passes through, a misspelt alias is refused; (g) `agent.spawn` of `Explore` without a model param reaches the bottom with `model === 'claude-sonnet-5-5'`, and with `model: 'opus'` given the param is untouched. + CHECK: cd mods/model-router && out=$(claude plugin test . 2>&1); rc=$?; echo "$out" | tail -n 5; [ $rc -eq 0 ] && [ "$(grep -cE '^\s*test\(' hooks/register.test.ts)" -ge 7 ] && echo PLUGIN-TEST-OK + EXPECT: PLUGIN-TEST-OK + EVIDENCE: MET exit=0 marker-found :: (pass) a tabled agent steps at its table effort [19.75ms] 11 pass 0 fail Ran 11 tests across 1 file. [0.47s] PLUGIN-TEST-OK +4. Hooks present, as `claude plugin validate` lists them: `session.start`, `command.run{command=route}`, `tool.call{tool=mcp__model-router__route}`, `tool.call{tool=Skill}`, `tool.call{tool=Agent}`, `skill.prompt`, `agent.spawn`, `turn.step`, `prompt.submit`, `turn.complete`, `ui.render{component=Spinner}`; every gating hook carries a fail-open `.catch` (validate prints no "gating hook without .catch"). + CHECK: cd mods/model-router && out=$(claude plugin validate . 2>&1) && for h in session.start 'command.run{command=route}' 'tool.call{tool=mcp__model-router__route}' 'tool.call{tool=Skill}' 'tool.call{tool=Agent}' skill.prompt agent.spawn turn.step prompt.submit turn.complete 'ui.render{component=Spinner}'; do echo "$out" | grep -qF -- "$h" || { echo "missing $h"; exit 1; }; done && ! echo "$out" | grep -q 'without .catch' && echo HOOKS-OK + EXPECT: HOOKS-OK + EVIDENCE: MET exit=0 marker-found :: HOOKS-OK +5. Code style: no line over 80 chars, no `any` type, no `import()`; `register.ts` imports only from `claude-code` and its own plugin files. + CHECK: cd mods/model-router/hooks && ! grep -nE '.{81,}' register.ts register.test.ts && ! grep -nE ':\s*any\b||as any\b' register.ts && ! grep -q 'import(' register.ts && [ "$(grep -cE "^import .* from '(claude-code|\./)" register.ts)" -eq "$(grep -c '^import ' register.ts)" ] && echo STYLE-OK + EXPECT: STYLE-OK + EVIDENCE: MET exit=0 marker-found :: STYLE-OK +6. Judged by reading: the mod never writes a model alias into a request (every `model` it sets at `agent.spawn` or `turn.step` passes through `resolveModel`); the Agent tool's `model`/`effort` params are never rewritten and an explicit `model` param is never overridden at spawn or at any step; an agent's model is written once, at spawn (a per-step model rewrite happens only after an in-agent `route` call and only while `e.model` still equals the spawn model); a fork (`e.fork`) and a workflow agent (`e.workflow`) are never re-modelled; every write from a sub-agent's Skill or `route` call lands on that agent's loop, never on main; main-loop model changes happen only when `mainModelSwitch` is true; a `.catch` on every gating hook fails open (pass-through or an in-place answer) so a mod failure never blocks a call; all state lives in the `register` closure and the defaults constant is never mutated; no function over 25 logic lines; the spike's `via` / `stepModel` / `agentsDefault` levers and the tool's `model` / `scope` params are gone. +Q (r2): `scope: agents` / `/route agents` / A: dropped (challenge r1, no requirement behind it; per-call Agent params and the table cover it). The route tool takes `phase`, `effort`, `clear` only. [orchestrator — simplicity] +Q (r2): agents table in wave 1 / A: built-ins only (Explore, Plan), matched for the engine provider; repo agents keep their frontmatter as the single writer until wave 2. [orchestrator — single source of truth] +Q (r2): `/route off` / A: added as the session kill switch (every hook passes through). [orchestrator — robustness] + +## FILE SCOPE +mods/model-router/.claude-plugin/plugin.json · mods/model-router/hooks/hooks.json · mods/model-router/hooks/register.ts · mods/model-router/hooks/register.test.ts diff --git a/.claude/tasks/plans/2026-10-08-model-router-w1a-1533.md b/.claude/tasks/plans/2026-10-08-model-router-w1a-1533.md new file mode 100644 index 0000000..dd88350 --- /dev/null +++ b/.claude/tasks/plans/2026-10-08-model-router-w1a-1533.md @@ -0,0 +1,347 @@ +# PLAN — model-router wave 1-A: the mod (dispatch-ready) — REVISED r3 + +r3 changes (confirmation pass, 4 MAJOR + minors): explicit Agent params are +frozen on the Loop and never overridden by in-agent `route`/`Skill(effort-*)`; +`pendingPrompt` honours `e.wait`; tests assert on the `main:` line of `show` +only, with full typed inputs; haiku gets `effort: undefined`; `userMain` +beats every turn route including a typed `/effort-` (the text says so); +`skill.prompt` acts only when no Skill call is in flight; route tool handles +`clear` first and answers "off" in place; `turn.step` catch is a generator; +windows keyed by resolved id; validate user entries BEFORE merging; truthful +answers; unpinned-skill reset accepted and documented. +Contract: .claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md +Wave plan + harness facts: .claude/tasks/plans/2026-10-08-model-router-mod.md +Base: the spike `~/.claude/dev-mods/385f7190-70f5-4bdd-b0d8-e4566cd412fd/model-router/hooks/register.ts` +(read it first; keep its proven hook shapes; drop its spike levers `via`, +`stepModel`, `agentsDefault`, the tool's `model`/`scope`/`mainModelSwitch` +params and the hardcoded PHASES). +API reference: `/.claude-plugin/types/claude-code/index.d.ts` (grep the +event or noun; `declare module 'claude-code/testing'` for the test kit) and +`/.claude-plugin/types/claude-code-tools/index.d.ts` (`Skill: {`, +`Agent: {`, and the Skill RESULT schema near line 5139). + +r2 changes (challenge round, 3 lenses): agents table = built-ins only; one +writer per agent model (spawn), no per-step model rewrite unless an in-agent +`route` call changed it and the engine did not fall back; every write goes +to the CALLING loop; `agentsNext` / `scope` / `/route agents` / tool `model` +dropped; Skill bridge answers in the tool's output shape; unconditional +skill-load reset (prompt route kept); config validated at load, refs +resolved once; state in the `register` closure, cloned defaults; `/route off` +kill switch; queued `ultrathink` promoted to its own turn; AC1/AC5 amended. + +## Files +- [ ] mods/model-router/.claude-plugin/plugin.json — `{ "name": "model-router", "version": "0.1.0", "description": "", "author": { "name": "bchanot" } }` +- [ ] mods/model-router/hooks/hooks.json — `{ "modules": ["./register.ts"] }` +- [ ] mods/model-router/hooks/register.ts — the hooks module (below) +- [ ] mods/model-router/hooks/register.test.ts — `claude plugin test` suite (below) +The engine lays `./tsconfig.json` and `.claude-plugin/types/` beside a loaded +mod; both are ignored by AC1 and gitignored in wave 1-B. Never commit them. + +## Config (one shape, defaults in code, optional override on disk) +```ts +type Level = 'low' | 'medium' | 'high' | 'xhigh' | 'max' +type Route = { model?: string; effort?: Level } // model = alias OR full id +type Config = { + models: Record // alias → full id + windows: Record // alias → context window (tokens) + phases: Record + agents: Record // built-in subagentType → phase name + skills: Record // skill name → phase name + prompt: { pattern: string; phase: string }[] // regex source, flag i + mainModelSwitch: boolean; verbose: boolean; spinner: boolean +} +``` +DEFAULT_CONFIG values: +- models: haiku→`claude-haiku-4-5-20251001`, sonnet→`claude-sonnet-5-5`, + opus→`claude-opus-5-5`, fable→`claude-fable-5-1`. +- windows: `claude-haiku-4-5-20251001`→200000 (keyed by FULL id; others + unknown: absent = no check). +- phases: plan {effort xhigh}, reflect {effort high}, orchestrate {effort medium}, + escalate {effort max}, judge {opus, xhigh}, implement {sonnet, medium}, + write {sonnet, medium}, verify {sonnet, xhigh}, explore {sonnet, medium}, + mechanical {haiku, low}. A phase without `model` keeps the loop's model. +- agents: Explore→explore, Plan→judge. NOTHING else in wave 1: every repo + agent keeps its frontmatter pin (the engine applies it); the pins move into + this table in wave 2, in the same change that deletes the frontmatter. +- skills: {} (wave 2 fills it). +- prompt: [{ pattern: '\\bultrathink\\b', phase: 'escalate' }]. +- mainModelSwitch false, verbose false, spinner true. (Pass B, user + 2026-10-08: verbose default off but ON for now through the override file, + spinner on, user `/route` sticky, Explore → sonnet/medium.) + +`loadConfig($, log)` → `Config` (never throws): +1. `home = await $.env.get('HOME')`; path `${home}/.claude/model-router.json`; + `$.fs.exists` then `$.fs.read`; `JSON.parse`. Any failure → the defaults. +2. `mergeConfig(D, u, log)`: FIXED merge, no recursion, VALIDATE EACH USER + ENTRY BEFORE IT REPLACES A DEFAULT (an invalid user `models.sonnet` is + dropped and the default kept, so the phases on `sonnet` stay valid): + for each table (`models`, `windows`, `phases`, `agents`, `skills`) take + `u.` only when it is a plain object, then per key: valid → over + the default, invalid → `log(...)` and keep the default. Scalars + (`mainModelSwitch`, `verbose`, `spinner`) taken only when boolean. + `prompt` taken only when an array; each rule validated or dropped. + Type guards on `unknown`, no `any`. +3. Validity rules (`log(...)` ALWAYS, not only verbose: a config error must + be seen once): + - models value: string matching `/^claude-[a-z0-9.-]+$/`; + - windows: key a full id (same regex), value a positive integer; + - phase: plain object; `effort` absent or in LEVELS; `model` absent, a + `models` key or a full id; at least one of the two; + - agents / skills value: a phase name (checked after phases merged); + - prompt rule: `{ pattern: string, phase: }` whose pattern + compiles (`new RegExp(p, 'i')` in try/catch). + Lookups use `Object.hasOwn`, never bare indexing on user keys. +4. Returns `structuredClone`d data: the defaults constant is never handed + out by reference. +`compileRules(cfg)` → `{ re: RegExp; phase: string }[]` once per load. +`resolveModel(cfg, name)`: `Object.hasOwn(cfg.models, name) ? cfg.models[name] : name`. +`isModelName(cfg, v)`: a `models` key or the full-id regex. `isLevel(v)`. + +## State — ONE object built inside `register`, passed to every helper +```ts +type Source = 'user' | 'model' | 'skill' | 'prompt' | 'slash' +type Routed = { phase: string; route: Route; source: Source } +type Loop = { + effort?: Level; model?: string // routed by the table or an in-agent call + spawnModel: string; frozen: boolean // engine's model at spawn; fork/workflow + explicitModel: boolean; explicitEffort: boolean // Agent params given → axis frozen +} +type State = { + cfg: Config; rules: Rule[] + userMain: Routed | null // /route by the user; sticky until /route clear + turnMain: Routed | null // tool / skill / prompt / slash; dropped at turn end + pendingPrompt: Routed | null // prompt rule typed mid-turn, promoted next turn + loops: Map // agentId → that loop's routing + explicitEffort: Map // Agent tool_use_id → explicit effort param + skillCalls: number // Skill tool calls in flight (hook 4 ± around next) + off: boolean // /route off: every hook passes through + lastMain: string // "model/effort" of the last main step (spinner) +} +``` +`newState(cfg)` builds it; `register` calls it once; `session.start` reloads +`cfg` + `rules` into it; `session.end` rebuilds it (`/clear` fires no +`session.start`, so sticky routes must not survive a clear). +Effective main route `mainRoute(st)`: `st.userMain ?? st.turnMain`. ONE +order, transitive: user `/route` (sticky) > the latest turn route (tool, +skill, slash, prompt all share `turnMain`; last writer wins) > session. A +typed `/effort-` while a sticky route is in force does not apply; its +text says so (hook 5). +Loop lookups: `loopOf(st, e.agentId)`; a missing entry is created on first +write as `{ spawnModel: '', frozen: false, explicitModel: false, +explicitEffort: false }`. In-agent writes (hooks 3 and 4) never set an axis +whose `explicit*` flag is true: an explicit Agent param wins for the whole +run. + +## Hooks +Rule for failures: hooks that only observe or rewrite carry +`.catch(($, e, next) => next(e))`; `turn.step` streams, so its catch is the +generator form `async function* ($, e, next) { return yield* next(e) }` (a +plain function there is a type error). The four hooks that ANSWER without `next` +(command.run, the route tool, the Skill `effort-*` bridge, `skill.prompt`) +carry a `.catch` that answers in place: `{ text: 'route failed ()' }`, +`{ result: 'route failed (); nothing routed' }`, and for the two skill +hooks `next(e)` (the skill then loads normally — a safe fallback). State is +mutated only AFTER the input validated. +1. `session.start`: `st.cfg = await loadConfig(...)`, `st.rules = compileRules`; + `registerTool($, st)` (`$.tool.register({ name: 'route', description, + inputSchema })`: properties `phase` (enum = Object.keys(st.cfg.phases)), + `effort` (enum LEVELS), `clear` (boolean); no `required`). The description + tells the model: declare the phase before a span changes nature; acts on + the calling loop only; no model choice here. `$.command.register({ name: + 'route', description, argumentHint: '[show|clear|off|on|reload|| + model= effort=|switch on|off|verbose on|off]', immediate: + true })` in try/catch (log on failure, keep going). `$.ui.status(statusLine(st))`. +2. `command.run {command:'route'}` → `{ text: handleCommand($, st, e.args) }`: + `show`/empty → `show(st)`; `clear` → userMain = turnMain = pendingPrompt = + null; `off` / `on` → st.off; `reload` → loadConfig + compileRules + + `registerTool` again (the phase enum follows the config) + 'config + reloaded' + show; `switch on|off` → cfg.mainModelSwitch; `verbose on|off`; + otherwise `parseRoute(st.cfg, args)`: a phase name, or tokens `model=` + / `effort=` / bare alias / bare level, each validated by `isModelName` + / `isLevel` → `st.userMain = { phase, route, source: 'user' }`; any + unknown token → error text listing the phases and the levels. Never + calls `next`. +3. `tool.call {tool:'mcp__model-router__route'}` → `handleRouteTool`, in + this order: (i) `st.off` → `{ result: 'model-router is off (/route on to + resume); nothing routed' }`; (ii) `clear` → main: `turnMain = null`; agent: + unset the loop's `effort`/`model` → `{ result: 'route cleared for ' }`; + (iii) validate: `phase` given and not a `phases` key → `{ deny: 'unknown + phase "

"; phases: …' }` (even with a valid `effort`); `effort` given and + not a level → deny naming the levels; neither given → deny. (iv) route = + `{ ...phases[phase], ...(effort ? { effort } : {}) }` (an explicit effort + overrides the phase's). Target = the CALLING loop: main → `turnMain = { + phase: phase ?? 'effort-' + effort, route, source: 'model' }`; agent → + `loop.effort = route.effort` unless `loop.explicitEffort`; `loop.model = + route.model` unless `loop.explicitModel` (applied at step only under the + fallback guard, never on a frozen loop). (v) Answer on the EFFECTIVE + outcome: main with a sticky `userMain` → `'recorded for this turn, + but a sticky /route is in force; it wins until /route + clear'`; otherwise `'routed to : effort , model '`, and when a model is part of + the route on main while `mainModelSwitch` is off, say `model unchanged + (switch off)`. Verbose → log. +4. `tool.call {tool:'Skill'}`: + a. `e.skill` matches `/^effort-(low|medium|high|xhigh|max)$/` → the + CALLING loop: main → `turnMain = { phase: e.skill, route: { ...st.turnMain?.route, effort }, source: 'skill' }` + (effort merged over the current turn route, last loaded wins); agent → + `loopOf(...).effort = level` unless `loop.explicitEffort` (entry created + if missing: an untabled agent's shift must still land). Answer WITHOUT + `next`, in the Skill tool's output shape (claude-code-tools ~5139; a + string result is refused and the skill would load): + `{ result: { success: true, commandName: e.skill, status: 'inline' }, + context: [] }` where `` states the effective outcome: + `'model-router: effort → for this loop from the next request on; the + effort- skill text was not loaded.'`, or when main has a sticky + `userMain`: `'model-router: effort- recorded, but a sticky /route + is in force and wins until /route clear.'`, or when the agent + axis is explicit: `'model-router: this agent was dispatched with an + explicit effort; the shift does not apply.'` + b. Any other skill: `st.skillCalls += 1` before `next(e)`, `-= 1` after + (try/finally). Main → `turnMain = null` when its source is 'model', + 'skill' or 'slash' (a 'prompt' route such as `ultrathink` stays unless + the skill has a table entry); `Object.hasOwn(cfg.skills, e.skill)` → + `turnMain = { phase, route, source: 'skill' }`. Agent → unset the loop's + non-explicit `effort`/`model`; table entry → `loop.effort = route.effort` + unless explicit (never model). Accepted change vs the legacy shifters: + loading an UNPINNED skill after a shift returns main to the harness + level (orchestrators already re-assert after a nested skill, + lib/effort-shift.md § Re-assert). Then `return next(e)`. +5. `skill.prompt {skill: /^effort-/}`: `st.skillCalls > 0` (reached through + the bridge's fallback or a Skill call) → `next(e)`. Otherwise it is a + user-typed `/effort-` (or a preload, unsupported: treated the same): + level parse; `turnMain = { phase: e.skill, route: { effort }, source: + 'slash' }`; return `{ text: + '\n' + e.text }` (prepend, never + replace: args ride in the text) where `` is `'Effort shifted to + by model-router for this turn.'` or, with a sticky `userMain`, `'Effort + recorded; the sticky /route wins until /route clear.'`. + Unknown suffix → `next(e)`. +6. `tool.call {tool:'Agent'}`: `isLevel(e.effort) && typeof e.tool_use_id === + 'string'` → `st.explicitEffort.set(e.tool_use_id, e.effort)`; always + `return next(e)` unchanged (params are never rewritten). +7. `agent.spawn`: `frozen = e.fork || e.workflow !== undefined`. Route: + `!frozen && e.provider.plugin === 'engine' && Object.hasOwn(cfg.agents, e.subagentType)` + → `cfg.phases[cfg.agents[e.subagentType]]`, else none. Model rewrite ONLY + when route?.model is set AND `e.model === undefined` (an explicit param + wins): `next({ ...e, model: resolveModel(cfg, route.model) })`, else + `next(e)`. On a non-deny result with `agentId`: `loops.set(agentId, { + spawnModel: result.model, frozen, explicitModel: e.model !== undefined, + explicitEffort: given !== undefined, effort: given ? undefined : route?.effort })` + where `given = explicitEffort.get(e.tool_use_id)` (then deleted). `model` + is NOT stored at spawn: the engine already runs the agent on it. Verbose + log `spawn : → `. Known limit, + documented in a comment: `provider.plugin === 'engine'` is the best + available test for a built-in at spawn; a user agent named `Explore` in a + foreign project would also match (wave 1 impact: sonnet/medium on it). +8. `turn.step` (async generator). `st.off` → log when verbose, `yield* next(e)`. + Agent loop (`e.agentId`): `loop = loops.get(...)`; `effort = loop?.effort ?? + e.effort`; `model = loop?.model && !loop.frozen && e.model === loop.spawnModel + ? resolveModel(cfg, loop.model) : e.model` (an engine fallback — `e.model` + differs from the spawn model — is never fought). Main: `set = mainRoute(st)`; + `effort = set?.route.effort ?? e.effort`; `model = set?.route.model && + cfg.mainModelSwitch && windowOk ? resolved : e.model`, where `resolved = + resolveModel(cfg, set.route.model)` and `windowOk` = no + `cfg.windows[resolved]` entry (windows are keyed by FULL id; the defaults + key haiku's full id) or `(await $.session.usage()).context.tokens` is a + number below it; an absent `tokens` or a failed `usage()` → no switch, + logged once per turn. If the model actually sent starts with + `claude-haiku`, send `effort: undefined` (omit it entirely; haiku takes + none and a hook-set effort on it is unproven). Main → `lastMain = + '/'`, `$.ui.status(statusLine(st))`. Verbose → + log before (`step : → `) and after (`answered by + `). `const r = yield* next(changed ? { ...e, model, effort } : e); return r`. +9. `prompt.submit`: `e.origin.kind !== 'composer'` → `next(e)`. First rule in + `st.rules` whose `re.test(e.text)` → `routed = { phase, route, source: 'prompt' }`; + `e.turnId !== undefined && e.wait` (typed mid-turn and asked to wait: it + belongs to the NEXT turn) → `st.pendingPrompt = routed`; otherwise + (idle, or delivered INTO the running turn) → `st.turnMain = routed`. + `return next(e)`. +10. `turn.complete`: `e.agentId` → `loops.delete(e.agentId)`. Main → + `turnMain = pendingPrompt; pendingPrompt = null; explicitEffort.clear(); + lastMain = ''`; `$.ui.status(statusLine(st))`. `return next(e)`. +11. `ui.render {component:'Spinner'}`: `cfg.spinner && lastMain` → + `next({ ...e, props: { ...e.props, suffix: ' · ' + lastMain + '…' } })` else `next(e)`. +12. `session.end`: `Object.assign(st, newState(st.cfg))` (keeps the loaded + config, drops every route and map). `return next(e)`. +`statusLine(st)`: `'route: ' + (st.off ? 'off' : describe(mainRoute(st)) )` +where `describe` = `' '` or `'session defaults'`, plus +`' · switch on'` when `cfg.mainModelSwitch`. +`show(st)`: main (effective, with its source), off/on, switch, verbose, +spinner, live loops count, phases as `name=/` +(resolved ids printed, so a wrong `models` entry is visible), config source +line (`defaults` or the override path). + +## Tests (register.test.ts, `import { test, expect } from 'claude-code/testing'`) +Read the kit's declarations first (`declare module 'claude-code/testing'`): +`test(name, async ($, on) => …)`; events are fired as calls on `$` with the +event's FULL input (the kit's `$` is `EngineCall = (e: Args)`, and +AC2 type-checks the test file): `$.command.run({ command: 'route', args: +'show', origin: { kind: 'composer' }, presentation: })`, `$.prompt.submit({ text: 'ultrathink please', wait: false, +origin: { kind: 'composer' } })`, `$.agent.spawn({ tool_use_id: 't1', +prompt: 'x', description: 'd', subagentType: 'Explore', provider: { plugin: +'engine', tier: 'core' }, parentModel: 'claude-fable-5-1', background: false, +fork: false })`, `$.tool.call({ tool: 'Skill', skill: 'effort-low' })`. Read +each input type and fill every required field; never relax a test to dodge +a type. Establish from the kit whether `session.start` fires at load; if +not, fire `$.session.start(...)` first in every test. An event whose hook +calls `next` needs a BOTTOM hook registered by the test through its `on` +(the kit's bottom throws otherwise), e.g. `on('prompt.submit', ($, e) => ({ +text: e.text }))`, `on('agent.spawn', ($, e) => ({ model: e.model, agentId: +'a1' }))`. A helper `mainLine(text)` returns the `main:` line of `show`; +EVERY assertion on a route reads that line only (the phases listing always +contains every id and level, so matching the whole text proves nothing). +Tests (one per contract item 3a-3g): +- 3a `Skill(effort-low)` via `$.tool.call`: resolves with `result.success === + true` and `result.commandName === 'effort-low'`; a bottom `on('tool.call', + { tool: 'Skill' })` registered by the test is NOT reached (flag); `mainLine` + contains `skill effort-low` and `low`. +- 3b route tool `{ phase: 'orchestrate' }` → `mainLine` contains `model + orchestrate` and `medium`. +- 3c `/route clear` → `mainLine` contains `session defaults`. +- 3d `/route bogus` → text contains `unknown` and every phase name. +- 3e `$.prompt.submit` with `ultrathink`, `wait: false`, no `turnId` → + `mainLine` contains `prompt escalate`. +- 3f `/route model=sonnet` → `mainLine` contains `claude-sonnet-5-5`; `/route + model=claude-x-9` → `mainLine` contains `claude-x-9`; `/route model=sonet` → + text contains `unknown`. +- 3g spawn path: `$.agent.spawn(...)` for `Explore` without `model`, bottom + hook captures `e.model === 'claude-sonnet-5-5'`; with `model: 'opus'` given + → captured `e.model === 'opus'`. +No fs, network or process in tests: the defaults path is the one exercised. + +## Edge cases +- `$.command.register` throws when `/route` is taken → log, keep the tool. +- `loadConfig` never throws out of `session.start`; invalid entries dropped with a log. +- `/clear` → `session.end` rebuilds the state; `/route reload` re-registers the tool. +- Remote agents raise no `turn.complete`; denied Agent calls never spawn: both + maps are bounded by `explicitEffort.clear()` at main turn end and `loops` + entries only for started agents (a leak of a few entries per turn is accepted). +- `loops.delete` at an agent's `turn.complete`: a resumed agent (SendMessage, + woken teammate) runs its later turns at the engine's effort. Accepted in + wave 1 (built-ins only); revisit with the pins in wave 2. +- Spawn-vs-first-step race: the loop entry is set after `next(e)` resolves, + so step 0 of a tabled built-in may run at the engine's effort. Accepted in + wave 1 (Explore/Plan only); wave 2 verifies the ordering before pins move. +- Unpinned agents (general-purpose, interviewer, client-handover-writer) no + longer inherit a shifted level: the bridge does not move the harness + level. Accepted: BDR-077 already requires explicit call-site params for + built-ins; the two inline-load agents run on main's own route. +- The word `any` must not appear as a TypeScript type in register.ts + (AC5 greps `: any`, ``, `as any`). + +## Disposition (STEP 0.6) +- honors BDR-066/076/077 (tiers): wave 1 touches only the two built-ins that + carry no pin (Explore sonnet/medium, Plan opus/xhigh); every repo agent + keeps its frontmatter as the single writer; explicit call-site params win. +- honors BDR-107/108: levels and aliases unchanged; aliases stay the config's + vocabulary, full ids are resolved by the mod (LRN-203, BLK-029). +- LRN-180/181 made moot: the bridge answers `Skill(effort-*)` itself, no + pairing rule. "Last loaded wins" holds among pinned skills and shifts; the + one behaviour change, accepted: loading an UNPINNED skill after a shift + returns main to the harness level (orchestrators re-assert after nested + skills already, lib/effort-shift.md § Re-assert). +- LRN-204: main-loop model switch behind `mainModelSwitch` (default false) and + a context-window guard. +- BDR-044 not contradicted: the mod routes model/effort, never skills. +- Deferred (minor, challenge r1): a ceiling on model-declared efforts; the + spawn-vs-step race. From 6dc2d748fccffa667c612dca63633958c02a233d Mon Sep 17 00:00:00 2001 From: bchanot Date: Thu, 8 Oct 2026 16:27:43 +0200 Subject: [PATCH 04/22] =?UTF-8?q?chore(memory):=20journal=20+=20TODO=20?= =?UTF-8?q?=E2=80=94=20model-router=20wave=201-A=20done,=20hardening=20+?= =?UTF-8?q?=201-B=20queued?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- .claude/memory/journal.md | 1 + .claude/tasks/TODO.md | 4 +++- 2 files changed, 4 insertions(+), 1 deletion(-) diff --git a/.claude/memory/journal.md b/.claude/memory/journal.md index 739c1be..5e77116 100644 --- a/.claude/memory/journal.md +++ b/.claude/memory/journal.md @@ -579,3 +579,4 @@ rules: ## 2026-10-08 - model-router mod, wave 0 spike (user ask: one mod routes model + effort per request, replaces effort-* shifters + pins). Plan `.claude/tasks/plans/2026-10-08-model-router-mod.md`, 4 decisions by AskUserQuestion (spike-first main-loop switch, `CLAUDE_CODE_PLUGIN_DIRS` load, migration wave 2, names model-router / route / /route), rule "pin = entry default, sub-tasks route finer". Spike in dev-mods, hot reload on: `turn.step` effort rewrite proven (transcript `effort` field is the oracle, not `CLAUDE_EFFORT`); sub-agent model at `agent.spawn` + effort per step by agentId proven; main-loop fable → sonnet-5-5/low for 3 steps then back: works, one cold-cache step per switch INTO a model, return free. Found: hook-side alias resolver stale (`sonnet` → `claude-sonnet-5`, 404; Agent tool enum resolves the same alias to 5-5) → mod writes full ids only. feature/model-router-mod open, nothing committed yet (plan + TODO + journal pending). +- model-router wave 1-A (/feat, user go): mod built in `mods/model-router/` (4 files, 834 + 202 lines, 11 plugin tests). Plan r1 → r3: 3 challengers (simplicity CONCERNS, robustness CONCERNS(6), correctness FATAL(8)) + 1 confirmation CONCERNS(4); converged on: agents table = built-ins only in wave 1 (frontmatter stays single writer), agent model written once at spawn, explicit Agent params frozen per loop, every sub-agent write on its own loop, Skill bridge answers in the Skill tool's OUTPUT schema (string result refused → skill would load), config validated before merge, state in closure, `/route off`. feater DONE first pass; GATE 0 MET; verifier CONFORME 6/6; security PASS (4 MEDIUM + 5 LOW parked in TODO for user go). Commits b721c94 (mod) + e8ca713 (contract/plan). Override `~/.claude/model-router.json` {verbose:true} written (pass B). Doc-sync deferred to 1-B. diff --git a/.claude/tasks/TODO.md b/.claude/tasks/TODO.md index 6eb8418..3050a16 100644 --- a/.claude/tasks/TODO.md +++ b/.claude/tasks/TODO.md @@ -5,7 +5,9 @@ Plan `.claude/tasks/plans/2026-10-08-model-router-mod.md`. Decisions 2026-10-08: model switch spike-first then flag off; load via `CLAUDE_CODE_PLUGIN_DIRS` + link.sh; migration of shifters/pins/model-gate in wave 2 after proof; names model-router / route / /route. - [x] W0 spike in dev-mods (hot reload): facts a-d established 2026-10-08 (plan file § Spike facts); e moved to W1.10 -- [ ] W1 core mod in `mods/model-router/` (config, route tool, /route, agents, skills, prompt rules, visibility, tests, install) +- [x] W1-A the mod in `mods/model-router/` (b721c94, contract `2026-10-08-model-router-w1a-1533`, plan r3): challenge 3 lenses + 1 confirmation (2 BLOCKER + 10 MAJOR closed by named changes), feater DONE first pass, GATE 0 MET 5/5, verifier CONFORME 6/6, security PASS (4 MEDIUM + 5 LOW reported, below) +- [ ] W1-A hardening (security report, user decision): `/route` command origin check (`composer` only); in-agent `route` must not re-model the agent (strip `route.model` for agent loops, drop the dead per-step model path, align the tool description); ReDoS caps on config patterns (length ≤ 200, test ≤ 4 KB); one log per session in every `.catch` + log the silent config drops; phase key charset `^[a-z][a-z0-9_-]{0,31}$`; `additionalProperties: false` on the tool schema + `typeof e.skill`; post-`next` bookkeeping in try/catch; config file size cap 64 KB. Deferred by design: effort ceiling for model-declared routes; window check beyond haiku. +- [ ] W1-B install + docs: `mods/` symlink + `CLAUDE_CODE_PLUGIN_DIRS` in settings.json env via link.sh, root `.gitignore` for the engine-laid `mods/*/tsconfig.json` + `.claude-plugin/types/`, doctor line, README/USAGE/CHANGELOG (doc-sync deferred here from the 1-A /feat run: nothing to document before the install exists), live test in session (swap the spike for the real mod) - [ ] W2 migration: 15 skills off `Skill(effort-*)`, remove shifters + effort-pins + model-gate, census repointed, docs - [ ] W3 optional: step heuristics, haiku classifier, quota-aware downgrade, A/B From 64702d50eaf46156f10f9ef30011e99e0254c459 Mon Sep 17 00:00:00 2001 From: bchanot Date: Thu, 8 Oct 2026 16:38:08 +0200 Subject: [PATCH 05/22] =?UTF-8?q?chore(memory):=20BDR-115=20+=20LRN-205/20?= =?UTF-8?q?6=20+=20EVAL-040=20=E2=80=94=20model-router=20architecture,=20t?= =?UTF-8?q?ool=20output=20schema,=20plugin=20test=20kit,=20challenge=20val?= =?UTF-8?q?ue?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- .claude/memory/decisions.md | 9 +++++++++ .claude/memory/evals.md | 9 +++++++++ .claude/memory/learnings.md | 8 ++++++++ .../2026-10-08-model-router-w1a-1533.md | 19 +++++++++++++++++++ 4 files changed, 45 insertions(+) diff --git a/.claude/memory/decisions.md b/.claude/memory/decisions.md index 4082f09..1dc23fb 100644 --- a/.claude/memory/decisions.md +++ b/.claude/memory/decisions.md @@ -132,6 +132,7 @@ rules: | BDR-108 | 2026-09-29 | Effort round: level on every skill next to its model pin (3 repo + 25 vendored via `lib/effort-pins.txt` re-applied after the LAST vendoring step of install + resync), design stack ONE level (high), model pins stay tier aliases: quality/price trade-off = tier × effort, never version | accepted | | BDR-109 | 2026-09-30 | Higgsfield pack: npm CLI `latest` + 8 upstream skills git-cloned into gitignored `skills-external/higgsfield-*`, OFF by default, in no profile; two toggles (`higgsfield` = allowlist of 7 media skills, `higgsfield-websites` = landing-page aid, never website create/deploy/publish); CLI presence by probe; routing on explicit ask | accepted | | BDR-110 | 2026-10-06 | Shell portability doctrine: native userland on macOS AND Linux, no Homebrew GNU tools on PATH; `lib/tests/portability-census.test.sh` locks deterministic GNU-only idioms | accepted | +| BDR-115 | 2026-10-08 | model-router mod: pin = entry default, sub-tasks route finer; one writer per axis; full ids from the mod table; built-ins-only agents table until frontmatter pins go; state in closure | accepted | --- @@ -1393,3 +1394,11 @@ Branch feature/user-writing-web-rules, UNMERGED (human gate). - **Alternatives rejected**: regen via `install-hook` (writes a LOCAL hooks-path entry) or `global-hooks` (writes the GLOBAL config when the value is missing — it was, [[LRN-200]]) → `emit-hook > file` only; temp file for the verb's stderr in a hook (fail-open on a full TMPDIR, predictable path) → `2>&1` capture ([[LRN-202]]); classification before the token cap (13 s flood) → cap first ([[LRN-201]]); shell check of the release version string (interpolation sink) → by reading. - **Gates**: D1 3 lenses + confirm FATAL(1) (global-config write) → emit-hook; feater; GATE 0 MET; verifier CONFORME; security PASS. D2 3 lenses + confirm FATAL(3) (escape alternative, same-quote-inside, `bare=$one` when unparsed); feater; GATE 0 MET; verifier CONFORME; security BLOCK(1) cap-after-fork → fixed (20k tokens 0.13 s) → CONFORME + PASS. D3 3 lenses + confirm CONCERNS(3); feater; CONFORME + PASS. - **Refs**: contracts/plans `2026-10-07-manual-push-failclosed-d1-1522`, `…-guard-residuals-d2-1526`, `…-prose-d3-1530`; commits 472cccb, 3c59333, 64ca0f8 (feature/manual-push-mode, UNMERGED). Supersedes the "fail-open on invalid value" line of [[BDR-111]]. Links [[BDR-112]], [[BDR-113]], [[LRN-114]], [[LRN-196]]. Residuals: TODO "post-run-D residuals" (soft_deny names only `false`; stale `.githooks/` in onboarded repos until reconcile). + +## BDR-115 — model-router mod: pin = entry default, sub-tasks route finer; one writer per axis; full ids; built-ins-only table until pins go [accepted] (2026-10-08) +- **Decision**: one function-hooks mod (`mods/model-router/`) routes model + effort per request. Rules: (1) a skill/agent pin = DEFAULT route of the run, never ceiling/floor; inside the run every sub-task routes to its phase, declared (`route` tool, `/route`, `ultrathink`) or derived (skill load, agent dispatch). (2) ONE writer per axis: `Skill(effort-*)` answered by the mod WITHOUT loading the skill (no pairing rule, no doublon); every write lands on the CALLING loop; explicit Agent params frozen per loop for the whole run; agent model written once at spawn. (3) hooks write FULL ids from `config.models` ([[LRN-203]]). (4) wave 1 agents table = built-ins only (Explore sonnet/medium, Plan opus/xhigh); repo agents keep frontmatter as single writer until wave 2 moves pins into the table and deletes frontmatter + shifters + effort-pins + model-gate. (5) precedence: user `/route` sticky > latest turn route (one slot, last writer wins; non-effort skill load resets it except a prompt rule) > session. (6) main-loop model switch behind `mainModelSwitch` (default off) + context-window guard ([[LRN-204]]). (7) state in `register` closure, defaults cloned, `/route off` kill switch, config validated before merge. Load: `CLAUDE_CODE_PLUGIN_DIRS` in settings env via link.sh (wave 1-B). +- **Why**: user 2026-10-08: existing pins + `effort-*` shifters built for this goal with older tools; some pins forced (one level for a skill doing many things). Mods expose `turn.step` model/effort rewrite + `agent.spawn` + `tool.call` answers = cleaner, unified, automatic, works headless too (BDR-107 gap). +- **Alternatives rejected**: agents table copying the 21 frontmatter pins in wave 1 (third source of truth, silent override of a frontmatter edit); `scope: agents` / `/route agents` bulk lever (no requirement, sixth precedence tier); `model` param on the model-facing tool (typo → LRN-203 404 class); per-step agent model rewrite (fights engine fallback, beat explicit params); `userConfig` (one config file instead); marketplace install (live symlink repo model); `source`-dependent skill-load reset (kept stale shifts). +- **Gates**: plan r1 → r3 through 3 challengers + 1 confirmation (2 BLOCKER + 10 MAJOR closed, [[EVAL-040]]); feater DONE first pass; GATE 0 MET; verifier CONFORME 6/6; security PASS + hardening round (criteria 7-11). +- **Refs**: plan `.claude/tasks/plans/2026-10-08-model-router-mod.md`, contract `.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md`, [[BDR-107]], [[BDR-108]], [[BLK-029]], [[LRN-205]], [[LRN-206]]. + diff --git a/.claude/memory/evals.md b/.claude/memory/evals.md index 9ece412..d0078a9 100644 --- a/.claude/memory/evals.md +++ b/.claude/memory/evals.md @@ -60,6 +60,7 @@ rules: | EVAL-037 | 2026-09-28 | correction of EVAL-035/036 counts: transcript records are per content block; deduped by message.id → main-loop thinking share 99.9%, thinking share of weighted cost 5.6%, sonnet think/msg 26→0.2, A/B requests 9→8 | conclusions hold (sharper: main-loop thinking 96.6%→99.9%, weighted-cost thinking corrected 8.4%→5.6%); effort-audit.py dedupes from a3b479e+ | | EVAL-038 | 2026-09-29 | correction of EVAL-037: 94 % of sub-agent usage records carry no `output_tokens_details` (Fable subs at xhigh read 0 thinking, impossible with always-on thinking) → sub-agent thinking UNMEASURED, not ≈0; main loop 100 % counted; weighted-cost split (61/39) still holds | `effort-audit.py` prints coverage + CAVEAT; cite the cost split only; agent effort pins stay unmeasured; a tier move on a price argument = judgment, not figure | | EVAL-039 | 2026-09-30 | ship-feature run higgsfield-pack: plan dry-run in scratch → 0 executor failure on 7 tasks; challenge found 7 MAJOR I missed; floor-guard caught 2 shellcheck suppressions of mine; final review found README/code gap | keep | +| EVAL-040 | 2026-10-08 | model-router w1a plan: 3 challengers + 1 confirmation found 2 BLOCKER + 14 MAJOR on a plan judged closed; executor then passed every gate first time | keep the round, never dispatch a mod plan without it | --- @@ -370,3 +371,11 @@ Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itse - **Method**: plan code dry-run in a scratch copy before the gate (suite per stage 0/5→5/0, 6/8→14/0, 14/1→15/0, 15/1→16/0, 4 mutation tests); 3 challengers + 1 confirmation; SDD per-task reviews; GATE 0/1/2 twice; final review on opus. - **Anomaly**: my first plan was green in dry-run and still wrong on 7 MAJOR points (shim vs binary, unbounded toggle probe, denylist membership, vacuous fixtures, askpass prompt): a dry-run proves the code does what I wrote, not that I wrote the right thing. Floor-guard flagged 2 `shellcheck disable=SC2016` I added to keep "shellcheck clean" green. Final review found the README promised drift reporting that the enabled state never reached. doc-syncer patch hit a shape escalation because I filed a script-comment edit under MINOR doc. One oracle of mine was shape-bound ([[LRN-188]]). - **Action**: keep the pre-gate dry-run (0 executor failure, 1 fix round in 7 tasks) AND the challenge (orthogonal finds); never silence a linter to satisfy a criterion, rewrite the line; doc patch plans carry public-doc paths only, script comments go as code commits. + +## EVAL-040 — model-router w1a: the challenge round caught what the author could not see +- **Date**: 2026-10-08 +- **Output checked**: plan `.claude/tasks/plans/2026-10-08-model-router-w1a-1533.md` r1, written after a successful spike with every harness fact in hand. +- **Method**: 3 blind opus challengers (simplicity CONCERNS(3), robustness CONCERNS(6), correctness FATAL(8)) + 1 confirmation (CONCERNS(4)); every BLOCKER/MAJOR closed by a named plan change (r2, r3); then feater, GATE 0, verifier, security. +- **Anomaly**: r1 carried 2 BLOCKER (explicit Agent `model` overridden at every step; Skill bridge answer shape refused by the output schema → doublon kept) + 10 MAJOR, all invisible to me: spike levers carried over as design (`agentsNext`, per-step model rewrite), a table copying 21 pins = third source of truth, writes on main from sub-agent loops. Confirmation found 4 more MAJOR (explicit params vs in-agent writes, `e.wait`, vacuous test assertions, haiku effort). Executor then DONE first pass, verifier CONFORME 6/6, security PASS: the plan was the whole risk. +- **Action**: a mod plan always goes through the full round + confirmation; test assertions must read the one line that carries the value; spike code is a FACT source, never a design source ([[BDR-115]], [[LRN-205]], [[LRN-206]]). + diff --git a/.claude/memory/learnings.md b/.claude/memory/learnings.md index c46d9a0..91110c6 100644 --- a/.claude/memory/learnings.md +++ b/.claude/memory/learnings.md @@ -1757,3 +1757,11 @@ Rule: when editing a doctrine file under structure locks, grep the test's lock s - **Context**: spike 2026-10-08, fable → `claude-sonnet-5-5`/low for 3 steps on ~260k context, then back. Conversation intact (tools, results, thinking blocks from another model in history: no error). First sonnet step cache_read 0 (full 260k billed), next steps 237k cached; return to fable step read 263k cached. - **Apply**: switch the main loop only for spans long enough to amortize one uncached read of the whole context (many mechanical steps), never per tool call; short mechanical work → small-context haiku sub-agent. Haiku 4.5 window 200k: a long main loop cannot go to haiku at all. Flag off by default in model-router. Links [[LRN-203]], [[BDR-107]]. +## LRN-205 — A hook answering `tool.call` in place of a built-in tool must return that tool's OUTPUT schema shape; a string result is refused and the tool runs anyway +- **Context**: model-router Skill bridge, plan r1: `{ result: '' }` for `Skill(effort-*)`. Challenger: Skill has output schema `{ success, commandName, status?, … }` (claude-code-tools index.d.ts ~5139); core validates a hook's answer against it (claude-code index.d.ts ~12641), wrong shape = hook skipped = skill loads = the exact doublon the bridge exists to remove. Fix: `{ result: { success: true, commandName: e.skill, status: 'inline' }, context: ['…'] }`; text for the model goes in `context`, never in `result`. +- **Apply**: before answering any `tool.call` without `next`, grep the tool's RESULT type in claude-code-tools and mirror it; put model-facing prose in `context`. Links [[BDR-115]], [[LRN-203]]. + +## LRN-206 — `claude plugin test` kit facts (2.1.294): nothing fires at load, inputs are the FULL event, a bottom hook is mandatory under every `next`, `turn.step` streams +- **Context**: model-router tests. Kit `$` is `EngineCall = (e: Args)`: `command.run` needs `origin` + `presentation`, `prompt.submit` needs `wait` + `origin`, `agent.spawn` needs `tool_use_id, description, provider, parentModel, background, fork`; the test file is type-checked with the hooks (tsc include). `session.start` does NOT fire at load → every test boots with a bottom `on('session.start')` + `$.session.start({ cwd, surface: null, isInteractive: false })`. A hook calling `next` hits the kit's bottom which throws unless the test registered one (`on('agent.spawn', ($, e) => ({ model: e.model, agentId: 'a1' }))`). `$.turn.step` returns a stream: drain with `for await` then await `.result` (awaiting `.result` alone runs no hook). No fs/network/process: defaults path only. Assert on the ONE line that carries the value (a `show()` listing every phase always contains every id and level). +- **Apply**: write the boot helper first, type every input from the declarations, never relax a test to dodge a type. Links [[BDR-115]]. + diff --git a/.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md b/.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md index 9c07039..3221849 100644 --- a/.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md +++ b/.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md @@ -43,5 +43,24 @@ Q (r2): `scope: agents` / `/route agents` / A: dropped (challenge r1, no require Q (r2): agents table in wave 1 / A: built-ins only (Explore, Plan), matched for the engine provider; repo agents keep their frontmatter as the single writer until wave 2. [orchestrator — single source of truth] Q (r2): `/route off` / A: added as the session kill switch (every hook passes through). [orchestrator — robustness] +## HARDENING ROUND (security gate 2026-10-08, user go "oui durcis") — criteria 7-11 +7. `/route` is user-only: `command.run` answers `{ text: 'route: user-only command' }` without acting when `e.origin.kind !== 'composer'`; a test proves it (origin `{ kind: 'plugin', name: 'x' }` or the kit's non-composer origin → text contains `user-only`, state unchanged). + CHECK: cd mods/model-router && grep -q "origin.kind" hooks/register.ts && grep -q "user-only" hooks/register.ts && grep -q "user-only" hooks/register.test.ts && echo ORIGIN-OK + EXPECT: ORIGIN-OK + EVIDENCE: pending +8. An in-agent `route` call or skill table entry never changes that agent's MODEL: the `Loop` type has no `model` axis and no `explicitModel` flag, `turn.step` on an agent loop rewrites `effort` only, the route tool's description says "effort only; the model of a sub-agent is fixed at spawn". A test proves it: a route tool call carrying `agentId: 'a1'` with `phase: 'judge'` followed by a `turn.step` for `a1` leaves `model` as given and sets `effort` to `xhigh`. + CHECK: cd mods/model-router && ! grep -qE "loop\.model|explicitModel|spawnModel" hooks/register.ts && grep -q "fixed at spawn" hooks/register.ts && echo NO-AGENT-MODEL-OK + EXPECT: NO-AGENT-MODEL-OK + EVIDENCE: pending +9. Config hardening: a prompt rule pattern longer than 200 chars is dropped (logged); `re.test` runs on at most the first 4096 chars of the prompt; phase keys must match `^[a-z][a-z0-9_-]{0,31}$` (others dropped, logged); the override file is refused above 65536 bytes (logged, defaults kept); the route tool schema carries `additionalProperties: false`; the Skill hook checks `typeof e.skill === 'string'`. + CHECK: cd mods/model-router && grep -q "additionalProperties: false" hooks/register.ts && grep -qE "\[a-z\]\[a-z0-9_-\]\{0,31\}" hooks/register.ts && grep -qE "4096|4_096" hooks/register.ts && grep -qE "65536|65_536|64 \* 1024" hooks/register.ts && grep -qE "200" hooks/register.ts && grep -q "typeof e.skill === 'string'" hooks/register.ts && echo CONFIG-HARDEN-OK + EXPECT: CONFIG-HARDEN-OK + EVIDENCE: pending +10. Visible fail-open: every `.catch` logs once per session per hook (`$.ui.log('model-router: failed (): routing skipped for this event')`, a `warned: Set` in the state) before passing through or answering; the three silent config drops (non-object top level, wrong-typed table, non-array `prompt`) log a line. + CHECK: cd mods/model-router && [ "$(grep -c '\.catch(' hooks/register.ts)" -ge 12 ] && grep -q "warned" hooks/register.ts && grep -q "routing skipped" hooks/register.ts && echo CATCH-LOG-OK + EXPECT: CATCH-LOG-OK + EVIDENCE: pending +11. Post-`next` bookkeeping (loop tracking and logs after `await next(...)` in `agent.spawn` and the Skill hook) runs inside its own try/catch so a logging failure can never make the `.catch` re-run `next`. Judged by reading, with criteria 1-6 still MET (validate, tsc, tests ≥ 13, style, AC6 minus the removed model axis). + ## FILE SCOPE mods/model-router/.claude-plugin/plugin.json · mods/model-router/hooks/hooks.json · mods/model-router/hooks/register.ts · mods/model-router/hooks/register.test.ts From 346d6aeab2894495c40e9d4d90afbbdd0bb56a6b Mon Sep 17 00:00:00 2001 From: bchanot Date: Thu, 8 Oct 2026 16:53:43 +0200 Subject: [PATCH 06/22] =?UTF-8?q?feat(mods):=20model-router=20hardening=20?= =?UTF-8?q?=E2=80=94=20user-only=20/route,=20effort-only=20agent=20routes,?= =?UTF-8?q?=20config=20caps,=20visible=20fail-open?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Security-gate round on the wave 1-A mod: /route answers only a composer origin; an in-agent route call can no longer change the agent's model (effort only, model fixed at spawn); config patterns capped (200 chars, 4096-char scan), phase keys restricted, override file refused above 64 KB, additionalProperties false on the tool schema; every .catch logs once per session; post-next bookkeeping isolated. 14 plugin tests, verifier 11/11. --- mods/model-router/hooks/register.test.ts | 38 ++++ mods/model-router/hooks/register.ts | 218 ++++++++++++++++------- 2 files changed, 196 insertions(+), 60 deletions(-) diff --git a/mods/model-router/hooks/register.test.ts b/mods/model-router/hooks/register.test.ts index 5c0c634..030970b 100644 --- a/mods/model-router/hooks/register.test.ts +++ b/mods/model-router/hooks/register.test.ts @@ -200,3 +200,41 @@ test('a tabled agent steps at its table effort', async ($, on) => { await runStep($, stepInput('a1')) expect(seen).toEqual([{ model: 'claude-fable-5-1', effort: 'medium' }]) }) + +test('/route from a non-composer origin is refused, state kept', async ( + $, on) => { + await boot($, on) + await route($, 'judge') + const out = await $.command.run({ + command: 'route', + args: 'mechanical', + origin: { kind: 'plugin', name: 'x' }, + presentation: { isFullscreen: false, columns: 80 }, + }) + expect(out.text).toContain('user-only') + expect(mainLine(await route($, 'show'))).toContain('user judge') +}) + +test('an in-agent route sets effort only, the model stays', async ($, on) => { + const seen: Seen[] = [] + recordSteps(on, seen) + await boot($, on) + await $.tool.call({ + tool: 'mcp__model-router__route', + phase: 'judge', + agentId: 'a1', + }) + await runStep($, { ...stepInput('a1'), model: 'claude-sonnet-5-5' }) + expect(seen).toEqual([{ model: 'claude-sonnet-5-5', effort: 'xhigh' }]) +}) + +test('a rule only scans the first 4096 chars of a prompt', async ($, on) => { + on('prompt.submit', ($, e) => ({ text: e.text })) + await boot($, on) + await $.prompt.submit({ + text: 'x'.repeat(5000) + ' ultrathink', + wait: false, + origin: { kind: 'composer' }, + }) + expect(mainLine(await route($, 'show'))).toContain('session defaults') +}) diff --git a/mods/model-router/hooks/register.ts b/mods/model-router/hooks/register.ts index 9bf4ae4..fa00a06 100644 --- a/mods/model-router/hooks/register.ts +++ b/mods/model-router/hooks/register.ts @@ -24,11 +24,7 @@ type Rule = { re: RegExp; phase: string } type Source = 'user' | 'model' | 'skill' | 'prompt' | 'slash' type Routed = { phase: string; route: Route; source: Source } type Loop = { - effort?: Level - model?: string // routed by the table or an in-agent call - spawnModel: string // the engine's model at spawn - frozen: boolean // fork or workflow agent: never re-modelled - explicitModel: boolean // Agent call gave a model: axis frozen + effort?: Level // an agent's model is fixed at spawn: effort is its only axis explicitEffort: boolean // Agent call gave an effort: axis frozen } type State = { @@ -44,6 +40,7 @@ type State = { off: boolean // /route off: every hook passes through lastMain: string // "model/effort" of the last main step (spinner) windowWarned: boolean // context-window warning already logged this turn + warned: Set // hooks whose fail-open was already logged } type Log = (text: string) => void type StepIn = Readonly @@ -62,6 +59,10 @@ const TOOL = 'mcp__model-router__route' const EFFORT_SKILL = /^effort-(low|medium|high|xhigh|max)$/ const OVERRIDE = '.claude/model-router.json' const HAIKU = 'claude-haiku' +const MAX_PATTERN = 200 // chars of a prompt-rule pattern +const MAX_PROMPT_SCAN = 4096 // chars of a prompt a rule is run against +const MAX_CONFIG_BYTES = 65536 // override file size +const PHASE_KEY = /^[a-z][a-z0-9_-]{0,31}$/ const DEFAULT_CONFIG: Config = { models: { @@ -123,7 +124,12 @@ function mergeTable( log: Log, ): Record { const out = { ...base } - if (!isRecord(user)) return out + if (!isRecord(user)) { + if (user !== undefined) { + log(`model-router: config ${name} ignored: not an object`) + } + return out + } for (const [key, value] of Object.entries(user)) { const ok = key === '__proto__' ? undefined : accept(key, value) if (ok === undefined) log(`model-router: config ${name}.${key} ignored`) @@ -141,8 +147,8 @@ const acceptWindow = (key: string, v: unknown): number | undefined => : undefined function acceptPhase(models: Record) { - return (_key: string, v: unknown): Route | undefined => { - if (!isRecord(v)) return undefined + return (key: string, v: unknown): Route | undefined => { + if (!PHASE_KEY.test(key) || !isRecord(v)) return undefined const route: Route = {} if (v.effort !== undefined) { if (!isLevel(v.effort)) return undefined @@ -158,7 +164,9 @@ function acceptPhase(models: Record) { function acceptPhaseRef(phases: Record) { return (_key: string, v: unknown): string | undefined => - typeof v === 'string' && hasKey(phases, v) ? v : undefined + typeof v === 'string' && PHASE_KEY.test(v) && hasKey(phases, v) + ? v + : undefined } function compiles(pattern: string): boolean { @@ -174,6 +182,7 @@ function acceptRule(phases: Record, v: unknown) { if (!isRecord(v)) return undefined const { pattern, phase } = v if (typeof pattern !== 'string' || typeof phase !== 'string') return undefined + if (pattern.length > MAX_PATTERN) return undefined return hasKey(phases, phase) && compiles(pattern) ? { pattern, phase } : undefined @@ -186,7 +195,12 @@ function mergePrompt( phases: Record, log: Log, ): PromptRule[] { - if (!Array.isArray(user)) return base + if (!Array.isArray(user)) { + if (user !== undefined) { + log('model-router: config prompt ignored: not a list') + } + return base + } const rules: PromptRule[] = [] for (const item of user as unknown[]) { const rule = acceptRule(phases, item) @@ -202,7 +216,10 @@ const pickBool = (v: unknown, fallback: boolean): boolean => /** Defaults overlaid with the user's entries, each validated first. */ function mergeConfig(user: unknown, log: Log): Config { const base = structuredClone(DEFAULT_CONFIG) - if (!isRecord(user)) return base + if (!isRecord(user)) { + if (user !== undefined) log('model-router: config ignored: not an object') + return base + } const models = mergeTable( base.models, user.models, 'models', acceptModel, log) const phases = mergeTable( @@ -222,6 +239,24 @@ function mergeConfig(user: unknown, log: Log): Config { } } +/** The file's text, or undefined (logged) when it exceeds the cap. */ +async function readCapped( + $: Api, + path: string, + log: Log, +): Promise { + const tooBig = + `model-router: ${OVERRIDE} over ${MAX_CONFIG_BYTES} bytes; defaults` + if ((await $.fs.stat(path)).size > MAX_CONFIG_BYTES) { + log(tooBig) + return undefined + } + const text = await $.fs.read(path) + if (text.length <= MAX_CONFIG_BYTES) return text + log(tooBig) + return undefined +} + async function readOverride( $: Api, log: Log, @@ -231,7 +266,8 @@ async function readOverride( if (!home) return undefined const path = `${home}/${OVERRIDE}` if (!(await $.fs.exists(path))) return undefined - return { path, data: JSON.parse(await $.fs.read(path)) } + const text = await readCapped($, path, log) + return text === undefined ? undefined : { path, data: JSON.parse(text) } } catch (err) { log(`model-router: ${OVERRIDE} unreadable (${String(err)}); defaults`) return undefined @@ -271,6 +307,28 @@ function newState(cfg: Config, source: string): State { off: false, lastMain: '', windowWarned: false, + warned: new Set(), + } +} + +/** Logs a hook's fail-open once per session; never throws itself. */ +function warnOnce(st: State, $: Api, hook: string, kind: string): void { + if (st.warned.has(hook)) return + st.warned.add(hook) + try { + $.ui.log(`model-router: ${hook} failed (${kind}): ` + + 'routing skipped for this event') + } catch { + // a failing log must not break the fail-open itself + } +} + +/** Runs post-`next` bookkeeping so its failure can never re-run `next`. */ +function safely(st: State, $: Api, hook: string, work: () => void): void { + try { + work() + } catch { + warnOnce(st, $, hook, 'bookkeeping') } } @@ -281,9 +339,6 @@ function loopOf(st: State, agentId: string): Loop { const known = st.loops.get(agentId) if (known) return known const fresh: Loop = { - spawnModel: '', - frozen: false, - explicitModel: false, explicitEffort: false, } st.loops.set(agentId, fresh) @@ -293,7 +348,6 @@ function loopOf(st: State, agentId: string): Loop { /** Writes a route on an agent loop, never on an axis given explicitly. */ function writeLoop(loop: Loop, route: Route): void { if (!loop.explicitEffort) loop.effort = route.effort - if (!loop.explicitModel) loop.model = route.model } function clearRoutes(st: State): void { @@ -439,10 +493,11 @@ async function registerTool($: Api, st: State): Promise { await $.tool.register({ name: 'route', description: - 'Declare the phase of the work ahead so the next model requests ' + - 'run at the effort (and model) it deserves. Call it before a span ' + - 'of work changes nature (planning, orchestrating, mechanical ' + - 'work). It acts on the calling loop only; no model choice here. ' + + 'Declare the phase of the work ahead so the next requests of THIS ' + + 'loop run at the right effort (and, on the main loop, the model ' + + 'when the switch is on). Call it before a span of work changes ' + + 'nature (planning, orchestrating, mechanical work). Effort only ' + + 'for a sub-agent; the model of a sub-agent is fixed at spawn. ' + `Phases: ${Object.keys(st.cfg.phases).join(', ')}.`, inputSchema: { type: 'object', @@ -451,6 +506,7 @@ async function registerTool($: Api, st: State): Promise { effort: { type: 'string', enum: [...LEVELS] }, clear: { type: 'boolean', description: 'drop this loop\'s route' }, }, + additionalProperties: false, }, }) } catch (err) { @@ -516,10 +572,12 @@ function routedText(st: State, agentId: string | undefined, p: Picked): string { } const loop = agentId === undefined ? undefined : st.loops.get(agentId) const effort = loop?.explicitEffort ? undefined : p.route.effort - const model = agentId === undefined && !st.cfg.mainModelSwitch - ? undefined - : loop?.explicitModel ? undefined : p.route.model - const modelNote = p.route.model && model === undefined ? ' (switch off)' : '' + const model = agentId === undefined && st.cfg.mainModelSwitch + ? p.route.model + : undefined + const modelNote = !p.route.model || model !== undefined + ? '' + : agentId === undefined ? ' (switch off)' : ' (fixed at spawn)' return `routed ${agentId === undefined ? 'main' : 'this agent'} to ` + `${p.phase}: effort ${effort ?? 'unchanged'}, model ` + `${model === undefined ? 'unchanged' : resolveModel(st.cfg, model)}` + @@ -619,16 +677,12 @@ function spawnRoute(cfg: Config, e: SpawnIn, frozen: boolean) { } function trackLoop(st: State, e: SpawnIn, started: { - model: string agentId?: string -}, route: Route | undefined, frozen: boolean): void { +}, route: Route | undefined): void { const given = st.explicitEffort.get(e.tool_use_id) st.explicitEffort.delete(e.tool_use_id) if (started.agentId === undefined) return st.loops.set(started.agentId, { - spawnModel: started.model, - frozen, - explicitModel: e.model !== undefined, explicitEffort: given !== undefined, effort: given ? undefined : route?.effort, }) @@ -638,13 +692,7 @@ function trackLoop(st: State, e: SpawnIn, started: { function agentPlan(st: State, e: StepIn): Plan { const loop = e.agentId === undefined ? undefined : st.loops.get(e.agentId) - const wanted = loop?.model - const reroute = loop && wanted !== undefined && !loop.frozen && - e.model === loop.spawnModel - return { - model: reroute ? resolveModel(st.cfg, wanted) : e.model, - effort: loop?.effort ?? e.effort, - } + return { model: e.model, effort: loop?.effort ?? e.effort } } /** True when the context still fits the target model's known window. */ @@ -719,14 +767,29 @@ function registerSession(on: On, st: State): void { await registerCommand($) refresh($, st) return next(e) - }).catch(($, e, next) => next(e)) + }).catch(($, e, next) => { + warnOnce(st, $, 'session.start', next.error.kind) + return next(e) + }) on('session.end', async ($, e, next) => { Object.assign(st, newState(st.cfg, st.source)) return next(e) - }).catch(($, e, next) => next(e)) - on('command.run', { command: 'route' }, async ($, e) => ({ - text: await handleCommand($, st, e.args), - })).catch(($, e, next) => ({ text: `route failed (${next.error.kind})` })) + }).catch(($, e, next) => { + warnOnce(st, $, 'session.end', next.error.kind) + return next(e) + }) +} + +function registerCommandHook(on: On, st: State): void { + on('command.run', { command: 'route' }, async ($, e) => { + if (e.origin.kind !== 'composer') { + return { text: 'route: user-only command' } + } + return { text: await handleCommand($, st, e.args) } + }).catch(($, e, next) => { + warnOnce(st, $, 'command.run', next.error.kind) + return { text: `route failed (${next.error.kind})` } + }) } function registerRouteTool(on: On, st: State): void { @@ -735,28 +798,36 @@ function registerRouteTool(on: On, st: State): void { refresh($, st) vlog($, st, `route ${e.agentId ?? 'main'}: ${JSON.stringify(out)}`) return out - }).catch(($, e, next) => ({ - result: `route failed (${next.error.kind}); nothing routed`, - })) + }).catch(($, e, next) => { + warnOnce(st, $, 'route tool', next.error.kind) + return { result: `route failed (${next.error.kind}); nothing routed` } + }) } function registerSkills(on: On, st: State): void { on('tool.call', { tool: 'Skill' }, async ($, e, next) => { - if (st.off) return next(e) - const level = EFFORT_SKILL.exec(e.skill)?.[1] - if (isLevel(level)) return effortBridge(st, e.agentId, e.skill, level) + const skill = typeof e.skill === 'string' ? e.skill : undefined + if (st.off || skill === undefined) return next(e) + const level = EFFORT_SKILL.exec(skill)?.[1] + if (isLevel(level)) return effortBridge(st, e.agentId, skill, level) st.skillCalls += 1 try { - onSkillLoad(st, e.skill, e.agentId) + safely(st, $, 'Skill', () => onSkillLoad(st, skill, e.agentId)) return await next(e) } finally { st.skillCalls -= 1 } - }).catch(($, e, next) => next(e)) + }).catch(($, e, next) => { + warnOnce(st, $, 'Skill', next.error.kind) + return next(e) + }) on('skill.prompt', async ($, e, next) => { if (st.off || st.skillCalls > 0) return next(e) return slashEffort(st, e.skill, e.text) ?? next(e) - }).catch(($, e, next) => next(e)) + }).catch(($, e, next) => { + warnOnce(st, $, 'skill.prompt', next.error.kind) + return next(e) + }) } function registerAgents(on: On, st: State): void { @@ -765,7 +836,13 @@ function registerAgents(on: On, st: State): void { st.explicitEffort.set(e.tool_use_id, e.effort) } return next(e) - }).catch(($, e, next) => next(e)) + }).catch(($, e, next) => { + warnOnce(st, $, 'Agent', next.error.kind) + return next(e) + }) +} + +function registerSpawn(on: On, st: State): void { on('agent.spawn', async ($, e, next) => { if (st.off) return next(e) const frozen = e.fork || e.workflow !== undefined @@ -775,11 +852,18 @@ function registerAgents(on: On, st: State): void { const started = await next(wanted === undefined ? e : { ...e, model: resolveModel(st.cfg, wanted) }) - if (started.deny !== undefined) return started - trackLoop(st, e, started, route, frozen) - vlog($, st, `spawn ${e.subagentType}: ${e.model ?? '-'} → ${started.model}`) + if (typeof started.agentId === 'string') { + safely(st, $, 'agent.spawn', () => { + trackLoop(st, e, started, route) + vlog($, st, + `spawn ${e.subagentType}: ${e.model ?? '-'} → ${started.model}`) + }) + } return started - }).catch(($, e, next) => next(e)) + }).catch(($, e, next) => { + warnOnce(st, $, 'agent.spawn', next.error.kind) + return next(e) + }) } function registerTurns(on: On, st: State): void { @@ -790,22 +874,28 @@ function registerTurns(on: On, st: State): void { if (e.agentId === undefined) noteMain($, st, plan) vlog($, st, stepLog(e, plan)) const result = yield* next(changed ? withPlan(e, plan) : e) - vlog($, st, `step ${e.index} answered by ${result.usage?.model ?? '?'}`) + safely(st, $, 'turn.step', () => vlog($, st, + `step ${e.index} answered by ${result.usage?.model ?? '?'}`)) return result }).catch(async function* ($, e, next) { + warnOnce(st, $, 'turn.step', next.error.kind) return yield* next(e) }) on('turn.complete', async ($, e, next) => { if (e.agentId !== undefined) st.loops.delete(e.agentId) else endMainTurn($, st) return next(e) - }).catch(($, e, next) => next(e)) + }).catch(($, e, next) => { + warnOnce(st, $, 'turn.complete', next.error.kind) + return next(e) + }) } function registerPrompt(on: On, st: State): void { on('prompt.submit', async ($, e, next) => { if (st.off || e.origin.kind !== 'composer') return next(e) - const rule = st.rules.find(r => r.re.test(e.text)) + const scanned = e.text.slice(0, MAX_PROMPT_SCAN) + const rule = st.rules.find(r => r.re.test(scanned)) const route = rule ? phaseRoute(st.cfg, rule.phase) : undefined if (rule && route) { const routed: Routed = { phase: rule.phase, route, source: 'prompt' } @@ -815,20 +905,28 @@ function registerPrompt(on: On, st: State): void { refresh($, st) } return next(e) - }).catch(($, e, next) => next(e)) + }).catch(($, e, next) => { + warnOnce(st, $, 'prompt.submit', next.error.kind) + return next(e) + }) on('ui.render', { component: 'Spinner' }, async ($, e, next) => { if (st.off || !st.cfg.spinner || !st.lastMain) return next(e) const suffix = ` · ${st.lastMain}…` return next({ ...e, props: { ...e.props, suffix } }) - }).catch(($, e, next) => next(e)) + }).catch(($, e, next) => { + warnOnce(st, $, 'ui.render', next.error.kind) + return next(e) + }) } export const register: Register = on => { const st = newState(mergeConfig(undefined, () => undefined), 'defaults') registerSession(on, st) + registerCommandHook(on, st) registerRouteTool(on, st) registerSkills(on, st) registerAgents(on, st) + registerSpawn(on, st) registerTurns(on, st) registerPrompt(on, st) } From ae0179f4912d28a1e16bc8ad3622b01f17a53113 Mon Sep 17 00:00:00 2001 From: bchanot Date: Thu, 8 Oct 2026 16:53:43 +0200 Subject: [PATCH 07/22] chore(tasks): model-router contract criteria 7-11, TODO hardening done + residuals, journal --- .claude/memory/journal.md | 1 + .claude/tasks/TODO.md | 3 ++- .../contracts/2026-10-08-model-router-w1a-1533.md | 12 ++++++------ 3 files changed, 9 insertions(+), 7 deletions(-) diff --git a/.claude/memory/journal.md b/.claude/memory/journal.md index 5e77116..adc0290 100644 --- a/.claude/memory/journal.md +++ b/.claude/memory/journal.md @@ -580,3 +580,4 @@ rules: ## 2026-10-08 - model-router mod, wave 0 spike (user ask: one mod routes model + effort per request, replaces effort-* shifters + pins). Plan `.claude/tasks/plans/2026-10-08-model-router-mod.md`, 4 decisions by AskUserQuestion (spike-first main-loop switch, `CLAUDE_CODE_PLUGIN_DIRS` load, migration wave 2, names model-router / route / /route), rule "pin = entry default, sub-tasks route finer". Spike in dev-mods, hot reload on: `turn.step` effort rewrite proven (transcript `effort` field is the oracle, not `CLAUDE_EFFORT`); sub-agent model at `agent.spawn` + effort per step by agentId proven; main-loop fable → sonnet-5-5/low for 3 steps then back: works, one cold-cache step per switch INTO a model, return free. Found: hook-side alias resolver stale (`sonnet` → `claude-sonnet-5`, 404; Agent tool enum resolves the same alias to 5-5) → mod writes full ids only. feature/model-router-mod open, nothing committed yet (plan + TODO + journal pending). - model-router wave 1-A (/feat, user go): mod built in `mods/model-router/` (4 files, 834 + 202 lines, 11 plugin tests). Plan r1 → r3: 3 challengers (simplicity CONCERNS, robustness CONCERNS(6), correctness FATAL(8)) + 1 confirmation CONCERNS(4); converged on: agents table = built-ins only in wave 1 (frontmatter stays single writer), agent model written once at spawn, explicit Agent params frozen per loop, every sub-agent write on its own loop, Skill bridge answers in the Skill tool's OUTPUT schema (string result refused → skill would load), config validated before merge, state in closure, `/route off`. feater DONE first pass; GATE 0 MET; verifier CONFORME 6/6; security PASS (4 MEDIUM + 5 LOW parked in TODO for user go). Commits b721c94 (mod) + e8ca713 (contract/plan). Override `~/.claude/model-router.json` {verbose:true} written (pass B). Doc-sync deferred to 1-B. +- model-router W1-A hardening (user go): fresh feater on contract criteria 7-11 → gap round (dead `Loop.frozen`, spawn returns `started` verbatim, tool description) → verifier CONFORME 11/11 → security PASS (1 MEDIUM residual: ReDoS size-bounded only, self-inflicted config; 5 LOW parked). Registries BDR-115, LRN-205, LRN-206, EVAL-040 written on user go (64702d5). Next: live swap of the real mod into the hot-reload folder, then W1-B install + docs. diff --git a/.claude/tasks/TODO.md b/.claude/tasks/TODO.md index 3050a16..b391e2e 100644 --- a/.claude/tasks/TODO.md +++ b/.claude/tasks/TODO.md @@ -6,7 +6,8 @@ model switch spike-first then flag off; load via `CLAUDE_CODE_PLUGIN_DIRS` + lin migration of shifters/pins/model-gate in wave 2 after proof; names model-router / route / /route. - [x] W0 spike in dev-mods (hot reload): facts a-d established 2026-10-08 (plan file § Spike facts); e moved to W1.10 - [x] W1-A the mod in `mods/model-router/` (b721c94, contract `2026-10-08-model-router-w1a-1533`, plan r3): challenge 3 lenses + 1 confirmation (2 BLOCKER + 10 MAJOR closed by named changes), feater DONE first pass, GATE 0 MET 5/5, verifier CONFORME 6/6, security PASS (4 MEDIUM + 5 LOW reported, below) -- [ ] W1-A hardening (security report, user decision): `/route` command origin check (`composer` only); in-agent `route` must not re-model the agent (strip `route.model` for agent loops, drop the dead per-step model path, align the tool description); ReDoS caps on config patterns (length ≤ 200, test ≤ 4 KB); one log per session in every `.catch` + log the silent config drops; phase key charset `^[a-z][a-z0-9_-]{0,31}$`; `additionalProperties: false` on the tool schema + `typeof e.skill`; post-`next` bookkeeping in try/catch; config file size cap 64 KB. Deferred by design: effort ceiling for model-declared routes; window check beyond haiku. +- [x] W1-A hardening (user go "oui durcis", contract criteria 7-11): `/route` composer-only; no agent model axis (effort only after spawn); pattern ≤ 200 / scan ≤ 4096 / phase keys `^[a-z][a-z0-9_-]{0,31}$` / config ≤ 64 KB / `additionalProperties: false` / `typeof e.skill`; `warnOnce` in all 14 catches + config-drop logs; `safely` around post-`next` bookkeeping. feater DONE, gap round (dead `Loop.frozen`, spawn returns `started` verbatim, tool description), GATE 0 MET 9/9, verifier CONFORME 11/11, security PASS. +- [ ] W1-A residuals (security, accepted, none exploitable from outside the user's own files): ReDoS is size-bounded only (`(a+)+$` in `~/.claude/model-router.json` + a 4 KB paste hangs the hook; fix = nested-quantifier rejection or a far lower scan cap); `stat().size` trusted (FIFO/device path in ~/.claude); unrestricted `models`/`agents`/`skills` KEYS echoed raw in logs (log flood); `agentId` read from the flat tool event (engine strip unverified); 5 closures without a kit test (no fs in the kit, LRN-206); "nothing routed" catch text after a state write. - [ ] W1-B install + docs: `mods/` symlink + `CLAUDE_CODE_PLUGIN_DIRS` in settings.json env via link.sh, root `.gitignore` for the engine-laid `mods/*/tsconfig.json` + `.claude-plugin/types/`, doctor line, README/USAGE/CHANGELOG (doc-sync deferred here from the 1-A /feat run: nothing to document before the install exists), live test in session (swap the spike for the real mod) - [ ] W2 migration: 15 skills off `Skill(effort-*)`, remove shifters + effort-pins + model-gate, census repointed, docs - [ ] W3 optional: step heuristics, haiku classifier, quota-aware downgrade, A/B diff --git a/.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md b/.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md index 3221849..e938038 100644 --- a/.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md +++ b/.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md @@ -29,7 +29,7 @@ Q: legacy `Skill(effort-*)` / A: answered by the mod without loading the skill ( 3. `claude plugin test mods/model-router` passes; the suite covers: (a) `Skill(effort-low)` via `$.tool.call` is answered without `next` in the Skill tool's output shape (`success`, `commandName`) and the route shows `low` on main; (b) the `route` tool with `phase: "orchestrate"` sets `medium` on main and `/route show` prints it; (c) `/route clear` drops it; (d) `/route bogus` returns an error text naming the phases; (e) a `prompt.submit` text holding `ultrathink` sets `escalate` on main; (f) `/route model=sonnet` shows `claude-sonnet-5-5`, a full id passes through, a misspelt alias is refused; (g) `agent.spawn` of `Explore` without a model param reaches the bottom with `model === 'claude-sonnet-5-5'`, and with `model: 'opus'` given the param is untouched. CHECK: cd mods/model-router && out=$(claude plugin test . 2>&1); rc=$?; echo "$out" | tail -n 5; [ $rc -eq 0 ] && [ "$(grep -cE '^\s*test\(' hooks/register.test.ts)" -ge 7 ] && echo PLUGIN-TEST-OK EXPECT: PLUGIN-TEST-OK - EVIDENCE: MET exit=0 marker-found :: (pass) a tabled agent steps at its table effort [19.75ms] 11 pass 0 fail Ran 11 tests across 1 file. [0.47s] PLUGIN-TEST-OK + EVIDENCE: MET exit=0 marker-found :: (pass) a rule only scans the first 4096 chars of a prompt [27.16ms] 14 pass 0 fail Ran 14 tests across 1 file. [0.60s] PLUGIN-TEST-OK 4. Hooks present, as `claude plugin validate` lists them: `session.start`, `command.run{command=route}`, `tool.call{tool=mcp__model-router__route}`, `tool.call{tool=Skill}`, `tool.call{tool=Agent}`, `skill.prompt`, `agent.spawn`, `turn.step`, `prompt.submit`, `turn.complete`, `ui.render{component=Spinner}`; every gating hook carries a fail-open `.catch` (validate prints no "gating hook without .catch"). CHECK: cd mods/model-router && out=$(claude plugin validate . 2>&1) && for h in session.start 'command.run{command=route}' 'tool.call{tool=mcp__model-router__route}' 'tool.call{tool=Skill}' 'tool.call{tool=Agent}' skill.prompt agent.spawn turn.step prompt.submit turn.complete 'ui.render{component=Spinner}'; do echo "$out" | grep -qF -- "$h" || { echo "missing $h"; exit 1; }; done && ! echo "$out" | grep -q 'without .catch' && echo HOOKS-OK EXPECT: HOOKS-OK @@ -43,23 +43,23 @@ Q (r2): `scope: agents` / `/route agents` / A: dropped (challenge r1, no require Q (r2): agents table in wave 1 / A: built-ins only (Explore, Plan), matched for the engine provider; repo agents keep their frontmatter as the single writer until wave 2. [orchestrator — single source of truth] Q (r2): `/route off` / A: added as the session kill switch (every hook passes through). [orchestrator — robustness] -## HARDENING ROUND (security gate 2026-10-08, user go "oui durcis") — criteria 7-11 +Hardening round (security gate 2026-10-08, user go "oui durcis") — criteria 7-11, same ledger: 7. `/route` is user-only: `command.run` answers `{ text: 'route: user-only command' }` without acting when `e.origin.kind !== 'composer'`; a test proves it (origin `{ kind: 'plugin', name: 'x' }` or the kit's non-composer origin → text contains `user-only`, state unchanged). CHECK: cd mods/model-router && grep -q "origin.kind" hooks/register.ts && grep -q "user-only" hooks/register.ts && grep -q "user-only" hooks/register.test.ts && echo ORIGIN-OK EXPECT: ORIGIN-OK - EVIDENCE: pending + EVIDENCE: MET exit=0 marker-found :: ORIGIN-OK 8. An in-agent `route` call or skill table entry never changes that agent's MODEL: the `Loop` type has no `model` axis and no `explicitModel` flag, `turn.step` on an agent loop rewrites `effort` only, the route tool's description says "effort only; the model of a sub-agent is fixed at spawn". A test proves it: a route tool call carrying `agentId: 'a1'` with `phase: 'judge'` followed by a `turn.step` for `a1` leaves `model` as given and sets `effort` to `xhigh`. CHECK: cd mods/model-router && ! grep -qE "loop\.model|explicitModel|spawnModel" hooks/register.ts && grep -q "fixed at spawn" hooks/register.ts && echo NO-AGENT-MODEL-OK EXPECT: NO-AGENT-MODEL-OK - EVIDENCE: pending + EVIDENCE: MET exit=0 marker-found :: NO-AGENT-MODEL-OK 9. Config hardening: a prompt rule pattern longer than 200 chars is dropped (logged); `re.test` runs on at most the first 4096 chars of the prompt; phase keys must match `^[a-z][a-z0-9_-]{0,31}$` (others dropped, logged); the override file is refused above 65536 bytes (logged, defaults kept); the route tool schema carries `additionalProperties: false`; the Skill hook checks `typeof e.skill === 'string'`. CHECK: cd mods/model-router && grep -q "additionalProperties: false" hooks/register.ts && grep -qE "\[a-z\]\[a-z0-9_-\]\{0,31\}" hooks/register.ts && grep -qE "4096|4_096" hooks/register.ts && grep -qE "65536|65_536|64 \* 1024" hooks/register.ts && grep -qE "200" hooks/register.ts && grep -q "typeof e.skill === 'string'" hooks/register.ts && echo CONFIG-HARDEN-OK EXPECT: CONFIG-HARDEN-OK - EVIDENCE: pending + EVIDENCE: MET exit=0 marker-found :: CONFIG-HARDEN-OK 10. Visible fail-open: every `.catch` logs once per session per hook (`$.ui.log('model-router: failed (): routing skipped for this event')`, a `warned: Set` in the state) before passing through or answering; the three silent config drops (non-object top level, wrong-typed table, non-array `prompt`) log a line. CHECK: cd mods/model-router && [ "$(grep -c '\.catch(' hooks/register.ts)" -ge 12 ] && grep -q "warned" hooks/register.ts && grep -q "routing skipped" hooks/register.ts && echo CATCH-LOG-OK EXPECT: CATCH-LOG-OK - EVIDENCE: pending + EVIDENCE: MET exit=0 marker-found :: CATCH-LOG-OK 11. Post-`next` bookkeeping (loop tracking and logs after `await next(...)` in `agent.spawn` and the Skill hook) runs inside its own try/catch so a logging failure can never make the `.catch` re-run `next`. Judged by reading, with criteria 1-6 still MET (validate, tsc, tests ≥ 13, style, AC6 minus the removed model axis). ## FILE SCOPE From 77ad7cf494f53cfc1915b18058c723afe42ad6a6 Mon Sep 17 00:00:00 2001 From: bchanot Date: Thu, 8 Oct 2026 18:37:32 +0200 Subject: [PATCH 08/22] =?UTF-8?q?chore(tasks):=20model-router=20W1-B=20spl?= =?UTF-8?q?it=20=E2=80=94=20floor=20+=20wiring=20contracts/plans,=20skills?= =?UTF-8?q?-dir=20loading=20decision,=20journal?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- .claude/memory/journal.md | 1 + .claude/tasks/TODO.md | 3 +- .../2026-10-08-model-router-floor-1835.md | 33 ++++++ .../2026-10-08-model-router-wiring-1835.md | 45 ++++++++ .../2026-10-08-model-router-floor-1835.md | 108 ++++++++++++++++++ .../plans/2026-10-08-model-router-mod.md | 21 ++++ .../2026-10-08-model-router-wiring-1835.md | 106 +++++++++++++++++ 7 files changed, 316 insertions(+), 1 deletion(-) create mode 100644 .claude/tasks/contracts/2026-10-08-model-router-floor-1835.md create mode 100644 .claude/tasks/contracts/2026-10-08-model-router-wiring-1835.md create mode 100644 .claude/tasks/plans/2026-10-08-model-router-floor-1835.md create mode 100644 .claude/tasks/plans/2026-10-08-model-router-wiring-1835.md diff --git a/.claude/memory/journal.md b/.claude/memory/journal.md index adc0290..0af5f38 100644 --- a/.claude/memory/journal.md +++ b/.claude/memory/journal.md @@ -581,3 +581,4 @@ rules: - model-router mod, wave 0 spike (user ask: one mod routes model + effort per request, replaces effort-* shifters + pins). Plan `.claude/tasks/plans/2026-10-08-model-router-mod.md`, 4 decisions by AskUserQuestion (spike-first main-loop switch, `CLAUDE_CODE_PLUGIN_DIRS` load, migration wave 2, names model-router / route / /route), rule "pin = entry default, sub-tasks route finer". Spike in dev-mods, hot reload on: `turn.step` effort rewrite proven (transcript `effort` field is the oracle, not `CLAUDE_EFFORT`); sub-agent model at `agent.spawn` + effort per step by agentId proven; main-loop fable → sonnet-5-5/low for 3 steps then back: works, one cold-cache step per switch INTO a model, return free. Found: hook-side alias resolver stale (`sonnet` → `claude-sonnet-5`, 404; Agent tool enum resolves the same alias to 5-5) → mod writes full ids only. feature/model-router-mod open, nothing committed yet (plan + TODO + journal pending). - model-router wave 1-A (/feat, user go): mod built in `mods/model-router/` (4 files, 834 + 202 lines, 11 plugin tests). Plan r1 → r3: 3 challengers (simplicity CONCERNS, robustness CONCERNS(6), correctness FATAL(8)) + 1 confirmation CONCERNS(4); converged on: agents table = built-ins only in wave 1 (frontmatter stays single writer), agent model written once at spawn, explicit Agent params frozen per loop, every sub-agent write on its own loop, Skill bridge answers in the Skill tool's OUTPUT schema (string result refused → skill would load), config validated before merge, state in closure, `/route off`. feater DONE first pass; GATE 0 MET; verifier CONFORME 6/6; security PASS (4 MEDIUM + 5 LOW parked in TODO for user go). Commits b721c94 (mod) + e8ca713 (contract/plan). Override `~/.claude/model-router.json` {verbose:true} written (pass B). Doc-sync deferred to 1-B. - model-router W1-A hardening (user go): fresh feater on contract criteria 7-11 → gap round (dead `Loop.frozen`, spawn returns `started` verbatim, tool description) → verifier CONFORME 11/11 → security PASS (1 MEDIUM residual: ReDoS size-bounded only, self-inflicted config; 5 LOW parked). Registries BDR-115, LRN-205, LRN-206, EVAL-040 written on user go (64702d5). Next: live swap of the real mod into the hot-reload folder, then W1-B install + docs. +- model-router live checks on Opus 5.5 (user /model, uncommitted settings.json change left to the user): Skill(effort-low) bridge answered in place → next request low; ultrathink turn ran max (engine base medium); Explore without params → claude-sonnet-5-5, 3 steps medium. Loading switched to tracked symlink skills/model-router → @skills-dir (PLUGIN_DIRS non-portable: absolute path, no $HOME expansion, tracked settings); isolated-HOME probe listed/enabled/loaded. User chose floor semantics for ultrathink + typed /effort-. W1-B split: B1 floor (register.ts) + B2 wiring (symlink, gitignore, mods suite, doctor, CLAUDE.md); 6 challengers in flight. Effort shifters skipped this turn: the bridge would overwrite the user's ultrathink until B1 lands. diff --git a/.claude/tasks/TODO.md b/.claude/tasks/TODO.md index b391e2e..d1b2620 100644 --- a/.claude/tasks/TODO.md +++ b/.claude/tasks/TODO.md @@ -8,7 +8,8 @@ migration of shifters/pins/model-gate in wave 2 after proof; names model-router - [x] W1-A the mod in `mods/model-router/` (b721c94, contract `2026-10-08-model-router-w1a-1533`, plan r3): challenge 3 lenses + 1 confirmation (2 BLOCKER + 10 MAJOR closed by named changes), feater DONE first pass, GATE 0 MET 5/5, verifier CONFORME 6/6, security PASS (4 MEDIUM + 5 LOW reported, below) - [x] W1-A hardening (user go "oui durcis", contract criteria 7-11): `/route` composer-only; no agent model axis (effort only after spawn); pattern ≤ 200 / scan ≤ 4096 / phase keys `^[a-z][a-z0-9_-]{0,31}$` / config ≤ 64 KB / `additionalProperties: false` / `typeof e.skill`; `warnOnce` in all 14 catches + config-drop logs; `safely` around post-`next` bookkeeping. feater DONE, gap round (dead `Loop.frozen`, spawn returns `started` verbatim, tool description), GATE 0 MET 9/9, verifier CONFORME 11/11, security PASS. - [ ] W1-A residuals (security, accepted, none exploitable from outside the user's own files): ReDoS is size-bounded only (`(a+)+$` in `~/.claude/model-router.json` + a 4 KB paste hangs the hook; fix = nested-quantifier rejection or a far lower scan cap); `stat().size` trusted (FIFO/device path in ~/.claude); unrestricted `models`/`agents`/`skills` KEYS echoed raw in logs (log flood); `agentId` read from the flat tool event (engine strip unverified); 5 closures without a kit test (no fs in the kit, LRN-206); "nothing routed" catch text after a state write. -- [ ] W1-B install + docs: `mods/` symlink + `CLAUDE_CODE_PLUGIN_DIRS` in settings.json env via link.sh, root `.gitignore` for the engine-laid `mods/*/tsconfig.json` + `.claude-plugin/types/`, doctor line, README/USAGE/CHANGELOG (doc-sync deferred here from the 1-A /feat run: nothing to document before the install exists), live test in session (swap the spike for the real mod) +- [ ] W1-B1 floor precedence (contract `2026-10-08-model-router-floor-1835`): `ultrathink` + typed `/effort-` = floor for the main turn (user choice 2026-10-08) +- [ ] W1-B2 wiring (contract `2026-10-08-model-router-wiring-1835`): tracked symlink `skills/model-router` → `../mods/model-router` (loads as `@skills-dir`; PLUGIN_DIRS dropped: absolute paths in tracked settings), `.gitignore` engine-laid files, `lib/tests/mods.test.sh`, doctor `── Mods ──`, CLAUDE.md `## mods/`; then doc-sync (README/USAGE/CHANGELOG), live swap (remove the hot-reload link, `/reload-plugins`), BDR-115 amendment - [ ] W2 migration: 15 skills off `Skill(effort-*)`, remove shifters + effort-pins + model-gate, census repointed, docs - [ ] W3 optional: step heuristics, haiku classifier, quota-aware downgrade, A/B diff --git a/.claude/tasks/contracts/2026-10-08-model-router-floor-1835.md b/.claude/tasks/contracts/2026-10-08-model-router-floor-1835.md new file mode 100644 index 0000000..53daccc --- /dev/null +++ b/.claude/tasks/contracts/2026-10-08-model-router-floor-1835.md @@ -0,0 +1,33 @@ +# CONTRACT — model-router-floor (wave 1-B1: user effort floor for the turn) +- date: 2026-10-08 | flow: feat | branch: feature/model-router-mod +- status: active + +## REQUEST (verbatim — IMMUTABLE) +AskUserQuestion 2026-10-08, question: "Aujourd'hui, `ultrathink` met le tour en max, mais si je déclare une phase (ex. orchestrate) puis charge un skill, ton max est perdu pour la suite du tour. Ça contredit notre règle « un choix explicite bat la phase déduite ». Quel sens donner à `ultrathink` et à un `/effort-x` tapé par toi ?" +User's answer: "Plancher pour le tour (Recommended)" — option text: "Ton niveau est un minimum pour tout le tour. Les routes du modèle et des skills peuvent monter au-dessus (escalade à max), jamais descendre en dessous. Il passe aussi par-dessus un /route sticky plus bas." +Session rule this fixes (wave plan, user-approved 2026-10-08): "an explicit per-call choice (Agent `model`/`effort` param, `/route`, `ultrathink`) beats the derived phase for that span". +User, same turn: "continu avec Opus en /ultrathink". + +## CLARIFICATIONS +Q: which loops does the floor cover? / A: the MAIN loop only ("le tour" = the user's turn); sub-agents keep their own routes and pins. [orchestrator — derived, stated to the user] +Q: model axis / A: the floor constrains effort only. A floor route's model (none by default: ultrathink → escalate carries none) applies on main only when neither the sticky nor the turn route names a model, and only with the switch on. [orchestrator — internal] +Q: numeric or absent engine effort / A: the floor level replaces it (the user's explicit level wins over an unknown budget); the haiku effort omission still applies after flooring. [orchestrator — internal] +Q: who clears the floor / A: main turn end (a queued prompt's floor is then promoted), `/route clear`, `/route off` (pass-through). A model `route({clear})`, a `Skill(effort-*)` or any skill load never touches it. [orchestrator — derived from "jamais descendre en dessous"] + +## ACCEPTANCE CRITERIA +1. Suite green with the new tests: `claude plugin test` passes with at least 19 `test(` calls; `claude plugin validate` passes with no warning; no line over 80 chars; no `any` type. + CHECK: cd mods/model-router && out=$(claude plugin test . 2>&1); rc=$?; echo "$out" | tail -n 3; [ $rc -eq 0 ] && [ "$(grep -cE '^\s*test\(' hooks/register.test.ts)" -ge 19 ] && v=$(claude plugin validate . 2>&1) && echo "$v" | grep -q 'Validation passed' && ! echo "$v" | grep -qi 'warning' && ! grep -nE '.{81,}' hooks/register.ts hooks/register.test.ts && ! grep -nE ':\s*any\b||as any\b' hooks/register.ts && echo FLOOR-SUITE-OK + EXPECT: FLOOR-SUITE-OK + EVIDENCE: pending +2. Type-check clean against this build's declarations. + CHECK: T=/Users/b.chanot/.claude/dev-mods/385f7190-70f5-4bdd-b0d8-e4566cd412fd/model-router/.claude-plugin/types; [ -d "$T" ] || T=/Users/b.chanot/Documents/claude/mods/model-router/.claude-plugin/types; W=$(mktemp -d) && printf '{"compilerOptions":{"target":"es2023","lib":["es2023"],"types":[],"module":"esnext","moduleResolution":"bundler","strict":true,"noUncheckedIndexedAccess":true,"noEmit":true,"skipLibCheck":true,"jsx":"react","jsxFactory":"h","jsxFragmentFactory":"Fragment"},"include":["%s/claude-code/index.d.ts","%s/claude-code-tools/index.d.ts","%s/hooks"]}' "$T" "$T" "$PWD/mods/model-router" > "$W/tsconfig.json" && (cd "$W" && npx --yes -p typescript@5 tsc -p tsconfig.json) && echo TSC-OK + EXPECT: TSC-OK + EVIDENCE: pending +3. The floor is its own slot: `turnFloor` is declared in `State`, initialised in `newState`, written by the prompt rule and by a typed `/effort-`, cleared by `/route clear` and at main turn end; the suite carries at least 6 tests whose name contains `floor`. + CHECK: cd mods/model-router/hooks && [ "$(grep -c 'turnFloor' register.ts)" -ge 6 ] && [ "$(grep -cE "^\s*test\('[^']*floor" register.test.ts)" -ge 6 ] && echo FLOOR-SLOT-OK + EXPECT: FLOOR-SLOT-OK + EVIDENCE: pending +4. Judged by reading: effective main effort = the higher (LEVELS order) of the floor and `(userMain ?? turnMain)?.route.effort ?? e.effort`, the floor replacing a numeric or absent engine effort; the floor never applies to a sub-agent step; `turnMain` only ever holds 'model' or 'skill' sources; the prompt rule writes `turnFloor` (or `pendingPrompt` when typed mid-turn with `wait`), promoted into `turnFloor` at main turn end; a non-effort skill load resets `turnMain` only; a model `route({clear})` clears `turnMain` only; every answer that the floor overrides says so truthfully (Skill bridge context, route tool text, `/effort-` text); `/route show` and the status line display the floor; every criterion of `.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md` still holds; no function over 25 logic lines. + +## FILE SCOPE +mods/model-router/hooks/register.ts · mods/model-router/hooks/register.test.ts diff --git a/.claude/tasks/contracts/2026-10-08-model-router-wiring-1835.md b/.claude/tasks/contracts/2026-10-08-model-router-wiring-1835.md new file mode 100644 index 0000000..fac021b --- /dev/null +++ b/.claude/tasks/contracts/2026-10-08-model-router-wiring-1835.md @@ -0,0 +1,45 @@ +# CONTRACT — model-router-wiring (wave 1-B2: active in every session, tests, doctor) +- date: 2026-10-08 | flow: feat | branch: feature/model-router-mod +- status: active + +## REQUEST (verbatim — IMMUTABLE) +User (fr, first message of the session): "Et le mieux que ca soit configurable et qu'on puisse l'installer et qu'il soit actif sur toutes les session en userscope". +AskUserQuestion 2026-10-08, question on the loading mechanism (CLAUDE_CODE_PLUGIN_DIRS needs an absolute path, no $HOME expansion in settings `env`, settings.json is tracked and shared): answer "1 . Mais ce n'est pas un skill on est d'accord ? C'est un mod. don cplus u plugin. ET du coup pourquoi pqs link directement ule dossier mod vers le .claude/mods directement via le link.sh ? Pourquoi passer par skills ?" (option 1 = "Lien sous skills/ (Recommended)"). +Explanation given to the user the same turn: a mod is a plugin, not a skill; Claude Code loads a plugin only from a marketplace, an absolute `CLAUDE_CODE_PLUGIN_DIRS` path, a claude.ai sync, or a plugin folder (`.claude-plugin/plugin.json`) under `~/.claude/skills/` (origin `@skills-dir`, loaded in place); a `~/.claude/mods` link alone loads nothing. +Same question batch, settings answer: "Je gère moi-même" (the user's uncommitted `/model` change to settings.json). + +## CLARIFICATIONS +Q: loading mechanism / A: `skills/` = relative symlink `../mods/`, tracked in git; loads as `@skills-dir`, in place, on every machine where link.sh links `~/.claude/skills`. No settings.json change, no link.sh change. Probe 2026-10-08 in an isolated HOME: listed, enabled, "Status: ✔ loaded". [gated 2026-10-08] +Q: settings.json / A: the working-tree change (model → opus, env block moved) stays untouched and out of every commit. [gated 2026-10-08] +Q: override convention / A: a mod's optional user config lives at `~/.claude/.json` (model-router already reads `~/.claude/model-router.json`). [orchestrator] +Q: doctor scope / A: per mod: the loading link resolves into the repo mod dir; `claude plugin list --json` lists `@skills-dir` enabled (warn, not fail, when absent, disabled, or `claude` missing); the override file parses as JSON when present. [orchestrator] + +## ACCEPTANCE CRITERIA +1. `skills/model-router` is a symlink whose target is exactly `../mods/model-router`, and git does not ignore it. + CHECK: [ -L skills/model-router ] && [ "$(readlink skills/model-router)" = "../mods/model-router" ] && ! git check-ignore -q skills/model-router && echo LINK-OK + EXPECT: LINK-OK + EVIDENCE: pending +2. Engine-laid files are ignored: a mod's root `tsconfig.json` and anything under its `.claude-plugin/types/`; the tracked mod files are not ignored. + CHECK: git check-ignore -q mods/model-router/tsconfig.json && git check-ignore -q mods/model-router/.claude-plugin/types/claude-code/index.d.ts && ! git check-ignore -q mods/model-router/hooks/register.ts && ! git check-ignore -q mods/model-router/.claude-plugin/plugin.json && echo IGNORE-OK + EXPECT: IGNORE-OK + EVIDENCE: pending +3. `lib/tests/mods.test.sh` passes on the repo and fails on a fixture that lacks the loading link (positive control through `MODS_ROOT`). + CHECK: make test suite=lib/tests/mods.test.sh >/dev/null 2>&1 && W=$(mktemp -d) && mkdir -p "$W/mods" "$W/skills" && cp -R mods/model-router "$W/mods/" && ! MODS_ROOT="$W" bash lib/tests/mods.test.sh >/dev/null 2>&1 && ln -s ../mods/model-router "$W/skills/model-router" && MODS_ROOT="$W" bash lib/tests/mods.test.sh >/dev/null 2>&1 && echo MODS-SUITE-OK + EXPECT: MODS-SUITE-OK + EVIDENCE: pending +4. `doctor.sh` prints a `── Mods ──` section with a ✓ line for the model-router loading link. + CHECK: out=$(bash doctor.sh 2>&1); echo "$out" | sed -n '/── Mods ──/,/^$/p' | grep -q '✓.*model-router' && echo DOCTOR-MODS-OK + EXPECT: DOCTOR-MODS-OK + EVIDENCE: pending +5. `CLAUDE.md` has a `## mods/` section naming the `skills/` relative symlink, the `@skills-dir` origin, why not `CLAUDE_CODE_PLUGIN_DIRS`, the gitignored engine-laid files, `~/.claude/.json`, the suite command and how to turn a mod off. + CHECK: grep -q '^## mods/' CLAUDE.md && grep -q '@skills-dir' CLAUDE.md && grep -q 'CLAUDE_CODE_PLUGIN_DIRS' CLAUDE.md && grep -q 'mods.test.sh' CLAUDE.md && grep -q '.json' CLAUDE.md && grep -q '@skills-dir": false' CLAUDE.md && echo CLAUDEMD-OK + EXPECT: CLAUDEMD-OK + EVIDENCE: pending +6. Health stack on the touched shell files, doctrine census green. + CHECK: shellcheck lib/tests/mods.test.sh doctor.sh && make test suite=lib/tests/doctrine-citers.test.sh >/dev/null 2>&1 && echo HEALTH-OK + EXPECT: HEALTH-OK + EVIDENCE: pending +7. Judged by reading: no change to settings.json, link.sh or any install script; the user's settings.json working-tree diff is untouched; the suite SKIPs (explicit SKIP line, exit 0 for that part) only the `claude`-dependent checks when `claude` is absent, and fails when no mod is found at all; doctor's new section never increments the core-link counter (`_LINK_PASS`) and never fails on a missing `claude`; the CLAUDE.md section is terse English matching the file's style. + +## FILE SCOPE +skills/model-router (new symlink) · .gitignore · lib/tests/mods.test.sh (new) · doctor.sh · CLAUDE.md diff --git a/.claude/tasks/plans/2026-10-08-model-router-floor-1835.md b/.claude/tasks/plans/2026-10-08-model-router-floor-1835.md new file mode 100644 index 0000000..6058c60 --- /dev/null +++ b/.claude/tasks/plans/2026-10-08-model-router-floor-1835.md @@ -0,0 +1,108 @@ +# PLAN — model-router wave 1-B1: user effort floor for the turn (dispatch-ready) +Contract: .claude/tasks/contracts/2026-10-08-model-router-floor-1835.md +Code: mods/model-router/hooks/register.ts (read it in full first) and +register.test.ts. API truth: the engine-laid declarations under +mods/model-router/.claude-plugin/types/ (claude-code/index.d.ts, +claude-code-tools/index.d.ts). + +## Why +Today the prompt rule (`ultrathink` → escalate) and a typed `/effort-` +share ONE slot (`turnMain`) with the model's `route` calls and the +`Skill(effort-*)` bridge: last writer wins, and a later skill load resets +the slot to the session default. A user's explicit level is therefore lost +mid-turn. The user chose FLOOR semantics: their level is a minimum for the +whole main turn; derived routes may go above it, never below; it also +lifts a lower sticky `/route`. + +## Precedence after the change (main loop only) +- model axis (unchanged order, floor last): `userMain ?? turnMain ?? turnFloor` + route's `model`, applied only with `mainModelSwitch` and the window guard. +- effort axis: `base = (userMain ?? turnMain)?.route.effort ?? e.effort`; + `effort = floored(base, turnFloor?.route.effort)`. +- `floored(effort, floor)`: no floor → `effort`; `effort` is a Level whose + LEVELS index ≥ the floor's → `effort`; otherwise (lower Level, a number, + or undefined) → `floor`. +- The haiku omission (`effort: undefined` when the model sent starts with + `claude-haiku`) still runs AFTER flooring. +- Sub-agent steps (`e.agentId` set) never read `turnFloor`. + +## Changes in register.ts (names as in the current file) +1. `State`: add `turnFloor: Routed | null` with the comment `user-explicit + level for this turn (prompt rule, typed /effort-): a floor, main loop + only`; reword the `turnMain` comment to `model route tool, skill table + row, Skill(effort-*) bridge; dropped at turn end`. `newState`: + `turnFloor: null`. +2. Helpers (new, small): `const rank = (l: Level): number => LEVELS.indexOf(l)`; + `function floored(effort: StepIn['effort'], floor: Level | undefined)` + per the rule above. A helper `floorLevel(st)` returning + `st.turnFloor?.route.effort` is allowed if it keeps call sites short. +3. `mainPlan`: compute `base` and `effort = floored(base, floorLevel(st))`; + `wanted = set?.route.model ?? st.turnFloor?.route.model`; the rest + (switch, window guard) unchanged. +4. `registerPrompt` / `prompt.submit`: the non-queued branch writes + `st.turnFloor = routed` (instead of `st.turnMain`); the queued branch + (`e.turnId !== undefined && e.wait`) keeps writing `st.pendingPrompt`. +5. `slashEffort`: write `st.turnFloor = { phase: skill, route: { effort: + level }, source: 'slash' }`; returned text line becomes `Effort floor + set by model-router for this turn: nothing below it runs.` + followed by `\n` + the original text (prepend, never replace). +6. `onSkillLoad` (main branch): `st.turnMain = null` unconditionally (the + slot no longer holds prompt or slash routes), then the table row as now. + `turnFloor` is never touched there. +7. `clearRoutes` (`/route clear`): also `st.turnFloor = null`. + `clearLoop` (model `route({clear})`, main branch): `turnMain` only, as now. +8. `endMainTurn`: `st.turnFloor = st.pendingPrompt; st.pendingPrompt = + null; st.turnMain = null;` then the existing resets. +9. Truthful answers (main branch only; agent branches unchanged): + - `effortBridge`: keep the sticky sentence when `st.userMain` is set; + else when the floor ranks above `level`: `model-router: + recorded, but the user's floor for this turn keeps main at ; + the skill text was not loaded.`; else the current sentence. + - `routedText`: keep the sticky branch; else when `p.route.effort` is + set and the floor ranks above it, print the effort as ` (user + floor; asked )`. +10. Display: `mainText` appends ` · floor ()` when `turnFloor` is + set (also after `main: session defaults`); `statusLine` appends + ` · floor `. + +## Tests in register.test.ts (keep every existing test; adapt only what +the new slot changes, e.g. the `ultrathink` test now expects the floor on +the `main:` line). Add at least six tests whose names contain `floor`, +using the existing boot helper, full typed inputs and bottom hooks, and +asserting on the `main:` line or on what the bottom `turn.step` hook +receives (drain the stream with `for await`, then `.result`): +- `floor: ultrathink survives a model route` — prompt `ultrathink` + (composer, `wait: false`, no `turnId`) then route tool `orchestrate` → + a main step with engine effort `high` reaches the bottom at `max`. +- `floor: a typed /effort-medium floors a lower route and allows a higher + one` — `$.skill.prompt({ skill: 'effort-medium', text: 'x' })` (no Skill + call in flight) → route tool `mechanical` → main step at `medium`; + then route tool `escalate` → main step at `max`. +- `floor: survives a skill load` — ultrathink, route tool `orchestrate`, + then a non-effort `Skill` call (bottom `tool.call` hook registered) → + main step at `max`, and `main:` line no longer names `orchestrate`. +- `floor: lifts a lower sticky route, then ends with the turn` — `/route + effort=low` then ultrathink → main step at `max`; fire a main + `turn.complete` → next main step at `low`. +- `floor: main only` — ultrathink, spawn `Explore` (bottom `agent.spawn` + hook returning an `agentId`), then a step for that `agentId` → reaches + the bottom at `medium`, not `max`. +- `floor: /route clear removes it` — ultrathink, `/route clear` → main + step keeps the engine effort. +Optional seventh: the bridge context line names the floor when it wins. + +## Constraints +- Style: ≤ 25 logic lines per function, 80 chars per line, no `any`, no + module-level mutable state, doc comments state intent. +- Do not touch: the agent axis, the spawn table, config loading, the + hardening (caps, warnOnce, safely), the route tool schema. +- Verify (paste outputs): `claude plugin validate .`, the contract's tsc + command, `claude plugin test .`, then from the repo root + `bash ~/.claude/lib/gates.sh run .claude/tasks/contracts/2026-10-08-model-router-floor-1835.md`. + +## Disposition +- honors BDR-115 (one writer per axis, calling-loop writes, truthful + answers) and the wave plan's routing rule (explicit user choice beats the + derived phase); supersedes the 1-A contract's one-slot precedence for + prompt and slash sources. +- LRN-206 (kit facts) applies to every new test. diff --git a/.claude/tasks/plans/2026-10-08-model-router-mod.md b/.claude/tasks/plans/2026-10-08-model-router-mod.md index 20f3832..b5ff7a6 100644 --- a/.claude/tasks/plans/2026-10-08-model-router-mod.md +++ b/.claude/tasks/plans/2026-10-08-model-router-mod.md @@ -95,6 +95,27 @@ does many different things inside one run. So: | verify | sonnet | xhigh | verifier, security-auditor | | mechanical | haiku | low | cp/mv, git bookkeeping, status collection, listing | +## Decisions 2026-10-08 (evening, user via AskUserQuestion) +- LOADING (supersedes "CLAUDE_CODE_PLUGIN_DIRS via settings.json + link.sh"): + PLUGIN_DIRS needs absolute paths, settings `env` has no `$HOME` + expansion, settings.json is tracked and shared across machines; a local + marketplace `add` writes an absolute path into settings.json too. Chosen: + tracked relative symlink `skills/` → `../mods/`; Claude Code + loads it as `@skills-dir`, in place (docs plugins/loading; probe in + an isolated HOME: listed, enabled, loaded). Repo scripts walking skills/ + glob `*/SKILL.md` or fixed names: unaffected. User asked why not + `~/.claude/mods`: Claude Code scans no such folder, a link there loads + nothing. +- PRECEDENCE: `ultrathink` (prompt rule) and a typed `/effort-` become a + FLOOR for the main turn: derived routes may go above, never below; it + lifts a lower sticky `/route`. Main loop only. +- settings.json: the user's uncommitted `/model` change (model → opus) is + theirs to manage; never staged. +- Live 2026-10-08: Skill(effort-low) bridge answered in place, next request + `low`; `ultrathink` turn on Opus ran at `max` (engine base `medium`); + Explore without params spawned on `claude-sonnet-5-5`, all 3 steps + `medium` (no spawn/step race observed). + ## Wave 0 — spike (dev-mods folder, hot reload, this session) - [x] W0.1 minimal mod: `/route` command, `route` tool, `turn.step` logging + rewrite, `agent.spawn` rewrite, `ultrathink` → max, spinner suffix diff --git a/.claude/tasks/plans/2026-10-08-model-router-wiring-1835.md b/.claude/tasks/plans/2026-10-08-model-router-wiring-1835.md new file mode 100644 index 0000000..b136601 --- /dev/null +++ b/.claude/tasks/plans/2026-10-08-model-router-wiring-1835.md @@ -0,0 +1,106 @@ +# PLAN — model-router wave 1-B2: active in every session, tests, doctor (dispatch-ready) +Contract: .claude/tasks/contracts/2026-10-08-model-router-wiring-1835.md +Repo root: /Users/b.chanot/Documents/claude (branch feature/model-router-mod). + +## Facts this plan rests on (verified 2026-10-08) +- Claude Code loads a folder holding `.claude-plugin/plugin.json` under + `~/.claude/skills/` as `@skills-dir`, in place, live at the next + session start or `/reload-plugins` (docs: plugins/loading "In-place and + copied plugins"; probe in an isolated HOME: listed, enabled, loaded). +- `~/.claude/skills` is already a symlink to the repo's `skills/` (link.sh). +- Repo scripts that walk `skills/` glob `*/SKILL.md` or fixed paths + (doctor.sh, lib/skill-routing-census.py, the census suites); + lib/profile.sh only moves entries named in a profile. An entry without + SKILL.md is never counted, moved or flagged. +- The engine lays `/tsconfig.json` (extends the types) and + `/.claude-plugin/types/` (own `.gitignore` holding `*`) when a mod + loads; today `mods/model-router/tsconfig.json` shows as untracked. +- settings.json carries the user's uncommitted `/model` change: never + stage, edit or restore it. + +## Files +- [ ] `skills/model-router` — new RELATIVE symlink: from the repo root, + `ln -s ../mods/model-router skills/model-router`. Nothing else in skills/. +- [ ] `.gitignore` — append a block: + ``` + # mods/: files the engine lays beside a loaded mod (editor types) + mods/*/tsconfig.json + mods/*/.claude-plugin/types/ + ``` + Check first that no existing pattern ignores `skills/model-router` or + the tracked mod files (contract AC1/AC2 oracles). +- [ ] `lib/tests/mods.test.sh` — new suite, style of the existing suites + (read lib/tests/effort-pins.test.sh first and mirror its header, helpers + and summary). Behaviour: + - `ROOT="${MODS_ROOT:-}"`. + - Collect `$ROOT/mods/*/.claude-plugin/plugin.json`; none → FAIL + ("no mod found") so the suite can never pass vacuously. + - Per mod dir ``: (1) the manifest `name` (python3 json, argv — + never string-spliced) equals the folder name; (2) `$ROOT/skills/` + is a symlink whose `readlink` is exactly `../mods/`; (3) when + `command -v claude` succeeds: `claude plugin validate "$ROOT/mods/"` + prints `Validation passed` and no `warning` (case-insensitive); + (4) same condition: `claude plugin test "$ROOT/mods/"` exits 0. + - `claude` absent → one `SKIP: claude CLI not found — validate/test not + run` line; checks (1)-(2) still run and decide the exit code. + - Exit 1 on any failure, 0 otherwise; one PASS/FAIL line per check and a + final count line. + - shellcheck clean. No network, no writes outside a `mktemp -d` if any + scratch is needed (none expected). +- [ ] `doctor.sh` — new section `── Mods ──`, placed right after the + "Vendored skills" section (read lines 120-160 first; mirror its + `echo ""` / heading / pass-warn-fail-info style). For each + `$REPO/mods/*/` holding `.claude-plugin/plugin.json` (`` = folder): + - link `$HOME/.claude/skills/`: `readlink -f` equal to + `$REPO/mods/` → `pass "mod : loading link ~/.claude/skills/"`; + missing → `fail "mod : ~/.claude/skills/ MISSING — git checkout skills/, then make link"`; + elsewhere → `warn`. Do NOT call `check_symlink` (it feeds the core-link + counter `_LINK_PASS` / `_EXPECTED_LINKS`). + - `command -v claude` → `claude plugin list --json` parsed with python3 + (argv/stdin, no splicing): id `@skills-dir` with `enabled: true` + → `pass "mod : loaded as @skills-dir"`; present but + disabled → `warn "... disabled (enabledPlugins \"@skills-dir\": false)"`; + absent → `warn "... not listed — new session or /reload-plugins"`. + `claude` missing → `info "claude CLI not found — load state not checked"`. + - `$HOME/.claude/.json` present → `python3 -m json.tool` (quiet) + → `pass "mod : override ~/.claude/.json parses"` or + `fail "... invalid JSON"`; absent → nothing. + - No mod at all → `info "no mods"`. +- [ ] `CLAUDE.md` (project, repo root) — new section `## mods/ — function-hooks + plugins (Claude Code mods)` placed after the graphify section, terse + English in the file's own style, at most ~14 lines, covering: what lives + in `mods//`; it loads through the tracked relative symlink + `skills/` → `../mods/` as `@skills-dir` (in place, live + at the next session or `/reload-plugins`); why not + `CLAUDE_CODE_PLUGIN_DIRS` (absolute path, settings `env` has no `$HOME` + expansion, settings.json is tracked) nor a local marketplace (its `add` + writes an absolute path into settings.json); engine-laid + `tsconfig.json` + `.claude-plugin/types/` are gitignored; optional user + config `~/.claude/.json`; tests `make test suite=lib/tests/mods.test.sh` + (validate + `claude plugin test`); turn a mod off with + `"@skills-dir": false` in `enabledPlugins`; a dev copy loaded with + `--plugin-dir` or the hot-reload folder shadows the skills-dir copy + (same name, session-only wins). + +## Verify (executor pastes outputs) +`ls -l skills/model-router`; `git check-ignore -v mods/model-router/tsconfig.json`; +`make test suite=lib/tests/mods.test.sh`; the contract AC3 positive control; +`bash doctor.sh | sed -n '/── Mods ──/,/^$/p'`; `shellcheck lib/tests/mods.test.sh doctor.sh`; +`git status --short` (settings.json still ` M`, untouched); then from the +repo root `bash ~/.claude/lib/gates.sh run .claude/tasks/contracts/2026-10-08-model-router-wiring-1835.md`. + +## Edge cases +- The engine-laid `mods/model-router/tsconfig.json` already exists on disk: + after the `.gitignore` change it must disappear from `git status`. +- doctor runs without `claude` on PATH (Linux box): info line, no failure. +- A second mod later: the suite and doctor loop over `mods/*/` already. +- A hot-reload or `--plugin-dir` copy of the same mod shadows the + skills-dir copy in that session; doctor reads `claude plugin list` from a + fresh process, which sees only the skills-dir copy. + +## Disposition +- honors BDR-115 (mod in `mods/`, single source); amends its "Load:" line + (PLUGIN_DIRS → skills-dir link), to be recorded at capitalize. +- honors the destructive-tools rule: no recursive delete, no transfer + tool; LRN-150/LRN-171 shell hygiene (`command grep` where a shim can + interfere is not needed here: plain bash). From 868a7f05153a41ef129c61f1c3748f2b2731ce20 Mon Sep 17 00:00:00 2001 From: bchanot Date: Fri, 9 Oct 2026 09:24:32 +0200 Subject: [PATCH 09/22] chore(tasks): model-router B1/B2 plans r2 after the 6-lens challenge round --- .../2026-10-08-model-router-floor-1835.md | 15 +++-- .../2026-10-08-model-router-wiring-1835.md | 20 ++++-- .../2026-10-08-model-router-floor-1835.md | 64 +++++++++++++++++++ .../2026-10-08-model-router-wiring-1835.md | 60 +++++++++++++++++ 4 files changed, 146 insertions(+), 13 deletions(-) diff --git a/.claude/tasks/contracts/2026-10-08-model-router-floor-1835.md b/.claude/tasks/contracts/2026-10-08-model-router-floor-1835.md index 53daccc..30943eb 100644 --- a/.claude/tasks/contracts/2026-10-08-model-router-floor-1835.md +++ b/.claude/tasks/contracts/2026-10-08-model-router-floor-1835.md @@ -10,24 +10,27 @@ User, same turn: "continu avec Opus en /ultrathink". ## CLARIFICATIONS Q: which loops does the floor cover? / A: the MAIN loop only ("le tour" = the user's turn); sub-agents keep their own routes and pins. [orchestrator — derived, stated to the user] -Q: model axis / A: the floor constrains effort only. A floor route's model (none by default: ultrathink → escalate carries none) applies on main only when neither the sticky nor the turn route names a model, and only with the switch on. [orchestrator — internal] +Q: model axis / A: one rule: `userMain?.route.model ?? turnMain?.route.model ?? turnFloor?.route.model`, applied only with the switch on (r2). [orchestrator — internal] +Q (r2): default vs minimum / A: the user's level is the turn's default when no sticky or turn route names an effort AND its minimum; a typed `/effort-low` therefore still lowers an unrouted turn. Derived from the chosen option ("Ton niveau est un minimum pour tout le tour") plus the challenge finding that a pure floor would make `/effort-low` a no-op. [orchestrator — r2] +Q (r2): mid-turn prompt / A: applied to the running turn AND kept for the next (`wait` ignored, the engine queues either way). [orchestrator — r2] +Q (r2): per-machine off switch / A: `"enabled": false` in `~/.claude/model-router.json` (untracked); an `enabledPlugins` entry would dirty the tracked settings.json on every machine. [orchestrator — r2, from the user's "configurable"] Q: numeric or absent engine effort / A: the floor level replaces it (the user's explicit level wins over an unknown budget); the haiku effort omission still applies after flooring. [orchestrator — internal] Q: who clears the floor / A: main turn end (a queued prompt's floor is then promoted), `/route clear`, `/route off` (pass-through). A model `route({clear})`, a `Skill(effort-*)` or any skill load never touches it. [orchestrator — derived from "jamais descendre en dessous"] ## ACCEPTANCE CRITERIA -1. Suite green with the new tests: `claude plugin test` passes with at least 19 `test(` calls; `claude plugin validate` passes with no warning; no line over 80 chars; no `any` type. - CHECK: cd mods/model-router && out=$(claude plugin test . 2>&1); rc=$?; echo "$out" | tail -n 3; [ $rc -eq 0 ] && [ "$(grep -cE '^\s*test\(' hooks/register.test.ts)" -ge 19 ] && v=$(claude plugin validate . 2>&1) && echo "$v" | grep -q 'Validation passed' && ! echo "$v" | grep -qi 'warning' && ! grep -nE '.{81,}' hooks/register.ts hooks/register.test.ts && ! grep -nE ':\s*any\b||as any\b' hooks/register.ts && echo FLOOR-SUITE-OK +1. Suite green with the new tests: `claude plugin test` passes with at least 22 `test(` calls; `claude plugin validate` passes with no warning; no line over 80 chars; no `any` type. + CHECK: cd mods/model-router && out=$(claude plugin test . 2>&1); rc=$?; echo "$out" | tail -n 3; [ $rc -eq 0 ] && [ "$(grep -cE '^\s*test\(' hooks/register.test.ts)" -ge 22 ] && v=$(claude plugin validate . 2>&1) && echo "$v" | grep -q 'Validation passed' && ! echo "$v" | grep -qi 'warning' && ! grep -nE '.{81,}' hooks/register.ts hooks/register.test.ts && ! grep -nE ':\s*any\b||as any\b' hooks/register.ts && echo FLOOR-SUITE-OK EXPECT: FLOOR-SUITE-OK EVIDENCE: pending 2. Type-check clean against this build's declarations. CHECK: T=/Users/b.chanot/.claude/dev-mods/385f7190-70f5-4bdd-b0d8-e4566cd412fd/model-router/.claude-plugin/types; [ -d "$T" ] || T=/Users/b.chanot/Documents/claude/mods/model-router/.claude-plugin/types; W=$(mktemp -d) && printf '{"compilerOptions":{"target":"es2023","lib":["es2023"],"types":[],"module":"esnext","moduleResolution":"bundler","strict":true,"noUncheckedIndexedAccess":true,"noEmit":true,"skipLibCheck":true,"jsx":"react","jsxFactory":"h","jsxFragmentFactory":"Fragment"},"include":["%s/claude-code/index.d.ts","%s/claude-code-tools/index.d.ts","%s/hooks"]}' "$T" "$T" "$PWD/mods/model-router" > "$W/tsconfig.json" && (cd "$W" && npx --yes -p typescript@5 tsc -p tsconfig.json) && echo TSC-OK EXPECT: TSC-OK EVIDENCE: pending -3. The floor is its own slot: `turnFloor` is declared in `State`, initialised in `newState`, written by the prompt rule and by a typed `/effort-`, cleared by `/route clear` and at main turn end; the suite carries at least 6 tests whose name contains `floor`. - CHECK: cd mods/model-router/hooks && [ "$(grep -c 'turnFloor' register.ts)" -ge 6 ] && [ "$(grep -cE "^\s*test\('[^']*floor" register.test.ts)" -ge 6 ] && echo FLOOR-SLOT-OK +3. The floor is its own slot: `turnFloor` is declared in `State`, initialised in `newState`, written by the prompt rule and by a typed `/effort-`, cleared by `/route clear` and at main turn end; the suite carries at least 8 tests whose name contains `floor`; one helper `mainEffort` decides the main effort; the config key `enabled` exists. + CHECK: cd mods/model-router/hooks && [ "$(grep -c 'turnFloor' register.ts)" -ge 6 ] && [ "$(grep -cE "^\s*test\('[^']*floor" register.test.ts)" -ge 8 ] && grep -q 'mainEffort' register.ts && grep -q 'enabled' register.ts && echo FLOOR-SLOT-OK EXPECT: FLOOR-SLOT-OK EVIDENCE: pending -4. Judged by reading: effective main effort = the higher (LEVELS order) of the floor and `(userMain ?? turnMain)?.route.effort ?? e.effort`, the floor replacing a numeric or absent engine effort; the floor never applies to a sub-agent step; `turnMain` only ever holds 'model' or 'skill' sources; the prompt rule writes `turnFloor` (or `pendingPrompt` when typed mid-turn with `wait`), promoted into `turnFloor` at main turn end; a non-effort skill load resets `turnMain` only; a model `route({clear})` clears `turnMain` only; every answer that the floor overrides says so truthfully (Skill bridge context, route tool text, `/effort-` text); `/route show` and the status line display the floor; every criterion of `.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md` still holds; no function over 25 logic lines. +4. Judged by reading: effective main effort comes from ONE helper `mainEffort` used by `mainPlan` and by every answer text: base `(userMain ?? turnMain)?.route.effort ?? turnFloor?.route.effort ?? e.effort`, then floored by `turnFloor` (LEVELS order; a numeric or absent value is replaced); the floor never applies to a sub-agent step; `turnMain` only ever holds 'model' or 'skill' sources; the prompt rule writes `turnFloor` (keeping the higher of two) and, when typed mid-turn (`turnId` set, `wait` ignored), also `pendingPrompt`, promoted into `turnFloor` at main turn end; `"enabled": false` in the override file makes every hook pass through after each config load (`/route on` re-enables for the session); a non-effort skill load resets `turnMain` only; a model `route({clear})` clears `turnMain` only; every answer that the floor overrides says so truthfully (Skill bridge context, route tool text, `/effort-` text); `/route show` and the status line display the floor; every criterion of `.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md` still holds; no function over 25 logic lines. ## FILE SCOPE mods/model-router/hooks/register.ts · mods/model-router/hooks/register.test.ts diff --git a/.claude/tasks/contracts/2026-10-08-model-router-wiring-1835.md b/.claude/tasks/contracts/2026-10-08-model-router-wiring-1835.md index fac021b..6f887de 100644 --- a/.claude/tasks/contracts/2026-10-08-model-router-wiring-1835.md +++ b/.claude/tasks/contracts/2026-10-08-model-router-wiring-1835.md @@ -19,18 +19,22 @@ Q: doctor scope / A: per mod: the loading link resolves into the repo mod dir; ` CHECK: [ -L skills/model-router ] && [ "$(readlink skills/model-router)" = "../mods/model-router" ] && ! git check-ignore -q skills/model-router && echo LINK-OK EXPECT: LINK-OK EVIDENCE: pending -2. Engine-laid files are ignored: a mod's root `tsconfig.json` and anything under its `.claude-plugin/types/`; the tracked mod files are not ignored. - CHECK: git check-ignore -q mods/model-router/tsconfig.json && git check-ignore -q mods/model-router/.claude-plugin/types/claude-code/index.d.ts && ! git check-ignore -q mods/model-router/hooks/register.ts && ! git check-ignore -q mods/model-router/.claude-plugin/plugin.json && echo IGNORE-OK +2. Engine-laid files are ignored for ANY mod: a mod's root `tsconfig.json` (root `.gitignore`; the `.claude-plugin/types/` folder ignores itself); the tracked mod files are not ignored. + CHECK: git check-ignore -q mods/model-router/tsconfig.json && git check-ignore -q mods/zz-future/tsconfig.json && ! git check-ignore -q mods/model-router/hooks/register.ts && ! git check-ignore -q mods/model-router/.claude-plugin/plugin.json && [ -z "$(git status --short mods/)" ] && echo IGNORE-OK EXPECT: IGNORE-OK EVIDENCE: pending -3. `lib/tests/mods.test.sh` passes on the repo and fails on a fixture that lacks the loading link (positive control through `MODS_ROOT`). - CHECK: make test suite=lib/tests/mods.test.sh >/dev/null 2>&1 && W=$(mktemp -d) && mkdir -p "$W/mods" "$W/skills" && cp -R mods/model-router "$W/mods/" && ! MODS_ROOT="$W" bash lib/tests/mods.test.sh >/dev/null 2>&1 && ln -s ../mods/model-router "$W/skills/model-router" && MODS_ROOT="$W" bash lib/tests/mods.test.sh >/dev/null 2>&1 && echo MODS-SUITE-OK +3. `lib/tests/mods.test.sh` passes on the repo, fails on a fixture that lacks the loading link, fails on an empty `mods/`, and SKIPs (exit 0) the CLI checks when `claude plugin test` is unavailable (probe by capability, PATH-shadowed `claude` in the control). + CHECK: make test suite=lib/tests/mods.test.sh >/dev/null 2>&1 && W=$(mktemp -d) && mkdir -p "$W/mods" "$W/skills" "$W/bin" && cp -R mods/model-router "$W/mods/" && ! MODS_ROOT="$W" bash lib/tests/mods.test.sh >/dev/null 2>&1 && ln -s ../mods/model-router "$W/skills/model-router" && MODS_ROOT="$W" bash lib/tests/mods.test.sh >/dev/null 2>&1 && printf '#!/bin/sh\nexit 1\n' > "$W/bin/claude" && chmod +x "$W/bin/claude" && PATH="$W/bin:$PATH" MODS_ROOT="$W" bash lib/tests/mods.test.sh 2>&1 | grep -q '^SKIP' && E=$(mktemp -d) && mkdir -p "$E/mods" "$E/skills" && ! MODS_ROOT="$E" bash lib/tests/mods.test.sh >/dev/null 2>&1 && echo MODS-SUITE-OK EXPECT: MODS-SUITE-OK EVIDENCE: pending -4. `doctor.sh` prints a `── Mods ──` section with a ✓ line for the model-router loading link. - CHECK: out=$(bash doctor.sh 2>&1); echo "$out" | sed -n '/── Mods ──/,/^$/p' | grep -q '✓.*model-router' && echo DOCTOR-MODS-OK +4. `doctor.sh` prints a `── Mods ──` section with a ✓ line for model-router; with the link absent (HOME pointed at a scratch `.claude` whose `skills/` lacks the link) the section prints an info line, doctor reaches its summary and exits 0 for that section's sake (no new error). + CHECK: out=$(bash doctor.sh 2>&1); echo "$out" | sed -n '/── Mods ──/,/^$/p' | grep -q '✓.*model-router' && H=$(mktemp -d) && mkdir -p "$H/.claude/skills" && o2=$(HOME="$H" bash doctor.sh 2>&1); echo "$o2" | sed -n '/── Mods ──/,/^$/p' | grep -qi 'not linked' && echo "$o2" | grep -q '═══' && echo DOCTOR-MODS-OK EXPECT: DOCTOR-MODS-OK EVIDENCE: pending +4b. The mod is enabled through the tracked link in a FRESH process: `claude plugin list --json` lists `model-router@skills-dir` with `enabled: true` (run after the dev-mods link is removed, see W6). + CHECK: claude plugin list --json 2>/dev/null | python3 -c 'import json,sys; rows=json.load(sys.stdin); ok=any(r.get("id")=="model-router@skills-dir" and r.get("enabled") is True for r in rows); sys.exit(0 if ok else 1)' && echo LOADED-OK + EXPECT: LOADED-OK + EVIDENCE: pending 5. `CLAUDE.md` has a `## mods/` section naming the `skills/` relative symlink, the `@skills-dir` origin, why not `CLAUDE_CODE_PLUGIN_DIRS`, the gitignored engine-laid files, `~/.claude/.json`, the suite command and how to turn a mod off. CHECK: grep -q '^## mods/' CLAUDE.md && grep -q '@skills-dir' CLAUDE.md && grep -q 'CLAUDE_CODE_PLUGIN_DIRS' CLAUDE.md && grep -q 'mods.test.sh' CLAUDE.md && grep -q '.json' CLAUDE.md && grep -q '@skills-dir": false' CLAUDE.md && echo CLAUDEMD-OK EXPECT: CLAUDEMD-OK @@ -39,7 +43,9 @@ Q: doctor scope / A: per mod: the loading link resolves into the repo mod dir; ` CHECK: shellcheck lib/tests/mods.test.sh doctor.sh && make test suite=lib/tests/doctrine-citers.test.sh >/dev/null 2>&1 && echo HEALTH-OK EXPECT: HEALTH-OK EVIDENCE: pending -7. Judged by reading: no change to settings.json, link.sh or any install script; the user's settings.json working-tree diff is untouched; the suite SKIPs (explicit SKIP line, exit 0 for that part) only the `claude`-dependent checks when `claude` is absent, and fails when no mod is found at all; doctor's new section never increments the core-link counter (`_LINK_PASS`) and never fails on a missing `claude`; the CLAUDE.md section is terse English matching the file's style. +7. Judged by reading: no change to settings.json, link.sh or any install script; the suite probes the CAPABILITY (`claude plugin test --help`), bounds every CLI call in time, captures `2>&1`, SKIPs with a reason, fails when no mod is found; doctor's section is fail-soft under `set -euo pipefail` (existence test before readlink, `-ef` comparison, one guarded `claude plugin list --json`, python exits 0 with `unknown` on any parse error), never increments `_LINK_PASS`, says "enabled" not "loaded", treats a missing link as info; the link step is idempotent; CLAUDE.md names the per-machine `"enabled": false` switch, the tracked-settings cost of `enabledPlugins`, and the dev-copy shadowing rule; the CLAUDE.md section is terse English matching the file's style. +Q (r2): ordering / A: this contract runs after the floor contract (`2026-10-08-model-router-floor-1835`) is committed and green. [orchestrator] +Q (r2): update-all `claude plugin update` over `@skills-dir` / A: accepted residual (one recurring warn), logged in TODO; out of FILE SCOPE. [orchestrator] ## FILE SCOPE skills/model-router (new symlink) · .gitignore · lib/tests/mods.test.sh (new) · doctor.sh · CLAUDE.md diff --git a/.claude/tasks/plans/2026-10-08-model-router-floor-1835.md b/.claude/tasks/plans/2026-10-08-model-router-floor-1835.md index 6058c60..8c3019b 100644 --- a/.claude/tasks/plans/2026-10-08-model-router-floor-1835.md +++ b/.claude/tasks/plans/2026-10-08-model-router-floor-1835.md @@ -106,3 +106,67 @@ Optional seventh: the bridge context line names the floor when it wins. derived phase); supersedes the 1-A contract's one-slot precedence for prompt and slash sources. - LRN-206 (kit facts) applies to every new test. + +## r2 — challenge round (3 lenses, 0 BLOCKER, 4 MAJOR): BINDING, overrides the sections above where they conflict +R1. ONE decision helper, used by `mainPlan` AND by every answer text: + `mainEffort(st, engine: StepIn['effort'])` → `{ effort, by }` with + `by` ∈ `'floor' | 'sticky' | 'turn' | 'engine'`. + `base = (st.userMain ?? st.turnMain)?.route.effort + ?? st.turnFloor?.route.effort ?? engine` + `effort = floored(base, st.turnFloor?.route.effort)`; `by = 'floor'` + when the floor raised or supplied the value, else the slot it came from. + The user's level is therefore BOTH the turn's default (when no sticky + or turn route names an effort) AND its minimum: a typed `/effort-low` + lowers a turn that has no route (engine `high` → `low`), and a route + can still go higher. No text function compares ranks on its own. +R2. Model axis, one rule written once (contract updated): + `st.userMain?.route.model ?? st.turnMain?.route.model ?? st.turnFloor?.route.model`, + switch and window guard unchanged. +R3. Prompt rule with `e.turnId !== undefined` (typed while a turn runs; + `wait` is IGNORED: the engine queues every mid-turn prompt either way): + write the floor NOW (higher of the existing floor and the new one) + AND set `pendingPrompt` to it, so the turn that reads the prompt has it + whichever it is. No `turnId` → write the floor (higher of two). + `endMainTurn` promotes `pendingPrompt` into `turnFloor`. Two floors in + one turn always keep the higher one (prompt rule and typed slash). +R4. Truthful texts, all phrased from `mainEffort` (main branch only): + - Skill bridge, route tool, typed `/effort-`: when `by === 'floor'` + and the result differs from what was asked, name the floor and its + source (`ultrathink rule` or `typed /effort-`) and add + `/route clear to drop it`; when `by === 'sticky'`, the sticky + sentence; the old fixed "sticky wins" sentences go. + - `/effort-` text: `Effort set by model-router for the main loop + this turn (minimum; a higher route still applies).` plus the floor + or sticky outcome when one changes it. + - main loop on a haiku model: print `effort - (haiku takes none)`. + - model `route({clear})` on main with a floor set: append `; user floor + () still holds — /route clear drops it`. +R5. Display: `mainText` / `statusLine` show ` · floor ` only when the + floor carries an effort AND the router is on; the `skill.prompt` hook + calls `refresh($, st)` after the slash write. +R6. Persistent per-machine off switch (wiring challenge, user's "configurable"): + config key `enabled: boolean` (default `true`) in + `~/.claude/model-router.json` (untracked, per machine). `false` → + `st.off = true` after every config load (session start, `/route + reload`); `/route on` re-enables for the session only; `show` and the + status line say `off (config)` vs `off`. Merged with `pickBool` like + the other scalars; `DEFAULT_CONFIG.enabled = true`. +R7. Tests (replace the list above where it differs): `runStep` takes a + full `TurnStepInput` (from 'claude-code'); every floor test steps with + engine effort `high` (or `xhigh`); every `test('…', async (` line ≤ 80 + chars with `floor` in the single-line name. At least 8 floor tests: + ultrathink survives a model route · typed /effort-medium clamps low, + lets max pass · typed /effort-low lowers an unrouted turn (engine high + → low) · survives a skill load · lifts a lower sticky then ends with + the turn · main only (agent step unaffected) · /route clear removes it + · mid-turn prompt (turnId + wait) is applied now AND promoted after the + main turn.complete · mandatory text test: sticky `/route effort=low`, + ultrathink, route tool `plan` → the answer names the floor. + `enabled: false` cannot be reached in the kit (no fs, LRN-206): cover + the off path through `/route off` and say so in a comment. +R8. Residuals accepted (logged in TODO, not built): floor expiry depends on + a main `turn.complete` reaching this mod (another plugin answering it + without `next` would keep it); `skill.prompt` cannot tell a typed + `/effort-` from a sub-agent preload (no agentId; no repo agent + preloads one); an incidental "ultrathink" in pasted text floors the + turn (mitigated by R4 naming the source and the `/route clear` hint). diff --git a/.claude/tasks/plans/2026-10-08-model-router-wiring-1835.md b/.claude/tasks/plans/2026-10-08-model-router-wiring-1835.md index b136601..9b05f82 100644 --- a/.claude/tasks/plans/2026-10-08-model-router-wiring-1835.md +++ b/.claude/tasks/plans/2026-10-08-model-router-wiring-1835.md @@ -98,6 +98,66 @@ repo root `bash ~/.claude/lib/gates.sh run .claude/tasks/contracts/2026-10-08-mo skills-dir copy in that session; doctor reads `claude plugin list` from a fresh process, which sees only the skills-dir copy. +## r2 — challenge round (3 lenses, 0 BLOCKER, 5 MAJOR): BINDING, overrides the sections above where they conflict +W1. ORDER: this plan runs AFTER the floor plan (B1) is committed and green + on the same branch: the suite and doctor test whatever register.ts is + on disk. +W2. `.gitignore`: add ONLY `mods/*/tsconfig.json` with the comment + `# mods/: the engine lays tsconfig.json beside a loaded mod; its + .claude-plugin/types/ ignores itself`. (The types folder carries its + own `.gitignore` holding `*`.) +W3. Link step idempotent: `[ -L skills/model-router ] || ln -s + ../mods/model-router skills/model-router` (a bare `ln -s` re-run would + create a nested link inside the mod). +W4. `lib/tests/mods.test.sh` fail-soft and bounded: + - capability probe, not presence: `command -v claude` AND `claude plugin + test --help >/dev/null 2>&1`; otherwise ONE `SKIP: claude plugin test + unavailable () — validate/test not run` line, checks (1)-(2) + still decide the exit code; + - `claude plugin validate` and `claude plugin test` captured with `2>&1`; + the validate verdict is the line matching `Validation passed`, with + `warning` searched only in that captured output; + - every CLI call bounded: `timeout 120` when available (coreutils / + `gtimeout`), else a background-and-wait guard; a timeout is a FAIL + naming it; + - no mod found → FAIL (never vacuous). +W5. doctor `── Mods ──` fail-soft under `set -euo pipefail`: + - `[ -L "$link" ] || [ -e "$link" ]` BEFORE any readlink; compare with + `[ "$link" -ef "$REPO/mods/" ]` (handles logical vs physical + repo paths), never string equality on `readlink -f`; + - a missing link is `info "mod : not linked (skills/ absent) + — git checkout skills/ if wanted"`, NOT `fail` (a user may + remove the link on purpose; doctor red forever would break + update-all's final doctor run); + - ONE `claude plugin list --json` call before the loop, inside + `if ! out=$(claude plugin list --json 2>/dev/null); then warn "mods: + claude plugin list failed — load state not checked"; out=""; fi`; the + python3 parse reads stdin, exits 0 always, prints `enabled|disabled| + absent|unknown` per name (any parse error → `unknown`); + - wording: `pass "mod : enabled as @skills-dir"` (not + "loaded": the list proves enablement, not a successful load); + `disabled` → warn naming `"@skills-dir": false`; `absent` → + `warn "mod : not listed as @skills-dir — run: claude plugin + validate mods/ (policy, manifest or name conflict)"` (a fresh + process rescans skills/, so a restart changes nothing); `unknown` → + warn "list output not understood"; + - `claude` missing → nothing (doctor's Prerequisites section already + fails on it); no override-file JSON check (the mod validates its own + config and logs at session start). +W6. CLAUDE.md `## mods/` also says: the only per-machine off switch is + `"enabled": false` in `~/.claude/.json` (untracked); an + `enabledPlugins` `"@skills-dir": false` entry works too but lands + in the TRACKED settings.json, so it dirties every machine's tree; and + that a hot-reload / `--plugin-dir` copy of the same name shadows the + skills-dir copy for that session (docs plugins/loading "Name + conflicts"), so the dev link in `~/.claude/dev-mods//` must + be removed before `/reload-plugins` is read as a test of the skills-dir + path. +W7. `update-all.sh` runs `claude plugin update` over every listed plugin + (lines ~606-618): a `@skills-dir` entry will produce one recurring + warn there. Accepted residual, logged in TODO (an update-all edit is + out of this contract's FILE SCOPE). + ## Disposition - honors BDR-115 (mod in `mods/`, single source); amends its "Load:" line (PLUGIN_DIRS → skills-dir link), to be recorded at capitalize. From 1ff608a68cf8509d64cbb8e50b4defb986c3ca18 Mon Sep 17 00:00:00 2001 From: bchanot Date: Fri, 9 Oct 2026 10:12:28 +0200 Subject: [PATCH 10/22] =?UTF-8?q?feat(mods):=20model-router=20user=20effor?= =?UTF-8?q?t=20floor=20=E2=80=94=20ultrathink=20and=20typed=20/effort-?= =?UTF-8?q?=20set=20the=20main=20turn's=20default=20and=20minimum?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit One decision helper (mainEffort) feeds the plan and every answer text; per-axis precedence (sticky > turn route > floor > engine); a mid-turn prompt floors the running turn and the next; per-machine kill switch "enabled": false in ~/.claude/model-router.json, kept across /clear and across a failed reload; typed /effort- attested at prompt.submit so a sub-agent preload cannot floor the main loop. 30 plugin tests. --- mods/model-router/hooks/register.test.ts | 216 +++++++++++++++++- mods/model-router/hooks/register.ts | 276 +++++++++++++++++++---- 2 files changed, 439 insertions(+), 53 deletions(-) diff --git a/mods/model-router/hooks/register.test.ts b/mods/model-router/hooks/register.test.ts index 030970b..b1074d5 100644 --- a/mods/model-router/hooks/register.test.ts +++ b/mods/model-router/hooks/register.test.ts @@ -1,6 +1,6 @@ import { test, expect } from 'claude-code/testing' import type { Engine } from 'claude-code/testing' -import type { On } from 'claude-code' +import type { On, TurnStepInput } from 'claude-code' /** Fires session.start so the mod loads its config and registers /route. */ async function boot($: Engine, on: On): Promise { @@ -84,7 +84,7 @@ test('/route bogus names the phases', async ($, on) => { } }) -test('ultrathink in a prompt sets escalate on main', async ($, on) => { +test('ultrathink in a prompt sets a max floor on main', async ($, on) => { on('prompt.submit', ($, e) => ({ text: e.text })) await boot($, on) await $.prompt.submit({ @@ -92,7 +92,7 @@ test('ultrathink in a prompt sets escalate on main', async ($, on) => { wait: false, origin: { kind: 'composer' }, }) - expect(mainLine(await route($, 'show'))).toContain('prompt escalate') + expect(mainLine(await route($, 'show'))).toContain('floor max') }) test('/route model: alias resolved, id passed, typo refused', async ( @@ -169,7 +169,7 @@ const stepInput = (agentId?: string) => ({ /** Streams one turn.step to its end; the hooks run as the chunks flow. */ async function runStep( $: Engine, - input: ReturnType, + input: TurnStepInput, ): Promise { const stream = $.turn.step(input) for await (const _chunk of stream) { @@ -238,3 +238,211 @@ test('a rule only scans the first 4096 chars of a prompt', async ($, on) => { }) expect(mainLine(await route($, 'show'))).toContain('session defaults') }) + +// ---- user effort floor ------------------------------------------------- + +const ROUTE_TOOL = 'mcp__model-router__route' + +/** A main (or agent) step as the engine would make it: engine effort high. */ +const highStep = (agentId?: string): TurnStepInput => ({ + ...stepInput(agentId), + effort: 'high' as const, +}) + +/** Types `ultrathink` in the composer, idle or over a running turn. */ +async function ultrathink($: Engine, turnId?: string): Promise { + await $.prompt.submit({ + text: 'ultrathink please', + wait: turnId !== undefined, + origin: { kind: 'composer' }, + ...(turnId === undefined ? {} : { turnId }), + }) +} + +/** Ends a main turn: the floor's life is bounded by this event. */ +async function endTurn($: Engine): Promise { + await $.turn.complete({ + turnId: 'u1', + answer: '', + durationMs: 1, + isAborted: false, + reason: 'answer', + }) +} + +/** One main step; returns the effort that reached the bottom hook. */ +async function stepEffort($: Engine, seen: Seen[]): Promise { + await runStep($, highStep()) + return seen[seen.length - 1]?.effort +} + +/** Boots with a bottom step recorder and prompt/turn hooks in place. */ +async function bootFloor($: Engine, on: On): Promise { + const seen: Seen[] = [] + recordSteps(on, seen) + on('prompt.submit', ($, e) => ({ text: e.text })) + on('turn.complete', () => ({ text: '' })) + on('tool.call', { tool: 'Skill' }, () => ({ + result: { success: true, commandName: 'other' }, + })) + on('agent.spawn', ($, e) => ({ + model: e.model ?? e.parentModel, + agentId: 'a1', + })) + await boot($, on) + return seen +} + +test('floor: ultrathink survives a model route', async ($, on) => { + const seen = await bootFloor($, on) + await ultrathink($) + await $.tool.call({ tool: ROUTE_TOOL, phase: 'orchestrate' }) + expect(await stepEffort($, seen)).toBe('max') +}) + +test('floor: typed /effort-medium clamps low, lets max pass', async ( + $, on) => { + const seen = await bootFloor($, on) + await $.skill.prompt({ skill: 'effort-medium', text: 'x' }) + await $.tool.call({ tool: ROUTE_TOOL, phase: 'mechanical' }) + expect(await stepEffort($, seen)).toBe('medium') + await $.tool.call({ tool: ROUTE_TOOL, phase: 'escalate' }) + expect(await stepEffort($, seen)).toBe('max') +}) + +test('floor: typed /effort-low lowers an unrouted turn', async ($, on) => { + const seen = await bootFloor($, on) + const out = await $.skill.prompt({ skill: 'effort-low', text: 'x' }) + expect(out.text).toContain('minimum') + expect(await stepEffort($, seen)).toBe('low') +}) + +test('floor: survives a skill load', async ($, on) => { + const seen = await bootFloor($, on) + await ultrathink($) + await $.tool.call({ tool: ROUTE_TOOL, phase: 'orchestrate' }) + await $.tool.call({ tool: 'Skill', skill: 'other' }) + expect(await stepEffort($, seen)).toBe('max') + expect(mainLine(await route($, 'show'))).not.toContain('orchestrate') +}) + +test('floor: lifts a lower sticky, then ends with the turn', async ( + $, on) => { + const seen = await bootFloor($, on) + await route($, 'effort=low') + await ultrathink($) + expect(await stepEffort($, seen)).toBe('max') + await endTurn($) + expect(await stepEffort($, seen)).toBe('low') +}) + +test('floor: main only, an agent step is unaffected', async ($, on) => { + const seen = await bootFloor($, on) + await $.agent.spawn(spawnInput()) + await ultrathink($) + await runStep($, highStep('a1')) + expect(seen[0]?.effort).toBe('medium') +}) + +test('floor: /route clear removes it', async ($, on) => { + const seen = await bootFloor($, on) + await ultrathink($) + expect(mainLine(await route($, 'show'))).toContain('floor max') + await route($, 'clear') + expect(await stepEffort($, seen)).toBe('high') + expect(mainLine(await route($, 'show'))).not.toContain('floor') +}) + +test('floor: a mid-turn prompt applies now and is kept for the next', async ( + $, on) => { + const seen = await bootFloor($, on) + await ultrathink($, 'u1') + expect(await stepEffort($, seen)).toBe('max') + await endTurn($) + expect(await stepEffort($, seen)).toBe('max') + await endTurn($) + expect(await stepEffort($, seen)).toBe('high') +}) + +test('floor: the route answer names the floor over a sticky', async ( + $, on) => { + await bootFloor($, on) + await route($, 'effort=low') + await ultrathink($) + const out = await $.tool.call({ tool: ROUTE_TOOL, phase: 'plan' }) + const text = JSON.stringify(out) + expect(text).toContain('user floor max (prompt rule escalate)') + expect(text).toContain('/route clear') +}) + +// `enabled: false` in ~/.claude/model-router.json cannot be reached here (the +// kit has no fs); the shared off path is covered through `/route off`. +test('floor: /route off keeps the floor from routing', async ($, on) => { + const seen = await bootFloor($, on) + await ultrathink($) + await route($, 'off') + expect(await stepEffort($, seen)).toBe('high') + expect(mainLine(await route($, 'show'))).not.toContain('floor') +}) + +test('per axis: a model-only sticky keeps the turn route effort', async ( + $, on) => { + const seen = await bootFloor($, on) + await route($, 'model=sonnet') + await $.tool.call({ tool: ROUTE_TOOL, phase: 'orchestrate' }) + expect(await stepEffort($, seen)).toBe('medium') +}) + +// The config-driven `offConfig` is re-applied in the same rebuild; only the +// session-only off is reachable here (the kit has no fs). +test('session.end rebuild drops a session-only /route off', async ( + $, on) => { + on('session.end', ($, e) => ({ sessionId: e.sessionId })) + await boot($, on) + await route($, 'off') + await $.session.end({ + reason: 'clear', + sessionId: 's1', + resume: { id: 's1' }, + }) + expect(await route($, 'show')).toContain('router: on') +}) + +// The kit has no fs: the reload read fails, so the previous cfg must stay. +test('reload with an unreadable override keeps the previous config', async ( + $, on) => { + await boot($, on) + await route($, 'switch on') + const out = await route($, 'reload') + expect(out).toContain('switch: on') + expect(await route($, 'show')).toContain('switch: on') +}) + +test('typed marker: a preload in a live agent is ignored', async ($, on) => { + const seen = await bootFloor($, on) + await $.agent.spawn(spawnInput()) + const out = await $.skill.prompt({ skill: 'effort-max', text: 'x' }) + expect(out.text?.startsWith('model-router: effort-max preload')).toBe(true) + expect(mainLine(await route($, 'show'))).not.toContain('floor') + expect(await stepEffort($, seen)).not.toBe('max') +}) + +test('typed marker: /effort-max seen at submit writes the floor', async ( + $, on) => { + await bootFloor($, on) + await $.agent.spawn(spawnInput()) + await $.prompt.submit({ + text: '/effort-max go', + wait: false, + origin: { kind: 'composer' }, + }) + await $.skill.prompt({ skill: 'effort-max', text: 'x' }) + expect(mainLine(await route($, 'show'))).toContain('floor max') +}) + +test('typed slash with no agent and no marker writes the floor', async ( + $, on) => { + await bootFloor($, on) + await $.skill.prompt({ skill: 'effort-max', text: 'x' }) + expect(mainLine(await route($, 'show'))).toContain('floor max') +}) diff --git a/mods/model-router/hooks/register.ts b/mods/model-router/hooks/register.ts index fa00a06..4d757b8 100644 --- a/mods/model-router/hooks/register.ts +++ b/mods/model-router/hooks/register.ts @@ -19,6 +19,7 @@ type Config = { mainModelSwitch: boolean verbose: boolean spinner: boolean + enabled: boolean // false: every hook passes through (per machine) } type Rule = { re: RegExp; phase: string } type Source = 'user' | 'model' | 'skill' | 'prompt' | 'slash' @@ -32,12 +33,17 @@ type State = { rules: Rule[] source: string // 'defaults' or the override path userMain: Routed | null // /route by the user, sticky until /route clear - turnMain: Routed | null // tool, skill, slash, prompt; dropped at turn end - pendingPrompt: Routed | null // typed mid-turn, promoted next turn + turnMain: Routed | null // model route tool, skill table row, Skill(effort-*) + // bridge; dropped at turn end + turnFloor: Routed | null // user-explicit level for this turn (prompt rule, + // typed /effort-): a floor, main loop only + pendingPrompt: Routed | null // typed mid-turn: the next turn's floor + typedSlash: boolean // one-shot: prompt.submit saw a typed /effort- loops: Map // agentId -> that loop's routing explicitEffort: Map // Agent tool_use_id -> effort param skillCalls: number // Skill tool calls in flight - off: boolean // /route off: every hook passes through + off: boolean // /route off or config: every hook passes through + offConfig: boolean // `off` comes from the config key `enabled` lastMain: string // "model/effort" of the last main step (spinner) windowWarned: boolean // context-window warning already logged this turn warned: Set // hooks whose fail-open was already logged @@ -52,6 +58,9 @@ type RouteInput = { agentId?: string } type Picked = { phase: string; route: Route } +type Effort = StepIn['effort'] +type EffortBy = 'floor' | 'sticky' | 'turn' | 'engine' +type Decision = { effort: Effort; by: EffortBy } const LEVELS: readonly Level[] = ['low', 'medium', 'high', 'xhigh', 'max'] const MODEL_ID = /^claude-[a-z0-9.-]+$/ @@ -91,6 +100,7 @@ const DEFAULT_CONFIG: Config = { mainModelSwitch: false, verbose: false, spinner: true, + enabled: true, } // ---- config ---------------------------------------------------------- @@ -213,6 +223,14 @@ function mergePrompt( const pickBool = (v: unknown, fallback: boolean): boolean => typeof v === 'boolean' ? v : fallback +/** The kill switch: a non-boolean value is dropped, and said so. */ +function pickEnabled(v: unknown, fallback: boolean, log: Log): boolean { + if (v !== undefined && typeof v !== 'boolean') { + log('model-router: config "enabled" is not a boolean; ignored') + } + return pickBool(v, fallback) +} + /** Defaults overlaid with the user's entries, each validated first. */ function mergeConfig(user: unknown, log: Log): Config { const base = structuredClone(DEFAULT_CONFIG) @@ -236,6 +254,7 @@ function mergeConfig(user: unknown, log: Log): Config { mainModelSwitch: pickBool(user.mainModelSwitch, base.mainModelSwitch), verbose: pickBool(user.verbose, base.verbose), spinner: pickBool(user.spinner, base.spinner), + enabled: pickEnabled(user.enabled, base.enabled, log), } } @@ -246,7 +265,7 @@ async function readCapped( log: Log, ): Promise { const tooBig = - `model-router: ${OVERRIDE} over ${MAX_CONFIG_BYTES} bytes; defaults` + `model-router: ${OVERRIDE} over ${MAX_CONFIG_BYTES} bytes` if ((await $.fs.stat(path)).size > MAX_CONFIG_BYTES) { log(tooBig) return undefined @@ -257,30 +276,37 @@ async function readCapped( return undefined } -async function readOverride( - $: Api, - log: Log, -): Promise<{ path: string; data: unknown } | undefined> { +type Override = { path: string; data: unknown } | 'absent' | 'failed' + +/** 'absent': no file (defaults apply). 'failed': present but unusable. */ +async function readOverride($: Api, log: Log): Promise { try { const home = await $.env.get('HOME') - if (!home) return undefined + if (!home) return 'absent' const path = `${home}/${OVERRIDE}` - if (!(await $.fs.exists(path))) return undefined + if (!(await $.fs.exists(path))) return 'absent' const text = await readCapped($, path, log) - return text === undefined ? undefined : { path, data: JSON.parse(text) } + if (text === undefined) return 'failed' + const data: unknown = JSON.parse(text) + if (isRecord(data)) return { path, data } + log(`model-router: ${OVERRIDE} is not an object`) + return 'failed' } catch (err) { - log(`model-router: ${OVERRIDE} unreadable (${String(err)}); defaults`) - return undefined + log(`model-router: ${OVERRIDE} unreadable (${String(err)})`) + return 'failed' } } -/** Never throws; a failed read or parse leaves the defaults. */ +/** Never throws; undefined when the override exists but cannot be used. */ async function loadConfig( $: Api, log: Log, -): Promise<{ cfg: Config; source: string }> { +): Promise<{ cfg: Config; source: string } | undefined> { const found = await readOverride($, log) - if (!found) return { cfg: mergeConfig(undefined, log), source: 'defaults' } + if (found === 'failed') return undefined + if (found === 'absent') { + return { cfg: mergeConfig(undefined, log), source: 'defaults' } + } return { cfg: mergeConfig(found.data, log), source: found.path } } @@ -300,11 +326,14 @@ function newState(cfg: Config, source: string): State { source, userMain: null, turnMain: null, + turnFloor: null, pendingPrompt: null, + typedSlash: false, loops: new Map(), explicitEffort: new Map(), skillCalls: 0, off: false, + offConfig: false, lastMain: '', windowWarned: false, warned: new Set(), @@ -335,6 +364,67 @@ function safely(st: State, $: Api, hook: string, work: () => void): void { /** The main loop's effective route: user /route > latest turn route. */ const mainRoute = (st: State): Routed | null => st.userMain ?? st.turnMain +const rank = (l: Level | undefined): number => + l === undefined ? -1 : LEVELS.indexOf(l) + +/** Lifts `effort` to `floor`; a lower level, a number or none is replaced. */ +function floored(effort: Effort, floor: Level | undefined): Effort { + if (floor === undefined) return effort + return isLevel(effort) && rank(effort) >= rank(floor) ? effort : floor +} + +/** Two floors in one turn: the higher level stays (a tie takes the new). */ +function higherFloor(cur: Routed | null, next: Routed): Routed { + return cur && rank(cur.route.effort) > rank(next.route.effort) ? cur : next +} + +/** + * The one decision of the main loop's effort. The user's floor is the turn's + * default (no sticky or turn route names an effort) and its minimum. + * `by` names who set the value: the floor when it raised or supplied it. + */ +function mainEffort(st: State, engine: Effort): Decision { + const sticky = st.userMain?.route.effort + const named = sticky ?? st.turnMain?.route.effort + const floor = st.turnFloor?.route.effort + const base = named ?? floor ?? engine + const effort = floored(base, floor) + if (floor !== undefined && (named === undefined || effort !== base)) { + return { effort, by: 'floor' } + } + if (named === undefined) return { effort, by: 'engine' } + return { effort, by: sticky === undefined ? 'turn' : 'sticky' } +} + +/** Model axis: sticky, then turn route, then the floor's own model. */ +const mainModel = (st: State): string | undefined => + st.userMain?.route.model ?? + st.turnMain?.route.model ?? + st.turnFloor?.route.model + +const floorSource = (f: Routed): string => + f.source === 'prompt' ? `prompt rule ${f.phase}` : `typed /${f.phase}` + +const floorWord = (f: Routed): string => + `user floor ${f.route.effort ?? '-'} (${floorSource(f)})` + +/** + * Why main will not run at `asked`, or '' when it will. Truthful tail of + * every answer that records an effort for the main loop. + */ +function mainNote(st: State, asked: Level | undefined): string { + const d = mainEffort(st, undefined) + if (d.by === 'floor') { + if (!st.turnFloor || asked === undefined || d.effort === asked) return '' + return `${floorWord(st.turnFloor)} keeps main at ${String(d.effort)}; ` + + '/route clear to drop it' + } + return d.by === 'sticky' && st.userMain + ? `a sticky /route ${st.userMain.phase} is in force and wins until ` + + '/route clear' + : '' +} + function loopOf(st: State, agentId: string): Loop { const known = st.loops.get(agentId) if (known) return known @@ -353,20 +443,56 @@ function writeLoop(loop: Loop, route: Route): void { function clearRoutes(st: State): void { st.userMain = null st.turnMain = null + st.turnFloor = null st.pendingPrompt = null } +/** Config `enabled: false` switches the router off; true lifts only that. */ +function applyEnabled(st: State, enabled: boolean): void { + if (!enabled) { + st.off = true + st.offConfig = true + } else if (st.offConfig) { + st.off = false + st.offConfig = false + } +} + +const routerWord = (st: State): string => + st.off ? (st.offConfig ? 'off (config)' : 'off') : 'on' + // ---- text ------------------------------------------------------------ const modelText = (cfg: Config, model: string | undefined): string => model === undefined ? '-' : resolveModel(cfg, model) +/** The floor's level when it carries one and the router is on. */ +function liveFloor(st: State): { f: Routed; level: Level } | undefined { + const f = st.turnFloor + const level = f?.route.effort + return st.off || !f || level === undefined ? undefined : { f, level } +} + +/** True when the switch would put the main loop on a haiku model. */ +function mainOnHaiku(st: State): boolean { + const model = mainModel(st) + return st.cfg.mainModelSwitch && model !== undefined && + resolveModel(st.cfg, model).startsWith(HAIKU) +} + +function effortWord(st: State): string { + if (mainOnHaiku(st)) return '- (haiku takes none)' + return String(mainEffort(st, undefined).effort ?? '-') +} + function mainText(st: State): string { const r = mainRoute(st) - if (!r) return 'main: session defaults' - const model = modelText(st.cfg, r.route.model) + const live = liveFloor(st) + const floor = live ? ` · floor ${live.level} (${live.f.phase})` : '' + if (!r) return 'main: session defaults' + floor + const model = modelText(st.cfg, mainModel(st)) return `main: ${r.source} ${r.phase} · model ${model} · effort ${ - r.route.effort ?? '-'}` + effortWord(st)}${floor}` } function phasesText(cfg: Config): string { @@ -381,7 +507,7 @@ function show(st: State): string { const flag = (b: boolean) => (b ? 'on' : 'off') return [ mainText(st), - `router: ${st.off ? 'off' : 'on'} · switch: ${flag(c.mainModelSwitch)} · ` + + `router: ${routerWord(st)} · switch: ${flag(c.mainModelSwitch)} · ` + `verbose: ${flag(c.verbose)} · spinner: ${flag(c.spinner)}`, `live loops: ${st.loops.size}`, phasesText(c), @@ -391,8 +517,12 @@ function show(st: State): string { function statusLine(st: State): string { const r = mainRoute(st) - const now = st.off ? 'off' : r ? `${r.source} ${r.phase}` : 'session defaults' - return `route: ${now}${st.cfg.mainModelSwitch ? ' · switch on' : ''}` + const now = st.off + ? routerWord(st) + : r ? `${r.source} ${r.phase}` : 'session defaults' + const floor = liveFloor(st) + return `route: ${now}${floor ? ` · floor ${floor.level}` : ''}${ + st.cfg.mainModelSwitch ? ' · switch on' : ''}` } const refresh = ($: Api, st: State): void => $.ui.status(statusLine(st)) @@ -450,12 +580,21 @@ function setUserRoute($: Api, st: State, args: string): string { return show(st) } -/** (Re)loads the config into the state and re-registers the route tool. */ +/** + * (Re)loads the config into the state and re-registers the route tool. + * An unusable override keeps the previous config: the kill switch fails + * closed, never back to the defaults. + */ async function reloadConfig($: Api, st: State): Promise { const loaded = await loadConfig($, text => $.ui.log(text)) - st.cfg = loaded.cfg - st.rules = compileRules(loaded.cfg) - st.source = loaded.source + if (loaded) { + st.cfg = loaded.cfg + st.rules = compileRules(loaded.cfg) + st.source = loaded.source + applyEnabled(st, loaded.cfg.enabled) + } else { + $.ui.log('model-router: override unreadable; keeping the previous config') + } await registerTool($, st) } @@ -472,6 +611,7 @@ async function handleCommand($: Api, st: State, args: string): Promise { case 'on': case 'off': st.off = head === 'off' + st.offConfig = false refresh($, st) return show(st) case 'reload': @@ -557,7 +697,11 @@ function applyRoute(st: State, agentId: string | undefined, p: Picked): void { function clearLoop(st: State, agentId: string | undefined): string { if (agentId === undefined) { st.turnMain = null - return 'route cleared for main' + const f = st.turnFloor + const held = f && f.route.effort !== undefined + ? `; ${floorWord(f)} still holds, /route clear drops it` + : '' + return 'route cleared for main' + held } const loop = st.loops.get(agentId) if (loop) writeLoop(loop, {}) @@ -566,10 +710,8 @@ function clearLoop(st: State, agentId: string | undefined): string { /** Truthful answer: states what the calling loop will actually do. */ function routedText(st: State, agentId: string | undefined, p: Picked): string { - if (agentId === undefined && st.userMain) { - return `recorded ${p.phase} for this turn, but a sticky /route ` + - `${st.userMain.phase} is in force; it wins until /route clear` - } + const note = agentId === undefined ? mainNote(st, p.route.effort) : '' + if (note) return `recorded ${p.phase} for this turn, but ${note}` const loop = agentId === undefined ? undefined : st.loops.get(agentId) const effort = loop?.explicitEffort ? undefined : p.route.effort const model = agentId === undefined && st.cfg.mainModelSwitch @@ -610,9 +752,10 @@ function effortBridge(st: State, agentId: string | undefined, skill: string, if (agentId === undefined) { const route = { ...st.turnMain?.route, effort: level } st.turnMain = { phase: skill, route, source: 'skill' } - return skillResult(skill, st.userMain - ? `model-router: ${skill} recorded, but a sticky /route ` + - `${st.userMain.phase} is in force and wins until /route clear.` + const note = mainNote(st, level) + return skillResult(skill, note + ? `model-router: ${skill} recorded, but ${note}; the ${skill} skill ` + + 'text was not loaded.' : `model-router: effort → ${level} for this loop from the next ` + `request on; the ${skill} skill text was not loaded.`) } @@ -631,7 +774,7 @@ function onSkillLoad(st: State, skill: string, agentId: string | undefined) { const table = hasKey(st.cfg.skills, skill) ? st.cfg.skills[skill] : undefined const route = table === undefined ? undefined : phaseRoute(st.cfg, table) if (agentId === undefined) { - if (st.turnMain && st.turnMain.source !== 'prompt') st.turnMain = null + st.turnMain = null if (table !== undefined && route) { st.turnMain = { phase: table, route, source: 'skill' } } @@ -641,18 +784,38 @@ function onSkillLoad(st: State, skill: string, agentId: string | undefined) { if (loop) writeLoop(loop, { effort: route?.effort }) } -/** A user-typed /effort-: prepends one line, args ride in the text. */ +/** A user-typed /effort-: a floor for the turn, prepends one line. */ function slashEffort(st: State, skill: string, text: string) { const level = EFFORT_SKILL.exec(skill)?.[1] if (!isLevel(level)) return undefined - st.turnMain = { phase: skill, route: { effort: level }, source: 'slash' } - const line = st.userMain - ? `Effort ${level} recorded; the sticky /route ${st.userMain.phase} ` + - 'wins until /route clear.' - : `Effort shifted to ${level} by model-router for this turn.` + const route: Route = { effort: level } + const slash: Routed = { phase: skill, route, source: 'slash' } + st.turnFloor = higherFloor(st.turnFloor, slash) + const note = mainNote(st, level) + const line = `Effort ${level} set by model-router for the main loop this ` + + 'turn (minimum; a higher route still applies).' + + (note ? ` But ${note}.` : '') return { text: line + '\n' + text } } +/** + * Floor write for a skill.prompt. Only a typed slash may write it: the + * marker from prompt.submit attests the typing. Without it, a live + * sub-agent means the prompt is a preload inside that agent: ignored. + */ +function guardedSlash(st: State, skill: string, text: string) { + if (!EFFORT_SKILL.test(skill)) return undefined + if (st.typedSlash) { + st.typedSlash = false + } else if (st.loops.size > 0) { + return { + text: `model-router: ${skill} preload inside a live sub-agent is ` + + 'ignored on the main loop.\n' + text, + } + } + return slashEffort(st, skill, text) +} + // ---- agents ---------------------------------------------------------- type SpawnIn = { @@ -714,9 +877,8 @@ async function windowOk($: Api, st: State, id: string): Promise { } async function mainPlan($: Api, st: State, e: StepIn): Promise { - const set = mainRoute(st) - const effort = set?.route.effort ?? e.effort - const wanted = set?.route.model + const { effort } = mainEffort(st, e.effort) + const wanted = mainModel(st) if (wanted === undefined || !st.cfg.mainModelSwitch) { return { model: e.model, effort } } @@ -751,8 +913,10 @@ function noteMain($: Api, st: State, plan: Plan): void { } function endMainTurn($: Api, st: State): void { - st.turnMain = st.pendingPrompt + st.turnFloor = st.pendingPrompt st.pendingPrompt = null + st.turnMain = null + st.typedSlash = false st.explicitEffort.clear() st.lastMain = '' st.windowWarned = false @@ -773,6 +937,9 @@ function registerSession(on: On, st: State): void { }) on('session.end', async ($, e, next) => { Object.assign(st, newState(st.cfg, st.source)) + // session.start never fires after /clear: re-apply the config's + // `enabled` so a config-disabled router (offConfig) stays off. + applyEnabled(st, st.cfg.enabled) return next(e) }).catch(($, e, next) => { warnOnce(st, $, 'session.end', next.error.kind) @@ -823,7 +990,10 @@ function registerSkills(on: On, st: State): void { }) on('skill.prompt', async ($, e, next) => { if (st.off || st.skillCalls > 0) return next(e) - return slashEffort(st, e.skill, e.text) ?? next(e) + const out = guardedSlash(st, e.skill, e.text) + if (!out) return next(e) + refresh($, st) + return out }).catch(($, e, next) => { warnOnce(st, $, 'skill.prompt', next.error.kind) return next(e) @@ -891,6 +1061,15 @@ function registerTurns(on: On, st: State): void { }) } +/** + * A prompt's level is a floor now; typed mid-turn (`wait` is ignored, the + * engine queues either way) it is also kept for the next turn. + */ +function floorFromPrompt(st: State, midTurn: boolean, routed: Routed): void { + st.turnFloor = higherFloor(st.turnFloor, routed) + if (midTurn) st.pendingPrompt = higherFloor(st.pendingPrompt, routed) +} + function registerPrompt(on: On, st: State): void { on('prompt.submit', async ($, e, next) => { if (st.off || e.origin.kind !== 'composer') return next(e) @@ -899,11 +1078,10 @@ function registerPrompt(on: On, st: State): void { const route = rule ? phaseRoute(st.cfg, rule.phase) : undefined if (rule && route) { const routed: Routed = { phase: rule.phase, route, source: 'prompt' } - // Typed mid-turn and asked to wait: it belongs to the NEXT turn. - if (e.turnId !== undefined && e.wait) st.pendingPrompt = routed - else st.turnMain = routed + floorFromPrompt(st, e.turnId !== undefined, routed) refresh($, st) } + if (e.text.trimStart().startsWith('/effort-')) st.typedSlash = true return next(e) }).catch(($, e, next) => { warnOnce(st, $, 'prompt.submit', next.error.kind) From 24e180ade0189c2e0184166a66d404e38caf6504 Mon Sep 17 00:00:00 2001 From: bchanot Date: Fri, 9 Oct 2026 10:13:03 +0200 Subject: [PATCH 11/22] =?UTF-8?q?chore(tasks):=20model-router=20B1=20done?= =?UTF-8?q?=20=E2=80=94=20contract=20evidence,=20TODO,=20journal?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- .claude/memory/journal.md | 1 + .claude/tasks/TODO.md | 3 ++- .../2026-10-08-model-router-floor-1835.md | 18 ++++++++++++++---- 3 files changed, 17 insertions(+), 5 deletions(-) diff --git a/.claude/memory/journal.md b/.claude/memory/journal.md index 0af5f38..c8e2ae8 100644 --- a/.claude/memory/journal.md +++ b/.claude/memory/journal.md @@ -582,3 +582,4 @@ rules: - model-router wave 1-A (/feat, user go): mod built in `mods/model-router/` (4 files, 834 + 202 lines, 11 plugin tests). Plan r1 → r3: 3 challengers (simplicity CONCERNS, robustness CONCERNS(6), correctness FATAL(8)) + 1 confirmation CONCERNS(4); converged on: agents table = built-ins only in wave 1 (frontmatter stays single writer), agent model written once at spawn, explicit Agent params frozen per loop, every sub-agent write on its own loop, Skill bridge answers in the Skill tool's OUTPUT schema (string result refused → skill would load), config validated before merge, state in closure, `/route off`. feater DONE first pass; GATE 0 MET; verifier CONFORME 6/6; security PASS (4 MEDIUM + 5 LOW parked in TODO for user go). Commits b721c94 (mod) + e8ca713 (contract/plan). Override `~/.claude/model-router.json` {verbose:true} written (pass B). Doc-sync deferred to 1-B. - model-router W1-A hardening (user go): fresh feater on contract criteria 7-11 → gap round (dead `Loop.frozen`, spawn returns `started` verbatim, tool description) → verifier CONFORME 11/11 → security PASS (1 MEDIUM residual: ReDoS size-bounded only, self-inflicted config; 5 LOW parked). Registries BDR-115, LRN-205, LRN-206, EVAL-040 written on user go (64702d5). Next: live swap of the real mod into the hot-reload folder, then W1-B install + docs. - model-router live checks on Opus 5.5 (user /model, uncommitted settings.json change left to the user): Skill(effort-low) bridge answered in place → next request low; ultrathink turn ran max (engine base medium); Explore without params → claude-sonnet-5-5, 3 steps medium. Loading switched to tracked symlink skills/model-router → @skills-dir (PLUGIN_DIRS non-portable: absolute path, no $HOME expansion, tracked settings); isolated-HOME probe listed/enabled/loaded. User chose floor semantics for ultrathink + typed /effort-. W1-B split: B1 floor (register.ts) + B2 wiring (symlink, gitignore, mods suite, doctor, CLAUDE.md); 6 challengers in flight. Effort shifters skipped this turn: the bridge would overwrite the user's ultrathink until B1 lands. +- model-router W1-B1 floor landed (1ff608a): plan r2 from 3 lenses (0 BLOCKER, 4 MAJOR: one decision helper, typed level = default + minimum, mid-turn prompt now + next, per-machine enabled:false), feater DONE, gap round (/clear lost enabled:false, 'ultrathink rule' label, per-axis effort base), hardening round (kill switch keeps previous cfg on failed reload, typed slash attested at prompt.submit vs sub-agent preload). 30 tests, verifier CONFORME 6/6, security PASS ×2 with parked residuals. B2 wiring dispatched next. diff --git a/.claude/tasks/TODO.md b/.claude/tasks/TODO.md index d1b2620..4e8ac1e 100644 --- a/.claude/tasks/TODO.md +++ b/.claude/tasks/TODO.md @@ -8,7 +8,8 @@ migration of shifters/pins/model-gate in wave 2 after proof; names model-router - [x] W1-A the mod in `mods/model-router/` (b721c94, contract `2026-10-08-model-router-w1a-1533`, plan r3): challenge 3 lenses + 1 confirmation (2 BLOCKER + 10 MAJOR closed by named changes), feater DONE first pass, GATE 0 MET 5/5, verifier CONFORME 6/6, security PASS (4 MEDIUM + 5 LOW reported, below) - [x] W1-A hardening (user go "oui durcis", contract criteria 7-11): `/route` composer-only; no agent model axis (effort only after spawn); pattern ≤ 200 / scan ≤ 4096 / phase keys `^[a-z][a-z0-9_-]{0,31}$` / config ≤ 64 KB / `additionalProperties: false` / `typeof e.skill`; `warnOnce` in all 14 catches + config-drop logs; `safely` around post-`next` bookkeeping. feater DONE, gap round (dead `Loop.frozen`, spawn returns `started` verbatim, tool description), GATE 0 MET 9/9, verifier CONFORME 11/11, security PASS. - [ ] W1-A residuals (security, accepted, none exploitable from outside the user's own files): ReDoS is size-bounded only (`(a+)+$` in `~/.claude/model-router.json` + a 4 KB paste hangs the hook; fix = nested-quantifier rejection or a far lower scan cap); `stat().size` trusted (FIFO/device path in ~/.claude); unrestricted `models`/`agents`/`skills` KEYS echoed raw in logs (log flood); `agentId` read from the flat tool event (engine strip unverified); 5 closures without a kit test (no fs in the kit, LRN-206); "nothing routed" catch text after a state write. -- [ ] W1-B1 floor precedence (contract `2026-10-08-model-router-floor-1835`): `ultrathink` + typed `/effort-` = floor for the main turn (user choice 2026-10-08) +- [x] W1-B1 floor precedence (contract `2026-10-08-model-router-floor-1835`, 2026-10-09): `ultrathink` + typed `/effort-` = the main turn's default AND minimum (r2 after 3 lenses: a pure floor made /effort-low a no-op); one helper `mainEffort`; per-axis precedence; mid-turn prompt floors now + next turn; `"enabled": false` per machine, kept across /clear and failed reload; typed slash attested. 30 tests; verifier CONFORME 6/6 (after 1 gap round); security PASS ×2 +- [ ] W1-B1 residuals (security 2026-10-09, accepted): first-load failure of the override falls to defaults (`enabled: true`); non-boolean `enabled` drops to the default on reload; `skill.prompt` preload guard is a heuristic (`loops.size > 0`; a preload during spawn or a non-composer `/effort-*` while idle still writes the floor); marker not bound to a valid level; `/route reload` answers "config reloaded" even when the previous config was kept; transient missing override lifts a config-set off; `String(err)` of a JSON parse in the local log. Display: `/route show` folds the floor into the effort while off; model-axis text with a model-only sticky and the switch on. - [ ] W1-B2 wiring (contract `2026-10-08-model-router-wiring-1835`): tracked symlink `skills/model-router` → `../mods/model-router` (loads as `@skills-dir`; PLUGIN_DIRS dropped: absolute paths in tracked settings), `.gitignore` engine-laid files, `lib/tests/mods.test.sh`, doctor `── Mods ──`, CLAUDE.md `## mods/`; then doc-sync (README/USAGE/CHANGELOG), live swap (remove the hot-reload link, `/reload-plugins`), BDR-115 amendment - [ ] W2 migration: 15 skills off `Skill(effort-*)`, remove shifters + effort-pins + model-gate, census repointed, docs - [ ] W3 optional: step heuristics, haiku classifier, quota-aware downgrade, A/B diff --git a/.claude/tasks/contracts/2026-10-08-model-router-floor-1835.md b/.claude/tasks/contracts/2026-10-08-model-router-floor-1835.md index 30943eb..450452d 100644 --- a/.claude/tasks/contracts/2026-10-08-model-router-floor-1835.md +++ b/.claude/tasks/contracts/2026-10-08-model-router-floor-1835.md @@ -21,16 +21,26 @@ Q: who clears the floor / A: main turn end (a queued prompt's floor is then prom 1. Suite green with the new tests: `claude plugin test` passes with at least 22 `test(` calls; `claude plugin validate` passes with no warning; no line over 80 chars; no `any` type. CHECK: cd mods/model-router && out=$(claude plugin test . 2>&1); rc=$?; echo "$out" | tail -n 3; [ $rc -eq 0 ] && [ "$(grep -cE '^\s*test\(' hooks/register.test.ts)" -ge 22 ] && v=$(claude plugin validate . 2>&1) && echo "$v" | grep -q 'Validation passed' && ! echo "$v" | grep -qi 'warning' && ! grep -nE '.{81,}' hooks/register.ts hooks/register.test.ts && ! grep -nE ':\s*any\b||as any\b' hooks/register.ts && echo FLOOR-SUITE-OK EXPECT: FLOOR-SUITE-OK - EVIDENCE: pending + EVIDENCE: MET exit=0 marker-found :: 30 pass 0 fail Ran 30 tests across 1 file. [1.05s] FLOOR-SUITE-OK 2. Type-check clean against this build's declarations. CHECK: T=/Users/b.chanot/.claude/dev-mods/385f7190-70f5-4bdd-b0d8-e4566cd412fd/model-router/.claude-plugin/types; [ -d "$T" ] || T=/Users/b.chanot/Documents/claude/mods/model-router/.claude-plugin/types; W=$(mktemp -d) && printf '{"compilerOptions":{"target":"es2023","lib":["es2023"],"types":[],"module":"esnext","moduleResolution":"bundler","strict":true,"noUncheckedIndexedAccess":true,"noEmit":true,"skipLibCheck":true,"jsx":"react","jsxFactory":"h","jsxFragmentFactory":"Fragment"},"include":["%s/claude-code/index.d.ts","%s/claude-code-tools/index.d.ts","%s/hooks"]}' "$T" "$T" "$PWD/mods/model-router" > "$W/tsconfig.json" && (cd "$W" && npx --yes -p typescript@5 tsc -p tsconfig.json) && echo TSC-OK EXPECT: TSC-OK - EVIDENCE: pending + EVIDENCE: MET exit=0 marker-found :: TSC-OK 3. The floor is its own slot: `turnFloor` is declared in `State`, initialised in `newState`, written by the prompt rule and by a typed `/effort-`, cleared by `/route clear` and at main turn end; the suite carries at least 8 tests whose name contains `floor`; one helper `mainEffort` decides the main effort; the config key `enabled` exists. CHECK: cd mods/model-router/hooks && [ "$(grep -c 'turnFloor' register.ts)" -ge 6 ] && [ "$(grep -cE "^\s*test\('[^']*floor" register.test.ts)" -ge 8 ] && grep -q 'mainEffort' register.ts && grep -q 'enabled' register.ts && echo FLOOR-SLOT-OK EXPECT: FLOOR-SLOT-OK - EVIDENCE: pending -4. Judged by reading: effective main effort comes from ONE helper `mainEffort` used by `mainPlan` and by every answer text: base `(userMain ?? turnMain)?.route.effort ?? turnFloor?.route.effort ?? e.effort`, then floored by `turnFloor` (LEVELS order; a numeric or absent value is replaced); the floor never applies to a sub-agent step; `turnMain` only ever holds 'model' or 'skill' sources; the prompt rule writes `turnFloor` (keeping the higher of two) and, when typed mid-turn (`turnId` set, `wait` ignored), also `pendingPrompt`, promoted into `turnFloor` at main turn end; `"enabled": false` in the override file makes every hook pass through after each config load (`/route on` re-enables for the session); a non-effort skill load resets `turnMain` only; a model `route({clear})` clears `turnMain` only; every answer that the floor overrides says so truthfully (Skill bridge context, route tool text, `/effort-` text); `/route show` and the status line display the floor; every criterion of `.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md` still holds; no function over 25 logic lines. + EVIDENCE: MET exit=0 marker-found :: FLOOR-SLOT-OK +4. Judged by reading: effective main effort comes from ONE helper `mainEffort` used by `mainPlan` and by every answer text: base `userMain?.route.effort ?? turnMain?.route.effort ?? turnFloor?.route.effort ?? e.effort` (per-axis, like the model rule: a model-only sticky never hides a turn route's effort; gap round 2026-10-09), then floored by `turnFloor` (LEVELS order; a numeric or absent value is replaced); the floor never applies to a sub-agent step; `turnMain` only ever holds 'model' or 'skill' sources; the prompt rule writes `turnFloor` (keeping the higher of two) and, when typed mid-turn (`turnId` set, `wait` ignored), also `pendingPrompt`, promoted into `turnFloor` at main turn end; `"enabled": false` in the override file makes every hook pass through after each config load AND survives `/clear` (`session.end` rebuilds the state but re-applies the config's `enabled`; `/route on` re-enables for the session); a prompt-rule floor is labelled by its matched phase (`prompt rule `), never by a fixed word; a non-effort skill load resets `turnMain` only; a model `route({clear})` clears `turnMain` only; every answer that the floor overrides says so truthfully (Skill bridge context, route tool text, `/effort-` text); `/route show` and the status line display the floor; every criterion of `.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md` still holds; no function over 25 logic lines. + +Hardening round (security gate 2026-10-09, 2 MEDIUM) — criteria 5-6, same ledger: +5. The kill switch fails closed: a failed override read on `/route reload` (unreadable, oversized, invalid JSON) keeps the PREVIOUS config (and therefore the previous `enabled`) instead of falling back to the defaults; a non-boolean `enabled` value is dropped WITH a log line; at session start with no previous config the defaults still apply. + CHECK: cd mods/model-router && grep -q "previous" hooks/register.ts && grep -qE "enabled.*(not a boolean|non-boolean|ignored)" hooks/register.ts && grep -qE "test\('[^']*(reload|previous|kill)" hooks/register.test.ts && echo KILL-CLOSED-OK + EXPECT: KILL-CLOSED-OK + EVIDENCE: MET exit=0 marker-found :: KILL-CLOSED-OK +6. `skill.prompt` writes the floor only for a typed `/effort-`: a one-shot marker set at `prompt.submit` (composer origin, text starting with `/effort-`) attests the typing; without the marker the write is refused while any sub-agent loop is live (a preload fires inside an agent's life), and accepted otherwise (no agent can be preloading); the refused case returns the text unchanged with a one-line note. Tests: preload simulation (spawned agent live, no marker → no floor), typed with marker → floor, typed with no marker and no agent → floor. + CHECK: cd mods/model-router && grep -q "slashMarker\|typedSlash" hooks/register.ts && [ "$(grep -cE "test\('[^']*(preload|marker|typed)" hooks/register.test.ts)" -ge 2 ] && echo SLASH-ATTEST-OK + EXPECT: SLASH-ATTEST-OK + EVIDENCE: MET exit=0 marker-found :: SLASH-ATTEST-OK ## FILE SCOPE mods/model-router/hooks/register.ts · mods/model-router/hooks/register.test.ts From 6430ac65ecaa4154e45f76a163f3f8d7f8a39356 Mon Sep 17 00:00:00 2001 From: bchanot Date: Fri, 9 Oct 2026 10:51:13 +0200 Subject: [PATCH 12/22] =?UTF-8?q?feat(mods):=20model-router=20active=20in?= =?UTF-8?q?=20every=20session=20=E2=80=94=20skills-dir=20link,=20mods=20su?= =?UTF-8?q?ite,=20doctor=20section,=20CLAUDE.md?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Tracked relative symlink skills/model-router -> ../mods/model-router: Claude Code loads the mod in place as model-router@skills-dir wherever link.sh links ~/.claude/skills (no CLAUDE_CODE_PLUGIN_DIRS: absolute paths in the tracked settings.json). Engine-laid mods/*/tsconfig.json gitignored. lib/tests/mods.test.sh: manifest name, link target, claude plugin validate and test per mod, capability-probed, time-bounded, SKIP with reason. doctor.sh: fail-soft Mods section (link by -ef, one guarded plugin list). CLAUDE.md: mods/ section (loading, per-machine enabled:false switch, dev-copy shadowing, tests). --- .gitignore | 4 ++ CLAUDE.md | 22 +++++++++++ doctor.sh | 58 ++++++++++++++++++++++++++++ lib/tests/mods.test.sh | 87 ++++++++++++++++++++++++++++++++++++++++++ skills/model-router | 1 + 5 files changed, 172 insertions(+) create mode 100644 lib/tests/mods.test.sh create mode 120000 skills/model-router diff --git a/.gitignore b/.gitignore index a45eacc..42c29f2 100644 --- a/.gitignore +++ b/.gitignore @@ -262,3 +262,7 @@ skills-external/.higgsfield-stage.*/ # ── gitflow standard socle (added by gitflow_init; additive, safe to edit) ── *.log !.claude/deploy/ + +# mods/: the engine lays tsconfig.json beside a loaded mod; its +# .claude-plugin/types/ ignores itself +mods/*/tsconfig.json diff --git a/CLAUDE.md b/CLAUDE.md index 5e0bfcf..c58878d 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -59,6 +59,28 @@ Gotcha, learned the hard way: `git rm --cached` keeps the working file, but if the branch you merge into still tracks it, the merge deletes it from disk. Untrack and merge, then restore with the command above. +## mods/ — function-hooks plugins (Claude Code mods) + +A mod lives in `mods//` (`.claude-plugin/plugin.json` + hooks). It +loads through the tracked relative symlink `skills/` -> `../mods/` +(`~/.claude/skills` links to `skills/`) as `@skills-dir`, in place, +live at the next session or `/reload-plugins`. New mod: `ln -s ../mods/ +skills/` from the repo root (guard with `[ -L ]`, a re-run nests a link). +Not `CLAUDE_CODE_PLUGIN_DIRS` (absolute path, settings `env` has no `$HOME` +expansion, settings.json is tracked), nor a local marketplace (`add` writes +an absolute path into settings.json). +- The engine lays `mods//tsconfig.json` and `.claude-plugin/types/`; + both are gitignored. +- Optional user config: `~/.claude/.json`. Its `"enabled": false` is the + per-machine off switch (untracked). `"@skills-dir": false` in + `enabledPlugins` also works but lands in the TRACKED settings.json and + dirties every machine's tree. +- A dev copy of the same name (`--plugin-dir`, hot-reload link in + `~/.claude/dev-mods//`) shadows the skills-dir copy for that + session: remove it before reading `/reload-plugins` as a test of the link. +- Tests: `make test suite=lib/tests/mods.test.sh` (manifest, link, + `claude plugin validate`, `claude plugin test`). `doctor.sh` has a Mods section. + ## Transient planning artifacts `docs/superpowers/specs/**` and `docs/superpowers/plans/**` are run-time diff --git a/doctor.sh b/doctor.sh index 63f281b..8bfe308 100644 --- a/doctor.sh +++ b/doctor.sh @@ -155,6 +155,64 @@ unset _dv_active_profile _dv_profile_file echo "" +# ──────────────────────────────────────────────────────────── +# 2c. Mods (mods// plugins, loaded through the tracked +# skills/ symlink as @skills-dir). Fail-soft: a missing link +# is info (the user may have removed it on purpose), never an error. +# ──────────────────────────────────────────────────────────── +echo "── Mods ──" + +# Prints enabled|disabled|absent|unknown for $1 read from the JSON on stdin; +# always exits 0 so a bad payload cannot abort doctor under set -e. +mod_state() { + python3 -c ' +import json, sys +try: + rows = json.load(sys.stdin) + row = [r for r in rows if r.get("id") == sys.argv[1] + "@skills-dir"] + print("absent" if not row else + "enabled" if row[0].get("enabled") is True else "disabled") +except Exception: + print("unknown") +' "$1" 2>/dev/null || true +} + +_mods_list="" +if command -v claude &>/dev/null; then + if ! _mods_list=$(claude plugin list --json 2>/dev/null); then + warn "mods: claude plugin list failed — load state not checked" + _mods_list="" + fi +fi + +_mods_seen=0 +for _mod_manifest in "$REPO"/mods/*/.claude-plugin/plugin.json; do + [ -f "$_mod_manifest" ] || continue + _mods_seen=$((_mods_seen + 1)) + _mod=$(basename "$(dirname "$(dirname "$_mod_manifest")")") + _mod_link="$HOME/.claude/skills/$_mod" + if ! { [ -L "$_mod_link" ] || [ -e "$_mod_link" ]; }; then + info "mod $_mod: not linked (skills/$_mod absent) — git checkout skills/$_mod if wanted" + continue + fi + if [ "$_mod_link" -ef "$REPO/mods/$_mod" ]; then + pass "mod $_mod: loading link ~/.claude/skills/$_mod" + else + warn "mod $_mod: ~/.claude/skills/$_mod does not resolve to $REPO/mods/$_mod" + fi + [ -n "$_mods_list" ] || continue + case "$(printf '%s' "$_mods_list" | mod_state "$_mod")" in + enabled) pass "mod $_mod: enabled as $_mod@skills-dir" ;; + disabled) warn "mod $_mod: disabled (\"$_mod@skills-dir\": false in enabledPlugins)" ;; + absent) warn "mod $_mod: not listed as @skills-dir — run: claude plugin validate mods/$_mod (policy, manifest or name conflict)" ;; + *) warn "mod $_mod: claude plugin list output not understood" ;; + esac +done +[ "$_mods_seen" -gt 0 ] || info "no mods" +unset _mods_list _mods_seen _mod_manifest _mod _mod_link + +echo "" + # ── Playwright browsers (read-only report; NOT nested under gstack — 2 of # the 3 registered installs are gsd-pi, not gstack) ── echo "── Playwright browsers ──" diff --git a/lib/tests/mods.test.sh b/lib/tests/mods.test.sh new file mode 100644 index 0000000..dcc7f9c --- /dev/null +++ b/lib/tests/mods.test.sh @@ -0,0 +1,87 @@ +#!/usr/bin/env bash +# lib/tests/mods.test.sh — every mods// plugin: the manifest name +# equals the folder, skills/ is the relative loading symlink +# ../mods/, and (when the CLI offers `claude plugin test`) the mod +# passes `claude plugin validate` without warning and `claude plugin test`. +# MODS_ROOT overrides the repo root (fixture controls). Fails when no mod +# is found, so it can never pass vacuously. +set -u +ROOT="${MODS_ROOT:-$(cd "$(dirname "$0")/../.." && pwd)}" +CLI_TIMEOUT=120 +pass=0; fail=0 +ok() { pass=$((pass+1)); echo "PASS $1"; } +ko() { fail=$((fail+1)); echo "FAIL $1"; } +check() { if [ "$2" = "$3" ]; then ok "$1"; else ko "$1: got[$2] want[$3]"; fi; } + +# bounded CMD...: stdout+stderr on stdout, rc 124 on timeout. +bounded() { + local t + t=$(command -v timeout || command -v gtimeout || true) + if [ -n "$t" ]; then "$t" "$CLI_TIMEOUT" "$@" 2>&1; return; fi + local out rc=0 pid i=0 + out=$(mktemp) || return 1 + "$@" >"$out" 2>&1 & pid=$! + while kill -0 "$pid" 2>/dev/null && [ "$i" -lt "$CLI_TIMEOUT" ]; do + sleep 1; i=$((i+1)) + done + if kill -0 "$pid" 2>/dev/null; then kill "$pid" 2>/dev/null; rc=124 + else wait "$pid" || rc=$?; fi + cat "$out"; rm -f "$out"; return "$rc" +} + +manifest_name() { + python3 -c 'import json,sys; print(json.load(open(sys.argv[1])).get("name",""))' \ + "$1" 2>/dev/null +} + +# cli_unavailable: prints the reason and returns 0 when the capability is missing. +cli_unavailable() { + command -v claude >/dev/null 2>&1 || { echo "claude not found"; return 0; } + local rc=0 + bounded claude plugin test --help >/dev/null || rc=$? + [ "$rc" -eq 124 ] && { echo "probe timed out after ${CLI_TIMEOUT}s"; return 0; } + [ "$rc" -ne 0 ] && { echo "no 'claude plugin test' command"; return 0; } + return 1 +} + +check_cli() { + local name="$1" dir="$2" out rc=0 + out=$(bounded claude plugin validate "$dir") || rc=$? + if [ "$rc" -eq 124 ]; then ko "$name: validate timed out after ${CLI_TIMEOUT}s" + elif ! printf '%s' "$out" | grep -q 'Validation passed'; then + ko "$name: validate did not pass: $(printf '%s' "$out" | head -3 | tr '\n' ' ')" + elif printf '%s' "$out" | grep -qi 'warning'; then + ko "$name: validate printed a warning" + else ok "$name: validate passed, no warning"; fi + rc=0 + out=$(bounded claude plugin test "$dir") || rc=$? + if [ "$rc" -eq 124 ]; then ko "$name: plugin test timed out after ${CLI_TIMEOUT}s" + elif [ "$rc" -ne 0 ]; then + ko "$name: plugin test rc=$rc: $(printf '%s' "$out" | tail -3 | tr '\n' ' ')" + else ok "$name: plugin test passed"; fi +} + +manifests=() +for m in "$ROOT"/mods/*/.claude-plugin/plugin.json; do + [ -f "$m" ] && manifests+=("$m") +done +if [ "${#manifests[@]}" -eq 0 ]; then ko "no mod found under $ROOT/mods"; fi + +use_cli=1 +if reason=$(cli_unavailable); then + use_cli=0 + echo "SKIP: claude plugin test unavailable ($reason) — validate/test not run" +fi + +for m in ${manifests[@]+"${manifests[@]}"}; do + dir="$(dirname "$(dirname "$m")")"; name="$(basename "$dir")" + check "$name: manifest name matches folder" "$(manifest_name "$m")" "$name" + link="$ROOT/skills/$name" + if [ -L "$link" ]; then + check "$name: loading link target" "$(readlink "$link")" "../mods/$name" + else ko "$name: skills/$name is not a symlink"; fi + [ "$use_cli" -eq 1 ] && check_cli "$name" "$dir" +done + +echo "mods: $pass pass, $fail fail" +[ "$fail" -eq 0 ] diff --git a/skills/model-router b/skills/model-router new file mode 120000 index 0000000..e051553 --- /dev/null +++ b/skills/model-router @@ -0,0 +1 @@ +../mods/model-router \ No newline at end of file From 3c44dd00d3f31b55060c79344c3cf19a2c6377a4 Mon Sep 17 00:00:00 2001 From: bchanot Date: Fri, 9 Oct 2026 10:51:51 +0200 Subject: [PATCH 13/22] =?UTF-8?q?chore(tasks):=20model-router=20B2=20done?= =?UTF-8?q?=20=E2=80=94=20contract=20evidence,=20TODO=20close-out=20queue?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- .claude/tasks/TODO.md | 4 +++- .../2026-10-08-model-router-wiring-1835.md | 16 ++++++++-------- 2 files changed, 11 insertions(+), 9 deletions(-) diff --git a/.claude/tasks/TODO.md b/.claude/tasks/TODO.md index 4e8ac1e..400a906 100644 --- a/.claude/tasks/TODO.md +++ b/.claude/tasks/TODO.md @@ -10,7 +10,9 @@ migration of shifters/pins/model-gate in wave 2 after proof; names model-router - [ ] W1-A residuals (security, accepted, none exploitable from outside the user's own files): ReDoS is size-bounded only (`(a+)+$` in `~/.claude/model-router.json` + a 4 KB paste hangs the hook; fix = nested-quantifier rejection or a far lower scan cap); `stat().size` trusted (FIFO/device path in ~/.claude); unrestricted `models`/`agents`/`skills` KEYS echoed raw in logs (log flood); `agentId` read from the flat tool event (engine strip unverified); 5 closures without a kit test (no fs in the kit, LRN-206); "nothing routed" catch text after a state write. - [x] W1-B1 floor precedence (contract `2026-10-08-model-router-floor-1835`, 2026-10-09): `ultrathink` + typed `/effort-` = the main turn's default AND minimum (r2 after 3 lenses: a pure floor made /effort-low a no-op); one helper `mainEffort`; per-axis precedence; mid-turn prompt floors now + next turn; `"enabled": false` per machine, kept across /clear and failed reload; typed slash attested. 30 tests; verifier CONFORME 6/6 (after 1 gap round); security PASS ×2 - [ ] W1-B1 residuals (security 2026-10-09, accepted): first-load failure of the override falls to defaults (`enabled: true`); non-boolean `enabled` drops to the default on reload; `skill.prompt` preload guard is a heuristic (`loops.size > 0`; a preload during spawn or a non-composer `/effort-*` while idle still writes the floor); marker not bound to a valid level; `/route reload` answers "config reloaded" even when the previous config was kept; transient missing override lifts a config-set off; `String(err)` of a JSON parse in the local log. Display: `/route show` folds the floor into the effort while off; model-axis text with a model-only sticky and the switch on. -- [ ] W1-B2 wiring (contract `2026-10-08-model-router-wiring-1835`): tracked symlink `skills/model-router` → `../mods/model-router` (loads as `@skills-dir`; PLUGIN_DIRS dropped: absolute paths in tracked settings), `.gitignore` engine-laid files, `lib/tests/mods.test.sh`, doctor `── Mods ──`, CLAUDE.md `## mods/`; then doc-sync (README/USAGE/CHANGELOG), live swap (remove the hot-reload link, `/reload-plugins`), BDR-115 amendment +- [x] W1-B2 wiring (contract `2026-10-08-model-router-wiring-1835`, 2026-10-09): tracked symlink `skills/model-router` → `../mods/model-router` loads as `model-router@skills-dir` (fresh-process `claude plugin list --json` proves it), `.gitignore` `mods/*/tsconfig.json`, `lib/tests/mods.test.sh` (4 checks, capability probe, bounded, SKIP), doctor `── Mods ──` fail-soft, CLAUDE.md `## mods/`. Plan r2 from 3 lenses (5 MAJOR), feater DONE, gap round (4b label + unbounded probe), verifier CONFORME 8/8, security PASS +- [ ] W1-B2 residuals (accepted): `claude plugin --help` returns 0 → the suite's capability probe can pass on a CLI without `plugin test` and then FAIL instead of SKIP (fail-closed; fix = grep the probe output for the test usage line); doctor `claude plugin list --json` unbounded; doctor `echo -e` helpers interpolate `$_mod` (tracked folder names only); fallback timeout guard orphans grandchildren (only without coreutils timeout); `update-all.sh` runs `claude plugin update` over `@skills-dir` → one recurring warn (needs an update-all edit) +- [ ] W1 close-out: doc-sync (README / USAGE / CHANGELOG), live swap (remove the hot-reload link in `~/.claude/dev-mods//`, user runs `/reload-plugins`), BDR-115 amendment (loading = skills-dir link; floor semantics), journal - [ ] W2 migration: 15 skills off `Skill(effort-*)`, remove shifters + effort-pins + model-gate, census repointed, docs - [ ] W3 optional: step heuristics, haiku classifier, quota-aware downgrade, A/B diff --git a/.claude/tasks/contracts/2026-10-08-model-router-wiring-1835.md b/.claude/tasks/contracts/2026-10-08-model-router-wiring-1835.md index 6f887de..8c32fac 100644 --- a/.claude/tasks/contracts/2026-10-08-model-router-wiring-1835.md +++ b/.claude/tasks/contracts/2026-10-08-model-router-wiring-1835.md @@ -18,31 +18,31 @@ Q: doctor scope / A: per mod: the loading link resolves into the repo mod dir; ` 1. `skills/model-router` is a symlink whose target is exactly `../mods/model-router`, and git does not ignore it. CHECK: [ -L skills/model-router ] && [ "$(readlink skills/model-router)" = "../mods/model-router" ] && ! git check-ignore -q skills/model-router && echo LINK-OK EXPECT: LINK-OK - EVIDENCE: pending + EVIDENCE: MET exit=0 marker-found :: LINK-OK 2. Engine-laid files are ignored for ANY mod: a mod's root `tsconfig.json` (root `.gitignore`; the `.claude-plugin/types/` folder ignores itself); the tracked mod files are not ignored. CHECK: git check-ignore -q mods/model-router/tsconfig.json && git check-ignore -q mods/zz-future/tsconfig.json && ! git check-ignore -q mods/model-router/hooks/register.ts && ! git check-ignore -q mods/model-router/.claude-plugin/plugin.json && [ -z "$(git status --short mods/)" ] && echo IGNORE-OK EXPECT: IGNORE-OK - EVIDENCE: pending + EVIDENCE: MET exit=0 marker-found :: IGNORE-OK 3. `lib/tests/mods.test.sh` passes on the repo, fails on a fixture that lacks the loading link, fails on an empty `mods/`, and SKIPs (exit 0) the CLI checks when `claude plugin test` is unavailable (probe by capability, PATH-shadowed `claude` in the control). CHECK: make test suite=lib/tests/mods.test.sh >/dev/null 2>&1 && W=$(mktemp -d) && mkdir -p "$W/mods" "$W/skills" "$W/bin" && cp -R mods/model-router "$W/mods/" && ! MODS_ROOT="$W" bash lib/tests/mods.test.sh >/dev/null 2>&1 && ln -s ../mods/model-router "$W/skills/model-router" && MODS_ROOT="$W" bash lib/tests/mods.test.sh >/dev/null 2>&1 && printf '#!/bin/sh\nexit 1\n' > "$W/bin/claude" && chmod +x "$W/bin/claude" && PATH="$W/bin:$PATH" MODS_ROOT="$W" bash lib/tests/mods.test.sh 2>&1 | grep -q '^SKIP' && E=$(mktemp -d) && mkdir -p "$E/mods" "$E/skills" && ! MODS_ROOT="$E" bash lib/tests/mods.test.sh >/dev/null 2>&1 && echo MODS-SUITE-OK EXPECT: MODS-SUITE-OK - EVIDENCE: pending + EVIDENCE: MET exit=0 marker-found :: MODS-SUITE-OK 4. `doctor.sh` prints a `── Mods ──` section with a ✓ line for model-router; with the link absent (HOME pointed at a scratch `.claude` whose `skills/` lacks the link) the section prints an info line, doctor reaches its summary and exits 0 for that section's sake (no new error). CHECK: out=$(bash doctor.sh 2>&1); echo "$out" | sed -n '/── Mods ──/,/^$/p' | grep -q '✓.*model-router' && H=$(mktemp -d) && mkdir -p "$H/.claude/skills" && o2=$(HOME="$H" bash doctor.sh 2>&1); echo "$o2" | sed -n '/── Mods ──/,/^$/p' | grep -qi 'not linked' && echo "$o2" | grep -q '═══' && echo DOCTOR-MODS-OK EXPECT: DOCTOR-MODS-OK - EVIDENCE: pending -4b. The mod is enabled through the tracked link in a FRESH process: `claude plugin list --json` lists `model-router@skills-dir` with `enabled: true` (run after the dev-mods link is removed, see W6). + EVIDENCE: MET exit=0 marker-found :: DOCTOR-MODS-OK +8. The mod is enabled through the tracked link in a FRESH process: `claude plugin list --json` lists `model-router@skills-dir` with `enabled: true` (run after the dev-mods link is removed, see W6). CHECK: claude plugin list --json 2>/dev/null | python3 -c 'import json,sys; rows=json.load(sys.stdin); ok=any(r.get("id")=="model-router@skills-dir" and r.get("enabled") is True for r in rows); sys.exit(0 if ok else 1)' && echo LOADED-OK EXPECT: LOADED-OK - EVIDENCE: pending + EVIDENCE: MET exit=0 marker-found :: LOADED-OK 5. `CLAUDE.md` has a `## mods/` section naming the `skills/` relative symlink, the `@skills-dir` origin, why not `CLAUDE_CODE_PLUGIN_DIRS`, the gitignored engine-laid files, `~/.claude/.json`, the suite command and how to turn a mod off. CHECK: grep -q '^## mods/' CLAUDE.md && grep -q '@skills-dir' CLAUDE.md && grep -q 'CLAUDE_CODE_PLUGIN_DIRS' CLAUDE.md && grep -q 'mods.test.sh' CLAUDE.md && grep -q '.json' CLAUDE.md && grep -q '@skills-dir": false' CLAUDE.md && echo CLAUDEMD-OK EXPECT: CLAUDEMD-OK - EVIDENCE: pending + EVIDENCE: MET exit=0 marker-found :: CLAUDEMD-OK 6. Health stack on the touched shell files, doctrine census green. CHECK: shellcheck lib/tests/mods.test.sh doctor.sh && make test suite=lib/tests/doctrine-citers.test.sh >/dev/null 2>&1 && echo HEALTH-OK EXPECT: HEALTH-OK - EVIDENCE: pending + EVIDENCE: MET exit=0 marker-found :: HEALTH-OK 7. Judged by reading: no change to settings.json, link.sh or any install script; the suite probes the CAPABILITY (`claude plugin test --help`), bounds every CLI call in time, captures `2>&1`, SKIPs with a reason, fails when no mod is found; doctor's section is fail-soft under `set -euo pipefail` (existence test before readlink, `-ef` comparison, one guarded `claude plugin list --json`, python exits 0 with `unknown` on any parse error), never increments `_LINK_PASS`, says "enabled" not "loaded", treats a missing link as info; the link step is idempotent; CLAUDE.md names the per-machine `"enabled": false` switch, the tracked-settings cost of `enabledPlugins`, and the dev-copy shadowing rule; the CLAUDE.md section is terse English matching the file's style. Q (r2): ordering / A: this contract runs after the floor contract (`2026-10-08-model-router-floor-1835`) is committed and green. [orchestrator] Q (r2): update-all `claude plugin update` over `@skills-dir` / A: accepted residual (one recurring warn), logged in TODO; out of FILE SCOPE. [orchestrator] From a6e200392ca9a5bc12c863924245a52be44db0c8 Mon Sep 17 00:00:00 2001 From: bchanot Date: Fri, 9 Oct 2026 10:52:28 +0200 Subject: [PATCH 14/22] chore(memory): BDR-115 amendment (skills-dir load, floor, kill switch) + journal B2 --- .claude/memory/decisions.md | 1 + .claude/memory/journal.md | 1 + 2 files changed, 2 insertions(+) diff --git a/.claude/memory/decisions.md b/.claude/memory/decisions.md index 1dc23fb..e3ddd31 100644 --- a/.claude/memory/decisions.md +++ b/.claude/memory/decisions.md @@ -1401,4 +1401,5 @@ Branch feature/user-writing-web-rules, UNMERGED (human gate). - **Alternatives rejected**: agents table copying the 21 frontmatter pins in wave 1 (third source of truth, silent override of a frontmatter edit); `scope: agents` / `/route agents` bulk lever (no requirement, sixth precedence tier); `model` param on the model-facing tool (typo → LRN-203 404 class); per-step agent model rewrite (fights engine fallback, beat explicit params); `userConfig` (one config file instead); marketplace install (live symlink repo model); `source`-dependent skill-load reset (kept stale shifts). - **Gates**: plan r1 → r3 through 3 challengers + 1 confirmation (2 BLOCKER + 10 MAJOR closed, [[EVAL-040]]); feater DONE first pass; GATE 0 MET; verifier CONFORME 6/6; security PASS + hardening round (criteria 7-11). - **Refs**: plan `.claude/tasks/plans/2026-10-08-model-router-mod.md`, contract `.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md`, [[BDR-107]], [[BDR-108]], [[BLK-029]], [[LRN-205]], [[LRN-206]]. +- **Amendment (2026-10-09, user decisions 2026-10-08 evening)**: (a) LOAD supersedes the "Load:" line: tracked relative symlink `skills/model-router` → `../mods/model-router`, loaded in place as `model-router@skills-dir` wherever link.sh links `~/.claude/skills`; `CLAUDE_CODE_PLUGIN_DIRS` dropped (absolute path, settings `env` has no `$HOME` expansion, settings.json tracked), local marketplace dropped (`add` writes an absolute path into settings.json). Proven by fresh-process `claude plugin list --json`. (b) PRECEDENCE amended: `ultrathink` and a typed `/effort-` are the main turn's DEFAULT and MINIMUM (floor slot `turnFloor`): sticky `/route` effort > turn route effort > floor > engine, then floored; per axis; mid-turn prompt floors the running turn and the next (`wait` ignored). Rationale: user "un choix explicite bat la phase déduite"; a pure floor made `/effort-low` a no-op (challenge finding). (c) Per-machine kill switch `"enabled": false` in the untracked `~/.claude/model-router.json` (survives `/clear`, a failed reload keeps the previous config); `enabledPlugins` would dirty the tracked settings.json on every machine. (d) Hardening: `/route` composer-only; agent loops effort-only (model fixed at spawn); config caps; typed slash attested at `prompt.submit`. Commits 346d6ae, 1ff608a, 6430ac6; contracts `2026-10-08-model-router-floor-1835`, `2026-10-08-model-router-wiring-1835`; residuals parked in TODO. diff --git a/.claude/memory/journal.md b/.claude/memory/journal.md index c8e2ae8..7ae00aa 100644 --- a/.claude/memory/journal.md +++ b/.claude/memory/journal.md @@ -583,3 +583,4 @@ rules: - model-router W1-A hardening (user go): fresh feater on contract criteria 7-11 → gap round (dead `Loop.frozen`, spawn returns `started` verbatim, tool description) → verifier CONFORME 11/11 → security PASS (1 MEDIUM residual: ReDoS size-bounded only, self-inflicted config; 5 LOW parked). Registries BDR-115, LRN-205, LRN-206, EVAL-040 written on user go (64702d5). Next: live swap of the real mod into the hot-reload folder, then W1-B install + docs. - model-router live checks on Opus 5.5 (user /model, uncommitted settings.json change left to the user): Skill(effort-low) bridge answered in place → next request low; ultrathink turn ran max (engine base medium); Explore without params → claude-sonnet-5-5, 3 steps medium. Loading switched to tracked symlink skills/model-router → @skills-dir (PLUGIN_DIRS non-portable: absolute path, no $HOME expansion, tracked settings); isolated-HOME probe listed/enabled/loaded. User chose floor semantics for ultrathink + typed /effort-. W1-B split: B1 floor (register.ts) + B2 wiring (symlink, gitignore, mods suite, doctor, CLAUDE.md); 6 challengers in flight. Effort shifters skipped this turn: the bridge would overwrite the user's ultrathink until B1 lands. - model-router W1-B1 floor landed (1ff608a): plan r2 from 3 lenses (0 BLOCKER, 4 MAJOR: one decision helper, typed level = default + minimum, mid-turn prompt now + next, per-machine enabled:false), feater DONE, gap round (/clear lost enabled:false, 'ultrathink rule' label, per-axis effort base), hardening round (kill switch keeps previous cfg on failed reload, typed slash attested at prompt.submit vs sub-agent preload). 30 tests, verifier CONFORME 6/6, security PASS ×2 with parked residuals. B2 wiring dispatched next. +- model-router W1-B2 landed (6430ac6): tracked symlink skills/model-router → ../mods/model-router (mode 120000), gitignore, lib/tests/mods.test.sh, doctor Mods section, CLAUDE.md § mods/. Plan r2 (3 lenses, 5 MAJOR), feater DONE, gap round (my `4b.` label unparsed by gates.sh + unbounded probe), verifier CONFORME 8/8, security PASS (LOW: `claude plugin --help` rc 0 weakens the SKIP probe, fail-closed). Dev hot-reload link removed from ~/.claude/dev-mods; user to run /reload-plugins. BDR-115 amended (load, floor, kill switch, hardening). Doc audit (opus) in flight. From b22f8947f9618bdcccf201fb1305d9bfd0052479 Mon Sep 17 00:00:00 2001 From: bchanot Date: Fri, 9 Oct 2026 11:04:10 +0200 Subject: [PATCH 15/22] =?UTF-8?q?docs:=20README=20effort=20routing=20+=20/?= =?UTF-8?q?route,=20USAGE,=20ARCHITECTURE=20mods/,=20CHANGELOG=20=E2=80=94?= =?UTF-8?q?=20model-router=20wave=201?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- ARCHITECTURE.md | 2 ++ CHANGELOG.md | 1 + README.md | 23 +++++++++++++++++++---- USAGE.md | 3 +++ 4 files changed, 25 insertions(+), 4 deletions(-) diff --git a/ARCHITECTURE.md b/ARCHITECTURE.md index 3c0d2d0..c0d616a 100644 --- a/ARCHITECTURE.md +++ b/ARCHITECTURE.md @@ -26,6 +26,7 @@ claude-config/ ├── rules/ # Rule files deployed to ~/.claude/rules (path-scoped or always-on) ├── agents/ # Execution units called by skills (never invoked directly) ├── skills/ # Entry points invoked via /skill-name +├── mods/ # Claude Code mods (function-hooks plugins), loaded through the skills/ symlink ├── skills-external/ # Vendored skill packs: gstack submodule, design skills, superpowers, agent-skills, MengTo scroll skills, 21st and Higgsfield packs (machine-owned copies gitignored) ├── templates/ # Per-project templates (CLAUDE.md, settings, memory registries, deploy runbook, gitignore) └── lib/ # Shared libs: gitflow, profiles, vendoring, effort pins, gates, archetypes, tests @@ -35,5 +36,6 @@ claude-config/ - `skills/` = entry points you invoke via `/skill-name` - `agents/` = execution units called by skills (never invoked directly by user) +- `mods/` = Claude Code mods (function-hooks plugins); each loads through the tracked symlink `skills/` as `@skills-dir`, live at the next session - `templates/` = symlinked to `~/.claude/templates/` — copy into projects via `/onboard` or manually - **Graphify** builds a knowledge graph of any codebase (`/graphify query`), producing a navigable wiki in `graphify-out/wiki/`. This map helps Claude understand project structure, find relevant code faster, and reason across files. Essential for large-scope tasks (multi-file features, complex bugs, architectural changes). Small tasks should skip it and read files directly. Proposed only from 200 tracked code files: the session-start banner informs, the user decides; nothing builds a graph without that go. diff --git a/CHANGELOG.md b/CHANGELOG.md index 26db705..4dac3e8 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -7,6 +7,7 @@ Format follows [Keep a Changelog](https://keepachangelog.com/) and this project ## [Unreleased] ### Added +- **model-router mod**: `mods/model-router/`, a Claude Code mod (function-hooks plugin), routes the effort of every main-loop request from a phase table, along with the model and effort of the built-in sub-agents (Explore on sonnet/medium, Plan on opus/xhigh). It answers `Skill(effort-*)` itself, so the five `effort-*` skills no longer load while it is on. `ultrathink` in a prompt and a typed `/effort-` set the main turn's default and minimum effort. The model gets a `route` tool and the user a `/route` command (`show|clear|off|on|reload||model= effort=|switch on|off|verbose on|off`). Optional per-machine config `~/.claude/model-router.json`, where `"enabled": false` turns it off on that machine. The spinner suffix and the status line show the route in force. It loads in every session through the tracked symlink `skills/model-router` (`model-router@skills-dir`). `make doctor` gains a Mods section; suite `make test suite=lib/tests/mods.test.sh`. Known limits: the main loop switches model only with `mainModelSwitch` on (default off, one cold-cache step per switch into another model), and the hooks send full model ids, so the `models` table has to follow new versions. - **Manual-push mode**: `git config gitflow.autopush false` (human-set) now stops every push the gitflow lib makes, not only the post-commit / post-merge hooks. `gitflow start` and `finish` branch, commit and merge locally and push nothing; `gitflow delete` leaves the `origin/` copy in place and prints `git push origin --delete
` for the user to run. `hooks/unpushed-guard.sh` stays silent at turn end in this mode and opens each session with one `ℹ manual push mode:` line counting the commits no remote holds across every local branch; an unparseable or unreadable `gitflow.autopush` value is treated as manual push mode too, and that line names it. `hooks/push-guard.sh` (PreToolUse, `Bash|Monitor`) refuses any `git push` Claude types while `gitflow.autopush` reads false in the session cwd or in a literal `-C`/`cd` directory the command names (global config counts outside a repo); the refusal tells the user to run it with `! git push`. It reads the mode through the same lib verb as every other reader and fails closed: an unparseable or unreadable value reads as manual, and an internal error, a missing `lib/gitflow.sh`, more than 20 distinct directory tokens in one command (capped before any token is classified), a `cd`/`-C` directory token mixing quoted and unquoted parts, or a payload jq cannot parse whose raw text looks like a push refuse the push (these pathological cases fire in auto mode too). Directory tokens are read as whole shell words, adjacent quoted segments and backslash escapes included. In manual mode it over-blocks any command where a `push` word follows a `git` token; the misses listed in its header fall to a new `autoMode.soft_deny` rule that no request in the turn clears. The session banner adds `🔒 push : manual (autopush=false) — ! git push` when the key reads false, and `🔒 push : manual (autopush bad) — ! git push` when the value is invalid. Skills read the mode through a new lib verb, `bash ~/.claude/lib/gitflow.sh push-mode`: it prints `auto`, `manual` or `invalid` (rc 0) and names an invalid value on stderr (printable characters only, 64 at most). It is the one reader a skill may call, since the `git config` read of the key is denied to Claude. Skills push nothing on their own, except the `/release-candidate` tag in auto-push mode on an explicit go. Every "on origin" or "not pushed" line they print comes from `git rev-list --count origin/
..
` read after the fact, with the complete `! git …` command when something is left for the user to push. An invalid value (anything but unset, true or false, or a read that fails) is manual push mode for every reader and is named where it is read (see Fixed). Tests: `lib/gitflow-test.sh` T11b (push-mode verb), T18m and T18q blocks, `lib/tests/unpushed-guard.test.sh` T10-T16, `lib/tests/push-guard.test.sh` (98 checks). ### Changed diff --git a/README.md b/README.md index 1c47c19..801d907 100644 --- a/README.md +++ b/README.md @@ -88,7 +88,7 @@ every call site. | doc-syncer | sonnet pin; audit mode dispatched `model="opus"` | two-mode: audit (drift judgment, opus) / patch (mechanical apply, sonnet) | | handover-doc-writer | sonnet pin; synthesize mode dispatched `model="opus"` | two-mode: synthesize (opus) / render (sonnet) — client deliverable | | interviewer, client-handover-writer | unpinned (inline-load = session model) | they ARE the main loop — a frontmatter pin would be inert | -| Explore (built-in) | inherit session (Fable/Opus) | search feeds reflection — kept on the big model, not pinned down | +| Explore, Plan (built-in) | model-router mod: Explore → sonnet/medium, Plan → opus/xhigh, set at spawn; an explicit `model=` on the call wins; mod off: inherit session | built-ins routed per phase by the mod | The pure-execution skills `/doc`, `/status`, `/commit-change`, `/release-candidate` **dispatch** their agent (instead of inline-loading it) @@ -108,12 +108,26 @@ entry level (`/status` low … `/ship-feature` xhigh); the vendored externals `lib/effort-pins.txt`, re-applied by `lib/effort-pins.sh` after every vendoring step. Orchestrators shift per phase through the `effort-low` … `effort-max` skills (`lib/effort-shift.md`, always sent with another tool -call: a lone Skill call applies nothing). Model pins stay tier aliases +call: a lone Skill call applies nothing (mod off)). Model pins stay tier aliases (`sonnet`, `opus`, `haiku`, `fable`): the latest version of a tier is also the cheapest or same-priced, so the quality/price trade-off is tier × effort, never version. Census `lib/tests/effort-routing.test.sh`; transcript audit `python3 lib/effort-audit.py`. +### model-router mod + +`mods/model-router/` is a Claude Code mod (a function-hooks plugin) that applies this table per request. It loads in every session through the tracked symlink `skills/model-router`, as `model-router@skills-dir`. + +- Main loop: every request gets the effort of the phase in force. The mod answers `Skill(effort-*)` itself and applies the level from the next request on, so the five `effort-*` skills no longer load while it is on. +- Built-in sub-agents: Explore runs on sonnet/medium, Plan on opus/xhigh. An explicit `model` on the Agent call wins. +- User floor: `ultrathink` in a prompt, or a typed `/effort-`, sets the main turn's default and minimum effort. +- `/route` (user command) shows or sets the route: `show`, `clear`, `off`, `on`, `reload`, a phase name, `model= effort=`, `switch on|off`, `verbose on|off`. The model sets routes through a `route` tool. +- The spinner suffix and the status line under the prompt show the route in force. + +Optional per-machine config: `~/.claude/model-router.json`. Keys: `models` (alias → full id), `windows` (context window per full id), `phases`, `agents`, `skills`, `prompt` (rules), `mainModelSwitch` (default `false`), `verbose` (default `false`), `spinner` (default `true`), `enabled` (default `true`; `false` turns the mod off on that machine). `/route reload` re-reads it. + +Limits: the main loop changes model only with `mainModelSwitch` on, and each switch into another model costs one cold-cache step. The hooks send full model ids, so the `models` table has to follow new model versions. + --- ## Install notes @@ -204,7 +218,8 @@ a different package, ships its own conflicting `graphify` bin) — see | `/profile` | Activate a skill profile (web / seo / web-full / full / max / backend / design / dev / qa / audit / minimal) (default: full) | | `/tour` | Grouped all-axes sweep — cleanup + security + reconcile + doc, fix and loop until clean | | `/site-motion` | Site-level motion: scroll engine choice, page transitions, pin/scrub sequencing across a page or Astro route (design stack) | -| `/effort-low` … `/effort-max` | Effort shifters the orchestrators send per phase; type `/effort-max` to re-run a stuck turn at maximum | +| `/effort-low` … `/effort-max` | Effort shifters the orchestrators send per phase (answered by the model-router mod when on); typed by you, they set the main turn's default and minimum effort | +| `/route` | model-router mod: show or set the main-loop route (show, clear, off, on, reload, , model=… effort=…, switch on\|off, verbose on\|off) | > This table lists personal skills. Gstack skills (investigate, review, retro, > office-hours, cso…) and marketplace plugins add many more — run @@ -452,7 +467,7 @@ make profile-reset # go to the default profile (full) make new-skill name=myskill # scaffold agent + skill files ``` -`doctor.sh` checks: symlinks, GStack submodule, vendored skills (curl-pinned externals in `plugins.lock.json` + `link.sh`'s `EXTERNAL_SKILLS`, per the active profile), Playwright browser cache, prerequisites (git, Node, Cargo, Python, Claude Code), plugins, permissions, token budget, config consistency, git hooks (global core.hooksPath + generated githooks/), scratchpad (TMPDIR quota), Higgsfield CLI and session, seo-data layer. +`doctor.sh` checks: symlinks, GStack submodule, vendored skills (curl-pinned externals in `plugins.lock.json` + `link.sh`'s `EXTERNAL_SKILLS`, per the active profile), Playwright browser cache, prerequisites (git, Node, Cargo, Python, Claude Code), plugins, permissions, token budget, config consistency, git hooks (global core.hooksPath + generated githooks/), mods (loading link and `@skills-dir` state), scratchpad (TMPDIR quota), Higgsfield CLI and session, seo-data layer. --- diff --git a/USAGE.md b/USAGE.md index 24a2c54..3a002d1 100644 --- a/USAGE.md +++ b/USAGE.md @@ -163,6 +163,7 @@ Tu veux... | `/pdf-translate` | Traduire un PDF vers une autre langue | Sortie HTML fidèle (images, layout, style préservés) | | `/impeccable` | Audit/polish design + détecteur anti-slop déterministe | 23 verbes ; `npx impeccable detect` (exit 0/2) | | `/tour` | Sweep groupé sur un ou plusieurs projets | Sécu + nettoyage + reconcile + doc, boucle jusqu'à un pass propre | +| `/route` | Voir ou fixer la route du mod model-router | show / clear / off / on / reload / / model=… effort=… / switch on\|off / verbose on\|off | | `/profile` | Changer le profil de skills | web / seo / web-full / full / max / backend / design / dev / qa / audit / minimal | > Cette table couvre les skills personnels principaux. Les plugins (gstack, @@ -185,6 +186,8 @@ la main relance un tour bloqué au maximum. Les skills externes vendorés `lib/effort-pins.txt`. Un skill chargé seul par Claude n'applique pas son niveau : il doit partir avec un autre appel d'outil dans le même message. +Avec le mod model-router (`mods/model-router/`, actif dans chaque session), le niveau suit la phase à chaque requête. Le mod répond lui-même à `Skill(effort-*)` : le niveau s'applique dès la requête suivante et le texte des skills `effort-*` n'est plus chargé. Écrire `ultrathink` dans un prompt, ou taper `/effort-`, fixe le niveau par défaut et le minimum du tour principal. Les sous-agents intégrés suivent leur route : Explore en sonnet/medium, Plan en opus/xhigh. `/route` affiche ou fixe la route (`/route show`, `/route clear`, `/route off`). La config par machine, optionnelle, vit dans `~/.claude/model-router.json` ; `"enabled": false` y coupe le mod sur cette machine. + ## Les plugins — décision rapide ``` From e79db7e6df8ff2bc6c2c7ba388f53adc4edc7a8a Mon Sep 17 00:00:00 2001 From: bchanot Date: Fri, 9 Oct 2026 11:16:32 +0200 Subject: [PATCH 16/22] =?UTF-8?q?chore(memory):=20model-router=20wave=201?= =?UTF-8?q?=20closed=20=E2=80=94=20TODO=20W2=20queued,=20journal?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- .claude/memory/journal.md | 1 + .claude/tasks/TODO.md | 3 ++- 2 files changed, 3 insertions(+), 1 deletion(-) diff --git a/.claude/memory/journal.md b/.claude/memory/journal.md index 7ae00aa..3ccf88d 100644 --- a/.claude/memory/journal.md +++ b/.claude/memory/journal.md @@ -584,3 +584,4 @@ rules: - model-router live checks on Opus 5.5 (user /model, uncommitted settings.json change left to the user): Skill(effort-low) bridge answered in place → next request low; ultrathink turn ran max (engine base medium); Explore without params → claude-sonnet-5-5, 3 steps medium. Loading switched to tracked symlink skills/model-router → @skills-dir (PLUGIN_DIRS non-portable: absolute path, no $HOME expansion, tracked settings); isolated-HOME probe listed/enabled/loaded. User chose floor semantics for ultrathink + typed /effort-. W1-B split: B1 floor (register.ts) + B2 wiring (symlink, gitignore, mods suite, doctor, CLAUDE.md); 6 challengers in flight. Effort shifters skipped this turn: the bridge would overwrite the user's ultrathink until B1 lands. - model-router W1-B1 floor landed (1ff608a): plan r2 from 3 lenses (0 BLOCKER, 4 MAJOR: one decision helper, typed level = default + minimum, mid-turn prompt now + next, per-machine enabled:false), feater DONE, gap round (/clear lost enabled:false, 'ultrathink rule' label, per-axis effort base), hardening round (kill switch keeps previous cfg on failed reload, typed slash attested at prompt.submit vs sub-agent preload). 30 tests, verifier CONFORME 6/6, security PASS ×2 with parked residuals. B2 wiring dispatched next. - model-router W1-B2 landed (6430ac6): tracked symlink skills/model-router → ../mods/model-router (mode 120000), gitignore, lib/tests/mods.test.sh, doctor Mods section, CLAUDE.md § mods/. Plan r2 (3 lenses, 5 MAJOR), feater DONE, gap round (my `4b.` label unparsed by gates.sh + unbounded probe), verifier CONFORME 8/8, security PASS (LOW: `claude plugin --help` rc 0 weakens the SKIP probe, fail-closed). Dev hot-reload link removed from ~/.claude/dev-mods; user to run /reload-plugins. BDR-115 amended (load, floor, kill switch, hardening). Doc audit (opus) in flight. +- model-router wave 1 CLOSED on feature/model-router-mod (15 commits ahead of develop, nothing pushed: manual mode). Docs: opus audit SIGNIFICANT (README Explore row false, no mention of the mod) → user go all 10 → first patch self-reverted by the MINOR-envelope oracle (plan carried MINOR labels; a new heading exceeds the envelope) → re-dispatched with SIGNIFICANT provenance → b22f894. Full `make test`: every suite green except the pre-existing env red design-tool-gate (21st CLI present, not hermetic). Open for the user: /reload-plugins here; merge decision (gitflow finish); settings.json own change; W2 migration queued. diff --git a/.claude/tasks/TODO.md b/.claude/tasks/TODO.md index 400a906..962d6d9 100644 --- a/.claude/tasks/TODO.md +++ b/.claude/tasks/TODO.md @@ -12,7 +12,8 @@ migration of shifters/pins/model-gate in wave 2 after proof; names model-router - [ ] W1-B1 residuals (security 2026-10-09, accepted): first-load failure of the override falls to defaults (`enabled: true`); non-boolean `enabled` drops to the default on reload; `skill.prompt` preload guard is a heuristic (`loops.size > 0`; a preload during spawn or a non-composer `/effort-*` while idle still writes the floor); marker not bound to a valid level; `/route reload` answers "config reloaded" even when the previous config was kept; transient missing override lifts a config-set off; `String(err)` of a JSON parse in the local log. Display: `/route show` folds the floor into the effort while off; model-axis text with a model-only sticky and the switch on. - [x] W1-B2 wiring (contract `2026-10-08-model-router-wiring-1835`, 2026-10-09): tracked symlink `skills/model-router` → `../mods/model-router` loads as `model-router@skills-dir` (fresh-process `claude plugin list --json` proves it), `.gitignore` `mods/*/tsconfig.json`, `lib/tests/mods.test.sh` (4 checks, capability probe, bounded, SKIP), doctor `── Mods ──` fail-soft, CLAUDE.md `## mods/`. Plan r2 from 3 lenses (5 MAJOR), feater DONE, gap round (4b label + unbounded probe), verifier CONFORME 8/8, security PASS - [ ] W1-B2 residuals (accepted): `claude plugin --help` returns 0 → the suite's capability probe can pass on a CLI without `plugin test` and then FAIL instead of SKIP (fail-closed; fix = grep the probe output for the test usage line); doctor `claude plugin list --json` unbounded; doctor `echo -e` helpers interpolate `$_mod` (tracked folder names only); fallback timeout guard orphans grandchildren (only without coreutils timeout); `update-all.sh` runs `claude plugin update` over `@skills-dir` → one recurring warn (needs an update-all edit) -- [ ] W1 close-out: doc-sync (README / USAGE / CHANGELOG), live swap (remove the hot-reload link in `~/.claude/dev-mods//`, user runs `/reload-plugins`), BDR-115 amendment (loading = skills-dir link; floor semantics), journal +- [x] W1 close-out (2026-10-09): doc-sync opus audit SIGNIFICANT → user go all 10 → patched (README effort routing + Explore row + /route, USAGE, ARCHITECTURE mods/, CHANGELOG Unreleased) b22f894; hot-reload link removed from `~/.claude/dev-mods//`; BDR-115 amended (a6e2003). Pending user: `/reload-plugins` in this session (new sessions load the skills-dir copy by themselves); merge decision on feature/model-router-mod (`gitflow finish`, human signal) +- [ ] W2 migration (after wave 1 proven in daily use): 15 skills `Skill(effort-*)` → `route` tool calls; remove `skills/effort-*`, `lib/effort-pins.txt/.sh`, install/update steps, `effort:` frontmatter on skills and agents; repo agents into the mod's `agents` table (verify the spawn/first-step ordering first); `lib/model-gate.md` + `lib/model-check.sh` → mod rule; census tests repointed; `lib/effort-shift.md` rewritten; docs - [ ] W2 migration: 15 skills off `Skill(effort-*)`, remove shifters + effort-pins + model-gate, census repointed, docs - [ ] W3 optional: step heuristics, haiku classifier, quota-aware downgrade, A/B From 2d8cd6bf4c040b6062ac338e074c91ef03494578 Mon Sep 17 00:00:00 2001 From: bchanot Date: Fri, 9 Oct 2026 12:40:19 +0200 Subject: [PATCH 17/22] chore(tasks): model-router W1-C contract + plan (absolute tiers, breaker fallback, derived phases) --- .claude/tasks/TODO.md | 1 + .../2026-10-09-model-router-tiers-1237.md | 37 ++++ .../plans/2026-10-08-model-router-mod.md | 15 ++ .../2026-10-09-model-router-tiers-1237.md | 188 ++++++++++++++++++ 4 files changed, 241 insertions(+) create mode 100644 .claude/tasks/contracts/2026-10-09-model-router-tiers-1237.md create mode 100644 .claude/tasks/plans/2026-10-09-model-router-tiers-1237.md diff --git a/.claude/tasks/TODO.md b/.claude/tasks/TODO.md index 962d6d9..321aaf4 100644 --- a/.claude/tasks/TODO.md +++ b/.claude/tasks/TODO.md @@ -13,6 +13,7 @@ migration of shifters/pins/model-gate in wave 2 after proof; names model-router - [x] W1-B2 wiring (contract `2026-10-08-model-router-wiring-1835`, 2026-10-09): tracked symlink `skills/model-router` → `../mods/model-router` loads as `model-router@skills-dir` (fresh-process `claude plugin list --json` proves it), `.gitignore` `mods/*/tsconfig.json`, `lib/tests/mods.test.sh` (4 checks, capability probe, bounded, SKIP), doctor `── Mods ──` fail-soft, CLAUDE.md `## mods/`. Plan r2 from 3 lenses (5 MAJOR), feater DONE, gap round (4b label + unbounded probe), verifier CONFORME 8/8, security PASS - [ ] W1-B2 residuals (accepted): `claude plugin --help` returns 0 → the suite's capability probe can pass on a CLI without `plugin test` and then FAIL instead of SKIP (fail-closed; fix = grep the probe output for the test usage line); doctor `claude plugin list --json` unbounded; doctor `echo -e` helpers interpolate `$_mod` (tracked folder names only); fallback timeout guard orphans grandchildren (only without coreutils timeout); `update-all.sh` runs `claude plugin update` over `@skills-dir` → one recurring warn (needs an update-all edit) - [x] W1 close-out (2026-10-09): doc-sync opus audit SIGNIFICANT → user go all 10 → patched (README effort routing + Explore row + /route, USAGE, ARCHITECTURE mods/, CHANGELOG Unreleased) b22f894; hot-reload link removed from `~/.claude/dev-mods//`; BDR-115 amended (a6e2003). Pending user: `/reload-plugins` in this session (new sessions load the skills-dir copy by themselves); merge decision on feature/model-router-mod (`gitflow finish`, human signal) +- [ ] W1-C adaptive tiers (user 2026-10-09: "un système logique et optimisé", no /route typed, works on a haiku session, falls back when fable has no credit): contract `2026-10-09-model-router-tiers-1237`, plan same slug. Phases name absolute tiers (best/big/work/cheap), availability breaker (turn error/refusal + engine auto switch, 15 min cooldown), fallback chain fable→opus→sonnet→haiku, main upgrade on / downgrade gated, prompt default rules (plan/reflect keywords FR+EN) + dispatch push/pop + optional classifier - [ ] W2 migration (after wave 1 proven in daily use): 15 skills `Skill(effort-*)` → `route` tool calls; remove `skills/effort-*`, `lib/effort-pins.txt/.sh`, install/update steps, `effort:` frontmatter on skills and agents; repo agents into the mod's `agents` table (verify the spawn/first-step ordering first); `lib/model-gate.md` + `lib/model-check.sh` → mod rule; census tests repointed; `lib/effort-shift.md` rewritten; docs - [ ] W2 migration: 15 skills off `Skill(effort-*)`, remove shifters + effort-pins + model-gate, census repointed, docs - [ ] W3 optional: step heuristics, haiku classifier, quota-aware downgrade, A/B diff --git a/.claude/tasks/contracts/2026-10-09-model-router-tiers-1237.md b/.claude/tasks/contracts/2026-10-09-model-router-tiers-1237.md new file mode 100644 index 0000000..57eb85d --- /dev/null +++ b/.claude/tasks/contracts/2026-10-09-model-router-tiers-1237.md @@ -0,0 +1,37 @@ +# CONTRACT — model-router-tiers (wave 1-C: absolute tiers, availability fallback, derived phases) +- date: 2026-10-09 | flow: feat | branch: feature/model-router-mod +- status: active + +## REQUEST (verbatim — IMMUTABLE) +User (fr): "j'aimerais ne pas avoir a reflechir a tout ca, donc ne pas lancer les /route moi meme par exemple. On vois le reflect, orchestratm escalate et plan session (fable) par du preincipe que la sessino est sur fable de base ? Si on est sur haiku j'aimerias que ca fonctionne aussi avec les models adapte. et d'ailleurs que se passe til quand on a plus de credit fable, cela passe sur opus xhigh ? il faut se fallback. ET surtout oui avec un systeme adaptatif et pour la reflexion pour trouver une solution ou des idees, il faut le meilleur, pour en faire le plan en se basant sur cette reflexion. Bref un systeme logique et optimise. /ultrathink . ensuite je reload puginm tu test, on commit puis on passe a la vague 2" + +## CLARIFICATIONS +Q: "session" phases / A: no phase keeps "the session model" any more: every phase names a TIER (`best`, `big`, `work`, `cheap`), an ordered list of aliases; the first AVAILABLE alias wins. plan/reflect/orchestrate/escalate → `best` (fable, then opus, then sonnet). A haiku session asked to plan runs the plan on fable. [user: "si on est sur haiku j'aimerais que ca fonctionne aussi avec les models adaptés"] +Q: what is "available" / A: the mod cannot read per-model quota (rateLimits are account windows: five_hour, seven_day). Availability = a circuit breaker: a model is DOWN for `cooldownMinutes` (default 15) after a main or agent turn ends with `reason: 'error'` or `'refusal'` on it, or when the engine itself switched away from it (`classic.PostModelSwitch`, `source: 'auto'`). A down model is skipped in every tier and in the `fallback` chain; `/route show` lists down models with their reset time; `/route reload` clears the breaker. [orchestrator — derived from the engine's declarations] +Q: no credit left on fable / A: main loop on a down model → next available alias of the `fallback` chain (`fable, opus, sonnet, haiku`), effort unchanged (plan stays xhigh → "opus xhigh"), always allowed (a down model yields nothing), logged and shown. [user: "il faut se fallback"] +Q: main-loop model moves / A: UPGRADE (phase tier ranks above the current model) allowed by default (`mainUpgrade: true`); DOWNGRADE (cheaper model) still gated by `mainModelSwitch` (default false: one cold-cache step per switch, and haiku's window may not fit); same rank → no switch. [orchestrator — cost/quality trade-off, stated to the user] +Q: no `/route` typed by the user / A: phases come from (1) prompt rules at turn start (keyword → phase as the turn's DEFAULT route, source 'prompt', overridable; `ultrathink` stays a FLOOR), (2) the model's own `route` calls and the skills table, (3) derived: an Agent dispatch from main pushes `orchestrate` and the previous route is restored when the last live agent ends, (4) optional classifier (`classifier: false` by default) that asks the engine's small model for a phase label when no rule matched. [user: "ne pas lancer les /route moi-même"] +Q: default prompt rules / A: floor: `\bultrathink\b` → escalate. Defaults: `\b(plan|planifie|planning|brainstorm|architecture|con[cç]ois|design)\b` → plan; `\b(pourquoi|why|explique|explain|analyse|analyze|comprendre|understand|review|audit)\b` → reflect. No default rule lowers a turn (no `mechanical` rule): lowering is explicit (route tool, skills). [orchestrator — conservative defaults, user-editable in the override] +Q: typed `/effort-` / A: unchanged (floor + default of the turn, B1). + +## ACCEPTANCE CRITERIA +1. Suite green: `claude plugin test` passes with at least 40 `test(` calls; `claude plugin validate` passes with no warning; no line over 80 chars; no `any` type. + CHECK: cd mods/model-router && out=$(claude plugin test . 2>&1); rc=$?; echo "$out" | tail -n 3; [ $rc -eq 0 ] && [ "$(grep -cE '^\s*test\(' hooks/register.test.ts)" -ge 40 ] && v=$(claude plugin validate . 2>&1) && echo "$v" | grep -q 'Validation passed' && ! echo "$v" | grep -qi 'warning' && ! grep -nE '.{81,}' hooks/register.ts hooks/register.test.ts && ! grep -nE ':\s*any\b||as any\b' hooks/register.ts && echo TIERS-SUITE-OK + EXPECT: TIERS-SUITE-OK + EVIDENCE: pending +2. Type-check clean against this build's declarations. + CHECK: T=/Users/b.chanot/Documents/claude/mods/model-router/.claude-plugin/types; W=$(mktemp -d) && printf '{"compilerOptions":{"target":"es2023","lib":["es2023"],"types":[],"module":"esnext","moduleResolution":"bundler","strict":true,"noUncheckedIndexedAccess":true,"noEmit":true,"skipLibCheck":true,"jsx":"react","jsxFactory":"h","jsxFragmentFactory":"Fragment"},"include":["%s/claude-code/index.d.ts","%s/claude-code-tools/index.d.ts","%s/hooks"]}' "$T" "$T" "$PWD/mods/model-router" > "$W/tsconfig.json" && (cd "$W" && npx --yes -p typescript@5 tsc -p tsconfig.json) && echo TSC-OK + EXPECT: TSC-OK + EVIDENCE: pending +3. Tiers in the config: `tiers` (best/big/work/cheap), `fallback`, `cooldownMinutes`, `mainUpgrade`, `classifier` exist in DEFAULT_CONFIG; every default phase names a tier, none a bare model; `/route show` prints the resolved model of each phase and a `down:` line. + CHECK: cd mods/model-router/hooks && grep -q "tiers:" register.ts && grep -q "fallback:" register.ts && grep -q "cooldownMinutes" register.ts && grep -q "mainUpgrade" register.ts && grep -q "classifier" register.ts && [ "$(awk '/^const DEFAULT_CONFIG/,/^}/' register.ts | grep -cE "tier: '(best|big|work|cheap)'")" -ge 10 ] && ! awk '/^const DEFAULT_CONFIG/,/^}/' register.ts | grep -qE "^\s+[a-z]+: \{ model: '" && grep -q "down:" register.ts && echo TIERS-CONFIG-OK + EXPECT: TIERS-CONFIG-OK + EVIDENCE: pending +4. Tests prove (names contain the quoted word): `tier` — a `plan` route on a session model `claude-haiku-4-5-20251001` makes the main step run on `claude-fable-5-1` at xhigh (upgrade, default on); `downgrade` — a `mechanical` route on a fable session leaves the model unchanged while `mainModelSwitch` is off; `fallback` — after a main `turn.complete` with `reason: 'error'` on fable, the next main step runs on `claude-opus-5-5` with its effort unchanged, and after `/route reload` fable is used again; `breaker` — a `classic.PostModelSwitch` with `source: 'auto'` from fable marks fable down and `/route show` lists it; `spawn` — `Explore` spawns on `claude-opus-5-5` while sonnet is down; `derived` — an Agent tool call from main sets `orchestrate` and the previous route (`plan`) is back after the last agent's `turn.complete`; `default rule` — a prompt "planifie la migration" sets the `plan` route as the turn default and a later `route` tool call overrides it; `floor` tests from B1 still pass. + CHECK: cd mods/model-router/hooks && for w in tier downgrade fallback breaker spawn derived "default rule"; do grep -qE "test\('[^']*$w" register.test.ts || { echo "missing test: $w"; exit 1; }; done && echo TIERS-TESTS-OK + EXPECT: TIERS-TESTS-OK + EVIDENCE: pending +5. Judged by reading: ONE resolver (`resolveModel`) turns a tier name, an alias or a full id into the first AVAILABLE full id (tier → list → skip down → alias → id; a bare alias or id passes through even when down, since explicit means explicit); rank = position in `fallback`; the main-loop decision is: route model wanted → upgrade allowed by `mainUpgrade`, downgrade gated by `mainModelSwitch` + window, same rank → keep; no route and current model down → next available in `fallback`; the breaker is fed only by `turn.complete` `reason` in `error|refusal` (main: the last main step's model; agent: the loop's spawn model) and by `classic.PostModelSwitch` `source: 'auto'` (from_model), never by an aborted turn; the derived `orchestrate` push/pop never overrides a route the model declared after the dispatch; prompt default rules write `turnMain` (source 'prompt'), floor rules write `turnFloor`; the classifier runs only when `classifier` is true and no rule matched, through `$.model.classify`, labels limited to the phase names plus `other`, any failure → no route; explicit Agent params still win; agent model fixed at spawn; every B1/1-A criterion still holds; no function over 25 logic lines; truthful texts name the resolved model and say "fallback" when the breaker chose it. + +## FILE SCOPE +mods/model-router/hooks/register.ts · mods/model-router/hooks/register.test.ts diff --git a/.claude/tasks/plans/2026-10-08-model-router-mod.md b/.claude/tasks/plans/2026-10-08-model-router-mod.md index b5ff7a6..6dab05a 100644 --- a/.claude/tasks/plans/2026-10-08-model-router-mod.md +++ b/.claude/tasks/plans/2026-10-08-model-router-mod.md @@ -116,6 +116,21 @@ does many different things inside one run. So: Explore without params spawned on `claude-sonnet-5-5`, all 3 steps `medium` (no spawn/step race observed). +## Decisions 2026-10-09 (user, after the wave-1 table was shown) +- No "session model" phase: every phase names an ABSOLUTE tier (best = fable, + opus, sonnet · big = opus, fable, sonnet · work = sonnet, opus · cheap = haiku, + sonnet); a haiku session asked to plan runs on fable. Reflection and planning + always on the best available model. +- Availability = circuit breaker (turn error/refusal, engine auto switch), not + quota reading (rateLimits are account windows). Fallback fable → opus → + sonnet → haiku, effort unchanged ("plus de crédit fable → opus xhigh"). +- The user never types /route: prompt default rules, dispatch push/pop + (orchestrate while agents run, previous route restored), skills table + (wave 2), optional classifier. Main upgrade allowed by default, downgrade + still gated (cold-cache cost). +- Sequence agreed: 1-C lands → user /reload-plugins → live test → commit → + wave 2. + ## Wave 0 — spike (dev-mods folder, hot reload, this session) - [x] W0.1 minimal mod: `/route` command, `route` tool, `turn.step` logging + rewrite, `agent.spawn` rewrite, `ultrathink` → max, spinner suffix diff --git a/.claude/tasks/plans/2026-10-09-model-router-tiers-1237.md b/.claude/tasks/plans/2026-10-09-model-router-tiers-1237.md new file mode 100644 index 0000000..639d822 --- /dev/null +++ b/.claude/tasks/plans/2026-10-09-model-router-tiers-1237.md @@ -0,0 +1,188 @@ +# PLAN — model-router wave 1-C: absolute tiers, availability fallback, derived phases (dispatch-ready) +Contract: .claude/tasks/contracts/2026-10-09-model-router-tiers-1237.md +Code: mods/model-router/hooks/register.ts (read in full) and register.test.ts. +API truth: mods/model-router/.claude-plugin/types/claude-code/index.d.ts +(TurnStepInput, TurnCompleteInput + TurnCompleteReason, SessionRateLimit, +`classic.PostModelSwitch` → PostModelSwitchHookInput, `$.session.model`, +`$.model.classify`, 'claude-code/testing'), .../claude-code-tools/index.d.ts. + +## Why (user, 2026-10-09) +1. "Session model" phases assumed Fable. On a haiku session, `plan` at xhigh on + haiku is wrong. Phases must name ABSOLUTE tiers. +2. When Fable has no credit left, routing must fall back (plan → opus xhigh). +3. The user never types `/route`: phases are derived (prompt wording, + dispatch spans, skills) or declared by the model. +4. Reflection and planning always get the best available model. + +## Engine facts (read in the declarations today) +- `SessionRateLimit.kind` ∈ five_hour | seven_day | spend_limit: account + windows, NOT per model → availability cannot be read from quotas. +- A dead request: `turn.step` result `usage: null`, `stopReason: null`; + `turn.complete` `reason: 'error'` (retries exhausted) or `'refusal'` + (refused, no fallback model); `'aborted'` = the user interrupted. +- `classic.PostModelSwitch` fires on every main-model change with + `from_model`, `to_model`, `source` ∈ command|picker|sdk|auto|resume, + `context_tokens`, `prompt_cache_warm`. +- `turn.step` `e.model` = the id the engine resolved (session's or a + fallback's); `$.session.model()` = the main loop's model as `/model` shows. +- `$.model.classify(text, labels, { model? })` → label | undefined, rejects + on failure; default = the engine's small fast model. + +## Config (additions; everything else unchanged) +```ts +type Route = { model?: string; effort?: Level } // model: TIER name, alias or full id +type PromptRule = { pattern: string; phase: string; mode?: 'floor' | 'default' } +type Config = { + …existing… + tiers: Record // tier → ordered alias preference + fallback: string[] // alias order, best first; = rank + cooldownMinutes: number // breaker hold + mainUpgrade: boolean // main loop may switch UP to a phase's tier + classifier: boolean // ask the small model when no rule matched +} +``` +DEFAULT_CONFIG changes: +- `tiers`: best `['fable','opus','sonnet']`, big `['opus','fable','sonnet']`, + work `['sonnet','opus']`, cheap `['haiku','sonnet']`. +- `fallback`: `['fable','opus','sonnet','haiku']`. `cooldownMinutes: 15`. + `mainUpgrade: true`. `classifier: false`. +- phases: plan `{ model: 'best', effort: 'xhigh' }`, reflect `{ best, high }`, + orchestrate `{ best, medium }`, escalate `{ best, max }`, judge `{ big, xhigh }`, + implement `{ work, medium }`, write `{ work, medium }`, verify `{ work, xhigh }`, + explore `{ work, medium }`, mechanical `{ cheap, low }`. (Keep the `model` + field name: a tier name is a model NAME the resolver understands; no new + `tier` field. AC3 greps `tier: '…'`?? NO: AC3 is written against + `model: 'best'`-style entries? → see AC3 note below.) +- prompt: `[{ pattern: '\\bultrathink\\b', phase: 'escalate', mode: 'floor' }, + { pattern: '\\b(plan|planifie|planning|brainstorm|architecture|con[cç]ois|design)\\b', phase: 'plan', mode: 'default' }, + { pattern: '\\b(pourquoi|why|explique|explain|analyse|analyze|comprendre|understand|review|audit)\\b', phase: 'reflect', mode: 'default' }]`. + `mode` absent → 'default'. +AC3 note for the executor: the contract's CHECK counts `tier: '(best|big|work|cheap)'` +in DEFAULT_CONFIG and refuses `model: '` entries there. So the PHASE type gets +an explicit `tier?: string` field: `type Route = { tier?: string; model?: string; +effort?: Level }`; a route resolves `tier` first, then `model`. Default phases +use `tier:`. `/route model=` keeps writing `model` (alias or id). The route +tool keeps `phase`/`effort`/`clear` only. +Validation: `tiers` values = non-empty arrays of `models` keys (bad entries +dropped, logged); `fallback` = array of `models` keys, deduplicated, non-empty +(else default); phase `tier` must be a `tiers` key; `cooldownMinutes` positive +integer; `mode` ∈ floor|default. + +## Resolution (ONE resolver, used by spawn, main plan and texts) +```ts +function availableIn(st, aliases: string[]): string | undefined + // first alias whose full id is not down (st.down.get(id) > now → down) +function resolveModel(st, name: string): string + // tier name → availableIn(tiers[name]) ?? first alias → id + // alias → id (explicit: never skipped when down); full id → itself +function resolveRoute(st, route: Route): string | undefined + // route.tier ? resolveModel(st, route.tier) : route.model ? resolveModel(st, route.model) : undefined +function rank(st, id: string): number + // index in cfg.fallback of the alias whose id prefixes `id` (strip "[1m]"); unknown → fallback.length +``` +`now` comes from `$.clock.now()` (read once per hook call that needs it). + +## Breaker (`st.down: Map` full id → until ms; `st.lastMainModel: string`) +- `turn.complete`: main (`e.agentId` undefined) with `reason` ∈ error|refusal → + `markDown(st, st.lastMainModel)`; agent with that reason → `markDown(loop.model)` + (Loop gains `model: string`, the model the engine reported at spawn, + `started.model`). `aborted`/`answer` → nothing. +- `classic.PostModelSwitch` with `e.source === 'auto'` → `markDown(e.from_model)`. +- `markDown` logs ALWAYS (not only verbose): `model-router: unavailable + until ; routing falls back`. `/route reload` and `session.end` clear + the map. `show()` gets a `down: until , …` or `down: none` line. + +## Main-loop model decision (`mainModel` rewritten) +``` +wanted = resolveRoute(st, (userMain ?? turnMain ?? turnFloor)?.route) // per-axis as today +cur = e.model +if cur is down and wanted is undefined → wanted = nextAvailable(st, cur) // fallback chain after cur's alias +if wanted undefined or sameRank(wanted, cur) → cur +if rank(wanted) < rank(cur) (better) → cfg.mainUpgrade ? wanted : cur +if rank(wanted) > rank(cur) (cheaper) → cfg.mainModelSwitch && windowOk ? wanted : cur +if cur is down and wanted defined → wanted (always: nothing to lose) +``` +`st.lastMainModel = plan.model` at every main step. `sameRank` compares +aliases (so `claude-fable-5-1` vs `claude-fable-5-1[1m]` never flips). +The log/status show `→ fallback` when the breaker chose the model. + +## Spawn (`spawnRoute` / `registerSpawn`) +`wanted = resolveRoute(st, route)`; explicit `e.model` still wins. Store +`loop.model = started.model`. (Explicit alias given by the caller while down: +left alone, explicit means explicit; note in the tool description.) + +## Derived phases (no user action) +D1. Dispatch push/pop: in the Agent `tool.call` hook, when `e.agentId` is + undefined (main) and `st.turnMain?.source !== 'model'` written AFTER the + dispatch… simpler rule: on a main Agent call, if `st.resumeMain` is + unset, `st.resumeMain = st.turnMain ?? NONE` and `st.turnMain = { phase: + 'orchestrate', route: phases.orchestrate, source: 'derived' }`. When an + agent's `turn.complete` leaves `st.loops` empty AND `st.turnMain?.source + === 'derived'` → `st.turnMain = st.resumeMain` (NONE → null), clear + `resumeMain`. A `route` call or skill load in between replaces turnMain + (source model/skill) so the pop is skipped and `resumeMain` cleared at + the next main `turn.complete` (endMainTurn clears both). Source type gains + `'derived'`. +D2. Prompt default rules (`mode: 'default'`): write `turnMain = { phase, + route, source: 'prompt' }` (NOT the floor) — overridable by routes and + skills; floor rules unchanged. Mid-turn prompt with a default rule → only + `pendingPrompt`-like handling for FLOOR rules stays; a default rule typed + mid-turn is ignored (the running turn has its own routes). +D3. Classifier: when `cfg.classifier` and no rule matched and the prompt is + composer-origin and idle (no `turnId`): `label = await + $.model.classify(e.text.slice(0, MAX_PROMPT_SCAN), [...phaseNames, + 'other'])` in try/catch; a phase label → default route (source 'prompt'); + anything else → nothing. Document the cost in the config comment. + +## Texts +`routedText`/`mainNote`/`show`: print the RESOLVED id and `(fallback)` when +the breaker skipped a better alias; `(tier best → claude-fable-5-1)`. +Spinner/status unchanged shape. + +## Tests (register.test.ts; names must contain the contract's words) +Reuse the boot helper; full typed inputs; bottom hooks (`agent.spawn`, +`turn.complete`, `classic.PostModelSwitch` — read its input type for the +required fields; `prompt.submit`). Engine effort `high` in steps. +- `tier: a plan route upgrades a haiku session to fable at xhigh` — route tool + `plan`, step with `model: 'claude-haiku-4-5-20251001'` → bottom sees + `claude-fable-5-1` and `xhigh`. +- `downgrade: mechanical on fable keeps the model while the switch is off`. +- `fallback: an error turn on fable moves the next main step to opus` — step on + fable (sets lastMainModel), `$.turn.complete({ reason: 'error', agentId + undefined, … })`, step on fable → bottom sees `claude-opus-5-5`, effort + unchanged; then `/route reload` → step on fable stays fable. +- `breaker: PostModelSwitch auto marks the old model down` — `$.classic.PostModelSwitch({ from_model: 'claude-fable-5-1', to_model: 'claude-opus-5-5', source: 'auto', … })` → `/route show` lists `claude-fable-5-1` under `down:`. +- `spawn: Explore goes to opus while sonnet is down` — mark sonnet down + through an agent error turn (spawn Explore via bottom hook returning a1, + `$.turn.complete({ agentId: 'a1', reason: 'error' })`), spawn again → bottom + `e.model === 'claude-opus-5-5'`. +- `derived: a dispatch pushes orchestrate and pops the previous plan route` + — route tool `plan`, `$.tool.call({ tool: 'Agent', … })` with a bottom hook, + `/route show` main line shows `derived orchestrate`; `$.turn.complete({ + agentId: 'a1', reason: 'answer' })` (loop registered via spawn) → main line + shows `model plan` again. +- `default rule: "planifie la migration" routes the turn to plan, a route call overrides`. +- Keep every B1 `floor` test and all earlier tests green (≥ 40 tests total). + +## Constraints +- ≤ 25 logic lines per function (extract helpers: `availableIn`, `rank`, + `markDown`, `nextAvailable`, `decideMain`, `pushOrchestrate`, `popOrchestrate`, + `applyDefaultRule`, `classifyPrompt`), 80 chars/line, no `any`, state in + the closure, fail-open `.catch` with `warnOnce` on every new hook + (`classic.PostModelSwitch`), the route tool schema unchanged. +- Do not touch: the hardening (caps, `safely`, attestation), the Skill + bridge, the floor slot semantics. +- Verify: validate, the contract's tsc CHECK, `claude plugin test .`, AC3/AC4 + greps, `gates.sh run` on the contract. + +## Disposition +- honors BDR-115 and its amendment (one resolver, calling-loop writes, + truthful texts, per-machine config); supersedes "a phase without model + keeps the loop's model" (every default phase now names a tier). +- honors BDR-076 (dispatched judgment on opus first: `big` = opus, fable, + sonnet) and BDR-066 (execution on sonnet: `work`). +- LRN-203: hooks still write full ids (the resolver's output). +- LRN-204: downgrade on main stays gated; upgrade accepted (quality over one + cold-cache step). +- Deferred: repo agents' frontmatter pins cannot fall back (the mod does not + see them in wave 1) → wave 2 moves them into the table with tiers. From 140c16a67f85071529da68b4dd65c15a120b0501 Mon Sep 17 00:00:00 2001 From: bchanot Date: Fri, 9 Oct 2026 12:55:14 +0200 Subject: [PATCH 18/22] chore(tasks): model-router W1-C plan r2 + contract amendments after the FATAL round --- .../2026-10-09-model-router-tiers-1237.md | 18 +- .../2026-10-09-model-router-tiers-1237.md | 164 ++++++++++++++++++ 2 files changed, 175 insertions(+), 7 deletions(-) diff --git a/.claude/tasks/contracts/2026-10-09-model-router-tiers-1237.md b/.claude/tasks/contracts/2026-10-09-model-router-tiers-1237.md index 57eb85d..62754b7 100644 --- a/.claude/tasks/contracts/2026-10-09-model-router-tiers-1237.md +++ b/.claude/tasks/contracts/2026-10-09-model-router-tiers-1237.md @@ -13,25 +13,29 @@ Q: main-loop model moves / A: UPGRADE (phase tier ranks above the current model) Q: no `/route` typed by the user / A: phases come from (1) prompt rules at turn start (keyword → phase as the turn's DEFAULT route, source 'prompt', overridable; `ultrathink` stays a FLOOR), (2) the model's own `route` calls and the skills table, (3) derived: an Agent dispatch from main pushes `orchestrate` and the previous route is restored when the last live agent ends, (4) optional classifier (`classifier: false` by default) that asks the engine's small model for a phase label when no rule matched. [user: "ne pas lancer les /route moi-même"] Q: default prompt rules / A: floor: `\bultrathink\b` → escalate. Defaults: `\b(plan|planifie|planning|brainstorm|architecture|con[cç]ois|design)\b` → plan; `\b(pourquoi|why|explique|explain|analyse|analyze|comprendre|understand|review|audit)\b` → reflect. No default rule lowers a turn (no `mechanical` rule): lowering is explicit (route tool, skills). [orchestrator — conservative defaults, user-editable in the override] Q: typed `/effort-` / A: unchanged (floor + default of the turn, B1). +Q (r2): breaker signal / A: `classic.StopFailure` error kinds `rate_limit | overloaded | billing_error | model_not_found` (+ the engine's own fallback detected at the step), NOT `turn.complete` reasons (a context-limit or network error is not unavailability); backoff 15 → 300 min; `/clear` keeps the marks (account-wide), `/route reload` and a user `/model` clear them. [orchestrator — challenge r1] +Q (r2): classifier / A: deferred to wave 2 (dead code by default, no test can reach it, no timeout on the call). [orchestrator — challenge r1] +Q (r2): upgrade cost / A: an upgrade is skipped above `upgradeMaxTokens` (default 200000 context tokens): one cold read of the whole context per switch into a model. [orchestrator — LRN-204] +Q (r2): superseded clauses / A: floor AC4 (`turnMain` sources gain 'derived' and 'prompt'), W1-A AC6 (the switch gates downgrades only), BDR-115 (6) (window guard on every switch). [orchestrator] ## ACCEPTANCE CRITERIA -1. Suite green: `claude plugin test` passes with at least 40 `test(` calls; `claude plugin validate` passes with no warning; no line over 80 chars; no `any` type. - CHECK: cd mods/model-router && out=$(claude plugin test . 2>&1); rc=$?; echo "$out" | tail -n 3; [ $rc -eq 0 ] && [ "$(grep -cE '^\s*test\(' hooks/register.test.ts)" -ge 40 ] && v=$(claude plugin validate . 2>&1) && echo "$v" | grep -q 'Validation passed' && ! echo "$v" | grep -qi 'warning' && ! grep -nE '.{81,}' hooks/register.ts hooks/register.test.ts && ! grep -nE ':\s*any\b||as any\b' hooks/register.ts && echo TIERS-SUITE-OK +1. Suite green: `claude plugin test` passes with at least 43 `test(` calls; `claude plugin validate` passes with no warning; no line over 80 chars; no `any` type. + CHECK: cd mods/model-router && out=$(claude plugin test . 2>&1); rc=$?; echo "$out" | tail -n 3; [ $rc -eq 0 ] && [ "$(grep -cE '^\s*test\(' hooks/register.test.ts)" -ge 43 ] && v=$(claude plugin validate . 2>&1) && echo "$v" | grep -q 'Validation passed' && ! echo "$v" | grep -qi 'warning' && ! grep -nE '.{81,}' hooks/register.ts hooks/register.test.ts && ! grep -nE ':\s*any\b||as any\b' hooks/register.ts && echo TIERS-SUITE-OK EXPECT: TIERS-SUITE-OK EVIDENCE: pending 2. Type-check clean against this build's declarations. CHECK: T=/Users/b.chanot/Documents/claude/mods/model-router/.claude-plugin/types; W=$(mktemp -d) && printf '{"compilerOptions":{"target":"es2023","lib":["es2023"],"types":[],"module":"esnext","moduleResolution":"bundler","strict":true,"noUncheckedIndexedAccess":true,"noEmit":true,"skipLibCheck":true,"jsx":"react","jsxFactory":"h","jsxFragmentFactory":"Fragment"},"include":["%s/claude-code/index.d.ts","%s/claude-code-tools/index.d.ts","%s/hooks"]}' "$T" "$T" "$PWD/mods/model-router" > "$W/tsconfig.json" && (cd "$W" && npx --yes -p typescript@5 tsc -p tsconfig.json) && echo TSC-OK EXPECT: TSC-OK EVIDENCE: pending -3. Tiers in the config: `tiers` (best/big/work/cheap), `fallback`, `cooldownMinutes`, `mainUpgrade`, `classifier` exist in DEFAULT_CONFIG; every default phase names a tier, none a bare model; `/route show` prints the resolved model of each phase and a `down:` line. - CHECK: cd mods/model-router/hooks && grep -q "tiers:" register.ts && grep -q "fallback:" register.ts && grep -q "cooldownMinutes" register.ts && grep -q "mainUpgrade" register.ts && grep -q "classifier" register.ts && [ "$(awk '/^const DEFAULT_CONFIG/,/^}/' register.ts | grep -cE "tier: '(best|big|work|cheap)'")" -ge 10 ] && ! awk '/^const DEFAULT_CONFIG/,/^}/' register.ts | grep -qE "^\s+[a-z]+: \{ model: '" && grep -q "down:" register.ts && echo TIERS-CONFIG-OK +3. Tiers in the config: `tiers` (best/big/work/cheap), `fallback`, `cooldownMinutes`, `mainUpgrade`, `upgradeMaxTokens` exist in DEFAULT_CONFIG; every default phase names a tier, none a bare model; `/route show` prints the resolved model of each phase and a `down:` line; no `classifier` (deferred to wave 2); no `Loop.model` / `spawnModel` / `loop.model` identifier (W1-A AC8). + CHECK: cd mods/model-router/hooks && grep -q "tiers:" register.ts && grep -q "fallback:" register.ts && grep -q "cooldownMinutes" register.ts && grep -q "mainUpgrade" register.ts && grep -q "upgradeMaxTokens" register.ts && ! grep -q "classifier" register.ts && [ "$(awk '/^const DEFAULT_CONFIG/,/^}/' register.ts | grep -cE "tier: '(best|big|work|cheap)'")" -ge 10 ] && ! awk '/^const DEFAULT_CONFIG/,/^}/' register.ts | grep -qE "model: '(haiku|sonnet|opus|fable)'" && ! grep -qE "loop\.model|explicitModel|spawnModel" register.ts && grep -q "down:" register.ts && echo TIERS-CONFIG-OK EXPECT: TIERS-CONFIG-OK EVIDENCE: pending -4. Tests prove (names contain the quoted word): `tier` — a `plan` route on a session model `claude-haiku-4-5-20251001` makes the main step run on `claude-fable-5-1` at xhigh (upgrade, default on); `downgrade` — a `mechanical` route on a fable session leaves the model unchanged while `mainModelSwitch` is off; `fallback` — after a main `turn.complete` with `reason: 'error'` on fable, the next main step runs on `claude-opus-5-5` with its effort unchanged, and after `/route reload` fable is used again; `breaker` — a `classic.PostModelSwitch` with `source: 'auto'` from fable marks fable down and `/route show` lists it; `spawn` — `Explore` spawns on `claude-opus-5-5` while sonnet is down; `derived` — an Agent tool call from main sets `orchestrate` and the previous route (`plan`) is back after the last agent's `turn.complete`; `default rule` — a prompt "planifie la migration" sets the `plan` route as the turn default and a later `route` tool call overrides it; `floor` tests from B1 still pass. - CHECK: cd mods/model-router/hooks && for w in tier downgrade fallback breaker spawn derived "default rule"; do grep -qE "test\('[^']*$w" register.test.ts || { echo "missing test: $w"; exit 1; }; done && echo TIERS-TESTS-OK +4. Tests prove (names contain the quoted word; plan r2 R14 lists them): `tier` — a `plan` route on a session model `claude-haiku-4-5-20251001` makes the main step run on `claude-fable-5-1` at xhigh (upgrade, default on); `downgrade` — a `mechanical` route on a fable session leaves the model unchanged while `mainModelSwitch` is off; `fallback` — with a `plan` route, a `classic.StopFailure` `rate_limit` on main after a fable step makes the next main step run on `claude-opus-5-5` at xhigh, and after `/route reload` fable is used again; `breaker` — an `invalid_request` failure never marks a model down, backoff expiry restores it, a `/model` command (`PostModelSwitch` source `command`) clears it; `engine fallback` — a step arriving on a model other than `$.session.model()` is never upgraded back and the session model is marked down; `unknown` — a session model absent from the table is never switched; `spawn` — `Explore` spawns on `claude-opus-5-5` while sonnet is down (agent StopFailure with `agent_id`); `derived` — a main Agent call sets `orchestrate` and the previous `plan` route is back when the spawned agent ends; a route declared after the dispatch is not overwritten by the pop; `default rule` — "planifie la migration" sets `plan` as the turn default and a later route call overrides it; a typed `/analyze …` and a `/effort-low pourquoi …` prompt get no default rule; `per axis` — a model-less sticky never hides a turn route's tier; `floor` tests from B1 still pass. + CHECK: cd mods/model-router/hooks && for w in tier downgrade fallback breaker "engine fallback" unknown spawn derived "default rule" "per axis"; do grep -qE "test\('[^']*$w" register.test.ts || { echo "missing test: $w"; exit 1; }; done && [ "$(grep -cE "test\('[^']*(breaker|derived|default rule)" register.test.ts)" -ge 7 ] && echo TIERS-TESTS-OK EXPECT: TIERS-TESTS-OK EVIDENCE: pending -5. Judged by reading: ONE resolver (`resolveModel`) turns a tier name, an alias or a full id into the first AVAILABLE full id (tier → list → skip down → alias → id; a bare alias or id passes through even when down, since explicit means explicit); rank = position in `fallback`; the main-loop decision is: route model wanted → upgrade allowed by `mainUpgrade`, downgrade gated by `mainModelSwitch` + window, same rank → keep; no route and current model down → next available in `fallback`; the breaker is fed only by `turn.complete` `reason` in `error|refusal` (main: the last main step's model; agent: the loop's spawn model) and by `classic.PostModelSwitch` `source: 'auto'` (from_model), never by an aborted turn; the derived `orchestrate` push/pop never overrides a route the model declared after the dispatch; prompt default rules write `turnMain` (source 'prompt'), floor rules write `turnFloor`; the classifier runs only when `classifier` is true and no rule matched, through `$.model.classify`, labels limited to the phase names plus `other`, any failure → no route; explicit Agent params still win; agent model fixed at spawn; every B1/1-A criterion still holds; no function over 25 logic lines; truthful texts name the resolved model and say "fallback" when the breaker chose it. +5. Judged by reading (plan r2 R1-R15 are binding): ONE resolver turns a tier name, an alias or a full id into an AVAILABLE canonical id (tier → list → skip down → id; an exhausted tier → the global chain; a bare alias or id passes through even when down); ids are canonical everywhere (`[1m]` stripped for comparison and carried on the replacement, alias → id, two-way prefix); ONE decision `decideMain(cur, wanted, ctx)` in the binding order (off → unknown cur: no switch → cur down: wanted or next available, windowOk → same alias: keep → better: `mainUpgrade` and `upgradeMaxTokens` → cheaper: `mainModelSwitch` and windowOk), used by `mainPlan` AND by every text; the engine's own fallback is respected (a step arriving off the session model marks the session model down and is never upgraded back); the breaker is fed only by `classic.StopFailure` errors `rate_limit | overloaded | billing_error | model_not_found` (main → the last main plan's model, agent → `agentModels`) and by the engine-fallback detection; `turn.complete` reasons and `PostModelSwitch` `auto` never mark (auto is logged); a user `/model` (`PostModelSwitch` command|picker|sdk) clears the target's mark; backoff 15 → 30 → 60 → 120 → 300 min per id, `model_not_found` until reload; the breaker survives `/clear` and is cleared by `/route reload` before the config read; the derived `orchestrate` push/pop tracks this turn's spawns and never overwrites a route the model declared after the dispatch; default prompt rules (two passes, absent `mode` = floor, `iu` flags, Unicode guards) write `turnMain` (source 'prompt') and are skipped for a leading `/`, for a prompt carrying a floor or a typed slash, and mid-turn; floor rules write `turnFloor`; no classifier; explicit Agent params still win; agent model fixed at spawn; every B1/1-A criterion still holds EXCEPT the three clauses R15 names (turnMain sources, the switch clause, the window-guard scope); no function over 25 logic lines; truthful texts come from `decideMain` and say `upgrade`, `fallback`, `switch off` or `unchanged`. ## FILE SCOPE mods/model-router/hooks/register.ts · mods/model-router/hooks/register.test.ts diff --git a/.claude/tasks/plans/2026-10-09-model-router-tiers-1237.md b/.claude/tasks/plans/2026-10-09-model-router-tiers-1237.md index 639d822..2e309c1 100644 --- a/.claude/tasks/plans/2026-10-09-model-router-tiers-1237.md +++ b/.claude/tasks/plans/2026-10-09-model-router-tiers-1237.md @@ -186,3 +186,167 @@ required fields; `prompt.submit`). Engine effort `high` in steps. cold-cache step). - Deferred: repo agents' frontmatter pins cannot fall back (the mod does not see them in wave 1) → wave 2 moves them into the table with tiers. + +## r2 — challenge round (3 lenses, all FATAL: 4 BLOCKER, 20 MAJOR): BINDING, overrides every section above where they conflict +R1. ONE phase field for the fallback-aware choice: `Route = { tier?: string; + model?: string; effort?: Level }`. Default phases use `tier:` only (plan, + reflect, orchestrate, escalate → best; judge → big; implement, write, + verify, explore → work; mechanical → cheap). `acceptPhase` refuses a route + carrying both `tier` and `model`, and refuses a `tiers` key that collides + with a `models` alias. The earlier "model: 'best'" drafts and the "no new + tier field" sentence are VOID. `/route model=` keeps writing + `model`; the route tool schema is unchanged. +R2. Ids: `canonical(st, id)` = strip a trailing `[1m]`, then alias → table id, + then two-way prefix match against the table ids (`id.startsWith(tableId) + || tableId.startsWith(id)`), else the id itself. `aliasOf(st, id)` and + `modelRank(st, id)` (= index of the alias in `fallback`, `undefined` when + unknown) work on canonical ids. The existing effort `rank` keeps its name. + Breaker keys, `agentModels` values and comparisons are canonical. When the + current main model carries `[1m]`, a resolved replacement carries `[1m]` + too (the long-context tier is a property of the session, not of the + alias); log the first time it happens (unverified live: see Verify). +R3. Model axis per slot: `routeModelName(route) = route.tier ?? route.model`; + the main model axis is the FIRST defined `routeModelName` across + userMain, turnMain, turnFloor (per-axis, like B1's effort). `resolveName` + turns that name into an available id: a `tiers` key → first alias of the + list not down → `models` id; a tier whose every alias is down → + `nextAvailable(st, cur)` (global chain) → may be `undefined` (keep cur); + an alias or full id → canonical id, never skipped (explicit means explicit). +R4. Main decision `decideMain(st, cur, wanted, ctx)` → `{ model, why }`, used + by `mainPlan` AND by every text (texts pass `cur = canonical(await + $.session.model())`); order is BINDING: + 1. `st.off` → cur. + 2. `rankCur = modelRank(cur)`; UNKNOWN cur (not in the table) → cur, log + once per session (`model-router: unknown to the models table; no + model switch`), the breaker still applies at step 3 if it is down. + 3. cur DOWN → `wanted` if defined and not down, else `nextAvailable(cur)`; + apply `windowOk`; if nothing fits → cur (why `fallback`). + 4. `wanted` undefined or `aliasOf(wanted) === aliasOf(cur)` → cur. + 5. `modelRank(wanted) < rankCur` (better) → `cfg.mainUpgrade && + ctx.tokens <= cfg.upgradeMaxTokens` ? wanted (why `upgrade`) : cur + (why `upgrade skipped: context tokens over ` or `switch off`). + 6. cheaper → `cfg.mainModelSwitch && windowOk` ? wanted (why `downgrade`) + : cur (why `switch off`). + New config scalar `upgradeMaxTokens` (default 200000): an upgrade pays a + cold read of the whole context on the new model (LRN-204); above the + threshold it is skipped and logged once per turn. `ctx.tokens` comes from + `$.session.usage()` read once per main step (fail → treat as 0). +R5. Engine fallback respected: at every main step `sess = canonical(await + $.session.model())`; when `canonical(e.model) !== sess`, the engine is on + a fallback → `markDown(sess, 'engine fallback')` and `cur = e.model` (the + router never upgrades back to the model the engine just left). +R6. Breaker inputs (replace the r1 list): (a) `classic.StopFailure` with + `error` ∈ rate_limit | overloaded | billing_error | model_not_found → + `markDown(target)` where target = `st.agentModels.get(e.agent_id)` when + `e.agent_id` is set, else `st.lastPlan?.model`; other errors (context + limit = invalid_request, server_error, auth, max_output_tokens…) → nothing; + (b) R5's engine-fallback detection; (c) `classic.PostModelSwitch`: source + `command | picker | sdk` → `st.down.delete(canonical(to_model))` and reset + its strikes (the user's explicit `/model` wins); source `auto` → LOG only + (`requested_model`, from, to), never a mark (unverified semantics). + `turn.complete` `reason` is NOT a breaker input any more (context-limit + and network errors are not availability); refusal → nothing. + Backoff per canonical id: strikes 1, 2, 3… → 15, 30, 60, 120, 300 min + (cap); `model_not_found` → until `/route reload`. `markDown` logs ALWAYS: + `model-router: unavailable () until ; routing falls + back`. Inert while `st.off`. + Lifecycle: `/route reload` clears `down` and strikes BEFORE loading the + config (whatever the read result); `session.end` (/clear) KEEPS `down`, + strikes and `agentModels` (availability is account-wide); expired + entries are pruned at the start of any hook that reads them, with `now` + read ONLY when `st.down.size > 0` (`$.clock.now()`), passed explicitly to + the helpers (no clock read in sync text functions: they receive the + pruned map). +R7. Agent models: `st.agentModels: Map` set at spawn + from `started.model` (canonicalized; an alias answered by a hook above is + mapped through the table); deleted with the loop. No `Loop.model`, + `spawnModel` or `loop.model` identifier anywhere (W1-A AC8 grep). + `spawnRoute` resolves `route.tier ?? route.model` through `resolveName` + (skips down aliases); explicit `e.model` still wins even when down. + Deferred (noted): agents without a table row and no explicit model follow + `parentModel`; forks always inherit; neither falls back in wave 1. +R8. Derived orchestrate (D1) made exact: state `pushed: { prev: Routed | null; + spawnIds: Set } | null`. In the main Agent `tool.call` hook: + before `next`, if `st.turnMain?.source` is not 'model' or 'skill' and + `st.pushed` is null → `st.pushed = { prev: st.turnMain, spawnIds: new Set() }` + and `st.turnMain = { phase: 'orchestrate', route: phases.orchestrate, + source: 'derived' }`; after `next` resolves: the spawned `agentId` (from + `st.spawnByCall: Map` filled at `agent.spawn`) is + added to `pushed.spawnIds`; if NO agent was registered for this + `tool_use_id` (foreground run already finished, or denied) → nothing to + wait for from this call. Pop rule: when `pushed.spawnIds` is empty after + the Agent call returned, or when the LAST id of `pushed.spawnIds` ends + (`turn.complete` with that agentId, deleted from the set), and + `st.turnMain?.source === 'derived'` → `st.turnMain = pushed.prev`, + `st.pushed = null`. A route/skill write in between (source model/skill) + replaces turnMain; the pop then only clears `pushed`. `endMainTurn` + clears `pushed` and `spawnByCall`. In `turn.complete` for an agent, delete + the loop and the maps FIRST, inside `safely`, before any other work. +R9. Prompt default rules (D2) made safe: rules scanned in two passes (floor + rules, then default rules), each pass first match; absent `mode` → + 'floor' (B1 override files keep their meaning). Default rules are SKIPPED + when the trimmed text starts with `/` (slash commands and skills route + themselves), when the same prompt carries a floor match or sets + `typedSlash` (the user's explicit level wins), or when typed mid-turn. + Patterns compile with flags `iu` and the defaults use Unicode-aware + guards instead of `\b`: `(? until + () …| none`, each phase as `name=→/`. + Existing test 3f (`/route model=sonnet` shows `claude-sonnet-5-5`) is + adapted: on the kit's session model the line reads `asked claude-sonnet-5-5, + keeps (switch off)`; the alias→id resolution is asserted on the + `asked` part. +R12. `st.lastPlan: Plan | null` replaces `lastMain` and `lastMainModel`; the + spinner text is derived at render; `endMainTurn` resets it. +R13. Config validation additions: `tiers` values non-empty arrays of alias + keys (bad entries dropped, logged), `fallback` deduplicated non-empty + alias list (else default, logged), `cooldownMinutes` and + `upgradeMaxTokens` positive integers, `mode` ∈ floor|default, a log at + load when `tiers.best[0] !== fallback[0]` (rank comes from `fallback` + alone). `mainModelSwitch` documented as DOWNGRADE-only in the Config + comment. +R14. Tests (≥ 43 total, names carry the contract words): keep all 30; add: + `tier` (plan on a haiku session → fable xhigh, with mock.clock installed + where the breaker is touched), `downgrade` (mechanical on fable keeps + fable, switch off), `fallback` (plan route + `$.classic.StopFailure({ + error: 'rate_limit', … })` on main after a fable step → next step + `claude-opus-5-5` at xhigh; `/route reload` → fable again), `breaker` + ×3 (an aborted/`invalid_request` failure never marks down; backoff expiry + via `mock.clock` advance restores fable; `/model` command + `PostModelSwitch source: 'command'` clears a down model), `engine fallback` + (`$.session.model` mocked/answered as fable while the step arrives on + opus → no upgrade back, fable marked down), `unknown` (cur + `claude-zz-9` never switches), `spawn` (Explore → opus while sonnet is + down through an agent StopFailure with `agent_id`), `derived` ×2 (push on + dispatch, pop when the spawned agent ends → plan back; a route call after + the dispatch is NOT overwritten by the pop), `default rule` ×3 (planifie + → plan then a route call overrides; `/analyze …` typed → no rule; + `/effort-low pourquoi …` → no default rule, floor low), `per axis` + (`/route effort=low` sticky + turn `plan` tier → model axis = best). + Read `mock.clock` and how `$.session.model` is answered in the kit + (a bottom `on('session.model', …)` hook) before writing them. +R15. Disposition, superseded clauses named: floor contract AC4 "`turnMain` + only ever holds 'model' or 'skill' sources" → now also 'derived' and + 'prompt'; W1-A AC6 "main-loop model changes happen only when + `mainModelSwitch` is true" → true for DOWNGRADES only; upgrades follow + `mainUpgrade` + `upgradeMaxTokens`, and the breaker/engine-fallback path + moves off a dead model unconditionally; BDR-115 (6) window guard → applied + to every switch (up, down, fallback) through `windowOk`. The tiers + contract AC5 reads "every B1/1-A criterion still holds EXCEPT the three + clauses above". +R16. Live verification after reload (orchestrator, not the executor): the + `[1m]` carry-over on a fallback id, `PostModelSwitch` `source: 'auto'` + semantics, `StopFailure` reaching the mod with `agent_id`. From 977be7cad899b8b4bc9bf2e11e672717e6d96e66 Mon Sep 17 00:00:00 2001 From: bchanot Date: Fri, 9 Oct 2026 13:22:07 +0200 Subject: [PATCH 19/22] chore(tasks): model-router W1-C plan r3 + r4 after two confirmation passes --- .../2026-10-09-model-router-tiers-1237.md | 6 +- .../2026-10-09-model-router-tiers-1237.md | 126 ++++++++++++++++++ 2 files changed, 129 insertions(+), 3 deletions(-) diff --git a/.claude/tasks/contracts/2026-10-09-model-router-tiers-1237.md b/.claude/tasks/contracts/2026-10-09-model-router-tiers-1237.md index 62754b7..8c55109 100644 --- a/.claude/tasks/contracts/2026-10-09-model-router-tiers-1237.md +++ b/.claude/tasks/contracts/2026-10-09-model-router-tiers-1237.md @@ -28,14 +28,14 @@ Q (r2): superseded clauses / A: floor AC4 (`turnMain` sources gain 'derived' and EXPECT: TSC-OK EVIDENCE: pending 3. Tiers in the config: `tiers` (best/big/work/cheap), `fallback`, `cooldownMinutes`, `mainUpgrade`, `upgradeMaxTokens` exist in DEFAULT_CONFIG; every default phase names a tier, none a bare model; `/route show` prints the resolved model of each phase and a `down:` line; no `classifier` (deferred to wave 2); no `Loop.model` / `spawnModel` / `loop.model` identifier (W1-A AC8). - CHECK: cd mods/model-router/hooks && grep -q "tiers:" register.ts && grep -q "fallback:" register.ts && grep -q "cooldownMinutes" register.ts && grep -q "mainUpgrade" register.ts && grep -q "upgradeMaxTokens" register.ts && ! grep -q "classifier" register.ts && [ "$(awk '/^const DEFAULT_CONFIG/,/^}/' register.ts | grep -cE "tier: '(best|big|work|cheap)'")" -ge 10 ] && ! awk '/^const DEFAULT_CONFIG/,/^}/' register.ts | grep -qE "model: '(haiku|sonnet|opus|fable)'" && ! grep -qE "loop\.model|explicitModel|spawnModel" register.ts && grep -q "down:" register.ts && echo TIERS-CONFIG-OK + CHECK: cd mods/model-router/hooks && D=$(awk '/^const DEFAULT_CONFIG/,/^}/' register.ts) && echo "$D" | grep -q "tiers:" && echo "$D" | grep -q "fallback:" && grep -q "cooldownMinutes" register.ts && grep -q "mainUpgrade" register.ts && grep -q "upgradeMaxTokens" register.ts && ! grep -q "classifier" register.ts && [ "$(awk '/^const DEFAULT_CONFIG/,/^}/' register.ts | grep -cE "tier: '(best|big|work|cheap)'")" -ge 10 ] && ! awk '/^const DEFAULT_CONFIG/,/^}/' register.ts | grep -qE "model: '(haiku|sonnet|opus|fable)'" && ! grep -qE "loop\.model|explicitModel|spawnModel" register.ts && grep -q "down:" register.ts && echo TIERS-CONFIG-OK EXPECT: TIERS-CONFIG-OK EVIDENCE: pending -4. Tests prove (names contain the quoted word; plan r2 R14 lists them): `tier` — a `plan` route on a session model `claude-haiku-4-5-20251001` makes the main step run on `claude-fable-5-1` at xhigh (upgrade, default on); `downgrade` — a `mechanical` route on a fable session leaves the model unchanged while `mainModelSwitch` is off; `fallback` — with a `plan` route, a `classic.StopFailure` `rate_limit` on main after a fable step makes the next main step run on `claude-opus-5-5` at xhigh, and after `/route reload` fable is used again; `breaker` — an `invalid_request` failure never marks a model down, backoff expiry restores it, a `/model` command (`PostModelSwitch` source `command`) clears it; `engine fallback` — a step arriving on a model other than `$.session.model()` is never upgraded back and the session model is marked down; `unknown` — a session model absent from the table is never switched; `spawn` — `Explore` spawns on `claude-opus-5-5` while sonnet is down (agent StopFailure with `agent_id`); `derived` — a main Agent call sets `orchestrate` and the previous `plan` route is back when the spawned agent ends; a route declared after the dispatch is not overwritten by the pop; `default rule` — "planifie la migration" sets `plan` as the turn default and a later route call overrides it; a typed `/analyze …` and a `/effort-low pourquoi …` prompt get no default rule; `per axis` — a model-less sticky never hides a turn route's tier; `floor` tests from B1 still pass. +4. Tests prove (names contain the quoted word; plan r2 R14 lists them): `tier` — a `plan` route on a session model `claude-haiku-4-5-20251001` makes the main step run on `claude-fable-5-1` at xhigh (upgrade, default on); `downgrade` — a `mechanical` route on a fable session leaves the model unchanged while `mainModelSwitch` is off; `fallback` — with a `plan` route, a `classic.StopFailure` `rate_limit` on main after a fable step makes the next main step run on `claude-opus-5-5` at xhigh, and after `/route reload` fable is used again; `breaker` — an `invalid_request` failure never marks a model down, backoff expiry restores it, a `/model` command (`PostModelSwitch` source `command`) clears it; `engine fallback` — a `PostModelSwitch` with source `auto` marks the model the engine left (one strike, idempotent within the hold) and a plan route does not go back to it; `unknown` — a session model absent from the table is never switched; `spawn` — `Explore` spawns on `claude-opus-5-5` while sonnet is down (agent StopFailure with `agent_id`); `derived` — a main Agent call sets `orchestrate` and the previous `plan` route is back when the spawned agent ends; a route declared after the dispatch is not overwritten by the pop; `default rule` — "planifie la migration" sets `plan` as the turn default and a later route call overrides it; a typed `/analyze …` and a `/effort-low pourquoi …` prompt get no default rule; `per axis` — a model-less sticky never hides a turn route's tier; `floor` tests from B1 still pass. CHECK: cd mods/model-router/hooks && for w in tier downgrade fallback breaker "engine fallback" unknown spawn derived "default rule" "per axis"; do grep -qE "test\('[^']*$w" register.test.ts || { echo "missing test: $w"; exit 1; }; done && [ "$(grep -cE "test\('[^']*(breaker|derived|default rule)" register.test.ts)" -ge 7 ] && echo TIERS-TESTS-OK EXPECT: TIERS-TESTS-OK EVIDENCE: pending -5. Judged by reading (plan r2 R1-R15 are binding): ONE resolver turns a tier name, an alias or a full id into an AVAILABLE canonical id (tier → list → skip down → id; an exhausted tier → the global chain; a bare alias or id passes through even when down); ids are canonical everywhere (`[1m]` stripped for comparison and carried on the replacement, alias → id, two-way prefix); ONE decision `decideMain(cur, wanted, ctx)` in the binding order (off → unknown cur: no switch → cur down: wanted or next available, windowOk → same alias: keep → better: `mainUpgrade` and `upgradeMaxTokens` → cheaper: `mainModelSwitch` and windowOk), used by `mainPlan` AND by every text; the engine's own fallback is respected (a step arriving off the session model marks the session model down and is never upgraded back); the breaker is fed only by `classic.StopFailure` errors `rate_limit | overloaded | billing_error | model_not_found` (main → the last main plan's model, agent → `agentModels`) and by the engine-fallback detection; `turn.complete` reasons and `PostModelSwitch` `auto` never mark (auto is logged); a user `/model` (`PostModelSwitch` command|picker|sdk) clears the target's mark; backoff 15 → 30 → 60 → 120 → 300 min per id, `model_not_found` until reload; the breaker survives `/clear` and is cleared by `/route reload` before the config read; the derived `orchestrate` push/pop tracks this turn's spawns and never overwrites a route the model declared after the dispatch; default prompt rules (two passes, absent `mode` = floor, `iu` flags, Unicode guards) write `turnMain` (source 'prompt') and are skipped for a leading `/`, for a prompt carrying a floor or a typed slash, and mid-turn; floor rules write `turnFloor`; no classifier; explicit Agent params still win; agent model fixed at spawn; every B1/1-A criterion still holds EXCEPT the three clauses R15 names (turnMain sources, the switch clause, the window-guard scope); no function over 25 logic lines; truthful texts come from `decideMain` and say `upgrade`, `fallback`, `switch off` or `unchanged`. +5. Judged by reading (plan r2 R1-R15, r3 S1-S11 and r4 T1-T8 are binding; r4 wins over r3, r3 over r2 where they conflict: two fields `turnModel` (sticky, per turn) and `lastPlan` (breaker target, kept), unrouted steps pass `e.model` verbatim, auto marks target the model actually sent and skip when the engine landed where the router was, `sessionModel` preserved across /clear and never prefix-matched when empty, tokens `number | undefined` with windowOk failing closed, episode strikes with `model_not_found` lengthening a hold; no per-step engine-fallback detection, `PostModelSwitch` auto DOES mark one idempotent strike, the main model is sticky within a turn, `sessionModel` cached at start and on PostModelSwitch, D1 for background dispatches via the Agent result status, breaker targets kept until replaced): ONE resolver turns a tier name, an alias or a full id into an AVAILABLE canonical id (tier → list → skip down → id; an exhausted tier → the global chain; a bare alias or id passes through even when down); ids are canonical everywhere (`[1m]` stripped for comparison and carried on the replacement, alias → id, two-way prefix); ONE decision `decideMain(cur, wanted, ctx)` in the binding order (off → unknown cur: no switch → cur down: wanted or next available, windowOk → same alias: keep → better: `mainUpgrade` and `upgradeMaxTokens` → cheaper: `mainModelSwitch` and windowOk), used by `mainPlan` AND by every text; the breaker is fed only by `classic.StopFailure` errors `rate_limit | overloaded | billing_error | model_not_found` (main → the last main plan's model, agent → `agentModels`) and by `PostModelSwitch` `source: 'auto'` (the model the engine left, one strike, idempotent while down); `turn.complete` reasons never mark; a user `/model` (`PostModelSwitch` command|picker|sdk) clears the target's mark; backoff 15 → 30 → 60 → 120 → 300 min per id, `model_not_found` until reload; the breaker survives `/clear` and is cleared by `/route reload` before the config read; the derived `orchestrate` push/pop tracks this turn's spawns and never overwrites a route the model declared after the dispatch; default prompt rules (two passes, absent `mode` = floor, `iu` flags, Unicode guards) write `turnMain` (source 'prompt') and are skipped for a leading `/`, for a prompt carrying a floor or a typed slash, and mid-turn; floor rules write `turnFloor`; no classifier; explicit Agent params still win; agent model fixed at spawn; every B1/1-A criterion still holds EXCEPT the three clauses R15 names (turnMain sources, the switch clause, the window-guard scope); no function over 25 logic lines; truthful texts come from `decideMain` and say `upgrade`, `fallback`, `switch off` or `unchanged`. ## FILE SCOPE mods/model-router/hooks/register.ts · mods/model-router/hooks/register.test.ts diff --git a/.claude/tasks/plans/2026-10-09-model-router-tiers-1237.md b/.claude/tasks/plans/2026-10-09-model-router-tiers-1237.md index 2e309c1..ac52a11 100644 --- a/.claude/tasks/plans/2026-10-09-model-router-tiers-1237.md +++ b/.claude/tasks/plans/2026-10-09-model-router-tiers-1237.md @@ -350,3 +350,129 @@ R15. Disposition, superseded clauses named: floor contract AC4 "`turnMain` R16. Live verification after reload (orchestrator, not the executor): the `[1m]` carry-over on a fallback id, `PostModelSwitch` `source: 'auto'` semantics, `StopFailure` reaching the mod with `agent_id`. + +## r3 — confirmation pass (FATAL(8): 1 BLOCKER, 6 MAJOR): BINDING over r2 where they conflict +S1. R5 (engine-fallback detection at every step) is REMOVED: no comparison of + `e.model` with `$.session.model()` at steps, no mark from it. The + engine's own fallback is learned ONLY through `classic.PostModelSwitch` + `source: 'auto'`, which now MARKS `canonical(from_model)` down with one + strike (15 min) when `from_model` is a table id, logging + `requested_model`, `to_model`. (R6(c) "auto → log only" is void.) A mark + is idempotent per episode: `markDown` on an id already down adds NO + strike and logs nothing; strikes count episodes (a mark after expiry). +S2. `st.sessionModel` (raw string) is read once at `session.start` through + `$.session.model()` inside try/catch ('' on failure) and refreshed in the + `PostModelSwitch` hook from `e.to_model` (any source). No other + `$.session.model()` call anywhere; texts use `st.sessionModel`. +S3. Within a turn the main model is STICKY once moved: `cur` for the decision + is `st.lastPlan?.model ?? e.model` (the model actually sent last; lastPlan + is reset at `endMainTurn` so each turn starts from the engine's model). + After an upgrade (plan → fable), a later cheaper phase in the same turn + (implement → work) goes through the CHEAPER branch against cur = fable: + gated by `mainModelSwitch` + windowOk, so no return trip and no second + cold read. After a fallback (fable down → opus), later steps stay on opus + for the turn. "Keep cur" returns the exact string last sent (`e.model` + verbatim on the first step), so `[1m]` is preserved; a resolved + replacement carries `[1m]` only when the raw session string carries it + AND the target alias is not haiku. +S4. decideMain spelled out (order binding): off → cur · unknown cur (no table + alias) → cur, logged once · cur down → first available of [wanted (if + a table id and not down), nextAvailable(cur)] that passes windowOk, else + cur · wanted undefined → cur · wanted unknown to the table (explicit full + id such as `claude-x-9`) → treated as CHEAPER (gated by `mainModelSwitch`, + windowOk) · same alias → cur · better → `mainUpgrade && tokens ≤ + upgradeMaxTokens && windowOk` ? wanted : cur · cheaper → `mainModelSwitch + && windowOk` ? wanted : cur. `ctx.tokens` from `$.session.usage()` read + once per main step (catch → 0); texts read it the same way (async), so a + text and the step agree. `nextAvailable(cur)` walks `fallback` from the + alias after cur's (unknown cur → from the top) skipping down ids; at + spawn, `nextAvailable` walks from the tier's last alias. + `model_not_found` marks show `until reload` in texts. +S5. Derived orchestrate (R8 rewritten): D1 affects BACKGROUND dispatches only. + In the main Agent `tool.call` hook: push as in R8 (source not model/skill, + `pushed` null) BEFORE `next`; after `next`: read the RESULT — `status === + 'async_launched'` → add `result.agentId` to `pushed.spawnIds`; any other + status or a deny → nothing to wait for from this call. Pop rule unchanged + (spawnIds empty after the call, or the last id's `turn.complete`); no + `spawnByCall` map. Documented: a foreground dispatch pushes and pops + inside one call, so no main step runs at orchestrate for it (fine: main + is blocked meanwhile). +S6. Breaker targets keep their value until replaced: `st.lastPlan` is NOT + reset at `endMainTurn` (only the spinner text is cleared via a separate + `st.spinner` string); `agentModels` entries are deleted at `session.end` + only, never at an agent's `turn.complete` (the StopFailure/turn.complete + order is unverified; R16 gains it). +S7. R9 patterns: compile with `iu`; on a SyntaxError retry with `i` (B1 + override files keep working); a pattern failing both is dropped, logged. +S8. R11: the route tool description is STATIC text (registered once): "the + main loop moves up to a phase's tier by itself (below the context cap), + down only with the switch on; a sub-agent's model is fixed at spawn". + `show`, `routedText`, `mainNote`, `statusLine` call `decideMain` with + `cur = st.lastPlan?.model ?? st.sessionModel` and the same tokens read. +S9. Tests, kit recipe (replaces R14 details): `boot(model = 'claude-fable-5-1')` + registers, before the first `$` call, bottom hooks `on('session.model', + () => ({ value: model }))` (answer shape per the Op results in the + declarations), `on('classic.StopFailure', ($, e) => )`, + `on('classic.PostModelSwitch', …)`, and installs `mock.clock(on)`; every + breaker test advances the mock clock. `derived` recipe: prompt + "planifie …" (source 'prompt'), then `$.tool.call({ tool: 'Agent', … })` + whose bottom hook returns `{ result: { status: 'async_launched', + agentId: 'a1', … } }` (read the Agent RESULT type for the required + fields) → `/route show` main line says `derived orchestrate`; then + `$.turn.complete({ agentId: 'a1', … })` → main line says `prompt plan`. + Second derived test: same, but a route tool call `reflect` after the + dispatch → the pop does not overwrite `model reflect`. `engine fallback` + test: `$.classic.PostModelSwitch({ from_model: 'claude-fable-5-1', + to_model: 'claude-opus-5-5', source: 'auto', … })` → fable listed under + `down:`; a plan route step does not go back to fable; a second auto + switch inside the hold adds no strike (show prints the same until). + Strikes test: expire (advance clock) → mark again → until doubles. +S10. AC3 fix: the `fallback:` and `tiers:` greps run inside the + DEFAULT_CONFIG awk range. +S11. R16 gains: the order of `classic.StopFailure` vs `turn.complete`; whether + the engine's fallback on a hook-rewritten request raises `PostModelSwitch`. + +## r4 — second confirmation (FATAL(4): 1 BLOCKER, 3 MAJOR): BINDING over r3 where they conflict; the last revision, executor dispatched on it +T1. Two fields, no contradiction: `st.turnModel: string | undefined` is the + STICKY cur, set ONLY when `decideMain` moved the model (upgrade, + downgrade or fallback), reset in `endMainTurn` and at `session.end`; + `st.lastPlan` (the plan actually sent last, breaker target) is KEPT across + turns and never used as cur. `cur = st.turnModel ?? e.model`. "Keep cur" + returns `st.turnModel` when set, else `e.model` VERBATIM: an unrouted step + never re-sends a model the router did not choose this turn, so an + engine fallback that lands in `e.model` is respected by construction. + S3's "lastPlan is reset at endMainTurn" is void (S6 stands). +T2. Auto switch marking (S1 refined): on `PostModelSwitch` `source: 'auto'`, + let `sent = canonical(st.lastPlan?.model)` and `to = canonical(to_model)`. + If `to === sent` → nothing (the engine landed where the router already + was, or the router's own rewrite surfaced as a switch). Else the mark + target is `sent` when it is a table id (the model actually sent), else + `canonical(from_model)` when THAT is a table id, else nothing. Always + log `from_model`, `to_model`, `requested_model`. R16 gains: which + `from_model` the event carries after a router upgrade, and whether a + router rewrite itself raises an `auto` switch. +T3. `st.sessionModel` is PRESERVED through the `session.end` rebuild (listed + with `down`, strikes, `agentModels`). `canonical()` never prefix-matches + an empty string or a string that does not start with `claude-`: both + map to UNKNOWN (returned unchanged, no table id). The `[1m]` carry reads + the raw string of the step (`e.model`, or `st.turnModel`), never + `sessionModel`. `sessionModel` is used by texts only; when it is '' or + unknown, texts print the engine word `session model` instead of an id. +T4. Tokens: `ctx.tokens: number | undefined` (undefined on a failed or absent + read). The upgrade cap treats undefined as 0 (upgrade allowed: the targets + are fable/opus, no window entry); `windowOk` treats undefined as NOT + fitting (fail closed, as today). +T5. Marks: strikes are per EPISODE (a mark on an id already down adds no + strike and no log), but a `model_not_found` arriving during a timed hold + LENGTHENS it to "until reload" (logged once). Auto marks use the same + episode backoff (15 → 30 → 60 → 120 → 300 min). +T6. S5 race: on an `async_launched` result with `st.pushed === null`, push + again first (if `turnMain?.source` still allows it), then add the id. +T7. Tests assert hold DURATIONS (minutes until, computed from the mock clock) + or the presence of the id under `down:`, never a literal `HH:MM`. + `show` prints `down: for min ()` (and `until reload`), + computed from the pruned map and the clock value passed in. +T8. R16 final list (live, orchestrator): StopFailure vs turn.complete order; + PostModelSwitch on a rewritten-request fallback and its `from_model`; + whether a router rewrite raises `auto`; `[1m]` carry validity on opus; + `$.session.model()` string form. From d0fa1001bb91171dadebbf33413623d62a7cdecd Mon Sep 17 00:00:00 2001 From: bchanot Date: Fri, 9 Oct 2026 15:08:49 +0200 Subject: [PATCH 20/22] =?UTF-8?q?feat(mods):=20model-router=20adaptive=20t?= =?UTF-8?q?iers=20=E2=80=94=20absolute=20tiers,=20availability=20breaker,?= =?UTF-8?q?=20derived=20phases?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Phases name absolute tiers (best fable>opus>sonnet, big opus>fable>sonnet, work sonnet>opus, cheap haiku>sonnet) resolved to the first available full id; per-model circuit breaker fed by StopFailure kinds (rate_limit, overloaded, billing_error, model_not_found) and PostModelSwitch auto, with episode backoff 15→300 min, cleared by a user /model or /route reload and kept across /clear; fallback chain fable→opus→sonnet→haiku with the effort unchanged; main loop upgrades to a phase's tier by itself under a context cap (fails closed on unknown usage), downgrades only with the switch on, sticky within a turn; derived orchestrate on background dispatches; prompt default rules (plan/reflect, Unicode guards, skipped on slash commands, floor matches and mid-turn). 58 plugin tests. --- mods/model-router/hooks/register.test.ts | 408 +++++++++- mods/model-router/hooks/register.ts | 995 +++++++++++++++++++---- 2 files changed, 1248 insertions(+), 155 deletions(-) diff --git a/mods/model-router/hooks/register.test.ts b/mods/model-router/hooks/register.test.ts index b1074d5..d85b9a2 100644 --- a/mods/model-router/hooks/register.test.ts +++ b/mods/model-router/hooks/register.test.ts @@ -1,11 +1,32 @@ -import { test, expect } from 'claude-code/testing' -import type { Engine } from 'claude-code/testing' +import { test, expect, mock } from 'claude-code/testing' +import type { Engine, MockClock } from 'claude-code/testing' import type { On, TurnStepInput } from 'claude-code' -/** Fires session.start so the mod loads its config and registers /route. */ -async function boot($: Engine, on: On): Promise { +const FABLE = 'claude-fable-5-1' +const OPUS = 'claude-opus-5-5' +const SONNET = 'claude-sonnet-5-5' +const HAIKU = 'claude-haiku-4-5-20251001' + +/** + * Fires session.start so the mod loads its config and registers /route. + * The bottom hooks the router reads are in place before the first `$` call: + * the session model, the classic failure and switch events, a mock clock, + * a small context (`tokens`; null = usage unanswered, size unknown). + */ +async function boot( + $: Engine, + on: On, + model: string = FABLE, + tokens: number | null = 1000, +): Promise { + const clock = mock.clock(on) + on('session.model', () => ({ value: model })) + on('classic.StopFailure', () => ({})) + on('classic.PostModelSwitch', () => ({})) + if (tokens !== null) usageOf(on, tokens) on('session.start', ($, e) => ({ cwd: e.cwd })) await $.session.start({ cwd: '/tmp', surface: null, isInteractive: false }) + return clock } /** Runs `/route ` as the user would type it. */ @@ -99,7 +120,7 @@ test('/route model: alias resolved, id passed, typo refused', async ( $, on) => { await boot($, on) const alias = await route($, 'model=sonnet') - expect(mainLine(alias)).toContain('claude-sonnet-5-5') + expect(mainLine(alias)).toContain('asked claude-sonnet-5-5, keeps ') const id = await route($, 'model=claude-x-9') expect(mainLine(id)).toContain('claude-x-9') expect(await route($, 'model=sonet')).toContain('unknown') @@ -446,3 +467,380 @@ test('typed slash with no agent and no marker writes the floor', async ( await $.skill.prompt({ skill: 'effort-max', text: 'x' }) expect(mainLine(await route($, 'show'))).toContain('floor max') }) + +// ---- tiers, breaker, derived phases ------------------------------------- + +type Rig = { seen: Seen[]; clock: MockClock; specs: (string | undefined)[] } + +type Failure = 'rate_limit' | 'overloaded' | 'invalid_request' + +/** Bottom hooks for a routed run, then the boot; `specs` records spawns. */ +async function bootRig( + $: Engine, + on: On, + model: string = FABLE, + tokens: number | null = 1000, +): Promise { + const seen: Seen[] = [] + const specs: (string | undefined)[] = [] + recordSteps(on, seen) + on('prompt.submit', ($, e) => ({ text: e.text })) + on('turn.complete', () => ({ text: '' })) + on('agent.spawn', ($, e) => { + specs.push(e.model) + return { model: e.model ?? e.parentModel, agentId: 'a1' } + }) + on('tool.call', { tool: 'Agent' }, () => ({ + result: { + status: 'async_launched' as const, + agentId: 'a1', + description: 'd', + prompt: 'p', + outputFile: '/tmp/o', + }, + })) + const clock = await boot($, on, model, tokens) + return { seen, clock, specs } +} + +/** One main step on `model`; returns what reached the bottom hook. */ +async function stepOn($: Engine, rig: Rig, model: string): Promise { + await runStep($, { ...highStep(), model }) + return rig.seen[rig.seen.length - 1] ?? { model: '', effort: undefined } +} + +const failWith = ($: Engine, error: Failure, agentId?: string) => + $.classic.StopFailure({ + error, + ...(agentId === undefined ? {} : { agent_id: agentId }), + }) + +const switchTo = ( + $: Engine, + from: string, + to: string, + source: 'command' | 'auto', +) => $.classic.PostModelSwitch({ + from_model: from, + to_model: to, + requested_model: null, + source, + context_tokens: 0, + prompt_cache_warm: false, + cache_ttl: '5m', + estimated_cache_write_usd: 0, + pricing: 'catalog', +}) + +const planRoute = ($: Engine) => + $.tool.call({ tool: ROUTE_TOOL, phase: 'plan' }) + +/** A bottom `session.usage` answering a context of `tokens`. */ +function usageOf(on: On, tokens: number): void { + on('session.usage', () => ({ + value: { + startedAt: 0, + context: { window: 1000000, tokens }, + rateLimits: [], + }, + })) +} + +test('tier: a plan route upgrades a haiku session to fable xhigh', async ( + $, on) => { + const rig = await bootRig($, on, HAIKU) + await planRoute($) + expect(await stepOn($, rig, HAIKU)).toEqual({ + model: FABLE, + effort: 'xhigh', + }) +}) + +test('tier: show prints the resolved id of each phase and a down line', async ( + $, on) => { + await bootRig($, on) + const text = await route($, 'show') + expect(text).toContain('plan=best→claude-fable-5-1/xhigh') + expect(text).toContain('judge=big→claude-opus-5-5/xhigh') + expect(text).toContain('down: none') +}) + +test('tier: an upgrade is skipped above the context cap', async ($, on) => { + const rig = await bootRig($, on, SONNET, 300000) + await planRoute($) + expect(await stepOn($, rig, SONNET)).toEqual({ + model: SONNET, + effort: 'xhigh', + }) +}) + +test('cap: unknown usage blocks the upgrade', async ($, on) => { + const rig = await bootRig($, on, HAIKU, null) + await planRoute($) + expect((await stepOn($, rig, HAIKU)).model).toBe(HAIKU) +}) + +test('cap: a down model leaves for a better wanted only under the cap', async ( + $, on) => { + const rig = await bootRig($, on, OPUS, 300000) + await planRoute($) + await stepOn($, rig, OPUS) + await failWith($, 'rate_limit') + expect((await stepOn($, rig, OPUS)).model).toBe(SONNET) +}) + +test('breaker: model_not_found on a non-table id marks nothing', async ( + $, on) => { + const rig = await bootRig($, on) + await planRoute($) + await stepOn($, rig, 'claude-zz-9') + await $.classic.StopFailure({ error: 'model_not_found' }) + expect(await route($, 'show')).toContain('down: none') +}) + +test('downgrade: mechanical on fable keeps the model, switch off', async ( + $, on) => { + const rig = await bootRig($, on) + await $.tool.call({ tool: ROUTE_TOOL, phase: 'mechanical' }) + expect(await stepOn($, rig, FABLE)).toEqual({ model: FABLE, effort: 'low' }) +}) + +test('downgrade: the switch on moves main to the cheap tier', async ( + $, on) => { + const rig = await bootRig($, on) + await route($, 'switch on') + await $.tool.call({ tool: ROUTE_TOOL, phase: 'mechanical' }) + const sent = await stepOn($, rig, FABLE) + expect(sent.model).toBe(HAIKU) + expect(sent.effort).toBeUndefined() +}) + +test('fallback: a rate limit on fable moves plan to opus xhigh', async ( + $, on) => { + const rig = await bootRig($, on) + await planRoute($) + expect((await stepOn($, rig, FABLE)).model).toBe(FABLE) + await failWith($, 'rate_limit') + expect(await stepOn($, rig, FABLE)).toEqual({ model: OPUS, effort: 'xhigh' }) + expect(await route($, 'show')).toContain('claude-fable-5-1 for 15 min') + await route($, 'reload') + expect((await stepOn($, rig, FABLE)).model).toBe(FABLE) +}) + +test('breaker: an invalid_request or a turn error never marks down', async ( + $, on) => { + const rig = await bootRig($, on) + await planRoute($) + await stepOn($, rig, FABLE) + await failWith($, 'invalid_request') + await $.turn.complete({ + turnId: 'u1', + answer: '', + durationMs: 1, + isAborted: false, + reason: 'error', + }) + expect((await stepOn($, rig, FABLE)).model).toBe(FABLE) + expect(await route($, 'show')).toContain('down: none') +}) + +test('breaker: the hold lapses with the clock, a second one doubles', async ( + $, on) => { + const rig = await bootRig($, on) + await planRoute($) + await stepOn($, rig, FABLE) + await failWith($, 'overloaded') + await rig.clock.advance(15 * 60000) + expect(await route($, 'show')).toContain('down: none') + await endTurn($) + await planRoute($) + expect((await stepOn($, rig, FABLE)).model).toBe(FABLE) + await failWith($, 'overloaded') + expect(await route($, 'show')).toContain('claude-fable-5-1 for 30 min') +}) + +test('breaker: a /model command clears the mark of its target', async ( + $, on) => { + const rig = await bootRig($, on) + await planRoute($) + await stepOn($, rig, FABLE) + await failWith($, 'rate_limit') + expect(await route($, 'show')).toContain('claude-fable-5-1 for') + await switchTo($, OPUS, FABLE, 'command') + expect(await route($, 'show')).toContain('down: none') + expect((await stepOn($, rig, FABLE)).model).toBe(FABLE) +}) + +test('breaker: model_not_found holds until the reload', async ($, on) => { + const rig = await bootRig($, on) + await planRoute($) + await stepOn($, rig, FABLE) + await $.classic.StopFailure({ error: 'model_not_found' }) + await rig.clock.advance(24 * 3600000) + expect(await route($, 'show')).toContain('until reload (model_not_found)') +}) + +test('engine fallback: an auto switch marks the model it left, once', async ( + $, on) => { + const rig = await bootRig($, on) + await switchTo($, FABLE, OPUS, 'auto') + const first = await route($, 'show') + expect(first).toContain('claude-fable-5-1 for 15 min (engine fallback)') + await planRoute($) + expect((await stepOn($, rig, OPUS)).model).toBe(OPUS) + await rig.clock.advance(60000) + await switchTo($, FABLE, OPUS, 'auto') + expect(await route($, 'show')).toContain('claude-fable-5-1 for 14 min') +}) + +test('unknown: a session model absent from the table is never switched', async ( + $, on) => { + const rig = await bootRig($, on) + await planRoute($) + expect((await stepOn($, rig, 'claude-zz-9')).model).toBe('claude-zz-9') +}) + +test('spawn: Explore goes to opus while sonnet is down', async ($, on) => { + const rig = await bootRig($, on) + await $.agent.spawn(spawnInput()) + await failWith($, 'overloaded', 'a1') + await $.agent.spawn(spawnInput()) + expect(rig.specs).toEqual([SONNET, OPUS]) +}) + +/** Types a plain prompt, idle, in the composer. */ +const typed = ($: Engine, text: string) => $.prompt.submit({ + text, + wait: false, + origin: { kind: 'composer' }, +}) + +const dispatch = ($: Engine) => + $.tool.call({ tool: 'Agent', description: 'd', prompt: 'p' }) + +const agentEnds = ($: Engine) => $.turn.complete({ + turnId: 'a1', + agentId: 'a1', + answer: '', + durationMs: 1, + isAborted: false, + reason: 'answer', +}) + +test('derived: a dispatch pushes orchestrate, the end pops the prompt', async ( + $, on) => { + await bootRig($, on) + await typed($, 'planifie la migration') + expect(mainLine(await route($, 'show'))).toContain('prompt plan') + await dispatch($) + expect(mainLine(await route($, 'show'))).toContain('derived orchestrate') + await agentEnds($) + expect(mainLine(await route($, 'show'))).toContain('prompt plan') +}) + +test('derived: a route declared after the dispatch survives the pop', async ( + $, on) => { + await bootRig($, on) + await typed($, 'planifie la migration') + await dispatch($) + await $.tool.call({ tool: ROUTE_TOOL, phase: 'reflect' }) + await agentEnds($) + expect(mainLine(await route($, 'show'))).toContain('model reflect') +}) + +test('default rule: planifie sets plan, a later route overrides', async ( + $, on) => { + const rig = await bootRig($, on, HAIKU) + await typed($, 'planifie la migration') + expect(await stepOn($, rig, HAIKU)).toEqual({ model: FABLE, effort: 'xhigh' }) + await $.tool.call({ tool: ROUTE_TOOL, phase: 'implement' }) + const line = mainLine(await route($, 'show')) + expect(line).toContain('model implement') + expect(line).not.toContain('prompt') +}) + +test('default rule: a typed slash command gets no default rule', async ( + $, on) => { + await bootRig($, on) + await typed($, '/analyze pourquoi ça plante') + expect(mainLine(await route($, 'show'))).toContain('session defaults') +}) + +test('default rule: /effort-low pourquoi sets the floor, no default', async ( + $, on) => { + await bootRig($, on) + await typed($, '/effort-low pourquoi ça plante') + await $.skill.prompt({ skill: 'effort-low', text: 'x' }) + const line = mainLine(await route($, 'show')) + expect(line).toContain('floor low') + expect(line).not.toContain('reflect') +}) + +test('per axis: a model-less sticky never hides a turn route tier', async ( + $, on) => { + const rig = await bootRig($, on, HAIKU) + await route($, 'effort=low') + await planRoute($) + expect(await stepOn($, rig, HAIKU)).toEqual({ model: FABLE, effort: 'low' }) +}) + +// ---- texts follow the decision, from the turn's model ------------------ + +test('text: effort-only route on a down model says fallback', async ( + $, on) => { + const rig = await bootRig($, on) + await stepOn($, rig, FABLE) + await failWith($, 'rate_limit') + const out = await $.tool.call({ tool: ROUTE_TOOL, effort: 'low' }) + expect(JSON.stringify(out)).toContain(`model ${OPUS} (fallback)`) +}) + +test('text: /route effort= on a down model names the fallback', async ( + $, on) => { + const rig = await bootRig($, on) + await stepOn($, rig, FABLE) + await failWith($, 'rate_limit') + await route($, 'effort=low') + expect(mainLine(await route($, 'show'))).toContain(`${OPUS} (fallback)`) +}) + +test('text: show names the session model after the turn ended', async ( + $, on) => { + const rig = await bootRig($, on, HAIKU) + await planRoute($) + expect((await stepOn($, rig, HAIKU)).model).toBe(FABLE) + await endTurn($) + await route($, 'mechanical') + const line = mainLine(await route($, 'show')) + expect(line).toContain(`model ${HAIKU}`) + expect(line).not.toContain(FABLE) +}) + +test('text: show reflects a /model command to opus', async ($, on) => { + const rig = await bootRig($, on) + await stepOn($, rig, FABLE) + await switchTo($, FABLE, OPUS, 'command') + await route($, 'mechanical') + const line = mainLine(await route($, 'show')) + expect(line).toContain(OPUS) + expect(line).not.toContain(FABLE) +}) + +test('text: an unknown session model prints "session model"', async ( + $, on) => { + await bootRig($, on, 'mystery-model') + await route($, 'plan') + expect(mainLine(await route($, 'show'))).toContain('keeps session model') +}) + +test('text: show names the model a floor upgrade moves to', async ( + $, on) => { + await bootRig($, on, HAIKU) + await $.prompt.submit({ + text: 'ultrathink please', + wait: false, + origin: { kind: 'composer' }, + }) + const line = mainLine(await route($, 'show')) + expect(line).toContain(`${FABLE} (upgrade)`) +}) diff --git a/mods/model-router/hooks/register.ts b/mods/model-router/hooks/register.ts index 4d757b8..09f2d3b 100644 --- a/mods/model-router/hooks/register.ts +++ b/mods/model-router/hooks/register.ts @@ -1,14 +1,22 @@ // model-router: routes each model request (main loop and sub-agents) to the -// model and effort its phase deserves. Phases, agents, skills and prompt -// rules come from DEFAULT_CONFIG, overridable by ~/.claude/model-router.json. +// model and effort its phase deserves. Phases name a TIER (an ordered list of +// aliases, the first available wins); a circuit breaker fed by the engine's +// own failure events marks a model down for a while. Phases, agents, skills +// and prompt rules come from DEFAULT_CONFIG, overridable by +// ~/.claude/model-router.json. // One writer per concern: the Agent tool's own params are never rewritten, // the spawn hook sets an agent's model once, turn.step sets efforts. import type { EngineInterface, On, Register, TurnStepInput } from 'claude-code' type Api = EngineInterface type Level = 'low' | 'medium' | 'high' | 'xhigh' | 'max' -type Route = { model?: string; effort?: Level } // model: alias or full id -type PromptRule = { pattern: string; phase: string } +type Route = { + tier?: string // a `tiers` key: the first AVAILABLE alias of its list + model?: string // alias or full id: explicit, never skipped when down + effort?: Level +} +type PromptMode = 'floor' | 'default' +type PromptRule = { pattern: string; phase: string; mode?: PromptMode } type Config = { models: Record // alias -> full id windows: Record // full id -> context window (tokens) @@ -16,14 +24,21 @@ type Config = { agents: Record // built-in subagentType -> phase skills: Record // skill name -> phase prompt: PromptRule[] - mainModelSwitch: boolean + tiers: Record // tier -> alias preference, best first + fallback: string[] // alias order, best first: the rank and the fallback chain + cooldownMinutes: number // breaker hold at the first strike + mainUpgrade: boolean // main may move UP to a phase's tier by itself + upgradeMaxTokens: number // above this context an upgrade is skipped + mainModelSwitch: boolean // gates DOWNGRADES only verbose: boolean spinner: boolean enabled: boolean // false: every hook passes through (per machine) } -type Rule = { re: RegExp; phase: string } -type Source = 'user' | 'model' | 'skill' | 'prompt' | 'slash' +type Rule = { re: RegExp; phase: string; mode: PromptMode } +type Source = 'user' | 'model' | 'skill' | 'prompt' | 'slash' | 'derived' type Routed = { phase: string; route: Route; source: Source } +type Hold = { until: number; reason: string } // until: ms, Infinity = reload +type Pushed = { prev: Routed | null; spawnIds: Set } type Loop = { effort?: Level // an agent's model is fixed at spawn: effort is its only axis explicitEffort: boolean // Agent call gave an effort: axis frozen @@ -34,7 +49,7 @@ type State = { source: string // 'defaults' or the override path userMain: Routed | null // /route by the user, sticky until /route clear turnMain: Routed | null // model route tool, skill table row, Skill(effort-*) - // bridge; dropped at turn end + // bridge, prompt default rule, derived orchestrate; dropped at turn end turnFloor: Routed | null // user-explicit level for this turn (prompt rule, // typed /effort-): a floor, main loop only pendingPrompt: Routed | null // typed mid-turn: the next turn's floor @@ -44,8 +59,16 @@ type State = { skillCalls: number // Skill tool calls in flight off: boolean // /route off or config: every hook passes through offConfig: boolean // `off` comes from the config key `enabled` - lastMain: string // "model/effort" of the last main step (spinner) - windowWarned: boolean // context-window warning already logged this turn + spinner: string // "model/effort" of the last main step of this turn + lastPlan: Plan | null // the plan last sent on main: the breaker's target + turnModel: string | undefined // sticky main model, set when the router moved + sessionModel: string // the main loop's model as /model shows it, '' unknown + down: Map // canonical id -> unavailable until (breaker) + strikes: Map // canonical id -> episodes, drives the backoff + agentModels: Map // agentId -> exact id sent at spawn + pushed: Pushed | null // derived orchestrate in force (background dispatch) + logged: Set // session-once log lines + turnLogged: Set // turn-once log lines warned: Set // hooks whose fail-open was already logged } type Log = (text: string) => void @@ -71,6 +94,16 @@ const HAIKU = 'claude-haiku' const MAX_PATTERN = 200 // chars of a prompt-rule pattern const MAX_PROMPT_SCAN = 4096 // chars of a prompt a rule is run against const MAX_CONFIG_BYTES = 65536 // override file size +const MAX_AGENT_MODELS = 256 // agent -> model entries kept (oldest dropped) +const PREFIX = 'claude-' +const ONE_M = /\[1m\]$/ // the long-context variant of a model id +// Hold = cooldownMinutes x step: 15 -> 30 -> 60 -> 120 -> 300 by default. +const BACKOFF_STEPS: readonly number[] = [1, 2, 4, 8, 20] +// StopFailure kinds that mean "this model cannot serve now". The others +// (context limit, network, auth...) are not availability. +const UNAVAILABLE: ReadonlySet = new Set([ + 'rate_limit', 'overloaded', 'billing_error', 'model_not_found', +]) const PHASE_KEY = /^[a-z][a-z0-9_-]{0,31}$/ const DEFAULT_CONFIG: Config = { @@ -81,22 +114,49 @@ const DEFAULT_CONFIG: Config = { fable: 'claude-fable-5-1', }, windows: { 'claude-haiku-4-5-20251001': 200000 }, + // A tier is an ordered alias list: the first one not down is used. + tiers: { + best: ['fable', 'opus', 'sonnet'], + big: ['opus', 'fable', 'sonnet'], + work: ['sonnet', 'opus'], + cheap: ['haiku', 'sonnet'], + }, + fallback: ['fable', 'opus', 'sonnet', 'haiku'], // rank = index, best first + cooldownMinutes: 15, + mainUpgrade: true, + upgradeMaxTokens: 200000, // an upgrade re-reads the whole context cold phases: { - plan: { effort: 'xhigh' }, - reflect: { effort: 'high' }, - orchestrate: { effort: 'medium' }, - escalate: { effort: 'max' }, - judge: { model: 'opus', effort: 'xhigh' }, - implement: { model: 'sonnet', effort: 'medium' }, - write: { model: 'sonnet', effort: 'medium' }, - verify: { model: 'sonnet', effort: 'xhigh' }, - explore: { model: 'sonnet', effort: 'medium' }, - mechanical: { model: 'haiku', effort: 'low' }, + plan: { tier: 'best', effort: 'xhigh' }, + reflect: { tier: 'best', effort: 'high' }, + orchestrate: { tier: 'best', effort: 'medium' }, + escalate: { tier: 'best', effort: 'max' }, + judge: { tier: 'big', effort: 'xhigh' }, + implement: { tier: 'work', effort: 'medium' }, + write: { tier: 'work', effort: 'medium' }, + verify: { tier: 'work', effort: 'xhigh' }, + explore: { tier: 'work', effort: 'medium' }, + mechanical: { tier: 'cheap', effort: 'low' }, }, // Built-ins only: repo agents keep their frontmatter pin (wave 2). agents: { Explore: 'explore', Plan: 'judge' }, skills: {}, - prompt: [{ pattern: '\\bultrathink\\b', phase: 'escalate' }], + // `floor` rules set the turn's minimum effort; `default` rules set the + // turn's route, which a route call or a skill overrides. Neither lowers. + prompt: [ + { pattern: '\\bultrathink\\b', phase: 'escalate', mode: 'floor' }, + { + pattern: '(?, v: unknown): v is string { return typeof v === 'string' && (hasKey(models, v) || isFullId(v)) } -function resolveModel(cfg: Config, name: string): string { - return hasKey(cfg.models, name) ? (cfg.models[name] ?? name) : name -} - function phaseRoute(cfg: Config, phase: string): Route | undefined { return hasKey(cfg.phases, phase) ? cfg.phases[phase] : undefined } @@ -156,9 +212,14 @@ const acceptWindow = (key: string, v: unknown): number | undefined => ? v : undefined -function acceptPhase(models: Record) { +/** A phase names a tier OR a model, never both; a tier must exist. */ +function acceptPhase( + models: Record, + tiers: Record, +) { return (key: string, v: unknown): Route | undefined => { if (!PHASE_KEY.test(key) || !isRecord(v)) return undefined + if (v.tier !== undefined && v.model !== undefined) return undefined const route: Route = {} if (v.effort !== undefined) { if (!isLevel(v.effort)) return undefined @@ -168,10 +229,66 @@ function acceptPhase(models: Record) { if (!isModelName(models, v.model)) return undefined route.model = v.model } - return route.effort || route.model ? route : undefined + if (v.tier !== undefined) { + if (typeof v.tier !== 'string' || !hasKey(tiers, v.tier)) return undefined + route.tier = v.tier + } + return route.effort || route.model || route.tier ? route : undefined } } +/** A tier: a non-empty alias list; unknown aliases dropped, duplicates too. */ +function acceptTier(models: Record, log: Log) { + return (key: string, v: unknown): string[] | undefined => { + // a tier named like an alias would make a bare name ambiguous + if (!PHASE_KEY.test(key) || hasKey(models, key)) return undefined + if (!Array.isArray(v)) return undefined + const items = v as unknown[] + const kept = items.filter( + (a): a is string => typeof a === 'string' && hasKey(models, a)) + if (kept.length < items.length) { + log(`model-router: config tiers.${key} entries dropped`) + } + return kept.length > 0 ? [...new Set(kept)] : undefined + } +} + +/** Appends the `models` aliases a user's chain omits, so all are ranked. */ +function withEveryAlias( + chain: string[], + models: Record, + log: Log, +): string[] { + const missing = Object.keys(models).filter(a => !chain.includes(a)) + if (missing.length > 0) { + log(`model-router: config fallback lacks ${missing.join(', ')}; appended`) + } + return [...chain, ...missing] +} + +/** The fallback chain: deduplicated aliases of `models`, else the default. */ +function pickFallback( + user: unknown, + base: string[], + models: Record, + log: Log, +): string[] { + if (user === undefined) return base + const items = Array.isArray(user) ? (user as unknown[]) : [] + const kept = items.filter( + (a): a is string => typeof a === 'string' && hasKey(models, a)) + if (kept.length > 0) return withEveryAlias([...new Set(kept)], models, log) + log('model-router: config fallback ignored: need a list of model aliases') + return base +} + +function pickPositive(v: unknown, fallback: number, name: string, log: Log) { + if (v === undefined) return fallback + if (typeof v === 'number' && Number.isInteger(v) && v > 0) return v + log(`model-router: config ${name} ignored: not a positive integer`) + return fallback +} + function acceptPhaseRef(phases: Record) { return (_key: string, v: unknown): string | undefined => typeof v === 'string' && PHASE_KEY.test(v) && hasKey(phases, v) @@ -179,22 +296,32 @@ function acceptPhaseRef(phases: Record) { : undefined } -function compiles(pattern: string): boolean { - try { - new RegExp(pattern, 'i') - return true - } catch { - return false +/** Unicode-aware first; a pattern only valid without `u` still works. */ +function buildRegex(pattern: string): RegExp | undefined { + for (const flags of ['iu', 'i']) { + try { + return new RegExp(pattern, flags) + } catch { + // retry with the next flag set + } } + return undefined } -function acceptRule(phases: Record, v: unknown) { +/** Absent `mode` = floor: override files written before modes keep meaning. */ +function acceptRule( + phases: Record, + v: unknown, +): PromptRule | undefined { if (!isRecord(v)) return undefined - const { pattern, phase } = v + const { pattern, phase, mode } = v if (typeof pattern !== 'string' || typeof phase !== 'string') return undefined + if (mode !== undefined && mode !== 'floor' && mode !== 'default') { + return undefined + } if (pattern.length > MAX_PATTERN) return undefined - return hasKey(phases, phase) && compiles(pattern) - ? { pattern, phase } + return hasKey(phases, phase) && buildRegex(pattern) + ? { pattern, phase, mode: mode ?? 'floor' } : undefined } @@ -231,6 +358,34 @@ function pickEnabled(v: unknown, fallback: boolean, log: Log): boolean { return pickBool(v, fallback) } +type Routing = Pick + +/** Tiers, fallback chain and breaker knobs, validated against `models`. */ +function mergeRouting( + base: Config, + user: Record, + models: Record, + log: Log, +): Routing { + const tiers = mergeTable( + base.tiers, user.tiers, 'tiers', acceptTier(models, log), log) + const fallback = pickFallback(user.fallback, base.fallback, models, log) + if (tiers.best?.[0] !== fallback[0]) { + log('model-router: tiers.best does not lead the fallback chain; rank ' + + 'comes from the fallback list alone') + } + return { + tiers, + fallback, + cooldownMinutes: pickPositive( + user.cooldownMinutes, base.cooldownMinutes, 'cooldownMinutes', log), + mainUpgrade: pickBool(user.mainUpgrade, base.mainUpgrade), + upgradeMaxTokens: pickPositive( + user.upgradeMaxTokens, base.upgradeMaxTokens, 'upgradeMaxTokens', log), + } +} + /** Defaults overlaid with the user's entries, each validated first. */ function mergeConfig(user: unknown, log: Log): Config { const base = structuredClone(DEFAULT_CONFIG) @@ -240,11 +395,13 @@ function mergeConfig(user: unknown, log: Log): Config { } const models = mergeTable( base.models, user.models, 'models', acceptModel, log) - const phases = mergeTable( - base.phases, user.phases, 'phases', acceptPhase(models), log) + const routing = mergeRouting(base, user, models, log) + const phases = mergeTable(base.phases, user.phases, 'phases', + acceptPhase(models, routing.tiers), log) const ref = acceptPhaseRef(phases) return { models, + ...routing, windows: mergeTable( base.windows, user.windows, 'windows', acceptWindow, log), phases, @@ -310,11 +467,15 @@ async function loadConfig( return { cfg: mergeConfig(found.data, log), source: found.path } } -function compileRules(cfg: Config): Rule[] { - return cfg.prompt.map(r => ({ - re: new RegExp(r.pattern, 'i'), - phase: r.phase, - })) +/** Compiles the prompt rules; one that fails to compile is dropped, logged. */ +function compileRules(cfg: Config, log: Log): Rule[] { + const rules: Rule[] = [] + for (const r of cfg.prompt) { + const re = buildRegex(r.pattern) + if (re) rules.push({ re, phase: r.phase, mode: r.mode ?? 'floor' }) + else log(`model-router: prompt rule for ${r.phase} does not compile`) + } + return rules } // ---- state ----------------------------------------------------------- @@ -322,7 +483,7 @@ function compileRules(cfg: Config): Rule[] { function newState(cfg: Config, source: string): State { return { cfg, - rules: compileRules(cfg), + rules: compileRules(cfg, () => undefined), source, userMain: null, turnMain: null, @@ -334,12 +495,34 @@ function newState(cfg: Config, source: string): State { skillCalls: 0, off: false, offConfig: false, - lastMain: '', - windowWarned: false, + spinner: '', + lastPlan: null, + turnModel: undefined, + sessionModel: '', + down: new Map(), + strikes: new Map(), + agentModels: new Map(), + pushed: null, + logged: new Set(), + turnLogged: new Set(), warned: new Set(), } } +/** + * /clear rebuilds the state but keeps what is account- or process-wide: + * the breaker (a model's quota outlives the conversation) and the model. + */ +function resetSession(st: State): void { + const kept = { + down: st.down, + strikes: st.strikes, + agentModels: st.agentModels, + sessionModel: st.sessionModel, + } + Object.assign(st, newState(st.cfg, st.source), kept) +} + /** Logs a hook's fail-open once per session; never throws itself. */ function warnOnce(st: State, $: Api, hook: string, kind: string): void { if (st.warned.has(hook)) return @@ -361,6 +544,289 @@ function safely(st: State, $: Api, hook: string, work: () => void): void { } } +// ---- models: ids, tiers, breaker, the main decision ------------------- + +/** The table id an alias, a `[1m]` variant or a dated id stands for. */ +function canonical(cfg: Config, id: string): string { + const bare = id.replace(ONE_M, '') + if (hasKey(cfg.models, bare)) return cfg.models[bare] ?? bare + if (!bare.startsWith(PREFIX) || bare.length <= PREFIX.length) return bare + const hit = Object.values(cfg.models).find( + known => bare.startsWith(known)) + return hit ?? bare +} + +/** The table alias of a canonical id; undefined for an id the table lacks. */ +const aliasOf = (cfg: Config, id: string): string | undefined => + Object.keys(cfg.models).find(alias => cfg.models[alias] === id) + +/** Position in the fallback chain, best first; undefined when unranked. */ +function modelRank(cfg: Config, id: string): number | undefined { + const alias = aliasOf(cfg, id) + const at = alias === undefined ? -1 : cfg.fallback.indexOf(alias) + return at < 0 ? undefined : at +} + +const routeName = (r: Route | null | undefined): string | undefined => + r?.tier ?? r?.model + +/** First alias of the list whose model is not down. */ +function availableIn(st: State, aliases: readonly string[]) { + for (const alias of aliases) { + const id = st.cfg.models[alias] + if (id !== undefined && !st.down.has(id)) return id + } + return undefined +} + +/** The first model not down in the fallback chain, after `after`'s place. */ +function nextAvailable(st: State, after: string | undefined) { + const { fallback, models } = st.cfg + const at = fallback.findIndex(alias => models[alias] === after) + return availableIn(st, fallback.slice(at + 1)) +} + +/** + * A tier, alias or id as an id to run on. A tier skips down aliases, then + * walks the global chain (from `after`, default the tier's last alias); an + * alias or id is explicit and never skipped. + */ +function resolveName(st: State, name: string, after?: string) { + const tier = hasKey(st.cfg.tiers, name) ? st.cfg.tiers[name] : undefined + if (tier === undefined) return canonical(st.cfg, name) + const last = tier[tier.length - 1] + const from = after ?? (last === undefined ? undefined : st.cfg.models[last]) + return availableIn(st, tier) ?? nextAvailable(st, from) +} + +function resolveRoute(st: State, route: Route | undefined) { + const name = routeName(route) + return name === undefined ? undefined : resolveName(st, name) +} + +/** Drops lapsed holds; the clock is read only while something is held. */ +async function prune($: Api, st: State): Promise { + if (st.down.size === 0) return 0 + const now = await $.clock.now() + for (const [id, hold] of st.down) { + if (hold.until <= now) st.down.delete(id) + } + return now +} + +const holdWord = (until: number, now: number): string => + until === Infinity + ? 'until reload' + : `for ${Math.max(1, Math.ceil((until - now) / 60000))} min` + +function holdMinutes(cfg: Config, strikes: number, reason: string): number { + if (reason === 'model_not_found') return Infinity + const step = BACKOFF_STEPS[Math.min(strikes, BACKOFF_STEPS.length) - 1] ?? 1 + return cfg.cooldownMinutes * step +} + +/** + * Marks a table model down for an episode. A mark on a model already down + * adds no strike and no log, except `model_not_found`, which lengthens a + * timed hold to "until reload". Inert while the router is off. + */ +function markDown( + $: Api, + st: State, + id: string, + reason: string, + now: number, +): void { + if (st.off || aliasOf(st.cfg, id) === undefined) return + const held = st.down.get(id) + const lengthen = reason === 'model_not_found' && held?.until !== Infinity + if (held && !lengthen) return + const strikes = (st.strikes.get(id) ?? 0) + (held ? 0 : 1) + st.strikes.set(id, strikes) + const until = now + holdMinutes(st.cfg, strikes, reason) * 60000 + st.down.set(id, { until, reason }) + $.ui.log(`model-router: ${id} unavailable (${reason}) ${ + holdWord(until, now)}; routing falls back`) +} + +/** `/route reload` and a user /model: the marks are stale. */ +function clearBreaker(st: State, id?: string): void { + if (id === undefined) { + st.down.clear() + st.strikes.clear() + } else { + st.down.delete(id) + st.strikes.delete(id) + } + st.turnModel = undefined +} + +type Call = { + model: string // exact string to send: `cur` verbatim when not moved + why: string + moved: boolean + log?: string // turn-once line + logKey?: string // its dedupe key when the text varies; default the text + once?: string // session-once line +} + +const keep = (cur: string, why = 'unchanged', log?: string): Call => + ({ model: cur, why, moved: false, ...(log === undefined ? {} : { log }) }) + +/** The replacement keeps the session's `[1m]` tier, haiku has none. */ +function moveTo(cur: string, target: string, why: string): Call { + const carried = ONE_M.test(cur) && !target.startsWith(HAIKU) + const model = carried ? `${target}[1m]` : target + const once = carried + ? `model-router: ${why} to ${model}: the [1m] variant is carried over` + : undefined + return { model, why, moved: true, ...(once === undefined ? {} : { once }) } +} + +/** True when the context still fits the target model's known window. */ +function fits(st: State, id: string, tokens: number | undefined): boolean { + const limit = hasKey(st.cfg.windows, id) ? st.cfg.windows[id] : undefined + return limit === undefined || (tokens !== undefined && tokens < limit) +} + +const noFit = (id: string): string => + `model-router: no switch to ${id}: context not known to fit` + +/** Why the upgrade cap blocks a move, or undefined. Unknown size = blocked. */ +function capBlock(st: State, tokens: number | undefined): string | undefined { + if (tokens === undefined) return 'context size unknown' + const max = st.cfg.upgradeMaxTokens + return tokens > max ? `context ${tokens} tokens over ${max}` : undefined +} + +/** True when `wanted` ranks above `id` in the fallback chain. */ +function ranksAbove(st: State, wanted: string, id: string): boolean { + const rw = modelRank(st.cfg, wanted) + const rc = modelRank(st.cfg, id) + return rw !== undefined && rc !== undefined && rw < rc +} + +/** + * `cur` is down: `wanted`, else the next model of the chain that fits. A + * better `wanted` is an upgrade and passes the same switch and cap. + */ +function leaveDown( + st: State, + cur: string, + wanted: string | undefined, + tokens: number | undefined, +): Call { + const id = canonical(st.cfg, cur) + const upOk = st.cfg.mainUpgrade && capBlock(st, tokens) === undefined + for (const option of [wanted, nextAvailable(st, id)]) { + if (option === undefined || st.down.has(option)) continue + if (aliasOf(st.cfg, option) === undefined) continue + if (option === wanted && !upOk && ranksAbove(st, option, id)) continue + if (fits(st, option, tokens)) return moveTo(cur, option, 'fallback') + } + return keep(cur, 'fallback unavailable') +} + +/** `wanted` ranks above `cur`: upgrade, under the switch and the cap. */ +function upgradeCall( + st: State, + cur: string, + wanted: string, + tokens: number | undefined, +): Call { + if (!st.cfg.mainUpgrade) return keep(cur, 'switch off') + const block = capBlock(st, tokens) + if (block !== undefined) { + const why = `upgrade skipped: ${block}` + const call = keep(cur, why, `model-router: ${why}`) + return { ...call, logKey: 'upgrade-skipped' } + } + if (!fits(st, wanted, tokens)) return keep(cur, 'no fit', noFit(wanted)) + return moveTo(cur, wanted, 'upgrade') +} + +/** `wanted` is cheaper (or unranked): only with the downgrade switch on. */ +function downgradeCall( + st: State, + cur: string, + wanted: string, + tokens: number | undefined, +): Call { + if (!st.cfg.mainModelSwitch) return keep(cur, 'switch off') + if (!fits(st, wanted, tokens)) return keep(cur, 'no fit', noFit(wanted)) + return moveTo(cur, wanted, 'downgrade') +} + +/** A model the table does not rank is never switched; said once. */ +function unknownCall(cur: string, id: string): Call { + const call = keep(cur, 'model unknown to the table') + if (id === '') return call + const once = `model-router: ${id} unknown to the models table; no switch` + return { ...call, once } +} + +/** + * The one decision of the main loop's model, in this order: router off, a + * model the table does not rank (never touched), `cur` down (leave it, + * always), no or same wanted model, a better one (upgrade), a cheaper one + * (downgrade). Used by the step AND by every text, so they agree. + */ +function decideMain( + st: State, + cur: string, + wanted: string | undefined, + tokens: number | undefined, +): Call { + const id = canonical(st.cfg, cur) + if (st.off) return keep(cur) + if (modelRank(st.cfg, id) === undefined) return unknownCall(cur, id) + if (st.down.has(id)) return leaveDown(st, cur, wanted, tokens) + if (wanted === undefined) return keep(cur) + if (aliasOf(st.cfg, wanted) === aliasOf(st.cfg, id)) return keep(cur) + return ranksAbove(st, wanted, id) + ? upgradeCall(st, cur, wanted, tokens) + : downgradeCall(st, cur, wanted, tokens) +} + +/** Model axis: sticky, then turn route, then the floor's own model. */ +const mainModel = (st: State): string | undefined => + routeName(st.userMain?.route) ?? + routeName(st.turnMain?.route) ?? + routeName(st.turnFloor?.route) + +async function readTokens($: Api): Promise { + try { + return (await $.session.usage()).context.tokens + } catch { + return undefined + } +} + +type Verdict = { call: Call; wanted: string | undefined } + +/** What the main loop would run on from `cur`: wanted model + decision. */ +async function decideFor($: Api, st: State, cur: string): Promise { + const name = mainModel(st) + const wanted = name === undefined + ? undefined + : resolveName(st, name, canonical(st.cfg, cur)) + const needsTokens = wanted !== undefined || st.down.size > 0 + const tokens = needsTokens ? await readTokens($) : undefined + return { call: decideMain(st, cur, wanted, tokens), wanted } +} + +function logCall($: Api, st: State, call: Call): void { + const key = call.logKey ?? call.log + if (call.log !== undefined && key !== undefined && !st.turnLogged.has(key)) { + st.turnLogged.add(key) + $.ui.log(call.log) + } + if (call.once !== undefined && !st.logged.has(call.once)) { + st.logged.add(call.once) + $.ui.log(call.once) + } +} + /** The main loop's effective route: user /route > latest turn route. */ const mainRoute = (st: State): Routed | null => st.userMain ?? st.turnMain @@ -396,12 +862,6 @@ function mainEffort(st: State, engine: Effort): Decision { return { effort, by: sticky === undefined ? 'turn' : 'sticky' } } -/** Model axis: sticky, then turn route, then the floor's own model. */ -const mainModel = (st: State): string | undefined => - st.userMain?.route.model ?? - st.turnMain?.route.model ?? - st.turnFloor?.route.model - const floorSource = (f: Routed): string => f.source === 'prompt' ? `prompt rule ${f.phase}` : `typed /${f.phase}` @@ -463,9 +923,6 @@ const routerWord = (st: State): string => // ---- text ------------------------------------------------------------ -const modelText = (cfg: Config, model: string | undefined): string => - model === undefined ? '-' : resolveModel(cfg, model) - /** The floor's level when it carries one and the router is on. */ function liveFloor(st: State): { f: Routed; level: Level } | undefined { const f = st.turnFloor @@ -473,44 +930,92 @@ function liveFloor(st: State): { f: Routed; level: Level } | undefined { return st.off || !f || level === undefined ? undefined : { f, level } } -/** True when the switch would put the main loop on a haiku model. */ -function mainOnHaiku(st: State): boolean { - const model = mainModel(st) - return st.cfg.mainModelSwitch && model !== undefined && - resolveModel(st.cfg, model).startsWith(HAIKU) +/** An id the table knows, else "session model": '' and foreign ids alike. */ +function idWord(st: State, model: string): string { + const known = aliasOf(st.cfg, canonical(st.cfg, model)) !== undefined + return known ? model : 'session model' } -function effortWord(st: State): string { - if (mainOnHaiku(st)) return '- (haiku takes none)' +/** The router left a down model, or wanted to and found nowhere to go. */ +const isFallback = (call: Call): boolean => call.why.startsWith('fallback') + +/** Model words of the main line, from the decision a step would take. */ +function modelWord(st: State, v: Verdict): string { + const id = idWord(st, v.call.model) + if (v.call.moved) return `${id} (${v.call.why})` + if (v.wanted === undefined) { + return isFallback(v.call) ? `${id} (${v.call.why})` : '-' + } + if (canonical(st.cfg, v.call.model) === v.wanted) return id + return `asked ${v.wanted}, keeps ${id} (${v.call.why})` +} + +/** The tier a route names, shown beside the model it resolved to. */ +function tierWord(st: State, word: string): string { + const name = mainModel(st) + const tier = name !== undefined && hasKey(st.cfg.tiers, name) + return tier ? `${word} [tier ${name}]` : word +} + +/** Haiku takes no effort: the main line says so when it will run there. */ +function effortWord(st: State, v: Verdict): string { + if (v.call.model.startsWith(HAIKU)) return '- (haiku takes none)' return String(mainEffort(st, undefined).effort ?? '-') } -function mainText(st: State): string { +function mainText(st: State, v: Verdict): string { const r = mainRoute(st) const live = liveFloor(st) const floor = live ? ` · floor ${live.level} (${live.f.phase})` : '' - if (!r) return 'main: session defaults' + floor - const model = modelText(st.cfg, mainModel(st)) + if (!r) { + const left = v.call.moved || isFallback(v.call) + ? ` · model ${modelWord(st, v)}` + : '' + return 'main: session defaults' + left + floor + } + const model = tierWord(st, modelWord(st, v)) return `main: ${r.source} ${r.phase} · model ${model} · effort ${ - effortWord(st)}${floor}` + effortWord(st, v)}${floor}` } -function phasesText(cfg: Config): string { - const entries = Object.entries(cfg.phases).map(([name, r]) => - `${name}=${r.model ? resolveModel(cfg, r.model) : 'session'}/${ - r.effort ?? 'session'}`) +/** `name=→/` for every phase. */ +function phasesText(st: State): string { + const entries = Object.entries(st.cfg.phases).map(([name, r]) => { + const asked = routeName(r) + const id = asked === undefined ? undefined : resolveName(st, asked) + const model = asked === undefined + ? 'session' + : asked === id ? asked : `${asked}→${id ?? '-'}` + return `${name}=${model}/${r.effort ?? 'session'}` + }) return `phases: ${entries.join(' ')}` } -function show(st: State): string { +function downText(st: State, now: number): string { + const held = [...st.down].map(([id, h]) => + `${id} ${holdWord(h.until, now)} (${h.reason})`) + return `down: ${held.length > 0 ? held.join(', ') : 'none'}` +} + +/** The verdict a text reports: what the next main step would decide. */ +async function snapshot($: Api, st: State) { + const now = await prune($, st) + const cur = st.turnModel ?? st.sessionModel + return { now, ...(await decideFor($, st, cur)) } +} + +async function show($: Api, st: State): Promise { const c = st.cfg const flag = (b: boolean) => (b ? 'on' : 'off') + const s = await snapshot($, st) return [ - mainText(st), - `router: ${routerWord(st)} · switch: ${flag(c.mainModelSwitch)} · ` + - `verbose: ${flag(c.verbose)} · spinner: ${flag(c.spinner)}`, + mainText(st, s), + `router: ${routerWord(st)} · switch: ${flag(c.mainModelSwitch)} ` + + `(downgrade) · upgrade: ${flag(c.mainUpgrade)} · verbose: ${ + flag(c.verbose)} · spinner: ${flag(c.spinner)}`, + downText(st, s.now), `live loops: ${st.loops.size}`, - phasesText(c), + phasesText(st), `config: ${st.source}`, ].join('\n') } @@ -565,19 +1070,24 @@ function parseRoute(cfg: Config, args: string): Routed | string { return { phase: 'custom', route, source: 'user' } } -function toggle(st: State, what: string, arg: string | undefined): string { +async function toggle( + $: Api, + st: State, + what: string, + arg: string | undefined, +): Promise { if (arg !== 'on' && arg !== 'off') return `usage: /route ${what} on|off` if (what === 'switch') st.cfg.mainModelSwitch = arg === 'on' else st.cfg.verbose = arg === 'on' - return show(st) + return show($, st) } -function setUserRoute($: Api, st: State, args: string): string { +async function setUserRoute($: Api, st: State, args: string) { const parsed = parseRoute(st.cfg, args) if (typeof parsed === 'string') return parsed st.userMain = parsed refresh($, st) - return show(st) + return show($, st) } /** @@ -589,7 +1099,7 @@ async function reloadConfig($: Api, st: State): Promise { const loaded = await loadConfig($, text => $.ui.log(text)) if (loaded) { st.cfg = loaded.cfg - st.rules = compileRules(loaded.cfg) + st.rules = compileRules(loaded.cfg, text => $.ui.log(text)) st.source = loaded.source applyEnabled(st, loaded.cfg.enabled) } else { @@ -598,29 +1108,35 @@ async function reloadConfig($: Api, st: State): Promise { await registerTool($, st) } +/** The breaker is cleared first, whatever the config read then does. */ +async function reloadCommand($: Api, st: State): Promise { + clearBreaker(st) + await reloadConfig($, st) + refresh($, st) + return 'config reloaded\n' + (await show($, st)) +} + async function handleCommand($: Api, st: State, args: string): Promise { const [head = '', ...rest] = args.trim().split(/\s+/) switch (head) { case '': case 'show': - return show(st) + return show($, st) case 'clear': clearRoutes(st) refresh($, st) - return 'route cleared\n' + show(st) + return 'route cleared\n' + (await show($, st)) case 'on': case 'off': st.off = head === 'off' st.offConfig = false refresh($, st) - return show(st) + return show($, st) case 'reload': - await reloadConfig($, st) - refresh($, st) - return 'config reloaded\n' + show(st) + return reloadCommand($, st) case 'switch': case 'verbose': - return toggle(st, head, rest[0]) + return toggle($, st, head, rest[0]) default: return setUserRoute($, st, args) } @@ -634,10 +1150,11 @@ async function registerTool($: Api, st: State): Promise { name: 'route', description: 'Declare the phase of the work ahead so the next requests of THIS ' + - 'loop run at the right effort (and, on the main loop, the model ' + - 'when the switch is on). Call it before a span of work changes ' + - 'nature (planning, orchestrating, mechanical work). Effort only ' + - 'for a sub-agent; the model of a sub-agent is fixed at spawn. ' + + 'loop run at the right effort and model. Call it before a span of ' + + 'work changes nature (planning, orchestrating, mechanical work). ' + + 'The main loop moves up to a phase\'s tier by itself (below the ' + + 'context cap), down only with the switch on; a sub-agent\'s model ' + + 'is fixed at spawn. ' + `Phases: ${Object.keys(st.cfg.phases).join(', ')}.`, inputSchema: { type: 'object', @@ -708,25 +1225,34 @@ function clearLoop(st: State, agentId: string | undefined): string { return 'route cleared for this agent' } +/** What the main loop's model will do, in the words of the decision. */ +async function mainAnswer($: Api, st: State): Promise { + const { call, wanted } = await snapshot($, st) + if (call.moved) return `${idWord(st, call.model)} (${call.why})` + const quiet = call.why === 'unchanged' || + (wanted === undefined && !isFallback(call)) + return quiet ? 'unchanged' : `unchanged (${call.why})` +} + /** Truthful answer: states what the calling loop will actually do. */ -function routedText(st: State, agentId: string | undefined, p: Picked): string { +async function routedText( + $: Api, + st: State, + agentId: string | undefined, + p: Picked, +): Promise { const note = agentId === undefined ? mainNote(st, p.route.effort) : '' if (note) return `recorded ${p.phase} for this turn, but ${note}` const loop = agentId === undefined ? undefined : st.loops.get(agentId) const effort = loop?.explicitEffort ? undefined : p.route.effort - const model = agentId === undefined && st.cfg.mainModelSwitch - ? p.route.model - : undefined - const modelNote = !p.route.model || model !== undefined - ? '' - : agentId === undefined ? ' (switch off)' : ' (fixed at spawn)' + const model = agentId === undefined + ? await mainAnswer($, st) + : routeName(p.route) ? 'unchanged (fixed at spawn)' : 'unchanged' return `routed ${agentId === undefined ? 'main' : 'this agent'} to ` + - `${p.phase}: effort ${effort ?? 'unchanged'}, model ` + - `${model === undefined ? 'unchanged' : resolveModel(st.cfg, model)}` + - modelNote + `${p.phase}: effort ${effort ?? 'unchanged'}, model ${model}` } -function handleRouteTool(st: State, e: RouteInput) { +async function handleRouteTool($: Api, st: State, e: RouteInput) { if (st.off) { return { result: 'model-router is off (/route on to resume); nothing routed', @@ -736,7 +1262,7 @@ function handleRouteTool(st: State, e: RouteInput) { const picked = pickRoute(st.cfg, e.phase, e.effort) if (typeof picked === 'string') return { deny: picked } applyRoute(st, e.agentId, picked) - return { result: routedText(st, e.agentId, picked) } + return { result: await routedText($, st, e.agentId, picked) } } // ---- skills ---------------------------------------------------------- @@ -839,6 +1365,11 @@ function spawnRoute(cfg: Config, e: SpawnIn, frozen: boolean) { return phase === undefined ? undefined : phaseRoute(cfg, phase) } +/** The model a spawn is rewritten to; an explicit `model` param wins. */ +function spawnTarget(st: State, e: SpawnIn, route: Route | undefined) { + return e.model === undefined ? resolveRoute(st, route) : undefined +} + function trackLoop(st: State, e: SpawnIn, started: { agentId?: string }, route: Route | undefined): void { @@ -851,6 +1382,62 @@ function trackLoop(st: State, e: SpawnIn, started: { }) } +/** The breaker's target for an agent's failure; oldest entries dropped. */ +function rememberAgent(st: State, agentId: string, model: string): void { + st.agentModels.set(agentId, model.replace(ONE_M, '')) + if (st.agentModels.size <= MAX_AGENT_MODELS) return + const oldest = st.agentModels.keys().next().value + if (oldest !== undefined) st.agentModels.delete(oldest) +} + +// ---- derived orchestrate --------------------------------------------- + +/** + * A main Agent dispatch is orchestration: the turn's route becomes + * `orchestrate` until the spawned agents end. A route the model or a skill + * declared is never replaced. + */ +function pushOrchestrate(st: State): void { + const src = st.turnMain?.source + const route = phaseRoute(st.cfg, 'orchestrate') + if (st.pushed !== null || !route) return + if (src === 'model' || src === 'skill') return + st.pushed = { prev: st.turnMain, spawnIds: new Set() } + st.turnMain = { phase: 'orchestrate', route: { ...route }, source: 'derived' } +} + +/** Restores the route from before the dispatch, unless one replaced it. */ +function popOrchestrate(st: State): void { + if (st.pushed && st.turnMain?.source === 'derived') { + st.turnMain = st.pushed.prev + } + st.pushed = null +} + +/** The agent id of a background launch, from the Agent tool's result. */ +function launchedId(result: unknown): string | undefined { + if (!isRecord(result) || result.status !== 'async_launched') return undefined + return typeof result.agentId === 'string' ? result.agentId : undefined +} + +/** After the Agent call: wait for a background agent, else pop at once. */ +function afterDispatch(st: State, result: unknown): void { + const id = launchedId(result) + if (id !== undefined) { + pushOrchestrate(st) + st.pushed?.spawnIds.add(id) + } else if (st.pushed && st.pushed.spawnIds.size === 0) { + popOrchestrate(st) + } +} + +/** An agent's loop ended: forget it, pop when it was the last awaited. */ +function endAgent(st: State, agentId: string): void { + st.loops.delete(agentId) + const awaited = st.pushed?.spawnIds + if (awaited?.delete(agentId) && awaited.size === 0) popOrchestrate(st) +} + // ---- turn steps ------------------------------------------------------ function agentPlan(st: State, e: StepIn): Plan { @@ -858,32 +1445,19 @@ function agentPlan(st: State, e: StepIn): Plan { return { model: e.model, effort: loop?.effort ?? e.effort } } -/** True when the context still fits the target model's known window. */ -async function windowOk($: Api, st: State, id: string): Promise { - const limit = hasKey(st.cfg.windows, id) ? st.cfg.windows[id] : undefined - if (limit === undefined) return true - let tokens: number | undefined - try { - tokens = (await $.session.usage()).context.tokens - } catch { - tokens = undefined - } - if (typeof tokens === 'number' && tokens < limit) return true - if (!st.windowWarned) { - st.windowWarned = true - $.ui.log(`model-router: no switch to ${id}: context not known to fit`) - } - return false -} - +/** + * The main loop's plan: effort from the floor/sticky/turn decision, model + * from `decideMain`. `cur` is the model the router moved this turn, else + * the engine's own (verbatim: an engine fallback is respected). + */ async function mainPlan($: Api, st: State, e: StepIn): Promise { const { effort } = mainEffort(st, e.effort) - const wanted = mainModel(st) - if (wanted === undefined || !st.cfg.mainModelSwitch) { - return { model: e.model, effort } - } - const id = resolveModel(st.cfg, wanted) - return { model: (await windowOk($, st, id)) ? id : e.model, effort } + await prune($, st) + const cur = st.turnModel ?? e.model + const { call } = await decideFor($, st, cur) + logCall($, st, call) + if (call.moved) st.turnModel = call.model + return { model: call.model, effort } } async function planStep($: Api, st: State, e: StepIn): Promise { @@ -908,7 +1482,8 @@ function stepLog(e: StepIn, plan: Plan): string { } function noteMain($: Api, st: State, plan: Plan): void { - st.lastMain = `${plan.model.replace(/^claude-/, '')}/${plan.effort ?? '-'}` + st.lastPlan = plan + st.spinner = `${plan.model.replace(/^claude-/, '')}/${plan.effort ?? '-'}` refresh($, st) } @@ -916,18 +1491,88 @@ function endMainTurn($: Api, st: State): void { st.turnFloor = st.pendingPrompt st.pendingPrompt = null st.turnMain = null + st.pushed = null + st.turnModel = undefined st.typedSlash = false st.explicitEffort.clear() - st.lastMain = '' - st.windowWarned = false + st.spinner = '' + st.turnLogged.clear() refresh($, st) } +// ---- breaker inputs -------------------------------------------------- + +/** The exact id a failure is charged to: the agent's, else main's plan. */ +function failureTarget(st: State, agentId: string | undefined) { + if (agentId !== undefined) return st.agentModels.get(agentId) + return st.lastPlan?.model.replace(ONE_M, '') +} + +/** An engine StopFailure of an availability kind marks its model down. */ +async function onStopFailure( + $: Api, + st: State, + e: { error: string; agent_id?: string }, +): Promise { + if (st.off || !UNAVAILABLE.has(e.error)) return + const sent = failureTarget(st, e.agent_id) + if (sent === undefined) return + if (e.error === 'model_not_found' && aliasOf(st.cfg, sent) === undefined) { + $.ui.log(`model-router: model_not_found for ${sent}: not a table id; ` + + 'no mark') + return + } + await prune($, st) + markDown($, st, canonical(st.cfg, sent), e.error, await $.clock.now()) +} + +type SwitchIn = { + from_model: string + to_model: string + requested_model: string | null + source: string +} + +/** + * The engine switched the model by itself. The model it left is marked, one + * strike, unless it landed where the router already was; the mark targets + * the model actually sent, else the model reported as left. + */ +async function onAutoSwitch($: Api, st: State, e: SwitchIn): Promise { + const cfg = st.cfg + const from = canonical(cfg, e.from_model) + const to = canonical(cfg, e.to_model) + $.ui.log(`model-router: engine switched ${from} → ${to} ` + + `(requested ${e.requested_model ?? '-'})`) + const sent = st.lastPlan ? canonical(cfg, st.lastPlan.model) : undefined + if (to === sent) return + const target = sent !== undefined && aliasOf(cfg, sent) ? sent : from + await prune($, st) + markDown($, st, target, 'engine fallback', await $.clock.now()) +} + +/** Any model change: keeps `sessionModel`; a user choice clears its mark. */ +async function onModelSwitch($: Api, st: State, e: SwitchIn): Promise { + st.sessionModel = e.to_model + if (st.off || e.source === 'resume') return + if (e.source === 'auto') await onAutoSwitch($, st, e) + else clearBreaker(st, canonical(st.cfg, e.to_model)) +} + // ---- registration ---------------------------------------------------- +async function readSessionModel($: Api): Promise { + try { + return await $.session.model() + } catch { + return '' + } +} + function registerSession(on: On, st: State): void { on('session.start', async ($, e, next) => { await reloadConfig($, st) + st.sessionModel = await readSessionModel($) await registerCommand($) refresh($, st) return next(e) @@ -936,7 +1581,7 @@ function registerSession(on: On, st: State): void { return next(e) }) on('session.end', async ($, e, next) => { - Object.assign(st, newState(st.cfg, st.source)) + resetSession(st) // session.start never fires after /clear: re-apply the config's // `enabled` so a config-disabled router (offConfig) stays off. applyEnabled(st, st.cfg.enabled) @@ -947,6 +1592,23 @@ function registerSession(on: On, st: State): void { }) } +function registerBreaker(on: On, st: State): void { + on('classic.StopFailure', async ($, e, next) => { + await onStopFailure($, st, e) + return next(e) + }).catch(($, e, next) => { + warnOnce(st, $, 'StopFailure', next.error.kind) + return next(e) + }) + on('classic.PostModelSwitch', async ($, e, next) => { + await onModelSwitch($, st, e) + return next(e) + }).catch(($, e, next) => { + warnOnce(st, $, 'PostModelSwitch', next.error.kind) + return next(e) + }) +} + function registerCommandHook(on: On, st: State): void { on('command.run', { command: 'route' }, async ($, e) => { if (e.origin.kind !== 'composer') { @@ -961,7 +1623,7 @@ function registerCommandHook(on: On, st: State): void { function registerRouteTool(on: On, st: State): void { on('tool.call', { tool: TOOL }, async ($, e) => { - const out = handleRouteTool(st, e) + const out = await handleRouteTool($, st, e) refresh($, st) vlog($, st, `route ${e.agentId ?? 'main'}: ${JSON.stringify(out)}`) return out @@ -1002,10 +1664,15 @@ function registerSkills(on: On, st: State): void { function registerAgents(on: On, st: State): void { on('tool.call', { tool: 'Agent' }, async ($, e, next) => { - if (!st.off && isLevel(e.effort) && typeof e.tool_use_id === 'string') { + if (st.off) return next(e) + if (isLevel(e.effort) && typeof e.tool_use_id === 'string') { st.explicitEffort.set(e.tool_use_id, e.effort) } - return next(e) + if (e.agentId !== undefined) return next(e) + pushOrchestrate(st) + const out = await next(e) + safely(st, $, 'Agent', () => afterDispatch(st, out.result)) + return out }).catch(($, e, next) => { warnOnce(st, $, 'Agent', next.error.kind) return next(e) @@ -1015,16 +1682,18 @@ function registerAgents(on: On, st: State): void { function registerSpawn(on: On, st: State): void { on('agent.spawn', async ($, e, next) => { if (st.off) return next(e) + await prune($, st) const frozen = e.fork || e.workflow !== undefined const route = spawnRoute(st.cfg, e, frozen) - // An explicit model param on the Agent call always wins. - const wanted = e.model === undefined ? route?.model : undefined + const wanted = spawnTarget(st, e, route) const started = await next(wanted === undefined ? e - : { ...e, model: resolveModel(st.cfg, wanted) }) + : { ...e, model: wanted }) if (typeof started.agentId === 'string') { + const agentId = started.agentId safely(st, $, 'agent.spawn', () => { trackLoop(st, e, started, route) + rememberAgent(st, agentId, started.model) vlog($, st, `spawn ${e.subagentType}: ${e.model ?? '-'} → ${started.model}`) }) @@ -1052,7 +1721,9 @@ function registerTurns(on: On, st: State): void { return yield* next(e) }) on('turn.complete', async ($, e, next) => { - if (e.agentId !== undefined) st.loops.delete(e.agentId) + const agentId = e.agentId + if (agentId !== undefined) safely(st, $, 'turn.complete', () => + endAgent(st, agentId)) else endMainTurn($, st) return next(e) }).catch(($, e, next) => { @@ -1070,26 +1741,49 @@ function floorFromPrompt(st: State, midTurn: boolean, routed: Routed): void { if (midTurn) st.pendingPrompt = higherFloor(st.pendingPrompt, routed) } +const firstRule = (st: State, mode: PromptMode, text: string) => + st.rules.find(r => r.mode === mode && r.re.test(text)) + +/** A default rule sets the idle turn's route; routes and skills override. */ +function defaultFromPrompt(st: State, scanned: string): void { + const rule = firstRule(st, 'default', scanned) + const route = rule ? phaseRoute(st.cfg, rule.phase) : undefined + if (rule && route) { + st.turnMain = { phase: rule.phase, route: { ...route }, source: 'prompt' } + } +} + +/** + * Floor rules first (the user's explicit minimum). Default rules run only + * on an idle, plain prompt: not mid-turn, not a slash command or skill + * (they route themselves), not one that already carries a floor. + */ +function routeFromPrompt(st: State, text: string, midTurn: boolean): void { + const scanned = text.slice(0, MAX_PROMPT_SCAN) + const floor = firstRule(st, 'floor', scanned) + const route = floor ? phaseRoute(st.cfg, floor.phase) : undefined + if (floor && route) { + const routed: Routed = { phase: floor.phase, route, source: 'prompt' } + floorFromPrompt(st, midTurn, routed) + } + const slash = text.trimStart().startsWith('/') + if (text.trimStart().startsWith('/effort-')) st.typedSlash = true + if (!midTurn && !slash && !floor) defaultFromPrompt(st, scanned) +} + function registerPrompt(on: On, st: State): void { on('prompt.submit', async ($, e, next) => { if (st.off || e.origin.kind !== 'composer') return next(e) - const scanned = e.text.slice(0, MAX_PROMPT_SCAN) - const rule = st.rules.find(r => r.re.test(scanned)) - const route = rule ? phaseRoute(st.cfg, rule.phase) : undefined - if (rule && route) { - const routed: Routed = { phase: rule.phase, route, source: 'prompt' } - floorFromPrompt(st, e.turnId !== undefined, routed) - refresh($, st) - } - if (e.text.trimStart().startsWith('/effort-')) st.typedSlash = true + routeFromPrompt(st, e.text, e.turnId !== undefined) + refresh($, st) return next(e) }).catch(($, e, next) => { warnOnce(st, $, 'prompt.submit', next.error.kind) return next(e) }) on('ui.render', { component: 'Spinner' }, async ($, e, next) => { - if (st.off || !st.cfg.spinner || !st.lastMain) return next(e) - const suffix = ` · ${st.lastMain}…` + if (st.off || !st.cfg.spinner || !st.spinner) return next(e) + const suffix = ` · ${st.spinner}…` return next({ ...e, props: { ...e.props, suffix } }) }).catch(($, e, next) => { warnOnce(st, $, 'ui.render', next.error.kind) @@ -1100,6 +1794,7 @@ function registerPrompt(on: On, st: State): void { export const register: Register = on => { const st = newState(mergeConfig(undefined, () => undefined), 'defaults') registerSession(on, st) + registerBreaker(on, st) registerCommandHook(on, st) registerRouteTool(on, st) registerSkills(on, st) From 6f31f49d7ceda8d9ea3dbbc50dfd33ed27682ef3 Mon Sep 17 00:00:00 2001 From: bchanot Date: Fri, 9 Oct 2026 15:09:07 +0200 Subject: [PATCH 21/22] =?UTF-8?q?chore(tasks):=20model-router=20W1-C=20don?= =?UTF-8?q?e=20=E2=80=94=20contract=20evidence,=20TODO=20(accepted=20MEDIU?= =?UTF-8?q?Ms,=20residuals,=20live=20checks),=20journal?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- .claude/memory/journal.md | 1 + .claude/tasks/TODO.md | 4 +++- .../2026-10-09-model-router-tiers-1237.md | 15 +++++++++++---- 3 files changed, 15 insertions(+), 5 deletions(-) diff --git a/.claude/memory/journal.md b/.claude/memory/journal.md index 3ccf88d..8de9176 100644 --- a/.claude/memory/journal.md +++ b/.claude/memory/journal.md @@ -585,3 +585,4 @@ rules: - model-router W1-B1 floor landed (1ff608a): plan r2 from 3 lenses (0 BLOCKER, 4 MAJOR: one decision helper, typed level = default + minimum, mid-turn prompt now + next, per-machine enabled:false), feater DONE, gap round (/clear lost enabled:false, 'ultrathink rule' label, per-axis effort base), hardening round (kill switch keeps previous cfg on failed reload, typed slash attested at prompt.submit vs sub-agent preload). 30 tests, verifier CONFORME 6/6, security PASS ×2 with parked residuals. B2 wiring dispatched next. - model-router W1-B2 landed (6430ac6): tracked symlink skills/model-router → ../mods/model-router (mode 120000), gitignore, lib/tests/mods.test.sh, doctor Mods section, CLAUDE.md § mods/. Plan r2 (3 lenses, 5 MAJOR), feater DONE, gap round (my `4b.` label unparsed by gates.sh + unbounded probe), verifier CONFORME 8/8, security PASS (LOW: `claude plugin --help` rc 0 weakens the SKIP probe, fail-closed). Dev hot-reload link removed from ~/.claude/dev-mods; user to run /reload-plugins. BDR-115 amended (load, floor, kill switch, hardening). Doc audit (opus) in flight. - model-router wave 1 CLOSED on feature/model-router-mod (15 commits ahead of develop, nothing pushed: manual mode). Docs: opus audit SIGNIFICANT (README Explore row false, no mention of the mod) → user go all 10 → first patch self-reverted by the MINOR-envelope oracle (plan carried MINOR labels; a new heading exceeds the envelope) → re-dispatched with SIGNIFICANT provenance → b22f894. Full `make test`: every suite green except the pre-existing env red design-tool-gate (21st CLI present, not hermetic). Open for the user: /reload-plugins here; merge decision (gitflow finish); settings.json own change; W2 migration queued. +- model-router W1-C adaptive tiers landed (d0fa100): plan r1→r4 through 3 lenses (all FATAL: 4 BLOCKER + 20 MAJOR) + 2 confirmations (1 BLOCKER each, in my own r2 then r3) → deviation from the one-confirmation cap, stated. Design: absolute tiers, StopFailure-kind breaker + PostModelSwitch auto, fallback chain, main upgrade under a 200k cap (fails closed), sticky turnModel, derived orchestrate (background dispatches), prompt default rules with skip rules; classifier deferred. feater DONE → 2 gap rounds (texts) → hardening (leaveDown gates, cap fail-closed, one-way prefix, log key) → verifier CONFORME, security PASS (LOW only). 58 tests. Lesson: I sent iteration history in a security brief; the auditor contract forbids it (blind scan) — scope only next time. Live checks pending after /reload-plugins (R16/T8). Next: user reload, live test, wave 2. diff --git a/.claude/tasks/TODO.md b/.claude/tasks/TODO.md index 321aaf4..76874d6 100644 --- a/.claude/tasks/TODO.md +++ b/.claude/tasks/TODO.md @@ -13,7 +13,9 @@ migration of shifters/pins/model-gate in wave 2 after proof; names model-router - [x] W1-B2 wiring (contract `2026-10-08-model-router-wiring-1835`, 2026-10-09): tracked symlink `skills/model-router` → `../mods/model-router` loads as `model-router@skills-dir` (fresh-process `claude plugin list --json` proves it), `.gitignore` `mods/*/tsconfig.json`, `lib/tests/mods.test.sh` (4 checks, capability probe, bounded, SKIP), doctor `── Mods ──` fail-soft, CLAUDE.md `## mods/`. Plan r2 from 3 lenses (5 MAJOR), feater DONE, gap round (4b label + unbounded probe), verifier CONFORME 8/8, security PASS - [ ] W1-B2 residuals (accepted): `claude plugin --help` returns 0 → the suite's capability probe can pass on a CLI without `plugin test` and then FAIL instead of SKIP (fail-closed; fix = grep the probe output for the test usage line); doctor `claude plugin list --json` unbounded; doctor `echo -e` helpers interpolate `$_mod` (tracked folder names only); fallback timeout guard orphans grandchildren (only without coreutils timeout); `update-all.sh` runs `claude plugin update` over `@skills-dir` → one recurring warn (needs an update-all edit) - [x] W1 close-out (2026-10-09): doc-sync opus audit SIGNIFICANT → user go all 10 → patched (README effort routing + Explore row + /route, USAGE, ARCHITECTURE mods/, CHANGELOG Unreleased) b22f894; hot-reload link removed from `~/.claude/dev-mods//`; BDR-115 amended (a6e2003). Pending user: `/reload-plugins` in this session (new sessions load the skills-dir copy by themselves); merge decision on feature/model-router-mod (`gitflow finish`, human signal) -- [ ] W1-C adaptive tiers (user 2026-10-09: "un système logique et optimisé", no /route typed, works on a haiku session, falls back when fable has no credit): contract `2026-10-09-model-router-tiers-1237`, plan same slug. Phases name absolute tiers (best/big/work/cheap), availability breaker (turn error/refusal + engine auto switch, 15 min cooldown), fallback chain fable→opus→sonnet→haiku, main upgrade on / downgrade gated, prompt default rules (plan/reflect keywords FR+EN) + dispatch push/pop + optional classifier +- [x] W1-C adaptive tiers (user 2026-10-09, contract `2026-10-09-model-router-tiers-1237`, plan r4 after 3 lenses + 2 confirmations: 6 BLOCKER + 29 MAJOR closed by named changes): phases name absolute tiers (best fable>opus>sonnet · big opus>fable>sonnet · work sonnet>opus · cheap haiku>sonnet); breaker fed by `classic.StopFailure` kinds rate_limit|overloaded|billing_error|model_not_found + `PostModelSwitch` auto (episode backoff 15→300 min, `/model` clears, `/clear` keeps, `/route reload` clears); fallback chain fable→opus→sonnet→haiku; main UPGRADE by default under `upgradeMaxTokens` 200k (fails closed on unknown usage), DOWNGRADE gated by `mainModelSwitch`; sticky `turnModel` per turn; derived `orchestrate` on background dispatches; prompt default rules (plan/reflect FR+EN, Unicode guards, skipped on `/…`, floor match, mid-turn). 58 tests; verifier CONFORME then 3 gap/hardening rounds; security PASS +- [ ] W1-C accepted-by-design (security 2026-10-09, MEDIUM, not coded around): (1) the model itself can raise the main loop to fable for the rest of a turn through the `route` tool (plan/reflect/escalate/judge) or a background dispatch (derived orchestrate = best tier) — bounded by the turn and `upgradeMaxTokens`, no sticky route is model-callable; (2) `ultrathink` and the default keyword rules now mean "best tier" (fable) at the phase's effort, so an incidental keyword in a pasted composer prompt costs a fable turn (origin composer only). Residuals: `PostModelSwitch` `auto` covers "other programmatic change" (a healthy model could be marked 15 min); strikes never decay inside a session; a `[1m]` variant's `model_not_found` marks the base model until reload; a `models` alias added by the override without a `fallback` key stays unranked (`withEveryAlias` not applied to the default chain) so `leaveDown` skips the upgrade gates for it; `canonical` prefix match has no segment boundary (`claude-sonnet-5-50` would map to sonnet); classifier deferred to W2 +- [ ] W1-C live verification (after `/reload-plugins`, R16/T8): StopFailure vs turn.complete order; `PostModelSwitch` on a rewritten-request fallback and its `from_model`; whether a router rewrite raises `auto`; `[1m]` carry validity on opus; `$.session.model()` string form; the derived orchestrate on a real background dispatch - [ ] W2 migration (after wave 1 proven in daily use): 15 skills `Skill(effort-*)` → `route` tool calls; remove `skills/effort-*`, `lib/effort-pins.txt/.sh`, install/update steps, `effort:` frontmatter on skills and agents; repo agents into the mod's `agents` table (verify the spawn/first-step ordering first); `lib/model-gate.md` + `lib/model-check.sh` → mod rule; census tests repointed; `lib/effort-shift.md` rewritten; docs - [ ] W2 migration: 15 skills off `Skill(effort-*)`, remove shifters + effort-pins + model-gate, census repointed, docs - [ ] W3 optional: step heuristics, haiku classifier, quota-aware downgrade, A/B diff --git a/.claude/tasks/contracts/2026-10-09-model-router-tiers-1237.md b/.claude/tasks/contracts/2026-10-09-model-router-tiers-1237.md index 8c55109..538ff43 100644 --- a/.claude/tasks/contracts/2026-10-09-model-router-tiers-1237.md +++ b/.claude/tasks/contracts/2026-10-09-model-router-tiers-1237.md @@ -22,20 +22,27 @@ Q (r2): superseded clauses / A: floor AC4 (`turnMain` sources gain 'derived' and 1. Suite green: `claude plugin test` passes with at least 43 `test(` calls; `claude plugin validate` passes with no warning; no line over 80 chars; no `any` type. CHECK: cd mods/model-router && out=$(claude plugin test . 2>&1); rc=$?; echo "$out" | tail -n 3; [ $rc -eq 0 ] && [ "$(grep -cE '^\s*test\(' hooks/register.test.ts)" -ge 43 ] && v=$(claude plugin validate . 2>&1) && echo "$v" | grep -q 'Validation passed' && ! echo "$v" | grep -qi 'warning' && ! grep -nE '.{81,}' hooks/register.ts hooks/register.test.ts && ! grep -nE ':\s*any\b||as any\b' hooks/register.ts && echo TIERS-SUITE-OK EXPECT: TIERS-SUITE-OK - EVIDENCE: pending + EVIDENCE: MET exit=0 marker-found :: 58 pass 0 fail Ran 58 tests across 1 file. [2.03s] TIERS-SUITE-OK 2. Type-check clean against this build's declarations. CHECK: T=/Users/b.chanot/Documents/claude/mods/model-router/.claude-plugin/types; W=$(mktemp -d) && printf '{"compilerOptions":{"target":"es2023","lib":["es2023"],"types":[],"module":"esnext","moduleResolution":"bundler","strict":true,"noUncheckedIndexedAccess":true,"noEmit":true,"skipLibCheck":true,"jsx":"react","jsxFactory":"h","jsxFragmentFactory":"Fragment"},"include":["%s/claude-code/index.d.ts","%s/claude-code-tools/index.d.ts","%s/hooks"]}' "$T" "$T" "$PWD/mods/model-router" > "$W/tsconfig.json" && (cd "$W" && npx --yes -p typescript@5 tsc -p tsconfig.json) && echo TSC-OK EXPECT: TSC-OK - EVIDENCE: pending + EVIDENCE: MET exit=0 marker-found :: TSC-OK 3. Tiers in the config: `tiers` (best/big/work/cheap), `fallback`, `cooldownMinutes`, `mainUpgrade`, `upgradeMaxTokens` exist in DEFAULT_CONFIG; every default phase names a tier, none a bare model; `/route show` prints the resolved model of each phase and a `down:` line; no `classifier` (deferred to wave 2); no `Loop.model` / `spawnModel` / `loop.model` identifier (W1-A AC8). CHECK: cd mods/model-router/hooks && D=$(awk '/^const DEFAULT_CONFIG/,/^}/' register.ts) && echo "$D" | grep -q "tiers:" && echo "$D" | grep -q "fallback:" && grep -q "cooldownMinutes" register.ts && grep -q "mainUpgrade" register.ts && grep -q "upgradeMaxTokens" register.ts && ! grep -q "classifier" register.ts && [ "$(awk '/^const DEFAULT_CONFIG/,/^}/' register.ts | grep -cE "tier: '(best|big|work|cheap)'")" -ge 10 ] && ! awk '/^const DEFAULT_CONFIG/,/^}/' register.ts | grep -qE "model: '(haiku|sonnet|opus|fable)'" && ! grep -qE "loop\.model|explicitModel|spawnModel" register.ts && grep -q "down:" register.ts && echo TIERS-CONFIG-OK EXPECT: TIERS-CONFIG-OK - EVIDENCE: pending + EVIDENCE: MET exit=0 marker-found :: TIERS-CONFIG-OK 4. Tests prove (names contain the quoted word; plan r2 R14 lists them): `tier` — a `plan` route on a session model `claude-haiku-4-5-20251001` makes the main step run on `claude-fable-5-1` at xhigh (upgrade, default on); `downgrade` — a `mechanical` route on a fable session leaves the model unchanged while `mainModelSwitch` is off; `fallback` — with a `plan` route, a `classic.StopFailure` `rate_limit` on main after a fable step makes the next main step run on `claude-opus-5-5` at xhigh, and after `/route reload` fable is used again; `breaker` — an `invalid_request` failure never marks a model down, backoff expiry restores it, a `/model` command (`PostModelSwitch` source `command`) clears it; `engine fallback` — a `PostModelSwitch` with source `auto` marks the model the engine left (one strike, idempotent within the hold) and a plan route does not go back to it; `unknown` — a session model absent from the table is never switched; `spawn` — `Explore` spawns on `claude-opus-5-5` while sonnet is down (agent StopFailure with `agent_id`); `derived` — a main Agent call sets `orchestrate` and the previous `plan` route is back when the spawned agent ends; a route declared after the dispatch is not overwritten by the pop; `default rule` — "planifie la migration" sets `plan` as the turn default and a later route call overrides it; a typed `/analyze …` and a `/effort-low pourquoi …` prompt get no default rule; `per axis` — a model-less sticky never hides a turn route's tier; `floor` tests from B1 still pass. CHECK: cd mods/model-router/hooks && for w in tier downgrade fallback breaker "engine fallback" unknown spawn derived "default rule" "per axis"; do grep -qE "test\('[^']*$w" register.test.ts || { echo "missing test: $w"; exit 1; }; done && [ "$(grep -cE "test\('[^']*(breaker|derived|default rule)" register.test.ts)" -ge 7 ] && echo TIERS-TESTS-OK EXPECT: TIERS-TESTS-OK - EVIDENCE: pending + EVIDENCE: MET exit=0 marker-found :: TIERS-TESTS-OK 5. Judged by reading (plan r2 R1-R15, r3 S1-S11 and r4 T1-T8 are binding; r4 wins over r3, r3 over r2 where they conflict: two fields `turnModel` (sticky, per turn) and `lastPlan` (breaker target, kept), unrouted steps pass `e.model` verbatim, auto marks target the model actually sent and skip when the engine landed where the router was, `sessionModel` preserved across /clear and never prefix-matched when empty, tokens `number | undefined` with windowOk failing closed, episode strikes with `model_not_found` lengthening a hold; no per-step engine-fallback detection, `PostModelSwitch` auto DOES mark one idempotent strike, the main model is sticky within a turn, `sessionModel` cached at start and on PostModelSwitch, D1 for background dispatches via the Agent result status, breaker targets kept until replaced): ONE resolver turns a tier name, an alias or a full id into an AVAILABLE canonical id (tier → list → skip down → id; an exhausted tier → the global chain; a bare alias or id passes through even when down); ids are canonical everywhere (`[1m]` stripped for comparison and carried on the replacement, alias → id, two-way prefix); ONE decision `decideMain(cur, wanted, ctx)` in the binding order (off → unknown cur: no switch → cur down: wanted or next available, windowOk → same alias: keep → better: `mainUpgrade` and `upgradeMaxTokens` → cheaper: `mainModelSwitch` and windowOk), used by `mainPlan` AND by every text; the breaker is fed only by `classic.StopFailure` errors `rate_limit | overloaded | billing_error | model_not_found` (main → the last main plan's model, agent → `agentModels`) and by `PostModelSwitch` `source: 'auto'` (the model the engine left, one strike, idempotent while down); `turn.complete` reasons never mark; a user `/model` (`PostModelSwitch` command|picker|sdk) clears the target's mark; backoff 15 → 30 → 60 → 120 → 300 min per id, `model_not_found` until reload; the breaker survives `/clear` and is cleared by `/route reload` before the config read; the derived `orchestrate` push/pop tracks this turn's spawns and never overwrites a route the model declared after the dispatch; default prompt rules (two passes, absent `mode` = floor, `iu` flags, Unicode guards) write `turnMain` (source 'prompt') and are skipped for a leading `/`, for a prompt carrying a floor or a typed slash, and mid-turn; floor rules write `turnFloor`; no classifier; explicit Agent params still win; agent model fixed at spawn; every B1/1-A criterion still holds EXCEPT the three clauses R15 names (turnMain sources, the switch clause, the window-guard scope); no function over 25 logic lines; truthful texts come from `decideMain` and say `upgrade`, `fallback`, `switch off` or `unchanged`. +Hardening round (security gate 2026-10-09, 3 MEDIUM + 1 LOW accepted) — criteria 6-7, same ledger: +6. (a) `leaveDown` runs a `wanted` that ranks ABOVE cur through the upgrade checks (`mainUpgrade`, `upgradeMaxTokens`, windowOk) before taking it; (b) the upgrade cap fails CLOSED: unknown tokens (`undefined`) → no upgrade, logged once per turn; the test boot answers `session.usage` with a small context so the upgrade tests still run, and one test proves an unanswered usage blocks the upgrade; (c) `canonical` keeps ONE prefix direction only (`bare.startsWith(tableId)`, which covers `[1m]` and dated variants) and `markDown` charges `model_not_found` to the exact id the request carried when that id is not itself a table id (no mark); (d) the once-per-turn log key for "upgrade skipped: context N tokens" is fixed (no N in the key). + CHECK: cd mods/model-router/hooks && ! grep -qE "known\.startsWith\(bare\)|tableId\.startsWith\(bare\)|id\.startsWith\(bare\)" register.ts && grep -qE "test\('[^']*(cap|unknown tokens|usage)" register.test.ts && grep -q "upgrade skipped" register.ts && echo HARDEN-C-OK + EXPECT: HARDEN-C-OK + EVIDENCE: MET exit=0 marker-found :: HARDEN-C-OK +7. Judged by reading: criteria 1-5 still hold (tests ≥ 55 + the new ones green, validate, tsc, style); the `PostModelSwitch` `auto` handling is unchanged and listed for live verification; the two accepted-by-design MEDIUMs (model-initiated upgrades within a turn; `ultrathink` → best tier at max) are recorded in TODO, not coded around. + ## FILE SCOPE mods/model-router/hooks/register.ts · mods/model-router/hooks/register.test.ts From a4f660d0b6adbb48e033b01f90e9dda6bc0055ad Mon Sep 17 00:00:00 2001 From: bchanot Date: Fri, 9 Oct 2026 15:14:55 +0200 Subject: [PATCH 22/22] chore(tasks): model-router W1-C live checks part 1 done, part 2 queued; journal --- .claude/memory/journal.md | 1 + .claude/tasks/TODO.md | 3 ++- 2 files changed, 3 insertions(+), 1 deletion(-) diff --git a/.claude/memory/journal.md b/.claude/memory/journal.md index 8de9176..51d8498 100644 --- a/.claude/memory/journal.md +++ b/.claude/memory/journal.md @@ -586,3 +586,4 @@ rules: - model-router W1-B2 landed (6430ac6): tracked symlink skills/model-router → ../mods/model-router (mode 120000), gitignore, lib/tests/mods.test.sh, doctor Mods section, CLAUDE.md § mods/. Plan r2 (3 lenses, 5 MAJOR), feater DONE, gap round (my `4b.` label unparsed by gates.sh + unbounded probe), verifier CONFORME 8/8, security PASS (LOW: `claude plugin --help` rc 0 weakens the SKIP probe, fail-closed). Dev hot-reload link removed from ~/.claude/dev-mods; user to run /reload-plugins. BDR-115 amended (load, floor, kill switch, hardening). Doc audit (opus) in flight. - model-router wave 1 CLOSED on feature/model-router-mod (15 commits ahead of develop, nothing pushed: manual mode). Docs: opus audit SIGNIFICANT (README Explore row false, no mention of the mod) → user go all 10 → first patch self-reverted by the MINOR-envelope oracle (plan carried MINOR labels; a new heading exceeds the envelope) → re-dispatched with SIGNIFICANT provenance → b22f894. Full `make test`: every suite green except the pre-existing env red design-tool-gate (21st CLI present, not hermetic). Open for the user: /reload-plugins here; merge decision (gitflow finish); settings.json own change; W2 migration queued. - model-router W1-C adaptive tiers landed (d0fa100): plan r1→r4 through 3 lenses (all FATAL: 4 BLOCKER + 20 MAJOR) + 2 confirmations (1 BLOCKER each, in my own r2 then r3) → deviation from the one-confirmation cap, stated. Design: absolute tiers, StopFailure-kind breaker + PostModelSwitch auto, fallback chain, main upgrade under a 200k cap (fails closed), sticky turnModel, derived orchestrate (background dispatches), prompt default rules with skip rules; classifier deferred. feater DONE → 2 gap rounds (texts) → hardening (leaveDown gates, cap fail-closed, one-way prefix, log key) → verifier CONFORME, security PASS (LOW only). 58 tests. Lesson: I sent iteration history in a security brief; the auditor contract forbids it (blind scan) — scope only next time. Live checks pending after /reload-plugins (R16/T8). Next: user reload, live test, wave 2. +- W1-C live checks after the user's /reload-plugins (skills-dir copy, 12 hooks): derived orchestrate real (main high→medium during a background Explore→high after), Explore sonnet/medium, route plan → xhigh, Skill(effort-low) bridge → low, both confirmed in engine records; /route show resolves 10 phases to full ids, down none; steps carry bare ids (no [1m]) → suffix carry inert here. Breaker/auto/StopFailure order wait for a real incident. Branch feature/model-router-mod: 22 commits ahead, unpushed (manual), merge = human signal. Wave 2 queued. diff --git a/.claude/tasks/TODO.md b/.claude/tasks/TODO.md index 76874d6..94b4b08 100644 --- a/.claude/tasks/TODO.md +++ b/.claude/tasks/TODO.md @@ -15,7 +15,8 @@ migration of shifters/pins/model-gate in wave 2 after proof; names model-router - [x] W1 close-out (2026-10-09): doc-sync opus audit SIGNIFICANT → user go all 10 → patched (README effort routing + Explore row + /route, USAGE, ARCHITECTURE mods/, CHANGELOG Unreleased) b22f894; hot-reload link removed from `~/.claude/dev-mods//`; BDR-115 amended (a6e2003). Pending user: `/reload-plugins` in this session (new sessions load the skills-dir copy by themselves); merge decision on feature/model-router-mod (`gitflow finish`, human signal) - [x] W1-C adaptive tiers (user 2026-10-09, contract `2026-10-09-model-router-tiers-1237`, plan r4 after 3 lenses + 2 confirmations: 6 BLOCKER + 29 MAJOR closed by named changes): phases name absolute tiers (best fable>opus>sonnet · big opus>fable>sonnet · work sonnet>opus · cheap haiku>sonnet); breaker fed by `classic.StopFailure` kinds rate_limit|overloaded|billing_error|model_not_found + `PostModelSwitch` auto (episode backoff 15→300 min, `/model` clears, `/clear` keeps, `/route reload` clears); fallback chain fable→opus→sonnet→haiku; main UPGRADE by default under `upgradeMaxTokens` 200k (fails closed on unknown usage), DOWNGRADE gated by `mainModelSwitch`; sticky `turnModel` per turn; derived `orchestrate` on background dispatches; prompt default rules (plan/reflect FR+EN, Unicode guards, skipped on `/…`, floor match, mid-turn). 58 tests; verifier CONFORME then 3 gap/hardening rounds; security PASS - [ ] W1-C accepted-by-design (security 2026-10-09, MEDIUM, not coded around): (1) the model itself can raise the main loop to fable for the rest of a turn through the `route` tool (plan/reflect/escalate/judge) or a background dispatch (derived orchestrate = best tier) — bounded by the turn and `upgradeMaxTokens`, no sticky route is model-callable; (2) `ultrathink` and the default keyword rules now mean "best tier" (fable) at the phase's effort, so an incidental keyword in a pasted composer prompt costs a fable turn (origin composer only). Residuals: `PostModelSwitch` `auto` covers "other programmatic change" (a healthy model could be marked 15 min); strikes never decay inside a session; a `[1m]` variant's `model_not_found` marks the base model until reload; a `models` alias added by the override without a `fallback` key stays unranked (`withEveryAlias` not applied to the default chain) so `leaveDown` skips the upgrade gates for it; `canonical` prefix match has no segment boundary (`claude-sonnet-5-50` would map to sonnet); classifier deferred to W2 -- [ ] W1-C live verification (after `/reload-plugins`, R16/T8): StopFailure vs turn.complete order; `PostModelSwitch` on a rewritten-request fallback and its `from_model`; whether a router rewrite raises `auto`; `[1m]` carry validity on opus; `$.session.model()` string form; the derived orchestrate on a real background dispatch +- [x] W1-C live verification part 1 (2026-10-09 after `/reload-plugins`, skills-dir copy loaded, 12 hooks): derived orchestrate on a real background Explore (main high → medium while it ran → high after its end); Explore on sonnet/medium; `route plan` → next request xhigh (engine record); `Skill(effort-low)` bridged → next request low (engine record); `/route show` resolves all 10 phases to full ids, `down: none`; steps arrive as bare `claude-fable-5-1` (no `[1m]`) so the suffix carry never fires on this session +- [ ] W1-C live verification part 2 (needs a real incident): StopFailure vs turn.complete order; `PostModelSwitch` `auto` semantics and its `from_model` after a router upgrade; whether a router rewrite raises `auto`; `[1m]` carry validity on opus (only on a session whose steps carry it) - [ ] W2 migration (after wave 1 proven in daily use): 15 skills `Skill(effort-*)` → `route` tool calls; remove `skills/effort-*`, `lib/effort-pins.txt/.sh`, install/update steps, `effort:` frontmatter on skills and agents; repo agents into the mod's `agents` table (verify the spawn/first-step ordering first); `lib/model-gate.md` + `lib/model-check.sh` → mod rule; census tests repointed; `lib/effort-shift.md` rewritten; docs - [ ] W2 migration: 15 skills off `Skill(effort-*)`, remove shifters + effort-pins + model-gate, census repointed, docs - [ ] W3 optional: step heuristics, haiku classifier, quota-aware downgrade, A/B