diff --git a/.claude/memory/journal.md b/.claude/memory/journal.md index 0af5f38..c8e2ae8 100644 --- a/.claude/memory/journal.md +++ b/.claude/memory/journal.md @@ -582,3 +582,4 @@ rules: - model-router wave 1-A (/feat, user go): mod built in `mods/model-router/` (4 files, 834 + 202 lines, 11 plugin tests). Plan r1 → r3: 3 challengers (simplicity CONCERNS, robustness CONCERNS(6), correctness FATAL(8)) + 1 confirmation CONCERNS(4); converged on: agents table = built-ins only in wave 1 (frontmatter stays single writer), agent model written once at spawn, explicit Agent params frozen per loop, every sub-agent write on its own loop, Skill bridge answers in the Skill tool's OUTPUT schema (string result refused → skill would load), config validated before merge, state in closure, `/route off`. feater DONE first pass; GATE 0 MET; verifier CONFORME 6/6; security PASS (4 MEDIUM + 5 LOW parked in TODO for user go). Commits b721c94 (mod) + e8ca713 (contract/plan). Override `~/.claude/model-router.json` {verbose:true} written (pass B). Doc-sync deferred to 1-B. - model-router W1-A hardening (user go): fresh feater on contract criteria 7-11 → gap round (dead `Loop.frozen`, spawn returns `started` verbatim, tool description) → verifier CONFORME 11/11 → security PASS (1 MEDIUM residual: ReDoS size-bounded only, self-inflicted config; 5 LOW parked). Registries BDR-115, LRN-205, LRN-206, EVAL-040 written on user go (64702d5). Next: live swap of the real mod into the hot-reload folder, then W1-B install + docs. - model-router live checks on Opus 5.5 (user /model, uncommitted settings.json change left to the user): Skill(effort-low) bridge answered in place → next request low; ultrathink turn ran max (engine base medium); Explore without params → claude-sonnet-5-5, 3 steps medium. Loading switched to tracked symlink skills/model-router → @skills-dir (PLUGIN_DIRS non-portable: absolute path, no $HOME expansion, tracked settings); isolated-HOME probe listed/enabled/loaded. User chose floor semantics for ultrathink + typed /effort-. W1-B split: B1 floor (register.ts) + B2 wiring (symlink, gitignore, mods suite, doctor, CLAUDE.md); 6 challengers in flight. Effort shifters skipped this turn: the bridge would overwrite the user's ultrathink until B1 lands. +- model-router W1-B1 floor landed (1ff608a): plan r2 from 3 lenses (0 BLOCKER, 4 MAJOR: one decision helper, typed level = default + minimum, mid-turn prompt now + next, per-machine enabled:false), feater DONE, gap round (/clear lost enabled:false, 'ultrathink rule' label, per-axis effort base), hardening round (kill switch keeps previous cfg on failed reload, typed slash attested at prompt.submit vs sub-agent preload). 30 tests, verifier CONFORME 6/6, security PASS ×2 with parked residuals. B2 wiring dispatched next. diff --git a/.claude/tasks/TODO.md b/.claude/tasks/TODO.md index d1b2620..4e8ac1e 100644 --- a/.claude/tasks/TODO.md +++ b/.claude/tasks/TODO.md @@ -8,7 +8,8 @@ migration of shifters/pins/model-gate in wave 2 after proof; names model-router - [x] W1-A the mod in `mods/model-router/` (b721c94, contract `2026-10-08-model-router-w1a-1533`, plan r3): challenge 3 lenses + 1 confirmation (2 BLOCKER + 10 MAJOR closed by named changes), feater DONE first pass, GATE 0 MET 5/5, verifier CONFORME 6/6, security PASS (4 MEDIUM + 5 LOW reported, below) - [x] W1-A hardening (user go "oui durcis", contract criteria 7-11): `/route` composer-only; no agent model axis (effort only after spawn); pattern ≤ 200 / scan ≤ 4096 / phase keys `^[a-z][a-z0-9_-]{0,31}$` / config ≤ 64 KB / `additionalProperties: false` / `typeof e.skill`; `warnOnce` in all 14 catches + config-drop logs; `safely` around post-`next` bookkeeping. feater DONE, gap round (dead `Loop.frozen`, spawn returns `started` verbatim, tool description), GATE 0 MET 9/9, verifier CONFORME 11/11, security PASS. - [ ] W1-A residuals (security, accepted, none exploitable from outside the user's own files): ReDoS is size-bounded only (`(a+)+$` in `~/.claude/model-router.json` + a 4 KB paste hangs the hook; fix = nested-quantifier rejection or a far lower scan cap); `stat().size` trusted (FIFO/device path in ~/.claude); unrestricted `models`/`agents`/`skills` KEYS echoed raw in logs (log flood); `agentId` read from the flat tool event (engine strip unverified); 5 closures without a kit test (no fs in the kit, LRN-206); "nothing routed" catch text after a state write. -- [ ] W1-B1 floor precedence (contract `2026-10-08-model-router-floor-1835`): `ultrathink` + typed `/effort-` = floor for the main turn (user choice 2026-10-08) +- [x] W1-B1 floor precedence (contract `2026-10-08-model-router-floor-1835`, 2026-10-09): `ultrathink` + typed `/effort-` = the main turn's default AND minimum (r2 after 3 lenses: a pure floor made /effort-low a no-op); one helper `mainEffort`; per-axis precedence; mid-turn prompt floors now + next turn; `"enabled": false` per machine, kept across /clear and failed reload; typed slash attested. 30 tests; verifier CONFORME 6/6 (after 1 gap round); security PASS ×2 +- [ ] W1-B1 residuals (security 2026-10-09, accepted): first-load failure of the override falls to defaults (`enabled: true`); non-boolean `enabled` drops to the default on reload; `skill.prompt` preload guard is a heuristic (`loops.size > 0`; a preload during spawn or a non-composer `/effort-*` while idle still writes the floor); marker not bound to a valid level; `/route reload` answers "config reloaded" even when the previous config was kept; transient missing override lifts a config-set off; `String(err)` of a JSON parse in the local log. Display: `/route show` folds the floor into the effort while off; model-axis text with a model-only sticky and the switch on. - [ ] W1-B2 wiring (contract `2026-10-08-model-router-wiring-1835`): tracked symlink `skills/model-router` → `../mods/model-router` (loads as `@skills-dir`; PLUGIN_DIRS dropped: absolute paths in tracked settings), `.gitignore` engine-laid files, `lib/tests/mods.test.sh`, doctor `── Mods ──`, CLAUDE.md `## mods/`; then doc-sync (README/USAGE/CHANGELOG), live swap (remove the hot-reload link, `/reload-plugins`), BDR-115 amendment - [ ] W2 migration: 15 skills off `Skill(effort-*)`, remove shifters + effort-pins + model-gate, census repointed, docs - [ ] W3 optional: step heuristics, haiku classifier, quota-aware downgrade, A/B diff --git a/.claude/tasks/contracts/2026-10-08-model-router-floor-1835.md b/.claude/tasks/contracts/2026-10-08-model-router-floor-1835.md index 30943eb..450452d 100644 --- a/.claude/tasks/contracts/2026-10-08-model-router-floor-1835.md +++ b/.claude/tasks/contracts/2026-10-08-model-router-floor-1835.md @@ -21,16 +21,26 @@ Q: who clears the floor / A: main turn end (a queued prompt's floor is then prom 1. Suite green with the new tests: `claude plugin test` passes with at least 22 `test(` calls; `claude plugin validate` passes with no warning; no line over 80 chars; no `any` type. CHECK: cd mods/model-router && out=$(claude plugin test . 2>&1); rc=$?; echo "$out" | tail -n 3; [ $rc -eq 0 ] && [ "$(grep -cE '^\s*test\(' hooks/register.test.ts)" -ge 22 ] && v=$(claude plugin validate . 2>&1) && echo "$v" | grep -q 'Validation passed' && ! echo "$v" | grep -qi 'warning' && ! grep -nE '.{81,}' hooks/register.ts hooks/register.test.ts && ! grep -nE ':\s*any\b||as any\b' hooks/register.ts && echo FLOOR-SUITE-OK EXPECT: FLOOR-SUITE-OK - EVIDENCE: pending + EVIDENCE: MET exit=0 marker-found :: 30 pass 0 fail Ran 30 tests across 1 file. [1.05s] FLOOR-SUITE-OK 2. Type-check clean against this build's declarations. CHECK: T=/Users/b.chanot/.claude/dev-mods/385f7190-70f5-4bdd-b0d8-e4566cd412fd/model-router/.claude-plugin/types; [ -d "$T" ] || T=/Users/b.chanot/Documents/claude/mods/model-router/.claude-plugin/types; W=$(mktemp -d) && printf '{"compilerOptions":{"target":"es2023","lib":["es2023"],"types":[],"module":"esnext","moduleResolution":"bundler","strict":true,"noUncheckedIndexedAccess":true,"noEmit":true,"skipLibCheck":true,"jsx":"react","jsxFactory":"h","jsxFragmentFactory":"Fragment"},"include":["%s/claude-code/index.d.ts","%s/claude-code-tools/index.d.ts","%s/hooks"]}' "$T" "$T" "$PWD/mods/model-router" > "$W/tsconfig.json" && (cd "$W" && npx --yes -p typescript@5 tsc -p tsconfig.json) && echo TSC-OK EXPECT: TSC-OK - EVIDENCE: pending + EVIDENCE: MET exit=0 marker-found :: TSC-OK 3. The floor is its own slot: `turnFloor` is declared in `State`, initialised in `newState`, written by the prompt rule and by a typed `/effort-`, cleared by `/route clear` and at main turn end; the suite carries at least 8 tests whose name contains `floor`; one helper `mainEffort` decides the main effort; the config key `enabled` exists. CHECK: cd mods/model-router/hooks && [ "$(grep -c 'turnFloor' register.ts)" -ge 6 ] && [ "$(grep -cE "^\s*test\('[^']*floor" register.test.ts)" -ge 8 ] && grep -q 'mainEffort' register.ts && grep -q 'enabled' register.ts && echo FLOOR-SLOT-OK EXPECT: FLOOR-SLOT-OK - EVIDENCE: pending -4. Judged by reading: effective main effort comes from ONE helper `mainEffort` used by `mainPlan` and by every answer text: base `(userMain ?? turnMain)?.route.effort ?? turnFloor?.route.effort ?? e.effort`, then floored by `turnFloor` (LEVELS order; a numeric or absent value is replaced); the floor never applies to a sub-agent step; `turnMain` only ever holds 'model' or 'skill' sources; the prompt rule writes `turnFloor` (keeping the higher of two) and, when typed mid-turn (`turnId` set, `wait` ignored), also `pendingPrompt`, promoted into `turnFloor` at main turn end; `"enabled": false` in the override file makes every hook pass through after each config load (`/route on` re-enables for the session); a non-effort skill load resets `turnMain` only; a model `route({clear})` clears `turnMain` only; every answer that the floor overrides says so truthfully (Skill bridge context, route tool text, `/effort-` text); `/route show` and the status line display the floor; every criterion of `.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md` still holds; no function over 25 logic lines. + EVIDENCE: MET exit=0 marker-found :: FLOOR-SLOT-OK +4. Judged by reading: effective main effort comes from ONE helper `mainEffort` used by `mainPlan` and by every answer text: base `userMain?.route.effort ?? turnMain?.route.effort ?? turnFloor?.route.effort ?? e.effort` (per-axis, like the model rule: a model-only sticky never hides a turn route's effort; gap round 2026-10-09), then floored by `turnFloor` (LEVELS order; a numeric or absent value is replaced); the floor never applies to a sub-agent step; `turnMain` only ever holds 'model' or 'skill' sources; the prompt rule writes `turnFloor` (keeping the higher of two) and, when typed mid-turn (`turnId` set, `wait` ignored), also `pendingPrompt`, promoted into `turnFloor` at main turn end; `"enabled": false` in the override file makes every hook pass through after each config load AND survives `/clear` (`session.end` rebuilds the state but re-applies the config's `enabled`; `/route on` re-enables for the session); a prompt-rule floor is labelled by its matched phase (`prompt rule `), never by a fixed word; a non-effort skill load resets `turnMain` only; a model `route({clear})` clears `turnMain` only; every answer that the floor overrides says so truthfully (Skill bridge context, route tool text, `/effort-` text); `/route show` and the status line display the floor; every criterion of `.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md` still holds; no function over 25 logic lines. + +Hardening round (security gate 2026-10-09, 2 MEDIUM) — criteria 5-6, same ledger: +5. The kill switch fails closed: a failed override read on `/route reload` (unreadable, oversized, invalid JSON) keeps the PREVIOUS config (and therefore the previous `enabled`) instead of falling back to the defaults; a non-boolean `enabled` value is dropped WITH a log line; at session start with no previous config the defaults still apply. + CHECK: cd mods/model-router && grep -q "previous" hooks/register.ts && grep -qE "enabled.*(not a boolean|non-boolean|ignored)" hooks/register.ts && grep -qE "test\('[^']*(reload|previous|kill)" hooks/register.test.ts && echo KILL-CLOSED-OK + EXPECT: KILL-CLOSED-OK + EVIDENCE: MET exit=0 marker-found :: KILL-CLOSED-OK +6. `skill.prompt` writes the floor only for a typed `/effort-`: a one-shot marker set at `prompt.submit` (composer origin, text starting with `/effort-`) attests the typing; without the marker the write is refused while any sub-agent loop is live (a preload fires inside an agent's life), and accepted otherwise (no agent can be preloading); the refused case returns the text unchanged with a one-line note. Tests: preload simulation (spawned agent live, no marker → no floor), typed with marker → floor, typed with no marker and no agent → floor. + CHECK: cd mods/model-router && grep -q "slashMarker\|typedSlash" hooks/register.ts && [ "$(grep -cE "test\('[^']*(preload|marker|typed)" hooks/register.test.ts)" -ge 2 ] && echo SLASH-ATTEST-OK + EXPECT: SLASH-ATTEST-OK + EVIDENCE: MET exit=0 marker-found :: SLASH-ATTEST-OK ## FILE SCOPE mods/model-router/hooks/register.ts · mods/model-router/hooks/register.test.ts