Compare commits
31
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
65dff0e768 | ||
|
|
bb56f3e41e | ||
|
|
862740da9d | ||
|
|
5e0e5c0bc4 | ||
|
|
ca9645833c | ||
|
|
b09e84497a | ||
|
|
efdd491d63 | ||
|
|
ff741e3a82 | ||
|
|
abbdf7926d | ||
|
|
a4f660d0b6 | ||
|
|
6f31f49d7c | ||
|
|
d0fa1001bb | ||
|
|
977be7cad8 | ||
|
|
140c16a67f | ||
|
|
2d8cd6bf4c | ||
|
|
e79db7e6df | ||
|
|
b22f8947f9 | ||
|
|
a6e200392c | ||
|
|
3c44dd00d3 | ||
|
|
6430ac65ec | ||
|
|
24e180ade0 | ||
|
|
1ff608a68c | ||
|
|
868a7f0515 | ||
|
|
77ad7cf494 | ||
|
|
ae0179f491 | ||
|
|
346d6aeab2 | ||
|
|
64702d50ea | ||
|
|
6dc2d748fc | ||
|
|
e8ca713d9e | ||
|
|
b721c94dcb | ||
|
|
b73d1b127e |
@@ -48,6 +48,7 @@ rules:
|
|||||||
| BLK-026 | 2026-10-06 | `make test` red on macOS: 13 suites, GNU-only idioms in suite + 7 libs (SIGPIPE under pipefail, `sed -i`, `wc` padding, `stat -c`, `realpath -m`, bare `timeout`, `grep -oP`) | resolved |
|
| BLK-026 | 2026-10-06 | `make test` red on macOS: 13 suites, GNU-only idioms in suite + 7 libs (SIGPIPE under pipefail, `sed -i`, `wc` padding, `stat -c`, `realpath -m`, bare `timeout`, `grep -oP`) | resolved |
|
||||||
| BLK-027 | 2026-10-06 | this machine never ran `make link`/`make plugin`: no global `core.hooksPath` → post-commit push never fired, branches landed ahead of upstream; 11 vendored skills + `~/.claude/.env` missing | resolved (link) / open (plugin) |
|
| BLK-027 | 2026-10-06 | this machine never ran `make link`/`make plugin`: no global `core.hooksPath` → post-commit push never fired, branches landed ahead of upstream; 11 vendored skills + `~/.claude/.env` missing | resolved (link) / open (plugin) |
|
||||||
| BLK-028 | 2026-10-06 | notify-attention on a VS Code client: bell + toast silent-degradation faults (merge of BLK-019 + BLK-020): terminalBell sound default off, ext hooks only terminals born after activation, Code muted in Windows mixer | resolved |
|
| BLK-028 | 2026-10-06 | notify-attention on a VS Code client: bell + toast silent-degradation faults (merge of BLK-019 + BLK-020): terminalBell sound default off, ext hooks only terminals born after activation, Code muted in Windows mixer | resolved |
|
||||||
|
| BLK-029 | 2026-10-08 | Claude Code mods (2.1.294): alias → id resolution for a model set by a hook lags the Agent tool's (`sonnet` → `claude-sonnet-5`, 404); Agent tool schema refuses full ids | upstream |
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -305,3 +306,11 @@ rules:
|
|||||||
- **Probe order (do FIRST, before server archaeology)**: fresh VS Code terminal, `printf '\a\a\033]777;notify;Test;hello\033\\'` → splits terminal path from client renderer; palette `Help: List Signal Sounds` → Terminal Bell preview bypasses terminal/BEL/hook/dtach/ext, isolates renderer audio in one step.
|
- **Probe order (do FIRST, before server archaeology)**: fresh VS Code terminal, `printf '\a\a\033]777;notify;Test;hello\033\\'` → splits terminal path from client renderer; palette `Help: List Signal Sounds` → Terminal Bell preview bypasses terminal/BEL/hook/dtach/ext, isolates renderer audio in one step.
|
||||||
- **Status**: resolved (BLK-019 2026-09-01, BLK-020 A+B 2026-09-02/03). Sources superseded by this entry; bodies kept for history.
|
- **Status**: resolved (BLK-019 2026-09-01, BLK-020 A+B 2026-09-02/03). Sources superseded by this entry; bodies kept for history.
|
||||||
- **Reference**: `~/.claude/hooks/notify-attention.sh` header documents the setting; [[LRN-145]] terminalSequence-not-/dev/tty; silent-degradation class [[LRN-047]]; sources [[BLK-019]], [[BLK-020]].
|
- **Reference**: `~/.claude/hooks/notify-attention.sh` header documents the setting; [[LRN-145]] terminalSequence-not-/dev/tty; silent-degradation class [[LRN-047]]; sources [[BLK-019]], [[BLK-020]].
|
||||||
|
|
||||||
|
## BLK-029 — Mods: model alias set by a hook resolves to a stale id (`sonnet` → `claude-sonnet-5`, 404) — 2026-10-08
|
||||||
|
- **Friction**: model-router spike. Sub-agent routed by `agent.spawn` or `tool.call Agent` rewrite with alias `sonnet`/`haiku` → "model_not_found HTTP 404, model sent to the API: claude-sonnet-5". Same alias typed by the model in the Agent tool param → `claude-sonnet-5-5`, OK.
|
||||||
|
- **Real cause**: two alias tables in the CLI (2.1.294): the Agent tool's is current, the function-hooks path's is stale. Not an access issue (`/model` lists all four tiers; explicit haiku/sonnet dispatches answered).
|
||||||
|
- **Solution**: hooks write full ids (`claude-sonnet-5-5`, `claude-haiku-4-5-20251001`, `claude-opus-5-5`, `claude-fable-5-1`) from the mod's own table; Agent tool param rewrite stays alias-only (schema enum) → the mod sets the model at `agent.spawn`, not at the param.
|
||||||
|
- **Status**: upstream (report to anthropics/claude-code with the request id `req_011Cfppp7VFt8z2Pd3zJrpUi`); workaround in model-router.
|
||||||
|
- **Reference**: [[LRN-203]], plan `.claude/tasks/plans/2026-10-08-model-router-mod.md`.
|
||||||
|
|
||||||
|
|||||||
@@ -132,6 +132,7 @@ rules:
|
|||||||
| BDR-108 | 2026-09-29 | Effort round: level on every skill next to its model pin (3 repo + 25 vendored via `lib/effort-pins.txt` re-applied after the LAST vendoring step of install + resync), design stack ONE level (high), model pins stay tier aliases: quality/price trade-off = tier × effort, never version | accepted |
|
| BDR-108 | 2026-09-29 | Effort round: level on every skill next to its model pin (3 repo + 25 vendored via `lib/effort-pins.txt` re-applied after the LAST vendoring step of install + resync), design stack ONE level (high), model pins stay tier aliases: quality/price trade-off = tier × effort, never version | accepted |
|
||||||
| BDR-109 | 2026-09-30 | Higgsfield pack: npm CLI `latest` + 8 upstream skills git-cloned into gitignored `skills-external/higgsfield-*`, OFF by default, in no profile; two toggles (`higgsfield` = allowlist of 7 media skills, `higgsfield-websites` = landing-page aid, never website create/deploy/publish); CLI presence by probe; routing on explicit ask | accepted |
|
| BDR-109 | 2026-09-30 | Higgsfield pack: npm CLI `latest` + 8 upstream skills git-cloned into gitignored `skills-external/higgsfield-*`, OFF by default, in no profile; two toggles (`higgsfield` = allowlist of 7 media skills, `higgsfield-websites` = landing-page aid, never website create/deploy/publish); CLI presence by probe; routing on explicit ask | accepted |
|
||||||
| BDR-110 | 2026-10-06 | Shell portability doctrine: native userland on macOS AND Linux, no Homebrew GNU tools on PATH; `lib/tests/portability-census.test.sh` locks deterministic GNU-only idioms | accepted |
|
| BDR-110 | 2026-10-06 | Shell portability doctrine: native userland on macOS AND Linux, no Homebrew GNU tools on PATH; `lib/tests/portability-census.test.sh` locks deterministic GNU-only idioms | accepted |
|
||||||
|
| BDR-115 | 2026-10-08 | model-router mod: pin = entry default, sub-tasks route finer; one writer per axis; full ids from the mod table; built-ins-only agents table until frontmatter pins go; state in closure | accepted |
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -1393,3 +1394,12 @@ Branch feature/user-writing-web-rules, UNMERGED (human gate).
|
|||||||
- **Alternatives rejected**: regen via `install-hook` (writes a LOCAL hooks-path entry) or `global-hooks` (writes the GLOBAL config when the value is missing — it was, [[LRN-200]]) → `emit-hook > file` only; temp file for the verb's stderr in a hook (fail-open on a full TMPDIR, predictable path) → `2>&1` capture ([[LRN-202]]); classification before the token cap (13 s flood) → cap first ([[LRN-201]]); shell check of the release version string (interpolation sink) → by reading.
|
- **Alternatives rejected**: regen via `install-hook` (writes a LOCAL hooks-path entry) or `global-hooks` (writes the GLOBAL config when the value is missing — it was, [[LRN-200]]) → `emit-hook > file` only; temp file for the verb's stderr in a hook (fail-open on a full TMPDIR, predictable path) → `2>&1` capture ([[LRN-202]]); classification before the token cap (13 s flood) → cap first ([[LRN-201]]); shell check of the release version string (interpolation sink) → by reading.
|
||||||
- **Gates**: D1 3 lenses + confirm FATAL(1) (global-config write) → emit-hook; feater; GATE 0 MET; verifier CONFORME; security PASS. D2 3 lenses + confirm FATAL(3) (escape alternative, same-quote-inside, `bare=$one` when unparsed); feater; GATE 0 MET; verifier CONFORME; security BLOCK(1) cap-after-fork → fixed (20k tokens 0.13 s) → CONFORME + PASS. D3 3 lenses + confirm CONCERNS(3); feater; CONFORME + PASS.
|
- **Gates**: D1 3 lenses + confirm FATAL(1) (global-config write) → emit-hook; feater; GATE 0 MET; verifier CONFORME; security PASS. D2 3 lenses + confirm FATAL(3) (escape alternative, same-quote-inside, `bare=$one` when unparsed); feater; GATE 0 MET; verifier CONFORME; security BLOCK(1) cap-after-fork → fixed (20k tokens 0.13 s) → CONFORME + PASS. D3 3 lenses + confirm CONCERNS(3); feater; CONFORME + PASS.
|
||||||
- **Refs**: contracts/plans `2026-10-07-manual-push-failclosed-d1-1522`, `…-guard-residuals-d2-1526`, `…-prose-d3-1530`; commits 472cccb, 3c59333, 64ca0f8 (feature/manual-push-mode, UNMERGED). Supersedes the "fail-open on invalid value" line of [[BDR-111]]. Links [[BDR-112]], [[BDR-113]], [[LRN-114]], [[LRN-196]]. Residuals: TODO "post-run-D residuals" (soft_deny names only `false`; stale `.githooks/` in onboarded repos until reconcile).
|
- **Refs**: contracts/plans `2026-10-07-manual-push-failclosed-d1-1522`, `…-guard-residuals-d2-1526`, `…-prose-d3-1530`; commits 472cccb, 3c59333, 64ca0f8 (feature/manual-push-mode, UNMERGED). Supersedes the "fail-open on invalid value" line of [[BDR-111]]. Links [[BDR-112]], [[BDR-113]], [[LRN-114]], [[LRN-196]]. Residuals: TODO "post-run-D residuals" (soft_deny names only `false`; stale `.githooks/` in onboarded repos until reconcile).
|
||||||
|
|
||||||
|
## BDR-115 — model-router mod: pin = entry default, sub-tasks route finer; one writer per axis; full ids; built-ins-only table until pins go [accepted] (2026-10-08)
|
||||||
|
- **Decision**: one function-hooks mod (`mods/model-router/`) routes model + effort per request. Rules: (1) a skill/agent pin = DEFAULT route of the run, never ceiling/floor; inside the run every sub-task routes to its phase, declared (`route` tool, `/route`, `ultrathink`) or derived (skill load, agent dispatch). (2) ONE writer per axis: `Skill(effort-*)` answered by the mod WITHOUT loading the skill (no pairing rule, no doublon); every write lands on the CALLING loop; explicit Agent params frozen per loop for the whole run; agent model written once at spawn. (3) hooks write FULL ids from `config.models` ([[LRN-203]]). (4) wave 1 agents table = built-ins only (Explore sonnet/medium, Plan opus/xhigh); repo agents keep frontmatter as single writer until wave 2 moves pins into the table and deletes frontmatter + shifters + effort-pins + model-gate. (5) precedence: user `/route` sticky > latest turn route (one slot, last writer wins; non-effort skill load resets it except a prompt rule) > session. (6) main-loop model switch behind `mainModelSwitch` (default off) + context-window guard ([[LRN-204]]). (7) state in `register` closure, defaults cloned, `/route off` kill switch, config validated before merge. Load: `CLAUDE_CODE_PLUGIN_DIRS` in settings env via link.sh (wave 1-B).
|
||||||
|
- **Why**: user 2026-10-08: existing pins + `effort-*` shifters built for this goal with older tools; some pins forced (one level for a skill doing many things). Mods expose `turn.step` model/effort rewrite + `agent.spawn` + `tool.call` answers = cleaner, unified, automatic, works headless too (BDR-107 gap).
|
||||||
|
- **Alternatives rejected**: agents table copying the 21 frontmatter pins in wave 1 (third source of truth, silent override of a frontmatter edit); `scope: agents` / `/route agents` bulk lever (no requirement, sixth precedence tier); `model` param on the model-facing tool (typo → LRN-203 404 class); per-step agent model rewrite (fights engine fallback, beat explicit params); `userConfig` (one config file instead); marketplace install (live symlink repo model); `source`-dependent skill-load reset (kept stale shifts).
|
||||||
|
- **Gates**: plan r1 → r3 through 3 challengers + 1 confirmation (2 BLOCKER + 10 MAJOR closed, [[EVAL-040]]); feater DONE first pass; GATE 0 MET; verifier CONFORME 6/6; security PASS + hardening round (criteria 7-11).
|
||||||
|
- **Refs**: plan `.claude/tasks/plans/2026-10-08-model-router-mod.md`, contract `.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md`, [[BDR-107]], [[BDR-108]], [[BLK-029]], [[LRN-205]], [[LRN-206]].
|
||||||
|
- **Amendment (2026-10-09, user decisions 2026-10-08 evening)**: (a) LOAD supersedes the "Load:" line: tracked relative symlink `skills/model-router` → `../mods/model-router`, loaded in place as `model-router@skills-dir` wherever link.sh links `~/.claude/skills`; `CLAUDE_CODE_PLUGIN_DIRS` dropped (absolute path, settings `env` has no `$HOME` expansion, settings.json tracked), local marketplace dropped (`add` writes an absolute path into settings.json). Proven by fresh-process `claude plugin list --json`. (b) PRECEDENCE amended: `ultrathink` and a typed `/effort-<l>` are the main turn's DEFAULT and MINIMUM (floor slot `turnFloor`): sticky `/route` effort > turn route effort > floor > engine, then floored; per axis; mid-turn prompt floors the running turn and the next (`wait` ignored). Rationale: user "un choix explicite bat la phase déduite"; a pure floor made `/effort-low` a no-op (challenge finding). (c) Per-machine kill switch `"enabled": false` in the untracked `~/.claude/model-router.json` (survives `/clear`, a failed reload keeps the previous config); `enabledPlugins` would dirty the tracked settings.json on every machine. (d) Hardening: `/route` composer-only; agent loops effort-only (model fixed at spawn); config caps; typed slash attested at `prompt.submit`. Commits 346d6ae, 1ff608a, 6430ac6; contracts `2026-10-08-model-router-floor-1835`, `2026-10-08-model-router-wiring-1835`; residuals parked in TODO.
|
||||||
|
|
||||||
|
|||||||
@@ -60,6 +60,8 @@ rules:
|
|||||||
| EVAL-037 | 2026-09-28 | correction of EVAL-035/036 counts: transcript records are per content block; deduped by message.id → main-loop thinking share 99.9%, thinking share of weighted cost 5.6%, sonnet think/msg 26→0.2, A/B requests 9→8 | conclusions hold (sharper: main-loop thinking 96.6%→99.9%, weighted-cost thinking corrected 8.4%→5.6%); effort-audit.py dedupes from a3b479e+ |
|
| EVAL-037 | 2026-09-28 | correction of EVAL-035/036 counts: transcript records are per content block; deduped by message.id → main-loop thinking share 99.9%, thinking share of weighted cost 5.6%, sonnet think/msg 26→0.2, A/B requests 9→8 | conclusions hold (sharper: main-loop thinking 96.6%→99.9%, weighted-cost thinking corrected 8.4%→5.6%); effort-audit.py dedupes from a3b479e+ |
|
||||||
| EVAL-038 | 2026-09-29 | correction of EVAL-037: 94 % of sub-agent usage records carry no `output_tokens_details` (Fable subs at xhigh read 0 thinking, impossible with always-on thinking) → sub-agent thinking UNMEASURED, not ≈0; main loop 100 % counted; weighted-cost split (61/39) still holds | `effort-audit.py` prints coverage + CAVEAT; cite the cost split only; agent effort pins stay unmeasured; a tier move on a price argument = judgment, not figure |
|
| EVAL-038 | 2026-09-29 | correction of EVAL-037: 94 % of sub-agent usage records carry no `output_tokens_details` (Fable subs at xhigh read 0 thinking, impossible with always-on thinking) → sub-agent thinking UNMEASURED, not ≈0; main loop 100 % counted; weighted-cost split (61/39) still holds | `effort-audit.py` prints coverage + CAVEAT; cite the cost split only; agent effort pins stay unmeasured; a tier move on a price argument = judgment, not figure |
|
||||||
| EVAL-039 | 2026-09-30 | ship-feature run higgsfield-pack: plan dry-run in scratch → 0 executor failure on 7 tasks; challenge found 7 MAJOR I missed; floor-guard caught 2 shellcheck suppressions of mine; final review found README/code gap | keep |
|
| EVAL-039 | 2026-09-30 | ship-feature run higgsfield-pack: plan dry-run in scratch → 0 executor failure on 7 tasks; challenge found 7 MAJOR I missed; floor-guard caught 2 shellcheck suppressions of mine; final review found README/code gap | keep |
|
||||||
|
| EVAL-040 | 2026-10-08 | model-router w1a plan: 3 challengers + 1 confirmation found 2 BLOCKER + 14 MAJOR on a plan judged closed; executor then passed every gate first time | keep the round, never dispatch a mod plan without it |
|
||||||
|
| EVAL-041 | 2026-10-09 | model-router W1-C plan: 3 lenses FATAL (4 BLOCKER + 20 MAJOR) then 2 confirmations each FATAL with a NEW BLOCKER in my own revision; executor DONE first pass, 3 short text/hardening rounds | one confirmation is not enough when a revision removes a whole mechanism; the plan carried the risk, the code almost none |
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -370,3 +372,18 @@ Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itse
|
|||||||
- **Method**: plan code dry-run in a scratch copy before the gate (suite per stage 0/5→5/0, 6/8→14/0, 14/1→15/0, 15/1→16/0, 4 mutation tests); 3 challengers + 1 confirmation; SDD per-task reviews; GATE 0/1/2 twice; final review on opus.
|
- **Method**: plan code dry-run in a scratch copy before the gate (suite per stage 0/5→5/0, 6/8→14/0, 14/1→15/0, 15/1→16/0, 4 mutation tests); 3 challengers + 1 confirmation; SDD per-task reviews; GATE 0/1/2 twice; final review on opus.
|
||||||
- **Anomaly**: my first plan was green in dry-run and still wrong on 7 MAJOR points (shim vs binary, unbounded toggle probe, denylist membership, vacuous fixtures, askpass prompt): a dry-run proves the code does what I wrote, not that I wrote the right thing. Floor-guard flagged 2 `shellcheck disable=SC2016` I added to keep "shellcheck clean" green. Final review found the README promised drift reporting that the enabled state never reached. doc-syncer patch hit a shape escalation because I filed a script-comment edit under MINOR doc. One oracle of mine was shape-bound ([[LRN-188]]).
|
- **Anomaly**: my first plan was green in dry-run and still wrong on 7 MAJOR points (shim vs binary, unbounded toggle probe, denylist membership, vacuous fixtures, askpass prompt): a dry-run proves the code does what I wrote, not that I wrote the right thing. Floor-guard flagged 2 `shellcheck disable=SC2016` I added to keep "shellcheck clean" green. Final review found the README promised drift reporting that the enabled state never reached. doc-syncer patch hit a shape escalation because I filed a script-comment edit under MINOR doc. One oracle of mine was shape-bound ([[LRN-188]]).
|
||||||
- **Action**: keep the pre-gate dry-run (0 executor failure, 1 fix round in 7 tasks) AND the challenge (orthogonal finds); never silence a linter to satisfy a criterion, rewrite the line; doc patch plans carry public-doc paths only, script comments go as code commits.
|
- **Action**: keep the pre-gate dry-run (0 executor failure, 1 fix round in 7 tasks) AND the challenge (orthogonal finds); never silence a linter to satisfy a criterion, rewrite the line; doc patch plans carry public-doc paths only, script comments go as code commits.
|
||||||
|
|
||||||
|
## EVAL-040 — model-router w1a: the challenge round caught what the author could not see
|
||||||
|
- **Date**: 2026-10-08
|
||||||
|
- **Output checked**: plan `.claude/tasks/plans/2026-10-08-model-router-w1a-1533.md` r1, written after a successful spike with every harness fact in hand.
|
||||||
|
- **Method**: 3 blind opus challengers (simplicity CONCERNS(3), robustness CONCERNS(6), correctness FATAL(8)) + 1 confirmation (CONCERNS(4)); every BLOCKER/MAJOR closed by a named plan change (r2, r3); then feater, GATE 0, verifier, security.
|
||||||
|
- **Anomaly**: r1 carried 2 BLOCKER (explicit Agent `model` overridden at every step; Skill bridge answer shape refused by the output schema → doublon kept) + 10 MAJOR, all invisible to me: spike levers carried over as design (`agentsNext`, per-step model rewrite), a table copying 21 pins = third source of truth, writes on main from sub-agent loops. Confirmation found 4 more MAJOR (explicit params vs in-agent writes, `e.wait`, vacuous test assertions, haiku effort). Executor then DONE first pass, verifier CONFORME 6/6, security PASS: the plan was the whole risk.
|
||||||
|
- **Action**: a mod plan always goes through the full round + confirmation; test assertions must read the one line that carries the value; spike code is a FACT source, never a design source ([[BDR-115]], [[LRN-205]], [[LRN-206]]).
|
||||||
|
|
||||||
|
## EVAL-041 — model-router W1-C: the plan was the whole risk, two confirmations were needed
|
||||||
|
- **Date**: 2026-10-09
|
||||||
|
- **Output checked**: plan `.claude/tasks/plans/2026-10-09-model-router-tiers-1237.md` r1 → r4 (absolute tiers, breaker, derived phases), written with every engine fact in hand.
|
||||||
|
- **Method**: 3 blind opus challengers (simplicity FATAL(6), correctness FATAL(11), robustness FATAL(11)) → r2; confirmation FATAL(8) with a NEW BLOCKER introduced by r2 (per-step engine-fallback detection climbing the backoff) → r3; second confirmation FATAL(4) with a NEW BLOCKER introduced by r3 (`lastPlan` reset vs kept) → r4; executor DONE first pass; verifier ECARTS ×2 on texts (3 gaps) + hardening (4 items) → CONFORME; security PASS.
|
||||||
|
- **Anomaly**: 6 BLOCKER + 29 MAJOR over four revisions, each confirmation found a flaw my own fix had introduced; the doctrine cap (one confirmation) would have shipped r2 with a 5-hour false outage. The executor never needed a re-dispatch for logic: all later rounds were text truthfulness and hardening.
|
||||||
|
- **Action**: when a revision REMOVES or REPLACES a mechanism, re-challenge once more (state the deviation); keep plan sections additive with an explicit precedence line (r4 > r3 > r2) so executors and verifiers read one law; name superseded clauses of prior contracts in the Disposition. Links [[EVAL-040]], [[LRN-207]], [[BDR-115]].
|
||||||
|
|
||||||
|
|||||||
@@ -577,3 +577,17 @@ rules:
|
|||||||
- Merge (user go "tu peux merge dans develop"): final full suite green (46 suites minus the declared env red) + Health Stack shellcheck clean on the branch tip → `gitflow finish feature manual-push-mode` → develop 669db06, pushed, branch removed local + origin. 19 commits (runs A, B, C1/C2, D1/D2/D3 + docs + memory). User answered: only pushes change; commits/branches/local merges untouched; invalid value now fail-closed everywhere. User plan: dotfiles installer prompts for `gitflow.autopush` (default false) — told them the gitconfig template also needs `core.hooksPath` (the install wiped it). Open: user probe `! git push --dry-run` under autopush=false; AC6 env red (design-tool-gate); post-run-D residuals in TODO.
|
- Merge (user go "tu peux merge dans develop"): final full suite green (46 suites minus the declared env red) + Health Stack shellcheck clean on the branch tip → `gitflow finish feature manual-push-mode` → develop 669db06, pushed, branch removed local + origin. 19 commits (runs A, B, C1/C2, D1/D2/D3 + docs + memory). User answered: only pushes change; commits/branches/local merges untouched; invalid value now fail-closed everywhere. User plan: dotfiles installer prompts for `gitflow.autopush` (default false) — told them the gitconfig template also needs `core.hooksPath` (the install wiped it). Open: user probe `! git push --dry-run` under autopush=false; AC6 env red (design-tool-gate); post-run-D residuals in TODO.
|
||||||
- User tested manual-push mode on their machine: works (no auto push under `false`, bang-prefixed dry-run passes). Prompt handed over for the dotfiles repo: gitconfig template gets `core.hooksPath = ~/.claude/githooks` + `[gitflow] autopush = @AUTOPUSH@` rendered from an install question (default false, true/false only, unrendered placeholder = render failure). BDR-112 amended.
|
- User tested manual-push mode on their machine: works (no auto push under `false`, bang-prefixed dry-run passes). Prompt handed over for the dotfiles repo: gitconfig template gets `core.hooksPath = ~/.claude/githooks` + `[gitflow] autopush = @AUTOPUSH@` rendered from an install question (default false, true/false only, unrendered placeholder = render failure). BDR-112 amended.
|
||||||
|
|
||||||
|
## 2026-10-08
|
||||||
|
- model-router mod, wave 0 spike (user ask: one mod routes model + effort per request, replaces effort-* shifters + pins). Plan `.claude/tasks/plans/2026-10-08-model-router-mod.md`, 4 decisions by AskUserQuestion (spike-first main-loop switch, `CLAUDE_CODE_PLUGIN_DIRS` load, migration wave 2, names model-router / route / /route), rule "pin = entry default, sub-tasks route finer". Spike in dev-mods, hot reload on: `turn.step` effort rewrite proven (transcript `effort` field is the oracle, not `CLAUDE_EFFORT`); sub-agent model at `agent.spawn` + effort per step by agentId proven; main-loop fable → sonnet-5-5/low for 3 steps then back: works, one cold-cache step per switch INTO a model, return free. Found: hook-side alias resolver stale (`sonnet` → `claude-sonnet-5`, 404; Agent tool enum resolves the same alias to 5-5) → mod writes full ids only. feature/model-router-mod open, nothing committed yet (plan + TODO + journal pending).
|
||||||
|
- model-router wave 1-A (/feat, user go): mod built in `mods/model-router/` (4 files, 834 + 202 lines, 11 plugin tests). Plan r1 → r3: 3 challengers (simplicity CONCERNS, robustness CONCERNS(6), correctness FATAL(8)) + 1 confirmation CONCERNS(4); converged on: agents table = built-ins only in wave 1 (frontmatter stays single writer), agent model written once at spawn, explicit Agent params frozen per loop, every sub-agent write on its own loop, Skill bridge answers in the Skill tool's OUTPUT schema (string result refused → skill would load), config validated before merge, state in closure, `/route off`. feater DONE first pass; GATE 0 MET; verifier CONFORME 6/6; security PASS (4 MEDIUM + 5 LOW parked in TODO for user go). Commits b721c94 (mod) + e8ca713 (contract/plan). Override `~/.claude/model-router.json` {verbose:true} written (pass B). Doc-sync deferred to 1-B.
|
||||||
|
- model-router W1-A hardening (user go): fresh feater on contract criteria 7-11 → gap round (dead `Loop.frozen`, spawn returns `started` verbatim, tool description) → verifier CONFORME 11/11 → security PASS (1 MEDIUM residual: ReDoS size-bounded only, self-inflicted config; 5 LOW parked). Registries BDR-115, LRN-205, LRN-206, EVAL-040 written on user go (64702d5). Next: live swap of the real mod into the hot-reload folder, then W1-B install + docs.
|
||||||
|
- model-router live checks on Opus 5.5 (user /model, uncommitted settings.json change left to the user): Skill(effort-low) bridge answered in place → next request low; ultrathink turn ran max (engine base medium); Explore without params → claude-sonnet-5-5, 3 steps medium. Loading switched to tracked symlink skills/model-router → @skills-dir (PLUGIN_DIRS non-portable: absolute path, no $HOME expansion, tracked settings); isolated-HOME probe listed/enabled/loaded. User chose floor semantics for ultrathink + typed /effort-<l>. W1-B split: B1 floor (register.ts) + B2 wiring (symlink, gitignore, mods suite, doctor, CLAUDE.md); 6 challengers in flight. Effort shifters skipped this turn: the bridge would overwrite the user's ultrathink until B1 lands.
|
||||||
|
- model-router W1-B1 floor landed (1ff608a): plan r2 from 3 lenses (0 BLOCKER, 4 MAJOR: one decision helper, typed level = default + minimum, mid-turn prompt now + next, per-machine enabled:false), feater DONE, gap round (/clear lost enabled:false, 'ultrathink rule' label, per-axis effort base), hardening round (kill switch keeps previous cfg on failed reload, typed slash attested at prompt.submit vs sub-agent preload). 30 tests, verifier CONFORME 6/6, security PASS ×2 with parked residuals. B2 wiring dispatched next.
|
||||||
|
- model-router W1-B2 landed (6430ac6): tracked symlink skills/model-router → ../mods/model-router (mode 120000), gitignore, lib/tests/mods.test.sh, doctor Mods section, CLAUDE.md § mods/. Plan r2 (3 lenses, 5 MAJOR), feater DONE, gap round (my `4b.` label unparsed by gates.sh + unbounded probe), verifier CONFORME 8/8, security PASS (LOW: `claude plugin <unknown> --help` rc 0 weakens the SKIP probe, fail-closed). Dev hot-reload link removed from ~/.claude/dev-mods; user to run /reload-plugins. BDR-115 amended (load, floor, kill switch, hardening). Doc audit (opus) in flight.
|
||||||
|
- model-router wave 1 CLOSED on feature/model-router-mod (15 commits ahead of develop, nothing pushed: manual mode). Docs: opus audit SIGNIFICANT (README Explore row false, no mention of the mod) → user go all 10 → first patch self-reverted by the MINOR-envelope oracle (plan carried MINOR labels; a new heading exceeds the envelope) → re-dispatched with SIGNIFICANT provenance → b22f894. Full `make test`: every suite green except the pre-existing env red design-tool-gate (21st CLI present, not hermetic). Open for the user: /reload-plugins here; merge decision (gitflow finish); settings.json own change; W2 migration queued.
|
||||||
|
- model-router W1-C adaptive tiers landed (d0fa100): plan r1→r4 through 3 lenses (all FATAL: 4 BLOCKER + 20 MAJOR) + 2 confirmations (1 BLOCKER each, in my own r2 then r3) → deviation from the one-confirmation cap, stated. Design: absolute tiers, StopFailure-kind breaker + PostModelSwitch auto, fallback chain, main upgrade under a 200k cap (fails closed), sticky turnModel, derived orchestrate (background dispatches), prompt default rules with skip rules; classifier deferred. feater DONE → 2 gap rounds (texts) → hardening (leaveDown gates, cap fail-closed, one-way prefix, log key) → verifier CONFORME, security PASS (LOW only). 58 tests. Lesson: I sent iteration history in a security brief; the auditor contract forbids it (blind scan) — scope only next time. Live checks pending after /reload-plugins (R16/T8). Next: user reload, live test, wave 2.
|
||||||
|
- W1-C live checks after the user's /reload-plugins (skills-dir copy, 12 hooks): derived orchestrate real (main high→medium during a background Explore→high after), Explore sonnet/medium, route plan → xhigh, Skill(effort-low) bridge → low, both confirmed in engine records; /route show resolves 10 phases to full ids, down none; steps carry bare ids (no [1m]) → suffix carry inert here. Breaker/auto/StopFailure order wait for a real incident. Branch feature/model-router-mod: 22 commits ahead, unpushed (manual), merge = human signal. Wave 2 queued.
|
||||||
|
- User go 'ok merge': full make test green (except env red design-tool-gate), shellcheck clean → gitflow finish feature model-router-mod → develop abbdf79 (22 commits: waves 0, 1-A, 1-B1, 1-B2, 1-C + docs + registries). Manual mode: develop NOT pushed, the user publishes by hand from the terminal. Branch removed locally. User will /clear before wave 2.
|
||||||
|
- /hotfix make-test-names-red-suites (user: '9 min pour un merge?'): measured from transcript timestamps, the merge took <20 s; 454 s went to make test run TWICE (full + per-suite sweep to name the red suite, because the aggregate rc is silent). Fix: Makefile test prints FAIL <suite> + summary, rc unchanged (GNU make returns 2 on a failed recipe; my first oracle expected 1). hotfixer DONE, GATE 0 MET 2/2, security PASS. Branch bugfix/make-test-names-red-suites UNMERGED (human signal). Method note: a mods/-only diff needs only the mods suite + doctrine census, the full run once before merge.
|
||||||
|
- User go 'all, merge le, et je clear': LRN-207/208/209 + EVAL-041 written (ca96458); gitflow finish bugfix/make-test-names-red-suites → develop 5e0e5c0, branch removed. develop 28 commits ahead of origin, NOT pushed (manual mode, user publishes). Only settings.json dirty (user's own /model change). Session closes; next: wave 2 of model-router (TODO W2 line).
|
||||||
|
- model-router W2-A landed (bb56f3e, feature/model-router-w2). User go 'lance la vague 2' + 4 pass-B answers: shifters deleted, pins reworked into phase rows (not copied), phases declared, slim gate. Plan r1→r4: 3 lenses (1 BLOCKER: mod off → agents inherit parent model → frontmatter kept as off-state floor), confirmation 1 FATAL (BLOCKER: route calls wiped the sticky slot → separate runMain), confirmation 2 CONCERNS (deviation: 2 confirmations, stated). Feater DONE + 4 gap rounds (A5 in-agent decision, 6+3+3 coverage tests, every one mutation-proven). GATE 0 MET; verifier ECARTS(3)/(1)/(1) all on coverage of criterion 3 clauses, never code → diagnosis at max: compound criterion → user accepted at the cap. Security PASS (2 MEDIUM pre-existing). Kit 58→88. Next: user /reload-plugins + live probe, then W2-B. Lesson for the next contract: one coverage clause per criterion, not twelve.
|
||||||
|
|||||||
@@ -1748,3 +1748,32 @@ Rule: when editing a doctrine file under structure locks, grep the test's lock s
|
|||||||
## LRN-202 — Reading a stderr-then-stdout verb from a hook: `out=$(cmd 2>&1)`, last line = word, prefix line = reason; no temp file; lib path absolute before any cd
|
## LRN-202 — Reading a stderr-then-stdout verb from a hook: `out=$(cmd 2>&1)`, last line = word, prefix line = reason; no temp file; lib path absolute before any cd
|
||||||
- **Context**: unpushed-guard plan used `2>"${TMPDIR:-/tmp}/x.$$"` + cat + rm: fail-OPEN when TMPDIR is full (`|| mode=auto`), predictable path, symlink-followable on shared /tmp, leaked on kill; `mode=$(cmd 2>&1 >/dev/null)` captures ONLY stderr. push-guard's `mktemp` variant added `set -u` trap hazards. The verb writes its stderr line BEFORE its stdout word in one process, so `${out##*$'\n'}` is the word and the `gitflow.sh push-mode:` line is the reason (select by prefix, not `head -1`: a bash startup warning could precede it). Resolve the lib path to an absolute one BEFORE the hook's `cd "$cwd"` (a relative invocation otherwise resolves into the target repo).
|
- **Context**: unpushed-guard plan used `2>"${TMPDIR:-/tmp}/x.$$"` + cat + rm: fail-OPEN when TMPDIR is full (`|| mode=auto`), predictable path, symlink-followable on shared /tmp, leaked on kill; `mode=$(cmd 2>&1 >/dev/null)` captures ONLY stderr. push-guard's `mktemp` variant added `set -u` trap hazards. The verb writes its stderr line BEFORE its stdout word in one process, so `${out##*$'\n'}` is the word and the `gitflow.sh push-mode:` line is the reason (select by prefix, not `head -1`: a bash startup warning could precede it). Resolve the lib path to an absolute one BEFORE the hook's `cd "$cwd"` (a relative invocation otherwise resolves into the target repo).
|
||||||
- **Apply**: hooks never touch temp files for a one-line capture; anything but the expected word is treated as the fail-closed state, never as the default. Links [[BDR-114]], [[LRN-196]], [[LRN-199]].
|
- **Apply**: hooks never touch temp files for a one-line capture; anything but the expected word is treated as the fail-closed state, never as the default. Links [[BDR-114]], [[LRN-196]], [[LRN-199]].
|
||||||
|
|
||||||
|
## LRN-203 — Mod hooks: model ALIAS set by a hook resolves through stale table (`sonnet` → `claude-sonnet-5`, 404); write full ids; oracle = transcript fields, not `CLAUDE_EFFORT`
|
||||||
|
- **Context**: model-router spike 2026-10-08, CLI 2.1.294. `agent.spawn` or `tool.call Agent` param rewrite with alias `sonnet` → API got `claude-sonnet-5`, HTTP 404 model_not_found. Same alias passed by the model in the Agent tool param → `claude-sonnet-5-5`, fine. Full id `claude-sonnet-5-5` from hook → fine, every step answered by 5-5. Agent tool schema enum refuses full ids, so full ids reach API only via hooks. Effort rewrite at `turn.step` proven by transcript record field `effort` (high → medium); `CLAUDE_EFFORT` env + `perTurnEffort` stay at turn setting, blind to per-request rewrite.
|
||||||
|
- **Apply**: any mod that sets a model carries its own alias → full-id table (one place to bump per tier release). Verify routing with transcript `message.model` + record `effort`, never env vars. Links [[BDR-108]] (aliases as pins: still right at the Agent-tool call site, wrong inside hooks), [[BLK-029]].
|
||||||
|
|
||||||
|
## LRN-204 — Main-loop model switch mid-turn works, costs one cold-cache step on the full context per switch INTO a model; return free (per-model cache, 1 h TTL)
|
||||||
|
- **Context**: spike 2026-10-08, fable → `claude-sonnet-5-5`/low for 3 steps on ~260k context, then back. Conversation intact (tools, results, thinking blocks from another model in history: no error). First sonnet step cache_read 0 (full 260k billed), next steps 237k cached; return to fable step read 263k cached.
|
||||||
|
- **Apply**: switch the main loop only for spans long enough to amortize one uncached read of the whole context (many mechanical steps), never per tool call; short mechanical work → small-context haiku sub-agent. Haiku 4.5 window 200k: a long main loop cannot go to haiku at all. Flag off by default in model-router. Links [[LRN-203]], [[BDR-107]].
|
||||||
|
|
||||||
|
## LRN-205 — A hook answering `tool.call` in place of a built-in tool must return that tool's OUTPUT schema shape; a string result is refused and the tool runs anyway
|
||||||
|
- **Context**: model-router Skill bridge, plan r1: `{ result: '<string>' }` for `Skill(effort-*)`. Challenger: Skill has output schema `{ success, commandName, status?, … }` (claude-code-tools index.d.ts ~5139); core validates a hook's answer against it (claude-code index.d.ts ~12641), wrong shape = hook skipped = skill loads = the exact doublon the bridge exists to remove. Fix: `{ result: { success: true, commandName: e.skill, status: 'inline' }, context: ['…'] }`; text for the model goes in `context`, never in `result`.
|
||||||
|
- **Apply**: before answering any `tool.call` without `next`, grep the tool's RESULT type in claude-code-tools and mirror it; put model-facing prose in `context`. Links [[BDR-115]], [[LRN-203]].
|
||||||
|
|
||||||
|
## LRN-206 — `claude plugin test` kit facts (2.1.294): nothing fires at load, inputs are the FULL event, a bottom hook is mandatory under every `next`, `turn.step` streams
|
||||||
|
- **Context**: model-router tests. Kit `$` is `EngineCall<E> = (e: Args<E>)`: `command.run` needs `origin` + `presentation`, `prompt.submit` needs `wait` + `origin`, `agent.spawn` needs `tool_use_id, description, provider, parentModel, background, fork`; the test file is type-checked with the hooks (tsc include). `session.start` does NOT fire at load → every test boots with a bottom `on('session.start')` + `$.session.start({ cwd, surface: null, isInteractive: false })`. A hook calling `next` hits the kit's bottom which throws unless the test registered one (`on('agent.spawn', ($, e) => ({ model: e.model, agentId: 'a1' }))`). `$.turn.step` returns a stream: drain with `for await` then await `.result` (awaiting `.result` alone runs no hook). No fs/network/process: defaults path only. Assert on the ONE line that carries the value (a `show()` listing every phase always contains every id and level).
|
||||||
|
- **Apply**: write the boot helper first, type every input from the declarations, never relax a test to dodge a type. Links [[BDR-115]].
|
||||||
|
|
||||||
|
## LRN-207 — Model availability = typed API errors (`classic.StopFailure` rate_limit|overloaded|billing_error|model_not_found) + `PostModelSwitch auto`; never the turn-end reason, never `rateLimits`; sticky state and breaker target = two fields
|
||||||
|
- **Context**: model-router W1-C. `turn.complete reason: error` covers context-limit and network errors → a breaker fed by it marked fable down 15 min on a context overflow (challenge r1). `rateLimits` kinds are account windows (five_hour, seven_day, spend_limit), not per model. A per-step "engine fallback detection" (`e.model` ≠ `$.session.model()`) re-marked the session model at EVERY step → backoff climbed to the 5 h cap in one turn (confirmation r2). One field used both as sticky cur and as breaker target contradicted itself (reset at turn end vs kept).
|
||||||
|
- **Apply**: feed a breaker only from typed error kinds that name unavailability; mark per EPISODE (idempotent while down), backoff 15→300 min, `model_not_found` until reload; a user `/model` clears, `/clear` keeps (account-wide). Keep `turnModel` (sticky, reset per turn) and `lastPlan` (breaker target, kept) separate; unrouted steps pass `e.model` verbatim so the engine's own fallback is respected. Links [[BDR-115]], [[LRN-204]].
|
||||||
|
|
||||||
|
## LRN-208 — Security-auditor brief = scope only; iteration history and claimed closures violate the blind-scan contract
|
||||||
|
- **Context**: W1-C gate 2026-10-09. I sent "previous PASS with 6 MEDIUM, closed: …" in the brief. Auditor: "The auditor contract says that is never sent and must be ignored … Please do not send it next time" (it scanned blind anyway). Same rule as the verifier (lib/verify-secure-loop.md: fresh, no history).
|
||||||
|
- **Apply**: SCOPE + a neutral CONTEXT of what the code does; never prior verdicts, closures or accepted residuals. Residuals live in TODO, the auditor rediscovers them (that is the point). Links [[LRN-083]].
|
||||||
|
|
||||||
|
## LRN-209 — Scope the test run to the diff; one full `make test` before the merge; `make test` now names the red suite; GNU make rc 2 on a failed recipe
|
||||||
|
- **Context**: 2026-10-09 the pre-merge check took 454 s = full `make test` (48 suites) + a per-suite re-run to NAME the red one (aggregate rc, permanent env red design-tool-gate). Three more full passes earlier that day for mods/-only diffs. Hotfix efdd491: the recipe prints `FAIL <suite>` + `all suites green` / `<n> suite(s) red: …`. My oracle expected rc 1; GNU make returns 2 when a recipe line fails.
|
||||||
|
- **Apply**: a diff confined to one component runs that component's suite + the doctrine census; the full suite runs ONCE before `gitflow finish`; read the FAIL lines, never re-run per suite. Oracles on `make` test `[ $rc -ne 0 ]`, not `-eq 1`. Links [[LRN-173]].
|
||||||
|
|
||||||
|
|||||||
@@ -1,5 +1,28 @@
|
|||||||
# TODO
|
# TODO
|
||||||
|
|
||||||
|
## 2026-10-08 — model-router mod: one mod routes model + effort per request (feature/model-router-mod)
|
||||||
|
Plan `.claude/tasks/plans/2026-10-08-model-router-mod.md`. Decisions 2026-10-08: main-loop
|
||||||
|
model switch spike-first then flag off; load via `CLAUDE_CODE_PLUGIN_DIRS` + link.sh;
|
||||||
|
migration of shifters/pins/model-gate in wave 2 after proof; names model-router / route / /route.
|
||||||
|
- [x] W0 spike in dev-mods (hot reload): facts a-d established 2026-10-08 (plan file § Spike facts); e moved to W1.10
|
||||||
|
- [x] W1-A the mod in `mods/model-router/` (b721c94, contract `2026-10-08-model-router-w1a-1533`, plan r3): challenge 3 lenses + 1 confirmation (2 BLOCKER + 10 MAJOR closed by named changes), feater DONE first pass, GATE 0 MET 5/5, verifier CONFORME 6/6, security PASS (4 MEDIUM + 5 LOW reported, below)
|
||||||
|
- [x] W1-A hardening (user go "oui durcis", contract criteria 7-11): `/route` composer-only; no agent model axis (effort only after spawn); pattern ≤ 200 / scan ≤ 4096 / phase keys `^[a-z][a-z0-9_-]{0,31}$` / config ≤ 64 KB / `additionalProperties: false` / `typeof e.skill`; `warnOnce` in all 14 catches + config-drop logs; `safely` around post-`next` bookkeeping. feater DONE, gap round (dead `Loop.frozen`, spawn returns `started` verbatim, tool description), GATE 0 MET 9/9, verifier CONFORME 11/11, security PASS.
|
||||||
|
- [ ] W1-A residuals (security, accepted, none exploitable from outside the user's own files): ReDoS is size-bounded only (`(a+)+$` in `~/.claude/model-router.json` + a 4 KB paste hangs the hook; fix = nested-quantifier rejection or a far lower scan cap); `stat().size` trusted (FIFO/device path in ~/.claude); unrestricted `models`/`agents`/`skills` KEYS echoed raw in logs (log flood); `agentId` read from the flat tool event (engine strip unverified); 5 closures without a kit test (no fs in the kit, LRN-206); "nothing routed" catch text after a state write.
|
||||||
|
- [x] W1-B1 floor precedence (contract `2026-10-08-model-router-floor-1835`, 2026-10-09): `ultrathink` + typed `/effort-<l>` = the main turn's default AND minimum (r2 after 3 lenses: a pure floor made /effort-low a no-op); one helper `mainEffort`; per-axis precedence; mid-turn prompt floors now + next turn; `"enabled": false` per machine, kept across /clear and failed reload; typed slash attested. 30 tests; verifier CONFORME 6/6 (after 1 gap round); security PASS ×2
|
||||||
|
- [ ] W1-B1 residuals (security 2026-10-09, accepted): first-load failure of the override falls to defaults (`enabled: true`); non-boolean `enabled` drops to the default on reload; `skill.prompt` preload guard is a heuristic (`loops.size > 0`; a preload during spawn or a non-composer `/effort-*` while idle still writes the floor); marker not bound to a valid level; `/route reload` answers "config reloaded" even when the previous config was kept; transient missing override lifts a config-set off; `String(err)` of a JSON parse in the local log. Display: `/route show` folds the floor into the effort while off; model-axis text with a model-only sticky and the switch on.
|
||||||
|
- [x] W1-B2 wiring (contract `2026-10-08-model-router-wiring-1835`, 2026-10-09): tracked symlink `skills/model-router` → `../mods/model-router` loads as `model-router@skills-dir` (fresh-process `claude plugin list --json` proves it), `.gitignore` `mods/*/tsconfig.json`, `lib/tests/mods.test.sh` (4 checks, capability probe, bounded, SKIP), doctor `── Mods ──` fail-soft, CLAUDE.md `## mods/`. Plan r2 from 3 lenses (5 MAJOR), feater DONE, gap round (4b label + unbounded probe), verifier CONFORME 8/8, security PASS
|
||||||
|
- [ ] W1-B2 residuals (accepted): `claude plugin <unknown> --help` returns 0 → the suite's capability probe can pass on a CLI without `plugin test` and then FAIL instead of SKIP (fail-closed; fix = grep the probe output for the test usage line); doctor `claude plugin list --json` unbounded; doctor `echo -e` helpers interpolate `$_mod` (tracked folder names only); fallback timeout guard orphans grandchildren (only without coreutils timeout); `update-all.sh` runs `claude plugin update` over `@skills-dir` → one recurring warn (needs an update-all edit)
|
||||||
|
- [x] W1 close-out (2026-10-09): doc-sync opus audit SIGNIFICANT → user go all 10 → patched (README effort routing + Explore row + /route, USAGE, ARCHITECTURE mods/, CHANGELOG Unreleased) b22f894; hot-reload link removed from `~/.claude/dev-mods/<session>/`; BDR-115 amended (a6e2003). Pending user: `/reload-plugins` in this session (new sessions load the skills-dir copy by themselves); merge decision on feature/model-router-mod (`gitflow finish`, human signal)
|
||||||
|
- [x] W1-C adaptive tiers (user 2026-10-09, contract `2026-10-09-model-router-tiers-1237`, plan r4 after 3 lenses + 2 confirmations: 6 BLOCKER + 29 MAJOR closed by named changes): phases name absolute tiers (best fable>opus>sonnet · big opus>fable>sonnet · work sonnet>opus · cheap haiku>sonnet); breaker fed by `classic.StopFailure` kinds rate_limit|overloaded|billing_error|model_not_found + `PostModelSwitch` auto (episode backoff 15→300 min, `/model` clears, `/clear` keeps, `/route reload` clears); fallback chain fable→opus→sonnet→haiku; main UPGRADE by default under `upgradeMaxTokens` 200k (fails closed on unknown usage), DOWNGRADE gated by `mainModelSwitch`; sticky `turnModel` per turn; derived `orchestrate` on background dispatches; prompt default rules (plan/reflect FR+EN, Unicode guards, skipped on `/…`, floor match, mid-turn). 58 tests; verifier CONFORME then 3 gap/hardening rounds; security PASS
|
||||||
|
- [ ] W1-C accepted-by-design (security 2026-10-09, MEDIUM, not coded around): (1) the model itself can raise the main loop to fable for the rest of a turn through the `route` tool (plan/reflect/escalate/judge) or a background dispatch (derived orchestrate = best tier) — bounded by the turn and `upgradeMaxTokens`, no sticky route is model-callable; (2) `ultrathink` and the default keyword rules now mean "best tier" (fable) at the phase's effort, so an incidental keyword in a pasted composer prompt costs a fable turn (origin composer only). Residuals: `PostModelSwitch` `auto` covers "other programmatic change" (a healthy model could be marked 15 min); strikes never decay inside a session; a `[1m]` variant's `model_not_found` marks the base model until reload; a `models` alias added by the override without a `fallback` key stays unranked (`withEveryAlias` not applied to the default chain) so `leaveDown` skips the upgrade gates for it; `canonical` prefix match has no segment boundary (`claude-sonnet-5-50` would map to sonnet); classifier deferred to W2
|
||||||
|
- [x] W1-C live verification part 1 (2026-10-09 after `/reload-plugins`, skills-dir copy loaded, 12 hooks): derived orchestrate on a real background Explore (main high → medium while it ran → high after its end); Explore on sonnet/medium; `route plan` → next request xhigh (engine record); `Skill(effort-low)` bridged → next request low (engine record); `/route show` resolves all 10 phases to full ids, `down: none`; steps arrive as bare `claude-fable-5-1` (no `[1m]`) so the suffix carry never fires on this session
|
||||||
|
- [ ] W1-C live verification part 2 (needs a real incident): StopFailure vs turn.complete order; `PostModelSwitch` `auto` semantics and its `from_model` after a router upgrade; whether a router rewrite raises `auto`; `[1m]` carry validity on opus (only on a session whose steps carry it)
|
||||||
|
- [x] W2-A mod (2026-10-09, bb56f3e on feature/model-router-w2, contract `2026-10-09-model-router-w2a-1546`, plan r4 `2026-10-09-model-router-w2-1546`): user decisions — effort-* skills deleted (W2-B), pins REWORKED not deleted (rows = phases by role; frontmatter `model:`/`effort:` kept as census-locked off-state floor after a robustness BLOCKER), orchestrators declare phases, slim model gate. Phases `write` work/high + `apply` work/low; 56 skill rows, 21 agent rows; agents' model at spawn within tier, upward only, project-defined agents skipped (agent.offer); typed slash → name-bound marker (composer|sdk|bridge) + pending slot + idle fallback; best-tier rows in a `runMain` slot surviving turn end; unrowed skill leaves the route; route answer always names the id; `null` override rows. 3 lenses + 2 confirmations (1 BLOCKER each round closed), feater + 4 gap rounds, GATE 0 MET, verifier 3× ECARTS on coverage clauses only → user accepted at the cap, security PASS (2 MEDIUM pre-existing: first-load kill switch fails open, ReDoS size-bounded). Kit 58 → 88 tests. Doc-sync skipped for A (mod-only, docs at B).
|
||||||
|
- [ ] W2 gate A→B (user): `/reload-plugins`, then the live probe of plan § Gate (4 points: typed `/status` marker vs fallback in the verbose log; a real rowed spawn line + `step 0 agent` effort = spawn/first-step ordering; sonnet session typed `/feat` self-check + route answer; probe 1 again with a background agent alive). `typed-marker` never seen → W2-B blocked, A4 re-planned.
|
||||||
|
- [ ] W2-B repo migration (plan § W2-B B0-B7 + STEP 6/7): bridge removal, `lib/effort-shift.md` rewrite, 15 citers → `route`, `effort=` on opus general-purpose dispatches, slim `lib/model-gate.md`, delete `model-check.sh` + `effort-pins.*` + their tests + install/update blocks, delete `skills/effort-*`, analyzer `effort: xhigh`, census rewrite (drift lock rows ↔ frontmatter), docs + registries (BDR-115 amendment, LRN typed-slash/run slot, EVAL)
|
||||||
|
- [ ] W2-A residuals (security, accepted): first-load failure of the override activates the router despite `enabled:false` (fix = treat a failed first load as off); ReDoS on a self-authored prompt pattern (size-bounded); error text in the local log; model alias keys unvalidated (PHASE_KEY would do); `offers` map uncapped; `__proto__`/`constructor` override keys untested. Known limits: `skillCalls`/`spawning` counters are global; offers keyed by name only; builtin `/effort` is not a lever inside a run.
|
||||||
|
- [ ] W3 optional: step heuristics, haiku classifier, quota-aware downgrade, A/B
|
||||||
|
|
||||||
## 2026-09-30 — Higgsfield pack: CLI + skills in the install process, off by default (feature/higgsfield-pack)
|
## 2026-09-30 — Higgsfield pack: CLI + skills in the install process, off by default (feature/higgsfield-pack)
|
||||||
Contract `.claude/tasks/contracts/2026-09-30-higgsfield-pack-1412.md`, spec + plan under
|
Contract `.claude/tasks/contracts/2026-09-30-higgsfield-pack-1412.md`, spec + plan under
|
||||||
`docs/superpowers/` (transient). Approved 2026-09-30: toggle pack off by default, two toggles,
|
`docs/superpowers/` (transient). Approved 2026-09-30: toggle pack off by default, two toggles,
|
||||||
|
|||||||
@@ -0,0 +1,46 @@
|
|||||||
|
# CONTRACT — model-router-floor (wave 1-B1: user effort floor for the turn)
|
||||||
|
- date: 2026-10-08 | flow: feat | branch: feature/model-router-mod
|
||||||
|
- status: active
|
||||||
|
|
||||||
|
## REQUEST (verbatim — IMMUTABLE)
|
||||||
|
AskUserQuestion 2026-10-08, question: "Aujourd'hui, `ultrathink` met le tour en max, mais si je déclare une phase (ex. orchestrate) puis charge un skill, ton max est perdu pour la suite du tour. Ça contredit notre règle « un choix explicite bat la phase déduite ». Quel sens donner à `ultrathink` et à un `/effort-x` tapé par toi ?"
|
||||||
|
User's answer: "Plancher pour le tour (Recommended)" — option text: "Ton niveau est un minimum pour tout le tour. Les routes du modèle et des skills peuvent monter au-dessus (escalade à max), jamais descendre en dessous. Il passe aussi par-dessus un /route sticky plus bas."
|
||||||
|
Session rule this fixes (wave plan, user-approved 2026-10-08): "an explicit per-call choice (Agent `model`/`effort` param, `/route`, `ultrathink`) beats the derived phase for that span".
|
||||||
|
User, same turn: "continu avec Opus en /ultrathink".
|
||||||
|
|
||||||
|
## CLARIFICATIONS
|
||||||
|
Q: which loops does the floor cover? / A: the MAIN loop only ("le tour" = the user's turn); sub-agents keep their own routes and pins. [orchestrator — derived, stated to the user]
|
||||||
|
Q: model axis / A: one rule: `userMain?.route.model ?? turnMain?.route.model ?? turnFloor?.route.model`, applied only with the switch on (r2). [orchestrator — internal]
|
||||||
|
Q (r2): default vs minimum / A: the user's level is the turn's default when no sticky or turn route names an effort AND its minimum; a typed `/effort-low` therefore still lowers an unrouted turn. Derived from the chosen option ("Ton niveau est un minimum pour tout le tour") plus the challenge finding that a pure floor would make `/effort-low` a no-op. [orchestrator — r2]
|
||||||
|
Q (r2): mid-turn prompt / A: applied to the running turn AND kept for the next (`wait` ignored, the engine queues either way). [orchestrator — r2]
|
||||||
|
Q (r2): per-machine off switch / A: `"enabled": false` in `~/.claude/model-router.json` (untracked); an `enabledPlugins` entry would dirty the tracked settings.json on every machine. [orchestrator — r2, from the user's "configurable"]
|
||||||
|
Q: numeric or absent engine effort / A: the floor level replaces it (the user's explicit level wins over an unknown budget); the haiku effort omission still applies after flooring. [orchestrator — internal]
|
||||||
|
Q: who clears the floor / A: main turn end (a queued prompt's floor is then promoted), `/route clear`, `/route off` (pass-through). A model `route({clear})`, a `Skill(effort-*)` or any skill load never touches it. [orchestrator — derived from "jamais descendre en dessous"]
|
||||||
|
|
||||||
|
## ACCEPTANCE CRITERIA
|
||||||
|
1. Suite green with the new tests: `claude plugin test` passes with at least 22 `test(` calls; `claude plugin validate` passes with no warning; no line over 80 chars; no `any` type.
|
||||||
|
CHECK: cd mods/model-router && out=$(claude plugin test . 2>&1); rc=$?; echo "$out" | tail -n 3; [ $rc -eq 0 ] && [ "$(grep -cE '^\s*test\(' hooks/register.test.ts)" -ge 22 ] && v=$(claude plugin validate . 2>&1) && echo "$v" | grep -q 'Validation passed' && ! echo "$v" | grep -qi 'warning' && ! grep -nE '.{81,}' hooks/register.ts hooks/register.test.ts && ! grep -nE ':\s*any\b|<any>|as any\b' hooks/register.ts && echo FLOOR-SUITE-OK
|
||||||
|
EXPECT: FLOOR-SUITE-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: 30 pass 0 fail Ran 30 tests across 1 file. [1.05s] FLOOR-SUITE-OK
|
||||||
|
2. Type-check clean against this build's declarations.
|
||||||
|
CHECK: T=/Users/b.chanot/.claude/dev-mods/385f7190-70f5-4bdd-b0d8-e4566cd412fd/model-router/.claude-plugin/types; [ -d "$T" ] || T=/Users/b.chanot/Documents/claude/mods/model-router/.claude-plugin/types; W=$(mktemp -d) && printf '{"compilerOptions":{"target":"es2023","lib":["es2023"],"types":[],"module":"esnext","moduleResolution":"bundler","strict":true,"noUncheckedIndexedAccess":true,"noEmit":true,"skipLibCheck":true,"jsx":"react","jsxFactory":"h","jsxFragmentFactory":"Fragment"},"include":["%s/claude-code/index.d.ts","%s/claude-code-tools/index.d.ts","%s/hooks"]}' "$T" "$T" "$PWD/mods/model-router" > "$W/tsconfig.json" && (cd "$W" && npx --yes -p typescript@5 tsc -p tsconfig.json) && echo TSC-OK
|
||||||
|
EXPECT: TSC-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: TSC-OK
|
||||||
|
3. The floor is its own slot: `turnFloor` is declared in `State`, initialised in `newState`, written by the prompt rule and by a typed `/effort-<l>`, cleared by `/route clear` and at main turn end; the suite carries at least 8 tests whose name contains `floor`; one helper `mainEffort` decides the main effort; the config key `enabled` exists.
|
||||||
|
CHECK: cd mods/model-router/hooks && [ "$(grep -c 'turnFloor' register.ts)" -ge 6 ] && [ "$(grep -cE "^\s*test\('[^']*floor" register.test.ts)" -ge 8 ] && grep -q 'mainEffort' register.ts && grep -q 'enabled' register.ts && echo FLOOR-SLOT-OK
|
||||||
|
EXPECT: FLOOR-SLOT-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: FLOOR-SLOT-OK
|
||||||
|
4. Judged by reading: effective main effort comes from ONE helper `mainEffort` used by `mainPlan` and by every answer text: base `userMain?.route.effort ?? turnMain?.route.effort ?? turnFloor?.route.effort ?? e.effort` (per-axis, like the model rule: a model-only sticky never hides a turn route's effort; gap round 2026-10-09), then floored by `turnFloor` (LEVELS order; a numeric or absent value is replaced); the floor never applies to a sub-agent step; `turnMain` only ever holds 'model' or 'skill' sources; the prompt rule writes `turnFloor` (keeping the higher of two) and, when typed mid-turn (`turnId` set, `wait` ignored), also `pendingPrompt`, promoted into `turnFloor` at main turn end; `"enabled": false` in the override file makes every hook pass through after each config load AND survives `/clear` (`session.end` rebuilds the state but re-applies the config's `enabled`; `/route on` re-enables for the session); a prompt-rule floor is labelled by its matched phase (`prompt rule <phase>`), never by a fixed word; a non-effort skill load resets `turnMain` only; a model `route({clear})` clears `turnMain` only; every answer that the floor overrides says so truthfully (Skill bridge context, route tool text, `/effort-<l>` text); `/route show` and the status line display the floor; every criterion of `.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md` still holds; no function over 25 logic lines.
|
||||||
|
|
||||||
|
Hardening round (security gate 2026-10-09, 2 MEDIUM) — criteria 5-6, same ledger:
|
||||||
|
5. The kill switch fails closed: a failed override read on `/route reload` (unreadable, oversized, invalid JSON) keeps the PREVIOUS config (and therefore the previous `enabled`) instead of falling back to the defaults; a non-boolean `enabled` value is dropped WITH a log line; at session start with no previous config the defaults still apply.
|
||||||
|
CHECK: cd mods/model-router && grep -q "previous" hooks/register.ts && grep -qE "enabled.*(not a boolean|non-boolean|ignored)" hooks/register.ts && grep -qE "test\('[^']*(reload|previous|kill)" hooks/register.test.ts && echo KILL-CLOSED-OK
|
||||||
|
EXPECT: KILL-CLOSED-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: KILL-CLOSED-OK
|
||||||
|
6. `skill.prompt` writes the floor only for a typed `/effort-<l>`: a one-shot marker set at `prompt.submit` (composer origin, text starting with `/effort-`) attests the typing; without the marker the write is refused while any sub-agent loop is live (a preload fires inside an agent's life), and accepted otherwise (no agent can be preloading); the refused case returns the text unchanged with a one-line note. Tests: preload simulation (spawned agent live, no marker → no floor), typed with marker → floor, typed with no marker and no agent → floor.
|
||||||
|
CHECK: cd mods/model-router && grep -q "slashMarker\|typedSlash" hooks/register.ts && [ "$(grep -cE "test\('[^']*(preload|marker|typed)" hooks/register.test.ts)" -ge 2 ] && echo SLASH-ATTEST-OK
|
||||||
|
EXPECT: SLASH-ATTEST-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: SLASH-ATTEST-OK
|
||||||
|
|
||||||
|
## FILE SCOPE
|
||||||
|
mods/model-router/hooks/register.ts · mods/model-router/hooks/register.test.ts
|
||||||
@@ -0,0 +1,66 @@
|
|||||||
|
# CONTRACT — model-router-w1a (wave 1-A: the mod itself)
|
||||||
|
- date: 2026-10-08 | flow: feat | branch: feature/model-router-mod
|
||||||
|
- status: active
|
||||||
|
|
||||||
|
## REQUEST (verbatim — IMMUTABLE)
|
||||||
|
Skill args: "model-router mod, wave 1-A: the mod itself under mods/model-router/ (plugin.json, hooks.json, register.ts, config.json alias→id + phases/agents/skills/prompt tables, register.test.ts); spike code in ~/.claude/dev-mods/385f7190-70f5-4bdd-b0d8-e4566cd412fd/model-router/ is the base; plan .claude/tasks/plans/2026-10-08-model-router-mod.md W1.1-W1.9"
|
||||||
|
User (fr, same session): "go pour le registre et go sur la vague 1". Earlier framing (verbatim excerpts): "repartir correctement chaque tache au model qui lui correspond […] Il faut que l'effort aussi soit en consequence […] plus propre, plus unifier et plus automatique (meme dans la discussion courante ou d'un agent on puisse switch d'un model / effort a un autre. Et le mieux que ca soit configurable et qu'on puisse l'installer et qu'il soit actif sur toutes les session en userscope"; "Il faut un pin pour le global, mais toutes les sous taches fait pas le routage donne au model correspondant"; "si la route modifie deja les efforts, alors les skills pour changer les efforts devienne inutile mais vont quand meme etre trigger. Ca fait doublon, des token pour rien used, et peut etre meme des conflits non ?"
|
||||||
|
|
||||||
|
## CLARIFICATIONS
|
||||||
|
Q: config.json as a 5th file? / A: no — defaults live in register.ts (`DEFAULT_CONFIG`), the optional user override is `~/.claude/model-router.json` (deep-merged); `claude plugin test` runs without fs, so the mod must work with no file at all. 4 files. [orchestrator — internal, derived from the test sandbox]
|
||||||
|
Q: userConfig (W1.8) / A: dropped for 1-A — `mainModelSwitch`, `verbose`, `spinner` are keys of the same config (one source), toggled live by `/route`. [orchestrator — internal]
|
||||||
|
Q: model ids / A: hooks always write FULL ids from `config.models` (alias → id); the Agent tool param is never rewritten (its schema accepts aliases only, and the hook-side alias resolver is stale, LRN-203 / BLK-029). [orchestrator — in-force learning]
|
||||||
|
Q: precedence / A: user `/route` (sticky until `/route clear`) > the latest turn-scoped route on main (model `route` tool, a skill load's table phase or a `Skill(effort-*)` shift, a typed `/effort-*`, a prompt rule: one slot, last writer wins; a non-effort skill load resets the slot except a prompt rule) > session settings. Explicit Agent-call `model`/`effort` params always win for that agent, for its whole run: an in-agent `route` call or `Skill(effort-*)` never touches an axis given explicitly. [orchestrator — derived from the user's "pin = entry default, sub-tasks route finer"; r3 after the confirmation challenge]
|
||||||
|
Q: verbose default / A: `verbose: false` in DEFAULT_CONFIG; for now the user wants it ON to watch the routing → after the build the orchestrator writes `~/.claude/model-router.json` with `{"verbose": true}` (user-home file, outside FILE SCOPE). [gated 2026-10-08]
|
||||||
|
Q: spinner suffix default / A: on (`spinner: true`). [gated 2026-10-08]
|
||||||
|
Q: `/route` typed by the user / A: sticky until `/route clear`, wins over model-declared routes. [gated 2026-10-08]
|
||||||
|
Q: Explore built-in / A: `explore` phase = sonnet / medium (supersedes the BDR-066 wave-3 inherit for Explore). [gated 2026-10-08]
|
||||||
|
Q: legacy `Skill(effort-*)` / A: answered by the mod without loading the skill (single writer, no pairing rule); `/effort-*` typed by the user → `skill.prompt` sets the same route and returns a one-line text. [user 2026-10-08: "doublon … conflits"]
|
||||||
|
|
||||||
|
## ACCEPTANCE CRITERIA
|
||||||
|
1. `mods/model-router/` holds exactly `.claude-plugin/plugin.json`, `hooks/hooks.json`, `hooks/register.ts`, `hooks/register.test.ts` (the engine-laid `.claude-plugin/types/` folder and `./tsconfig.json` are ignored, never committed); `claude plugin validate` passes with no warning.
|
||||||
|
CHECK: cd mods/model-router && [ "$(find . -type f | grep -v '/.claude-plugin/types/' | grep -v '^./tsconfig.json$' | sort | tr '\n' ' ')" = "./.claude-plugin/plugin.json ./hooks/hooks.json ./hooks/register.test.ts ./hooks/register.ts " ] && out=$(claude plugin validate . 2>&1) && echo "$out" | grep -q 'Validation passed' && ! echo "$out" | grep -qi 'warning' && echo FILES-VALIDATE-OK
|
||||||
|
EXPECT: FILES-VALIDATE-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: FILES-VALIDATE-OK
|
||||||
|
2. Type-check clean against this build's declarations (the engine-laid copy beside the spike mod).
|
||||||
|
CHECK: T=/Users/b.chanot/.claude/dev-mods/385f7190-70f5-4bdd-b0d8-e4566cd412fd/model-router/.claude-plugin/types; W=$(mktemp -d) && printf '{"compilerOptions":{"target":"es2023","lib":["es2023"],"types":[],"module":"esnext","moduleResolution":"bundler","strict":true,"noUncheckedIndexedAccess":true,"noEmit":true,"skipLibCheck":true,"jsx":"react","jsxFactory":"h","jsxFragmentFactory":"Fragment"},"include":["%s/claude-code/index.d.ts","%s/claude-code-tools/index.d.ts","%s/hooks"]}' "$T" "$T" "$PWD/mods/model-router" > "$W/tsconfig.json" && (cd "$W" && npx --yes -p typescript@5 tsc -p tsconfig.json) && echo TSC-OK
|
||||||
|
EXPECT: TSC-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: TSC-OK
|
||||||
|
3. `claude plugin test mods/model-router` passes; the suite covers: (a) `Skill(effort-low)` via `$.tool.call` is answered without `next` in the Skill tool's output shape (`success`, `commandName`) and the route shows `low` on main; (b) the `route` tool with `phase: "orchestrate"` sets `medium` on main and `/route show` prints it; (c) `/route clear` drops it; (d) `/route bogus` returns an error text naming the phases; (e) a `prompt.submit` text holding `ultrathink` sets `escalate` on main; (f) `/route model=sonnet` shows `claude-sonnet-5-5`, a full id passes through, a misspelt alias is refused; (g) `agent.spawn` of `Explore` without a model param reaches the bottom with `model === 'claude-sonnet-5-5'`, and with `model: 'opus'` given the param is untouched.
|
||||||
|
CHECK: cd mods/model-router && out=$(claude plugin test . 2>&1); rc=$?; echo "$out" | tail -n 5; [ $rc -eq 0 ] && [ "$(grep -cE '^\s*test\(' hooks/register.test.ts)" -ge 7 ] && echo PLUGIN-TEST-OK
|
||||||
|
EXPECT: PLUGIN-TEST-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: (pass) a rule only scans the first 4096 chars of a prompt [27.16ms] 14 pass 0 fail Ran 14 tests across 1 file. [0.60s] PLUGIN-TEST-OK
|
||||||
|
4. Hooks present, as `claude plugin validate` lists them: `session.start`, `command.run{command=route}`, `tool.call{tool=mcp__model-router__route}`, `tool.call{tool=Skill}`, `tool.call{tool=Agent}`, `skill.prompt`, `agent.spawn`, `turn.step`, `prompt.submit`, `turn.complete`, `ui.render{component=Spinner}`; every gating hook carries a fail-open `.catch` (validate prints no "gating hook without .catch").
|
||||||
|
CHECK: cd mods/model-router && out=$(claude plugin validate . 2>&1) && for h in session.start 'command.run{command=route}' 'tool.call{tool=mcp__model-router__route}' 'tool.call{tool=Skill}' 'tool.call{tool=Agent}' skill.prompt agent.spawn turn.step prompt.submit turn.complete 'ui.render{component=Spinner}'; do echo "$out" | grep -qF -- "$h" || { echo "missing $h"; exit 1; }; done && ! echo "$out" | grep -q 'without .catch' && echo HOOKS-OK
|
||||||
|
EXPECT: HOOKS-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: HOOKS-OK
|
||||||
|
5. Code style: no line over 80 chars, no `any` type, no `import()`; `register.ts` imports only from `claude-code` and its own plugin files.
|
||||||
|
CHECK: cd mods/model-router/hooks && ! grep -nE '.{81,}' register.ts register.test.ts && ! grep -nE ':\s*any\b|<any>|as any\b' register.ts && ! grep -q 'import(' register.ts && [ "$(grep -cE "^import .* from '(claude-code|\./)" register.ts)" -eq "$(grep -c '^import ' register.ts)" ] && echo STYLE-OK
|
||||||
|
EXPECT: STYLE-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: STYLE-OK
|
||||||
|
6. Judged by reading: the mod never writes a model alias into a request (every `model` it sets at `agent.spawn` or `turn.step` passes through `resolveModel`); the Agent tool's `model`/`effort` params are never rewritten and an explicit `model` param is never overridden at spawn or at any step; an agent's model is written once, at spawn (a per-step model rewrite happens only after an in-agent `route` call and only while `e.model` still equals the spawn model); a fork (`e.fork`) and a workflow agent (`e.workflow`) are never re-modelled; every write from a sub-agent's Skill or `route` call lands on that agent's loop, never on main; main-loop model changes happen only when `mainModelSwitch` is true; a `.catch` on every gating hook fails open (pass-through or an in-place answer) so a mod failure never blocks a call; all state lives in the `register` closure and the defaults constant is never mutated; no function over 25 logic lines; the spike's `via` / `stepModel` / `agentsDefault` levers and the tool's `model` / `scope` params are gone.
|
||||||
|
Q (r2): `scope: agents` / `/route agents` / A: dropped (challenge r1, no requirement behind it; per-call Agent params and the table cover it). The route tool takes `phase`, `effort`, `clear` only. [orchestrator — simplicity]
|
||||||
|
Q (r2): agents table in wave 1 / A: built-ins only (Explore, Plan), matched for the engine provider; repo agents keep their frontmatter as the single writer until wave 2. [orchestrator — single source of truth]
|
||||||
|
Q (r2): `/route off` / A: added as the session kill switch (every hook passes through). [orchestrator — robustness]
|
||||||
|
|
||||||
|
Hardening round (security gate 2026-10-08, user go "oui durcis") — criteria 7-11, same ledger:
|
||||||
|
7. `/route` is user-only: `command.run` answers `{ text: 'route: user-only command' }` without acting when `e.origin.kind !== 'composer'`; a test proves it (origin `{ kind: 'plugin', name: 'x' }` or the kit's non-composer origin → text contains `user-only`, state unchanged).
|
||||||
|
CHECK: cd mods/model-router && grep -q "origin.kind" hooks/register.ts && grep -q "user-only" hooks/register.ts && grep -q "user-only" hooks/register.test.ts && echo ORIGIN-OK
|
||||||
|
EXPECT: ORIGIN-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: ORIGIN-OK
|
||||||
|
8. An in-agent `route` call or skill table entry never changes that agent's MODEL: the `Loop` type has no `model` axis and no `explicitModel` flag, `turn.step` on an agent loop rewrites `effort` only, the route tool's description says "effort only; the model of a sub-agent is fixed at spawn". A test proves it: a route tool call carrying `agentId: 'a1'` with `phase: 'judge'` followed by a `turn.step` for `a1` leaves `model` as given and sets `effort` to `xhigh`.
|
||||||
|
CHECK: cd mods/model-router && ! grep -qE "loop\.model|explicitModel|spawnModel" hooks/register.ts && grep -q "fixed at spawn" hooks/register.ts && echo NO-AGENT-MODEL-OK
|
||||||
|
EXPECT: NO-AGENT-MODEL-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: NO-AGENT-MODEL-OK
|
||||||
|
9. Config hardening: a prompt rule pattern longer than 200 chars is dropped (logged); `re.test` runs on at most the first 4096 chars of the prompt; phase keys must match `^[a-z][a-z0-9_-]{0,31}$` (others dropped, logged); the override file is refused above 65536 bytes (logged, defaults kept); the route tool schema carries `additionalProperties: false`; the Skill hook checks `typeof e.skill === 'string'`.
|
||||||
|
CHECK: cd mods/model-router && grep -q "additionalProperties: false" hooks/register.ts && grep -qE "\[a-z\]\[a-z0-9_-\]\{0,31\}" hooks/register.ts && grep -qE "4096|4_096" hooks/register.ts && grep -qE "65536|65_536|64 \* 1024" hooks/register.ts && grep -qE "200" hooks/register.ts && grep -q "typeof e.skill === 'string'" hooks/register.ts && echo CONFIG-HARDEN-OK
|
||||||
|
EXPECT: CONFIG-HARDEN-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: CONFIG-HARDEN-OK
|
||||||
|
10. Visible fail-open: every `.catch` logs once per session per hook (`$.ui.log('model-router: <hook> failed (<kind>): routing skipped for this event')`, a `warned: Set<string>` in the state) before passing through or answering; the three silent config drops (non-object top level, wrong-typed table, non-array `prompt`) log a line.
|
||||||
|
CHECK: cd mods/model-router && [ "$(grep -c '\.catch(' hooks/register.ts)" -ge 12 ] && grep -q "warned" hooks/register.ts && grep -q "routing skipped" hooks/register.ts && echo CATCH-LOG-OK
|
||||||
|
EXPECT: CATCH-LOG-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: CATCH-LOG-OK
|
||||||
|
11. Post-`next` bookkeeping (loop tracking and logs after `await next(...)` in `agent.spawn` and the Skill hook) runs inside its own try/catch so a logging failure can never make the `.catch` re-run `next`. Judged by reading, with criteria 1-6 still MET (validate, tsc, tests ≥ 13, style, AC6 minus the removed model axis).
|
||||||
|
|
||||||
|
## FILE SCOPE
|
||||||
|
mods/model-router/.claude-plugin/plugin.json · mods/model-router/hooks/hooks.json · mods/model-router/hooks/register.ts · mods/model-router/hooks/register.test.ts
|
||||||
@@ -0,0 +1,51 @@
|
|||||||
|
# CONTRACT — model-router-wiring (wave 1-B2: active in every session, tests, doctor)
|
||||||
|
- date: 2026-10-08 | flow: feat | branch: feature/model-router-mod
|
||||||
|
- status: active
|
||||||
|
|
||||||
|
## REQUEST (verbatim — IMMUTABLE)
|
||||||
|
User (fr, first message of the session): "Et le mieux que ca soit configurable et qu'on puisse l'installer et qu'il soit actif sur toutes les session en userscope".
|
||||||
|
AskUserQuestion 2026-10-08, question on the loading mechanism (CLAUDE_CODE_PLUGIN_DIRS needs an absolute path, no $HOME expansion in settings `env`, settings.json is tracked and shared): answer "1 . Mais ce n'est pas un skill on est d'accord ? C'est un mod. don cplus u plugin. ET du coup pourquoi pqs link directement ule dossier mod vers le .claude/mods directement via le link.sh ? Pourquoi passer par skills ?" (option 1 = "Lien sous skills/ (Recommended)").
|
||||||
|
Explanation given to the user the same turn: a mod is a plugin, not a skill; Claude Code loads a plugin only from a marketplace, an absolute `CLAUDE_CODE_PLUGIN_DIRS` path, a claude.ai sync, or a plugin folder (`.claude-plugin/plugin.json`) under `~/.claude/skills/` (origin `@skills-dir`, loaded in place); a `~/.claude/mods` link alone loads nothing.
|
||||||
|
Same question batch, settings answer: "Je gère moi-même" (the user's uncommitted `/model` change to settings.json).
|
||||||
|
|
||||||
|
## CLARIFICATIONS
|
||||||
|
Q: loading mechanism / A: `skills/<name>` = relative symlink `../mods/<name>`, tracked in git; loads as `<name>@skills-dir`, in place, on every machine where link.sh links `~/.claude/skills`. No settings.json change, no link.sh change. Probe 2026-10-08 in an isolated HOME: listed, enabled, "Status: ✔ loaded". [gated 2026-10-08]
|
||||||
|
Q: settings.json / A: the working-tree change (model → opus, env block moved) stays untouched and out of every commit. [gated 2026-10-08]
|
||||||
|
Q: override convention / A: a mod's optional user config lives at `~/.claude/<name>.json` (model-router already reads `~/.claude/model-router.json`). [orchestrator]
|
||||||
|
Q: doctor scope / A: per mod: the loading link resolves into the repo mod dir; `claude plugin list --json` lists `<name>@skills-dir` enabled (warn, not fail, when absent, disabled, or `claude` missing); the override file parses as JSON when present. [orchestrator]
|
||||||
|
|
||||||
|
## ACCEPTANCE CRITERIA
|
||||||
|
1. `skills/model-router` is a symlink whose target is exactly `../mods/model-router`, and git does not ignore it.
|
||||||
|
CHECK: [ -L skills/model-router ] && [ "$(readlink skills/model-router)" = "../mods/model-router" ] && ! git check-ignore -q skills/model-router && echo LINK-OK
|
||||||
|
EXPECT: LINK-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: LINK-OK
|
||||||
|
2. Engine-laid files are ignored for ANY mod: a mod's root `tsconfig.json` (root `.gitignore`; the `.claude-plugin/types/` folder ignores itself); the tracked mod files are not ignored.
|
||||||
|
CHECK: git check-ignore -q mods/model-router/tsconfig.json && git check-ignore -q mods/zz-future/tsconfig.json && ! git check-ignore -q mods/model-router/hooks/register.ts && ! git check-ignore -q mods/model-router/.claude-plugin/plugin.json && [ -z "$(git status --short mods/)" ] && echo IGNORE-OK
|
||||||
|
EXPECT: IGNORE-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: IGNORE-OK
|
||||||
|
3. `lib/tests/mods.test.sh` passes on the repo, fails on a fixture that lacks the loading link, fails on an empty `mods/`, and SKIPs (exit 0) the CLI checks when `claude plugin test` is unavailable (probe by capability, PATH-shadowed `claude` in the control).
|
||||||
|
CHECK: make test suite=lib/tests/mods.test.sh >/dev/null 2>&1 && W=$(mktemp -d) && mkdir -p "$W/mods" "$W/skills" "$W/bin" && cp -R mods/model-router "$W/mods/" && ! MODS_ROOT="$W" bash lib/tests/mods.test.sh >/dev/null 2>&1 && ln -s ../mods/model-router "$W/skills/model-router" && MODS_ROOT="$W" bash lib/tests/mods.test.sh >/dev/null 2>&1 && printf '#!/bin/sh\nexit 1\n' > "$W/bin/claude" && chmod +x "$W/bin/claude" && PATH="$W/bin:$PATH" MODS_ROOT="$W" bash lib/tests/mods.test.sh 2>&1 | grep -q '^SKIP' && E=$(mktemp -d) && mkdir -p "$E/mods" "$E/skills" && ! MODS_ROOT="$E" bash lib/tests/mods.test.sh >/dev/null 2>&1 && echo MODS-SUITE-OK
|
||||||
|
EXPECT: MODS-SUITE-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: MODS-SUITE-OK
|
||||||
|
4. `doctor.sh` prints a `── Mods ──` section with a ✓ line for model-router; with the link absent (HOME pointed at a scratch `.claude` whose `skills/` lacks the link) the section prints an info line, doctor reaches its summary and exits 0 for that section's sake (no new error).
|
||||||
|
CHECK: out=$(bash doctor.sh 2>&1); echo "$out" | sed -n '/── Mods ──/,/^$/p' | grep -q '✓.*model-router' && H=$(mktemp -d) && mkdir -p "$H/.claude/skills" && o2=$(HOME="$H" bash doctor.sh 2>&1); echo "$o2" | sed -n '/── Mods ──/,/^$/p' | grep -qi 'not linked' && echo "$o2" | grep -q '═══' && echo DOCTOR-MODS-OK
|
||||||
|
EXPECT: DOCTOR-MODS-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: DOCTOR-MODS-OK
|
||||||
|
8. The mod is enabled through the tracked link in a FRESH process: `claude plugin list --json` lists `model-router@skills-dir` with `enabled: true` (run after the dev-mods link is removed, see W6).
|
||||||
|
CHECK: claude plugin list --json 2>/dev/null | python3 -c 'import json,sys; rows=json.load(sys.stdin); ok=any(r.get("id")=="model-router@skills-dir" and r.get("enabled") is True for r in rows); sys.exit(0 if ok else 1)' && echo LOADED-OK
|
||||||
|
EXPECT: LOADED-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: LOADED-OK
|
||||||
|
5. `CLAUDE.md` has a `## mods/` section naming the `skills/<name>` relative symlink, the `@skills-dir` origin, why not `CLAUDE_CODE_PLUGIN_DIRS`, the gitignored engine-laid files, `~/.claude/<name>.json`, the suite command and how to turn a mod off.
|
||||||
|
CHECK: grep -q '^## mods/' CLAUDE.md && grep -q '@skills-dir' CLAUDE.md && grep -q 'CLAUDE_CODE_PLUGIN_DIRS' CLAUDE.md && grep -q 'mods.test.sh' CLAUDE.md && grep -q '<name>.json' CLAUDE.md && grep -q '@skills-dir": false' CLAUDE.md && echo CLAUDEMD-OK
|
||||||
|
EXPECT: CLAUDEMD-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: CLAUDEMD-OK
|
||||||
|
6. Health stack on the touched shell files, doctrine census green.
|
||||||
|
CHECK: shellcheck lib/tests/mods.test.sh doctor.sh && make test suite=lib/tests/doctrine-citers.test.sh >/dev/null 2>&1 && echo HEALTH-OK
|
||||||
|
EXPECT: HEALTH-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: HEALTH-OK
|
||||||
|
7. Judged by reading: no change to settings.json, link.sh or any install script; the suite probes the CAPABILITY (`claude plugin test --help`), bounds every CLI call in time, captures `2>&1`, SKIPs with a reason, fails when no mod is found; doctor's section is fail-soft under `set -euo pipefail` (existence test before readlink, `-ef` comparison, one guarded `claude plugin list --json`, python exits 0 with `unknown` on any parse error), never increments `_LINK_PASS`, says "enabled" not "loaded", treats a missing link as info; the link step is idempotent; CLAUDE.md names the per-machine `"enabled": false` switch, the tracked-settings cost of `enabledPlugins`, and the dev-copy shadowing rule; the CLAUDE.md section is terse English matching the file's style.
|
||||||
|
Q (r2): ordering / A: this contract runs after the floor contract (`2026-10-08-model-router-floor-1835`) is committed and green. [orchestrator]
|
||||||
|
Q (r2): update-all `claude plugin update` over `@skills-dir` / A: accepted residual (one recurring warn), logged in TODO; out of FILE SCOPE. [orchestrator]
|
||||||
|
|
||||||
|
## FILE SCOPE
|
||||||
|
skills/model-router (new symlink) · .gitignore · lib/tests/mods.test.sh (new) · doctor.sh · CLAUDE.md
|
||||||
@@ -0,0 +1,24 @@
|
|||||||
|
# CONTRACT — make-test-names-red-suites
|
||||||
|
- date: 2026-10-09 | flow: hotfix | branch: bugfix/make-test-names-red-suites
|
||||||
|
- status: active
|
||||||
|
|
||||||
|
## REQUEST (verbatim — IMMUTABLE)
|
||||||
|
Skill args: "Makefile `test` target: print `FAIL <suite>` for every red suite and a final summary line (`<n> suite(s) red: <names>` or `all suites green`) so a full `make test` names the failing suites itself; exit code unchanged (1 on any red). Today only `== <suite>` headers print and the aggregate rc forces a second per-suite run to find the red one."
|
||||||
|
User (fr): "c'est quand meme long 9 min pour faire un merge non ? … ou est le bottlneck ?" → measured: the pre-merge `make test` (454 s) was a full run PLUS a per-suite re-run to name the red suite; "oui vas y fait le maintenant".
|
||||||
|
|
||||||
|
## CLARIFICATIONS
|
||||||
|
Q: wording / A: given by the request: `FAIL <suite>` per red suite right after it runs, then one summary line `<n> suite(s) red: <names>` or `all suites green`. [user]
|
||||||
|
Q: exit code / A: unchanged: 1 when any suite is red, 0 otherwise. [user]
|
||||||
|
|
||||||
|
## ACCEPTANCE CRITERIA
|
||||||
|
1. Symptom gone: a run with one red suite prints `FAIL <that suite>` and `1 suite(s) red: <that suite>` and exits non-zero (GNU make reports a failed recipe as 2); a run with only green suites prints `all suites green` and exits 0. Checked on a two-suite fixture through `make test suite="<green> <red>"`-style invocations (the `suite` variable already accepts a list).
|
||||||
|
CHECK: cd /Users/b.chanot/Documents/claude && W=$(mktemp -d) && printf '#!/usr/bin/env bash\nexit 0\n' > "$W/green.test.sh" && printf '#!/usr/bin/env bash\nexit 1\n' > "$W/red.test.sh" && out=$(make test suite="$W/green.test.sh $W/red.test.sh" 2>&1); rc=$?; [ $rc -ne 0 ] && echo "$out" | grep -q "^FAIL $W/red.test.sh" && echo "$out" | grep -q "1 suite(s) red: $W/red.test.sh" && out2=$(make test suite="$W/green.test.sh" 2>&1); rc2=$?; [ $rc2 -eq 0 ] && echo "$out2" | grep -q "all suites green" && echo SUMMARY-OK
|
||||||
|
EXPECT: SUMMARY-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: SUMMARY-OK
|
||||||
|
2. Build/tests green: the Makefile still runs the real suites (`make test suite=lib/tests/mods.test.sh` exits 0 and prints `all suites green`); `make -n test` parses.
|
||||||
|
CHECK: cd /Users/b.chanot/Documents/claude && make -n test >/dev/null && out=$(make test suite=lib/tests/mods.test.sh 2>&1); rc=$?; [ $rc -eq 0 ] && echo "$out" | grep -q "all suites green" && echo REAL-SUITE-OK
|
||||||
|
EXPECT: REAL-SUITE-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: REAL-SUITE-OK
|
||||||
|
|
||||||
|
## FILE SCOPE
|
||||||
|
Makefile
|
||||||
@@ -0,0 +1,48 @@
|
|||||||
|
# CONTRACT — model-router-tiers (wave 1-C: absolute tiers, availability fallback, derived phases)
|
||||||
|
- date: 2026-10-09 | flow: feat | branch: feature/model-router-mod
|
||||||
|
- status: active
|
||||||
|
|
||||||
|
## REQUEST (verbatim — IMMUTABLE)
|
||||||
|
User (fr): "j'aimerais ne pas avoir a reflechir a tout ca, donc ne pas lancer les /route moi meme par exemple. On vois le reflect, orchestratm escalate et plan session (fable) par du preincipe que la sessino est sur fable de base ? Si on est sur haiku j'aimerias que ca fonctionne aussi avec les models adapte. et d'ailleurs que se passe til quand on a plus de credit fable, cela passe sur opus xhigh ? il faut se fallback. ET surtout oui avec un systeme adaptatif et pour la reflexion pour trouver une solution ou des idees, il faut le meilleur, pour en faire le plan en se basant sur cette reflexion. Bref un systeme logique et optimise. /ultrathink . ensuite je reload puginm tu test, on commit puis on passe a la vague 2"
|
||||||
|
|
||||||
|
## CLARIFICATIONS
|
||||||
|
Q: "session" phases / A: no phase keeps "the session model" any more: every phase names a TIER (`best`, `big`, `work`, `cheap`), an ordered list of aliases; the first AVAILABLE alias wins. plan/reflect/orchestrate/escalate → `best` (fable, then opus, then sonnet). A haiku session asked to plan runs the plan on fable. [user: "si on est sur haiku j'aimerais que ca fonctionne aussi avec les models adaptés"]
|
||||||
|
Q: what is "available" / A: the mod cannot read per-model quota (rateLimits are account windows: five_hour, seven_day). Availability = a circuit breaker: a model is DOWN for `cooldownMinutes` (default 15) after a main or agent turn ends with `reason: 'error'` or `'refusal'` on it, or when the engine itself switched away from it (`classic.PostModelSwitch`, `source: 'auto'`). A down model is skipped in every tier and in the `fallback` chain; `/route show` lists down models with their reset time; `/route reload` clears the breaker. [orchestrator — derived from the engine's declarations]
|
||||||
|
Q: no credit left on fable / A: main loop on a down model → next available alias of the `fallback` chain (`fable, opus, sonnet, haiku`), effort unchanged (plan stays xhigh → "opus xhigh"), always allowed (a down model yields nothing), logged and shown. [user: "il faut se fallback"]
|
||||||
|
Q: main-loop model moves / A: UPGRADE (phase tier ranks above the current model) allowed by default (`mainUpgrade: true`); DOWNGRADE (cheaper model) still gated by `mainModelSwitch` (default false: one cold-cache step per switch, and haiku's window may not fit); same rank → no switch. [orchestrator — cost/quality trade-off, stated to the user]
|
||||||
|
Q: no `/route` typed by the user / A: phases come from (1) prompt rules at turn start (keyword → phase as the turn's DEFAULT route, source 'prompt', overridable; `ultrathink` stays a FLOOR), (2) the model's own `route` calls and the skills table, (3) derived: an Agent dispatch from main pushes `orchestrate` and the previous route is restored when the last live agent ends, (4) optional classifier (`classifier: false` by default) that asks the engine's small model for a phase label when no rule matched. [user: "ne pas lancer les /route moi-même"]
|
||||||
|
Q: default prompt rules / A: floor: `\bultrathink\b` → escalate. Defaults: `\b(plan|planifie|planning|brainstorm|architecture|con[cç]ois|design)\b` → plan; `\b(pourquoi|why|explique|explain|analyse|analyze|comprendre|understand|review|audit)\b` → reflect. No default rule lowers a turn (no `mechanical` rule): lowering is explicit (route tool, skills). [orchestrator — conservative defaults, user-editable in the override]
|
||||||
|
Q: typed `/effort-<l>` / A: unchanged (floor + default of the turn, B1).
|
||||||
|
Q (r2): breaker signal / A: `classic.StopFailure` error kinds `rate_limit | overloaded | billing_error | model_not_found` (+ the engine's own fallback detected at the step), NOT `turn.complete` reasons (a context-limit or network error is not unavailability); backoff 15 → 300 min; `/clear` keeps the marks (account-wide), `/route reload` and a user `/model` clear them. [orchestrator — challenge r1]
|
||||||
|
Q (r2): classifier / A: deferred to wave 2 (dead code by default, no test can reach it, no timeout on the call). [orchestrator — challenge r1]
|
||||||
|
Q (r2): upgrade cost / A: an upgrade is skipped above `upgradeMaxTokens` (default 200000 context tokens): one cold read of the whole context per switch into a model. [orchestrator — LRN-204]
|
||||||
|
Q (r2): superseded clauses / A: floor AC4 (`turnMain` sources gain 'derived' and 'prompt'), W1-A AC6 (the switch gates downgrades only), BDR-115 (6) (window guard on every switch). [orchestrator]
|
||||||
|
|
||||||
|
## ACCEPTANCE CRITERIA
|
||||||
|
1. Suite green: `claude plugin test` passes with at least 43 `test(` calls; `claude plugin validate` passes with no warning; no line over 80 chars; no `any` type.
|
||||||
|
CHECK: cd mods/model-router && out=$(claude plugin test . 2>&1); rc=$?; echo "$out" | tail -n 3; [ $rc -eq 0 ] && [ "$(grep -cE '^\s*test\(' hooks/register.test.ts)" -ge 43 ] && v=$(claude plugin validate . 2>&1) && echo "$v" | grep -q 'Validation passed' && ! echo "$v" | grep -qi 'warning' && ! grep -nE '.{81,}' hooks/register.ts hooks/register.test.ts && ! grep -nE ':\s*any\b|<any>|as any\b' hooks/register.ts && echo TIERS-SUITE-OK
|
||||||
|
EXPECT: TIERS-SUITE-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: 58 pass 0 fail Ran 58 tests across 1 file. [2.03s] TIERS-SUITE-OK
|
||||||
|
2. Type-check clean against this build's declarations.
|
||||||
|
CHECK: T=/Users/b.chanot/Documents/claude/mods/model-router/.claude-plugin/types; W=$(mktemp -d) && printf '{"compilerOptions":{"target":"es2023","lib":["es2023"],"types":[],"module":"esnext","moduleResolution":"bundler","strict":true,"noUncheckedIndexedAccess":true,"noEmit":true,"skipLibCheck":true,"jsx":"react","jsxFactory":"h","jsxFragmentFactory":"Fragment"},"include":["%s/claude-code/index.d.ts","%s/claude-code-tools/index.d.ts","%s/hooks"]}' "$T" "$T" "$PWD/mods/model-router" > "$W/tsconfig.json" && (cd "$W" && npx --yes -p typescript@5 tsc -p tsconfig.json) && echo TSC-OK
|
||||||
|
EXPECT: TSC-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: TSC-OK
|
||||||
|
3. Tiers in the config: `tiers` (best/big/work/cheap), `fallback`, `cooldownMinutes`, `mainUpgrade`, `upgradeMaxTokens` exist in DEFAULT_CONFIG; every default phase names a tier, none a bare model; `/route show` prints the resolved model of each phase and a `down:` line; no `classifier` (deferred to wave 2); no `Loop.model` / `spawnModel` / `loop.model` identifier (W1-A AC8).
|
||||||
|
CHECK: cd mods/model-router/hooks && D=$(awk '/^const DEFAULT_CONFIG/,/^}/' register.ts) && echo "$D" | grep -q "tiers:" && echo "$D" | grep -q "fallback:" && grep -q "cooldownMinutes" register.ts && grep -q "mainUpgrade" register.ts && grep -q "upgradeMaxTokens" register.ts && ! grep -q "classifier" register.ts && [ "$(awk '/^const DEFAULT_CONFIG/,/^}/' register.ts | grep -cE "tier: '(best|big|work|cheap)'")" -ge 10 ] && ! awk '/^const DEFAULT_CONFIG/,/^}/' register.ts | grep -qE "model: '(haiku|sonnet|opus|fable)'" && ! grep -qE "loop\.model|explicitModel|spawnModel" register.ts && grep -q "down:" register.ts && echo TIERS-CONFIG-OK
|
||||||
|
EXPECT: TIERS-CONFIG-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: TIERS-CONFIG-OK
|
||||||
|
4. Tests prove (names contain the quoted word; plan r2 R14 lists them): `tier` — a `plan` route on a session model `claude-haiku-4-5-20251001` makes the main step run on `claude-fable-5-1` at xhigh (upgrade, default on); `downgrade` — a `mechanical` route on a fable session leaves the model unchanged while `mainModelSwitch` is off; `fallback` — with a `plan` route, a `classic.StopFailure` `rate_limit` on main after a fable step makes the next main step run on `claude-opus-5-5` at xhigh, and after `/route reload` fable is used again; `breaker` — an `invalid_request` failure never marks a model down, backoff expiry restores it, a `/model` command (`PostModelSwitch` source `command`) clears it; `engine fallback` — a `PostModelSwitch` with source `auto` marks the model the engine left (one strike, idempotent within the hold) and a plan route does not go back to it; `unknown` — a session model absent from the table is never switched; `spawn` — `Explore` spawns on `claude-opus-5-5` while sonnet is down (agent StopFailure with `agent_id`); `derived` — a main Agent call sets `orchestrate` and the previous `plan` route is back when the spawned agent ends; a route declared after the dispatch is not overwritten by the pop; `default rule` — "planifie la migration" sets `plan` as the turn default and a later route call overrides it; a typed `/analyze …` and a `/effort-low pourquoi …` prompt get no default rule; `per axis` — a model-less sticky never hides a turn route's tier; `floor` tests from B1 still pass.
|
||||||
|
CHECK: cd mods/model-router/hooks && for w in tier downgrade fallback breaker "engine fallback" unknown spawn derived "default rule" "per axis"; do grep -qE "test\('[^']*$w" register.test.ts || { echo "missing test: $w"; exit 1; }; done && [ "$(grep -cE "test\('[^']*(breaker|derived|default rule)" register.test.ts)" -ge 7 ] && echo TIERS-TESTS-OK
|
||||||
|
EXPECT: TIERS-TESTS-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: TIERS-TESTS-OK
|
||||||
|
5. Judged by reading (plan r2 R1-R15, r3 S1-S11 and r4 T1-T8 are binding; r4 wins over r3, r3 over r2 where they conflict: two fields `turnModel` (sticky, per turn) and `lastPlan` (breaker target, kept), unrouted steps pass `e.model` verbatim, auto marks target the model actually sent and skip when the engine landed where the router was, `sessionModel` preserved across /clear and never prefix-matched when empty, tokens `number | undefined` with windowOk failing closed, episode strikes with `model_not_found` lengthening a hold; no per-step engine-fallback detection, `PostModelSwitch` auto DOES mark one idempotent strike, the main model is sticky within a turn, `sessionModel` cached at start and on PostModelSwitch, D1 for background dispatches via the Agent result status, breaker targets kept until replaced): ONE resolver turns a tier name, an alias or a full id into an AVAILABLE canonical id (tier → list → skip down → id; an exhausted tier → the global chain; a bare alias or id passes through even when down); ids are canonical everywhere (`[1m]` stripped for comparison and carried on the replacement, alias → id, two-way prefix); ONE decision `decideMain(cur, wanted, ctx)` in the binding order (off → unknown cur: no switch → cur down: wanted or next available, windowOk → same alias: keep → better: `mainUpgrade` and `upgradeMaxTokens` → cheaper: `mainModelSwitch` and windowOk), used by `mainPlan` AND by every text; the breaker is fed only by `classic.StopFailure` errors `rate_limit | overloaded | billing_error | model_not_found` (main → the last main plan's model, agent → `agentModels`) and by `PostModelSwitch` `source: 'auto'` (the model the engine left, one strike, idempotent while down); `turn.complete` reasons never mark; a user `/model` (`PostModelSwitch` command|picker|sdk) clears the target's mark; backoff 15 → 30 → 60 → 120 → 300 min per id, `model_not_found` until reload; the breaker survives `/clear` and is cleared by `/route reload` before the config read; the derived `orchestrate` push/pop tracks this turn's spawns and never overwrites a route the model declared after the dispatch; default prompt rules (two passes, absent `mode` = floor, `iu` flags, Unicode guards) write `turnMain` (source 'prompt') and are skipped for a leading `/`, for a prompt carrying a floor or a typed slash, and mid-turn; floor rules write `turnFloor`; no classifier; explicit Agent params still win; agent model fixed at spawn; every B1/1-A criterion still holds EXCEPT the three clauses R15 names (turnMain sources, the switch clause, the window-guard scope); no function over 25 logic lines; truthful texts come from `decideMain` and say `upgrade`, `fallback`, `switch off` or `unchanged`.
|
||||||
|
|
||||||
|
Hardening round (security gate 2026-10-09, 3 MEDIUM + 1 LOW accepted) — criteria 6-7, same ledger:
|
||||||
|
6. (a) `leaveDown` runs a `wanted` that ranks ABOVE cur through the upgrade checks (`mainUpgrade`, `upgradeMaxTokens`, windowOk) before taking it; (b) the upgrade cap fails CLOSED: unknown tokens (`undefined`) → no upgrade, logged once per turn; the test boot answers `session.usage` with a small context so the upgrade tests still run, and one test proves an unanswered usage blocks the upgrade; (c) `canonical` keeps ONE prefix direction only (`bare.startsWith(tableId)`, which covers `[1m]` and dated variants) and `markDown` charges `model_not_found` to the exact id the request carried when that id is not itself a table id (no mark); (d) the once-per-turn log key for "upgrade skipped: context N tokens" is fixed (no N in the key).
|
||||||
|
CHECK: cd mods/model-router/hooks && ! grep -qE "known\.startsWith\(bare\)|tableId\.startsWith\(bare\)|id\.startsWith\(bare\)" register.ts && grep -qE "test\('[^']*(cap|unknown tokens|usage)" register.test.ts && grep -q "upgrade skipped" register.ts && echo HARDEN-C-OK
|
||||||
|
EXPECT: HARDEN-C-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: HARDEN-C-OK
|
||||||
|
7. Judged by reading: criteria 1-5 still hold (tests ≥ 55 + the new ones green, validate, tsc, style); the `PostModelSwitch` `auto` handling is unchanged and listed for live verification; the two accepted-by-design MEDIUMs (model-initiated upgrades within a turn; `ultrathink` → best tier at max) are recorded in TODO, not coded around.
|
||||||
|
|
||||||
|
## FILE SCOPE
|
||||||
|
mods/model-router/hooks/register.ts · mods/model-router/hooks/register.test.ts
|
||||||
@@ -0,0 +1,44 @@
|
|||||||
|
# CONTRACT — model-router-w2a
|
||||||
|
- date: 2026-10-09 | flow: feat | branch: feature/model-router-w2 (to start off develop)
|
||||||
|
- status: active
|
||||||
|
- wave: W2-A (mod, plan r2). W2-B (repo migration + docs/registries as its STEP 6/7) gets its own contract after the live probe.
|
||||||
|
|
||||||
|
## REQUEST (verbatim — IMMUTABLE)
|
||||||
|
tu peux lancer la vague 2
|
||||||
|
(TODO W2 line: 15 skills `Skill(effort-*)` → `route` tool calls; remove `skills/effort-*`, `lib/effort-pins.txt/.sh`, install/update steps, `effort:` frontmatter on skills and agents; repo agents into the mod's `agents` table (verify the spawn/first-step ordering first); `lib/model-gate.md` + `lib/model-check.sh` → mod rule; census tests repointed; `lib/effort-shift.md` rewritten; docs)
|
||||||
|
|
||||||
|
## CLARIFICATIONS
|
||||||
|
(pass A: none — outcome, scope and constraints come from the TODO W2 line + plan `2026-10-08-model-router-mod.md` § Wave 2)
|
||||||
|
Q: keep the five `effort-*` skills as typed floor levers? / A: "supprimer, j'utilise les commandes builtin si j'ai besoin" → deleted; the mod drops its `Skill(effort-*)` bridge and typed-slash floor code; `ultrathink` floor stays [gated 2026-10-09]
|
||||||
|
Q: table rows = bare level next to phases, or phases only? / A: "le pin était là car pas encore de système de routage fiable; maintenant qu'il y en a un, soit on le supprime, soit tu les remanies pour que ça fonctionne encore mieux" → reworked: rows are PHASES by role (plan/reflect/implement/write/verify/judge/mechanical), no level rows; agents routed on both axes by the mod while on (model written at spawn from the phase tier, upward only); the tracked `model:`/`effort:` frontmatter STAYS as the census-locked off-state floor (plan r3, confirmation 1/8) [gated 2026-10-09]
|
||||||
|
Q: orchestrators declare phases or effort only? / A: phases [gated 2026-10-09]
|
||||||
|
Q: (mid-run, feater NEED-DECISION, CLASS internal) a skill WITHOUT a row loaded INSIDE a sub-agent: keep the old reset of that loop's effort, or leave the loop untouched like main (A5)? / A: leave the agent loop untouched too — one rule for both loops, an unrowed helper skill (find-docs) never wipes an agent row's effort [gated 2026-10-09]
|
||||||
|
Q: GATE 1 cap (3 × ECARTS on test coverage of criterion 3 clauses, code correct ×3; last rows landed by a feater, unverified) / A: user "Accepter et commiter W2-A" — accepted on the three reports + 88-test kit suite; diagnosis: compound coverage criterion, not the code [gated 2026-10-09]
|
||||||
|
Q: model gate fate? / A: slim include (self-check + STOP when the mod could not raise), `lib/model-check.sh` + test deleted [gated 2026-10-09]
|
||||||
|
|
||||||
|
## ACCEPTANCE CRITERIA
|
||||||
|
1. The mod's kit suite is green, including the new tests of criteria 3-5.
|
||||||
|
CHECK: cd mods/model-router && out="$(claude plugin test . 2>&1)" && printf '%s\n' "$out" | grep -qE '[0-9]+ pass' && printf '%s\n' "$out" | grep -qE '(^|[^0-9])0 fail' && ! printf '%s\n' "$out" | grep -qE '[1-9][0-9]* fail' && echo W2A-TESTS-GREEN
|
||||||
|
EXPECT: W2A-TESTS-GREEN
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: W2A-TESTS-GREEN
|
||||||
|
2. `DEFAULT_CONFIG.phases` has `write` at work/high and a new `apply` at work/low; `DEFAULT_CONFIG.skills` and `.agents` carry exactly the phase rows of plan r2 § Row tables (56 skill rows, 21 agent rows + Explore/Plan), every value a phase key; the "Built-ins only" comment is gone.
|
||||||
|
CHECK: bash .claude/tasks/contracts/w2a-table-census.sh mods/model-router/hooks/register.ts
|
||||||
|
EXPECT: W2A-TABLE-COMPLETE
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: W2A-TABLE-COMPLETE
|
||||||
|
3. A user-typed slash of a rowed skill routes main to its row (name-bound marker from `prompt.submit`, origins composer|sdk|bridge, pending slot mid-turn, or no live/spawning sub-agent); a preload inside a live sub-agent without the marker leaves main; a skill WITHOUT a row leaves the route in force; a best-tier skill row lives in a separate `runMain` slot that survives `turn.complete` and later `route` calls, and is dropped by `/route clear`, `/route off`, a user-typed non-best rowed skill, or a user `/model`; the `Skill(effort-*)` bridge is not sticky. Covered by kit tests.
|
||||||
|
4. Any rowed agent spawn (whatever `provider.plugin`) with `e.model` undefined gets the row's model written at spawn: full id resolved within the tier only and ranked ≥ the tier head (upward only); otherwise no write + one deduped log line; the row's effort applies on its loop from its first step unless the call carried `effort`; `fork`/`workflow` spawns and agents whose `agent.offer` source is a project definition are untouched; an override `agents: { name: null }` drops a row; the `route` tool answer always names the id the next main step runs on. Covered by kit tests.
|
||||||
|
5. Typed-floor code removed: no `slashEffort`, `guardedSlash`, `'slash'` Source, `typed /effort-` floor word in `register.ts`; `EFFORT_SKILL`/`effortBridge` (tool.call bridge) and the `ultrathink` floor kept until W2-B; `/route` unchanged (existing tests still green).
|
||||||
|
CHECK: ! grep -qE "slashEffort|guardedSlash|'slash'|typed /effort-" mods/model-router/hooks/register.ts && grep -q 'turnFloor' mods/model-router/hooks/register.ts && grep -q 'effortBridge' mods/model-router/hooks/register.ts && echo W2A-DEAD-CODE-GONE
|
||||||
|
EXPECT: W2A-DEAD-CODE-GONE
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: W2A-DEAD-CODE-GONE
|
||||||
|
7. `claude plugin validate` passes.
|
||||||
|
CHECK: cd mods/model-router && claude plugin validate . 2>&1 | grep -qi 'valid' && echo W2A-VALID
|
||||||
|
EXPECT: W2A-VALID
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: W2A-VALID
|
||||||
|
6. `make test suite=lib/tests/mods.test.sh` green.
|
||||||
|
CHECK: make test suite=lib/tests/mods.test.sh 2>&1 | tail -5 | grep -q 'mods' && ! make test suite=lib/tests/mods.test.sh 2>&1 | grep -q '^FAIL' && echo W2A-MODS-SUITE
|
||||||
|
EXPECT: W2A-MODS-SUITE
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: W2A-MODS-SUITE
|
||||||
|
|
||||||
|
## FILE SCOPE
|
||||||
|
mods/model-router/hooks/register.ts, mods/model-router/hooks/register.test.ts, .claude/tasks/contracts/w2a-table-census.sh (oracle, written by the orchestrator)
|
||||||
Executable
+56
@@ -0,0 +1,56 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# GATE 0 oracle, contract 2026-10-09-model-router-w2a: DEFAULT_CONFIG.skills
|
||||||
|
# and .agents in register.ts hold exactly the plan's phase rows (plan
|
||||||
|
# 2026-10-09-model-router-w2 § Row tables), every value a declared phase.
|
||||||
|
# Prints W2A-TABLE-COMPLETE only when every assertion passes.
|
||||||
|
set -u
|
||||||
|
F="${1:?register.ts path}"
|
||||||
|
python3 - "$F" <<'PY'
|
||||||
|
import re, sys
|
||||||
|
src = open(sys.argv[1]).read()
|
||||||
|
m = re.search(r'const DEFAULT_CONFIG: Config = \{(.*?)\n\}\n', src, re.S)
|
||||||
|
if not m: print("no DEFAULT_CONFIG block"); sys.exit(1)
|
||||||
|
cfg = m.group(1)
|
||||||
|
def block(name):
|
||||||
|
b = re.search(r'\n ' + name + r': \{([^\n]*)\},', cfg) \
|
||||||
|
or re.search(r'\n ' + name + r': \{(.*?)\n \}', cfg, re.S)
|
||||||
|
if not b: print(f"no {name} block"); sys.exit(1)
|
||||||
|
rows = re.findall(r"(?:^|[{,])\s*'?([A-Za-z0-9_-]+)'?:\s*'([a-z]+)'", b.group(1), re.M)
|
||||||
|
return dict(rows)
|
||||||
|
phases = set(
|
||||||
|
re.findall(r"^\s*([a-z]+): \{ tier:", re.search(r'\n phases: \{(.*?)\n \}', cfg, re.S).group(1), re.M))
|
||||||
|
want_skills = {}
|
||||||
|
for ph, names in {
|
||||||
|
'plan': 'ship-feature init-project onboard tour audit-delta analyze code-clean client-handover brainstorming writing-plans requesting-code-review 21st-ui-review',
|
||||||
|
'reflect': 'feat hotfix bugfix refactor web-validate harden seo geo site-motion frontend-design emil-design-eng design-motion-principles 21st-ui-build scroll-world-storytelling build-threejs-scroll-worlds scroll-scrubbed-visual-sequence scroll-scrubbed-word-reveal scroll-progress-timeline subagent-driven-development writing-skills deprecation-and-migration 21st-ai 21st-ui-explore',
|
||||||
|
'implement': 'gitflow prune-memory pdf-translate ci-cd-and-automation observability-and-instrumentation test-driven-development',
|
||||||
|
'apply': 'commit-change release-candidate doc capitalize close reconcile deploy',
|
||||||
|
'mechanical': 'status profile plugin-check skills-perso using-git-worktrees 21st-cli-use 21st-registry 21st-design-sync',
|
||||||
|
}.items():
|
||||||
|
for n in names.split(): want_skills[n] = ph
|
||||||
|
want_agents = {'Explore': 'explore', 'Plan': 'judge'}
|
||||||
|
for ph, names in {
|
||||||
|
'implement': 'feater bugfixer code-cleaner scaffolder onboarder',
|
||||||
|
'write': 'commit-changer doc-syncer handover-doc-writer refactorer',
|
||||||
|
'apply': 'hotfixer release-executor plugin-probe validator-analyzer',
|
||||||
|
'verify': 'verifier security-auditor',
|
||||||
|
'judge': 'plan-challenger plugin-advisor seo-analyzer geo-analyzer analyzer',
|
||||||
|
'mechanical': 'status-reporter',
|
||||||
|
}.items():
|
||||||
|
for n in names.split(): want_agents[n] = ph
|
||||||
|
ok = True
|
||||||
|
for name, want in (('skills', want_skills), ('agents', want_agents)):
|
||||||
|
got = block(name)
|
||||||
|
for k in sorted(set(want) | set(got)):
|
||||||
|
if want.get(k) != got.get(k):
|
||||||
|
ok = False; print(f"{name}.{k}: want {want.get(k)} got {got.get(k)}")
|
||||||
|
bad = {k: v for k, v in got.items() if v not in phases}
|
||||||
|
if bad: ok = False; print(f"{name}: undeclared phases {bad}")
|
||||||
|
if len(want_skills) != 56 or len(want_agents) != 23:
|
||||||
|
ok = False; print(f"oracle self-check: {len(want_skills)} skill rows, {len(want_agents)} agent rows")
|
||||||
|
if 'Built-ins only' in cfg: ok = False; print("stale comment 'Built-ins only'")
|
||||||
|
ph = re.search(r'\n phases: \{(.*?)\n \}', cfg, re.S).group(1)
|
||||||
|
for want in ("write: { tier: 'work', effort: 'high' }", "apply: { tier: 'work', effort: 'low' }"):
|
||||||
|
if want not in ph: ok = False; print(f"phases: missing {want}")
|
||||||
|
print("W2A-TABLE-COMPLETE" if ok else "W2A-TABLE-INCOMPLETE"); sys.exit(0 if ok else 1)
|
||||||
|
PY
|
||||||
@@ -0,0 +1,172 @@
|
|||||||
|
# PLAN — model-router wave 1-B1: user effort floor for the turn (dispatch-ready)
|
||||||
|
Contract: .claude/tasks/contracts/2026-10-08-model-router-floor-1835.md
|
||||||
|
Code: mods/model-router/hooks/register.ts (read it in full first) and
|
||||||
|
register.test.ts. API truth: the engine-laid declarations under
|
||||||
|
mods/model-router/.claude-plugin/types/ (claude-code/index.d.ts,
|
||||||
|
claude-code-tools/index.d.ts).
|
||||||
|
|
||||||
|
## Why
|
||||||
|
Today the prompt rule (`ultrathink` → escalate) and a typed `/effort-<l>`
|
||||||
|
share ONE slot (`turnMain`) with the model's `route` calls and the
|
||||||
|
`Skill(effort-*)` bridge: last writer wins, and a later skill load resets
|
||||||
|
the slot to the session default. A user's explicit level is therefore lost
|
||||||
|
mid-turn. The user chose FLOOR semantics: their level is a minimum for the
|
||||||
|
whole main turn; derived routes may go above it, never below; it also
|
||||||
|
lifts a lower sticky `/route`.
|
||||||
|
|
||||||
|
## Precedence after the change (main loop only)
|
||||||
|
- model axis (unchanged order, floor last): `userMain ?? turnMain ?? turnFloor`
|
||||||
|
route's `model`, applied only with `mainModelSwitch` and the window guard.
|
||||||
|
- effort axis: `base = (userMain ?? turnMain)?.route.effort ?? e.effort`;
|
||||||
|
`effort = floored(base, turnFloor?.route.effort)`.
|
||||||
|
- `floored(effort, floor)`: no floor → `effort`; `effort` is a Level whose
|
||||||
|
LEVELS index ≥ the floor's → `effort`; otherwise (lower Level, a number,
|
||||||
|
or undefined) → `floor`.
|
||||||
|
- The haiku omission (`effort: undefined` when the model sent starts with
|
||||||
|
`claude-haiku`) still runs AFTER flooring.
|
||||||
|
- Sub-agent steps (`e.agentId` set) never read `turnFloor`.
|
||||||
|
|
||||||
|
## Changes in register.ts (names as in the current file)
|
||||||
|
1. `State`: add `turnFloor: Routed | null` with the comment `user-explicit
|
||||||
|
level for this turn (prompt rule, typed /effort-<l>): a floor, main loop
|
||||||
|
only`; reword the `turnMain` comment to `model route tool, skill table
|
||||||
|
row, Skill(effort-*) bridge; dropped at turn end`. `newState`:
|
||||||
|
`turnFloor: null`.
|
||||||
|
2. Helpers (new, small): `const rank = (l: Level): number => LEVELS.indexOf(l)`;
|
||||||
|
`function floored(effort: StepIn['effort'], floor: Level | undefined)`
|
||||||
|
per the rule above. A helper `floorLevel(st)` returning
|
||||||
|
`st.turnFloor?.route.effort` is allowed if it keeps call sites short.
|
||||||
|
3. `mainPlan`: compute `base` and `effort = floored(base, floorLevel(st))`;
|
||||||
|
`wanted = set?.route.model ?? st.turnFloor?.route.model`; the rest
|
||||||
|
(switch, window guard) unchanged.
|
||||||
|
4. `registerPrompt` / `prompt.submit`: the non-queued branch writes
|
||||||
|
`st.turnFloor = routed` (instead of `st.turnMain`); the queued branch
|
||||||
|
(`e.turnId !== undefined && e.wait`) keeps writing `st.pendingPrompt`.
|
||||||
|
5. `slashEffort`: write `st.turnFloor = { phase: skill, route: { effort:
|
||||||
|
level }, source: 'slash' }`; returned text line becomes `Effort floor
|
||||||
|
<level> set by model-router for this turn: nothing below it runs.`
|
||||||
|
followed by `\n` + the original text (prepend, never replace).
|
||||||
|
6. `onSkillLoad` (main branch): `st.turnMain = null` unconditionally (the
|
||||||
|
slot no longer holds prompt or slash routes), then the table row as now.
|
||||||
|
`turnFloor` is never touched there.
|
||||||
|
7. `clearRoutes` (`/route clear`): also `st.turnFloor = null`.
|
||||||
|
`clearLoop` (model `route({clear})`, main branch): `turnMain` only, as now.
|
||||||
|
8. `endMainTurn`: `st.turnFloor = st.pendingPrompt; st.pendingPrompt =
|
||||||
|
null; st.turnMain = null;` then the existing resets.
|
||||||
|
9. Truthful answers (main branch only; agent branches unchanged):
|
||||||
|
- `effortBridge`: keep the sticky sentence when `st.userMain` is set;
|
||||||
|
else when the floor ranks above `level`: `model-router: <skill>
|
||||||
|
recorded, but the user's floor <f> for this turn keeps main at <f>;
|
||||||
|
the <skill> skill text was not loaded.`; else the current sentence.
|
||||||
|
- `routedText`: keep the sticky branch; else when `p.route.effort` is
|
||||||
|
set and the floor ranks above it, print the effort as `<f> (user
|
||||||
|
floor; asked <asked>)`.
|
||||||
|
10. Display: `mainText` appends ` · floor <f> (<phase>)` when `turnFloor` is
|
||||||
|
set (also after `main: session defaults`); `statusLine` appends
|
||||||
|
` · floor <f>`.
|
||||||
|
|
||||||
|
## Tests in register.test.ts (keep every existing test; adapt only what
|
||||||
|
the new slot changes, e.g. the `ultrathink` test now expects the floor on
|
||||||
|
the `main:` line). Add at least six tests whose names contain `floor`,
|
||||||
|
using the existing boot helper, full typed inputs and bottom hooks, and
|
||||||
|
asserting on the `main:` line or on what the bottom `turn.step` hook
|
||||||
|
receives (drain the stream with `for await`, then `.result`):
|
||||||
|
- `floor: ultrathink survives a model route` — prompt `ultrathink`
|
||||||
|
(composer, `wait: false`, no `turnId`) then route tool `orchestrate` →
|
||||||
|
a main step with engine effort `high` reaches the bottom at `max`.
|
||||||
|
- `floor: a typed /effort-medium floors a lower route and allows a higher
|
||||||
|
one` — `$.skill.prompt({ skill: 'effort-medium', text: 'x' })` (no Skill
|
||||||
|
call in flight) → route tool `mechanical` → main step at `medium`;
|
||||||
|
then route tool `escalate` → main step at `max`.
|
||||||
|
- `floor: survives a skill load` — ultrathink, route tool `orchestrate`,
|
||||||
|
then a non-effort `Skill` call (bottom `tool.call` hook registered) →
|
||||||
|
main step at `max`, and `main:` line no longer names `orchestrate`.
|
||||||
|
- `floor: lifts a lower sticky route, then ends with the turn` — `/route
|
||||||
|
effort=low` then ultrathink → main step at `max`; fire a main
|
||||||
|
`turn.complete` → next main step at `low`.
|
||||||
|
- `floor: main only` — ultrathink, spawn `Explore` (bottom `agent.spawn`
|
||||||
|
hook returning an `agentId`), then a step for that `agentId` → reaches
|
||||||
|
the bottom at `medium`, not `max`.
|
||||||
|
- `floor: /route clear removes it` — ultrathink, `/route clear` → main
|
||||||
|
step keeps the engine effort.
|
||||||
|
Optional seventh: the bridge context line names the floor when it wins.
|
||||||
|
|
||||||
|
## Constraints
|
||||||
|
- Style: ≤ 25 logic lines per function, 80 chars per line, no `any`, no
|
||||||
|
module-level mutable state, doc comments state intent.
|
||||||
|
- Do not touch: the agent axis, the spawn table, config loading, the
|
||||||
|
hardening (caps, warnOnce, safely), the route tool schema.
|
||||||
|
- Verify (paste outputs): `claude plugin validate .`, the contract's tsc
|
||||||
|
command, `claude plugin test .`, then from the repo root
|
||||||
|
`bash ~/.claude/lib/gates.sh run .claude/tasks/contracts/2026-10-08-model-router-floor-1835.md`.
|
||||||
|
|
||||||
|
## Disposition
|
||||||
|
- honors BDR-115 (one writer per axis, calling-loop writes, truthful
|
||||||
|
answers) and the wave plan's routing rule (explicit user choice beats the
|
||||||
|
derived phase); supersedes the 1-A contract's one-slot precedence for
|
||||||
|
prompt and slash sources.
|
||||||
|
- LRN-206 (kit facts) applies to every new test.
|
||||||
|
|
||||||
|
## r2 — challenge round (3 lenses, 0 BLOCKER, 4 MAJOR): BINDING, overrides the sections above where they conflict
|
||||||
|
R1. ONE decision helper, used by `mainPlan` AND by every answer text:
|
||||||
|
`mainEffort(st, engine: StepIn['effort'])` → `{ effort, by }` with
|
||||||
|
`by` ∈ `'floor' | 'sticky' | 'turn' | 'engine'`.
|
||||||
|
`base = (st.userMain ?? st.turnMain)?.route.effort
|
||||||
|
?? st.turnFloor?.route.effort ?? engine`
|
||||||
|
`effort = floored(base, st.turnFloor?.route.effort)`; `by = 'floor'`
|
||||||
|
when the floor raised or supplied the value, else the slot it came from.
|
||||||
|
The user's level is therefore BOTH the turn's default (when no sticky
|
||||||
|
or turn route names an effort) AND its minimum: a typed `/effort-low`
|
||||||
|
lowers a turn that has no route (engine `high` → `low`), and a route
|
||||||
|
can still go higher. No text function compares ranks on its own.
|
||||||
|
R2. Model axis, one rule written once (contract updated):
|
||||||
|
`st.userMain?.route.model ?? st.turnMain?.route.model ?? st.turnFloor?.route.model`,
|
||||||
|
switch and window guard unchanged.
|
||||||
|
R3. Prompt rule with `e.turnId !== undefined` (typed while a turn runs;
|
||||||
|
`wait` is IGNORED: the engine queues every mid-turn prompt either way):
|
||||||
|
write the floor NOW (higher of the existing floor and the new one)
|
||||||
|
AND set `pendingPrompt` to it, so the turn that reads the prompt has it
|
||||||
|
whichever it is. No `turnId` → write the floor (higher of two).
|
||||||
|
`endMainTurn` promotes `pendingPrompt` into `turnFloor`. Two floors in
|
||||||
|
one turn always keep the higher one (prompt rule and typed slash).
|
||||||
|
R4. Truthful texts, all phrased from `mainEffort` (main branch only):
|
||||||
|
- Skill bridge, route tool, typed `/effort-<l>`: when `by === 'floor'`
|
||||||
|
and the result differs from what was asked, name the floor and its
|
||||||
|
source (`ultrathink rule` or `typed /effort-<l>`) and add
|
||||||
|
`/route clear to drop it`; when `by === 'sticky'`, the sticky
|
||||||
|
sentence; the old fixed "sticky wins" sentences go.
|
||||||
|
- `/effort-<l>` text: `Effort <l> set by model-router for the main loop
|
||||||
|
this turn (minimum; a higher route still applies).` plus the floor
|
||||||
|
or sticky outcome when one changes it.
|
||||||
|
- main loop on a haiku model: print `effort - (haiku takes none)`.
|
||||||
|
- model `route({clear})` on main with a floor set: append `; user floor
|
||||||
|
<f> (<source>) still holds — /route clear drops it`.
|
||||||
|
R5. Display: `mainText` / `statusLine` show ` · floor <f>` only when the
|
||||||
|
floor carries an effort AND the router is on; the `skill.prompt` hook
|
||||||
|
calls `refresh($, st)` after the slash write.
|
||||||
|
R6. Persistent per-machine off switch (wiring challenge, user's "configurable"):
|
||||||
|
config key `enabled: boolean` (default `true`) in
|
||||||
|
`~/.claude/model-router.json` (untracked, per machine). `false` →
|
||||||
|
`st.off = true` after every config load (session start, `/route
|
||||||
|
reload`); `/route on` re-enables for the session only; `show` and the
|
||||||
|
status line say `off (config)` vs `off`. Merged with `pickBool` like
|
||||||
|
the other scalars; `DEFAULT_CONFIG.enabled = true`.
|
||||||
|
R7. Tests (replace the list above where it differs): `runStep` takes a
|
||||||
|
full `TurnStepInput` (from 'claude-code'); every floor test steps with
|
||||||
|
engine effort `high` (or `xhigh`); every `test('…', async (` line ≤ 80
|
||||||
|
chars with `floor` in the single-line name. At least 8 floor tests:
|
||||||
|
ultrathink survives a model route · typed /effort-medium clamps low,
|
||||||
|
lets max pass · typed /effort-low lowers an unrouted turn (engine high
|
||||||
|
→ low) · survives a skill load · lifts a lower sticky then ends with
|
||||||
|
the turn · main only (agent step unaffected) · /route clear removes it
|
||||||
|
· mid-turn prompt (turnId + wait) is applied now AND promoted after the
|
||||||
|
main turn.complete · mandatory text test: sticky `/route effort=low`,
|
||||||
|
ultrathink, route tool `plan` → the answer names the floor.
|
||||||
|
`enabled: false` cannot be reached in the kit (no fs, LRN-206): cover
|
||||||
|
the off path through `/route off` and say so in a comment.
|
||||||
|
R8. Residuals accepted (logged in TODO, not built): floor expiry depends on
|
||||||
|
a main `turn.complete` reaching this mod (another plugin answering it
|
||||||
|
without `next` would keep it); `skill.prompt` cannot tell a typed
|
||||||
|
`/effort-<l>` from a sub-agent preload (no agentId; no repo agent
|
||||||
|
preloads one); an incidental "ultrathink" in pasted text floors the
|
||||||
|
turn (mitigated by R4 naming the source and the `/route clear` hint).
|
||||||
@@ -0,0 +1,195 @@
|
|||||||
|
# PLAN — model-router mod (feature/model-router-mod)
|
||||||
|
|
||||||
|
User ask 2026-10-08: one Claude Code mod that routes every request to the
|
||||||
|
model and effort its task deserves, main loop and sub-agents alike, declared
|
||||||
|
or automatic, configurable, loaded in every session (user scope). Replaces
|
||||||
|
the five `effort-*` shifter skills, the `effort:` / `model:` frontmatter
|
||||||
|
pins, `lib/effort-pins.txt` + `.sh` and `lib/model-gate.md` once proven.
|
||||||
|
|
||||||
|
Decisions taken 2026-10-08 (user, AskUserQuestion):
|
||||||
|
- Main-loop MODEL switch: spike first, then behind a userConfig flag, off by
|
||||||
|
default, applied only on explicit declaration. Main-loop EFFORT always routed.
|
||||||
|
- Loading: `CLAUDE_CODE_PLUGIN_DIRS` in settings.json `env`, mod lives in the
|
||||||
|
repo under `mods/model-router/`, symlinked by link.sh. No marketplace.
|
||||||
|
- Migration: wave 2, after wave 1 is proven. Mod = single source of truth.
|
||||||
|
- Names: mod `model-router`, tool `route` (model sees `mcp__model-router__route`),
|
||||||
|
command `/route`, config `~/.claude/model-router.json`.
|
||||||
|
|
||||||
|
## Routing rule (user, 2026-10-08): pin = entry default, sub-tasks route finer
|
||||||
|
The existing pins and `effort-*` shifters were built for this same goal with
|
||||||
|
the tools of their time; the mod replaces them (or nearly). A pin is not to
|
||||||
|
be contested one by one, but some were forced: a skill pinned to one level
|
||||||
|
does many different things inside one run. So:
|
||||||
|
- the entry pin of a skill or agent (today frontmatter / effort-pins.txt,
|
||||||
|
tomorrow the config table) = the DEFAULT route of the run, never a ceiling
|
||||||
|
or a floor;
|
||||||
|
- inside the run every sub-task routes to its own phase: declared by the
|
||||||
|
skill through the `route` tool (replaces `Skill(effort-*)` + pairing rule),
|
||||||
|
or derived by the mod (Agent dispatch → orchestrate, Skill load → its
|
||||||
|
entry phase, Read/Grep result → comprehension level, bookkeeping tail →
|
||||||
|
mechanical);
|
||||||
|
- an explicit per-call choice (Agent `model`/`effort` param, `/route`,
|
||||||
|
`ultrathink`) beats the derived phase for that span;
|
||||||
|
- wave 1 keeps the frontmatter pins as the entry defaults (the mod reads the
|
||||||
|
same values), wave 2 moves them into the config table and deletes the
|
||||||
|
frontmatter + shifters. Skills get their intra-run `route` calls in wave 2
|
||||||
|
(the 15 `Skill(effort-*)` citers first).
|
||||||
|
|
||||||
|
## Spike facts so far (2026-10-08)
|
||||||
|
- (a) main-loop effort rewrite at `turn.step` reaches the API: transcript
|
||||||
|
records flip `effort: high` → `medium` after a `route` call. `CLAUDE_EFFORT`
|
||||||
|
is NOT an oracle (turn-level setting); the transcript `effort` field is.
|
||||||
|
- (c) sub-agent steps are visible by `agentId`; an explicit Agent `model`
|
||||||
|
param resolves fine (sonnet → claude-sonnet-5-5, haiku → claude-haiku-4-5-20251001;
|
||||||
|
haiku steps carry `effort: undefined`, no effort on that model).
|
||||||
|
- ROOT CAUSE of the 404 (T1-T3, 2026-10-08): a model set by a hook as an
|
||||||
|
ALIAS is resolved by a stale table (`sonnet` -> `claude-sonnet-5`, 404);
|
||||||
|
the Agent tool's own enum resolves the same alias to `claude-sonnet-5-5`.
|
||||||
|
T1 param rewrite + alias: 404. T2 spawn rewrite + alias: 404. T3b spawn
|
||||||
|
rewrite + full id `claude-sonnet-5-5`: OK, every step answered by
|
||||||
|
claude-sonnet-5-5 at effort medium (turn.step by agentId also OK).
|
||||||
|
Rule for the mod: ALWAYS write full model ids from its own alias -> id
|
||||||
|
table in the config (one place to bump when a tier ships). The Agent tool
|
||||||
|
schema accepts only aliases, so full ids can only come from the hooks.
|
||||||
|
To report upstream: hook-side alias resolution lags the tool's.
|
||||||
|
- T4 main-loop model switch (fable -> claude-sonnet-5-5/low, switch on): WORKS.
|
||||||
|
Steps 11 and 12 answered by claude-sonnet-5-5 at effort low, transcript
|
||||||
|
records agree, thinking still produced (4.6 s), conversation intact (897
|
||||||
|
messages, tools and results carried across). COST: the first step after a
|
||||||
|
switch read 0 cached tokens on a ~260k context (cache is per model), the
|
||||||
|
next step read 237k. Switching BACK to fable at step 14 read 263k cached
|
||||||
|
tokens: the fable cache survived three sonnet steps (per-model caches,
|
||||||
|
1 h TTL), so the return is free. Every switch INTO another model pays one
|
||||||
|
cold-cache step on the full context. Consequence for the design: main-loop
|
||||||
|
model switches only for spans long enough to amortize (many mechanical
|
||||||
|
steps), never per tool call; short mechanical work goes to a haiku
|
||||||
|
sub-agent whose context is small. Flag stays off by default.
|
||||||
|
- Open: T3a (param rewrite + full id), fable/opus ids for the table,
|
||||||
|
switching back mid-turn, behavior with thinking blocks from another model
|
||||||
|
in history (no error seen), headless `-p` run.
|
||||||
|
|
||||||
|
## Harness facts (types 2.1.292, CLI 2.1.294)
|
||||||
|
- `turn.step` (async generator) rewrites `model` and `effort` per request;
|
||||||
|
`e.agentId` set inside a sub-agent loop. Pinned: turn, index, messageCount.
|
||||||
|
- `agent.spawn` rewrites `model` (not effort); result carries `agentId`.
|
||||||
|
- `tool.call {tool:'Agent'}` sees and rewrites the call's `model` / `effort`
|
||||||
|
params; `tool.call {tool:'Skill'}` names the skill loading.
|
||||||
|
- `$.tool.register` / `$.command.register` (`immediate: true` runs mid-turn).
|
||||||
|
- `$.model.classify(text, labels)`, `$.session.usage().rateLimits`.
|
||||||
|
- Hooks run under `claude -p` too (closes the BDR-107 headless gap).
|
||||||
|
- Hook budget 10 s own code; `$` calls do not count.
|
||||||
|
- Honest limit: a hook cannot know what the NEXT request will decide to do.
|
||||||
|
Routing = declared phase (skill table, `route` tool, `/route`, prompt
|
||||||
|
rules) + conservative after-the-fact heuristics on the following step.
|
||||||
|
|
||||||
|
## Phase table (proposal, config-driven)
|
||||||
|
| phase | model | effort | when |
|
||||||
|
|---|---|---|---|
|
||||||
|
| plan | session | xhigh | brainstorm, plan, architecture, challenge synthesis, audit verdict |
|
||||||
|
| reflect | session | high | diagnosis, reading to understand, contract, review |
|
||||||
|
| orchestrate | session | medium | between dispatches, reading a report |
|
||||||
|
| escalate | session | max | `ultrathink`, stuck loop, STOP relaunch |
|
||||||
|
| judge | opus | xhigh | dispatched challengers, analyzers, audits (BDR-076) |
|
||||||
|
| implement | sonnet | medium | code from a closed plan (feater, bugfixer, …) |
|
||||||
|
| write | sonnet | medium | docs, prose from decided content |
|
||||||
|
| verify | sonnet | xhigh | verifier, security-auditor |
|
||||||
|
| mechanical | haiku | low | cp/mv, git bookkeeping, status collection, listing |
|
||||||
|
|
||||||
|
## Decisions 2026-10-08 (evening, user via AskUserQuestion)
|
||||||
|
- LOADING (supersedes "CLAUDE_CODE_PLUGIN_DIRS via settings.json + link.sh"):
|
||||||
|
PLUGIN_DIRS needs absolute paths, settings `env` has no `$HOME`
|
||||||
|
expansion, settings.json is tracked and shared across machines; a local
|
||||||
|
marketplace `add` writes an absolute path into settings.json too. Chosen:
|
||||||
|
tracked relative symlink `skills/<name>` → `../mods/<name>`; Claude Code
|
||||||
|
loads it as `<name>@skills-dir`, in place (docs plugins/loading; probe in
|
||||||
|
an isolated HOME: listed, enabled, loaded). Repo scripts walking skills/
|
||||||
|
glob `*/SKILL.md` or fixed names: unaffected. User asked why not
|
||||||
|
`~/.claude/mods`: Claude Code scans no such folder, a link there loads
|
||||||
|
nothing.
|
||||||
|
- PRECEDENCE: `ultrathink` (prompt rule) and a typed `/effort-<l>` become a
|
||||||
|
FLOOR for the main turn: derived routes may go above, never below; it
|
||||||
|
lifts a lower sticky `/route`. Main loop only.
|
||||||
|
- settings.json: the user's uncommitted `/model` change (model → opus) is
|
||||||
|
theirs to manage; never staged.
|
||||||
|
- Live 2026-10-08: Skill(effort-low) bridge answered in place, next request
|
||||||
|
`low`; `ultrathink` turn on Opus ran at `max` (engine base `medium`);
|
||||||
|
Explore without params spawned on `claude-sonnet-5-5`, all 3 steps
|
||||||
|
`medium` (no spawn/step race observed).
|
||||||
|
|
||||||
|
## Decisions 2026-10-09 (user, after the wave-1 table was shown)
|
||||||
|
- No "session model" phase: every phase names an ABSOLUTE tier (best = fable,
|
||||||
|
opus, sonnet · big = opus, fable, sonnet · work = sonnet, opus · cheap = haiku,
|
||||||
|
sonnet); a haiku session asked to plan runs on fable. Reflection and planning
|
||||||
|
always on the best available model.
|
||||||
|
- Availability = circuit breaker (turn error/refusal, engine auto switch), not
|
||||||
|
quota reading (rateLimits are account windows). Fallback fable → opus →
|
||||||
|
sonnet → haiku, effort unchanged ("plus de crédit fable → opus xhigh").
|
||||||
|
- The user never types /route: prompt default rules, dispatch push/pop
|
||||||
|
(orchestrate while agents run, previous route restored), skills table
|
||||||
|
(wave 2), optional classifier. Main upgrade allowed by default, downgrade
|
||||||
|
still gated (cold-cache cost).
|
||||||
|
- Sequence agreed: 1-C lands → user /reload-plugins → live test → commit →
|
||||||
|
wave 2.
|
||||||
|
|
||||||
|
## Wave 0 — spike (dev-mods folder, hot reload, this session)
|
||||||
|
- [x] W0.1 minimal mod: `/route` command, `route` tool, `turn.step` logging +
|
||||||
|
rewrite, `agent.spawn` rewrite, `ultrathink` → max, spinner suffix
|
||||||
|
- [x] W0.2 `claude plugin validate` clean; type-check with the header tsconfig
|
||||||
|
- [x] W0.3 facts to establish, each with its evidence (usage.model, CLAUDE_EFFORT,
|
||||||
|
debug log): (a) effort rewrite on main loop takes effect; (b) model rewrite
|
||||||
|
on main loop mid-turn: works / breaks (thinking signatures, cache, tools);
|
||||||
|
(c) sub-agent model via spawn + effort via step by agentId; (d) `/route`
|
||||||
|
immediate mid-turn; (e) load via `CLAUDE_CODE_PLUGIN_DIRS` from settings env → moved to W1.10
|
||||||
|
- [x] W0.4 record facts → journal + BDR draft; freeze wave 1 scope (facts in this file; registries pending user go)
|
||||||
|
|
||||||
|
## Wave 1 — core (repo `mods/model-router/`)
|
||||||
|
- [ ] W1.1 config loader: `~/.claude/model-router.json` (phases, agents,
|
||||||
|
skills, prompt rules, defaults); schema check; `/route reload`
|
||||||
|
- [ ] W1.2 state: per-loop phase (main + agentId map), reset at `turn.start`
|
||||||
|
to the prompt-derived phase; explicit > table > heuristic
|
||||||
|
- [ ] W1.3 `route` tool + `/route [phase|show|reload|clear]` (immediate)
|
||||||
|
- [ ] W1.4 agents: `tool.call Agent` param rewrite + `agent.spawn` model +
|
||||||
|
`turn.step` effort by agentId; covers built-ins (Explore, Plan, general-purpose)
|
||||||
|
- [ ] W1.5 skills: `tool.call Skill` → phase from the skills table; any
|
||||||
|
skill load RESETS the main route to that skill's entry phase (table,
|
||||||
|
else the frontmatter `effort:` the harness just applied), so no route
|
||||||
|
declared earlier in the turn survives a skill change silently
|
||||||
|
- [ ] W1.5b single-writer bridge for the legacy shifters (user, 2026-10-08:
|
||||||
|
doublon + silent one-way conflict): `tool.call {tool:'Skill', skill:
|
||||||
|
/^effort-/}` answers WITHOUT `next` (skill text never loaded, no pairing
|
||||||
|
rule) and translates the level into a route on the calling loop;
|
||||||
|
`skill.prompt {skill:/^effort-/}` does the same for a user-typed
|
||||||
|
`/effort-max` and returns a one-line text. The 15 citers keep working
|
||||||
|
untouched until wave 2 rewrites them to `route`. Rule: every effort
|
||||||
|
change goes through the mod's state; frontmatter values are inputs.
|
||||||
|
- [ ] W1.6 prompt rules: `ultrathink` → escalate; keyword → phase (effort up
|
||||||
|
only; never a main-loop model change without declaration)
|
||||||
|
- [ ] W1.7 visibility: Spinner suffix `· <model>/<effort>`, `$.ui.status`,
|
||||||
|
`$.ui.log` when verbose, `/route show`
|
||||||
|
- [ ] W1.8 userConfig: `mainLoopModelSwitch` (false), `verbose` (false),
|
||||||
|
`classifier` (false)
|
||||||
|
- [ ] W1.9 tests `*.test.ts` under `claude plugin test`; `claude plugin validate`
|
||||||
|
- [ ] W1.10 install: `mods/` symlink + `CLAUDE_CODE_PLUGIN_DIRS` in settings.json
|
||||||
|
env via link.sh; doctor line; README/USAGE/CHANGELOG
|
||||||
|
- [ ] W1.11 contract + GATE 0 + fresh verifier + security gate; `make test`
|
||||||
|
|
||||||
|
## Wave 2 — migration (after wave 1 proven)
|
||||||
|
- [ ] W2.1 15 skills `Skill(effort-*)` → `route` tool calls (lib/effort-shift.md rewritten)
|
||||||
|
- [ ] W2.2 remove `skills/effort-*`, `lib/effort-pins.txt`, `lib/effort-pins.sh`,
|
||||||
|
install/update steps, `effort:` frontmatter on skills and agents
|
||||||
|
- [ ] W2.3 `lib/model-gate.md` + `lib/model-check.sh` → mod rule (reflect on a
|
||||||
|
small model → raise); census tests repointed to the config table
|
||||||
|
- [ ] W2.4 docs + CHANGELOG + registries (BDR, LRN, EVAL via effort-audit.py)
|
||||||
|
|
||||||
|
## Wave 3 — optional
|
||||||
|
- [ ] W3.1 per-step heuristics (Read/Grep → +1 level next step; Agent return → orchestrate)
|
||||||
|
- [ ] W3.2 haiku classifier on `prompt.submit` (`$.model.classify`)
|
||||||
|
- [ ] W3.3 quota-aware downgrade from `$.session.usage().rateLimits`
|
||||||
|
- [ ] W3.4 A/B via `lib/effort-audit.py`
|
||||||
|
|
||||||
|
## Risks
|
||||||
|
- Main-loop model switch mid-turn unproven (W0.3b decides).
|
||||||
|
- A mod bug cuts all routing at once: fail-open (`.catch` → `next(e)`), never deny.
|
||||||
|
- Two sources of truth during wave 1 (pins + mod): mod must agree with the
|
||||||
|
pins until wave 2 removes them.
|
||||||
|
- API early access: types change between releases; pin the CLI version in the README.
|
||||||
@@ -0,0 +1,347 @@
|
|||||||
|
# PLAN — model-router wave 1-A: the mod (dispatch-ready) — REVISED r3
|
||||||
|
|
||||||
|
r3 changes (confirmation pass, 4 MAJOR + minors): explicit Agent params are
|
||||||
|
frozen on the Loop and never overridden by in-agent `route`/`Skill(effort-*)`;
|
||||||
|
`pendingPrompt` honours `e.wait`; tests assert on the `main:` line of `show`
|
||||||
|
only, with full typed inputs; haiku gets `effort: undefined`; `userMain`
|
||||||
|
beats every turn route including a typed `/effort-<l>` (the text says so);
|
||||||
|
`skill.prompt` acts only when no Skill call is in flight; route tool handles
|
||||||
|
`clear` first and answers "off" in place; `turn.step` catch is a generator;
|
||||||
|
windows keyed by resolved id; validate user entries BEFORE merging; truthful
|
||||||
|
answers; unpinned-skill reset accepted and documented.
|
||||||
|
Contract: .claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md
|
||||||
|
Wave plan + harness facts: .claude/tasks/plans/2026-10-08-model-router-mod.md
|
||||||
|
Base: the spike `~/.claude/dev-mods/385f7190-70f5-4bdd-b0d8-e4566cd412fd/model-router/hooks/register.ts`
|
||||||
|
(read it first; keep its proven hook shapes; drop its spike levers `via`,
|
||||||
|
`stepModel`, `agentsDefault`, the tool's `model`/`scope`/`mainModelSwitch`
|
||||||
|
params and the hardcoded PHASES).
|
||||||
|
API reference: `<spike>/.claude-plugin/types/claude-code/index.d.ts` (grep the
|
||||||
|
event or noun; `declare module 'claude-code/testing'` for the test kit) and
|
||||||
|
`<spike>/.claude-plugin/types/claude-code-tools/index.d.ts` (`Skill: {`,
|
||||||
|
`Agent: {`, and the Skill RESULT schema near line 5139).
|
||||||
|
|
||||||
|
r2 changes (challenge round, 3 lenses): agents table = built-ins only; one
|
||||||
|
writer per agent model (spawn), no per-step model rewrite unless an in-agent
|
||||||
|
`route` call changed it and the engine did not fall back; every write goes
|
||||||
|
to the CALLING loop; `agentsNext` / `scope` / `/route agents` / tool `model`
|
||||||
|
dropped; Skill bridge answers in the tool's output shape; unconditional
|
||||||
|
skill-load reset (prompt route kept); config validated at load, refs
|
||||||
|
resolved once; state in the `register` closure, cloned defaults; `/route off`
|
||||||
|
kill switch; queued `ultrathink` promoted to its own turn; AC1/AC5 amended.
|
||||||
|
|
||||||
|
## Files
|
||||||
|
- [ ] mods/model-router/.claude-plugin/plugin.json — `{ "name": "model-router", "version": "0.1.0", "description": "<one line>", "author": { "name": "bchanot" } }`
|
||||||
|
- [ ] mods/model-router/hooks/hooks.json — `{ "modules": ["./register.ts"] }`
|
||||||
|
- [ ] mods/model-router/hooks/register.ts — the hooks module (below)
|
||||||
|
- [ ] mods/model-router/hooks/register.test.ts — `claude plugin test` suite (below)
|
||||||
|
The engine lays `./tsconfig.json` and `.claude-plugin/types/` beside a loaded
|
||||||
|
mod; both are ignored by AC1 and gitignored in wave 1-B. Never commit them.
|
||||||
|
|
||||||
|
## Config (one shape, defaults in code, optional override on disk)
|
||||||
|
```ts
|
||||||
|
type Level = 'low' | 'medium' | 'high' | 'xhigh' | 'max'
|
||||||
|
type Route = { model?: string; effort?: Level } // model = alias OR full id
|
||||||
|
type Config = {
|
||||||
|
models: Record<string, string> // alias → full id
|
||||||
|
windows: Record<string, number> // alias → context window (tokens)
|
||||||
|
phases: Record<string, Route>
|
||||||
|
agents: Record<string, string> // built-in subagentType → phase name
|
||||||
|
skills: Record<string, string> // skill name → phase name
|
||||||
|
prompt: { pattern: string; phase: string }[] // regex source, flag i
|
||||||
|
mainModelSwitch: boolean; verbose: boolean; spinner: boolean
|
||||||
|
}
|
||||||
|
```
|
||||||
|
DEFAULT_CONFIG values:
|
||||||
|
- models: haiku→`claude-haiku-4-5-20251001`, sonnet→`claude-sonnet-5-5`,
|
||||||
|
opus→`claude-opus-5-5`, fable→`claude-fable-5-1`.
|
||||||
|
- windows: `claude-haiku-4-5-20251001`→200000 (keyed by FULL id; others
|
||||||
|
unknown: absent = no check).
|
||||||
|
- phases: plan {effort xhigh}, reflect {effort high}, orchestrate {effort medium},
|
||||||
|
escalate {effort max}, judge {opus, xhigh}, implement {sonnet, medium},
|
||||||
|
write {sonnet, medium}, verify {sonnet, xhigh}, explore {sonnet, medium},
|
||||||
|
mechanical {haiku, low}. A phase without `model` keeps the loop's model.
|
||||||
|
- agents: Explore→explore, Plan→judge. NOTHING else in wave 1: every repo
|
||||||
|
agent keeps its frontmatter pin (the engine applies it); the pins move into
|
||||||
|
this table in wave 2, in the same change that deletes the frontmatter.
|
||||||
|
- skills: {} (wave 2 fills it).
|
||||||
|
- prompt: [{ pattern: '\\bultrathink\\b', phase: 'escalate' }].
|
||||||
|
- mainModelSwitch false, verbose false, spinner true. (Pass B, user
|
||||||
|
2026-10-08: verbose default off but ON for now through the override file,
|
||||||
|
spinner on, user `/route` sticky, Explore → sonnet/medium.)
|
||||||
|
|
||||||
|
`loadConfig($, log)` → `Config` (never throws):
|
||||||
|
1. `home = await $.env.get('HOME')`; path `${home}/.claude/model-router.json`;
|
||||||
|
`$.fs.exists` then `$.fs.read`; `JSON.parse`. Any failure → the defaults.
|
||||||
|
2. `mergeConfig(D, u, log)`: FIXED merge, no recursion, VALIDATE EACH USER
|
||||||
|
ENTRY BEFORE IT REPLACES A DEFAULT (an invalid user `models.sonnet` is
|
||||||
|
dropped and the default kept, so the phases on `sonnet` stay valid):
|
||||||
|
for each table (`models`, `windows`, `phases`, `agents`, `skills`) take
|
||||||
|
`u.<table>` only when it is a plain object, then per key: valid → over
|
||||||
|
the default, invalid → `log(...)` and keep the default. Scalars
|
||||||
|
(`mainModelSwitch`, `verbose`, `spinner`) taken only when boolean.
|
||||||
|
`prompt` taken only when an array; each rule validated or dropped.
|
||||||
|
Type guards on `unknown`, no `any`.
|
||||||
|
3. Validity rules (`log(...)` ALWAYS, not only verbose: a config error must
|
||||||
|
be seen once):
|
||||||
|
- models value: string matching `/^claude-[a-z0-9.-]+$/`;
|
||||||
|
- windows: key a full id (same regex), value a positive integer;
|
||||||
|
- phase: plain object; `effort` absent or in LEVELS; `model` absent, a
|
||||||
|
`models` key or a full id; at least one of the two;
|
||||||
|
- agents / skills value: a phase name (checked after phases merged);
|
||||||
|
- prompt rule: `{ pattern: string, phase: <phase name> }` whose pattern
|
||||||
|
compiles (`new RegExp(p, 'i')` in try/catch).
|
||||||
|
Lookups use `Object.hasOwn`, never bare indexing on user keys.
|
||||||
|
4. Returns `structuredClone`d data: the defaults constant is never handed
|
||||||
|
out by reference.
|
||||||
|
`compileRules(cfg)` → `{ re: RegExp; phase: string }[]` once per load.
|
||||||
|
`resolveModel(cfg, name)`: `Object.hasOwn(cfg.models, name) ? cfg.models[name] : name`.
|
||||||
|
`isModelName(cfg, v)`: a `models` key or the full-id regex. `isLevel(v)`.
|
||||||
|
|
||||||
|
## State — ONE object built inside `register`, passed to every helper
|
||||||
|
```ts
|
||||||
|
type Source = 'user' | 'model' | 'skill' | 'prompt' | 'slash'
|
||||||
|
type Routed = { phase: string; route: Route; source: Source }
|
||||||
|
type Loop = {
|
||||||
|
effort?: Level; model?: string // routed by the table or an in-agent call
|
||||||
|
spawnModel: string; frozen: boolean // engine's model at spawn; fork/workflow
|
||||||
|
explicitModel: boolean; explicitEffort: boolean // Agent params given → axis frozen
|
||||||
|
}
|
||||||
|
type State = {
|
||||||
|
cfg: Config; rules: Rule[]
|
||||||
|
userMain: Routed | null // /route by the user; sticky until /route clear
|
||||||
|
turnMain: Routed | null // tool / skill / prompt / slash; dropped at turn end
|
||||||
|
pendingPrompt: Routed | null // prompt rule typed mid-turn, promoted next turn
|
||||||
|
loops: Map<string, Loop> // agentId → that loop's routing
|
||||||
|
explicitEffort: Map<string, Level> // Agent tool_use_id → explicit effort param
|
||||||
|
skillCalls: number // Skill tool calls in flight (hook 4 ± around next)
|
||||||
|
off: boolean // /route off: every hook passes through
|
||||||
|
lastMain: string // "model/effort" of the last main step (spinner)
|
||||||
|
}
|
||||||
|
```
|
||||||
|
`newState(cfg)` builds it; `register` calls it once; `session.start` reloads
|
||||||
|
`cfg` + `rules` into it; `session.end` rebuilds it (`/clear` fires no
|
||||||
|
`session.start`, so sticky routes must not survive a clear).
|
||||||
|
Effective main route `mainRoute(st)`: `st.userMain ?? st.turnMain`. ONE
|
||||||
|
order, transitive: user `/route` (sticky) > the latest turn route (tool,
|
||||||
|
skill, slash, prompt all share `turnMain`; last writer wins) > session. A
|
||||||
|
typed `/effort-<l>` while a sticky route is in force does not apply; its
|
||||||
|
text says so (hook 5).
|
||||||
|
Loop lookups: `loopOf(st, e.agentId)`; a missing entry is created on first
|
||||||
|
write as `{ spawnModel: '', frozen: false, explicitModel: false,
|
||||||
|
explicitEffort: false }`. In-agent writes (hooks 3 and 4) never set an axis
|
||||||
|
whose `explicit*` flag is true: an explicit Agent param wins for the whole
|
||||||
|
run.
|
||||||
|
|
||||||
|
## Hooks
|
||||||
|
Rule for failures: hooks that only observe or rewrite carry
|
||||||
|
`.catch(($, e, next) => next(e))`; `turn.step` streams, so its catch is the
|
||||||
|
generator form `async function* ($, e, next) { return yield* next(e) }` (a
|
||||||
|
plain function there is a type error). The four hooks that ANSWER without `next`
|
||||||
|
(command.run, the route tool, the Skill `effort-*` bridge, `skill.prompt`)
|
||||||
|
carry a `.catch` that answers in place: `{ text: 'route failed (<kind>)' }`,
|
||||||
|
`{ result: 'route failed (<kind>); nothing routed' }`, and for the two skill
|
||||||
|
hooks `next(e)` (the skill then loads normally — a safe fallback). State is
|
||||||
|
mutated only AFTER the input validated.
|
||||||
|
1. `session.start`: `st.cfg = await loadConfig(...)`, `st.rules = compileRules`;
|
||||||
|
`registerTool($, st)` (`$.tool.register({ name: 'route', description,
|
||||||
|
inputSchema })`: properties `phase` (enum = Object.keys(st.cfg.phases)),
|
||||||
|
`effort` (enum LEVELS), `clear` (boolean); no `required`). The description
|
||||||
|
tells the model: declare the phase before a span changes nature; acts on
|
||||||
|
the calling loop only; no model choice here. `$.command.register({ name:
|
||||||
|
'route', description, argumentHint: '[show|clear|off|on|reload|<phase>|
|
||||||
|
model=<alias|id> effort=<level>|switch on|off|verbose on|off]', immediate:
|
||||||
|
true })` in try/catch (log on failure, keep going). `$.ui.status(statusLine(st))`.
|
||||||
|
2. `command.run {command:'route'}` → `{ text: handleCommand($, st, e.args) }`:
|
||||||
|
`show`/empty → `show(st)`; `clear` → userMain = turnMain = pendingPrompt =
|
||||||
|
null; `off` / `on` → st.off; `reload` → loadConfig + compileRules +
|
||||||
|
`registerTool` again (the phase enum follows the config) + 'config
|
||||||
|
reloaded' + show; `switch on|off` → cfg.mainModelSwitch; `verbose on|off`;
|
||||||
|
otherwise `parseRoute(st.cfg, args)`: a phase name, or tokens `model=<m>`
|
||||||
|
/ `effort=<l>` / bare alias / bare level, each validated by `isModelName`
|
||||||
|
/ `isLevel` → `st.userMain = { phase, route, source: 'user' }`; any
|
||||||
|
unknown token → error text listing the phases and the levels. Never
|
||||||
|
calls `next`.
|
||||||
|
3. `tool.call {tool:'mcp__model-router__route'}` → `handleRouteTool`, in
|
||||||
|
this order: (i) `st.off` → `{ result: 'model-router is off (/route on to
|
||||||
|
resume); nothing routed' }`; (ii) `clear` → main: `turnMain = null`; agent:
|
||||||
|
unset the loop's `effort`/`model` → `{ result: 'route cleared for <loop>' }`;
|
||||||
|
(iii) validate: `phase` given and not a `phases` key → `{ deny: 'unknown
|
||||||
|
phase "<p>"; phases: …' }` (even with a valid `effort`); `effort` given and
|
||||||
|
not a level → deny naming the levels; neither given → deny. (iv) route =
|
||||||
|
`{ ...phases[phase], ...(effort ? { effort } : {}) }` (an explicit effort
|
||||||
|
overrides the phase's). Target = the CALLING loop: main → `turnMain = {
|
||||||
|
phase: phase ?? 'effort-' + effort, route, source: 'model' }`; agent →
|
||||||
|
`loop.effort = route.effort` unless `loop.explicitEffort`; `loop.model =
|
||||||
|
route.model` unless `loop.explicitModel` (applied at step only under the
|
||||||
|
fallback guard, never on a frozen loop). (v) Answer on the EFFECTIVE
|
||||||
|
outcome: main with a sticky `userMain` → `'recorded <phase> for this turn,
|
||||||
|
but a sticky /route <userMain.phase> is in force; it wins until /route
|
||||||
|
clear'`; otherwise `'routed <main|this agent> to <phase>: effort <level|
|
||||||
|
unchanged>, model <resolved id|unchanged>'`, and when a model is part of
|
||||||
|
the route on main while `mainModelSwitch` is off, say `model unchanged
|
||||||
|
(switch off)`. Verbose → log.
|
||||||
|
4. `tool.call {tool:'Skill'}`:
|
||||||
|
a. `e.skill` matches `/^effort-(low|medium|high|xhigh|max)$/` → the
|
||||||
|
CALLING loop: main → `turnMain = { phase: e.skill, route: { ...st.turnMain?.route, effort }, source: 'skill' }`
|
||||||
|
(effort merged over the current turn route, last loaded wins); agent →
|
||||||
|
`loopOf(...).effort = level` unless `loop.explicitEffort` (entry created
|
||||||
|
if missing: an untabled agent's shift must still land). Answer WITHOUT
|
||||||
|
`next`, in the Skill tool's output shape (claude-code-tools ~5139; a
|
||||||
|
string result is refused and the skill would load):
|
||||||
|
`{ result: { success: true, commandName: e.skill, status: 'inline' },
|
||||||
|
context: [<line>] }` where `<line>` states the effective outcome:
|
||||||
|
`'model-router: effort → <l> for this loop from the next request on; the
|
||||||
|
effort-<l> skill text was not loaded.'`, or when main has a sticky
|
||||||
|
`userMain`: `'model-router: effort-<l> recorded, but a sticky /route
|
||||||
|
<phase> is in force and wins until /route clear.'`, or when the agent
|
||||||
|
axis is explicit: `'model-router: this agent was dispatched with an
|
||||||
|
explicit effort; the shift does not apply.'`
|
||||||
|
b. Any other skill: `st.skillCalls += 1` before `next(e)`, `-= 1` after
|
||||||
|
(try/finally). Main → `turnMain = null` when its source is 'model',
|
||||||
|
'skill' or 'slash' (a 'prompt' route such as `ultrathink` stays unless
|
||||||
|
the skill has a table entry); `Object.hasOwn(cfg.skills, e.skill)` →
|
||||||
|
`turnMain = { phase, route, source: 'skill' }`. Agent → unset the loop's
|
||||||
|
non-explicit `effort`/`model`; table entry → `loop.effort = route.effort`
|
||||||
|
unless explicit (never model). Accepted change vs the legacy shifters:
|
||||||
|
loading an UNPINNED skill after a shift returns main to the harness
|
||||||
|
level (orchestrators already re-assert after a nested skill,
|
||||||
|
lib/effort-shift.md § Re-assert). Then `return next(e)`.
|
||||||
|
5. `skill.prompt {skill: /^effort-/}`: `st.skillCalls > 0` (reached through
|
||||||
|
the bridge's fallback or a Skill call) → `next(e)`. Otherwise it is a
|
||||||
|
user-typed `/effort-<l>` (or a preload, unsupported: treated the same):
|
||||||
|
level parse; `turnMain = { phase: e.skill, route: { effort }, source:
|
||||||
|
'slash' }`; return `{ text: <line> + '\n' + e.text }` (prepend, never
|
||||||
|
replace: args ride in the text) where `<line>` is `'Effort shifted to <l>
|
||||||
|
by model-router for this turn.'` or, with a sticky `userMain`, `'Effort
|
||||||
|
<l> recorded; the sticky /route <phase> wins until /route clear.'`.
|
||||||
|
Unknown suffix → `next(e)`.
|
||||||
|
6. `tool.call {tool:'Agent'}`: `isLevel(e.effort) && typeof e.tool_use_id ===
|
||||||
|
'string'` → `st.explicitEffort.set(e.tool_use_id, e.effort)`; always
|
||||||
|
`return next(e)` unchanged (params are never rewritten).
|
||||||
|
7. `agent.spawn`: `frozen = e.fork || e.workflow !== undefined`. Route:
|
||||||
|
`!frozen && e.provider.plugin === 'engine' && Object.hasOwn(cfg.agents, e.subagentType)`
|
||||||
|
→ `cfg.phases[cfg.agents[e.subagentType]]`, else none. Model rewrite ONLY
|
||||||
|
when route?.model is set AND `e.model === undefined` (an explicit param
|
||||||
|
wins): `next({ ...e, model: resolveModel(cfg, route.model) })`, else
|
||||||
|
`next(e)`. On a non-deny result with `agentId`: `loops.set(agentId, {
|
||||||
|
spawnModel: result.model, frozen, explicitModel: e.model !== undefined,
|
||||||
|
explicitEffort: given !== undefined, effort: given ? undefined : route?.effort })`
|
||||||
|
where `given = explicitEffort.get(e.tool_use_id)` (then deleted). `model`
|
||||||
|
is NOT stored at spawn: the engine already runs the agent on it. Verbose
|
||||||
|
log `spawn <type>: <e.model ?? '-'> → <result.model>`. Known limit,
|
||||||
|
documented in a comment: `provider.plugin === 'engine'` is the best
|
||||||
|
available test for a built-in at spawn; a user agent named `Explore` in a
|
||||||
|
foreign project would also match (wave 1 impact: sonnet/medium on it).
|
||||||
|
8. `turn.step` (async generator). `st.off` → log when verbose, `yield* next(e)`.
|
||||||
|
Agent loop (`e.agentId`): `loop = loops.get(...)`; `effort = loop?.effort ??
|
||||||
|
e.effort`; `model = loop?.model && !loop.frozen && e.model === loop.spawnModel
|
||||||
|
? resolveModel(cfg, loop.model) : e.model` (an engine fallback — `e.model`
|
||||||
|
differs from the spawn model — is never fought). Main: `set = mainRoute(st)`;
|
||||||
|
`effort = set?.route.effort ?? e.effort`; `model = set?.route.model &&
|
||||||
|
cfg.mainModelSwitch && windowOk ? resolved : e.model`, where `resolved =
|
||||||
|
resolveModel(cfg, set.route.model)` and `windowOk` = no
|
||||||
|
`cfg.windows[resolved]` entry (windows are keyed by FULL id; the defaults
|
||||||
|
key haiku's full id) or `(await $.session.usage()).context.tokens` is a
|
||||||
|
number below it; an absent `tokens` or a failed `usage()` → no switch,
|
||||||
|
logged once per turn. If the model actually sent starts with
|
||||||
|
`claude-haiku`, send `effort: undefined` (omit it entirely; haiku takes
|
||||||
|
none and a hook-set effort on it is unproven). Main → `lastMain =
|
||||||
|
'<model without claude->/<effort>'`, `$.ui.status(statusLine(st))`. Verbose →
|
||||||
|
log before (`step <i> <loop>: <from> → <to>`) and after (`answered by
|
||||||
|
<usage.model>`). `const r = yield* next(changed ? { ...e, model, effort } : e); return r`.
|
||||||
|
9. `prompt.submit`: `e.origin.kind !== 'composer'` → `next(e)`. First rule in
|
||||||
|
`st.rules` whose `re.test(e.text)` → `routed = { phase, route, source: 'prompt' }`;
|
||||||
|
`e.turnId !== undefined && e.wait` (typed mid-turn and asked to wait: it
|
||||||
|
belongs to the NEXT turn) → `st.pendingPrompt = routed`; otherwise
|
||||||
|
(idle, or delivered INTO the running turn) → `st.turnMain = routed`.
|
||||||
|
`return next(e)`.
|
||||||
|
10. `turn.complete`: `e.agentId` → `loops.delete(e.agentId)`. Main →
|
||||||
|
`turnMain = pendingPrompt; pendingPrompt = null; explicitEffort.clear();
|
||||||
|
lastMain = ''`; `$.ui.status(statusLine(st))`. `return next(e)`.
|
||||||
|
11. `ui.render {component:'Spinner'}`: `cfg.spinner && lastMain` →
|
||||||
|
`next({ ...e, props: { ...e.props, suffix: ' · ' + lastMain + '…' } })` else `next(e)`.
|
||||||
|
12. `session.end`: `Object.assign(st, newState(st.cfg))` (keeps the loaded
|
||||||
|
config, drops every route and map). `return next(e)`.
|
||||||
|
`statusLine(st)`: `'route: ' + (st.off ? 'off' : describe(mainRoute(st)) )`
|
||||||
|
where `describe` = `'<source> <phase>'` or `'session defaults'`, plus
|
||||||
|
`' · switch on'` when `cfg.mainModelSwitch`.
|
||||||
|
`show(st)`: main (effective, with its source), off/on, switch, verbose,
|
||||||
|
spinner, live loops count, phases as `name=<resolved id|session>/<effort>`
|
||||||
|
(resolved ids printed, so a wrong `models` entry is visible), config source
|
||||||
|
line (`defaults` or the override path).
|
||||||
|
|
||||||
|
## Tests (register.test.ts, `import { test, expect } from 'claude-code/testing'`)
|
||||||
|
Read the kit's declarations first (`declare module 'claude-code/testing'`):
|
||||||
|
`test(name, async ($, on) => …)`; events are fired as calls on `$` with the
|
||||||
|
event's FULL input (the kit's `$` is `EngineCall<E> = (e: Args<E>)`, and
|
||||||
|
AC2 type-checks the test file): `$.command.run({ command: 'route', args:
|
||||||
|
'show', origin: { kind: 'composer' }, presentation: <a valid value from the
|
||||||
|
types> })`, `$.prompt.submit({ text: 'ultrathink please', wait: false,
|
||||||
|
origin: { kind: 'composer' } })`, `$.agent.spawn({ tool_use_id: 't1',
|
||||||
|
prompt: 'x', description: 'd', subagentType: 'Explore', provider: { plugin:
|
||||||
|
'engine', tier: 'core' }, parentModel: 'claude-fable-5-1', background: false,
|
||||||
|
fork: false })`, `$.tool.call({ tool: 'Skill', skill: 'effort-low' })`. Read
|
||||||
|
each input type and fill every required field; never relax a test to dodge
|
||||||
|
a type. Establish from the kit whether `session.start` fires at load; if
|
||||||
|
not, fire `$.session.start(...)` first in every test. An event whose hook
|
||||||
|
calls `next` needs a BOTTOM hook registered by the test through its `on`
|
||||||
|
(the kit's bottom throws otherwise), e.g. `on('prompt.submit', ($, e) => ({
|
||||||
|
text: e.text }))`, `on('agent.spawn', ($, e) => ({ model: e.model, agentId:
|
||||||
|
'a1' }))`. A helper `mainLine(text)` returns the `main:` line of `show`;
|
||||||
|
EVERY assertion on a route reads that line only (the phases listing always
|
||||||
|
contains every id and level, so matching the whole text proves nothing).
|
||||||
|
Tests (one per contract item 3a-3g):
|
||||||
|
- 3a `Skill(effort-low)` via `$.tool.call`: resolves with `result.success ===
|
||||||
|
true` and `result.commandName === 'effort-low'`; a bottom `on('tool.call',
|
||||||
|
{ tool: 'Skill' })` registered by the test is NOT reached (flag); `mainLine`
|
||||||
|
contains `skill effort-low` and `low`.
|
||||||
|
- 3b route tool `{ phase: 'orchestrate' }` → `mainLine` contains `model
|
||||||
|
orchestrate` and `medium`.
|
||||||
|
- 3c `/route clear` → `mainLine` contains `session defaults`.
|
||||||
|
- 3d `/route bogus` → text contains `unknown` and every phase name.
|
||||||
|
- 3e `$.prompt.submit` with `ultrathink`, `wait: false`, no `turnId` →
|
||||||
|
`mainLine` contains `prompt escalate`.
|
||||||
|
- 3f `/route model=sonnet` → `mainLine` contains `claude-sonnet-5-5`; `/route
|
||||||
|
model=claude-x-9` → `mainLine` contains `claude-x-9`; `/route model=sonet` →
|
||||||
|
text contains `unknown`.
|
||||||
|
- 3g spawn path: `$.agent.spawn(...)` for `Explore` without `model`, bottom
|
||||||
|
hook captures `e.model === 'claude-sonnet-5-5'`; with `model: 'opus'` given
|
||||||
|
→ captured `e.model === 'opus'`.
|
||||||
|
No fs, network or process in tests: the defaults path is the one exercised.
|
||||||
|
|
||||||
|
## Edge cases
|
||||||
|
- `$.command.register` throws when `/route` is taken → log, keep the tool.
|
||||||
|
- `loadConfig` never throws out of `session.start`; invalid entries dropped with a log.
|
||||||
|
- `/clear` → `session.end` rebuilds the state; `/route reload` re-registers the tool.
|
||||||
|
- Remote agents raise no `turn.complete`; denied Agent calls never spawn: both
|
||||||
|
maps are bounded by `explicitEffort.clear()` at main turn end and `loops`
|
||||||
|
entries only for started agents (a leak of a few entries per turn is accepted).
|
||||||
|
- `loops.delete` at an agent's `turn.complete`: a resumed agent (SendMessage,
|
||||||
|
woken teammate) runs its later turns at the engine's effort. Accepted in
|
||||||
|
wave 1 (built-ins only); revisit with the pins in wave 2.
|
||||||
|
- Spawn-vs-first-step race: the loop entry is set after `next(e)` resolves,
|
||||||
|
so step 0 of a tabled built-in may run at the engine's effort. Accepted in
|
||||||
|
wave 1 (Explore/Plan only); wave 2 verifies the ordering before pins move.
|
||||||
|
- Unpinned agents (general-purpose, interviewer, client-handover-writer) no
|
||||||
|
longer inherit a shifted level: the bridge does not move the harness
|
||||||
|
level. Accepted: BDR-077 already requires explicit call-site params for
|
||||||
|
built-ins; the two inline-load agents run on main's own route.
|
||||||
|
- The word `any` must not appear as a TypeScript type in register.ts
|
||||||
|
(AC5 greps `: any`, `<any>`, `as any`).
|
||||||
|
|
||||||
|
## Disposition (STEP 0.6)
|
||||||
|
- honors BDR-066/076/077 (tiers): wave 1 touches only the two built-ins that
|
||||||
|
carry no pin (Explore sonnet/medium, Plan opus/xhigh); every repo agent
|
||||||
|
keeps its frontmatter as the single writer; explicit call-site params win.
|
||||||
|
- honors BDR-107/108: levels and aliases unchanged; aliases stay the config's
|
||||||
|
vocabulary, full ids are resolved by the mod (LRN-203, BLK-029).
|
||||||
|
- LRN-180/181 made moot: the bridge answers `Skill(effort-*)` itself, no
|
||||||
|
pairing rule. "Last loaded wins" holds among pinned skills and shifts; the
|
||||||
|
one behaviour change, accepted: loading an UNPINNED skill after a shift
|
||||||
|
returns main to the harness level (orchestrators re-assert after nested
|
||||||
|
skills already, lib/effort-shift.md § Re-assert).
|
||||||
|
- LRN-204: main-loop model switch behind `mainModelSwitch` (default false) and
|
||||||
|
a context-window guard.
|
||||||
|
- BDR-044 not contradicted: the mod routes model/effort, never skills.
|
||||||
|
- Deferred (minor, challenge r1): a ceiling on model-declared efforts; the
|
||||||
|
spawn-vs-step race.
|
||||||
@@ -0,0 +1,166 @@
|
|||||||
|
# PLAN — model-router wave 1-B2: active in every session, tests, doctor (dispatch-ready)
|
||||||
|
Contract: .claude/tasks/contracts/2026-10-08-model-router-wiring-1835.md
|
||||||
|
Repo root: /Users/b.chanot/Documents/claude (branch feature/model-router-mod).
|
||||||
|
|
||||||
|
## Facts this plan rests on (verified 2026-10-08)
|
||||||
|
- Claude Code loads a folder holding `.claude-plugin/plugin.json` under
|
||||||
|
`~/.claude/skills/` as `<name>@skills-dir`, in place, live at the next
|
||||||
|
session start or `/reload-plugins` (docs: plugins/loading "In-place and
|
||||||
|
copied plugins"; probe in an isolated HOME: listed, enabled, loaded).
|
||||||
|
- `~/.claude/skills` is already a symlink to the repo's `skills/` (link.sh).
|
||||||
|
- Repo scripts that walk `skills/` glob `*/SKILL.md` or fixed paths
|
||||||
|
(doctor.sh, lib/skill-routing-census.py, the census suites);
|
||||||
|
lib/profile.sh only moves entries named in a profile. An entry without
|
||||||
|
SKILL.md is never counted, moved or flagged.
|
||||||
|
- The engine lays `<mod>/tsconfig.json` (extends the types) and
|
||||||
|
`<mod>/.claude-plugin/types/` (own `.gitignore` holding `*`) when a mod
|
||||||
|
loads; today `mods/model-router/tsconfig.json` shows as untracked.
|
||||||
|
- settings.json carries the user's uncommitted `/model` change: never
|
||||||
|
stage, edit or restore it.
|
||||||
|
|
||||||
|
## Files
|
||||||
|
- [ ] `skills/model-router` — new RELATIVE symlink: from the repo root,
|
||||||
|
`ln -s ../mods/model-router skills/model-router`. Nothing else in skills/.
|
||||||
|
- [ ] `.gitignore` — append a block:
|
||||||
|
```
|
||||||
|
# mods/: files the engine lays beside a loaded mod (editor types)
|
||||||
|
mods/*/tsconfig.json
|
||||||
|
mods/*/.claude-plugin/types/
|
||||||
|
```
|
||||||
|
Check first that no existing pattern ignores `skills/model-router` or
|
||||||
|
the tracked mod files (contract AC1/AC2 oracles).
|
||||||
|
- [ ] `lib/tests/mods.test.sh` — new suite, style of the existing suites
|
||||||
|
(read lib/tests/effort-pins.test.sh first and mirror its header, helpers
|
||||||
|
and summary). Behaviour:
|
||||||
|
- `ROOT="${MODS_ROOT:-<repo root from the script path>}"`.
|
||||||
|
- Collect `$ROOT/mods/*/.claude-plugin/plugin.json`; none → FAIL
|
||||||
|
("no mod found") so the suite can never pass vacuously.
|
||||||
|
- Per mod dir `<name>`: (1) the manifest `name` (python3 json, argv —
|
||||||
|
never string-spliced) equals the folder name; (2) `$ROOT/skills/<name>`
|
||||||
|
is a symlink whose `readlink` is exactly `../mods/<name>`; (3) when
|
||||||
|
`command -v claude` succeeds: `claude plugin validate "$ROOT/mods/<name>"`
|
||||||
|
prints `Validation passed` and no `warning` (case-insensitive);
|
||||||
|
(4) same condition: `claude plugin test "$ROOT/mods/<name>"` exits 0.
|
||||||
|
- `claude` absent → one `SKIP: claude CLI not found — validate/test not
|
||||||
|
run` line; checks (1)-(2) still run and decide the exit code.
|
||||||
|
- Exit 1 on any failure, 0 otherwise; one PASS/FAIL line per check and a
|
||||||
|
final count line.
|
||||||
|
- shellcheck clean. No network, no writes outside a `mktemp -d` if any
|
||||||
|
scratch is needed (none expected).
|
||||||
|
- [ ] `doctor.sh` — new section `── Mods ──`, placed right after the
|
||||||
|
"Vendored skills" section (read lines 120-160 first; mirror its
|
||||||
|
`echo ""` / heading / pass-warn-fail-info style). For each
|
||||||
|
`$REPO/mods/*/` holding `.claude-plugin/plugin.json` (`<name>` = folder):
|
||||||
|
- link `$HOME/.claude/skills/<name>`: `readlink -f` equal to
|
||||||
|
`$REPO/mods/<name>` → `pass "mod <name>: loading link ~/.claude/skills/<name>"`;
|
||||||
|
missing → `fail "mod <name>: ~/.claude/skills/<name> MISSING — git checkout skills/<name>, then make link"`;
|
||||||
|
elsewhere → `warn`. Do NOT call `check_symlink` (it feeds the core-link
|
||||||
|
counter `_LINK_PASS` / `_EXPECTED_LINKS`).
|
||||||
|
- `command -v claude` → `claude plugin list --json` parsed with python3
|
||||||
|
(argv/stdin, no splicing): id `<name>@skills-dir` with `enabled: true`
|
||||||
|
→ `pass "mod <name>: loaded as <name>@skills-dir"`; present but
|
||||||
|
disabled → `warn "... disabled (enabledPlugins \"<name>@skills-dir\": false)"`;
|
||||||
|
absent → `warn "... not listed — new session or /reload-plugins"`.
|
||||||
|
`claude` missing → `info "claude CLI not found — load state not checked"`.
|
||||||
|
- `$HOME/.claude/<name>.json` present → `python3 -m json.tool` (quiet)
|
||||||
|
→ `pass "mod <name>: override ~/.claude/<name>.json parses"` or
|
||||||
|
`fail "... invalid JSON"`; absent → nothing.
|
||||||
|
- No mod at all → `info "no mods"`.
|
||||||
|
- [ ] `CLAUDE.md` (project, repo root) — new section `## mods/ — function-hooks
|
||||||
|
plugins (Claude Code mods)` placed after the graphify section, terse
|
||||||
|
English in the file's own style, at most ~14 lines, covering: what lives
|
||||||
|
in `mods/<name>/`; it loads through the tracked relative symlink
|
||||||
|
`skills/<name>` → `../mods/<name>` as `<name>@skills-dir` (in place, live
|
||||||
|
at the next session or `/reload-plugins`); why not
|
||||||
|
`CLAUDE_CODE_PLUGIN_DIRS` (absolute path, settings `env` has no `$HOME`
|
||||||
|
expansion, settings.json is tracked) nor a local marketplace (its `add`
|
||||||
|
writes an absolute path into settings.json); engine-laid
|
||||||
|
`tsconfig.json` + `.claude-plugin/types/` are gitignored; optional user
|
||||||
|
config `~/.claude/<name>.json`; tests `make test suite=lib/tests/mods.test.sh`
|
||||||
|
(validate + `claude plugin test`); turn a mod off with
|
||||||
|
`"<name>@skills-dir": false` in `enabledPlugins`; a dev copy loaded with
|
||||||
|
`--plugin-dir` or the hot-reload folder shadows the skills-dir copy
|
||||||
|
(same name, session-only wins).
|
||||||
|
|
||||||
|
## Verify (executor pastes outputs)
|
||||||
|
`ls -l skills/model-router`; `git check-ignore -v mods/model-router/tsconfig.json`;
|
||||||
|
`make test suite=lib/tests/mods.test.sh`; the contract AC3 positive control;
|
||||||
|
`bash doctor.sh | sed -n '/── Mods ──/,/^$/p'`; `shellcheck lib/tests/mods.test.sh doctor.sh`;
|
||||||
|
`git status --short` (settings.json still ` M`, untouched); then from the
|
||||||
|
repo root `bash ~/.claude/lib/gates.sh run .claude/tasks/contracts/2026-10-08-model-router-wiring-1835.md`.
|
||||||
|
|
||||||
|
## Edge cases
|
||||||
|
- The engine-laid `mods/model-router/tsconfig.json` already exists on disk:
|
||||||
|
after the `.gitignore` change it must disappear from `git status`.
|
||||||
|
- doctor runs without `claude` on PATH (Linux box): info line, no failure.
|
||||||
|
- A second mod later: the suite and doctor loop over `mods/*/` already.
|
||||||
|
- A hot-reload or `--plugin-dir` copy of the same mod shadows the
|
||||||
|
skills-dir copy in that session; doctor reads `claude plugin list` from a
|
||||||
|
fresh process, which sees only the skills-dir copy.
|
||||||
|
|
||||||
|
## r2 — challenge round (3 lenses, 0 BLOCKER, 5 MAJOR): BINDING, overrides the sections above where they conflict
|
||||||
|
W1. ORDER: this plan runs AFTER the floor plan (B1) is committed and green
|
||||||
|
on the same branch: the suite and doctor test whatever register.ts is
|
||||||
|
on disk.
|
||||||
|
W2. `.gitignore`: add ONLY `mods/*/tsconfig.json` with the comment
|
||||||
|
`# mods/: the engine lays tsconfig.json beside a loaded mod; its
|
||||||
|
.claude-plugin/types/ ignores itself`. (The types folder carries its
|
||||||
|
own `.gitignore` holding `*`.)
|
||||||
|
W3. Link step idempotent: `[ -L skills/model-router ] || ln -s
|
||||||
|
../mods/model-router skills/model-router` (a bare `ln -s` re-run would
|
||||||
|
create a nested link inside the mod).
|
||||||
|
W4. `lib/tests/mods.test.sh` fail-soft and bounded:
|
||||||
|
- capability probe, not presence: `command -v claude` AND `claude plugin
|
||||||
|
test --help >/dev/null 2>&1`; otherwise ONE `SKIP: claude plugin test
|
||||||
|
unavailable (<reason>) — validate/test not run` line, checks (1)-(2)
|
||||||
|
still decide the exit code;
|
||||||
|
- `claude plugin validate` and `claude plugin test` captured with `2>&1`;
|
||||||
|
the validate verdict is the line matching `Validation passed`, with
|
||||||
|
`warning` searched only in that captured output;
|
||||||
|
- every CLI call bounded: `timeout 120` when available (coreutils /
|
||||||
|
`gtimeout`), else a background-and-wait guard; a timeout is a FAIL
|
||||||
|
naming it;
|
||||||
|
- no mod found → FAIL (never vacuous).
|
||||||
|
W5. doctor `── Mods ──` fail-soft under `set -euo pipefail`:
|
||||||
|
- `[ -L "$link" ] || [ -e "$link" ]` BEFORE any readlink; compare with
|
||||||
|
`[ "$link" -ef "$REPO/mods/<name>" ]` (handles logical vs physical
|
||||||
|
repo paths), never string equality on `readlink -f`;
|
||||||
|
- a missing link is `info "mod <name>: not linked (skills/<name> absent)
|
||||||
|
— git checkout skills/<name> if wanted"`, NOT `fail` (a user may
|
||||||
|
remove the link on purpose; doctor red forever would break
|
||||||
|
update-all's final doctor run);
|
||||||
|
- ONE `claude plugin list --json` call before the loop, inside
|
||||||
|
`if ! out=$(claude plugin list --json 2>/dev/null); then warn "mods:
|
||||||
|
claude plugin list failed — load state not checked"; out=""; fi`; the
|
||||||
|
python3 parse reads stdin, exits 0 always, prints `enabled|disabled|
|
||||||
|
absent|unknown` per name (any parse error → `unknown`);
|
||||||
|
- wording: `pass "mod <name>: enabled as <name>@skills-dir"` (not
|
||||||
|
"loaded": the list proves enablement, not a successful load);
|
||||||
|
`disabled` → warn naming `"<name>@skills-dir": false`; `absent` →
|
||||||
|
`warn "mod <name>: not listed as @skills-dir — run: claude plugin
|
||||||
|
validate mods/<name> (policy, manifest or name conflict)"` (a fresh
|
||||||
|
process rescans skills/, so a restart changes nothing); `unknown` →
|
||||||
|
warn "list output not understood";
|
||||||
|
- `claude` missing → nothing (doctor's Prerequisites section already
|
||||||
|
fails on it); no override-file JSON check (the mod validates its own
|
||||||
|
config and logs at session start).
|
||||||
|
W6. CLAUDE.md `## mods/` also says: the only per-machine off switch is
|
||||||
|
`"enabled": false` in `~/.claude/<name>.json` (untracked); an
|
||||||
|
`enabledPlugins` `"<name>@skills-dir": false` entry works too but lands
|
||||||
|
in the TRACKED settings.json, so it dirties every machine's tree; and
|
||||||
|
that a hot-reload / `--plugin-dir` copy of the same name shadows the
|
||||||
|
skills-dir copy for that session (docs plugins/loading "Name
|
||||||
|
conflicts"), so the dev link in `~/.claude/dev-mods/<session>/` must
|
||||||
|
be removed before `/reload-plugins` is read as a test of the skills-dir
|
||||||
|
path.
|
||||||
|
W7. `update-all.sh` runs `claude plugin update` over every listed plugin
|
||||||
|
(lines ~606-618): a `@skills-dir` entry will produce one recurring
|
||||||
|
warn there. Accepted residual, logged in TODO (an update-all edit is
|
||||||
|
out of this contract's FILE SCOPE).
|
||||||
|
|
||||||
|
## Disposition
|
||||||
|
- honors BDR-115 (mod in `mods/`, single source); amends its "Load:" line
|
||||||
|
(PLUGIN_DIRS → skills-dir link), to be recorded at capitalize.
|
||||||
|
- honors the destructive-tools rule: no recursive delete, no transfer
|
||||||
|
tool; LRN-150/LRN-171 shell hygiene (`command grep` where a shim can
|
||||||
|
interfere is not needed here: plain bash).
|
||||||
@@ -0,0 +1,478 @@
|
|||||||
|
# PLAN — model-router wave 1-C: absolute tiers, availability fallback, derived phases (dispatch-ready)
|
||||||
|
Contract: .claude/tasks/contracts/2026-10-09-model-router-tiers-1237.md
|
||||||
|
Code: mods/model-router/hooks/register.ts (read in full) and register.test.ts.
|
||||||
|
API truth: mods/model-router/.claude-plugin/types/claude-code/index.d.ts
|
||||||
|
(TurnStepInput, TurnCompleteInput + TurnCompleteReason, SessionRateLimit,
|
||||||
|
`classic.PostModelSwitch` → PostModelSwitchHookInput, `$.session.model`,
|
||||||
|
`$.model.classify`, 'claude-code/testing'), .../claude-code-tools/index.d.ts.
|
||||||
|
|
||||||
|
## Why (user, 2026-10-09)
|
||||||
|
1. "Session model" phases assumed Fable. On a haiku session, `plan` at xhigh on
|
||||||
|
haiku is wrong. Phases must name ABSOLUTE tiers.
|
||||||
|
2. When Fable has no credit left, routing must fall back (plan → opus xhigh).
|
||||||
|
3. The user never types `/route`: phases are derived (prompt wording,
|
||||||
|
dispatch spans, skills) or declared by the model.
|
||||||
|
4. Reflection and planning always get the best available model.
|
||||||
|
|
||||||
|
## Engine facts (read in the declarations today)
|
||||||
|
- `SessionRateLimit.kind` ∈ five_hour | seven_day | spend_limit: account
|
||||||
|
windows, NOT per model → availability cannot be read from quotas.
|
||||||
|
- A dead request: `turn.step` result `usage: null`, `stopReason: null`;
|
||||||
|
`turn.complete` `reason: 'error'` (retries exhausted) or `'refusal'`
|
||||||
|
(refused, no fallback model); `'aborted'` = the user interrupted.
|
||||||
|
- `classic.PostModelSwitch` fires on every main-model change with
|
||||||
|
`from_model`, `to_model`, `source` ∈ command|picker|sdk|auto|resume,
|
||||||
|
`context_tokens`, `prompt_cache_warm`.
|
||||||
|
- `turn.step` `e.model` = the id the engine resolved (session's or a
|
||||||
|
fallback's); `$.session.model()` = the main loop's model as `/model` shows.
|
||||||
|
- `$.model.classify(text, labels, { model? })` → label | undefined, rejects
|
||||||
|
on failure; default = the engine's small fast model.
|
||||||
|
|
||||||
|
## Config (additions; everything else unchanged)
|
||||||
|
```ts
|
||||||
|
type Route = { model?: string; effort?: Level } // model: TIER name, alias or full id
|
||||||
|
type PromptRule = { pattern: string; phase: string; mode?: 'floor' | 'default' }
|
||||||
|
type Config = {
|
||||||
|
…existing…
|
||||||
|
tiers: Record<string, string[]> // tier → ordered alias preference
|
||||||
|
fallback: string[] // alias order, best first; = rank
|
||||||
|
cooldownMinutes: number // breaker hold
|
||||||
|
mainUpgrade: boolean // main loop may switch UP to a phase's tier
|
||||||
|
classifier: boolean // ask the small model when no rule matched
|
||||||
|
}
|
||||||
|
```
|
||||||
|
DEFAULT_CONFIG changes:
|
||||||
|
- `tiers`: best `['fable','opus','sonnet']`, big `['opus','fable','sonnet']`,
|
||||||
|
work `['sonnet','opus']`, cheap `['haiku','sonnet']`.
|
||||||
|
- `fallback`: `['fable','opus','sonnet','haiku']`. `cooldownMinutes: 15`.
|
||||||
|
`mainUpgrade: true`. `classifier: false`.
|
||||||
|
- phases: plan `{ model: 'best', effort: 'xhigh' }`, reflect `{ best, high }`,
|
||||||
|
orchestrate `{ best, medium }`, escalate `{ best, max }`, judge `{ big, xhigh }`,
|
||||||
|
implement `{ work, medium }`, write `{ work, medium }`, verify `{ work, xhigh }`,
|
||||||
|
explore `{ work, medium }`, mechanical `{ cheap, low }`. (Keep the `model`
|
||||||
|
field name: a tier name is a model NAME the resolver understands; no new
|
||||||
|
`tier` field. AC3 greps `tier: '…'`?? NO: AC3 is written against
|
||||||
|
`model: 'best'`-style entries? → see AC3 note below.)
|
||||||
|
- prompt: `[{ pattern: '\\bultrathink\\b', phase: 'escalate', mode: 'floor' },
|
||||||
|
{ pattern: '\\b(plan|planifie|planning|brainstorm|architecture|con[cç]ois|design)\\b', phase: 'plan', mode: 'default' },
|
||||||
|
{ pattern: '\\b(pourquoi|why|explique|explain|analyse|analyze|comprendre|understand|review|audit)\\b', phase: 'reflect', mode: 'default' }]`.
|
||||||
|
`mode` absent → 'default'.
|
||||||
|
AC3 note for the executor: the contract's CHECK counts `tier: '(best|big|work|cheap)'`
|
||||||
|
in DEFAULT_CONFIG and refuses `model: '` entries there. So the PHASE type gets
|
||||||
|
an explicit `tier?: string` field: `type Route = { tier?: string; model?: string;
|
||||||
|
effort?: Level }`; a route resolves `tier` first, then `model`. Default phases
|
||||||
|
use `tier:`. `/route model=<x>` keeps writing `model` (alias or id). The route
|
||||||
|
tool keeps `phase`/`effort`/`clear` only.
|
||||||
|
Validation: `tiers` values = non-empty arrays of `models` keys (bad entries
|
||||||
|
dropped, logged); `fallback` = array of `models` keys, deduplicated, non-empty
|
||||||
|
(else default); phase `tier` must be a `tiers` key; `cooldownMinutes` positive
|
||||||
|
integer; `mode` ∈ floor|default.
|
||||||
|
|
||||||
|
## Resolution (ONE resolver, used by spawn, main plan and texts)
|
||||||
|
```ts
|
||||||
|
function availableIn(st, aliases: string[]): string | undefined
|
||||||
|
// first alias whose full id is not down (st.down.get(id) > now → down)
|
||||||
|
function resolveModel(st, name: string): string
|
||||||
|
// tier name → availableIn(tiers[name]) ?? first alias → id
|
||||||
|
// alias → id (explicit: never skipped when down); full id → itself
|
||||||
|
function resolveRoute(st, route: Route): string | undefined
|
||||||
|
// route.tier ? resolveModel(st, route.tier) : route.model ? resolveModel(st, route.model) : undefined
|
||||||
|
function rank(st, id: string): number
|
||||||
|
// index in cfg.fallback of the alias whose id prefixes `id` (strip "[1m]"); unknown → fallback.length
|
||||||
|
```
|
||||||
|
`now` comes from `$.clock.now()` (read once per hook call that needs it).
|
||||||
|
|
||||||
|
## Breaker (`st.down: Map<string, number>` full id → until ms; `st.lastMainModel: string`)
|
||||||
|
- `turn.complete`: main (`e.agentId` undefined) with `reason` ∈ error|refusal →
|
||||||
|
`markDown(st, st.lastMainModel)`; agent with that reason → `markDown(loop.model)`
|
||||||
|
(Loop gains `model: string`, the model the engine reported at spawn,
|
||||||
|
`started.model`). `aborted`/`answer` → nothing.
|
||||||
|
- `classic.PostModelSwitch` with `e.source === 'auto'` → `markDown(e.from_model)`.
|
||||||
|
- `markDown` logs ALWAYS (not only verbose): `model-router: <id> unavailable
|
||||||
|
until <HH:MM>; routing falls back`. `/route reload` and `session.end` clear
|
||||||
|
the map. `show()` gets a `down: <id> until <HH:MM>, …` or `down: none` line.
|
||||||
|
|
||||||
|
## Main-loop model decision (`mainModel` rewritten)
|
||||||
|
```
|
||||||
|
wanted = resolveRoute(st, (userMain ?? turnMain ?? turnFloor)?.route) // per-axis as today
|
||||||
|
cur = e.model
|
||||||
|
if cur is down and wanted is undefined → wanted = nextAvailable(st, cur) // fallback chain after cur's alias
|
||||||
|
if wanted undefined or sameRank(wanted, cur) → cur
|
||||||
|
if rank(wanted) < rank(cur) (better) → cfg.mainUpgrade ? wanted : cur
|
||||||
|
if rank(wanted) > rank(cur) (cheaper) → cfg.mainModelSwitch && windowOk ? wanted : cur
|
||||||
|
if cur is down and wanted defined → wanted (always: nothing to lose)
|
||||||
|
```
|
||||||
|
`st.lastMainModel = plan.model` at every main step. `sameRank` compares
|
||||||
|
aliases (so `claude-fable-5-1` vs `claude-fable-5-1[1m]` never flips).
|
||||||
|
The log/status show `→ fallback` when the breaker chose the model.
|
||||||
|
|
||||||
|
## Spawn (`spawnRoute` / `registerSpawn`)
|
||||||
|
`wanted = resolveRoute(st, route)`; explicit `e.model` still wins. Store
|
||||||
|
`loop.model = started.model`. (Explicit alias given by the caller while down:
|
||||||
|
left alone, explicit means explicit; note in the tool description.)
|
||||||
|
|
||||||
|
## Derived phases (no user action)
|
||||||
|
D1. Dispatch push/pop: in the Agent `tool.call` hook, when `e.agentId` is
|
||||||
|
undefined (main) and `st.turnMain?.source !== 'model'` written AFTER the
|
||||||
|
dispatch… simpler rule: on a main Agent call, if `st.resumeMain` is
|
||||||
|
unset, `st.resumeMain = st.turnMain ?? NONE` and `st.turnMain = { phase:
|
||||||
|
'orchestrate', route: phases.orchestrate, source: 'derived' }`. When an
|
||||||
|
agent's `turn.complete` leaves `st.loops` empty AND `st.turnMain?.source
|
||||||
|
=== 'derived'` → `st.turnMain = st.resumeMain` (NONE → null), clear
|
||||||
|
`resumeMain`. A `route` call or skill load in between replaces turnMain
|
||||||
|
(source model/skill) so the pop is skipped and `resumeMain` cleared at
|
||||||
|
the next main `turn.complete` (endMainTurn clears both). Source type gains
|
||||||
|
`'derived'`.
|
||||||
|
D2. Prompt default rules (`mode: 'default'`): write `turnMain = { phase,
|
||||||
|
route, source: 'prompt' }` (NOT the floor) — overridable by routes and
|
||||||
|
skills; floor rules unchanged. Mid-turn prompt with a default rule → only
|
||||||
|
`pendingPrompt`-like handling for FLOOR rules stays; a default rule typed
|
||||||
|
mid-turn is ignored (the running turn has its own routes).
|
||||||
|
D3. Classifier: when `cfg.classifier` and no rule matched and the prompt is
|
||||||
|
composer-origin and idle (no `turnId`): `label = await
|
||||||
|
$.model.classify(e.text.slice(0, MAX_PROMPT_SCAN), [...phaseNames,
|
||||||
|
'other'])` in try/catch; a phase label → default route (source 'prompt');
|
||||||
|
anything else → nothing. Document the cost in the config comment.
|
||||||
|
|
||||||
|
## Texts
|
||||||
|
`routedText`/`mainNote`/`show`: print the RESOLVED id and `(fallback)` when
|
||||||
|
the breaker skipped a better alias; `(tier best → claude-fable-5-1)`.
|
||||||
|
Spinner/status unchanged shape.
|
||||||
|
|
||||||
|
## Tests (register.test.ts; names must contain the contract's words)
|
||||||
|
Reuse the boot helper; full typed inputs; bottom hooks (`agent.spawn`,
|
||||||
|
`turn.complete`, `classic.PostModelSwitch` — read its input type for the
|
||||||
|
required fields; `prompt.submit`). Engine effort `high` in steps.
|
||||||
|
- `tier: a plan route upgrades a haiku session to fable at xhigh` — route tool
|
||||||
|
`plan`, step with `model: 'claude-haiku-4-5-20251001'` → bottom sees
|
||||||
|
`claude-fable-5-1` and `xhigh`.
|
||||||
|
- `downgrade: mechanical on fable keeps the model while the switch is off`.
|
||||||
|
- `fallback: an error turn on fable moves the next main step to opus` — step on
|
||||||
|
fable (sets lastMainModel), `$.turn.complete({ reason: 'error', agentId
|
||||||
|
undefined, … })`, step on fable → bottom sees `claude-opus-5-5`, effort
|
||||||
|
unchanged; then `/route reload` → step on fable stays fable.
|
||||||
|
- `breaker: PostModelSwitch auto marks the old model down` — `$.classic.PostModelSwitch({ from_model: 'claude-fable-5-1', to_model: 'claude-opus-5-5', source: 'auto', … })` → `/route show` lists `claude-fable-5-1` under `down:`.
|
||||||
|
- `spawn: Explore goes to opus while sonnet is down` — mark sonnet down
|
||||||
|
through an agent error turn (spawn Explore via bottom hook returning a1,
|
||||||
|
`$.turn.complete({ agentId: 'a1', reason: 'error' })`), spawn again → bottom
|
||||||
|
`e.model === 'claude-opus-5-5'`.
|
||||||
|
- `derived: a dispatch pushes orchestrate and pops the previous plan route`
|
||||||
|
— route tool `plan`, `$.tool.call({ tool: 'Agent', … })` with a bottom hook,
|
||||||
|
`/route show` main line shows `derived orchestrate`; `$.turn.complete({
|
||||||
|
agentId: 'a1', reason: 'answer' })` (loop registered via spawn) → main line
|
||||||
|
shows `model plan` again.
|
||||||
|
- `default rule: "planifie la migration" routes the turn to plan, a route call overrides`.
|
||||||
|
- Keep every B1 `floor` test and all earlier tests green (≥ 40 tests total).
|
||||||
|
|
||||||
|
## Constraints
|
||||||
|
- ≤ 25 logic lines per function (extract helpers: `availableIn`, `rank`,
|
||||||
|
`markDown`, `nextAvailable`, `decideMain`, `pushOrchestrate`, `popOrchestrate`,
|
||||||
|
`applyDefaultRule`, `classifyPrompt`), 80 chars/line, no `any`, state in
|
||||||
|
the closure, fail-open `.catch` with `warnOnce` on every new hook
|
||||||
|
(`classic.PostModelSwitch`), the route tool schema unchanged.
|
||||||
|
- Do not touch: the hardening (caps, `safely`, attestation), the Skill
|
||||||
|
bridge, the floor slot semantics.
|
||||||
|
- Verify: validate, the contract's tsc CHECK, `claude plugin test .`, AC3/AC4
|
||||||
|
greps, `gates.sh run` on the contract.
|
||||||
|
|
||||||
|
## Disposition
|
||||||
|
- honors BDR-115 and its amendment (one resolver, calling-loop writes,
|
||||||
|
truthful texts, per-machine config); supersedes "a phase without model
|
||||||
|
keeps the loop's model" (every default phase now names a tier).
|
||||||
|
- honors BDR-076 (dispatched judgment on opus first: `big` = opus, fable,
|
||||||
|
sonnet) and BDR-066 (execution on sonnet: `work`).
|
||||||
|
- LRN-203: hooks still write full ids (the resolver's output).
|
||||||
|
- LRN-204: downgrade on main stays gated; upgrade accepted (quality over one
|
||||||
|
cold-cache step).
|
||||||
|
- Deferred: repo agents' frontmatter pins cannot fall back (the mod does not
|
||||||
|
see them in wave 1) → wave 2 moves them into the table with tiers.
|
||||||
|
|
||||||
|
## r2 — challenge round (3 lenses, all FATAL: 4 BLOCKER, 20 MAJOR): BINDING, overrides every section above where they conflict
|
||||||
|
R1. ONE phase field for the fallback-aware choice: `Route = { tier?: string;
|
||||||
|
model?: string; effort?: Level }`. Default phases use `tier:` only (plan,
|
||||||
|
reflect, orchestrate, escalate → best; judge → big; implement, write,
|
||||||
|
verify, explore → work; mechanical → cheap). `acceptPhase` refuses a route
|
||||||
|
carrying both `tier` and `model`, and refuses a `tiers` key that collides
|
||||||
|
with a `models` alias. The earlier "model: 'best'" drafts and the "no new
|
||||||
|
tier field" sentence are VOID. `/route model=<alias|id>` keeps writing
|
||||||
|
`model`; the route tool schema is unchanged.
|
||||||
|
R2. Ids: `canonical(st, id)` = strip a trailing `[1m]`, then alias → table id,
|
||||||
|
then two-way prefix match against the table ids (`id.startsWith(tableId)
|
||||||
|
|| tableId.startsWith(id)`), else the id itself. `aliasOf(st, id)` and
|
||||||
|
`modelRank(st, id)` (= index of the alias in `fallback`, `undefined` when
|
||||||
|
unknown) work on canonical ids. The existing effort `rank` keeps its name.
|
||||||
|
Breaker keys, `agentModels` values and comparisons are canonical. When the
|
||||||
|
current main model carries `[1m]`, a resolved replacement carries `[1m]`
|
||||||
|
too (the long-context tier is a property of the session, not of the
|
||||||
|
alias); log the first time it happens (unverified live: see Verify).
|
||||||
|
R3. Model axis per slot: `routeModelName(route) = route.tier ?? route.model`;
|
||||||
|
the main model axis is the FIRST defined `routeModelName` across
|
||||||
|
userMain, turnMain, turnFloor (per-axis, like B1's effort). `resolveName`
|
||||||
|
turns that name into an available id: a `tiers` key → first alias of the
|
||||||
|
list not down → `models` id; a tier whose every alias is down →
|
||||||
|
`nextAvailable(st, cur)` (global chain) → may be `undefined` (keep cur);
|
||||||
|
an alias or full id → canonical id, never skipped (explicit means explicit).
|
||||||
|
R4. Main decision `decideMain(st, cur, wanted, ctx)` → `{ model, why }`, used
|
||||||
|
by `mainPlan` AND by every text (texts pass `cur = canonical(await
|
||||||
|
$.session.model())`); order is BINDING:
|
||||||
|
1. `st.off` → cur.
|
||||||
|
2. `rankCur = modelRank(cur)`; UNKNOWN cur (not in the table) → cur, log
|
||||||
|
once per session (`model-router: <id> unknown to the models table; no
|
||||||
|
model switch`), the breaker still applies at step 3 if it is down.
|
||||||
|
3. cur DOWN → `wanted` if defined and not down, else `nextAvailable(cur)`;
|
||||||
|
apply `windowOk`; if nothing fits → cur (why `fallback`).
|
||||||
|
4. `wanted` undefined or `aliasOf(wanted) === aliasOf(cur)` → cur.
|
||||||
|
5. `modelRank(wanted) < rankCur` (better) → `cfg.mainUpgrade &&
|
||||||
|
ctx.tokens <= cfg.upgradeMaxTokens` ? wanted (why `upgrade`) : cur
|
||||||
|
(why `upgrade skipped: context <n> tokens over <max>` or `switch off`).
|
||||||
|
6. cheaper → `cfg.mainModelSwitch && windowOk` ? wanted (why `downgrade`)
|
||||||
|
: cur (why `switch off`).
|
||||||
|
New config scalar `upgradeMaxTokens` (default 200000): an upgrade pays a
|
||||||
|
cold read of the whole context on the new model (LRN-204); above the
|
||||||
|
threshold it is skipped and logged once per turn. `ctx.tokens` comes from
|
||||||
|
`$.session.usage()` read once per main step (fail → treat as 0).
|
||||||
|
R5. Engine fallback respected: at every main step `sess = canonical(await
|
||||||
|
$.session.model())`; when `canonical(e.model) !== sess`, the engine is on
|
||||||
|
a fallback → `markDown(sess, 'engine fallback')` and `cur = e.model` (the
|
||||||
|
router never upgrades back to the model the engine just left).
|
||||||
|
R6. Breaker inputs (replace the r1 list): (a) `classic.StopFailure` with
|
||||||
|
`error` ∈ rate_limit | overloaded | billing_error | model_not_found →
|
||||||
|
`markDown(target)` where target = `st.agentModels.get(e.agent_id)` when
|
||||||
|
`e.agent_id` is set, else `st.lastPlan?.model`; other errors (context
|
||||||
|
limit = invalid_request, server_error, auth, max_output_tokens…) → nothing;
|
||||||
|
(b) R5's engine-fallback detection; (c) `classic.PostModelSwitch`: source
|
||||||
|
`command | picker | sdk` → `st.down.delete(canonical(to_model))` and reset
|
||||||
|
its strikes (the user's explicit `/model` wins); source `auto` → LOG only
|
||||||
|
(`requested_model`, from, to), never a mark (unverified semantics).
|
||||||
|
`turn.complete` `reason` is NOT a breaker input any more (context-limit
|
||||||
|
and network errors are not availability); refusal → nothing.
|
||||||
|
Backoff per canonical id: strikes 1, 2, 3… → 15, 30, 60, 120, 300 min
|
||||||
|
(cap); `model_not_found` → until `/route reload`. `markDown` logs ALWAYS:
|
||||||
|
`model-router: <id> unavailable (<reason>) until <HH:MM>; routing falls
|
||||||
|
back`. Inert while `st.off`.
|
||||||
|
Lifecycle: `/route reload` clears `down` and strikes BEFORE loading the
|
||||||
|
config (whatever the read result); `session.end` (/clear) KEEPS `down`,
|
||||||
|
strikes and `agentModels` (availability is account-wide); expired
|
||||||
|
entries are pruned at the start of any hook that reads them, with `now`
|
||||||
|
read ONLY when `st.down.size > 0` (`$.clock.now()`), passed explicitly to
|
||||||
|
the helpers (no clock read in sync text functions: they receive the
|
||||||
|
pruned map).
|
||||||
|
R7. Agent models: `st.agentModels: Map<agentId, canonicalId>` set at spawn
|
||||||
|
from `started.model` (canonicalized; an alias answered by a hook above is
|
||||||
|
mapped through the table); deleted with the loop. No `Loop.model`,
|
||||||
|
`spawnModel` or `loop.model` identifier anywhere (W1-A AC8 grep).
|
||||||
|
`spawnRoute` resolves `route.tier ?? route.model` through `resolveName`
|
||||||
|
(skips down aliases); explicit `e.model` still wins even when down.
|
||||||
|
Deferred (noted): agents without a table row and no explicit model follow
|
||||||
|
`parentModel`; forks always inherit; neither falls back in wave 1.
|
||||||
|
R8. Derived orchestrate (D1) made exact: state `pushed: { prev: Routed | null;
|
||||||
|
spawnIds: Set<string> } | null`. In the main Agent `tool.call` hook:
|
||||||
|
before `next`, if `st.turnMain?.source` is not 'model' or 'skill' and
|
||||||
|
`st.pushed` is null → `st.pushed = { prev: st.turnMain, spawnIds: new Set() }`
|
||||||
|
and `st.turnMain = { phase: 'orchestrate', route: phases.orchestrate,
|
||||||
|
source: 'derived' }`; after `next` resolves: the spawned `agentId` (from
|
||||||
|
`st.spawnByCall: Map<tool_use_id, agentId>` filled at `agent.spawn`) is
|
||||||
|
added to `pushed.spawnIds`; if NO agent was registered for this
|
||||||
|
`tool_use_id` (foreground run already finished, or denied) → nothing to
|
||||||
|
wait for from this call. Pop rule: when `pushed.spawnIds` is empty after
|
||||||
|
the Agent call returned, or when the LAST id of `pushed.spawnIds` ends
|
||||||
|
(`turn.complete` with that agentId, deleted from the set), and
|
||||||
|
`st.turnMain?.source === 'derived'` → `st.turnMain = pushed.prev`,
|
||||||
|
`st.pushed = null`. A route/skill write in between (source model/skill)
|
||||||
|
replaces turnMain; the pop then only clears `pushed`. `endMainTurn`
|
||||||
|
clears `pushed` and `spawnByCall`. In `turn.complete` for an agent, delete
|
||||||
|
the loop and the maps FIRST, inside `safely`, before any other work.
|
||||||
|
R9. Prompt default rules (D2) made safe: rules scanned in two passes (floor
|
||||||
|
rules, then default rules), each pass first match; absent `mode` →
|
||||||
|
'floor' (B1 override files keep their meaning). Default rules are SKIPPED
|
||||||
|
when the trimmed text starts with `/` (slash commands and skills route
|
||||||
|
themselves), when the same prompt carries a floor match or sets
|
||||||
|
`typedSlash` (the user's explicit level wins), or when typed mid-turn.
|
||||||
|
Patterns compile with flags `iu` and the defaults use Unicode-aware
|
||||||
|
guards instead of `\b`: `(?<![\p{L}\p{N}-])(plan|planifie|planning|
|
||||||
|
brainstorm|architecture|con[cç]ois|design)(?![\p{L}\p{N}-])` and the
|
||||||
|
reflect list likewise; the validator requires the pattern to compile
|
||||||
|
with `iu`. A default-rule route is written to `turnMain` (source
|
||||||
|
'prompt'); it never lowers (no cheap/work default rule shipped).
|
||||||
|
R10. Classifier (D3) DEFERRED to wave 2: no `classifier` key, no code.
|
||||||
|
R11. Texts: `routedText`, `effortBridge`/`mainNote`, `slashEffort`, `show`,
|
||||||
|
`statusLine`, the route tool description and `mainOnHaiku` derive their
|
||||||
|
MODEL words from `decideMain` with `cur = canonical(await
|
||||||
|
$.session.model())` (hooks are async; `show` becomes async — the
|
||||||
|
command hook awaits it); they print the decided id and `why`
|
||||||
|
(`upgrade`, `fallback`, `switch off`, `unchanged`). The tool description
|
||||||
|
says: "the main loop moves UP to a phase's tier by itself, DOWN only with
|
||||||
|
the switch on; a sub-agent's model is fixed at spawn". `show` prints:
|
||||||
|
`upgrade: on|off`, `switch (downgrade): on|off`, `down: <id> until <HH:MM>
|
||||||
|
(<reason>) …| none`, each phase as `name=<tier or model>→<resolved id>/<effort>`.
|
||||||
|
Existing test 3f (`/route model=sonnet` shows `claude-sonnet-5-5`) is
|
||||||
|
adapted: on the kit's session model the line reads `asked claude-sonnet-5-5,
|
||||||
|
keeps <cur> (switch off)`; the alias→id resolution is asserted on the
|
||||||
|
`asked` part.
|
||||||
|
R12. `st.lastPlan: Plan | null` replaces `lastMain` and `lastMainModel`; the
|
||||||
|
spinner text is derived at render; `endMainTurn` resets it.
|
||||||
|
R13. Config validation additions: `tiers` values non-empty arrays of alias
|
||||||
|
keys (bad entries dropped, logged), `fallback` deduplicated non-empty
|
||||||
|
alias list (else default, logged), `cooldownMinutes` and
|
||||||
|
`upgradeMaxTokens` positive integers, `mode` ∈ floor|default, a log at
|
||||||
|
load when `tiers.best[0] !== fallback[0]` (rank comes from `fallback`
|
||||||
|
alone). `mainModelSwitch` documented as DOWNGRADE-only in the Config
|
||||||
|
comment.
|
||||||
|
R14. Tests (≥ 43 total, names carry the contract words): keep all 30; add:
|
||||||
|
`tier` (plan on a haiku session → fable xhigh, with mock.clock installed
|
||||||
|
where the breaker is touched), `downgrade` (mechanical on fable keeps
|
||||||
|
fable, switch off), `fallback` (plan route + `$.classic.StopFailure({
|
||||||
|
error: 'rate_limit', … })` on main after a fable step → next step
|
||||||
|
`claude-opus-5-5` at xhigh; `/route reload` → fable again), `breaker`
|
||||||
|
×3 (an aborted/`invalid_request` failure never marks down; backoff expiry
|
||||||
|
via `mock.clock` advance restores fable; `/model` command
|
||||||
|
`PostModelSwitch source: 'command'` clears a down model), `engine fallback`
|
||||||
|
(`$.session.model` mocked/answered as fable while the step arrives on
|
||||||
|
opus → no upgrade back, fable marked down), `unknown` (cur
|
||||||
|
`claude-zz-9` never switches), `spawn` (Explore → opus while sonnet is
|
||||||
|
down through an agent StopFailure with `agent_id`), `derived` ×2 (push on
|
||||||
|
dispatch, pop when the spawned agent ends → plan back; a route call after
|
||||||
|
the dispatch is NOT overwritten by the pop), `default rule` ×3 (planifie
|
||||||
|
→ plan then a route call overrides; `/analyze …` typed → no rule;
|
||||||
|
`/effort-low pourquoi …` → no default rule, floor low), `per axis`
|
||||||
|
(`/route effort=low` sticky + turn `plan` tier → model axis = best).
|
||||||
|
Read `mock.clock` and how `$.session.model` is answered in the kit
|
||||||
|
(a bottom `on('session.model', …)` hook) before writing them.
|
||||||
|
R15. Disposition, superseded clauses named: floor contract AC4 "`turnMain`
|
||||||
|
only ever holds 'model' or 'skill' sources" → now also 'derived' and
|
||||||
|
'prompt'; W1-A AC6 "main-loop model changes happen only when
|
||||||
|
`mainModelSwitch` is true" → true for DOWNGRADES only; upgrades follow
|
||||||
|
`mainUpgrade` + `upgradeMaxTokens`, and the breaker/engine-fallback path
|
||||||
|
moves off a dead model unconditionally; BDR-115 (6) window guard → applied
|
||||||
|
to every switch (up, down, fallback) through `windowOk`. The tiers
|
||||||
|
contract AC5 reads "every B1/1-A criterion still holds EXCEPT the three
|
||||||
|
clauses above".
|
||||||
|
R16. Live verification after reload (orchestrator, not the executor): the
|
||||||
|
`[1m]` carry-over on a fallback id, `PostModelSwitch` `source: 'auto'`
|
||||||
|
semantics, `StopFailure` reaching the mod with `agent_id`.
|
||||||
|
|
||||||
|
## r3 — confirmation pass (FATAL(8): 1 BLOCKER, 6 MAJOR): BINDING over r2 where they conflict
|
||||||
|
S1. R5 (engine-fallback detection at every step) is REMOVED: no comparison of
|
||||||
|
`e.model` with `$.session.model()` at steps, no mark from it. The
|
||||||
|
engine's own fallback is learned ONLY through `classic.PostModelSwitch`
|
||||||
|
`source: 'auto'`, which now MARKS `canonical(from_model)` down with one
|
||||||
|
strike (15 min) when `from_model` is a table id, logging
|
||||||
|
`requested_model`, `to_model`. (R6(c) "auto → log only" is void.) A mark
|
||||||
|
is idempotent per episode: `markDown` on an id already down adds NO
|
||||||
|
strike and logs nothing; strikes count episodes (a mark after expiry).
|
||||||
|
S2. `st.sessionModel` (raw string) is read once at `session.start` through
|
||||||
|
`$.session.model()` inside try/catch ('' on failure) and refreshed in the
|
||||||
|
`PostModelSwitch` hook from `e.to_model` (any source). No other
|
||||||
|
`$.session.model()` call anywhere; texts use `st.sessionModel`.
|
||||||
|
S3. Within a turn the main model is STICKY once moved: `cur` for the decision
|
||||||
|
is `st.lastPlan?.model ?? e.model` (the model actually sent last; lastPlan
|
||||||
|
is reset at `endMainTurn` so each turn starts from the engine's model).
|
||||||
|
After an upgrade (plan → fable), a later cheaper phase in the same turn
|
||||||
|
(implement → work) goes through the CHEAPER branch against cur = fable:
|
||||||
|
gated by `mainModelSwitch` + windowOk, so no return trip and no second
|
||||||
|
cold read. After a fallback (fable down → opus), later steps stay on opus
|
||||||
|
for the turn. "Keep cur" returns the exact string last sent (`e.model`
|
||||||
|
verbatim on the first step), so `[1m]` is preserved; a resolved
|
||||||
|
replacement carries `[1m]` only when the raw session string carries it
|
||||||
|
AND the target alias is not haiku.
|
||||||
|
S4. decideMain spelled out (order binding): off → cur · unknown cur (no table
|
||||||
|
alias) → cur, logged once · cur down → first available of [wanted (if
|
||||||
|
a table id and not down), nextAvailable(cur)] that passes windowOk, else
|
||||||
|
cur · wanted undefined → cur · wanted unknown to the table (explicit full
|
||||||
|
id such as `claude-x-9`) → treated as CHEAPER (gated by `mainModelSwitch`,
|
||||||
|
windowOk) · same alias → cur · better → `mainUpgrade && tokens ≤
|
||||||
|
upgradeMaxTokens && windowOk` ? wanted : cur · cheaper → `mainModelSwitch
|
||||||
|
&& windowOk` ? wanted : cur. `ctx.tokens` from `$.session.usage()` read
|
||||||
|
once per main step (catch → 0); texts read it the same way (async), so a
|
||||||
|
text and the step agree. `nextAvailable(cur)` walks `fallback` from the
|
||||||
|
alias after cur's (unknown cur → from the top) skipping down ids; at
|
||||||
|
spawn, `nextAvailable` walks from the tier's last alias.
|
||||||
|
`model_not_found` marks show `until reload` in texts.
|
||||||
|
S5. Derived orchestrate (R8 rewritten): D1 affects BACKGROUND dispatches only.
|
||||||
|
In the main Agent `tool.call` hook: push as in R8 (source not model/skill,
|
||||||
|
`pushed` null) BEFORE `next`; after `next`: read the RESULT — `status ===
|
||||||
|
'async_launched'` → add `result.agentId` to `pushed.spawnIds`; any other
|
||||||
|
status or a deny → nothing to wait for from this call. Pop rule unchanged
|
||||||
|
(spawnIds empty after the call, or the last id's `turn.complete`); no
|
||||||
|
`spawnByCall` map. Documented: a foreground dispatch pushes and pops
|
||||||
|
inside one call, so no main step runs at orchestrate for it (fine: main
|
||||||
|
is blocked meanwhile).
|
||||||
|
S6. Breaker targets keep their value until replaced: `st.lastPlan` is NOT
|
||||||
|
reset at `endMainTurn` (only the spinner text is cleared via a separate
|
||||||
|
`st.spinner` string); `agentModels` entries are deleted at `session.end`
|
||||||
|
only, never at an agent's `turn.complete` (the StopFailure/turn.complete
|
||||||
|
order is unverified; R16 gains it).
|
||||||
|
S7. R9 patterns: compile with `iu`; on a SyntaxError retry with `i` (B1
|
||||||
|
override files keep working); a pattern failing both is dropped, logged.
|
||||||
|
S8. R11: the route tool description is STATIC text (registered once): "the
|
||||||
|
main loop moves up to a phase's tier by itself (below the context cap),
|
||||||
|
down only with the switch on; a sub-agent's model is fixed at spawn".
|
||||||
|
`show`, `routedText`, `mainNote`, `statusLine` call `decideMain` with
|
||||||
|
`cur = st.lastPlan?.model ?? st.sessionModel` and the same tokens read.
|
||||||
|
S9. Tests, kit recipe (replaces R14 details): `boot(model = 'claude-fable-5-1')`
|
||||||
|
registers, before the first `$` call, bottom hooks `on('session.model',
|
||||||
|
() => ({ value: model }))` (answer shape per the Op results in the
|
||||||
|
declarations), `on('classic.StopFailure', ($, e) => <passthrough result>)`,
|
||||||
|
`on('classic.PostModelSwitch', …)`, and installs `mock.clock(on)`; every
|
||||||
|
breaker test advances the mock clock. `derived` recipe: prompt
|
||||||
|
"planifie …" (source 'prompt'), then `$.tool.call({ tool: 'Agent', … })`
|
||||||
|
whose bottom hook returns `{ result: { status: 'async_launched',
|
||||||
|
agentId: 'a1', … } }` (read the Agent RESULT type for the required
|
||||||
|
fields) → `/route show` main line says `derived orchestrate`; then
|
||||||
|
`$.turn.complete({ agentId: 'a1', … })` → main line says `prompt plan`.
|
||||||
|
Second derived test: same, but a route tool call `reflect` after the
|
||||||
|
dispatch → the pop does not overwrite `model reflect`. `engine fallback`
|
||||||
|
test: `$.classic.PostModelSwitch({ from_model: 'claude-fable-5-1',
|
||||||
|
to_model: 'claude-opus-5-5', source: 'auto', … })` → fable listed under
|
||||||
|
`down:`; a plan route step does not go back to fable; a second auto
|
||||||
|
switch inside the hold adds no strike (show prints the same until).
|
||||||
|
Strikes test: expire (advance clock) → mark again → until doubles.
|
||||||
|
S10. AC3 fix: the `fallback:` and `tiers:` greps run inside the
|
||||||
|
DEFAULT_CONFIG awk range.
|
||||||
|
S11. R16 gains: the order of `classic.StopFailure` vs `turn.complete`; whether
|
||||||
|
the engine's fallback on a hook-rewritten request raises `PostModelSwitch`.
|
||||||
|
|
||||||
|
## r4 — second confirmation (FATAL(4): 1 BLOCKER, 3 MAJOR): BINDING over r3 where they conflict; the last revision, executor dispatched on it
|
||||||
|
T1. Two fields, no contradiction: `st.turnModel: string | undefined` is the
|
||||||
|
STICKY cur, set ONLY when `decideMain` moved the model (upgrade,
|
||||||
|
downgrade or fallback), reset in `endMainTurn` and at `session.end`;
|
||||||
|
`st.lastPlan` (the plan actually sent last, breaker target) is KEPT across
|
||||||
|
turns and never used as cur. `cur = st.turnModel ?? e.model`. "Keep cur"
|
||||||
|
returns `st.turnModel` when set, else `e.model` VERBATIM: an unrouted step
|
||||||
|
never re-sends a model the router did not choose this turn, so an
|
||||||
|
engine fallback that lands in `e.model` is respected by construction.
|
||||||
|
S3's "lastPlan is reset at endMainTurn" is void (S6 stands).
|
||||||
|
T2. Auto switch marking (S1 refined): on `PostModelSwitch` `source: 'auto'`,
|
||||||
|
let `sent = canonical(st.lastPlan?.model)` and `to = canonical(to_model)`.
|
||||||
|
If `to === sent` → nothing (the engine landed where the router already
|
||||||
|
was, or the router's own rewrite surfaced as a switch). Else the mark
|
||||||
|
target is `sent` when it is a table id (the model actually sent), else
|
||||||
|
`canonical(from_model)` when THAT is a table id, else nothing. Always
|
||||||
|
log `from_model`, `to_model`, `requested_model`. R16 gains: which
|
||||||
|
`from_model` the event carries after a router upgrade, and whether a
|
||||||
|
router rewrite itself raises an `auto` switch.
|
||||||
|
T3. `st.sessionModel` is PRESERVED through the `session.end` rebuild (listed
|
||||||
|
with `down`, strikes, `agentModels`). `canonical()` never prefix-matches
|
||||||
|
an empty string or a string that does not start with `claude-`: both
|
||||||
|
map to UNKNOWN (returned unchanged, no table id). The `[1m]` carry reads
|
||||||
|
the raw string of the step (`e.model`, or `st.turnModel`), never
|
||||||
|
`sessionModel`. `sessionModel` is used by texts only; when it is '' or
|
||||||
|
unknown, texts print the engine word `session model` instead of an id.
|
||||||
|
T4. Tokens: `ctx.tokens: number | undefined` (undefined on a failed or absent
|
||||||
|
read). The upgrade cap treats undefined as 0 (upgrade allowed: the targets
|
||||||
|
are fable/opus, no window entry); `windowOk` treats undefined as NOT
|
||||||
|
fitting (fail closed, as today).
|
||||||
|
T5. Marks: strikes are per EPISODE (a mark on an id already down adds no
|
||||||
|
strike and no log), but a `model_not_found` arriving during a timed hold
|
||||||
|
LENGTHENS it to "until reload" (logged once). Auto marks use the same
|
||||||
|
episode backoff (15 → 30 → 60 → 120 → 300 min).
|
||||||
|
T6. S5 race: on an `async_launched` result with `st.pushed === null`, push
|
||||||
|
again first (if `turnMain?.source` still allows it), then add the id.
|
||||||
|
T7. Tests assert hold DURATIONS (minutes until, computed from the mock clock)
|
||||||
|
or the presence of the id under `down:`, never a literal `HH:MM`.
|
||||||
|
`show` prints `down: <id> for <n> min (<reason>)` (and `until reload`),
|
||||||
|
computed from the pruned map and the clock value passed in.
|
||||||
|
T8. R16 final list (live, orchestrator): StopFailure vs turn.complete order;
|
||||||
|
PostModelSwitch on a rewritten-request fallback and its `from_model`;
|
||||||
|
whether a router rewrite raises `auto`; `[1m]` carry validity on opus;
|
||||||
|
`$.session.model()` string form.
|
||||||
@@ -0,0 +1,345 @@
|
|||||||
|
# PLAN r4 — model-router wave 2 (migration), 2026-10-09
|
||||||
|
|
||||||
|
r1 → r2 after 3 lenses (simplicity CONCERNS(5), correctness CONCERNS(10),
|
||||||
|
robustness FATAL(8)); r2 → r3 after the confirmation pass (robustness
|
||||||
|
FATAL(8): 1 BLOCKER); r3 → r4 after a second confirmation (correctness
|
||||||
|
CONCERNS(3), wording-level). Every BLOCKER/MAJOR is closed by a NAMED change
|
||||||
|
(§ Challenge ledger). Mod = the router while on; the tracked frontmatter
|
||||||
|
(`model:` + `effort:` on agents, `effort:` on skills) STAYS as the
|
||||||
|
off-state floor, census-locked equal to the rows (r3).
|
||||||
|
Two sub-runs on `feature/model-router-w2`: W2-A (mod) then, after a user
|
||||||
|
`/reload-plugins` + live probe, W2-B (repo migration). Docs + registries =
|
||||||
|
W2-B's own STEP 6/7 (feat pipeline tail); W2-A runs no doc-sync (deviation,
|
||||||
|
stated: a mod-only diff has no public doc of its own until B lands).
|
||||||
|
|
||||||
|
## Decisions (pass B, user 2026-10-09) — amended by the challenge
|
||||||
|
- D1 the five `effort-*` skills are DELETED in W2-B. The mod's
|
||||||
|
`Skill(effort-*)` bridge is deleted in W2-B too (same commit as the citers,
|
||||||
|
robustness 7: edits are live on the symlinked tree). The typed `/effort-`
|
||||||
|
floor code goes in W2-A (A1). The `ultrathink` floor stays.
|
||||||
|
- D2 pins reworked, not deleted (r3): rows are PHASES by ROLE, no bare
|
||||||
|
levels; the mod routes agents on both axes while on (model written at
|
||||||
|
spawn from the tier, only UPWARD in rank; effort per step; explicit
|
||||||
|
Agent params win). The tracked frontmatter stays as the OFF-STATE FLOOR
|
||||||
|
(mod off / `enabled:false` / unloaded / a hook failing open → the engine
|
||||||
|
applies the frontmatter as today: never the parent model, never session
|
||||||
|
effort on an xhigh gate agent; r2 robustness 1 + confirmation 8). The
|
||||||
|
census locks every row equal to its frontmatter (tier head == `model:`
|
||||||
|
alias, phase effort == `effort:`), so there is one declared value, two
|
||||||
|
carriers. Deleted: the five shifters, `lib/effort-pins.*` (vendored
|
||||||
|
levels become rows; off-state = session level for them, as before
|
||||||
|
BDR-108), the witness script.
|
||||||
|
- D3 the orchestrators declare PHASES: dispatch span → `route(phase=
|
||||||
|
"orchestrate")`; own level high → `reflect`, xhigh → `plan`; bookkeeping
|
||||||
|
tail → `route(phase="apply")` (work/low); escalation → `escalate`.
|
||||||
|
Judgment dispatches of BUILT-INS (`general-purpose` with `model="opus"`
|
||||||
|
or `"fable"`) carry an explicit `effort=` param (a main route never
|
||||||
|
reaches a child; correctness 8). A best-tier skill row SURVIVES the end
|
||||||
|
of the turn in its own slot (A6 `runMain`): a run spans prose gates;
|
||||||
|
turn-scoped routes (`route` calls, prompt rules, the bridge) never
|
||||||
|
touch it.
|
||||||
|
- D4 `lib/model-gate.md` = the mod rule; the witness is the `route` tool's
|
||||||
|
own answer (correctness 7): self-check big → silent; small → call
|
||||||
|
`route(phase=<entry phase>)` and STOP unless the answer names a fable or
|
||||||
|
opus id; tool absent or "is off" → STOP with the remedy. `model-check.sh`
|
||||||
|
+ its test deleted. The route answer ALWAYS names the id the next main
|
||||||
|
step runs on (A3d). Relaunch levers in STOP texts: `ultrathink` in the
|
||||||
|
relaunch prompt (turn floor) or `/route effort=max` (sticky, `/route
|
||||||
|
clear` after). Builtin `/effort` is NOT a lever inside a run (rows and
|
||||||
|
routes rank above the engine effort; r2's engine-effort detector dropped
|
||||||
|
as unsafe, confirmation 4) — documented, named to the user.
|
||||||
|
|
||||||
|
## Phase table (A2a) — two rows added to `phases`
|
||||||
|
| phase | tier | effort | role |
|
||||||
|
| plan | best | xhigh | brainstorm, plan, architecture, audit verdict |
|
||||||
|
| reflect | best | high | diagnosis, review, contract, day-to-day orchestrator entry |
|
||||||
|
| orchestrate | best | medium | between dispatches |
|
||||||
|
| escalate | best | max | stuck, cap reached |
|
||||||
|
| judge | big | xhigh | dispatched challengers, analyzers, audits |
|
||||||
|
| implement | work | medium | code from a closed plan |
|
||||||
|
| write | work | **high** (was medium) | docs, commits, refactors, handover prose on sonnet at high (BDR-107 "high judgment on sonnet") |
|
||||||
|
| verify | work | xhigh | verifier, security-auditor |
|
||||||
|
| explore | work | medium | Explore |
|
||||||
|
| **apply** (new) | work | low | low appliers (BDR-107): small fixes, release mechanics, probes, validators; bookkeeping tail of the main loop |
|
||||||
|
| mechanical | cheap | low | listing, status, profile toggles |
|
||||||
|
|
||||||
|
## Row tables (A2b)
|
||||||
|
skills (56):
|
||||||
|
- plan: ship-feature init-project onboard tour audit-delta analyze
|
||||||
|
code-clean client-handover brainstorming writing-plans
|
||||||
|
requesting-code-review 21st-ui-review
|
||||||
|
- reflect: feat hotfix bugfix refactor web-validate harden seo geo
|
||||||
|
site-motion frontend-design emil-design-eng design-motion-principles
|
||||||
|
21st-ui-build scroll-world-storytelling build-threejs-scroll-worlds
|
||||||
|
scroll-scrubbed-visual-sequence scroll-scrubbed-word-reveal
|
||||||
|
scroll-progress-timeline subagent-driven-development writing-skills
|
||||||
|
deprecation-and-migration 21st-ai 21st-ui-explore
|
||||||
|
- implement: gitflow prune-memory pdf-translate ci-cd-and-automation
|
||||||
|
observability-and-instrumentation test-driven-development
|
||||||
|
- apply: commit-change release-candidate doc capitalize close reconcile
|
||||||
|
deploy (work tier: never haiku on main even with the switch on,
|
||||||
|
robustness 15)
|
||||||
|
- mechanical: status profile plugin-check skills-perso using-git-worktrees
|
||||||
|
21st-cli-use 21st-registry 21st-design-sync
|
||||||
|
agents (21 + 2 built-ins), model = frontmatter alias, unchanged everywhere;
|
||||||
|
`effort:` frontmatter updated where the row differs (analyzer → xhigh):
|
||||||
|
- implement (sonnet/medium): feater bugfixer code-cleaner scaffolder onboarder
|
||||||
|
- write (sonnet/high): commit-changer doc-syncer handover-doc-writer refactorer
|
||||||
|
- apply (sonnet/low): hotfixer release-executor plugin-probe validator-analyzer
|
||||||
|
- verify (sonnet/xhigh): verifier security-auditor
|
||||||
|
- judge (opus/xhigh): plan-challenger plugin-advisor seo-analyzer
|
||||||
|
geo-analyzer analyzer (the ONE level delta: high → xhigh, BDR-108 rung
|
||||||
|
"audit before validation"; named at the gate)
|
||||||
|
- mechanical (haiku/low): status-reporter
|
||||||
|
- built-ins: Explore explore, Plan judge
|
||||||
|
- no row: interviewer, client-handover-writer (inline-load), impeccable-*
|
||||||
|
(gitignored vendor output, untouched)
|
||||||
|
Deltas named at the gate: analyzer effort; best-tier raise now also reaches
|
||||||
|
refactor + the design/vendored reflect skills on a small session
|
||||||
|
(simplicity 6: the raise is the point of the mod).
|
||||||
|
|
||||||
|
## W2-A — mod (contract 2026-10-09-model-router-w2a-1546)
|
||||||
|
- [ ] A1 `register.ts`: delete the typed-floor code — `slashEffort`,
|
||||||
|
`guardedSlash`, `Source` member `'slash'`, `floorSource`'s `typed /`
|
||||||
|
branch, the State comments on the typed `/effort-<l>` floor. KEEP
|
||||||
|
`EFFORT_SKILL` + `effortBridge` (tool.call bridge, removed in W2-B)
|
||||||
|
and `turnFloor`/`higherFloor`/`floorWord` (ultrathink).
|
||||||
|
- [ ] A2 `DEFAULT_CONFIG`: phases per § Phase table (`write` high, `apply`
|
||||||
|
new); `skills` + `agents` per § Row tables; the "Built-ins only"
|
||||||
|
comment → "phase = role; one row per routed repo skill/agent (wave 2);
|
||||||
|
a project-level agent of the same name shadows its row (A3)".
|
||||||
|
- [ ] A3 spawn: (a) explicit model = `e.model !== undefined` at spawn (the
|
||||||
|
engine does not pre-fill the frontmatter; no tool_use_id map);
|
||||||
|
explicit effort as today (`explicitEffort`). (b) `spawnRoute`:
|
||||||
|
`frozen` → none; row lookup for every provider; a row is SKIPPED when
|
||||||
|
the agent's definition `source`, recorded per `subagentType` from the
|
||||||
|
`agent.offer` event, is `projectSettings` or `localSettings` (a
|
||||||
|
foreign repo's own `verifier.md`); `userSettings`, `built-in`,
|
||||||
|
`plugin` or no record → the row applies (fail-open on routing, as
|
||||||
|
W1). Known limit, stated in a comment: the record is keyed by name
|
||||||
|
only (an offer fired inside a sub-agent with another cwd overwrites it). (c) `spawnTarget` for a rowed spawn
|
||||||
|
without explicit model: resolve WITHIN the tier only, and write only
|
||||||
|
an alias ranked ≥ the tier head in `fallback` (agents move UP, never
|
||||||
|
below their frontmatter alias): `big` with opus down → fable; opus +
|
||||||
|
fable down → no write + one log line (deduped per agent+tier per
|
||||||
|
turn) "model-router: <agent> tier <t> down, frontmatter model kept".
|
||||||
|
(d) `routedText`/`mainAnswer` always name the id the next main step
|
||||||
|
runs on (`model <id> (<why>)`), the sticky/floor note APPENDED, never
|
||||||
|
substituted (the gate reads this answer).
|
||||||
|
- [ ] A4 typed slash: `typedSlash: string | null` = the first token of a
|
||||||
|
slash prompt at `prompt.submit` when `origin.kind` ∈ {composer, sdk,
|
||||||
|
bridge} (allowlist; floor/default rules stay composer-only), stored
|
||||||
|
only when it is a `cfg.skills` key; mid-turn → `pendingSlash`,
|
||||||
|
promoted at `endMainTurn`. `prompt.submit` also records
|
||||||
|
`st.promptAllowed` (origin in the allowlist) for the turn it opens
|
||||||
|
(pending slot for a mid-turn prompt, like the marker). `skill.prompt`
|
||||||
|
on main (`skillCalls === 0`): apply the row when `e.skill ===
|
||||||
|
typedSlash` (then null it) OR when `promptAllowed && loops.size === 0
|
||||||
|
&& spawning === 0` (`spawning` = a counter held from spawn-hook entry
|
||||||
|
to after `next`); otherwise `next(e)`. Verbose log names which path
|
||||||
|
fired (`typed-marker` / `typed-fallback`).
|
||||||
|
- [ ] A5 `onSkillLoad`: a skill WITHOUT a row leaves `turnMain` untouched
|
||||||
|
(find-docs / gstack / plugin skills mid-run no longer clear the
|
||||||
|
run's route; correctness 11); a rowed skill replaces it. Same rule
|
||||||
|
INSIDE a sub-agent (gated 2026-10-09, feater NEED-DECISION): an
|
||||||
|
unrowed skill leaves the agent loop's effort untouched; a rowed one
|
||||||
|
writes it (test: `feater` loop at medium, `Skill(find-docs)` with
|
||||||
|
that agentId → next step still medium).
|
||||||
|
- [ ] A6 run slot: new `runMain: Routed | null`. A rowed skill load ON
|
||||||
|
MAIN whose phase tier is `best` writes BOTH `turnMain` (source
|
||||||
|
`skill`, as today) and `runMain`; a non-best row loaded by the
|
||||||
|
model's `Skill` tool writes `turnMain` only (helper skills such as
|
||||||
|
`using-git-worktrees` inside SDD never drop the run); a USER-TYPED
|
||||||
|
rowed skill (marker path) replaces `runMain` with its row when best,
|
||||||
|
drops it when not. Precedence `userMain > turnMain > runMain > floor
|
||||||
|
> engine` (per axis, as today); `mainRoute` (statusline, spinner,
|
||||||
|
planStep) includes it; `endMainTurn` leaves it; `pushOrchestrate`
|
||||||
|
writes `turnMain` when empty exactly as today (derived orchestrate
|
||||||
|
overrides the run default for the dispatch span, like a `route
|
||||||
|
orchestrate` call); cleared by `/route clear` (text: "run slot
|
||||||
|
dropped" when one held), `/route off`, a user `/model` switch
|
||||||
|
(`PostModelSwitch` source command|picker|sdk); `route(clear=true)`
|
||||||
|
from the model clears `turnMain` only and its answer says "run
|
||||||
|
<phase> still holds". A `Skill` call inside a sub-agent never touches
|
||||||
|
it. `/route show` prints `run <phase>` when no turn route is in force.
|
||||||
|
- [ ] A7 (dropped in r3: engine-effort detector, confirmation 4).
|
||||||
|
- [ ] A8 `mergeTable` accepts `null` in the override's `skills`/`agents`
|
||||||
|
to drop a default row (robustness 4).
|
||||||
|
- [ ] A9 `register.test.ts`: delete the typed `/effort-*` floor tests;
|
||||||
|
keep the bridge test; add — typed `/feat` (prompt.submit `/feat x`,
|
||||||
|
origin composer, then skill.prompt feat) → `main: skill reflect`,
|
||||||
|
`effort high`, `[tier best]`; origin `channel`: not armed AND the
|
||||||
|
idle fallback refused (main untouched); skill.prompt `feat` with a
|
||||||
|
live sub-agent and no marker → main untouched; same with no live loop
|
||||||
|
and an allowed origin → routed; a mid-turn `/status` → pending,
|
||||||
|
applied after turn.complete; `Skill(find-docs)` (no row) after
|
||||||
|
`route plan` keeps plan — the existing test 'floor: survives a skill
|
||||||
|
load' (register.test.ts:341-348) is REWRITTEN to assert orchestrate
|
||||||
|
KEPT (A5 behavior, authorized here); run slot: `Skill(feat)`, then
|
||||||
|
`route orchestrate`, then `turn.complete` → `/route show` back to
|
||||||
|
`run reflect`; `Skill(using-git-worktrees)` (mechanical, model path)
|
||||||
|
keeps `run reflect` under a mechanical turn route; a TYPED `/status`
|
||||||
|
drops the run slot; `/route clear` drops it; `/route off` drops it;
|
||||||
|
PostModelSwitch source `command` drops it; a `Skill(feat)` with an
|
||||||
|
agentId leaves `runMain`; `Skill(effort-low)` bridge then
|
||||||
|
`turn.complete` → session defaults (not sticky); `feater` spawn (provider
|
||||||
|
`{plugin:'engine',tier:'core'}`, no model) → `started.model` = sonnet
|
||||||
|
full id and the first `turn.step` of that agentId at medium;
|
||||||
|
`plan-challenger` with opus down → fable id; opus AND fable down →
|
||||||
|
no write, one log line; explicit `model`/`effort` win; `fork: true`
|
||||||
|
untouched; a project-source `agent.offer` record → no write; route
|
||||||
|
answer text names the id with a floor in force; override `agents: {
|
||||||
|
verifier: null }` drops the row (bottom `fs`/`env` mocked like
|
||||||
|
`session.model`; if the kit refuses, the case is dropped and said so).
|
||||||
|
- Disposition: honors BDR-115 (one writer per axis while on, full ids,
|
||||||
|
config tables, closure state; rule 4 closed, rule 5 amended by A5/A6),
|
||||||
|
BDR-107/108 (roles + levels preserved except analyzer), BDR-076/077
|
||||||
|
(judge rows = opus, off-state floor kept), LRN-205, LRN-206, LRN-207
|
||||||
|
(breaker feeds the tier resolution).
|
||||||
|
|
||||||
|
## Gate between A and B — live probe (user runs `/reload-plugins`)
|
||||||
|
Evidence into the W2-B contract: (1) typed `/status` → `/route show` main
|
||||||
|
`skill mechanical`; (2) a real `plugin-probe` or `feater` dispatch with
|
||||||
|
verbose on → spawn log line (provider shape, model id written) + `step 0
|
||||||
|
agent …` effort line = spawn/first-step ordering fact the TODO asks for;
|
||||||
|
(3) on a SONNET session (`/model sonnet`, then back): typed `/feat` → the
|
||||||
|
self-check wording of the system prompt + the `route` answer id (gate
|
||||||
|
witness fact); (4) probe (1) again with a background Explore alive (marker
|
||||||
|
path vs fallback path in the verbose log). The verbose log line (`typed-marker` /
|
||||||
|
`typed-fallback`, `spawn … → <id>`, `step 0 agent …`) is the witness for
|
||||||
|
(1), (2), (4), not `/route show` after the turn (turn-scoped routes are
|
||||||
|
gone by then). Decision rules: (2) step 0 BEFORE the spawn bookkeeping →
|
||||||
|
step 0 runs on the frontmatter `effort:` (kept, D2), recorded as a known
|
||||||
|
limit; (4) `typed-marker` never seen → W2-B blocked, A4 re-planned (the
|
||||||
|
`prompt.submit` text fact does not hold). Only then W2-B.
|
||||||
|
|
||||||
|
## W2-B — repo migration (own contract)
|
||||||
|
- [ ] B0 mod: delete `EFFORT_SKILL`, `effortBridge`, the Skill-hook branch,
|
||||||
|
their test; `agents/`+`skills/` rows unchanged.
|
||||||
|
- [ ] B1 `lib/effort-shift.md` rewritten (~40 lines): route doctrine, the
|
||||||
|
wiring points in `route` terms (D3), judgment built-ins carry
|
||||||
|
`effort=`, sticky skill route + `/route clear`, no pairing rule,
|
||||||
|
headless OK (hooks run under -p), levers = `ultrathink` / `/route
|
||||||
|
effort=max` (builtin `/effort` is not a lever inside a run), last ROWED
|
||||||
|
skill loaded wins
|
||||||
|
(an unrowed one changes nothing).
|
||||||
|
- [ ] B2 citers `Skill(effort-*)` → `route`: ship-feature 9, init-project
|
||||||
|
5, feat 4, bugfix 4, web-validate 3, seo 3, hotfix 3, geo 3,
|
||||||
|
verify-secure-loop 3, harden 2, code-clean 2, audit-delta 2, onboard
|
||||||
|
1, client-handover-writer 1; EFFORT SHIFTS header lines in the 13
|
||||||
|
orchestrators + tour:33; `effort=` added to every `general-purpose`
|
||||||
|
judgment dispatch (`model="opus"`: ship-feature, init-project,
|
||||||
|
onboard ×7, tour; `model: "fable"` skill-runners in
|
||||||
|
client-handover-writer); STOP/remedy texts (challenge-plan.md:62,
|
||||||
|
verify-secure-loop.md:109) → the D4 levers; the 21 prose sites
|
||||||
|
stating "sonnet by frontmatter pin" / "`model: opus`-pinned,
|
||||||
|
session-independent" (challenge-plan:46, feat:146, bugfix:160,
|
||||||
|
hotfix:138, code-clean:175, seo:463, 6 agents…) reworded "routed by
|
||||||
|
the model-router row (frontmatter = off-state floor)".
|
||||||
|
- [ ] B3 `lib/model-gate.md` → D4 (≤ 25 lines, keeps `model: "fable"` for
|
||||||
|
skill-runners and the dispatch-tier table); delete
|
||||||
|
`lib/model-check.sh`, `lib/tests/model-check.test.sh`.
|
||||||
|
- [ ] B4 delete `lib/effort-pins.txt`, `lib/effort-pins.sh`,
|
||||||
|
`lib/tests/effort-pins.test.sh`; remove the re-apply blocks + comments
|
||||||
|
in `install-plugins.sh` (941-943, 1008 comment, 1131-1136) and
|
||||||
|
`update-all.sh` (472, 567-572); `lib/tests/higgsfield.test.sh`
|
||||||
|
:324 + the `before-pins` check (:330-356) dropped.
|
||||||
|
- [ ] B5 `skills/effort-*` deleted (5 dirs; `profile`/catalog lists
|
||||||
|
grepped). Tracked frontmatter KEPT and ALIGNED: `agents/analyzer.md`
|
||||||
|
`effort: xhigh`; the vendored tracked `design-motion-principles`
|
||||||
|
keeps `effort: high`; impeccable-* (gitignored) untouched; every
|
||||||
|
other `model:`/`effort:` value already equals its row.
|
||||||
|
- [ ] B6 census: `lib/tests/effort-routing.test.sh` rewritten — every
|
||||||
|
tracked `effort:` equals its row's phase effort and every agent
|
||||||
|
`model:` alias equals its row's tier head (drift lock, both
|
||||||
|
directions, parsed from `register.ts`; haiku rows exempt from the
|
||||||
|
effort direction: no effort on haiku); enumeration = tracked
|
||||||
|
`SKILL.md` under `skills/` + `skills-external/` (`git ls-files`),
|
||||||
|
no-row list = graphify, model-router, find-docs, impeccable, the
|
||||||
|
agents interviewer/client-handover-writer/impeccable-*; no
|
||||||
|
`Skill(effort-` anywhere in
|
||||||
|
skills/agents/lib; D3 wiring markers (orchestrate in the 11, apply
|
||||||
|
tail in the 5, escalate ×3 in verify-secure-loop + ship-feature,
|
||||||
|
`effort=` on every `model="opus"` general-purpose dispatch); every
|
||||||
|
tracked skill/agent (minus the no-row list) has a row in
|
||||||
|
`register.ts` with the EXPECTED PHASE (per-name asserts, not mere
|
||||||
|
presence); `model-routing.test.sh` :96-97 model-gate locks
|
||||||
|
updated (the 18 `model:` locks, loops-light.test.sh:74,79 and
|
||||||
|
plan-challenger.test.sh:20 stay valid: frontmatter kept); `CLAUDE.global.md` Design
|
||||||
|
work lines 299-301 → "rows in the mod, an unrowed member changes
|
||||||
|
nothing".
|
||||||
|
- [ ] B7 doctrine-citers census; full `make test` once before merge.
|
||||||
|
- [ ] STEP 6/7 of the B run: doc-syncer audit → README/USAGE/ARCHITECTURE/
|
||||||
|
CHANGELOG; BDR-115 amendment (wave 2 closed, D1-D4, A5/A6), LRN
|
||||||
|
(typed-slash marker + sticky skill route), EVAL on the run; TODO W2
|
||||||
|
lines checked.
|
||||||
|
|
||||||
|
## Challenge ledger (r1 → r2)
|
||||||
|
- robustness 1 BLOCKER (mod off → agents inherit the parent model) →
|
||||||
|
D2: `model:` frontmatter kept as off-state floor; A3c row overrides it
|
||||||
|
while on.
|
||||||
|
- correctness 1/2, robustness 2 (marker unbound, ordering unproven) → A4
|
||||||
|
name-bound marker + pending slot + loops-size fallback; live probe gate.
|
||||||
|
- correctness 3 (show text) → A9 asserts `skill reflect`.
|
||||||
|
- correctness 4 (56) → fixed everywhere.
|
||||||
|
- correctness 5, simplicity 1/2, robustness 13 (level deltas) → `write`
|
||||||
|
high + `apply` phase; one delta left (analyzer), named.
|
||||||
|
- correctness 6 (raise lasts one turn) → A6 sticky skill route.
|
||||||
|
- correctness 7, robustness 5, simplicity 8 (gate witness) → D4 route answer.
|
||||||
|
- correctness 8, robustness 11, simplicity 7 (point 5) → explicit `effort=`.
|
||||||
|
- correctness 9, robustness 6 (levers) → D4 texts (r2's A7 engine-effort
|
||||||
|
detector dropped again in r3).
|
||||||
|
- correctness 10, robustness 9/10 (census) → B6 per-phase asserts, suites
|
||||||
|
listed; no `effort="high"` call-site edits (write = high).
|
||||||
|
- correctness 11 (unrowed skill clears) → A5.
|
||||||
|
- correctness 12, robustness 12, simplicity 11 (counts, impeccable) → B5
|
||||||
|
from `git ls-files`; impeccable untouched.
|
||||||
|
- correctness 13 (dead code) → A1 list; contract criterion 5 grep widened.
|
||||||
|
- correctness 14 (override test needs fs) → A9 bottom mocks or drop, stated.
|
||||||
|
- correctness 15, robustness 4/8 (provider shape, collisions) → A3b
|
||||||
|
source from `agent.offer` + A8 null rows; tests use `engine/core`.
|
||||||
|
- robustness 3 (haiku via fallback chain) → A3c tier-only + log.
|
||||||
|
- robustness 7 (live edits mid-migration) → bridge stays until B0; probe
|
||||||
|
gate before B.
|
||||||
|
- robustness 14 (tour:33, install comment) → B2/B4.
|
||||||
|
- robustness 15 (deploy on haiku) → `apply` rows.
|
||||||
|
- simplicity 3 (tail) → kept as `route(phase="apply")` (phases only; the
|
||||||
|
switch-on cost is the user's opt-in). [kept, reasoned]
|
||||||
|
- simplicity 4 (W2-C) → folded into W2-B STEP 6/7.
|
||||||
|
- simplicity 12 (two parsers) → contract oracle = one-shot values; B6 =
|
||||||
|
durable per-phase census. [kept, reasoned]
|
||||||
|
|
||||||
|
## Confirmation ledger (r2 → r3)
|
||||||
|
- conf 1 BLOCKER (route calls wipe the sticky slot) → A6 separate `runMain`
|
||||||
|
slot, turn writers never touch it.
|
||||||
|
- conf 2 (low/haiku leaks across turns) → only best-tier rows enter `runMain`.
|
||||||
|
- conf 3 (bridge sticky in the A→B window) → bridge writes `turnMain` only.
|
||||||
|
- conf 4 (engine-effort detector) → A7 dropped; builtin `/effort` named as
|
||||||
|
a non-lever to the user; levers = `ultrathink`, `/route effort=max`.
|
||||||
|
- conf 5 (route answer without id) → A3d always names the id, note appended;
|
||||||
|
probe (3) on a sonnet session.
|
||||||
|
- conf 6 (fs/cwd shadow unreliable) → A3b definition source from
|
||||||
|
`agent.offer`; fs dropped.
|
||||||
|
- conf 7 (judge → sonnet in-tier) → A3c rank ≥ tier head, else no write + log.
|
||||||
|
- conf 8 (off-state quality drop, step-0 ordering) → D2 frontmatter kept
|
||||||
|
as floor on both axes, census-locked; probe decision rule written.
|
||||||
|
- conf 9 (pre-fill premise) → explicit = `e.model !== undefined`, no map.
|
||||||
|
- conf 10 (origin denylist) → allowlist composer|sdk|bridge.
|
||||||
|
- conf 11 (preload before trackLoop) → spawning counter; log names the path.
|
||||||
|
- conf 12 (kit has no fs) → fs dropped. conf 13 → deduped log.
|
||||||
|
- conf 14 (skill rows shadow) → accepted, named: skill rows affect main
|
||||||
|
only; a foreign skill of the same name gets the row for its turn.
|
||||||
|
- conf 15 (artefact drift) → contract clarification amended; status-reporter
|
||||||
|
listed under mechanical.
|
||||||
|
|
||||||
|
## Confirmation 2 ledger (r3 → r4)
|
||||||
|
- conf2 1 (helper skill drops the run) → A6: model-loaded non-best rows
|
||||||
|
write turnMain only; typed rowed skills manage runMain.
|
||||||
|
- conf2 2 (test :347 goes red) → A9 authorizes its rewrite.
|
||||||
|
- conf2 3 (fallback ignores origin) → A4 `promptAllowed` flag on the fallback.
|
||||||
|
- conf2 4/5/6 (runMain wiring) → A6 names both slots, mainRoute, clearLoop
|
||||||
|
text, pushOrchestrate, `/model` + `/route off` + sub-agent cases in A9.
|
||||||
|
- conf2 7 (probe witnesses) → gate: verbose lines + decision rules.
|
||||||
|
- conf2 8 (grep) → contract criterion 5 narrowed.
|
||||||
|
- conf2 9 (source tokens) → A3b names projectSettings|localSettings + limit.
|
||||||
|
- conf2 10/11 (census) → B6 enumeration, haiku exemption, cites fixed.
|
||||||
@@ -262,3 +262,7 @@ skills-external/.higgsfield-stage.*/
|
|||||||
# ── gitflow standard socle (added by gitflow_init; additive, safe to edit) ──
|
# ── gitflow standard socle (added by gitflow_init; additive, safe to edit) ──
|
||||||
*.log
|
*.log
|
||||||
!.claude/deploy/
|
!.claude/deploy/
|
||||||
|
|
||||||
|
# mods/: the engine lays tsconfig.json beside a loaded mod; its
|
||||||
|
# .claude-plugin/types/ ignores itself
|
||||||
|
mods/*/tsconfig.json
|
||||||
|
|||||||
@@ -26,6 +26,7 @@ claude-config/
|
|||||||
├── rules/ # Rule files deployed to ~/.claude/rules (path-scoped or always-on)
|
├── rules/ # Rule files deployed to ~/.claude/rules (path-scoped or always-on)
|
||||||
├── agents/ # Execution units called by skills (never invoked directly)
|
├── agents/ # Execution units called by skills (never invoked directly)
|
||||||
├── skills/ # Entry points invoked via /skill-name
|
├── skills/ # Entry points invoked via /skill-name
|
||||||
|
├── mods/ # Claude Code mods (function-hooks plugins), loaded through the skills/<name> symlink
|
||||||
├── skills-external/ # Vendored skill packs: gstack submodule, design skills, superpowers, agent-skills, MengTo scroll skills, 21st and Higgsfield packs (machine-owned copies gitignored)
|
├── skills-external/ # Vendored skill packs: gstack submodule, design skills, superpowers, agent-skills, MengTo scroll skills, 21st and Higgsfield packs (machine-owned copies gitignored)
|
||||||
├── templates/ # Per-project templates (CLAUDE.md, settings, memory registries, deploy runbook, gitignore)
|
├── templates/ # Per-project templates (CLAUDE.md, settings, memory registries, deploy runbook, gitignore)
|
||||||
└── lib/ # Shared libs: gitflow, profiles, vendoring, effort pins, gates, archetypes, tests
|
└── lib/ # Shared libs: gitflow, profiles, vendoring, effort pins, gates, archetypes, tests
|
||||||
@@ -35,5 +36,6 @@ claude-config/
|
|||||||
|
|
||||||
- `skills/` = entry points you invoke via `/skill-name`
|
- `skills/` = entry points you invoke via `/skill-name`
|
||||||
- `agents/` = execution units called by skills (never invoked directly by user)
|
- `agents/` = execution units called by skills (never invoked directly by user)
|
||||||
|
- `mods/` = Claude Code mods (function-hooks plugins); each loads through the tracked symlink `skills/<name>` as `<name>@skills-dir`, live at the next session
|
||||||
- `templates/` = symlinked to `~/.claude/templates/` — copy into projects via `/onboard` or manually
|
- `templates/` = symlinked to `~/.claude/templates/` — copy into projects via `/onboard` or manually
|
||||||
- **Graphify** builds a knowledge graph of any codebase (`/graphify query`), producing a navigable wiki in `graphify-out/wiki/`. This map helps Claude understand project structure, find relevant code faster, and reason across files. Essential for large-scope tasks (multi-file features, complex bugs, architectural changes). Small tasks should skip it and read files directly. Proposed only from 200 tracked code files: the session-start banner informs, the user decides; nothing builds a graph without that go.
|
- **Graphify** builds a knowledge graph of any codebase (`/graphify query`), producing a navigable wiki in `graphify-out/wiki/`. This map helps Claude understand project structure, find relevant code faster, and reason across files. Essential for large-scope tasks (multi-file features, complex bugs, architectural changes). Small tasks should skip it and read files directly. Proposed only from 200 tracked code files: the session-start banner informs, the user decides; nothing builds a graph without that go.
|
||||||
|
|||||||
@@ -7,6 +7,7 @@ Format follows [Keep a Changelog](https://keepachangelog.com/) and this project
|
|||||||
## [Unreleased]
|
## [Unreleased]
|
||||||
|
|
||||||
### Added
|
### Added
|
||||||
|
- **model-router mod**: `mods/model-router/`, a Claude Code mod (function-hooks plugin), routes the effort of every main-loop request from a phase table, along with the model and effort of the built-in sub-agents (Explore on sonnet/medium, Plan on opus/xhigh). It answers `Skill(effort-*)` itself, so the five `effort-*` skills no longer load while it is on. `ultrathink` in a prompt and a typed `/effort-<level>` set the main turn's default and minimum effort. The model gets a `route` tool and the user a `/route` command (`show|clear|off|on|reload|<phase>|model=<alias|id> effort=<level>|switch on|off|verbose on|off`). Optional per-machine config `~/.claude/model-router.json`, where `"enabled": false` turns it off on that machine. The spinner suffix and the status line show the route in force. It loads in every session through the tracked symlink `skills/model-router` (`model-router@skills-dir`). `make doctor` gains a Mods section; suite `make test suite=lib/tests/mods.test.sh`. Known limits: the main loop switches model only with `mainModelSwitch` on (default off, one cold-cache step per switch into another model), and the hooks send full model ids, so the `models` table has to follow new versions.
|
||||||
- **Manual-push mode**: `git config gitflow.autopush false` (human-set) now stops every push the gitflow lib makes, not only the post-commit / post-merge hooks. `gitflow start` and `finish` branch, commit and merge locally and push nothing; `gitflow delete` leaves the `origin/` copy in place and prints `git push origin --delete <br>` for the user to run. `hooks/unpushed-guard.sh` stays silent at turn end in this mode and opens each session with one `ℹ manual push mode:` line counting the commits no remote holds across every local branch; an unparseable or unreadable `gitflow.autopush` value is treated as manual push mode too, and that line names it. `hooks/push-guard.sh` (PreToolUse, `Bash|Monitor`) refuses any `git push` Claude types while `gitflow.autopush` reads false in the session cwd or in a literal `-C`/`cd` directory the command names (global config counts outside a repo); the refusal tells the user to run it with `! git push`. It reads the mode through the same lib verb as every other reader and fails closed: an unparseable or unreadable value reads as manual, and an internal error, a missing `lib/gitflow.sh`, more than 20 distinct directory tokens in one command (capped before any token is classified), a `cd`/`-C` directory token mixing quoted and unquoted parts, or a payload jq cannot parse whose raw text looks like a push refuse the push (these pathological cases fire in auto mode too). Directory tokens are read as whole shell words, adjacent quoted segments and backslash escapes included. In manual mode it over-blocks any command where a `push` word follows a `git` token; the misses listed in its header fall to a new `autoMode.soft_deny` rule that no request in the turn clears. The session banner adds `🔒 push : manual (autopush=false) — ! git push` when the key reads false, and `🔒 push : manual (autopush bad) — ! git push` when the value is invalid. Skills read the mode through a new lib verb, `bash ~/.claude/lib/gitflow.sh push-mode`: it prints `auto`, `manual` or `invalid` (rc 0) and names an invalid value on stderr (printable characters only, 64 at most). It is the one reader a skill may call, since the `git config` read of the key is denied to Claude. Skills push nothing on their own, except the `/release-candidate` tag in auto-push mode on an explicit go. Every "on origin" or "not pushed" line they print comes from `git rev-list --count origin/<br>..<br>` read after the fact, with the complete `! git …` command when something is left for the user to push. An invalid value (anything but unset, true or false, or a read that fails) is manual push mode for every reader and is named where it is read (see Fixed). Tests: `lib/gitflow-test.sh` T11b (push-mode verb), T18m and T18q blocks, `lib/tests/unpushed-guard.test.sh` T10-T16, `lib/tests/push-guard.test.sh` (98 checks).
|
- **Manual-push mode**: `git config gitflow.autopush false` (human-set) now stops every push the gitflow lib makes, not only the post-commit / post-merge hooks. `gitflow start` and `finish` branch, commit and merge locally and push nothing; `gitflow delete` leaves the `origin/` copy in place and prints `git push origin --delete <br>` for the user to run. `hooks/unpushed-guard.sh` stays silent at turn end in this mode and opens each session with one `ℹ manual push mode:` line counting the commits no remote holds across every local branch; an unparseable or unreadable `gitflow.autopush` value is treated as manual push mode too, and that line names it. `hooks/push-guard.sh` (PreToolUse, `Bash|Monitor`) refuses any `git push` Claude types while `gitflow.autopush` reads false in the session cwd or in a literal `-C`/`cd` directory the command names (global config counts outside a repo); the refusal tells the user to run it with `! git push`. It reads the mode through the same lib verb as every other reader and fails closed: an unparseable or unreadable value reads as manual, and an internal error, a missing `lib/gitflow.sh`, more than 20 distinct directory tokens in one command (capped before any token is classified), a `cd`/`-C` directory token mixing quoted and unquoted parts, or a payload jq cannot parse whose raw text looks like a push refuse the push (these pathological cases fire in auto mode too). Directory tokens are read as whole shell words, adjacent quoted segments and backslash escapes included. In manual mode it over-blocks any command where a `push` word follows a `git` token; the misses listed in its header fall to a new `autoMode.soft_deny` rule that no request in the turn clears. The session banner adds `🔒 push : manual (autopush=false) — ! git push` when the key reads false, and `🔒 push : manual (autopush bad) — ! git push` when the value is invalid. Skills read the mode through a new lib verb, `bash ~/.claude/lib/gitflow.sh push-mode`: it prints `auto`, `manual` or `invalid` (rc 0) and names an invalid value on stderr (printable characters only, 64 at most). It is the one reader a skill may call, since the `git config` read of the key is denied to Claude. Skills push nothing on their own, except the `/release-candidate` tag in auto-push mode on an explicit go. Every "on origin" or "not pushed" line they print comes from `git rev-list --count origin/<br>..<br>` read after the fact, with the complete `! git …` command when something is left for the user to push. An invalid value (anything but unset, true or false, or a read that fails) is manual push mode for every reader and is named where it is read (see Fixed). Tests: `lib/gitflow-test.sh` T11b (push-mode verb), T18m and T18q blocks, `lib/tests/unpushed-guard.test.sh` T10-T16, `lib/tests/push-guard.test.sh` (98 checks).
|
||||||
|
|
||||||
### Changed
|
### Changed
|
||||||
|
|||||||
@@ -59,6 +59,28 @@ Gotcha, learned the hard way: `git rm --cached` keeps the working file,
|
|||||||
but if the branch you merge into still tracks it, the merge deletes it
|
but if the branch you merge into still tracks it, the merge deletes it
|
||||||
from disk. Untrack and merge, then restore with the command above.
|
from disk. Untrack and merge, then restore with the command above.
|
||||||
|
|
||||||
|
## mods/ — function-hooks plugins (Claude Code mods)
|
||||||
|
|
||||||
|
A mod lives in `mods/<name>/` (`.claude-plugin/plugin.json` + hooks). It
|
||||||
|
loads through the tracked relative symlink `skills/<name>` -> `../mods/<name>`
|
||||||
|
(`~/.claude/skills` links to `skills/`) as `<name>@skills-dir`, in place,
|
||||||
|
live at the next session or `/reload-plugins`. New mod: `ln -s ../mods/<name>
|
||||||
|
skills/<name>` from the repo root (guard with `[ -L ]`, a re-run nests a link).
|
||||||
|
Not `CLAUDE_CODE_PLUGIN_DIRS` (absolute path, settings `env` has no `$HOME`
|
||||||
|
expansion, settings.json is tracked), nor a local marketplace (`add` writes
|
||||||
|
an absolute path into settings.json).
|
||||||
|
- The engine lays `mods/<name>/tsconfig.json` and `.claude-plugin/types/`;
|
||||||
|
both are gitignored.
|
||||||
|
- Optional user config: `~/.claude/<name>.json`. Its `"enabled": false` is the
|
||||||
|
per-machine off switch (untracked). `"<name>@skills-dir": false` in
|
||||||
|
`enabledPlugins` also works but lands in the TRACKED settings.json and
|
||||||
|
dirties every machine's tree.
|
||||||
|
- A dev copy of the same name (`--plugin-dir`, hot-reload link in
|
||||||
|
`~/.claude/dev-mods/<session>/`) shadows the skills-dir copy for that
|
||||||
|
session: remove it before reading `/reload-plugins` as a test of the link.
|
||||||
|
- Tests: `make test suite=lib/tests/mods.test.sh` (manifest, link,
|
||||||
|
`claude plugin validate`, `claude plugin test`). `doctor.sh` has a Mods section.
|
||||||
|
|
||||||
## Transient planning artifacts
|
## Transient planning artifacts
|
||||||
|
|
||||||
`docs/superpowers/specs/**` and `docs/superpowers/plans/**` are run-time
|
`docs/superpowers/specs/**` and `docs/superpowers/plans/**` are run-time
|
||||||
|
|||||||
@@ -34,12 +34,15 @@ test: ## Run deterministic tests hermetically (one: make test suite=lib/tests/x.
|
|||||||
@# fire inside the throwaway repos the suites build. The export lives
|
@# fire inside the throwaway repos the suites build. The export lives
|
||||||
@# HERE so nobody has to type the (denied) env-prefix form by hand.
|
@# HERE so nobody has to type the (denied) env-prefix form by hand.
|
||||||
@export GIT_CONFIG_GLOBAL=/dev/null GIT_CONFIG_SYSTEM=/dev/null; \
|
@export GIT_CONFIG_GLOBAL=/dev/null GIT_CONFIG_SYSTEM=/dev/null; \
|
||||||
fail=0; for t in $(or $(suite),$(SUITES)); do \
|
fail=0; red=""; for t in $(or $(suite),$(SUITES)); do \
|
||||||
echo "== $$t"; \
|
echo "== $$t"; \
|
||||||
case "$$(basename "$$t")" in \
|
case "$$(basename "$$t")" in \
|
||||||
run-release-candidate.sh) RC_WORK=$$(mktemp -d) RC_TAG=1 bash "$$t" || fail=1 ;; \
|
run-release-candidate.sh) RC_WORK=$$(mktemp -d) RC_TAG=1 bash "$$t" || { fail=1; red="$$red $$t"; echo "FAIL $$t"; } ;; \
|
||||||
*) bash "$$t" || fail=1 ;; \
|
*) bash "$$t" || { fail=1; red="$$red $$t"; echo "FAIL $$t"; } ;; \
|
||||||
esac; done; exit $$fail
|
esac; done; \
|
||||||
|
if [ $$fail -eq 0 ]; then echo "all suites green"; \
|
||||||
|
else echo "$$(echo $$red | wc -w | tr -d ' ') suite(s) red:$$red"; fi; \
|
||||||
|
exit $$fail
|
||||||
|
|
||||||
scan-secrets: ## Gitleaks sweep: this repo's history + ~/.claude. Extra repos: make scan-secrets repos="path1 path2"
|
scan-secrets: ## Gitleaks sweep: this repo's history + ~/.claude. Extra repos: make scan-secrets repos="path1 path2"
|
||||||
@command -v gitleaks >/dev/null 2>&1 || { echo "gitleaks not installed — https://github.com/gitleaks/gitleaks"; exit 1; }
|
@command -v gitleaks >/dev/null 2>&1 || { echo "gitleaks not installed — https://github.com/gitleaks/gitleaks"; exit 1; }
|
||||||
|
|||||||
@@ -88,7 +88,7 @@ every call site.
|
|||||||
| doc-syncer | sonnet pin; audit mode dispatched `model="opus"` | two-mode: audit (drift judgment, opus) / patch (mechanical apply, sonnet) |
|
| doc-syncer | sonnet pin; audit mode dispatched `model="opus"` | two-mode: audit (drift judgment, opus) / patch (mechanical apply, sonnet) |
|
||||||
| handover-doc-writer | sonnet pin; synthesize mode dispatched `model="opus"` | two-mode: synthesize (opus) / render (sonnet) — client deliverable |
|
| handover-doc-writer | sonnet pin; synthesize mode dispatched `model="opus"` | two-mode: synthesize (opus) / render (sonnet) — client deliverable |
|
||||||
| interviewer, client-handover-writer | unpinned (inline-load = session model) | they ARE the main loop — a frontmatter pin would be inert |
|
| interviewer, client-handover-writer | unpinned (inline-load = session model) | they ARE the main loop — a frontmatter pin would be inert |
|
||||||
| Explore (built-in) | inherit session (Fable/Opus) | search feeds reflection — kept on the big model, not pinned down |
|
| Explore, Plan (built-in) | model-router mod: Explore → sonnet/medium, Plan → opus/xhigh, set at spawn; an explicit `model=` on the call wins; mod off: inherit session | built-ins routed per phase by the mod |
|
||||||
|
|
||||||
The pure-execution skills `/doc`, `/status`, `/commit-change`,
|
The pure-execution skills `/doc`, `/status`, `/commit-change`,
|
||||||
`/release-candidate` **dispatch** their agent (instead of inline-loading it)
|
`/release-candidate` **dispatch** their agent (instead of inline-loading it)
|
||||||
@@ -108,12 +108,26 @@ entry level (`/status` low … `/ship-feature` xhigh); the vendored externals
|
|||||||
`lib/effort-pins.txt`, re-applied by `lib/effort-pins.sh` after every
|
`lib/effort-pins.txt`, re-applied by `lib/effort-pins.sh` after every
|
||||||
vendoring step. Orchestrators shift per phase through the `effort-low` …
|
vendoring step. Orchestrators shift per phase through the `effort-low` …
|
||||||
`effort-max` skills (`lib/effort-shift.md`, always sent with another tool
|
`effort-max` skills (`lib/effort-shift.md`, always sent with another tool
|
||||||
call: a lone Skill call applies nothing). Model pins stay tier aliases
|
call: a lone Skill call applies nothing (mod off)). Model pins stay tier aliases
|
||||||
(`sonnet`, `opus`, `haiku`, `fable`): the latest version of a tier is also
|
(`sonnet`, `opus`, `haiku`, `fable`): the latest version of a tier is also
|
||||||
the cheapest or same-priced, so the quality/price trade-off is tier × effort,
|
the cheapest or same-priced, so the quality/price trade-off is tier × effort,
|
||||||
never version. Census `lib/tests/effort-routing.test.sh`; transcript audit
|
never version. Census `lib/tests/effort-routing.test.sh`; transcript audit
|
||||||
`python3 lib/effort-audit.py`.
|
`python3 lib/effort-audit.py`.
|
||||||
|
|
||||||
|
### model-router mod
|
||||||
|
|
||||||
|
`mods/model-router/` is a Claude Code mod (a function-hooks plugin) that applies this table per request. It loads in every session through the tracked symlink `skills/model-router`, as `model-router@skills-dir`.
|
||||||
|
|
||||||
|
- Main loop: every request gets the effort of the phase in force. The mod answers `Skill(effort-*)` itself and applies the level from the next request on, so the five `effort-*` skills no longer load while it is on.
|
||||||
|
- Built-in sub-agents: Explore runs on sonnet/medium, Plan on opus/xhigh. An explicit `model` on the Agent call wins.
|
||||||
|
- User floor: `ultrathink` in a prompt, or a typed `/effort-<level>`, sets the main turn's default and minimum effort.
|
||||||
|
- `/route` (user command) shows or sets the route: `show`, `clear`, `off`, `on`, `reload`, a phase name, `model=<alias|id> effort=<level>`, `switch on|off`, `verbose on|off`. The model sets routes through a `route` tool.
|
||||||
|
- The spinner suffix and the status line under the prompt show the route in force.
|
||||||
|
|
||||||
|
Optional per-machine config: `~/.claude/model-router.json`. Keys: `models` (alias → full id), `windows` (context window per full id), `phases`, `agents`, `skills`, `prompt` (rules), `mainModelSwitch` (default `false`), `verbose` (default `false`), `spinner` (default `true`), `enabled` (default `true`; `false` turns the mod off on that machine). `/route reload` re-reads it.
|
||||||
|
|
||||||
|
Limits: the main loop changes model only with `mainModelSwitch` on, and each switch into another model costs one cold-cache step. The hooks send full model ids, so the `models` table has to follow new model versions.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## Install notes
|
## Install notes
|
||||||
@@ -204,7 +218,8 @@ a different package, ships its own conflicting `graphify` bin) — see
|
|||||||
| `/profile` | Activate a skill profile (web / seo / web-full / full / max / backend / design / dev / qa / audit / minimal) (default: full) |
|
| `/profile` | Activate a skill profile (web / seo / web-full / full / max / backend / design / dev / qa / audit / minimal) (default: full) |
|
||||||
| `/tour` | Grouped all-axes sweep — cleanup + security + reconcile + doc, fix and loop until clean |
|
| `/tour` | Grouped all-axes sweep — cleanup + security + reconcile + doc, fix and loop until clean |
|
||||||
| `/site-motion` | Site-level motion: scroll engine choice, page transitions, pin/scrub sequencing across a page or Astro route (design stack) |
|
| `/site-motion` | Site-level motion: scroll engine choice, page transitions, pin/scrub sequencing across a page or Astro route (design stack) |
|
||||||
| `/effort-low` … `/effort-max` | Effort shifters the orchestrators send per phase; type `/effort-max` to re-run a stuck turn at maximum |
|
| `/effort-low` … `/effort-max` | Effort shifters the orchestrators send per phase (answered by the model-router mod when on); typed by you, they set the main turn's default and minimum effort |
|
||||||
|
| `/route` | model-router mod: show or set the main-loop route (show, clear, off, on, reload, <phase>, model=… effort=…, switch on\|off, verbose on\|off) |
|
||||||
|
|
||||||
> This table lists personal skills. Gstack skills (investigate, review, retro,
|
> This table lists personal skills. Gstack skills (investigate, review, retro,
|
||||||
> office-hours, cso…) and marketplace plugins add many more — run
|
> office-hours, cso…) and marketplace plugins add many more — run
|
||||||
@@ -452,7 +467,7 @@ make profile-reset # go to the default profile (full)
|
|||||||
make new-skill name=myskill # scaffold agent + skill files
|
make new-skill name=myskill # scaffold agent + skill files
|
||||||
```
|
```
|
||||||
|
|
||||||
`doctor.sh` checks: symlinks, GStack submodule, vendored skills (curl-pinned externals in `plugins.lock.json` + `link.sh`'s `EXTERNAL_SKILLS`, per the active profile), Playwright browser cache, prerequisites (git, Node, Cargo, Python, Claude Code), plugins, permissions, token budget, config consistency, git hooks (global core.hooksPath + generated githooks/), scratchpad (TMPDIR quota), Higgsfield CLI and session, seo-data layer.
|
`doctor.sh` checks: symlinks, GStack submodule, vendored skills (curl-pinned externals in `plugins.lock.json` + `link.sh`'s `EXTERNAL_SKILLS`, per the active profile), Playwright browser cache, prerequisites (git, Node, Cargo, Python, Claude Code), plugins, permissions, token budget, config consistency, git hooks (global core.hooksPath + generated githooks/), mods (loading link and `@skills-dir` state), scratchpad (TMPDIR quota), Higgsfield CLI and session, seo-data layer.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
@@ -163,6 +163,7 @@ Tu veux...
|
|||||||
| `/pdf-translate` | Traduire un PDF vers une autre langue | Sortie HTML fidèle (images, layout, style préservés) |
|
| `/pdf-translate` | Traduire un PDF vers une autre langue | Sortie HTML fidèle (images, layout, style préservés) |
|
||||||
| `/impeccable` | Audit/polish design + détecteur anti-slop déterministe | 23 verbes ; `npx impeccable detect` (exit 0/2) |
|
| `/impeccable` | Audit/polish design + détecteur anti-slop déterministe | 23 verbes ; `npx impeccable detect` (exit 0/2) |
|
||||||
| `/tour` | Sweep groupé sur un ou plusieurs projets | Sécu + nettoyage + reconcile + doc, boucle jusqu'à un pass propre |
|
| `/tour` | Sweep groupé sur un ou plusieurs projets | Sécu + nettoyage + reconcile + doc, boucle jusqu'à un pass propre |
|
||||||
|
| `/route` | Voir ou fixer la route du mod model-router | show / clear / off / on / reload / <phase> / model=… effort=… / switch on\|off / verbose on\|off |
|
||||||
| `/profile` | Changer le profil de skills | web / seo / web-full / full / max / backend / design / dev / qa / audit / minimal |
|
| `/profile` | Changer le profil de skills | web / seo / web-full / full / max / backend / design / dev / qa / audit / minimal |
|
||||||
|
|
||||||
> Cette table couvre les skills personnels principaux. Les plugins (gstack,
|
> Cette table couvre les skills personnels principaux. Les plugins (gstack,
|
||||||
@@ -185,6 +186,8 @@ la main relance un tour bloqué au maximum. Les skills externes vendorés
|
|||||||
`lib/effort-pins.txt`. Un skill chargé seul par Claude n'applique pas son
|
`lib/effort-pins.txt`. Un skill chargé seul par Claude n'applique pas son
|
||||||
niveau : il doit partir avec un autre appel d'outil dans le même message.
|
niveau : il doit partir avec un autre appel d'outil dans le même message.
|
||||||
|
|
||||||
|
Avec le mod model-router (`mods/model-router/`, actif dans chaque session), le niveau suit la phase à chaque requête. Le mod répond lui-même à `Skill(effort-*)` : le niveau s'applique dès la requête suivante et le texte des skills `effort-*` n'est plus chargé. Écrire `ultrathink` dans un prompt, ou taper `/effort-<niveau>`, fixe le niveau par défaut et le minimum du tour principal. Les sous-agents intégrés suivent leur route : Explore en sonnet/medium, Plan en opus/xhigh. `/route` affiche ou fixe la route (`/route show`, `/route clear`, `/route off`). La config par machine, optionnelle, vit dans `~/.claude/model-router.json` ; `"enabled": false` y coupe le mod sur cette machine.
|
||||||
|
|
||||||
## Les plugins — décision rapide
|
## Les plugins — décision rapide
|
||||||
|
|
||||||
```
|
```
|
||||||
|
|||||||
@@ -155,6 +155,64 @@ unset _dv_active_profile _dv_profile_file
|
|||||||
|
|
||||||
echo ""
|
echo ""
|
||||||
|
|
||||||
|
# ────────────────────────────────────────────────────────────
|
||||||
|
# 2c. Mods (mods/<name>/ plugins, loaded through the tracked
|
||||||
|
# skills/<name> symlink as <name>@skills-dir). Fail-soft: a missing link
|
||||||
|
# is info (the user may have removed it on purpose), never an error.
|
||||||
|
# ────────────────────────────────────────────────────────────
|
||||||
|
echo "── Mods ──"
|
||||||
|
|
||||||
|
# Prints enabled|disabled|absent|unknown for $1 read from the JSON on stdin;
|
||||||
|
# always exits 0 so a bad payload cannot abort doctor under set -e.
|
||||||
|
mod_state() {
|
||||||
|
python3 -c '
|
||||||
|
import json, sys
|
||||||
|
try:
|
||||||
|
rows = json.load(sys.stdin)
|
||||||
|
row = [r for r in rows if r.get("id") == sys.argv[1] + "@skills-dir"]
|
||||||
|
print("absent" if not row else
|
||||||
|
"enabled" if row[0].get("enabled") is True else "disabled")
|
||||||
|
except Exception:
|
||||||
|
print("unknown")
|
||||||
|
' "$1" 2>/dev/null || true
|
||||||
|
}
|
||||||
|
|
||||||
|
_mods_list=""
|
||||||
|
if command -v claude &>/dev/null; then
|
||||||
|
if ! _mods_list=$(claude plugin list --json 2>/dev/null); then
|
||||||
|
warn "mods: claude plugin list failed — load state not checked"
|
||||||
|
_mods_list=""
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
|
||||||
|
_mods_seen=0
|
||||||
|
for _mod_manifest in "$REPO"/mods/*/.claude-plugin/plugin.json; do
|
||||||
|
[ -f "$_mod_manifest" ] || continue
|
||||||
|
_mods_seen=$((_mods_seen + 1))
|
||||||
|
_mod=$(basename "$(dirname "$(dirname "$_mod_manifest")")")
|
||||||
|
_mod_link="$HOME/.claude/skills/$_mod"
|
||||||
|
if ! { [ -L "$_mod_link" ] || [ -e "$_mod_link" ]; }; then
|
||||||
|
info "mod $_mod: not linked (skills/$_mod absent) — git checkout skills/$_mod if wanted"
|
||||||
|
continue
|
||||||
|
fi
|
||||||
|
if [ "$_mod_link" -ef "$REPO/mods/$_mod" ]; then
|
||||||
|
pass "mod $_mod: loading link ~/.claude/skills/$_mod"
|
||||||
|
else
|
||||||
|
warn "mod $_mod: ~/.claude/skills/$_mod does not resolve to $REPO/mods/$_mod"
|
||||||
|
fi
|
||||||
|
[ -n "$_mods_list" ] || continue
|
||||||
|
case "$(printf '%s' "$_mods_list" | mod_state "$_mod")" in
|
||||||
|
enabled) pass "mod $_mod: enabled as $_mod@skills-dir" ;;
|
||||||
|
disabled) warn "mod $_mod: disabled (\"$_mod@skills-dir\": false in enabledPlugins)" ;;
|
||||||
|
absent) warn "mod $_mod: not listed as @skills-dir — run: claude plugin validate mods/$_mod (policy, manifest or name conflict)" ;;
|
||||||
|
*) warn "mod $_mod: claude plugin list output not understood" ;;
|
||||||
|
esac
|
||||||
|
done
|
||||||
|
[ "$_mods_seen" -gt 0 ] || info "no mods"
|
||||||
|
unset _mods_list _mods_seen _mod_manifest _mod _mod_link
|
||||||
|
|
||||||
|
echo ""
|
||||||
|
|
||||||
# ── Playwright browsers (read-only report; NOT nested under gstack — 2 of
|
# ── Playwright browsers (read-only report; NOT nested under gstack — 2 of
|
||||||
# the 3 registered installs are gsd-pi, not gstack) ──
|
# the 3 registered installs are gsd-pi, not gstack) ──
|
||||||
echo "── Playwright browsers ──"
|
echo "── Playwright browsers ──"
|
||||||
|
|||||||
@@ -0,0 +1,87 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# lib/tests/mods.test.sh — every mods/<name>/ plugin: the manifest name
|
||||||
|
# equals the folder, skills/<name> is the relative loading symlink
|
||||||
|
# ../mods/<name>, and (when the CLI offers `claude plugin test`) the mod
|
||||||
|
# passes `claude plugin validate` without warning and `claude plugin test`.
|
||||||
|
# MODS_ROOT overrides the repo root (fixture controls). Fails when no mod
|
||||||
|
# is found, so it can never pass vacuously.
|
||||||
|
set -u
|
||||||
|
ROOT="${MODS_ROOT:-$(cd "$(dirname "$0")/../.." && pwd)}"
|
||||||
|
CLI_TIMEOUT=120
|
||||||
|
pass=0; fail=0
|
||||||
|
ok() { pass=$((pass+1)); echo "PASS $1"; }
|
||||||
|
ko() { fail=$((fail+1)); echo "FAIL $1"; }
|
||||||
|
check() { if [ "$2" = "$3" ]; then ok "$1"; else ko "$1: got[$2] want[$3]"; fi; }
|
||||||
|
|
||||||
|
# bounded CMD...: stdout+stderr on stdout, rc 124 on timeout.
|
||||||
|
bounded() {
|
||||||
|
local t
|
||||||
|
t=$(command -v timeout || command -v gtimeout || true)
|
||||||
|
if [ -n "$t" ]; then "$t" "$CLI_TIMEOUT" "$@" 2>&1; return; fi
|
||||||
|
local out rc=0 pid i=0
|
||||||
|
out=$(mktemp) || return 1
|
||||||
|
"$@" >"$out" 2>&1 & pid=$!
|
||||||
|
while kill -0 "$pid" 2>/dev/null && [ "$i" -lt "$CLI_TIMEOUT" ]; do
|
||||||
|
sleep 1; i=$((i+1))
|
||||||
|
done
|
||||||
|
if kill -0 "$pid" 2>/dev/null; then kill "$pid" 2>/dev/null; rc=124
|
||||||
|
else wait "$pid" || rc=$?; fi
|
||||||
|
cat "$out"; rm -f "$out"; return "$rc"
|
||||||
|
}
|
||||||
|
|
||||||
|
manifest_name() {
|
||||||
|
python3 -c 'import json,sys; print(json.load(open(sys.argv[1])).get("name",""))' \
|
||||||
|
"$1" 2>/dev/null
|
||||||
|
}
|
||||||
|
|
||||||
|
# cli_unavailable: prints the reason and returns 0 when the capability is missing.
|
||||||
|
cli_unavailable() {
|
||||||
|
command -v claude >/dev/null 2>&1 || { echo "claude not found"; return 0; }
|
||||||
|
local rc=0
|
||||||
|
bounded claude plugin test --help >/dev/null || rc=$?
|
||||||
|
[ "$rc" -eq 124 ] && { echo "probe timed out after ${CLI_TIMEOUT}s"; return 0; }
|
||||||
|
[ "$rc" -ne 0 ] && { echo "no 'claude plugin test' command"; return 0; }
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
|
||||||
|
check_cli() {
|
||||||
|
local name="$1" dir="$2" out rc=0
|
||||||
|
out=$(bounded claude plugin validate "$dir") || rc=$?
|
||||||
|
if [ "$rc" -eq 124 ]; then ko "$name: validate timed out after ${CLI_TIMEOUT}s"
|
||||||
|
elif ! printf '%s' "$out" | grep -q 'Validation passed'; then
|
||||||
|
ko "$name: validate did not pass: $(printf '%s' "$out" | head -3 | tr '\n' ' ')"
|
||||||
|
elif printf '%s' "$out" | grep -qi 'warning'; then
|
||||||
|
ko "$name: validate printed a warning"
|
||||||
|
else ok "$name: validate passed, no warning"; fi
|
||||||
|
rc=0
|
||||||
|
out=$(bounded claude plugin test "$dir") || rc=$?
|
||||||
|
if [ "$rc" -eq 124 ]; then ko "$name: plugin test timed out after ${CLI_TIMEOUT}s"
|
||||||
|
elif [ "$rc" -ne 0 ]; then
|
||||||
|
ko "$name: plugin test rc=$rc: $(printf '%s' "$out" | tail -3 | tr '\n' ' ')"
|
||||||
|
else ok "$name: plugin test passed"; fi
|
||||||
|
}
|
||||||
|
|
||||||
|
manifests=()
|
||||||
|
for m in "$ROOT"/mods/*/.claude-plugin/plugin.json; do
|
||||||
|
[ -f "$m" ] && manifests+=("$m")
|
||||||
|
done
|
||||||
|
if [ "${#manifests[@]}" -eq 0 ]; then ko "no mod found under $ROOT/mods"; fi
|
||||||
|
|
||||||
|
use_cli=1
|
||||||
|
if reason=$(cli_unavailable); then
|
||||||
|
use_cli=0
|
||||||
|
echo "SKIP: claude plugin test unavailable ($reason) — validate/test not run"
|
||||||
|
fi
|
||||||
|
|
||||||
|
for m in ${manifests[@]+"${manifests[@]}"}; do
|
||||||
|
dir="$(dirname "$(dirname "$m")")"; name="$(basename "$dir")"
|
||||||
|
check "$name: manifest name matches folder" "$(manifest_name "$m")" "$name"
|
||||||
|
link="$ROOT/skills/$name"
|
||||||
|
if [ -L "$link" ]; then
|
||||||
|
check "$name: loading link target" "$(readlink "$link")" "../mods/$name"
|
||||||
|
else ko "$name: skills/$name is not a symlink"; fi
|
||||||
|
[ "$use_cli" -eq 1 ] && check_cli "$name" "$dir"
|
||||||
|
done
|
||||||
|
|
||||||
|
echo "mods: $pass pass, $fail fail"
|
||||||
|
[ "$fail" -eq 0 ]
|
||||||
@@ -0,0 +1,8 @@
|
|||||||
|
{
|
||||||
|
"name": "model-router",
|
||||||
|
"version": "0.1.0",
|
||||||
|
"description": "Routes every model request (main loop and sub-agents) to the model and effort its phase deserves, from a user config, a route tool, /route and skill or prompt rules.",
|
||||||
|
"author": {
|
||||||
|
"name": "bchanot"
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
{ "modules": ["./register.ts"] }
|
||||||
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
Symlink
+1
@@ -0,0 +1 @@
|
|||||||
|
../mods/model-router
|
||||||
Reference in New Issue
Block a user