Merge feature/model-router-confirm into develop

This commit is contained in:
bchanot
2026-10-11 16:20:01 +02:00
24 changed files with 4640 additions and 193 deletions
+11
View File
@@ -133,6 +133,7 @@ rules:
| BDR-109 | 2026-09-30 | Higgsfield pack: npm CLI `latest` + 8 upstream skills git-cloned into gitignored `skills-external/higgsfield-*`, OFF by default, in no profile; two toggles (`higgsfield` = allowlist of 7 media skills, `higgsfield-websites` = landing-page aid, never website create/deploy/publish); CLI presence by probe; routing on explicit ask | accepted | | BDR-109 | 2026-09-30 | Higgsfield pack: npm CLI `latest` + 8 upstream skills git-cloned into gitignored `skills-external/higgsfield-*`, OFF by default, in no profile; two toggles (`higgsfield` = allowlist of 7 media skills, `higgsfield-websites` = landing-page aid, never website create/deploy/publish); CLI presence by probe; routing on explicit ask | accepted |
| BDR-110 | 2026-10-06 | Shell portability doctrine: native userland on macOS AND Linux, no Homebrew GNU tools on PATH; `lib/tests/portability-census.test.sh` locks deterministic GNU-only idioms | accepted | | BDR-110 | 2026-10-06 | Shell portability doctrine: native userland on macOS AND Linux, no Homebrew GNU tools on PATH; `lib/tests/portability-census.test.sh` locks deterministic GNU-only idioms | accepted |
| BDR-115 | 2026-10-08 | model-router mod: pin = entry default, sub-tasks route finer; one writer per axis; full ids from the mod table; built-ins-only agents table until frontmatter pins go; state in closure | accepted | | BDR-115 | 2026-10-08 | model-router mod: pin = entry default, sub-tasks route finer; one writer per axis; full ids from the mod table; built-ins-only agents table until frontmatter pins go; state in closure | accepted |
| BDR-116 | 2026-10-11 | model-router first-use confirmation: tracked routing.json = source of phases + rows + decisions; engine dialog at first use; project exceptions in user scope keyed by normalized remote; model never writes; census tolerates a decided row | accepted |
--- ---
@@ -1403,3 +1404,13 @@ Branch feature/user-writing-web-rules, UNMERGED (human gate).
- **Refs**: plan `.claude/tasks/plans/2026-10-08-model-router-mod.md`, contract `.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md`, [[BDR-107]], [[BDR-108]], [[BLK-029]], [[LRN-205]], [[LRN-206]]. - **Refs**: plan `.claude/tasks/plans/2026-10-08-model-router-mod.md`, contract `.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md`, [[BDR-107]], [[BDR-108]], [[BLK-029]], [[LRN-205]], [[LRN-206]].
- **Amendment (2026-10-09, user decisions 2026-10-08 evening)**: (a) LOAD supersedes the "Load:" line: tracked relative symlink `skills/model-router` → `../mods/model-router`, loaded in place as `model-router@skills-dir` wherever link.sh links `~/.claude/skills`; `CLAUDE_CODE_PLUGIN_DIRS` dropped (absolute path, settings `env` has no `$HOME` expansion, settings.json tracked), local marketplace dropped (`add` writes an absolute path into settings.json). Proven by fresh-process `claude plugin list --json`. (b) PRECEDENCE amended: `ultrathink` and a typed `/effort-<l>` are the main turn's DEFAULT and MINIMUM (floor slot `turnFloor`): sticky `/route` effort > turn route effort > floor > engine, then floored; per axis; mid-turn prompt floors the running turn and the next (`wait` ignored). Rationale: user "un choix explicite bat la phase déduite"; a pure floor made `/effort-low` a no-op (challenge finding). (c) Per-machine kill switch `"enabled": false` in the untracked `~/.claude/model-router.json` (survives `/clear`, a failed reload keeps the previous config); `enabledPlugins` would dirty the tracked settings.json on every machine. (d) Hardening: `/route` composer-only; agent loops effort-only (model fixed at spawn); config caps; typed slash attested at `prompt.submit`. Commits 346d6ae, 1ff608a, 6430ac6; contracts `2026-10-08-model-router-floor-1835`, `2026-10-08-model-router-wiring-1835`; residuals parked in TODO. - **Amendment (2026-10-09, user decisions 2026-10-08 evening)**: (a) LOAD supersedes the "Load:" line: tracked relative symlink `skills/model-router` → `../mods/model-router`, loaded in place as `model-router@skills-dir` wherever link.sh links `~/.claude/skills`; `CLAUDE_CODE_PLUGIN_DIRS` dropped (absolute path, settings `env` has no `$HOME` expansion, settings.json tracked), local marketplace dropped (`add` writes an absolute path into settings.json). Proven by fresh-process `claude plugin list --json`. (b) PRECEDENCE amended: `ultrathink` and a typed `/effort-<l>` are the main turn's DEFAULT and MINIMUM (floor slot `turnFloor`): sticky `/route` effort > turn route effort > floor > engine, then floored; per axis; mid-turn prompt floors the running turn and the next (`wait` ignored). Rationale: user "un choix explicite bat la phase déduite"; a pure floor made `/effort-low` a no-op (challenge finding). (c) Per-machine kill switch `"enabled": false` in the untracked `~/.claude/model-router.json` (survives `/clear`, a failed reload keeps the previous config); `enabledPlugins` would dirty the tracked settings.json on every machine. (d) Hardening: `/route` composer-only; agent loops effort-only (model fixed at spawn); config caps; typed slash attested at `prompt.submit`. Commits 346d6ae, 1ff608a, 6430ac6; contracts `2026-10-08-model-router-floor-1835`, `2026-10-08-model-router-wiring-1835`; residuals parked in TODO.
- **Amendment 2 (2026-10-10, wave 2 closed, commits bb56f3e + 1f2d33b)**: (a) rule 4 closed: rows for every repo skill (56) + agent (21) + Explore/Plan, PHASES by role (plan/reflect/orchestrate/escalate best · judge big · implement/write/verify/explore/apply work · mechanical cheap; `write` work/high + `apply` work/low added). User: pins "deleted or reworked", not copied → rows are the live source, tracked `model:`/`effort:` frontmatter KEPT as off-state floor, census-locked equal to rows (`lib/tests/effort-routing.test.sh`). Robustness BLOCKER closed: mod off → agents would inherit parent model. (b) Rule 5 amended: unrowed skill load changes nothing; best-tier skill row lives in `runMain` slot surviving turn end (precedence userMain > turnMain > runMain > floor > engine), dropped by `/route clear`, `/route off`, user `/model`, typed non-best skill; turn writers (route tool, prompt rules) never touch it. (c) Agents: model written at spawn WITHIN tier, upward only (never below frontmatter alias); explicit Agent params win; project-defined agent (agent.offer source projectSettings|localSettings) skipped. (d) Typed slash: name-bound marker at prompt.submit (composer|sdk|bridge) + pending slot + idle fallback (no live/spawning loop). (e) Shifters `effort-*`, `effort-pins.*`, `model-check.sh` DELETED; orchestrators call `mcp__model-router__route` per phase; gate witness = route answer id, remedy `/route on`; builtin `/effort` not a lever inside a run (levers `ultrathink`, `/route effort=max`). (f) SemVer: typed `/effort-*` removal = breaking → next release 3.0.0. Plan `.claude/tasks/plans/2026-10-09-model-router-w2-1546.md` r4, contracts `2026-10-09-model-router-w2a-1546`, `2026-10-10-model-router-w2b-1045`. Links [[LRN-210]], [[LRN-211]], [[EVAL-042]]. - **Amendment 2 (2026-10-10, wave 2 closed, commits bb56f3e + 1f2d33b)**: (a) rule 4 closed: rows for every repo skill (56) + agent (21) + Explore/Plan, PHASES by role (plan/reflect/orchestrate/escalate best · judge big · implement/write/verify/explore/apply work · mechanical cheap; `write` work/high + `apply` work/low added). User: pins "deleted or reworked", not copied → rows are the live source, tracked `model:`/`effort:` frontmatter KEPT as off-state floor, census-locked equal to rows (`lib/tests/effort-routing.test.sh`). Robustness BLOCKER closed: mod off → agents would inherit parent model. (b) Rule 5 amended: unrowed skill load changes nothing; best-tier skill row lives in `runMain` slot surviving turn end (precedence userMain > turnMain > runMain > floor > engine), dropped by `/route clear`, `/route off`, user `/model`, typed non-best skill; turn writers (route tool, prompt rules) never touch it. (c) Agents: model written at spawn WITHIN tier, upward only (never below frontmatter alias); explicit Agent params win; project-defined agent (agent.offer source projectSettings|localSettings) skipped. (d) Typed slash: name-bound marker at prompt.submit (composer|sdk|bridge) + pending slot + idle fallback (no live/spawning loop). (e) Shifters `effort-*`, `effort-pins.*`, `model-check.sh` DELETED; orchestrators call `mcp__model-router__route` per phase; gate witness = route answer id, remedy `/route on`; builtin `/effort` not a lever inside a run (levers `ultrathink`, `/route effort=max`). (f) SemVer: typed `/effort-*` removal = breaking → next release 3.0.0. Plan `.claude/tasks/plans/2026-10-09-model-router-w2-1546.md` r4, contracts `2026-10-09-model-router-w2a-1546`, `2026-10-10-model-router-w2b-1045`. Links [[LRN-210]], [[LRN-211]], [[EVAL-042]].
## BDR-116 — model-router first-use confirmation: tracked `routing.json` = source of phases, rows and decisions; engine dialog once per row/phase; project exceptions in user scope; model never writes [accepted] (2026-10-11)
- **Decision**: (1) `mods/model-router/routing.json` (tracked, reached from every project through the plugin dir `$.plugin.root` = the existing `~/.claude/skills/model-router` link) is the single source: 11 phases (full routes), 56 skill rows, 23 agent rows, `confirmed` (kind/name → phase), `changed` (from/to for an Everywhere change), `projects[<key>]` exceptions, `ask`. `DEFAULT_CONFIG` keeps only the phases as fallback; rows `{}`. (2) First use of a rowed typed skill (T1), a rowed agent spawn with no explicit model (T2) or a main-loop phase declared via the route tool (T3) opens `$.ui.ask`: options Later / Keep / 2 alt phases (Other = phase name; T3 Later/Keep); a change asks Everywhere (row moved + `changed`) or This project only (`projects[key]`, confirmed = base row so no other project asks). Keep endorses the phase too. One dialog in flight; concurrent uses route unasked; headless/dismissed/garbage = Later; never inside a sub-agent; pre-ask re-read of the file (other sessions' decisions seen). (3) Project key = origin remote normalized (`new URL` host+path or strict scp regex; userinfo never read; any remaining `@`/`:`/empty host → no key); no `local:` path keys; no remote → Everywhere/Later only. The project tree's `.claude/model-router.json` is NEVER read (a cloned repo must not re-route the user's gate agents). (4) Writers = dialog answers + composer `/route ask on|off` only; serialized chain; output capped 64 KB; file never created; later unreadable → previous config kept (kill-switch rule); layers routing.json < `~/.claude/model-router.json` (its `ask` wins). (5) Census reads rows + phases from the file, locks DEFAULT phases == file, tolerates a drift only for a `changed` row whose frontmatter == `from` and row == `to` (WARN); Keep-only drift still FAILs. (6) `/route pending`, `/route ask on|off`; `set`/`confirm`/show suffix deferred.
- **Why**: user 2026-10-10: see in practice whether routing fits ("prompt qui demande de confirmer… mémoire de ce qu'on décide… met à jour la table… exception pour ce projet… persistant sur tous les projets, se redéploie comme le mod"). Tracked file = deploys with the mod via git; user-scope exceptions = security (robustness lens: project-tree layer let a cloned repo downgrade security-auditor to haiku).
- **Alternatives rejected**: `$.store` (machine-local, re-asks per machine); project file in the tree (security hole + writes into worktrees); shared-promise dedupe for parallel spawns (hook budget: awaiters time out); `confirmed` dates (git log dates them); dialog edits of phases (blast radius: a phase is shared by many rows); per-layer degrade after first load (one transient read failure wiped 79 rows + the kill switch); `local:<realpath>` keys (home path in a tracked file).
- **Gates**: plan r1 → r4 (simplicity CONCERNS(3), correctness FATAL(8) with the census BLOCKER, robustness CONCERNS(8) after a network-killed first run, confirmation CONCERNS(6)); feater DONE + 4 rounds; GATE 0 MET; verifier ECARTS(3)/(2) → CONFORME, re-verify after security ECARTS(1) → fixtures; security BLOCK(1) real (password with `@`/`/` stored in the key) → fixed, PASS. Kit 86 → 190 tests. Live T2 dialogs answered by the user from the hot-loaded working-tree mod ([[LRN-212]]).
- **Refs**: contract `.claude/tasks/contracts/2026-10-10-model-router-w3a-confirm-1201.md`, plan `.claude/tasks/plans/2026-10-10-model-router-w3a-confirm-1201.md`, commit 22455c0, [[BDR-115]], [[LRN-210]], [[LRN-211]], [[EVAL-043]].
- **Amendment (2026-10-11, wave 3-B, commit 604a6c4)**: clause (2) superseded: the dialog carries CONTEXT (skill description = first sentence of its SKILL.md frontmatter, five YAML forms, hardened read: name allowlist, stat kind file, size cap, `clean()` Cc/Cf + code points; agent description from `agent.offer`, bounded cache; phase `about` line, new 120-char field on the 11 routing.json phases, kept in memory not on Route) and the REAL next model id + effort. Options Later / Keep / Change (T3 Later / Keep). Change = model (4 fixed alias labels with tier role) → effort among those the phases headed by that alias offer (ascending, ≤ 4, one → skipped, zero → Later) → scope; the pair maps to an existing PHASE (rows stay phase names; row's current phase wins a tie; same phase = Keep); a pair no phase offers = a new phase by hand. Other anywhere = Later + toast (no echo). A main-row change toasts the real decision + `/route switch on` hint on a held downgrade. Three lenses rejected inline-route rows and dialog phase edits (2 BLOCKERs). First real use: user moved security-auditor to `judge` (opus/xhigh) through it; floor aligned (`model: opus`, 2026-10-11). Links [[EVAL-044]].
- **Amendment (2026-10-11, wave 3-C, commit cd8d72f)**: clause (6) extended: `/route forget <name|all|projects>` (composer only) clears decisions in the tracked file through the same writer: a name in every table (`confirmed`, `changed` → row restored to its recorded `from` when it is a string, the row still equals `to` and `from` is a known phase; `projects[*]` pruned), asked again at once; `all`/`projects` confirm through the engine dialog (Cancel first, anything else = cancelled) and hold the single-dialog slot; nothing written when nothing to forget; the answer names a restored row's file + values to realign (a floor aligned to the decision makes the census FAIL after the restore). No per-user decisions file (user: "seulement le forget"): another user inherits the shared memory and may forget. Residuals (security LOW/MEDIUM): raw error text in the answer, file fragments unsanitized in the answer, `floorFile` name not allowlisted, `__proto__` plan/apply mismatch, `changed.phases` restorable. Links [[EVAL-044]].
+17
View File
@@ -63,6 +63,8 @@ rules:
| EVAL-040 | 2026-10-08 | model-router w1a plan: 3 challengers + 1 confirmation found 2 BLOCKER + 14 MAJOR on a plan judged closed; executor then passed every gate first time | keep the round, never dispatch a mod plan without it | | EVAL-040 | 2026-10-08 | model-router w1a plan: 3 challengers + 1 confirmation found 2 BLOCKER + 14 MAJOR on a plan judged closed; executor then passed every gate first time | keep the round, never dispatch a mod plan without it |
| EVAL-041 | 2026-10-09 | model-router W1-C plan: 3 lenses FATAL (4 BLOCKER + 20 MAJOR) then 2 confirmations each FATAL with a NEW BLOCKER in my own revision; executor DONE first pass, 3 short text/hardening rounds | one confirmation is not enough when a revision removes a whole mechanism; the plan carried the risk, the code almost none | | EVAL-041 | 2026-10-09 | model-router W1-C plan: 3 lenses FATAL (4 BLOCKER + 20 MAJOR) then 2 confirmations each FATAL with a NEW BLOCKER in my own revision; executor DONE first pass, 3 short text/hardening rounds | one confirmation is not enough when a revision removes a whole mechanism; the plan carried the risk, the code almost none |
| EVAL-042 | 2026-10-10 | model-router W2 plan r1 → r4: 3 lenses (1 BLOCKER), 2 confirmations (1 BLOCKER then 0); W2-A verifier 3× ECARTS on coverage clauses only, W2-B ECARTS(7) → CONFORME; gate A→B read from engine records | second confirmation paid again (BLOCKER on my own r2 slot); compound coverage criterion = endless ECARTS; engine jsonl replaces the live log | | EVAL-042 | 2026-10-10 | model-router W2 plan r1 → r4: 3 lenses (1 BLOCKER), 2 confirmations (1 BLOCKER then 0); W2-A verifier 3× ECARTS on coverage clauses only, W2-B ECARTS(7) → CONFORME; gate A→B read from engine records | second confirmation paid again (BLOCKER on my own r2 slot); compound coverage criterion = endless ECARTS; engine jsonl replaces the live log |
| EVAL-043 | 2026-10-11 | model-router W3-A: 3 lenses + 1 confirmation (census BLOCKER, project-tree layer dropped), feater DONE + 4 rounds, verifier 3× then re-verify after a REAL security BLOCK (credential fragment in the tracked key) | challenge + security gates both earned their cost; coverage-shaped criteria still cost 3 verifier rounds; live mod side effects misread as a test leak |
| EVAL-044 | 2026-10-11 | model-router W3-B (dialog with context): 3 lenses converged on 2 BLOCKERs of my r1 (inline rows, phase edits) + 1 confirmation; feater DONE + 2 coverage rounds; verifier 3× to CONFORME; security PASS; user's first real decision through the new dialog | the lenses pay most when the plan adds a data shape; coverage criteria still cost rounds; live dialogs during the gates must be announced and their data reconciled before commit |
--- ---
@@ -394,3 +396,18 @@ Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itse
- **Method**: 3 blind lenses (simplicity CONCERNS(5), correctness CONCERNS(10), robustness FATAL(8) → BLOCKER: mod off = agents inherit parent model); confirmation 1 robustness FATAL(8) → BLOCKER introduced by r2 (route calls wiped the sticky slot); confirmation 2 correctness CONCERNS(3), no BLOCKER → r4. W2-A: feater DONE + 4 rounds (1 internal decision, 3 coverage), GATE 0 MET, verifier ECARTS(3)/(1)/(1) all coverage, user accepted at cap; security PASS. W2-B: feater DONE first pass, verifier ECARTS(7) (5 FLOOR items = planned deletions needing a CLARIFICATIONS line, 2 prose, 1 scope add) → CONFORME 10/10; security PASS; full `make test` once (env red only). Gate A→B: 3/4 probes answered from engine jsonl, probe 4 (typed skill with a live agent) unobserved, recorded as a limit. - **Method**: 3 blind lenses (simplicity CONCERNS(5), correctness CONCERNS(10), robustness FATAL(8) → BLOCKER: mod off = agents inherit parent model); confirmation 1 robustness FATAL(8) → BLOCKER introduced by r2 (route calls wiped the sticky slot); confirmation 2 correctness CONCERNS(3), no BLOCKER → r4. W2-A: feater DONE + 4 rounds (1 internal decision, 3 coverage), GATE 0 MET, verifier ECARTS(3)/(1)/(1) all coverage, user accepted at cap; security PASS. W2-B: feater DONE first pass, verifier ECARTS(7) (5 FLOOR items = planned deletions needing a CLARIFICATIONS line, 2 prose, 1 scope add) → CONFORME 10/10; security PASS; full `make test` once (env red only). Gate A→B: 3/4 probes answered from engine jsonl, probe 4 (typed skill with a live agent) unobserved, recorded as a limit.
- **Anomaly**: the user's rework answer ("delete or rework, not copy") changed the design mid-plan; r2's own fix carried a BLOCKER again (as in EVAL-041). Coverage criterion: 3 verifiers, 0 defects. FLOOR guard needs the test deletions named in CLARIFICATIONS, not only in criteria. - **Anomaly**: the user's rework answer ("delete or rework, not copy") changed the design mid-plan; r2's own fix carried a BLOCKER again (as in EVAL-041). Coverage criterion: 3 verifiers, 0 defects. FLOOR guard needs the test deletions named in CLARIFICATIONS, not only in criteria.
- **Action**: keep the "second confirmation after a mechanism change" rule; write coverage criteria one clause each (LRN-210); when a plan deletes tests, write the authorizing CLARIFICATIONS line BEFORE the first verifier. Links [[EVAL-041]], [[LRN-210]], [[LRN-211]], [[BDR-115]]. - **Action**: keep the "second confirmation after a mechanism change" rule; write coverage criteria one clause each (LRN-210); when a plan deletes tests, write the authorizing CLARIFICATIONS line BEFORE the first verifier. Links [[EVAL-041]], [[LRN-210]], [[LRN-211]], [[BDR-115]].
## EVAL-043 — model-router W3-A: the gates caught two real defects, the coverage criteria still cost three verifier rounds
- **Date**: 2026-10-11
- **Output checked**: plan `.claude/tasks/plans/2026-10-10-model-router-w3a-confirm-1201.md` r1 → r4; diff 22455c0 (register.ts, routing.json, 190 kit tests, census).
- **Method**: 3 lenses (simplicity CONCERNS(3), correctness FATAL(8): BLOCKER = an Everywhere decision turns the census red; robustness: first run died on a DNS error, fresh re-dispatch on r2 CONCERNS(8): project-tree layer = security hole, shared-promise dedupe burns hook budgets, per-layer degrade wipes rows); confirmation correctness CONCERNS(6) → r4. Feater DONE first pass (169 tests); verifier ECARTS(3) → feater → ECARTS(2) → feater → CONFORME; security BLOCK(1): `normalizeRemote` kept a password tail when it held `@` or `/` → fixed (strict URL/scp parsing, no `local:` keys, output cap) → re-verify ECARTS(1) (guard fixtures) → feater → re-scan PASS. Full `make test` once (env red only).
- **Anomaly**: the live hot-loaded mod answered-by-user dialogs were first misread as a kit write leak (LRN-212). A GATE 0 CHECK written as a multi-line heredoc is not runnable by gates.sh (one line or a script). Criterion 3 again bundled ~15 clauses: three verifier rounds on coverage, zero code defects from them (LRN-210 not yet applied by me).
- **Action**: write GATE 0 oracles as scripts from the start; split coverage criteria per clause BEFORE the first verifier; when a mod is under development, announce the live side effects at each dispatch. Links [[EVAL-042]], [[BDR-116]], [[LRN-212]].
## EVAL-044 — model-router W3-B: the lenses killed a second row format before it existed; the live dialog answered the request's own example
- **Date**: 2026-10-11
- **Output checked**: plan `.claude/tasks/plans/2026-10-11-model-router-w3b-dialog-1240.md` r1 → r3; diff 604a6c4 (232 kit tests), data commits 32b71d2 (security-auditor → judge) and the floor alignment.
- **Method**: simplicity CONCERNS(3), robustness FATAL(5), correctness FATAL(8): all three rejected inline `{model, effort}` rows (string tables everywhere, breaker semantics, PHASE_KEY on row names) and dialog edits of phases (census lock, blast radius, BDR-116) → r2 phases-only with efforts derived from the alias's phases; confirmation CONCERNS(2) → r3 (test migration list, read-hardening tests, ordering, clean() order). Feater DONE first pass (229 tests, 26 migrated, 41 new, 9 mutation proofs); verifier ECARTS(2) coverage → ECARTS(1) non-discriminating HOME test (the engine resolves a relative path against the plugin root) → CONFORME; security PASS (4 LOW). GATE 0 heredoc CHECK again not runnable → script (second time: write oracles as scripts from the start).
- **Anomaly**: my r1 proposed a second row shape to satisfy "model then effort" literally; the derived-effort list satisfies it with zero new data shapes. During the gates the user answered the NEW dialog live: security-auditor → opus/xhigh (the request's example), plan-challenger → verify (reset on user go); my reset of `changed` wiped a legitimate entry and left the committed file census-red for one commit.
- **Action**: when a request says "choose X then Y", look for the existing shape that already encodes (X, Y) before adding one; reconcile routing.json decisions one by one (never `changed = {}`), and run the census before every data commit. Links [[EVAL-043]], [[BDR-116]], [[LRN-212]].
+5
View File
@@ -594,3 +594,8 @@ rules:
## 2026-10-10 ## 2026-10-10
- model-router W2-B landed (1f2d33b) + gate A→B. User /reload-plugins + typed /status; I read the engine jsonl instead of the UI log: main low on typed /status, analyzer opus/xhigh at step 0 → ordering + row-over-frontmatter proven; probe 4 unobserved (user away), limit recorded. Route tool checked live (`mcp__model-router__route`, deferred → ToolSearch once; answer names the id); session itself routed through it (reflect → orchestrate). W2-B: contract 10 criteria, feater DONE first pass (43 files), verifier ECARTS(7): 5 FLOOR items (planned test deletions → CLARIFICATIONS line), `/route on` wording, false CLAUDE.global.md sentence, plan-challenger scope add → CONFORME 10/10; I folded 3 observations by hand (SDD sonnet implementers `effort="medium"`, run-slot droppers, 80-col). Security PASS (push-guard false positive on a grep pattern; it also bit my registry heredoc). Full make test green (env red only). Doc audit SIGNIFICANT → user: apply all P1-P8, SemVer BREAKING → 3.0.0, registries all → BDR-115 amendment 2, LRN-210/211, EVAL-042. Pending: doc commit, memory commit, user merge decision (gitflow finish, human signal), publish by hand. - model-router W2-B landed (1f2d33b) + gate A→B. User /reload-plugins + typed /status; I read the engine jsonl instead of the UI log: main low on typed /status, analyzer opus/xhigh at step 0 → ordering + row-over-frontmatter proven; probe 4 unobserved (user away), limit recorded. Route tool checked live (`mcp__model-router__route`, deferred → ToolSearch once; answer names the id); session itself routed through it (reflect → orchestrate). W2-B: contract 10 criteria, feater DONE first pass (43 files), verifier ECARTS(7): 5 FLOOR items (planned test deletions → CLARIFICATIONS line), `/route on` wording, false CLAUDE.global.md sentence, plan-challenger scope add → CONFORME 10/10; I folded 3 observations by hand (SDD sonnet implementers `effort="medium"`, run-slot droppers, 80-col). Security PASS (push-guard false positive on a grep pattern; it also bit my registry heredoc). Full make test green (env red only). Doc audit SIGNIFICANT → user: apply all P1-P8, SemVer BREAKING → 3.0.0, registries all → BDR-115 amendment 2, LRN-210/211, EVAL-042. Pending: doc commit, memory commit, user merge decision (gitflow finish, human signal), publish by hand.
## 2026-10-11
- model-router W3-A (first-use confirmation) landed 22455c0 on feature/model-router-confirm. User asked for it after the W2 merge; 3 pass-B answers. Plan r1→r4: correctness BLOCKER (Everywhere decision = census red → `changed` WARN exemption), robustness (first challenger killed by DNS, fresh one: project-tree layer dropped for security, single dialog in flight, keep-previous after first load), confirmation CONCERNS(6). Feater DONE (169 tests) + 4 rounds. Verifier ECARTS(3)/(2)/CONFORME; security BLOCK(1) REAL: password with `@`/`/` landed in the tracked key → strict parsing, no `local:` keys, output cap → re-verify ECARTS(1) fixtures → PASS. The working-tree mod was hot-loaded by the engine: my own dispatches opened T2 dialogs the user answered (verifier Keep, feater Keep, security-auditor → implement → user reset to verify); first misread as a test leak (LRN-212). GATE 0 oracle as heredoc not runnable → script. Full make test green (env red only). Docs P1-P13 user-approved (patch in flight); BDR-116, LRN-212, EVAL-043 written. Next: doc commit, memory commit, user merge + publish by hand; T1/T3 live checks open.
- model-router W3-B (dialog with context) landed 604a6c4 + data 32b71d2 + floor chore. User asked after W3-A: explain what is decided, OK/change/later, model then effort then scope. Plan r1→r3: all three lenses killed my inline-row format and the T3 phase edit (2 BLOCKERs avoided); efforts derived from the phases of the chosen model. Feater DONE first pass (229 tests), verifier ECARTS(2)→(1)→CONFORME, security PASS. GATE 0 heredoc CHECK not runnable again → script. Live during the gates: user answered the NEW dialog for security-auditor → opus/xhigh (the request's example; kept, floor aligned to opus, model-routing lock updated), plan-challenger → verify (reset on user go), orchestrate Keep (T3 check done). My `changed = {}` reset wiped a legit entry → census red for one commit → restored. Docs P1-P6 approved (patch in flight), BDR-116 amendment + EVAL-044 written. Next: doc commit, memory commit, user merge of feature/model-router-confirm (W3-A + W3-B) + publish by hand; T1 live check still open.
- model-router W3-C `/route forget` landed cd8d72f (+ data commit: escalate Keep). User asked whether decisions can be cleared and what another user can do; chose the command only, no per-user file. Plan r1→r3: three lenses CONCERNS (false `st.mem` premise, kind order, docs scope, census red after restoring an aligned floor, no confirmation on `all`), confirmation CONCERNS(3). Feater DONE; verifier ECARTS(2) with a REAL typo (`ForgetPlan`), then coverage only ×2 → cap → user accepted; security PASS (1 MEDIUM + 4 LOW parked). My own `route(escalate)` call opened the T3 dialog the user Kept. Branch feature/model-router-confirm now carries W3-A/B/C (14 commits), unmerged, unpushed; merge on signal; full make test running.
+5
View File
@@ -1784,3 +1784,8 @@ Rule: when editing a doctrine file under structure locks, grep the test's lock s
## LRN-211 — Typed-slash routing: name-bound marker + idle fallback; run slot separate from turn routes; engine records are the live oracle ## LRN-211 — Typed-slash routing: name-bound marker + idle fallback; run slot separate from turn routes; engine records are the live oracle
- **Context**: mod needs "user typed /feat" from `skill.prompt`, which carries no origin. `prompt.submit` sees the raw `/name` first (composer|sdk|bridge): store the NAME (not a boolean; a bare flag leaked to the next preload), pending slot when mid-turn, consume only on the matching `skill.prompt`; fallback = no live/spawning loop AND allowed origin. Sticky run route in the SAME slot as turn routes was wiped by the first `route()` call (confirmation BLOCKER) → separate `runMain`, best-tier rows only (work/cheap rows leak low effort across turns). Live facts read from `~/.claude/projects/<repo>/<session>.jsonl` (+ `subagents/agent-*.jsonl`): `effort` + `message.model` per step. Typed `/status` → main low; analyzer step 0 opus/xhigh with frontmatter high (spawn bookkeeping precedes step 0; row beats frontmatter). - **Context**: mod needs "user typed /feat" from `skill.prompt`, which carries no origin. `prompt.submit` sees the raw `/name` first (composer|sdk|bridge): store the NAME (not a boolean; a bare flag leaked to the next preload), pending slot when mid-turn, consume only on the matching `skill.prompt`; fallback = no live/spawning loop AND allowed origin. Sticky run route in the SAME slot as turn routes was wiped by the first `route()` call (confirmation BLOCKER) → separate `runMain`, best-tier rows only (work/cheap rows leak low effort across turns). Live facts read from `~/.claude/projects/<repo>/<session>.jsonl` (+ `subagents/agent-*.jsonl`): `effort` + `message.model` per step. Typed `/status` → main low; analyzer step 0 opus/xhigh with frontmatter high (spawn bookkeeping precedes step 0; row beats frontmatter).
- **Apply**: hook-side user-intent markers: bind to a name, add a pending slot for mid-turn, keep an ordering-independent fallback. Two lifetimes = two slots, never one slot with a source tag. Verify engine behaviour in the transcript jsonl, not in `$.ui.log`. Links [[BDR-115]], [[LRN-206]]. - **Apply**: hook-side user-intent markers: bind to a name, add a pending slot for mid-turn, keep an ordering-independent fallback. Two lifetimes = two slots, never one slot with a source tag. Verify engine behaviour in the transcript jsonl, not in `$.ui.log`. Links [[BDR-115]], [[LRN-206]].
## LRN-212 — The engine hot-reloads a mod's hooks module from the working tree: unverified code runs live during its own feature run; dialogs answered there are facts; the kit never touches the disk
- **Context**: W3-A, 2026-10-10. Mid-run, `routing.json` gained decisions nobody expected (verifier → judge, feater → judge). First read: a kit test wrote the real file. Truth (feater, transcript timestamps = file mtimes to the second): the engine had reloaded `register.ts` from the working tree through the `skills/model-router` link, the live mod opened the T2 dialog on MY verifier/feater spawns, the user answered them in the terminal. The kit cannot write: an unmocked `fs.write` is refused ("no implementation"), a throwing mock lands on the same refusal; the real file's sha was stable across 10+ suite runs.
- **Apply**: while a mod is under development in this repo, its working-tree code is LIVE in every session (no `/reload-plugins` needed): expect its side effects (dialogs, writes) during the gates; tell the user what dialogs will pop and what to answer; reset the data file deliberately at the gate. Blame the kit last: check mtimes against the transcript before assuming a test leak. Live answers = criterion evidence (record them `[gated]`). Links [[BDR-116]], [[LRN-206]], [[LRN-211]].
+12
View File
@@ -23,6 +23,18 @@ migration of shifters/pins/model-gate in wave 2 after proof; names model-router
- [ ] (was) W2 gate A→B (user): `/reload-plugins`, then the live probe of plan § Gate (4 points: typed `/status` marker vs fallback in the verbose log; a real rowed spawn line + `step 0 agent` effort = spawn/first-step ordering; sonnet session typed `/feat` self-check + route answer; probe 1 again with a background agent alive). `typed-marker` never seen → W2-B blocked, A4 re-planned. - [ ] (was) W2 gate A→B (user): `/reload-plugins`, then the live probe of plan § Gate (4 points: typed `/status` marker vs fallback in the verbose log; a real rowed spawn line + `step 0 agent` effort = spawn/first-step ordering; sonnet session typed `/feat` self-check + route answer; probe 1 again with a background agent alive). `typed-marker` never seen → W2-B blocked, A4 re-planned.
- [x] (was) W2-B repo migration (plan § W2-B B0-B7 + STEP 6/7): bridge removal, `lib/effort-shift.md` rewrite, 15 citers → `route`, `effort=` on opus general-purpose dispatches, slim `lib/model-gate.md`, delete `model-check.sh` + `effort-pins.*` + their tests + install/update blocks, delete `skills/effort-*`, analyzer `effort: xhigh`, census rewrite (drift lock rows ↔ frontmatter), docs + registries (BDR-115 amendment, LRN typed-slash/run slot, EVAL) - [x] (was) W2-B repo migration (plan § W2-B B0-B7 + STEP 6/7): bridge removal, `lib/effort-shift.md` rewrite, 15 citers → `route`, `effort=` on opus general-purpose dispatches, slim `lib/model-gate.md`, delete `model-check.sh` + `effort-pins.*` + their tests + install/update blocks, delete `skills/effort-*`, analyzer `effort: xhigh`, census rewrite (drift lock rows ↔ frontmatter), docs + registries (BDR-115 amendment, LRN typed-slash/run slot, EVAL)
- [ ] W2-A residuals (security, accepted): first-load failure of the override activates the router despite `enabled:false` (fix = treat a failed first load as off); ReDoS on a self-authored prompt pattern (size-bounded); error text in the local log; model alias keys unvalidated (PHASE_KEY would do); `offers` map uncapped; `__proto__`/`constructor` override keys untested. Known limits: `skillCalls`/`spawning` counters are global; offers keyed by name only; builtin `/effort` is not a lever inside a run. - [ ] W2-A residuals (security, accepted): first-load failure of the override activates the router despite `enabled:false` (fix = treat a failed first load as off); ReDoS on a self-authored prompt pattern (size-bounded); error text in the local log; model alias keys unvalidated (PHASE_KEY would do); `offers` map uncapped; `__proto__`/`constructor` override keys untested. Known limits: `skillCalls`/`spawning` counters are global; offers keyed by name only; builtin `/effort` is not a lever inside a run.
## 2026-10-10 — model-router wave 3-A: first-use route confirmation + decision memory (feature/model-router-confirm)
Plan `.claude/tasks/plans/2026-10-10-model-router-w3a-confirm-1201.md` r4, contract `2026-10-10-model-router-w3a-confirm-1201`. User decisions 2026-10-10: tracked routing.json reached through the plugin dir (no new link); blocking `$.ui.ask` dialog at first use; per skill row, agent row, main phase.
- [x] W3-A landed (22455c0, 2026-10-11): routing.json = phases + rows + decisions; T1/T2/T3 dialogs; project exceptions `projects[<normalized remote>]` in the tracked file (project-tree layer dropped: security); writers = dialog + `/route ask`; `/route pending`; census from the file with the `changed` WARN exemption; kit 86 → 190. Gates: 3 lenses + 1 confirmation (BLOCKER census), feater + 4 rounds, verifier CONFORME + re-verify, security BLOCK(1) credential leak fixed → PASS, full make test green (design-tool-gate env red). Live: T2 dialogs answered by the user (verifier Keep, feater Keep, security-auditor → implement then reset to verify on user go).
- [x] W3-B dialog with context (2026-10-11, 604a6c4, contract `2026-10-11-model-router-w3b-dialog-1240`, plan r3): descriptions (skill frontmatter / agent.offer / phase `about`), real id + effort in the question, Later / Keep / Change, model → effort (derived from the phases of that model) → scope, rows stay phases, T3 Later/Keep. 3 lenses (2 BLOCKERs on my r1 avoided) + 1 confirmation, feater + 2 rounds, verifier CONFORME (3rd), security PASS (4 LOW: symlinked SKILL.md read, TOCTOU, raw key in a log, any plugin hooking AskUserQuestion could answer). Live decisions: security-auditor → judge (kept, floor aligned `model: opus`), plan-challenger → verify (reset on user go), orchestrate Keep (T3 live check done).
- [ ] W3-B residuals: a Change answered Everywhere can store a machine-only phase name (defined in `~/.claude/model-router.json`) into the tracked file (security note, correctness); SKILL.md symlink follow (LOW); `descs` cap and "(custom)" label untested; the security-auditor's `changed.from = verify` entry stays in routing.json although the floor now equals the row (harmless; prune at release).
- [ ] trace the "auto mode: use Bash/sed instead of Edit/Write" instruction block that sub-agents see attributed to the model-router MCP instructions (security-auditor 2026-10-11): it is NOT in mods/model-router source; likely the harness's own auto-mode text rendered next to the mod's MCP block. Confirm the origin.
- [x] W3-C `/route forget` (2026-10-11, cd8d72f, contract `2026-10-11-model-router-w3c-forget-1457`, plan r3): name/all/projects, guarded restore, confirmation dialog for all/projects, no-write paths, docs. 3 lenses + 1 confirmation, feater + 3 rounds (1 real defect: `ForgetPlan` typo), verifier at the cap on coverage (user accepted), security PASS. Kit 232 → 284. Live: orchestrate + escalate T3 Keeps seen.
- [ ] W3-C residuals (security, accepted): `forget not applied (${String(err)})` echoes the raw writer error (send to the log); file-sourced `from`/`to`/keys unsanitized in the answer (clean() them); `floorFile` builds a display path from an unvalidated key (SKILL_NAME check); `__proto__` accepted by `restoreOf` while `setKey` refuses it (plan/apply mismatch); `changed.phases` entries restorable (limit the scan to ROWS). Live smoke of `/route forget` from the terminal still unseen (criterion 7).
- [ ] W3-A live checks still open: T1 (typed `/feat` etc.) dialog (T3 seen live 2026-10-11: orchestrate Keep); a parallel same-agent dispatch answered after > 10 s; what an unanswered dialog resolves to on the terminal (must be Later or nothing written). Record in the contract as `[gated]` when seen.
- [ ] W3-A residuals (security LOW, accepted): non-atomic cross-process write (two sessions, crash mid-write → file reads as failed, previous config kept); `host/path` of a private remote in the tracked file; a persistent write failure re-asks every use (`asked.delete` in the catch); `out.length` vs byte size at the cap; `changePatch` scope 'project' with a vanished key writes Everywhere (cwd change mid-dialog). Deferred: `/route set`, `/route confirm`, `(unconfirmed)` in show, dialog edits of phases, frontmatter auto-alignment, session-start warning when routing.json is dirty, per-machine `projects`.
- [ ] routing.json ships `confirmed` for verifier/feater (user Keeps) and phases verify/implement: every clone inherits them (by design: decisions travel). Revisit at release if unwanted.
- [ ] W2 residuals: probe 4 unobserved (above); the push-guard hook denies a read-only grep (or a heredoc) whose TEXT contains the push verb next to `git` (security-auditor + orchestrator 2026-10-10, false positives, reworded); ~27 loose "sonnet pin"/"opus pin" shorthand sites kept (true by census); README config key list omits `tiers`/`fallback`/`cooldownMinutes`/`mainUpgrade`/`upgradeMaxTokens` (doc audit item 6, pre-existing); MIGRATION.md "Upgrading to 3.0.0" at release time; `skillCalls`/`spawning` counters global; offers keyed by name. - [ ] W2 residuals: probe 4 unobserved (above); the push-guard hook denies a read-only grep (or a heredoc) whose TEXT contains the push verb next to `git` (security-auditor + orchestrator 2026-10-10, false positives, reworded); ~27 loose "sonnet pin"/"opus pin" shorthand sites kept (true by census); README config key list omits `tiers`/`fallback`/`cooldownMinutes`/`mainUpgrade`/`upgradeMaxTokens` (doc audit item 6, pre-existing); MIGRATION.md "Upgrading to 3.0.0" at release time; `skillCalls`/`spawning` counters global; offers keyed by name.
- [ ] W3 optional: step heuristics, haiku classifier, quota-aware downgrade, A/B - [ ] W3 optional: step heuristics, haiku classifier, quota-aware downgrade, A/B
@@ -0,0 +1,40 @@
# CONTRACT — model-router-w3a-confirm
- date: 2026-10-10 | flow: feat | branch: feature/model-router-confirm (to start off develop)
- status: active
## REQUEST (verbatim — IMMUTABLE)
une fois fini j'aimerais que pour les premiere fois, les premiers switch de model et de'effort, on est un prompt qui demqnde de confirmer si on utilise bien ce model ou si moi j'en recommande un autre. Ca permet de voir si le routing correspond bien a nos besoin dans la pratique. sois faire un truc qui se souvient en userscop (si ca demande de confirmer la route sur un autre projet pour tel tache, que ca ne ele redemande jamais sur un autre projet. une memoire de ce qu'on decide, et ca met a jour la table de routing existante si il y a des changements, et si on fait un changement demander si c'est une exception pour ce projet ou non. Faire un truc qui demande au debut, mais qui est persistant une fois demander sur tout les projet et qui se redploi automatiquement comment tout le mod.
## CLARIFICATIONS
Q: where does the decision memory live? / A: "le 1 [fichier suivi dans le repo], mais il faut qu'il soit lisible et écrivable par tous les projets, donc le déployer (ln -s) dans le .claude du home à l'install, comme le reste" → tracked `mods/model-router/routing.json`, reached from every project through the EXISTING link `~/.claude/skills/model-router` → `../mods/model-router` (no new link: the plugin's own directory, `$.plugin`, resolves to it) [gated 2026-10-10]
Q: how is the question asked? / A: blocking dialog at the moment of the switch (`$.ui.ask`, the engine's AskUserQuestion), options Keep / Change / Later; never in headless; `/route ask off` cuts it [gated 2026-10-10]
Q: granularity? / A: per skill row, per agent row, per main-loop phase declared through the route tool; each once, across projects [gated 2026-10-10]
Q: public names (orchestrator default, user may veto): `/route pending`, `/route ask on|off` (`/route set` and `/route confirm` deferred after the simplicity lens); dialog texts in English like the rest of the mod; project exceptions live in `routing.json` under `projects[<normalized repo key>]` (r3: the project-tree file `<project>/.claude/model-router.json` was dropped after the robustness lens showed a cloned repo could re-route the user's gate agents and that writes would land in foreign trees) [stated 2026-10-10, user may veto]
Q: live T2 dialogs answered during the run (verifier → judge, feater → judge, Everywhere; the working-tree mod was hot-loaded by the engine without /reload-plugins) / A: user "Restaurer les deux" → routing.json reset to the shipped rows, confirmed/changed emptied; criterion 9 evidence: T2 dialog seen and written twice (texts and scope question confirmed live) [gated 2026-10-10]
Q: live decisions during the gates (verifier Keep, feater Keep, security-auditor → implement Everywhere) / A: user "Revenir à verify" for security-auditor (row reset, its confirmed/changed entries removed); the two Keeps stay. Criterion 9 evidence: T2 dialog seen 5 times live (Keep, Change + scope, dismissed/Later), the mod hot-loaded from the working tree by the engine [gated 2026-10-11]
Q: security gate BLOCK(1) credential fragment in the remote key / A: fixed (strict URL/scp parsing, credentials never read), `local:` path keys removed (no remote → Everywhere/Later only), output size cap; re-verify ECARTS(1) on guard fixtures → closed; re-scan PASS [gated 2026-10-11]
## ACCEPTANCE CRITERIA
1. `mods/model-router/routing.json` (tracked) is the single source of the phase table (11 full routes) and of the skill and agent rows (56 skill rows, 21 agent rows + Explore/Plan of plan r4 § Row tables), plus `confirmed` (skills/agents/phases → the confirmed phase) and `ask` (boolean); every row value is a phase key of that table; `DEFAULT_CONFIG.skills` and `.agents` in `register.ts` are `{}` (its `phases` stay as the fallback when the file is unreadable); rows and phases arrive from the file at load (session.start, or lazily once after a `/reload-plugins`), on `/route reload`, after each write, after `/clear` and on a cwd change.
CHECK: bash .claude/tasks/contracts/w3a-routing-file.sh
EXPECT: W3A-ROUTING-FILE
EVIDENCE: MET exit=0 marker-found :: W3A-ROUTING-FILE
2. Config layers, later wins: `routing.json` (phases, rows, then its `projects[<repo key>]` rows for the current repo) < `~/.claude/model-router.json` (machine override, its `ask: false` honored); a `.claude/model-router.json` inside the project tree is NEVER read; phases merge first (an invalid or missing routing.json phase falls back to the code default by name, logged), rows validate against the final table; a failed layer is skipped only at the first load, afterwards a failed read keeps the whole previous config; session toggles survive a rebuild; `/clear` keeps the config and re-reads at the next prompt. One kit test per clause.
3. First use asks once, one dialog: a typed skill with a row (main, allowed origin), a rowed agent spawn without an explicit `model` (`parentAgentId` undefined), a main-loop phase declared through the route tool → `$.ui.ask` whose question carries the REAL next model id and effort (decideFor/mainEffort on main, spawnTarget + explicit effort on spawn) and the options Later / Keep / <alt phase> / <alt phase> (T3: Later / Keep, only for a phase-key call); the file is re-read right before the dialog and the key re-checked as decided (another session's decision is seen); Keep writes `confirmed.<kind>.<name> = phase` and `confirmed.phases[phase]` and applies the row; an alt or a valid "Other" phase asks the scope (Everywhere → `routing.json` row + `changed.<kind>.<name> = {from, to}`; This project only → `routing.json` `projects[<normalized repo key>]` row, base row untouched, `confirmed` = the base row so no other project asks; one serialized write; a key failure offers Everywhere only), the new route applies to the current decision at once; any other answer (Later, dismissed, rejected in headless, unknown phase) → default applied, nothing written, not asked again this session; at most one dialog in flight: a concurrent use (same or another key) applies its current route unasked; a phase endorsed by a Keep is not asked at T3; a row that exists in `projects[key]` or in the machine override counts as decided; never asked inside a sub-agent (`parentAgentId` set), with an explicit `model` param, with the mod off, with `ask` false, or while routing.json is unreadable. One kit test per clause.
4. Writes happen only from a dialog answer or the composer `/route ask` command (never from the route tool, a prompt rule, or any model-originated event); only to `${$.plugin.root}/routing.json`, never into a project tree; the writer refuses when the file is absent or unparsable (toast, answer = Later) and never creates it; read-modify-write of the whole JSON (2-space, key order phases, skills, agents, projects, confirmed, changed, ask), serialized through one promise chain; a write or rebuild failure logs once, applies the default and leaves the key unasked; after a write the config is rebuilt (swapped only on a successful read) and a toast says "routing.json updated: commit it from the config repo (chore branch)". One kit test per clause + reading.
5. `/route pending` lists the rows and phases not yet confirmed (asked this session first, then the rest); `/route ask on|off` toggles asking and writes `ask`; `/route reload` re-reads the three layers. (`set`, `confirm`, a show suffix: deferred.) One kit test per command.
6. The kit suite and the mods suite are green; `claude plugin validate` passes.
CHECK: cd mods/model-router && out="$(claude plugin test . 2>&1)" && printf '%s\n' "$out" | grep -qE '[0-9]+ pass' && ! printf '%s\n' "$out" | grep -qE '[1-9][0-9]* fail' && claude plugin validate . 2>&1 | grep -q 'passed' && cd ../.. && make test suite=lib/tests/mods.test.sh 2>&1 | grep -q 'all suites green' && echo W3A-MOD-GREEN
EXPECT: W3A-MOD-GREEN
EVIDENCE: MET exit=0 marker-found :: W3A-MOD-GREEN
7. `lib/tests/effort-routing.test.sh` reads rows AND phases from `routing.json` (only the tier heads still come from `register.ts`), locks `DEFAULT_CONFIG.phases` equal to the file's phases, keeps the drift lock except for a row with a `changed.<kind>.<name>` entry whose `from` route equals the frontmatter and whose `to` equals the row (then a `WARN floor drift` line, no FAIL; a Keep-only drift, an empty or unknown frontmatter value, a hand edit after a change still FAIL; flip-tested), and is green; `lib/effort-shift.md` names the first-use dialog, `/route pending` and `/route ask` in ≤ 6 added lines.
CHECK: grep -q 'routing.json' lib/tests/effort-routing.test.sh && [ "$(grep -c 'register.ts' lib/tests/effort-routing.test.sh)" -le 3 ] && grep -q 'floor drift' lib/tests/effort-routing.test.sh && bash lib/tests/effort-routing.test.sh >/dev/null 2>&1 && grep -q '/route pending' lib/effort-shift.md && [ "$(wc -l < lib/effort-shift.md)" -le 66 ] && echo W3A-CENSUS-DOC
EXPECT: W3A-CENSUS-DOC
EVIDENCE: MET exit=0 marker-found :: W3A-CENSUS-DOC
8. Hook budget: no dialog or file write on the `turn.step` path; a `$.ui.ask` in flight never blocks a second, unrelated spawn of another agent name beyond the dialog itself; the ask is awaited outside `safely` in the owning hook with a `.catch` → Later. Judged by reading.
9. Live checks after the user's `/reload-plugins` (EVIDENCE lines added by the orchestrator; the answers committed on the branch before finish): one dialog at each of the three sites; a parallel same-agent dispatch answered after more than 10 s routes the unowned spawn without a timeout; an unanswered dialog on the terminal resolves to "Later" or to nothing written.
## FILE SCOPE
mods/model-router/routing.json (new), mods/model-router/hooks/register.ts, mods/model-router/hooks/register.test.ts, lib/tests/effort-routing.test.sh, lib/effort-shift.md; .claude/tasks/contracts/w3a-routing-file.sh (oracle, orchestrator)
@@ -0,0 +1,38 @@
# CONTRACT — model-router-w3b-dialog
- date: 2026-10-11 | flow: feat | branch: feature/model-router-confirm (continues W3-A before its merge)
- status: active
## REQUEST (verbatim — IMMUTABLE)
j'aimerais qu'on vois pour que quand je le prompt de demande de confirmation ou changement pour un model et un effort, il faudrait qu'il soit un peu plus complet. par exemple expliquer brievement ce que fait ce pour quoi on demqnde de choisir plutot que juste le nom. car juste verifier en vrai ca peut etre pleins de chose, ou dire feature : auditor. car c'est pas tres clair on sait pas pour quoi on prend la decision. Surtout que lam auditor security je croism ca proposait sonnet alors que nonm un audit securite ca devrait etre le plus performant du model. bref metre un peu de contextm dire oui ca me vam ou changer et si on choisis de changerm proposer quel model puis quel effort, puis pour userscope et projet ca c'est bien, et l'enregistrer.
## CLARIFICATIONS
(pass A: none — outcome, scope and flow given in the request; W3-A contract and plan r4 give the base)
Q: public shape (orchestrator default after three lenses, user may veto): ASK1 options Later / Keep / Change; ASK2 model = the four aliases labelled with their tier role (`fable (best)`, `opus (big)`, `sonnet (work)`, `haiku (cheap)`; Other = Later); ASK3 effort = the efforts offered by the phases headed by that model (fable: medium high xhigh max; opus: xhigh, skipped; sonnet: low medium high xhigh; haiku: low, skipped); the pair maps to that phase (rows stay phase names; a pair no phase offers is added by hand as a new phase in routing.json); ASK4 scope as W3-A; T3 (main-loop phase) = Later / Keep with context, no edit (BDR-116) [stated 2026-10-11]
Q: live decisions during this run (plan-challenger judge → verify Everywhere; orchestrate phase Keep = the T3 live check) / A: user "Revenir à judge" → row reset, its confirmed/changed entries removed; orchestrate Keep kept [gated 2026-10-11]
Q: live W3-B dialog answered during the security scan: security-auditor → Change → opus → xhigh → Everywhere = row `judge` (the request's own example; the first T2 answer through the new dialog) / A: kept; my reset of `changed` had wiped its from/to entry (census red at 604a6c4) → restored in the follow-up data commit [gated 2026-10-11]
Q: GATE 1 ran 3 verifiers (ECARTS(2) coverage → ECARTS(1) one non-discriminating test → CONFORME), security PASS (4 LOW) [recorded 2026-10-11]
## ACCEPTANCE CRITERIA
1. ASK1 text carries context: for a skill row "`/<name>` — <description ≤ 140 chars from the SKILL.md frontmatter, five YAML forms, first sentence>. Routed to <phase> (<about>): <model id> at <effort>. OK?"; for an agent row "`<name>` — <description from agent.offer ≤ 140 chars>. Routed to …"; for a phase "Phase `<name>` — <about>; used by <n> rows [+ its non-row consumers]: <model id> at <effort>. OK?"; descriptions and `about` pass `clean()` (Cc/Cf stripped, newlines folded, code-point truncation with "…"); a missing or unreadable description, an unset HOME or a name outside `^[A-Za-z0-9][A-Za-z0-9._:-]*$` degrades to the name alone (never blocks, never reads outside `${HOME}/.claude/skills/<name>/SKILL.md`, stat kind `file`, size-capped). Options: Later / Keep / Change (T3: Later / Keep); Other at ASK1 = Later + toast. One kit test per clause.
2. Change flow (rows only): ASK2 model = four fixed labels (Other or unknown = Later + toast); ASK3 effort = the efforts of the phases headed by the chosen alias, sorted ascending by level (≤ 4 labels, one → skipped, zero → Later + toast; Other = Later); target phase = the row's current phase when it matches, else the first matching phase in `cfg.phases` order; same phase as now = Keep (no `changed` entry); ASK4 as W3-A ([Everywhere, This project only], or [Everywhere, Later] with no repo key); any non-matching answer = Later, nothing written; storage as W3-A (`changed` from/to strings, `confirmed` = to); the new route applies to the current decision at once; after a main-row change the toast reports the real decision and the `/route switch on` hint when the pick is a downgrade. T3 never changes anything. One kit test per clause.
3. Rows stay phase names in every layer; `ALTS`, the alt options and the free-text phase branch of the W3-A dialog are removed (no dead code); the W3-A dialog tests are migrated to the new answers and every existing test stays green.
CHECK: ! grep -qE 'ALTS\b' mods/model-router/hooks/register.ts && echo W3B-NO-ALTS
EXPECT: W3B-NO-ALTS
EVIDENCE: MET exit=0 marker-found :: W3B-NO-ALTS
4. `routing.json` phases gain an `about` string (≤ 120 chars) for the 11 phases, loaded into a separate map (never on `Route`, never in `DEFAULT_CONFIG`), a non-string or overlong value dropped + logged once; descriptions never read from the project tree.
CHECK: bash .claude/tasks/contracts/w3b-about.sh
EXPECT: W3B-ABOUT
EVIDENCE: MET exit=0 marker-found :: W3B-ABOUT
5. Census unchanged and green (rows are phases; the DEFAULT == file phase lock keeps matching with `about` present); `lib/effort-shift.md` dialog lines updated (≤ 4 lines changed, Later/Keep/Change and the model-then-effort flow named).
CHECK: bash lib/tests/effort-routing.test.sh >/dev/null 2>&1 && grep -q 'Change' lib/effort-shift.md && [ "$(wc -l < lib/effort-shift.md)" -le 66 ] && echo W3B-CENSUS
EXPECT: W3B-CENSUS
EVIDENCE: MET exit=0 marker-found :: W3B-CENSUS
6. Kit suite, mods suite and validate green.
CHECK: cd mods/model-router && out="$(claude plugin test . 2>&1)" && printf '%s\n' "$out" | grep -qE '[0-9]+ pass' && ! printf '%s\n' "$out" | grep -qE '[1-9][0-9]* fail' && claude plugin validate . 2>&1 | grep -q 'passed' && cd ../.. && make test suite=lib/tests/mods.test.sh 2>&1 | grep -q 'all suites green' && echo W3B-MOD-GREEN
EXPECT: W3B-MOD-GREEN
EVIDENCE: MET exit=0 marker-found :: W3B-MOD-GREEN
7. Security posture unchanged: writers, paths, caps, no model-originated write; descriptions truncated and control characters stripped before they enter a question. Judged by reading + one kit test (a description with control characters and 500 chars → ≤ 140 clean chars).
## FILE SCOPE
mods/model-router/routing.json, mods/model-router/hooks/register.ts, mods/model-router/hooks/register.test.ts, lib/tests/effort-routing.test.sh, lib/effort-shift.md; .claude/tasks/contracts/w3b-about.sh (oracle, orchestrator)
@@ -0,0 +1,32 @@
# CONTRACT — model-router-w3c-forget
- date: 2026-10-11 | flow: feat | branch: feature/model-router-confirm (continues W3-A/B before their merge)
- status: active
## REQUEST (verbatim — IMMUTABLE)
d'ailleurs est-ce quil y a une comande pour clear les decision prise ? si jamais cest un user aure que moi si il peut clean me decisions de routage ?
[orchestrator proposal: `/route forget <nom|all|projects>` + an optional per-user decisions file] → "seulement le forget"
## CLARIFICATIONS
Q: scope / A: `/route forget` only; no per-user decisions file (the tracked routing.json stays the shared memory; another user inherits and may forget) [gated 2026-10-11]
Q: public shape (orchestrator default, user may veto): `/route forget <name>` (a skill row, an agent row or a phase), `/route forget all`, `/route forget projects`; forgetting a changed row RESTORES it to its shipped phase (`changed.from`); the frontmatter floors are never touched (the answer says so when a changed row is restored) [stated 2026-10-11]
Q: GATE 1 cap (verifier ECARTS(2) incl. one real defect `ForgetPlan` → `Plan`, then ECARTS(2)/(2) coverage only; 284 kit tests) / A: user "Accepter et commiter"; security PASS (1 MEDIUM + 4 LOW parked in TODO); live T3 Keeps on orchestrate and escalate seen during the run [gated 2026-10-11]
## ACCEPTANCE CRITERIA
1. `/route forget <name>` (composer only, exactly one argument): for every kind (skills, agents, phases) removes `confirmed.<kind>.<name>`; restores a `changed` row to its `from` only when `from` is a string, the row exists, equals `changed.to` and `from` is a known phase (else the entry is kept and the answer says why); removes every `projects[*].<kind>.<name>` (emptied tables pruned); drops the key(s) from this session's asked set; nothing removed and nothing asked → "nothing to forget for <name>" with NO write; key only asked this session → reset, no write; refused while a dialog is open; two words → usage. One kit test per clause (folded where one setup proves two).
2. `/route forget all` and `/route forget projects` ASK FIRST (engine dialog, options Cancel / Forget, Cancel first; anything else, dismissed or headless = cancelled, nothing written; skipped when there is nothing to forget; the forget holds the single-dialog slot while asking): `all` empties `confirmed`, `changed` (rows restored under the same guards) and `projects` and clears the asked set; `projects` empties only `projects`. One kit test per clause.
3. Writes go through the existing writer only (`writeRouting`: serialized, size-capped, refused when the file is absent or unparsable with the existing toast, config rebuilt after the write); the model (route tool, prompt rules, sub-agents) can never trigger a forget; a non-composer origin is refused like the other `/route` writes. Judged by reading + one kit test (route tool cannot forget) + the existing origin test extended.
4. One answer formatter: "forgot <label>: <c> confirmed, <m> row(s) restored, <k> project exception(s) removed"; a restored row adds "(<kind>.<name> → <from>: <alias> at <effort>; if <agents/<name>.md | skills/<name>/SKILL.md> was aligned to <to>, set model: <alias>, effort: <effort> and its lock in lib/tests/model-routing.test.sh, then `make test`)" (no file clause for a built-in agent: Explore, Plan); "; ~/.claude/model-router.json still sets <name> and wins here" when the machine override pins it; "; current run may keep <phase>: /route clear to apply" when the run slot came from a restored skill row [wording per plan r3, gated 2026-10-11]; write outcomes: refused → "nothing saved", rebuild failure after a landed write → "saved, config not rebuilt: /route reload". Write outcomes: refused → "nothing saved", rebuild failure after a landed write (dedicated error) → "saved, config not rebuilt: /route reload", any other failure → "forget not applied". One kit test per text.
5. `lib/effort-shift.md` names `/route forget` in ≤ 2 changed lines; README.md, USAGE.md and CHANGELOG.md no longer claim that only a dialog answer or `/route ask` writes routing.json (one clause each naming `/route forget`); the register.ts writer comment says the same; `/route pending` lists a forgotten key again.
CHECK: grep -q 'forget <name|all|projects>' mods/model-router/hooks/register.ts && grep -q '/route forget' lib/effort-shift.md && [ "$(wc -l < lib/effort-shift.md)" -le 66 ] && grep -q 'route forget' README.md && grep -q 'route forget' USAGE.md && grep -q 'route forget' CHANGELOG.md && ! grep -q 'Only a dialog answer or `/route ask on|off` writes' README.md && echo W3C-DOC
EXPECT: W3C-DOC
EVIDENCE: MET exit=0 marker-found :: W3C-DOC
6. Kit suite, mods suite and validate green; the census stays green on the shipped file.
CHECK: cd mods/model-router && out="$(claude plugin test . 2>&1)" && printf '%s\n' "$out" | grep -qE '[0-9]+ pass' && ! printf '%s\n' "$out" | grep -qE '[1-9][0-9]* fail' && claude plugin validate . 2>&1 | grep -q 'passed' && cd ../.. && make test suite=lib/tests/mods.test.sh 2>&1 | grep -q 'all suites green' && bash lib/tests/effort-routing.test.sh >/dev/null 2>&1 && echo W3C-GREEN
EXPECT: W3C-GREEN
EVIDENCE: MET exit=0 marker-found :: W3C-GREEN
7. Live smoke (orchestrator + user, after the gates, recorded `[gated]`): one `/route forget` typed in the terminal opens its dialog (`$.ui.ask` from an immediate command) and the answer lands in routing.json.
## FILE SCOPE
mods/model-router/hooks/register.ts, mods/model-router/hooks/register.test.ts, lib/effort-shift.md, README.md (one clause), USAGE.md (one clause), CHANGELOG.md (one clause under Unreleased)
+16
View File
@@ -0,0 +1,16 @@
#!/usr/bin/env bash
# GATE 0 oracle, contract 2026-10-10-model-router-w3a-confirm: routing.json
# is the single source of phases + rows; register.ts rows are empty.
set -u
python3 - <<'PY'
import json,re,sys
r=json.load(open('mods/model-router/routing.json'))
src=open('mods/model-router/hooks/register.ts').read()
ph=set(re.findall(r"^\s+([a-z]+): \{ tier:", re.search(r"phases: \{(.*?)\n \}", src, re.S).group(1), re.M))
jp=set(r.get('phases',{}).keys())
ok = jp==ph and len(r['skills'])==56 and len(r['agents'])==23
ok = ok and all(v in jp for v in list(r['skills'].values())+list(r['agents'].values()))
ok = ok and isinstance(r.get('confirmed'),dict) and isinstance(r.get('ask'),bool) and 'version' not in r
ok = ok and bool(re.search(r"\n skills: \{\},", src)) and bool(re.search(r"\n agents: \{\},", src))
print('W3A-ROUTING-FILE' if ok else 'W3A-ROUTING-BAD'); sys.exit(0 if ok else 1)
PY
+9
View File
@@ -0,0 +1,9 @@
#!/usr/bin/env bash
# GATE 0 oracle, contract w3b: every routing.json phase carries a 1-120 char `about`.
set -u
python3 -I - <<'PY'
import json,sys
r=json.load(open('mods/model-router/routing.json'))
ok=all(isinstance(v.get('about'),str) and 0<len(v['about'])<=120 for v in r['phases'].values())
print('W3B-ABOUT' if ok else 'W3B-ABOUT-MISSING'); sys.exit(0 if ok else 1)
PY
@@ -0,0 +1,288 @@
# PLAN r4 — model-router wave 3-A: first-use route confirmation + decision memory + project exceptions (2026-10-10)
r1 → r2 after simplicity CONCERNS(3) + correctness FATAL(8); r2 → r3 after
robustness CONCERNS(8) (fresh dispatch on r2 after the first died on a
network error); r3 → r4 after the confirmation pass (correctness
CONCERNS(6), no BLOCKER). Precedence: r4 > r3 > r2 > r1 where they differ. Contract
`.claude/tasks/contracts/2026-10-10-model-router-w3a-confirm-1201.md`.
Facts (kit types, checked 2026-10-10): `$.ui.ask(question, {options,
header})` opens the engine's AskUserQuestion dialog, resolves to the chosen
label or the "Other" free text, REJECTS when dismissed and in a `-p` run;
`$.fs.read/stat/exists/write` take absolute paths, write creates
directories; `$.session.root()` = the project root (follows /cd);
`$.plugin.root` = the plugin directory, absolute (routing.json =
`${$.plugin.root}/routing.json`, reached through the existing
`~/.claude/skills/model-router` link); `$.ui.toast`; `$` calls stop the
hook budget clock. The kit has no fs: tests mock bottom `fs.*`, `env.get`,
`session.root` and `tool.call AskUserQuestion` hooks.
## Data (one tracked source for rows AND phases; r3)
- `mods/model-router/routing.json` (tracked):
`{ "phases": {11 full routes}, "skills": {56}, "agents": {23},
"projects": { "<repo key>": { "skills": {}, "agents": {} } },
"confirmed": { "skills": {}, "agents": {}, "phases": {} },
"changed": { "skills": {}, "agents": {} }, "ask": true }`
`confirmed.<kind>.<name>` = the routing.json phase endorsed by a Keep;
`changed.<kind>.<name>` = `{ "from": <shipped phase>, "to": <phase> }` for
an "Everywhere" change (the census reads `from`). `projects[key]` = the
"This project only" exceptions. Key (r4) = the origin remote normalized:
scheme and userinfo stripped, host lowercased, `.git` and trailing slash
removed (`github.com/acme/app`); no remote → `local:<root realpath>`
(meaningful on that machine only, stated in the toast); `session.repo`
or the realpath failing → the `projects` layer is skipped and ASK2
offers Everywhere only. Resolved once per decision and re-resolved on
`classic.CwdChanged` (rebuild). No version key, no dates. Never a URL
with credentials in the file.
- NEVER read from the project tree: `<root>/.claude/model-router.json` is
not a layer (a cloned repo must not route the user's agents). Layers,
later wins: routing.json (phases, rows, then `projects[key]` rows) <
`~/.claude/model-router.json` (machine override; its `ask: false` also
honored). Phases merge first, rows validate against the final table.
- Phases: routing.json `phases` replaces `DEFAULT_CONFIG.phases` route by
route; an invalid or missing phase falls back to `DEFAULT_CONFIG.phases
[name]` with one log line (a typo never removes a phase); alts offered
by the dialog are filtered against the final table.
- Loading: full read at session.start, and LAZILY once when the State was
never loaded (first prompt.submit / skill.prompt / agent.spawn / route
tool after a `/reload-plugins`, which re-runs register() without
session.start); a failed layer is skipped and logged at that first load
only; afterwards any failed read, and a routing.json that went missing,
keeps the WHOLE previous config (today's kill-switch rule), and no ask
runs while routing.json is unreadable. `/route show` prints `config:
routing.json` once loaded. `resetSession` keeps its current set (cfg
included); the layers are re-read at the next `prompt.submit` and the
cfg swapped only after a fully successful read. Session toggles
(`/route switch`, `verbose`) live in State, re-applied after any rebuild,
and added to resetSession's kept set (they survive /clear as today).
- `DEFAULT_CONFIG.skills/.agents` = `{}`; the file is the source. A
missing routing.json → DEFAULT phases, no rows, one log, no asks, no
writes (the writer never creates the file).
## Dialog (main loop only; `ask` true; mod on; never inside an agent)
- Trigger sites and guards:
T1 typed skill with a row: `applyTypedSkill` made async, awaited in the
skill.prompt hook OUTSIDE `safely`: route FIRST (`routeMainBySkill`,
typed=true), build the text from the route now in force, ask, and on
a change re-run `routeMainBySkill` with the recomputed `skillRow`
(the model's `Skill` path never asks).
T2 rowed agent spawn: in `registerSpawn` before `spawnTarget`, awaited;
guards `e.parentAgentId === undefined`, not frozen, not shadowed, and
`e.model === undefined` (an explicit model is not a routing decision).
T3 route tool phase on main (`handleRouteTool`, no agentId, and only
when `e.phase` is a phase key — never for an effort-only call):
confirm-only.
- One dialog (ASK1), 4 options: question text built from the REAL decision:
main → `decideFor`/`mainEffort` like `mainRoutedText` ("First route for
/feat: reflect, next step claude-fable-5-1 at high. Keep it?"); spawn →
`spawnTarget` + the explicit `effort` param when given ("First dispatch
of feater: implement, claude-sonnet-5-5 at medium. Keep it?"); T3 →
"First use of orchestrate on the main loop: claude-fable-5-1 at medium.
Keep it?" with options ["Later", "Keep"] only.
Options T1/T2: ["Later", "Keep", <altA>, <altB>] (Later FIRST: an idle
auto-pick, if any surface does one, must never write), header
"model-router";
alts = the first two of [plan, reflect, implement, apply] (skills) or
[judge, implement, verify, apply] (agents) minus the current phase;
"Other" free text = a phase key of the final table (else toast "unknown
phase <x>, default kept" → Later); an Other equal to the current phase
= Keep. Any answer that is not exactly Keep, an alt, or a valid Other →
Later (covers dismissed, rejected, auto-resolved idle answers). ASK2
accepts only its two labels; anything else = Later.
Before opening any dialog the mod re-reads routing.json (the pre-ask
read) and re-checks `ask` AND "decided" for the key from that fresh
content, so a decision taken in another live session is seen.
- Keep → `confirmed.<kind>.<name> = <routing.json phase>` AND
`confirmed.phases[<phase>] = <phase>` (a Keep endorses the phase: no
duplicate T3 later); the row applies. "Decided" = `confirmed` equals the
routing.json row, OR a `projects[key]` / machine-override row exists for
that name (presence = decided; never compared across layers).
- Change (alt/Other) → ASK2 scope ["Everywhere", "This project only"].
Everywhere → routing.json row + `changed.<kind>.<name> = {from, to}`
(`from` kept from the FIRST entry when one exists, `to` = the new row) +
`confirmed.<kind>.<name> = to`; This project only → `projects[key]
.<kind>.<name> = phase` (routing.json row untouched) + `confirmed.<kind>
.<name> = the BASE routing.json row` (so no other project asks again). One file, one serialized write.
Then rebuild; the new route applies to the current decision at once
(T1: recompute `skillRow`; T2: recompute `spawnRoute` and pass it to the
spawn bookkeeping). A write or rebuild failure → log once, default row
applied, key left unasked (never escapes to the hook's `.catch`).
- Later → default applies, nothing written.
- One dialog in flight per session, ever (`st.asking: string | null`):
a use that finds a dialog open (same key or another) applies its current
route and stays unasked (asked later); the owner alone runs ASK1, ASK2
and the write. A parallel spawn never awaits a promise it did not create
(hook budget: only the owner's `$` call stops the clock). Asked-this-
session = `Set<key>` in State, dropped by resetSession.
- Never on `turn.step`. The ask is awaited in the owning hook with
`.catch(() => 'Later')`.
## Commands (composer only, as `/route` today)
- `/route pending` → keys not yet confirmed (asked-this-session first).
- `/route ask on|off` → `ask` in routing.json (write path below); the
flag is re-read (one `$.fs.read`) right before each ask so another
session's change is seen.
- `/route reload` → re-reads the three layers. (`set`, `confirm`, a show
suffix: deferred to TODO; the dialog is the writer.)
## Writes
- `writeRouting($, st, patch)`: read-modify-write of the whole JSON
(2-space; key order phases, skills, agents, projects, confirmed, changed,
ask), serialized through ONE promise chain in State; refuses (toast,
answer treated as Later) when routing.json is absent or unparsable:
never creates the file; size cap as the override, `readCapped` takes
the file label; after a write: rebuild (swap only on a successful read)
+ toast "routing.json updated: commit it from the config repo (chore
branch, `gitflow.sh start chore …`)".
- Writers: dialog answers and composer `/route ask` ONLY. The route tool
(model), prompt rules and sub-agents never write. Nothing is ever
written into a project tree.
## Census and floors (BLOCKER closed)
- `lib/tests/effort-routing.test.sh`: rows AND phases read from
routing.json (python3 json); tier heads still parsed from `register.ts`
`tiers` (unchanged by this wave; the only register.ts read left); a new
lock: `DEFAULT_CONFIG.phases` values == routing.json phases (the fallback
never applies a stale route). The drift lock stays for every row EXCEPT
one with a `changed.<kind>.<name>` entry whose `from` phase route equals
the frontmatter AND whose `to` equals the current row: then one `WARN
floor drift: <file> <value> vs row <to> (<value>)` line, no FAIL. A
Keep-only row, an empty or unknown frontmatter value, a hand edit after
a change (`to` ≠ row), or a drift not matching `from` still FAILs.
- BDR-115 amendment 2(a) "census-locked equal" → amended to "equal unless
the row is a confirmed user decision (floor may lag; WARN)" at STEP 7.
## Steps
- [ ] S1 routing.json from the current DEFAULT_CONFIG (phases + rows +
Explore/Plan), `projects`/`confirmed`/`changed` empty, `ask` true;
DEFAULT rows → `{}`; `loadLayers` (two files + `projects[key]`),
first-load degrade then keep-previous rule, phase fallback by name,
session toggles in State, resetSession kept set unchanged + re-read
at the next prompt.submit.
- [ ] S2 ask engine: `askFirst($, st, kind, name, text, options)` with the
single-dialog guard + asked set; `decide(answer)`; `applyDecision`
(Keep / change + scope) → serialized write → rebuild → recompute,
wrapped so a failure logs once and applies the default.
- [ ] S3 wiring T1 (async applyTypedSkill), T2 (registerSpawn), T3
(handleRouteTool, confirm-only).
- [ ] S4 commands `pending`, `ask on|off`; `reload` reads three layers.
- [ ] S5 census (rows + phases from JSON, drift lock with the decision
exception) + `lib/effort-shift.md` ≤ 6 added lines (dialog, `/route
pending`, `/route ask`).
- [ ] S6 kit tests. First: `boot()`/`bootRun()` gain a path-aware fixture
(bottom `fs.exists`/`fs.stat` (+ `realPath`)/`fs.read`/`fs.write`,
`env.get` HOME, `session.root`, `session.repo`) serving an inline
routing.json with `ask: false`
unless a test opts in (`boot($, on, {ask: true, project: {...},
home: {...}})`); the existing `{verifier: null}` test becomes
path-aware; all 86 tests green again. Then one test per clause:
layer precedence (routing.json rows < projects[key] < machine
override, null drop; a `.claude/model-router.json` in the project
tree is NEVER read; machine `ask:false` honored); typed `/feat` first use → AskUserQuestion mock
sees the id + "high" → "Keep" → fs.write of routing.json captured
with confirmed.skills.feat = reflect AND confirmed.phases.reflect →
second typed `/feat` → no ask; alt "implement" → ASK2 → "Everywhere"
→ routing.json row + main routed implement now; "This project only"
→ project file written, routing.json row untouched, confirmed in
routing.json; "Later"/reject/free text garbage → default, no write,
no second ask this session; agent first spawn → ask → Keep → model
written; explicit `model` param → no ask; two parallel spawns → one
ask, the second spawn routed on the current row without waiting; parentAgentId set → no ask; route tool phase first use → ask
(Keep/Later) → Keep → confirmed.phases; phase already endorsed by a
T1 Keep → no T3 ask; `/route ask off` → no asks + write; `/route
pending` text; routing.json unreadable at start → DEFAULT phases, empty
rows, one log, no asks; unreadable at a later rebuild → previous cfg
kept; missing file → `/route ask off` refused with a toast, file not
created; a phase typo in routing.json → DEFAULT phase by name + log,
alts filtered; `/route switch on` survives a rebuild; writes are
serialized (two decisions, one file, both present); row edited by
hand after a Keep → asked again; census flip-tests: confirmed-only
drift FAILs, `changed.from` drift WARNs, empty `model:` FAILs.
- [ ] S7 live checks after the user's /reload-plugins (EVIDENCE lines,
answers committed on the branch before finish): `/route show`
prints `config: routing.json` right after the reload (lazy load);
one dialog at each of the three sites; a parallel same-agent dispatch answered after
more than 10 s (the unowned spawn must be routed, not timed out);
what an unanswered dialog resolves to on the terminal (and on sdk/
bridge if reachable): any auto-pick must land on "Later".
- Disposition: honors BDR-115 (user writers only; full ids in texts;
amendment 2(a) to be amended), BDR-107/108 (phases unchanged), LRN-206
(kit facts), LRN-210 (one clause per test), LRN-211 (typed path only).
- Deferred (TODO): `/route set`, `/route confirm`, show suffix, dialog
edits of phases, frontmatter auto-alignment, a session-start warning
when routing.json is dirty, per-machine `projects` (today one tracked
map keyed by repo).
## Challenge ledger (r1 → r2)
- correctness 1 BLOCKER + simplicity 3 (census red on a decision) → drift
lock with the confirmed-decision exception (WARN), BDR-115 2(a) amended.
- correctness 2/5, simplicity 2 (partial phase overrides, blast radius) →
phases move to routing.json as a full table, dialog never edits phases,
T3 confirm-only, project layer rows only.
- correctness 3, simplicity 10 (kit loses rows) → S6 path-aware fixture first.
- correctness 4 (texts from the real decision) → ASK1 texts from
decideFor/mainEffort and spawnTarget; T2 skipped on explicit model.
- correctness 6, simplicity 9 (sync site) → async applyTypedSkill awaited
in the hook.
- correctness 7 (root == HOME) → layer 3 skipped, ASK2 skipped.
- correctness 8/9/10/16/17 (layers, /clear, validation order, stale root,
concurrent writes) → re-read from disk on every rebuild, per-layer
degrade, phases first, root re-resolved, one write chain.
- correctness 11/12 (Other, counts), simplicity 4/7 → one 4-option dialog,
2 alts, non-matching answers = Later.
- correctness 13 (confirmed location) → always routing.json.
- correctness 14/15 (parentAgentId, fs.stat label) → written in.
- correctness 18 (no live proof) → S7.
- simplicity 1 (commands) → pending + ask only; set/confirm deferred.
- simplicity 5/6/8 → confirmed = phase, no version, one Map, Keep endorses
the phase.
- simplicity 11 → project layer rows only.
## Challenge ledger (r2 → r3, robustness)
- rob 1/9 (shared promise burns the awaiters' budget; stacked dialogs) →
one dialog in flight, owner-only; concurrent uses apply the current
route unasked; S7 >10 s check.
- rob 2 (per-layer degrade wipes rows, reopens the kill switch) → degrade
at first load only, then keep-previous; no asks while unreadable.
- rob 3 (a phase typo removes the phase, gate STOPs everywhere) → fallback
to DEFAULT phase by name + log; alts filtered.
- rob 4 (writer creates a missing file) → writer refuses, never creates.
- rob 5 (global confirmed vs layered rows re-asks forever) → confirmed
compared with the routing.json row only; layer rows = decided by
presence; one file, one write.
- rob 6 (project-tree layer = security hole + foreign writes) → layer
removed; exceptions in routing.json `projects[key]`; nothing written in
a project tree.
- rob 7 (drift exception too wide) → exemption only for `changed.from`
matches; Keep-only, empty and unknown values FAIL.
- rob 8 (/clear drops cfg/kill switch/breaker) → kept set unchanged;
re-read at next prompt.submit, swap on success.
- rob 10 (session toggles reverted by rebuilds) → toggles in State.
- rob 11 (dirty tracked repo, S7 answers) → toast names the chore flow;
S7 answers committed before finish; session-start warning deferred.
- rob 12 (idle auto-pick) → "Later" first; S7 verifies.
- rob 13 (`ask` only in the tracked file) → machine override `ask`
honored; flag re-read before each ask.
- rob 14 (write failure escapes) → applyDecision wrapped.
## Confirmation ledger (r3 → r4, correctness)
- conf 1 (tier heads) → census keeps a tiers parser on register.ts; AC7
CHECK narrowed to the rows/phases parser.
- conf 2 (rows lost after /reload-plugins) → lazy load once; S7 line.
- conf 3 (T1 text before routing) → route first, text, ask, re-route.
- conf 4 (exception re-asks elsewhere) → confirmed = base row on "This
project only"; test X then Y.
- conf 5 (multi-session re-ask) → pre-ask read re-checks decided.
- conf 6 (raw remote URL / machine path as key) → normalized key, no
userinfo, `local:` fallback, key failure skips projects only.
- conf 7 (leftovers, un-gated move of the exception file) → S6/contract
fixed; named to the user at the gate.
- conf 8 (fixture) → session.repo + realPath mocked; key failure scoped.
- conf 9 (/cd) → rebuild on `classic.CwdChanged`.
- conf 10 (changed.from on a second change) → first `from` kept, exempt
only when row == to.
- conf 11 (unspecified answers) → ASK2 two labels only; Other == current
= Keep.
- conf 12 (DEFAULT phases drift) → census lock DEFAULT == routing.json.
- conf 13 (effort-only route call) → T3 only for a phase key.
- conf 14 (later absence) → treated as failed, previous kept.
- conf 15 (toggles across /clear) → added to the kept set.
@@ -0,0 +1,170 @@
# PLAN r3 — model-router wave 3-B: a dialog with context, model then effort on change (2026-10-11)
r1 → r2 after simplicity CONCERNS(3), robustness FATAL(5), correctness
FATAL(8): all three rejected inline-route rows and dialog edits of phases.
r2 → r3 after the confirmation (correctness CONCERNS(2)). Precedence r3 > r2 > r1. Contract `.claude/tasks/contracts/2026-10-11-model-router-w3b-dialog-1240.md`.
Base = W3-A (22455c0, plan r4, BDR-116). Facts: `agent.offer` payload carries
`description` and `source`; `$.ui.ask` takes 2-4 labels + Other; `cfg.models`
maps the four aliases to ids; each default alias heads one tier (best fable,
big opus, work sonnet, cheap haiku); `readCapped` exists (size cap, stat).
## Decisions
- D1 rows stay PHASE NAMES everywhere (routing.json, projects, override).
No inline route rows. The dialog's model+effort choice maps to a phase.
- D2 the dialog never edits a phase: T3 offers Later / Keep only, with
context. Phases change by hand in routing.json + DEFAULT_CONFIG (census
lock kept). BDR-116 unchanged.
- D3 one description rule for every kind, hardened read, no cache.
## Texts (ASK1)
- `clean(s, n)`: FIRST fold every `\s` and `\p{Z}` (incl. U+2028/2029) to
a space, THEN strip `\p{Cc}` and `\p{Cf}`, collapse spaces, trim; when
longer than n code points keep n-1 and append "…" (total ≤ n). Applied to
every description and `about` (no answer echo anywhere).
- Leads are KEPT as the first words (existing assertions stay valid):
"First route for /feat — <desc>. Routed to reflect (<about>): claude-fable-5-1
at high. OK?", "First dispatch of security-auditor — <desc>. Routed to
verify (<about>): claude-sonnet-5-5 at xhigh. OK?", "First use of
orchestrate on the main loop — <about>; used by …: … OK?". Without a
description the " — <desc>" part is omitted; without `about` the
parenthesis is omitted.
- Skill row desc = `skillDesc($, st, name)`: name must match
`^[A-Za-z0-9][A-Za-z0-9._:-]*$` (else name only); `readCapped` of
`${HOME}/.claude/skills/<name>/SKILL.md` (stat kind must be `file`, size
cap as the override; HOME from `$.env.get`, unset → name only); parse the
frontmatter `description:` in its five forms (plain, `|`, `>`/`>-`,
single-quoted with `''`, double-quoted): unquote, join the block lines
with spaces, cut at the first sentence end (`. ` / `.` at end), then
clean(…, 140); any failure → name only. Read at ask time, once (the key
is asked at most once per session). The kind check lives in a
skill-specific wrapper around `readCapped` (config reads untouched).
- Agent row: the offer description goes through the SAME rule (sentence
cut + clean 140) WHEN STORED at `agent.offer` (bounded cache
`Map<type, {source, description}>`); absent → name only.
- Phase (T3): "First use of orchestrate on the main loop — <about>; used by
<n> rows[, the derived route of dispatches][, <k> prompt rule(s)]:
claude-fable-5-1 at medium. OK?"; consumers derived: `orchestrate` →
"the derived route of dispatches" (hardcoded, pushOrchestrate), prompt
rules counted from `cfg.prompt.filter(r => r.phase === name)` (floor
rules named "floor rule"). Options ["Later", "Keep"].
- Row options: ["Later", "Keep", "Change"]. At EVERY step the `askOr` LATER
(dismissed, headless) is checked first and is silent; any other
non-label answer = Later + toast "answer not recognised, default kept"
(no echo). `ALTS`, the alt slice in `optionsFor` and the free-text phase
branch of `decide` are removed.
- Main-row texts keep the real-decision rule (decideFor/mainEffort); spawn
texts keep spawnTarget + explicit effort.
## Change flow (rows only)
- ASK2 "Which model for `<name>`?": four FIXED labels from the default
aliases, each with its tier role ("fable (best)", "opus (big)", "sonnet
(work)", "haiku (cheap)"; role = the tier headed by the alias in
`cfg.tiers`, "(custom)" when none); Other or an unknown label = Later +
toast.
- ASK3 "Which effort on <alias>?": the efforts of the phases (in
`cfg.phases`) whose tier head is the chosen alias, deduplicated and
sorted ASCENDING by LEVELS (fable: medium high xhigh max; opus: xhigh;
sonnet: low medium high xhigh; haiku: low) → at most 4 labels (the
first 4); exactly one → ASK3 skipped; zero → Later + toast "no phase
uses <alias>"; Other = Later.
- Target phase = the row's CURRENT phase when it matches (alias, effort),
else the first matching phase in `cfg.phases` order. Same phase as now →
treated as Keep (no `changed` entry).
- ASK4 scope as W3-A: ["Everywhere", "This project only"], or ["Everywhere",
"Later"] when there is no repo key.
- After a CHANGE on a MAIN row (T1, never on Keep), computed in
`confirmSkill` from `skillRow` + `decideFor`: toast "`/<name>` now runs
<id> at <effort>" plus ", main moves up only: `/route switch on` allows
a downgrade" when the chosen alias ranks below the session model and the
switch is off; the new route is re-picked into the current decision
(T1/T2 as W3-A).
- Storage unchanged from W3-A: `changed.<kind>.<name> = {from (first), to}`,
`confirmed.<kind>.<name> = to` (strings), `projects[key]` rows.
## `about`
- routing.json: `phases.<name>.about` (≤ 120 chars, clean) on the 11
phases; loaded into `mem.about: Record<phase, string>` (part of the
memory rebuilt with the config, so it survives /clear with cfg; never on
`Route`, never in DEFAULT_CONFIG; the census `code_phase` regex
untouched); a non-string or overlong `about` is dropped + logged once;
a phase without `about` → no parenthesis. acceptPhase ignores the key.
## Steps
- [ ] S1 routing.json `about` ×11 (one line each from the W2 phase table);
loader + map + validation.
- [ ] S2 descriptions: offer cache `{source, description}`; `skillDesc`
(allowlist, readCapped, kind check, YAML forms, sentence cut); `clean`.
- [ ] S3 ASK1 texts + options Later/Keep/Change; Other → Later + toast;
dead code removed (`ALTS`, alt slice, free-text branch).
- [ ] S4 change flow ASK2 → ASK3 (derived efforts) → ASK4; target-phase
rule; same-phase = Keep; main toast with the real decision.
- [ ] S5 `lib/effort-shift.md` dialog lines (≤ 4 changed).
- [ ] S6a migrate the W3-A dialog tests: every `asks([...])` answering a
phase name at ASK1 becomes Keep or a Change chain (ASK1 → ASK2 →
[ASK3] → ASK4; the `asked[n]` indexes shift accordingly); option-list
assertions become Later/Keep/Change (T3 Later/Keep); the lead-text
assertions (:2266 "First dispatch of feater", :2395 "First use of
orchestrate on the main loop") stay valid by construction; :1914
expects the new toast "answer not recognised, default kept"; the two
free-text tests (:1871 "a free-text phase counts as an alt", :1878
"equal to the current phase is a Keep") are DELETED (authorized:
they test the removed branch); "a phase name answer is a Later" kept.
World/helper extensions: `offerOf` takes a description; the World
gains a per-path stat `kinds` override and a HOME override (unset).
- [ ] S6b new tests, one per clause: skill text with each YAML form (plain,
`|`, `>-`, 'sq' with `''`, "dq"), missing SKILL.md → name only, a bad
name (`../x`) → name only, stat kind ≠ file → name only, a SKILL.md
over the cap → name only, HOME unset → name only, a 500-char
description with control, bidi and U+2028 chars → ≤ 140 clean code
points ending in "…"; a Keep on a main row → no "now runs" toast;
a Change → the toast with the switch hint; zero-effort alias → Later
+ toast; dismissed ASK2 → silent Later; agent text from the offer
description; T3 text with `about` and consumers; ASK2 labels; ASK3
derived efforts per alias (fable 4, opus → skipped, sonnet 4, haiku →
skipped); security-auditor Change → opus → (skipped) → Everywhere →
row `judge`, spawn on claude-opus-5-5 at xhigh now; feater Change →
sonnet → high → `write`; Explore Change → sonnet → medium → stays
`explore` (current phase wins) → treated as Keep, no `changed`;
`/feat` Change → sonnet → medium → `implement`, toast names "moves up
only" with the switch off; Other at ASK1/ASK2/ASK3 → Later + toast,
nothing written; project scope with the derived phase; `about`
overlong → dropped + logged.
- Disposition: honors BDR-116 rules (1), (3)-(6) unchanged (rows = phases,
dialog never edits phases, writers/paths/caps as is); AMENDS its clause
(2): options become Later/Keep/Change with a model-then-effort chain,
Other = Later (successor amendment written at CAPITALIZE, append-only); BDR-115 (full ids from `cfg.models`),
LRN-210 (one clause per test), LRN-212 (live mod: announce dialogs).
- Docs (README/USAGE/CHANGELOG lines on the dialog) drift after this wave:
STEP 6 doc-sync, out of this contract's scope.
## Challenge ledger (r1 → r2)
- simp 1/2, rob 2/3/4/7/11/12, corr 1/2/3/9/10/12 (inline rows) → D1:
rows stay phases; efforts derived from the alias's phases; target-phase
rule with tie-break; same-phase = Keep.
- simp 3, rob 1 BLOCKER/5/6/13, corr 4/5 (T3 phase edits) → D2: T3
Later/Keep with context; BDR-116 unchanged.
- simp 6, rob 8/9, corr 8 (descriptions) → D3 single rule, five YAML
forms, hardened read, allowlist, clean() with Cc/Cf and code points.
- simp 7, corr 14 (`about` on Route) → separate map, validated at load.
- simp 8, corr 13 (ASK2 Other, aliases) → fixed four labels, Other = Later.
- simp 9 (dead code) → S3 removal.
- rob 10, corr 6 (main model pick is a placebo) → toast with the real
decision and the switch hint.
- corr 7 (W3-A tests) → S6a migration step.
- corr 11 (no-key scope) → W3-A behaviour kept ([Everywhere, Later]).
- corr 15 (docs, AC5 oracle) → doc-sync noted; contract AC5 reworded.
## Confirmation ledger (r2 → r3)
- conf 1 (S6a under-listed) → leads kept by construction; toast text,
deleted free-text tests, `asked[n]` shifts listed.
- conf 2 (read-hardening clauses untested) → S6b tests + World extensions.
- conf 3 (ASK3 order) → ascending by LEVELS; tie-break in `cfg.phases` order.
- conf 4 (clean order, length) → fold first, strip, n-1 + "…".
- conf 5 (dismissed ASK2/3 toasted) → LATER checked first, silent.
- conf 6 (zero efforts, texts, echo) → Later + toast; ASK2/ASK3 texts; no echo.
- conf 7 (toast on Keep) → Change only, computed in confirmSkill.
- conf 8 (BDR-116 clause 2) → amendment at CAPITALIZE.
- conf 9 (`about` across /clear) → in `mem`, rebuilt with cfg.
- conf 10 (agents vs skills rule) → same rule at offer time, bounded.
- conf 11 (hardcoded consumers) → derived from cfg.prompt.
- conf 12 (kind check placement) → skill wrapper around readCapped.
@@ -0,0 +1,155 @@
# PLAN r3 — model-router wave 3-C: `/route forget` (2026-10-11)
r1 → r2 after simplicity CONCERNS(3), correctness CONCERNS(4), robustness
CONCERNS(4); r2 → r3 after the confirmation (correctness CONCERNS(3)).
Precedence r3 > r2 > r1. Contract
`.claude/tasks/contracts/2026-10-11-model-router-w3c-forget-1457.md`.
Base W3-B (604a6c4 + 32b71d2 + 938c3d1). Corrected facts: `Memory` holds
ok/ask/askLocal/base/confirmed/about/local/key only (no `changed`, no
`projects`); the composer gate lives in `registerCommandHook`, not in
`handleCommand`; `writeNow` always writes once the patch ran and toasts
UPDATED, returns false on refusal, throws when the rebuild fails after a
landed write; `show` prints no command list (the only visible list is the
`argumentHint`); `routeMainBySkill` copies a route into `turnMain`/`runMain`
and a rebuild never touches those slots; the shipped file holds one changed
row (security-auditor verify → judge) whose frontmatter floor was aligned
to `opus`: restoring it without re-aligning the floor makes the census
FAIL (no `changed` entry = no WARN exemption).
## Behaviour (composer `/route forget <arg>`, exactly ONE argument)
- Parsing: no arg or more than one → "usage: /route forget <name|all|projects>".
Reserved words `all` and `projects`; a row or phase named like them
cannot be forgotten by name (answer says "edit routing.json by hand").
- Guard: while a dialog is open (`st.asking !== null`) → "answer the open
dialog first", nothing written. The forget confirmation TAKES the slot
itself (`st.asking = 'forget:<target>'` in a try/finally around `askOr`)
so no first-use dialog opens under it and two forgets cannot stack.
- The decision runs on the FRESH file, twice: a PRE-read (after awaiting
`st.writes`, so a landed Keep is visible) decides the no-write paths and
the confirmation counts; the PATCH itself re-runs `forgetPlan(file,
target)` on `writeNow`'s own freshly read `file` (the closure exports the
counts actually applied; the answer reports those). Pre-read missing or
unparsable → the writer's refusal path ("routing.json is missing or
unreadable: nothing saved"); nothing to remove, no refused restore and
the key(s) not in `st.asked` → "nothing to forget for <name>" / "nothing
to forget", NO write; only a refused restore → its reason, NO write;
key(s) only in `st.asked` → reset them, NO write, "asked again at the
next use, nothing was saved" (when the key is decided by the machine
override or a project row, say "still decided by …" instead).
- `<name>`: for EVERY kind in KINDS (skills, agents, phases): delete
`confirmed.<kind>.<name>`; restore `changed.<kind>.<name>` → row =
`from` ONLY when `from` is a string, the row exists, the row equals
`changed.to`, and `from` is a phase of `file.phases` or
`DEFAULT_CONFIG.phases` (else the entry is kept and the answer says
why); delete `projects[k].<kind>.<name>` for every k, pruning an emptied
kind table then an emptied `projects[k]`; then `st.asked.delete(key)`.
- `all`: the same over every entry of `confirmed`, `changed`, `projects`;
`projects`: only `projects` emptied. Both ASK FIRST through `askOr`:
"Forget <n> decision(s): <m> row(s) restored, <k> project exception(s)
in <r> repo(s)?" options ["Cancel", "Forget"] (Cancel first; any other
answer, a dismissed dialog or headless = cancelled, nothing written);
skipped when every count is zero ("nothing to forget").
- Write: ONE `writeRouting` patch (the closure carries the counts out);
outcomes: false → "nothing saved" (the writer's toast already explains);
true → asked keys deleted, config rebuilt; a rebuild failure after a
landed write is signalled by a dedicated error class (`RebuildFailed`,
thrown by `writeNow` there and only there) → "saved, config not rebuilt:
/route reload" (asked keys untouched); any other throw (patch, fs.write)
→ "forget not applied (<err kind>)", asked keys untouched.
- Answer, one formatter (used by every form, per restored row): "forgot
<label>: <c> confirmed, <m> row(s) restored, <k> project exception(s)
removed[, <j> restore(s) kept: <reasons>]" + per restored row " (<kind>.
<name> → <from>: <alias> at <effort>; if <file> was aligned to <to>, set
model: <alias>, effort: <effort> and its lock in
lib/tests/model-routing.test.sh, then `make test`)" where `<file>` =
`agents/<name>.md` or `skills/<name>/SKILL.md` (no file clause for a
built-in agent such as Explore/Plan), alias = head of the phase's tier in
the SHIPPED `DEFAULT_CONFIG.tiers` (or the phase's model), effort from
the phase;
+ "; ~/.claude/model-router.json still sets <name> and wins here" when
`mem.local` has it, checked only after a successful rebuild;
+ "; current run may keep <phase>: /route clear to apply" when a
restored SKILL row's `to` phase equals the phase of `runMain` or
`turnMain` with source `run`/`skill`;
+ "; other live sessions see it after /route reload".
- `all` prunes emptied kind tables in `confirmed`, `changed`, `projects`
(a kept unrestorable `changed` entry stays and is reported).
- `decisions` count = confirmed entries + changed entries + project rows
(an entry counted once per table it sits in).
- The route tool has no `forget`; sub-agents cannot run commands; the
origin gate is the existing one in `registerCommandHook`.
## Docs and comments (same commit)
- `argumentHint` gains `forget <name|all|projects>` (locked by a grep in
the contract CHECK, not by a kit test); register.ts comment "Only a
dialog answer or `/route ask` writes" → "… or `/route ask` / `/route
forget` writes".
- README.md: the `/route` subcommand list (~138) gains `forget
<name|all|projects>` and the writers sentence (~140) names it;
USAGE.md: the sentence "`/route pending` liste …, `/route ask off|on`
…" (~190) gains `/route forget <nom|all|projects>` (efface une
décision, la ligne revient à la phase livrée); CHANGELOG.md: the
argument form (~10) and the writers sentence (~11); `lib/effort-shift.md`
commands line.
## Steps
- [ ] S1 `forgetPlan(file, target)` (pure over the fresh file data:
removals, restores, counts, skipped restores with reasons) and
`forgetCommand($, st, args)` (arity, reserved words, asking guard,
read, no-write paths, confirmation for all/projects, write via
`writeRouting` with a patch that applies the plan, outcome branches,
asked reset, answer); `handleCommand` case `forget`; argumentHint.
- [ ] S2 docs/comments above.
- [ ] S3 kit tests, folded: forget a confirmed skill → entry gone, asked
again in the same session, `/route pending` lists it; forget a changed
agent → row restored, entry gone, the next spawn on the restored model,
answer carries the realign clause with the file and values; forget a
phase → confirmed.phases gone, T3 asks again; forget a name with a
project exception in two repos → both rows removed, emptied keys
pruned; undecided row → "nothing to forget", no write; asked-only key
→ reset, no write; unknown name → "nothing to forget for", no write;
arity (two words → usage); reserved word collision answer; open dialog
→ refused; `all` cancelled → nothing written; `all` confirmed →
everything empty, counts in the answer; `projects` → only projects
emptied; restore refused (row ≠ to / from unknown) → entry kept, reason;
override-pinned row → the "wins here" clause; restored skill with a run
slot → the "/route clear" clause; missing file → "nothing saved";
rebuild throw → "saved, config not rebuilt"; route tool `clear` leaves
decisions; the existing origin test extended with `forget x`;
`failWrite` → "forget not applied"; (argumentHint: grep lock in the
contract, no kit test).
- [ ] S4 live smoke after the feater (orchestrator, user): one `/route
forget projects` (or a name) typed in the terminal with the dialog
answered → recorded `[gated]` in the contract (`$.ui.ask` from an
`immediate` command is unverified live).
- Disposition: honors BDR-116 (user-only writers, same path/caps; the
shared tracked memory stays the user's choice), LRN-210 (one clause per
test, folded where one setup proves two clauses), LRN-212.
## Challenge ledger (r1 → r2)
- simp 1, corr 1, rob 2 (false `st.mem` premise) → plan over the fresh
file, closure counts, pre-read no-write paths.
- simp 2, corr 2, rob 7 (kind order, reserved words, no-op writes) → all
kinds, removed-count rule, arity, reserved-word answer.
- simp 3, corr 4/9/10, rob (docs) → docs/comment scope, AC4 = argumentHint.
- simp 4 (projects form) → kept, one-liner, with the confirmation.
- simp 5/6, corr 3/7, rob 3/9 (restore + floors + validation) → guarded
restore, realign clause with file + values, one formatter.
- corr 5, rob 8 (override-pinned) → "wins here" clause.
- corr 6, rob 6 (dialog in flight) → refuse while asking.
- corr 8, rob 4 (run slot) → "/route clear" clause.
- rob 1 (no confirmation) → askOr confirmation for all/projects, Cancel first.
- rob 5 (three write outcomes) → branches.
- simp 7 (test folds) → S3 folded.
## Confirmation ledger (r2 → r3)
- conf 1 (throw mapping) → `RebuildFailed` class; other throws "not applied".
- conf 2 (slot not taken) → forget takes `st.asking` around its askOr.
- conf 3 (stale plan at write) → patch recomputes on its own file; pre-read
awaits `st.writes`; applied counts reported.
- conf 4 (run-slot rule) → kind skills, source run/skill, phase == `to`.
- conf 5 (realign clause) → conditional, shipped tiers, no file for built-ins.
- conf 6 (argumentHint test) → grep lock in the contract CHECK.
- conf 7 (live dialog from a command) → S4 smoke.
- conf 8 (ambiguities a-f) → sentences added.
- conf 9 (docs anchors) → README list + USAGE anchor + CHANGELOG form.
+1 -1
View File
@@ -36,6 +36,6 @@ claude-config/
- `skills/` = entry points you invoke via `/skill-name` - `skills/` = entry points you invoke via `/skill-name`
- `agents/` = execution units called by skills (never invoked directly by user) - `agents/` = execution units called by skills (never invoked directly by user)
- `mods/` = Claude Code mods (function-hooks plugins); each loads through the tracked symlink `skills/<name>` as `<name>@skills-dir`, live at the next session - `mods/` = Claude Code mods (function-hooks plugins); each loads through the tracked symlink `skills/<name>` as `<name>@skills-dir`, live at the next session; `mods/model-router/routing.json` (tracked) holds the phase table (tier, effort and the `about` line the first-use dialog shows), every skill and agent row and the first-use decisions
- `templates/` = symlinked to `~/.claude/templates/` — copy into projects via `/onboard` or manually - `templates/` = symlinked to `~/.claude/templates/` — copy into projects via `/onboard` or manually
- **Graphify** builds a knowledge graph of any codebase (`/graphify query`), producing a navigable wiki in `graphify-out/wiki/`. This map helps Claude understand project structure, find relevant code faster, and reason across files. Essential for large-scope tasks (multi-file features, complex bugs, architectural changes). Small tasks should skip it and read files directly. Proposed only from 200 tracked code files: the session-start banner informs, the user decides; nothing builds a graph without that go. - **Graphify** builds a knowledge graph of any codebase (`/graphify query`), producing a navigable wiki in `graphify-out/wiki/`. This map helps Claude understand project structure, find relevant code faster, and reason across files. Essential for large-scope tasks (multi-file features, complex bugs, architectural changes). Small tasks should skip it and read files directly. Proposed only from 200 tracked code files: the session-start banner informs, the user decides; nothing builds a graph without that go.
+4 -2
View File
@@ -7,11 +7,13 @@ Format follows [Keep a Changelog](https://keepachangelog.com/) and this project
## [Unreleased] ## [Unreleased]
### Added ### Added
- **model-router mod**: `mods/model-router/`, a Claude Code mod (function-hooks plugin), routes every repo skill and agent from phase rows (`plan`, `reflect`, `orchestrate`, `escalate`, `judge`, `implement`, `write`, `verify`, `explore`, `apply`, `mechanical`; built-ins Explore on sonnet/medium, Plan on opus/xhigh). A typed skill routes the main loop to its row, and a best-tier row holds across turns in a run slot until `/route clear`, `/route off`, a user `/model` or a typed skill on a non-best row. Agents get their row's model at spawn (within the tier, upward only) and its effort on every step; explicit Agent params win. Orchestrators declare their phases through the `mcp__model-router__route` tool. `ultrathink` in a prompt sets the turn's minimum effort and `/route effort=max` holds until `/route clear`; the built-in `/effort` is not a lever inside a run. `/route show` names the run slot (`main: run <phase>`), and a `null` row in the override drops a default row. The model gets a `route` tool and the user a `/route` command (`show|clear|off|on|reload|<phase>|model=<alias|id> effort=<level>|switch on|off|verbose on|off`). Optional per-machine config `~/.claude/model-router.json`, where `"enabled": false` turns it off on that machine. The spinner suffix and the status line show the route in force. It loads in every session through the tracked symlink `skills/model-router` (`model-router@skills-dir`). `make doctor` gains a Mods section; suite `make test suite=lib/tests/mods.test.sh`. Known limits: the main loop switches model only with `mainModelSwitch` on (default off, one cold-cache step per switch into another model), and the hooks send full model ids, so the `models` table has to follow new versions. - **model-router mod**: `mods/model-router/`, a Claude Code mod (function-hooks plugin), routes every repo skill and agent from phase rows (`plan`, `reflect`, `orchestrate`, `escalate`, `judge`, `implement`, `write`, `verify`, `explore`, `apply`, `mechanical`; built-ins Explore on sonnet/medium, Plan on opus/xhigh). A typed skill routes the main loop to its row, and a best-tier row holds across turns in a run slot until `/route clear`, `/route off`, a user `/model` or a typed skill on a non-best row. Agents get their row's model at spawn (within the tier, upward only) and its effort on every step; explicit Agent params win. Orchestrators declare their phases through the `mcp__model-router__route` tool. `ultrathink` in a prompt sets the turn's minimum effort and `/route effort=max` holds until `/route clear`; the built-in `/effort` is not a lever inside a run. `/route show` names the run slot (`main: run <phase>`), and a `null` row in the override drops a default row. The model gets a `route` tool and the user a `/route` command (`show|clear|off|on|reload|pending|ask on|off|forget <name|all|projects>|<phase>|model=<alias|id> effort=<level>|switch on|off|verbose on|off`). Optional per-machine config `~/.claude/model-router.json`, where `"enabled": false` turns it off on that machine. The spinner suffix and the status line show the route in force. It loads in every session through the tracked symlink `skills/model-router` (`model-router@skills-dir`). `make doctor` gains a Mods section; suite `make test suite=lib/tests/mods.test.sh`. Known limits: the main loop switches model only with `mainModelSwitch` on (default off, one cold-cache step per switch into another model), and the hooks send full model ids, so the `models` table has to follow new versions.
- **model-router first-use confirmation**: `mods/model-router/routing.json` (tracked) is the single source of the phase table and of every skill and agent row, and keeps the decisions: `confirmed`, `changed` (`from`/`to`), `projects` exceptions keyed by the origin remote reduced to `host/path` (credentials and local paths never stored), and `ask`. The first use of a rowed typed skill, a rowed agent spawn or a main-loop phase declared through the `route` tool opens a dialog with context: the skill's description (first sentence of its `SKILL.md` frontmatter) or the agent's, the phase with its `about` line (a new field on each of the 11 phases) and the model id and effort the next step really runs on. A row offers Later, Keep or Change; a main-loop phase Later or Keep. Change asks the model (fable, opus, sonnet, haiku with their tier), then the effort among those the phases of that model use (skipped when there is only one), then Everywhere or This project only; the pair maps to an existing phase (rows stay phase names, the same phase counts as Keep, a pair no phase offers is added by hand as a new phase). A skill-row change toasts the model and effort it now runs on, with the `/route switch on` hint when the main loop holds back a downgrade; a free-text answer counts as Later. One dialog at a time, never in headless (`-p`) runs or inside a sub-agent. Only a dialog answer, `/route ask on|off` or `/route forget` (which takes decisions back; a changed row returns to its shipped phase, the frontmatter floors stay yours to realign) writes the file (serialized, 64 KiB cap, never created when absent); each write asks you to commit it from the config repo. `/route pending` lists unconfirmed rows and phases. Layers: routing.json < `~/.claude/model-router.json` (its `ask` wins); a project's `.claude/model-router.json` is never read. Tests: `mods/model-router/hooks/register.test.ts`.
- **Manual-push mode**: `git config gitflow.autopush false` (human-set) now stops every push the gitflow lib makes, not only the post-commit / post-merge hooks. `gitflow start` and `finish` branch, commit and merge locally and push nothing; `gitflow delete` leaves the `origin/` copy in place and prints `git push origin --delete <br>` for the user to run. `hooks/unpushed-guard.sh` stays silent at turn end in this mode and opens each session with one `ℹ manual push mode:` line counting the commits no remote holds across every local branch; an unparseable or unreadable `gitflow.autopush` value is treated as manual push mode too, and that line names it. `hooks/push-guard.sh` (PreToolUse, `Bash|Monitor`) refuses any `git push` Claude types while `gitflow.autopush` reads false in the session cwd or in a literal `-C`/`cd` directory the command names (global config counts outside a repo); the refusal tells the user to run it with `! git push`. It reads the mode through the same lib verb as every other reader and fails closed: an unparseable or unreadable value reads as manual, and an internal error, a missing `lib/gitflow.sh`, more than 20 distinct directory tokens in one command (capped before any token is classified), a `cd`/`-C` directory token mixing quoted and unquoted parts, or a payload jq cannot parse whose raw text looks like a push refuse the push (these pathological cases fire in auto mode too). Directory tokens are read as whole shell words, adjacent quoted segments and backslash escapes included. In manual mode it over-blocks any command where a `push` word follows a `git` token; the misses listed in its header fall to a new `autoMode.soft_deny` rule that no request in the turn clears. The session banner adds `🔒 push : manual (autopush=false) — ! git push` when the key reads false, and `🔒 push : manual (autopush bad) — ! git push` when the value is invalid. Skills read the mode through a new lib verb, `bash ~/.claude/lib/gitflow.sh push-mode`: it prints `auto`, `manual` or `invalid` (rc 0) and names an invalid value on stderr (printable characters only, 64 at most). It is the one reader a skill may call, since the `git config` read of the key is denied to Claude. Skills push nothing on their own, except the `/release-candidate` tag in auto-push mode on an explicit go. Every "on origin" or "not pushed" line they print comes from `git rev-list --count origin/<br>..<br>` read after the fact, with the complete `! git …` command when something is left for the user to push. An invalid value (anything but unset, true or false, or a read that fails) is manual push mode for every reader and is named where it is read (see Fixed). Tests: `lib/gitflow-test.sh` T11b (push-mode verb), T18m and T18q blocks, `lib/tests/unpushed-guard.test.sh` T10-T16, `lib/tests/push-guard.test.sh` (98 checks). - **Manual-push mode**: `git config gitflow.autopush false` (human-set) now stops every push the gitflow lib makes, not only the post-commit / post-merge hooks. `gitflow start` and `finish` branch, commit and merge locally and push nothing; `gitflow delete` leaves the `origin/` copy in place and prints `git push origin --delete <br>` for the user to run. `hooks/unpushed-guard.sh` stays silent at turn end in this mode and opens each session with one `ℹ manual push mode:` line counting the commits no remote holds across every local branch; an unparseable or unreadable `gitflow.autopush` value is treated as manual push mode too, and that line names it. `hooks/push-guard.sh` (PreToolUse, `Bash|Monitor`) refuses any `git push` Claude types while `gitflow.autopush` reads false in the session cwd or in a literal `-C`/`cd` directory the command names (global config counts outside a repo); the refusal tells the user to run it with `! git push`. It reads the mode through the same lib verb as every other reader and fails closed: an unparseable or unreadable value reads as manual, and an internal error, a missing `lib/gitflow.sh`, more than 20 distinct directory tokens in one command (capped before any token is classified), a `cd`/`-C` directory token mixing quoted and unquoted parts, or a payload jq cannot parse whose raw text looks like a push refuse the push (these pathological cases fire in auto mode too). Directory tokens are read as whole shell words, adjacent quoted segments and backslash escapes included. In manual mode it over-blocks any command where a `push` word follows a `git` token; the misses listed in its header fall to a new `autoMode.soft_deny` rule that no request in the turn clears. The session banner adds `🔒 push : manual (autopush=false) — ! git push` when the key reads false, and `🔒 push : manual (autopush bad) — ! git push` when the value is invalid. Skills read the mode through a new lib verb, `bash ~/.claude/lib/gitflow.sh push-mode`: it prints `auto`, `manual` or `invalid` (rc 0) and names an invalid value on stderr (printable characters only, 64 at most). It is the one reader a skill may call, since the `git config` read of the key is denied to Claude. Skills push nothing on their own, except the `/release-candidate` tag in auto-push mode on an explicit go. Every "on origin" or "not pushed" line they print comes from `git rev-list --count origin/<br>..<br>` read after the fact, with the complete `! git …` command when something is left for the user to push. An invalid value (anything but unset, true or false, or a read that fails) is manual push mode for every reader and is named where it is read (see Fixed). Tests: `lib/gitflow-test.sh` T11b (push-mode verb), T18m and T18q blocks, `lib/tests/unpushed-guard.test.sh` T10-T16, `lib/tests/push-guard.test.sh` (98 checks).
### Changed ### Changed
- `lib/model-gate.md` calls the model-router `route` tool and takes its answer as the witness; with the mod off the gate stops and names `/route on`. The `model:` / `effort:` frontmatter of agents and skills is now the off-state floor, census-locked equal to the rows (`lib/tests/effort-routing.test.sh`). `analyzer` effort goes from high to xhigh. Built-in judgment dispatches carry an explicit `effort=`. - `security-auditor` runs on the `judge` row (opus/xhigh) by a first-use decision; its frontmatter floor is aligned (`model: opus`).
- `lib/model-gate.md` calls the model-router `route` tool and takes its answer as the witness; with the mod off the gate stops and names `/route on`. The `model:` / `effort:` frontmatter of agents and skills is now the off-state floor, census-locked equal to the rows (`lib/tests/effort-routing.test.sh`; a row changed through the first-use dialog passes with a `WARN floor drift` line while its frontmatter still holds the shipped value). `analyzer` effort goes from high to xhigh. Built-in judgment dispatches carry an explicit `effort=`.
- `settings.json` denies every write form of the human-only `gitflow.*` keys (18 entries): any `git … config` spelling, section remove/rename, `git -c`, the git config env overrides, and Edit/Write of `.git/config`, `.gitconfig` and `~/.config/git/config`. Side effect: Claude can no longer read `gitflow.autopush` through `git config` either; hooks and `lib/gitflow.sh` still read it. The `hard_deny` rule on routing around a guardrail now names PreToolUse hook refusals. - `settings.json` denies every write form of the human-only `gitflow.*` keys (18 entries): any `git … config` spelling, section remove/rename, `git -c`, the git config env overrides, and Edit/Write of `.git/config`, `.gitconfig` and `~/.config/git/config`. Side effect: Claude can no longer read `gitflow.autopush` through `git config` either; hooks and `lib/gitflow.sh` still read it. The `hard_deny` rule on routing around a guardrail now names PreToolUse hook refusals.
- `gitflow start` and `finish` warn on stderr when a base is behind origin and cannot fast-forward, instead of a silent `git pull --ff-only || true` (T18l, T18n). - `gitflow start` and `finish` warn on stderr when a base is behind origin and cannot fast-forward, instead of a silent `git pull --ff-only || true` (T18l, T18n).
- `/close` (`/capitalize` STEP 5C) no longer runs its own push of develop: `gitflow finish` already pushes develop in auto-push mode (BDR-095). The closing line reports the real state, read after the merge: pushed, manual push mode with the `! git push origin develop` to run, not on origin, push failed, or an invalid `gitflow.autopush` value named and treated as manual push mode. A finish whose merge landed but whose branch delete failed (rc 5/2/6) still reports the push state. - `/close` (`/capitalize` STEP 5C) no longer runs its own push of develop: `gitflow finish` already pushes develop in auto-push mode (BDR-095). The closing line reports the real state, read after the merge: pushed, manual push mode with the `! git push origin develop` to run, not on origin, push failed, or an invalid `gitflow.autopush` value named and treated as manual push mode. A finish whose merge landed but whose branch delete failed (rc 5/2/6) still reports the push state.
+13 -7
View File
@@ -83,7 +83,8 @@ inherits silently: typed agents run on their model-router row, their
| Agent | Model | Tier | | Agent | Model | Tier |
|---|---|---| |---|---|---|
| feater, hotfixer, bugfixer | sonnet (pinned) | executors — code from a closed plan (feat), fix from a closed diagnosis (bugfix), fix-bundle appliers | | feater, hotfixer, bugfixer | sonnet (pinned) | executors — code from a closed plan (feat), fix from a closed diagnosis (bugfix), fix-bundle appliers |
| verifier, security-auditor | sonnet (pinned) | fresh gates (≤3×/loop) | | verifier | sonnet (pinned) | fresh gate (≤3×/loop) |
| security-auditor | opus (judge row: opus/xhigh, the user's first-use decision; frontmatter floor aligned) | fresh gate (≤3×/loop) |
| commit-changer, release-executor, code-cleaner | sonnet (pinned) | dispatched execution — grouping+commit / release spans / approved cleanup (audit + approval gates stay in the dispatcher) | | commit-changer, release-executor, code-cleaner | sonnet (pinned) | dispatched execution — grouping+commit / release spans / approved cleanup (audit + approval gates stay in the dispatcher) |
| onboarder, scaffolder, refactorer, validator-analyzer, plugin-probe | sonnet (pinned) | workers — config generation, scaffold, refactor, deterministic W3C/WCAG runner, mechanical plugin probe | | onboarder, scaffolder, refactorer, validator-analyzer, plugin-probe | sonnet (pinned) | workers — config generation, scaffold, refactor, deterministic W3C/WCAG runner, mechanical plugin probe |
| status-reporter | haiku (pinned) | mechanical collector | | status-reporter | haiku (pinned) | mechanical collector |
@@ -104,14 +105,17 @@ children are dispatched `model:"fable"` (they carry reflection).
## Effort routing (BDR-107, BDR-108) ## Effort routing (BDR-107, BDR-108)
Second axis of the same table: how hard each phase thinks. Session default Second axis of the same table: how hard each phase thinks. Session default
`high`. The model-router mod (below) holds the live source: one phase row `high`. The live source is `mods/model-router/routing.json`, tracked with the
per repo skill and agent, each phase naming a tier and an effort level model-router mod (below): the phase table and one phase row per repo skill
and agent, each phase naming a tier and an effort level
(`plan` best/xhigh, `reflect` best/high, `orchestrate` best/medium, (`plan` best/xhigh, `reflect` best/high, `orchestrate` best/medium,
`escalate` best/max, `judge` big/xhigh, `implement` work/medium, `write` `escalate` best/max, `judge` big/xhigh, `implement` work/medium, `write`
work/high, `verify` work/xhigh, `explore` work/medium, `apply` work/low, work/high, `verify` work/xhigh, `explore` work/medium, `apply` work/low,
`mechanical` cheap/low). The `model:` and `effort:` frontmatter of every `mechanical` cheap/low). The `model:` and `effort:` frontmatter of every
typed agent and user-invoked skill stays as the off-state floor, kept equal typed agent and user-invoked skill stays as the off-state floor, kept equal
to the rows by the census `lib/tests/effort-routing.test.sh`: low to the rows by the census `lib/tests/effort-routing.test.sh` (a row
changed through the first-use dialog passes with a `WARN floor drift` line
until its frontmatter follows): low
appliers, medium executors, high writers, xhigh judgment and gates (none appliers, medium executors, high writers, xhigh judgment and gates (none
on haiku, which rejects the parameter); `/status` low … `/ship-feature` on haiku, which rejects the parameter); `/status` low … `/ship-feature`
xhigh. The vendored externals (design stack, superpowers, agent-skills, xhigh. The vendored externals (design stack, superpowers, agent-skills,
@@ -131,10 +135,12 @@ effort, never version. Transcript audit `python3 lib/effort-audit.py`.
- Main loop: every request gets the route in force. A typed skill with a row routes the main loop to it; a best-tier row (`plan`, `reflect`, `orchestrate`, `escalate`) holds across turns in a run slot until `/route clear`, `/route off`, a user `/model` or a typed skill on a non-best row. A skill without a row leaves the route as it is. - Main loop: every request gets the route in force. A typed skill with a row routes the main loop to it; a best-tier row (`plan`, `reflect`, `orchestrate`, `escalate`) holds across turns in a run slot until `/route clear`, `/route off`, a user `/model` or a typed skill on a non-best row. A skill without a row leaves the route as it is.
- Sub-agents: a routed agent gets its row's model at spawn, within its tier and only upward from its frontmatter model, and the row's effort on every step. Explicit `model` / `effort` params on the Agent call win; a project-defined agent of the same name keeps its own definition. Built-ins: Explore runs on sonnet/medium, Plan on opus/xhigh. - Sub-agents: a routed agent gets its row's model at spawn, within its tier and only upward from its frontmatter model, and the row's effort on every step. Explicit `model` / `effort` params on the Agent call win; a project-defined agent of the same name keeps its own definition. Built-ins: Explore runs on sonnet/medium, Plan on opus/xhigh.
- Levers inside a run: `ultrathink` in a prompt sets the turn's minimum effort; `/route effort=max` holds until `/route clear`. The built-in `/effort` is not a lever inside a run, rows and routes outrank it. - Levers inside a run: `ultrathink` in a prompt sets the turn's minimum effort; `/route effort=max` holds until `/route clear`. The built-in `/effort` is not a lever inside a run, rows and routes outrank it.
- `/route` (user command) shows or sets the route: `show`, `clear`, `off`, `on`, `reload`, a phase name, `model=<alias|id> effort=<level>`, `switch on|off`, `verbose on|off`. `/route show` names the run slot when one holds (`main: run <phase>`). The model sets routes through a `route` tool. - `/route` (user command) shows or sets the route: `show`, `clear`, `off`, `on`, `reload`, `pending` (rows and phases not confirmed yet), `ask on|off` (first-use dialog on or off), `forget <name|all|projects>` (takes decisions back: a changed row returns to its shipped phase; `all` and `projects` ask first), a phase name, `model=<alias|id> effort=<level>`, `switch on|off`, `verbose on|off`. `/route show` names the run slot when one holds (`main: run <phase>`). The model sets routes through a `route` tool.
- First use: the first time a rowed skill is typed, a rowed agent is spawned or a phase is declared on the main loop through the `route` tool, the mod asks once whether the route is right. The question gives context: the skill's description (first sentence of its `SKILL.md` frontmatter) or the agent's, the phase with its `about` line, and the model id and effort the next step really runs on; a declared phase also says what uses it (rows, prompt rules). A row offers Later, Keep or Change; a declared phase offers Later or Keep. Change asks the model (fable, opus, sonnet, haiku, each shown with the tier it heads), then the effort among those the phases on that model use (shipped table: fable medium, high, xhigh or max; sonnet low, medium, high or xhigh; opus and haiku have one level each, so no question), then Everywhere or This project only. Rows stay phase names, so the pair maps to an existing phase: the row's current phase wins a tie and the same phase counts as Keep; a pair no phase offers needs a new phase added by hand in routing.json. Everywhere moves the row and records the shipped phase under `changed`; This project only stores an exception under `projects`, keyed by the origin remote reduced to `host/path` (no credentials, no local paths; a remote that cannot be read that way offers no project choice). Keep lands under `confirmed`; Later, a dismissed dialog or a free-text answer asks again next session. After a skill row changes, a toast names the model and effort that skill now runs on, plus the `/route switch on` hint when the main loop holds back a downgrade. One dialog at a time, never in a headless (`-p`) run, never inside a sub-agent.
- Only a dialog answer, `/route ask on|off` or `/route forget` writes `routing.json`, never the model. Writes are serialized, capped at 64 KiB, and refused when the file is missing (it is never created). Each write leaves the config repo dirty; a toast reminds you to commit it from there.
- The spinner suffix and the status line under the prompt show the route in force. - The spinner suffix and the status line under the prompt show the route in force.
Optional per-machine config: `~/.claude/model-router.json`. Keys: `models` (alias → full id), `windows` (context window per full id), `phases`, `agents`, `skills`, `prompt` (rules), `mainModelSwitch` (default `false`), `verbose` (default `false`), `spinner` (default `true`), `enabled` (default `true`; `false` turns the mod off on that machine). A `null` value in `agents` or `skills` drops a default row. `/route reload` re-reads it. Config layers: `mods/model-router/routing.json` (tracked: phases with their tier, effort and `about` line, rows, decisions, `ask`), then the optional per-machine `~/.claude/model-router.json`, which wins. A `.claude/model-router.json` inside a project is never read. Machine keys: `models` (alias → full id), `windows` (context window per full id), `tiers` (ordered alias lists per tier: `best` fable>opus>sonnet, `big` opus>fable>sonnet, `work` sonnet>opus, `cheap` haiku>sonnet; the first alias not down is used), `fallback` (the rank order of the aliases, best first, used by the breaker), `cooldownMinutes` (how long a model stays marked down after an availability error, 15 by default, doubling per episode up to 300), `mainUpgrade` (default `true`: the main loop may move up to a phase's tier), `upgradeMaxTokens` (default 200000: no main-loop upgrade above this context size, since an upgrade re-reads the whole context cold), `phases`, `agents`, `skills`, `prompt` (rules), `mainModelSwitch` (default `false`), `verbose` (default `false`), `spinner` (default `true`), `enabled` (default `true`; `false` turns the mod off on that machine), `ask` (overrides routing.json's `ask` on that machine). A `null` value in `agents` or `skills` drops a row. Edit phases and rows by hand in routing.json (a phase's `about`, 120 characters at most, is read from there only and shown in the dialog; a model and effort pair no phase offers needs a new phase there); without it the mod runs the code's default phases, with no rows and no dialog. `/route reload` re-reads both files.
Limits: the main loop changes model only with `mainModelSwitch` on, and each switch into another model costs one cold-cache step. The hooks send full model ids, so the `models` table has to follow new model versions. Limits: the main loop changes model only with `mainModelSwitch` on, and each switch into another model costs one cold-cache step. The hooks send full model ids, so the `models` table has to follow new model versions.
@@ -228,7 +234,7 @@ a different package, ships its own conflicting `graphify` bin) — see
| `/profile` | Activate a skill profile (web / seo / web-full / full / max / backend / design / dev / qa / audit / minimal) (default: full) | | `/profile` | Activate a skill profile (web / seo / web-full / full / max / backend / design / dev / qa / audit / minimal) (default: full) |
| `/tour` | Grouped all-axes sweep — cleanup + security + reconcile + doc, fix and loop until clean | | `/tour` | Grouped all-axes sweep — cleanup + security + reconcile + doc, fix and loop until clean |
| `/site-motion` | Site-level motion: scroll engine choice, page transitions, pin/scrub sequencing across a page or Astro route (design stack) | | `/site-motion` | Site-level motion: scroll engine choice, page transitions, pin/scrub sequencing across a page or Astro route (design stack) |
| `/route` | model-router mod: show or set the main-loop route (show, clear, off, on, reload, <phase>, model=… effort=…, switch on\|off, verbose on\|off) | | `/route` | model-router mod: show or set the main-loop route (show, clear, off, on, reload, pending, ask on\|off, <phase>, model=… effort=…, switch on\|off, verbose on\|off) |
> This table lists personal skills. Gstack skills (investigate, review, retro, > This table lists personal skills. Gstack skills (investigate, review, retro,
> office-hours, cso…) and marketplace plugins add many more — run > office-hours, cso…) and marketplace plugins add many more — run
+6 -4
View File
@@ -163,7 +163,7 @@ Tu veux...
| `/pdf-translate` | Traduire un PDF vers une autre langue | Sortie HTML fidèle (images, layout, style préservés) | | `/pdf-translate` | Traduire un PDF vers une autre langue | Sortie HTML fidèle (images, layout, style préservés) |
| `/impeccable` | Audit/polish design + détecteur anti-slop déterministe | 23 verbes ; `npx impeccable detect` (exit 0/2) | | `/impeccable` | Audit/polish design + détecteur anti-slop déterministe | 23 verbes ; `npx impeccable detect` (exit 0/2) |
| `/tour` | Sweep groupé sur un ou plusieurs projets | Sécu + nettoyage + reconcile + doc, boucle jusqu'à un pass propre | | `/tour` | Sweep groupé sur un ou plusieurs projets | Sécu + nettoyage + reconcile + doc, boucle jusqu'à un pass propre |
| `/route` | Voir ou fixer la route du mod model-router | show / clear / off / on / reload / <phase> / model=… effort=… / switch on\|off / verbose on\|off | | `/route` | Voir ou fixer la route du mod model-router | show / clear / off / on / reload / pending / ask on\|off / <phase> / model=… effort=… / switch on\|off / verbose on\|off |
| `/profile` | Changer le profil de skills | web / seo / web-full / full / max / backend / design / dev / qa / audit / minimal | | `/profile` | Changer le profil de skills | web / seo / web-full / full / max / backend / design / dev / qa / audit / minimal |
> Cette table couvre les skills personnels principaux. Les plugins (gstack, > Cette table couvre les skills personnels principaux. Les plugins (gstack,
@@ -174,8 +174,8 @@ Tu veux...
### Niveau d'effort ### Niveau d'effort
Chaque commande a une ligne de phase dans le mod model-router Chaque commande a une ligne de phase dans `mods/model-router/routing.json`,
(`mods/model-router/`, actif dans chaque session), qui fixe son niveau de le fichier suivi du mod model-router (actif dans chaque session), qui fixe son niveau de
réflexion : low pour la tenue de registre (`/status`, `/close`, réflexion : low pour la tenue de registre (`/status`, `/close`,
`/commit-change`), medium pour le courant (`/gitflow`, `/prune-memory`), `/commit-change`), medium pour le courant (`/gitflow`, `/prune-memory`),
high pour un fix ou un refactor (`/feat`, `/hotfix`, `/bugfix`, high pour un fix ou un refactor (`/feat`, `/hotfix`, `/bugfix`,
@@ -187,7 +187,9 @@ l'outil `mcp__model-router__route` (`lib/effort-shift.md`). Les skills
externes vendorés (pile design, superpowers, agent-skills, skills scroll externes vendorés (pile design, superpowers, agent-skills, skills scroll
MengTo, 21st) ont aussi leur ligne. MengTo, 21st) ont aussi leur ligne.
Taper un skill qui a une ligne route la boucle principale dessus. Une ligne du tier best (plan, reflect, orchestrate, escalate) tient d'un tour à l'autre pendant tout le run, jusqu'à `/route clear`, `/route off`, un `/model` tapé ou un skill d'un autre tier tapé. Pour relancer un tour bloqué : `ultrathink` dans le prompt (plancher du tour) ou `/route effort=max` (tient jusqu'à `/route clear`). Le `/effort` intégré n'a pas d'effet dans un run : les lignes et les routes passent devant. Les sous-agents reçoivent le modèle de leur ligne au lancement (dans leur tier, jamais en dessous de leur frontmatter) et son niveau à chaque étape ; un `model` ou `effort` explicite sur l'appel gagne. Explore tourne en sonnet/medium, Plan en opus/xhigh. `/route show` affiche la route en cours, slot de run compris (`main: run <phase>`). La config par machine, optionnelle, vit dans `~/.claude/model-router.json` ; `"enabled": false` y coupe le mod sur cette machine, et une ligne à `null` y retire une ligne par défaut. Taper un skill qui a une ligne route la boucle principale dessus. Une ligne du tier best (plan, reflect, orchestrate, escalate) tient d'un tour à l'autre pendant tout le run, jusqu'à `/route clear`, `/route off`, un `/model` tapé ou un skill d'un autre tier tapé. Pour relancer un tour bloqué : `ultrathink` dans le prompt (plancher du tour) ou `/route effort=max` (tient jusqu'à `/route clear`). Le `/effort` intégré n'a pas d'effet dans un run : les lignes et les routes passent devant. Les sous-agents reçoivent le modèle de leur ligne au lancement (dans leur tier, jamais en dessous de leur frontmatter) et son niveau à chaque étape ; un `model` ou `effort` explicite sur l'appel gagne. Explore tourne en sonnet/medium, Plan en opus/xhigh. `/route show` affiche la route en cours, slot de run compris (`main: run <phase>`). Première utilisation : la première fois qu'un skill avec ligne est tapé, qu'un agent avec ligne est lancé ou qu'une phase est déclarée sur la boucle principale via l'outil `route`, le mod demande une fois si la route convient. La question donne le contexte : la description du skill (première phrase du frontmatter de son `SKILL.md`) ou celle de l'agent, la phase avec sa ligne `about` (champ de la phase dans `routing.json`), puis le modèle et le niveau réels de l'étape suivante ; une phase déclarée dit aussi qui l'utilise. Pour une ligne : Later, Keep ou Change ; pour une phase déclarée : Later ou Keep. Change demande le modèle (fable, opus, sonnet, haiku, chacun avec son tier), puis le niveau parmi ceux des phases de ce modèle (fable : medium, high, xhigh ou max ; sonnet : low, medium, high ou xhigh ; opus et haiku n'en ont qu'un, la question est sautée), puis Everywhere ou This project only (exception rangée sous le remote origin du dépôt réduit à `host/path`, jamais d'identifiants). Une ligne reste un nom de phase : le couple modèle et niveau désigne une phase existante, et la même phase qu'avant vaut Keep. Un couple qu'aucune phase n'offre s'ajoute à la main comme nouvelle phase dans `routing.json`. Après un changement sur un skill, un toast donne le modèle et le niveau réels, et rappelle `/route switch on` quand la boucle principale refuse de descendre. La réponse est écrite dans `routing.json` ; Later, un texte libre (Other) ou un dialogue fermé redemande à une session suivante. Un seul dialogue à la fois, jamais en headless (`-p`) ni dans un sous-agent, et le modèle n'écrit jamais ce fichier. Chaque réponse laisse le repo de config modifié : commite-le depuis ce repo. `/route pending` liste ce qui reste à confirmer, `/route ask off|on` coupe ou rallume le dialogue, `/route forget <nom|all|projects>` efface une décision (une ligne changée revient à la phase livrée ; `all` et `projects` demandent d'abord).
La config par machine, optionnelle, vit dans `~/.claude/model-router.json` et passe devant `routing.json` ; `"enabled": false` y coupe le mod sur cette machine, `"ask": false` y coupe le dialogue, une ligne à `null` y retire une ligne. Un `.claude/model-router.json` dans un projet n'est jamais lu. Les phases et les lignes se modifient à la main dans `routing.json` ; `/route reload` relit les deux fichiers.
## Les plugins — décision rapide ## Les plugins — décision rapide
+1 -1
View File
@@ -2,7 +2,7 @@
name: security-auditor name: security-auditor
description: 'SAST security gate — runs the pinned semgrep rulesets + the CLAUDE.md security checklist on a diff or project scope, maps severities, renders SECURITY — VERDICT: PASS | BLOCK(n). Blocks HIGH/CRITICAL only, reports the rest. Never fixes code. Fresh dispatch, no iteration history.' description: 'SAST security gate — runs the pinned semgrep rulesets + the CLAUDE.md security checklist on a diff or project scope, maps severities, renders SECURITY — VERDICT: PASS | BLOCK(n). Blocks HIGH/CRITICAL only, reports the rest. Never fixes code. Fresh dispatch, no iteration history.'
tools: Read, Grep, Glob, Bash, Write tools: Read, Grep, Glob, Bash, Write
model: sonnet model: opus
effort: xhigh effort: xhigh
--- ---
+5
View File
@@ -48,6 +48,11 @@ Builtin `/effort` is NOT a lever inside a run: rows and routes outrank it.
typed marker (unverified live 2026-10-10). typed marker (unverified live 2026-10-10).
- Headless (`-p`, SDK) runs the hooks, so routing works there too. - Headless (`-p`, SDK) runs the hooks, so routing works there too.
- Mod off: typed agents fall back to their `model:`/`effort:` frontmatter. - Mod off: typed agents fall back to their `model:`/`effort:` frontmatter.
- First use of a row asks once, with context (Later, Keep or Change: model,
then effort, then Everywhere or this project only); the answer is kept in
`mods/model-router/routing.json`. `/route pending` lists what is still
unconfirmed, `/route ask off|on` toggles the dialog, `/route forget
<name|all|projects>` takes a decision back.
Measure the split any time: `python3 ~/.claude/lib/effort-audit.py` Measure the split any time: `python3 ~/.claude/lib/effort-audit.py`
(thinking/output/cache tokens per scope, model and effort). (thinking/output/cache tokens per scope, model and effort).
+117 -52
View File
@@ -1,12 +1,16 @@
#!/usr/bin/env bash #!/usr/bin/env bash
# lib/tests/effort-routing.test.sh — wave-2 census of the model-router rows. # lib/tests/effort-routing.test.sh — wave-2 census of the model-router rows.
# Drift lock: every tracked skill/agent row in mods/model-router/hooks/ # Drift lock: every tracked skill/agent row of mods/model-router/routing.json
# register.ts equals its frontmatter (the off-state floor), the D3 wiring # equals its frontmatter (the off-state floor), except a row the user changed
# markers sit in the orchestrators, no shifter citer survives. # through the first-use dialog (WARN floor drift: the floor may lag a
# confirmed decision); the D3 wiring markers sit in the orchestrators, no
# shifter citer survives.
# shellcheck disable=SC2015,SC2016 # A && ok || ko is deliberate (ok/ko never fail) # shellcheck disable=SC2015,SC2016 # A && ok || ko is deliberate (ok/ko never fail)
set -u set -u
R="$(cd "$(dirname "$0")/../.." && pwd)" R="$(cd "$(dirname "$0")/../.." && pwd)"
REG="$R/mods/model-router/hooks/register.ts" MOD="$R/mods/model-router"
RJ="$MOD/routing.json"
REG="$MOD/hooks/register.ts"
pass=0; fail=0 pass=0; fail=0
ok() { pass=$((pass+1)); } ok() { pass=$((pass+1)); }
ko() { fail=$((fail+1)); printf 'FAIL %s\n' "$1"; } ko() { fail=$((fail+1)); printf 'FAIL %s\n' "$1"; }
@@ -16,84 +20,145 @@ lacks() { if grep -qF "$2" "$R/$1"; then ko "$1 must NOT contain: $2"; else ok;
fm() { awk 'NR==1&&/^---$/{p=1;next} p&&/^---$/{exit} p' "$1"; } fm() { awk 'NR==1&&/^---$/{p=1;next} p&&/^---$/{exit} p' "$1"; }
fm_val() { fm "$1" | grep -E "^$2: [a-z]+$" | head -1 | cut -d' ' -f2; } fm_val() { fm "$1" | grep -E "^$2: [a-z]+$" | head -1 | cut -d' ' -f2; }
# ── register.ts parsers (awk/sed on the DEFAULT_CONFIG literal) ────────── # ── routing.json (rows, phases, decisions) and the tier heads ────────────
# rows <file> <agents|skills> -> "name phase" per row # dump <routing.json> -> "skills|agents <name> <phase>", "phases <p> <tier>
rows() { # <effort>", "changed <kind> <name> <from> <to>" lines
awk -v s="$2" '$0 ~ "^ "s": \\{"{f=1;next} f&&/^ \},?$/{f=0} f' "$1" \ dump() {
| grep -v '^ *//' | grep -oE "('[^']+'|[A-Za-z0-9_-]+): '[a-z]+'" \ python3 -I - "$1" <<'PY'
| sed -E "s/'//g; s/: / /" import json, sys
} d = json.load(open(sys.argv[1]))
# phase_effort <file> <phase> -> "<tier> <effort>" for kind in ("skills", "agents"):
phase_effort() { for k, v in d.get(kind, {}).items():
awk '/^ phases: \{/{f=1;next} f&&/^ \},?$/{f=0} f' "$1" \ print(kind, k, v)
| sed -nE "s/^ *$2: \{ tier: '([a-z]+)', effort: '([a-z]+)' \},?$/\1 \2/p" for k, v in d.get("phases", {}).items():
print("phases", k, v.get("tier", ""), v.get("effort", ""))
for kind in ("skills", "agents"):
for k, v in d.get("changed", {}).get(kind, {}).items():
print("changed", kind, k, v["from"], v["to"])
PY
} }
# lookups on a dump text: row_of D kind name · phase_effort D phase ·
# changed_of D kind name -> "from to"
row_of() { printf '%s\n' "$1" | awk -v k="$2" -v n="$3" '$1==k&&$2==n{print $3}'; }
phase_effort() { printf '%s\n' "$1" | awk -v p="$2" '$1=="phases"&&$2==p{print $3, $4}'; }
changed_of() { printf '%s\n' "$1" | awk -v k="$2" -v n="$3" '$1=="changed"&&$2==k&&$3==n{print $4, $5}'; }
# tier_head <file> <tier> -> first alias of the tier list # tier_head <file> <tier> -> first alias of the tier list
tier_head() { tier_head() {
awk '/^ tiers: \{/{f=1;next} f&&/^ \},?$/{f=0} f' "$1" \ awk '/^ tiers: \{/{f=1;next} f&&/^ \},?$/{f=0} f' "$1" \
| sed -nE "s/^ *$2: \['([a-z]+)'.*$/\1/p" | sed -nE "s/^ *$2: \['([a-z]+)'.*$/\1/p"
} }
row_of() { rows "$REG" "$1" | awk -v n="$2" '$1==n{print $2}'; } # the fallback copy of the phases inside the mod's code: "<tier> <effort>"
code_phase() {
awk '/^ phases: \{/{f=1;next} f&&/^ \},?$/{f=0} f' "$REG" \
| sed -nE "s/^ *$1: \{ tier: '([a-z]+)', effort: '([a-z]+)' \},?$/\1 \2/p"
}
# field_of <dump> <phase> <effort|model> -> that phase's frontmatter value
field_of() {
local pe tier
pe="$(phase_effort "$1" "$2")"
if [ "$3" = effort ]; then printf '%s\n' "${pe#* }"; return; fi
tier="${pe% *}"; tier_head "$REG" "$tier"
}
# verdict <dump> <kind> <name> <effort|model> <got> <phase>: ok when the
# frontmatter equals the row's phase; warn when the row was changed through
# the dialog (changed.from is the shipped phase, changed.to the row) and the
# frontmatter still holds the shipped value; fail otherwise (a Keep alone,
# an empty or unknown value, a hand edit after a change)
verdict() {
local D="$1" kind="$2" name="$3" field="$4" got="$5" phase="$6" ch from fromval
[ -n "$got" ] || { echo fail; return; }
[ "$got" = "$(field_of "$D" "$phase" "$field")" ] && { echo ok; return; }
ch="$(changed_of "$D" "$kind" "$name")"; from="${ch% *}"
[ -n "$ch" ] && [ "${ch#* }" = "$phase" ] || { echo fail; return; }
fromval="$(field_of "$D" "$from" "$field")"
[ -n "$fromval" ] && [ "$got" = "$fromval" ] && echo warn || echo fail
}
# floor_check <dump> <kind> <name> <effort|model> <got> <phase> <label>
floor_check() {
case "$(verdict "$1" "$2" "$3" "$4" "$5" "$6")" in
ok) ok;;
warn) printf 'WARN floor drift: %s %s vs row %s (%s)\n' "$7" "$5" "$6" "$4"; ok;;
*) ko "$7: $4 '$5' != row $6 ($(field_of "$1" "$6" "$4"))";;
esac
}
RD="$(dump "$RJ")"
# ── flip-test: the parsers read a fixture, reject a missing key ────────── # ── flip-test: the readers parse a fixture and reject the wrong verdicts ──
FIX="$(mktemp -d)"; trap 'rm -rf "$FIX"' EXIT FIX="$(mktemp -d)"; trap 'rm -rf "$FIX"' EXIT
cat > "$FIX/reg.ts" <<'FX' cat > "$FIX/reg.ts" <<'FX'
tiers: { tiers: {
big: ['opus', 'fable'], big: ['opus', 'fable'],
}, },
phases: {
judge: { tier: 'big', effort: 'xhigh' },
},
agents: {
// judge
Plan: 'judge', 'plan-challenger': 'judge',
},
skills: {
'ship-feature': 'plan', doc: 'apply',
},
FX FX
[ "$(rows "$FIX/reg.ts" agents | tr '\n' ,)" = "Plan judge,plan-challenger judge," ] \ cat > "$FIX/routing.json" <<'FX'
&& ok || ko "flip: agents rows misparsed" {
[ "$(rows "$FIX/reg.ts" skills | tr '\n' ,)" = "ship-feature plan,doc apply," ] \ "phases": {
&& ok || ko "flip: skills rows misparsed" "judge": { "tier": "big", "effort": "xhigh" },
[ "$(phase_effort "$FIX/reg.ts" judge)" = "big xhigh" ] && ok || ko "flip: phase" "reflect": { "tier": "best", "effort": "high" },
[ -z "$(phase_effort "$FIX/reg.ts" nothere)" ] && ok || ko "flip: ghost phase" "implement": { "tier": "work", "effort": "medium" }
},
"skills": { "feat": "implement", "doc": "reflect" },
"agents": { "Plan": "judge" },
"changed": { "skills": { "feat": { "from": "reflect", "to": "implement" } } },
"ask": true
}
FX
FD="$(dump "$FIX/routing.json")"
[ "$(row_of "$FD" agents Plan)" = judge ] && ok || ko "flip: agents row misread"
[ "$(row_of "$FD" skills feat)" = implement ] && ok || ko "flip: skills row misread"
[ -z "$(row_of "$FD" skills ghost)" ] && ok || ko "flip: ghost row"
[ "$(phase_effort "$FD" judge)" = "big xhigh" ] && ok || ko "flip: phase"
[ -z "$(phase_effort "$FD" nothere)" ] && ok || ko "flip: ghost phase"
[ "$(tier_head "$FIX/reg.ts" big)" = "opus" ] && ok || ko "flip: tier head" [ "$(tier_head "$FIX/reg.ts" big)" = "opus" ] && ok || ko "flip: tier head"
[ "$(rows "$REG" skills | wc -l)" -gt 40 ] && ok || ko "register.ts: skills rows not parsed" [ "$(changed_of "$FD" skills feat)" = "reflect implement" ] && ok || ko "flip: changed entry"
[ "$(rows "$REG" agents | wc -l)" -gt 15 ] && ok || ko "register.ts: agents rows not parsed" [ -z "$(changed_of "$FD" skills doc)" ] && ok || ko "flip: ghost changed entry"
# the drift verdicts (high = reflect's effort, medium = implement's)
[ "$(verdict "$FD" skills doc effort high reflect)" = ok ] && ok || ko "flip: equal must pass"
[ "$(verdict "$FD" skills doc effort medium reflect)" = fail ] && ok || ko "flip: a Keep-only drift must FAIL"
[ "$(verdict "$FD" skills feat effort high implement)" = warn ] && ok || ko "flip: changed.from drift must WARN"
[ "$(verdict "$FD" skills feat effort '' implement)" = fail ] && ok || ko "flip: an empty value must FAIL"
[ "$(verdict "$FD" skills feat effort banana implement)" = fail ] && ok || ko "flip: an unknown value must FAIL"
[ "$(verdict "$FD" skills feat effort low implement)" = fail ] && ok || ko "flip: a drift off the shipped value must FAIL"
[ "$(verdict "$FD" skills feat effort high reflect)" = ok ] && ok || ko "flip: row back to its shipped phase passes"
[ "$(verdict "$FD" skills feat effort high judge)" = fail ] && ok || ko "flip: a hand edit after a change must FAIL"
[ "$(echo "$RD" | grep -c '^skills ')" -gt 40 ] && ok || ko "routing.json: skills rows not read"
[ "$(echo "$RD" | grep -c '^agents ')" -gt 15 ] && ok || ko "routing.json: agents rows not read"
# ── the code's fallback phases equal routing.json's (never a stale route) ─
n_code="$(awk '/^ phases: \{/{f=1;next} f&&/^ \},?$/{f=0} f' "$REG" | grep -c "tier:")"
n_json="$(echo "$RD" | grep -c '^phases ')"
[ "$n_code" -eq "$n_json" ] && ok || ko "phases: $n_code in code vs $n_json in routing.json"
while read -r _ p tier effort; do
[ "$(code_phase "$p")" = "$tier $effort" ] && ok \
|| ko "phase $p: code '$(code_phase "$p")' != routing.json '$tier $effort'"
done < <(echo "$RD" | grep '^phases ')
# ── (b) tracked skills: row exists, frontmatter effort equals the row ──── # ── (b) tracked skills: row exists, frontmatter effort equals the row ────
NO_ROW_SKILLS=" find-docs graphify impeccable model-router " NO_ROW_SKILLS=" find-docs graphify impeccable model-router "
check_skill() { check_skill() {
local f="$1" name phase want got local f="$1" name phase
name="$(basename "$(dirname "$f")")" name="$(basename "$(dirname "$f")")"
case "$NO_ROW_SKILLS" in *" $name "*) return;; esac case "$NO_ROW_SKILLS" in *" $name "*) return;; esac
phase="$(row_of skills "$name")" phase="$(row_of "$RD" skills "$name")"
[ -n "$phase" ] || { ko "skills/$name: no row in register.ts"; return; } [ -n "$phase" ] || { ko "skills/$name: no row in routing.json"; return; }
want="$(phase_effort "$REG" "$phase" | cut -d' ' -f2)" [ -n "$(fm_val "$R/$f" effort)" ] || { ko "skills/$name: routed skill without effort:"; return; }
got="$(fm_val "$R/$f" effort)" floor_check "$RD" skills "$name" effort "$(fm_val "$R/$f" effort)" "$phase" "skills/$name"
[ -n "$got" ] || { ko "skills/$name: routed skill without effort:"; return; }
[ "$got" = "$want" ] && ok || ko "skills/$name: effort $got != row $phase ($want)"
} }
while IFS= read -r f; do check_skill "$f"; done < <( while IFS= read -r f; do check_skill "$f"; done < <(
cd "$R" && git ls-files 'skills/*/SKILL.md' 'skills-external/*/SKILL.md') cd "$R" && git ls-files 'skills/*/SKILL.md' 'skills-external/*/SKILL.md')
# ── (c) tracked agents with a row: tier head == model:, effort == effort: ─ # ── (c) tracked agents with a row: tier head == model:, effort == effort: ─
NO_ROW_AGENTS=" interviewer client-handover-writer " NO_ROW_AGENTS=" interviewer client-handover-writer "
tier_alias() { tier_head "$REG" "$(phase_effort "$REG" "$1" | cut -d' ' -f1)"; }
check_agent() { check_agent() {
local f="$1" name phase alias want got local f="$1" name phase
name="$(basename "$f" .md)" name="$(basename "$f" .md)"
case "$NO_ROW_AGENTS" in *" $name "*) return;; esac case "$NO_ROW_AGENTS" in *" $name "*) return;; esac
case "$name" in impeccable-*) return;; esac case "$name" in impeccable-*) return;; esac
phase="$(row_of agents "$name")" phase="$(row_of "$RD" agents "$name")"
[ -n "$phase" ] || { ko "agents/$name: no row in register.ts"; return; } [ -n "$phase" ] || { ko "agents/$name: no row in routing.json"; return; }
alias="$(tier_alias "$phase")"; got="$(fm_val "$R/$f" model)" floor_check "$RD" agents "$name" model "$(fm_val "$R/$f" model)" "$phase" "agents/$name"
[ "$got" = "$alias" ] && ok || ko "agents/$name: model $got != row $phase ($alias)" [ "$(field_of "$RD" "$phase" model)" = haiku ] && return
[ "$alias" = haiku ] && return floor_check "$RD" agents "$name" effort "$(fm_val "$R/$f" effort)" "$phase" "agents/$name"
want="$(phase_effort "$REG" "$phase" | cut -d' ' -f2)"
got="$(fm_val "$R/$f" effort)"
[ "$got" = "$want" ] && ok || ko "agents/$name: effort $got != row $phase ($want)"
} }
while IFS= read -r f; do check_agent "$f"; done < <( while IFS= read -r f; do check_agent "$f"; done < <(
cd "$R" && git ls-files 'agents/*.md' | grep -E '^agents/[^/]+\.md$' \ cd "$R" && git ls-files 'agents/*.md' | grep -E '^agents/[^/]+\.md$' \
+1 -1
View File
@@ -21,7 +21,7 @@ done
has "agents/feater.md" 'model: sonnet' has "agents/feater.md" 'model: sonnet'
has "agents/hotfixer.md" 'model: sonnet' has "agents/hotfixer.md" 'model: sonnet'
has "agents/verifier.md" 'model: sonnet' has "agents/verifier.md" 'model: sonnet'
has "agents/security-auditor.md" 'model: sonnet' has "agents/security-auditor.md" 'model: opus' # judge row since 2026-10-11 (user first-use decision)
has "agents/analyzer.md" 'model: opus' has "agents/analyzer.md" 'model: opus'
# 4) /feat executor shape # 4) /feat executor shape
has "skills/feat/SKILL.md" 'subagent_type="feater"' has "skills/feat/SKILL.md" 'subagent_type="feater"'
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+167
View File
@@ -0,0 +1,167 @@
{
"phases": {
"plan": {
"tier": "best",
"effort": "xhigh",
"about": "design and architecture: the deepest thinking, before any code exists"
},
"reflect": {
"tier": "best",
"effort": "high",
"about": "analysis, review and synthesis: reasoning about what already exists"
},
"orchestrate": {
"tier": "best",
"effort": "medium",
"about": "dispatching and coordinating sub-agents: mostly handing out work"
},
"escalate": {
"tier": "best",
"effort": "max",
"about": "a stuck problem or a judged need: maximum effort, diagnosis only"
},
"judge": {
"tier": "big",
"effort": "xhigh",
"about": "independent judgment of work: audits, challenges and verdicts"
},
"implement": {
"tier": "work",
"effort": "medium",
"about": "writing the code of an already planned change"
},
"write": {
"tier": "work",
"effort": "high",
"about": "writing prose and docs, or a careful edit that needs polish"
},
"verify": {
"tier": "work",
"effort": "xhigh",
"about": "checking a result against its spec: tests, security review, gates"
},
"explore": {
"tier": "work",
"effort": "medium",
"about": "reading and searching the codebase to answer a question"
},
"apply": {
"tier": "work",
"effort": "low",
"about": "bookkeeping: commits, memory entries and other small edits"
},
"mechanical": {
"tier": "cheap",
"effort": "low",
"about": "scripted or trivial work needing no judgment, on the cheapest model"
}
},
"skills": {
"ship-feature": "plan",
"init-project": "plan",
"onboard": "plan",
"tour": "plan",
"audit-delta": "plan",
"analyze": "plan",
"code-clean": "plan",
"client-handover": "plan",
"brainstorming": "plan",
"writing-plans": "plan",
"requesting-code-review": "plan",
"21st-ui-review": "plan",
"feat": "reflect",
"hotfix": "reflect",
"bugfix": "reflect",
"refactor": "reflect",
"web-validate": "reflect",
"harden": "reflect",
"seo": "reflect",
"geo": "reflect",
"site-motion": "reflect",
"frontend-design": "reflect",
"emil-design-eng": "reflect",
"design-motion-principles": "reflect",
"21st-ui-build": "reflect",
"scroll-world-storytelling": "reflect",
"build-threejs-scroll-worlds": "reflect",
"scroll-scrubbed-visual-sequence": "reflect",
"scroll-scrubbed-word-reveal": "reflect",
"scroll-progress-timeline": "reflect",
"subagent-driven-development": "reflect",
"writing-skills": "reflect",
"deprecation-and-migration": "reflect",
"21st-ai": "reflect",
"21st-ui-explore": "reflect",
"gitflow": "implement",
"prune-memory": "implement",
"pdf-translate": "implement",
"ci-cd-and-automation": "implement",
"observability-and-instrumentation": "implement",
"test-driven-development": "implement",
"commit-change": "apply",
"release-candidate": "apply",
"doc": "apply",
"capitalize": "apply",
"close": "apply",
"reconcile": "apply",
"deploy": "apply",
"status": "mechanical",
"profile": "mechanical",
"plugin-check": "mechanical",
"skills-perso": "mechanical",
"using-git-worktrees": "mechanical",
"21st-cli-use": "mechanical",
"21st-registry": "mechanical",
"21st-design-sync": "mechanical"
},
"agents": {
"Explore": "explore",
"Plan": "judge",
"plan-challenger": "judge",
"plugin-advisor": "judge",
"seo-analyzer": "judge",
"geo-analyzer": "judge",
"analyzer": "judge",
"feater": "implement",
"bugfixer": "implement",
"code-cleaner": "implement",
"scaffolder": "implement",
"onboarder": "implement",
"commit-changer": "write",
"doc-syncer": "write",
"handover-doc-writer": "write",
"refactorer": "write",
"hotfixer": "apply",
"release-executor": "apply",
"plugin-probe": "apply",
"validator-analyzer": "apply",
"verifier": "verify",
"security-auditor": "judge",
"status-reporter": "mechanical"
},
"projects": {},
"confirmed": {
"agents": {
"verifier": "verify",
"feater": "implement",
"doc-syncer": "write",
"security-auditor": "judge"
},
"phases": {
"verify": "verify",
"implement": "implement",
"write": "write",
"orchestrate": "orchestrate",
"escalate": "escalate"
}
},
"changed": {
"agents": {
"security-auditor": {
"from": "verify",
"to": "judge"
}
}
},
"ask": true
}