diff --git a/.claude/memory/decisions.md b/.claude/memory/decisions.md index de2ecf2..db80814 100644 --- a/.claude/memory/decisions.md +++ b/.claude/memory/decisions.md @@ -133,6 +133,7 @@ rules: | BDR-109 | 2026-09-30 | Higgsfield pack: npm CLI `latest` + 8 upstream skills git-cloned into gitignored `skills-external/higgsfield-*`, OFF by default, in no profile; two toggles (`higgsfield` = allowlist of 7 media skills, `higgsfield-websites` = landing-page aid, never website create/deploy/publish); CLI presence by probe; routing on explicit ask | accepted | | BDR-110 | 2026-10-06 | Shell portability doctrine: native userland on macOS AND Linux, no Homebrew GNU tools on PATH; `lib/tests/portability-census.test.sh` locks deterministic GNU-only idioms | accepted | | BDR-115 | 2026-10-08 | model-router mod: pin = entry default, sub-tasks route finer; one writer per axis; full ids from the mod table; built-ins-only agents table until frontmatter pins go; state in closure | accepted | +| BDR-116 | 2026-10-11 | model-router first-use confirmation: tracked routing.json = source of phases + rows + decisions; engine dialog at first use; project exceptions in user scope keyed by normalized remote; model never writes; census tolerates a decided row | accepted | --- @@ -1403,3 +1404,13 @@ Branch feature/user-writing-web-rules, UNMERGED (human gate). - **Refs**: plan `.claude/tasks/plans/2026-10-08-model-router-mod.md`, contract `.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md`, [[BDR-107]], [[BDR-108]], [[BLK-029]], [[LRN-205]], [[LRN-206]]. - **Amendment (2026-10-09, user decisions 2026-10-08 evening)**: (a) LOAD supersedes the "Load:" line: tracked relative symlink `skills/model-router` → `../mods/model-router`, loaded in place as `model-router@skills-dir` wherever link.sh links `~/.claude/skills`; `CLAUDE_CODE_PLUGIN_DIRS` dropped (absolute path, settings `env` has no `$HOME` expansion, settings.json tracked), local marketplace dropped (`add` writes an absolute path into settings.json). Proven by fresh-process `claude plugin list --json`. (b) PRECEDENCE amended: `ultrathink` and a typed `/effort-` are the main turn's DEFAULT and MINIMUM (floor slot `turnFloor`): sticky `/route` effort > turn route effort > floor > engine, then floored; per axis; mid-turn prompt floors the running turn and the next (`wait` ignored). Rationale: user "un choix explicite bat la phase déduite"; a pure floor made `/effort-low` a no-op (challenge finding). (c) Per-machine kill switch `"enabled": false` in the untracked `~/.claude/model-router.json` (survives `/clear`, a failed reload keeps the previous config); `enabledPlugins` would dirty the tracked settings.json on every machine. (d) Hardening: `/route` composer-only; agent loops effort-only (model fixed at spawn); config caps; typed slash attested at `prompt.submit`. Commits 346d6ae, 1ff608a, 6430ac6; contracts `2026-10-08-model-router-floor-1835`, `2026-10-08-model-router-wiring-1835`; residuals parked in TODO. - **Amendment 2 (2026-10-10, wave 2 closed, commits bb56f3e + 1f2d33b)**: (a) rule 4 closed: rows for every repo skill (56) + agent (21) + Explore/Plan, PHASES by role (plan/reflect/orchestrate/escalate best · judge big · implement/write/verify/explore/apply work · mechanical cheap; `write` work/high + `apply` work/low added). User: pins "deleted or reworked", not copied → rows are the live source, tracked `model:`/`effort:` frontmatter KEPT as off-state floor, census-locked equal to rows (`lib/tests/effort-routing.test.sh`). Robustness BLOCKER closed: mod off → agents would inherit parent model. (b) Rule 5 amended: unrowed skill load changes nothing; best-tier skill row lives in `runMain` slot surviving turn end (precedence userMain > turnMain > runMain > floor > engine), dropped by `/route clear`, `/route off`, user `/model`, typed non-best skill; turn writers (route tool, prompt rules) never touch it. (c) Agents: model written at spawn WITHIN tier, upward only (never below frontmatter alias); explicit Agent params win; project-defined agent (agent.offer source projectSettings|localSettings) skipped. (d) Typed slash: name-bound marker at prompt.submit (composer|sdk|bridge) + pending slot + idle fallback (no live/spawning loop). (e) Shifters `effort-*`, `effort-pins.*`, `model-check.sh` DELETED; orchestrators call `mcp__model-router__route` per phase; gate witness = route answer id, remedy `/route on`; builtin `/effort` not a lever inside a run (levers `ultrathink`, `/route effort=max`). (f) SemVer: typed `/effort-*` removal = breaking → next release 3.0.0. Plan `.claude/tasks/plans/2026-10-09-model-router-w2-1546.md` r4, contracts `2026-10-09-model-router-w2a-1546`, `2026-10-10-model-router-w2b-1045`. Links [[LRN-210]], [[LRN-211]], [[EVAL-042]]. + +## BDR-116 — model-router first-use confirmation: tracked `routing.json` = source of phases, rows and decisions; engine dialog once per row/phase; project exceptions in user scope; model never writes [accepted] (2026-10-11) +- **Decision**: (1) `mods/model-router/routing.json` (tracked, reached from every project through the plugin dir `$.plugin.root` = the existing `~/.claude/skills/model-router` link) is the single source: 11 phases (full routes), 56 skill rows, 23 agent rows, `confirmed` (kind/name → phase), `changed` (from/to for an Everywhere change), `projects[]` exceptions, `ask`. `DEFAULT_CONFIG` keeps only the phases as fallback; rows `{}`. (2) First use of a rowed typed skill (T1), a rowed agent spawn with no explicit model (T2) or a main-loop phase declared via the route tool (T3) opens `$.ui.ask`: options Later / Keep / 2 alt phases (Other = phase name; T3 Later/Keep); a change asks Everywhere (row moved + `changed`) or This project only (`projects[key]`, confirmed = base row so no other project asks). Keep endorses the phase too. One dialog in flight; concurrent uses route unasked; headless/dismissed/garbage = Later; never inside a sub-agent; pre-ask re-read of the file (other sessions' decisions seen). (3) Project key = origin remote normalized (`new URL` host+path or strict scp regex; userinfo never read; any remaining `@`/`:`/empty host → no key); no `local:` path keys; no remote → Everywhere/Later only. The project tree's `.claude/model-router.json` is NEVER read (a cloned repo must not re-route the user's gate agents). (4) Writers = dialog answers + composer `/route ask on|off` only; serialized chain; output capped 64 KB; file never created; later unreadable → previous config kept (kill-switch rule); layers routing.json < `~/.claude/model-router.json` (its `ask` wins). (5) Census reads rows + phases from the file, locks DEFAULT phases == file, tolerates a drift only for a `changed` row whose frontmatter == `from` and row == `to` (WARN); Keep-only drift still FAILs. (6) `/route pending`, `/route ask on|off`; `set`/`confirm`/show suffix deferred. +- **Why**: user 2026-10-10: see in practice whether routing fits ("prompt qui demande de confirmer… mémoire de ce qu'on décide… met à jour la table… exception pour ce projet… persistant sur tous les projets, se redéploie comme le mod"). Tracked file = deploys with the mod via git; user-scope exceptions = security (robustness lens: project-tree layer let a cloned repo downgrade security-auditor to haiku). +- **Alternatives rejected**: `$.store` (machine-local, re-asks per machine); project file in the tree (security hole + writes into worktrees); shared-promise dedupe for parallel spawns (hook budget: awaiters time out); `confirmed` dates (git log dates them); dialog edits of phases (blast radius: a phase is shared by many rows); per-layer degrade after first load (one transient read failure wiped 79 rows + the kill switch); `local:` keys (home path in a tracked file). +- **Gates**: plan r1 → r4 (simplicity CONCERNS(3), correctness FATAL(8) with the census BLOCKER, robustness CONCERNS(8) after a network-killed first run, confirmation CONCERNS(6)); feater DONE + 4 rounds; GATE 0 MET; verifier ECARTS(3)/(2) → CONFORME, re-verify after security ECARTS(1) → fixtures; security BLOCK(1) real (password with `@`/`/` stored in the key) → fixed, PASS. Kit 86 → 190 tests. Live T2 dialogs answered by the user from the hot-loaded working-tree mod ([[LRN-212]]). +- **Refs**: contract `.claude/tasks/contracts/2026-10-10-model-router-w3a-confirm-1201.md`, plan `.claude/tasks/plans/2026-10-10-model-router-w3a-confirm-1201.md`, commit 22455c0, [[BDR-115]], [[LRN-210]], [[LRN-211]], [[EVAL-043]]. +- **Amendment (2026-10-11, wave 3-B, commit 604a6c4)**: clause (2) superseded: the dialog carries CONTEXT (skill description = first sentence of its SKILL.md frontmatter, five YAML forms, hardened read: name allowlist, stat kind file, size cap, `clean()` Cc/Cf + code points; agent description from `agent.offer`, bounded cache; phase `about` line, new 120-char field on the 11 routing.json phases, kept in memory not on Route) and the REAL next model id + effort. Options Later / Keep / Change (T3 Later / Keep). Change = model (4 fixed alias labels with tier role) → effort among those the phases headed by that alias offer (ascending, ≤ 4, one → skipped, zero → Later) → scope; the pair maps to an existing PHASE (rows stay phase names; row's current phase wins a tie; same phase = Keep); a pair no phase offers = a new phase by hand. Other anywhere = Later + toast (no echo). A main-row change toasts the real decision + `/route switch on` hint on a held downgrade. Three lenses rejected inline-route rows and dialog phase edits (2 BLOCKERs). First real use: user moved security-auditor to `judge` (opus/xhigh) through it; floor aligned (`model: opus`, 2026-10-11). Links [[EVAL-044]]. +- **Amendment (2026-10-11, wave 3-C, commit cd8d72f)**: clause (6) extended: `/route forget ` (composer only) clears decisions in the tracked file through the same writer: a name in every table (`confirmed`, `changed` → row restored to its recorded `from` when it is a string, the row still equals `to` and `from` is a known phase; `projects[*]` pruned), asked again at once; `all`/`projects` confirm through the engine dialog (Cancel first, anything else = cancelled) and hold the single-dialog slot; nothing written when nothing to forget; the answer names a restored row's file + values to realign (a floor aligned to the decision makes the census FAIL after the restore). No per-user decisions file (user: "seulement le forget"): another user inherits the shared memory and may forget. Residuals (security LOW/MEDIUM): raw error text in the answer, file fragments unsanitized in the answer, `floorFile` name not allowlisted, `__proto__` plan/apply mismatch, `changed.phases` restorable. Links [[EVAL-044]]. + diff --git a/.claude/memory/evals.md b/.claude/memory/evals.md index 5fad2a1..1d0222e 100644 --- a/.claude/memory/evals.md +++ b/.claude/memory/evals.md @@ -63,6 +63,8 @@ rules: | EVAL-040 | 2026-10-08 | model-router w1a plan: 3 challengers + 1 confirmation found 2 BLOCKER + 14 MAJOR on a plan judged closed; executor then passed every gate first time | keep the round, never dispatch a mod plan without it | | EVAL-041 | 2026-10-09 | model-router W1-C plan: 3 lenses FATAL (4 BLOCKER + 20 MAJOR) then 2 confirmations each FATAL with a NEW BLOCKER in my own revision; executor DONE first pass, 3 short text/hardening rounds | one confirmation is not enough when a revision removes a whole mechanism; the plan carried the risk, the code almost none | | EVAL-042 | 2026-10-10 | model-router W2 plan r1 → r4: 3 lenses (1 BLOCKER), 2 confirmations (1 BLOCKER then 0); W2-A verifier 3× ECARTS on coverage clauses only, W2-B ECARTS(7) → CONFORME; gate A→B read from engine records | second confirmation paid again (BLOCKER on my own r2 slot); compound coverage criterion = endless ECARTS; engine jsonl replaces the live log | +| EVAL-043 | 2026-10-11 | model-router W3-A: 3 lenses + 1 confirmation (census BLOCKER, project-tree layer dropped), feater DONE + 4 rounds, verifier 3× then re-verify after a REAL security BLOCK (credential fragment in the tracked key) | challenge + security gates both earned their cost; coverage-shaped criteria still cost 3 verifier rounds; live mod side effects misread as a test leak | +| EVAL-044 | 2026-10-11 | model-router W3-B (dialog with context): 3 lenses converged on 2 BLOCKERs of my r1 (inline rows, phase edits) + 1 confirmation; feater DONE + 2 coverage rounds; verifier 3× to CONFORME; security PASS; user's first real decision through the new dialog | the lenses pay most when the plan adds a data shape; coverage criteria still cost rounds; live dialogs during the gates must be announced and their data reconciled before commit | --- @@ -394,3 +396,18 @@ Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itse - **Method**: 3 blind lenses (simplicity CONCERNS(5), correctness CONCERNS(10), robustness FATAL(8) → BLOCKER: mod off = agents inherit parent model); confirmation 1 robustness FATAL(8) → BLOCKER introduced by r2 (route calls wiped the sticky slot); confirmation 2 correctness CONCERNS(3), no BLOCKER → r4. W2-A: feater DONE + 4 rounds (1 internal decision, 3 coverage), GATE 0 MET, verifier ECARTS(3)/(1)/(1) all coverage, user accepted at cap; security PASS. W2-B: feater DONE first pass, verifier ECARTS(7) (5 FLOOR items = planned deletions needing a CLARIFICATIONS line, 2 prose, 1 scope add) → CONFORME 10/10; security PASS; full `make test` once (env red only). Gate A→B: 3/4 probes answered from engine jsonl, probe 4 (typed skill with a live agent) unobserved, recorded as a limit. - **Anomaly**: the user's rework answer ("delete or rework, not copy") changed the design mid-plan; r2's own fix carried a BLOCKER again (as in EVAL-041). Coverage criterion: 3 verifiers, 0 defects. FLOOR guard needs the test deletions named in CLARIFICATIONS, not only in criteria. - **Action**: keep the "second confirmation after a mechanism change" rule; write coverage criteria one clause each (LRN-210); when a plan deletes tests, write the authorizing CLARIFICATIONS line BEFORE the first verifier. Links [[EVAL-041]], [[LRN-210]], [[LRN-211]], [[BDR-115]]. + +## EVAL-043 — model-router W3-A: the gates caught two real defects, the coverage criteria still cost three verifier rounds +- **Date**: 2026-10-11 +- **Output checked**: plan `.claude/tasks/plans/2026-10-10-model-router-w3a-confirm-1201.md` r1 → r4; diff 22455c0 (register.ts, routing.json, 190 kit tests, census). +- **Method**: 3 lenses (simplicity CONCERNS(3), correctness FATAL(8): BLOCKER = an Everywhere decision turns the census red; robustness: first run died on a DNS error, fresh re-dispatch on r2 CONCERNS(8): project-tree layer = security hole, shared-promise dedupe burns hook budgets, per-layer degrade wipes rows); confirmation correctness CONCERNS(6) → r4. Feater DONE first pass (169 tests); verifier ECARTS(3) → feater → ECARTS(2) → feater → CONFORME; security BLOCK(1): `normalizeRemote` kept a password tail when it held `@` or `/` → fixed (strict URL/scp parsing, no `local:` keys, output cap) → re-verify ECARTS(1) (guard fixtures) → feater → re-scan PASS. Full `make test` once (env red only). +- **Anomaly**: the live hot-loaded mod answered-by-user dialogs were first misread as a kit write leak (LRN-212). A GATE 0 CHECK written as a multi-line heredoc is not runnable by gates.sh (one line or a script). Criterion 3 again bundled ~15 clauses: three verifier rounds on coverage, zero code defects from them (LRN-210 not yet applied by me). +- **Action**: write GATE 0 oracles as scripts from the start; split coverage criteria per clause BEFORE the first verifier; when a mod is under development, announce the live side effects at each dispatch. Links [[EVAL-042]], [[BDR-116]], [[LRN-212]]. + +## EVAL-044 — model-router W3-B: the lenses killed a second row format before it existed; the live dialog answered the request's own example +- **Date**: 2026-10-11 +- **Output checked**: plan `.claude/tasks/plans/2026-10-11-model-router-w3b-dialog-1240.md` r1 → r3; diff 604a6c4 (232 kit tests), data commits 32b71d2 (security-auditor → judge) and the floor alignment. +- **Method**: simplicity CONCERNS(3), robustness FATAL(5), correctness FATAL(8): all three rejected inline `{model, effort}` rows (string tables everywhere, breaker semantics, PHASE_KEY on row names) and dialog edits of phases (census lock, blast radius, BDR-116) → r2 phases-only with efforts derived from the alias's phases; confirmation CONCERNS(2) → r3 (test migration list, read-hardening tests, ordering, clean() order). Feater DONE first pass (229 tests, 26 migrated, 41 new, 9 mutation proofs); verifier ECARTS(2) coverage → ECARTS(1) non-discriminating HOME test (the engine resolves a relative path against the plugin root) → CONFORME; security PASS (4 LOW). GATE 0 heredoc CHECK again not runnable → script (second time: write oracles as scripts from the start). +- **Anomaly**: my r1 proposed a second row shape to satisfy "model then effort" literally; the derived-effort list satisfies it with zero new data shapes. During the gates the user answered the NEW dialog live: security-auditor → opus/xhigh (the request's example), plan-challenger → verify (reset on user go); my reset of `changed` wiped a legitimate entry and left the committed file census-red for one commit. +- **Action**: when a request says "choose X then Y", look for the existing shape that already encodes (X, Y) before adding one; reconcile routing.json decisions one by one (never `changed = {}`), and run the census before every data commit. Links [[EVAL-043]], [[BDR-116]], [[LRN-212]]. + diff --git a/.claude/memory/journal.md b/.claude/memory/journal.md index ac278f4..06772d6 100644 --- a/.claude/memory/journal.md +++ b/.claude/memory/journal.md @@ -594,3 +594,8 @@ rules: ## 2026-10-10 - model-router W2-B landed (1f2d33b) + gate A→B. User /reload-plugins + typed /status; I read the engine jsonl instead of the UI log: main low on typed /status, analyzer opus/xhigh at step 0 → ordering + row-over-frontmatter proven; probe 4 unobserved (user away), limit recorded. Route tool checked live (`mcp__model-router__route`, deferred → ToolSearch once; answer names the id); session itself routed through it (reflect → orchestrate). W2-B: contract 10 criteria, feater DONE first pass (43 files), verifier ECARTS(7): 5 FLOOR items (planned test deletions → CLARIFICATIONS line), `/route on` wording, false CLAUDE.global.md sentence, plan-challenger scope add → CONFORME 10/10; I folded 3 observations by hand (SDD sonnet implementers `effort="medium"`, run-slot droppers, 80-col). Security PASS (push-guard false positive on a grep pattern; it also bit my registry heredoc). Full make test green (env red only). Doc audit SIGNIFICANT → user: apply all P1-P8, SemVer BREAKING → 3.0.0, registries all → BDR-115 amendment 2, LRN-210/211, EVAL-042. Pending: doc commit, memory commit, user merge decision (gitflow finish, human signal), publish by hand. + +## 2026-10-11 +- model-router W3-A (first-use confirmation) landed 22455c0 on feature/model-router-confirm. User asked for it after the W2 merge; 3 pass-B answers. Plan r1→r4: correctness BLOCKER (Everywhere decision = census red → `changed` WARN exemption), robustness (first challenger killed by DNS, fresh one: project-tree layer dropped for security, single dialog in flight, keep-previous after first load), confirmation CONCERNS(6). Feater DONE (169 tests) + 4 rounds. Verifier ECARTS(3)/(2)/CONFORME; security BLOCK(1) REAL: password with `@`/`/` landed in the tracked key → strict parsing, no `local:` keys, output cap → re-verify ECARTS(1) fixtures → PASS. The working-tree mod was hot-loaded by the engine: my own dispatches opened T2 dialogs the user answered (verifier Keep, feater Keep, security-auditor → implement → user reset to verify); first misread as a test leak (LRN-212). GATE 0 oracle as heredoc not runnable → script. Full make test green (env red only). Docs P1-P13 user-approved (patch in flight); BDR-116, LRN-212, EVAL-043 written. Next: doc commit, memory commit, user merge + publish by hand; T1/T3 live checks open. +- model-router W3-B (dialog with context) landed 604a6c4 + data 32b71d2 + floor chore. User asked after W3-A: explain what is decided, OK/change/later, model then effort then scope. Plan r1→r3: all three lenses killed my inline-row format and the T3 phase edit (2 BLOCKERs avoided); efforts derived from the phases of the chosen model. Feater DONE first pass (229 tests), verifier ECARTS(2)→(1)→CONFORME, security PASS. GATE 0 heredoc CHECK not runnable again → script. Live during the gates: user answered the NEW dialog for security-auditor → opus/xhigh (the request's example; kept, floor aligned to opus, model-routing lock updated), plan-challenger → verify (reset on user go), orchestrate Keep (T3 check done). My `changed = {}` reset wiped a legit entry → census red for one commit → restored. Docs P1-P6 approved (patch in flight), BDR-116 amendment + EVAL-044 written. Next: doc commit, memory commit, user merge of feature/model-router-confirm (W3-A + W3-B) + publish by hand; T1 live check still open. +- model-router W3-C `/route forget` landed cd8d72f (+ data commit: escalate Keep). User asked whether decisions can be cleared and what another user can do; chose the command only, no per-user file. Plan r1→r3: three lenses CONCERNS (false `st.mem` premise, kind order, docs scope, census red after restoring an aligned floor, no confirmation on `all`), confirmation CONCERNS(3). Feater DONE; verifier ECARTS(2) with a REAL typo (`ForgetPlan`), then coverage only ×2 → cap → user accepted; security PASS (1 MEDIUM + 4 LOW parked). My own `route(escalate)` call opened the T3 dialog the user Kept. Branch feature/model-router-confirm now carries W3-A/B/C (14 commits), unmerged, unpushed; merge on signal; full make test running. diff --git a/.claude/memory/learnings.md b/.claude/memory/learnings.md index 34b18a7..78c4e5c 100644 --- a/.claude/memory/learnings.md +++ b/.claude/memory/learnings.md @@ -1784,3 +1784,8 @@ Rule: when editing a doctrine file under structure locks, grep the test's lock s ## LRN-211 — Typed-slash routing: name-bound marker + idle fallback; run slot separate from turn routes; engine records are the live oracle - **Context**: mod needs "user typed /feat" from `skill.prompt`, which carries no origin. `prompt.submit` sees the raw `/name` first (composer|sdk|bridge): store the NAME (not a boolean; a bare flag leaked to the next preload), pending slot when mid-turn, consume only on the matching `skill.prompt`; fallback = no live/spawning loop AND allowed origin. Sticky run route in the SAME slot as turn routes was wiped by the first `route()` call (confirmation BLOCKER) → separate `runMain`, best-tier rows only (work/cheap rows leak low effort across turns). Live facts read from `~/.claude/projects//.jsonl` (+ `subagents/agent-*.jsonl`): `effort` + `message.model` per step. Typed `/status` → main low; analyzer step 0 opus/xhigh with frontmatter high (spawn bookkeeping precedes step 0; row beats frontmatter). - **Apply**: hook-side user-intent markers: bind to a name, add a pending slot for mid-turn, keep an ordering-independent fallback. Two lifetimes = two slots, never one slot with a source tag. Verify engine behaviour in the transcript jsonl, not in `$.ui.log`. Links [[BDR-115]], [[LRN-206]]. + +## LRN-212 — The engine hot-reloads a mod's hooks module from the working tree: unverified code runs live during its own feature run; dialogs answered there are facts; the kit never touches the disk +- **Context**: W3-A, 2026-10-10. Mid-run, `routing.json` gained decisions nobody expected (verifier → judge, feater → judge). First read: a kit test wrote the real file. Truth (feater, transcript timestamps = file mtimes to the second): the engine had reloaded `register.ts` from the working tree through the `skills/model-router` link, the live mod opened the T2 dialog on MY verifier/feater spawns, the user answered them in the terminal. The kit cannot write: an unmocked `fs.write` is refused ("no implementation"), a throwing mock lands on the same refusal; the real file's sha was stable across 10+ suite runs. +- **Apply**: while a mod is under development in this repo, its working-tree code is LIVE in every session (no `/reload-plugins` needed): expect its side effects (dialogs, writes) during the gates; tell the user what dialogs will pop and what to answer; reset the data file deliberately at the gate. Blame the kit last: check mtimes against the transcript before assuming a test leak. Live answers = criterion evidence (record them `[gated]`). Links [[BDR-116]], [[LRN-206]], [[LRN-211]]. + diff --git a/.claude/tasks/TODO.md b/.claude/tasks/TODO.md index d7c3756..73b575e 100644 --- a/.claude/tasks/TODO.md +++ b/.claude/tasks/TODO.md @@ -23,6 +23,18 @@ migration of shifters/pins/model-gate in wave 2 after proof; names model-router - [ ] (was) W2 gate A→B (user): `/reload-plugins`, then the live probe of plan § Gate (4 points: typed `/status` marker vs fallback in the verbose log; a real rowed spawn line + `step 0 agent` effort = spawn/first-step ordering; sonnet session typed `/feat` self-check + route answer; probe 1 again with a background agent alive). `typed-marker` never seen → W2-B blocked, A4 re-planned. - [x] (was) W2-B repo migration (plan § W2-B B0-B7 + STEP 6/7): bridge removal, `lib/effort-shift.md` rewrite, 15 citers → `route`, `effort=` on opus general-purpose dispatches, slim `lib/model-gate.md`, delete `model-check.sh` + `effort-pins.*` + their tests + install/update blocks, delete `skills/effort-*`, analyzer `effort: xhigh`, census rewrite (drift lock rows ↔ frontmatter), docs + registries (BDR-115 amendment, LRN typed-slash/run slot, EVAL) - [ ] W2-A residuals (security, accepted): first-load failure of the override activates the router despite `enabled:false` (fix = treat a failed first load as off); ReDoS on a self-authored prompt pattern (size-bounded); error text in the local log; model alias keys unvalidated (PHASE_KEY would do); `offers` map uncapped; `__proto__`/`constructor` override keys untested. Known limits: `skillCalls`/`spawning` counters are global; offers keyed by name only; builtin `/effort` is not a lever inside a run. +## 2026-10-10 — model-router wave 3-A: first-use route confirmation + decision memory (feature/model-router-confirm) +Plan `.claude/tasks/plans/2026-10-10-model-router-w3a-confirm-1201.md` r4, contract `2026-10-10-model-router-w3a-confirm-1201`. User decisions 2026-10-10: tracked routing.json reached through the plugin dir (no new link); blocking `$.ui.ask` dialog at first use; per skill row, agent row, main phase. +- [x] W3-A landed (22455c0, 2026-10-11): routing.json = phases + rows + decisions; T1/T2/T3 dialogs; project exceptions `projects[]` in the tracked file (project-tree layer dropped: security); writers = dialog + `/route ask`; `/route pending`; census from the file with the `changed` WARN exemption; kit 86 → 190. Gates: 3 lenses + 1 confirmation (BLOCKER census), feater + 4 rounds, verifier CONFORME + re-verify, security BLOCK(1) credential leak fixed → PASS, full make test green (design-tool-gate env red). Live: T2 dialogs answered by the user (verifier Keep, feater Keep, security-auditor → implement then reset to verify on user go). +- [x] W3-B dialog with context (2026-10-11, 604a6c4, contract `2026-10-11-model-router-w3b-dialog-1240`, plan r3): descriptions (skill frontmatter / agent.offer / phase `about`), real id + effort in the question, Later / Keep / Change, model → effort (derived from the phases of that model) → scope, rows stay phases, T3 Later/Keep. 3 lenses (2 BLOCKERs on my r1 avoided) + 1 confirmation, feater + 2 rounds, verifier CONFORME (3rd), security PASS (4 LOW: symlinked SKILL.md read, TOCTOU, raw key in a log, any plugin hooking AskUserQuestion could answer). Live decisions: security-auditor → judge (kept, floor aligned `model: opus`), plan-challenger → verify (reset on user go), orchestrate Keep (T3 live check done). +- [ ] W3-B residuals: a Change answered Everywhere can store a machine-only phase name (defined in `~/.claude/model-router.json`) into the tracked file (security note, correctness); SKILL.md symlink follow (LOW); `descs` cap and "(custom)" label untested; the security-auditor's `changed.from = verify` entry stays in routing.json although the floor now equals the row (harmless; prune at release). +- [ ] trace the "auto mode: use Bash/sed instead of Edit/Write" instruction block that sub-agents see attributed to the model-router MCP instructions (security-auditor 2026-10-11): it is NOT in mods/model-router source; likely the harness's own auto-mode text rendered next to the mod's MCP block. Confirm the origin. +- [x] W3-C `/route forget` (2026-10-11, cd8d72f, contract `2026-10-11-model-router-w3c-forget-1457`, plan r3): name/all/projects, guarded restore, confirmation dialog for all/projects, no-write paths, docs. 3 lenses + 1 confirmation, feater + 3 rounds (1 real defect: `ForgetPlan` typo), verifier at the cap on coverage (user accepted), security PASS. Kit 232 → 284. Live: orchestrate + escalate T3 Keeps seen. +- [ ] W3-C residuals (security, accepted): `forget not applied (${String(err)})` echoes the raw writer error (send to the log); file-sourced `from`/`to`/keys unsanitized in the answer (clean() them); `floorFile` builds a display path from an unvalidated key (SKILL_NAME check); `__proto__` accepted by `restoreOf` while `setKey` refuses it (plan/apply mismatch); `changed.phases` entries restorable (limit the scan to ROWS). Live smoke of `/route forget` from the terminal still unseen (criterion 7). +- [ ] W3-A live checks still open: T1 (typed `/feat` etc.) dialog (T3 seen live 2026-10-11: orchestrate Keep); a parallel same-agent dispatch answered after > 10 s; what an unanswered dialog resolves to on the terminal (must be Later or nothing written). Record in the contract as `[gated]` when seen. +- [ ] W3-A residuals (security LOW, accepted): non-atomic cross-process write (two sessions, crash mid-write → file reads as failed, previous config kept); `host/path` of a private remote in the tracked file; a persistent write failure re-asks every use (`asked.delete` in the catch); `out.length` vs byte size at the cap; `changePatch` scope 'project' with a vanished key writes Everywhere (cwd change mid-dialog). Deferred: `/route set`, `/route confirm`, `(unconfirmed)` in show, dialog edits of phases, frontmatter auto-alignment, session-start warning when routing.json is dirty, per-machine `projects`. +- [ ] routing.json ships `confirmed` for verifier/feater (user Keeps) and phases verify/implement: every clone inherits them (by design: decisions travel). Revisit at release if unwanted. + - [ ] W2 residuals: probe 4 unobserved (above); the push-guard hook denies a read-only grep (or a heredoc) whose TEXT contains the push verb next to `git` (security-auditor + orchestrator 2026-10-10, false positives, reworded); ~27 loose "sonnet pin"/"opus pin" shorthand sites kept (true by census); README config key list omits `tiers`/`fallback`/`cooldownMinutes`/`mainUpgrade`/`upgradeMaxTokens` (doc audit item 6, pre-existing); MIGRATION.md "Upgrading to 3.0.0" at release time; `skillCalls`/`spawning` counters global; offers keyed by name. - [ ] W3 optional: step heuristics, haiku classifier, quota-aware downgrade, A/B diff --git a/.claude/tasks/contracts/2026-10-10-model-router-w3a-confirm-1201.md b/.claude/tasks/contracts/2026-10-10-model-router-w3a-confirm-1201.md new file mode 100644 index 0000000..ba2c27f --- /dev/null +++ b/.claude/tasks/contracts/2026-10-10-model-router-w3a-confirm-1201.md @@ -0,0 +1,40 @@ +# CONTRACT — model-router-w3a-confirm +- date: 2026-10-10 | flow: feat | branch: feature/model-router-confirm (to start off develop) +- status: active + +## REQUEST (verbatim — IMMUTABLE) +une fois fini j'aimerais que pour les premiere fois, les premiers switch de model et de'effort, on est un prompt qui demqnde de confirmer si on utilise bien ce model ou si moi j'en recommande un autre. Ca permet de voir si le routing correspond bien a nos besoin dans la pratique. sois faire un truc qui se souvient en userscop (si ca demande de confirmer la route sur un autre projet pour tel tache, que ca ne ele redemande jamais sur un autre projet. une memoire de ce qu'on decide, et ca met a jour la table de routing existante si il y a des changements, et si on fait un changement demander si c'est une exception pour ce projet ou non. Faire un truc qui demande au debut, mais qui est persistant une fois demander sur tout les projet et qui se redploi automatiquement comment tout le mod. + +## CLARIFICATIONS +Q: where does the decision memory live? / A: "le 1 [fichier suivi dans le repo], mais il faut qu'il soit lisible et écrivable par tous les projets, donc le déployer (ln -s) dans le .claude du home à l'install, comme le reste" → tracked `mods/model-router/routing.json`, reached from every project through the EXISTING link `~/.claude/skills/model-router` → `../mods/model-router` (no new link: the plugin's own directory, `$.plugin`, resolves to it) [gated 2026-10-10] +Q: how is the question asked? / A: blocking dialog at the moment of the switch (`$.ui.ask`, the engine's AskUserQuestion), options Keep / Change / Later; never in headless; `/route ask off` cuts it [gated 2026-10-10] +Q: granularity? / A: per skill row, per agent row, per main-loop phase declared through the route tool; each once, across projects [gated 2026-10-10] +Q: public names (orchestrator default, user may veto): `/route pending`, `/route ask on|off` (`/route set` and `/route confirm` deferred after the simplicity lens); dialog texts in English like the rest of the mod; project exceptions live in `routing.json` under `projects[]` (r3: the project-tree file `/.claude/model-router.json` was dropped after the robustness lens showed a cloned repo could re-route the user's gate agents and that writes would land in foreign trees) [stated 2026-10-10, user may veto] + +Q: live T2 dialogs answered during the run (verifier → judge, feater → judge, Everywhere; the working-tree mod was hot-loaded by the engine without /reload-plugins) / A: user "Restaurer les deux" → routing.json reset to the shipped rows, confirmed/changed emptied; criterion 9 evidence: T2 dialog seen and written twice (texts and scope question confirmed live) [gated 2026-10-10] + +Q: live decisions during the gates (verifier Keep, feater Keep, security-auditor → implement Everywhere) / A: user "Revenir à verify" for security-auditor (row reset, its confirmed/changed entries removed); the two Keeps stay. Criterion 9 evidence: T2 dialog seen 5 times live (Keep, Change + scope, dismissed/Later), the mod hot-loaded from the working tree by the engine [gated 2026-10-11] +Q: security gate BLOCK(1) credential fragment in the remote key / A: fixed (strict URL/scp parsing, credentials never read), `local:` path keys removed (no remote → Everywhere/Later only), output size cap; re-verify ECARTS(1) on guard fixtures → closed; re-scan PASS [gated 2026-10-11] + +## ACCEPTANCE CRITERIA +1. `mods/model-router/routing.json` (tracked) is the single source of the phase table (11 full routes) and of the skill and agent rows (56 skill rows, 21 agent rows + Explore/Plan of plan r4 § Row tables), plus `confirmed` (skills/agents/phases → the confirmed phase) and `ask` (boolean); every row value is a phase key of that table; `DEFAULT_CONFIG.skills` and `.agents` in `register.ts` are `{}` (its `phases` stay as the fallback when the file is unreadable); rows and phases arrive from the file at load (session.start, or lazily once after a `/reload-plugins`), on `/route reload`, after each write, after `/clear` and on a cwd change. + CHECK: bash .claude/tasks/contracts/w3a-routing-file.sh + EXPECT: W3A-ROUTING-FILE + EVIDENCE: MET exit=0 marker-found :: W3A-ROUTING-FILE +2. Config layers, later wins: `routing.json` (phases, rows, then its `projects[]` rows for the current repo) < `~/.claude/model-router.json` (machine override, its `ask: false` honored); a `.claude/model-router.json` inside the project tree is NEVER read; phases merge first (an invalid or missing routing.json phase falls back to the code default by name, logged), rows validate against the final table; a failed layer is skipped only at the first load, afterwards a failed read keeps the whole previous config; session toggles survive a rebuild; `/clear` keeps the config and re-reads at the next prompt. One kit test per clause. +3. First use asks once, one dialog: a typed skill with a row (main, allowed origin), a rowed agent spawn without an explicit `model` (`parentAgentId` undefined), a main-loop phase declared through the route tool → `$.ui.ask` whose question carries the REAL next model id and effort (decideFor/mainEffort on main, spawnTarget + explicit effort on spawn) and the options Later / Keep / / (T3: Later / Keep, only for a phase-key call); the file is re-read right before the dialog and the key re-checked as decided (another session's decision is seen); Keep writes `confirmed.. = phase` and `confirmed.phases[phase]` and applies the row; an alt or a valid "Other" phase asks the scope (Everywhere → `routing.json` row + `changed.. = {from, to}`; This project only → `routing.json` `projects[]` row, base row untouched, `confirmed` = the base row so no other project asks; one serialized write; a key failure offers Everywhere only), the new route applies to the current decision at once; any other answer (Later, dismissed, rejected in headless, unknown phase) → default applied, nothing written, not asked again this session; at most one dialog in flight: a concurrent use (same or another key) applies its current route unasked; a phase endorsed by a Keep is not asked at T3; a row that exists in `projects[key]` or in the machine override counts as decided; never asked inside a sub-agent (`parentAgentId` set), with an explicit `model` param, with the mod off, with `ask` false, or while routing.json is unreadable. One kit test per clause. +4. Writes happen only from a dialog answer or the composer `/route ask` command (never from the route tool, a prompt rule, or any model-originated event); only to `${$.plugin.root}/routing.json`, never into a project tree; the writer refuses when the file is absent or unparsable (toast, answer = Later) and never creates it; read-modify-write of the whole JSON (2-space, key order phases, skills, agents, projects, confirmed, changed, ask), serialized through one promise chain; a write or rebuild failure logs once, applies the default and leaves the key unasked; after a write the config is rebuilt (swapped only on a successful read) and a toast says "routing.json updated: commit it from the config repo (chore branch)". One kit test per clause + reading. +5. `/route pending` lists the rows and phases not yet confirmed (asked this session first, then the rest); `/route ask on|off` toggles asking and writes `ask`; `/route reload` re-reads the three layers. (`set`, `confirm`, a show suffix: deferred.) One kit test per command. +6. The kit suite and the mods suite are green; `claude plugin validate` passes. + CHECK: cd mods/model-router && out="$(claude plugin test . 2>&1)" && printf '%s\n' "$out" | grep -qE '[0-9]+ pass' && ! printf '%s\n' "$out" | grep -qE '[1-9][0-9]* fail' && claude plugin validate . 2>&1 | grep -q 'passed' && cd ../.. && make test suite=lib/tests/mods.test.sh 2>&1 | grep -q 'all suites green' && echo W3A-MOD-GREEN + EXPECT: W3A-MOD-GREEN + EVIDENCE: MET exit=0 marker-found :: W3A-MOD-GREEN +7. `lib/tests/effort-routing.test.sh` reads rows AND phases from `routing.json` (only the tier heads still come from `register.ts`), locks `DEFAULT_CONFIG.phases` equal to the file's phases, keeps the drift lock except for a row with a `changed..` entry whose `from` route equals the frontmatter and whose `to` equals the row (then a `WARN floor drift` line, no FAIL; a Keep-only drift, an empty or unknown frontmatter value, a hand edit after a change still FAIL; flip-tested), and is green; `lib/effort-shift.md` names the first-use dialog, `/route pending` and `/route ask` in ≤ 6 added lines. + CHECK: grep -q 'routing.json' lib/tests/effort-routing.test.sh && [ "$(grep -c 'register.ts' lib/tests/effort-routing.test.sh)" -le 3 ] && grep -q 'floor drift' lib/tests/effort-routing.test.sh && bash lib/tests/effort-routing.test.sh >/dev/null 2>&1 && grep -q '/route pending' lib/effort-shift.md && [ "$(wc -l < lib/effort-shift.md)" -le 66 ] && echo W3A-CENSUS-DOC + EXPECT: W3A-CENSUS-DOC + EVIDENCE: MET exit=0 marker-found :: W3A-CENSUS-DOC +8. Hook budget: no dialog or file write on the `turn.step` path; a `$.ui.ask` in flight never blocks a second, unrelated spawn of another agent name beyond the dialog itself; the ask is awaited outside `safely` in the owning hook with a `.catch` → Later. Judged by reading. +9. Live checks after the user's `/reload-plugins` (EVIDENCE lines added by the orchestrator; the answers committed on the branch before finish): one dialog at each of the three sites; a parallel same-agent dispatch answered after more than 10 s routes the unowned spawn without a timeout; an unanswered dialog on the terminal resolves to "Later" or to nothing written. + +## FILE SCOPE +mods/model-router/routing.json (new), mods/model-router/hooks/register.ts, mods/model-router/hooks/register.test.ts, lib/tests/effort-routing.test.sh, lib/effort-shift.md; .claude/tasks/contracts/w3a-routing-file.sh (oracle, orchestrator) diff --git a/.claude/tasks/contracts/2026-10-11-model-router-w3b-dialog-1240.md b/.claude/tasks/contracts/2026-10-11-model-router-w3b-dialog-1240.md new file mode 100644 index 0000000..465eb04 --- /dev/null +++ b/.claude/tasks/contracts/2026-10-11-model-router-w3b-dialog-1240.md @@ -0,0 +1,38 @@ +# CONTRACT — model-router-w3b-dialog +- date: 2026-10-11 | flow: feat | branch: feature/model-router-confirm (continues W3-A before its merge) +- status: active + +## REQUEST (verbatim — IMMUTABLE) +j'aimerais qu'on vois pour que quand je le prompt de demande de confirmation ou changement pour un model et un effort, il faudrait qu'il soit un peu plus complet. par exemple expliquer brievement ce que fait ce pour quoi on demqnde de choisir plutot que juste le nom. car juste verifier en vrai ca peut etre pleins de chose, ou dire feature : auditor. car c'est pas tres clair on sait pas pour quoi on prend la decision. Surtout que lam auditor security je croism ca proposait sonnet alors que nonm un audit securite ca devrait etre le plus performant du model. bref metre un peu de contextm dire oui ca me vam ou changer et si on choisis de changerm proposer quel model puis quel effort, puis pour userscope et projet ca c'est bien, et l'enregistrer. + +## CLARIFICATIONS +(pass A: none — outcome, scope and flow given in the request; W3-A contract and plan r4 give the base) +Q: public shape (orchestrator default after three lenses, user may veto): ASK1 options Later / Keep / Change; ASK2 model = the four aliases labelled with their tier role (`fable (best)`, `opus (big)`, `sonnet (work)`, `haiku (cheap)`; Other = Later); ASK3 effort = the efforts offered by the phases headed by that model (fable: medium high xhigh max; opus: xhigh, skipped; sonnet: low medium high xhigh; haiku: low, skipped); the pair maps to that phase (rows stay phase names; a pair no phase offers is added by hand as a new phase in routing.json); ASK4 scope as W3-A; T3 (main-loop phase) = Later / Keep with context, no edit (BDR-116) [stated 2026-10-11] + +Q: live decisions during this run (plan-challenger judge → verify Everywhere; orchestrate phase Keep = the T3 live check) / A: user "Revenir à judge" → row reset, its confirmed/changed entries removed; orchestrate Keep kept [gated 2026-10-11] +Q: live W3-B dialog answered during the security scan: security-auditor → Change → opus → xhigh → Everywhere = row `judge` (the request's own example; the first T2 answer through the new dialog) / A: kept; my reset of `changed` had wiped its from/to entry (census red at 604a6c4) → restored in the follow-up data commit [gated 2026-10-11] +Q: GATE 1 ran 3 verifiers (ECARTS(2) coverage → ECARTS(1) one non-discriminating test → CONFORME), security PASS (4 LOW) [recorded 2026-10-11] + +## ACCEPTANCE CRITERIA +1. ASK1 text carries context: for a skill row "`/` — . Routed to (): at . OK?"; for an agent row "`` — . Routed to …"; for a phase "Phase `` — ; used by rows [+ its non-row consumers]: at . OK?"; descriptions and `about` pass `clean()` (Cc/Cf stripped, newlines folded, code-point truncation with "…"); a missing or unreadable description, an unset HOME or a name outside `^[A-Za-z0-9][A-Za-z0-9._:-]*$` degrades to the name alone (never blocks, never reads outside `${HOME}/.claude/skills//SKILL.md`, stat kind `file`, size-capped). Options: Later / Keep / Change (T3: Later / Keep); Other at ASK1 = Later + toast. One kit test per clause. +2. Change flow (rows only): ASK2 model = four fixed labels (Other or unknown = Later + toast); ASK3 effort = the efforts of the phases headed by the chosen alias, sorted ascending by level (≤ 4 labels, one → skipped, zero → Later + toast; Other = Later); target phase = the row's current phase when it matches, else the first matching phase in `cfg.phases` order; same phase as now = Keep (no `changed` entry); ASK4 as W3-A ([Everywhere, This project only], or [Everywhere, Later] with no repo key); any non-matching answer = Later, nothing written; storage as W3-A (`changed` from/to strings, `confirmed` = to); the new route applies to the current decision at once; after a main-row change the toast reports the real decision and the `/route switch on` hint when the pick is a downgrade. T3 never changes anything. One kit test per clause. +3. Rows stay phase names in every layer; `ALTS`, the alt options and the free-text phase branch of the W3-A dialog are removed (no dead code); the W3-A dialog tests are migrated to the new answers and every existing test stays green. + CHECK: ! grep -qE 'ALTS\b' mods/model-router/hooks/register.ts && echo W3B-NO-ALTS + EXPECT: W3B-NO-ALTS + EVIDENCE: MET exit=0 marker-found :: W3B-NO-ALTS +4. `routing.json` phases gain an `about` string (≤ 120 chars) for the 11 phases, loaded into a separate map (never on `Route`, never in `DEFAULT_CONFIG`), a non-string or overlong value dropped + logged once; descriptions never read from the project tree. + CHECK: bash .claude/tasks/contracts/w3b-about.sh + EXPECT: W3B-ABOUT + EVIDENCE: MET exit=0 marker-found :: W3B-ABOUT +5. Census unchanged and green (rows are phases; the DEFAULT == file phase lock keeps matching with `about` present); `lib/effort-shift.md` dialog lines updated (≤ 4 lines changed, Later/Keep/Change and the model-then-effort flow named). + CHECK: bash lib/tests/effort-routing.test.sh >/dev/null 2>&1 && grep -q 'Change' lib/effort-shift.md && [ "$(wc -l < lib/effort-shift.md)" -le 66 ] && echo W3B-CENSUS + EXPECT: W3B-CENSUS + EVIDENCE: MET exit=0 marker-found :: W3B-CENSUS +6. Kit suite, mods suite and validate green. + CHECK: cd mods/model-router && out="$(claude plugin test . 2>&1)" && printf '%s\n' "$out" | grep -qE '[0-9]+ pass' && ! printf '%s\n' "$out" | grep -qE '[1-9][0-9]* fail' && claude plugin validate . 2>&1 | grep -q 'passed' && cd ../.. && make test suite=lib/tests/mods.test.sh 2>&1 | grep -q 'all suites green' && echo W3B-MOD-GREEN + EXPECT: W3B-MOD-GREEN + EVIDENCE: MET exit=0 marker-found :: W3B-MOD-GREEN +7. Security posture unchanged: writers, paths, caps, no model-originated write; descriptions truncated and control characters stripped before they enter a question. Judged by reading + one kit test (a description with control characters and 500 chars → ≤ 140 clean chars). + +## FILE SCOPE +mods/model-router/routing.json, mods/model-router/hooks/register.ts, mods/model-router/hooks/register.test.ts, lib/tests/effort-routing.test.sh, lib/effort-shift.md; .claude/tasks/contracts/w3b-about.sh (oracle, orchestrator) diff --git a/.claude/tasks/contracts/2026-10-11-model-router-w3c-forget-1457.md b/.claude/tasks/contracts/2026-10-11-model-router-w3c-forget-1457.md new file mode 100644 index 0000000..7267127 --- /dev/null +++ b/.claude/tasks/contracts/2026-10-11-model-router-w3c-forget-1457.md @@ -0,0 +1,32 @@ +# CONTRACT — model-router-w3c-forget +- date: 2026-10-11 | flow: feat | branch: feature/model-router-confirm (continues W3-A/B before their merge) +- status: active + +## REQUEST (verbatim — IMMUTABLE) +d'ailleurs est-ce quil y a une comande pour clear les decision prise ? si jamais cest un user aure que moi si il peut clean me decisions de routage ? +[orchestrator proposal: `/route forget ` + an optional per-user decisions file] → "seulement le forget" + +## CLARIFICATIONS +Q: scope / A: `/route forget` only; no per-user decisions file (the tracked routing.json stays the shared memory; another user inherits and may forget) [gated 2026-10-11] +Q: public shape (orchestrator default, user may veto): `/route forget ` (a skill row, an agent row or a phase), `/route forget all`, `/route forget projects`; forgetting a changed row RESTORES it to its shipped phase (`changed.from`); the frontmatter floors are never touched (the answer says so when a changed row is restored) [stated 2026-10-11] + +Q: GATE 1 cap (verifier ECARTS(2) incl. one real defect `ForgetPlan` → `Plan`, then ECARTS(2)/(2) coverage only; 284 kit tests) / A: user "Accepter et commiter"; security PASS (1 MEDIUM + 4 LOW parked in TODO); live T3 Keeps on orchestrate and escalate seen during the run [gated 2026-10-11] + +## ACCEPTANCE CRITERIA +1. `/route forget ` (composer only, exactly one argument): for every kind (skills, agents, phases) removes `confirmed..`; restores a `changed` row to its `from` only when `from` is a string, the row exists, equals `changed.to` and `from` is a known phase (else the entry is kept and the answer says why); removes every `projects[*]..` (emptied tables pruned); drops the key(s) from this session's asked set; nothing removed and nothing asked → "nothing to forget for " with NO write; key only asked this session → reset, no write; refused while a dialog is open; two words → usage. One kit test per clause (folded where one setup proves two). +2. `/route forget all` and `/route forget projects` ASK FIRST (engine dialog, options Cancel / Forget, Cancel first; anything else, dismissed or headless = cancelled, nothing written; skipped when there is nothing to forget; the forget holds the single-dialog slot while asking): `all` empties `confirmed`, `changed` (rows restored under the same guards) and `projects` and clears the asked set; `projects` empties only `projects`. One kit test per clause. +3. Writes go through the existing writer only (`writeRouting`: serialized, size-capped, refused when the file is absent or unparsable with the existing toast, config rebuilt after the write); the model (route tool, prompt rules, sub-agents) can never trigger a forget; a non-composer origin is refused like the other `/route` writes. Judged by reading + one kit test (route tool cannot forget) + the existing origin test extended. +4. One answer formatter: "forgot