chore(memory): BDR-116 + LRN-212 + EVAL-043 + journal/TODO/contract w3a — feat model-router wave 3-A

This commit is contained in:
bchanot
2026-10-11 11:57:56 +02:00
parent 5bab61e13a
commit f2f404a002
8 changed files with 377 additions and 0 deletions
+9
View File
@@ -133,6 +133,7 @@ rules:
| BDR-109 | 2026-09-30 | Higgsfield pack: npm CLI `latest` + 8 upstream skills git-cloned into gitignored `skills-external/higgsfield-*`, OFF by default, in no profile; two toggles (`higgsfield` = allowlist of 7 media skills, `higgsfield-websites` = landing-page aid, never website create/deploy/publish); CLI presence by probe; routing on explicit ask | accepted |
| BDR-110 | 2026-10-06 | Shell portability doctrine: native userland on macOS AND Linux, no Homebrew GNU tools on PATH; `lib/tests/portability-census.test.sh` locks deterministic GNU-only idioms | accepted |
| BDR-115 | 2026-10-08 | model-router mod: pin = entry default, sub-tasks route finer; one writer per axis; full ids from the mod table; built-ins-only agents table until frontmatter pins go; state in closure | accepted |
| BDR-116 | 2026-10-11 | model-router first-use confirmation: tracked routing.json = source of phases + rows + decisions; engine dialog at first use; project exceptions in user scope keyed by normalized remote; model never writes; census tolerates a decided row | accepted |
---
@@ -1403,3 +1404,11 @@ Branch feature/user-writing-web-rules, UNMERGED (human gate).
- **Refs**: plan `.claude/tasks/plans/2026-10-08-model-router-mod.md`, contract `.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md`, [[BDR-107]], [[BDR-108]], [[BLK-029]], [[LRN-205]], [[LRN-206]].
- **Amendment (2026-10-09, user decisions 2026-10-08 evening)**: (a) LOAD supersedes the "Load:" line: tracked relative symlink `skills/model-router` → `../mods/model-router`, loaded in place as `model-router@skills-dir` wherever link.sh links `~/.claude/skills`; `CLAUDE_CODE_PLUGIN_DIRS` dropped (absolute path, settings `env` has no `$HOME` expansion, settings.json tracked), local marketplace dropped (`add` writes an absolute path into settings.json). Proven by fresh-process `claude plugin list --json`. (b) PRECEDENCE amended: `ultrathink` and a typed `/effort-<l>` are the main turn's DEFAULT and MINIMUM (floor slot `turnFloor`): sticky `/route` effort > turn route effort > floor > engine, then floored; per axis; mid-turn prompt floors the running turn and the next (`wait` ignored). Rationale: user "un choix explicite bat la phase déduite"; a pure floor made `/effort-low` a no-op (challenge finding). (c) Per-machine kill switch `"enabled": false` in the untracked `~/.claude/model-router.json` (survives `/clear`, a failed reload keeps the previous config); `enabledPlugins` would dirty the tracked settings.json on every machine. (d) Hardening: `/route` composer-only; agent loops effort-only (model fixed at spawn); config caps; typed slash attested at `prompt.submit`. Commits 346d6ae, 1ff608a, 6430ac6; contracts `2026-10-08-model-router-floor-1835`, `2026-10-08-model-router-wiring-1835`; residuals parked in TODO.
- **Amendment 2 (2026-10-10, wave 2 closed, commits bb56f3e + 1f2d33b)**: (a) rule 4 closed: rows for every repo skill (56) + agent (21) + Explore/Plan, PHASES by role (plan/reflect/orchestrate/escalate best · judge big · implement/write/verify/explore/apply work · mechanical cheap; `write` work/high + `apply` work/low added). User: pins "deleted or reworked", not copied → rows are the live source, tracked `model:`/`effort:` frontmatter KEPT as off-state floor, census-locked equal to rows (`lib/tests/effort-routing.test.sh`). Robustness BLOCKER closed: mod off → agents would inherit parent model. (b) Rule 5 amended: unrowed skill load changes nothing; best-tier skill row lives in `runMain` slot surviving turn end (precedence userMain > turnMain > runMain > floor > engine), dropped by `/route clear`, `/route off`, user `/model`, typed non-best skill; turn writers (route tool, prompt rules) never touch it. (c) Agents: model written at spawn WITHIN tier, upward only (never below frontmatter alias); explicit Agent params win; project-defined agent (agent.offer source projectSettings|localSettings) skipped. (d) Typed slash: name-bound marker at prompt.submit (composer|sdk|bridge) + pending slot + idle fallback (no live/spawning loop). (e) Shifters `effort-*`, `effort-pins.*`, `model-check.sh` DELETED; orchestrators call `mcp__model-router__route` per phase; gate witness = route answer id, remedy `/route on`; builtin `/effort` not a lever inside a run (levers `ultrathink`, `/route effort=max`). (f) SemVer: typed `/effort-*` removal = breaking → next release 3.0.0. Plan `.claude/tasks/plans/2026-10-09-model-router-w2-1546.md` r4, contracts `2026-10-09-model-router-w2a-1546`, `2026-10-10-model-router-w2b-1045`. Links [[LRN-210]], [[LRN-211]], [[EVAL-042]].
## BDR-116 — model-router first-use confirmation: tracked `routing.json` = source of phases, rows and decisions; engine dialog once per row/phase; project exceptions in user scope; model never writes [accepted] (2026-10-11)
- **Decision**: (1) `mods/model-router/routing.json` (tracked, reached from every project through the plugin dir `$.plugin.root` = the existing `~/.claude/skills/model-router` link) is the single source: 11 phases (full routes), 56 skill rows, 23 agent rows, `confirmed` (kind/name → phase), `changed` (from/to for an Everywhere change), `projects[<key>]` exceptions, `ask`. `DEFAULT_CONFIG` keeps only the phases as fallback; rows `{}`. (2) First use of a rowed typed skill (T1), a rowed agent spawn with no explicit model (T2) or a main-loop phase declared via the route tool (T3) opens `$.ui.ask`: options Later / Keep / 2 alt phases (Other = phase name; T3 Later/Keep); a change asks Everywhere (row moved + `changed`) or This project only (`projects[key]`, confirmed = base row so no other project asks). Keep endorses the phase too. One dialog in flight; concurrent uses route unasked; headless/dismissed/garbage = Later; never inside a sub-agent; pre-ask re-read of the file (other sessions' decisions seen). (3) Project key = origin remote normalized (`new URL` host+path or strict scp regex; userinfo never read; any remaining `@`/`:`/empty host → no key); no `local:` path keys; no remote → Everywhere/Later only. The project tree's `.claude/model-router.json` is NEVER read (a cloned repo must not re-route the user's gate agents). (4) Writers = dialog answers + composer `/route ask on|off` only; serialized chain; output capped 64 KB; file never created; later unreadable → previous config kept (kill-switch rule); layers routing.json < `~/.claude/model-router.json` (its `ask` wins). (5) Census reads rows + phases from the file, locks DEFAULT phases == file, tolerates a drift only for a `changed` row whose frontmatter == `from` and row == `to` (WARN); Keep-only drift still FAILs. (6) `/route pending`, `/route ask on|off`; `set`/`confirm`/show suffix deferred.
- **Why**: user 2026-10-10: see in practice whether routing fits ("prompt qui demande de confirmer… mémoire de ce qu'on décide… met à jour la table… exception pour ce projet… persistant sur tous les projets, se redéploie comme le mod"). Tracked file = deploys with the mod via git; user-scope exceptions = security (robustness lens: project-tree layer let a cloned repo downgrade security-auditor to haiku).
- **Alternatives rejected**: `$.store` (machine-local, re-asks per machine); project file in the tree (security hole + writes into worktrees); shared-promise dedupe for parallel spawns (hook budget: awaiters time out); `confirmed` dates (git log dates them); dialog edits of phases (blast radius: a phase is shared by many rows); per-layer degrade after first load (one transient read failure wiped 79 rows + the kill switch); `local:<realpath>` keys (home path in a tracked file).
- **Gates**: plan r1 → r4 (simplicity CONCERNS(3), correctness FATAL(8) with the census BLOCKER, robustness CONCERNS(8) after a network-killed first run, confirmation CONCERNS(6)); feater DONE + 4 rounds; GATE 0 MET; verifier ECARTS(3)/(2) → CONFORME, re-verify after security ECARTS(1) → fixtures; security BLOCK(1) real (password with `@`/`/` stored in the key) → fixed, PASS. Kit 86 → 190 tests. Live T2 dialogs answered by the user from the hot-loaded working-tree mod ([[LRN-212]]).
- **Refs**: contract `.claude/tasks/contracts/2026-10-10-model-router-w3a-confirm-1201.md`, plan `.claude/tasks/plans/2026-10-10-model-router-w3a-confirm-1201.md`, commit 22455c0, [[BDR-115]], [[LRN-210]], [[LRN-211]], [[EVAL-043]].
+9
View File
@@ -63,6 +63,7 @@ rules:
| EVAL-040 | 2026-10-08 | model-router w1a plan: 3 challengers + 1 confirmation found 2 BLOCKER + 14 MAJOR on a plan judged closed; executor then passed every gate first time | keep the round, never dispatch a mod plan without it |
| EVAL-041 | 2026-10-09 | model-router W1-C plan: 3 lenses FATAL (4 BLOCKER + 20 MAJOR) then 2 confirmations each FATAL with a NEW BLOCKER in my own revision; executor DONE first pass, 3 short text/hardening rounds | one confirmation is not enough when a revision removes a whole mechanism; the plan carried the risk, the code almost none |
| EVAL-042 | 2026-10-10 | model-router W2 plan r1 → r4: 3 lenses (1 BLOCKER), 2 confirmations (1 BLOCKER then 0); W2-A verifier 3× ECARTS on coverage clauses only, W2-B ECARTS(7) → CONFORME; gate A→B read from engine records | second confirmation paid again (BLOCKER on my own r2 slot); compound coverage criterion = endless ECARTS; engine jsonl replaces the live log |
| EVAL-043 | 2026-10-11 | model-router W3-A: 3 lenses + 1 confirmation (census BLOCKER, project-tree layer dropped), feater DONE + 4 rounds, verifier 3× then re-verify after a REAL security BLOCK (credential fragment in the tracked key) | challenge + security gates both earned their cost; coverage-shaped criteria still cost 3 verifier rounds; live mod side effects misread as a test leak |
---
@@ -394,3 +395,11 @@ Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itse
- **Method**: 3 blind lenses (simplicity CONCERNS(5), correctness CONCERNS(10), robustness FATAL(8) → BLOCKER: mod off = agents inherit parent model); confirmation 1 robustness FATAL(8) → BLOCKER introduced by r2 (route calls wiped the sticky slot); confirmation 2 correctness CONCERNS(3), no BLOCKER → r4. W2-A: feater DONE + 4 rounds (1 internal decision, 3 coverage), GATE 0 MET, verifier ECARTS(3)/(1)/(1) all coverage, user accepted at cap; security PASS. W2-B: feater DONE first pass, verifier ECARTS(7) (5 FLOOR items = planned deletions needing a CLARIFICATIONS line, 2 prose, 1 scope add) → CONFORME 10/10; security PASS; full `make test` once (env red only). Gate A→B: 3/4 probes answered from engine jsonl, probe 4 (typed skill with a live agent) unobserved, recorded as a limit.
- **Anomaly**: the user's rework answer ("delete or rework, not copy") changed the design mid-plan; r2's own fix carried a BLOCKER again (as in EVAL-041). Coverage criterion: 3 verifiers, 0 defects. FLOOR guard needs the test deletions named in CLARIFICATIONS, not only in criteria.
- **Action**: keep the "second confirmation after a mechanism change" rule; write coverage criteria one clause each (LRN-210); when a plan deletes tests, write the authorizing CLARIFICATIONS line BEFORE the first verifier. Links [[EVAL-041]], [[LRN-210]], [[LRN-211]], [[BDR-115]].
## EVAL-043 — model-router W3-A: the gates caught two real defects, the coverage criteria still cost three verifier rounds
- **Date**: 2026-10-11
- **Output checked**: plan `.claude/tasks/plans/2026-10-10-model-router-w3a-confirm-1201.md` r1 → r4; diff 22455c0 (register.ts, routing.json, 190 kit tests, census).
- **Method**: 3 lenses (simplicity CONCERNS(3), correctness FATAL(8): BLOCKER = an Everywhere decision turns the census red; robustness: first run died on a DNS error, fresh re-dispatch on r2 CONCERNS(8): project-tree layer = security hole, shared-promise dedupe burns hook budgets, per-layer degrade wipes rows); confirmation correctness CONCERNS(6) → r4. Feater DONE first pass (169 tests); verifier ECARTS(3) → feater → ECARTS(2) → feater → CONFORME; security BLOCK(1): `normalizeRemote` kept a password tail when it held `@` or `/` → fixed (strict URL/scp parsing, no `local:` keys, output cap) → re-verify ECARTS(1) (guard fixtures) → feater → re-scan PASS. Full `make test` once (env red only).
- **Anomaly**: the live hot-loaded mod answered-by-user dialogs were first misread as a kit write leak (LRN-212). A GATE 0 CHECK written as a multi-line heredoc is not runnable by gates.sh (one line or a script). Criterion 3 again bundled ~15 clauses: three verifier rounds on coverage, zero code defects from them (LRN-210 not yet applied by me).
- **Action**: write GATE 0 oracles as scripts from the start; split coverage criteria per clause BEFORE the first verifier; when a mod is under development, announce the live side effects at each dispatch. Links [[EVAL-042]], [[BDR-116]], [[LRN-212]].
+3
View File
@@ -594,3 +594,6 @@ rules:
## 2026-10-10
- model-router W2-B landed (1f2d33b) + gate A→B. User /reload-plugins + typed /status; I read the engine jsonl instead of the UI log: main low on typed /status, analyzer opus/xhigh at step 0 → ordering + row-over-frontmatter proven; probe 4 unobserved (user away), limit recorded. Route tool checked live (`mcp__model-router__route`, deferred → ToolSearch once; answer names the id); session itself routed through it (reflect → orchestrate). W2-B: contract 10 criteria, feater DONE first pass (43 files), verifier ECARTS(7): 5 FLOOR items (planned test deletions → CLARIFICATIONS line), `/route on` wording, false CLAUDE.global.md sentence, plan-challenger scope add → CONFORME 10/10; I folded 3 observations by hand (SDD sonnet implementers `effort="medium"`, run-slot droppers, 80-col). Security PASS (push-guard false positive on a grep pattern; it also bit my registry heredoc). Full make test green (env red only). Doc audit SIGNIFICANT → user: apply all P1-P8, SemVer BREAKING → 3.0.0, registries all → BDR-115 amendment 2, LRN-210/211, EVAL-042. Pending: doc commit, memory commit, user merge decision (gitflow finish, human signal), publish by hand.
## 2026-10-11
- model-router W3-A (first-use confirmation) landed 22455c0 on feature/model-router-confirm. User asked for it after the W2 merge; 3 pass-B answers. Plan r1→r4: correctness BLOCKER (Everywhere decision = census red → `changed` WARN exemption), robustness (first challenger killed by DNS, fresh one: project-tree layer dropped for security, single dialog in flight, keep-previous after first load), confirmation CONCERNS(6). Feater DONE (169 tests) + 4 rounds. Verifier ECARTS(3)/(2)/CONFORME; security BLOCK(1) REAL: password with `@`/`/` landed in the tracked key → strict parsing, no `local:` keys, output cap → re-verify ECARTS(1) fixtures → PASS. The working-tree mod was hot-loaded by the engine: my own dispatches opened T2 dialogs the user answered (verifier Keep, feater Keep, security-auditor → implement → user reset to verify); first misread as a test leak (LRN-212). GATE 0 oracle as heredoc not runnable → script. Full make test green (env red only). Docs P1-P13 user-approved (patch in flight); BDR-116, LRN-212, EVAL-043 written. Next: doc commit, memory commit, user merge + publish by hand; T1/T3 live checks open.
+5
View File
@@ -1784,3 +1784,8 @@ Rule: when editing a doctrine file under structure locks, grep the test's lock s
## LRN-211 — Typed-slash routing: name-bound marker + idle fallback; run slot separate from turn routes; engine records are the live oracle
- **Context**: mod needs "user typed /feat" from `skill.prompt`, which carries no origin. `prompt.submit` sees the raw `/name` first (composer|sdk|bridge): store the NAME (not a boolean; a bare flag leaked to the next preload), pending slot when mid-turn, consume only on the matching `skill.prompt`; fallback = no live/spawning loop AND allowed origin. Sticky run route in the SAME slot as turn routes was wiped by the first `route()` call (confirmation BLOCKER) → separate `runMain`, best-tier rows only (work/cheap rows leak low effort across turns). Live facts read from `~/.claude/projects/<repo>/<session>.jsonl` (+ `subagents/agent-*.jsonl`): `effort` + `message.model` per step. Typed `/status` → main low; analyzer step 0 opus/xhigh with frontmatter high (spawn bookkeeping precedes step 0; row beats frontmatter).
- **Apply**: hook-side user-intent markers: bind to a name, add a pending slot for mid-turn, keep an ordering-independent fallback. Two lifetimes = two slots, never one slot with a source tag. Verify engine behaviour in the transcript jsonl, not in `$.ui.log`. Links [[BDR-115]], [[LRN-206]].
## LRN-212 — The engine hot-reloads a mod's hooks module from the working tree: unverified code runs live during its own feature run; dialogs answered there are facts; the kit never touches the disk
- **Context**: W3-A, 2026-10-10. Mid-run, `routing.json` gained decisions nobody expected (verifier → judge, feater → judge). First read: a kit test wrote the real file. Truth (feater, transcript timestamps = file mtimes to the second): the engine had reloaded `register.ts` from the working tree through the `skills/model-router` link, the live mod opened the T2 dialog on MY verifier/feater spawns, the user answered them in the terminal. The kit cannot write: an unmocked `fs.write` is refused ("no implementation"), a throwing mock lands on the same refusal; the real file's sha was stable across 10+ suite runs.
- **Apply**: while a mod is under development in this repo, its working-tree code is LIVE in every session (no `/reload-plugins` needed): expect its side effects (dialogs, writes) during the gates; tell the user what dialogs will pop and what to answer; reset the data file deliberately at the gate. Blame the kit last: check mtimes against the transcript before assuming a test leak. Live answers = criterion evidence (record them `[gated]`). Links [[BDR-116]], [[LRN-206]], [[LRN-211]].