chore(memory): BDR-116 + LRN-212 + EVAL-043 + journal/TODO/contract w3a — feat model-router wave 3-A

This commit is contained in:
bchanot
2026-10-11 11:57:56 +02:00
parent 5bab61e13a
commit f2f404a002
8 changed files with 377 additions and 0 deletions
+9
View File
@@ -63,6 +63,7 @@ rules:
| EVAL-040 | 2026-10-08 | model-router w1a plan: 3 challengers + 1 confirmation found 2 BLOCKER + 14 MAJOR on a plan judged closed; executor then passed every gate first time | keep the round, never dispatch a mod plan without it |
| EVAL-041 | 2026-10-09 | model-router W1-C plan: 3 lenses FATAL (4 BLOCKER + 20 MAJOR) then 2 confirmations each FATAL with a NEW BLOCKER in my own revision; executor DONE first pass, 3 short text/hardening rounds | one confirmation is not enough when a revision removes a whole mechanism; the plan carried the risk, the code almost none |
| EVAL-042 | 2026-10-10 | model-router W2 plan r1 → r4: 3 lenses (1 BLOCKER), 2 confirmations (1 BLOCKER then 0); W2-A verifier 3× ECARTS on coverage clauses only, W2-B ECARTS(7) → CONFORME; gate A→B read from engine records | second confirmation paid again (BLOCKER on my own r2 slot); compound coverage criterion = endless ECARTS; engine jsonl replaces the live log |
| EVAL-043 | 2026-10-11 | model-router W3-A: 3 lenses + 1 confirmation (census BLOCKER, project-tree layer dropped), feater DONE + 4 rounds, verifier 3× then re-verify after a REAL security BLOCK (credential fragment in the tracked key) | challenge + security gates both earned their cost; coverage-shaped criteria still cost 3 verifier rounds; live mod side effects misread as a test leak |
---
@@ -394,3 +395,11 @@ Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itse
- **Method**: 3 blind lenses (simplicity CONCERNS(5), correctness CONCERNS(10), robustness FATAL(8) → BLOCKER: mod off = agents inherit parent model); confirmation 1 robustness FATAL(8) → BLOCKER introduced by r2 (route calls wiped the sticky slot); confirmation 2 correctness CONCERNS(3), no BLOCKER → r4. W2-A: feater DONE + 4 rounds (1 internal decision, 3 coverage), GATE 0 MET, verifier ECARTS(3)/(1)/(1) all coverage, user accepted at cap; security PASS. W2-B: feater DONE first pass, verifier ECARTS(7) (5 FLOOR items = planned deletions needing a CLARIFICATIONS line, 2 prose, 1 scope add) → CONFORME 10/10; security PASS; full `make test` once (env red only). Gate A→B: 3/4 probes answered from engine jsonl, probe 4 (typed skill with a live agent) unobserved, recorded as a limit.
- **Anomaly**: the user's rework answer ("delete or rework, not copy") changed the design mid-plan; r2's own fix carried a BLOCKER again (as in EVAL-041). Coverage criterion: 3 verifiers, 0 defects. FLOOR guard needs the test deletions named in CLARIFICATIONS, not only in criteria.
- **Action**: keep the "second confirmation after a mechanism change" rule; write coverage criteria one clause each (LRN-210); when a plan deletes tests, write the authorizing CLARIFICATIONS line BEFORE the first verifier. Links [[EVAL-041]], [[LRN-210]], [[LRN-211]], [[BDR-115]].
## EVAL-043 — model-router W3-A: the gates caught two real defects, the coverage criteria still cost three verifier rounds
- **Date**: 2026-10-11
- **Output checked**: plan `.claude/tasks/plans/2026-10-10-model-router-w3a-confirm-1201.md` r1 → r4; diff 22455c0 (register.ts, routing.json, 190 kit tests, census).
- **Method**: 3 lenses (simplicity CONCERNS(3), correctness FATAL(8): BLOCKER = an Everywhere decision turns the census red; robustness: first run died on a DNS error, fresh re-dispatch on r2 CONCERNS(8): project-tree layer = security hole, shared-promise dedupe burns hook budgets, per-layer degrade wipes rows); confirmation correctness CONCERNS(6) → r4. Feater DONE first pass (169 tests); verifier ECARTS(3) → feater → ECARTS(2) → feater → CONFORME; security BLOCK(1): `normalizeRemote` kept a password tail when it held `@` or `/` → fixed (strict URL/scp parsing, no `local:` keys, output cap) → re-verify ECARTS(1) (guard fixtures) → feater → re-scan PASS. Full `make test` once (env red only).
- **Anomaly**: the live hot-loaded mod answered-by-user dialogs were first misread as a kit write leak (LRN-212). A GATE 0 CHECK written as a multi-line heredoc is not runnable by gates.sh (one line or a script). Criterion 3 again bundled ~15 clauses: three verifier rounds on coverage, zero code defects from them (LRN-210 not yet applied by me).
- **Action**: write GATE 0 oracles as scripts from the start; split coverage criteria per clause BEFORE the first verifier; when a mod is under development, announce the live side effects at each dispatch. Links [[EVAL-042]], [[BDR-116]], [[LRN-212]].