chore(memory): LRN-207/208/209 + EVAL-041 — availability signal, blind security brief, scoped test runs, W1-C challenge value
This commit is contained in:
@@ -61,6 +61,7 @@ rules:
|
|||||||
| EVAL-038 | 2026-09-29 | correction of EVAL-037: 94 % of sub-agent usage records carry no `output_tokens_details` (Fable subs at xhigh read 0 thinking, impossible with always-on thinking) → sub-agent thinking UNMEASURED, not ≈0; main loop 100 % counted; weighted-cost split (61/39) still holds | `effort-audit.py` prints coverage + CAVEAT; cite the cost split only; agent effort pins stay unmeasured; a tier move on a price argument = judgment, not figure |
|
| EVAL-038 | 2026-09-29 | correction of EVAL-037: 94 % of sub-agent usage records carry no `output_tokens_details` (Fable subs at xhigh read 0 thinking, impossible with always-on thinking) → sub-agent thinking UNMEASURED, not ≈0; main loop 100 % counted; weighted-cost split (61/39) still holds | `effort-audit.py` prints coverage + CAVEAT; cite the cost split only; agent effort pins stay unmeasured; a tier move on a price argument = judgment, not figure |
|
||||||
| EVAL-039 | 2026-09-30 | ship-feature run higgsfield-pack: plan dry-run in scratch → 0 executor failure on 7 tasks; challenge found 7 MAJOR I missed; floor-guard caught 2 shellcheck suppressions of mine; final review found README/code gap | keep |
|
| EVAL-039 | 2026-09-30 | ship-feature run higgsfield-pack: plan dry-run in scratch → 0 executor failure on 7 tasks; challenge found 7 MAJOR I missed; floor-guard caught 2 shellcheck suppressions of mine; final review found README/code gap | keep |
|
||||||
| EVAL-040 | 2026-10-08 | model-router w1a plan: 3 challengers + 1 confirmation found 2 BLOCKER + 14 MAJOR on a plan judged closed; executor then passed every gate first time | keep the round, never dispatch a mod plan without it |
|
| EVAL-040 | 2026-10-08 | model-router w1a plan: 3 challengers + 1 confirmation found 2 BLOCKER + 14 MAJOR on a plan judged closed; executor then passed every gate first time | keep the round, never dispatch a mod plan without it |
|
||||||
|
| EVAL-041 | 2026-10-09 | model-router W1-C plan: 3 lenses FATAL (4 BLOCKER + 20 MAJOR) then 2 confirmations each FATAL with a NEW BLOCKER in my own revision; executor DONE first pass, 3 short text/hardening rounds | one confirmation is not enough when a revision removes a whole mechanism; the plan carried the risk, the code almost none |
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -379,3 +380,10 @@ Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itse
|
|||||||
- **Anomaly**: r1 carried 2 BLOCKER (explicit Agent `model` overridden at every step; Skill bridge answer shape refused by the output schema → doublon kept) + 10 MAJOR, all invisible to me: spike levers carried over as design (`agentsNext`, per-step model rewrite), a table copying 21 pins = third source of truth, writes on main from sub-agent loops. Confirmation found 4 more MAJOR (explicit params vs in-agent writes, `e.wait`, vacuous test assertions, haiku effort). Executor then DONE first pass, verifier CONFORME 6/6, security PASS: the plan was the whole risk.
|
- **Anomaly**: r1 carried 2 BLOCKER (explicit Agent `model` overridden at every step; Skill bridge answer shape refused by the output schema → doublon kept) + 10 MAJOR, all invisible to me: spike levers carried over as design (`agentsNext`, per-step model rewrite), a table copying 21 pins = third source of truth, writes on main from sub-agent loops. Confirmation found 4 more MAJOR (explicit params vs in-agent writes, `e.wait`, vacuous test assertions, haiku effort). Executor then DONE first pass, verifier CONFORME 6/6, security PASS: the plan was the whole risk.
|
||||||
- **Action**: a mod plan always goes through the full round + confirmation; test assertions must read the one line that carries the value; spike code is a FACT source, never a design source ([[BDR-115]], [[LRN-205]], [[LRN-206]]).
|
- **Action**: a mod plan always goes through the full round + confirmation; test assertions must read the one line that carries the value; spike code is a FACT source, never a design source ([[BDR-115]], [[LRN-205]], [[LRN-206]]).
|
||||||
|
|
||||||
|
## EVAL-041 — model-router W1-C: the plan was the whole risk, two confirmations were needed
|
||||||
|
- **Date**: 2026-10-09
|
||||||
|
- **Output checked**: plan `.claude/tasks/plans/2026-10-09-model-router-tiers-1237.md` r1 → r4 (absolute tiers, breaker, derived phases), written with every engine fact in hand.
|
||||||
|
- **Method**: 3 blind opus challengers (simplicity FATAL(6), correctness FATAL(11), robustness FATAL(11)) → r2; confirmation FATAL(8) with a NEW BLOCKER introduced by r2 (per-step engine-fallback detection climbing the backoff) → r3; second confirmation FATAL(4) with a NEW BLOCKER introduced by r3 (`lastPlan` reset vs kept) → r4; executor DONE first pass; verifier ECARTS ×2 on texts (3 gaps) + hardening (4 items) → CONFORME; security PASS.
|
||||||
|
- **Anomaly**: 6 BLOCKER + 29 MAJOR over four revisions, each confirmation found a flaw my own fix had introduced; the doctrine cap (one confirmation) would have shipped r2 with a 5-hour false outage. The executor never needed a re-dispatch for logic: all later rounds were text truthfulness and hardening.
|
||||||
|
- **Action**: when a revision REMOVES or REPLACES a mechanism, re-challenge once more (state the deviation); keep plan sections additive with an explicit precedence line (r4 > r3 > r2) so executors and verifiers read one law; name superseded clauses of prior contracts in the Disposition. Links [[EVAL-040]], [[LRN-207]], [[BDR-115]].
|
||||||
|
|
||||||
|
|||||||
@@ -1765,3 +1765,15 @@ Rule: when editing a doctrine file under structure locks, grep the test's lock s
|
|||||||
- **Context**: model-router tests. Kit `$` is `EngineCall<E> = (e: Args<E>)`: `command.run` needs `origin` + `presentation`, `prompt.submit` needs `wait` + `origin`, `agent.spawn` needs `tool_use_id, description, provider, parentModel, background, fork`; the test file is type-checked with the hooks (tsc include). `session.start` does NOT fire at load → every test boots with a bottom `on('session.start')` + `$.session.start({ cwd, surface: null, isInteractive: false })`. A hook calling `next` hits the kit's bottom which throws unless the test registered one (`on('agent.spawn', ($, e) => ({ model: e.model, agentId: 'a1' }))`). `$.turn.step` returns a stream: drain with `for await` then await `.result` (awaiting `.result` alone runs no hook). No fs/network/process: defaults path only. Assert on the ONE line that carries the value (a `show()` listing every phase always contains every id and level).
|
- **Context**: model-router tests. Kit `$` is `EngineCall<E> = (e: Args<E>)`: `command.run` needs `origin` + `presentation`, `prompt.submit` needs `wait` + `origin`, `agent.spawn` needs `tool_use_id, description, provider, parentModel, background, fork`; the test file is type-checked with the hooks (tsc include). `session.start` does NOT fire at load → every test boots with a bottom `on('session.start')` + `$.session.start({ cwd, surface: null, isInteractive: false })`. A hook calling `next` hits the kit's bottom which throws unless the test registered one (`on('agent.spawn', ($, e) => ({ model: e.model, agentId: 'a1' }))`). `$.turn.step` returns a stream: drain with `for await` then await `.result` (awaiting `.result` alone runs no hook). No fs/network/process: defaults path only. Assert on the ONE line that carries the value (a `show()` listing every phase always contains every id and level).
|
||||||
- **Apply**: write the boot helper first, type every input from the declarations, never relax a test to dodge a type. Links [[BDR-115]].
|
- **Apply**: write the boot helper first, type every input from the declarations, never relax a test to dodge a type. Links [[BDR-115]].
|
||||||
|
|
||||||
|
## LRN-207 — Model availability = typed API errors (`classic.StopFailure` rate_limit|overloaded|billing_error|model_not_found) + `PostModelSwitch auto`; never the turn-end reason, never `rateLimits`; sticky state and breaker target = two fields
|
||||||
|
- **Context**: model-router W1-C. `turn.complete reason: error` covers context-limit and network errors → a breaker fed by it marked fable down 15 min on a context overflow (challenge r1). `rateLimits` kinds are account windows (five_hour, seven_day, spend_limit), not per model. A per-step "engine fallback detection" (`e.model` ≠ `$.session.model()`) re-marked the session model at EVERY step → backoff climbed to the 5 h cap in one turn (confirmation r2). One field used both as sticky cur and as breaker target contradicted itself (reset at turn end vs kept).
|
||||||
|
- **Apply**: feed a breaker only from typed error kinds that name unavailability; mark per EPISODE (idempotent while down), backoff 15→300 min, `model_not_found` until reload; a user `/model` clears, `/clear` keeps (account-wide). Keep `turnModel` (sticky, reset per turn) and `lastPlan` (breaker target, kept) separate; unrouted steps pass `e.model` verbatim so the engine's own fallback is respected. Links [[BDR-115]], [[LRN-204]].
|
||||||
|
|
||||||
|
## LRN-208 — Security-auditor brief = scope only; iteration history and claimed closures violate the blind-scan contract
|
||||||
|
- **Context**: W1-C gate 2026-10-09. I sent "previous PASS with 6 MEDIUM, closed: …" in the brief. Auditor: "The auditor contract says that is never sent and must be ignored … Please do not send it next time" (it scanned blind anyway). Same rule as the verifier (lib/verify-secure-loop.md: fresh, no history).
|
||||||
|
- **Apply**: SCOPE + a neutral CONTEXT of what the code does; never prior verdicts, closures or accepted residuals. Residuals live in TODO, the auditor rediscovers them (that is the point). Links [[LRN-083]].
|
||||||
|
|
||||||
|
## LRN-209 — Scope the test run to the diff; one full `make test` before the merge; `make test` now names the red suite; GNU make rc 2 on a failed recipe
|
||||||
|
- **Context**: 2026-10-09 the pre-merge check took 454 s = full `make test` (48 suites) + a per-suite re-run to NAME the red one (aggregate rc, permanent env red design-tool-gate). Three more full passes earlier that day for mods/-only diffs. Hotfix efdd491: the recipe prints `FAIL <suite>` + `all suites green` / `<n> suite(s) red: …`. My oracle expected rc 1; GNU make returns 2 when a recipe line fails.
|
||||||
|
- **Apply**: a diff confined to one component runs that component's suite + the doctrine census; the full suite runs ONCE before `gitflow finish`; read the FAIL lines, never re-run per suite. Oracles on `make` test `[ $rc -ne 0 ]`, not `-eq 1`. Links [[LRN-173]].
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user