chore(memory): BDR-115 + LRN-205/206 + EVAL-040 — model-router architecture, tool output schema, plugin test kit, challenge value
This commit is contained in:
@@ -132,6 +132,7 @@ rules:
|
||||
| BDR-108 | 2026-09-29 | Effort round: level on every skill next to its model pin (3 repo + 25 vendored via `lib/effort-pins.txt` re-applied after the LAST vendoring step of install + resync), design stack ONE level (high), model pins stay tier aliases: quality/price trade-off = tier × effort, never version | accepted |
|
||||
| BDR-109 | 2026-09-30 | Higgsfield pack: npm CLI `latest` + 8 upstream skills git-cloned into gitignored `skills-external/higgsfield-*`, OFF by default, in no profile; two toggles (`higgsfield` = allowlist of 7 media skills, `higgsfield-websites` = landing-page aid, never website create/deploy/publish); CLI presence by probe; routing on explicit ask | accepted |
|
||||
| BDR-110 | 2026-10-06 | Shell portability doctrine: native userland on macOS AND Linux, no Homebrew GNU tools on PATH; `lib/tests/portability-census.test.sh` locks deterministic GNU-only idioms | accepted |
|
||||
| BDR-115 | 2026-10-08 | model-router mod: pin = entry default, sub-tasks route finer; one writer per axis; full ids from the mod table; built-ins-only agents table until frontmatter pins go; state in closure | accepted |
|
||||
|
||||
---
|
||||
|
||||
@@ -1393,3 +1394,11 @@ Branch feature/user-writing-web-rules, UNMERGED (human gate).
|
||||
- **Alternatives rejected**: regen via `install-hook` (writes a LOCAL hooks-path entry) or `global-hooks` (writes the GLOBAL config when the value is missing — it was, [[LRN-200]]) → `emit-hook > file` only; temp file for the verb's stderr in a hook (fail-open on a full TMPDIR, predictable path) → `2>&1` capture ([[LRN-202]]); classification before the token cap (13 s flood) → cap first ([[LRN-201]]); shell check of the release version string (interpolation sink) → by reading.
|
||||
- **Gates**: D1 3 lenses + confirm FATAL(1) (global-config write) → emit-hook; feater; GATE 0 MET; verifier CONFORME; security PASS. D2 3 lenses + confirm FATAL(3) (escape alternative, same-quote-inside, `bare=$one` when unparsed); feater; GATE 0 MET; verifier CONFORME; security BLOCK(1) cap-after-fork → fixed (20k tokens 0.13 s) → CONFORME + PASS. D3 3 lenses + confirm CONCERNS(3); feater; CONFORME + PASS.
|
||||
- **Refs**: contracts/plans `2026-10-07-manual-push-failclosed-d1-1522`, `…-guard-residuals-d2-1526`, `…-prose-d3-1530`; commits 472cccb, 3c59333, 64ca0f8 (feature/manual-push-mode, UNMERGED). Supersedes the "fail-open on invalid value" line of [[BDR-111]]. Links [[BDR-112]], [[BDR-113]], [[LRN-114]], [[LRN-196]]. Residuals: TODO "post-run-D residuals" (soft_deny names only `false`; stale `.githooks/` in onboarded repos until reconcile).
|
||||
|
||||
## BDR-115 — model-router mod: pin = entry default, sub-tasks route finer; one writer per axis; full ids; built-ins-only table until pins go [accepted] (2026-10-08)
|
||||
- **Decision**: one function-hooks mod (`mods/model-router/`) routes model + effort per request. Rules: (1) a skill/agent pin = DEFAULT route of the run, never ceiling/floor; inside the run every sub-task routes to its phase, declared (`route` tool, `/route`, `ultrathink`) or derived (skill load, agent dispatch). (2) ONE writer per axis: `Skill(effort-*)` answered by the mod WITHOUT loading the skill (no pairing rule, no doublon); every write lands on the CALLING loop; explicit Agent params frozen per loop for the whole run; agent model written once at spawn. (3) hooks write FULL ids from `config.models` ([[LRN-203]]). (4) wave 1 agents table = built-ins only (Explore sonnet/medium, Plan opus/xhigh); repo agents keep frontmatter as single writer until wave 2 moves pins into the table and deletes frontmatter + shifters + effort-pins + model-gate. (5) precedence: user `/route` sticky > latest turn route (one slot, last writer wins; non-effort skill load resets it except a prompt rule) > session. (6) main-loop model switch behind `mainModelSwitch` (default off) + context-window guard ([[LRN-204]]). (7) state in `register` closure, defaults cloned, `/route off` kill switch, config validated before merge. Load: `CLAUDE_CODE_PLUGIN_DIRS` in settings env via link.sh (wave 1-B).
|
||||
- **Why**: user 2026-10-08: existing pins + `effort-*` shifters built for this goal with older tools; some pins forced (one level for a skill doing many things). Mods expose `turn.step` model/effort rewrite + `agent.spawn` + `tool.call` answers = cleaner, unified, automatic, works headless too (BDR-107 gap).
|
||||
- **Alternatives rejected**: agents table copying the 21 frontmatter pins in wave 1 (third source of truth, silent override of a frontmatter edit); `scope: agents` / `/route agents` bulk lever (no requirement, sixth precedence tier); `model` param on the model-facing tool (typo → LRN-203 404 class); per-step agent model rewrite (fights engine fallback, beat explicit params); `userConfig` (one config file instead); marketplace install (live symlink repo model); `source`-dependent skill-load reset (kept stale shifts).
|
||||
- **Gates**: plan r1 → r3 through 3 challengers + 1 confirmation (2 BLOCKER + 10 MAJOR closed, [[EVAL-040]]); feater DONE first pass; GATE 0 MET; verifier CONFORME 6/6; security PASS + hardening round (criteria 7-11).
|
||||
- **Refs**: plan `.claude/tasks/plans/2026-10-08-model-router-mod.md`, contract `.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md`, [[BDR-107]], [[BDR-108]], [[BLK-029]], [[LRN-205]], [[LRN-206]].
|
||||
|
||||
|
||||
@@ -60,6 +60,7 @@ rules:
|
||||
| EVAL-037 | 2026-09-28 | correction of EVAL-035/036 counts: transcript records are per content block; deduped by message.id → main-loop thinking share 99.9%, thinking share of weighted cost 5.6%, sonnet think/msg 26→0.2, A/B requests 9→8 | conclusions hold (sharper: main-loop thinking 96.6%→99.9%, weighted-cost thinking corrected 8.4%→5.6%); effort-audit.py dedupes from a3b479e+ |
|
||||
| EVAL-038 | 2026-09-29 | correction of EVAL-037: 94 % of sub-agent usage records carry no `output_tokens_details` (Fable subs at xhigh read 0 thinking, impossible with always-on thinking) → sub-agent thinking UNMEASURED, not ≈0; main loop 100 % counted; weighted-cost split (61/39) still holds | `effort-audit.py` prints coverage + CAVEAT; cite the cost split only; agent effort pins stay unmeasured; a tier move on a price argument = judgment, not figure |
|
||||
| EVAL-039 | 2026-09-30 | ship-feature run higgsfield-pack: plan dry-run in scratch → 0 executor failure on 7 tasks; challenge found 7 MAJOR I missed; floor-guard caught 2 shellcheck suppressions of mine; final review found README/code gap | keep |
|
||||
| EVAL-040 | 2026-10-08 | model-router w1a plan: 3 challengers + 1 confirmation found 2 BLOCKER + 14 MAJOR on a plan judged closed; executor then passed every gate first time | keep the round, never dispatch a mod plan without it |
|
||||
|
||||
---
|
||||
|
||||
@@ -370,3 +371,11 @@ Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itse
|
||||
- **Method**: plan code dry-run in a scratch copy before the gate (suite per stage 0/5→5/0, 6/8→14/0, 14/1→15/0, 15/1→16/0, 4 mutation tests); 3 challengers + 1 confirmation; SDD per-task reviews; GATE 0/1/2 twice; final review on opus.
|
||||
- **Anomaly**: my first plan was green in dry-run and still wrong on 7 MAJOR points (shim vs binary, unbounded toggle probe, denylist membership, vacuous fixtures, askpass prompt): a dry-run proves the code does what I wrote, not that I wrote the right thing. Floor-guard flagged 2 `shellcheck disable=SC2016` I added to keep "shellcheck clean" green. Final review found the README promised drift reporting that the enabled state never reached. doc-syncer patch hit a shape escalation because I filed a script-comment edit under MINOR doc. One oracle of mine was shape-bound ([[LRN-188]]).
|
||||
- **Action**: keep the pre-gate dry-run (0 executor failure, 1 fix round in 7 tasks) AND the challenge (orthogonal finds); never silence a linter to satisfy a criterion, rewrite the line; doc patch plans carry public-doc paths only, script comments go as code commits.
|
||||
|
||||
## EVAL-040 — model-router w1a: the challenge round caught what the author could not see
|
||||
- **Date**: 2026-10-08
|
||||
- **Output checked**: plan `.claude/tasks/plans/2026-10-08-model-router-w1a-1533.md` r1, written after a successful spike with every harness fact in hand.
|
||||
- **Method**: 3 blind opus challengers (simplicity CONCERNS(3), robustness CONCERNS(6), correctness FATAL(8)) + 1 confirmation (CONCERNS(4)); every BLOCKER/MAJOR closed by a named plan change (r2, r3); then feater, GATE 0, verifier, security.
|
||||
- **Anomaly**: r1 carried 2 BLOCKER (explicit Agent `model` overridden at every step; Skill bridge answer shape refused by the output schema → doublon kept) + 10 MAJOR, all invisible to me: spike levers carried over as design (`agentsNext`, per-step model rewrite), a table copying 21 pins = third source of truth, writes on main from sub-agent loops. Confirmation found 4 more MAJOR (explicit params vs in-agent writes, `e.wait`, vacuous test assertions, haiku effort). Executor then DONE first pass, verifier CONFORME 6/6, security PASS: the plan was the whole risk.
|
||||
- **Action**: a mod plan always goes through the full round + confirmation; test assertions must read the one line that carries the value; spike code is a FACT source, never a design source ([[BDR-115]], [[LRN-205]], [[LRN-206]]).
|
||||
|
||||
|
||||
@@ -1757,3 +1757,11 @@ Rule: when editing a doctrine file under structure locks, grep the test's lock s
|
||||
- **Context**: spike 2026-10-08, fable → `claude-sonnet-5-5`/low for 3 steps on ~260k context, then back. Conversation intact (tools, results, thinking blocks from another model in history: no error). First sonnet step cache_read 0 (full 260k billed), next steps 237k cached; return to fable step read 263k cached.
|
||||
- **Apply**: switch the main loop only for spans long enough to amortize one uncached read of the whole context (many mechanical steps), never per tool call; short mechanical work → small-context haiku sub-agent. Haiku 4.5 window 200k: a long main loop cannot go to haiku at all. Flag off by default in model-router. Links [[LRN-203]], [[BDR-107]].
|
||||
|
||||
## LRN-205 — A hook answering `tool.call` in place of a built-in tool must return that tool's OUTPUT schema shape; a string result is refused and the tool runs anyway
|
||||
- **Context**: model-router Skill bridge, plan r1: `{ result: '<string>' }` for `Skill(effort-*)`. Challenger: Skill has output schema `{ success, commandName, status?, … }` (claude-code-tools index.d.ts ~5139); core validates a hook's answer against it (claude-code index.d.ts ~12641), wrong shape = hook skipped = skill loads = the exact doublon the bridge exists to remove. Fix: `{ result: { success: true, commandName: e.skill, status: 'inline' }, context: ['…'] }`; text for the model goes in `context`, never in `result`.
|
||||
- **Apply**: before answering any `tool.call` without `next`, grep the tool's RESULT type in claude-code-tools and mirror it; put model-facing prose in `context`. Links [[BDR-115]], [[LRN-203]].
|
||||
|
||||
## LRN-206 — `claude plugin test` kit facts (2.1.294): nothing fires at load, inputs are the FULL event, a bottom hook is mandatory under every `next`, `turn.step` streams
|
||||
- **Context**: model-router tests. Kit `$` is `EngineCall<E> = (e: Args<E>)`: `command.run` needs `origin` + `presentation`, `prompt.submit` needs `wait` + `origin`, `agent.spawn` needs `tool_use_id, description, provider, parentModel, background, fork`; the test file is type-checked with the hooks (tsc include). `session.start` does NOT fire at load → every test boots with a bottom `on('session.start')` + `$.session.start({ cwd, surface: null, isInteractive: false })`. A hook calling `next` hits the kit's bottom which throws unless the test registered one (`on('agent.spawn', ($, e) => ({ model: e.model, agentId: 'a1' }))`). `$.turn.step` returns a stream: drain with `for await` then await `.result` (awaiting `.result` alone runs no hook). No fs/network/process: defaults path only. Assert on the ONE line that carries the value (a `show()` listing every phase always contains every id and level).
|
||||
- **Apply**: write the boot helper first, type every input from the declarations, never relax a test to dodge a type. Links [[BDR-115]].
|
||||
|
||||
|
||||
@@ -43,5 +43,24 @@ Q (r2): `scope: agents` / `/route agents` / A: dropped (challenge r1, no require
|
||||
Q (r2): agents table in wave 1 / A: built-ins only (Explore, Plan), matched for the engine provider; repo agents keep their frontmatter as the single writer until wave 2. [orchestrator — single source of truth]
|
||||
Q (r2): `/route off` / A: added as the session kill switch (every hook passes through). [orchestrator — robustness]
|
||||
|
||||
## HARDENING ROUND (security gate 2026-10-08, user go "oui durcis") — criteria 7-11
|
||||
7. `/route` is user-only: `command.run` answers `{ text: 'route: user-only command' }` without acting when `e.origin.kind !== 'composer'`; a test proves it (origin `{ kind: 'plugin', name: 'x' }` or the kit's non-composer origin → text contains `user-only`, state unchanged).
|
||||
CHECK: cd mods/model-router && grep -q "origin.kind" hooks/register.ts && grep -q "user-only" hooks/register.ts && grep -q "user-only" hooks/register.test.ts && echo ORIGIN-OK
|
||||
EXPECT: ORIGIN-OK
|
||||
EVIDENCE: pending
|
||||
8. An in-agent `route` call or skill table entry never changes that agent's MODEL: the `Loop` type has no `model` axis and no `explicitModel` flag, `turn.step` on an agent loop rewrites `effort` only, the route tool's description says "effort only; the model of a sub-agent is fixed at spawn". A test proves it: a route tool call carrying `agentId: 'a1'` with `phase: 'judge'` followed by a `turn.step` for `a1` leaves `model` as given and sets `effort` to `xhigh`.
|
||||
CHECK: cd mods/model-router && ! grep -qE "loop\.model|explicitModel|spawnModel" hooks/register.ts && grep -q "fixed at spawn" hooks/register.ts && echo NO-AGENT-MODEL-OK
|
||||
EXPECT: NO-AGENT-MODEL-OK
|
||||
EVIDENCE: pending
|
||||
9. Config hardening: a prompt rule pattern longer than 200 chars is dropped (logged); `re.test` runs on at most the first 4096 chars of the prompt; phase keys must match `^[a-z][a-z0-9_-]{0,31}$` (others dropped, logged); the override file is refused above 65536 bytes (logged, defaults kept); the route tool schema carries `additionalProperties: false`; the Skill hook checks `typeof e.skill === 'string'`.
|
||||
CHECK: cd mods/model-router && grep -q "additionalProperties: false" hooks/register.ts && grep -qE "\[a-z\]\[a-z0-9_-\]\{0,31\}" hooks/register.ts && grep -qE "4096|4_096" hooks/register.ts && grep -qE "65536|65_536|64 \* 1024" hooks/register.ts && grep -qE "200" hooks/register.ts && grep -q "typeof e.skill === 'string'" hooks/register.ts && echo CONFIG-HARDEN-OK
|
||||
EXPECT: CONFIG-HARDEN-OK
|
||||
EVIDENCE: pending
|
||||
10. Visible fail-open: every `.catch` logs once per session per hook (`$.ui.log('model-router: <hook> failed (<kind>): routing skipped for this event')`, a `warned: Set<string>` in the state) before passing through or answering; the three silent config drops (non-object top level, wrong-typed table, non-array `prompt`) log a line.
|
||||
CHECK: cd mods/model-router && [ "$(grep -c '\.catch(' hooks/register.ts)" -ge 12 ] && grep -q "warned" hooks/register.ts && grep -q "routing skipped" hooks/register.ts && echo CATCH-LOG-OK
|
||||
EXPECT: CATCH-LOG-OK
|
||||
EVIDENCE: pending
|
||||
11. Post-`next` bookkeeping (loop tracking and logs after `await next(...)` in `agent.spawn` and the Skill hook) runs inside its own try/catch so a logging failure can never make the `.catch` re-run `next`. Judged by reading, with criteria 1-6 still MET (validate, tsc, tests ≥ 13, style, AC6 minus the removed model axis).
|
||||
|
||||
## FILE SCOPE
|
||||
mods/model-router/.claude-plugin/plugin.json · mods/model-router/hooks/hooks.json · mods/model-router/hooks/register.ts · mods/model-router/hooks/register.test.ts
|
||||
|
||||
Reference in New Issue
Block a user