chore(memory): BDR-115 + LRN-205/206 + EVAL-040 — model-router architecture, tool output schema, plugin test kit, challenge value

This commit is contained in:
bchanot
2026-10-08 16:38:08 +02:00
parent 6dc2d748fc
commit 64702d50ea
4 changed files with 45 additions and 0 deletions
+8
View File
@@ -1757,3 +1757,11 @@ Rule: when editing a doctrine file under structure locks, grep the test's lock s
- **Context**: spike 2026-10-08, fable → `claude-sonnet-5-5`/low for 3 steps on ~260k context, then back. Conversation intact (tools, results, thinking blocks from another model in history: no error). First sonnet step cache_read 0 (full 260k billed), next steps 237k cached; return to fable step read 263k cached.
- **Apply**: switch the main loop only for spans long enough to amortize one uncached read of the whole context (many mechanical steps), never per tool call; short mechanical work → small-context haiku sub-agent. Haiku 4.5 window 200k: a long main loop cannot go to haiku at all. Flag off by default in model-router. Links [[LRN-203]], [[BDR-107]].
## LRN-205 — A hook answering `tool.call` in place of a built-in tool must return that tool's OUTPUT schema shape; a string result is refused and the tool runs anyway
- **Context**: model-router Skill bridge, plan r1: `{ result: '<string>' }` for `Skill(effort-*)`. Challenger: Skill has output schema `{ success, commandName, status?, … }` (claude-code-tools index.d.ts ~5139); core validates a hook's answer against it (claude-code index.d.ts ~12641), wrong shape = hook skipped = skill loads = the exact doublon the bridge exists to remove. Fix: `{ result: { success: true, commandName: e.skill, status: 'inline' }, context: ['…'] }`; text for the model goes in `context`, never in `result`.
- **Apply**: before answering any `tool.call` without `next`, grep the tool's RESULT type in claude-code-tools and mirror it; put model-facing prose in `context`. Links [[BDR-115]], [[LRN-203]].
## LRN-206 — `claude plugin test` kit facts (2.1.294): nothing fires at load, inputs are the FULL event, a bottom hook is mandatory under every `next`, `turn.step` streams
- **Context**: model-router tests. Kit `$` is `EngineCall<E> = (e: Args<E>)`: `command.run` needs `origin` + `presentation`, `prompt.submit` needs `wait` + `origin`, `agent.spawn` needs `tool_use_id, description, provider, parentModel, background, fork`; the test file is type-checked with the hooks (tsc include). `session.start` does NOT fire at load → every test boots with a bottom `on('session.start')` + `$.session.start({ cwd, surface: null, isInteractive: false })`. A hook calling `next` hits the kit's bottom which throws unless the test registered one (`on('agent.spawn', ($, e) => ({ model: e.model, agentId: 'a1' }))`). `$.turn.step` returns a stream: drain with `for await` then await `.result` (awaiting `.result` alone runs no hook). No fs/network/process: defaults path only. Assert on the ONE line that carries the value (a `show()` listing every phase always contains every id and level).
- **Apply**: write the boot helper first, type every input from the declarations, never relax a test to dodge a type. Links [[BDR-115]].