Files
claude/lib/effort-shift.md
T
bchanot 22455c051f feat(model-router): wave 3-A — first-use route confirmation, decision memory, project exceptions
routing.json (tracked, reached through the plugin directory) is now the
single source of the phase table and of every skill/agent row, plus the
decisions: confirmed rows/phases, changed rows (from/to) and projects
exceptions keyed by a normalized git remote (credentials never stored, no
machine paths). First use of a rowed typed skill, a rowed agent spawn or a
main-loop phase opens the engine's dialog (Later / Keep / two alternative
phases; Other = a phase name); a change asks Everywhere or This project
only. One dialog at a time, never in headless, never inside an agent,
never written by the model: only a dialog answer or /route ask writes,
serialized, size-capped, never creating the file. Layers: routing.json <
~/.claude/model-router.json; the project tree is never read. /route
pending, /route ask on|off. Census reads rows and phases from the file and
tolerates a user-changed row (WARN). Kit suite 86 → 190 tests.

Contract .claude/tasks/contracts/2026-10-10-model-router-w3a-confirm-1201.md,
plan r4: 3 lenses + 1 confirmation, feater + 4 rounds, GATE 0 MET,
verifier CONFORME then re-verify after security, security BLOCK(1) fixed
then PASS. Live: T2 dialogs answered by the user from the hot-loaded mod.
2026-10-11 11:48:25 +02:00

63 lines
2.9 KiB
Markdown

# Route doctrine — phase-level model and effort on the main loop (BDR-107)
Shared include, companion of `lib/model-gate.md`: the gate fixes WHICH model
reflects, the model-router mod (`mods/model-router`) fixes HOW HARD each
phase thinks. Rungs: low (fix a line, run a script) · medium (day-to-day) ·
high (refactor, resisting bug) · xhigh (architecture, audit before
validation) · max (stuck error, judged need).
## The tool
`mcp__model-router__route` (params `phase` | `effort` | `clear`). It is a
deferred tool: when not loaded, run
`ToolSearch("select:mcp__model-router__route")` once per session. A route
applies from the next request on, paired with
another tool call or not (pairing only saves a request). The answer always
names the id and effort main runs on. A skill with a row routes itself on
load; a skill without one changes nothing, the last ROWED skill wins.
## Wiring points
1. Dispatch span starts → `route(phase="orchestrate")`, sent with the
dispatch.
2. Reflection resumes (challenge synthesis, verdict, plan revision) →
`route(phase="reflect")` or `"plan"` per the skill's own level; the line
before every `lib/challenge-plan.md` call.
3. Bookkeeping tail (memory commit, doc commit) → `route(phase="apply")`.
4. Escalation → `route(phase="escalate")`: verify-secure loop caps and
ship-feature STEP 4b. Not automatic: the challenge fail-safe and "gone
WRONG → STOP"; their STOP text names the levers below.
5. Built-in judgment dispatch (`general-purpose` `model="opus"`, `model:
"fable"` skill-runners) → explicit `effort=` on the Agent call (`xhigh`
for opus reviewers, `high` for fable runners). A main route never
reaches a child. Typed agents run on their row, never on a shift.
6. After a prose gate that ends the turn, the resumed reflection phase
starts with its own route call.
## Run slot and levers
A best-tier skill row survives the end of the turn (a run spans prose
gates); `/route clear`, `/route off` and a user `/model` drop it. Levers for a
relaunch: `ultrathink` in the prompt (turn floor) or `/route effort=max`
(sticky, `/route clear` after).
Builtin `/effort` is NOT a lever inside a run: rows and routes outrank it.
## Limits
- A skill typed while a background agent is live routes only through the
typed marker (unverified live 2026-10-10).
- Headless (`-p`, SDK) runs the hooks, so routing works there too.
- Mod off: typed agents fall back to their `model:`/`effort:` frontmatter.
- First use of a row asks once (Keep, Later or another phase, then
Everywhere or this project only); the answer is kept in
`mods/model-router/routing.json`. `/route pending` lists what is still
unconfirmed, `/route ask off|on` toggles the dialog.
Measure the split any time: `python3 ~/.claude/lib/effort-audit.py`
(thinking/output/cache tokens per scope, model and effort).
## Never
- A route inside a dispatched agent: its row rules there.
- Max is for diagnosis, not for retrying the same fix harder.