Merge feature/model-router-w2 into develop

This commit is contained in:
bchanot
2026-10-10 11:43:59 +02:00
56 changed files with 1668 additions and 1012 deletions
+1 -1
View File
@@ -1402,4 +1402,4 @@ Branch feature/user-writing-web-rules, UNMERGED (human gate).
- **Gates**: plan r1 → r3 through 3 challengers + 1 confirmation (2 BLOCKER + 10 MAJOR closed, [[EVAL-040]]); feater DONE first pass; GATE 0 MET; verifier CONFORME 6/6; security PASS + hardening round (criteria 7-11). - **Gates**: plan r1 → r3 through 3 challengers + 1 confirmation (2 BLOCKER + 10 MAJOR closed, [[EVAL-040]]); feater DONE first pass; GATE 0 MET; verifier CONFORME 6/6; security PASS + hardening round (criteria 7-11).
- **Refs**: plan `.claude/tasks/plans/2026-10-08-model-router-mod.md`, contract `.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md`, [[BDR-107]], [[BDR-108]], [[BLK-029]], [[LRN-205]], [[LRN-206]]. - **Refs**: plan `.claude/tasks/plans/2026-10-08-model-router-mod.md`, contract `.claude/tasks/contracts/2026-10-08-model-router-w1a-1533.md`, [[BDR-107]], [[BDR-108]], [[BLK-029]], [[LRN-205]], [[LRN-206]].
- **Amendment (2026-10-09, user decisions 2026-10-08 evening)**: (a) LOAD supersedes the "Load:" line: tracked relative symlink `skills/model-router` → `../mods/model-router`, loaded in place as `model-router@skills-dir` wherever link.sh links `~/.claude/skills`; `CLAUDE_CODE_PLUGIN_DIRS` dropped (absolute path, settings `env` has no `$HOME` expansion, settings.json tracked), local marketplace dropped (`add` writes an absolute path into settings.json). Proven by fresh-process `claude plugin list --json`. (b) PRECEDENCE amended: `ultrathink` and a typed `/effort-<l>` are the main turn's DEFAULT and MINIMUM (floor slot `turnFloor`): sticky `/route` effort > turn route effort > floor > engine, then floored; per axis; mid-turn prompt floors the running turn and the next (`wait` ignored). Rationale: user "un choix explicite bat la phase déduite"; a pure floor made `/effort-low` a no-op (challenge finding). (c) Per-machine kill switch `"enabled": false` in the untracked `~/.claude/model-router.json` (survives `/clear`, a failed reload keeps the previous config); `enabledPlugins` would dirty the tracked settings.json on every machine. (d) Hardening: `/route` composer-only; agent loops effort-only (model fixed at spawn); config caps; typed slash attested at `prompt.submit`. Commits 346d6ae, 1ff608a, 6430ac6; contracts `2026-10-08-model-router-floor-1835`, `2026-10-08-model-router-wiring-1835`; residuals parked in TODO. - **Amendment (2026-10-09, user decisions 2026-10-08 evening)**: (a) LOAD supersedes the "Load:" line: tracked relative symlink `skills/model-router` → `../mods/model-router`, loaded in place as `model-router@skills-dir` wherever link.sh links `~/.claude/skills`; `CLAUDE_CODE_PLUGIN_DIRS` dropped (absolute path, settings `env` has no `$HOME` expansion, settings.json tracked), local marketplace dropped (`add` writes an absolute path into settings.json). Proven by fresh-process `claude plugin list --json`. (b) PRECEDENCE amended: `ultrathink` and a typed `/effort-<l>` are the main turn's DEFAULT and MINIMUM (floor slot `turnFloor`): sticky `/route` effort > turn route effort > floor > engine, then floored; per axis; mid-turn prompt floors the running turn and the next (`wait` ignored). Rationale: user "un choix explicite bat la phase déduite"; a pure floor made `/effort-low` a no-op (challenge finding). (c) Per-machine kill switch `"enabled": false` in the untracked `~/.claude/model-router.json` (survives `/clear`, a failed reload keeps the previous config); `enabledPlugins` would dirty the tracked settings.json on every machine. (d) Hardening: `/route` composer-only; agent loops effort-only (model fixed at spawn); config caps; typed slash attested at `prompt.submit`. Commits 346d6ae, 1ff608a, 6430ac6; contracts `2026-10-08-model-router-floor-1835`, `2026-10-08-model-router-wiring-1835`; residuals parked in TODO.
- **Amendment 2 (2026-10-10, wave 2 closed, commits bb56f3e + 1f2d33b)**: (a) rule 4 closed: rows for every repo skill (56) + agent (21) + Explore/Plan, PHASES by role (plan/reflect/orchestrate/escalate best · judge big · implement/write/verify/explore/apply work · mechanical cheap; `write` work/high + `apply` work/low added). User: pins "deleted or reworked", not copied → rows are the live source, tracked `model:`/`effort:` frontmatter KEPT as off-state floor, census-locked equal to rows (`lib/tests/effort-routing.test.sh`). Robustness BLOCKER closed: mod off → agents would inherit parent model. (b) Rule 5 amended: unrowed skill load changes nothing; best-tier skill row lives in `runMain` slot surviving turn end (precedence userMain > turnMain > runMain > floor > engine), dropped by `/route clear`, `/route off`, user `/model`, typed non-best skill; turn writers (route tool, prompt rules) never touch it. (c) Agents: model written at spawn WITHIN tier, upward only (never below frontmatter alias); explicit Agent params win; project-defined agent (agent.offer source projectSettings|localSettings) skipped. (d) Typed slash: name-bound marker at prompt.submit (composer|sdk|bridge) + pending slot + idle fallback (no live/spawning loop). (e) Shifters `effort-*`, `effort-pins.*`, `model-check.sh` DELETED; orchestrators call `mcp__model-router__route` per phase; gate witness = route answer id, remedy `/route on`; builtin `/effort` not a lever inside a run (levers `ultrathink`, `/route effort=max`). (f) SemVer: typed `/effort-*` removal = breaking → next release 3.0.0. Plan `.claude/tasks/plans/2026-10-09-model-router-w2-1546.md` r4, contracts `2026-10-09-model-router-w2a-1546`, `2026-10-10-model-router-w2b-1045`. Links [[LRN-210]], [[LRN-211]], [[EVAL-042]].
+7
View File
@@ -62,6 +62,7 @@ rules:
| EVAL-039 | 2026-09-30 | ship-feature run higgsfield-pack: plan dry-run in scratch → 0 executor failure on 7 tasks; challenge found 7 MAJOR I missed; floor-guard caught 2 shellcheck suppressions of mine; final review found README/code gap | keep | | EVAL-039 | 2026-09-30 | ship-feature run higgsfield-pack: plan dry-run in scratch → 0 executor failure on 7 tasks; challenge found 7 MAJOR I missed; floor-guard caught 2 shellcheck suppressions of mine; final review found README/code gap | keep |
| EVAL-040 | 2026-10-08 | model-router w1a plan: 3 challengers + 1 confirmation found 2 BLOCKER + 14 MAJOR on a plan judged closed; executor then passed every gate first time | keep the round, never dispatch a mod plan without it | | EVAL-040 | 2026-10-08 | model-router w1a plan: 3 challengers + 1 confirmation found 2 BLOCKER + 14 MAJOR on a plan judged closed; executor then passed every gate first time | keep the round, never dispatch a mod plan without it |
| EVAL-041 | 2026-10-09 | model-router W1-C plan: 3 lenses FATAL (4 BLOCKER + 20 MAJOR) then 2 confirmations each FATAL with a NEW BLOCKER in my own revision; executor DONE first pass, 3 short text/hardening rounds | one confirmation is not enough when a revision removes a whole mechanism; the plan carried the risk, the code almost none | | EVAL-041 | 2026-10-09 | model-router W1-C plan: 3 lenses FATAL (4 BLOCKER + 20 MAJOR) then 2 confirmations each FATAL with a NEW BLOCKER in my own revision; executor DONE first pass, 3 short text/hardening rounds | one confirmation is not enough when a revision removes a whole mechanism; the plan carried the risk, the code almost none |
| EVAL-042 | 2026-10-10 | model-router W2 plan r1 → r4: 3 lenses (1 BLOCKER), 2 confirmations (1 BLOCKER then 0); W2-A verifier 3× ECARTS on coverage clauses only, W2-B ECARTS(7) → CONFORME; gate A→B read from engine records | second confirmation paid again (BLOCKER on my own r2 slot); compound coverage criterion = endless ECARTS; engine jsonl replaces the live log |
--- ---
@@ -387,3 +388,9 @@ Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itse
- **Anomaly**: 6 BLOCKER + 29 MAJOR over four revisions, each confirmation found a flaw my own fix had introduced; the doctrine cap (one confirmation) would have shipped r2 with a 5-hour false outage. The executor never needed a re-dispatch for logic: all later rounds were text truthfulness and hardening. - **Anomaly**: 6 BLOCKER + 29 MAJOR over four revisions, each confirmation found a flaw my own fix had introduced; the doctrine cap (one confirmation) would have shipped r2 with a 5-hour false outage. The executor never needed a re-dispatch for logic: all later rounds were text truthfulness and hardening.
- **Action**: when a revision REMOVES or REPLACES a mechanism, re-challenge once more (state the deviation); keep plan sections additive with an explicit precedence line (r4 > r3 > r2) so executors and verifiers read one law; name superseded clauses of prior contracts in the Disposition. Links [[EVAL-040]], [[LRN-207]], [[BDR-115]]. - **Action**: when a revision REMOVES or REPLACES a mechanism, re-challenge once more (state the deviation); keep plan sections additive with an explicit precedence line (r4 > r3 > r2) so executors and verifiers read one law; name superseded clauses of prior contracts in the Disposition. Links [[EVAL-040]], [[LRN-207]], [[BDR-115]].
## EVAL-042 — model-router W2: challenge value held on the second confirmation; coverage-shaped criterion burned three verifiers
- **Date**: 2026-10-10
- **Output checked**: plan `.claude/tasks/plans/2026-10-09-model-router-w2-1546.md` r1 → r4; W2-A diff (register.ts, 88 kit tests); W2-B diff (43 files, 140-check census).
- **Method**: 3 blind lenses (simplicity CONCERNS(5), correctness CONCERNS(10), robustness FATAL(8) → BLOCKER: mod off = agents inherit parent model); confirmation 1 robustness FATAL(8) → BLOCKER introduced by r2 (route calls wiped the sticky slot); confirmation 2 correctness CONCERNS(3), no BLOCKER → r4. W2-A: feater DONE + 4 rounds (1 internal decision, 3 coverage), GATE 0 MET, verifier ECARTS(3)/(1)/(1) all coverage, user accepted at cap; security PASS. W2-B: feater DONE first pass, verifier ECARTS(7) (5 FLOOR items = planned deletions needing a CLARIFICATIONS line, 2 prose, 1 scope add) → CONFORME 10/10; security PASS; full `make test` once (env red only). Gate A→B: 3/4 probes answered from engine jsonl, probe 4 (typed skill with a live agent) unobserved, recorded as a limit.
- **Anomaly**: the user's rework answer ("delete or rework, not copy") changed the design mid-plan; r2's own fix carried a BLOCKER again (as in EVAL-041). Coverage criterion: 3 verifiers, 0 defects. FLOOR guard needs the test deletions named in CLARIFICATIONS, not only in criteria.
- **Action**: keep the "second confirmation after a mechanism change" rule; write coverage criteria one clause each (LRN-210); when a plan deletes tests, write the authorizing CLARIFICATIONS line BEFORE the first verifier. Links [[EVAL-041]], [[LRN-210]], [[LRN-211]], [[BDR-115]].
+4
View File
@@ -590,3 +590,7 @@ rules:
- User go 'ok merge': full make test green (except env red design-tool-gate), shellcheck clean → gitflow finish feature model-router-mod → develop abbdf79 (22 commits: waves 0, 1-A, 1-B1, 1-B2, 1-C + docs + registries). Manual mode: develop NOT pushed, the user publishes by hand from the terminal. Branch removed locally. User will /clear before wave 2. - User go 'ok merge': full make test green (except env red design-tool-gate), shellcheck clean → gitflow finish feature model-router-mod → develop abbdf79 (22 commits: waves 0, 1-A, 1-B1, 1-B2, 1-C + docs + registries). Manual mode: develop NOT pushed, the user publishes by hand from the terminal. Branch removed locally. User will /clear before wave 2.
- /hotfix make-test-names-red-suites (user: '9 min pour un merge?'): measured from transcript timestamps, the merge took <20 s; 454 s went to make test run TWICE (full + per-suite sweep to name the red suite, because the aggregate rc is silent). Fix: Makefile test prints FAIL <suite> + summary, rc unchanged (GNU make returns 2 on a failed recipe; my first oracle expected 1). hotfixer DONE, GATE 0 MET 2/2, security PASS. Branch bugfix/make-test-names-red-suites UNMERGED (human signal). Method note: a mods/-only diff needs only the mods suite + doctrine census, the full run once before merge. - /hotfix make-test-names-red-suites (user: '9 min pour un merge?'): measured from transcript timestamps, the merge took <20 s; 454 s went to make test run TWICE (full + per-suite sweep to name the red suite, because the aggregate rc is silent). Fix: Makefile test prints FAIL <suite> + summary, rc unchanged (GNU make returns 2 on a failed recipe; my first oracle expected 1). hotfixer DONE, GATE 0 MET 2/2, security PASS. Branch bugfix/make-test-names-red-suites UNMERGED (human signal). Method note: a mods/-only diff needs only the mods suite + doctrine census, the full run once before merge.
- User go 'all, merge le, et je clear': LRN-207/208/209 + EVAL-041 written (ca96458); gitflow finish bugfix/make-test-names-red-suites → develop 5e0e5c0, branch removed. develop 28 commits ahead of origin, NOT pushed (manual mode, user publishes). Only settings.json dirty (user's own /model change). Session closes; next: wave 2 of model-router (TODO W2 line). - User go 'all, merge le, et je clear': LRN-207/208/209 + EVAL-041 written (ca96458); gitflow finish bugfix/make-test-names-red-suites → develop 5e0e5c0, branch removed. develop 28 commits ahead of origin, NOT pushed (manual mode, user publishes). Only settings.json dirty (user's own /model change). Session closes; next: wave 2 of model-router (TODO W2 line).
- model-router W2-A landed (bb56f3e, feature/model-router-w2). User go 'lance la vague 2' + 4 pass-B answers: shifters deleted, pins reworked into phase rows (not copied), phases declared, slim gate. Plan r1→r4: 3 lenses (1 BLOCKER: mod off → agents inherit parent model → frontmatter kept as off-state floor), confirmation 1 FATAL (BLOCKER: route calls wiped the sticky slot → separate runMain), confirmation 2 CONCERNS (deviation: 2 confirmations, stated). Feater DONE + 4 gap rounds (A5 in-agent decision, 6+3+3 coverage tests, every one mutation-proven). GATE 0 MET; verifier ECARTS(3)/(1)/(1) all on coverage of criterion 3 clauses, never code → diagnosis at max: compound criterion → user accepted at the cap. Security PASS (2 MEDIUM pre-existing). Kit 58→88. Next: user /reload-plugins + live probe, then W2-B. Lesson for the next contract: one coverage clause per criterion, not twelve.
## 2026-10-10
- model-router W2-B landed (1f2d33b) + gate A→B. User /reload-plugins + typed /status; I read the engine jsonl instead of the UI log: main low on typed /status, analyzer opus/xhigh at step 0 → ordering + row-over-frontmatter proven; probe 4 unobserved (user away), limit recorded. Route tool checked live (`mcp__model-router__route`, deferred → ToolSearch once; answer names the id); session itself routed through it (reflect → orchestrate). W2-B: contract 10 criteria, feater DONE first pass (43 files), verifier ECARTS(7): 5 FLOOR items (planned test deletions → CLARIFICATIONS line), `/route on` wording, false CLAUDE.global.md sentence, plan-challenger scope add → CONFORME 10/10; I folded 3 observations by hand (SDD sonnet implementers `effort="medium"`, run-slot droppers, 80-col). Security PASS (push-guard false positive on a grep pattern; it also bit my registry heredoc). Full make test green (env red only). Doc audit SIGNIFICANT → user: apply all P1-P8, SemVer BREAKING → 3.0.0, registries all → BDR-115 amendment 2, LRN-210/211, EVAL-042. Pending: doc commit, memory commit, user merge decision (gitflow finish, human signal), publish by hand.
+7
View File
@@ -1777,3 +1777,10 @@ Rule: when editing a doctrine file under structure locks, grep the test's lock s
- **Context**: 2026-10-09 the pre-merge check took 454 s = full `make test` (48 suites) + a per-suite re-run to NAME the red one (aggregate rc, permanent env red design-tool-gate). Three more full passes earlier that day for mods/-only diffs. Hotfix efdd491: the recipe prints `FAIL <suite>` + `all suites green` / `<n> suite(s) red: …`. My oracle expected rc 1; GNU make returns 2 when a recipe line fails. - **Context**: 2026-10-09 the pre-merge check took 454 s = full `make test` (48 suites) + a per-suite re-run to NAME the red one (aggregate rc, permanent env red design-tool-gate). Three more full passes earlier that day for mods/-only diffs. Hotfix efdd491: the recipe prints `FAIL <suite>` + `all suites green` / `<n> suite(s) red: …`. My oracle expected rc 1; GNU make returns 2 when a recipe line fails.
- **Apply**: a diff confined to one component runs that component's suite + the doctrine census; the full suite runs ONCE before `gitflow finish`; read the FAIL lines, never re-run per suite. Oracles on `make` test `[ $rc -ne 0 ]`, not `-eq 1`. Links [[LRN-173]]. - **Apply**: a diff confined to one component runs that component's suite + the doctrine census; the full suite runs ONCE before `gitflow finish`; read the FAIL lines, never re-run per suite. Oracles on `make` test `[ $rc -ne 0 ]`, not `-eq 1`. Links [[LRN-173]].
## LRN-210 — One coverage clause per contract criterion; a compound "covered by kit tests" criterion yields ECARTS forever with zero defects
- **Context**: W2-A contract criterion 3 bundled ~12 clauses under one "Covered by kit tests". Three fresh verifiers: ECARTS(3), (1), (1), each a NEW untested guard, code judged correct 3×. Cap hit, diagnosis at max, user accepted. 4 feater rounds added 30 tests (58 → 88), every one mutation-proven.
- **Apply**: split coverage criteria: one clause = one criterion with its own test name; or make the oracle a mutation script (`CHECK:` flips the guard, expects a red test). Verifier reads clauses literally: write only what one test can prove. Links [[LRN-209]], [[EVAL-042]].
## LRN-211 — Typed-slash routing: name-bound marker + idle fallback; run slot separate from turn routes; engine records are the live oracle
- **Context**: mod needs "user typed /feat" from `skill.prompt`, which carries no origin. `prompt.submit` sees the raw `/name` first (composer|sdk|bridge): store the NAME (not a boolean; a bare flag leaked to the next preload), pending slot when mid-turn, consume only on the matching `skill.prompt`; fallback = no live/spawning loop AND allowed origin. Sticky run route in the SAME slot as turn routes was wiped by the first `route()` call (confirmation BLOCKER) → separate `runMain`, best-tier rows only (work/cheap rows leak low effort across turns). Live facts read from `~/.claude/projects/<repo>/<session>.jsonl` (+ `subagents/agent-*.jsonl`): `effort` + `message.model` per step. Typed `/status` → main low; analyzer step 0 opus/xhigh with frontmatter high (spawn bookkeeping precedes step 0; row beats frontmatter).
- **Apply**: hook-side user-intent markers: bind to a name, add a pending slot for mid-turn, keep an ordering-independent fallback. Two lifetimes = two slots, never one slot with a source tag. Verify engine behaviour in the transcript jsonl, not in `$.ui.log`. Links [[BDR-115]], [[LRN-206]].
+7 -2
View File
@@ -17,8 +17,13 @@ migration of shifters/pins/model-gate in wave 2 after proof; names model-router
- [ ] W1-C accepted-by-design (security 2026-10-09, MEDIUM, not coded around): (1) the model itself can raise the main loop to fable for the rest of a turn through the `route` tool (plan/reflect/escalate/judge) or a background dispatch (derived orchestrate = best tier) — bounded by the turn and `upgradeMaxTokens`, no sticky route is model-callable; (2) `ultrathink` and the default keyword rules now mean "best tier" (fable) at the phase's effort, so an incidental keyword in a pasted composer prompt costs a fable turn (origin composer only). Residuals: `PostModelSwitch` `auto` covers "other programmatic change" (a healthy model could be marked 15 min); strikes never decay inside a session; a `[1m]` variant's `model_not_found` marks the base model until reload; a `models` alias added by the override without a `fallback` key stays unranked (`withEveryAlias` not applied to the default chain) so `leaveDown` skips the upgrade gates for it; `canonical` prefix match has no segment boundary (`claude-sonnet-5-50` would map to sonnet); classifier deferred to W2 - [ ] W1-C accepted-by-design (security 2026-10-09, MEDIUM, not coded around): (1) the model itself can raise the main loop to fable for the rest of a turn through the `route` tool (plan/reflect/escalate/judge) or a background dispatch (derived orchestrate = best tier) — bounded by the turn and `upgradeMaxTokens`, no sticky route is model-callable; (2) `ultrathink` and the default keyword rules now mean "best tier" (fable) at the phase's effort, so an incidental keyword in a pasted composer prompt costs a fable turn (origin composer only). Residuals: `PostModelSwitch` `auto` covers "other programmatic change" (a healthy model could be marked 15 min); strikes never decay inside a session; a `[1m]` variant's `model_not_found` marks the base model until reload; a `models` alias added by the override without a `fallback` key stays unranked (`withEveryAlias` not applied to the default chain) so `leaveDown` skips the upgrade gates for it; `canonical` prefix match has no segment boundary (`claude-sonnet-5-50` would map to sonnet); classifier deferred to W2
- [x] W1-C live verification part 1 (2026-10-09 after `/reload-plugins`, skills-dir copy loaded, 12 hooks): derived orchestrate on a real background Explore (main high → medium while it ran → high after its end); Explore on sonnet/medium; `route plan` → next request xhigh (engine record); `Skill(effort-low)` bridged → next request low (engine record); `/route show` resolves all 10 phases to full ids, `down: none`; steps arrive as bare `claude-fable-5-1` (no `[1m]`) so the suffix carry never fires on this session - [x] W1-C live verification part 1 (2026-10-09 after `/reload-plugins`, skills-dir copy loaded, 12 hooks): derived orchestrate on a real background Explore (main high → medium while it ran → high after its end); Explore on sonnet/medium; `route plan` → next request xhigh (engine record); `Skill(effort-low)` bridged → next request low (engine record); `/route show` resolves all 10 phases to full ids, `down: none`; steps arrive as bare `claude-fable-5-1` (no `[1m]`) so the suffix carry never fires on this session
- [ ] W1-C live verification part 2 (needs a real incident): StopFailure vs turn.complete order; `PostModelSwitch` `auto` semantics and its `from_model` after a router upgrade; whether a router rewrite raises `auto`; `[1m]` carry validity on opus (only on a session whose steps carry it) - [ ] W1-C live verification part 2 (needs a real incident): StopFailure vs turn.complete order; `PostModelSwitch` `auto` semantics and its `from_model` after a router upgrade; whether a router rewrite raises `auto`; `[1m]` carry validity on opus (only on a session whose steps carry it)
- [ ] W2 migration (after wave 1 proven in daily use): 15 skills `Skill(effort-*)` → `route` tool calls; remove `skills/effort-*`, `lib/effort-pins.txt/.sh`, install/update steps, `effort:` frontmatter on skills and agents; repo agents into the mod's `agents` table (verify the spawn/first-step ordering first); `lib/model-gate.md` + `lib/model-check.sh` → mod rule; census tests repointed; `lib/effort-shift.md` rewritten; docs - [x] W2-A mod (2026-10-09, bb56f3e on feature/model-router-w2, contract `2026-10-09-model-router-w2a-1546`, plan r4 `2026-10-09-model-router-w2-1546`): user decisions — effort-* skills deleted (W2-B), pins REWORKED not deleted (rows = phases by role; frontmatter `model:`/`effort:` kept as census-locked off-state floor after a robustness BLOCKER), orchestrators declare phases, slim model gate. Phases `write` work/high + `apply` work/low; 56 skill rows, 21 agent rows; agents' model at spawn within tier, upward only, project-defined agents skipped (agent.offer); typed slash → name-bound marker (composer|sdk|bridge) + pending slot + idle fallback; best-tier rows in a `runMain` slot surviving turn end; unrowed skill leaves the route; route answer always names the id; `null` override rows. 3 lenses + 2 confirmations (1 BLOCKER each round closed), feater + 4 gap rounds, GATE 0 MET, verifier 3× ECARTS on coverage clauses only → user accepted at the cap, security PASS (2 MEDIUM pre-existing: first-load kill switch fails open, ReDoS size-bounded). Kit 58 → 88 tests. Doc-sync skipped for A (mod-only, docs at B).
- [ ] W2 migration: 15 skills off `Skill(effort-*)`, remove shifters + effort-pins + model-gate, census repointed, docs - [x] W2 gate A→B (2026-10-10, engine records): typed `/status` → main low then medium during the dispatch; analyzer opus/xhigh from step 0 (frontmatter high) = spawn bookkeeping first, row beats frontmatter; status-reporter haiku step 0. Probe 3 dropped (gate always calls route). Probe 4 (typed skill while a background agent is live: marker vs fallback) NOT observed → limit in `lib/effort-shift.md`; re-check in the records whenever the user types a skill with an agent alive.
- [x] W2-B repo migration (2026-10-10, 1f2d33b, contract `2026-10-10-model-router-w2b-1045`): 15 citers → `mcp__model-router__route` per phase, `effort=` on built-in judgment dispatches + SDD sonnet implementers (`effort="medium"`), `lib/effort-shift.md` + `lib/model-gate.md` rewritten (witness = route answer, `/route on`), deletions (effort-*, effort-pins.*, model-check.sh, tests, installer blocks), bridge removed from the mod, census rewritten (140 checks, drift lock rows ↔ frontmatter), analyzer xhigh, CLAUDE.global.md Design-work line. Verifier ECARTS(7) → CONFORME 10/10, security PASS, full make test green (design-tool-gate env red). Docs P1-P8 user-approved; SemVer: breaking → next release 3.0.0 (MIGRATION section at release).
- [ ] (was) W2 gate A→B (user): `/reload-plugins`, then the live probe of plan § Gate (4 points: typed `/status` marker vs fallback in the verbose log; a real rowed spawn line + `step 0 agent` effort = spawn/first-step ordering; sonnet session typed `/feat` self-check + route answer; probe 1 again with a background agent alive). `typed-marker` never seen → W2-B blocked, A4 re-planned.
- [x] (was) W2-B repo migration (plan § W2-B B0-B7 + STEP 6/7): bridge removal, `lib/effort-shift.md` rewrite, 15 citers → `route`, `effort=` on opus general-purpose dispatches, slim `lib/model-gate.md`, delete `model-check.sh` + `effort-pins.*` + their tests + install/update blocks, delete `skills/effort-*`, analyzer `effort: xhigh`, census rewrite (drift lock rows ↔ frontmatter), docs + registries (BDR-115 amendment, LRN typed-slash/run slot, EVAL)
- [ ] W2-A residuals (security, accepted): first-load failure of the override activates the router despite `enabled:false` (fix = treat a failed first load as off); ReDoS on a self-authored prompt pattern (size-bounded); error text in the local log; model alias keys unvalidated (PHASE_KEY would do); `offers` map uncapped; `__proto__`/`constructor` override keys untested. Known limits: `skillCalls`/`spawning` counters are global; offers keyed by name only; builtin `/effort` is not a lever inside a run.
- [ ] W2 residuals: probe 4 unobserved (above); the push-guard hook denies a read-only grep (or a heredoc) whose TEXT contains the push verb next to `git` (security-auditor + orchestrator 2026-10-10, false positives, reworded); ~27 loose "sonnet pin"/"opus pin" shorthand sites kept (true by census); README config key list omits `tiers`/`fallback`/`cooldownMinutes`/`mainUpgrade`/`upgradeMaxTokens` (doc audit item 6, pre-existing); MIGRATION.md "Upgrading to 3.0.0" at release time; `skillCalls`/`spawning` counters global; offers keyed by name.
- [ ] W3 optional: step heuristics, haiku classifier, quota-aware downgrade, A/B - [ ] W3 optional: step heuristics, haiku classifier, quota-aware downgrade, A/B
## 2026-09-30 — Higgsfield pack: CLI + skills in the install process, off by default (feature/higgsfield-pack) ## 2026-09-30 — Higgsfield pack: CLI + skills in the install process, off by default (feature/higgsfield-pack)
@@ -0,0 +1,44 @@
# CONTRACT — model-router-w2a
- date: 2026-10-09 | flow: feat | branch: feature/model-router-w2 (to start off develop)
- status: active
- wave: W2-A (mod, plan r2). W2-B (repo migration + docs/registries as its STEP 6/7) gets its own contract after the live probe.
## REQUEST (verbatim — IMMUTABLE)
tu peux lancer la vague 2
(TODO W2 line: 15 skills `Skill(effort-*)` → `route` tool calls; remove `skills/effort-*`, `lib/effort-pins.txt/.sh`, install/update steps, `effort:` frontmatter on skills and agents; repo agents into the mod's `agents` table (verify the spawn/first-step ordering first); `lib/model-gate.md` + `lib/model-check.sh` → mod rule; census tests repointed; `lib/effort-shift.md` rewritten; docs)
## CLARIFICATIONS
(pass A: none — outcome, scope and constraints come from the TODO W2 line + plan `2026-10-08-model-router-mod.md` § Wave 2)
Q: keep the five `effort-*` skills as typed floor levers? / A: "supprimer, j'utilise les commandes builtin si j'ai besoin" → deleted; the mod drops its `Skill(effort-*)` bridge and typed-slash floor code; `ultrathink` floor stays [gated 2026-10-09]
Q: table rows = bare level next to phases, or phases only? / A: "le pin était là car pas encore de système de routage fiable; maintenant qu'il y en a un, soit on le supprime, soit tu les remanies pour que ça fonctionne encore mieux" → reworked: rows are PHASES by role (plan/reflect/implement/write/verify/judge/mechanical), no level rows; agents routed on both axes by the mod while on (model written at spawn from the phase tier, upward only); the tracked `model:`/`effort:` frontmatter STAYS as the census-locked off-state floor (plan r3, confirmation 1/8) [gated 2026-10-09]
Q: orchestrators declare phases or effort only? / A: phases [gated 2026-10-09]
Q: (mid-run, feater NEED-DECISION, CLASS internal) a skill WITHOUT a row loaded INSIDE a sub-agent: keep the old reset of that loop's effort, or leave the loop untouched like main (A5)? / A: leave the agent loop untouched too — one rule for both loops, an unrowed helper skill (find-docs) never wipes an agent row's effort [gated 2026-10-09]
Q: GATE 1 cap (3 × ECARTS on test coverage of criterion 3 clauses, code correct ×3; last rows landed by a feater, unverified) / A: user "Accepter et commiter W2-A" — accepted on the three reports + 88-test kit suite; diagnosis: compound coverage criterion, not the code [gated 2026-10-09]
Q: model gate fate? / A: slim include (self-check + STOP when the mod could not raise), `lib/model-check.sh` + test deleted [gated 2026-10-09]
## ACCEPTANCE CRITERIA
1. The mod's kit suite is green, including the new tests of criteria 3-5.
CHECK: cd mods/model-router && out="$(claude plugin test . 2>&1)" && printf '%s\n' "$out" | grep -qE '[0-9]+ pass' && printf '%s\n' "$out" | grep -qE '(^|[^0-9])0 fail' && ! printf '%s\n' "$out" | grep -qE '[1-9][0-9]* fail' && echo W2A-TESTS-GREEN
EXPECT: W2A-TESTS-GREEN
EVIDENCE: MET exit=0 marker-found :: W2A-TESTS-GREEN
2. `DEFAULT_CONFIG.phases` has `write` at work/high and a new `apply` at work/low; `DEFAULT_CONFIG.skills` and `.agents` carry exactly the phase rows of plan r2 § Row tables (56 skill rows, 21 agent rows + Explore/Plan), every value a phase key; the "Built-ins only" comment is gone.
CHECK: bash .claude/tasks/contracts/w2a-table-census.sh mods/model-router/hooks/register.ts
EXPECT: W2A-TABLE-COMPLETE
EVIDENCE: MET exit=0 marker-found :: W2A-TABLE-COMPLETE
3. A user-typed slash of a rowed skill routes main to its row (name-bound marker from `prompt.submit`, origins composer|sdk|bridge, pending slot mid-turn, or no live/spawning sub-agent); a preload inside a live sub-agent without the marker leaves main; a skill WITHOUT a row leaves the route in force; a best-tier skill row lives in a separate `runMain` slot that survives `turn.complete` and later `route` calls, and is dropped by `/route clear`, `/route off`, a user-typed non-best rowed skill, or a user `/model`; the `Skill(effort-*)` bridge is not sticky. Covered by kit tests.
4. Any rowed agent spawn (whatever `provider.plugin`) with `e.model` undefined gets the row's model written at spawn: full id resolved within the tier only and ranked ≥ the tier head (upward only); otherwise no write + one deduped log line; the row's effort applies on its loop from its first step unless the call carried `effort`; `fork`/`workflow` spawns and agents whose `agent.offer` source is a project definition are untouched; an override `agents: { name: null }` drops a row; the `route` tool answer always names the id the next main step runs on. Covered by kit tests.
5. Typed-floor code removed: no `slashEffort`, `guardedSlash`, `'slash'` Source, `typed /effort-` floor word in `register.ts`; `EFFORT_SKILL`/`effortBridge` (tool.call bridge) and the `ultrathink` floor kept until W2-B; `/route` unchanged (existing tests still green).
CHECK: ! grep -qE "slashEffort|guardedSlash|'slash'|typed /effort-" mods/model-router/hooks/register.ts && grep -q 'turnFloor' mods/model-router/hooks/register.ts && grep -q 'effortBridge' mods/model-router/hooks/register.ts && echo W2A-DEAD-CODE-GONE
EXPECT: W2A-DEAD-CODE-GONE
EVIDENCE: MET exit=0 marker-found :: W2A-DEAD-CODE-GONE
7. `claude plugin validate` passes.
CHECK: cd mods/model-router && claude plugin validate . 2>&1 | grep -qi 'valid' && echo W2A-VALID
EXPECT: W2A-VALID
EVIDENCE: MET exit=0 marker-found :: W2A-VALID
6. `make test suite=lib/tests/mods.test.sh` green.
CHECK: make test suite=lib/tests/mods.test.sh 2>&1 | tail -5 | grep -q 'mods' && ! make test suite=lib/tests/mods.test.sh 2>&1 | grep -q '^FAIL' && echo W2A-MODS-SUITE
EXPECT: W2A-MODS-SUITE
EVIDENCE: MET exit=0 marker-found :: W2A-MODS-SUITE
## FILE SCOPE
mods/model-router/hooks/register.ts, mods/model-router/hooks/register.test.ts, .claude/tasks/contracts/w2a-table-census.sh (oracle, written by the orchestrator)
@@ -0,0 +1,55 @@
# CONTRACT — model-router-w2b
- date: 2026-10-10 | flow: feat | branch: feature/model-router-w2
- status: active
- wave: W2-B (repo migration), plan r4 `.claude/tasks/plans/2026-10-09-model-router-w2-1546.md` § D1-D4, § Row tables, § W2-B B0-B7. Follows W2-A (bb56f3e) and the A→B gate.
## REQUEST (verbatim — IMMUTABLE)
tu peux lancer la vague 2
(TODO W2 line: 15 skills `Skill(effort-*)` → `route` tool calls; remove `skills/effort-*`, `lib/effort-pins.txt/.sh`, install/update steps, `effort:` frontmatter on skills and agents; repo agents into the mod's `agents` table (verify the spawn/first-step ordering first); `lib/model-gate.md` + `lib/model-check.sh` → mod rule; census tests repointed; `lib/effort-shift.md` rewritten; docs)
## CLARIFICATIONS
Q: pass B 2026-10-09 (see the W2-A contract) / A: effort-* deleted; pins reworked into phase rows with the tracked frontmatter kept as census-locked off-state floor; orchestrators declare phases; slim model gate [gated 2026-10-09]
Q: A→B gate evidence (engine records 2026-10-10) / A: typed `/status` → main effort low then medium during the dispatch (row applied to a typed slash); analyzer spawn → opus/xhigh from step 0 while its frontmatter still said high (spawn bookkeeping precedes the first step; the row beats the frontmatter); status-reporter haiku from step 0. Probe 4 (typed skill while a background agent is live: marker vs fallback) NOT observed → recorded as a limit in `lib/effort-shift.md`, re-checked later in the records; not blocking [gated 2026-10-10]
Q: slim gate design / A: the include ALWAYS calls the route tool at entry (`mcp__model-router__route` phase = the skill's row) and reads the id in the answer; no self-check wording dependency; the tool is deferred for the model → `ToolSearch("select:mcp__model-router__route")` once per session when it is not loaded [gated 2026-10-10]
Q: FLOOR items (verifier 1): deletion of `lib/tests/effort-pins.test.sh` and `lib/tests/model-check.test.sh`, removal of the two `Skill(effort-low)` bridge kit tests, the from-scratch rewrite of `lib/tests/effort-routing.test.sh` (has/ok/ko helpers, its `# shellcheck disable=SC2015,SC2016` directive carried over from HEAD) / A: all authorized — they are the plan's B0 and B4 steps, accepted at pass B (shifters, pins and witness deleted with their tests) [gated 2026-10-10]
Q: scope addition (verifier 1, criterion 9a): `agents/plan-challenger.md:110-112` twins `lib/challenge-plan.md:46` ("`model: opus`-pinned in its frontmatter, session-independent") / A: added to FILE SCOPE, same rewording (row + off-state floor); the ~20 loose "sonnet pin"/"opus pin" shorthand sites stay (true: alias == row tier by census) [gated 2026-10-10, orchestrator — minor, user may veto]
## ACCEPTANCE CRITERIA
1. No `Skill(effort-` call and no `EFFORT SHIFTS:` header line remains in tracked skills, agents and lib (the mod and CHANGELOG excepted).
CHECK: ! git grep -nE 'Skill\(effort-|EFFORT SHIFTS:' -- skills agents lib ':!mods' >/dev/null && echo W2B-NO-SHIFTER-CITER
EXPECT: W2B-NO-SHIFTER-CITER
EVIDENCE: MET exit=0 marker-found :: W2B-NO-SHIFTER-CITER
2. Deleted: `skills/effort-{low,medium,high,xhigh,max}`, `lib/effort-pins.txt`, `lib/effort-pins.sh`, `lib/model-check.sh`, `lib/tests/effort-pins.test.sh`, `lib/tests/model-check.test.sh`; no tracked reference to `effort-pins`, `apply_effort_pins` or `model-check.sh` outside CHANGELOG.md, `.claude/`, README.md and USAGE.md (the two docs are the STEP 6 doc-sync targets [gated 2026-10-10]).
CHECK: for p in skills/effort-low skills/effort-medium skills/effort-high skills/effort-xhigh skills/effort-max lib/effort-pins.txt lib/effort-pins.sh lib/model-check.sh lib/tests/effort-pins.test.sh lib/tests/model-check.test.sh; do [ ! -e "$p" ] || { echo "still here: $p"; exit 1; }; done && ! git grep -nE 'effort-pins|apply_effort_pins|model-check\.sh' -- . ':!CHANGELOG.md' ':!.claude' ':!README.md' ':!USAGE.md' >/dev/null && echo W2B-DELETIONS-DONE
EXPECT: W2B-DELETIONS-DONE
EVIDENCE: MET exit=0 marker-found :: W2B-DELETIONS-DONE
3. Mod: the `Skill(effort-*)` bridge is gone (`EFFORT_SKILL`, `effortBridge` absent), the kit suite and the mods suite are green.
CHECK: ! grep -qE 'EFFORT_SKILL|effortBridge' mods/model-router/hooks/register.ts && cd mods/model-router && out="$(claude plugin test . 2>&1)" && printf '%s\n' "$out" | grep -qE '[0-9]+ pass' && ! printf '%s\n' "$out" | grep -qE '[1-9][0-9]* fail' && cd ../.. && make test suite=lib/tests/mods.test.sh 2>&1 | grep -q 'all suites green' && echo W2B-MOD-GREEN
EXPECT: W2B-MOD-GREEN
EVIDENCE: MET exit=0 marker-found :: W2B-MOD-GREEN
4. `lib/tests/effort-routing.test.sh` is rewritten as the wave-2 census (drift lock rows ↔ frontmatter parsed from `register.ts`, no `effort:`-less routed skill, D3 wiring markers, no shifter citers) and is green; `lib/tests/model-routing.test.sh` and `lib/tests/higgsfield.test.sh` are green.
CHECK: grep -q 'register.ts' lib/tests/effort-routing.test.sh && bash lib/tests/effort-routing.test.sh >/dev/null 2>&1 && bash lib/tests/model-routing.test.sh >/dev/null 2>&1 && bash lib/tests/higgsfield.test.sh >/dev/null 2>&1 && echo W2B-CENSUS-GREEN
EXPECT: W2B-CENSUS-GREEN
EVIDENCE: MET exit=0 marker-found :: W2B-CENSUS-GREEN
5. Wiring follows plan D3 at every former shifter site: dispatch span → `mcp__model-router__route` phase orchestrate; own level high → reflect, xhigh → plan; bookkeeping tail → apply; escalation (verify-secure-loop ×3, ship-feature error recovery) → escalate; judgment dispatches of built-ins (`general-purpose` `model="opus"` in ship-feature, init-project, onboard, tour; `model: "fable"` skill-runners in client-handover-writer) carry an explicit `effort=` param instead of a shift; the "turn reset" re-assert sites become a route call at the first step of the resumed turn. Judged by reading each site against the plan.
6. `lib/model-gate.md` is the mod rule, ≤ 30 lines: always call `mcp__model-router__route` at entry with the skill's row phase, STOP unless the answer names a fable or opus id (tool missing → ToolSearch once; "is off" → STOP with the remedy); keeps the dispatch-tier table with `model: "fable"` for skill-runners; no `model-check` reference.
CHECK: [ "$(wc -l < lib/model-gate.md)" -le 30 ] && grep -q 'mcp__model-router__route' lib/model-gate.md && grep -q 'model: "fable"' lib/model-gate.md && ! grep -q 'model-check' lib/model-gate.md && echo W2B-GATE-SLIM
EXPECT: W2B-GATE-SLIM
EVIDENCE: MET exit=0 marker-found :: W2B-GATE-SLIM
7. `lib/effort-shift.md` is rewritten (≤ 60 lines) as the route doctrine: the tool name and the ToolSearch note, the D3 wiring points, explicit `effort=` for built-in judgment dispatches, the run slot and `/route clear`, levers (`ultrathink`, `/route effort=max`; builtin `/effort` is not a lever inside a run), last rowed skill wins and an unrowed one changes nothing, the probe-4 limit, headless OK, `effort-audit.py`.
CHECK: [ "$(wc -l < lib/effort-shift.md)" -le 60 ] && grep -q 'mcp__model-router__route' lib/effort-shift.md && grep -q 'ToolSearch' lib/effort-shift.md && grep -q 'effort-audit.py' lib/effort-shift.md && ! grep -q 'Skill(effort' lib/effort-shift.md && echo W2B-DOCTRINE
EXPECT: W2B-DOCTRINE
EVIDENCE: MET exit=0 marker-found :: W2B-DOCTRINE
8. Frontmatter aligned to the rows: `agents/analyzer.md` carries `effort: xhigh`; every other tracked `effort:`/`model:` value equals its row (criterion 4 census proves it); `CLAUDE.global.md` Design-work lines no longer cite `lib/effort-pins.txt` or the lone-Skill-call rule.
CHECK: grep -q '^effort: xhigh' agents/analyzer.md && ! grep -q 'effort-pins' CLAUDE.global.md && ! grep -q 'lone Skill call' CLAUDE.global.md && echo W2B-FRONTMATTER
EXPECT: W2B-FRONTMATTER
EVIDENCE: MET exit=0 marker-found :: W2B-FRONTMATTER
9. Prose sites that stated pin semantics ("sonnet by frontmatter pin", "`model: opus`-pinned in their frontmatter, session-independent", STOP texts naming `/effort-max`) are reworded to the row/off-state-floor semantics and the D4 levers. Judged by reading.
10. The suites the diff touches are green (doctrine-citers, loops-light, plan-challenger, skill-routing-census, profile-census, portability-census, gitflow-test if present) and `shellcheck *.sh hooks/*.sh lib/*.sh` is clean; the full `make test` runs once before the merge (LRN-209), the orchestrator reports it.
CHECK: for t in doctrine-citers loops-light plan-challenger skill-routing-census profile-census portability-census; do [ -f "lib/tests/$t.test.sh" ] || continue; bash "lib/tests/$t.test.sh" >/dev/null 2>&1 || { echo "red: $t"; exit 1; }; done; shellcheck ./*.sh hooks/*.sh lib/*.sh && echo W2B-SUITE-GREEN
EXPECT: W2B-SUITE-GREEN
EVIDENCE: MET exit=0 marker-found :: W2B-SUITE-GREEN
## FILE SCOPE
mods/model-router/hooks/register.ts, mods/model-router/hooks/register.test.ts; lib/effort-shift.md, lib/model-gate.md, lib/verify-secure-loop.md, lib/challenge-plan.md; skills/{ship-feature,init-project,feat,bugfix,web-validate,seo,hotfix,geo,harden,code-clean,audit-delta,onboard,tour,client-handover,analyze,commit-change}/SKILL.md; agents/client-handover-writer.md, agents/analyzer.md, agents/plan-challenger.md (prose only), agents/{commit-changer,handover-doc-writer,geo-analyzer,doc-syncer,seo-analyzer}.md (prose only); deletions listed in criterion 2; install-plugins.sh, update-all.sh; lib/tests/{effort-routing,model-routing,higgsfield}.test.sh; CLAUDE.global.md (Design-work lines only).
+56
View File
@@ -0,0 +1,56 @@
#!/usr/bin/env bash
# GATE 0 oracle, contract 2026-10-09-model-router-w2a: DEFAULT_CONFIG.skills
# and .agents in register.ts hold exactly the plan's phase rows (plan
# 2026-10-09-model-router-w2 § Row tables), every value a declared phase.
# Prints W2A-TABLE-COMPLETE only when every assertion passes.
set -u
F="${1:?register.ts path}"
python3 - "$F" <<'PY'
import re, sys
src = open(sys.argv[1]).read()
m = re.search(r'const DEFAULT_CONFIG: Config = \{(.*?)\n\}\n', src, re.S)
if not m: print("no DEFAULT_CONFIG block"); sys.exit(1)
cfg = m.group(1)
def block(name):
b = re.search(r'\n ' + name + r': \{([^\n]*)\},', cfg) \
or re.search(r'\n ' + name + r': \{(.*?)\n \}', cfg, re.S)
if not b: print(f"no {name} block"); sys.exit(1)
rows = re.findall(r"(?:^|[{,])\s*'?([A-Za-z0-9_-]+)'?:\s*'([a-z]+)'", b.group(1), re.M)
return dict(rows)
phases = set(
re.findall(r"^\s*([a-z]+): \{ tier:", re.search(r'\n phases: \{(.*?)\n \}', cfg, re.S).group(1), re.M))
want_skills = {}
for ph, names in {
'plan': 'ship-feature init-project onboard tour audit-delta analyze code-clean client-handover brainstorming writing-plans requesting-code-review 21st-ui-review',
'reflect': 'feat hotfix bugfix refactor web-validate harden seo geo site-motion frontend-design emil-design-eng design-motion-principles 21st-ui-build scroll-world-storytelling build-threejs-scroll-worlds scroll-scrubbed-visual-sequence scroll-scrubbed-word-reveal scroll-progress-timeline subagent-driven-development writing-skills deprecation-and-migration 21st-ai 21st-ui-explore',
'implement': 'gitflow prune-memory pdf-translate ci-cd-and-automation observability-and-instrumentation test-driven-development',
'apply': 'commit-change release-candidate doc capitalize close reconcile deploy',
'mechanical': 'status profile plugin-check skills-perso using-git-worktrees 21st-cli-use 21st-registry 21st-design-sync',
}.items():
for n in names.split(): want_skills[n] = ph
want_agents = {'Explore': 'explore', 'Plan': 'judge'}
for ph, names in {
'implement': 'feater bugfixer code-cleaner scaffolder onboarder',
'write': 'commit-changer doc-syncer handover-doc-writer refactorer',
'apply': 'hotfixer release-executor plugin-probe validator-analyzer',
'verify': 'verifier security-auditor',
'judge': 'plan-challenger plugin-advisor seo-analyzer geo-analyzer analyzer',
'mechanical': 'status-reporter',
}.items():
for n in names.split(): want_agents[n] = ph
ok = True
for name, want in (('skills', want_skills), ('agents', want_agents)):
got = block(name)
for k in sorted(set(want) | set(got)):
if want.get(k) != got.get(k):
ok = False; print(f"{name}.{k}: want {want.get(k)} got {got.get(k)}")
bad = {k: v for k, v in got.items() if v not in phases}
if bad: ok = False; print(f"{name}: undeclared phases {bad}")
if len(want_skills) != 56 or len(want_agents) != 23:
ok = False; print(f"oracle self-check: {len(want_skills)} skill rows, {len(want_agents)} agent rows")
if 'Built-ins only' in cfg: ok = False; print("stale comment 'Built-ins only'")
ph = re.search(r'\n phases: \{(.*?)\n \}', cfg, re.S).group(1)
for want in ("write: { tier: 'work', effort: 'high' }", "apply: { tier: 'work', effort: 'low' }"):
if want not in ph: ok = False; print(f"phases: missing {want}")
print("W2A-TABLE-COMPLETE" if ok else "W2A-TABLE-INCOMPLETE"); sys.exit(0 if ok else 1)
PY
@@ -0,0 +1,345 @@
# PLAN r4 — model-router wave 2 (migration), 2026-10-09
r1 → r2 after 3 lenses (simplicity CONCERNS(5), correctness CONCERNS(10),
robustness FATAL(8)); r2 → r3 after the confirmation pass (robustness
FATAL(8): 1 BLOCKER); r3 → r4 after a second confirmation (correctness
CONCERNS(3), wording-level). Every BLOCKER/MAJOR is closed by a NAMED change
(§ Challenge ledger). Mod = the router while on; the tracked frontmatter
(`model:` + `effort:` on agents, `effort:` on skills) STAYS as the
off-state floor, census-locked equal to the rows (r3).
Two sub-runs on `feature/model-router-w2`: W2-A (mod) then, after a user
`/reload-plugins` + live probe, W2-B (repo migration). Docs + registries =
W2-B's own STEP 6/7 (feat pipeline tail); W2-A runs no doc-sync (deviation,
stated: a mod-only diff has no public doc of its own until B lands).
## Decisions (pass B, user 2026-10-09) — amended by the challenge
- D1 the five `effort-*` skills are DELETED in W2-B. The mod's
`Skill(effort-*)` bridge is deleted in W2-B too (same commit as the citers,
robustness 7: edits are live on the symlinked tree). The typed `/effort-`
floor code goes in W2-A (A1). The `ultrathink` floor stays.
- D2 pins reworked, not deleted (r3): rows are PHASES by ROLE, no bare
levels; the mod routes agents on both axes while on (model written at
spawn from the tier, only UPWARD in rank; effort per step; explicit
Agent params win). The tracked frontmatter stays as the OFF-STATE FLOOR
(mod off / `enabled:false` / unloaded / a hook failing open → the engine
applies the frontmatter as today: never the parent model, never session
effort on an xhigh gate agent; r2 robustness 1 + confirmation 8). The
census locks every row equal to its frontmatter (tier head == `model:`
alias, phase effort == `effort:`), so there is one declared value, two
carriers. Deleted: the five shifters, `lib/effort-pins.*` (vendored
levels become rows; off-state = session level for them, as before
BDR-108), the witness script.
- D3 the orchestrators declare PHASES: dispatch span → `route(phase=
"orchestrate")`; own level high → `reflect`, xhigh → `plan`; bookkeeping
tail → `route(phase="apply")` (work/low); escalation → `escalate`.
Judgment dispatches of BUILT-INS (`general-purpose` with `model="opus"`
or `"fable"`) carry an explicit `effort=` param (a main route never
reaches a child; correctness 8). A best-tier skill row SURVIVES the end
of the turn in its own slot (A6 `runMain`): a run spans prose gates;
turn-scoped routes (`route` calls, prompt rules, the bridge) never
touch it.
- D4 `lib/model-gate.md` = the mod rule; the witness is the `route` tool's
own answer (correctness 7): self-check big → silent; small → call
`route(phase=<entry phase>)` and STOP unless the answer names a fable or
opus id; tool absent or "is off" → STOP with the remedy. `model-check.sh`
+ its test deleted. The route answer ALWAYS names the id the next main
step runs on (A3d). Relaunch levers in STOP texts: `ultrathink` in the
relaunch prompt (turn floor) or `/route effort=max` (sticky, `/route
clear` after). Builtin `/effort` is NOT a lever inside a run (rows and
routes rank above the engine effort; r2's engine-effort detector dropped
as unsafe, confirmation 4) — documented, named to the user.
## Phase table (A2a) — two rows added to `phases`
| phase | tier | effort | role |
| plan | best | xhigh | brainstorm, plan, architecture, audit verdict |
| reflect | best | high | diagnosis, review, contract, day-to-day orchestrator entry |
| orchestrate | best | medium | between dispatches |
| escalate | best | max | stuck, cap reached |
| judge | big | xhigh | dispatched challengers, analyzers, audits |
| implement | work | medium | code from a closed plan |
| write | work | **high** (was medium) | docs, commits, refactors, handover prose on sonnet at high (BDR-107 "high judgment on sonnet") |
| verify | work | xhigh | verifier, security-auditor |
| explore | work | medium | Explore |
| **apply** (new) | work | low | low appliers (BDR-107): small fixes, release mechanics, probes, validators; bookkeeping tail of the main loop |
| mechanical | cheap | low | listing, status, profile toggles |
## Row tables (A2b)
skills (56):
- plan: ship-feature init-project onboard tour audit-delta analyze
code-clean client-handover brainstorming writing-plans
requesting-code-review 21st-ui-review
- reflect: feat hotfix bugfix refactor web-validate harden seo geo
site-motion frontend-design emil-design-eng design-motion-principles
21st-ui-build scroll-world-storytelling build-threejs-scroll-worlds
scroll-scrubbed-visual-sequence scroll-scrubbed-word-reveal
scroll-progress-timeline subagent-driven-development writing-skills
deprecation-and-migration 21st-ai 21st-ui-explore
- implement: gitflow prune-memory pdf-translate ci-cd-and-automation
observability-and-instrumentation test-driven-development
- apply: commit-change release-candidate doc capitalize close reconcile
deploy (work tier: never haiku on main even with the switch on,
robustness 15)
- mechanical: status profile plugin-check skills-perso using-git-worktrees
21st-cli-use 21st-registry 21st-design-sync
agents (21 + 2 built-ins), model = frontmatter alias, unchanged everywhere;
`effort:` frontmatter updated where the row differs (analyzer → xhigh):
- implement (sonnet/medium): feater bugfixer code-cleaner scaffolder onboarder
- write (sonnet/high): commit-changer doc-syncer handover-doc-writer refactorer
- apply (sonnet/low): hotfixer release-executor plugin-probe validator-analyzer
- verify (sonnet/xhigh): verifier security-auditor
- judge (opus/xhigh): plan-challenger plugin-advisor seo-analyzer
geo-analyzer analyzer (the ONE level delta: high → xhigh, BDR-108 rung
"audit before validation"; named at the gate)
- mechanical (haiku/low): status-reporter
- built-ins: Explore explore, Plan judge
- no row: interviewer, client-handover-writer (inline-load), impeccable-*
(gitignored vendor output, untouched)
Deltas named at the gate: analyzer effort; best-tier raise now also reaches
refactor + the design/vendored reflect skills on a small session
(simplicity 6: the raise is the point of the mod).
## W2-A — mod (contract 2026-10-09-model-router-w2a-1546)
- [ ] A1 `register.ts`: delete the typed-floor code — `slashEffort`,
`guardedSlash`, `Source` member `'slash'`, `floorSource`'s `typed /`
branch, the State comments on the typed `/effort-<l>` floor. KEEP
`EFFORT_SKILL` + `effortBridge` (tool.call bridge, removed in W2-B)
and `turnFloor`/`higherFloor`/`floorWord` (ultrathink).
- [ ] A2 `DEFAULT_CONFIG`: phases per § Phase table (`write` high, `apply`
new); `skills` + `agents` per § Row tables; the "Built-ins only"
comment → "phase = role; one row per routed repo skill/agent (wave 2);
a project-level agent of the same name shadows its row (A3)".
- [ ] A3 spawn: (a) explicit model = `e.model !== undefined` at spawn (the
engine does not pre-fill the frontmatter; no tool_use_id map);
explicit effort as today (`explicitEffort`). (b) `spawnRoute`:
`frozen` → none; row lookup for every provider; a row is SKIPPED when
the agent's definition `source`, recorded per `subagentType` from the
`agent.offer` event, is `projectSettings` or `localSettings` (a
foreign repo's own `verifier.md`); `userSettings`, `built-in`,
`plugin` or no record → the row applies (fail-open on routing, as
W1). Known limit, stated in a comment: the record is keyed by name
only (an offer fired inside a sub-agent with another cwd overwrites it). (c) `spawnTarget` for a rowed spawn
without explicit model: resolve WITHIN the tier only, and write only
an alias ranked ≥ the tier head in `fallback` (agents move UP, never
below their frontmatter alias): `big` with opus down → fable; opus +
fable down → no write + one log line (deduped per agent+tier per
turn) "model-router: <agent> tier <t> down, frontmatter model kept".
(d) `routedText`/`mainAnswer` always name the id the next main step
runs on (`model <id> (<why>)`), the sticky/floor note APPENDED, never
substituted (the gate reads this answer).
- [ ] A4 typed slash: `typedSlash: string | null` = the first token of a
slash prompt at `prompt.submit` when `origin.kind` ∈ {composer, sdk,
bridge} (allowlist; floor/default rules stay composer-only), stored
only when it is a `cfg.skills` key; mid-turn → `pendingSlash`,
promoted at `endMainTurn`. `prompt.submit` also records
`st.promptAllowed` (origin in the allowlist) for the turn it opens
(pending slot for a mid-turn prompt, like the marker). `skill.prompt`
on main (`skillCalls === 0`): apply the row when `e.skill ===
typedSlash` (then null it) OR when `promptAllowed && loops.size === 0
&& spawning === 0` (`spawning` = a counter held from spawn-hook entry
to after `next`); otherwise `next(e)`. Verbose log names which path
fired (`typed-marker` / `typed-fallback`).
- [ ] A5 `onSkillLoad`: a skill WITHOUT a row leaves `turnMain` untouched
(find-docs / gstack / plugin skills mid-run no longer clear the
run's route; correctness 11); a rowed skill replaces it. Same rule
INSIDE a sub-agent (gated 2026-10-09, feater NEED-DECISION): an
unrowed skill leaves the agent loop's effort untouched; a rowed one
writes it (test: `feater` loop at medium, `Skill(find-docs)` with
that agentId → next step still medium).
- [ ] A6 run slot: new `runMain: Routed | null`. A rowed skill load ON
MAIN whose phase tier is `best` writes BOTH `turnMain` (source
`skill`, as today) and `runMain`; a non-best row loaded by the
model's `Skill` tool writes `turnMain` only (helper skills such as
`using-git-worktrees` inside SDD never drop the run); a USER-TYPED
rowed skill (marker path) replaces `runMain` with its row when best,
drops it when not. Precedence `userMain > turnMain > runMain > floor
> engine` (per axis, as today); `mainRoute` (statusline, spinner,
planStep) includes it; `endMainTurn` leaves it; `pushOrchestrate`
writes `turnMain` when empty exactly as today (derived orchestrate
overrides the run default for the dispatch span, like a `route
orchestrate` call); cleared by `/route clear` (text: "run slot
dropped" when one held), `/route off`, a user `/model` switch
(`PostModelSwitch` source command|picker|sdk); `route(clear=true)`
from the model clears `turnMain` only and its answer says "run
<phase> still holds". A `Skill` call inside a sub-agent never touches
it. `/route show` prints `run <phase>` when no turn route is in force.
- [ ] A7 (dropped in r3: engine-effort detector, confirmation 4).
- [ ] A8 `mergeTable` accepts `null` in the override's `skills`/`agents`
to drop a default row (robustness 4).
- [ ] A9 `register.test.ts`: delete the typed `/effort-*` floor tests;
keep the bridge test; add — typed `/feat` (prompt.submit `/feat x`,
origin composer, then skill.prompt feat) → `main: skill reflect`,
`effort high`, `[tier best]`; origin `channel`: not armed AND the
idle fallback refused (main untouched); skill.prompt `feat` with a
live sub-agent and no marker → main untouched; same with no live loop
and an allowed origin → routed; a mid-turn `/status` → pending,
applied after turn.complete; `Skill(find-docs)` (no row) after
`route plan` keeps plan — the existing test 'floor: survives a skill
load' (register.test.ts:341-348) is REWRITTEN to assert orchestrate
KEPT (A5 behavior, authorized here); run slot: `Skill(feat)`, then
`route orchestrate`, then `turn.complete` → `/route show` back to
`run reflect`; `Skill(using-git-worktrees)` (mechanical, model path)
keeps `run reflect` under a mechanical turn route; a TYPED `/status`
drops the run slot; `/route clear` drops it; `/route off` drops it;
PostModelSwitch source `command` drops it; a `Skill(feat)` with an
agentId leaves `runMain`; `Skill(effort-low)` bridge then
`turn.complete` → session defaults (not sticky); `feater` spawn (provider
`{plugin:'engine',tier:'core'}`, no model) → `started.model` = sonnet
full id and the first `turn.step` of that agentId at medium;
`plan-challenger` with opus down → fable id; opus AND fable down →
no write, one log line; explicit `model`/`effort` win; `fork: true`
untouched; a project-source `agent.offer` record → no write; route
answer text names the id with a floor in force; override `agents: {
verifier: null }` drops the row (bottom `fs`/`env` mocked like
`session.model`; if the kit refuses, the case is dropped and said so).
- Disposition: honors BDR-115 (one writer per axis while on, full ids,
config tables, closure state; rule 4 closed, rule 5 amended by A5/A6),
BDR-107/108 (roles + levels preserved except analyzer), BDR-076/077
(judge rows = opus, off-state floor kept), LRN-205, LRN-206, LRN-207
(breaker feeds the tier resolution).
## Gate between A and B — live probe (user runs `/reload-plugins`)
Evidence into the W2-B contract: (1) typed `/status` → `/route show` main
`skill mechanical`; (2) a real `plugin-probe` or `feater` dispatch with
verbose on → spawn log line (provider shape, model id written) + `step 0
agent …` effort line = spawn/first-step ordering fact the TODO asks for;
(3) on a SONNET session (`/model sonnet`, then back): typed `/feat` → the
self-check wording of the system prompt + the `route` answer id (gate
witness fact); (4) probe (1) again with a background Explore alive (marker
path vs fallback path in the verbose log). The verbose log line (`typed-marker` /
`typed-fallback`, `spawn … → <id>`, `step 0 agent …`) is the witness for
(1), (2), (4), not `/route show` after the turn (turn-scoped routes are
gone by then). Decision rules: (2) step 0 BEFORE the spawn bookkeeping →
step 0 runs on the frontmatter `effort:` (kept, D2), recorded as a known
limit; (4) `typed-marker` never seen → W2-B blocked, A4 re-planned (the
`prompt.submit` text fact does not hold). Only then W2-B.
## W2-B — repo migration (own contract)
- [ ] B0 mod: delete `EFFORT_SKILL`, `effortBridge`, the Skill-hook branch,
their test; `agents/`+`skills/` rows unchanged.
- [ ] B1 `lib/effort-shift.md` rewritten (~40 lines): route doctrine, the
wiring points in `route` terms (D3), judgment built-ins carry
`effort=`, sticky skill route + `/route clear`, no pairing rule,
headless OK (hooks run under -p), levers = `ultrathink` / `/route
effort=max` (builtin `/effort` is not a lever inside a run), last ROWED
skill loaded wins
(an unrowed one changes nothing).
- [ ] B2 citers `Skill(effort-*)` → `route`: ship-feature 9, init-project
5, feat 4, bugfix 4, web-validate 3, seo 3, hotfix 3, geo 3,
verify-secure-loop 3, harden 2, code-clean 2, audit-delta 2, onboard
1, client-handover-writer 1; EFFORT SHIFTS header lines in the 13
orchestrators + tour:33; `effort=` added to every `general-purpose`
judgment dispatch (`model="opus"`: ship-feature, init-project,
onboard ×7, tour; `model: "fable"` skill-runners in
client-handover-writer); STOP/remedy texts (challenge-plan.md:62,
verify-secure-loop.md:109) → the D4 levers; the 21 prose sites
stating "sonnet by frontmatter pin" / "`model: opus`-pinned,
session-independent" (challenge-plan:46, feat:146, bugfix:160,
hotfix:138, code-clean:175, seo:463, 6 agents…) reworded "routed by
the model-router row (frontmatter = off-state floor)".
- [ ] B3 `lib/model-gate.md` → D4 (≤ 25 lines, keeps `model: "fable"` for
skill-runners and the dispatch-tier table); delete
`lib/model-check.sh`, `lib/tests/model-check.test.sh`.
- [ ] B4 delete `lib/effort-pins.txt`, `lib/effort-pins.sh`,
`lib/tests/effort-pins.test.sh`; remove the re-apply blocks + comments
in `install-plugins.sh` (941-943, 1008 comment, 1131-1136) and
`update-all.sh` (472, 567-572); `lib/tests/higgsfield.test.sh`
:324 + the `before-pins` check (:330-356) dropped.
- [ ] B5 `skills/effort-*` deleted (5 dirs; `profile`/catalog lists
grepped). Tracked frontmatter KEPT and ALIGNED: `agents/analyzer.md`
`effort: xhigh`; the vendored tracked `design-motion-principles`
keeps `effort: high`; impeccable-* (gitignored) untouched; every
other `model:`/`effort:` value already equals its row.
- [ ] B6 census: `lib/tests/effort-routing.test.sh` rewritten — every
tracked `effort:` equals its row's phase effort and every agent
`model:` alias equals its row's tier head (drift lock, both
directions, parsed from `register.ts`; haiku rows exempt from the
effort direction: no effort on haiku); enumeration = tracked
`SKILL.md` under `skills/` + `skills-external/` (`git ls-files`),
no-row list = graphify, model-router, find-docs, impeccable, the
agents interviewer/client-handover-writer/impeccable-*; no
`Skill(effort-` anywhere in
skills/agents/lib; D3 wiring markers (orchestrate in the 11, apply
tail in the 5, escalate ×3 in verify-secure-loop + ship-feature,
`effort=` on every `model="opus"` general-purpose dispatch); every
tracked skill/agent (minus the no-row list) has a row in
`register.ts` with the EXPECTED PHASE (per-name asserts, not mere
presence); `model-routing.test.sh` :96-97 model-gate locks
updated (the 18 `model:` locks, loops-light.test.sh:74,79 and
plan-challenger.test.sh:20 stay valid: frontmatter kept); `CLAUDE.global.md` Design
work lines 299-301 → "rows in the mod, an unrowed member changes
nothing".
- [ ] B7 doctrine-citers census; full `make test` once before merge.
- [ ] STEP 6/7 of the B run: doc-syncer audit → README/USAGE/ARCHITECTURE/
CHANGELOG; BDR-115 amendment (wave 2 closed, D1-D4, A5/A6), LRN
(typed-slash marker + sticky skill route), EVAL on the run; TODO W2
lines checked.
## Challenge ledger (r1 → r2)
- robustness 1 BLOCKER (mod off → agents inherit the parent model) →
D2: `model:` frontmatter kept as off-state floor; A3c row overrides it
while on.
- correctness 1/2, robustness 2 (marker unbound, ordering unproven) → A4
name-bound marker + pending slot + loops-size fallback; live probe gate.
- correctness 3 (show text) → A9 asserts `skill reflect`.
- correctness 4 (56) → fixed everywhere.
- correctness 5, simplicity 1/2, robustness 13 (level deltas) → `write`
high + `apply` phase; one delta left (analyzer), named.
- correctness 6 (raise lasts one turn) → A6 sticky skill route.
- correctness 7, robustness 5, simplicity 8 (gate witness) → D4 route answer.
- correctness 8, robustness 11, simplicity 7 (point 5) → explicit `effort=`.
- correctness 9, robustness 6 (levers) → D4 texts (r2's A7 engine-effort
detector dropped again in r3).
- correctness 10, robustness 9/10 (census) → B6 per-phase asserts, suites
listed; no `effort="high"` call-site edits (write = high).
- correctness 11 (unrowed skill clears) → A5.
- correctness 12, robustness 12, simplicity 11 (counts, impeccable) → B5
from `git ls-files`; impeccable untouched.
- correctness 13 (dead code) → A1 list; contract criterion 5 grep widened.
- correctness 14 (override test needs fs) → A9 bottom mocks or drop, stated.
- correctness 15, robustness 4/8 (provider shape, collisions) → A3b
source from `agent.offer` + A8 null rows; tests use `engine/core`.
- robustness 3 (haiku via fallback chain) → A3c tier-only + log.
- robustness 7 (live edits mid-migration) → bridge stays until B0; probe
gate before B.
- robustness 14 (tour:33, install comment) → B2/B4.
- robustness 15 (deploy on haiku) → `apply` rows.
- simplicity 3 (tail) → kept as `route(phase="apply")` (phases only; the
switch-on cost is the user's opt-in). [kept, reasoned]
- simplicity 4 (W2-C) → folded into W2-B STEP 6/7.
- simplicity 12 (two parsers) → contract oracle = one-shot values; B6 =
durable per-phase census. [kept, reasoned]
## Confirmation ledger (r2 → r3)
- conf 1 BLOCKER (route calls wipe the sticky slot) → A6 separate `runMain`
slot, turn writers never touch it.
- conf 2 (low/haiku leaks across turns) → only best-tier rows enter `runMain`.
- conf 3 (bridge sticky in the A→B window) → bridge writes `turnMain` only.
- conf 4 (engine-effort detector) → A7 dropped; builtin `/effort` named as
a non-lever to the user; levers = `ultrathink`, `/route effort=max`.
- conf 5 (route answer without id) → A3d always names the id, note appended;
probe (3) on a sonnet session.
- conf 6 (fs/cwd shadow unreliable) → A3b definition source from
`agent.offer`; fs dropped.
- conf 7 (judge → sonnet in-tier) → A3c rank ≥ tier head, else no write + log.
- conf 8 (off-state quality drop, step-0 ordering) → D2 frontmatter kept
as floor on both axes, census-locked; probe decision rule written.
- conf 9 (pre-fill premise) → explicit = `e.model !== undefined`, no map.
- conf 10 (origin denylist) → allowlist composer|sdk|bridge.
- conf 11 (preload before trackLoop) → spawning counter; log names the path.
- conf 12 (kit has no fs) → fs dropped. conf 13 → deduped log.
- conf 14 (skill rows shadow) → accepted, named: skill rows affect main
only; a foreign skill of the same name gets the row for its turn.
- conf 15 (artefact drift) → contract clarification amended; status-reporter
listed under mechanical.
## Confirmation 2 ledger (r3 → r4)
- conf2 1 (helper skill drops the run) → A6: model-loaded non-best rows
write turnMain only; typed rowed skills manage runMain.
- conf2 2 (test :347 goes red) → A9 authorizes its rewrite.
- conf2 3 (fallback ignores origin) → A4 `promptAllowed` flag on the fallback.
- conf2 4/5/6 (runMain wiring) → A6 names both slots, mainRoute, clearLoop
text, pushOrchestrate, `/model` + `/route off` + sub-agent cases in A9.
- conf2 7 (probe witnesses) → gate: verbose lines + decision rules.
- conf2 8 (grep) → contract criterion 5 narrowed.
- conf2 9 (source tokens) → A3b names projectSettings|localSettings + limit.
- conf2 10/11 (census) → B6 enumeration, haiku exemption, cites fixed.
+1 -1
View File
@@ -29,7 +29,7 @@ claude-config/
├── mods/ # Claude Code mods (function-hooks plugins), loaded through the skills/<name> symlink ├── mods/ # Claude Code mods (function-hooks plugins), loaded through the skills/<name> symlink
├── skills-external/ # Vendored skill packs: gstack submodule, design skills, superpowers, agent-skills, MengTo scroll skills, 21st and Higgsfield packs (machine-owned copies gitignored) ├── skills-external/ # Vendored skill packs: gstack submodule, design skills, superpowers, agent-skills, MengTo scroll skills, 21st and Higgsfield packs (machine-owned copies gitignored)
├── templates/ # Per-project templates (CLAUDE.md, settings, memory registries, deploy runbook, gitignore) ├── templates/ # Per-project templates (CLAUDE.md, settings, memory registries, deploy runbook, gitignore)
└── lib/ # Shared libs: gitflow, profiles, vendoring, effort pins, gates, archetypes, tests └── lib/ # Shared libs: gitflow, profiles, vendoring, route doctrine, gates, archetypes, tests
``` ```
## Architecture principles ## Architecture principles
+5 -1
View File
@@ -7,10 +7,11 @@ Format follows [Keep a Changelog](https://keepachangelog.com/) and this project
## [Unreleased] ## [Unreleased]
### Added ### Added
- **model-router mod**: `mods/model-router/`, a Claude Code mod (function-hooks plugin), routes the effort of every main-loop request from a phase table, along with the model and effort of the built-in sub-agents (Explore on sonnet/medium, Plan on opus/xhigh). It answers `Skill(effort-*)` itself, so the five `effort-*` skills no longer load while it is on. `ultrathink` in a prompt and a typed `/effort-<level>` set the main turn's default and minimum effort. The model gets a `route` tool and the user a `/route` command (`show|clear|off|on|reload|<phase>|model=<alias|id> effort=<level>|switch on|off|verbose on|off`). Optional per-machine config `~/.claude/model-router.json`, where `"enabled": false` turns it off on that machine. The spinner suffix and the status line show the route in force. It loads in every session through the tracked symlink `skills/model-router` (`model-router@skills-dir`). `make doctor` gains a Mods section; suite `make test suite=lib/tests/mods.test.sh`. Known limits: the main loop switches model only with `mainModelSwitch` on (default off, one cold-cache step per switch into another model), and the hooks send full model ids, so the `models` table has to follow new versions. - **model-router mod**: `mods/model-router/`, a Claude Code mod (function-hooks plugin), routes every repo skill and agent from phase rows (`plan`, `reflect`, `orchestrate`, `escalate`, `judge`, `implement`, `write`, `verify`, `explore`, `apply`, `mechanical`; built-ins Explore on sonnet/medium, Plan on opus/xhigh). A typed skill routes the main loop to its row, and a best-tier row holds across turns in a run slot until `/route clear`, `/route off`, a user `/model` or a typed skill on a non-best row. Agents get their row's model at spawn (within the tier, upward only) and its effort on every step; explicit Agent params win. Orchestrators declare their phases through the `mcp__model-router__route` tool. `ultrathink` in a prompt sets the turn's minimum effort and `/route effort=max` holds until `/route clear`; the built-in `/effort` is not a lever inside a run. `/route show` names the run slot (`main: run <phase>`), and a `null` row in the override drops a default row. The model gets a `route` tool and the user a `/route` command (`show|clear|off|on|reload|<phase>|model=<alias|id> effort=<level>|switch on|off|verbose on|off`). Optional per-machine config `~/.claude/model-router.json`, where `"enabled": false` turns it off on that machine. The spinner suffix and the status line show the route in force. It loads in every session through the tracked symlink `skills/model-router` (`model-router@skills-dir`). `make doctor` gains a Mods section; suite `make test suite=lib/tests/mods.test.sh`. Known limits: the main loop switches model only with `mainModelSwitch` on (default off, one cold-cache step per switch into another model), and the hooks send full model ids, so the `models` table has to follow new versions.
- **Manual-push mode**: `git config gitflow.autopush false` (human-set) now stops every push the gitflow lib makes, not only the post-commit / post-merge hooks. `gitflow start` and `finish` branch, commit and merge locally and push nothing; `gitflow delete` leaves the `origin/` copy in place and prints `git push origin --delete <br>` for the user to run. `hooks/unpushed-guard.sh` stays silent at turn end in this mode and opens each session with one `ℹ manual push mode:` line counting the commits no remote holds across every local branch; an unparseable or unreadable `gitflow.autopush` value is treated as manual push mode too, and that line names it. `hooks/push-guard.sh` (PreToolUse, `Bash|Monitor`) refuses any `git push` Claude types while `gitflow.autopush` reads false in the session cwd or in a literal `-C`/`cd` directory the command names (global config counts outside a repo); the refusal tells the user to run it with `! git push`. It reads the mode through the same lib verb as every other reader and fails closed: an unparseable or unreadable value reads as manual, and an internal error, a missing `lib/gitflow.sh`, more than 20 distinct directory tokens in one command (capped before any token is classified), a `cd`/`-C` directory token mixing quoted and unquoted parts, or a payload jq cannot parse whose raw text looks like a push refuse the push (these pathological cases fire in auto mode too). Directory tokens are read as whole shell words, adjacent quoted segments and backslash escapes included. In manual mode it over-blocks any command where a `push` word follows a `git` token; the misses listed in its header fall to a new `autoMode.soft_deny` rule that no request in the turn clears. The session banner adds `🔒 push : manual (autopush=false) — ! git push` when the key reads false, and `🔒 push : manual (autopush bad) — ! git push` when the value is invalid. Skills read the mode through a new lib verb, `bash ~/.claude/lib/gitflow.sh push-mode`: it prints `auto`, `manual` or `invalid` (rc 0) and names an invalid value on stderr (printable characters only, 64 at most). It is the one reader a skill may call, since the `git config` read of the key is denied to Claude. Skills push nothing on their own, except the `/release-candidate` tag in auto-push mode on an explicit go. Every "on origin" or "not pushed" line they print comes from `git rev-list --count origin/<br>..<br>` read after the fact, with the complete `! git …` command when something is left for the user to push. An invalid value (anything but unset, true or false, or a read that fails) is manual push mode for every reader and is named where it is read (see Fixed). Tests: `lib/gitflow-test.sh` T11b (push-mode verb), T18m and T18q blocks, `lib/tests/unpushed-guard.test.sh` T10-T16, `lib/tests/push-guard.test.sh` (98 checks). - **Manual-push mode**: `git config gitflow.autopush false` (human-set) now stops every push the gitflow lib makes, not only the post-commit / post-merge hooks. `gitflow start` and `finish` branch, commit and merge locally and push nothing; `gitflow delete` leaves the `origin/` copy in place and prints `git push origin --delete <br>` for the user to run. `hooks/unpushed-guard.sh` stays silent at turn end in this mode and opens each session with one `ℹ manual push mode:` line counting the commits no remote holds across every local branch; an unparseable or unreadable `gitflow.autopush` value is treated as manual push mode too, and that line names it. `hooks/push-guard.sh` (PreToolUse, `Bash|Monitor`) refuses any `git push` Claude types while `gitflow.autopush` reads false in the session cwd or in a literal `-C`/`cd` directory the command names (global config counts outside a repo); the refusal tells the user to run it with `! git push`. It reads the mode through the same lib verb as every other reader and fails closed: an unparseable or unreadable value reads as manual, and an internal error, a missing `lib/gitflow.sh`, more than 20 distinct directory tokens in one command (capped before any token is classified), a `cd`/`-C` directory token mixing quoted and unquoted parts, or a payload jq cannot parse whose raw text looks like a push refuse the push (these pathological cases fire in auto mode too). Directory tokens are read as whole shell words, adjacent quoted segments and backslash escapes included. In manual mode it over-blocks any command where a `push` word follows a `git` token; the misses listed in its header fall to a new `autoMode.soft_deny` rule that no request in the turn clears. The session banner adds `🔒 push : manual (autopush=false) — ! git push` when the key reads false, and `🔒 push : manual (autopush bad) — ! git push` when the value is invalid. Skills read the mode through a new lib verb, `bash ~/.claude/lib/gitflow.sh push-mode`: it prints `auto`, `manual` or `invalid` (rc 0) and names an invalid value on stderr (printable characters only, 64 at most). It is the one reader a skill may call, since the `git config` read of the key is denied to Claude. Skills push nothing on their own, except the `/release-candidate` tag in auto-push mode on an explicit go. Every "on origin" or "not pushed" line they print comes from `git rev-list --count origin/<br>..<br>` read after the fact, with the complete `! git …` command when something is left for the user to push. An invalid value (anything but unset, true or false, or a read that fails) is manual push mode for every reader and is named where it is read (see Fixed). Tests: `lib/gitflow-test.sh` T11b (push-mode verb), T18m and T18q blocks, `lib/tests/unpushed-guard.test.sh` T10-T16, `lib/tests/push-guard.test.sh` (98 checks).
### Changed ### Changed
- `lib/model-gate.md` calls the model-router `route` tool and takes its answer as the witness; with the mod off the gate stops and names `/route on`. The `model:` / `effort:` frontmatter of agents and skills is now the off-state floor, census-locked equal to the rows (`lib/tests/effort-routing.test.sh`). `analyzer` effort goes from high to xhigh. Built-in judgment dispatches carry an explicit `effort=`.
- `settings.json` denies every write form of the human-only `gitflow.*` keys (18 entries): any `git … config` spelling, section remove/rename, `git -c`, the git config env overrides, and Edit/Write of `.git/config`, `.gitconfig` and `~/.config/git/config`. Side effect: Claude can no longer read `gitflow.autopush` through `git config` either; hooks and `lib/gitflow.sh` still read it. The `hard_deny` rule on routing around a guardrail now names PreToolUse hook refusals. - `settings.json` denies every write form of the human-only `gitflow.*` keys (18 entries): any `git … config` spelling, section remove/rename, `git -c`, the git config env overrides, and Edit/Write of `.git/config`, `.gitconfig` and `~/.config/git/config`. Side effect: Claude can no longer read `gitflow.autopush` through `git config` either; hooks and `lib/gitflow.sh` still read it. The `hard_deny` rule on routing around a guardrail now names PreToolUse hook refusals.
- `gitflow start` and `finish` warn on stderr when a base is behind origin and cannot fast-forward, instead of a silent `git pull --ff-only || true` (T18l, T18n). - `gitflow start` and `finish` warn on stderr when a base is behind origin and cannot fast-forward, instead of a silent `git pull --ff-only || true` (T18l, T18n).
- `/close` (`/capitalize` STEP 5C) no longer runs its own push of develop: `gitflow finish` already pushes develop in auto-push mode (BDR-095). The closing line reports the real state, read after the merge: pushed, manual push mode with the `! git push origin develop` to run, not on origin, push failed, or an invalid `gitflow.autopush` value named and treated as manual push mode. A finish whose merge landed but whose branch delete failed (rc 5/2/6) still reports the push state. - `/close` (`/capitalize` STEP 5C) no longer runs its own push of develop: `gitflow finish` already pushes develop in auto-push mode (BDR-095). The closing line reports the real state, read after the merge: pushed, manual push mode with the `! git push origin develop` to run, not on origin, push failed, or an invalid `gitflow.autopush` value named and treated as manual push mode. A finish whose merge landed but whose branch delete failed (rc 5/2/6) still reports the push state.
@@ -18,6 +19,9 @@ Format follows [Keep a Changelog](https://keepachangelog.com/) and this project
- `/release-candidate` STEP 6: in manual push mode, with an invalid mode value, or when main or develop is not on origin, Claude pushes nothing and prints one command for the user, `! git push --atomic origin main develop v<X.Y.Z>`. The tag-push question remains for auto-push mode with both branches on origin. The version must match `^[0-9]+\.[0-9]+\.[0-9]+$` before it enters a command or tag; `release-executor` checks it too and blocks on anything else. - `/release-candidate` STEP 6: in manual push mode, with an invalid mode value, or when main or develop is not on origin, Claude pushes nothing and prints one command for the user, `! git push --atomic origin main develop v<X.Y.Z>`. The tag-push question remains for auto-push mode with both branches on origin. The version must match `^[0-9]+\.[0-9]+\.[0-9]+$` before it enters a command or tag; `release-executor` checks it too and blocks on anything else.
- `/tour`: each summary row says `on origin` or `local only` with the `! git -C "<project>" push -u origin <branch>` to run. The tour never pushes or retries. - `/tour`: each summary row says `on origin` or `local only` with the `! git -C "<project>" push -u origin <branch>` to run. The tour never pushes or retries.
### Removed
- The five `effort-low` … `effort-max` skills, `lib/effort-pins.txt`, `lib/effort-pins.sh` and `lib/model-check.sh`, with their tests and the effort-pin re-apply steps of `install-plugins.sh` and `update-all.sh`. The model-router rows replace them. A typed `/effort-<level>` no longer exists: use `ultrathink` or `/route effort=max`. Breaking: the next release is 3.0.0.
### Fixed ### Fixed
- `gitflow delete` (and `finish`) land on the base that contains the branch and drop the branch's upstream before `git branch -d`, so a branch whose upstream lags (manual-push mode) is deleted instead of refused by git (T18k). - `gitflow delete` (and `finish`) land on the base that contains the branch and drop the branch's upstream before `git branch -d`, so a branch whose upstream lags (manual-push mode) is deleted instead of refused by git (T18k).
- An invalid `gitflow.autopush` value (not a boolean, or a config read that fails) no longer pushes. The post-commit / post-merge hooks and every push site of `lib/gitflow.sh` (`start`, `finish`, the `origin/` cleanup of `delete`) read it as auto and pushed; they now push nothing and say why. Each hook run prints `gitflow post-commit: gitflow.autopush unreadable (git rc <n>) — NOT pushed, treated as manual push mode; fix the value by hand` (post-merge likewise), and the lib passes through the `push-mode` verb's line, `gitflow.sh push-mode: gitflow.autopush='<value>' is not a boolean (git rc <n>)`. A repo with its own committed `.githooks/` (onboarded projects) keeps running its old hooks, which still push on an invalid value, until a session start runs `reconcile-hooks` and rewrites them; commit that refresh so other clones get it. Tests: `lib/gitflow-test.sh` T18q block. - An invalid `gitflow.autopush` value (not a boolean, or a config read that fails) no longer pushes. The post-commit / post-merge hooks and every push site of `lib/gitflow.sh` (`start`, `finish`, the `origin/` cleanup of `delete`) read it as auto and pushed; they now push nothing and say why. Each hook run prints `gitflow post-commit: gitflow.autopush unreadable (git rc <n>) — NOT pushed, treated as manual push mode; fix the value by hand` (post-merge likewise), and the lib passes through the `push-mode` verb's line, `gitflow.sh push-mode: gitflow.autopush='<value>' is not a boolean (git rc <n>)`. A repo with its own committed `.githooks/` (onboarded projects) keeps running its old hooks, which still push on an invalid value, until a session start runs `reconcile-hooks` and rewrites them; commit that refresh so other clones get it. Tests: `lib/gitflow-test.sh` T18q block.
+3 -4
View File
@@ -295,10 +295,9 @@ design routing; the design-toolchain hook reinforces it.
- Design system / brand → design-consultation first, then the build tools. - Design system / brand → design-consultation first, then the build tools.
- Review / audit → design-review + emil-design-eng + design-motion-principles - Review / audit → design-review + emil-design-eng + design-motion-principles
+ /impeccable audit|critique + `impeccable detect` floor. + /impeccable audit|critique + `impeccable detect` floor.
- Load the stack paired with the first Read of the target file, never - The vendored stack members route to `reflect` (high) through their
alone (a lone Skill call applies no effort, `lib/effort-shift.md`); every model-router rows (`mods/model-router`); the last rowed skill loaded
vendored member pins `high`, one level per stack (`lib/effort-pins.txt`); wins and an unrowed member (plugin, gstack) changes nothing.
plugin and gstack members run at the level in force.
Scope doubt → ask or default to Build, never silently skip. Gate: light Scope doubt → ask or default to Build, never silently skip. Gate: light
skills run `~/.claude/lib/design-gate.md`, orchestrators plugin-check. 21st = skills run `~/.claude/lib/design-gate.md`, orchestrators plugin-check. 21st =
CLI (`npm i -g @21st-dev/cli`, `21st login`), no MCP, no key; search free, CLI (`npm i -g @21st-dev/cli`, `21st login`), no MCP, no key; search free,
+31 -22
View File
@@ -71,10 +71,14 @@ commands, settings, secrets, maintenance.
Doctrine: the session model (Fable) does main-loop reflection ONLY — Doctrine: the session model (Fable) does main-loop reflection ONLY —
brainstorm, plan, contract, audit judgment, gates, loop decisions — enforced brainstorm, plan, contract, audit judgment, gates, loop decisions — enforced
by a blocking gate (`lib/model-gate.md` + `lib/model-check.sh`) at the entry by a blocking gate (`lib/model-gate.md`) at the entry
of the 15 reflection skills (the orchestrators plus `/analyze`). Nothing dispatched inherits silently: of the 15 reflection skills (the orchestrators plus `/analyze`): the skill
typed agents carry a frontmatter pin, built-ins get an explicit `model=` at calls the model-router `route` tool and its answer, which names the model
every call site. id the main loop runs on, is the witness. A non-Fable/Opus id or the mod
off stops the skill (`/route on` resumes the mod). Nothing dispatched
inherits silently: typed agents run on their model-router row, their
`model:` frontmatter being the off-state floor; built-ins get an explicit
`model=` at every call site.
| Agent | Model | Tier | | Agent | Model | Tier |
|---|---|---| |---|---|---|
@@ -100,31 +104,37 @@ children are dispatched `model:"fable"` (they carry reflection).
## Effort routing (BDR-107, BDR-108) ## Effort routing (BDR-107, BDR-108)
Second axis of the same table: how hard each phase thinks. Session default Second axis of the same table: how hard each phase thinks. Session default
`high`. Every typed agent carries an `effort:` pin next to its `model:` (low `high`. The model-router mod (below) holds the live source: one phase row
appliers, medium executors, high judgment, xhigh challengers and gates; none per repo skill and agent, each phase naming a tier and an effort level
on haiku, which rejects the parameter). Every user-invoked skill carries an (`plan` best/xhigh, `reflect` best/high, `orchestrate` best/medium,
entry level (`/status` low … `/ship-feature` xhigh); the vendored externals `escalate` best/max, `judge` big/xhigh, `implement` work/medium, `write`
(design stack, superpowers, agent-skills, MengTo scroll skills, 21st) get theirs from work/high, `verify` work/xhigh, `explore` work/medium, `apply` work/low,
`lib/effort-pins.txt`, re-applied by `lib/effort-pins.sh` after every `mechanical` cheap/low). The `model:` and `effort:` frontmatter of every
vendoring step. Orchestrators shift per phase through the `effort-low` … typed agent and user-invoked skill stays as the off-state floor, kept equal
`effort-max` skills (`lib/effort-shift.md`, always sent with another tool to the rows by the census `lib/tests/effort-routing.test.sh`: low
call: a lone Skill call applies nothing (mod off)). Model pins stay tier aliases appliers, medium executors, high writers, xhigh judgment and gates (none
on haiku, which rejects the parameter); `/status` low … `/ship-feature`
xhigh. The vendored externals (design stack, superpowers, agent-skills,
MengTo scroll skills, 21st) get their level from their row only.
Orchestrators declare each phase through the `mcp__model-router__route`
tool (`lib/effort-shift.md`): `orchestrate` at a dispatch span, `reflect`
or `plan` when reflection resumes, `apply` at the bookkeeping tail,
`escalate` at the verify-secure caps. Model pins stay tier aliases
(`sonnet`, `opus`, `haiku`, `fable`): the latest version of a tier is also (`sonnet`, `opus`, `haiku`, `fable`): the latest version of a tier is also
the cheapest or same-priced, so the quality/price trade-off is tier × effort, the cheapest or same-priced, so the quality/price trade-off is tier ×
never version. Census `lib/tests/effort-routing.test.sh`; transcript audit effort, never version. Transcript audit `python3 lib/effort-audit.py`.
`python3 lib/effort-audit.py`.
### model-router mod ### model-router mod
`mods/model-router/` is a Claude Code mod (a function-hooks plugin) that applies this table per request. It loads in every session through the tracked symlink `skills/model-router`, as `model-router@skills-dir`. `mods/model-router/` is a Claude Code mod (a function-hooks plugin) that applies this table per request. It loads in every session through the tracked symlink `skills/model-router`, as `model-router@skills-dir`.
- Main loop: every request gets the effort of the phase in force. The mod answers `Skill(effort-*)` itself and applies the level from the next request on, so the five `effort-*` skills no longer load while it is on. - Main loop: every request gets the route in force. A typed skill with a row routes the main loop to it; a best-tier row (`plan`, `reflect`, `orchestrate`, `escalate`) holds across turns in a run slot until `/route clear`, `/route off`, a user `/model` or a typed skill on a non-best row. A skill without a row leaves the route as it is.
- Built-in sub-agents: Explore runs on sonnet/medium, Plan on opus/xhigh. An explicit `model` on the Agent call wins. - Sub-agents: a routed agent gets its row's model at spawn, within its tier and only upward from its frontmatter model, and the row's effort on every step. Explicit `model` / `effort` params on the Agent call win; a project-defined agent of the same name keeps its own definition. Built-ins: Explore runs on sonnet/medium, Plan on opus/xhigh.
- User floor: `ultrathink` in a prompt, or a typed `/effort-<level>`, sets the main turn's default and minimum effort. - Levers inside a run: `ultrathink` in a prompt sets the turn's minimum effort; `/route effort=max` holds until `/route clear`. The built-in `/effort` is not a lever inside a run, rows and routes outrank it.
- `/route` (user command) shows or sets the route: `show`, `clear`, `off`, `on`, `reload`, a phase name, `model=<alias|id> effort=<level>`, `switch on|off`, `verbose on|off`. The model sets routes through a `route` tool. - `/route` (user command) shows or sets the route: `show`, `clear`, `off`, `on`, `reload`, a phase name, `model=<alias|id> effort=<level>`, `switch on|off`, `verbose on|off`. `/route show` names the run slot when one holds (`main: run <phase>`). The model sets routes through a `route` tool.
- The spinner suffix and the status line under the prompt show the route in force. - The spinner suffix and the status line under the prompt show the route in force.
Optional per-machine config: `~/.claude/model-router.json`. Keys: `models` (alias → full id), `windows` (context window per full id), `phases`, `agents`, `skills`, `prompt` (rules), `mainModelSwitch` (default `false`), `verbose` (default `false`), `spinner` (default `true`), `enabled` (default `true`; `false` turns the mod off on that machine). `/route reload` re-reads it. Optional per-machine config: `~/.claude/model-router.json`. Keys: `models` (alias → full id), `windows` (context window per full id), `phases`, `agents`, `skills`, `prompt` (rules), `mainModelSwitch` (default `false`), `verbose` (default `false`), `spinner` (default `true`), `enabled` (default `true`; `false` turns the mod off on that machine). A `null` value in `agents` or `skills` drops a default row. `/route reload` re-reads it.
Limits: the main loop changes model only with `mainModelSwitch` on, and each switch into another model costs one cold-cache step. The hooks send full model ids, so the `models` table has to follow new model versions. Limits: the main loop changes model only with `mainModelSwitch` on, and each switch into another model costs one cold-cache step. The hooks send full model ids, so the `models` table has to follow new model versions.
@@ -218,7 +228,6 @@ a different package, ships its own conflicting `graphify` bin) — see
| `/profile` | Activate a skill profile (web / seo / web-full / full / max / backend / design / dev / qa / audit / minimal) (default: full) | | `/profile` | Activate a skill profile (web / seo / web-full / full / max / backend / design / dev / qa / audit / minimal) (default: full) |
| `/tour` | Grouped all-axes sweep — cleanup + security + reconcile + doc, fix and loop until clean | | `/tour` | Grouped all-axes sweep — cleanup + security + reconcile + doc, fix and loop until clean |
| `/site-motion` | Site-level motion: scroll engine choice, page transitions, pin/scrub sequencing across a page or Astro route (design stack) | | `/site-motion` | Site-level motion: scroll engine choice, page transitions, pin/scrub sequencing across a page or Astro route (design stack) |
| `/effort-low` … `/effort-max` | Effort shifters the orchestrators send per phase (answered by the model-router mod when on); typed by you, they set the main turn's default and minimum effort |
| `/route` | model-router mod: show or set the main-loop route (show, clear, off, on, reload, <phase>, model=… effort=…, switch on\|off, verbose on\|off) | | `/route` | model-router mod: show or set the main-loop route (show, clear, off, on, reload, <phase>, model=… effort=…, switch on\|off, verbose on\|off) |
> This table lists personal skills. Gstack skills (investigate, review, retro, > This table lists personal skills. Gstack skills (investigate, review, retro,
+12 -11
View File
@@ -174,19 +174,20 @@ Tu veux...
### Niveau d'effort ### Niveau d'effort
Chaque commande démarre à un niveau de réflexion fixé dans son frontmatter Chaque commande a une ligne de phase dans le mod model-router
(`effort:`) : low pour la tenue de registre (`/status`, `/close`, (`mods/model-router/`, actif dans chaque session), qui fixe son niveau de
réflexion : low pour la tenue de registre (`/status`, `/close`,
`/commit-change`), medium pour le courant (`/gitflow`, `/prune-memory`), `/commit-change`), medium pour le courant (`/gitflow`, `/prune-memory`),
high pour un fix ou un refactor (`/feat`, `/hotfix`, `/bugfix`, `/refactor`, high pour un fix ou un refactor (`/feat`, `/hotfix`, `/bugfix`,
audits avec fix), xhigh pour l'architecture et l'audit avant validation `/refactor`, audits avec fix), xhigh pour l'architecture et l'audit avant
(`/ship-feature`, `/onboard`, `/analyze`). Les orchestrateurs décalent validation (`/ship-feature`, `/onboard`, `/analyze`). Le frontmatter
ensuite le niveau par phase (`lib/effort-shift.md`), et `/effort-max` tapé à (`model:`, `effort:`) garde les mêmes valeurs et sert de plancher quand le
la main relance un tour bloqué au maximum. Les skills externes vendorés mod est coupé. Les orchestrateurs déclarent ensuite chaque phase via
(pile design, superpowers, agent-skills, skills scroll MengTo, 21st) reçoivent leur niveau de l'outil `mcp__model-router__route` (`lib/effort-shift.md`). Les skills
`lib/effort-pins.txt`. Un skill chargé seul par Claude n'applique pas son externes vendorés (pile design, superpowers, agent-skills, skills scroll
niveau : il doit partir avec un autre appel d'outil dans le même message. MengTo, 21st) ont aussi leur ligne.
Avec le mod model-router (`mods/model-router/`, actif dans chaque session), le niveau suit la phase à chaque requête. Le mod répond lui-même à `Skill(effort-*)` : le niveau s'applique dès la requête suivante et le texte des skills `effort-*` n'est plus chargé. Écrire `ultrathink` dans un prompt, ou taper `/effort-<niveau>`, fixe le niveau par défaut et le minimum du tour principal. Les sous-agents intégrés suivent leur route : Explore en sonnet/medium, Plan en opus/xhigh. `/route` affiche ou fixe la route (`/route show`, `/route clear`, `/route off`). La config par machine, optionnelle, vit dans `~/.claude/model-router.json` ; `"enabled": false` y coupe le mod sur cette machine. Taper un skill qui a une ligne route la boucle principale dessus. Une ligne du tier best (plan, reflect, orchestrate, escalate) tient d'un tour à l'autre pendant tout le run, jusqu'à `/route clear`, `/route off`, un `/model` tapé ou un skill d'un autre tier tapé. Pour relancer un tour bloqué : `ultrathink` dans le prompt (plancher du tour) ou `/route effort=max` (tient jusqu'à `/route clear`). Le `/effort` intégré n'a pas d'effet dans un run : les lignes et les routes passent devant. Les sous-agents reçoivent le modèle de leur ligne au lancement (dans leur tier, jamais en dessous de leur frontmatter) et son niveau à chaque étape ; un `model` ou `effort` explicite sur l'appel gagne. Explore tourne en sonnet/medium, Plan en opus/xhigh. `/route show` affiche la route en cours, slot de run compris (`main: run <phase>`). La config par machine, optionnelle, vit dans `~/.claude/model-router.json` ; `"enabled": false` y coupe le mod sur cette machine, et une ligne à `null` y retire une ligne par défaut.
## Les plugins — décision rapide ## Les plugins — décision rapide
+1 -1
View File
@@ -3,7 +3,7 @@ name: analyzer
description: Analyze code, codebase, or problem before any modification. Produces a factual report without proposing solutions. Use proactively before any refactoring, design, or implementation. description: Analyze code, codebase, or problem before any modification. Produces a factual report without proposing solutions. Use proactively before any refactoring, design, or implementation.
tools: Read, Grep, Glob, Bash tools: Read, Grep, Glob, Bash
model: opus model: opus
effort: high effort: xhigh
memory: project memory: project
--- ---
+10 -10
View File
@@ -97,7 +97,7 @@ Parse `$ARGUMENTS` for optional flags:
--- ---
EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route).
## STEP 1 — PRE-FLIGHT ## STEP 1 — PRE-FLIGHT
@@ -227,7 +227,7 @@ Store `DEPLOYED_URL` for STEP 7. If empty, ask user during STEP 6.
--- ---
## STEP 3 — BASELINE AUDITS (parallel) ## STEP 3 — BASELINE AUDITS (parallel)
First: `Skill(effort-high)` (effort-shift: judgment dispatch; the fable skill-runners are built-ins and inherit the level in force; high is the entry level of the audits they run). Every skill-runner dispatch below carries an explicit `effort="high"` (route: built-ins never inherit a main route; high is the entry level of the audits they run).
Goal: capture `SCORE_*_BEFORE` so the client doc shows the delta. Goal: capture `SCORE_*_BEFORE` so the client doc shows the delta.
@@ -262,10 +262,10 @@ the gate.
**Model routing (BDR-077):** EVERY `general-purpose` skill-runner dispatch in **Model routing (BDR-077):** EVERY `general-purpose` skill-runner dispatch in
this pipeline (initial audits, fix-loop re-dispatches, commit-change, this pipeline (initial audits, fix-loop re-dispatches, commit-change,
web-validate) carries `model: "fable"` — the child hosts gated orchestration web-validate) carries `model: "fable"` and `effort="high"` — the child hosts gated
on the pipeline's behalf; it must never inherit the session model. orchestration on the pipeline's behalf; it must never inherit the session model.
For web projects, dispatch in **a single message with two parallel Agent calls** (each with `model: "fable"`): For web projects, dispatch in **a single message with two parallel Agent calls** (each with `model: "fable"`, `effort="high"`):
| Audit (web) | Subagent | Prompt template | | Audit (web) | Subagent | Prompt template |
|---------------|-------------------|-----------------| |---------------|-------------------|-----------------|
@@ -424,7 +424,7 @@ for CSO) in the format
### Re-dispatch prompt template (SEO + GEO loop) ### Re-dispatch prompt template (SEO + GEO loop)
Send to `general-purpose` subagent (`model: "fable"`): Send to `general-purpose` subagent (`model: "fable"`, `effort="high"`):
> Read `~/.claude/skills/seo/SKILL.md` and re-run it on this project in > Read `~/.claude/skills/seo/SKILL.md` and re-run it on this project in
> **conservative (audit-only) intervention mode**: re-score and leave both > **conservative (audit-only) intervention mode**: re-score and leave both
@@ -455,7 +455,7 @@ Send to `general-purpose` subagent (`model: "fable"`):
### Re-dispatch prompt template (HARDEN loop) ### Re-dispatch prompt template (HARDEN loop)
Send to `general-purpose` subagent (`model: "fable"`): Send to `general-purpose` subagent (`model: "fable"`, `effort="high"`):
> Read `~/.claude/skills/harden/SKILL.md` and re-run it with `--fix` ONLY > Read `~/.claude/skills/harden/SKILL.md` and re-run it with `--fix` ONLY
> up to the bundle: stop at `READY TO APPLY — awaiting dispatcher > up to the bundle: stop at `READY TO APPLY — awaiting dispatcher
@@ -468,7 +468,7 @@ Send to `general-purpose` subagent (`model: "fable"`):
### Re-dispatch prompt template (CSO loop — non-web only) ### Re-dispatch prompt template (CSO loop — non-web only)
Send to `general-purpose` subagent (`model: "fable"`): Send to `general-purpose` subagent (`model: "fable"`, `effort="high"`):
> Read `~/.claude/skills/cso/SKILL.md` and re-run it in **daily mode**. > Read `~/.claude/skills/cso/SKILL.md` and re-run it in **daily mode**.
> Previous score: **`<SCORE_CSO_PREVIOUS>`/20** — below threshold. > Previous score: **`<SCORE_CSO_PREVIOUS>`/20** — below threshold.
@@ -557,7 +557,7 @@ listed changes by hand before deploy." Continue to STEP 6.
If `PENDING_CHANGES` non-empty → invoke /commit-change skill via subagent: If `PENDING_CHANGES` non-empty → invoke /commit-change skill via subagent:
> Dispatch `general-purpose` subagent (`model: "fable"`). Prompt: > Dispatch `general-purpose` subagent (`model: "fable"`, `effort="high"`). Prompt:
> >
> "Read `~/.claude/skills/commit-change/SKILL.md` and execute. All pending > "Read `~/.claude/skills/commit-change/SKILL.md` and execute. All pending
> changes were produced by the client-handover ship pipeline during the > changes were produced by the client-handover ship pipeline during the
@@ -677,7 +677,7 @@ Skip if `VALIDATE_SKIPPED=true` or `PROJECT_TYPE != web` (in either case
ensure `VALIDATE_SKIPPED=true` is set so the gate logic in STEP 8 treats ensure `VALIDATE_SKIPPED=true` is set so the gate logic in STEP 8 treats
VALIDATE as not-applicable rather than failed). VALIDATE as not-applicable rather than failed).
Dispatch `general-purpose` subagent (`model: "fable"`): Dispatch `general-purpose` subagent (`model: "fable"`, `effort="high"`):
> Read `~/.claude/skills/web-validate/SKILL.md` and execute against the > Read `~/.claude/skills/web-validate/SKILL.md` and execute against the
> deployed URL: `<DEPLOYED_URL>`. Audit W3C HTML validity (validator.nu), > deployed URL: `<DEPLOYED_URL>`. Audit W3C HTML validity (validator.nu),
+1 -1
View File
@@ -10,7 +10,7 @@ effort: high
> MODEL (BDR-077): `MODE: propose` is dispatched with `model="opus"` (the > MODEL (BDR-077): `MODE: propose` is dispatched with `model="opus"` (the
> call-site override — narrative reconstruction + capitalize routing are > call-site override — narrative reconstruction + capitalize routing are
> judgment); `MODE: apply` runs on the sonnet frontmatter pin (mechanical > judgment); `MODE: apply` runs on its sonnet row (frontmatter = off-state floor; mechanical
> staging/committing of an approved plan). > staging/committing of an approved plan).
Reconstruct the development narrative from a working directory. The goal Reconstruct the development narrative from a working directory. The goal
+2 -2
View File
@@ -60,8 +60,8 @@ audit, report, and patch.
Parse `$ARGUMENTS`: Parse `$ARGUMENTS`:
- **`MODE: patch`** — the dispatcher approved a PATCH PLAN and re-dispatches - **`MODE: patch`** — the dispatcher approved a PATCH PLAN and re-dispatches
this agent to APPLY it. Jump to MODE: PATCH section. Runs on the sonnet this agent to APPLY it. Jump to MODE: PATCH section. Runs on its sonnet
frontmatter pin. row (frontmatter = off-state floor).
- **`MODE: audit`** (or no explicit MODE — audit is the default) — analysis - **`MODE: audit`** (or no explicit MODE — audit is the default) — analysis
half, dispatched with `model: "opus"` (judgment tier; the call-site half, dispatched with `model: "opus"` (judgment tier; the call-site
override takes precedence over the sonnet pin). **READ-ONLY: Write and override takes precedence over the sonnet pin). **READ-ONLY: Write and
+1 -1
View File
@@ -104,7 +104,7 @@ Mirror of seo-analyzer's pipeline contract. Parse the MODE line:
to the run-scoped, gitignored `.audit/geo-signals-<RUNID>.md`, terminated to the run-scoped, gitignored `.audit/geo-signals-<RUNID>.md`, terminated
by `COLLECTION COMPLETE — RUNID: <RUNID>`; emit a `COLLECT REPORT` by `COLLECTION COMPLETE — RUNID: <RUNID>`; emit a `COLLECT REPORT`
(`STATUS`, RUNID, COVERAGE counts) and STOP. (`STATUS`, RUNID, COVERAGE counts) and STOP.
- **`MODE: judge`** — opus frontmatter pin. Fail-closed load of - **`MODE: judge`** — opus row (frontmatter = off-state floor). Fail-closed load of
`.audit/geo-signals-<RUNID>.md` (absent / RUNID mismatch / missing `.audit/geo-signals-<RUNID>.md` (absent / RUNID mismatch / missing
sentinel → `GEO JUDGE — VERDICT: ERROR(<reason>)`, STOP — never score sentinel → `GEO JUDGE — VERDICT: ERROR(<reason>)`, STOP — never score
stale or partial signals). Then STEP 6-12 (schema, entity — including stale or partial signals). Then STEP 6-12 (schema, entity — including
+1 -1
View File
@@ -57,7 +57,7 @@ The parent dispatches this agent TWICE, with the FULL PACKAGE both times
(`STATUS: DONE | BLOCKED`, RUNID, phase-cluster count, per-chapter word (`STATUS: DONE | BLOCKED`, RUNID, phase-cluster count, per-chapter word
counts) and STOPS — STEP 13-16, the final MD, HTML and PDF are NEVER counts) and STOPS — STEP 13-16, the final MD, HTML and PDF are NEVER
this mode's job. this mode's job.
- **`MODE: render`** — runs on the sonnet frontmatter pin. FIRST loads the - **`MODE: render`** — runs on its sonnet row (frontmatter = off-state floor). FIRST loads the
draft: absent file, RUNID mismatch, or missing `DRAFT COMPLETE` sentinel draft: absent file, RUNID mismatch, or missing `DRAFT COMPLETE` sentinel
→ `STATUS: BLOCKED` naming the cause (fail closed — never synthesize a → `STATUS: BLOCKED` naming the cause (fail closed — never synthesize a
missing draft, never render a partial one). Then runs STEP 13 → 14 → missing draft, never render a partial one). Then runs STEP 13 → 14 →
+5 -5
View File
@@ -106,11 +106,11 @@ the MAIN loop, never here):
- Dispatch THREE fresh challengers IN PARALLEL, one per lens - Dispatch THREE fresh challengers IN PARALLEL, one per lens
(correctness / robustness / simplicity), each blind to the others. (correctness / robustness / simplicity), each blind to the others.
- MODEL (BDR-076, supersedes the BDR-066 inherit): plan critique is AUDIT - MODEL (BDR-076, supersedes the BDR-066 inherit): plan critique is AUDIT
JUDGMENT, not a procedural gate — the challenger is `model: opus`-pinned in JUDGMENT, not a procedural gate — the challenger is routed to the `judge`
its frontmatter (big tier, session-independent; the session model stays on row (opus) by the model-router, the `model: opus` frontmatter being the off-state floor (big
the inline loop). Never `model: "sonnet"` — a silent judgment downgrade. tier; the session model stays on the inline loop). Never `model: "sonnet"` —
(Contrast the verifier, Sonnet-pinned only because it is oracle-anchored to a a silent judgment downgrade. (Contrast the verifier on its `verify` row and
contract.) the executors on their `implement` row, both sonnet, oracle-anchored to a contract.)
- FAIL-SAFE — never fail open: a malformed/empty verdict, a missing `PROOF`, or - FAIL-SAFE — never fail open: a malformed/empty verdict, a missing `PROOF`, or
a dead challenger → retry ONCE fresh; a 2nd failure → escalate to the human and a dead challenger → retry ONCE fresh; a 2nd failure → escalate to the human and
NAME the lens. Never report "plan challenged" on a silently dropped lens (same NAME the lens. Never report "plan challenged" on a silently dropped lens (same
+1 -1
View File
@@ -38,7 +38,7 @@ The dispatcher (/seo) runs this agent as a 3-stage pipeline; /harden and
terminated by the line `COLLECTION COMPLETE — RUNID: <RUNID>`, then terminated by the line `COLLECTION COMPLETE — RUNID: <RUNID>`, then
emits a short `COLLECT REPORT` (`STATUS: DONE | BLOCKED`, RUNID, emits a short `COLLECT REPORT` (`STATUS: DONE | BLOCKED`, RUNID,
COVERAGE counts) and STOPS. No scoring, no findings, no bundle. COVERAGE counts) and STOPS. No scoring, no findings, no bundle.
- **`MODE: judge`** — runs on the opus frontmatter pin (audit judgment). - **`MODE: judge`** — runs on its opus row (audit judgment; frontmatter = off-state floor).
FIRST loads `.audit/seo-signals-<RUNID>.md`: absent, RUNID mismatch, or FIRST loads `.audit/seo-signals-<RUNID>.md`: absent, RUNID mismatch, or
missing `COLLECTION COMPLETE` sentinel → emit missing `COLLECTION COMPLETE` sentinel → emit
`SEO JUDGE — VERDICT: ERROR(<reason>)` and STOP (fail closed — NEVER `SEO JUDGE — VERDICT: ERROR(<reason>)` and STOP (fail closed — NEVER
+1 -13
View File
@@ -937,10 +937,6 @@ for _ext_skill in "${EXT_SKILL_NAMES[@]}"; do
done done
echo "" echo ""
# Effort pins (BDR-107, BDR-108): every vendored external gets its entry
# level from lib/effort-pins.txt, re-applied ONCE after the last vendoring
# step (the 21st pack, STEP 8.7) — see apply_effort_pins there.
# ============================================================ # ============================================================
# STEP 8.5 — EXTERNAL SKILLS (npx skills add …) # STEP 8.5 — EXTERNAL SKILLS (npx skills add …)
# ============================================================ # ============================================================
@@ -1005,8 +1001,7 @@ echo ""
# profile: `lib/toggle-external.sh enable higgsfield` turns the media skills # profile: `lib/toggle-external.sh enable higgsfield` turns the media skills
# on, `enable higgsfield-websites` the landing-page aid. Keeping it out of # on, `enable higgsfield-websites` the landing-page aid. Keeping it out of
# link.sh and of every profile is what stops a re-run from re-enabling it # link.sh and of every profile is what stops a re-run from re-enabling it
# (BDR-093). This step runs before Step 8.7 so the effort pins are still # (BDR-093).
# re-applied after the last vendoring step (BDR-108).
echo "── Step 8.6: Higgsfield CLI + skill pack ───────────────────" echo "── Step 8.6: Higgsfield CLI + skill pack ───────────────────"
echo "" echo ""
# shellcheck source=lib/higgsfield-skills.sh disable=SC1091 # shellcheck source=lib/higgsfield-skills.sh disable=SC1091
@@ -1128,13 +1123,6 @@ if command -v 21st &>/dev/null; then
rm -rf "$TFD_STAGE" rm -rf "$TFD_STAGE"
fi fi
# Effort pins (BDR-107, BDR-108): the vendored externals carry no `effort:`
# upstream and every vendoring step above rewrites SKILL.md. Re-apply the
# entry levels from lib/effort-pins.txt once, after the LAST such step.
# shellcheck source=lib/effort-pins.sh disable=SC1091
source "$REPO/lib/effort-pins.sh"
apply_effort_pins "$REPO" || warn "effort pins: map lines rejected — fix lib/effort-pins.txt"
# Auth — detect, then offer login ONLY in an interactive TTY. A non-interactive # Auth — detect, then offer login ONLY in an interactive TTY. A non-interactive
# run (CI / headless / re-run) must never open a browser or block on OAuth. # run (CI / headless / re-run) must never open a browser or block on OAuth.
# Search and logo lookup are free; retrieving component code and 21st AI need # Search and logo lookup are free; retrieving component code and 21st AI need
+6 -4
View File
@@ -43,8 +43,9 @@ Agent(subagent_type="plan-challenger", description="challenge:<lens>", prompt=""
``` ```
**MODEL (BDR-076, supersedes the BDR-066 inherit):** plan critique is AUDIT **MODEL (BDR-076, supersedes the BDR-066 inherit):** plan critique is AUDIT
JUDGMENT — the challengers are `model: opus`-pinned in their frontmatter: a big JUDGMENT — the challengers are routed to the `judge` row (opus) by the model-router,
tier, session-independent, off the session model. The session model (Fable) the `model: opus` frontmatter being the off-state floor: a big tier, off
the session model. The session model (Fable)
keeps only this loop — synthesis, RE-THINK, gate. Never sonnet: that would keeps only this loop — synthesis, RE-THINK, gate. Never sonnet: that would
silently downgrade the judgment. (The executor gates stay sonnet.) silently downgrade the judgment. (The executor gates stay sonnet.)
@@ -59,8 +60,9 @@ silently downgrade the judgment. (The executor gates stay sonnet.)
A challenger that returns a malformed/empty verdict, a missing `PROOF`, or dies → A challenger that returns a malformed/empty verdict, a missing `PROOF`, or dies →
retry ONCE with a fresh challenger; a 2nd failure on that lens → STOP and escalate retry ONCE with a fresh challenger; a 2nd failure on that lens → STOP and escalate
(the STOP text names the level reached, `$CLAUDE_EFFORT`, and suggests `/effort-max` (the STOP text names the level reached, `$CLAUDE_EFFORT`, and suggests relaunching with
for the relaunch; no shift here: a mute challenger is an infrastructure failure) `ultrathink` in the prompt (turn floor) or `/route effort=max` (sticky,
`/route clear` after); no shift here: a mute challenger is an infrastructure failure)
to the human, NAMING the lens. Never carry "plan challenged" into the gate on a to the human, NAMING the lens. Never carry "plan challenged" into the gate on a
silently dropped lens (`verify-secure-loop.md`: "a mute verifier is NEVER a PASS"). silently dropped lens (`verify-secure-loop.md`: "a mute verifier is NEVER a PASS").
-120
View File
@@ -1,120 +0,0 @@
#!/usr/bin/env bash
# lib/effort-pins.sh — re-apply the entry effort level on vendored skills
# (BDR-107 second axis, extended to every vendored external by BDR-108).
# Upstream copies carry no `effort:` and every vendoring step rewrites
# SKILL.md, so the level lives in lib/effort-pins.txt and this helper puts
# it back after the last vendoring step of install-plugins.sh and
# update-all.sh. Idempotent: same level → untouched, other level →
# replaced inside the frontmatter only, skill not vendored → skipped,
# malformed map line → rejected loudly, never applied. Hardenings: a map
# whose last line lacks a newline is still read; a SKILL.md whose frontmatter
# never closes is skipped untouched; the level is re-read after every write
# and a mismatch counts as failed; the write goes through a mktemp sibling
# removed on any failure and on INT/TERM (previous traps restored, never an
# EXIT trap: the installer owns one); the rejected map line is printed
# shell-quoted so a caller's `echo -e` cannot interpret it. Placement inside
# the frontmatter has no effect on the harness, which reads the key anywhere.
#
# Usage: source it, then `apply_effort_pins [repo-root]`
# or standalone: bash lib/effort-pins.sh [repo-root]
# Exit 1 when at least one map line was rejected or a skill failed.
EFFORT_PINS_REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
EFFORT_PIN_LEVEL_RE='^(low|medium|high|xhigh|max)$'
EFFORT_PIN_NAME_RE='^[A-Za-z0-9][A-Za-z0-9._-]*$'
# Callers (install-plugins.sh, update-all.sh) define these; standalone
# runs get plain fallbacks.
declare -F ok >/dev/null || ok() { printf ' ok %s\n' "$*"; }
declare -F info >/dev/null || info() { printf ' info %s\n' "$*"; }
declare -F err >/dev/null || err() { printf ' ERR %s\n' "$*" >&2; }
# _effort_pin_current <skill-file> → prints the frontmatter effort, if any
_effort_pin_current() {
awk 'NR==1&&/^---$/{p=1;next} p&&/^---$/{exit} p' "$1" \
| sed -n 's/^effort: //p' | head -1
}
# _effort_pin_closed <skill-file> → rc 0 when the frontmatter has a closing ---
_effort_pin_closed() {
awk 'NR==1&&/^---$/{p=1;next} p&&/^---$/{f=1;exit} END{exit !f}' "$1"
}
# _effort_pin_traps_restore <saved> — drop the INT/TERM handlers set for the
# write and re-install the caller's saved ones. No exit-time handler here.
_effort_pin_traps_restore() {
trap - INT TERM
[ -z "$1" ] || eval "$1"
}
# _effort_pin_write <skill-file> <name> <level> — replace the frontmatter
# `effort:` line, or insert one after `name: <name>` (before the closing
# `---` when the frontmatter has no name line). Body lines never change.
# Writes a mktemp sibling then renames; any failure leaves no temp behind.
_effort_pin_write() {
local file="$1" name="$2" level="$3" tmp prev rc
tmp="$(mktemp "$file.XXXXXX")" || return 1
prev="$(trap -p INT TERM)"
trap 'rm -f "$tmp"; exit 130' INT TERM
cp -p "$file" "$tmp" && awk -v n="$name" -v lvl="$level" '
NR==1 && /^---$/ { fm=1; print; next }
fm && /^---$/ {
if (!done) { print "effort: " lvl; done=1 }
fm=0; print; next
}
fm && /^effort: / { if (!done) { print "effort: " lvl; done=1 }; next }
fm && $0 == "name: " n { print; if (!done) { print "effort: " lvl; done=1 }; next }
{ print }
' "$file" > "$tmp" && mv "$tmp" "$file"; rc=$?
[ "$rc" -eq 0 ] || rm -f "$tmp"
_effort_pin_traps_restore "$prev"
return "$rc"
}
# _effort_pin_apply_one <file> <name> <level> → rc 0 applied, 2 already at
# level, 1 failed (err line printed, file untouched or write rolled back)
_effort_pin_apply_one() {
local file="$1" name="$2" level="$3"
if ! _effort_pin_closed "$file"; then
err "effort-pins: $file: frontmatter never closed — skipped"; return 1
fi
[ "$(_effort_pin_current "$file")" = "$level" ] && return 2
if ! _effort_pin_write "$file" "$name" "$level"; then
err "effort-pins: $file: write failed"; return 1
fi
if [ "$(_effort_pin_current "$file")" != "$level" ]; then
err "effort-pins: $file: level not applied after write"
return 1
fi
return 0
}
# apply_effort_pins [repo-root] — walk the map, pin every vendored skill
apply_effort_pins() {
local repo="${1:-$EFFORT_PINS_REPO}" map name level rest file rc
local applied=0 kept=0 rejected=0 failed=0
map="$repo/lib/effort-pins.txt"
[ -f "$map" ] || { err "effort-pins: map missing: $map"; return 1; }
while read -r name level rest || [ -n "$name" ]; do
case "$name" in ''|'#'*) continue ;; esac
if [ -n "$rest" ] || ! [[ "$name" =~ $EFFORT_PIN_NAME_RE ]] \
|| ! [[ "$level" =~ $EFFORT_PIN_LEVEL_RE ]]; then
err "effort-pins: rejected map line $(printf '%q' "$name $level $rest")"
rejected=$((rejected + 1)); continue
fi
file="$repo/skills-external/$name/SKILL.md"
[ -f "$file" ] || continue
_effort_pin_apply_one "$file" "$name" "$level"; rc=$?
case "$rc" in
0) applied=$((applied + 1)) ;;
2) kept=$((kept + 1)) ;;
*) failed=$((failed + 1)) ;;
esac
done < "$map"
ok "effort-pins: $applied applied, $kept already at level, $failed failed"
[ "$rejected" -eq 0 ] && [ "$failed" -eq 0 ]
}
if [[ "${BASH_SOURCE[0]}" == "$0" ]]; then
apply_effort_pins "$@"
fi
-46
View File
@@ -1,46 +0,0 @@
# lib/effort-pins.txt — entry effort level of the vendored skills
# (skills-external/<name>/SKILL.md). Upstream copies carry no `effort:` and
# every resync rewrites SKILL.md, so the pin lives here and
# lib/effort-pins.sh re-applies it after the last vendoring step of
# install-plugins.sh and update-all.sh. One line = `<skill> <level>`,
# level in low|medium|high|xhigh|max. The census
# lib/tests/effort-routing.test.sh checks every vendored file against this
# map. Rungs (BDR-107, BDR-108): low = fix a line, run a script · medium =
# day-to-day · high = refactor, resisting bug · xhigh = architecture, audit
# before validation · max = stuck.
#
# superpowers (obra/superpowers, plugins.lock.json "superpowers")
brainstorming xhigh
writing-plans xhigh
requesting-code-review xhigh
subagent-driven-development high
writing-skills high
test-driven-development medium
using-git-worktrees low
#
# agent-skills (addyosmani/agent-skills, plugins.lock.json "agent-skills")
deprecation-and-migration high
ci-cd-and-automation medium
observability-and-instrumentation medium
#
# design stack — ONE level for every member: these skills load stacked in a
# single UI build and the last loaded wins (lib/effort-shift.md), so two
# levels in the stack would make the effort depend on load order.
# skills/site-motion (repo-authored) pins the same level in its frontmatter.
frontend-design high
emil-design-eng high
design-motion-principles high
21st-ui-build high
scroll-world-storytelling high
build-threejs-scroll-worlds high
scroll-scrubbed-visual-sequence high
scroll-scrubbed-word-reveal high
scroll-progress-timeline high
#
# 21st pack (`21st skills install`): tooling low, generation high, critique xhigh
21st-cli-use low
21st-registry low
21st-design-sync low
21st-ai high
21st-ui-explore high
21st-ui-review xhigh
+47 -74
View File
@@ -1,85 +1,58 @@
# Effort shift — phase-level reasoning effort on the main loop (BDR-107) # Route doctrine — phase-level model and effort on the main loop (BDR-107)
Shared include, companion of `lib/model-gate.md`: the gate fixes WHICH model Shared include, companion of `lib/model-gate.md`: the gate fixes WHICH model
reflects, this include fixes HOW HARD each phase thinks. The rungs are the reflects, the model-router mod (`mods/model-router`) fixes HOW HARD each
user's: low (fix a line, run a script) · medium (day-to-day) · high phase thinks. Rungs: low (fix a line, run a script) · medium (day-to-day) ·
(refactor, resisting bug) · xhigh (architecture, audit before validation) · high (refactor, resisting bug) · xhigh (architecture, audit before
max (stuck error, judged need). validation) · max (stuck error, judged need).
## Mechanics (verified on Claude Code 2.1.283) ## The tool
- **Pairing rule**: a `Skill(effort-<level>)` call applies its effort only `mcp__model-router__route` (params `phase` | `effort` | `clear`). It is a
when the same assistant message carries at least one other tool call deferred tool: when not loaded, run
after it; a lone Skill call is a no-op. Send the shift together with the `ToolSearch("select:mcp__model-router__route")` once per session. A route
step's first tool call, shift first. That paired call already runs at the applies from the next request on, paired with
new level: pair a downward shift with a pinned-agent dispatch or a another tool call or not (pairing only saves a request). The answer always
Read/Bash, never with a built-in judgment dispatch (`general-purpose`, names the id and effort main runs on. A skill with a row routes itself on
`model: "opus"`), which would inherit it. load; a skill without one changes nothing, the last ROWED skill wins.
- Re-loading a shifter already loaded in the conversation re-applies its
effort (the harness only dedupes the skill text), so bounce-back ## Wiring points
sequences such as medium → max → medium work.
- A skill's `effort:` frontmatter applies from the moment it loads to the 1. Dispatch span starts → `route(phase="orchestrate")`, sent with the
end of the turn: on the user's `/skill` unconditionally, and on a dispatch.
`Skill(...)` call by Claude only under the pairing rule above (a skill 2. Reflection resumes (challenge synthesis, verdict, plan revision) →
Claude loads alone, such as `brainstorming` or `writing-plans`, applies `route(phase="reflect")` or `"plan"` per the skill's own level; the line
nothing). Last loaded wins, both directions. The prompt cache survives a before every `lib/challenge-plan.md` call.
shift. 3. Bookkeeping tail (memory commit, doc commit) → `route(phase="apply")`.
- **Stacked skills share one level**: skills that load together in one 4. Escalation → `route(phase="escalate")`: verify-secure loop caps and
build (the design stack) all pin the same level, since the last loaded ship-feature STEP 4b. Not automatic: the challenge fail-safe and "gone
wins. Vendored externals get their level from `lib/effort-pins.txt`, WRONG → STOP"; their STOP text names the levers below.
re-applied by `lib/effort-pins.sh` after every vendoring step; repo 5. Built-in judgment dispatch (`general-purpose` `model="opus"`, `model:
skills carry it in their frontmatter. "fable"` skill-runners) → explicit `effort=` on the Agent call (`xhigh`
- Dispatched agents run on their own `effort:` pin, never on a shift. for opus reviewers, `high` for fable runners). A main route never
Unpinned agents inherit the level in force at dispatch. reaches a child. Typed agents run on their row, never on a shift.
- Headless sessions (`-p`, `claude agents`, SDK) ignore skill-level effort: 6. After a prose gate that ends the turn, the resumed reflection phase
the run stays at the session level. `CLAUDE_CODE_EFFORT_LEVEL` beats every starts with its own route call.
frontmatter; keep it unset (the session banner warns).
## Run slot and levers
A best-tier skill row survives the end of the turn (a run spans prose
gates); `/route clear`, `/route off` and a user `/model` drop it. Levers for a
relaunch: `ultrathink` in the prompt (turn floor) or `/route effort=max`
(sticky, `/route clear` after).
Builtin `/effort` is NOT a lever inside a run: rows and routes outrank it.
## Limits
- A skill typed while a background agent is live routes only through the
typed marker (unverified live 2026-10-10).
- Headless (`-p`, SDK) runs the hooks, so routing works there too.
- Mod off: typed agents fall back to their `model:`/`effort:` frontmatter.
Measure the split any time: `python3 ~/.claude/lib/effort-audit.py` Measure the split any time: `python3 ~/.claude/lib/effort-audit.py`
(thinking/output/cache tokens per scope, model and effort). (thinking/output/cache tokens per scope, model and effort).
## Shifters
`Skill(effort-low)` · `Skill(effort-medium)` · `Skill(effort-high)` ·
`Skill(effort-xhigh)` · `Skill(effort-max)`. One tool call, one-line body,
always sent with another tool call (Pairing rule).
Typed by the user, `/effort-max` is a turn-scoped max: the relaunch lever
after a STOP. `ultrathink` only adds an in-context nudge; the API level
does not move.
## Wiring — per orchestrator
1. A dispatch span starts (executor, collector, fan-out) →
`Skill(effort-medium)`.
2. Reflection resumes after a dispatch span (challenge synthesis, verdict,
plan revision) → `Skill(effort-<the skill's own level>)`. Concretely:
the line before every `lib/challenge-plan.md` call.
3. The bookkeeping tail (memory commit, doc commit) → `Skill(effort-low)`.
4. Escalation → `Skill(effort-max)`, then the skill's own level again once
the diagnosis is produced. Automatic points: verify-secure loop caps
(GATE 0 floor, GATE 1 conformity, GATE 2 security) and ship-feature
STEP 4b. Not automatic, by doctrine: the challenge fail-safe (a mute
challenger is an infrastructure failure) and "gone WRONG → STOP" (STOP
precedes any further reasoning); their STOP text names the level
reached and suggests `/effort-max` for the relaunch.
5. Before any built-in or unpinned dispatch that carries judgment (a
`general-purpose` with `model: "opus"` or `"fable"`, the code reviewer
of requesting-code-review, a skill-runner) → `Skill(effort-<own level>)`
paired with that dispatch: built-ins inherit the level in force, and a
medium set earlier in the span would downgrade them.
## Re-assert
- After any nested `Skill(...)` whose frontmatter carries a different
effort (feat → commit-change), reload the orchestrator's own level.
- After a prose gate that ends the turn, the resumed turn runs at the
session level. If the resumed phase is reflection, its first step is
`Skill(effort-<own level>)`; dispatch and orchestration phases need
nothing.
## Never ## Never
- A shift never inside a dispatched agent: pins rule there. - A route inside a dispatched agent: its row rules there.
- Max is for diagnosis, not for retrying the same fix harder. - Max is for diagnosis, not for retrying the same fix harder.
- A medium shift never precedes a judgment dispatch in the same span
without an own-level shift paired with that dispatch.
-33
View File
@@ -1,33 +0,0 @@
#!/usr/bin/env bash
# lib/model-check.sh — classify the persisted session model: big | small | unknown
#
# Witness for lib/model-gate.md (reflection requires a big model). Reads the
# "model" key of the user-scope settings (the file /model rewrites — LRN-098).
# Override the source with MODEL_CHECK_SETTINGS (tests use fixtures).
#
# stdout : <class>:<raw> (raw = value found, empty if none)
# exit : 0 = big (fable/opus) · 2 = small (sonnet/haiku) · 3 = unknown
set -u
SETTINGS="${MODEL_CHECK_SETTINGS:-$HOME/.claude/settings.json}"
raw=""
if [ -f "$SETTINGS" ]; then
raw="$(python3 - "$SETTINGS" 2>/dev/null <<'PY'
import json, sys
try:
v = json.load(open(sys.argv[1])).get("model", "")
print(v if isinstance(v, str) else "")
except Exception:
print("")
PY
)"
fi
norm="$(printf '%s' "$raw" | tr '[:upper:]' '[:lower:]')"
case "$norm" in
*opusplan*) printf 'unknown:%s\n' "$raw"; exit 3 ;; # opus-for-plan, sonnet otherwise — ambiguous
*fable*|*opus*) printf 'big:%s\n' "$raw"; exit 0 ;;
*sonnet*|*haiku*) printf 'small:%s\n' "$raw"; exit 2 ;;
*) printf 'unknown:%s\n' "$raw"; exit 3 ;;
esac
+23 -46
View File
@@ -1,53 +1,30 @@
# Model gate — reflection requires a big model (BLOCKING) # Model gate — reflection requires a big model (BLOCKING)
Shared include. Runs FIRST in any orchestrator whose reflection — Shared include, runs FIRST in an orchestrator whose reflection executes
brainstorming, planning, contract, audit judgment, loop decisions — inline (BDR-066). The witness is the model-router mod's own route tool.
executes inline or in inherit-model subagents. Sonnet-pinned executors are
not what this gate protects; it protects the thinking around them (BDR-066).
## 1. Self-check ## Entry call
ALWAYS call `mcp__model-router__route` with `phase` = the skill's row phase
(reflect or plan), no self-check shortcut. Tool not loaded (deferred) →
`ToolSearch("select:mcp__model-router__route")` once per session, then
call. The answer names the id main runs on next.
Your system prompt names the model powering this session. Fable or Opus → | answer | action |
big. Sonnet, Haiku, anything else → small. |---|---|
| names a fable or opus id | proceed, SILENT |
| names sonnet, haiku, anything else; "is off"; tool absent | **STOP** |
## 2. Witness — deterministic check **STOP means**: print exactly `⛔ MODEL GATE — session on <model>.
Reflection steps of this skill require Fable or Opus. Switch with /model,
then relaunch the skill.` (mod off: say so, `/route on` resumes it), then
end
the turn. No later step runs, no agent is dispatched, nothing is edited.
bash "$HOME/.claude/lib/model-check.sh" ## Dispatch tiers (BDR-077 — no inherit)
The gate guards the MAIN loop only. Typed agents are routed by their
Output `<class>:<raw>`; exit 0 = big, 2 = small, 3 = unknown. The witness model-router row (`model:` frontmatter = off-state floor). Built-ins
reads the PERSISTED model (settings.json — the file `/model` rewrites,
LRN-098). It can lag reality (session launched with `--model`, settings not
yet rewritten) — that is why the self-check exists alongside it.
## 3. Verdict
| self-check | witness | action |
|---|---|---|
| big | big (0) | proceed, SILENT — the nominal path prints nothing |
| small | any | **STOP** |
| big | small (2) | disagreement — **STOP**, surface BOTH values; the user confirms or relaunches |
| big | unknown (3) | fail-visible: print `model gate: witness unknown (<raw>) — self-check says <model>` and ask the user to confirm before continuing (BDR-025: unknown never silently passes) |
**STOP means**: print exactly
⛔ MODEL GATE — session on <model>. Reflection steps of this skill
require Fable or Opus. Switch with /model, then relaunch the skill.
then end the turn. No later step runs, no agent is dispatched, nothing is
edited.
## 4. Dispatch tiers (BDR-077 — no inherit)
The gate guards the MAIN loop only. Dispatched work NEVER inherits the
session model: typed agents run on their frontmatter pin; built-ins
(general-purpose / Explore / Plan) carry an explicit `model=` at every call (general-purpose / Explore / Plan) carry an explicit `model=` at every call
site — `model: "fable"` when the child performs reflection/orchestration on site: `model: "fable"` when the child reflects/orchestrates for the main
the main loop's behalf (skill-runners), otherwise its complexity tier loop (skill-runners), else its tier (opus = dispatched judgment, sonnet =
(opus = dispatched judgment, sonnet = execution/collection, haiku = short execution, haiku = mechanical probes). A built-in judgment dispatch also
mechanical probes). carries an explicit `effort=` (`lib/effort-shift.md`).
Effort is the second axis of the same table (BDR-107): every typed agent
carries an `effort:` pin next to `model:`, and the main loop shifts per phase
through `lib/effort-shift.md`. No typed agent inherits either axis;
built-ins inherit the effort in force at dispatch, so an orchestrator shifts
before dispatching them (`lib/effort-shift.md`, wiring point 5).
-129
View File
@@ -1,129 +0,0 @@
#!/usr/bin/env bash
# lib/tests/effort-pins.test.sh — lib/effort-pins.sh's apply_effort_pins():
# insert after `name:`, keep an equal level untouched, replace a different
# level inside the frontmatter only (a prose `effort:` in the body stays),
# skip a skill not vendored, insert before the closing `---` when the
# frontmatter has no name line, run idempotently, reject a bad level, a
# traversal name and a three-field line before writing anything, and
# parse the real map without error; hardening: last map line without a
# newline, unterminated frontmatter, CRLF file and read-only directory. All on a throwaway fixture repo.
set -u
ROOT="$(cd "$(dirname "$0")/../.." && pwd)"
LIB="$ROOT/lib/effort-pins.sh"
pass=0; fail=0
check() { if [ "$2" = "$3" ]; then pass=$((pass+1)); echo "PASS $1"
else fail=$((fail+1)); echo "FAIL $1: got[$2] want[$3]"; fi; }
fm_effort() { awk 'NR==1&&/^---$/{p=1;next} p&&/^---$/{exit} p' "$1" \
| sed -n 's/^effort: //p' | head -1; }
WORK="$(mktemp -d)" || exit 1; trap 'rm -rf "$WORK"' EXIT
REPO="$WORK/repo"; EXT="$REPO/skills-external"
mkdir -p "$REPO/lib" "$EXT/alpha" "$EXT/beta" "$EXT/gamma" "$EXT/noname"
printf -- '---\nname: alpha\ndescription: a\n---\nbody\n' > "$EXT/alpha/SKILL.md"
printf -- '---\nname: beta\neffort: low\n---\nprose says effort: max here\n' > "$EXT/beta/SKILL.md"
printf -- '---\nname: gamma\neffort: low\n---\nbody\n' > "$EXT/gamma/SKILL.md"
printf -- '---\ndescription: no name line\n---\nbody\n' > "$EXT/noname/SKILL.md"
printf '# map\nalpha high\nbeta medium\ngamma low\nghost xhigh\nnoname low\n' > "$REPO/lib/effort-pins.txt"
gamma_before="$(cat "$EXT/gamma/SKILL.md")"
bash "$LIB" "$REPO" >/dev/null 2>&1; check T1-rc-clean "$?" 0
check T2-insert-after-name "$(sed -n '3p' "$EXT/alpha/SKILL.md")" "effort: high"
check T3-replace-in-frontmatter "$(fm_effort "$EXT/beta/SKILL.md")" "medium"
check T3b-body-prose-untouched "$(grep -c 'effort: max' "$EXT/beta/SKILL.md")" 1
check T3c-single-effort-line "$(grep -c '^effort:' "$EXT/beta/SKILL.md")" 1
check T4-equal-level-untouched "$(cat "$EXT/gamma/SKILL.md")" "$gamma_before"
check T5-missing-skill-skipped "$([ -e "$EXT/ghost" ] && echo created || echo absent)" absent
check T6-no-name-inserts-before-closing "$(sed -n '3p' "$EXT/noname/SKILL.md")" "effort: low"
check T6b-no-name-still-frontmatter "$(fm_effort "$EXT/noname/SKILL.md")" "low"
snap="$(cat "$EXT"/*/SKILL.md)"
bash "$LIB" "$REPO" >/dev/null 2>&1
check T7-idempotent "$(cat "$EXT"/*/SKILL.md)" "$snap"
check T7b-no-tmp-left "$(find "$EXT" -name '*.tmp' | wc -l | tr -d ' ')" 0
# rejections: nothing written, rc 1
for bad in 'alpha turbo' '../evil high' 'alpha high extra'; do
printf '%s\n' "$bad" > "$REPO/lib/effort-pins.txt"
out="$(bash "$LIB" "$REPO" 2>&1)"; rc=$?
check "T8-rejected[$bad]-rc" "$rc" 1
check "T8-rejected[$bad]-named" "$(printf '%s' "$out" | grep -c 'rejected map line')" 1
done
check T8b-tree-unchanged-after-rejections "$(cat "$EXT"/*/SKILL.md)" "$snap"
check T8c-no-evil-dir "$([ -e "$WORK/evil" ] && echo created || echo absent)" absent
# the real map parses: fixture repo with the real map and no vendored skill
mkdir -p "$WORK/real/lib" "$WORK/real/skills-external"
cp "$ROOT/lib/effort-pins.txt" "$WORK/real/lib/"
out="$(bash "$LIB" "$WORK/real" 2>&1)"; check T9-real-map-parses "$?" 0
check T9b-real-map-nothing-applied "$(printf '%s' "$out" | grep -c '0 applied, 0 already')" 1
check T10-missing-map-rc "$(bash "$LIB" "$WORK/nowhere" >/dev/null 2>&1; echo $?)" 1
# hardening: each case in its own fixture repo
mkrepo() { R="$WORK/$1"; mkdir -p "$R/lib" "$R/skills-external/$2"; }
mkrepo h11 alpha; mkdir "$WORK/h11/skills-external/beta"
printf -- '---\nname: alpha\n---\nb\n' > "$WORK/h11/skills-external/alpha/SKILL.md"
printf -- '---\nname: beta\n---\nb\n' > "$WORK/h11/skills-external/beta/SKILL.md"
printf 'alpha high\nbeta low' > "$WORK/h11/lib/effort-pins.txt"
bash "$LIB" "$WORK/h11" >/dev/null 2>&1
check T11-last-line-no-newline "$(fm_effort "$WORK/h11/skills-external/beta/SKILL.md")" low
mkrepo h12 open; f12="$WORK/h12/skills-external/open/SKILL.md"
printf -- '---\nname: open\nbody effort: max\n' > "$f12"; b12="$(cat "$f12")"
printf 'open high\n' > "$WORK/h12/lib/effort-pins.txt"
out="$(bash "$LIB" "$WORK/h12" 2>&1)"; rc=$?
check T12-unterminated-frontmatter-skipped \
"$rc|$(cat "$f12" | cmp -s - <(printf '%s\n' "$b12") && echo same)|$(printf '%s' "$out" | grep -c "ERR .*$f12")" "1|same|1"
mkrepo h13 crlf
printf -- '---\r\nname: crlf\r\n---\r\nbody\r\n' > "$WORK/h13/skills-external/crlf/SKILL.md"
printf 'crlf high\n' > "$WORK/h13/lib/effort-pins.txt"
out="$(bash "$LIB" "$WORK/h13" 2>&1)"; rc=$?
check T13-crlf-file-rejected \
"$rc|$(printf '%s' "$out" | grep -c 'ERR ')|$(printf '%s' "$out" | grep -c ' 0 applied, ')" "1|1|1"
# T13b: the post-write re-read branch, reached with a no-op write stub
mkrepo h13b nowrite; f13b="$WORK/h13b/skills-external/nowrite/SKILL.md"
printf -- '---\nname: nowrite\n---\nb\n' > "$f13b"
out="$(bash -c 'source "$1"; _effort_pin_write() { return 0; }
_effort_pin_apply_one "$2" nowrite high' _ "$LIB" "$f13b" 2>&1)"; rc=$?
check T13b-reread-mismatch-fails \
"$rc|$(printf '%s' "$out" | grep -c 'level not applied after write')" "1|1"
if [ "${EFFORT_PINS_TEST_FAKE_ROOT:-0}" = 1 ] || [ "$(id -u)" -eq 0 ]; then
echo "SKIP T14-write-failure-no-temp: chmod bits ignored as root"
else
mkrepo h14 ro; d14="$WORK/h14/skills-external/ro"
printf -- '---\nname: ro\n---\nb\n' > "$d14/SKILL.md"
printf 'ro high\n' > "$WORK/h14/lib/effort-pins.txt"
chmod 555 "$d14"; out="$(bash "$LIB" "$WORK/h14" 2>&1)"; rc=$?; chmod 755 "$d14"
check T14-write-failure-no-temp \
"$rc|$(printf '%s' "$out" | grep -c 'ERR ')|$(find "$d14" -name 'SKILL.md.*' | wc -l | tr -d ' ')" "1|1|0"
fi
# T15: SIGINT during the awk write removes the temp sibling, exit 130
mkrepo h15 sig; d15="$WORK/h15/skills-external/sig"
printf -- '---\nname: sig\n---\nb\n' > "$d15/SKILL.md"
bash -c 'source "$1"; awk() { kill -INT $$; sleep 2; }
_effort_pin_write "$2" sig high' _ "$LIB" "$d15/SKILL.md" >/dev/null 2>&1
rc=$?
check T15-sigint-removes-temp \
"$rc|$(find "$d15" -name 'SKILL.md.*' | wc -l | tr -d ' ')" "130|0"
# T15b: previous INT trap restored on a normal return, no EXIT trap set
mkrepo h15b tr; d15b="$WORK/h15b/skills-external/tr"
printf -- '---\nname: tr\n---\nb\n' > "$d15b/SKILL.md"
out="$(bash -c 'source "$1"; trap "echo prev" INT
_effort_pin_write "$2" tr high
printf "INT:%s\n" "$(trap -p INT)"; printf "EXIT:%s\n" "$(trap -p EXIT)"' \
_ "$LIB" "$d15b/SKILL.md" 2>&1)"
check T15b-traps-restored \
"$(printf '%s' "$out" | grep -c "^INT:trap -- 'echo prev' SIGINT")|$(printf '%s' "$out" | grep -c '^EXIT:$')" "1|1"
# T16: a literal backslash-t in a map line is printed shell-quoted
mkrepo h16 q
printf 'bad\\tname high\n' > "$WORK/h16/lib/effort-pins.txt"
out="$(bash "$LIB" "$WORK/h16" 2>&1)"; rc=$?
check T16-rejected-line-quoted \
"$rc|$(printf '%s' "$out" | grep -cF 'bad\\tname')" "1|1"
echo "effort-pins: $pass pass, $fail fail"
[ "$fail" -eq 0 ]
+128 -112
View File
@@ -1,9 +1,12 @@
#!/usr/bin/env bash #!/usr/bin/env bash
# lib/tests/effort-routing.test.sh — census: effort tiering (BDR-107) # lib/tests/effort-routing.test.sh — wave-2 census of the model-router rows.
# agent pins, skill entry levels, shifter skills, orchestrator wiring, settings. # Drift lock: every tracked skill/agent row in mods/model-router/hooks/
# shellcheck disable=SC2015,SC2016 # A && ok || ko is deliberate (ok/ko never fail); '$REPO' locks are literal source text # register.ts equals its frontmatter (the off-state floor), the D3 wiring
# markers sit in the orchestrators, no shifter citer survives.
# shellcheck disable=SC2015,SC2016 # A && ok || ko is deliberate (ok/ko never fail)
set -u set -u
R="$(cd "$(dirname "$0")/../.." && pwd)" R="$(cd "$(dirname "$0")/../.." && pwd)"
REG="$R/mods/model-router/hooks/register.ts"
pass=0; fail=0 pass=0; fail=0
ok() { pass=$((pass+1)); } ok() { pass=$((pass+1)); }
ko() { fail=$((fail+1)); printf 'FAIL %s\n' "$1"; } ko() { fail=$((fail+1)); printf 'FAIL %s\n' "$1"; }
@@ -11,120 +14,133 @@ has() { if grep -qF "$2" "$R/$1"; then ok; else ko "$1 missing: $2"; fi; }
lacks() { if grep -qF "$2" "$R/$1"; then ko "$1 must NOT contain: $2"; else ok; fi; } lacks() { if grep -qF "$2" "$R/$1"; then ko "$1 must NOT contain: $2"; else ok; fi; }
# frontmatter = the lines between the first two '---' lines # frontmatter = the lines between the first two '---' lines
fm() { awk 'NR==1&&/^---$/{p=1;next} p&&/^---$/{exit} p' "$1"; } fm() { awk 'NR==1&&/^---$/{p=1;next} p&&/^---$/{exit} p' "$1"; }
fm_effort() { fm "$1" | grep -E '^effort: (low|medium|high|xhigh|max)$' | head -1 | cut -d' ' -f2; } fm_val() { fm "$1" | grep -E "^$2: [a-z]+$" | head -1 | cut -d' ' -f2; }
fm_has_effort() {
got="$(fm_effort "$R/$1")" # ── register.ts parsers (awk/sed on the DEFAULT_CONFIG literal) ──────────
if [ "$got" = "$2" ]; then ok; else ko "$1 frontmatter effort must be '$2', got '${got:-none}'"; fi # rows <file> <agents|skills> -> "name phase" per row
rows() {
awk -v s="$2" '$0 ~ "^ "s": \\{"{f=1;next} f&&/^ \},?$/{f=0} f' "$1" \
| grep -v '^ *//' | grep -oE "('[^']+'|[A-Za-z0-9_-]+): '[a-z]+'" \
| sed -E "s/'//g; s/: / /"
} }
fm_no_effort() { if fm "$R/$1" | grep -q '^effort:'; then ko "$1 must NOT pin effort"; else ok; fi; } # phase_effort <file> <phase> -> "<tier> <effort>"
phase_effort() {
awk '/^ phases: \{/{f=1;next} f&&/^ \},?$/{f=0} f' "$1" \
| sed -nE "s/^ *$2: \{ tier: '([a-z]+)', effort: '([a-z]+)' \},?$/\1 \2/p"
}
# tier_head <file> <tier> -> first alias of the tier list
tier_head() {
awk '/^ tiers: \{/{f=1;next} f&&/^ \},?$/{f=0} f' "$1" \
| sed -nE "s/^ *$2: \['([a-z]+)'.*$/\1/p"
}
row_of() { rows "$REG" "$1" | awk -v n="$2" '$1==n{print $2}'; }
# ── flip-test: the frontmatter reader must accept a valid level and reject an invalid one # ── flip-test: the parsers read a fixture, reject a missing key ──────────
FIX="$(mktemp -d)"; trap 'rm -rf "$FIX"' EXIT FIX="$(mktemp -d)"; trap 'rm -rf "$FIX"' EXIT
printf -- '---\nname: good\neffort: xhigh\n---\nbody with effort: low in prose\n' > "$FIX/good.md" cat > "$FIX/reg.ts" <<'FX'
printf -- '---\nname: bad\neffort: turbo\n---\n' > "$FIX/bad.md" tiers: {
[ "$(fm_effort "$FIX/good.md")" = "xhigh" ] && ok || ko "flip: valid level not read" big: ['opus', 'fable'],
[ -z "$(fm_effort "$FIX/bad.md")" ] && ok || ko "flip: invalid level accepted" },
[ "$(fm "$FIX/good.md" | grep -c 'prose')" -eq 0 ] && ok || ko "flip: body leaked into frontmatter" phases: {
judge: { tier: 'big', effort: 'xhigh' },
},
agents: {
// judge
Plan: 'judge', 'plan-challenger': 'judge',
},
skills: {
'ship-feature': 'plan', doc: 'apply',
},
FX
[ "$(rows "$FIX/reg.ts" agents | tr '\n' ,)" = "Plan judge,plan-challenger judge," ] \
&& ok || ko "flip: agents rows misparsed"
[ "$(rows "$FIX/reg.ts" skills | tr '\n' ,)" = "ship-feature plan,doc apply," ] \
&& ok || ko "flip: skills rows misparsed"
[ "$(phase_effort "$FIX/reg.ts" judge)" = "big xhigh" ] && ok || ko "flip: phase"
[ -z "$(phase_effort "$FIX/reg.ts" nothere)" ] && ok || ko "flip: ghost phase"
[ "$(tier_head "$FIX/reg.ts" big)" = "opus" ] && ok || ko "flip: tier head"
[ "$(rows "$REG" skills | wc -l)" -gt 40 ] && ok || ko "register.ts: skills rows not parsed"
[ "$(rows "$REG" agents | wc -l)" -gt 15 ] && ok || ko "register.ts: agents rows not parsed"
# ── 1) session default (spec D1) # ── (b) tracked skills: row exists, frontmatter effort equals the row ────
NO_ROW_SKILLS=" find-docs graphify impeccable model-router "
check_skill() {
local f="$1" name phase want got
name="$(basename "$(dirname "$f")")"
case "$NO_ROW_SKILLS" in *" $name "*) return;; esac
phase="$(row_of skills "$name")"
[ -n "$phase" ] || { ko "skills/$name: no row in register.ts"; return; }
want="$(phase_effort "$REG" "$phase" | cut -d' ' -f2)"
got="$(fm_val "$R/$f" effort)"
[ -n "$got" ] || { ko "skills/$name: routed skill without effort:"; return; }
[ "$got" = "$want" ] && ok || ko "skills/$name: effort $got != row $phase ($want)"
}
while IFS= read -r f; do check_skill "$f"; done < <(
cd "$R" && git ls-files 'skills/*/SKILL.md' 'skills-external/*/SKILL.md')
# ── (c) tracked agents with a row: tier head == model:, effort == effort: ─
NO_ROW_AGENTS=" interviewer client-handover-writer "
tier_alias() { tier_head "$REG" "$(phase_effort "$REG" "$1" | cut -d' ' -f1)"; }
check_agent() {
local f="$1" name phase alias want got
name="$(basename "$f" .md)"
case "$NO_ROW_AGENTS" in *" $name "*) return;; esac
case "$name" in impeccable-*) return;; esac
phase="$(row_of agents "$name")"
[ -n "$phase" ] || { ko "agents/$name: no row in register.ts"; return; }
alias="$(tier_alias "$phase")"; got="$(fm_val "$R/$f" model)"
[ "$got" = "$alias" ] && ok || ko "agents/$name: model $got != row $phase ($alias)"
[ "$alias" = haiku ] && return
want="$(phase_effort "$REG" "$phase" | cut -d' ' -f2)"
got="$(fm_val "$R/$f" effort)"
[ "$got" = "$want" ] && ok || ko "agents/$name: effort $got != row $phase ($want)"
}
while IFS= read -r f; do check_agent "$f"; done < <(
cd "$R" && git ls-files 'agents/*.md' | grep -E '^agents/[^/]+\.md$' \
| grep -v '/README\.md$')
# ── (d) no shifter citer in skills, agents, lib ──────────────────────────
if (cd "$R" && git grep -qE 'Skill\(effort-|EFFORT SHIFT[S]:' -- skills agents lib ':!lib/tests'); then
ko "a shifter-skill or shift-header citer survives in skills/agents/lib"
else ok; fi
# paths split so this file itself matches no deleted-name grep
for s in skills/effort-low skills/effort-max lib/effort-""pins.txt lib/model-""check.sh; do
[ ! -e "$R/$s" ] && ok || ko "$s must be deleted"
done
# ── (e) D3 wiring markers ────────────────────────────────────────────────
mark() { # mark <phase> <skills...>
local ph="$1" s; shift
for s in "$@"; do has "skills/$s/SKILL.md" "route(phase=\"$ph\")"; done
}
mark orchestrate feat hotfix bugfix ship-feature init-project code-clean seo geo harden web-validate audit-delta
mark apply feat hotfix bugfix ship-feature init-project
mark reflect feat hotfix bugfix seo geo harden web-validate
mark plan ship-feature init-project onboard code-clean audit-delta
mark escalate ship-feature
[ "$(grep -o 'route(phase="escalate")' "$R/lib/verify-secure-loop.md" | wc -l)" -eq 3 ] \
&& ok || ko "verify-secure-loop.md: escalate route must appear 3 times"
for s in ship-feature init-project; do has "skills/$s/SKILL.md" 'effort="xhigh"'; done
has "skills/tour/SKILL.md" 'effort="xhigh"'
n_opus="$(grep -c 'model="opus",$' "$R/skills/onboard/SKILL.md")"
n_eff="$(grep -c 'effort="xhigh",$' "$R/skills/onboard/SKILL.md")"
[ "$n_opus" -ge 7 ] && [ "$n_opus" -eq "$n_eff" ] && ok \
|| ko "onboard: $n_opus model=\"opus\" dispatches vs $n_eff effort=\"xhigh\""
bad="$(grep 'model: "fable"' "$R/agents/client-handover-writer.md" | grep -vc 'effort="high"')"
[ "$bad" -eq 0 ] && ok || ko "client-handover-writer: $bad model: \"fable\" line(s) without effort=\"high\""
for s in ship-feature init-project feat bugfix web-validate seo hotfix geo harden code-clean audit-delta tour onboard; do
has "skills/$s/SKILL.md" 'ROUTING: follow $HOME/.claude/lib/effort-shift.md'
done
has "agents/client-handover-writer.md" 'ROUTING: follow $HOME/.claude/lib/effort-shift.md'
# ── (f) doctrine includes ────────────────────────────────────────────────
has "lib/model-gate.md" 'mcp__model-router__route'
has "lib/effort-shift.md" 'ToolSearch'
has "CLAUDE.global.md" 'route to `reflect` (high) through their'
# ── (g) session default and hooks (unchanged locks) ──────────────────────
has "settings.json" '"effortLevel": "high"' has "settings.json" '"effortLevel": "high"'
# ── 2) hooks: env-var warning + live effort in the statusline (spec D1, D5)
has "hooks/session-start.sh" 'CLAUDE_CODE_EFFORT_LEVEL' has "hooks/session-start.sh" 'CLAUDE_CODE_EFFORT_LEVEL'
has "hooks/statusline.sh" 'CLAUDE_EFFORT' has "hooks/statusline.sh" 'CLAUDE_EFFORT'
# ── 3) agent pins (spec D2): one effort per agent file, judgment mode wins on mode-based agents printf 'PASS=%s FAIL=%s\n' "$pass" "$fail"; [ "$fail" -eq 0 ]
for a in hotfixer release-executor plugin-probe validator-analyzer; do fm_has_effort "agents/$a.md" low; done
for a in feater bugfixer code-cleaner onboarder scaffolder; do fm_has_effort "agents/$a.md" medium; done
for a in refactorer analyzer commit-changer doc-syncer handover-doc-writer; do fm_has_effort "agents/$a.md" high; done
for a in plan-challenger plugin-advisor verifier security-auditor seo-analyzer geo-analyzer; do fm_has_effort "agents/$a.md" xhigh; done
for a in interviewer client-handover-writer status-reporter; do fm_no_effort "agents/$a.md"; done
has "skills/init-project/SKILL.md" 'pin sonnet, effort medium'
# ── 4) skill entry levels (spec D3): the user's invocation sets the run's level
for s in status commit-change release-candidate doc capitalize close reconcile deploy profile plugin-check; do fm_has_effort "skills/$s/SKILL.md" low; done
for s in gitflow prune-memory; do fm_has_effort "skills/$s/SKILL.md" medium; done
for s in feat hotfix bugfix refactor web-validate harden seo geo; do fm_has_effort "skills/$s/SKILL.md" high; done
for s in ship-feature init-project onboard tour audit-delta analyze code-clean client-handover; do fm_has_effort "skills/$s/SKILL.md" xhigh; done
# BDR-108 round: the three repo skills that had no level
fm_has_effort "skills/skills-perso/SKILL.md" low
fm_has_effort "skills/pdf-translate/SKILL.md" medium
fm_has_effort "skills/site-motion/SKILL.md" high
# ── 9) vendored externals carry the level of lib/effort-pins.txt (BDR-108). The files live in
# skills-external/ (gitignored, machine-owned): the durable artifact is the map + the re-apply
# after the last vendoring step of install-plugins.sh AND update-all.sh; a skill not vendored
# yet SKIPs visibly (fresh clone before make plugin).
while read -r s lvl _; do
case "$s" in ''|'#'*) continue ;; esac
if [ -f "$R/skills-external/$s/SKILL.md" ]; then fm_has_effort "skills-external/$s/SKILL.md" "$lvl"
else printf 'SKIP skills-external/%s/SKILL.md not vendored yet (run make plugin)\n' "$s"; fi
done < "$R/lib/effort-pins.txt"
has "lib/effort-pins.txt" 'brainstorming xhigh'; has "lib/effort-pins.txt" 'writing-plans xhigh'
has "install-plugins.sh" 'apply_effort_pins "$REPO"'; has "update-all.sh" 'apply_effort_pins "$REPO"'
lacks "install-plugins.sh" 'for _s in brainstorming writing-plans; do'
ln_last() { grep -n "$2" "$R/$1" | tail -1 | cut -d: -f1; }
[ "$(ln_last install-plugins.sh 'apply_effort_pins "$REPO"')" -gt "$(ln_last install-plugins.sh 'rm -rf "$TFD_STAGE"')" ] \
&& ok || ko "install-plugins.sh: effort pins must be re-applied after the 21st pack refresh"
pins_ln=$(ln_last update-all.sh 'apply_effort_pins "$REPO"')
[ "$pins_ln" -gt "$(ln_last update-all.sh 'skills-external/$_tfd_name')" ] \
&& [ "$pins_ln" -gt "$(ln_last update-all.sh 'vendor_pinned_skills superpowers refresh')" ] \
&& ok || ko "update-all.sh: effort pins must be re-applied after the last vendoring step (21st pack)"
[ -x "$R/lib/effort-pins.sh" ] && ok || ko "lib/effort-pins.sh missing or not executable"
# 9b) design stack = ONE level (last loaded wins); site-motion (repo skill) pins the same one
stack_levels() { awk '/^# design stack/{f=1;next} f&&/^#$/{f=0} f&&!/^#/&&NF==2{print $2}' "$R/lib/effort-pins.txt" | sort -u; }
[ "$(stack_levels | wc -l)" -eq 1 ] && ok || ko "design stack must share ONE level in lib/effort-pins.txt (got: $(stack_levels | tr '\n' ' '))"
[ "$(stack_levels | wc -l)" -ge 1 ] && fm_has_effort "skills/site-motion/SKILL.md" "$(stack_levels | head -1)"
has "lib/effort-shift.md" 'Stacked skills share one level'
has "CLAUDE.global.md" 'lib/effort-pins.txt'
# ── 5) shifter skills + include (spec D4)
for l in low medium high xhigh max; do fm_has_effort "skills/effort-$l/SKILL.md" "$l"; has "skills/effort-$l/SKILL.md" "name: effort-$l"; done
has "lib/effort-shift.md" 'Headless sessions'
has "lib/effort-shift.md" 'Skill(effort-max)'
has "lib/effort-shift.md" 'never inside a dispatched agent'
has "lib/model-gate.md" 'lib/effort-shift.md'
# ── 6) orchestrator wiring (spec D4)
for s in feat hotfix bugfix ship-feature init-project onboard tour code-clean seo geo harden web-validate audit-delta; do
has "skills/$s/SKILL.md" 'lib/effort-shift.md'; has "skills/$s/SKILL.md" 'a lone Skill call is a no-op'; done
for s in feat hotfix bugfix ship-feature init-project code-clean seo geo harden web-validate audit-delta; do
has "skills/$s/SKILL.md" 'Skill(effort-medium)'; done
lacks "skills/onboard/SKILL.md" 'Skill(effort-medium)'; lacks "skills/tour/SKILL.md" 'Skill(effort-medium)'
has "agents/client-handover-writer.md" 'lib/effort-shift.md'; lacks "agents/client-handover-writer.md" 'Skill(effort-medium)'; has "agents/client-handover-writer.md" 'Skill(effort-high)'
for s in feat hotfix bugfix; do has "skills/$s/SKILL.md" 'Skill(effort-high)'; done
for s in ship-feature init-project onboard code-clean audit-delta; do has "skills/$s/SKILL.md" 'Skill(effort-xhigh)'; done
for s in seo geo harden web-validate; do has "skills/$s/SKILL.md" 'Skill(effort-high)'; done
for s in feat hotfix bugfix ship-feature init-project; do has "skills/$s/SKILL.md" 'Skill(effort-low)'; done
has "skills/feat/SKILL.md" 'effort-shift: nested commit-change'
# ── 6b) pairing rule documented (R11)
has "lib/effort-shift.md" 'lone Skill call is a no-op'
has "lib/effort-shift.md" 're-applies its'
[ "$(grep -c 'a lone Skill call is a no-op' "$R/skills/feat/SKILL.md")" -ge 1 ] && ok || ko "feat INC line must carry the pairing rule"
# ── 7) escalation at max (spec D4)
[ "$(grep -c 'Skill(effort-max)' "$R/lib/verify-secure-loop.md")" -eq 3 ] && ok || ko "verify-secure-loop.md must shift to max at its 3 caps"
has "skills/ship-feature/SKILL.md" 'Skill(effort-max)'
has "lib/challenge-plan.md" '/effort-max'
has "lib/verify-secure-loop.md" '/effort-max'
# ── 8) turn-reset re-assert after a prose gate followed by reflection
has "skills/bugfix/SKILL.md" 'effort-shift: turn reset'
# ── 11) audit tooling
has "lib/effort-shift.md" 'effort-audit.py'
[ -x "$R/lib/effort-audit.py" ] && ok || ko "lib/effort-audit.py missing or not executable"
# ── 6c) judgment dispatches re-raised, planning re-asserts, stronger locks (final review I1/I2/M5)
for s in ship-feature init-project; do has "skills/$s/SKILL.md" 'effort-shift: judgment dispatch'; has "skills/$s/SKILL.md" 'effort-shift: turn reset'; done
has "agents/client-handover-writer.md" 'effort-shift: judgment dispatch'
has "lib/effort-shift.md" 'Before any built-in or unpinned dispatch'
has "lib/model-gate.md" 'built-ins inherit the effort in force'
has "skills/ship-feature/SKILL.md" 'effort-shift: error recovery'
for s in feat hotfix bugfix seo geo harden web-validate ship-feature init-project onboard code-clean audit-delta; do has "skills/$s/SKILL.md" 'effort-shift: own level before the challenge'; done
has "update-all.sh" 'source "$REPO/lib/effort-pins.sh"'
# ── summary (later tasks insert their locks ABOVE this line)
printf 'effort-routing census: %d pass, %d fail\n' "$pass" "$fail"
[ "$fail" -eq 0 ]
+5 -11
View File
@@ -321,24 +321,20 @@ expect link-sh "$(count link.sh higgsfield)" 0
expect profile-sh "$(count lib/profile.sh higgsfield)" 0 expect profile-sh "$(count lib/profile.sh higgsfield)" 0
expect profiles \ expect profiles \
"$(cat "$ROOT"/lib/profiles/*.profile | grep -cF higgsfield)" 0 "$(cat "$ROOT"/lib/profiles/*.profile | grep -cF higgsfield)" 0
expect pins-map "$(count lib/effort-pins.txt higgsfield)" 0
verdict OFF_BY_DEFAULT_WIRING verdict OFF_BY_DEFAULT_WIRING
# ln_first / ln_last <file> <fixed string> — line number of a match. # ln_first / ln_last <file> <fixed string> — line number of a match.
ln_first() { grep -nF -- "$2" "$ROOT/$1" | head -1 | cut -d: -f1; } ln_first() { grep -nF -- "$2" "$ROOT/$1" | head -1 | cut -d: -f1; }
ln_last() { grep -nF -- "$2" "$ROOT/$1" | tail -1 | cut -d: -f1; } ln_last() { grep -nF -- "$2" "$ROOT/$1" | tail -1 | cut -d: -f1; }
PINS="apply_effort_pins \"\$REPO\""
# install-plugins.sh: the sync sits in Step 8.6, before the effort pins # install-plugins.sh: the sync sits in Step 8.6; the CLI is proven by a
# (BDR-108); the CLI is proven by a probe, not by its shim; every login # probe, not by its shim; every login offer tests stdin alone (stdout is
# offer tests stdin alone (stdout is the tee pipe). # the tee pipe).
sync_ln="$(ln_last install-plugins.sh 'higgsfield_sync_skills')" sync_ln="$(ln_last install-plugins.sh 'higgsfield_sync_skills')"
expect after-8.5 "$(yn test "$sync_ln" -gt \ expect after-8.5 "$(yn test "$sync_ln" -gt \
"$(ln_first install-plugins.sh 'Step 8.5: External skills')")" yes "$(ln_first install-plugins.sh 'Step 8.5: External skills')")" yes
expect before-8.7 "$(yn test "$sync_ln" -lt \ expect before-8.7 "$(yn test "$sync_ln" -lt \
"$(ln_first install-plugins.sh 'Step 8.7: 21st.dev')")" yes "$(ln_first install-plugins.sh 'Step 8.7: 21st.dev')")" yes
expect before-pins "$(yn test "$sync_ln" -lt \
"$(ln_last install-plugins.sh "$PINS")")" yes
expect probe-gates \ expect probe-gates \
"$(yn test "$(count install-plugins.sh 'if higgsfield_cli_ok')" -ge 3)" yes "$(yn test "$(count install-plugins.sh 'if higgsfield_cli_ok')" -ge 3)" yes
expect control "$(echo 'if [ -t 0 ] && [ -t 1 ]; then' | grep -cF -- '-t 1')" 1 expect control "$(echo 'if [ -t 0 ] && [ -t 1 ]; then' | grep -cF -- '-t 1')" 1
@@ -347,14 +343,12 @@ expect stdin-tests \
"$(yn test "$(count install-plugins.sh '[ -t 0 ]')" -ge 3)" yes "$(yn test "$(count install-plugins.sh '[ -t 0 ]')" -ge 3)" yes
verdict INSTALL_WIRING verdict INSTALL_WIRING
# update-all.sh: refresh before the 21st block and before the pins re-apply, # update-all.sh: refresh before the 21st block, and the updated CLI is
# and the updated CLI is proven by the probe, after the npm call. # proven by the probe, after the npm call.
NPM_UP="npm install -g \"\$HF_PKG\"" NPM_UP="npm install -g \"\$HF_PKG\""
sync_ln="$(ln_last update-all.sh 'higgsfield_sync_skills')" sync_ln="$(ln_last update-all.sh 'higgsfield_sync_skills')"
expect before-21st "$(yn test "$sync_ln" -lt \ expect before-21st "$(yn test "$sync_ln" -lt \
"$(ln_first update-all.sh '7.4. Update the 21st.dev')")" yes "$(ln_first update-all.sh '7.4. Update the 21st.dev')")" yes
expect before-pins "$(yn test "$sync_ln" -lt \
"$(ln_last update-all.sh "$PINS")")" yes
expect probe-after-npm "$(yn test \ expect probe-after-npm "$(yn test \
"$(ln_first update-all.sh 'higgsfield_cli_ok')" -gt \ "$(ln_first update-all.sh 'higgsfield_cli_ok')" -gt \
"$(ln_last update-all.sh "$NPM_UP")")" yes "$(ln_last update-all.sh "$NPM_UP")")" yes
-25
View File
@@ -1,25 +0,0 @@
#!/usr/bin/env bash
# lib/tests/model-check.test.sh — flip-tests for lib/model-check.sh (LRN-096)
set -u
S="$(cd "$(dirname "$0")/../.." && pwd)/lib/model-check.sh"
pass=0; fail=0
check() { if [ "$2" = "$3" ]; then pass=$((pass+1)); else fail=$((fail+1));
printf 'FAIL %s: got[%s] want[%s]\n' "$1" "$2" "$3"; fi; }
T="$(mktemp -d)"; trap 'rm -rf "$T"' EXIT
fx() { printf '{"model": "%s"}' "$1" > "$T/s.json"; }
run() { MODEL_CHECK_SETTINGS="$T/s.json" bash "$S" >"$T/out" 2>&1; echo "$?"; }
fx 'claude-fable-5[1m]'; check T1-fable-exit "$(run)" 0
check T1-fable-class "$(cut -d: -f1 <"$T/out")" big
fx 'claude-opus-4-8'; check T2-opus "$(run)" 0
fx 'claude-sonnet-5'; check T3-sonnet "$(run)" 2
fx 'claude-haiku-4-5-20251001'; check T4-haiku "$(run)" 2
fx 'opusplan'; check T5-opusplan "$(run)" 3
fx 'gpt-9-mega'; check T6-foreign "$(run)" 3
printf '{"no_model": true}' > "$T/s.json"; check T7-no-key "$(run)" 3
printf '{broken' > "$T/s.json"; check T8-malformed "$(run)" 3
check T9-missing-file "$(MODEL_CHECK_SETTINGS="$T/absent.json" bash "$S" >/dev/null 2>&1; echo $?)" 3
printf 'model-check: %d pass, %d fail\n' "$pass" "$fail"
[ "$fail" -eq 0 ]
+5 -4
View File
@@ -35,7 +35,7 @@ single `GATES — VERDICT:` line:
- `UNMET(n)` → hand the dev the CONTRACT path + the `NOT-MET` rows verbatim, - `UNMET(n)` → hand the dev the CONTRACT path + the `NOT-MET` rows verbatim,
nothing else; re-run GATE 0. **No verifier is dispatched** — a red build or nothing else; re-run GATE 0. **No verifier is dispatched** — a red build or
a red suite is not a judgement call, and paying an LLM to discover it is a red suite is not a judgement call, and paying an LLM to discover it is
waste. **Max 3 floor iterations** → `Skill(effort-max)` (effort-shift: cap reached, diagnose at max before escalating; send it in the same message as the first tool call that gathers the escalation evidence), then STOP + human escalation with the rows. waste. **Max 3 floor iterations** → `mcp__model-router__route(phase="escalate")` (route: cap reached, diagnose at max before escalating; send it with the first tool call that gathers the escalation evidence), then STOP + human escalation with the rows.
- `ABANDONED(n)` → floor green but a handoff stands. Continue to GATE 1; the - `ABANDONED(n)` → floor green but a handoff stands. Continue to GATE 1; the
verifier surfaces it and its `ABANDONED(n)` verdict routes to the human verifier surfaces it and its `ABANDONED(n)` verdict routes to the human
gate. gate.
@@ -74,7 +74,7 @@ Parse its single `VERIFY — VERDICT:` line:
lines (NOT-MET / out-of-scope), nothing else: re-dispatch a FRESH executor lines (NOT-MET / out-of-scope), nothing else: re-dispatch a FRESH executor
with those inputs only, never redo the fix by hand. Then re-run GATE 0 and with those inputs only, never redo the fix by hand. Then re-run GATE 0 and
re-dispatch a FRESH verifier. Repeat. re-dispatch a FRESH verifier. Repeat.
**Max 3 conformity iterations** → `Skill(effort-max)` (effort-shift: cap reached, diagnose at max before escalating; send it in the same message as the first tool call that gathers the escalation evidence), then STOP + human escalation with the **Max 3 conformity iterations** → `mcp__model-router__route(phase="escalate")` (route: cap reached, diagnose at max before escalating; send it with the first tool call that gathers the escalation evidence), then STOP + human escalation with the
CRITERIA table (the contract-vs-realized diff). CRITERIA table (the contract-vs-realized diff).
- `ABANDONED(n)` → direct human gate, never a dev loop (a dev cannot close - `ABANDONED(n)` → direct human gate, never a dev loop (a dev cannot close
what was proven impossible). The human lifts the abandonment or accepts what was proven impossible). The human lifts the abandonment or accepts
@@ -104,9 +104,10 @@ Parse its single `SECURITY — VERDICT:` line:
(re-dispatch a FRESH executor, never fix by hand). Then re-run GATE 0, then (re-dispatch a FRESH executor, never fix by hand). Then re-run GATE 0, then
**re-verify the REQUEST first** (GATE 1, fresh verifier) — a security fix **re-verify the REQUEST first** (GATE 1, fresh verifier) — a security fix
can drift the behavior — **then re-run GATE 2** (fresh auditor), in that can drift the behavior — **then re-run GATE 2** (fresh auditor), in that
order. **Max 3 security iterations** → `Skill(effort-max)` (effort-shift: cap reached, diagnose at max before escalating; send it in the same message as the first tool call that gathers the escalation evidence), then STOP + human escalation with the order. **Max 3 security iterations** → `mcp__model-router__route(phase="escalate")` (route: cap reached, diagnose at max before escalating; send it with the first tool call that gathers the escalation evidence), then STOP + human escalation with the
BLOCKING table. Every STOP text names the level reached (`$CLAUDE_EFFORT`) BLOCKING table. Every STOP text names the level reached (`$CLAUDE_EFFORT`)
and suggests `/effort-max` for the relaunch. and suggests relaunching with `ultrathink` in the prompt (turn floor) or
`/route effort=max` (sticky, `/route clear` after).
- `DEGRADED` (semgrep absent) → does NOT block on the tool's absence; surface - `DEGRADED` (semgrep absent) → does NOT block on the tool's absence; surface
the checklist result + recommend `make plugin`. A DEGRADED run that still the checklist result + recommend `make plugin`. A DEGRADED run that still
BLOCKs (grep-caught secret/injection) blocks like any other. BLOCKs (grep-caught secret/injection) blocks like any other.
+459 -75
View File
@@ -23,6 +23,8 @@ async function boot(
on('session.model', () => ({ value: model })) on('session.model', () => ({ value: model }))
on('classic.StopFailure', () => ({})) on('classic.StopFailure', () => ({}))
on('classic.PostModelSwitch', () => ({})) on('classic.PostModelSwitch', () => ({}))
on('skill.prompt', ($, e) => ({ text: e.text }))
on('agent.offer', () => ({ isOffered: true }))
if (tokens !== null) usageOf(on, tokens) if (tokens !== null) usageOf(on, tokens)
on('session.start', ($, e) => ({ cwd: e.cwd })) on('session.start', ($, e) => ({ cwd: e.cwd }))
await $.session.start({ cwd: '/tmp', surface: null, isInteractive: false }) await $.session.start({ cwd: '/tmp', surface: null, isInteractive: false })
@@ -57,24 +59,6 @@ const spawnInput = (model?: string) => ({
...(model === undefined ? {} : { model }), ...(model === undefined ? {} : { model }),
}) })
test('Skill(effort-low) is answered without next, route shows low', async (
$, on) => {
let reached = false
on('tool.call', { tool: 'Skill' }, () => {
reached = true
return { result: { success: true, commandName: 'bottom' } }
})
await boot($, on)
const out = await $.tool.call({ tool: 'Skill', skill: 'effort-low' })
expect(out).toMatchObject({
result: { success: true, commandName: 'effort-low' },
})
expect(reached).toBe(false)
const line = mainLine(await route($, 'show'))
expect(line).toContain('skill effort-low')
expect(line).toContain('effort low')
})
test('route tool with phase orchestrate sets medium on main', async ( test('route tool with phase orchestrate sets medium on main', async (
$, on) => { $, on) => {
await boot($, on) await boot($, on)
@@ -321,30 +305,14 @@ test('floor: ultrathink survives a model route', async ($, on) => {
expect(await stepEffort($, seen)).toBe('max') expect(await stepEffort($, seen)).toBe('max')
}) })
test('floor: typed /effort-medium clamps low, lets max pass', async ( test('floor: survives a skill load; an unrowed skill keeps the route', async (
$, on) => { $, on) => {
const seen = await bootFloor($, on)
await $.skill.prompt({ skill: 'effort-medium', text: 'x' })
await $.tool.call({ tool: ROUTE_TOOL, phase: 'mechanical' })
expect(await stepEffort($, seen)).toBe('medium')
await $.tool.call({ tool: ROUTE_TOOL, phase: 'escalate' })
expect(await stepEffort($, seen)).toBe('max')
})
test('floor: typed /effort-low lowers an unrouted turn', async ($, on) => {
const seen = await bootFloor($, on)
const out = await $.skill.prompt({ skill: 'effort-low', text: 'x' })
expect(out.text).toContain('minimum')
expect(await stepEffort($, seen)).toBe('low')
})
test('floor: survives a skill load', async ($, on) => {
const seen = await bootFloor($, on) const seen = await bootFloor($, on)
await ultrathink($) await ultrathink($)
await $.tool.call({ tool: ROUTE_TOOL, phase: 'orchestrate' }) await $.tool.call({ tool: ROUTE_TOOL, phase: 'orchestrate' })
await $.tool.call({ tool: 'Skill', skill: 'other' }) await $.tool.call({ tool: 'Skill', skill: 'other' })
expect(await stepEffort($, seen)).toBe('max') expect(await stepEffort($, seen)).toBe('max')
expect(mainLine(await route($, 'show'))).not.toContain('orchestrate') expect(mainLine(await route($, 'show'))).toContain('model orchestrate')
}) })
test('floor: lifts a lower sticky, then ends with the turn', async ( test('floor: lifts a lower sticky, then ends with the turn', async (
@@ -439,35 +407,6 @@ test('reload with an unreadable override keeps the previous config', async (
expect(await route($, 'show')).toContain('switch: on') expect(await route($, 'show')).toContain('switch: on')
}) })
test('typed marker: a preload in a live agent is ignored', async ($, on) => {
const seen = await bootFloor($, on)
await $.agent.spawn(spawnInput())
const out = await $.skill.prompt({ skill: 'effort-max', text: 'x' })
expect(out.text?.startsWith('model-router: effort-max preload')).toBe(true)
expect(mainLine(await route($, 'show'))).not.toContain('floor')
expect(await stepEffort($, seen)).not.toBe('max')
})
test('typed marker: /effort-max seen at submit writes the floor', async (
$, on) => {
await bootFloor($, on)
await $.agent.spawn(spawnInput())
await $.prompt.submit({
text: '/effort-max go',
wait: false,
origin: { kind: 'composer' },
})
await $.skill.prompt({ skill: 'effort-max', text: 'x' })
expect(mainLine(await route($, 'show'))).toContain('floor max')
})
test('typed slash with no agent and no marker writes the floor', async (
$, on) => {
await bootFloor($, on)
await $.skill.prompt({ skill: 'effort-max', text: 'x' })
expect(mainLine(await route($, 'show'))).toContain('floor max')
})
// ---- tiers, breaker, derived phases ------------------------------------- // ---- tiers, breaker, derived phases -------------------------------------
type Rig = { seen: Seen[]; clock: MockClock; specs: (string | undefined)[] } type Rig = { seen: Seen[]; clock: MockClock; specs: (string | undefined)[] }
@@ -766,16 +705,6 @@ test('default rule: a typed slash command gets no default rule', async (
expect(mainLine(await route($, 'show'))).toContain('session defaults') expect(mainLine(await route($, 'show'))).toContain('session defaults')
}) })
test('default rule: /effort-low pourquoi sets the floor, no default', async (
$, on) => {
await bootRig($, on)
await typed($, '/effort-low pourquoi ça plante')
await $.skill.prompt({ skill: 'effort-low', text: 'x' })
const line = mainLine(await route($, 'show'))
expect(line).toContain('floor low')
expect(line).not.toContain('reflect')
})
test('per axis: a model-less sticky never hides a turn route tier', async ( test('per axis: a model-less sticky never hides a turn route tier', async (
$, on) => { $, on) => {
const rig = await bootRig($, on, HAIKU) const rig = await bootRig($, on, HAIKU)
@@ -844,3 +773,458 @@ test('text: show names the model a floor upgrade moves to', async (
const line = mainLine(await route($, 'show')) const line = mainLine(await route($, 'show'))
expect(line).toContain(`${FABLE} (upgrade)`) expect(line).toContain(`${FABLE} (upgrade)`)
}) })
// ---- wave 2: skill rows, run slot, agent rows ---------------------------
const skillPrompt = ($: Engine, skill: string) =>
$.skill.prompt({ skill, text: 'x' })
const loadSkill = ($: Engine, skill: string, agentId?: string) =>
$.tool.call({
tool: 'Skill',
skill,
...(agentId === undefined ? {} : { agentId }),
})
/** bootRig plus a bottom Skill tool, so a model-loaded skill goes through. */
async function bootRun($: Engine, on: On, model: string = FABLE) {
on('tool.call', { tool: 'Skill' }, () => ({
result: { success: true, commandName: 'bottom' },
}))
return bootRig($, on, model)
}
const spawnOf = (agent: string, extra: object = {}) =>
({ ...spawnInput(), subagentType: agent, ...extra })
const offerOf = ($: Engine, agent: string, source: string) =>
$.agent.offer({
agent,
description: 'd',
source,
provider: { plugin: 'engine', tier: 'core' },
})
/**
* A ui.log recorder. The kit refuses a bottom hook for a void event, so the
* hook sits on top and swallows the missing implementation below it.
*/
function recordLogs(on: On): string[] {
const lines: string[] = []
on('ui.log', async ($, e, next) => {
lines.push(e.text)
try {
await next(e)
} catch {
// nothing implements ui.log below the recorder in the kit
}
})
return lines
}
test('typed /feat routes main to its row, best tier', async ($, on) => {
await bootRun($, on)
await typed($, '/feat add a thing')
await skillPrompt($, 'feat')
const line = mainLine(await route($, 'show'))
expect(line).toContain('skill reflect')
expect(line).toContain('effort high')
expect(line).toContain('[tier best]')
})
test('typed slash: a foreign origin arms nothing, fallback refused', async (
$, on) => {
await bootRun($, on)
await $.prompt.submit({
text: '/feat x',
wait: false,
origin: { kind: 'channel', server: 'slack' },
})
await skillPrompt($, 'feat')
expect(mainLine(await route($, 'show'))).toContain('session defaults')
})
test('typed slash: a preload in a live sub-agent leaves main alone', async (
$, on) => {
await bootRun($, on)
await typed($, 'hello')
await $.agent.spawn(spawnOf('feater'))
await skillPrompt($, 'feat')
expect(mainLine(await route($, 'show'))).toContain('session defaults')
})
test('typed slash: no live loop and an allowed origin still routes', async (
$, on) => {
await bootRun($, on)
await typed($, 'hello')
await skillPrompt($, 'feat')
expect(mainLine(await route($, 'show'))).toContain('skill reflect')
})
test('typed slash: an sdk origin arms the marker', async ($, on) => {
await bootRun($, on)
await $.prompt.submit({
text: '/status',
wait: false,
origin: { kind: 'sdk' },
})
await $.agent.spawn(spawnOf('feater'))
await skillPrompt($, 'status')
expect(mainLine(await route($, 'show'))).toContain('skill mechanical')
})
test('typed slash: a mid-turn /status waits for the turn end', async (
$, on) => {
await bootRun($, on)
await $.prompt.submit({
text: '/status',
wait: true,
origin: { kind: 'composer' },
turnId: 'u1',
})
await skillPrompt($, 'status')
expect(mainLine(await route($, 'show'))).toContain('session defaults')
await endTurn($)
await skillPrompt($, 'status')
expect(mainLine(await route($, 'show'))).toContain('skill mechanical')
})
test('skill without a row keeps the route; a rowed one replaces it', async (
$, on) => {
await bootRun($, on)
await $.tool.call({ tool: ROUTE_TOOL, phase: 'plan' })
await loadSkill($, 'find-docs')
expect(mainLine(await route($, 'show'))).toContain('model plan')
await loadSkill($, 'gitflow')
expect(mainLine(await route($, 'show'))).toContain('skill implement')
})
test('run slot: survives route calls and the turn end', async ($, on) => {
await bootRun($, on)
await loadSkill($, 'feat')
await $.tool.call({ tool: ROUTE_TOOL, phase: 'orchestrate' })
await endTurn($)
expect(mainLine(await route($, 'show'))).toContain('run reflect')
await $.tool.call({ tool: ROUTE_TOOL, phase: 'plan' })
await endTurn($)
expect(mainLine(await route($, 'show'))).toContain('run reflect')
})
test('run slot: a model-loaded helper skill keeps it', async ($, on) => {
await bootRun($, on)
await loadSkill($, 'feat')
await loadSkill($, 'using-git-worktrees')
expect(mainLine(await route($, 'show'))).toContain('skill mechanical')
await endTurn($)
expect(mainLine(await route($, 'show'))).toContain('run reflect')
})
test('run slot: a typed non-best skill drops it', async ($, on) => {
await bootRun($, on)
await loadSkill($, 'feat')
await endTurn($)
await typed($, '/status')
await skillPrompt($, 'status')
await endTurn($)
expect(mainLine(await route($, 'show'))).toContain('session defaults')
})
test('run slot: /route clear and /route off drop it', async ($, on) => {
await bootRun($, on)
await loadSkill($, 'feat')
expect(await route($, 'clear')).toContain('run slot dropped')
await endTurn($)
expect(mainLine(await route($, 'show'))).toContain('session defaults')
await loadSkill($, 'feat')
await route($, 'off')
await route($, 'on')
await endTurn($)
expect(mainLine(await route($, 'show'))).toContain('session defaults')
})
test('run slot: a user /model switch drops it, an auto one does not', async (
$, on) => {
await bootRun($, on)
await loadSkill($, 'feat')
await endTurn($)
await switchTo($, FABLE, OPUS, 'auto')
expect(mainLine(await route($, 'show'))).toContain('run reflect')
await switchTo($, FABLE, OPUS, 'command')
expect(mainLine(await route($, 'show'))).toContain('session defaults')
})
test('run slot: a Skill call inside a sub-agent never touches it', async (
$, on) => {
await bootRun($, on)
await loadSkill($, 'feat', 'a1')
await endTurn($)
expect(mainLine(await route($, 'show'))).toContain('session defaults')
})
test('run slot: route(clear) clears the turn and names the run', async (
$, on) => {
await bootRun($, on)
await loadSkill($, 'feat')
const out = await $.tool.call({ tool: ROUTE_TOOL, clear: true })
expect(JSON.stringify(out)).toContain('run reflect still holds')
})
test('route answer names the id even with a floor in force', async (
$, on) => {
await bootRun($, on)
await ultrathink($)
const out = await $.tool.call({ tool: ROUTE_TOOL, phase: 'plan' })
const text = JSON.stringify(out)
expect(text).toContain(`model ${FABLE}`)
expect(text).toContain('user floor max')
})
test('agent row: feater spawns on sonnet and steps at medium', async (
$, on) => {
const rig = await bootRun($, on)
const out = await $.agent.spawn(spawnOf('feater'))
expect(rig.specs).toEqual([SONNET])
expect(out.model).toBe(SONNET)
await runStep($, { ...highStep('a1'), model: SONNET })
expect(rig.seen[rig.seen.length - 1]).toEqual({
model: SONNET,
effort: 'medium',
})
})
/** Spawn feater, load a skill inside it, return the effort of its step. */
async function effortAfterSkill($: Engine, on: On, skill: string) {
const rig = await bootRun($, on)
await $.agent.spawn(spawnOf('feater'))
await loadSkill($, skill, 'a1')
await runStep($, { ...highStep('a1'), model: SONNET })
return rig.seen[rig.seen.length - 1].effort
}
test('agent skill: an unrowed one leaves the loop effort', async ($, on) => {
expect(await effortAfterSkill($, on, 'find-docs')).toBe('medium')
})
test('agent skill: a rowed one writes the loop effort', async ($, on) => {
expect(await effortAfterSkill($, on, 'status')).toBe('low')
})
test('agent row: judge moves up to fable when opus is down', async (
$, on) => {
const rig = await bootRun($, on)
await switchTo($, OPUS, SONNET, 'auto')
await $.agent.spawn(spawnOf('plan-challenger'))
expect(rig.specs).toEqual([FABLE])
})
test('agent row: opus and fable down, no write and one log line', async (
$, on) => {
const logs = recordLogs(on)
const rig = await bootRun($, on)
await switchTo($, OPUS, SONNET, 'auto')
await switchTo($, FABLE, SONNET, 'auto')
await $.agent.spawn(spawnOf('plan-challenger'))
await $.agent.spawn(spawnOf('plan-challenger'))
expect(rig.specs).toEqual([undefined, undefined])
const down = logs.filter(l => l.includes('frontmatter model kept'))
expect(down).toEqual([
'model-router: plan-challenger tier big down, frontmatter model kept',
])
})
test('agent row: explicit model, explicit effort and fork win', async (
$, on) => {
const rig = await bootRun($, on)
await $.agent.spawn(spawnOf('feater', { model: 'opus' }))
await $.agent.spawn(spawnOf('feater', { fork: true }))
expect(rig.specs).toEqual(['opus', undefined])
await $.tool.call({
tool: 'Agent',
description: 'd',
prompt: 'p',
effort: 'low',
tool_use_id: 't1',
})
await $.agent.spawn(spawnOf('feater'))
await runStep($, { ...highStep('a1'), model: SONNET })
expect(rig.seen[rig.seen.length - 1]?.effort).toBe('high')
})
test('agent row: a project definition keeps its own model', async ($, on) => {
const rig = await bootRun($, on)
await offerOf($, 'verifier', 'projectSettings')
await offerOf($, 'feater', 'userSettings')
await $.agent.spawn(spawnOf('verifier'))
await $.agent.spawn(spawnOf('feater'))
expect(rig.specs).toEqual([undefined, SONNET])
})
test('agent row: an override null drops the row', async ($, on) => {
on('env.get', () => ({ value: '/home/t' }))
on('fs.exists', () => ({ value: true }))
on('fs.stat', () => ({
value: { kind: 'file' as const, size: 40, mtimeMs: 0, isLink: false },
}))
on('fs.read', () => ({ value: '{"agents":{"verifier":null}}' }))
const rig = await bootRun($, on)
await $.agent.spawn(spawnOf('verifier'))
await $.agent.spawn(spawnOf('feater'))
expect(rig.specs).toEqual([undefined, SONNET])
})
test('show lists the write and apply phases', async ($, on) => {
await bootRun($, on)
const text = await route($, 'show')
expect(text).toContain('apply=work→claude-sonnet-5-5/low')
expect(text).toContain('write=work→claude-sonnet-5-5/high')
})
// ---- wave 2: typed-slash origins, spawning counter, spawn exclusions ----
/** A promise the test settles by hand, to hold a hook in flight. */
function deferred() {
let release: () => void = () => {}
const gate = new Promise<void>(r => { release = r })
return { gate, release }
}
test('typed slash: an in-flight spawn refuses the fallback', async (
$, on) => {
const hold = deferred()
const entered = deferred()
on('prompt.submit', ($, e) => ({ text: e.text }))
on('agent.spawn', async ($, e) => {
entered.release()
await hold.gate
return { model: e.model ?? e.parentModel, agentId: 'a1' }
})
await boot($, on)
await typed($, 'hello')
const spawned = $.agent.spawn(spawnOf('feater'))
await entered.gate
await skillPrompt($, 'feat')
expect(mainLine(await route($, 'show'))).toContain('session defaults')
hold.release()
await spawned
})
test('typed slash: a bridge origin arms the marker', async ($, on) => {
await bootRun($, on)
await $.prompt.submit({
text: '/feat x',
wait: false,
origin: { kind: 'bridge' },
})
await $.agent.spawn(spawnOf('feater'))
await skillPrompt($, 'feat')
expect(mainLine(await route($, 'show'))).toContain('skill reflect')
})
test('typed slash: a promoted mid-turn marker routes past a live loop', async (
$, on) => {
await bootRun($, on)
await $.agent.spawn(spawnOf('feater'))
await $.prompt.submit({
text: '/status',
wait: true,
origin: { kind: 'composer' },
turnId: 'u1',
})
await endTurn($)
await skillPrompt($, 'status')
expect(mainLine(await route($, 'show'))).toContain('skill mechanical')
})
test('agent row: any provider plugin applies, a local definition not', async (
$, on) => {
const rig = await bootRun($, on)
const provider = { plugin: 'some-plugin', tier: 'core' as const }
await $.agent.spawn(spawnOf('feater', { provider }))
await offerOf($, 'verifier', 'localSettings')
await $.agent.spawn(spawnOf('verifier', { provider }))
expect(rig.specs).toEqual([SONNET, undefined])
})
test('agent row: a workflow spawn is untouched, model and effort', async (
$, on) => {
const rig = await bootRun($, on)
await $.agent.spawn(spawnOf('feater', { workflow: {} }))
expect(rig.specs).toEqual([undefined])
await runStep($, { ...highStep('a1'), model: SONNET })
expect(rig.seen[rig.seen.length - 1]?.effort).toBe('high')
})
test('typed slash: the marker binds to its skill name only', async (
$, on) => {
await bootRun($, on)
await typed($, '/status')
await $.agent.spawn(spawnOf('feater'))
await skillPrompt($, 'feat')
expect(mainLine(await route($, 'show'))).toContain('session defaults')
})
test('typed slash: the marker is one-shot', async ($, on) => {
await bootRun($, on)
await typed($, '/status')
await $.agent.spawn(spawnOf('feater'))
await skillPrompt($, 'status')
expect(mainLine(await route($, 'show'))).toContain('skill mechanical')
await route($, 'clear')
await skillPrompt($, 'status')
expect(mainLine(await route($, 'show'))).toContain('session defaults')
})
test('skill.prompt nested in a Skill call leaves the run slot', async (
$, on) => {
on('tool.call', { tool: 'Skill' }, async (_api, e) => {
await skillPrompt($, 'status')
return { result: { success: true, commandName: String(e.skill) } }
})
await bootRig($, on, FABLE)
await typed($, '/status')
await loadSkill($, 'feat')
await endTurn($)
expect(mainLine(await route($, 'show'))).toContain('run reflect')
})
/** Submits a prompt from `kind` over a running turn (mid-turn). */
const queued = ($: Engine, kind: string, text: string) =>
$.prompt.submit({
text,
wait: true,
origin: { kind },
turnId: 'u1',
})
test('idle fallback: a mid-turn foreign prompt never arms it', async (
$, on) => {
await bootRun($, on)
await queued($, 'channel', 'hello')
await endTurn($)
await skillPrompt($, 'status')
expect(mainLine(await route($, 'show'))).toContain('session defaults')
})
test('idle fallback: a mid-turn composer prompt promotes the allowance', async (
$, on) => {
await bootRun($, on)
await $.prompt.submit({
text: 'hello',
wait: false,
origin: { kind: 'channel' },
})
await queued($, 'composer', 'hello')
await endTurn($)
await skillPrompt($, 'status')
expect(mainLine(await route($, 'show'))).toContain('skill mechanical')
})
test('spawning: a finished spawn no longer counts as a live loop', async (
$, on) => {
await bootRun($, on)
await $.agent.spawn(spawnOf('feater'))
await agentEnds($)
await typed($, 'hello')
await skillPrompt($, 'status')
expect(mainLine(await route($, 'show'))).toContain('skill mechanical')
})
+318 -147
View File
@@ -21,7 +21,7 @@ type Config = {
models: Record<string, string> // alias -> full id models: Record<string, string> // alias -> full id
windows: Record<string, number> // full id -> context window (tokens) windows: Record<string, number> // full id -> context window (tokens)
phases: Record<string, Route> phases: Record<string, Route>
agents: Record<string, string> // built-in subagentType -> phase agents: Record<string, string> // subagentType -> phase
skills: Record<string, string> // skill name -> phase skills: Record<string, string> // skill name -> phase
prompt: PromptRule[] prompt: PromptRule[]
tiers: Record<string, string[]> // tier -> alias preference, best first tiers: Record<string, string[]> // tier -> alias preference, best first
@@ -35,7 +35,7 @@ type Config = {
enabled: boolean // false: every hook passes through (per machine) enabled: boolean // false: every hook passes through (per machine)
} }
type Rule = { re: RegExp; phase: string; mode: PromptMode } type Rule = { re: RegExp; phase: string; mode: PromptMode }
type Source = 'user' | 'model' | 'skill' | 'prompt' | 'slash' | 'derived' type Source = 'user' | 'model' | 'skill' | 'prompt' | 'run' | 'derived'
type Routed = { phase: string; route: Route; source: Source } type Routed = { phase: string; route: Route; source: Source }
type Hold = { until: number; reason: string } // until: ms, Infinity = reload type Hold = { until: number; reason: string } // until: ms, Infinity = reload
type Pushed = { prev: Routed | null; spawnIds: Set<string> } type Pushed = { prev: Routed | null; spawnIds: Set<string> }
@@ -48,12 +48,19 @@ type State = {
rules: Rule[] rules: Rule[]
source: string // 'defaults' or the override path source: string // 'defaults' or the override path
userMain: Routed | null // /route by the user, sticky until /route clear userMain: Routed | null // /route by the user, sticky until /route clear
turnMain: Routed | null // model route tool, skill table row, Skill(effort-*) turnMain: Routed | null // model route tool, skill table row, prompt
// bridge, prompt default rule, derived orchestrate; dropped at turn end // default rule, derived orchestrate; dropped at turn end
turnFloor: Routed | null // user-explicit level for this turn (prompt rule, runMain: Routed | null // best-tier skill row: spans the turns of a run;
// typed /effort-<l>): a floor, main loop only // only a skill, /route clear|off or a user /model write or drop it
turnFloor: Routed | null // user-explicit level for this turn (prompt rule):
// a floor, main loop only
pendingPrompt: Routed | null // typed mid-turn: the next turn's floor pendingPrompt: Routed | null // typed mid-turn: the next turn's floor
typedSlash: boolean // one-shot: prompt.submit saw a typed /effort-<l> typedSlash: string | null // the rowed skill the user typed at prompt.submit
pendingSlash: string | null // same, typed mid-turn: promoted at turn end
promptAllowed: boolean // this turn's prompt came from a typing origin
pendingAllowed: boolean // same, for the mid-turn prompt
spawning: number // agent.spawn hooks in flight (preloads fire inside)
offers: Map<string, string> // subagentType -> definition source
loops: Map<string, Loop> // agentId -> that loop's routing loops: Map<string, Loop> // agentId -> that loop's routing
explicitEffort: Map<string, Level> // Agent tool_use_id -> effort param explicitEffort: Map<string, Level> // Agent tool_use_id -> effort param
skillCalls: number // Skill tool calls in flight skillCalls: number // Skill tool calls in flight
@@ -88,7 +95,19 @@ type Decision = { effort: Effort; by: EffortBy }
const LEVELS: readonly Level[] = ['low', 'medium', 'high', 'xhigh', 'max'] const LEVELS: readonly Level[] = ['low', 'medium', 'high', 'xhigh', 'max']
const MODEL_ID = /^claude-[a-z0-9.-]+$/ const MODEL_ID = /^claude-[a-z0-9.-]+$/
const TOOL = 'mcp__model-router__route' const TOOL = 'mcp__model-router__route'
const EFFORT_SKILL = /^effort-(low|medium|high|xhigh|max)$/ const BEST = 'best' // the tier whose skill rows hold for a whole run
// Origins that are a person typing: only these arm a typed slash.
const TYPED_ORIGINS: ReadonlySet<string> = new Set([
'composer', 'sdk', 'bridge',
])
// Definition sources of a foreign repo's own agent: its row is skipped.
const PROJECT_SOURCES: ReadonlySet<string> = new Set([
'projectSettings', 'localSettings',
])
// PostModelSwitch sources a person chose (auto and resume are not).
const USER_SWITCH: ReadonlySet<string> = new Set([
'command', 'picker', 'sdk',
])
const OVERRIDE = '.claude/model-router.json' const OVERRIDE = '.claude/model-router.json'
const HAIKU = 'claude-haiku' const HAIKU = 'claude-haiku'
const MAX_PATTERN = 200 // chars of a prompt-rule pattern const MAX_PATTERN = 200 // chars of a prompt-rule pattern
@@ -132,14 +151,69 @@ const DEFAULT_CONFIG: Config = {
escalate: { tier: 'best', effort: 'max' }, escalate: { tier: 'best', effort: 'max' },
judge: { tier: 'big', effort: 'xhigh' }, judge: { tier: 'big', effort: 'xhigh' },
implement: { tier: 'work', effort: 'medium' }, implement: { tier: 'work', effort: 'medium' },
write: { tier: 'work', effort: 'medium' }, write: { tier: 'work', effort: 'high' },
verify: { tier: 'work', effort: 'xhigh' }, verify: { tier: 'work', effort: 'xhigh' },
explore: { tier: 'work', effort: 'medium' }, explore: { tier: 'work', effort: 'medium' },
apply: { tier: 'work', effort: 'low' },
mechanical: { tier: 'cheap', effort: 'low' }, mechanical: { tier: 'cheap', effort: 'low' },
}, },
// Built-ins only: repo agents keep their frontmatter pin (wave 2). // phase = role; one row per routed repo skill/agent (wave 2); a
agents: { Explore: 'explore', Plan: 'judge' }, // project-level agent of the same name shadows its row (see spawnRoute).
skills: {}, agents: {
// explore
Explore: 'explore',
// judge
Plan: 'judge', 'plan-challenger': 'judge', 'plugin-advisor': 'judge',
'seo-analyzer': 'judge', 'geo-analyzer': 'judge', analyzer: 'judge',
// implement
feater: 'implement', bugfixer: 'implement', 'code-cleaner': 'implement',
scaffolder: 'implement', onboarder: 'implement',
// write
'commit-changer': 'write', 'doc-syncer': 'write',
'handover-doc-writer': 'write', refactorer: 'write',
// apply
hotfixer: 'apply', 'release-executor': 'apply', 'plugin-probe': 'apply',
'validator-analyzer': 'apply',
// verify
verifier: 'verify', 'security-auditor': 'verify',
// mechanical
'status-reporter': 'mechanical',
},
skills: {
// plan
'ship-feature': 'plan', 'init-project': 'plan', onboard: 'plan',
tour: 'plan', 'audit-delta': 'plan', analyze: 'plan',
'code-clean': 'plan', 'client-handover': 'plan', brainstorming: 'plan',
'writing-plans': 'plan', 'requesting-code-review': 'plan',
'21st-ui-review': 'plan',
// reflect
feat: 'reflect', hotfix: 'reflect', bugfix: 'reflect',
refactor: 'reflect', 'web-validate': 'reflect', harden: 'reflect',
seo: 'reflect', geo: 'reflect', 'site-motion': 'reflect',
'frontend-design': 'reflect', 'emil-design-eng': 'reflect',
'design-motion-principles': 'reflect', '21st-ui-build': 'reflect',
'scroll-world-storytelling': 'reflect',
'build-threejs-scroll-worlds': 'reflect',
'scroll-scrubbed-visual-sequence': 'reflect',
'scroll-scrubbed-word-reveal': 'reflect',
'scroll-progress-timeline': 'reflect',
'subagent-driven-development': 'reflect', 'writing-skills': 'reflect',
'deprecation-and-migration': 'reflect', '21st-ai': 'reflect',
'21st-ui-explore': 'reflect',
// implement
gitflow: 'implement', 'prune-memory': 'implement',
'pdf-translate': 'implement', 'ci-cd-and-automation': 'implement',
'observability-and-instrumentation': 'implement',
'test-driven-development': 'implement',
// apply
'commit-change': 'apply', 'release-candidate': 'apply', doc: 'apply',
capitalize: 'apply', close: 'apply', reconcile: 'apply', deploy: 'apply',
// mechanical
status: 'mechanical', profile: 'mechanical', 'plugin-check': 'mechanical',
'skills-perso': 'mechanical', 'using-git-worktrees': 'mechanical',
'21st-cli-use': 'mechanical', '21st-registry': 'mechanical',
'21st-design-sync': 'mechanical',
},
// `floor` rules set the turn's minimum effort; `default` rules set the // `floor` rules set the turn's minimum effort; `default` rules set the
// turn's route, which a route call or a skill overrides. Neither lowers. // turn's route, which a route call or a skill overrides. Neither lowers.
prompt: [ prompt: [
@@ -386,6 +460,25 @@ function mergeRouting(
} }
} }
/** A row table: `name: null` in the override drops a default row. */
function mergeRows(
base: Record<string, string>,
user: unknown,
name: string,
ref: (key: string, value: unknown) => string | undefined,
log: Log,
): Record<string, string> {
const entries = isRecord(user) ? Object.entries(user) : []
const kept = isRecord(user)
? Object.fromEntries(entries.filter(([, v]) => v !== null))
: user
const rows = mergeTable(base, kept, name, ref, log)
for (const [key, value] of entries) {
if (value === null && key !== '__proto__') delete rows[key]
}
return rows
}
/** Defaults overlaid with the user's entries, each validated first. */ /** Defaults overlaid with the user's entries, each validated first. */
function mergeConfig(user: unknown, log: Log): Config { function mergeConfig(user: unknown, log: Log): Config {
const base = structuredClone(DEFAULT_CONFIG) const base = structuredClone(DEFAULT_CONFIG)
@@ -405,8 +498,8 @@ function mergeConfig(user: unknown, log: Log): Config {
windows: mergeTable( windows: mergeTable(
base.windows, user.windows, 'windows', acceptWindow, log), base.windows, user.windows, 'windows', acceptWindow, log),
phases, phases,
agents: mergeTable(base.agents, user.agents, 'agents', ref, log), agents: mergeRows(base.agents, user.agents, 'agents', ref, log),
skills: mergeTable(base.skills, user.skills, 'skills', ref, log), skills: mergeRows(base.skills, user.skills, 'skills', ref, log),
prompt: mergePrompt(base.prompt, user.prompt, phases, log), prompt: mergePrompt(base.prompt, user.prompt, phases, log),
mainModelSwitch: pickBool(user.mainModelSwitch, base.mainModelSwitch), mainModelSwitch: pickBool(user.mainModelSwitch, base.mainModelSwitch),
verbose: pickBool(user.verbose, base.verbose), verbose: pickBool(user.verbose, base.verbose),
@@ -487,9 +580,15 @@ function newState(cfg: Config, source: string): State {
source, source,
userMain: null, userMain: null,
turnMain: null, turnMain: null,
runMain: null,
turnFloor: null, turnFloor: null,
pendingPrompt: null, pendingPrompt: null,
typedSlash: false, typedSlash: null,
pendingSlash: null,
promptAllowed: false,
pendingAllowed: false,
spawning: 0,
offers: new Map(),
loops: new Map(), loops: new Map(),
explicitEffort: new Map(), explicitEffort: new Map(),
skillCalls: 0, skillCalls: 0,
@@ -511,7 +610,8 @@ function newState(cfg: Config, source: string): State {
/** /**
* /clear rebuilds the state but keeps what is account- or process-wide: * /clear rebuilds the state but keeps what is account- or process-wide:
* the breaker (a model's quota outlives the conversation) and the model. * the breaker (a model's quota outlives the conversation), the model and
* the agent definitions' sources (the listing is not offered again).
*/ */
function resetSession(st: State): void { function resetSession(st: State): void {
const kept = { const kept = {
@@ -519,6 +619,7 @@ function resetSession(st: State): void {
strikes: st.strikes, strikes: st.strikes,
agentModels: st.agentModels, agentModels: st.agentModels,
sessionModel: st.sessionModel, sessionModel: st.sessionModel,
offers: st.offers,
} }
Object.assign(st, newState(st.cfg, st.source), kept) Object.assign(st, newState(st.cfg, st.source), kept)
} }
@@ -788,10 +889,11 @@ function decideMain(
: downgradeCall(st, cur, wanted, tokens) : downgradeCall(st, cur, wanted, tokens)
} }
/** Model axis: sticky, then turn route, then the floor's own model. */ /** Model axis: sticky, turn route, run slot, then the floor's own model. */
const mainModel = (st: State): string | undefined => const mainModel = (st: State): string | undefined =>
routeName(st.userMain?.route) ?? routeName(st.userMain?.route) ??
routeName(st.turnMain?.route) ?? routeName(st.turnMain?.route) ??
routeName(st.runMain?.route) ??
routeName(st.turnFloor?.route) routeName(st.turnFloor?.route)
async function readTokens($: Api): Promise<number | undefined> { async function readTokens($: Api): Promise<number | undefined> {
@@ -827,8 +929,9 @@ function logCall($: Api, st: State, call: Call): void {
} }
} }
/** The main loop's effective route: user /route > latest turn route. */ /** The main loop's effective route: user /route > turn route > run slot. */
const mainRoute = (st: State): Routed | null => st.userMain ?? st.turnMain const mainRoute = (st: State): Routed | null =>
st.userMain ?? st.turnMain ?? st.runMain
const rank = (l: Level | undefined): number => const rank = (l: Level | undefined): number =>
l === undefined ? -1 : LEVELS.indexOf(l) l === undefined ? -1 : LEVELS.indexOf(l)
@@ -851,7 +954,8 @@ function higherFloor(cur: Routed | null, next: Routed): Routed {
*/ */
function mainEffort(st: State, engine: Effort): Decision { function mainEffort(st: State, engine: Effort): Decision {
const sticky = st.userMain?.route.effort const sticky = st.userMain?.route.effort
const named = sticky ?? st.turnMain?.route.effort const named = sticky ?? st.turnMain?.route.effort ??
st.runMain?.route.effort
const floor = st.turnFloor?.route.effort const floor = st.turnFloor?.route.effort
const base = named ?? floor ?? engine const base = named ?? floor ?? engine
const effort = floored(base, floor) const effort = floored(base, floor)
@@ -862,11 +966,8 @@ function mainEffort(st: State, engine: Effort): Decision {
return { effort, by: sticky === undefined ? 'turn' : 'sticky' } return { effort, by: sticky === undefined ? 'turn' : 'sticky' }
} }
const floorSource = (f: Routed): string =>
f.source === 'prompt' ? `prompt rule ${f.phase}` : `typed /${f.phase}`
const floorWord = (f: Routed): string => const floorWord = (f: Routed): string =>
`user floor ${f.route.effort ?? '-'} (${floorSource(f)})` `user floor ${f.route.effort ?? '-'} (prompt rule ${f.phase})`
/** /**
* Why main will not run at `asked`, or '' when it will. Truthful tail of * Why main will not run at `asked`, or '' when it will. Truthful tail of
@@ -903,6 +1004,7 @@ function writeLoop(loop: Loop, route: Route): void {
function clearRoutes(st: State): void { function clearRoutes(st: State): void {
st.userMain = null st.userMain = null
st.turnMain = null st.turnMain = null
st.runMain = null
st.turnFloor = null st.turnFloor = null
st.pendingPrompt = null st.pendingPrompt = null
} }
@@ -1122,14 +1224,18 @@ async function handleCommand($: Api, st: State, args: string): Promise<string> {
case '': case '':
case 'show': case 'show':
return show($, st) return show($, st)
case 'clear': case 'clear': {
const run = st.runMain !== null
clearRoutes(st) clearRoutes(st)
refresh($, st) refresh($, st)
return 'route cleared\n' + (await show($, st)) return 'route cleared' + (run ? '; run slot dropped' : '') +
'\n' + (await show($, st))
}
case 'on': case 'on':
case 'off': case 'off':
st.off = head === 'off' st.off = head === 'off'
st.offConfig = false st.offConfig = false
if (st.off) st.runMain = null
refresh($, st) refresh($, st)
return show($, st) return show($, st)
case 'reload': case 'reload':
@@ -1218,20 +1324,28 @@ function clearLoop(st: State, agentId: string | undefined): string {
const held = f && f.route.effort !== undefined const held = f && f.route.effort !== undefined
? `; ${floorWord(f)} still holds, /route clear drops it` ? `; ${floorWord(f)} still holds, /route clear drops it`
: '' : ''
return 'route cleared for main' + held const run = st.runMain
? `; run ${st.runMain.phase} still holds`
: ''
return 'route cleared for main' + held + run
} }
const loop = st.loops.get(agentId) const loop = st.loops.get(agentId)
if (loop) writeLoop(loop, {}) if (loop) writeLoop(loop, {})
return 'route cleared for this agent' return 'route cleared for this agent'
} }
/** What the main loop's model will do, in the words of the decision. */ /** The id the next main step runs on, with the decision's reason. */
async function mainAnswer($: Api, st: State): Promise<string> { async function mainAnswer($: Api, st: State): Promise<string> {
const { call, wanted } = await snapshot($, st) const { call } = await snapshot($, st)
if (call.moved) return `${idWord(st, call.model)} (${call.why})` return `${idWord(st, call.model)} (${call.why})`
const quiet = call.why === 'unchanged' || }
(wanted === undefined && !isFallback(call))
return quiet ? 'unchanged' : `unchanged (${call.why})` /** Main's answer: effort and model asked, the sticky/floor note appended. */
async function mainRoutedText($: Api, st: State, p: Picked) {
const base = `routed main to ${p.phase}: effort ${
p.route.effort ?? 'unchanged'}, model ${await mainAnswer($, st)}`
const note = mainNote(st, p.route.effort)
return note ? `${base}; but ${note}` : base
} }
/** Truthful answer: states what the calling loop will actually do. */ /** Truthful answer: states what the calling loop will actually do. */
@@ -1241,15 +1355,11 @@ async function routedText(
agentId: string | undefined, agentId: string | undefined,
p: Picked, p: Picked,
): Promise<string> { ): Promise<string> {
const note = agentId === undefined ? mainNote(st, p.route.effort) : '' if (agentId === undefined) return mainRoutedText($, st, p)
if (note) return `recorded ${p.phase} for this turn, but ${note}` const explicit = st.loops.get(agentId)?.explicitEffort
const loop = agentId === undefined ? undefined : st.loops.get(agentId) const model = routeName(p.route) ? 'unchanged (fixed at spawn)' : 'unchanged'
const effort = loop?.explicitEffort ? undefined : p.route.effort return `routed this agent to ${p.phase}: effort ${
const model = agentId === undefined explicit ? 'unchanged' : p.route.effort ?? 'unchanged'}, model ${model}`
? await mainAnswer($, st)
: routeName(p.route) ? 'unchanged (fixed at spawn)' : 'unchanged'
return `routed ${agentId === undefined ? 'main' : 'this agent'} to ` +
`${p.phase}: effort ${effort ?? 'unchanged'}, model ${model}`
} }
async function handleRouteTool($: Api, st: State, e: RouteInput) { async function handleRouteTool($: Api, st: State, e: RouteInput) {
@@ -1267,79 +1377,59 @@ async function handleRouteTool($: Api, st: State, e: RouteInput) {
// ---- skills ---------------------------------------------------------- // ---- skills ----------------------------------------------------------
const skillResult = (skill: string, line: string) => ({ /** A skill's table row: its phase and that phase's route, if both exist. */
result: { success: true, commandName: skill, status: 'inline' as const }, function skillRow(st: State, skill: string): Picked | undefined {
context: [line], const phase = hasKey(st.cfg.skills, skill) ? st.cfg.skills[skill] : undefined
}) const route = phase === undefined ? undefined : phaseRoute(st.cfg, phase)
return phase !== undefined && route ? { phase, route } : undefined
/** Answers Skill(effort-<l>) in place: one writer, the skill never loads. */
function effortBridge(st: State, agentId: string | undefined, skill: string,
level: Level) {
if (agentId === undefined) {
const route = { ...st.turnMain?.route, effort: level }
st.turnMain = { phase: skill, route, source: 'skill' }
const note = mainNote(st, level)
return skillResult(skill, note
? `model-router: ${skill} recorded, but ${note}; the ${skill} skill ` +
'text was not loaded.'
: `model-router: effort → ${level} for this loop from the next ` +
`request on; the ${skill} skill text was not loaded.`)
}
const loop = loopOf(st, agentId)
if (loop.explicitEffort) {
return skillResult(skill, 'model-router: this agent was dispatched with ' +
'an explicit effort; the shift does not apply.')
}
loop.effort = level
return skillResult(skill, `model-router: effort → ${level} for this ` +
`loop from the next request on; the ${skill} skill text was not loaded.`)
}
/** A non-effort skill load: resets the loop's route, applies its table row. */
function onSkillLoad(st: State, skill: string, agentId: string | undefined) {
const table = hasKey(st.cfg.skills, skill) ? st.cfg.skills[skill] : undefined
const route = table === undefined ? undefined : phaseRoute(st.cfg, table)
if (agentId === undefined) {
st.turnMain = null
if (table !== undefined && route) {
st.turnMain = { phase: table, route, source: 'skill' }
}
return
}
const loop = route ? loopOf(st, agentId) : st.loops.get(agentId)
if (loop) writeLoop(loop, { effort: route?.effort })
}
/** A user-typed /effort-<l>: a floor for the turn, prepends one line. */
function slashEffort(st: State, skill: string, text: string) {
const level = EFFORT_SKILL.exec(skill)?.[1]
if (!isLevel(level)) return undefined
const route: Route = { effort: level }
const slash: Routed = { phase: skill, route, source: 'slash' }
st.turnFloor = higherFloor(st.turnFloor, slash)
const note = mainNote(st, level)
const line = `Effort ${level} set by model-router for the main loop this ` +
'turn (minimum; a higher route still applies).' +
(note ? ` But ${note}.` : '')
return { text: line + '\n' + text }
} }
/** /**
* Floor write for a skill.prompt. Only a typed slash may write it: the * A rowed skill on main sets the turn route; a best-tier row also holds the
* marker from prompt.submit attests the typing. Without it, a live * run slot. A typed non-best row drops the run; one the model loads (a
* sub-agent means the prompt is a preload inside that agent: ignored. * helper skill inside a run) leaves it.
*/ */
function guardedSlash(st: State, skill: string, text: string) { function routeMainBySkill(st: State, row: Picked, typed: boolean): void {
if (!EFFORT_SKILL.test(skill)) return undefined const turn: Routed = { phase: row.phase, route: { ...row.route },
if (st.typedSlash) { source: 'skill' }
st.typedSlash = false st.turnMain = turn
} else if (st.loops.size > 0) { if (row.route.tier === BEST) st.runMain = { ...turn, source: 'run' }
return { else if (typed) st.runMain = null
text: `model-router: ${skill} preload inside a live sub-agent is ` + }
'ignored on the main loop.\n' + text,
} /** A model-loaded skill: its row applies; an unrowed one changes nothing. */
function onSkillLoad(st: State, skill: string, agentId: string | undefined) {
const row = skillRow(st, skill)
if (agentId === undefined) {
if (row) routeMainBySkill(st, row, false)
return
} }
return slashEffort(st, skill, text) if (row) writeLoop(loopOf(st, agentId), row.route)
}
type TypedPath = 'typed-marker' | 'typed-fallback'
/**
* Whether a skill.prompt on main is the user's own typing. The marker from
* prompt.submit names the skill; failing that, an allowed-origin prompt with
* no live or spawning sub-agent (a preload fires inside one) still counts.
*/
function typedPath(st: State, skill: string): TypedPath | undefined {
if (st.typedSlash === skill) {
st.typedSlash = null
return 'typed-marker'
}
const idle = st.loops.size === 0 && st.spawning === 0
return st.promptAllowed && idle ? 'typed-fallback' : undefined
}
function applyTypedSkill($: Api, st: State, skill: string): void {
const row = skillRow(st, skill)
const path = typedPath(st, skill)
if (row === undefined || path === undefined) return
routeMainBySkill(st, row, true)
refresh($, st)
vlog($, st, `skill ${skill}: ${path} → ${row.phase}`)
} }
// ---- agents ---------------------------------------------------------- // ---- agents ----------------------------------------------------------
@@ -1347,27 +1437,60 @@ function guardedSlash(st: State, skill: string, text: string) {
type SpawnIn = { type SpawnIn = {
tool_use_id: string tool_use_id: string
subagentType: string subagentType: string
provider: { plugin: string }
model?: string model?: string
fork: boolean fork: boolean
workflow?: unknown workflow?: unknown
} }
/** /**
* The table row of a built-in agent. Known limit: provider.plugin === * The table row of an agent, for every provider. Skipped for a fork, a
* 'engine' is the best built-in test at spawn; a user agent named Explore * workflow agent, and an agent whose definition came from the project (a
* in a foreign project also matches (wave 1: sonnet/medium on it). * foreign repo's own verifier.md). Known limit: the source is recorded by
* name only, so an offer fired inside a sub-agent with another cwd
* overwrites it; no record = the row applies (fail-open on routing).
*/ */
function spawnRoute(cfg: Config, e: SpawnIn, frozen: boolean) { function spawnRoute(st: State, e: SpawnIn, frozen: boolean) {
if (frozen || e.provider.plugin !== 'engine') return undefined if (frozen || !hasKey(st.cfg.agents, e.subagentType)) return undefined
if (!hasKey(cfg.agents, e.subagentType)) return undefined if (PROJECT_SOURCES.has(st.offers.get(e.subagentType) ?? '')) {
const phase = cfg.agents[e.subagentType] return undefined
return phase === undefined ? undefined : phaseRoute(cfg, phase) }
const phase = st.cfg.agents[e.subagentType]
return phase === undefined ? undefined : phaseRoute(st.cfg, phase)
}
/** Once per turn: a spawn whose tier has nothing at or above its head. */
function tierDownLog($: Api, st: State, agent: string, tier: string): void {
const text = `model-router: ${agent} tier ${tier} down, ` +
'frontmatter model kept'
const key = `spawn-down:${agent}:${tier}`
if (st.turnLogged.has(key)) return
st.turnLogged.add(key)
$.ui.log(text)
}
/**
* The first model of the route's tier that is not down, only when it ranks
* at or above the tier's head: an agent moves UP from its frontmatter alias,
* never below. Nothing qualifies: no write, the frontmatter model runs.
*/
function inTier($: Api, st: State, agent: string, tier: string) {
const aliases = hasKey(st.cfg.tiers, tier) ? st.cfg.tiers[tier] : undefined
const head = aliases?.[0]
const headId = head === undefined ? undefined : st.cfg.models[head]
if (aliases === undefined || headId === undefined) return undefined
const id = availableIn(st, aliases)
if (id !== undefined && (id === headId || ranksAbove(st, id, headId))) {
return id
}
tierDownLog($, st, agent, tier)
return undefined
} }
/** The model a spawn is rewritten to; an explicit `model` param wins. */ /** The model a spawn is rewritten to; an explicit `model` param wins. */
function spawnTarget(st: State, e: SpawnIn, route: Route | undefined) { function spawnTarget($: Api, st: State, e: SpawnIn, route: Route | undefined) {
return e.model === undefined ? resolveRoute(st, route) : undefined if (e.model !== undefined || route === undefined) return undefined
if (route.tier === undefined) return resolveRoute(st, route)
return inTier($, st, e.subagentType, route.tier)
} }
function trackLoop(st: State, e: SpawnIn, started: { function trackLoop(st: State, e: SpawnIn, started: {
@@ -1493,7 +1616,10 @@ function endMainTurn($: Api, st: State): void {
st.turnMain = null st.turnMain = null
st.pushed = null st.pushed = null
st.turnModel = undefined st.turnModel = undefined
st.typedSlash = false st.typedSlash = st.pendingSlash
st.promptAllowed = st.pendingAllowed
st.pendingSlash = null
st.pendingAllowed = false
st.explicitEffort.clear() st.explicitEffort.clear()
st.spinner = '' st.spinner = ''
st.turnLogged.clear() st.turnLogged.clear()
@@ -1554,6 +1680,7 @@ async function onAutoSwitch($: Api, st: State, e: SwitchIn): Promise<void> {
/** Any model change: keeps `sessionModel`; a user choice clears its mark. */ /** Any model change: keeps `sessionModel`; a user choice clears its mark. */
async function onModelSwitch($: Api, st: State, e: SwitchIn): Promise<void> { async function onModelSwitch($: Api, st: State, e: SwitchIn): Promise<void> {
st.sessionModel = e.to_model st.sessionModel = e.to_model
if (USER_SWITCH.has(e.source)) st.runMain = null
if (st.off || e.source === 'resume') return if (st.off || e.source === 'resume') return
if (e.source === 'auto') await onAutoSwitch($, st, e) if (e.source === 'auto') await onAutoSwitch($, st, e)
else clearBreaker(st, canonical(st.cfg, e.to_model)) else clearBreaker(st, canonical(st.cfg, e.to_model))
@@ -1637,8 +1764,6 @@ function registerSkills(on: On, st: State): void {
on('tool.call', { tool: 'Skill' }, async ($, e, next) => { on('tool.call', { tool: 'Skill' }, async ($, e, next) => {
const skill = typeof e.skill === 'string' ? e.skill : undefined const skill = typeof e.skill === 'string' ? e.skill : undefined
if (st.off || skill === undefined) return next(e) if (st.off || skill === undefined) return next(e)
const level = EFFORT_SKILL.exec(skill)?.[1]
if (isLevel(level)) return effortBridge(st, e.agentId, skill, level)
st.skillCalls += 1 st.skillCalls += 1
try { try {
safely(st, $, 'Skill', () => onSkillLoad(st, skill, e.agentId)) safely(st, $, 'Skill', () => onSkillLoad(st, skill, e.agentId))
@@ -1652,10 +1777,8 @@ function registerSkills(on: On, st: State): void {
}) })
on('skill.prompt', async ($, e, next) => { on('skill.prompt', async ($, e, next) => {
if (st.off || st.skillCalls > 0) return next(e) if (st.off || st.skillCalls > 0) return next(e)
const out = guardedSlash(st, e.skill, e.text) safely(st, $, 'skill.prompt', () => applyTypedSkill($, st, e.skill))
if (!out) return next(e) return next(e)
refresh($, st)
return out
}).catch(($, e, next) => { }).catch(($, e, next) => {
warnOnce(st, $, 'skill.prompt', next.error.kind) warnOnce(st, $, 'skill.prompt', next.error.kind)
return next(e) return next(e)
@@ -1679,26 +1802,47 @@ function registerAgents(on: On, st: State): void {
}) })
} }
/** Bookkeeping once the agent started: its loop, its model, a log line. */
function afterSpawn(
$: Api,
st: State,
e: SpawnIn,
started: { agentId?: unknown; model: string },
route: Route | undefined,
): void {
if (typeof started.agentId !== 'string') return
const agentId = started.agentId
safely(st, $, 'agent.spawn', () => {
trackLoop(st, e, { agentId }, route)
rememberAgent(st, agentId, started.model)
vlog($, st,
`spawn ${e.subagentType}: ${e.model ?? '-'} → ${started.model}`)
})
}
function registerSpawn(on: On, st: State): void { function registerSpawn(on: On, st: State): void {
on('agent.offer', async ($, e, next) => {
if (!st.off) st.offers.set(e.agent, e.source)
return next(e)
}).catch(($, e, next) => {
warnOnce(st, $, 'agent.offer', next.error.kind)
return next(e)
})
on('agent.spawn', async ($, e, next) => { on('agent.spawn', async ($, e, next) => {
if (st.off) return next(e) if (st.off) return next(e)
await prune($, st) st.spawning += 1
const frozen = e.fork || e.workflow !== undefined try {
const route = spawnRoute(st.cfg, e, frozen) await prune($, st)
const wanted = spawnTarget(st, e, route) const route = spawnRoute(st, e, e.fork || e.workflow !== undefined)
const started = await next(wanted === undefined const wanted = spawnTarget($, st, e, route)
? e const started = await next(wanted === undefined
: { ...e, model: wanted }) ? e
if (typeof started.agentId === 'string') { : { ...e, model: wanted })
const agentId = started.agentId afterSpawn($, st, e, started, route)
safely(st, $, 'agent.spawn', () => { return started
trackLoop(st, e, started, route) } finally {
rememberAgent(st, agentId, started.model) st.spawning = Math.max(0, st.spawning - 1)
vlog($, st,
`spawn ${e.subagentType}: ${e.model ?? '-'} → ${started.model}`)
})
} }
return started
}).catch(($, e, next) => { }).catch(($, e, next) => {
warnOnce(st, $, 'agent.spawn', next.error.kind) warnOnce(st, $, 'agent.spawn', next.error.kind)
return next(e) return next(e)
@@ -1767,14 +1911,41 @@ function routeFromPrompt(st: State, text: string, midTurn: boolean): void {
floorFromPrompt(st, midTurn, routed) floorFromPrompt(st, midTurn, routed)
} }
const slash = text.trimStart().startsWith('/') const slash = text.trimStart().startsWith('/')
if (text.trimStart().startsWith('/effort-')) st.typedSlash = true
if (!midTurn && !slash && !floor) defaultFromPrompt(st, scanned) if (!midTurn && !slash && !floor) defaultFromPrompt(st, scanned)
} }
/** The rowed skill a slash prompt names, else null. */
function typedSkill(st: State, text: string): string | null {
const typed = text.trimStart()
if (!typed.startsWith('/')) return null
const name = typed.slice(1).split(/\s/, 1)[0] ?? ''
return hasKey(st.cfg.skills, name) ? name : null
}
/**
* Records what skill.prompt needs to tell a typed slash from a preload: the
* skill named and whether a person typed it. A mid-turn prompt waits in the
* pending slots for the turn end; a foreign one leaves them alone. A later
* queued prompt replaces an earlier one (a shortcut: one pending slot).
*/
function noteTyped(st: State, kind: string, text: string, midTurn: boolean) {
const allowed = TYPED_ORIGINS.has(kind)
const slash = allowed ? typedSkill(st, text) : null
if (!midTurn) {
st.typedSlash = slash
st.promptAllowed = allowed
} else if (allowed) {
st.pendingSlash = slash
st.pendingAllowed = true
}
}
function registerPrompt(on: On, st: State): void { function registerPrompt(on: On, st: State): void {
on('prompt.submit', async ($, e, next) => { on('prompt.submit', async ($, e, next) => {
if (st.off || e.origin.kind !== 'composer') return next(e) if (st.off) return next(e)
routeFromPrompt(st, e.text, e.turnId !== undefined) const midTurn = e.turnId !== undefined
noteTyped(st, e.origin.kind, e.text, midTurn)
if (e.origin.kind === 'composer') routeFromPrompt(st, e.text, midTurn)
refresh($, st) refresh($, st)
return next(e) return next(e)
}).catch(($, e, next) => { }).catch(($, e, next) => {
+3 -3
View File
@@ -29,7 +29,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit
judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the
gate prints the remedy; end the turn — no later step, no dispatch. Nominal gate prints the remedy; end the turn — no later step, no dispatch. Nominal
(big) path is silent. (big) path is silent.
EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route).
Audit only what changed since the last run, on the axes the user picks. Audit only what changed since the last run, on the axes the user picks.
Per axis: **audit → approval gate → fix → re-verify → marker update**, Per axis: **audit → approval gate → fix → re-verify → marker update**,
@@ -170,7 +170,7 @@ Then show the user the same compact table inline.
### 3b-bis. CHALLENGE THE PROPOSALS (before the gate) ### 3b-bis. CHALLENGE THE PROPOSALS (before the gate)
`Skill(effort-xhigh)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). `mcp__model-router__route(phase="plan")` first (route: own level before the challenge; send it with the challenger dispatch).
This axis' findings + proposed fixes are a proposal set worth attacking before This axis' findings + proposed fixes are a proposal set worth attacking before
the human gate. Persist THIS axis' finding list (not the whole append-only the human gate. Persist THIS axis' finding list (not the whole append-only
report) to `.claude/tasks/plans/<date>-<axis>-<HHMM>.md`, then run report) to `.claude/tasks/plans/<date>-<axis>-<HHMM>.md`, then run
@@ -256,7 +256,7 @@ Then offer to capitalize (per CLAUDE.md): recurring finding patterns →
below on the same delta (the SAST is a deterministic floor, the reasoned below on the same delta (the SAST is a deterministic floor, the reasoned
pass covers what grep/rules miss): pass covers what grep/rules miss):
``` ```
Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message mcp__model-router__route(phase="orchestrate") # route: dispatch span starts; send with the Agent call below in ONE message
Agent(subagent_type="security-auditor", description="audit-delta security — semgrep SAST", Agent(subagent_type="security-auditor", description="audit-delta security — semgrep SAST",
prompt="MODE: audit\nSCOPE: <delta file list>\nREPORT: .claude/audits/.audit-delta-semgrep.md\nFollow agents/security-auditor.md exactly. Pinned rulesets, no login. Write ONLY to REPORT. End with REPORT_WRITTEN: <path>.") prompt="MODE: audit\nSCOPE: <delta file list>\nREPORT: .claude/audits/.audit-delta-semgrep.md\nFollow agents/security-auditor.md exactly. Pinned rulesets, no login. Write ONLY to REPORT. End with REPORT_WRITTEN: <path>.")
``` ```
+6 -6
View File
@@ -26,7 +26,7 @@ allowed-tools:
MODEL GATE (blocking): run `$HOME/.claude/lib/model-gate.md` BEFORE any MODEL GATE (blocking): run `$HOME/.claude/lib/model-gate.md` BEFORE any
step below. Verdict `small` → STOP — print the gate's remedy, end the step below. Verdict `small` → STOP — print the gate's remedy, end the
turn, dispatch nothing. turn, dispatch nothing.
EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route).
## REQUEST ## REQUEST
$ARGUMENTS $ARGUMENTS
@@ -119,14 +119,14 @@ RISK: <low/medium — what could go wrong>
obvious fix. obvious fix.
- If the fix is significant (>10 lines, multiple files, - If the fix is significant (>10 lines, multiple files,
behavior change): wait for user approval. behavior change): wait for user approval.
On resume: `Skill(effort-high)` first, sent with the next tool call (effort-shift: turn reset). On resume: `mcp__model-router__route(phase="reflect")` first (route: resumed turn).
- Then run pass B of `$HOME/.claude/lib/contract-interview.md` against the - Then run pass B of `$HOME/.claude/lib/contract-interview.md` against the
FIX PLAN: every VISIBLE / PUBLIC NAME / SCOPE choice it settles that the FIX PLAN: every VISIBLE / PUBLIC NAME / SCOPE choice it settles that the
bug report left open → one batch of questions, before STEP 3b. The trivial bug report left open → one batch of questions, before STEP 3b. The trivial
fast-path is not exempt: a 1-line fix with a visible choice still asks. fast-path is not exempt: a 1-line fix with a visible choice still asks.
## STEP 3b — CHALLENGE THE FIX PLAN (before the contract) ## STEP 3b — CHALLENGE THE FIX PLAN (before the contract)
`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). `mcp__model-router__route(phase="reflect")` first (route: own level before the challenge; send it with the challenger dispatch).
Unless the fix is the trivial 1-2 line case STEP 3 already fast-paths, the Unless the fix is the trivial 1-2 line case STEP 3 already fast-paths, the
DIAGNOSIS + FIX PLAN is a reflection worth attacking before it hardens into a DIAGNOSIS + FIX PLAN is a reflection worth attacking before it hardens into a
contract. Persist it to `.claude/tasks/plans/<date>-<slug>-<HHMM>.md`, then run contract. Persist it to `.claude/tasks/plans/<date>-<slug>-<HHMM>.md`, then run
@@ -157,10 +157,10 @@ branch it's a no-op (commit in place). Never `finish`.
## STEP 5 — DISPATCH EXECUTOR ## STEP 5 — DISPATCH EXECUTOR
Dispatch the executor — sonnet by frontmatter pin, do not override: Dispatch the executor — sonnet by its row (frontmatter = off-state floor), do not override:
``` ```
Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message mcp__model-router__route(phase="orchestrate") # route: dispatch span starts; send with the Agent call below in ONE message
Agent(subagent_type="bugfixer") Agent(subagent_type="bugfixer")
prompt: "CONTRACT: <path from STEP 3.5> prompt: "CONTRACT: <path from STEP 3.5>
DIAGNOSIS: <ROOT CAUSE + EVIDENCE from STEP 3> DIAGNOSIS: <ROOT CAUSE + EVIDENCE from STEP 3>
@@ -282,7 +282,7 @@ A bugfix with an understood root cause is almost always worth one entry:
If the bug was trivial and the root cause not transferable → skip with `CAPITALIZE: trivial, skip`. If the bug was trivial and the root cause not transferable → skip with `CAPITALIZE: trivial, skip`.
`Skill(effort-low)` first (effort-shift: bookkeeping tail; send it in the same message as the memory-commit command). `mcp__model-router__route(phase="apply")` first (route: bookkeeping tail; send it with the memory-commit command).
**Then commit the memory** — follow `$HOME/.claude/lib/capitalize-commit.md`: it **Then commit the memory** — follow `$HOME/.claude/lib/capitalize-commit.md`: it
surgically commits what capitalize just wrote (`.claude/memory` + `.claude/tasks` surgically commits what capitalize just wrote (`.claude/memory` + `.claude/tasks`
+4 -4
View File
@@ -26,7 +26,7 @@ allowed-tools:
MODEL GATE (blocking): run `$HOME/.claude/lib/model-gate.md` BEFORE any MODEL GATE (blocking): run `$HOME/.claude/lib/model-gate.md` BEFORE any
step below. Verdict `small` → STOP — print the gate's remedy, end the step below. Verdict `small` → STOP — print the gate's remedy, end the
turn, dispatch nothing. turn, dispatch nothing.
EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route).
## TARGET ## TARGET
$ARGUMENTS $ARGUMENTS
@@ -122,7 +122,7 @@ TOTALS: <N blocking, N warn, N info>
If no issues found: report clean state and stop. If no issues found: report clean state and stop.
## STEP 3b — CHALLENGE THE SCOPE (before approval) ## STEP 3b — CHALLENGE THE SCOPE (before approval)
`Skill(effort-xhigh)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). `mcp__model-router__route(phase="plan")` first (route: own level before the challenge; send it with the challenger dispatch).
The STEP 3 report is the proposed cleanup scope — worth attacking before the The STEP 3 report is the proposed cleanup scope — worth attacking before the
human approves it. It is still inline, so FIRST persist it to human approves it. It is still inline, so FIRST persist it to
`.claude/tasks/plans/<date>-<slug>-<HHMM>.md` (STEP 3 report format, one item `.claude/tasks/plans/<date>-<slug>-<HHMM>.md` (STEP 3 report format, one item
@@ -172,10 +172,10 @@ is approved, stop — no dispatch.
severity — proposed fix`. This is the executor's scope-of-work on severity — proposed fix`. This is the executor's scope-of-work on
disk — named, auditable, the same contract discipline as the dev disk — named, auditable, the same contract discipline as the dev
gates (verifier reads its contract from disk). gates (verifier reads its contract from disk).
2. **Dispatch the executor** — sonnet by frontmatter pin, do not override: 2. **Dispatch the executor** — sonnet by its row (frontmatter = off-state floor), do not override:
``` ```
Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message mcp__model-router__route(phase="orchestrate") # route: dispatch span starts; send with the Agent call below in ONE message
Agent(subagent_type="code-cleaner") Agent(subagent_type="code-cleaner")
prompt: "SCOPE: .claude/audits/CODE-CLEAN-SCOPE.md prompt: "SCOPE: .claude/audits/CODE-CLEAN-SCOPE.md
APPROVED: <the approved item list, incl. any per-item exported-symbol clears> APPROVED: <the approved item list, incl. any per-item exported-symbol clears>
+1 -1
View File
@@ -73,7 +73,7 @@ $ARGUMENTS"
(`model="opus"` — BDR-077: propose = narrative reconstruction + capitalize (`model="opus"` — BDR-077: propose = narrative reconstruction + capitalize
routing, judgment tier; the call-site override takes precedence over the routing, judgment tier; the call-site override takes precedence over the
sonnet frontmatter pin. Apply, STEP 4, stays on the pin.) sonnet row (frontmatter = off-state floor). Apply, STEP 4, stays on its row.)
Read the returned `COMMIT PLAN` + `EDGE CASES` + `CAPITALIZE CANDIDATES`, Read the returned `COMMIT PLAN` + `EDGE CASES` + `CAPITALIZE CANDIDATES`,
terminated by `READY TO APPLY — awaiting dispatcher confirmation`. terminated by `READY TO APPLY — awaiting dispatcher confirmation`.
-6
View File
@@ -1,6 +0,0 @@
---
name: effort-high
description: Investigation shift. Deeper reasoning for diagnosis, LOCATE, contract drafting, refactor judgement inside feat, hotfix and bugfix runs.
effort: high
---
Effort shifted to high for the rest of this turn (lib/effort-shift.md). Continue with the caller's next step.
-6
View File
@@ -1,6 +0,0 @@
---
name: effort-low
description: Bookkeeping shift. Lowers reasoning to the cheapest level for the rest of the turn: journal lines, memory commits, capitalize, release bookkeeping, status output.
effort: low
---
Effort shifted to low for the rest of this turn (lib/effort-shift.md). Continue with the caller's next step.
-6
View File
@@ -1,6 +0,0 @@
---
name: effort-max
description: Escalation shift. Maximum reasoning when a verify or security loop hits its cap, a gate fails twice, or error recovery starts in ship-feature.
effort: max
---
Effort shifted to max for the rest of this turn (lib/effort-shift.md). Continue with the caller's next step.
-6
View File
@@ -1,6 +0,0 @@
---
name: effort-medium
description: Orchestration shift. Standard reasoning between two dispatches: read a subagent report, pick the next step, relay a gate verdict, route a branch.
effort: medium
---
Effort shifted to medium for the rest of this turn (lib/effort-shift.md). Continue with the caller's next step.
-6
View File
@@ -1,6 +0,0 @@
---
name: effort-xhigh
description: Reflection shift. Deep reasoning for brainstorm, planning, challenge synthesis and audit verdicts before a human validation gate.
effort: xhigh
---
Effort shifted to xhigh for the rest of this turn (lib/effort-shift.md). Continue with the caller's next step.
+6 -6
View File
@@ -26,7 +26,7 @@ allowed-tools:
MODEL GATE (blocking): run `$HOME/.claude/lib/model-gate.md` BEFORE any MODEL GATE (blocking): run `$HOME/.claude/lib/model-gate.md` BEFORE any
step below. Verdict `small` → STOP — print the gate's remedy, end the step below. Verdict `small` → STOP — print the gate's remedy, end the
turn, dispatch nothing. turn, dispatch nothing.
EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route).
## REQUEST ## REQUEST
$ARGUMENTS $ARGUMENTS
@@ -124,7 +124,7 @@ in the contract's CLARIFICATIONS `[gated]` and in the plan. A choice that
surfaces only during execution comes back as `NEED-DECISION` (STEP 3). surfaces only during execution comes back as `NEED-DECISION` (STEP 3).
## STEP 1b — CHALLENGE THE PLAN (before branching) ## STEP 1b — CHALLENGE THE PLAN (before branching)
`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). `mcp__model-router__route(phase="reflect")` first (route: own level before the challenge; send it with the challenger dispatch).
The STEP 1 plan is a reflection worth attacking before a branch is spent on it. The STEP 1 plan is a reflection worth attacking before a branch is spent on it.
Persist it to `.claude/tasks/plans/<date>-<slug>-<HHMM>.md`, then run Persist it to `.claude/tasks/plans/<date>-<slug>-<HHMM>.md`, then run
`$HOME/.claude/lib/challenge-plan.md` with `PLAN` = that file, `KIND` = `build-plan`, `$HOME/.claude/lib/challenge-plan.md` with `PLAN` = that file, `KIND` = `build-plan`,
@@ -143,10 +143,10 @@ branch it's a no-op (commit in place). Never `finish`.
## STEP 3 — DISPATCH EXECUTOR ## STEP 3 — DISPATCH EXECUTOR
Dispatch the executor — sonnet by frontmatter pin, do not override: Dispatch the executor — sonnet by its row (frontmatter = off-state floor), do not override:
``` ```
Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message mcp__model-router__route(phase="orchestrate") # route: dispatch span starts; send with the Agent call below in ONE message
Agent(subagent_type="feater") Agent(subagent_type="feater")
prompt: "CONTRACT: <path from STEP 0.7> prompt: "CONTRACT: <path from STEP 0.7>
PLAN: <the STEP 1 checklist + approach bullets + edge cases, verbatim> PLAN: <the STEP 1 checklist + approach bullets + edge cases, verbatim>
@@ -203,7 +203,7 @@ test), consider splitting into 2-3 atomic commits grouped by logical
unit — or run `/commit-change` on the pending work (it dispatches the unit — or run `/commit-change` on the pending work (it dispatches the
commit-changer (propose opus / apply sonnet, BDR-077); never inline-load the bare agent, it is now a commit-changer (propose opus / apply sonnet, BDR-077); never inline-load the bare agent, it is now a
propose/apply executor). propose/apply executor).
Then `Skill(effort-high)` (effort-shift: nested commit-change loaded at low; reload feat's level, sent with the next tool call). Then `mcp__model-router__route(phase="reflect")` (route: nested commit-change loaded at apply; reload feat's level).
Print summary: Print summary:
``` ```
@@ -254,7 +254,7 @@ Always append a 1-line entry to today's heading in `.claude/memory/journal.md`.
If no substantive capture candidate → skip with `CAPITALIZE: nothing to log`. If no substantive capture candidate → skip with `CAPITALIZE: nothing to log`.
`Skill(effort-low)` first (effort-shift: bookkeeping tail; send it in the same message as the memory-commit command). `mcp__model-router__route(phase="apply")` first (route: bookkeeping tail; send it with the memory-commit command).
**Then commit the memory** — follow `$HOME/.claude/lib/capitalize-commit.md`: it **Then commit the memory** — follow `$HOME/.claude/lib/capitalize-commit.md`: it
surgically commits what capitalize just wrote (`.claude/memory` + `.claude/tasks` surgically commits what capitalize just wrote (`.claude/memory` + `.claude/tasks`
+4 -4
View File
@@ -29,7 +29,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit
judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the
gate prints the remedy; end the turn — no later step, no dispatch. Nominal gate prints the remedy; end the turn — no later step, no dispatch. Nominal
(big) path is silent. (big) path is silent.
EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route).
Dispatches the `geo-analyzer` subagent (audit + fix bundle), then applies Dispatches the `geo-analyzer` subagent (audit + fix bundle), then applies
the bundle from THIS main loop at **L1** — same shape as `/web-validate` the bundle from THIS main loop at **L1** — same shape as `/web-validate`
@@ -47,7 +47,7 @@ every phase (LRN-126). Clean `.audit/geo-signals-<RUNID>.md` after apply.
**A — collect (sonnet):** **A — collect (sonnet):**
``` ```
Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message mcp__model-router__route(phase="orchestrate") # route: dispatch span starts; send with the Agent call below in ONE message
Agent(subagent_type="geo-analyzer", model="sonnet") Agent(subagent_type="geo-analyzer", model="sonnet")
prompt: "MODE: collect prompt: "MODE: collect
RUNID: <RUNID> RUNID: <RUNID>
@@ -86,7 +86,7 @@ your bundle."
``` ```
## STEP 1b — CHALLENGE THE FIX BUNDLE (advisory, before apply) ## STEP 1b — CHALLENGE THE FIX BUNDLE (advisory, before apply)
`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). `mcp__model-router__route(phase="reflect")` first (route: own level before the challenge; send it with the challenger dispatch).
The analyzer returned a `## FIX BUNDLE` — worth attacking before any edit lands. The analyzer returned a `## FIX BUNDLE` — worth attacking before any edit lands.
**Skip if intervention mode = conservative** (nothing is applied). Else persist the **Skip if intervention mode = conservative** (nothing is applied). Else persist the
bundle verbatim to `.claude/tasks/plans/<date>-<slug>-<HHMM>.md`, then run bundle verbatim to `.claude/tasks/plans/<date>-<slug>-<HHMM>.md`, then run
@@ -118,7 +118,7 @@ intent, not header wording: **AUTO** = no-confirmation items (G1–G4/G6);
For each AUTO item, dispatch its `applier` at L1, passing the item verbatim: For each AUTO item, dispatch its `applier` at L1, passing the item verbatim:
``` ```
Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message mcp__model-router__route(phase="orchestrate") # route: dispatch span starts; send with the Agent call below in ONE message
Agent(subagent_type="hotfixer") # or "feater" per the item's applier Agent(subagent_type="hotfixer") # or "feater" per the item's applier
prompt: "<paste the bundle item: files, concern, current, expected, prompt: "<paste the bundle item: files, concern, current, expected,
framework note + shared-file discipline>. framework note + shared-file discipline>.
+3 -3
View File
@@ -29,7 +29,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit
judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the
gate prints the remedy; end the turn — no later step, no dispatch. Nominal gate prints the remedy; end the turn — no later step, no dispatch. Nominal
(big) path is silent. (big) path is silent.
EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route).
This skill orchestrates a narrow-scope hardening audit: TLS + security This skill orchestrates a narrow-scope hardening audit: TLS + security
headers + redirects + canonical + custom 404 + server configs. It headers + redirects + canonical + custom 404 + server configs. It
@@ -261,7 +261,7 @@ seo-analyzer will run in parallel.
Spawn a single seo-analyzer subagent with an explicit IN/OUT scope list. Spawn a single seo-analyzer subagent with an explicit IN/OUT scope list.
``` ```
Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message mcp__model-router__route(phase="orchestrate") # route: dispatch span starts; send with the Agent call below in ONE message
Agent( Agent(
subagent_type="seo-analyzer", subagent_type="seo-analyzer",
description="harden — narrow-scope web hardening audit", description="harden — narrow-scope web hardening audit",
@@ -522,7 +522,7 @@ Extract the score and critical-alert count from `.claude/audits/HARDEN.md` for t
--- ---
## STEP 2b — CHALLENGE THE FIX BUNDLE (MODE=fix only, advisory) ## STEP 2b — CHALLENGE THE FIX BUNDLE (MODE=fix only, advisory)
`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). `mcp__model-router__route(phase="reflect")` first (route: own level before the challenge; send it with the challenger dispatch).
Skip if MODE=audit (no bundle exists). Else, before the STEP 3 gate, harden the bundle: Skip if MODE=audit (no bundle exists). Else, before the STEP 3 gate, harden the bundle:
extract the `## 8. Fix bundle` section from HARDEN.md to extract the `## 8. Fix bundle` section from HARDEN.md to
`.claude/tasks/plans/<date>-<slug>-<HHMM>.md` (a clean, blind-judgeable artifact), then run `.claude/tasks/plans/<date>-<slug>-<HHMM>.md` (a clean, blind-judgeable artifact), then run
+5 -5
View File
@@ -24,7 +24,7 @@ allowed-tools:
MODEL GATE (blocking): run `$HOME/.claude/lib/model-gate.md` BEFORE any MODEL GATE (blocking): run `$HOME/.claude/lib/model-gate.md` BEFORE any
step below. Verdict `small` → STOP — print the gate's remedy, end the step below. Verdict `small` → STOP — print the gate's remedy, end the
turn, dispatch nothing. turn, dispatch nothing.
EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route).
## REQUEST ## REQUEST
$ARGUMENTS $ARGUMENTS
@@ -94,7 +94,7 @@ point. Run it ONLY when the settled fix touches control flow or behaviour — an
off-by-one, a wrong operator/variable, a behaviour-changing config value, or a off-by-one, a wrong operator/variable, a behaviour-changing config value, or a
missing import that alters execution. In doubt → it is probably a `/bugfix`. missing import that alters execution. In doubt → it is probably a `/bugfix`.
`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). `mcp__model-router__route(phase="reflect")` first (route: own level before the challenge; send it with the challenger dispatch).
For a logic fix: persist the STEP 1 located fix (root cause + the exact edit) to For a logic fix: persist the STEP 1 located fix (root cause + the exact edit) to
`.claude/tasks/plans/<date>-<slug>-<HHMM>.md`, then run `.claude/tasks/plans/<date>-<slug>-<HHMM>.md`, then run
`$HOME/.claude/lib/challenge-plan.md` with `PLAN` = that file, `KIND` = `$HOME/.claude/lib/challenge-plan.md` with `PLAN` = that file, `KIND` =
@@ -135,10 +135,10 @@ mentioned: STOP and ask `"working tree dirty: stash and continue, or abort?"`.
## STEP 3 — DISPATCH EXECUTOR ## STEP 3 — DISPATCH EXECUTOR
Dispatch the executor — sonnet by frontmatter pin, do not override: Dispatch the executor — sonnet by its row (frontmatter = off-state floor), do not override:
``` ```
Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message mcp__model-router__route(phase="orchestrate") # route: dispatch span starts; send with the Agent call below in ONE message
Agent(subagent_type="hotfixer") Agent(subagent_type="hotfixer")
prompt: "CONTRACT: <path from STEP 1.7> prompt: "CONTRACT: <path from STEP 1.7>
LOCATED: <file(s) found in STEP 1 + the confirmed root cause> LOCATED: <file(s) found in STEP 1 + the confirmed root cause>
@@ -236,7 +236,7 @@ Always append a 1-line entry to today's heading in `.claude/memory/journal.md` (
**Language rule**: the journal line and any proposed BLK/LRN entries are ALWAYS written English AND caveman — fragments, articles dropped, code/IDs/quoted errors verbatim — per CLAUDE.md "Memory registries" (Always English, always caveman). **Language rule**: the journal line and any proposed BLK/LRN entries are ALWAYS written English AND caveman — fragments, articles dropped, code/IDs/quoted errors verbatim — per CLAUDE.md "Memory registries" (Always English, always caveman).
`Skill(effort-low)` first (effort-shift: bookkeeping tail; send it in the same message as the memory-commit command). `mcp__model-router__route(phase="apply")` first (route: bookkeeping tail; send it with the memory-commit command).
**Then commit the memory** — follow `$HOME/.claude/lib/capitalize-commit.md`: it **Then commit the memory** — follow `$HOME/.claude/lib/capitalize-commit.md`: it
surgically commits what capitalize just wrote (`.claude/memory` + `.claude/tasks` surgically commits what capitalize just wrote (`.claude/memory` + `.claude/tasks`
+8 -8
View File
@@ -14,7 +14,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit
judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the
gate prints the remedy; end the turn — no later step, no dispatch. Nominal gate prints the remedy; end the turn — no later step, no dispatch. Nominal
(big) path is silent. (big) path is silent.
EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route).
## REQUEST ## REQUEST
$ARGUMENTS $ARGUMENTS
@@ -181,12 +181,12 @@ This is the deterministic scaffold commit owner (closes BLK-010). The MVP is
implemented on a `feature/*` branch off `develop` (STEP 8). implemented on a `feature/*` branch off `develop` (STEP 8).
## STEP 6 — PLAN ## STEP 6 — PLAN
`Skill(effort-xhigh)` first, sent with the next tool call (effort-shift: turn reset; gate #1 ended the turn and the vendored `writing-plans` pin applies only when the user invokes it). `mcp__model-router__route(phase="plan")` first (route: resumed turn; gate #1 ended the turn).
Invoke `writing-plans` (vendored superpowers skill) with BRIEF + skeleton. Invoke `writing-plans` (vendored superpowers skill) with BRIEF + skeleton.
Granular tasks (2-5 min each), exact file paths, TDD: tests before code. Granular tasks (2-5 min each), exact file paths, TDD: tests before code.
## STEP 6b — CHALLENGE THE PLAN (before the gate) ## STEP 6b — CHALLENGE THE PLAN (before the gate)
`Skill(effort-xhigh)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). `mcp__model-router__route(phase="plan")` first (route: own level before the challenge; send it with the challenger dispatch).
Before the human sees the implementation plan, harden it. Run Before the human sees the implementation plan, harden it. Run
`$HOME/.claude/lib/challenge-plan.md` with `PLAN` = the plan STEP 6 wrote under `$HOME/.claude/lib/challenge-plan.md` with `PLAN` = the plan STEP 6 wrote under
`docs/superpowers/plans/`, `KIND` = `build-plan`, `SCOPE` = the skeleton + task file `docs/superpowers/plans/`, `KIND` = `build-plan`, `SCOPE` = the skeleton + task file
@@ -212,7 +212,7 @@ Approve and start? (yes / request changes)
Changes → back to STEP 6. Approved → continue. Changes → back to STEP 6. Approved → continue.
## STEP 8 — IMPLEMENT ## STEP 8 — IMPLEMENT
First: `Skill(effort-medium)` (effort-shift: dispatch span starts; send it in the same message as this step's first dispatch). First: `mcp__model-router__route(phase="orchestrate")` (route: dispatch span starts; send it with this step's first dispatch).
Start the MVP feature branch off develop, then implement on it: Start the MVP feature branch off develop, then implement on it:
```bash ```bash
bash "$HOME/.claude/lib/gitflow.sh" start feature mvp bash "$HOME/.claude/lib/gitflow.sh" start feature mvp
@@ -224,7 +224,8 @@ Invoke `subagent-driven-development` (vendored superpowers skill) for the per-ta
finishing-a-development-branch", stop and return. finishing-a-development-branch", stop and return.
**Model routing (BDR-066):** every subagent dispatched under SDD — per-task **Model routing (BDR-066):** every subagent dispatched under SDD — per-task
implementers AND its reviewers — MUST carry `model: "sonnet"` in the Agent implementers AND its reviewers — MUST carry `model: "sonnet"` and
`effort="medium"` (the `implement` level; a main route never reaches a child) in the Agent
call. The plan is closed; execution and plan-conformity review are sonnet call. The plan is closed; execution and plan-conformity review are sonnet
work. Reflection (task decomposition, review verdict arbitration) stays in work. Reflection (task decomposition, review verdict arbitration) stays in
this loop. this loop.
@@ -261,9 +262,8 @@ against the founding contract. Distinct axis from STEP 10 code review
([[LRN-095]]) — both run. ([[LRN-095]]) — both run.
## STEP 10 — CODE REVIEW ## STEP 10 — CODE REVIEW
`Skill(effort-xhigh)` first, sent with the review dispatch (effort-shift: judgment dispatch; the reviewer is a built-in and inherits the level in force).
Invoke `requesting-code-review` (vendored superpowers skill). **Model routing (BDR-077):** the Invoke `requesting-code-review` (vendored superpowers skill). **Model routing (BDR-077):** the
review subagent it dispatches MUST carry `model: "opus"` in the Agent call — review subagent it dispatches MUST carry `model: "opus"` and `effort="xhigh"` in the Agent call —
craft review is dispatched judgment, never inherited from the session. Fix craft review is dispatched judgment, never inherited from the session. Fix
all CRITICAL before proceeding. all CRITICAL before proceeding.
@@ -316,7 +316,7 @@ articles dropped, code/IDs/quoted errors verbatim — per CLAUDE.md "Memory
registries" (Always English, always caveman). The gate may mirror the user's registries" (Always English, always caveman). The gate may mirror the user's
language; entries must not. language; entries must not.
`Skill(effort-low)` first (effort-shift: bookkeeping tail; send it in the same message as the memory-commit command). `mcp__model-router__route(phase="apply")` first (route: bookkeeping tail; send it with the memory-commit command).
**Then commit the memory** — follow `$HOME/.claude/lib/capitalize-commit.md`: it **Then commit the memory** — follow `$HOME/.claude/lib/capitalize-commit.md`: it
surgically commits the approved founding decisions (`.claude/memory` + surgically commits the approved founding decisions (`.claude/memory` +
+10 -3
View File
@@ -14,7 +14,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit
judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the
gate prints the remedy; end the turn — no later step, no dispatch. Nominal gate prints the remedy; end the turn — no later step, no dispatch. Nominal
(big) path is silent. (big) path is silent.
EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route).
## REQUEST ## REQUEST
$ARGUMENTS $ARGUMENTS
@@ -368,7 +368,7 @@ Lire le bloc `audit_stack:` du fichier `~/.claude/lib/project-archetypes/<archet
| Entry | Action | Livraison | | Entry | Action | Livraison |
|---|---|---| |---|---|---|
| `analyze` | Déjà fait en STEP 5 | L3a | | `analyze` | Déjà fait en STEP 5 | L3a |
| `code-clean` | Spawn subagent `general-purpose` (audit-only, `model="opus"` — BDR-076: dispatched audits off the session model) | L3a | | `code-clean` | Spawn subagent `general-purpose` (audit-only, `model="opus"`, `effort="xhigh"` — BDR-076: dispatched audits off the session model) | L3a |
| `cso` | Si gstack ON → Skill(cso). Sinon → Agent general-purpose avec checklist OWASP + deps audit | L3a | | `cso` | Si gstack ON → Skill(cso). Sinon → Agent general-purpose avec checklist OWASP + deps audit | L3a |
| `doc` | Spawn subagent `doc-syncer` (auto-mode OFF, report-only) | L3a | | `doc` | Spawn subagent `doc-syncer` (auto-mode OFF, report-only) | L3a |
| `seo` | Subagents seo-analyzer + geo-analyzer en parallèle | L3b | | `seo` | Subagents seo-analyzer + geo-analyzer en parallèle | L3b |
@@ -385,6 +385,7 @@ Lancer EN PARALLÈLE (un seul message, plusieurs Agent calls) les audits corresp
Agent( Agent(
subagent_type="general-purpose", subagent_type="general-purpose",
model="opus", model="opus",
effort="xhigh",
description="Onboard — code-clean audit only (read-only, opus)", description="Onboard — code-clean audit only (read-only, opus)",
prompt=""" prompt="""
AUDIT-ONLY mode — NO fixes, NO refactoring, NO file modifications. AUDIT-ONLY mode — NO fixes, NO refactoring, NO file modifications.
@@ -422,6 +423,7 @@ bash $HOME/.claude/lib/toggle-external.sh list 2>/dev/null | grep -E "^gstack\s+
Agent( Agent(
subagent_type="general-purpose", subagent_type="general-purpose",
model="opus", model="opus",
effort="xhigh",
description="Onboard — security audit fallback (archetype-adaptive)", description="Onboard — security audit fallback (archetype-adaptive)",
prompt=""" prompt="""
READ-ONLY security audit. No file modifications. READ-ONLY security audit. No file modifications.
@@ -546,6 +548,7 @@ flux de dev sont deux formes distinctes ([[BDR-050]] pipeline dev ≠ audit).
Agent( Agent(
subagent_type="doc-syncer", subagent_type="doc-syncer",
model="opus", model="opus",
effort="xhigh",
description="Onboard — doc drift audit only", description="Onboard — doc drift audit only",
prompt=""" prompt="""
MODE: audit — REPORT-ONLY, NO edits, NO auto-sync (no patch dispatch MODE: audit — REPORT-ONLY, NO edits, NO auto-sync (no patch dispatch
@@ -666,6 +669,7 @@ Si le skill ne supporte pas `--output`, capturer la sortie et écrire à la main
Agent( Agent(
subagent_type="general-purpose", subagent_type="general-purpose",
model="opus", model="opus",
effort="xhigh",
description="Onboard — static design review fallback", description="Onboard — static design review fallback",
prompt=""" prompt="""
AUDIT-ONLY mode — NO edits. Static design review du code UI. AUDIT-ONLY mode — NO edits. Static design review du code UI.
@@ -710,6 +714,7 @@ Puis parser le JSON Lighthouse (scores perf/a11y/bp/seo/pwa + top opportunities)
Agent( Agent(
subagent_type="general-purpose", subagent_type="general-purpose",
model="opus", model="opus",
effort="xhigh",
description="Onboard — static perf audit", description="Onboard — static perf audit",
prompt=""" prompt="""
AUDIT-ONLY mode — NO edits. AUDIT-ONLY mode — NO edits.
@@ -753,6 +758,7 @@ Parser axe-core résultats (violations, incomplete, inapplicable, passes) → `.
Agent( Agent(
subagent_type="general-purpose", subagent_type="general-purpose",
model="opus", model="opus",
effort="xhigh",
description="Onboard — static a11y audit", description="Onboard — static a11y audit",
prompt=""" prompt="""
AUDIT-ONLY mode — NO edits. AUDIT-ONLY mode — NO edits.
@@ -799,6 +805,7 @@ Spawn un subagent synthétiseur (isolé, chargé uniquement du contenu de `.onbo
Agent( Agent(
subagent_type="general-purpose", subagent_type="general-purpose",
model="opus", model="opus",
effort="xhigh",
description="Onboard — synthèse vers .claude/audits/", description="Onboard — synthèse vers .claude/audits/",
prompt=""" prompt="""
Lire tous les fichiers de <PROJECT_ROOT>/.onboard-audit/ : Lire tous les fichiers de <PROJECT_ROOT>/.onboard-audit/ :
@@ -892,7 +899,7 @@ Vérifier que les 4 fichiers `.claude/audits/ONBOARD_REPORT.md`, `.claude/audits
--- ---
## STEP 7b — CHALLENGE THE PROPOSALS (before the human gate) ## STEP 7b — CHALLENGE THE PROPOSALS (before the human gate)
`Skill(effort-xhigh)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). `mcp__model-router__route(phase="plan")` first (route: own level before the challenge; send it with the challenger dispatch).
The 4 audit files are on disk; `AUDIT_PROPOSALS.md` is the artifact worth The 4 audit files are on disk; `AUDIT_PROPOSALS.md` is the artifact worth
attacking before the human spends a gate on it. Run attacking before the human spends a gate on it. Run
`$HOME/.claude/lib/challenge-plan.md` with `PLAN` = `$HOME/.claude/lib/challenge-plan.md` with `PLAN` =
+5 -5
View File
@@ -30,7 +30,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit
judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the
gate prints the remedy; end the turn — no later step, no dispatch. Nominal gate prints the remedy; end the turn — no later step, no dispatch. Nominal
(big) path is silent. (big) path is silent.
EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route).
This skill orchestrates TWO specialist agents running in parallel, then This skill orchestrates TWO specialist agents running in parallel, then
merges their output into a single `.claude/audits/SEO.md` report. It is the main merges their output into a single `.claude/audits/SEO.md` report. It is the main
@@ -325,7 +325,7 @@ templating.
**PHASE A — collect (both domains, one message):** **PHASE A — collect (both domains, one message):**
``` ```
Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message mcp__model-router__route(phase="orchestrate") # route: dispatch span starts; send with the Agent call below in ONE message
Agent(subagent_type="seo-analyzer", model="sonnet") Agent(subagent_type="seo-analyzer", model="sonnet")
prompt: """ prompt: """
MODE: collect MODE: collect
@@ -460,7 +460,7 @@ sentinel, emit the COLLECT REPORT, stop. No scoring, no bundle.
``` ```
**PHASE B — judge (both domains, one message, AFTER both COLLECT REPORTs **PHASE B — judge (both domains, one message, AFTER both COLLECT REPORTs
are DONE):** no `model=` override — the opus frontmatter pins apply. are DONE):** no `model=` override — the opus rows apply (frontmatter = off-state floor).
``` ```
Agent(subagent_type="seo-analyzer") Agent(subagent_type="seo-analyzer")
@@ -510,7 +510,7 @@ the reports."
``` ```
## STEP 1b — CHALLENGE THE FIX BUNDLE (advisory, before apply) ## STEP 1b — CHALLENGE THE FIX BUNDLE (advisory, before apply)
`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). `mcp__model-router__route(phase="reflect")` first (route: own level before the challenge; send it with the challenger dispatch).
Both envelopes now carry a `## FIX BUNDLE` — worth attacking before any edit lands. Both envelopes now carry a `## FIX BUNDLE` — worth attacking before any edit lands.
**Skip if intervention mode = conservative** (nothing is applied). Else persist both **Skip if intervention mode = conservative** (nothing is applied). Else persist both
bundles (seo + geo, verbatim) to `.claude/tasks/plans/<date>-<slug>-<HHMM>.md`, then run bundles (seo + geo, verbatim) to `.claude/tasks/plans/<date>-<slug>-<HHMM>.md`, then run
@@ -558,7 +558,7 @@ The two bundles may touch the same shared template (meta vs JSON-LD). Apply
For each AUTO item, dispatch its `applier` at L1, passing the item verbatim: For each AUTO item, dispatch its `applier` at L1, passing the item verbatim:
``` ```
Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message mcp__model-router__route(phase="orchestrate") # route: dispatch span starts; send with the Agent call below in ONE message
Agent(subagent_type="hotfixer") # or "feater" per the item's applier Agent(subagent_type="hotfixer") # or "feater" per the item's applier
prompt: "<paste the bundle item: files, concern, current, expected, prompt: "<paste the bundle item: files, concern, current, expected,
framework note + shared-file discipline>. framework note + shared-file discipline>.
+13 -13
View File
@@ -14,7 +14,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit
judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the
gate prints the remedy; end the turn — no later step, no dispatch. Nominal gate prints the remedy; end the turn — no later step, no dispatch. Nominal
(big) path is silent. (big) path is silent.
EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route).
## REQUEST ## REQUEST
$ARGUMENTS $ARGUMENTS
@@ -114,10 +114,10 @@ Inject ONLY what constrains: the NON-BINDING count does NOT enter the brainstorm
(the injection inherits the OUTPUT filter — detail what binds, drop what doesn't). (the injection inherits the OUTPUT filter — detail what binds, drop what doesn't).
Consumption = INPUT INJECTION (we can't modify the external skill; we control its input). Consumption = INPUT INJECTION (we can't modify the external skill; we control its input).
Refine request into validated design via Socratic questioning. Don't proceed until design approved. Refine request into validated design via Socratic questioning. Don't proceed until design approved.
Turns after a user reply run at the session level until a tool call is paired with `Skill(effort-xhigh)` (effort-shift: turn reset). The first step of every turn after a user reply is `mcp__model-router__route(phase="plan")` (route: resumed turn).
## STEP 2 — PLAN ## STEP 2 — PLAN
`Skill(effort-xhigh)` first, sent with the next tool call (effort-shift: turn reset; brainstorm turns after a user reply run at the session level, and the vendored `brainstorming` pin applies only when the user invokes it). `mcp__model-router__route(phase="plan")` first (route: resumed turn after the brainstorm gate).
Invoke `writing-plans` (vendored superpowers skill) with the validated design AND the 0d digest: every task Invoke `writing-plans` (vendored superpowers skill) with the validated design AND the 0d digest: every task
must be consistent with the in-force constraints; where a task implements or affects one, must be consistent with the in-force constraints; where a task implements or affects one,
note the ID inline. Break design into tasks (2-5 min each). Each task: exact file paths, full code, verification steps. note the ID inline. Break design into tasks (2-5 min each). Each task: exact file paths, full code, verification steps.
@@ -127,7 +127,7 @@ request nor the STEP 1 brainstorm settled (check the contract's CLARIFICATIONS
first) → one batch before STEP 2b; answers append to the contract `[gated]`. first) → one batch before STEP 2b; answers append to the contract `[gated]`.
## STEP 2b — CHALLENGE THE PLAN (adversarial, before the gate) ## STEP 2b — CHALLENGE THE PLAN (adversarial, before the gate)
`Skill(effort-xhigh)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). `mcp__model-router__route(phase="plan")` first (route: own level before the challenge; send it with the challenger dispatch).
Before the human sees the plan, harden it. Run `$HOME/.claude/lib/challenge-plan.md`: Before the human sees the plan, harden it. Run `$HOME/.claude/lib/challenge-plan.md`:
- `PLAN` = the plan STEP 2 wrote under `docs/superpowers/plans/` - `PLAN` = the plan STEP 2 wrote under `docs/superpowers/plans/`
- `KIND` = `build-plan` - `KIND` = `build-plan`
@@ -174,7 +174,7 @@ judges the diff against this ENRICHED contract, not the STEP 0e seed — so a
criterion the design introduced is verified, not lost. criterion the design introduced is verified, not lost.
## STEP 4 — IMPLEMENT ## STEP 4 — IMPLEMENT
First: `Skill(effort-medium)` (effort-shift: dispatch span starts; send it in the same message as this step's first dispatch). First: `mcp__model-router__route(phase="orchestrate")` (route: dispatch span starts; send it with this step's first dispatch).
Start the feature branch off develop, then implement on it: Start the feature branch off develop, then implement on it:
```bash ```bash
bash "$HOME/.claude/lib/gitflow.sh" start feature <name> bash "$HOME/.claude/lib/gitflow.sh" start feature <name>
@@ -186,14 +186,15 @@ Invoke `subagent-driven-development` (vendored superpowers skill) for the per-ta
finishing-a-development-branch", stop and return. finishing-a-development-branch", stop and return.
**Model routing (BDR-066):** every subagent dispatched under SDD — per-task **Model routing (BDR-066):** every subagent dispatched under SDD — per-task
implementers AND its reviewers — MUST carry `model: "sonnet"` in the Agent implementers AND its reviewers — MUST carry `model: "sonnet"` and
`effort="medium"` (the `implement` level; a main route never reaches a child) in the Agent
call. The plan is closed; execution and plan-conformity review are sonnet call. The plan is closed; execution and plan-conformity review are sonnet
work. Reflection (task decomposition, review verdict arbitration) stays in work. Reflection (task decomposition, review verdict arbitration) stays in
this loop. this loop.
## STEP 4b — ERROR RECOVERY (if STEP 4 fails) ## STEP 4b — ERROR RECOVERY (if STEP 4 fails)
If a subagent returns a build error, failing test, or type error: If a subagent returns a build error, failing test, or type error:
1. `Skill(effort-max)` (effort-shift: error recovery; send it in the same message as the Read of the analyzer file below), then load 1. `mcp__model-router__route(phase="escalate")` (route: error recovery; send it with the Read of the analyzer file below), then load
`$HOME/.claude/agents/analyzer.md` in DEBUG MODE on the exact error output. `$HOME/.claude/agents/analyzer.md` in DEBUG MODE on the exact error output.
Produce: root cause hypotheses (ordered), affected files, what NOT to touch. Produce: root cause hypotheses (ordered), affected files, what NOT to touch.
2. Present gate: 2. Present gate:
@@ -210,10 +211,10 @@ OPTIONS :
C) Abort feature — preserve work done so far C) Abort feature — preserve work done so far
``` ```
3. Wait for user choice. Do NOT auto-fix. Do NOT proceed without explicit approval. 3. Wait for user choice. Do NOT auto-fix. Do NOT proceed without explicit approval.
4. On resume the turn is at the session level (effort-shift: turn reset). 4. On resume: `mcp__model-router__route(phase="plan")` first (route: resumed turn).
If A → `Skill(effort-medium)` sent with the re-dispatch, apply minimal fix, re-run STEP 4 for the failed task only. Max 2 retry attempts. If A → `mcp__model-router__route(phase="orchestrate")` sent with the re-dispatch, apply minimal fix, re-run STEP 4 for the failed task only. Max 2 retry attempts.
If still failing after 2 → fall back to options B or C. If still failing after 2 → fall back to options B or C.
If B or C → `Skill(effort-xhigh)` first, sent with the next tool call. If B or C → continue at the plan level set above.
If B → before skipping: scan remaining task list for tasks that depend on the failed task If B → before skipping: scan remaining task list for tasks that depend on the failed task
(look for references to the same file or function in subsequent tasks). (look for references to the same file or function in subsequent tasks).
If dependents found → present: "Tasks [N, M] depend on the skipped task. If dependents found → present: "Tasks [N, M] depend on the skipped task.
@@ -243,9 +244,8 @@ conformity + security vs. craft/design) — both run, neither subsumes the
other ([[LRN-095]]). other ([[LRN-095]]).
## STEP 6 — CODE REVIEW ## STEP 6 — CODE REVIEW
`Skill(effort-xhigh)` first, sent with the review dispatch (effort-shift: judgment dispatch; the reviewer is a built-in and inherits the level in force).
Invoke `requesting-code-review` (vendored superpowers skill). **Model routing (BDR-077):** the Invoke `requesting-code-review` (vendored superpowers skill). **Model routing (BDR-077):** the
review subagent it dispatches MUST carry `model: "opus"` in the Agent call — review subagent it dispatches MUST carry `model: "opus"` and `effort="xhigh"` in the Agent call —
craft review is dispatched judgment, never inherited from the session. Fix craft review is dispatched judgment, never inherited from the session. Fix
all CRITICAL before proceeding. all CRITICAL before proceeding.
@@ -277,7 +277,7 @@ Feature shipped implies at least one design decision worth capturing. Run this B
If nothing substantive to log → print `CAPITALIZE: nothing substantive to log` and skip. If nothing substantive to log → print `CAPITALIZE: nothing substantive to log` and skip.
`Skill(effort-low)` first (effort-shift: bookkeeping tail; send it in the same message as the memory-commit command). `mcp__model-router__route(phase="apply")` first (route: bookkeeping tail; send it with the memory-commit command).
**Then commit the memory** — follow `$HOME/.claude/lib/capitalize-commit.md`: it **Then commit the memory** — follow `$HOME/.claude/lib/capitalize-commit.md`: it
surgically commits what capitalize just wrote (`.claude/memory` + `.claude/tasks` surgically commits what capitalize just wrote (`.claude/memory` + `.claude/tasks`
+3 -3
View File
@@ -30,7 +30,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit
judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the
gate prints the remedy; end the turn — no later step, no dispatch. Nominal gate prints the remedy; end the turn — no later step, no dispatch. Nominal
(big) path is silent. (big) path is silent.
EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route).
One pipeline per project: **security → clean → re-verify → reconcile → One pipeline per project: **security → clean → re-verify → reconcile →
doc → convergence re-audit**, looping until a full pass applies zero new doc → convergence re-audit**, looping until a full pass applies zero new
@@ -169,8 +169,8 @@ honestly in the summary. Never loop past 3.
### Phase B — CLEAN ### Phase B — CLEAN
1. Dispatch a read-only cleanup audit (analyzer — opus-pinned, BDR-076 — 1. Dispatch a read-only cleanup audit (analyzer — opus row, BDR-076 —
or general-purpose with `model="opus"`; NOT the sonnet code-cleaner, or general-purpose with `model="opus"`, `effort="xhigh"`; NOT the sonnet code-cleaner,
which is now a fix executor): dead code, unused imports/exports, which is now a fix executor): dead code, unused imports/exports,
commented-out blocks, stale flags, norm violations. Findings as commented-out blocks, stale flags, norm violations. Findings as
`id | file:line | finding | proposed fix`. `id | file:line | finding | proposed fix`.
+4 -4
View File
@@ -28,7 +28,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit
judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the
gate prints the remedy; end the turn — no later step, no dispatch. Nominal gate prints the remedy; end the turn — no later step, no dispatch. Nominal
(big) path is silent. (big) path is silent.
EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route).
This skill orchestrates a narrow-scope standards audit : This skill orchestrates a narrow-scope standards audit :
@@ -180,7 +180,7 @@ Spawn a single `validator-analyzer` subagent with explicit scope and
collected context : collected context :
``` ```
Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message mcp__model-router__route(phase="orchestrate") # route: dispatch span starts; send with the Agent call below in ONE message
Agent( Agent(
subagent_type="validator-analyzer", subagent_type="validator-analyzer",
description="validate — W3C HTML + CSS + WCAG audit", description="validate — W3C HTML + CSS + WCAG audit",
@@ -255,7 +255,7 @@ grep -c '^### \[Critique\]' .claude/audits/VALIDATE.md
--- ---
## STEP 2b — CHALLENGE THE FIX BUNDLE (MODE=fix only, advisory) ## STEP 2b — CHALLENGE THE FIX BUNDLE (MODE=fix only, advisory)
`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). `mcp__model-router__route(phase="reflect")` first (route: own level before the challenge; send it with the challenger dispatch).
Skip if MODE=audit (no bundle exists). Else, before the STEP 3 gate, harden the bundle: Skip if MODE=audit (no bundle exists). Else, before the STEP 3 gate, harden the bundle:
extract the `## 5. Fix bundle` section from VALIDATE.md to extract the `## 5. Fix bundle` section from VALIDATE.md to
`.claude/tasks/plans/<date>-<slug>-<HHMM>.md` (a clean, blind-judgeable artifact), then run `.claude/tasks/plans/<date>-<slug>-<HHMM>.md` (a clean, blind-judgeable artifact), then run
@@ -313,7 +313,7 @@ Options :
share files: share files:
``` ```
Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message mcp__model-router__route(phase="orchestrate") # route: dispatch span starts; send with the Agent call below in ONE message
Agent(subagent_type="hotfixer") Agent(subagent_type="hotfixer")
prompt: "<paste the file-group's bundle items: file, issue, current, prompt: "<paste the file-group's bundle items: file, issue, current,
expected fix>. expected fix>.
+1 -11
View File
@@ -468,8 +468,7 @@ fi
# ── 7.3b. Update the Higgsfield CLI + skill pack ── # ── 7.3b. Update the Higgsfield CLI + skill pack ──
# CLI: global npm bin. Skills: re-cloned by lib/higgsfield-skills.sh, which # CLI: global npm bin. Skills: re-cloned by lib/higgsfield-skills.sh, which
# replaces the SOURCE under skills-external/ only: a pack parked in # replaces the SOURCE under skills-external/ only: a pack parked in
# skills-disabled/ (symlinks to those sources) stays parked. Runs before the # skills-disabled/ (symlinks to those sources) stays parked.
# effort-pins re-apply below (BDR-108).
echo "" echo ""
echo "── Updating Higgsfield CLI + skill pack..." echo "── Updating Higgsfield CLI + skill pack..."
if ! command -v higgsfield &>/dev/null; then if ! command -v higgsfield &>/dev/null; then
@@ -562,15 +561,6 @@ print(d.get('21st',{}).get('version','latest'))
rm -rf "$TFD_STAGE" rm -rf "$TFD_STAGE"
fi fi
# Effort pins (BDR-107, BDR-108): every refresh above rewrites SKILL.md and
# drops the `effort:` line; the 21st pack refresh is the last step that rewrites
# a SKILL.md, so the entry levels of lib/effort-pins.txt go back here.
echo ""
echo "── Re-applying effort pins on the vendored skills..."
# shellcheck source=lib/effort-pins.sh disable=SC1091
source "$REPO/lib/effort-pins.sh"
apply_effort_pins "$REPO" || warn "effort pins: map lines rejected — fix lib/effort-pins.txt"
# ── 7.5. Update external skills (npx skills) ── # ── 7.5. Update external skills (npx skills) ──
echo "" echo ""
echo "── Updating external skills (npx skills)..." echo "── Updating external skills (npx skills)..."