diff --git a/CLAUDE.global.md b/CLAUDE.global.md index c784bb1..b3c5419 100644 --- a/CLAUDE.global.md +++ b/CLAUDE.global.md @@ -295,10 +295,9 @@ design routing; the design-toolchain hook reinforces it. - Design system / brand → design-consultation first, then the build tools. - Review / audit → design-review + emil-design-eng + design-motion-principles + /impeccable audit|critique + `impeccable detect` floor. -- Load the stack paired with the first Read of the target file, never - alone (a lone Skill call applies no effort, `lib/effort-shift.md`); every - vendored member pins `high`, one level per stack (`lib/effort-pins.txt`); - plugin and gstack members run at the level in force. +- The vendored stack members route to `reflect` (high) through their + model-router rows (`mods/model-router`); the last rowed skill loaded + wins and an unrowed member (plugin, gstack) changes nothing. Scope doubt → ask or default to Build, never silently skip. Gate: light skills run `~/.claude/lib/design-gate.md`, orchestrators plugin-check. 21st = CLI (`npm i -g @21st-dev/cli`, `21st login`), no MCP, no key; search free, diff --git a/agents/analyzer.md b/agents/analyzer.md index 84b7f0a..5b94959 100644 --- a/agents/analyzer.md +++ b/agents/analyzer.md @@ -3,7 +3,7 @@ name: analyzer description: Analyze code, codebase, or problem before any modification. Produces a factual report without proposing solutions. Use proactively before any refactoring, design, or implementation. tools: Read, Grep, Glob, Bash model: opus -effort: high +effort: xhigh memory: project --- diff --git a/agents/client-handover-writer.md b/agents/client-handover-writer.md index c608bed..c444a2f 100644 --- a/agents/client-handover-writer.md +++ b/agents/client-handover-writer.md @@ -97,7 +97,7 @@ Parse `$ARGUMENTS` for optional flags: --- -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. +ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route). ## STEP 1 — PRE-FLIGHT @@ -227,7 +227,7 @@ Store `DEPLOYED_URL` for STEP 7. If empty, ask user during STEP 6. --- ## STEP 3 — BASELINE AUDITS (parallel) -First: `Skill(effort-high)` (effort-shift: judgment dispatch; the fable skill-runners are built-ins and inherit the level in force; high is the entry level of the audits they run). +Every skill-runner dispatch below carries an explicit `effort="high"` (route: built-ins never inherit a main route; high is the entry level of the audits they run). Goal: capture `SCORE_*_BEFORE` so the client doc shows the delta. @@ -262,10 +262,10 @@ the gate. **Model routing (BDR-077):** EVERY `general-purpose` skill-runner dispatch in this pipeline (initial audits, fix-loop re-dispatches, commit-change, -web-validate) carries `model: "fable"` — the child hosts gated orchestration -on the pipeline's behalf; it must never inherit the session model. +web-validate) carries `model: "fable"` and `effort="high"` — the child hosts gated +orchestration on the pipeline's behalf; it must never inherit the session model. -For web projects, dispatch in **a single message with two parallel Agent calls** (each with `model: "fable"`): +For web projects, dispatch in **a single message with two parallel Agent calls** (each with `model: "fable"`, `effort="high"`): | Audit (web) | Subagent | Prompt template | |---------------|-------------------|-----------------| @@ -424,7 +424,7 @@ for CSO) in the format ### Re-dispatch prompt template (SEO + GEO loop) -Send to `general-purpose` subagent (`model: "fable"`): +Send to `general-purpose` subagent (`model: "fable"`, `effort="high"`): > Read `~/.claude/skills/seo/SKILL.md` and re-run it on this project in > **conservative (audit-only) intervention mode**: re-score and leave both @@ -455,7 +455,7 @@ Send to `general-purpose` subagent (`model: "fable"`): ### Re-dispatch prompt template (HARDEN loop) -Send to `general-purpose` subagent (`model: "fable"`): +Send to `general-purpose` subagent (`model: "fable"`, `effort="high"`): > Read `~/.claude/skills/harden/SKILL.md` and re-run it with `--fix` ONLY > up to the bundle: stop at `READY TO APPLY — awaiting dispatcher @@ -468,7 +468,7 @@ Send to `general-purpose` subagent (`model: "fable"`): ### Re-dispatch prompt template (CSO loop — non-web only) -Send to `general-purpose` subagent (`model: "fable"`): +Send to `general-purpose` subagent (`model: "fable"`, `effort="high"`): > Read `~/.claude/skills/cso/SKILL.md` and re-run it in **daily mode**. > Previous score: **``/20** — below threshold. @@ -557,7 +557,7 @@ listed changes by hand before deploy." Continue to STEP 6. If `PENDING_CHANGES` non-empty → invoke /commit-change skill via subagent: -> Dispatch `general-purpose` subagent (`model: "fable"`). Prompt: +> Dispatch `general-purpose` subagent (`model: "fable"`, `effort="high"`). Prompt: > > "Read `~/.claude/skills/commit-change/SKILL.md` and execute. All pending > changes were produced by the client-handover ship pipeline during the @@ -677,7 +677,7 @@ Skip if `VALIDATE_SKIPPED=true` or `PROJECT_TYPE != web` (in either case ensure `VALIDATE_SKIPPED=true` is set so the gate logic in STEP 8 treats VALIDATE as not-applicable rather than failed). -Dispatch `general-purpose` subagent (`model: "fable"`): +Dispatch `general-purpose` subagent (`model: "fable"`, `effort="high"`): > Read `~/.claude/skills/web-validate/SKILL.md` and execute against the > deployed URL: ``. Audit W3C HTML validity (validator.nu), diff --git a/agents/commit-changer.md b/agents/commit-changer.md index b74d68b..af624a0 100644 --- a/agents/commit-changer.md +++ b/agents/commit-changer.md @@ -10,7 +10,7 @@ effort: high > MODEL (BDR-077): `MODE: propose` is dispatched with `model="opus"` (the > call-site override — narrative reconstruction + capitalize routing are -> judgment); `MODE: apply` runs on the sonnet frontmatter pin (mechanical +> judgment); `MODE: apply` runs on its sonnet row (frontmatter = off-state floor; mechanical > staging/committing of an approved plan). Reconstruct the development narrative from a working directory. The goal diff --git a/agents/doc-syncer.md b/agents/doc-syncer.md index 08bf745..7fbaaaf 100644 --- a/agents/doc-syncer.md +++ b/agents/doc-syncer.md @@ -60,8 +60,8 @@ audit, report, and patch. Parse `$ARGUMENTS`: - **`MODE: patch`** — the dispatcher approved a PATCH PLAN and re-dispatches - this agent to APPLY it. Jump to MODE: PATCH section. Runs on the sonnet - frontmatter pin. + this agent to APPLY it. Jump to MODE: PATCH section. Runs on its sonnet + row (frontmatter = off-state floor). - **`MODE: audit`** (or no explicit MODE — audit is the default) — analysis half, dispatched with `model: "opus"` (judgment tier; the call-site override takes precedence over the sonnet pin). **READ-ONLY: Write and diff --git a/agents/geo-analyzer.md b/agents/geo-analyzer.md index 22f37d9..77d845b 100644 --- a/agents/geo-analyzer.md +++ b/agents/geo-analyzer.md @@ -104,7 +104,7 @@ Mirror of seo-analyzer's pipeline contract. Parse the MODE line: to the run-scoped, gitignored `.audit/geo-signals-.md`, terminated by `COLLECTION COMPLETE — RUNID: `; emit a `COLLECT REPORT` (`STATUS`, RUNID, COVERAGE counts) and STOP. -- **`MODE: judge`** — opus frontmatter pin. Fail-closed load of +- **`MODE: judge`** — opus row (frontmatter = off-state floor). Fail-closed load of `.audit/geo-signals-.md` (absent / RUNID mismatch / missing sentinel → `GEO JUDGE — VERDICT: ERROR()`, STOP — never score stale or partial signals). Then STEP 6-12 (schema, entity — including diff --git a/agents/handover-doc-writer.md b/agents/handover-doc-writer.md index 79b63dc..b2b0e79 100644 --- a/agents/handover-doc-writer.md +++ b/agents/handover-doc-writer.md @@ -57,7 +57,7 @@ The parent dispatches this agent TWICE, with the FULL PACKAGE both times (`STATUS: DONE | BLOCKED`, RUNID, phase-cluster count, per-chapter word counts) and STOPS — STEP 13-16, the final MD, HTML and PDF are NEVER this mode's job. -- **`MODE: render`** — runs on the sonnet frontmatter pin. FIRST loads the +- **`MODE: render`** — runs on its sonnet row (frontmatter = off-state floor). FIRST loads the draft: absent file, RUNID mismatch, or missing `DRAFT COMPLETE` sentinel → `STATUS: BLOCKED` naming the cause (fail closed — never synthesize a missing draft, never render a partial one). Then runs STEP 13 → 14 → diff --git a/agents/plan-challenger.md b/agents/plan-challenger.md index 58a8426..f09e1e0 100644 --- a/agents/plan-challenger.md +++ b/agents/plan-challenger.md @@ -106,11 +106,11 @@ the MAIN loop, never here): - Dispatch THREE fresh challengers IN PARALLEL, one per lens (correctness / robustness / simplicity), each blind to the others. - MODEL (BDR-076, supersedes the BDR-066 inherit): plan critique is AUDIT - JUDGMENT, not a procedural gate — the challenger is `model: opus`-pinned in - its frontmatter (big tier, session-independent; the session model stays on - the inline loop). Never `model: "sonnet"` — a silent judgment downgrade. - (Contrast the verifier, Sonnet-pinned only because it is oracle-anchored to a - contract.) + JUDGMENT, not a procedural gate — the challenger is routed to the `judge` + row (opus) by the model-router, the `model: opus` frontmatter being the off-state floor (big + tier; the session model stays on the inline loop). Never `model: "sonnet"` — + a silent judgment downgrade. (Contrast the verifier on its `verify` row and + the executors on their `implement` row, both sonnet, oracle-anchored to a contract.) - FAIL-SAFE — never fail open: a malformed/empty verdict, a missing `PROOF`, or a dead challenger → retry ONCE fresh; a 2nd failure → escalate to the human and NAME the lens. Never report "plan challenged" on a silently dropped lens (same diff --git a/agents/seo-analyzer.md b/agents/seo-analyzer.md index 9d5d580..7613567 100644 --- a/agents/seo-analyzer.md +++ b/agents/seo-analyzer.md @@ -38,7 +38,7 @@ The dispatcher (/seo) runs this agent as a 3-stage pipeline; /harden and terminated by the line `COLLECTION COMPLETE — RUNID: `, then emits a short `COLLECT REPORT` (`STATUS: DONE | BLOCKED`, RUNID, COVERAGE counts) and STOPS. No scoring, no findings, no bundle. -- **`MODE: judge`** — runs on the opus frontmatter pin (audit judgment). +- **`MODE: judge`** — runs on its opus row (audit judgment; frontmatter = off-state floor). FIRST loads `.audit/seo-signals-.md`: absent, RUNID mismatch, or missing `COLLECTION COMPLETE` sentinel → emit `SEO JUDGE — VERDICT: ERROR()` and STOP (fail closed — NEVER diff --git a/install-plugins.sh b/install-plugins.sh index c590355..3b6227a 100644 --- a/install-plugins.sh +++ b/install-plugins.sh @@ -937,10 +937,6 @@ for _ext_skill in "${EXT_SKILL_NAMES[@]}"; do done echo "" -# Effort pins (BDR-107, BDR-108): every vendored external gets its entry -# level from lib/effort-pins.txt, re-applied ONCE after the last vendoring -# step (the 21st pack, STEP 8.7) — see apply_effort_pins there. - # ============================================================ # STEP 8.5 — EXTERNAL SKILLS (npx skills add …) # ============================================================ @@ -1005,8 +1001,7 @@ echo "" # profile: `lib/toggle-external.sh enable higgsfield` turns the media skills # on, `enable higgsfield-websites` the landing-page aid. Keeping it out of # link.sh and of every profile is what stops a re-run from re-enabling it -# (BDR-093). This step runs before Step 8.7 so the effort pins are still -# re-applied after the last vendoring step (BDR-108). +# (BDR-093). echo "── Step 8.6: Higgsfield CLI + skill pack ───────────────────" echo "" # shellcheck source=lib/higgsfield-skills.sh disable=SC1091 @@ -1128,13 +1123,6 @@ if command -v 21st &>/dev/null; then rm -rf "$TFD_STAGE" fi -# Effort pins (BDR-107, BDR-108): the vendored externals carry no `effort:` -# upstream and every vendoring step above rewrites SKILL.md. Re-apply the -# entry levels from lib/effort-pins.txt once, after the LAST such step. -# shellcheck source=lib/effort-pins.sh disable=SC1091 -source "$REPO/lib/effort-pins.sh" -apply_effort_pins "$REPO" || warn "effort pins: map lines rejected — fix lib/effort-pins.txt" - # Auth — detect, then offer login ONLY in an interactive TTY. A non-interactive # run (CI / headless / re-run) must never open a browser or block on OAuth. # Search and logo lookup are free; retrieving component code and 21st AI need diff --git a/lib/challenge-plan.md b/lib/challenge-plan.md index 66d4cb1..56d454c 100644 --- a/lib/challenge-plan.md +++ b/lib/challenge-plan.md @@ -43,8 +43,9 @@ Agent(subagent_type="plan-challenger", description="challenge:", prompt="" ``` **MODEL (BDR-076, supersedes the BDR-066 inherit):** plan critique is AUDIT -JUDGMENT — the challengers are `model: opus`-pinned in their frontmatter: a big -tier, session-independent, off the session model. The session model (Fable) +JUDGMENT — the challengers are routed to the `judge` row (opus) by the model-router, +the `model: opus` frontmatter being the off-state floor: a big tier, off +the session model. The session model (Fable) keeps only this loop — synthesis, RE-THINK, gate. Never sonnet: that would silently downgrade the judgment. (The executor gates stay sonnet.) @@ -59,8 +60,9 @@ silently downgrade the judgment. (The executor gates stay sonnet.) A challenger that returns a malformed/empty verdict, a missing `PROOF`, or dies → retry ONCE with a fresh challenger; a 2nd failure on that lens → STOP and escalate -(the STOP text names the level reached, `$CLAUDE_EFFORT`, and suggests `/effort-max` -for the relaunch; no shift here: a mute challenger is an infrastructure failure) +(the STOP text names the level reached, `$CLAUDE_EFFORT`, and suggests relaunching with +`ultrathink` in the prompt (turn floor) or `/route effort=max` (sticky, +`/route clear` after); no shift here: a mute challenger is an infrastructure failure) to the human, NAMING the lens. Never carry "plan challenged" into the gate on a silently dropped lens (`verify-secure-loop.md`: "a mute verifier is NEVER a PASS"). diff --git a/lib/effort-pins.sh b/lib/effort-pins.sh deleted file mode 100755 index c8cc4d7..0000000 --- a/lib/effort-pins.sh +++ /dev/null @@ -1,120 +0,0 @@ -#!/usr/bin/env bash -# lib/effort-pins.sh — re-apply the entry effort level on vendored skills -# (BDR-107 second axis, extended to every vendored external by BDR-108). -# Upstream copies carry no `effort:` and every vendoring step rewrites -# SKILL.md, so the level lives in lib/effort-pins.txt and this helper puts -# it back after the last vendoring step of install-plugins.sh and -# update-all.sh. Idempotent: same level → untouched, other level → -# replaced inside the frontmatter only, skill not vendored → skipped, -# malformed map line → rejected loudly, never applied. Hardenings: a map -# whose last line lacks a newline is still read; a SKILL.md whose frontmatter -# never closes is skipped untouched; the level is re-read after every write -# and a mismatch counts as failed; the write goes through a mktemp sibling -# removed on any failure and on INT/TERM (previous traps restored, never an -# EXIT trap: the installer owns one); the rejected map line is printed -# shell-quoted so a caller's `echo -e` cannot interpret it. Placement inside -# the frontmatter has no effect on the harness, which reads the key anywhere. -# -# Usage: source it, then `apply_effort_pins [repo-root]` -# or standalone: bash lib/effort-pins.sh [repo-root] -# Exit 1 when at least one map line was rejected or a skill failed. - -EFFORT_PINS_REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" -EFFORT_PIN_LEVEL_RE='^(low|medium|high|xhigh|max)$' -EFFORT_PIN_NAME_RE='^[A-Za-z0-9][A-Za-z0-9._-]*$' - -# Callers (install-plugins.sh, update-all.sh) define these; standalone -# runs get plain fallbacks. -declare -F ok >/dev/null || ok() { printf ' ok %s\n' "$*"; } -declare -F info >/dev/null || info() { printf ' info %s\n' "$*"; } -declare -F err >/dev/null || err() { printf ' ERR %s\n' "$*" >&2; } - -# _effort_pin_current → prints the frontmatter effort, if any -_effort_pin_current() { - awk 'NR==1&&/^---$/{p=1;next} p&&/^---$/{exit} p' "$1" \ - | sed -n 's/^effort: //p' | head -1 -} - -# _effort_pin_closed → rc 0 when the frontmatter has a closing --- -_effort_pin_closed() { - awk 'NR==1&&/^---$/{p=1;next} p&&/^---$/{f=1;exit} END{exit !f}' "$1" -} - -# _effort_pin_traps_restore — drop the INT/TERM handlers set for the -# write and re-install the caller's saved ones. No exit-time handler here. -_effort_pin_traps_restore() { - trap - INT TERM - [ -z "$1" ] || eval "$1" -} - -# _effort_pin_write — replace the frontmatter -# `effort:` line, or insert one after `name: ` (before the closing -# `---` when the frontmatter has no name line). Body lines never change. -# Writes a mktemp sibling then renames; any failure leaves no temp behind. -_effort_pin_write() { - local file="$1" name="$2" level="$3" tmp prev rc - tmp="$(mktemp "$file.XXXXXX")" || return 1 - prev="$(trap -p INT TERM)" - trap 'rm -f "$tmp"; exit 130' INT TERM - cp -p "$file" "$tmp" && awk -v n="$name" -v lvl="$level" ' - NR==1 && /^---$/ { fm=1; print; next } - fm && /^---$/ { - if (!done) { print "effort: " lvl; done=1 } - fm=0; print; next - } - fm && /^effort: / { if (!done) { print "effort: " lvl; done=1 }; next } - fm && $0 == "name: " n { print; if (!done) { print "effort: " lvl; done=1 }; next } - { print } - ' "$file" > "$tmp" && mv "$tmp" "$file"; rc=$? - [ "$rc" -eq 0 ] || rm -f "$tmp" - _effort_pin_traps_restore "$prev" - return "$rc" -} - -# _effort_pin_apply_one → rc 0 applied, 2 already at -# level, 1 failed (err line printed, file untouched or write rolled back) -_effort_pin_apply_one() { - local file="$1" name="$2" level="$3" - if ! _effort_pin_closed "$file"; then - err "effort-pins: $file: frontmatter never closed — skipped"; return 1 - fi - [ "$(_effort_pin_current "$file")" = "$level" ] && return 2 - if ! _effort_pin_write "$file" "$name" "$level"; then - err "effort-pins: $file: write failed"; return 1 - fi - if [ "$(_effort_pin_current "$file")" != "$level" ]; then - err "effort-pins: $file: level not applied after write" - return 1 - fi - return 0 -} - -# apply_effort_pins [repo-root] — walk the map, pin every vendored skill -apply_effort_pins() { - local repo="${1:-$EFFORT_PINS_REPO}" map name level rest file rc - local applied=0 kept=0 rejected=0 failed=0 - map="$repo/lib/effort-pins.txt" - [ -f "$map" ] || { err "effort-pins: map missing: $map"; return 1; } - while read -r name level rest || [ -n "$name" ]; do - case "$name" in ''|'#'*) continue ;; esac - if [ -n "$rest" ] || ! [[ "$name" =~ $EFFORT_PIN_NAME_RE ]] \ - || ! [[ "$level" =~ $EFFORT_PIN_LEVEL_RE ]]; then - err "effort-pins: rejected map line $(printf '%q' "$name $level $rest")" - rejected=$((rejected + 1)); continue - fi - file="$repo/skills-external/$name/SKILL.md" - [ -f "$file" ] || continue - _effort_pin_apply_one "$file" "$name" "$level"; rc=$? - case "$rc" in - 0) applied=$((applied + 1)) ;; - 2) kept=$((kept + 1)) ;; - *) failed=$((failed + 1)) ;; - esac - done < "$map" - ok "effort-pins: $applied applied, $kept already at level, $failed failed" - [ "$rejected" -eq 0 ] && [ "$failed" -eq 0 ] -} - -if [[ "${BASH_SOURCE[0]}" == "$0" ]]; then - apply_effort_pins "$@" -fi diff --git a/lib/effort-pins.txt b/lib/effort-pins.txt deleted file mode 100644 index 00b90dd..0000000 --- a/lib/effort-pins.txt +++ /dev/null @@ -1,46 +0,0 @@ -# lib/effort-pins.txt — entry effort level of the vendored skills -# (skills-external//SKILL.md). Upstream copies carry no `effort:` and -# every resync rewrites SKILL.md, so the pin lives here and -# lib/effort-pins.sh re-applies it after the last vendoring step of -# install-plugins.sh and update-all.sh. One line = ` `, -# level in low|medium|high|xhigh|max. The census -# lib/tests/effort-routing.test.sh checks every vendored file against this -# map. Rungs (BDR-107, BDR-108): low = fix a line, run a script · medium = -# day-to-day · high = refactor, resisting bug · xhigh = architecture, audit -# before validation · max = stuck. -# -# superpowers (obra/superpowers, plugins.lock.json "superpowers") -brainstorming xhigh -writing-plans xhigh -requesting-code-review xhigh -subagent-driven-development high -writing-skills high -test-driven-development medium -using-git-worktrees low -# -# agent-skills (addyosmani/agent-skills, plugins.lock.json "agent-skills") -deprecation-and-migration high -ci-cd-and-automation medium -observability-and-instrumentation medium -# -# design stack — ONE level for every member: these skills load stacked in a -# single UI build and the last loaded wins (lib/effort-shift.md), so two -# levels in the stack would make the effort depend on load order. -# skills/site-motion (repo-authored) pins the same level in its frontmatter. -frontend-design high -emil-design-eng high -design-motion-principles high -21st-ui-build high -scroll-world-storytelling high -build-threejs-scroll-worlds high -scroll-scrubbed-visual-sequence high -scroll-scrubbed-word-reveal high -scroll-progress-timeline high -# -# 21st pack (`21st skills install`): tooling low, generation high, critique xhigh -21st-cli-use low -21st-registry low -21st-design-sync low -21st-ai high -21st-ui-explore high -21st-ui-review xhigh diff --git a/lib/effort-shift.md b/lib/effort-shift.md index fe92b1a..dc487d0 100644 --- a/lib/effort-shift.md +++ b/lib/effort-shift.md @@ -1,85 +1,58 @@ -# Effort shift — phase-level reasoning effort on the main loop (BDR-107) +# Route doctrine — phase-level model and effort on the main loop (BDR-107) Shared include, companion of `lib/model-gate.md`: the gate fixes WHICH model -reflects, this include fixes HOW HARD each phase thinks. The rungs are the -user's: low (fix a line, run a script) · medium (day-to-day) · high -(refactor, resisting bug) · xhigh (architecture, audit before validation) · -max (stuck error, judged need). +reflects, the model-router mod (`mods/model-router`) fixes HOW HARD each +phase thinks. Rungs: low (fix a line, run a script) · medium (day-to-day) · +high (refactor, resisting bug) · xhigh (architecture, audit before +validation) · max (stuck error, judged need). -## Mechanics (verified on Claude Code 2.1.283) +## The tool -- **Pairing rule**: a `Skill(effort-)` call applies its effort only - when the same assistant message carries at least one other tool call - after it; a lone Skill call is a no-op. Send the shift together with the - step's first tool call, shift first. That paired call already runs at the - new level: pair a downward shift with a pinned-agent dispatch or a - Read/Bash, never with a built-in judgment dispatch (`general-purpose`, - `model: "opus"`), which would inherit it. -- Re-loading a shifter already loaded in the conversation re-applies its - effort (the harness only dedupes the skill text), so bounce-back - sequences such as medium → max → medium work. -- A skill's `effort:` frontmatter applies from the moment it loads to the - end of the turn: on the user's `/skill` unconditionally, and on a - `Skill(...)` call by Claude only under the pairing rule above (a skill - Claude loads alone, such as `brainstorming` or `writing-plans`, applies - nothing). Last loaded wins, both directions. The prompt cache survives a - shift. -- **Stacked skills share one level**: skills that load together in one - build (the design stack) all pin the same level, since the last loaded - wins. Vendored externals get their level from `lib/effort-pins.txt`, - re-applied by `lib/effort-pins.sh` after every vendoring step; repo - skills carry it in their frontmatter. -- Dispatched agents run on their own `effort:` pin, never on a shift. - Unpinned agents inherit the level in force at dispatch. -- Headless sessions (`-p`, `claude agents`, SDK) ignore skill-level effort: - the run stays at the session level. `CLAUDE_CODE_EFFORT_LEVEL` beats every - frontmatter; keep it unset (the session banner warns). +`mcp__model-router__route` (params `phase` | `effort` | `clear`). It is a +deferred tool: when not loaded, run +`ToolSearch("select:mcp__model-router__route")` once per session. A route +applies from the next request on, paired with +another tool call or not (pairing only saves a request). The answer always +names the id and effort main runs on. A skill with a row routes itself on +load; a skill without one changes nothing, the last ROWED skill wins. + +## Wiring points + +1. Dispatch span starts → `route(phase="orchestrate")`, sent with the + dispatch. +2. Reflection resumes (challenge synthesis, verdict, plan revision) → + `route(phase="reflect")` or `"plan"` per the skill's own level; the line + before every `lib/challenge-plan.md` call. +3. Bookkeeping tail (memory commit, doc commit) → `route(phase="apply")`. +4. Escalation → `route(phase="escalate")`: verify-secure loop caps and + ship-feature STEP 4b. Not automatic: the challenge fail-safe and "gone + WRONG → STOP"; their STOP text names the levers below. +5. Built-in judgment dispatch (`general-purpose` `model="opus"`, `model: + "fable"` skill-runners) → explicit `effort=` on the Agent call (`xhigh` + for opus reviewers, `high` for fable runners). A main route never + reaches a child. Typed agents run on their row, never on a shift. +6. After a prose gate that ends the turn, the resumed reflection phase + starts with its own route call. + +## Run slot and levers + +A best-tier skill row survives the end of the turn (a run spans prose +gates); `/route clear`, `/route off` and a user `/model` drop it. Levers for a +relaunch: `ultrathink` in the prompt (turn floor) or `/route effort=max` +(sticky, `/route clear` after). +Builtin `/effort` is NOT a lever inside a run: rows and routes outrank it. + +## Limits + +- A skill typed while a background agent is live routes only through the + typed marker (unverified live 2026-10-10). +- Headless (`-p`, SDK) runs the hooks, so routing works there too. +- Mod off: typed agents fall back to their `model:`/`effort:` frontmatter. Measure the split any time: `python3 ~/.claude/lib/effort-audit.py` (thinking/output/cache tokens per scope, model and effort). -## Shifters - -`Skill(effort-low)` · `Skill(effort-medium)` · `Skill(effort-high)` · -`Skill(effort-xhigh)` · `Skill(effort-max)`. One tool call, one-line body, -always sent with another tool call (Pairing rule). -Typed by the user, `/effort-max` is a turn-scoped max: the relaunch lever -after a STOP. `ultrathink` only adds an in-context nudge; the API level -does not move. - -## Wiring — per orchestrator - -1. A dispatch span starts (executor, collector, fan-out) → - `Skill(effort-medium)`. -2. Reflection resumes after a dispatch span (challenge synthesis, verdict, - plan revision) → `Skill(effort-)`. Concretely: - the line before every `lib/challenge-plan.md` call. -3. The bookkeeping tail (memory commit, doc commit) → `Skill(effort-low)`. -4. Escalation → `Skill(effort-max)`, then the skill's own level again once - the diagnosis is produced. Automatic points: verify-secure loop caps - (GATE 0 floor, GATE 1 conformity, GATE 2 security) and ship-feature - STEP 4b. Not automatic, by doctrine: the challenge fail-safe (a mute - challenger is an infrastructure failure) and "gone WRONG → STOP" (STOP - precedes any further reasoning); their STOP text names the level - reached and suggests `/effort-max` for the relaunch. -5. Before any built-in or unpinned dispatch that carries judgment (a - `general-purpose` with `model: "opus"` or `"fable"`, the code reviewer - of requesting-code-review, a skill-runner) → `Skill(effort-)` - paired with that dispatch: built-ins inherit the level in force, and a - medium set earlier in the span would downgrade them. - -## Re-assert - -- After any nested `Skill(...)` whose frontmatter carries a different - effort (feat → commit-change), reload the orchestrator's own level. -- After a prose gate that ends the turn, the resumed turn runs at the - session level. If the resumed phase is reflection, its first step is - `Skill(effort-)`; dispatch and orchestration phases need - nothing. - ## Never -- A shift never inside a dispatched agent: pins rule there. +- A route inside a dispatched agent: its row rules there. - Max is for diagnosis, not for retrying the same fix harder. -- A medium shift never precedes a judgment dispatch in the same span - without an own-level shift paired with that dispatch. diff --git a/lib/model-check.sh b/lib/model-check.sh deleted file mode 100644 index 1cd3cf5..0000000 --- a/lib/model-check.sh +++ /dev/null @@ -1,33 +0,0 @@ -#!/usr/bin/env bash -# lib/model-check.sh — classify the persisted session model: big | small | unknown -# -# Witness for lib/model-gate.md (reflection requires a big model). Reads the -# "model" key of the user-scope settings (the file /model rewrites — LRN-098). -# Override the source with MODEL_CHECK_SETTINGS (tests use fixtures). -# -# stdout : : (raw = value found, empty if none) -# exit : 0 = big (fable/opus) · 2 = small (sonnet/haiku) · 3 = unknown -set -u - -SETTINGS="${MODEL_CHECK_SETTINGS:-$HOME/.claude/settings.json}" - -raw="" -if [ -f "$SETTINGS" ]; then - raw="$(python3 - "$SETTINGS" 2>/dev/null <<'PY' -import json, sys -try: - v = json.load(open(sys.argv[1])).get("model", "") - print(v if isinstance(v, str) else "") -except Exception: - print("") -PY -)" -fi - -norm="$(printf '%s' "$raw" | tr '[:upper:]' '[:lower:]')" -case "$norm" in - *opusplan*) printf 'unknown:%s\n' "$raw"; exit 3 ;; # opus-for-plan, sonnet otherwise — ambiguous - *fable*|*opus*) printf 'big:%s\n' "$raw"; exit 0 ;; - *sonnet*|*haiku*) printf 'small:%s\n' "$raw"; exit 2 ;; - *) printf 'unknown:%s\n' "$raw"; exit 3 ;; -esac diff --git a/lib/model-gate.md b/lib/model-gate.md index 874cd23..85a5ab7 100644 --- a/lib/model-gate.md +++ b/lib/model-gate.md @@ -1,53 +1,30 @@ # Model gate — reflection requires a big model (BLOCKING) -Shared include. Runs FIRST in any orchestrator whose reflection — -brainstorming, planning, contract, audit judgment, loop decisions — -executes inline or in inherit-model subagents. Sonnet-pinned executors are -not what this gate protects; it protects the thinking around them (BDR-066). +Shared include, runs FIRST in an orchestrator whose reflection executes +inline (BDR-066). The witness is the model-router mod's own route tool. -## 1. Self-check +## Entry call +ALWAYS call `mcp__model-router__route` with `phase` = the skill's row phase +(reflect or plan), no self-check shortcut. Tool not loaded (deferred) → +`ToolSearch("select:mcp__model-router__route")` once per session, then +call. The answer names the id main runs on next. -Your system prompt names the model powering this session. Fable or Opus → -big. Sonnet, Haiku, anything else → small. +| answer | action | +|---|---| +| names a fable or opus id | proceed, SILENT | +| names sonnet, haiku, anything else; "is off"; tool absent | **STOP** | -## 2. Witness — deterministic check +**STOP means**: print exactly `⛔ MODEL GATE — session on . +Reflection steps of this skill require Fable or Opus. Switch with /model, +then relaunch the skill.` (mod off: say so, `/route on` resumes it), then +end +the turn. No later step runs, no agent is dispatched, nothing is edited. - bash "$HOME/.claude/lib/model-check.sh" - -Output `:`; exit 0 = big, 2 = small, 3 = unknown. The witness -reads the PERSISTED model (settings.json — the file `/model` rewrites, -LRN-098). It can lag reality (session launched with `--model`, settings not -yet rewritten) — that is why the self-check exists alongside it. - -## 3. Verdict - -| self-check | witness | action | -|---|---|---| -| big | big (0) | proceed, SILENT — the nominal path prints nothing | -| small | any | **STOP** | -| big | small (2) | disagreement — **STOP**, surface BOTH values; the user confirms or relaunches | -| big | unknown (3) | fail-visible: print `model gate: witness unknown () — self-check says ` and ask the user to confirm before continuing (BDR-025: unknown never silently passes) | - -**STOP means**: print exactly - - ⛔ MODEL GATE — session on . Reflection steps of this skill - require Fable or Opus. Switch with /model, then relaunch the skill. - -then end the turn. No later step runs, no agent is dispatched, nothing is -edited. - -## 4. Dispatch tiers (BDR-077 — no inherit) - -The gate guards the MAIN loop only. Dispatched work NEVER inherits the -session model: typed agents run on their frontmatter pin; built-ins +## Dispatch tiers (BDR-077 — no inherit) +The gate guards the MAIN loop only. Typed agents are routed by their +model-router row (`model:` frontmatter = off-state floor). Built-ins (general-purpose / Explore / Plan) carry an explicit `model=` at every call -site — `model: "fable"` when the child performs reflection/orchestration on -the main loop's behalf (skill-runners), otherwise its complexity tier -(opus = dispatched judgment, sonnet = execution/collection, haiku = short -mechanical probes). - -Effort is the second axis of the same table (BDR-107): every typed agent -carries an `effort:` pin next to `model:`, and the main loop shifts per phase -through `lib/effort-shift.md`. No typed agent inherits either axis; -built-ins inherit the effort in force at dispatch, so an orchestrator shifts -before dispatching them (`lib/effort-shift.md`, wiring point 5). +site: `model: "fable"` when the child reflects/orchestrates for the main +loop (skill-runners), else its tier (opus = dispatched judgment, sonnet = +execution, haiku = mechanical probes). A built-in judgment dispatch also +carries an explicit `effort=` (`lib/effort-shift.md`). diff --git a/lib/tests/effort-pins.test.sh b/lib/tests/effort-pins.test.sh deleted file mode 100755 index 329136c..0000000 --- a/lib/tests/effort-pins.test.sh +++ /dev/null @@ -1,129 +0,0 @@ -#!/usr/bin/env bash -# lib/tests/effort-pins.test.sh — lib/effort-pins.sh's apply_effort_pins(): -# insert after `name:`, keep an equal level untouched, replace a different -# level inside the frontmatter only (a prose `effort:` in the body stays), -# skip a skill not vendored, insert before the closing `---` when the -# frontmatter has no name line, run idempotently, reject a bad level, a -# traversal name and a three-field line before writing anything, and -# parse the real map without error; hardening: last map line without a -# newline, unterminated frontmatter, CRLF file and read-only directory. All on a throwaway fixture repo. -set -u -ROOT="$(cd "$(dirname "$0")/../.." && pwd)" -LIB="$ROOT/lib/effort-pins.sh" -pass=0; fail=0 -check() { if [ "$2" = "$3" ]; then pass=$((pass+1)); echo "PASS $1" - else fail=$((fail+1)); echo "FAIL $1: got[$2] want[$3]"; fi; } -fm_effort() { awk 'NR==1&&/^---$/{p=1;next} p&&/^---$/{exit} p' "$1" \ - | sed -n 's/^effort: //p' | head -1; } - -WORK="$(mktemp -d)" || exit 1; trap 'rm -rf "$WORK"' EXIT -REPO="$WORK/repo"; EXT="$REPO/skills-external" -mkdir -p "$REPO/lib" "$EXT/alpha" "$EXT/beta" "$EXT/gamma" "$EXT/noname" -printf -- '---\nname: alpha\ndescription: a\n---\nbody\n' > "$EXT/alpha/SKILL.md" -printf -- '---\nname: beta\neffort: low\n---\nprose says effort: max here\n' > "$EXT/beta/SKILL.md" -printf -- '---\nname: gamma\neffort: low\n---\nbody\n' > "$EXT/gamma/SKILL.md" -printf -- '---\ndescription: no name line\n---\nbody\n' > "$EXT/noname/SKILL.md" -printf '# map\nalpha high\nbeta medium\ngamma low\nghost xhigh\nnoname low\n' > "$REPO/lib/effort-pins.txt" -gamma_before="$(cat "$EXT/gamma/SKILL.md")" - -bash "$LIB" "$REPO" >/dev/null 2>&1; check T1-rc-clean "$?" 0 -check T2-insert-after-name "$(sed -n '3p' "$EXT/alpha/SKILL.md")" "effort: high" -check T3-replace-in-frontmatter "$(fm_effort "$EXT/beta/SKILL.md")" "medium" -check T3b-body-prose-untouched "$(grep -c 'effort: max' "$EXT/beta/SKILL.md")" 1 -check T3c-single-effort-line "$(grep -c '^effort:' "$EXT/beta/SKILL.md")" 1 -check T4-equal-level-untouched "$(cat "$EXT/gamma/SKILL.md")" "$gamma_before" -check T5-missing-skill-skipped "$([ -e "$EXT/ghost" ] && echo created || echo absent)" absent -check T6-no-name-inserts-before-closing "$(sed -n '3p' "$EXT/noname/SKILL.md")" "effort: low" -check T6b-no-name-still-frontmatter "$(fm_effort "$EXT/noname/SKILL.md")" "low" -snap="$(cat "$EXT"/*/SKILL.md)" -bash "$LIB" "$REPO" >/dev/null 2>&1 -check T7-idempotent "$(cat "$EXT"/*/SKILL.md)" "$snap" -check T7b-no-tmp-left "$(find "$EXT" -name '*.tmp' | wc -l | tr -d ' ')" 0 - -# rejections: nothing written, rc 1 -for bad in 'alpha turbo' '../evil high' 'alpha high extra'; do - printf '%s\n' "$bad" > "$REPO/lib/effort-pins.txt" - out="$(bash "$LIB" "$REPO" 2>&1)"; rc=$? - check "T8-rejected[$bad]-rc" "$rc" 1 - check "T8-rejected[$bad]-named" "$(printf '%s' "$out" | grep -c 'rejected map line')" 1 -done -check T8b-tree-unchanged-after-rejections "$(cat "$EXT"/*/SKILL.md)" "$snap" -check T8c-no-evil-dir "$([ -e "$WORK/evil" ] && echo created || echo absent)" absent - -# the real map parses: fixture repo with the real map and no vendored skill -mkdir -p "$WORK/real/lib" "$WORK/real/skills-external" -cp "$ROOT/lib/effort-pins.txt" "$WORK/real/lib/" -out="$(bash "$LIB" "$WORK/real" 2>&1)"; check T9-real-map-parses "$?" 0 -check T9b-real-map-nothing-applied "$(printf '%s' "$out" | grep -c '0 applied, 0 already')" 1 -check T10-missing-map-rc "$(bash "$LIB" "$WORK/nowhere" >/dev/null 2>&1; echo $?)" 1 - -# hardening: each case in its own fixture repo -mkrepo() { R="$WORK/$1"; mkdir -p "$R/lib" "$R/skills-external/$2"; } -mkrepo h11 alpha; mkdir "$WORK/h11/skills-external/beta" -printf -- '---\nname: alpha\n---\nb\n' > "$WORK/h11/skills-external/alpha/SKILL.md" -printf -- '---\nname: beta\n---\nb\n' > "$WORK/h11/skills-external/beta/SKILL.md" -printf 'alpha high\nbeta low' > "$WORK/h11/lib/effort-pins.txt" -bash "$LIB" "$WORK/h11" >/dev/null 2>&1 -check T11-last-line-no-newline "$(fm_effort "$WORK/h11/skills-external/beta/SKILL.md")" low - -mkrepo h12 open; f12="$WORK/h12/skills-external/open/SKILL.md" -printf -- '---\nname: open\nbody effort: max\n' > "$f12"; b12="$(cat "$f12")" -printf 'open high\n' > "$WORK/h12/lib/effort-pins.txt" -out="$(bash "$LIB" "$WORK/h12" 2>&1)"; rc=$? -check T12-unterminated-frontmatter-skipped \ - "$rc|$(cat "$f12" | cmp -s - <(printf '%s\n' "$b12") && echo same)|$(printf '%s' "$out" | grep -c "ERR .*$f12")" "1|same|1" - -mkrepo h13 crlf -printf -- '---\r\nname: crlf\r\n---\r\nbody\r\n' > "$WORK/h13/skills-external/crlf/SKILL.md" -printf 'crlf high\n' > "$WORK/h13/lib/effort-pins.txt" -out="$(bash "$LIB" "$WORK/h13" 2>&1)"; rc=$? -check T13-crlf-file-rejected \ - "$rc|$(printf '%s' "$out" | grep -c 'ERR ')|$(printf '%s' "$out" | grep -c ' 0 applied, ')" "1|1|1" - -# T13b: the post-write re-read branch, reached with a no-op write stub -mkrepo h13b nowrite; f13b="$WORK/h13b/skills-external/nowrite/SKILL.md" -printf -- '---\nname: nowrite\n---\nb\n' > "$f13b" -out="$(bash -c 'source "$1"; _effort_pin_write() { return 0; } - _effort_pin_apply_one "$2" nowrite high' _ "$LIB" "$f13b" 2>&1)"; rc=$? -check T13b-reread-mismatch-fails \ - "$rc|$(printf '%s' "$out" | grep -c 'level not applied after write')" "1|1" - -if [ "${EFFORT_PINS_TEST_FAKE_ROOT:-0}" = 1 ] || [ "$(id -u)" -eq 0 ]; then - echo "SKIP T14-write-failure-no-temp: chmod bits ignored as root" -else - mkrepo h14 ro; d14="$WORK/h14/skills-external/ro" - printf -- '---\nname: ro\n---\nb\n' > "$d14/SKILL.md" - printf 'ro high\n' > "$WORK/h14/lib/effort-pins.txt" - chmod 555 "$d14"; out="$(bash "$LIB" "$WORK/h14" 2>&1)"; rc=$?; chmod 755 "$d14" - check T14-write-failure-no-temp \ - "$rc|$(printf '%s' "$out" | grep -c 'ERR ')|$(find "$d14" -name 'SKILL.md.*' | wc -l | tr -d ' ')" "1|1|0" -fi - -# T15: SIGINT during the awk write removes the temp sibling, exit 130 -mkrepo h15 sig; d15="$WORK/h15/skills-external/sig" -printf -- '---\nname: sig\n---\nb\n' > "$d15/SKILL.md" -bash -c 'source "$1"; awk() { kill -INT $$; sleep 2; } - _effort_pin_write "$2" sig high' _ "$LIB" "$d15/SKILL.md" >/dev/null 2>&1 -rc=$? -check T15-sigint-removes-temp \ - "$rc|$(find "$d15" -name 'SKILL.md.*' | wc -l | tr -d ' ')" "130|0" - -# T15b: previous INT trap restored on a normal return, no EXIT trap set -mkrepo h15b tr; d15b="$WORK/h15b/skills-external/tr" -printf -- '---\nname: tr\n---\nb\n' > "$d15b/SKILL.md" -out="$(bash -c 'source "$1"; trap "echo prev" INT - _effort_pin_write "$2" tr high - printf "INT:%s\n" "$(trap -p INT)"; printf "EXIT:%s\n" "$(trap -p EXIT)"' \ - _ "$LIB" "$d15b/SKILL.md" 2>&1)" -check T15b-traps-restored \ - "$(printf '%s' "$out" | grep -c "^INT:trap -- 'echo prev' SIGINT")|$(printf '%s' "$out" | grep -c '^EXIT:$')" "1|1" - -# T16: a literal backslash-t in a map line is printed shell-quoted -mkrepo h16 q -printf 'bad\\tname high\n' > "$WORK/h16/lib/effort-pins.txt" -out="$(bash "$LIB" "$WORK/h16" 2>&1)"; rc=$? -check T16-rejected-line-quoted \ - "$rc|$(printf '%s' "$out" | grep -cF 'bad\\tname')" "1|1" - -echo "effort-pins: $pass pass, $fail fail" -[ "$fail" -eq 0 ] diff --git a/lib/tests/effort-routing.test.sh b/lib/tests/effort-routing.test.sh index e6edeed..e2b9214 100755 --- a/lib/tests/effort-routing.test.sh +++ b/lib/tests/effort-routing.test.sh @@ -1,9 +1,12 @@ #!/usr/bin/env bash -# lib/tests/effort-routing.test.sh — census: effort tiering (BDR-107) -# agent pins, skill entry levels, shifter skills, orchestrator wiring, settings. -# shellcheck disable=SC2015,SC2016 # A && ok || ko is deliberate (ok/ko never fail); '$REPO' locks are literal source text +# lib/tests/effort-routing.test.sh — wave-2 census of the model-router rows. +# Drift lock: every tracked skill/agent row in mods/model-router/hooks/ +# register.ts equals its frontmatter (the off-state floor), the D3 wiring +# markers sit in the orchestrators, no shifter citer survives. +# shellcheck disable=SC2015,SC2016 # A && ok || ko is deliberate (ok/ko never fail) set -u R="$(cd "$(dirname "$0")/../.." && pwd)" +REG="$R/mods/model-router/hooks/register.ts" pass=0; fail=0 ok() { pass=$((pass+1)); } ko() { fail=$((fail+1)); printf 'FAIL %s\n' "$1"; } @@ -11,120 +14,133 @@ has() { if grep -qF "$2" "$R/$1"; then ok; else ko "$1 missing: $2"; fi; } lacks() { if grep -qF "$2" "$R/$1"; then ko "$1 must NOT contain: $2"; else ok; fi; } # frontmatter = the lines between the first two '---' lines fm() { awk 'NR==1&&/^---$/{p=1;next} p&&/^---$/{exit} p' "$1"; } -fm_effort() { fm "$1" | grep -E '^effort: (low|medium|high|xhigh|max)$' | head -1 | cut -d' ' -f2; } -fm_has_effort() { - got="$(fm_effort "$R/$1")" - if [ "$got" = "$2" ]; then ok; else ko "$1 frontmatter effort must be '$2', got '${got:-none}'"; fi +fm_val() { fm "$1" | grep -E "^$2: [a-z]+$" | head -1 | cut -d' ' -f2; } + +# ── register.ts parsers (awk/sed on the DEFAULT_CONFIG literal) ────────── +# rows -> "name phase" per row +rows() { + awk -v s="$2" '$0 ~ "^ "s": \\{"{f=1;next} f&&/^ \},?$/{f=0} f' "$1" \ + | grep -v '^ *//' | grep -oE "('[^']+'|[A-Za-z0-9_-]+): '[a-z]+'" \ + | sed -E "s/'//g; s/: / /" } -fm_no_effort() { if fm "$R/$1" | grep -q '^effort:'; then ko "$1 must NOT pin effort"; else ok; fi; } +# phase_effort -> " " +phase_effort() { + awk '/^ phases: \{/{f=1;next} f&&/^ \},?$/{f=0} f' "$1" \ + | sed -nE "s/^ *$2: \{ tier: '([a-z]+)', effort: '([a-z]+)' \},?$/\1 \2/p" +} +# tier_head -> first alias of the tier list +tier_head() { + awk '/^ tiers: \{/{f=1;next} f&&/^ \},?$/{f=0} f' "$1" \ + | sed -nE "s/^ *$2: \['([a-z]+)'.*$/\1/p" +} +row_of() { rows "$REG" "$1" | awk -v n="$2" '$1==n{print $2}'; } -# ── flip-test: the frontmatter reader must accept a valid level and reject an invalid one +# ── flip-test: the parsers read a fixture, reject a missing key ────────── FIX="$(mktemp -d)"; trap 'rm -rf "$FIX"' EXIT -printf -- '---\nname: good\neffort: xhigh\n---\nbody with effort: low in prose\n' > "$FIX/good.md" -printf -- '---\nname: bad\neffort: turbo\n---\n' > "$FIX/bad.md" -[ "$(fm_effort "$FIX/good.md")" = "xhigh" ] && ok || ko "flip: valid level not read" -[ -z "$(fm_effort "$FIX/bad.md")" ] && ok || ko "flip: invalid level accepted" -[ "$(fm "$FIX/good.md" | grep -c 'prose')" -eq 0 ] && ok || ko "flip: body leaked into frontmatter" +cat > "$FIX/reg.ts" <<'FX' + tiers: { + big: ['opus', 'fable'], + }, + phases: { + judge: { tier: 'big', effort: 'xhigh' }, + }, + agents: { + // judge + Plan: 'judge', 'plan-challenger': 'judge', + }, + skills: { + 'ship-feature': 'plan', doc: 'apply', + }, +FX +[ "$(rows "$FIX/reg.ts" agents | tr '\n' ,)" = "Plan judge,plan-challenger judge," ] \ + && ok || ko "flip: agents rows misparsed" +[ "$(rows "$FIX/reg.ts" skills | tr '\n' ,)" = "ship-feature plan,doc apply," ] \ + && ok || ko "flip: skills rows misparsed" +[ "$(phase_effort "$FIX/reg.ts" judge)" = "big xhigh" ] && ok || ko "flip: phase" +[ -z "$(phase_effort "$FIX/reg.ts" nothere)" ] && ok || ko "flip: ghost phase" +[ "$(tier_head "$FIX/reg.ts" big)" = "opus" ] && ok || ko "flip: tier head" +[ "$(rows "$REG" skills | wc -l)" -gt 40 ] && ok || ko "register.ts: skills rows not parsed" +[ "$(rows "$REG" agents | wc -l)" -gt 15 ] && ok || ko "register.ts: agents rows not parsed" -# ── 1) session default (spec D1) +# ── (b) tracked skills: row exists, frontmatter effort equals the row ──── +NO_ROW_SKILLS=" find-docs graphify impeccable model-router " +check_skill() { + local f="$1" name phase want got + name="$(basename "$(dirname "$f")")" + case "$NO_ROW_SKILLS" in *" $name "*) return;; esac + phase="$(row_of skills "$name")" + [ -n "$phase" ] || { ko "skills/$name: no row in register.ts"; return; } + want="$(phase_effort "$REG" "$phase" | cut -d' ' -f2)" + got="$(fm_val "$R/$f" effort)" + [ -n "$got" ] || { ko "skills/$name: routed skill without effort:"; return; } + [ "$got" = "$want" ] && ok || ko "skills/$name: effort $got != row $phase ($want)" +} +while IFS= read -r f; do check_skill "$f"; done < <( + cd "$R" && git ls-files 'skills/*/SKILL.md' 'skills-external/*/SKILL.md') + +# ── (c) tracked agents with a row: tier head == model:, effort == effort: ─ +NO_ROW_AGENTS=" interviewer client-handover-writer " +tier_alias() { tier_head "$REG" "$(phase_effort "$REG" "$1" | cut -d' ' -f1)"; } +check_agent() { + local f="$1" name phase alias want got + name="$(basename "$f" .md)" + case "$NO_ROW_AGENTS" in *" $name "*) return;; esac + case "$name" in impeccable-*) return;; esac + phase="$(row_of agents "$name")" + [ -n "$phase" ] || { ko "agents/$name: no row in register.ts"; return; } + alias="$(tier_alias "$phase")"; got="$(fm_val "$R/$f" model)" + [ "$got" = "$alias" ] && ok || ko "agents/$name: model $got != row $phase ($alias)" + [ "$alias" = haiku ] && return + want="$(phase_effort "$REG" "$phase" | cut -d' ' -f2)" + got="$(fm_val "$R/$f" effort)" + [ "$got" = "$want" ] && ok || ko "agents/$name: effort $got != row $phase ($want)" +} +while IFS= read -r f; do check_agent "$f"; done < <( + cd "$R" && git ls-files 'agents/*.md' | grep -E '^agents/[^/]+\.md$' \ + | grep -v '/README\.md$') + +# ── (d) no shifter citer in skills, agents, lib ────────────────────────── +if (cd "$R" && git grep -qE 'Skill\(effort-|EFFORT SHIFT[S]:' -- skills agents lib ':!lib/tests'); then + ko "a shifter-skill or shift-header citer survives in skills/agents/lib" +else ok; fi +# paths split so this file itself matches no deleted-name grep +for s in skills/effort-low skills/effort-max lib/effort-""pins.txt lib/model-""check.sh; do + [ ! -e "$R/$s" ] && ok || ko "$s must be deleted" +done + +# ── (e) D3 wiring markers ──────────────────────────────────────────────── +mark() { # mark + local ph="$1" s; shift + for s in "$@"; do has "skills/$s/SKILL.md" "route(phase=\"$ph\")"; done +} +mark orchestrate feat hotfix bugfix ship-feature init-project code-clean seo geo harden web-validate audit-delta +mark apply feat hotfix bugfix ship-feature init-project +mark reflect feat hotfix bugfix seo geo harden web-validate +mark plan ship-feature init-project onboard code-clean audit-delta +mark escalate ship-feature +[ "$(grep -o 'route(phase="escalate")' "$R/lib/verify-secure-loop.md" | wc -l)" -eq 3 ] \ + && ok || ko "verify-secure-loop.md: escalate route must appear 3 times" +for s in ship-feature init-project; do has "skills/$s/SKILL.md" 'effort="xhigh"'; done +has "skills/tour/SKILL.md" 'effort="xhigh"' +n_opus="$(grep -c 'model="opus",$' "$R/skills/onboard/SKILL.md")" +n_eff="$(grep -c 'effort="xhigh",$' "$R/skills/onboard/SKILL.md")" +[ "$n_opus" -ge 7 ] && [ "$n_opus" -eq "$n_eff" ] && ok \ + || ko "onboard: $n_opus model=\"opus\" dispatches vs $n_eff effort=\"xhigh\"" +bad="$(grep 'model: "fable"' "$R/agents/client-handover-writer.md" | grep -vc 'effort="high"')" +[ "$bad" -eq 0 ] && ok || ko "client-handover-writer: $bad model: \"fable\" line(s) without effort=\"high\"" +for s in ship-feature init-project feat bugfix web-validate seo hotfix geo harden code-clean audit-delta tour onboard; do + has "skills/$s/SKILL.md" 'ROUTING: follow $HOME/.claude/lib/effort-shift.md' +done +has "agents/client-handover-writer.md" 'ROUTING: follow $HOME/.claude/lib/effort-shift.md' + +# ── (f) doctrine includes ──────────────────────────────────────────────── +has "lib/model-gate.md" 'mcp__model-router__route' +has "lib/effort-shift.md" 'ToolSearch' +has "CLAUDE.global.md" 'route to `reflect` (high) through their' + +# ── (g) session default and hooks (unchanged locks) ────────────────────── has "settings.json" '"effortLevel": "high"' - -# ── 2) hooks: env-var warning + live effort in the statusline (spec D1, D5) has "hooks/session-start.sh" 'CLAUDE_CODE_EFFORT_LEVEL' has "hooks/statusline.sh" 'CLAUDE_EFFORT' -# ── 3) agent pins (spec D2): one effort per agent file, judgment mode wins on mode-based agents -for a in hotfixer release-executor plugin-probe validator-analyzer; do fm_has_effort "agents/$a.md" low; done -for a in feater bugfixer code-cleaner onboarder scaffolder; do fm_has_effort "agents/$a.md" medium; done -for a in refactorer analyzer commit-changer doc-syncer handover-doc-writer; do fm_has_effort "agents/$a.md" high; done -for a in plan-challenger plugin-advisor verifier security-auditor seo-analyzer geo-analyzer; do fm_has_effort "agents/$a.md" xhigh; done -for a in interviewer client-handover-writer status-reporter; do fm_no_effort "agents/$a.md"; done -has "skills/init-project/SKILL.md" 'pin sonnet, effort medium' - -# ── 4) skill entry levels (spec D3): the user's invocation sets the run's level -for s in status commit-change release-candidate doc capitalize close reconcile deploy profile plugin-check; do fm_has_effort "skills/$s/SKILL.md" low; done -for s in gitflow prune-memory; do fm_has_effort "skills/$s/SKILL.md" medium; done -for s in feat hotfix bugfix refactor web-validate harden seo geo; do fm_has_effort "skills/$s/SKILL.md" high; done -for s in ship-feature init-project onboard tour audit-delta analyze code-clean client-handover; do fm_has_effort "skills/$s/SKILL.md" xhigh; done -# BDR-108 round: the three repo skills that had no level -fm_has_effort "skills/skills-perso/SKILL.md" low -fm_has_effort "skills/pdf-translate/SKILL.md" medium -fm_has_effort "skills/site-motion/SKILL.md" high - -# ── 9) vendored externals carry the level of lib/effort-pins.txt (BDR-108). The files live in -# skills-external/ (gitignored, machine-owned): the durable artifact is the map + the re-apply -# after the last vendoring step of install-plugins.sh AND update-all.sh; a skill not vendored -# yet SKIPs visibly (fresh clone before make plugin). -while read -r s lvl _; do - case "$s" in ''|'#'*) continue ;; esac - if [ -f "$R/skills-external/$s/SKILL.md" ]; then fm_has_effort "skills-external/$s/SKILL.md" "$lvl" - else printf 'SKIP skills-external/%s/SKILL.md not vendored yet (run make plugin)\n' "$s"; fi -done < "$R/lib/effort-pins.txt" -has "lib/effort-pins.txt" 'brainstorming xhigh'; has "lib/effort-pins.txt" 'writing-plans xhigh' -has "install-plugins.sh" 'apply_effort_pins "$REPO"'; has "update-all.sh" 'apply_effort_pins "$REPO"' -lacks "install-plugins.sh" 'for _s in brainstorming writing-plans; do' -ln_last() { grep -n "$2" "$R/$1" | tail -1 | cut -d: -f1; } -[ "$(ln_last install-plugins.sh 'apply_effort_pins "$REPO"')" -gt "$(ln_last install-plugins.sh 'rm -rf "$TFD_STAGE"')" ] \ - && ok || ko "install-plugins.sh: effort pins must be re-applied after the 21st pack refresh" -pins_ln=$(ln_last update-all.sh 'apply_effort_pins "$REPO"') -[ "$pins_ln" -gt "$(ln_last update-all.sh 'skills-external/$_tfd_name')" ] \ - && [ "$pins_ln" -gt "$(ln_last update-all.sh 'vendor_pinned_skills superpowers refresh')" ] \ - && ok || ko "update-all.sh: effort pins must be re-applied after the last vendoring step (21st pack)" -[ -x "$R/lib/effort-pins.sh" ] && ok || ko "lib/effort-pins.sh missing or not executable" -# 9b) design stack = ONE level (last loaded wins); site-motion (repo skill) pins the same one -stack_levels() { awk '/^# design stack/{f=1;next} f&&/^#$/{f=0} f&&!/^#/&&NF==2{print $2}' "$R/lib/effort-pins.txt" | sort -u; } -[ "$(stack_levels | wc -l)" -eq 1 ] && ok || ko "design stack must share ONE level in lib/effort-pins.txt (got: $(stack_levels | tr '\n' ' '))" -[ "$(stack_levels | wc -l)" -ge 1 ] && fm_has_effort "skills/site-motion/SKILL.md" "$(stack_levels | head -1)" -has "lib/effort-shift.md" 'Stacked skills share one level' -has "CLAUDE.global.md" 'lib/effort-pins.txt' - -# ── 5) shifter skills + include (spec D4) -for l in low medium high xhigh max; do fm_has_effort "skills/effort-$l/SKILL.md" "$l"; has "skills/effort-$l/SKILL.md" "name: effort-$l"; done -has "lib/effort-shift.md" 'Headless sessions' -has "lib/effort-shift.md" 'Skill(effort-max)' -has "lib/effort-shift.md" 'never inside a dispatched agent' -has "lib/model-gate.md" 'lib/effort-shift.md' - -# ── 6) orchestrator wiring (spec D4) -for s in feat hotfix bugfix ship-feature init-project onboard tour code-clean seo geo harden web-validate audit-delta; do - has "skills/$s/SKILL.md" 'lib/effort-shift.md'; has "skills/$s/SKILL.md" 'a lone Skill call is a no-op'; done -for s in feat hotfix bugfix ship-feature init-project code-clean seo geo harden web-validate audit-delta; do - has "skills/$s/SKILL.md" 'Skill(effort-medium)'; done -lacks "skills/onboard/SKILL.md" 'Skill(effort-medium)'; lacks "skills/tour/SKILL.md" 'Skill(effort-medium)' -has "agents/client-handover-writer.md" 'lib/effort-shift.md'; lacks "agents/client-handover-writer.md" 'Skill(effort-medium)'; has "agents/client-handover-writer.md" 'Skill(effort-high)' -for s in feat hotfix bugfix; do has "skills/$s/SKILL.md" 'Skill(effort-high)'; done -for s in ship-feature init-project onboard code-clean audit-delta; do has "skills/$s/SKILL.md" 'Skill(effort-xhigh)'; done -for s in seo geo harden web-validate; do has "skills/$s/SKILL.md" 'Skill(effort-high)'; done -for s in feat hotfix bugfix ship-feature init-project; do has "skills/$s/SKILL.md" 'Skill(effort-low)'; done -has "skills/feat/SKILL.md" 'effort-shift: nested commit-change' - -# ── 6b) pairing rule documented (R11) -has "lib/effort-shift.md" 'lone Skill call is a no-op' -has "lib/effort-shift.md" 're-applies its' -[ "$(grep -c 'a lone Skill call is a no-op' "$R/skills/feat/SKILL.md")" -ge 1 ] && ok || ko "feat INC line must carry the pairing rule" - -# ── 7) escalation at max (spec D4) -[ "$(grep -c 'Skill(effort-max)' "$R/lib/verify-secure-loop.md")" -eq 3 ] && ok || ko "verify-secure-loop.md must shift to max at its 3 caps" -has "skills/ship-feature/SKILL.md" 'Skill(effort-max)' -has "lib/challenge-plan.md" '/effort-max' -has "lib/verify-secure-loop.md" '/effort-max' - -# ── 8) turn-reset re-assert after a prose gate followed by reflection -has "skills/bugfix/SKILL.md" 'effort-shift: turn reset' - -# ── 11) audit tooling -has "lib/effort-shift.md" 'effort-audit.py' -[ -x "$R/lib/effort-audit.py" ] && ok || ko "lib/effort-audit.py missing or not executable" - -# ── 6c) judgment dispatches re-raised, planning re-asserts, stronger locks (final review I1/I2/M5) -for s in ship-feature init-project; do has "skills/$s/SKILL.md" 'effort-shift: judgment dispatch'; has "skills/$s/SKILL.md" 'effort-shift: turn reset'; done -has "agents/client-handover-writer.md" 'effort-shift: judgment dispatch' -has "lib/effort-shift.md" 'Before any built-in or unpinned dispatch' -has "lib/model-gate.md" 'built-ins inherit the effort in force' -has "skills/ship-feature/SKILL.md" 'effort-shift: error recovery' -for s in feat hotfix bugfix seo geo harden web-validate ship-feature init-project onboard code-clean audit-delta; do has "skills/$s/SKILL.md" 'effort-shift: own level before the challenge'; done -has "update-all.sh" 'source "$REPO/lib/effort-pins.sh"' - -# ── summary (later tasks insert their locks ABOVE this line) -printf 'effort-routing census: %d pass, %d fail\n' "$pass" "$fail" -[ "$fail" -eq 0 ] +printf 'PASS=%s FAIL=%s\n' "$pass" "$fail"; [ "$fail" -eq 0 ] diff --git a/lib/tests/higgsfield.test.sh b/lib/tests/higgsfield.test.sh index f6ab348..689218b 100644 --- a/lib/tests/higgsfield.test.sh +++ b/lib/tests/higgsfield.test.sh @@ -321,24 +321,20 @@ expect link-sh "$(count link.sh higgsfield)" 0 expect profile-sh "$(count lib/profile.sh higgsfield)" 0 expect profiles \ "$(cat "$ROOT"/lib/profiles/*.profile | grep -cF higgsfield)" 0 -expect pins-map "$(count lib/effort-pins.txt higgsfield)" 0 verdict OFF_BY_DEFAULT_WIRING # ln_first / ln_last — line number of a match. ln_first() { grep -nF -- "$2" "$ROOT/$1" | head -1 | cut -d: -f1; } ln_last() { grep -nF -- "$2" "$ROOT/$1" | tail -1 | cut -d: -f1; } -PINS="apply_effort_pins \"\$REPO\"" -# install-plugins.sh: the sync sits in Step 8.6, before the effort pins -# (BDR-108); the CLI is proven by a probe, not by its shim; every login -# offer tests stdin alone (stdout is the tee pipe). +# install-plugins.sh: the sync sits in Step 8.6; the CLI is proven by a +# probe, not by its shim; every login offer tests stdin alone (stdout is +# the tee pipe). sync_ln="$(ln_last install-plugins.sh 'higgsfield_sync_skills')" expect after-8.5 "$(yn test "$sync_ln" -gt \ "$(ln_first install-plugins.sh 'Step 8.5: External skills')")" yes expect before-8.7 "$(yn test "$sync_ln" -lt \ "$(ln_first install-plugins.sh 'Step 8.7: 21st.dev')")" yes -expect before-pins "$(yn test "$sync_ln" -lt \ - "$(ln_last install-plugins.sh "$PINS")")" yes expect probe-gates \ "$(yn test "$(count install-plugins.sh 'if higgsfield_cli_ok')" -ge 3)" yes expect control "$(echo 'if [ -t 0 ] && [ -t 1 ]; then' | grep -cF -- '-t 1')" 1 @@ -347,14 +343,12 @@ expect stdin-tests \ "$(yn test "$(count install-plugins.sh '[ -t 0 ]')" -ge 3)" yes verdict INSTALL_WIRING -# update-all.sh: refresh before the 21st block and before the pins re-apply, -# and the updated CLI is proven by the probe, after the npm call. +# update-all.sh: refresh before the 21st block, and the updated CLI is +# proven by the probe, after the npm call. NPM_UP="npm install -g \"\$HF_PKG\"" sync_ln="$(ln_last update-all.sh 'higgsfield_sync_skills')" expect before-21st "$(yn test "$sync_ln" -lt \ "$(ln_first update-all.sh '7.4. Update the 21st.dev')")" yes -expect before-pins "$(yn test "$sync_ln" -lt \ - "$(ln_last update-all.sh "$PINS")")" yes expect probe-after-npm "$(yn test \ "$(ln_first update-all.sh 'higgsfield_cli_ok')" -gt \ "$(ln_last update-all.sh "$NPM_UP")")" yes diff --git a/lib/tests/model-check.test.sh b/lib/tests/model-check.test.sh deleted file mode 100644 index 040a8f5..0000000 --- a/lib/tests/model-check.test.sh +++ /dev/null @@ -1,25 +0,0 @@ -#!/usr/bin/env bash -# lib/tests/model-check.test.sh — flip-tests for lib/model-check.sh (LRN-096) -set -u -S="$(cd "$(dirname "$0")/../.." && pwd)/lib/model-check.sh" -pass=0; fail=0 -check() { if [ "$2" = "$3" ]; then pass=$((pass+1)); else fail=$((fail+1)); - printf 'FAIL %s: got[%s] want[%s]\n' "$1" "$2" "$3"; fi; } -T="$(mktemp -d)"; trap 'rm -rf "$T"' EXIT - -fx() { printf '{"model": "%s"}' "$1" > "$T/s.json"; } -run() { MODEL_CHECK_SETTINGS="$T/s.json" bash "$S" >"$T/out" 2>&1; echo "$?"; } - -fx 'claude-fable-5[1m]'; check T1-fable-exit "$(run)" 0 -check T1-fable-class "$(cut -d: -f1 <"$T/out")" big -fx 'claude-opus-4-8'; check T2-opus "$(run)" 0 -fx 'claude-sonnet-5'; check T3-sonnet "$(run)" 2 -fx 'claude-haiku-4-5-20251001'; check T4-haiku "$(run)" 2 -fx 'opusplan'; check T5-opusplan "$(run)" 3 -fx 'gpt-9-mega'; check T6-foreign "$(run)" 3 -printf '{"no_model": true}' > "$T/s.json"; check T7-no-key "$(run)" 3 -printf '{broken' > "$T/s.json"; check T8-malformed "$(run)" 3 -check T9-missing-file "$(MODEL_CHECK_SETTINGS="$T/absent.json" bash "$S" >/dev/null 2>&1; echo $?)" 3 - -printf 'model-check: %d pass, %d fail\n' "$pass" "$fail" -[ "$fail" -eq 0 ] diff --git a/lib/verify-secure-loop.md b/lib/verify-secure-loop.md index 2540670..78fc446 100644 --- a/lib/verify-secure-loop.md +++ b/lib/verify-secure-loop.md @@ -35,7 +35,7 @@ single `GATES — VERDICT:` line: - `UNMET(n)` → hand the dev the CONTRACT path + the `NOT-MET` rows verbatim, nothing else; re-run GATE 0. **No verifier is dispatched** — a red build or a red suite is not a judgement call, and paying an LLM to discover it is - waste. **Max 3 floor iterations** → `Skill(effort-max)` (effort-shift: cap reached, diagnose at max before escalating; send it in the same message as the first tool call that gathers the escalation evidence), then STOP + human escalation with the rows. + waste. **Max 3 floor iterations** → `mcp__model-router__route(phase="escalate")` (route: cap reached, diagnose at max before escalating; send it with the first tool call that gathers the escalation evidence), then STOP + human escalation with the rows. - `ABANDONED(n)` → floor green but a handoff stands. Continue to GATE 1; the verifier surfaces it and its `ABANDONED(n)` verdict routes to the human gate. @@ -74,7 +74,7 @@ Parse its single `VERIFY — VERDICT:` line: lines (NOT-MET / out-of-scope), nothing else: re-dispatch a FRESH executor with those inputs only, never redo the fix by hand. Then re-run GATE 0 and re-dispatch a FRESH verifier. Repeat. - **Max 3 conformity iterations** → `Skill(effort-max)` (effort-shift: cap reached, diagnose at max before escalating; send it in the same message as the first tool call that gathers the escalation evidence), then STOP + human escalation with the + **Max 3 conformity iterations** → `mcp__model-router__route(phase="escalate")` (route: cap reached, diagnose at max before escalating; send it with the first tool call that gathers the escalation evidence), then STOP + human escalation with the CRITERIA table (the contract-vs-realized diff). - `ABANDONED(n)` → direct human gate, never a dev loop (a dev cannot close what was proven impossible). The human lifts the abandonment or accepts @@ -104,9 +104,10 @@ Parse its single `SECURITY — VERDICT:` line: (re-dispatch a FRESH executor, never fix by hand). Then re-run GATE 0, then **re-verify the REQUEST first** (GATE 1, fresh verifier) — a security fix can drift the behavior — **then re-run GATE 2** (fresh auditor), in that - order. **Max 3 security iterations** → `Skill(effort-max)` (effort-shift: cap reached, diagnose at max before escalating; send it in the same message as the first tool call that gathers the escalation evidence), then STOP + human escalation with the + order. **Max 3 security iterations** → `mcp__model-router__route(phase="escalate")` (route: cap reached, diagnose at max before escalating; send it with the first tool call that gathers the escalation evidence), then STOP + human escalation with the BLOCKING table. Every STOP text names the level reached (`$CLAUDE_EFFORT`) - and suggests `/effort-max` for the relaunch. + and suggests relaunching with `ultrathink` in the prompt (turn floor) or + `/route effort=max` (sticky, `/route clear` after). - `DEGRADED` (semgrep absent) → does NOT block on the tool's absence; surface the checklist result + recommend `make plugin`. A DEGRADED run that still BLOCKs (grep-caught secret/injection) blocks like any other. diff --git a/mods/model-router/hooks/register.test.ts b/mods/model-router/hooks/register.test.ts index bebc952..7e8c9b4 100644 --- a/mods/model-router/hooks/register.test.ts +++ b/mods/model-router/hooks/register.test.ts @@ -59,24 +59,6 @@ const spawnInput = (model?: string) => ({ ...(model === undefined ? {} : { model }), }) -test('Skill(effort-low) is answered without next, route shows low', async ( - $, on) => { - let reached = false - on('tool.call', { tool: 'Skill' }, () => { - reached = true - return { result: { success: true, commandName: 'bottom' } } - }) - await boot($, on) - const out = await $.tool.call({ tool: 'Skill', skill: 'effort-low' }) - expect(out).toMatchObject({ - result: { success: true, commandName: 'effort-low' }, - }) - expect(reached).toBe(false) - const line = mainLine(await route($, 'show')) - expect(line).toContain('skill effort-low') - expect(line).toContain('effort low') -}) - test('route tool with phase orchestrate sets medium on main', async ( $, on) => { await boot($, on) @@ -987,14 +969,6 @@ test('run slot: route(clear) clears the turn and names the run', async ( expect(JSON.stringify(out)).toContain('run reflect still holds') }) -test('Skill(effort-low) bridge is not sticky', async ($, on) => { - await bootRun($, on) - await loadSkill($, 'effort-low') - expect(mainLine(await route($, 'show'))).toContain('effort low') - await endTurn($) - expect(mainLine(await route($, 'show'))).toContain('session defaults') -}) - test('route answer names the id even with a floor in force', async ( $, on) => { await bootRun($, on) diff --git a/mods/model-router/hooks/register.ts b/mods/model-router/hooks/register.ts index 8107cc5..a8ef430 100644 --- a/mods/model-router/hooks/register.ts +++ b/mods/model-router/hooks/register.ts @@ -48,8 +48,8 @@ type State = { rules: Rule[] source: string // 'defaults' or the override path userMain: Routed | null // /route by the user, sticky until /route clear - turnMain: Routed | null // model route tool, skill table row, Skill(effort-*) - // bridge, prompt default rule, derived orchestrate; dropped at turn end + turnMain: Routed | null // model route tool, skill table row, prompt + // default rule, derived orchestrate; dropped at turn end runMain: Routed | null // best-tier skill row: spans the turns of a run; // only a skill, /route clear|off or a user /model write or drop it turnFloor: Routed | null // user-explicit level for this turn (prompt rule): @@ -95,7 +95,6 @@ type Decision = { effort: Effort; by: EffortBy } const LEVELS: readonly Level[] = ['low', 'medium', 'high', 'xhigh', 'max'] const MODEL_ID = /^claude-[a-z0-9.-]+$/ const TOOL = 'mcp__model-router__route' -const EFFORT_SKILL = /^effort-(low|medium|high|xhigh|max)$/ const BEST = 'best' // the tier whose skill rows hold for a whole run // Origins that are a person typing: only these arm a typed slash. const TYPED_ORIGINS: ReadonlySet = new Set([ @@ -1378,34 +1377,6 @@ async function handleRouteTool($: Api, st: State, e: RouteInput) { // ---- skills ---------------------------------------------------------- -const skillResult = (skill: string, line: string) => ({ - result: { success: true, commandName: skill, status: 'inline' as const }, - context: [line], -}) - -/** Answers Skill(effort-) in place: one writer, the skill never loads. */ -function effortBridge(st: State, agentId: string | undefined, skill: string, - level: Level) { - if (agentId === undefined) { - const route = { ...st.turnMain?.route, effort: level } - st.turnMain = { phase: skill, route, source: 'skill' } - const note = mainNote(st, level) - return skillResult(skill, note - ? `model-router: ${skill} recorded, but ${note}; the ${skill} skill ` + - 'text was not loaded.' - : `model-router: effort → ${level} for this loop from the next ` + - `request on; the ${skill} skill text was not loaded.`) - } - const loop = loopOf(st, agentId) - if (loop.explicitEffort) { - return skillResult(skill, 'model-router: this agent was dispatched with ' + - 'an explicit effort; the shift does not apply.') - } - loop.effort = level - return skillResult(skill, `model-router: effort → ${level} for this ` + - `loop from the next request on; the ${skill} skill text was not loaded.`) -} - /** A skill's table row: its phase and that phase's route, if both exist. */ function skillRow(st: State, skill: string): Picked | undefined { const phase = hasKey(st.cfg.skills, skill) ? st.cfg.skills[skill] : undefined @@ -1793,8 +1764,6 @@ function registerSkills(on: On, st: State): void { on('tool.call', { tool: 'Skill' }, async ($, e, next) => { const skill = typeof e.skill === 'string' ? e.skill : undefined if (st.off || skill === undefined) return next(e) - const level = EFFORT_SKILL.exec(skill)?.[1] - if (isLevel(level)) return effortBridge(st, e.agentId, skill, level) st.skillCalls += 1 try { safely(st, $, 'Skill', () => onSkillLoad(st, skill, e.agentId)) diff --git a/skills/audit-delta/SKILL.md b/skills/audit-delta/SKILL.md index 4a13098..39de32d 100644 --- a/skills/audit-delta/SKILL.md +++ b/skills/audit-delta/SKILL.md @@ -29,7 +29,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. +ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route). Audit only what changed since the last run, on the axes the user picks. Per axis: **audit → approval gate → fix → re-verify → marker update**, @@ -170,7 +170,7 @@ Then show the user the same compact table inline. ### 3b-bis. CHALLENGE THE PROPOSALS (before the gate) -`Skill(effort-xhigh)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). +`mcp__model-router__route(phase="plan")` first (route: own level before the challenge; send it with the challenger dispatch). This axis' findings + proposed fixes are a proposal set worth attacking before the human gate. Persist THIS axis' finding list (not the whole append-only report) to `.claude/tasks/plans/--.md`, then run @@ -256,7 +256,7 @@ Then offer to capitalize (per CLAUDE.md): recurring finding patterns → below on the same delta (the SAST is a deterministic floor, the reasoned pass covers what grep/rules miss): ``` - Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message + mcp__model-router__route(phase="orchestrate") # route: dispatch span starts; send with the Agent call below in ONE message Agent(subagent_type="security-auditor", description="audit-delta security — semgrep SAST", prompt="MODE: audit\nSCOPE: \nREPORT: .claude/audits/.audit-delta-semgrep.md\nFollow agents/security-auditor.md exactly. Pinned rulesets, no login. Write ONLY to REPORT. End with REPORT_WRITTEN: .") ``` diff --git a/skills/bugfix/SKILL.md b/skills/bugfix/SKILL.md index 84051fd..25f3f3f 100644 --- a/skills/bugfix/SKILL.md +++ b/skills/bugfix/SKILL.md @@ -26,7 +26,7 @@ allowed-tools: MODEL GATE (blocking): run `$HOME/.claude/lib/model-gate.md` BEFORE any step below. Verdict `small` → STOP — print the gate's remedy, end the turn, dispatch nothing. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. +ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route). ## REQUEST $ARGUMENTS @@ -119,14 +119,14 @@ RISK: obvious fix. - If the fix is significant (>10 lines, multiple files, behavior change): wait for user approval. - On resume: `Skill(effort-high)` first, sent with the next tool call (effort-shift: turn reset). + On resume: `mcp__model-router__route(phase="reflect")` first (route: resumed turn). - Then run pass B of `$HOME/.claude/lib/contract-interview.md` against the FIX PLAN: every VISIBLE / PUBLIC NAME / SCOPE choice it settles that the bug report left open → one batch of questions, before STEP 3b. The trivial fast-path is not exempt: a 1-line fix with a visible choice still asks. ## STEP 3b — CHALLENGE THE FIX PLAN (before the contract) -`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). +`mcp__model-router__route(phase="reflect")` first (route: own level before the challenge; send it with the challenger dispatch). Unless the fix is the trivial 1-2 line case STEP 3 already fast-paths, the DIAGNOSIS + FIX PLAN is a reflection worth attacking before it hardens into a contract. Persist it to `.claude/tasks/plans/--.md`, then run @@ -157,10 +157,10 @@ branch it's a no-op (commit in place). Never `finish`. ## STEP 5 — DISPATCH EXECUTOR -Dispatch the executor — sonnet by frontmatter pin, do not override: +Dispatch the executor — sonnet by its row (frontmatter = off-state floor), do not override: ``` -Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message +mcp__model-router__route(phase="orchestrate") # route: dispatch span starts; send with the Agent call below in ONE message Agent(subagent_type="bugfixer") prompt: "CONTRACT: DIAGNOSIS: @@ -282,7 +282,7 @@ A bugfix with an understood root cause is almost always worth one entry: If the bug was trivial and the root cause not transferable → skip with `CAPITALIZE: trivial, skip`. -`Skill(effort-low)` first (effort-shift: bookkeeping tail; send it in the same message as the memory-commit command). +`mcp__model-router__route(phase="apply")` first (route: bookkeeping tail; send it with the memory-commit command). **Then commit the memory** — follow `$HOME/.claude/lib/capitalize-commit.md`: it surgically commits what capitalize just wrote (`.claude/memory` + `.claude/tasks` diff --git a/skills/code-clean/SKILL.md b/skills/code-clean/SKILL.md index 72458cb..e3c429c 100644 --- a/skills/code-clean/SKILL.md +++ b/skills/code-clean/SKILL.md @@ -26,7 +26,7 @@ allowed-tools: MODEL GATE (blocking): run `$HOME/.claude/lib/model-gate.md` BEFORE any step below. Verdict `small` → STOP — print the gate's remedy, end the turn, dispatch nothing. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. +ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route). ## TARGET $ARGUMENTS @@ -122,7 +122,7 @@ TOTALS: If no issues found: report clean state and stop. ## STEP 3b — CHALLENGE THE SCOPE (before approval) -`Skill(effort-xhigh)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). +`mcp__model-router__route(phase="plan")` first (route: own level before the challenge; send it with the challenger dispatch). The STEP 3 report is the proposed cleanup scope — worth attacking before the human approves it. It is still inline, so FIRST persist it to `.claude/tasks/plans/--.md` (STEP 3 report format, one item @@ -172,10 +172,10 @@ is approved, stop — no dispatch. severity — proposed fix`. This is the executor's scope-of-work on disk — named, auditable, the same contract discipline as the dev gates (verifier reads its contract from disk). -2. **Dispatch the executor** — sonnet by frontmatter pin, do not override: +2. **Dispatch the executor** — sonnet by its row (frontmatter = off-state floor), do not override: ``` - Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message + mcp__model-router__route(phase="orchestrate") # route: dispatch span starts; send with the Agent call below in ONE message Agent(subagent_type="code-cleaner") prompt: "SCOPE: .claude/audits/CODE-CLEAN-SCOPE.md APPROVED: diff --git a/skills/commit-change/SKILL.md b/skills/commit-change/SKILL.md index 876b162..33f703b 100644 --- a/skills/commit-change/SKILL.md +++ b/skills/commit-change/SKILL.md @@ -73,7 +73,7 @@ $ARGUMENTS" (`model="opus"` — BDR-077: propose = narrative reconstruction + capitalize routing, judgment tier; the call-site override takes precedence over the -sonnet frontmatter pin. Apply, STEP 4, stays on the pin.) +sonnet row (frontmatter = off-state floor). Apply, STEP 4, stays on its row.) Read the returned `COMMIT PLAN` + `EDGE CASES` + `CAPITALIZE CANDIDATES`, terminated by `READY TO APPLY — awaiting dispatcher confirmation`. diff --git a/skills/effort-high/SKILL.md b/skills/effort-high/SKILL.md deleted file mode 100644 index cbeaeaa..0000000 --- a/skills/effort-high/SKILL.md +++ /dev/null @@ -1,6 +0,0 @@ ---- -name: effort-high -description: Investigation shift. Deeper reasoning for diagnosis, LOCATE, contract drafting, refactor judgement inside feat, hotfix and bugfix runs. -effort: high ---- -Effort shifted to high for the rest of this turn (lib/effort-shift.md). Continue with the caller's next step. diff --git a/skills/effort-low/SKILL.md b/skills/effort-low/SKILL.md deleted file mode 100644 index 16da6a4..0000000 --- a/skills/effort-low/SKILL.md +++ /dev/null @@ -1,6 +0,0 @@ ---- -name: effort-low -description: Bookkeeping shift. Lowers reasoning to the cheapest level for the rest of the turn: journal lines, memory commits, capitalize, release bookkeeping, status output. -effort: low ---- -Effort shifted to low for the rest of this turn (lib/effort-shift.md). Continue with the caller's next step. diff --git a/skills/effort-max/SKILL.md b/skills/effort-max/SKILL.md deleted file mode 100644 index 06de425..0000000 --- a/skills/effort-max/SKILL.md +++ /dev/null @@ -1,6 +0,0 @@ ---- -name: effort-max -description: Escalation shift. Maximum reasoning when a verify or security loop hits its cap, a gate fails twice, or error recovery starts in ship-feature. -effort: max ---- -Effort shifted to max for the rest of this turn (lib/effort-shift.md). Continue with the caller's next step. diff --git a/skills/effort-medium/SKILL.md b/skills/effort-medium/SKILL.md deleted file mode 100644 index 6115c86..0000000 --- a/skills/effort-medium/SKILL.md +++ /dev/null @@ -1,6 +0,0 @@ ---- -name: effort-medium -description: Orchestration shift. Standard reasoning between two dispatches: read a subagent report, pick the next step, relay a gate verdict, route a branch. -effort: medium ---- -Effort shifted to medium for the rest of this turn (lib/effort-shift.md). Continue with the caller's next step. diff --git a/skills/effort-xhigh/SKILL.md b/skills/effort-xhigh/SKILL.md deleted file mode 100644 index 6a60abd..0000000 --- a/skills/effort-xhigh/SKILL.md +++ /dev/null @@ -1,6 +0,0 @@ ---- -name: effort-xhigh -description: Reflection shift. Deep reasoning for brainstorm, planning, challenge synthesis and audit verdicts before a human validation gate. -effort: xhigh ---- -Effort shifted to xhigh for the rest of this turn (lib/effort-shift.md). Continue with the caller's next step. diff --git a/skills/feat/SKILL.md b/skills/feat/SKILL.md index 4b10386..aecbde7 100644 --- a/skills/feat/SKILL.md +++ b/skills/feat/SKILL.md @@ -26,7 +26,7 @@ allowed-tools: MODEL GATE (blocking): run `$HOME/.claude/lib/model-gate.md` BEFORE any step below. Verdict `small` → STOP — print the gate's remedy, end the turn, dispatch nothing. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. +ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route). ## REQUEST $ARGUMENTS @@ -124,7 +124,7 @@ in the contract's CLARIFICATIONS `[gated]` and in the plan. A choice that surfaces only during execution comes back as `NEED-DECISION` (STEP 3). ## STEP 1b — CHALLENGE THE PLAN (before branching) -`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). +`mcp__model-router__route(phase="reflect")` first (route: own level before the challenge; send it with the challenger dispatch). The STEP 1 plan is a reflection worth attacking before a branch is spent on it. Persist it to `.claude/tasks/plans/--.md`, then run `$HOME/.claude/lib/challenge-plan.md` with `PLAN` = that file, `KIND` = `build-plan`, @@ -143,10 +143,10 @@ branch it's a no-op (commit in place). Never `finish`. ## STEP 3 — DISPATCH EXECUTOR -Dispatch the executor — sonnet by frontmatter pin, do not override: +Dispatch the executor — sonnet by its row (frontmatter = off-state floor), do not override: ``` -Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message +mcp__model-router__route(phase="orchestrate") # route: dispatch span starts; send with the Agent call below in ONE message Agent(subagent_type="feater") prompt: "CONTRACT: PLAN: @@ -203,7 +203,7 @@ test), consider splitting into 2-3 atomic commits grouped by logical unit — or run `/commit-change` on the pending work (it dispatches the commit-changer (propose opus / apply sonnet, BDR-077); never inline-load the bare agent, it is now a propose/apply executor). -Then `Skill(effort-high)` (effort-shift: nested commit-change loaded at low; reload feat's level, sent with the next tool call). +Then `mcp__model-router__route(phase="reflect")` (route: nested commit-change loaded at apply; reload feat's level). Print summary: ``` @@ -254,7 +254,7 @@ Always append a 1-line entry to today's heading in `.claude/memory/journal.md`. If no substantive capture candidate → skip with `CAPITALIZE: nothing to log`. -`Skill(effort-low)` first (effort-shift: bookkeeping tail; send it in the same message as the memory-commit command). +`mcp__model-router__route(phase="apply")` first (route: bookkeeping tail; send it with the memory-commit command). **Then commit the memory** — follow `$HOME/.claude/lib/capitalize-commit.md`: it surgically commits what capitalize just wrote (`.claude/memory` + `.claude/tasks` diff --git a/skills/geo/SKILL.md b/skills/geo/SKILL.md index ee76ad9..7a5533a 100644 --- a/skills/geo/SKILL.md +++ b/skills/geo/SKILL.md @@ -29,7 +29,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. +ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route). Dispatches the `geo-analyzer` subagent (audit + fix bundle), then applies the bundle from THIS main loop at **L1** — same shape as `/web-validate` @@ -47,7 +47,7 @@ every phase (LRN-126). Clean `.audit/geo-signals-.md` after apply. **A — collect (sonnet):** ``` -Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message +mcp__model-router__route(phase="orchestrate") # route: dispatch span starts; send with the Agent call below in ONE message Agent(subagent_type="geo-analyzer", model="sonnet") prompt: "MODE: collect RUNID: @@ -86,7 +86,7 @@ your bundle." ``` ## STEP 1b — CHALLENGE THE FIX BUNDLE (advisory, before apply) -`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). +`mcp__model-router__route(phase="reflect")` first (route: own level before the challenge; send it with the challenger dispatch). The analyzer returned a `## FIX BUNDLE` — worth attacking before any edit lands. **Skip if intervention mode = conservative** (nothing is applied). Else persist the bundle verbatim to `.claude/tasks/plans/--.md`, then run @@ -118,7 +118,7 @@ intent, not header wording: **AUTO** = no-confirmation items (G1–G4/G6); For each AUTO item, dispatch its `applier` at L1, passing the item verbatim: ``` -Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message +mcp__model-router__route(phase="orchestrate") # route: dispatch span starts; send with the Agent call below in ONE message Agent(subagent_type="hotfixer") # or "feater" per the item's applier prompt: ". diff --git a/skills/harden/SKILL.md b/skills/harden/SKILL.md index d6d53e3..27968fa 100644 --- a/skills/harden/SKILL.md +++ b/skills/harden/SKILL.md @@ -29,7 +29,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. +ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route). This skill orchestrates a narrow-scope hardening audit: TLS + security headers + redirects + canonical + custom 404 + server configs. It @@ -261,7 +261,7 @@ seo-analyzer will run in parallel. Spawn a single seo-analyzer subagent with an explicit IN/OUT scope list. ``` -Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message +mcp__model-router__route(phase="orchestrate") # route: dispatch span starts; send with the Agent call below in ONE message Agent( subagent_type="seo-analyzer", description="harden — narrow-scope web hardening audit", @@ -522,7 +522,7 @@ Extract the score and critical-alert count from `.claude/audits/HARDEN.md` for t --- ## STEP 2b — CHALLENGE THE FIX BUNDLE (MODE=fix only, advisory) -`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). +`mcp__model-router__route(phase="reflect")` first (route: own level before the challenge; send it with the challenger dispatch). Skip if MODE=audit (no bundle exists). Else, before the STEP 3 gate, harden the bundle: extract the `## 8. Fix bundle` section from HARDEN.md to `.claude/tasks/plans/--.md` (a clean, blind-judgeable artifact), then run diff --git a/skills/hotfix/SKILL.md b/skills/hotfix/SKILL.md index 61530b1..9ed5296 100644 --- a/skills/hotfix/SKILL.md +++ b/skills/hotfix/SKILL.md @@ -24,7 +24,7 @@ allowed-tools: MODEL GATE (blocking): run `$HOME/.claude/lib/model-gate.md` BEFORE any step below. Verdict `small` → STOP — print the gate's remedy, end the turn, dispatch nothing. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. +ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route). ## REQUEST $ARGUMENTS @@ -94,7 +94,7 @@ point. Run it ONLY when the settled fix touches control flow or behaviour — an off-by-one, a wrong operator/variable, a behaviour-changing config value, or a missing import that alters execution. In doubt → it is probably a `/bugfix`. -`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). +`mcp__model-router__route(phase="reflect")` first (route: own level before the challenge; send it with the challenger dispatch). For a logic fix: persist the STEP 1 located fix (root cause + the exact edit) to `.claude/tasks/plans/--.md`, then run `$HOME/.claude/lib/challenge-plan.md` with `PLAN` = that file, `KIND` = @@ -135,10 +135,10 @@ mentioned: STOP and ask `"working tree dirty: stash and continue, or abort?"`. ## STEP 3 — DISPATCH EXECUTOR -Dispatch the executor — sonnet by frontmatter pin, do not override: +Dispatch the executor — sonnet by its row (frontmatter = off-state floor), do not override: ``` -Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message +mcp__model-router__route(phase="orchestrate") # route: dispatch span starts; send with the Agent call below in ONE message Agent(subagent_type="hotfixer") prompt: "CONTRACT: LOCATED: @@ -236,7 +236,7 @@ Always append a 1-line entry to today's heading in `.claude/memory/journal.md` ( **Language rule**: the journal line and any proposed BLK/LRN entries are ALWAYS written English AND caveman — fragments, articles dropped, code/IDs/quoted errors verbatim — per CLAUDE.md "Memory registries" (Always English, always caveman). -`Skill(effort-low)` first (effort-shift: bookkeeping tail; send it in the same message as the memory-commit command). +`mcp__model-router__route(phase="apply")` first (route: bookkeeping tail; send it with the memory-commit command). **Then commit the memory** — follow `$HOME/.claude/lib/capitalize-commit.md`: it surgically commits what capitalize just wrote (`.claude/memory` + `.claude/tasks` diff --git a/skills/init-project/SKILL.md b/skills/init-project/SKILL.md index a7bc5ff..fba9456 100644 --- a/skills/init-project/SKILL.md +++ b/skills/init-project/SKILL.md @@ -14,7 +14,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. +ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route). ## REQUEST $ARGUMENTS @@ -181,12 +181,12 @@ This is the deterministic scaffold commit owner (closes BLK-010). The MVP is implemented on a `feature/*` branch off `develop` (STEP 8). ## STEP 6 — PLAN -`Skill(effort-xhigh)` first, sent with the next tool call (effort-shift: turn reset; gate #1 ended the turn and the vendored `writing-plans` pin applies only when the user invokes it). +`mcp__model-router__route(phase="plan")` first (route: resumed turn; gate #1 ended the turn). Invoke `writing-plans` (vendored superpowers skill) with BRIEF + skeleton. Granular tasks (2-5 min each), exact file paths, TDD: tests before code. ## STEP 6b — CHALLENGE THE PLAN (before the gate) -`Skill(effort-xhigh)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). +`mcp__model-router__route(phase="plan")` first (route: own level before the challenge; send it with the challenger dispatch). Before the human sees the implementation plan, harden it. Run `$HOME/.claude/lib/challenge-plan.md` with `PLAN` = the plan STEP 6 wrote under `docs/superpowers/plans/`, `KIND` = `build-plan`, `SCOPE` = the skeleton + task file @@ -212,7 +212,7 @@ Approve and start? (yes / request changes) Changes → back to STEP 6. Approved → continue. ## STEP 8 — IMPLEMENT -First: `Skill(effort-medium)` (effort-shift: dispatch span starts; send it in the same message as this step's first dispatch). +First: `mcp__model-router__route(phase="orchestrate")` (route: dispatch span starts; send it with this step's first dispatch). Start the MVP feature branch off develop, then implement on it: ```bash bash "$HOME/.claude/lib/gitflow.sh" start feature mvp @@ -224,7 +224,8 @@ Invoke `subagent-driven-development` (vendored superpowers skill) for the per-ta finishing-a-development-branch", stop and return. **Model routing (BDR-066):** every subagent dispatched under SDD — per-task -implementers AND its reviewers — MUST carry `model: "sonnet"` in the Agent +implementers AND its reviewers — MUST carry `model: "sonnet"` and +`effort="medium"` (the `implement` level; a main route never reaches a child) in the Agent call. The plan is closed; execution and plan-conformity review are sonnet work. Reflection (task decomposition, review verdict arbitration) stays in this loop. @@ -261,9 +262,8 @@ against the founding contract. Distinct axis from STEP 10 code review ([[LRN-095]]) — both run. ## STEP 10 — CODE REVIEW -`Skill(effort-xhigh)` first, sent with the review dispatch (effort-shift: judgment dispatch; the reviewer is a built-in and inherits the level in force). Invoke `requesting-code-review` (vendored superpowers skill). **Model routing (BDR-077):** the -review subagent it dispatches MUST carry `model: "opus"` in the Agent call — +review subagent it dispatches MUST carry `model: "opus"` and `effort="xhigh"` in the Agent call — craft review is dispatched judgment, never inherited from the session. Fix all CRITICAL before proceeding. @@ -316,7 +316,7 @@ articles dropped, code/IDs/quoted errors verbatim — per CLAUDE.md "Memory registries" (Always English, always caveman). The gate may mirror the user's language; entries must not. -`Skill(effort-low)` first (effort-shift: bookkeeping tail; send it in the same message as the memory-commit command). +`mcp__model-router__route(phase="apply")` first (route: bookkeeping tail; send it with the memory-commit command). **Then commit the memory** — follow `$HOME/.claude/lib/capitalize-commit.md`: it surgically commits the approved founding decisions (`.claude/memory` + diff --git a/skills/onboard/SKILL.md b/skills/onboard/SKILL.md index bf7bc43..a4f10b8 100644 --- a/skills/onboard/SKILL.md +++ b/skills/onboard/SKILL.md @@ -14,7 +14,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. +ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route). ## REQUEST $ARGUMENTS @@ -368,7 +368,7 @@ Lire le bloc `audit_stack:` du fichier `~/.claude/lib/project-archetypes//dev/null | grep -E "^gstack\s+ Agent( subagent_type="general-purpose", model="opus", + effort="xhigh", description="Onboard — security audit fallback (archetype-adaptive)", prompt=""" READ-ONLY security audit. No file modifications. @@ -546,6 +548,7 @@ flux de dev sont deux formes distinctes ([[BDR-050]] pipeline dev ≠ audit). Agent( subagent_type="doc-syncer", model="opus", + effort="xhigh", description="Onboard — doc drift audit only", prompt=""" MODE: audit — REPORT-ONLY, NO edits, NO auto-sync (no patch dispatch @@ -666,6 +669,7 @@ Si le skill ne supporte pas `--output`, capturer la sortie et écrire à la main Agent( subagent_type="general-purpose", model="opus", + effort="xhigh", description="Onboard — static design review fallback", prompt=""" AUDIT-ONLY mode — NO edits. Static design review du code UI. @@ -710,6 +714,7 @@ Puis parser le JSON Lighthouse (scores perf/a11y/bp/seo/pwa + top opportunities) Agent( subagent_type="general-purpose", model="opus", + effort="xhigh", description="Onboard — static perf audit", prompt=""" AUDIT-ONLY mode — NO edits. @@ -753,6 +758,7 @@ Parser axe-core résultats (violations, incomplete, inapplicable, passes) → `. Agent( subagent_type="general-purpose", model="opus", + effort="xhigh", description="Onboard — static a11y audit", prompt=""" AUDIT-ONLY mode — NO edits. @@ -799,6 +805,7 @@ Spawn un subagent synthétiseur (isolé, chargé uniquement du contenu de `.onbo Agent( subagent_type="general-purpose", model="opus", + effort="xhigh", description="Onboard — synthèse vers .claude/audits/", prompt=""" Lire tous les fichiers de /.onboard-audit/ : @@ -892,7 +899,7 @@ Vérifier que les 4 fichiers `.claude/audits/ONBOARD_REPORT.md`, `.claude/audits --- ## STEP 7b — CHALLENGE THE PROPOSALS (before the human gate) -`Skill(effort-xhigh)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). +`mcp__model-router__route(phase="plan")` first (route: own level before the challenge; send it with the challenger dispatch). The 4 audit files are on disk; `AUDIT_PROPOSALS.md` is the artifact worth attacking before the human spends a gate on it. Run `$HOME/.claude/lib/challenge-plan.md` with `PLAN` = diff --git a/skills/seo/SKILL.md b/skills/seo/SKILL.md index ea3a82e..cb4c77e 100644 --- a/skills/seo/SKILL.md +++ b/skills/seo/SKILL.md @@ -30,7 +30,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. +ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route). This skill orchestrates TWO specialist agents running in parallel, then merges their output into a single `.claude/audits/SEO.md` report. It is the main @@ -325,7 +325,7 @@ templating. **PHASE A — collect (both domains, one message):** ``` -Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message +mcp__model-router__route(phase="orchestrate") # route: dispatch span starts; send with the Agent call below in ONE message Agent(subagent_type="seo-analyzer", model="sonnet") prompt: """ MODE: collect @@ -460,7 +460,7 @@ sentinel, emit the COLLECT REPORT, stop. No scoring, no bundle. ``` **PHASE B — judge (both domains, one message, AFTER both COLLECT REPORTs -are DONE):** no `model=` override — the opus frontmatter pins apply. +are DONE):** no `model=` override — the opus rows apply (frontmatter = off-state floor). ``` Agent(subagent_type="seo-analyzer") @@ -510,7 +510,7 @@ the reports." ``` ## STEP 1b — CHALLENGE THE FIX BUNDLE (advisory, before apply) -`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). +`mcp__model-router__route(phase="reflect")` first (route: own level before the challenge; send it with the challenger dispatch). Both envelopes now carry a `## FIX BUNDLE` — worth attacking before any edit lands. **Skip if intervention mode = conservative** (nothing is applied). Else persist both bundles (seo + geo, verbatim) to `.claude/tasks/plans/--.md`, then run @@ -558,7 +558,7 @@ The two bundles may touch the same shared template (meta vs JSON-LD). Apply For each AUTO item, dispatch its `applier` at L1, passing the item verbatim: ``` -Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message +mcp__model-router__route(phase="orchestrate") # route: dispatch span starts; send with the Agent call below in ONE message Agent(subagent_type="hotfixer") # or "feater" per the item's applier prompt: ". diff --git a/skills/ship-feature/SKILL.md b/skills/ship-feature/SKILL.md index a941f35..6ae8ab6 100644 --- a/skills/ship-feature/SKILL.md +++ b/skills/ship-feature/SKILL.md @@ -14,7 +14,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. +ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route). ## REQUEST $ARGUMENTS @@ -114,10 +114,10 @@ Inject ONLY what constrains: the NON-BINDING count does NOT enter the brainstorm (the injection inherits the OUTPUT filter — detail what binds, drop what doesn't). Consumption = INPUT INJECTION (we can't modify the external skill; we control its input). Refine request into validated design via Socratic questioning. Don't proceed until design approved. -Turns after a user reply run at the session level until a tool call is paired with `Skill(effort-xhigh)` (effort-shift: turn reset). +The first step of every turn after a user reply is `mcp__model-router__route(phase="plan")` (route: resumed turn). ## STEP 2 — PLAN -`Skill(effort-xhigh)` first, sent with the next tool call (effort-shift: turn reset; brainstorm turns after a user reply run at the session level, and the vendored `brainstorming` pin applies only when the user invokes it). +`mcp__model-router__route(phase="plan")` first (route: resumed turn after the brainstorm gate). Invoke `writing-plans` (vendored superpowers skill) with the validated design AND the 0d digest: every task must be consistent with the in-force constraints; where a task implements or affects one, note the ID inline. Break design into tasks (2-5 min each). Each task: exact file paths, full code, verification steps. @@ -127,7 +127,7 @@ request nor the STEP 1 brainstorm settled (check the contract's CLARIFICATIONS first) → one batch before STEP 2b; answers append to the contract `[gated]`. ## STEP 2b — CHALLENGE THE PLAN (adversarial, before the gate) -`Skill(effort-xhigh)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). +`mcp__model-router__route(phase="plan")` first (route: own level before the challenge; send it with the challenger dispatch). Before the human sees the plan, harden it. Run `$HOME/.claude/lib/challenge-plan.md`: - `PLAN` = the plan STEP 2 wrote under `docs/superpowers/plans/` - `KIND` = `build-plan` @@ -174,7 +174,7 @@ judges the diff against this ENRICHED contract, not the STEP 0e seed — so a criterion the design introduced is verified, not lost. ## STEP 4 — IMPLEMENT -First: `Skill(effort-medium)` (effort-shift: dispatch span starts; send it in the same message as this step's first dispatch). +First: `mcp__model-router__route(phase="orchestrate")` (route: dispatch span starts; send it with this step's first dispatch). Start the feature branch off develop, then implement on it: ```bash bash "$HOME/.claude/lib/gitflow.sh" start feature @@ -186,14 +186,15 @@ Invoke `subagent-driven-development` (vendored superpowers skill) for the per-ta finishing-a-development-branch", stop and return. **Model routing (BDR-066):** every subagent dispatched under SDD — per-task -implementers AND its reviewers — MUST carry `model: "sonnet"` in the Agent +implementers AND its reviewers — MUST carry `model: "sonnet"` and +`effort="medium"` (the `implement` level; a main route never reaches a child) in the Agent call. The plan is closed; execution and plan-conformity review are sonnet work. Reflection (task decomposition, review verdict arbitration) stays in this loop. ## STEP 4b — ERROR RECOVERY (if STEP 4 fails) If a subagent returns a build error, failing test, or type error: -1. `Skill(effort-max)` (effort-shift: error recovery; send it in the same message as the Read of the analyzer file below), then load +1. `mcp__model-router__route(phase="escalate")` (route: error recovery; send it with the Read of the analyzer file below), then load `$HOME/.claude/agents/analyzer.md` in DEBUG MODE on the exact error output. Produce: root cause hypotheses (ordered), affected files, what NOT to touch. 2. Present gate: @@ -210,10 +211,10 @@ OPTIONS : C) Abort feature — preserve work done so far ``` 3. Wait for user choice. Do NOT auto-fix. Do NOT proceed without explicit approval. -4. On resume the turn is at the session level (effort-shift: turn reset). - If A → `Skill(effort-medium)` sent with the re-dispatch, apply minimal fix, re-run STEP 4 for the failed task only. Max 2 retry attempts. +4. On resume: `mcp__model-router__route(phase="plan")` first (route: resumed turn). + If A → `mcp__model-router__route(phase="orchestrate")` sent with the re-dispatch, apply minimal fix, re-run STEP 4 for the failed task only. Max 2 retry attempts. If still failing after 2 → fall back to options B or C. - If B or C → `Skill(effort-xhigh)` first, sent with the next tool call. + If B or C → continue at the plan level set above. If B → before skipping: scan remaining task list for tasks that depend on the failed task (look for references to the same file or function in subsequent tasks). If dependents found → present: "Tasks [N, M] depend on the skipped task. @@ -243,9 +244,8 @@ conformity + security vs. craft/design) — both run, neither subsumes the other ([[LRN-095]]). ## STEP 6 — CODE REVIEW -`Skill(effort-xhigh)` first, sent with the review dispatch (effort-shift: judgment dispatch; the reviewer is a built-in and inherits the level in force). Invoke `requesting-code-review` (vendored superpowers skill). **Model routing (BDR-077):** the -review subagent it dispatches MUST carry `model: "opus"` in the Agent call — +review subagent it dispatches MUST carry `model: "opus"` and `effort="xhigh"` in the Agent call — craft review is dispatched judgment, never inherited from the session. Fix all CRITICAL before proceeding. @@ -277,7 +277,7 @@ Feature shipped implies at least one design decision worth capturing. Run this B If nothing substantive to log → print `CAPITALIZE: nothing substantive to log` and skip. -`Skill(effort-low)` first (effort-shift: bookkeeping tail; send it in the same message as the memory-commit command). +`mcp__model-router__route(phase="apply")` first (route: bookkeeping tail; send it with the memory-commit command). **Then commit the memory** — follow `$HOME/.claude/lib/capitalize-commit.md`: it surgically commits what capitalize just wrote (`.claude/memory` + `.claude/tasks` diff --git a/skills/tour/SKILL.md b/skills/tour/SKILL.md index a486a4d..f62fa78 100644 --- a/skills/tour/SKILL.md +++ b/skills/tour/SKILL.md @@ -30,7 +30,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. +ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route). One pipeline per project: **security → clean → re-verify → reconcile → doc → convergence re-audit**, looping until a full pass applies zero new @@ -169,8 +169,8 @@ honestly in the summary. Never loop past 3. ### Phase B — CLEAN -1. Dispatch a read-only cleanup audit (analyzer — opus-pinned, BDR-076 — - or general-purpose with `model="opus"`; NOT the sonnet code-cleaner, +1. Dispatch a read-only cleanup audit (analyzer — opus row, BDR-076 — + or general-purpose with `model="opus"`, `effort="xhigh"`; NOT the sonnet code-cleaner, which is now a fix executor): dead code, unused imports/exports, commented-out blocks, stale flags, norm violations. Findings as `id | file:line | finding | proposed fix`. diff --git a/skills/web-validate/SKILL.md b/skills/web-validate/SKILL.md index 0901830..8b6405b 100644 --- a/skills/web-validate/SKILL.md +++ b/skills/web-validate/SKILL.md @@ -28,7 +28,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. +ROUTING: follow $HOME/.claude/lib/effort-shift.md (phases via mcp__model-router__route). This skill orchestrates a narrow-scope standards audit : @@ -180,7 +180,7 @@ Spawn a single `validator-analyzer` subagent with explicit scope and collected context : ``` -Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message +mcp__model-router__route(phase="orchestrate") # route: dispatch span starts; send with the Agent call below in ONE message Agent( subagent_type="validator-analyzer", description="validate — W3C HTML + CSS + WCAG audit", @@ -255,7 +255,7 @@ grep -c '^### \[Critique\]' .claude/audits/VALIDATE.md --- ## STEP 2b — CHALLENGE THE FIX BUNDLE (MODE=fix only, advisory) -`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). +`mcp__model-router__route(phase="reflect")` first (route: own level before the challenge; send it with the challenger dispatch). Skip if MODE=audit (no bundle exists). Else, before the STEP 3 gate, harden the bundle: extract the `## 5. Fix bundle` section from VALIDATE.md to `.claude/tasks/plans/--.md` (a clean, blind-judgeable artifact), then run @@ -313,7 +313,7 @@ Options : share files: ``` - Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message + mcp__model-router__route(phase="orchestrate") # route: dispatch span starts; send with the Agent call below in ONE message Agent(subagent_type="hotfixer") prompt: ". diff --git a/update-all.sh b/update-all.sh index f83d907..ba76c77 100644 --- a/update-all.sh +++ b/update-all.sh @@ -468,8 +468,7 @@ fi # ── 7.3b. Update the Higgsfield CLI + skill pack ── # CLI: global npm bin. Skills: re-cloned by lib/higgsfield-skills.sh, which # replaces the SOURCE under skills-external/ only: a pack parked in -# skills-disabled/ (symlinks to those sources) stays parked. Runs before the -# effort-pins re-apply below (BDR-108). +# skills-disabled/ (symlinks to those sources) stays parked. echo "" echo "── Updating Higgsfield CLI + skill pack..." if ! command -v higgsfield &>/dev/null; then @@ -562,15 +561,6 @@ print(d.get('21st',{}).get('version','latest')) rm -rf "$TFD_STAGE" fi -# Effort pins (BDR-107, BDR-108): every refresh above rewrites SKILL.md and -# drops the `effort:` line; the 21st pack refresh is the last step that rewrites -# a SKILL.md, so the entry levels of lib/effort-pins.txt go back here. -echo "" -echo "── Re-applying effort pins on the vendored skills..." -# shellcheck source=lib/effort-pins.sh disable=SC1091 -source "$REPO/lib/effort-pins.sh" -apply_effort_pins "$REPO" || warn "effort pins: map lines rejected — fix lib/effort-pins.txt" - # ── 7.5. Update external skills (npx skills) ── echo "" echo "── Updating external skills (npx skills)..."