feat(model-router): wave 2-B — orchestrators declare phases, shifters and pins removed, frontmatter = off-state floor

The 15 Skill(effort-*) citers now call mcp__model-router__route per phase
(orchestrate at a dispatch span, reflect/plan for the skill's own level,
apply at the bookkeeping tail, escalate at the verify-secure caps); built-in
judgment dispatches carry an explicit effort= param. lib/effort-shift.md is
the route doctrine, lib/model-gate.md the mod rule (route answer = witness,
/route on as remedy). Deleted: skills/effort-*, lib/effort-pins.txt/.sh,
lib/model-check.sh, their tests, the installers' re-apply blocks. The mod
drops its Skill(effort-*) bridge. The tracked model:/effort: frontmatter
stays as the off-state floor, census-locked equal to the rows
(lib/tests/effort-routing.test.sh rewritten, 140 checks; analyzer → xhigh).

Contract .claude/tasks/contracts/2026-10-10-model-router-w2b-1045.md, plan
r4 § W2-B: GATE 0 MET, verifier ECARTS(7) then CONFORME 10/10, security
PASS, full make test green (design-tool-gate env red only).
This commit is contained in:
bchanot
2026-10-10 11:24:54 +02:00
parent 65dff0e768
commit 1f2d33b7a6
43 changed files with 318 additions and 811 deletions
+6 -4
View File
@@ -43,8 +43,9 @@ Agent(subagent_type="plan-challenger", description="challenge:<lens>", prompt=""
```
**MODEL (BDR-076, supersedes the BDR-066 inherit):** plan critique is AUDIT
JUDGMENT — the challengers are `model: opus`-pinned in their frontmatter: a big
tier, session-independent, off the session model. The session model (Fable)
JUDGMENT — the challengers are routed to the `judge` row (opus) by the model-router,
the `model: opus` frontmatter being the off-state floor: a big tier, off
the session model. The session model (Fable)
keeps only this loop — synthesis, RE-THINK, gate. Never sonnet: that would
silently downgrade the judgment. (The executor gates stay sonnet.)
@@ -59,8 +60,9 @@ silently downgrade the judgment. (The executor gates stay sonnet.)
A challenger that returns a malformed/empty verdict, a missing `PROOF`, or dies →
retry ONCE with a fresh challenger; a 2nd failure on that lens → STOP and escalate
(the STOP text names the level reached, `$CLAUDE_EFFORT`, and suggests `/effort-max`
for the relaunch; no shift here: a mute challenger is an infrastructure failure)
(the STOP text names the level reached, `$CLAUDE_EFFORT`, and suggests relaunching with
`ultrathink` in the prompt (turn floor) or `/route effort=max` (sticky,
`/route clear` after); no shift here: a mute challenger is an infrastructure failure)
to the human, NAMING the lens. Never carry "plan challenged" into the gate on a
silently dropped lens (`verify-secure-loop.md`: "a mute verifier is NEVER a PASS").
-120
View File
@@ -1,120 +0,0 @@
#!/usr/bin/env bash
# lib/effort-pins.sh — re-apply the entry effort level on vendored skills
# (BDR-107 second axis, extended to every vendored external by BDR-108).
# Upstream copies carry no `effort:` and every vendoring step rewrites
# SKILL.md, so the level lives in lib/effort-pins.txt and this helper puts
# it back after the last vendoring step of install-plugins.sh and
# update-all.sh. Idempotent: same level → untouched, other level →
# replaced inside the frontmatter only, skill not vendored → skipped,
# malformed map line → rejected loudly, never applied. Hardenings: a map
# whose last line lacks a newline is still read; a SKILL.md whose frontmatter
# never closes is skipped untouched; the level is re-read after every write
# and a mismatch counts as failed; the write goes through a mktemp sibling
# removed on any failure and on INT/TERM (previous traps restored, never an
# EXIT trap: the installer owns one); the rejected map line is printed
# shell-quoted so a caller's `echo -e` cannot interpret it. Placement inside
# the frontmatter has no effect on the harness, which reads the key anywhere.
#
# Usage: source it, then `apply_effort_pins [repo-root]`
# or standalone: bash lib/effort-pins.sh [repo-root]
# Exit 1 when at least one map line was rejected or a skill failed.
EFFORT_PINS_REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
EFFORT_PIN_LEVEL_RE='^(low|medium|high|xhigh|max)$'
EFFORT_PIN_NAME_RE='^[A-Za-z0-9][A-Za-z0-9._-]*$'
# Callers (install-plugins.sh, update-all.sh) define these; standalone
# runs get plain fallbacks.
declare -F ok >/dev/null || ok() { printf ' ok %s\n' "$*"; }
declare -F info >/dev/null || info() { printf ' info %s\n' "$*"; }
declare -F err >/dev/null || err() { printf ' ERR %s\n' "$*" >&2; }
# _effort_pin_current <skill-file> → prints the frontmatter effort, if any
_effort_pin_current() {
awk 'NR==1&&/^---$/{p=1;next} p&&/^---$/{exit} p' "$1" \
| sed -n 's/^effort: //p' | head -1
}
# _effort_pin_closed <skill-file> → rc 0 when the frontmatter has a closing ---
_effort_pin_closed() {
awk 'NR==1&&/^---$/{p=1;next} p&&/^---$/{f=1;exit} END{exit !f}' "$1"
}
# _effort_pin_traps_restore <saved> — drop the INT/TERM handlers set for the
# write and re-install the caller's saved ones. No exit-time handler here.
_effort_pin_traps_restore() {
trap - INT TERM
[ -z "$1" ] || eval "$1"
}
# _effort_pin_write <skill-file> <name> <level> — replace the frontmatter
# `effort:` line, or insert one after `name: <name>` (before the closing
# `---` when the frontmatter has no name line). Body lines never change.
# Writes a mktemp sibling then renames; any failure leaves no temp behind.
_effort_pin_write() {
local file="$1" name="$2" level="$3" tmp prev rc
tmp="$(mktemp "$file.XXXXXX")" || return 1
prev="$(trap -p INT TERM)"
trap 'rm -f "$tmp"; exit 130' INT TERM
cp -p "$file" "$tmp" && awk -v n="$name" -v lvl="$level" '
NR==1 && /^---$/ { fm=1; print; next }
fm && /^---$/ {
if (!done) { print "effort: " lvl; done=1 }
fm=0; print; next
}
fm && /^effort: / { if (!done) { print "effort: " lvl; done=1 }; next }
fm && $0 == "name: " n { print; if (!done) { print "effort: " lvl; done=1 }; next }
{ print }
' "$file" > "$tmp" && mv "$tmp" "$file"; rc=$?
[ "$rc" -eq 0 ] || rm -f "$tmp"
_effort_pin_traps_restore "$prev"
return "$rc"
}
# _effort_pin_apply_one <file> <name> <level> → rc 0 applied, 2 already at
# level, 1 failed (err line printed, file untouched or write rolled back)
_effort_pin_apply_one() {
local file="$1" name="$2" level="$3"
if ! _effort_pin_closed "$file"; then
err "effort-pins: $file: frontmatter never closed — skipped"; return 1
fi
[ "$(_effort_pin_current "$file")" = "$level" ] && return 2
if ! _effort_pin_write "$file" "$name" "$level"; then
err "effort-pins: $file: write failed"; return 1
fi
if [ "$(_effort_pin_current "$file")" != "$level" ]; then
err "effort-pins: $file: level not applied after write"
return 1
fi
return 0
}
# apply_effort_pins [repo-root] — walk the map, pin every vendored skill
apply_effort_pins() {
local repo="${1:-$EFFORT_PINS_REPO}" map name level rest file rc
local applied=0 kept=0 rejected=0 failed=0
map="$repo/lib/effort-pins.txt"
[ -f "$map" ] || { err "effort-pins: map missing: $map"; return 1; }
while read -r name level rest || [ -n "$name" ]; do
case "$name" in ''|'#'*) continue ;; esac
if [ -n "$rest" ] || ! [[ "$name" =~ $EFFORT_PIN_NAME_RE ]] \
|| ! [[ "$level" =~ $EFFORT_PIN_LEVEL_RE ]]; then
err "effort-pins: rejected map line $(printf '%q' "$name $level $rest")"
rejected=$((rejected + 1)); continue
fi
file="$repo/skills-external/$name/SKILL.md"
[ -f "$file" ] || continue
_effort_pin_apply_one "$file" "$name" "$level"; rc=$?
case "$rc" in
0) applied=$((applied + 1)) ;;
2) kept=$((kept + 1)) ;;
*) failed=$((failed + 1)) ;;
esac
done < "$map"
ok "effort-pins: $applied applied, $kept already at level, $failed failed"
[ "$rejected" -eq 0 ] && [ "$failed" -eq 0 ]
}
if [[ "${BASH_SOURCE[0]}" == "$0" ]]; then
apply_effort_pins "$@"
fi
-46
View File
@@ -1,46 +0,0 @@
# lib/effort-pins.txt — entry effort level of the vendored skills
# (skills-external/<name>/SKILL.md). Upstream copies carry no `effort:` and
# every resync rewrites SKILL.md, so the pin lives here and
# lib/effort-pins.sh re-applies it after the last vendoring step of
# install-plugins.sh and update-all.sh. One line = `<skill> <level>`,
# level in low|medium|high|xhigh|max. The census
# lib/tests/effort-routing.test.sh checks every vendored file against this
# map. Rungs (BDR-107, BDR-108): low = fix a line, run a script · medium =
# day-to-day · high = refactor, resisting bug · xhigh = architecture, audit
# before validation · max = stuck.
#
# superpowers (obra/superpowers, plugins.lock.json "superpowers")
brainstorming xhigh
writing-plans xhigh
requesting-code-review xhigh
subagent-driven-development high
writing-skills high
test-driven-development medium
using-git-worktrees low
#
# agent-skills (addyosmani/agent-skills, plugins.lock.json "agent-skills")
deprecation-and-migration high
ci-cd-and-automation medium
observability-and-instrumentation medium
#
# design stack — ONE level for every member: these skills load stacked in a
# single UI build and the last loaded wins (lib/effort-shift.md), so two
# levels in the stack would make the effort depend on load order.
# skills/site-motion (repo-authored) pins the same level in its frontmatter.
frontend-design high
emil-design-eng high
design-motion-principles high
21st-ui-build high
scroll-world-storytelling high
build-threejs-scroll-worlds high
scroll-scrubbed-visual-sequence high
scroll-scrubbed-word-reveal high
scroll-progress-timeline high
#
# 21st pack (`21st skills install`): tooling low, generation high, critique xhigh
21st-cli-use low
21st-registry low
21st-design-sync low
21st-ai high
21st-ui-explore high
21st-ui-review xhigh
+47 -74
View File
@@ -1,85 +1,58 @@
# Effort shift — phase-level reasoning effort on the main loop (BDR-107)
# Route doctrine — phase-level model and effort on the main loop (BDR-107)
Shared include, companion of `lib/model-gate.md`: the gate fixes WHICH model
reflects, this include fixes HOW HARD each phase thinks. The rungs are the
user's: low (fix a line, run a script) · medium (day-to-day) · high
(refactor, resisting bug) · xhigh (architecture, audit before validation) ·
max (stuck error, judged need).
reflects, the model-router mod (`mods/model-router`) fixes HOW HARD each
phase thinks. Rungs: low (fix a line, run a script) · medium (day-to-day) ·
high (refactor, resisting bug) · xhigh (architecture, audit before
validation) · max (stuck error, judged need).
## Mechanics (verified on Claude Code 2.1.283)
## The tool
- **Pairing rule**: a `Skill(effort-<level>)` call applies its effort only
when the same assistant message carries at least one other tool call
after it; a lone Skill call is a no-op. Send the shift together with the
step's first tool call, shift first. That paired call already runs at the
new level: pair a downward shift with a pinned-agent dispatch or a
Read/Bash, never with a built-in judgment dispatch (`general-purpose`,
`model: "opus"`), which would inherit it.
- Re-loading a shifter already loaded in the conversation re-applies its
effort (the harness only dedupes the skill text), so bounce-back
sequences such as medium → max → medium work.
- A skill's `effort:` frontmatter applies from the moment it loads to the
end of the turn: on the user's `/skill` unconditionally, and on a
`Skill(...)` call by Claude only under the pairing rule above (a skill
Claude loads alone, such as `brainstorming` or `writing-plans`, applies
nothing). Last loaded wins, both directions. The prompt cache survives a
shift.
- **Stacked skills share one level**: skills that load together in one
build (the design stack) all pin the same level, since the last loaded
wins. Vendored externals get their level from `lib/effort-pins.txt`,
re-applied by `lib/effort-pins.sh` after every vendoring step; repo
skills carry it in their frontmatter.
- Dispatched agents run on their own `effort:` pin, never on a shift.
Unpinned agents inherit the level in force at dispatch.
- Headless sessions (`-p`, `claude agents`, SDK) ignore skill-level effort:
the run stays at the session level. `CLAUDE_CODE_EFFORT_LEVEL` beats every
frontmatter; keep it unset (the session banner warns).
`mcp__model-router__route` (params `phase` | `effort` | `clear`). It is a
deferred tool: when not loaded, run
`ToolSearch("select:mcp__model-router__route")` once per session. A route
applies from the next request on, paired with
another tool call or not (pairing only saves a request). The answer always
names the id and effort main runs on. A skill with a row routes itself on
load; a skill without one changes nothing, the last ROWED skill wins.
## Wiring points
1. Dispatch span starts → `route(phase="orchestrate")`, sent with the
dispatch.
2. Reflection resumes (challenge synthesis, verdict, plan revision) →
`route(phase="reflect")` or `"plan"` per the skill's own level; the line
before every `lib/challenge-plan.md` call.
3. Bookkeeping tail (memory commit, doc commit) → `route(phase="apply")`.
4. Escalation → `route(phase="escalate")`: verify-secure loop caps and
ship-feature STEP 4b. Not automatic: the challenge fail-safe and "gone
WRONG → STOP"; their STOP text names the levers below.
5. Built-in judgment dispatch (`general-purpose` `model="opus"`, `model:
"fable"` skill-runners) → explicit `effort=` on the Agent call (`xhigh`
for opus reviewers, `high` for fable runners). A main route never
reaches a child. Typed agents run on their row, never on a shift.
6. After a prose gate that ends the turn, the resumed reflection phase
starts with its own route call.
## Run slot and levers
A best-tier skill row survives the end of the turn (a run spans prose
gates); `/route clear`, `/route off` and a user `/model` drop it. Levers for a
relaunch: `ultrathink` in the prompt (turn floor) or `/route effort=max`
(sticky, `/route clear` after).
Builtin `/effort` is NOT a lever inside a run: rows and routes outrank it.
## Limits
- A skill typed while a background agent is live routes only through the
typed marker (unverified live 2026-10-10).
- Headless (`-p`, SDK) runs the hooks, so routing works there too.
- Mod off: typed agents fall back to their `model:`/`effort:` frontmatter.
Measure the split any time: `python3 ~/.claude/lib/effort-audit.py`
(thinking/output/cache tokens per scope, model and effort).
## Shifters
`Skill(effort-low)` · `Skill(effort-medium)` · `Skill(effort-high)` ·
`Skill(effort-xhigh)` · `Skill(effort-max)`. One tool call, one-line body,
always sent with another tool call (Pairing rule).
Typed by the user, `/effort-max` is a turn-scoped max: the relaunch lever
after a STOP. `ultrathink` only adds an in-context nudge; the API level
does not move.
## Wiring — per orchestrator
1. A dispatch span starts (executor, collector, fan-out) →
`Skill(effort-medium)`.
2. Reflection resumes after a dispatch span (challenge synthesis, verdict,
plan revision) → `Skill(effort-<the skill's own level>)`. Concretely:
the line before every `lib/challenge-plan.md` call.
3. The bookkeeping tail (memory commit, doc commit) → `Skill(effort-low)`.
4. Escalation → `Skill(effort-max)`, then the skill's own level again once
the diagnosis is produced. Automatic points: verify-secure loop caps
(GATE 0 floor, GATE 1 conformity, GATE 2 security) and ship-feature
STEP 4b. Not automatic, by doctrine: the challenge fail-safe (a mute
challenger is an infrastructure failure) and "gone WRONG → STOP" (STOP
precedes any further reasoning); their STOP text names the level
reached and suggests `/effort-max` for the relaunch.
5. Before any built-in or unpinned dispatch that carries judgment (a
`general-purpose` with `model: "opus"` or `"fable"`, the code reviewer
of requesting-code-review, a skill-runner) → `Skill(effort-<own level>)`
paired with that dispatch: built-ins inherit the level in force, and a
medium set earlier in the span would downgrade them.
## Re-assert
- After any nested `Skill(...)` whose frontmatter carries a different
effort (feat → commit-change), reload the orchestrator's own level.
- After a prose gate that ends the turn, the resumed turn runs at the
session level. If the resumed phase is reflection, its first step is
`Skill(effort-<own level>)`; dispatch and orchestration phases need
nothing.
## Never
- A shift never inside a dispatched agent: pins rule there.
- A route inside a dispatched agent: its row rules there.
- Max is for diagnosis, not for retrying the same fix harder.
- A medium shift never precedes a judgment dispatch in the same span
without an own-level shift paired with that dispatch.
-33
View File
@@ -1,33 +0,0 @@
#!/usr/bin/env bash
# lib/model-check.sh — classify the persisted session model: big | small | unknown
#
# Witness for lib/model-gate.md (reflection requires a big model). Reads the
# "model" key of the user-scope settings (the file /model rewrites — LRN-098).
# Override the source with MODEL_CHECK_SETTINGS (tests use fixtures).
#
# stdout : <class>:<raw> (raw = value found, empty if none)
# exit : 0 = big (fable/opus) · 2 = small (sonnet/haiku) · 3 = unknown
set -u
SETTINGS="${MODEL_CHECK_SETTINGS:-$HOME/.claude/settings.json}"
raw=""
if [ -f "$SETTINGS" ]; then
raw="$(python3 - "$SETTINGS" 2>/dev/null <<'PY'
import json, sys
try:
v = json.load(open(sys.argv[1])).get("model", "")
print(v if isinstance(v, str) else "")
except Exception:
print("")
PY
)"
fi
norm="$(printf '%s' "$raw" | tr '[:upper:]' '[:lower:]')"
case "$norm" in
*opusplan*) printf 'unknown:%s\n' "$raw"; exit 3 ;; # opus-for-plan, sonnet otherwise — ambiguous
*fable*|*opus*) printf 'big:%s\n' "$raw"; exit 0 ;;
*sonnet*|*haiku*) printf 'small:%s\n' "$raw"; exit 2 ;;
*) printf 'unknown:%s\n' "$raw"; exit 3 ;;
esac
+23 -46
View File
@@ -1,53 +1,30 @@
# Model gate — reflection requires a big model (BLOCKING)
Shared include. Runs FIRST in any orchestrator whose reflection —
brainstorming, planning, contract, audit judgment, loop decisions —
executes inline or in inherit-model subagents. Sonnet-pinned executors are
not what this gate protects; it protects the thinking around them (BDR-066).
Shared include, runs FIRST in an orchestrator whose reflection executes
inline (BDR-066). The witness is the model-router mod's own route tool.
## 1. Self-check
## Entry call
ALWAYS call `mcp__model-router__route` with `phase` = the skill's row phase
(reflect or plan), no self-check shortcut. Tool not loaded (deferred) →
`ToolSearch("select:mcp__model-router__route")` once per session, then
call. The answer names the id main runs on next.
Your system prompt names the model powering this session. Fable or Opus →
big. Sonnet, Haiku, anything else → small.
| answer | action |
|---|---|
| names a fable or opus id | proceed, SILENT |
| names sonnet, haiku, anything else; "is off"; tool absent | **STOP** |
## 2. Witness — deterministic check
**STOP means**: print exactly `⛔ MODEL GATE — session on <model>.
Reflection steps of this skill require Fable or Opus. Switch with /model,
then relaunch the skill.` (mod off: say so, `/route on` resumes it), then
end
the turn. No later step runs, no agent is dispatched, nothing is edited.
bash "$HOME/.claude/lib/model-check.sh"
Output `<class>:<raw>`; exit 0 = big, 2 = small, 3 = unknown. The witness
reads the PERSISTED model (settings.json — the file `/model` rewrites,
LRN-098). It can lag reality (session launched with `--model`, settings not
yet rewritten) — that is why the self-check exists alongside it.
## 3. Verdict
| self-check | witness | action |
|---|---|---|
| big | big (0) | proceed, SILENT — the nominal path prints nothing |
| small | any | **STOP** |
| big | small (2) | disagreement — **STOP**, surface BOTH values; the user confirms or relaunches |
| big | unknown (3) | fail-visible: print `model gate: witness unknown (<raw>) — self-check says <model>` and ask the user to confirm before continuing (BDR-025: unknown never silently passes) |
**STOP means**: print exactly
⛔ MODEL GATE — session on <model>. Reflection steps of this skill
require Fable or Opus. Switch with /model, then relaunch the skill.
then end the turn. No later step runs, no agent is dispatched, nothing is
edited.
## 4. Dispatch tiers (BDR-077 — no inherit)
The gate guards the MAIN loop only. Dispatched work NEVER inherits the
session model: typed agents run on their frontmatter pin; built-ins
## Dispatch tiers (BDR-077 — no inherit)
The gate guards the MAIN loop only. Typed agents are routed by their
model-router row (`model:` frontmatter = off-state floor). Built-ins
(general-purpose / Explore / Plan) carry an explicit `model=` at every call
site — `model: "fable"` when the child performs reflection/orchestration on
the main loop's behalf (skill-runners), otherwise its complexity tier
(opus = dispatched judgment, sonnet = execution/collection, haiku = short
mechanical probes).
Effort is the second axis of the same table (BDR-107): every typed agent
carries an `effort:` pin next to `model:`, and the main loop shifts per phase
through `lib/effort-shift.md`. No typed agent inherits either axis;
built-ins inherit the effort in force at dispatch, so an orchestrator shifts
before dispatching them (`lib/effort-shift.md`, wiring point 5).
site: `model: "fable"` when the child reflects/orchestrates for the main
loop (skill-runners), else its tier (opus = dispatched judgment, sonnet =
execution, haiku = mechanical probes). A built-in judgment dispatch also
carries an explicit `effort=` (`lib/effort-shift.md`).
-129
View File
@@ -1,129 +0,0 @@
#!/usr/bin/env bash
# lib/tests/effort-pins.test.sh — lib/effort-pins.sh's apply_effort_pins():
# insert after `name:`, keep an equal level untouched, replace a different
# level inside the frontmatter only (a prose `effort:` in the body stays),
# skip a skill not vendored, insert before the closing `---` when the
# frontmatter has no name line, run idempotently, reject a bad level, a
# traversal name and a three-field line before writing anything, and
# parse the real map without error; hardening: last map line without a
# newline, unterminated frontmatter, CRLF file and read-only directory. All on a throwaway fixture repo.
set -u
ROOT="$(cd "$(dirname "$0")/../.." && pwd)"
LIB="$ROOT/lib/effort-pins.sh"
pass=0; fail=0
check() { if [ "$2" = "$3" ]; then pass=$((pass+1)); echo "PASS $1"
else fail=$((fail+1)); echo "FAIL $1: got[$2] want[$3]"; fi; }
fm_effort() { awk 'NR==1&&/^---$/{p=1;next} p&&/^---$/{exit} p' "$1" \
| sed -n 's/^effort: //p' | head -1; }
WORK="$(mktemp -d)" || exit 1; trap 'rm -rf "$WORK"' EXIT
REPO="$WORK/repo"; EXT="$REPO/skills-external"
mkdir -p "$REPO/lib" "$EXT/alpha" "$EXT/beta" "$EXT/gamma" "$EXT/noname"
printf -- '---\nname: alpha\ndescription: a\n---\nbody\n' > "$EXT/alpha/SKILL.md"
printf -- '---\nname: beta\neffort: low\n---\nprose says effort: max here\n' > "$EXT/beta/SKILL.md"
printf -- '---\nname: gamma\neffort: low\n---\nbody\n' > "$EXT/gamma/SKILL.md"
printf -- '---\ndescription: no name line\n---\nbody\n' > "$EXT/noname/SKILL.md"
printf '# map\nalpha high\nbeta medium\ngamma low\nghost xhigh\nnoname low\n' > "$REPO/lib/effort-pins.txt"
gamma_before="$(cat "$EXT/gamma/SKILL.md")"
bash "$LIB" "$REPO" >/dev/null 2>&1; check T1-rc-clean "$?" 0
check T2-insert-after-name "$(sed -n '3p' "$EXT/alpha/SKILL.md")" "effort: high"
check T3-replace-in-frontmatter "$(fm_effort "$EXT/beta/SKILL.md")" "medium"
check T3b-body-prose-untouched "$(grep -c 'effort: max' "$EXT/beta/SKILL.md")" 1
check T3c-single-effort-line "$(grep -c '^effort:' "$EXT/beta/SKILL.md")" 1
check T4-equal-level-untouched "$(cat "$EXT/gamma/SKILL.md")" "$gamma_before"
check T5-missing-skill-skipped "$([ -e "$EXT/ghost" ] && echo created || echo absent)" absent
check T6-no-name-inserts-before-closing "$(sed -n '3p' "$EXT/noname/SKILL.md")" "effort: low"
check T6b-no-name-still-frontmatter "$(fm_effort "$EXT/noname/SKILL.md")" "low"
snap="$(cat "$EXT"/*/SKILL.md)"
bash "$LIB" "$REPO" >/dev/null 2>&1
check T7-idempotent "$(cat "$EXT"/*/SKILL.md)" "$snap"
check T7b-no-tmp-left "$(find "$EXT" -name '*.tmp' | wc -l | tr -d ' ')" 0
# rejections: nothing written, rc 1
for bad in 'alpha turbo' '../evil high' 'alpha high extra'; do
printf '%s\n' "$bad" > "$REPO/lib/effort-pins.txt"
out="$(bash "$LIB" "$REPO" 2>&1)"; rc=$?
check "T8-rejected[$bad]-rc" "$rc" 1
check "T8-rejected[$bad]-named" "$(printf '%s' "$out" | grep -c 'rejected map line')" 1
done
check T8b-tree-unchanged-after-rejections "$(cat "$EXT"/*/SKILL.md)" "$snap"
check T8c-no-evil-dir "$([ -e "$WORK/evil" ] && echo created || echo absent)" absent
# the real map parses: fixture repo with the real map and no vendored skill
mkdir -p "$WORK/real/lib" "$WORK/real/skills-external"
cp "$ROOT/lib/effort-pins.txt" "$WORK/real/lib/"
out="$(bash "$LIB" "$WORK/real" 2>&1)"; check T9-real-map-parses "$?" 0
check T9b-real-map-nothing-applied "$(printf '%s' "$out" | grep -c '0 applied, 0 already')" 1
check T10-missing-map-rc "$(bash "$LIB" "$WORK/nowhere" >/dev/null 2>&1; echo $?)" 1
# hardening: each case in its own fixture repo
mkrepo() { R="$WORK/$1"; mkdir -p "$R/lib" "$R/skills-external/$2"; }
mkrepo h11 alpha; mkdir "$WORK/h11/skills-external/beta"
printf -- '---\nname: alpha\n---\nb\n' > "$WORK/h11/skills-external/alpha/SKILL.md"
printf -- '---\nname: beta\n---\nb\n' > "$WORK/h11/skills-external/beta/SKILL.md"
printf 'alpha high\nbeta low' > "$WORK/h11/lib/effort-pins.txt"
bash "$LIB" "$WORK/h11" >/dev/null 2>&1
check T11-last-line-no-newline "$(fm_effort "$WORK/h11/skills-external/beta/SKILL.md")" low
mkrepo h12 open; f12="$WORK/h12/skills-external/open/SKILL.md"
printf -- '---\nname: open\nbody effort: max\n' > "$f12"; b12="$(cat "$f12")"
printf 'open high\n' > "$WORK/h12/lib/effort-pins.txt"
out="$(bash "$LIB" "$WORK/h12" 2>&1)"; rc=$?
check T12-unterminated-frontmatter-skipped \
"$rc|$(cat "$f12" | cmp -s - <(printf '%s\n' "$b12") && echo same)|$(printf '%s' "$out" | grep -c "ERR .*$f12")" "1|same|1"
mkrepo h13 crlf
printf -- '---\r\nname: crlf\r\n---\r\nbody\r\n' > "$WORK/h13/skills-external/crlf/SKILL.md"
printf 'crlf high\n' > "$WORK/h13/lib/effort-pins.txt"
out="$(bash "$LIB" "$WORK/h13" 2>&1)"; rc=$?
check T13-crlf-file-rejected \
"$rc|$(printf '%s' "$out" | grep -c 'ERR ')|$(printf '%s' "$out" | grep -c ' 0 applied, ')" "1|1|1"
# T13b: the post-write re-read branch, reached with a no-op write stub
mkrepo h13b nowrite; f13b="$WORK/h13b/skills-external/nowrite/SKILL.md"
printf -- '---\nname: nowrite\n---\nb\n' > "$f13b"
out="$(bash -c 'source "$1"; _effort_pin_write() { return 0; }
_effort_pin_apply_one "$2" nowrite high' _ "$LIB" "$f13b" 2>&1)"; rc=$?
check T13b-reread-mismatch-fails \
"$rc|$(printf '%s' "$out" | grep -c 'level not applied after write')" "1|1"
if [ "${EFFORT_PINS_TEST_FAKE_ROOT:-0}" = 1 ] || [ "$(id -u)" -eq 0 ]; then
echo "SKIP T14-write-failure-no-temp: chmod bits ignored as root"
else
mkrepo h14 ro; d14="$WORK/h14/skills-external/ro"
printf -- '---\nname: ro\n---\nb\n' > "$d14/SKILL.md"
printf 'ro high\n' > "$WORK/h14/lib/effort-pins.txt"
chmod 555 "$d14"; out="$(bash "$LIB" "$WORK/h14" 2>&1)"; rc=$?; chmod 755 "$d14"
check T14-write-failure-no-temp \
"$rc|$(printf '%s' "$out" | grep -c 'ERR ')|$(find "$d14" -name 'SKILL.md.*' | wc -l | tr -d ' ')" "1|1|0"
fi
# T15: SIGINT during the awk write removes the temp sibling, exit 130
mkrepo h15 sig; d15="$WORK/h15/skills-external/sig"
printf -- '---\nname: sig\n---\nb\n' > "$d15/SKILL.md"
bash -c 'source "$1"; awk() { kill -INT $$; sleep 2; }
_effort_pin_write "$2" sig high' _ "$LIB" "$d15/SKILL.md" >/dev/null 2>&1
rc=$?
check T15-sigint-removes-temp \
"$rc|$(find "$d15" -name 'SKILL.md.*' | wc -l | tr -d ' ')" "130|0"
# T15b: previous INT trap restored on a normal return, no EXIT trap set
mkrepo h15b tr; d15b="$WORK/h15b/skills-external/tr"
printf -- '---\nname: tr\n---\nb\n' > "$d15b/SKILL.md"
out="$(bash -c 'source "$1"; trap "echo prev" INT
_effort_pin_write "$2" tr high
printf "INT:%s\n" "$(trap -p INT)"; printf "EXIT:%s\n" "$(trap -p EXIT)"' \
_ "$LIB" "$d15b/SKILL.md" 2>&1)"
check T15b-traps-restored \
"$(printf '%s' "$out" | grep -c "^INT:trap -- 'echo prev' SIGINT")|$(printf '%s' "$out" | grep -c '^EXIT:$')" "1|1"
# T16: a literal backslash-t in a map line is printed shell-quoted
mkrepo h16 q
printf 'bad\\tname high\n' > "$WORK/h16/lib/effort-pins.txt"
out="$(bash "$LIB" "$WORK/h16" 2>&1)"; rc=$?
check T16-rejected-line-quoted \
"$rc|$(printf '%s' "$out" | grep -cF 'bad\\tname')" "1|1"
echo "effort-pins: $pass pass, $fail fail"
[ "$fail" -eq 0 ]
+128 -112
View File
@@ -1,9 +1,12 @@
#!/usr/bin/env bash
# lib/tests/effort-routing.test.sh — census: effort tiering (BDR-107)
# agent pins, skill entry levels, shifter skills, orchestrator wiring, settings.
# shellcheck disable=SC2015,SC2016 # A && ok || ko is deliberate (ok/ko never fail); '$REPO' locks are literal source text
# lib/tests/effort-routing.test.sh — wave-2 census of the model-router rows.
# Drift lock: every tracked skill/agent row in mods/model-router/hooks/
# register.ts equals its frontmatter (the off-state floor), the D3 wiring
# markers sit in the orchestrators, no shifter citer survives.
# shellcheck disable=SC2015,SC2016 # A && ok || ko is deliberate (ok/ko never fail)
set -u
R="$(cd "$(dirname "$0")/../.." && pwd)"
REG="$R/mods/model-router/hooks/register.ts"
pass=0; fail=0
ok() { pass=$((pass+1)); }
ko() { fail=$((fail+1)); printf 'FAIL %s\n' "$1"; }
@@ -11,120 +14,133 @@ has() { if grep -qF "$2" "$R/$1"; then ok; else ko "$1 missing: $2"; fi; }
lacks() { if grep -qF "$2" "$R/$1"; then ko "$1 must NOT contain: $2"; else ok; fi; }
# frontmatter = the lines between the first two '---' lines
fm() { awk 'NR==1&&/^---$/{p=1;next} p&&/^---$/{exit} p' "$1"; }
fm_effort() { fm "$1" | grep -E '^effort: (low|medium|high|xhigh|max)$' | head -1 | cut -d' ' -f2; }
fm_has_effort() {
got="$(fm_effort "$R/$1")"
if [ "$got" = "$2" ]; then ok; else ko "$1 frontmatter effort must be '$2', got '${got:-none}'"; fi
fm_val() { fm "$1" | grep -E "^$2: [a-z]+$" | head -1 | cut -d' ' -f2; }
# ── register.ts parsers (awk/sed on the DEFAULT_CONFIG literal) ──────────
# rows <file> <agents|skills> -> "name phase" per row
rows() {
awk -v s="$2" '$0 ~ "^ "s": \\{"{f=1;next} f&&/^ \},?$/{f=0} f' "$1" \
| grep -v '^ *//' | grep -oE "('[^']+'|[A-Za-z0-9_-]+): '[a-z]+'" \
| sed -E "s/'//g; s/: / /"
}
fm_no_effort() { if fm "$R/$1" | grep -q '^effort:'; then ko "$1 must NOT pin effort"; else ok; fi; }
# phase_effort <file> <phase> -> "<tier> <effort>"
phase_effort() {
awk '/^ phases: \{/{f=1;next} f&&/^ \},?$/{f=0} f' "$1" \
| sed -nE "s/^ *$2: \{ tier: '([a-z]+)', effort: '([a-z]+)' \},?$/\1 \2/p"
}
# tier_head <file> <tier> -> first alias of the tier list
tier_head() {
awk '/^ tiers: \{/{f=1;next} f&&/^ \},?$/{f=0} f' "$1" \
| sed -nE "s/^ *$2: \['([a-z]+)'.*$/\1/p"
}
row_of() { rows "$REG" "$1" | awk -v n="$2" '$1==n{print $2}'; }
# ── flip-test: the frontmatter reader must accept a valid level and reject an invalid one
# ── flip-test: the parsers read a fixture, reject a missing key ──────────
FIX="$(mktemp -d)"; trap 'rm -rf "$FIX"' EXIT
printf -- '---\nname: good\neffort: xhigh\n---\nbody with effort: low in prose\n' > "$FIX/good.md"
printf -- '---\nname: bad\neffort: turbo\n---\n' > "$FIX/bad.md"
[ "$(fm_effort "$FIX/good.md")" = "xhigh" ] && ok || ko "flip: valid level not read"
[ -z "$(fm_effort "$FIX/bad.md")" ] && ok || ko "flip: invalid level accepted"
[ "$(fm "$FIX/good.md" | grep -c 'prose')" -eq 0 ] && ok || ko "flip: body leaked into frontmatter"
cat > "$FIX/reg.ts" <<'FX'
tiers: {
big: ['opus', 'fable'],
},
phases: {
judge: { tier: 'big', effort: 'xhigh' },
},
agents: {
// judge
Plan: 'judge', 'plan-challenger': 'judge',
},
skills: {
'ship-feature': 'plan', doc: 'apply',
},
FX
[ "$(rows "$FIX/reg.ts" agents | tr '\n' ,)" = "Plan judge,plan-challenger judge," ] \
&& ok || ko "flip: agents rows misparsed"
[ "$(rows "$FIX/reg.ts" skills | tr '\n' ,)" = "ship-feature plan,doc apply," ] \
&& ok || ko "flip: skills rows misparsed"
[ "$(phase_effort "$FIX/reg.ts" judge)" = "big xhigh" ] && ok || ko "flip: phase"
[ -z "$(phase_effort "$FIX/reg.ts" nothere)" ] && ok || ko "flip: ghost phase"
[ "$(tier_head "$FIX/reg.ts" big)" = "opus" ] && ok || ko "flip: tier head"
[ "$(rows "$REG" skills | wc -l)" -gt 40 ] && ok || ko "register.ts: skills rows not parsed"
[ "$(rows "$REG" agents | wc -l)" -gt 15 ] && ok || ko "register.ts: agents rows not parsed"
# ── 1) session default (spec D1)
# ── (b) tracked skills: row exists, frontmatter effort equals the row ────
NO_ROW_SKILLS=" find-docs graphify impeccable model-router "
check_skill() {
local f="$1" name phase want got
name="$(basename "$(dirname "$f")")"
case "$NO_ROW_SKILLS" in *" $name "*) return;; esac
phase="$(row_of skills "$name")"
[ -n "$phase" ] || { ko "skills/$name: no row in register.ts"; return; }
want="$(phase_effort "$REG" "$phase" | cut -d' ' -f2)"
got="$(fm_val "$R/$f" effort)"
[ -n "$got" ] || { ko "skills/$name: routed skill without effort:"; return; }
[ "$got" = "$want" ] && ok || ko "skills/$name: effort $got != row $phase ($want)"
}
while IFS= read -r f; do check_skill "$f"; done < <(
cd "$R" && git ls-files 'skills/*/SKILL.md' 'skills-external/*/SKILL.md')
# ── (c) tracked agents with a row: tier head == model:, effort == effort: ─
NO_ROW_AGENTS=" interviewer client-handover-writer "
tier_alias() { tier_head "$REG" "$(phase_effort "$REG" "$1" | cut -d' ' -f1)"; }
check_agent() {
local f="$1" name phase alias want got
name="$(basename "$f" .md)"
case "$NO_ROW_AGENTS" in *" $name "*) return;; esac
case "$name" in impeccable-*) return;; esac
phase="$(row_of agents "$name")"
[ -n "$phase" ] || { ko "agents/$name: no row in register.ts"; return; }
alias="$(tier_alias "$phase")"; got="$(fm_val "$R/$f" model)"
[ "$got" = "$alias" ] && ok || ko "agents/$name: model $got != row $phase ($alias)"
[ "$alias" = haiku ] && return
want="$(phase_effort "$REG" "$phase" | cut -d' ' -f2)"
got="$(fm_val "$R/$f" effort)"
[ "$got" = "$want" ] && ok || ko "agents/$name: effort $got != row $phase ($want)"
}
while IFS= read -r f; do check_agent "$f"; done < <(
cd "$R" && git ls-files 'agents/*.md' | grep -E '^agents/[^/]+\.md$' \
| grep -v '/README\.md$')
# ── (d) no shifter citer in skills, agents, lib ──────────────────────────
if (cd "$R" && git grep -qE 'Skill\(effort-|EFFORT SHIFT[S]:' -- skills agents lib ':!lib/tests'); then
ko "a shifter-skill or shift-header citer survives in skills/agents/lib"
else ok; fi
# paths split so this file itself matches no deleted-name grep
for s in skills/effort-low skills/effort-max lib/effort-""pins.txt lib/model-""check.sh; do
[ ! -e "$R/$s" ] && ok || ko "$s must be deleted"
done
# ── (e) D3 wiring markers ────────────────────────────────────────────────
mark() { # mark <phase> <skills...>
local ph="$1" s; shift
for s in "$@"; do has "skills/$s/SKILL.md" "route(phase=\"$ph\")"; done
}
mark orchestrate feat hotfix bugfix ship-feature init-project code-clean seo geo harden web-validate audit-delta
mark apply feat hotfix bugfix ship-feature init-project
mark reflect feat hotfix bugfix seo geo harden web-validate
mark plan ship-feature init-project onboard code-clean audit-delta
mark escalate ship-feature
[ "$(grep -o 'route(phase="escalate")' "$R/lib/verify-secure-loop.md" | wc -l)" -eq 3 ] \
&& ok || ko "verify-secure-loop.md: escalate route must appear 3 times"
for s in ship-feature init-project; do has "skills/$s/SKILL.md" 'effort="xhigh"'; done
has "skills/tour/SKILL.md" 'effort="xhigh"'
n_opus="$(grep -c 'model="opus",$' "$R/skills/onboard/SKILL.md")"
n_eff="$(grep -c 'effort="xhigh",$' "$R/skills/onboard/SKILL.md")"
[ "$n_opus" -ge 7 ] && [ "$n_opus" -eq "$n_eff" ] && ok \
|| ko "onboard: $n_opus model=\"opus\" dispatches vs $n_eff effort=\"xhigh\""
bad="$(grep 'model: "fable"' "$R/agents/client-handover-writer.md" | grep -vc 'effort="high"')"
[ "$bad" -eq 0 ] && ok || ko "client-handover-writer: $bad model: \"fable\" line(s) without effort=\"high\""
for s in ship-feature init-project feat bugfix web-validate seo hotfix geo harden code-clean audit-delta tour onboard; do
has "skills/$s/SKILL.md" 'ROUTING: follow $HOME/.claude/lib/effort-shift.md'
done
has "agents/client-handover-writer.md" 'ROUTING: follow $HOME/.claude/lib/effort-shift.md'
# ── (f) doctrine includes ────────────────────────────────────────────────
has "lib/model-gate.md" 'mcp__model-router__route'
has "lib/effort-shift.md" 'ToolSearch'
has "CLAUDE.global.md" 'route to `reflect` (high) through their'
# ── (g) session default and hooks (unchanged locks) ──────────────────────
has "settings.json" '"effortLevel": "high"'
# ── 2) hooks: env-var warning + live effort in the statusline (spec D1, D5)
has "hooks/session-start.sh" 'CLAUDE_CODE_EFFORT_LEVEL'
has "hooks/statusline.sh" 'CLAUDE_EFFORT'
# ── 3) agent pins (spec D2): one effort per agent file, judgment mode wins on mode-based agents
for a in hotfixer release-executor plugin-probe validator-analyzer; do fm_has_effort "agents/$a.md" low; done
for a in feater bugfixer code-cleaner onboarder scaffolder; do fm_has_effort "agents/$a.md" medium; done
for a in refactorer analyzer commit-changer doc-syncer handover-doc-writer; do fm_has_effort "agents/$a.md" high; done
for a in plan-challenger plugin-advisor verifier security-auditor seo-analyzer geo-analyzer; do fm_has_effort "agents/$a.md" xhigh; done
for a in interviewer client-handover-writer status-reporter; do fm_no_effort "agents/$a.md"; done
has "skills/init-project/SKILL.md" 'pin sonnet, effort medium'
# ── 4) skill entry levels (spec D3): the user's invocation sets the run's level
for s in status commit-change release-candidate doc capitalize close reconcile deploy profile plugin-check; do fm_has_effort "skills/$s/SKILL.md" low; done
for s in gitflow prune-memory; do fm_has_effort "skills/$s/SKILL.md" medium; done
for s in feat hotfix bugfix refactor web-validate harden seo geo; do fm_has_effort "skills/$s/SKILL.md" high; done
for s in ship-feature init-project onboard tour audit-delta analyze code-clean client-handover; do fm_has_effort "skills/$s/SKILL.md" xhigh; done
# BDR-108 round: the three repo skills that had no level
fm_has_effort "skills/skills-perso/SKILL.md" low
fm_has_effort "skills/pdf-translate/SKILL.md" medium
fm_has_effort "skills/site-motion/SKILL.md" high
# ── 9) vendored externals carry the level of lib/effort-pins.txt (BDR-108). The files live in
# skills-external/ (gitignored, machine-owned): the durable artifact is the map + the re-apply
# after the last vendoring step of install-plugins.sh AND update-all.sh; a skill not vendored
# yet SKIPs visibly (fresh clone before make plugin).
while read -r s lvl _; do
case "$s" in ''|'#'*) continue ;; esac
if [ -f "$R/skills-external/$s/SKILL.md" ]; then fm_has_effort "skills-external/$s/SKILL.md" "$lvl"
else printf 'SKIP skills-external/%s/SKILL.md not vendored yet (run make plugin)\n' "$s"; fi
done < "$R/lib/effort-pins.txt"
has "lib/effort-pins.txt" 'brainstorming xhigh'; has "lib/effort-pins.txt" 'writing-plans xhigh'
has "install-plugins.sh" 'apply_effort_pins "$REPO"'; has "update-all.sh" 'apply_effort_pins "$REPO"'
lacks "install-plugins.sh" 'for _s in brainstorming writing-plans; do'
ln_last() { grep -n "$2" "$R/$1" | tail -1 | cut -d: -f1; }
[ "$(ln_last install-plugins.sh 'apply_effort_pins "$REPO"')" -gt "$(ln_last install-plugins.sh 'rm -rf "$TFD_STAGE"')" ] \
&& ok || ko "install-plugins.sh: effort pins must be re-applied after the 21st pack refresh"
pins_ln=$(ln_last update-all.sh 'apply_effort_pins "$REPO"')
[ "$pins_ln" -gt "$(ln_last update-all.sh 'skills-external/$_tfd_name')" ] \
&& [ "$pins_ln" -gt "$(ln_last update-all.sh 'vendor_pinned_skills superpowers refresh')" ] \
&& ok || ko "update-all.sh: effort pins must be re-applied after the last vendoring step (21st pack)"
[ -x "$R/lib/effort-pins.sh" ] && ok || ko "lib/effort-pins.sh missing or not executable"
# 9b) design stack = ONE level (last loaded wins); site-motion (repo skill) pins the same one
stack_levels() { awk '/^# design stack/{f=1;next} f&&/^#$/{f=0} f&&!/^#/&&NF==2{print $2}' "$R/lib/effort-pins.txt" | sort -u; }
[ "$(stack_levels | wc -l)" -eq 1 ] && ok || ko "design stack must share ONE level in lib/effort-pins.txt (got: $(stack_levels | tr '\n' ' '))"
[ "$(stack_levels | wc -l)" -ge 1 ] && fm_has_effort "skills/site-motion/SKILL.md" "$(stack_levels | head -1)"
has "lib/effort-shift.md" 'Stacked skills share one level'
has "CLAUDE.global.md" 'lib/effort-pins.txt'
# ── 5) shifter skills + include (spec D4)
for l in low medium high xhigh max; do fm_has_effort "skills/effort-$l/SKILL.md" "$l"; has "skills/effort-$l/SKILL.md" "name: effort-$l"; done
has "lib/effort-shift.md" 'Headless sessions'
has "lib/effort-shift.md" 'Skill(effort-max)'
has "lib/effort-shift.md" 'never inside a dispatched agent'
has "lib/model-gate.md" 'lib/effort-shift.md'
# ── 6) orchestrator wiring (spec D4)
for s in feat hotfix bugfix ship-feature init-project onboard tour code-clean seo geo harden web-validate audit-delta; do
has "skills/$s/SKILL.md" 'lib/effort-shift.md'; has "skills/$s/SKILL.md" 'a lone Skill call is a no-op'; done
for s in feat hotfix bugfix ship-feature init-project code-clean seo geo harden web-validate audit-delta; do
has "skills/$s/SKILL.md" 'Skill(effort-medium)'; done
lacks "skills/onboard/SKILL.md" 'Skill(effort-medium)'; lacks "skills/tour/SKILL.md" 'Skill(effort-medium)'
has "agents/client-handover-writer.md" 'lib/effort-shift.md'; lacks "agents/client-handover-writer.md" 'Skill(effort-medium)'; has "agents/client-handover-writer.md" 'Skill(effort-high)'
for s in feat hotfix bugfix; do has "skills/$s/SKILL.md" 'Skill(effort-high)'; done
for s in ship-feature init-project onboard code-clean audit-delta; do has "skills/$s/SKILL.md" 'Skill(effort-xhigh)'; done
for s in seo geo harden web-validate; do has "skills/$s/SKILL.md" 'Skill(effort-high)'; done
for s in feat hotfix bugfix ship-feature init-project; do has "skills/$s/SKILL.md" 'Skill(effort-low)'; done
has "skills/feat/SKILL.md" 'effort-shift: nested commit-change'
# ── 6b) pairing rule documented (R11)
has "lib/effort-shift.md" 'lone Skill call is a no-op'
has "lib/effort-shift.md" 're-applies its'
[ "$(grep -c 'a lone Skill call is a no-op' "$R/skills/feat/SKILL.md")" -ge 1 ] && ok || ko "feat INC line must carry the pairing rule"
# ── 7) escalation at max (spec D4)
[ "$(grep -c 'Skill(effort-max)' "$R/lib/verify-secure-loop.md")" -eq 3 ] && ok || ko "verify-secure-loop.md must shift to max at its 3 caps"
has "skills/ship-feature/SKILL.md" 'Skill(effort-max)'
has "lib/challenge-plan.md" '/effort-max'
has "lib/verify-secure-loop.md" '/effort-max'
# ── 8) turn-reset re-assert after a prose gate followed by reflection
has "skills/bugfix/SKILL.md" 'effort-shift: turn reset'
# ── 11) audit tooling
has "lib/effort-shift.md" 'effort-audit.py'
[ -x "$R/lib/effort-audit.py" ] && ok || ko "lib/effort-audit.py missing or not executable"
# ── 6c) judgment dispatches re-raised, planning re-asserts, stronger locks (final review I1/I2/M5)
for s in ship-feature init-project; do has "skills/$s/SKILL.md" 'effort-shift: judgment dispatch'; has "skills/$s/SKILL.md" 'effort-shift: turn reset'; done
has "agents/client-handover-writer.md" 'effort-shift: judgment dispatch'
has "lib/effort-shift.md" 'Before any built-in or unpinned dispatch'
has "lib/model-gate.md" 'built-ins inherit the effort in force'
has "skills/ship-feature/SKILL.md" 'effort-shift: error recovery'
for s in feat hotfix bugfix seo geo harden web-validate ship-feature init-project onboard code-clean audit-delta; do has "skills/$s/SKILL.md" 'effort-shift: own level before the challenge'; done
has "update-all.sh" 'source "$REPO/lib/effort-pins.sh"'
# ── summary (later tasks insert their locks ABOVE this line)
printf 'effort-routing census: %d pass, %d fail\n' "$pass" "$fail"
[ "$fail" -eq 0 ]
printf 'PASS=%s FAIL=%s\n' "$pass" "$fail"; [ "$fail" -eq 0 ]
+5 -11
View File
@@ -321,24 +321,20 @@ expect link-sh "$(count link.sh higgsfield)" 0
expect profile-sh "$(count lib/profile.sh higgsfield)" 0
expect profiles \
"$(cat "$ROOT"/lib/profiles/*.profile | grep -cF higgsfield)" 0
expect pins-map "$(count lib/effort-pins.txt higgsfield)" 0
verdict OFF_BY_DEFAULT_WIRING
# ln_first / ln_last <file> <fixed string> — line number of a match.
ln_first() { grep -nF -- "$2" "$ROOT/$1" | head -1 | cut -d: -f1; }
ln_last() { grep -nF -- "$2" "$ROOT/$1" | tail -1 | cut -d: -f1; }
PINS="apply_effort_pins \"\$REPO\""
# install-plugins.sh: the sync sits in Step 8.6, before the effort pins
# (BDR-108); the CLI is proven by a probe, not by its shim; every login
# offer tests stdin alone (stdout is the tee pipe).
# install-plugins.sh: the sync sits in Step 8.6; the CLI is proven by a
# probe, not by its shim; every login offer tests stdin alone (stdout is
# the tee pipe).
sync_ln="$(ln_last install-plugins.sh 'higgsfield_sync_skills')"
expect after-8.5 "$(yn test "$sync_ln" -gt \
"$(ln_first install-plugins.sh 'Step 8.5: External skills')")" yes
expect before-8.7 "$(yn test "$sync_ln" -lt \
"$(ln_first install-plugins.sh 'Step 8.7: 21st.dev')")" yes
expect before-pins "$(yn test "$sync_ln" -lt \
"$(ln_last install-plugins.sh "$PINS")")" yes
expect probe-gates \
"$(yn test "$(count install-plugins.sh 'if higgsfield_cli_ok')" -ge 3)" yes
expect control "$(echo 'if [ -t 0 ] && [ -t 1 ]; then' | grep -cF -- '-t 1')" 1
@@ -347,14 +343,12 @@ expect stdin-tests \
"$(yn test "$(count install-plugins.sh '[ -t 0 ]')" -ge 3)" yes
verdict INSTALL_WIRING
# update-all.sh: refresh before the 21st block and before the pins re-apply,
# and the updated CLI is proven by the probe, after the npm call.
# update-all.sh: refresh before the 21st block, and the updated CLI is
# proven by the probe, after the npm call.
NPM_UP="npm install -g \"\$HF_PKG\""
sync_ln="$(ln_last update-all.sh 'higgsfield_sync_skills')"
expect before-21st "$(yn test "$sync_ln" -lt \
"$(ln_first update-all.sh '7.4. Update the 21st.dev')")" yes
expect before-pins "$(yn test "$sync_ln" -lt \
"$(ln_last update-all.sh "$PINS")")" yes
expect probe-after-npm "$(yn test \
"$(ln_first update-all.sh 'higgsfield_cli_ok')" -gt \
"$(ln_last update-all.sh "$NPM_UP")")" yes
-25
View File
@@ -1,25 +0,0 @@
#!/usr/bin/env bash
# lib/tests/model-check.test.sh — flip-tests for lib/model-check.sh (LRN-096)
set -u
S="$(cd "$(dirname "$0")/../.." && pwd)/lib/model-check.sh"
pass=0; fail=0
check() { if [ "$2" = "$3" ]; then pass=$((pass+1)); else fail=$((fail+1));
printf 'FAIL %s: got[%s] want[%s]\n' "$1" "$2" "$3"; fi; }
T="$(mktemp -d)"; trap 'rm -rf "$T"' EXIT
fx() { printf '{"model": "%s"}' "$1" > "$T/s.json"; }
run() { MODEL_CHECK_SETTINGS="$T/s.json" bash "$S" >"$T/out" 2>&1; echo "$?"; }
fx 'claude-fable-5[1m]'; check T1-fable-exit "$(run)" 0
check T1-fable-class "$(cut -d: -f1 <"$T/out")" big
fx 'claude-opus-4-8'; check T2-opus "$(run)" 0
fx 'claude-sonnet-5'; check T3-sonnet "$(run)" 2
fx 'claude-haiku-4-5-20251001'; check T4-haiku "$(run)" 2
fx 'opusplan'; check T5-opusplan "$(run)" 3
fx 'gpt-9-mega'; check T6-foreign "$(run)" 3
printf '{"no_model": true}' > "$T/s.json"; check T7-no-key "$(run)" 3
printf '{broken' > "$T/s.json"; check T8-malformed "$(run)" 3
check T9-missing-file "$(MODEL_CHECK_SETTINGS="$T/absent.json" bash "$S" >/dev/null 2>&1; echo $?)" 3
printf 'model-check: %d pass, %d fail\n' "$pass" "$fail"
[ "$fail" -eq 0 ]
+5 -4
View File
@@ -35,7 +35,7 @@ single `GATES — VERDICT:` line:
- `UNMET(n)` → hand the dev the CONTRACT path + the `NOT-MET` rows verbatim,
nothing else; re-run GATE 0. **No verifier is dispatched** — a red build or
a red suite is not a judgement call, and paying an LLM to discover it is
waste. **Max 3 floor iterations** → `Skill(effort-max)` (effort-shift: cap reached, diagnose at max before escalating; send it in the same message as the first tool call that gathers the escalation evidence), then STOP + human escalation with the rows.
waste. **Max 3 floor iterations** → `mcp__model-router__route(phase="escalate")` (route: cap reached, diagnose at max before escalating; send it with the first tool call that gathers the escalation evidence), then STOP + human escalation with the rows.
- `ABANDONED(n)` → floor green but a handoff stands. Continue to GATE 1; the
verifier surfaces it and its `ABANDONED(n)` verdict routes to the human
gate.
@@ -74,7 +74,7 @@ Parse its single `VERIFY — VERDICT:` line:
lines (NOT-MET / out-of-scope), nothing else: re-dispatch a FRESH executor
with those inputs only, never redo the fix by hand. Then re-run GATE 0 and
re-dispatch a FRESH verifier. Repeat.
**Max 3 conformity iterations** → `Skill(effort-max)` (effort-shift: cap reached, diagnose at max before escalating; send it in the same message as the first tool call that gathers the escalation evidence), then STOP + human escalation with the
**Max 3 conformity iterations** → `mcp__model-router__route(phase="escalate")` (route: cap reached, diagnose at max before escalating; send it with the first tool call that gathers the escalation evidence), then STOP + human escalation with the
CRITERIA table (the contract-vs-realized diff).
- `ABANDONED(n)` → direct human gate, never a dev loop (a dev cannot close
what was proven impossible). The human lifts the abandonment or accepts
@@ -104,9 +104,10 @@ Parse its single `SECURITY — VERDICT:` line:
(re-dispatch a FRESH executor, never fix by hand). Then re-run GATE 0, then
**re-verify the REQUEST first** (GATE 1, fresh verifier) — a security fix
can drift the behavior — **then re-run GATE 2** (fresh auditor), in that
order. **Max 3 security iterations** → `Skill(effort-max)` (effort-shift: cap reached, diagnose at max before escalating; send it in the same message as the first tool call that gathers the escalation evidence), then STOP + human escalation with the
order. **Max 3 security iterations** → `mcp__model-router__route(phase="escalate")` (route: cap reached, diagnose at max before escalating; send it with the first tool call that gathers the escalation evidence), then STOP + human escalation with the
BLOCKING table. Every STOP text names the level reached (`$CLAUDE_EFFORT`)
and suggests `/effort-max` for the relaunch.
and suggests relaunching with `ultrathink` in the prompt (turn floor) or
`/route effort=max` (sticky, `/route clear` after).
- `DEGRADED` (semgrep absent) → does NOT block on the tool's absence; surface
the checklist result + recommend `make plugin`. A DEGRADED run that still
BLOCKs (grep-caught secret/injection) blocks like any other.