From 5367b29188a3e84ef4a3da22685834fd10f5d28f Mon Sep 17 00:00:00 2001 From: bastien Date: Mon, 28 Sep 2026 18:17:34 +0200 Subject: [PATCH 01/23] =?UTF-8?q?docs(spec):=20effort=20tiering=20design?= =?UTF-8?q?=20=E2=80=94=20session=20high,=20agent=20pins,=20skill=20effort?= =?UTF-8?q?,=20phase=20shifts,=20max=20at=20loop=20caps?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- .../specs/2026-09-28-effort-tiering-design.md | 241 ++++++++++++++++++ 1 file changed, 241 insertions(+) create mode 100644 docs/superpowers/specs/2026-09-28-effort-tiering-design.md diff --git a/docs/superpowers/specs/2026-09-28-effort-tiering-design.md b/docs/superpowers/specs/2026-09-28-effort-tiering-design.md new file mode 100644 index 0000000..05fab60 --- /dev/null +++ b/docs/superpowers/specs/2026-09-28-effort-tiering-design.md @@ -0,0 +1,241 @@ +# Effort tiering — design + +Date: 2026-09-28 · Branch: `feature/effort-tiering` · Status: draft for review + +## 1. Intent + +Adapt the reasoning effort along a development run, not hold the whole +session at `xhigh`. The user's five-rung scale is the contract: + +| Rung | User definition | Examples | +|---|---|---| +| low | fix a line, rename a file, run a script | journal, commit, release bookkeeping | +| medium | day-to-day work | implement a closed plan, orchestrate between dispatches | +| high | a refactor, a bug that resists | investigation, diagnosis, contract drafting | +| xhigh | architecture, audit before validation | brainstorm, plan, challenge synthesis, gates | +| max | a stuck error, an error that cannot be recovered, or judged need | loop caps, error recovery | + +Automatic wherever the harness allows it. Where it does not, the user gets a +one-keystroke lever, never a silent default. + +Effort is a second axis on the BDR-077 routing table: BDR-077 fixed WHICH +MODEL runs each role and forbade inherit; this design fixes HOW HARD it +thinks, with the same no-inherit principle. + +## 2. What the harness allows (verified on Claude Code 2.1.283, 2026-09-28) + +Sources: code.claude.com/docs (model-config, skills, sub-agents, hooks), +the CHANGELOG (2.1.120, 2.1.149, 2.1.267, 2.1.280) and live probes in this +repo. + +| Mechanism | Verified behaviour | Evidence | +|---|---|---| +| Session level | Resolution order: `CLAUDE_CODE_EFFORT_LEVEL` env > `--effort` / `/effort` > settings (`modelSettings` per model, else top-level `effortLevel`) > model default (`high` on Fable 5.1). `max` is session-only, never persisted. `/effort auto` clears the per-model saved level only; a top-level `effortLevel` still applies. | docs | +| Subagent frontmatter `effort:` | Applied to the subagent. Absent → **inherits the session level**. | built-in on sonnet printed `xhigh`; impeccable agent pinned `medium` printed `medium` | +| Skill frontmatter `effort:`, user-typed `/skill` | Applied for the **rest of the turn**, AskUserQuestion included. | headless `/effort-probe-low`: every request at `low` | +| Skill frontmatter `effort:`, loaded by Claude through the Skill tool, **interactive** session | Applied for the rest of the turn. Last loaded skill wins, up and down. | this session: `xhigh` → probe max → `$CLAUDE_EFFORT=max`, request records `effort=max` → probe xhigh → back to `xhigh` | +| Same, **headless** (`-p`) | **Not applied** (neither `effort:` nor `model:`). | three `-p` runs, transcript effort unchanged | +| Prompt cache on a mid-turn shift | **Preserved** on Fable 5.1: first request at max read 206,996 cached tokens, wrote 1,164. | this session | +| Agent tool call site | No `effort` parameter (only `model`). One agent file = one effort. | tool schema | +| Hooks | Read `$CLAUDE_EFFORT` / `effort.level`; **cannot change** the level. | docs | +| `ultrathink` keyword | In-context nudge only; the effort sent to the API is unchanged. | docs | +| Env var | `CLAUDE_CODE_EFFORT_LEVEL` beats every frontmatter override. Unset on this machine. | docs + `env` | + +## 3. What the numbers say (6 days of local transcripts, all projects, 10,955 requests) + +Weights relative to input price: output ×5, cache read ×0.1, cache write ×1.25. + +| Item | Share | +|---|---| +| Cache reads (context re-read per request) | 53 % of weighted spend | +| All output tokens | 16 % | +| of which thinking | 8 % | +| Thinking located in the main loop | 97 % of thinking | +| Mean thinking per request: Fable main loop / sonnet subagent at xhigh | 1,430 / 26 tokens | +| Mean cached context per main-loop request | ~320 k tokens | + +Consequences. Executors barely think even at xhigh: pinning them is about +explicitness and future models (Opus 5.5 "thinks more per turn at a given +level"), not savings today. The direct lever of effort is single-digit +percent; the indirect lever (fewer steps at lower effort → fewer requests → +fewer cache reads) is unmeasured and gets an A/B in §9. The dominant cost is +main-loop context size, out of scope here (see `/capitalize`, `/clear`). + +## 4. Decisions + +### D1. Session default `high` +`settings.json` `effortLevel`: `xhigh` → `high`, explicit rather than +deleted: the statusline reads the key, and LRN-139 wants a visible value to +sweep at every model bump. Interactive chat outside a skill runs at the +model default; the user raises with `/effort xhigh` (session) or the new +`/effort-max` shifter (turn, see D4). `CLAUDE_CODE_EFFORT_LEVEL` must stay +unset (it would silence every override below); the session-start banner +warns if it is set. + +### D2. Agent pins (approach A) — repo-authored agents only + +| effort | Agents | +|---|---| +| low | hotfixer, release-executor, plugin-probe, validator-analyzer | +| medium | feater, bugfixer, code-cleaner, onboarder, scaffolder (was `high`; citer `skills/init-project/SKILL.md:98` updated) | +| high | refactorer, analyzer, commit-changer, doc-syncer, handover-doc-writer | +| xhigh | plan-challenger, plugin-advisor, verifier, security-auditor, seo-analyzer, geo-analyzer | +| none | interviewer, client-handover-writer (inline-load only, a pin would be inert and misleading, BDR-076 precedent); status-reporter (haiku, no effort support); `impeccable-*` (vendored) | + +Rules. One effort per agent file, so a mode-based agent (BDR-077) pins the +level of its **judgment** mode and its mechanical modes over-tier: the +fail-safe direction, and free on sonnet per §3. Built-ins (Explore, +general-purpose, Plan) cannot be pinned at the call site and inherit the +main loop's current level; Explore on Fable thinks ~1 token per request, +so no wrapper agent is created. Verifier and security-auditor sit at xhigh +by the user's own definition ("audit before validation"); on sonnet the +cost difference is nil. + +### D3. Skill frontmatter effort (approach B) — the run's entry level +Applies from the user's invocation for the rest of the turn. + +| effort | Skills | +|---|---| +| low | status, commit-change, release-candidate, doc, capitalize, close, reconcile, deploy, profile, plugin-check | +| medium | gitflow, prune-memory, find-docs | +| high | feat, hotfix, bugfix, refactor, web-validate, harden, seo, geo | +| xhigh | ship-feature, init-project, onboard, tour, audit-delta, analyze, code-clean, client-handover, spec, skillify, brainstorming, writing-plans | +| unlisted | session default, by design (gstack and plugin skills are external; graphify is machine-owned) | + +`brainstorming` and `writing-plans` are vendored superpowers skills: the +one-line patch drifts from upstream at each resync; a census lock (§7) +makes the loss loud. + +A skill loaded by Claude as a sub-step (feat → commit-change) also shifts +the level for the rest of the turn (interactive, §2), so orchestrators +re-assert their own level after any nested Skill call whose level differs +(D4 protocol). + +### D4. Phase shifts inside a run (approach C) +Five one-line skills, no body beyond a sentence, user-invocable: +`effort-low`, `effort-medium`, `effort-high`, `effort-xhigh`, `effort-max`. +Descriptions as pre-validated against the routing census (pairwise +similarity ≤ 0.03). Protocol in a shared include `lib/effort-shift.md`, +mirroring `lib/model-gate.md`: + +- A shift is a `Skill(effort-)` call on the main loop. Never inside a + dispatched agent (agents run on their pin). One tool round-trip, + cache-safe (§2). +- Orchestrator wiring, three points each: `effort-medium` when the plan is + closed and the dispatch phase starts; `effort-low` before the + capitalize / journal / doc-commit tail; `effort-max` at an escalation + point, then the skill's own level again once the diagnosis is produced. +- Re-assert the skill's own level after any nested `Skill(...)` call whose + frontmatter carries a different effort (D3): the nested level would + otherwise hold for the rest of the turn. +- **Escalation points (automatic max)**: verify-secure loop GATE 1 cap + (3 conformity rounds) and GATE 2 cap (3 security rounds), before the + human-escalation table is composed; ship-feature STEP 4b, so the + inline analyzer DEBUG read runs at max. Full conversation context is the + asset here; a fresh diagnoser agent was considered and dropped (YAGNI: + no context, one more agent, same effort). +- **Not automatic, by doctrine**: the challenge fail-safe (a mute + challenger is an infrastructure failure, not a reasoning problem) and + the "gone WRONG → STOP" rule (STOP precedes any further reasoning). Both + STOP messages name the level reached and suggest `/effort-max` for the + relaunch: a turn-scoped max the user gets by typing one command. +- **Turn reset**: a prose gate that ends the turn (model-gate STOP, loop + cap STOP, and the four prose gates found in bugfix, ship-feature ×2, + init-project) drops the resumed turn to the session level. The plan + audits each such gate: if the resumed phase is reflection, the resume + step re-asserts with `Skill(effort-xhigh)`; if it is dispatch or + orchestration, session `high` is adequate and nothing is added. +- **Headless limitation**: `-p`, `claude agents` and SDK sessions ignore + skill-level effort (§2); runs there stay at the session level. Documented + in the include, no mitigation. + +### D5. Visibility +`hooks/statusline.sh` shows `$CLAUDE_EFFORT` when set (the live level, +shifts included) and falls back to the settings key. `/tasks` already shows +each subagent's effort (2.1.243). + +## 5. Alternatives rejected + +- **Keep xhigh, pin executors only**: executors think ~26 tokens per + request; the burn is in the main loop (§3). +- **Escalation diagnoser agent (`model: fable`, `effort: max`)**: chosen + before the interactive probe proved C viable; dropped because the + main-loop shift keeps the full failure context and adds no agent. +- **Move reflection into `model: fable` skill-runner children with + `effort: xhigh`, session at medium**: loses conversation context and + interactivity (BDR-077 retention criteria), heavy re-architecture for a + lever C delivers in five one-line files. +- **Rewrite settings.json mid-run to shift effort**: global side effect on + every session, LRN-098 drift class, fights the harness. +- **`maxEffortLevel` cap on sonnet**: pins already bound each agent; a cap + would hide a mis-pin instead of failing it in the census. + +## 6. Files touched + +| Area | Change | +|---|---| +| `settings.json` | `effortLevel` → `high` (curated config: read the diff, LRN-098) | +| `agents/*.md` (20) | `effort:` line per D2; `skills/init-project/SKILL.md:98` citer | +| `skills/*/SKILL.md` (33) | `effort:` line per D3, including the two vendored superpowers skills | +| `skills/effort-{low,medium,high,xhigh,max}/SKILL.md` | new, frontmatter + one sentence | +| `lib/effort-shift.md` | new include: protocol, wiring points, escalation, turn reset, headless note | +| `lib/model-gate.md` §4 | one paragraph: effort is the second axis, pointer to the include | +| `lib/verify-secure-loop.md` | `Skill(effort-max)` before each cap's human-escalation table; STOP text names the level | +| `lib/challenge-plan.md` | STOP text names the level, suggests `/effort-max` | +| orchestrator SKILL.md (feat, hotfix, bugfix, ship-feature, init-project, onboard, tour, code-clean, seo, geo, harden, web-validate, client-handover, audit-delta) | include line + the three wiring points; ship-feature 4b max | +| `hooks/statusline.sh`, `hooks/session-start.sh` | live effort display; env-var warning | +| `lib/tests/effort-routing.test.sh` | new census suite (§7) | +| `CHANGELOG.md`, `.claude/memory/*` | release note; BDR + LRN + EVAL + journal (§8) | + +## 7. Tests and census (`make test`) + +New suite `lib/tests/effort-routing.test.sh`, `grep -qF` locks in the +`model-routing.test.sh` style, flip-tested first (BDR-100): + +1. Every repo-authored agent outside the "none" list has `effort: ` + in its first 10 frontmatter lines, level in the allowed set; the "none" + list has no `effort:`. +2. Tier locks per D2 (one `has` per agent). +3. Skill locks per D3 (one `has` per skill), including `brainstorming` and + `writing-plans` (the resync alarm). +4. The five shifter skills exist with the exact `name:` and `effort:`. +5. `lib/effort-shift.md` is included by every orchestrator in the §6 list; + `verify-secure-loop.md` and `ship-feature/SKILL.md` contain the + `Skill(effort-max)` lock. +6. `settings.json` `effortLevel` is `high`. +7. `skill-routing-census` stays green with the five new descriptions + (pre-validated). +8. `doctrine-citers` stays green: no new `CLAUDE.md "…"` citation; the + doctrine lives in `lib/`. + +Per-wave smoke, planted input, disk-verified (BDR-077 precedent): +W1 dispatch a pinned agent that echoes `$CLAUDE_EFFORT`; W2 invoke `/status` +and read `effort=low` in the transcript; W3 run a skill through a shift +and read the request sequence; W4 statusline shows the live level. + +## 8. Rollout + +Four waves on `feature/effort-tiering`, one commit each, smoke as merge +gate, human signal for `gitflow finish`: + +- W1 settings + agent pins + test suite + model-gate paragraph. +- W2 skill frontmatter (D3) + superpowers patch. +- W3 shifter skills + `lib/effort-shift.md` + orchestrator wiring + + escalation points + turn-reset audit. +- W4 statusline + banner warning + CHANGELOG + registries. + +Registries: BDR (effort tiering, this spec's decisions and rejected +alternatives), LRN (skill effort applies on user invocation and on +interactive Skill-tool loads, not in `-p`; shifts are cache-safe), EVAL +(the §3 measurement and its method), journal line. + +## 9. Measurement after rollout + +A/B on a repeatable skill run (`/reconcile` on this repo, session `high` +vs `xhigh`): requests, output tokens, thinking tokens, wall time, from the +transcript. Records whether the indirect lever exists. Goes to EVAL. + +## 10. Out of scope + +Main-loop context size (the 53 %), gstack and plugin skills, the +`impeccable-*` agents, `graphify` (machine-owned), headless sessions. From 854b74e9a47b5eef0cdecdec8332be151d7b6d12 Mon Sep 17 00:00:00 2001 From: bastien Date: Mon, 28 Sep 2026 18:24:34 +0200 Subject: [PATCH 02/23] =?UTF-8?q?chore(memory):=20LRN-179=20+=20EVAL-035?= =?UTF-8?q?=20=E2=80=94=20effort=20spike=20facts,=20thinking-share=20measu?= =?UTF-8?q?rement?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- .claude/memory/evals.md | 8 ++++++++ .claude/memory/learnings.md | 5 +++++ 2 files changed, 13 insertions(+) diff --git a/.claude/memory/evals.md b/.claude/memory/evals.md index 509f95d..07573d0 100644 --- a/.claude/memory/evals.md +++ b/.claude/memory/evals.md @@ -55,6 +55,7 @@ rules: | EVAL-032 | 2026-09-27 | 4 parallel feater executors, one tree, gate loop: verifier caught a vacuous test, security caught a partial-write; my oracles wrong twice | keep same-tree parallel dispatch with disjoint FILE SCOPE + orchestrator-owned shared files; blind verifier stays; measure oracles on precedents | | EVAL-033 | 2026-09-28 | case 7: 2 analyzers + 2 executors + 3 re-dispatches; verifiers caught shape, convention and my wrong count; security caught an env override | brief names the scratchpad path explicitly (3 /tmp leftovers); keep blind verifiers; count claims get an artifact | | EVAL-034 | 2026-09-28 | catalog prune + 21st gate: two challenge rounds each found what r3 missed (nested SKILL.md, fixture cp lists, in-session export); my ledgers failed twice (heredoc CHECKs); 5 executors DONE first pass; verifier gap = tool false positive | keep the confirmation pass on any plan that changed materially; one-line CHECKs; grep fixture cp lists before a `source` | +| EVAL-035 | 2026-09-28 | thinking-share measurement, 6 days of transcripts (10,955 requests): thinking = 8 % of weighted spend, 97 % of it in the main loop; sonnet subagents at xhigh think 26 tok/request; cache reads = 53 % | pins = explicitness not savings; main-loop effort + context size are the levers; A/B after rollout | --- @@ -330,3 +331,10 @@ Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itse - **Result**: prune — challengers closed 8 MAJOR at r3, the confirmation pass still found 1 BLOCKER (nested SKILL.md in browser-skills/openclaw/node_modules) + 3 MAJOR (setup's global symlink, update-all 3rd copy, fixture cp lists); executors 4/4 DONE first pass; GATE 0 UNMET(4) = my heredoc CHECKs ([[LRN-176]]); verifier ECARTS(1) = floor-guard false positive ([[BLK-023]]), CONFORME at iteration 2; security PASS. 21st gate — three lenses: my shared-helper reflex = BLOCKER ×2 ([[LRN-178]]), my `export TWENTYFIRST_TOKEN` remedy = MAJOR (env does not persist); confirmation pass pinned the diagnostic format; executor DONE first pass, CONFORME 7/7, PASS. - **Anomalies**: (1) both times the confirmation pass found real defects after "all MAJOR closed" → r3 is not a stopping point; (2) every gate failure of the day was mine (ledger format, tool pattern), none the executors'; (3) verifier and challengers each re-ran the live oracles themselves (link.sh, `set full`, the gate) — cheap, decisive; (4) the user's rule ("full ⊇ every profile") arrived at pass B and inverted a settled plan step: pass B before challenge is the right order. - **Action**: keep the single confirmation pass mandatory when a plan changed materially; contract CHECKs one line, files under `.oracles/`; grep fixture `cp` lists before any new `source`; run the live oracle once by hand before dispatching the verifier. + +## EVAL-035 — effort burn measured, premise corrected: subagents don't think, the main loop does +- **Date**: 2026-09-28 +- **Output checked**: my hypothesis "executors inherit xhigh → that is the burn" vs `effort_split2.py` (scratchpad) over `~/.claude/projects/*`: main jsonl + `*/subagents/*.jsonl`, `isSidechain` split; weights output ×5, cache read ×0.1, cache write ×1.25. +- **Result**: main loop 67 % of weighted spend, 97 % of thinking (Fable 1,430 think-tok/request); sonnet subagents 5,268 requests at xhigh, 26 think-tok/request; thinking = 8 % of spend, all output 16 %, cache reads 53 % (main-loop context ~320 k tok/request). Window 6 days only. Indirect effect of effort (fewer steps → fewer requests) unmeasured. +- **Anomaly**: design was framed around executor pins; one script inverted it before any edit. Measure before routing. +- **Action**: pins stay (explicitness, future models); main-loop skill effort + phase shifts carry the savings; A/B `/reconcile` high vs xhigh after rollout; context size = bigger lever, separate track. diff --git a/.claude/memory/learnings.md b/.claude/memory/learnings.md index 269bfcf..b90abfa 100644 --- a/.claude/memory/learnings.md +++ b/.claude/memory/learnings.md @@ -198,6 +198,7 @@ rules: | LRN-176 | 2026-09-28 | gates.sh `CHECK:` is single-line: a heredoc body reads as prose, the oracle runs `python3 -` on empty stdin and lands NOT-MET "marker absent", never ERROR; multi-line oracle → `.oracles/*.py` | writing contract oracles longer than one line | | LRN-177 | 2026-09-28 | gstack skills hardcode `~/.claude/skills/gstack/` (83 paths: bin, scripts, ETHOS.md, */sections, review/specialists, make-pdf/dist, freeze/bin…); only bin + browse/dist were linked → dead skills and vacuous hooks (exit 127); ./setup plants a global symlink; whole-dir link exposes nested SKILL.md; `apply` is additive, `set` parks | any gstack wiring change, any "gstack skill fails" report | | LRN-178 | 2026-09-28 | a top-level `source` added to a lib breaks every hermetic suite that copies that lib alone into a fixture; grep the `cp` lists before adding one, or source lazily inside the branch that needs it | adding `source` to profile.sh / toggle-external.sh / any lib the suites copy | +| LRN-179 | 2026-09-28 | Skill `effort:` frontmatter shifts the MAIN LOOP for the rest of the turn on user slash invocation AND on interactive Skill-tool loads (last loaded wins, both directions, prompt cache kept); NOT applied in `-p`/headless; agent pins always honoured, unpinned agents inherit session | effort tiering; any skill or agent that must think more or less than the session | --- @@ -1640,3 +1641,7 @@ Rule: when editing a doctrine file under structure locks, grep the test's lock s ## LRN-178 — before a new top-level `source`, grep the fixture `cp` lists - **Context**: twice in one day. E1b's `source gstack-removed.sh` in profile.sh/toggle-external.sh needed a `cp` line in three suites (profile-default, profile-set-managed, toggle-external-repo-resolution) — caught by the confirmation challenger, fixed in scope. My 21st helper plan would have added a second top-level `source` to toggle-external.sh with no fixture update → four suites red under `set -euo pipefail`; two challengers flagged it as BLOCKER, the helper was dropped. - **Apply**: `grep -n "cp .*lib/" lib/tests/*.sh` before adding a `source` to a lib; either widen every fixture copy in the same change or source lazily inside the one branch that needs it. Prefer the inline predicate when only one caller needs the new semantics ([[BDR-105]]). + +## LRN-179 — skill `effort:` shifts the main loop for the rest of the turn, interactive only +- **Context**: effort-tiering spike 2026-09-28, Claude Code 2.1.283, Fable 5.1. Probes = `$CLAUDE_EFFORT` in Bash + transcript `effort` field per request. User-typed `/probe-low` → whole turn `low`. Skill-tool load in interactive session → `max` then `xhigh`, last loaded wins, both directions; first request after the switch read 206,996 cached tokens, wrote 1,164 (cache kept). Three `-p` runs: neither `effort:` nor `model:` skill frontmatter applied via Skill tool. Agent pin honoured (impeccable `medium`), unpinned built-in on sonnet inherited `xhigh`. Docs agent claimed "ultrathink keyword does not exist": wrong, docs = in-context nudge, API effort unchanged. Harness claims get verified against the harness ([[LRN-046]]). +- **Apply**: main-loop effort per phase = `Skill(effort-)` on the main loop, never inside a dispatched agent; headless runs stay at session level; keep `CLAUDE_CODE_EFFORT_LEVEL` unset (beats every frontmatter). Spec `docs/superpowers/specs/2026-09-28-effort-tiering-design.md`. From 4b722e05c96a849a864176b7aa03b735a45d390b Mon Sep 17 00:00:00 2001 From: bastien Date: Mon, 28 Sep 2026 18:35:38 +0200 Subject: [PATCH 03/23] docs(plan): effort tiering implementation plan, 11 tasks in 4 waves; TODO section --- .claude/tasks/TODO.md | 9 + .../plans/2026-09-28-effort-tiering.md | 1099 +++++++++++++++++ 2 files changed, 1108 insertions(+) create mode 100644 docs/superpowers/plans/2026-09-28-effort-tiering.md diff --git a/.claude/tasks/TODO.md b/.claude/tasks/TODO.md index 23fccc5..4932005 100644 --- a/.claude/tasks/TODO.md +++ b/.claude/tasks/TODO.md @@ -1,5 +1,14 @@ # TODO +## 2026-09-28 — effort tiering: session high, agent pins, skill levels, phase shifts (feature/effort-tiering) +Spec `docs/superpowers/specs/2026-09-28-effort-tiering-design.md`, plan +`docs/superpowers/plans/2026-09-28-effort-tiering.md`. Approved 2026-09-28: session +high, A+B+C, max on the main loop at the loop caps + ship-feature 4b, superpowers patch. +- [ ] W1 settings high + banner warning + statusline live level + 20 agent pins + census suite (Tasks 1-3) +- [ ] W2 33 skill entry levels + superpowers xhigh with resync re-apply (Tasks 4, 9) +- [ ] W3 five shifters + lib/effort-shift.md + orchestrator wiring + max at caps/4b + gate audit (Tasks 5-8) +- [ ] W4 BDR id + CHANGELOG + EVAL A/B + journal + audit script (Tasks 10-11) + ## 2026-09-28 — tier 2: vendor 7 superpowers skills, drop the plugin (feature/superpowers-vendored) User go "fais le tier 2" (decision 2026-09-28, batch 1). Contract `.claude/tasks/contracts/2026-09-28-superpowers-vendored-1357.md`. diff --git a/docs/superpowers/plans/2026-09-28-effort-tiering.md b/docs/superpowers/plans/2026-09-28-effort-tiering.md new file mode 100644 index 0000000..edfbfe2 --- /dev/null +++ b/docs/superpowers/plans/2026-09-28-effort-tiering.md @@ -0,0 +1,1099 @@ +# Effort Tiering Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Route reasoning effort per role and per phase (low → max) across the claude-config orchestrators, instead of one session-wide `xhigh`. + +**Architecture:** Effort becomes the second axis of the BDR-077 routing table. Three layers, each a one-line frontmatter mechanism the harness already honours: `effort:` pins on the 20 repo-authored agents (dispatched work), `effort:` on the 33 user-invoked skills (the run's entry level), and five empty "shifter" skills the orchestrators load at phase boundaries (`Skill(effort-)`), including `max` at the loop caps and ship-feature error recovery. A census suite locks every value. + +**Tech Stack:** bash, GNU sed, python3 stdlib, jq, shellcheck, `make test` (suites under `lib/tests/*.test.sh` are auto-discovered). + +**Spec:** `docs/superpowers/specs/2026-09-28-effort-tiering-design.md` (read it first; every decision below is argued there, §2 holds the harness evidence, §3 the measurement). + +## Global Constraints + +- Claude Code ≥ 2.1.267 on the executing machine (skill/agent `effort:` honoured on Fable); the spike ran on 2.1.283. +- `CLAUDE_CODE_EFFORT_LEVEL` must be unset in every session that runs a smoke: it silences every frontmatter override. +- Branch `feature/effort-tiering` (exists, off develop). Every task ends with a commit on it; never `--no-verify`; never commit on develop/main; `gitflow finish` only on the human's signal. +- `make test` green after every task; `shellcheck *.sh hooks/*.sh lib/*.sh` clean after any shell edit. +- Allowed effort values, exactly: `low`, `medium`, `high`, `xhigh`, `max`. `max` never in `settings.json` (harness rejects it there). +- Never edit: `agents/impeccable-*.md`, `skills/graphify/**`, anything under `skills-external/`. The two vendored superpowers files edited (`skills/brainstorming/SKILL.md`, `skills/writing-plans/SKILL.md`) get a resync re-apply in Task 9. +- No new `CLAUDE.md "…"` citations anywhere (doctrine-citers census); the doctrine lives in `lib/`. +- `BDR-NEXT` is a literal token used in lib text and CHANGELOG until Task 10 computes the real id and replaces it. It must not survive Task 10. +- Shell: functions ≤ 25 logic lines, 80-char lines. Registry entries: English, caveman. +- Commit messages: no attribution lines. + +## Review Focus + +1. A run launched headless (`claude -p`, `claude agents`, SDK) never shifts: the include must say so, and the census locks that sentence (Task 5). +2. `CLAUDE_CODE_EFFORT_LEVEL` exported in the user's shell silently disables every pin and shift: the session banner must warn, tested by running the hook with the variable set (Task 2). +3. A nested skill at a different level (feat → commit-change at low) leaves the rest of the run at low: feat must re-assert `Skill(effort-high)` right after, locked by the census (Task 6). +4. An `effort:` on a model without effort support (haiku) is at best ignored: `status-reporter` must carry none, locked (Task 3). +5. A superpowers resync overwrites the vendored frontmatter: the census lock alarms, and install-plugins re-applies it (Task 9). + +--- + +### Task 1: Census suite skeleton with flip-test and the settings lock + +**Files:** +- Create: `lib/tests/effort-routing.test.sh` +- Test: itself (`make test suite=lib/tests/effort-routing.test.sh`) + +**Interfaces:** +- Produces: helpers `has`, `lacks`, `fm`, `fm_effort`, `fm_has_effort `, `fm_no_effort `, counters `pass`/`fail`, final line `effort-routing census: N pass, M fail`. Later tasks append `has`/`fm_has_effort` lines to this file, above the summary block. + +- [ ] **Step 1: Write the suite with the flip-test and one real lock that fails today** + +```bash +cat > lib/tests/effort-routing.test.sh <<'EOF' +#!/usr/bin/env bash +# lib/tests/effort-routing.test.sh — census: effort tiering (BDR-NEXT) +# agent pins, skill entry levels, shifter skills, orchestrator wiring, settings. +set -u +R="$(cd "$(dirname "$0")/../.." && pwd)" +pass=0; fail=0 +ok() { pass=$((pass+1)); } +ko() { fail=$((fail+1)); printf 'FAIL %s\n' "$1"; } +has() { if grep -qF "$2" "$R/$1"; then ok; else ko "$1 missing: $2"; fi; } +lacks() { if grep -qF "$2" "$R/$1"; then ko "$1 must NOT contain: $2"; else ok; fi; } +# frontmatter = the lines between the first two '---' lines +fm() { awk 'NR==1&&/^---$/{p=1;next} p&&/^---$/{exit} p' "$1"; } +fm_effort() { fm "$1" | grep -E '^effort: (low|medium|high|xhigh|max)$' | head -1 | cut -d' ' -f2; } +fm_has_effort() { + got="$(fm_effort "$R/$1")" + if [ "$got" = "$2" ]; then ok; else ko "$1 frontmatter effort must be '$2', got '${got:-none}'"; fi +} +fm_no_effort() { if fm "$R/$1" | grep -q '^effort:'; then ko "$1 must NOT pin effort"; else ok; fi; } + +# ── flip-test: the frontmatter reader must accept a valid level and reject an invalid one +FIX="$(mktemp -d)"; trap 'rm -rf "$FIX"' EXIT +printf -- '---\nname: good\neffort: xhigh\n---\nbody with effort: low in prose\n' > "$FIX/good.md" +printf -- '---\nname: bad\neffort: turbo\n---\n' > "$FIX/bad.md" +[ "$(fm_effort "$FIX/good.md")" = "xhigh" ] && ok || ko "flip: valid level not read" +[ -z "$(fm_effort "$FIX/bad.md")" ] && ok || ko "flip: invalid level accepted" +[ "$(fm "$FIX/good.md" | grep -c 'prose')" -eq 0 ] && ok || ko "flip: body leaked into frontmatter" + +# ── 1) session default (spec D1) +has "settings.json" '"effortLevel": "high"' + +# ── summary (later tasks insert their locks ABOVE this line) +printf 'effort-routing census: %d pass, %d fail\n' "$pass" "$fail" +[ "$fail" -eq 0 ] +EOF +chmod +x lib/tests/effort-routing.test.sh +``` + +- [ ] **Step 2: Run it, expect the flip-test to pass and the settings lock to fail** + +Run: `make test suite=lib/tests/effort-routing.test.sh` +Expected: `FAIL settings.json missing: "effortLevel": "high"` then `effort-routing census: 3 pass, 1 fail`, non-zero exit. + +- [ ] **Step 3: Shellcheck** + +Run: `shellcheck lib/tests/effort-routing.test.sh` +Expected: no output. + +- [ ] **Step 4: Commit** + +```bash +git add lib/tests/effort-routing.test.sh +git commit -m "test(effort): census suite skeleton with flip-test and settings lock" +``` + +--- + +### Task 2: Session default `high`, banner warning, live effort in the statusline + +**Files:** +- Modify: `settings.json` (line with `"effortLevel"`) +- Modify: `hooks/session-start.sh` (after the line `unset _claude_real _repo_dir`, and after the banner's closing box line) +- Modify: `hooks/statusline.sh:36-41` (the `EFFORT=` block) +- Test: `lib/tests/effort-routing.test.sh` + +**Interfaces:** +- Produces: banner line `⚠️ CLAUDE_CODE_EFFORT_LEVEL= set: skill/agent effort pins ignored` when the variable is exported; statusline `effort: `. + +- [ ] **Step 1: Add the two hook locks to the census (above the summary block)** + +```bash +python3 - <<'PY' +p="lib/tests/effort-routing.test.sh"; s=open(p).read() +s=s.replace("# ── summary", """# ── 2) hooks: env-var warning + live effort in the statusline (spec D1, D5) +has "hooks/session-start.sh" 'CLAUDE_CODE_EFFORT_LEVEL' +has "hooks/statusline.sh" 'CLAUDE_EFFORT' + +# ── summary""") +open(p,"w").write(s) +PY +``` + +- [ ] **Step 2: Run the suite, expect 3 failures (settings + two hooks)** + +Run: `make test suite=lib/tests/effort-routing.test.sh` +Expected: three `FAIL` lines, non-zero exit. + +- [ ] **Step 3: Set the session default to `high`** + +```bash +sed -i 's/"effortLevel": "xhigh"/"effortLevel": "high"/' settings.json +git diff settings.json +``` +Expected diff: exactly one changed line. If anything else moved (LRN-098: `/effort` and `/model` rewrite this file), `git checkout settings.json` and redo the sed. + +- [ ] **Step 4: Banner warning in session-start.sh** + +Insert after the line `unset _claude_real _repo_dir`: + +```bash +python3 - <<'PY' +p="hooks/session-start.sh"; L=open(p).read().split("\n") +i=L.index("unset _claude_real _repo_dir") +L[i+1:i+1]=[ +"", +"# Effort tiering (BDR-NEXT): this env var beats every skill/agent `effort:` pin.", +"EFFORT_WARN=\"\"", +"if [ -n \"${CLAUDE_CODE_EFFORT_LEVEL:-}\" ]; then", +" EFFORT_WARN=\"⚠️ CLAUDE_CODE_EFFORT_LEVEL=${CLAUDE_CODE_EFFORT_LEVEL} set: skill/agent effort pins ignored\"", +"fi", +] +open(p,"w").write("\n".join(L)) +PY +grep -n '└' hooks/session-start.sh +``` +Expected: two lines, both plain `echo` statements (an early fix-hint box near l.32, the main banner near l.255). The python below inserts after the last one: + +```bash +python3 - <<'PY' +p="hooks/session-start.sh"; L=open(p).read().split("\n") +i=max(k for k,l in enumerate(L) if "└" in l) +L.insert(i+1, '[ -n "$EFFORT_WARN" ] && printf \'%s\\n\' "$EFFORT_WARN"') +open(p,"w").write("\n".join(L)) +PY +``` + +- [ ] **Step 5: Live effort in the statusline** + +Replace the block from the comment `# Effort level from settings.json` through the `fi` that closes `if [ -f "$REPO/settings.json" ]; then`: + +```bash +python3 - <<'PY' +p="hooks/statusline.sh"; s=open(p).read() +old_start=s.index("# Effort level from settings.json") +old_end=s.index("fi\n", s.index('jq -r \'.effortLevel', old_start))+3 +new='''# Effort level: the live value when the harness exports it (skill/agent +# `effort:` shifts included, BDR-NEXT), else the persisted settings.json key +# (.effortLevel — set by /effort or manual edit; symlinked into ~/.claude). +EFFORT="${CLAUDE_EFFORT:-}" +if [ -z "$EFFORT" ] && [ -f "$REPO/settings.json" ]; then + EFFORT=$(jq -r '.effortLevel // "?"' "$REPO/settings.json" 2>/dev/null) +fi +[ -z "$EFFORT" ] && EFFORT="?" +''' +open(p,"w").write(s[:old_start]+new+s[old_end:]) +PY +sed -n 34,46p hooks/statusline.sh +``` +Expected: the new block, no duplicate `EFFORT=` lines. + +- [ ] **Step 6: Run the hook with the variable set, then the suites** + +```bash +CLAUDE_CODE_EFFORT_LEVEL=medium bash hooks/session-start.sh 2>/dev/null | grep -c 'effort pins ignored' +bash hooks/session-start.sh 2>/dev/null | grep -c 'effort pins ignored' +shellcheck hooks/session-start.sh hooks/statusline.sh +make test suite=lib/tests/effort-routing.test.sh +``` +Expected: `1`, then `0`, shellcheck silent, census all pass. + +- [ ] **Step 7: Full test run and commit** + +Run: `make test` +Expected: every suite green (curated-config-guard accepts a hand edit of settings.json). + +```bash +git add settings.json hooks/session-start.sh hooks/statusline.sh lib/tests/effort-routing.test.sh +git commit -m "feat(effort): session default high, env-var warning, live effort in statusline" +``` + +--- + +### Task 3: Agent effort pins (spec D2) + +**Files:** +- Modify: 20 files `agents/.md` (line 5 is `model: sonnet` or `model: opus` in every one of them; `agents/scaffolder.md` already has `effort: high` on line 6) +- Modify: `skills/init-project/SKILL.md` (the sentence `(pin sonnet, effort high — BDR-077`) +- Test: `lib/tests/effort-routing.test.sh` + +**Interfaces:** +- Produces: `effort: ` on line 6 of each pinned agent. + +- [ ] **Step 1: Add the 23 locks to the census (above the summary block)** + +```bash +python3 - <<'PY' +p="lib/tests/effort-routing.test.sh"; s=open(p).read() +locks="""# ── 3) agent pins (spec D2): one effort per agent file, judgment mode wins on mode-based agents +for a in hotfixer release-executor plugin-probe validator-analyzer; do fm_has_effort "agents/$a.md" low; done +for a in feater bugfixer code-cleaner onboarder scaffolder; do fm_has_effort "agents/$a.md" medium; done +for a in refactorer analyzer commit-changer doc-syncer handover-doc-writer; do fm_has_effort "agents/$a.md" high; done +for a in plan-challenger plugin-advisor verifier security-auditor seo-analyzer geo-analyzer; do fm_has_effort "agents/$a.md" xhigh; done +for a in interviewer client-handover-writer status-reporter; do fm_no_effort "agents/$a.md"; done +has "skills/init-project/SKILL.md" 'pin sonnet, effort medium' + +# ── summary""" +open(p,"w").write(s.replace("# ── summary", locks)) +PY +make test suite=lib/tests/effort-routing.test.sh | tail -1 +``` +Expected: `... 20 fail` (19 pins + the citer; the three `fm_no_effort` pass already). + +- [ ] **Step 2: Apply the pins** + +```bash +pin() { L=$1; shift; for a in "$@"; do + sed -i "0,/^model: \(sonnet\|opus\)$/s//&\neffort: $L/" "agents/$a.md"; done; } +pin low hotfixer release-executor plugin-probe validator-analyzer +pin medium feater bugfixer code-cleaner onboarder +sed -i 's/^effort: high$/effort: medium/' agents/scaffolder.md +pin high refactorer analyzer commit-changer doc-syncer handover-doc-writer +pin xhigh plan-challenger plugin-advisor verifier security-auditor seo-analyzer geo-analyzer +sed -i 's/(pin sonnet, effort high — BDR-077/(pin sonnet, effort medium — BDR-077/' skills/init-project/SKILL.md +grep -c '^effort:' agents/*.md | grep -v ':0' | grep -v impeccable | wc -l +``` +Expected: `20`. + +- [ ] **Step 3: Suites** + +Run: `make test suite=lib/tests/effort-routing.test.sh && make test suite=lib/tests/model-routing.test.sh` +Expected: both green (the model locks read `model: sonnet` on line 5, untouched). + +- [ ] **Step 4: Smoke, planted input, disk-verified** + +From an interactive session in this repo, dispatch: +``` +Agent(subagent_type="release-executor", description="effort pin smoke", + prompt="Diagnostic only, no release work. Run exactly one bash command and report its raw output: echo CLAUDE_EFFORT=$CLAUDE_EFFORT") +``` +Expected report: `CLAUDE_EFFORT=low`. Then read the subagent transcript: +```bash +f=$(ls -t ~/.claude/projects/-home-bchanot-Documents-claude/*/subagents/*.jsonl | head -1) +grep -o '"effort":"[a-z]*"' "$f" | sort | uniq -c +``` +Expected: only `"effort":"low"`. + +- [ ] **Step 5: Commit** + +```bash +git add agents/*.md skills/init-project/SKILL.md lib/tests/effort-routing.test.sh +git commit -m "feat(effort): pin effort on the 20 repo-authored agents (BDR-077 second axis)" +``` + +--- + +### Task 4: Skill entry levels (spec D3) with a before/after measurement + +**Files:** +- Modify: 31 files `skills//SKILL.md` (line 2 is `name: ` in every one of them; the two superpowers files are Task 9) +- Test: `lib/tests/effort-routing.test.sh` + +**Interfaces:** +- Produces: `effort: ` on line 3 of each listed skill; the `lvl` helper reused by Task 9. + +- [ ] **Step 1: Baseline measurement BEFORE the change (spec §9)** + +```bash +S=/tmp/claude-1000/-home-bchanot-Documents-claude/e593bc78-da6b-469b-9d0c-08d1a4aa8373/scratchpad +mkdir -p "$S"; cd ~/Documents/claude +claude -p "/reconcile" --output-format json --allowedTools "Read" "Grep" "Glob" "Bash(git status:*)" "Bash(git log:*)" > "$S/ab-before.json" 2>/dev/null +python3 - "$S/ab-before.json" <<'PY' +import json,sys,glob,os +d=json.load(open(sys.argv[1])); sid=d["session_id"] +P=os.path.expanduser("~/.claude/projects/-home-bchanot-Documents-claude") +n=o=t=0; eff=set() +for line in open(f"{P}/{sid}.jsonl", errors="ignore"): + r=json.loads(line) + if r.get("type")!="assistant": continue + u=r["message"].get("usage") or {}; n+=1; o+=u.get("output_tokens",0) + t+=(u.get("output_tokens_details") or {}).get("thinking_tokens",0); eff.add(r.get("effort")) +print(f"BEFORE requests={n} output={o} thinking={t} effort={eff} duration_ms={d.get('duration_ms')}") +PY +``` +Expected: one line, `effort={'high'}` (session default after Task 2). Keep the line for Task 10. + +- [ ] **Step 2: Add the 31 locks to the census** + +```bash +python3 - <<'PY' +p="lib/tests/effort-routing.test.sh"; s=open(p).read() +locks="""# ── 4) skill entry levels (spec D3): the user's invocation sets the run's level +for s in status commit-change release-candidate doc capitalize close reconcile deploy profile plugin-check; do fm_has_effort "skills/$s/SKILL.md" low; done +for s in gitflow prune-memory find-docs; do fm_has_effort "skills/$s/SKILL.md" medium; done +for s in feat hotfix bugfix refactor web-validate harden seo geo; do fm_has_effort "skills/$s/SKILL.md" high; done +for s in ship-feature init-project onboard tour audit-delta analyze code-clean client-handover spec skillify; do fm_has_effort "skills/$s/SKILL.md" xhigh; done + +# ── summary""" +open(p,"w").write(s.replace("# ── summary", locks)) +PY +make test suite=lib/tests/effort-routing.test.sh | tail -1 +``` +Expected: `... 31 fail`. + +- [ ] **Step 3: Apply the levels** + +```bash +lvl() { L=$1; shift; for s in "$@"; do + sed -i "0,/^name: $s\$/s//&\neffort: $L/" "skills/$s/SKILL.md"; done; } +lvl low status commit-change release-candidate doc capitalize close reconcile deploy profile plugin-check +lvl medium gitflow prune-memory find-docs +lvl high feat hotfix bugfix refactor web-validate harden seo geo +lvl xhigh ship-feature init-project onboard tour audit-delta analyze code-clean client-handover spec skillify +grep -l '^effort:' skills/*/SKILL.md | wc -l +``` +Expected: `31`. + +- [ ] **Step 4: Suites** + +Run: `make test suite=lib/tests/effort-routing.test.sh && make test suite=lib/tests/skill-routing-census.test.sh` +Expected: both green. + +- [ ] **Step 5: Measurement AFTER, same command as Step 1 with `ab-after.json` and the label `AFTER`** + +Expected: `effort={'low'}`. Record both lines in the commit body; Task 10 turns them into the EVAL. + +- [ ] **Step 6: Commit** + +```bash +git add skills/*/SKILL.md lib/tests/effort-routing.test.sh +git commit -m "feat(effort): entry effort level on the 31 user-invoked skills (spec D3)" \ + -m "A/B /reconcile headless — " +``` + +--- + +### Task 5: Shifter skills, `lib/effort-shift.md`, model-gate paragraph (spec D4) + +**Files:** +- Create: `skills/effort-low/SKILL.md`, `skills/effort-medium/SKILL.md`, `skills/effort-high/SKILL.md`, `skills/effort-xhigh/SKILL.md`, `skills/effort-max/SKILL.md` +- Create: `lib/effort-shift.md` +- Modify: `lib/model-gate.md` (end of §4) +- Test: `lib/tests/effort-routing.test.sh`, `lib/tests/skill-routing-census.test.sh` + +**Interfaces:** +- Produces: skill names `effort-low`, `effort-medium`, `effort-high`, `effort-xhigh`, `effort-max`; the include path `$HOME/.claude/lib/effort-shift.md`; the wiring vocabulary Tasks 6-8 insert: `Skill(effort-)` lines with a trailing `# effort-shift: ` comment. + +- [ ] **Step 1: Locks (above the summary block)** + +```bash +python3 - <<'PY' +p="lib/tests/effort-routing.test.sh"; s=open(p).read() +locks="""# ── 5) shifter skills + include (spec D4) +for l in low medium high xhigh max; do fm_has_effort "skills/effort-$l/SKILL.md" "$l"; has "skills/effort-$l/SKILL.md" "name: effort-$l"; done +has "lib/effort-shift.md" 'Headless sessions' +has "lib/effort-shift.md" 'Skill(effort-max)' +has "lib/effort-shift.md" 'never inside a dispatched agent' +has "lib/model-gate.md" 'lib/effort-shift.md' + +# ── summary""" +open(p,"w").write(s.replace("# ── summary", locks)) +PY +``` + +- [ ] **Step 2: Create the five shifters (descriptions pre-validated against the routing census, pairwise ≤ 0.03)** + +```bash +mk() { mkdir -p "skills/effort-$1"; printf '%s\n' '---' "name: effort-$1" "description: $2" "effort: $1" '---' \ + "Effort shifted to $1 for the rest of this turn (lib/effort-shift.md). Continue with the caller's next step." \ + > "skills/effort-$1/SKILL.md"; } +mk low "Bookkeeping shift. Lowers reasoning to the cheapest level for the rest of the turn: journal lines, memory commits, capitalize, release bookkeeping, status output." +mk medium "Orchestration shift. Standard reasoning between two dispatches: read a subagent report, pick the next step, relay a gate verdict, route a branch." +mk high "Investigation shift. Deeper reasoning for diagnosis, LOCATE, contract drafting, refactor judgement inside feat, hotfix and bugfix runs." +mk xhigh "Reflection shift. Deep reasoning for brainstorm, planning, challenge synthesis and audit verdicts before a human validation gate." +mk max "Escalation shift. Maximum reasoning when a verify or security loop hits its cap, a gate fails twice, or error recovery starts in ship-feature." +``` + +- [ ] **Step 3: Write the include** + +```bash +cat > lib/effort-shift.md <<'EOF' +# Effort shift — phase-level reasoning effort on the main loop (BDR-NEXT) + +Shared include, companion of `lib/model-gate.md`: the gate fixes WHICH model +reflects, this include fixes HOW HARD each phase thinks. The rungs are the +user's: low (fix a line, run a script) · medium (day-to-day) · high +(refactor, resisting bug) · xhigh (architecture, audit before validation) · +max (stuck error, judged need). + +## Mechanics (verified on Claude Code 2.1.283) + +- A skill's `effort:` frontmatter applies from the moment it loads to the + end of the turn: on the user's `/skill` and on a `Skill(...)` call by + Claude in an interactive session. Last loaded wins, both directions. The + prompt cache survives a shift. +- Dispatched agents run on their own `effort:` pin, never on a shift. + Unpinned agents inherit the level in force at dispatch. +- Headless sessions (`-p`, `claude agents`, SDK) ignore skill-level effort: + the run stays at the session level. `CLAUDE_CODE_EFFORT_LEVEL` beats every + frontmatter; keep it unset (the session banner warns). + +## Shifters + +`Skill(effort-low)` · `Skill(effort-medium)` · `Skill(effort-high)` · +`Skill(effort-xhigh)` · `Skill(effort-max)`. One tool call, one-line body. +Typed by the user, `/effort-max` is a turn-scoped max: the relaunch lever +after a STOP. `ultrathink` only adds an in-context nudge; the API level +does not move. + +## Wiring — per orchestrator + +1. A dispatch span starts (executor, collector, fan-out) → + `Skill(effort-medium)`. +2. Reflection resumes after a dispatch span (challenge synthesis, verdict, + plan revision) → `Skill(effort-)`. Concretely: + the line before every `lib/challenge-plan.md` call. +3. The bookkeeping tail (memory commit, doc commit) → `Skill(effort-low)`. +4. Escalation → `Skill(effort-max)`, then the skill's own level again once + the diagnosis is produced. Automatic points: verify-secure loop caps + (GATE 0 floor, GATE 1 conformity, GATE 2 security) and ship-feature + STEP 4b. Not automatic, by doctrine: the challenge fail-safe (a mute + challenger is an infrastructure failure) and "gone WRONG → STOP" (STOP + precedes any further reasoning); their STOP text names the level + reached and suggests `/effort-max` for the relaunch. + +## Re-assert + +- After any nested `Skill(...)` whose frontmatter carries a different + effort (feat → commit-change), reload the orchestrator's own level. +- After a prose gate that ends the turn, the resumed turn runs at the + session level. If the resumed phase is reflection, its first step is + `Skill(effort-)`; dispatch and orchestration phases need + nothing. + +## Never + +- A shift never inside a dispatched agent: pins rule there. +- Max is for diagnosis, not for retrying the same fix harder. +EOF +``` + +- [ ] **Step 4: Model-gate paragraph (append to §4)** + +```bash +cat >> lib/model-gate.md <<'EOF' + +Effort is the second axis of the same table (BDR-NEXT): every typed agent +carries an `effort:` pin next to `model:`, and the main loop shifts per phase +through `lib/effort-shift.md`. Nothing dispatched inherits either axis. +EOF +``` + +- [ ] **Step 5: Suites** + +Run: `make test suite=lib/tests/effort-routing.test.sh && make test suite=lib/tests/skill-routing-census.test.sh && make test suite=lib/tests/profile-census.test.sh` +Expected: all green; the routing census prints no `effort-` pair under WARN. + +- [ ] **Step 6: Main-session smoke (interactive session only, cannot be headless)** + +In an interactive session in this repo, after `/reload-skills`: call `Skill(effort-max)`, then Bash `echo $CLAUDE_EFFORT`, then `Skill(effort-xhigh)`, then the echo again. +Expected: `max`, then `xhigh`. Transcript check for the cache: +```bash +f=~/.claude/projects/-home-bchanot-Documents-claude/$(ls -t ~/.claude/projects/-home-bchanot-Documents-claude/ | grep jsonl | head -1) +python3 -c " +import json,sys +for l in open('$f',errors='ignore'): + r=json.loads(l) + if r.get('type')=='assistant': + u=r['message'].get('usage') or {}; print(r.get('effort'), u.get('cache_creation_input_tokens'), u.get('cache_read_input_tokens'))" | tail -6 +``` +Expected: the first `max` row has `cache_creation` in the low thousands and `cache_read` unchanged from the row before (no cache bust). + +- [ ] **Step 7: Commit** + +```bash +git add skills/effort-*/SKILL.md lib/effort-shift.md lib/model-gate.md lib/tests/effort-routing.test.sh +git commit -m "feat(effort): five shifter skills, lib/effort-shift.md, model-gate second axis" +``` + +--- + +### Task 6: Orchestrator wiring (include line, medium at dispatch, own level at challenge, low at the tail, nested re-assert) + +**Files:** +- Modify: `skills/{feat,hotfix,bugfix,ship-feature,init-project,onboard,tour,code-clean,seo,geo,harden,web-validate,audit-delta}/SKILL.md`, `agents/client-handover-writer.md` (the client-handover skill loads this agent inline; its dispatches live there) +- Test: `lib/tests/effort-routing.test.sh` + +**Interfaces:** +- Consumes: shifter names and the include path from Task 5. +- Produces: helpers `ins_before`, `ins_after`, `ins_after_para` (local to this task's shell). + +- [ ] **Step 1: Locks (above the summary block)** + +```bash +python3 - <<'PY' +p="lib/tests/effort-routing.test.sh"; s=open(p).read() +locks="""# ── 6) orchestrator wiring (spec D4) +for s in feat hotfix bugfix ship-feature init-project onboard tour code-clean seo geo harden web-validate audit-delta; do + has "skills/$s/SKILL.md" 'lib/effort-shift.md'; has "skills/$s/SKILL.md" 'Skill(effort-medium)'; done +has "agents/client-handover-writer.md" 'lib/effort-shift.md'; has "agents/client-handover-writer.md" 'Skill(effort-medium)' +for s in feat hotfix bugfix; do has "skills/$s/SKILL.md" 'Skill(effort-high)'; done +for s in ship-feature init-project onboard code-clean audit-delta; do has "skills/$s/SKILL.md" 'Skill(effort-xhigh)'; done +for s in seo geo harden web-validate; do has "skills/$s/SKILL.md" 'Skill(effort-high)'; done +for s in feat hotfix bugfix ship-feature init-project; do has "skills/$s/SKILL.md" 'Skill(effort-low)'; done +has "skills/feat/SKILL.md" 'effort-shift: nested commit-change' + +# ── summary""" +open(p,"w").write(s.replace("# ── summary", locks)) +PY +``` + +- [ ] **Step 2: Define the three insertion helpers (exact-string anchors, first occurrence)** + +```bash +ins_before() { python3 - "$1" "$2" "$3" <<'PY' +import sys; f,a,t=sys.argv[1:]; L=open(f).read().split("\n") +i=next(k for k,l in enumerate(L) if a in l); L[i:i]=t.split("\\n"); open(f,"w").write("\n".join(L)) +PY +} +ins_after() { python3 - "$1" "$2" "$3" <<'PY' +import sys; f,a,t=sys.argv[1:]; L=open(f).read().split("\n") +i=next(k for k,l in enumerate(L) if a in l); L[i+1:i+1]=t.split("\\n"); open(f,"w").write("\n".join(L)) +PY +} +ins_after_para() { python3 - "$1" "$2" "$3" <<'PY' +import sys; f,a,t=sys.argv[1:]; L=open(f).read().split("\n") +i=next(k for k,l in enumerate(L) if a in l) +j=next(k for k in range(i,len(L)) if L[k].strip()=="") +L[j:j]=t.split("\\n"); open(f,"w").write("\n".join(L)) +PY +} +INC='EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation.' +``` +`next(...)` raises `StopIteration` when an anchor is absent: that is the intended failure, fix the anchor rather than the helper. + +- [ ] **Step 3: Include line, after the model-gate paragraph, in all 13 skills and the writer agent** + +```bash +for s in feat hotfix bugfix ship-feature init-project onboard tour code-clean seo geo harden web-validate audit-delta; do + ins_after_para "skills/$s/SKILL.md" 'lib/model-gate.md' "$INC"; done +ins_after_para agents/client-handover-writer.md 'model: "fable"' "$INC" +grep -c 'lib/effort-shift.md' skills/*/SKILL.md agents/client-handover-writer.md | grep -v ':0' | wc -l +``` +Expected: `14`. + +- [ ] **Step 4: Medium at the first executor/collector dispatch (anchors verified in the repo on 2026-09-28)** + +```bash +M='Skill(effort-medium) # effort-shift: dispatch span starts' +ins_before skills/feat/SKILL.md 'Agent(subagent_type="feater")' "$M" +ins_before skills/hotfix/SKILL.md 'Agent(subagent_type="hotfixer")' "$M" +ins_before skills/bugfix/SKILL.md 'Agent(subagent_type="bugfixer")' "$M" +ins_before skills/code-clean/SKILL.md 'Agent(subagent_type="code-cleaner")' "$M" +ins_before skills/seo/SKILL.md 'Agent(subagent_type="seo-analyzer", model="sonnet")' "$M" +ins_before skills/geo/SKILL.md 'Agent(subagent_type="geo-analyzer", model="sonnet")' "$M" +ins_before skills/web-validate/SKILL.md 'Agent(' "$M" +ins_before skills/harden/SKILL.md 'Agent(' "$M" +ins_after skills/ship-feature/SKILL.md '## STEP 4 — IMPLEMENT' "First: \`Skill(effort-medium)\` (effort-shift: dispatch span starts)." +ins_after skills/init-project/SKILL.md '## STEP 8 — IMPLEMENT' "First: \`Skill(effort-medium)\` (effort-shift: dispatch span starts)." +ins_before skills/onboard/SKILL.md 'Agent(subagent_type="onboarder")' "\`Skill(effort-medium)\` first (effort-shift: dispatch span starts)." +ins_after agents/client-handover-writer.md '## STEP 3 — BASELINE AUDITS' "First: \`Skill(effort-medium)\` (effort-shift: dispatch span starts)." +ins_before skills/tour/SKILL.md 'Agent(subagent_type="general-purpose",' "$M" +ins_before skills/audit-delta/SKILL.md 'Agent(subagent_type="security-auditor", description="audit-delta security' "$M" +``` +Anchors verified 2026-09-28: `web-validate` and `harden` open their first dispatch with a bare `Agent(` line (l.181 and l.262), the first `Agent(` in each file; `seo` (l.326) and `geo` (l.48) open the collect dispatch with the full call line, unique as first occurrence. `MODE: collect` is not an anchor: it sits inside the prompt string. + +- [ ] **Step 5: Own level before every challenge-plan call (reflection resumes)** + +```bash +for s in feat hotfix bugfix seo geo harden web-validate; do + ins_before "skills/$s/SKILL.md" 'lib/challenge-plan.md' "\`Skill(effort-high)\` first (effort-shift: reflection resumes)."; done +for s in ship-feature init-project onboard code-clean audit-delta; do + ins_before "skills/$s/SKILL.md" 'lib/challenge-plan.md' "\`Skill(effort-xhigh)\` first (effort-shift: reflection resumes)."; done +``` +`seo`, `geo` and `web-validate` dispatch their applier after the challenge (seo l.557, geo l.117, web-validate l.312; the first `Agent(subagent_type="hotfixer")` in each file), so a second medium shift goes there: +```bash +for s in seo geo web-validate; do ins_before "skills/$s/SKILL.md" 'Agent(subagent_type="hotfixer")' "$M"; done +``` +`harden` applies inline in its STEP 3 (main loop, after the user's confirmation): no applier dispatch, the own-level shift before its challenge line is its last shift. + +- [ ] **Step 6: Low at the bookkeeping tail (the five skills with a memory-commit include)** + +```bash +for s in feat hotfix bugfix ship-feature init-project; do + ins_before "skills/$s/SKILL.md" 'lib/capitalize-commit.md' "\`Skill(effort-low)\` first (effort-shift: bookkeeping tail).\\n"; done +``` + +- [ ] **Step 7: Nested re-assert in feat (commit-change runs at low)** + +```bash +grep -n -E 'commit-change' skills/feat/SKILL.md +``` +Expected: two consecutive prose lines near l.199 (`… or run \`/commit-change\` on the pending work (it dispatches the …`). The sentence continues, so append at the end of that paragraph, not after the line: +```bash +ins_after_para skills/feat/SKILL.md '/commit-change' "Then \`Skill(effort-high)\` (effort-shift: nested commit-change loaded at low; reload feat's level)." +``` + +- [ ] **Step 8: Suites, then a read-through of each edited file around the insertions** + +Run: `make test suite=lib/tests/effort-routing.test.sh && make test suite=lib/tests/model-routing.test.sh && make test suite=lib/tests/loops-light.test.sh` +Expected: all green (the model-routing and loops-light locks match single lines that this task never splits). + +```bash +git diff -U1 -- skills agents | grep -E '^\+' | grep -v '^+++' | wc -l +``` +Expected: about 40 added lines, none inside a YAML frontmatter block (every insertion sits below the second `---`). + +- [ ] **Step 9: Commit** + +```bash +git add skills/*/SKILL.md agents/client-handover-writer.md lib/tests/effort-routing.test.sh +git commit -m "feat(effort): wire phase shifts in the 13 orchestrators and the handover writer" +``` + +--- + +### Task 7: Escalation points at max (loop caps, ship-feature 4b) and STOP texts + +**Files:** +- Modify: `lib/verify-secure-loop.md` (the three `**Max 3 … iterations** → STOP + human escalation` sentences, lines 38, 77, 107 on 2026-09-28) +- Modify: `skills/ship-feature/SKILL.md` (STEP 4b, `1. Load \`$HOME/.claude/agents/analyzer.md\` in DEBUG MODE`, and step 4 `If A →` / `If B →`) +- Modify: `lib/challenge-plan.md` (line `retry ONCE with a fresh challenger; a 2nd failure on that lens → STOP and escalate`) +- Test: `lib/tests/effort-routing.test.sh` + +**Interfaces:** +- Consumes: `Skill(effort-max)` and `/effort-max` from Task 5. + +- [ ] **Step 1: Locks** + +```bash +python3 - <<'PY' +p="lib/tests/effort-routing.test.sh"; s=open(p).read() +locks="""# ── 7) escalation at max (spec D4) +[ "$(grep -c 'Skill(effort-max)' "$R/lib/verify-secure-loop.md")" -eq 3 ] && ok || ko "verify-secure-loop.md must shift to max at its 3 caps" +has "skills/ship-feature/SKILL.md" 'Skill(effort-max)' +has "lib/challenge-plan.md" '/effort-max' +has "lib/verify-secure-loop.md" '/effort-max' + +# ── summary""" +open(p,"w").write(s.replace("# ── summary", locks)) +PY +``` + +- [ ] **Step 2: Loop caps** + +```bash +python3 - <<'PY' +p="lib/verify-secure-loop.md"; s=open(p).read() +for cap in ("floor","conformity","security"): + old=f"**Max 3 {cap} iterations** → STOP + human escalation" + new=(f"**Max 3 {cap} iterations** → `Skill(effort-max)` (effort-shift: cap reached, " + f"diagnose at max before escalating), then STOP + human escalation") + assert s.count(old)==1, cap; s=s.replace(old,new) +s=s.replace("STOP + human escalation with the\n BLOCKING table.", + "STOP + human escalation with the\n BLOCKING table. Every STOP text names the level reached (`$CLAUDE_EFFORT`)\n and suggests `/effort-max` for the relaunch.") +open(p,"w").write(s) +PY +grep -c 'Skill(effort-max)' lib/verify-secure-loop.md; grep -c '/effort-max' lib/verify-secure-loop.md +``` +Expected: `3` and `1`. + +- [ ] **Step 3: ship-feature 4b** + +```bash +python3 - <<'PY' +p="skills/ship-feature/SKILL.md"; s=open(p).read() +old="1. Load `$HOME/.claude/agents/analyzer.md` in DEBUG MODE on the exact error output." +assert s.count(old)==1 +s=s.replace(old, "1. `Skill(effort-max)` (effort-shift: error recovery), then load\n `$HOME/.claude/agents/analyzer.md` in DEBUG MODE on the exact error output.") +old2="4. If A → apply minimal fix, re-run STEP 4 for the failed task only." +assert s.count(old2)==1 +s=s.replace(old2, "4. On resume the turn is at the session level (effort-shift: turn reset).\n If A → `Skill(effort-medium)`, apply minimal fix, re-run STEP 4 for the failed task only.") +old3=" If B → before skipping:" +assert s.count(old3)==1 +s=s.replace(old3, " If B or C → `Skill(effort-xhigh)` first.\n If B → before skipping:") +open(p,"w").write(s) +PY +``` + +- [ ] **Step 4: Challenge fail-safe STOP text** + +```bash +python3 - <<'PY' +p="lib/challenge-plan.md"; s=open(p).read() +old="retry ONCE with a fresh challenger; a 2nd failure on that lens → STOP and escalate" +assert s.count(old)==1 +s=s.replace(old, old+"\n(the STOP text names the level reached, `$CLAUDE_EFFORT`, and suggests `/effort-max`\nfor the relaunch; no shift here: a mute challenger is an infrastructure failure)") +open(p,"w").write(s) +PY +``` + +- [ ] **Step 5: Suites and commit** + +Run: `make test` +Expected: green. + +```bash +git add lib/verify-secure-loop.md lib/challenge-plan.md skills/ship-feature/SKILL.md lib/tests/effort-routing.test.sh +git commit -m "feat(effort): max at the verify-secure caps and ship-feature 4b; STOP texts suggest /effort-max" +``` + +--- + +### Task 8: Turn-ending gate audit and re-assert + +**Files:** +- Modify: `skills/bugfix/SKILL.md` (the gate `behavior change): wait for user approval.` before pass B of the contract interview) +- Possibly modify: any other orchestrator where the audit below finds a prose gate followed by reflection +- Test: `lib/tests/effort-routing.test.sh` + +- [ ] **Step 1: Lock** + +```bash +python3 - <<'PY' +p="lib/tests/effort-routing.test.sh"; s=open(p).read() +locks="""# ── 8) turn-reset re-assert after a prose gate followed by reflection +has "skills/bugfix/SKILL.md" 'effort-shift: turn reset' + +# ── summary""" +open(p,"w").write(s.replace("# ── summary", locks)) +PY +``` + +- [ ] **Step 2: Audit every prose gate** + +```bash +grep -n -i -E "end the turn|end your turn|wait for (the )?(user|human)|STOP and wait|wait for user" \ + skills/{feat,hotfix,bugfix,ship-feature,init-project,onboard,tour,code-clean,seo,geo,harden,web-validate,audit-delta}/SKILL.md \ + lib/contract-interview.md lib/challenge-plan.md lib/plugin-gate.md lib/verify-secure-loop.md +``` +Known on 2026-09-28: `bugfix:119` (resume = contract pass B, reflection → re-assert), `ship-feature:205` (handled in Task 7), `ship-feature:14` and `init-project:14` (model-gate STOP, the run ends → nothing). Classify every other hit the same way: model-gate STOP or loop-cap STOP → nothing; resume into dispatch/orchestration → nothing; resume into reflection → re-assert with the skill's own level. + +- [ ] **Step 3: Re-assert in bugfix** + +```bash +python3 - <<'PY' +p="skills/bugfix/SKILL.md"; s=open(p).read() +old=" behavior change): wait for user approval.\n" +assert s.count(old)==1 +s=s.replace(old, old+" On resume: `Skill(effort-high)` first (effort-shift: turn reset).\n") +open(p,"w").write(s) +PY +``` +Apply the same one-line pattern to any other reflection resume found in Step 2, with that skill's level. + +- [ ] **Step 4: Manual verification of the reset itself (interactive, once)** + +In an interactive session: type `/effort-max`, wait for the reply, then send a plain message such as `echo test` and read the transcript: +```bash +f=~/.claude/projects/-home-bchanot-Documents-claude/$(ls -t ~/.claude/projects/-home-bchanot-Documents-claude/ | grep jsonl | head -1) +grep -o '"effort":"[a-z]*"' "$f" | tail -4 +``` +Expected: `max` rows for the first turn, `high` for the second (the session default from Task 2). + +- [ ] **Step 5: Suites and commit** + +Run: `make test` +```bash +git add skills/*/SKILL.md lib/tests/effort-routing.test.sh +git commit -m "feat(effort): re-assert the skill level after prose gates that end the turn" +``` + +--- + +### Task 9: Vendored superpowers patch with resync re-apply + +**Files:** +- Modify: `skills/brainstorming/SKILL.md`, `skills/writing-plans/SKILL.md` (line 2 `name: …`) +- Modify: `install-plugins.sh` (end of the STEP 8e block that vendors the 7 superpowers skills) +- Test: `lib/tests/effort-routing.test.sh` + +- [ ] **Step 1: Locks** + +```bash +python3 - <<'PY' +p="lib/tests/effort-routing.test.sh"; s=open(p).read() +locks="""# ── 9) vendored superpowers carry xhigh; a resync that drops it fails here (spec D3) +for s in brainstorming writing-plans; do fm_has_effort "skills/$s/SKILL.md" xhigh; done +has "install-plugins.sh" 'effort: xhigh' + +# ── summary""" +open(p,"w").write(s.replace("# ── summary", locks)) +PY +``` + +- [ ] **Step 2: Patch the two files (same `lvl` helper as Task 4)** + +```bash +lvl() { L=$1; shift; for s in "$@"; do + sed -i "0,/^name: $s\$/s//&\neffort: $L/" "skills/$s/SKILL.md"; done; } +lvl xhigh brainstorming writing-plans +sed -n 1,4p skills/brainstorming/SKILL.md skills/writing-plans/SKILL.md +``` +Expected: `effort: xhigh` on line 3 of both. + +- [ ] **Step 3: Re-apply after every resync in install-plugins.sh** + +```bash +grep -n -i 'STEP 8e' install-plugins.sh +``` +Expected: three hits on 2026-09-28: a cross-reference comment near l.535, the heading `# ── Step 8e: Agent Skills …` near l.908, and its `echo` near l.915. The block ends where the `# ====` banner of STEP 8.5 begins (near l.937). Insert the re-apply right before that banner: +```bash +python3 - <<'PY' +p="install-plugins.sh"; L=open(p).read().split("\n") +i=next(k for k,l in enumerate(L) if l.startswith("# ── Step 8e:")) +j=next(k for k in range(i+1,len(L)) if L[k].startswith("# ====")) # the STEP 8.5 banner +L[j:j]=[ +"# Effort tiering (BDR-NEXT): the vendored brainstorming/writing-plans carry an", +"# effort pin upstream lacks; re-apply after every resync (census lock in", +"# lib/tests/effort-routing.test.sh alarms if this ever stops working).", +"for _s in brainstorming writing-plans; do", +" _f=\"$(cd \"$(dirname \"$0\")\" && pwd)/skills/$_s/SKILL.md\"", +" if [ -f \"$_f\" ] && ! grep -q '^effort:' \"$_f\"; then", +" sed -i \"0,/^name: $_s\\$/s//&\\neffort: xhigh/\" \"$_f\"", +" fi", +"done", +"unset _s _f", +"", +] +open(p,"w").write("\n".join(L)) +PY +shellcheck install-plugins.sh +``` +Expected: shellcheck silent, and `sed -n '/^unset _s _f/,+2p' install-plugins.sh` shows the blank line then the `# ====` banner of STEP 8.5. + +- [ ] **Step 4: Prove the re-apply works** + +```bash +sed -i '/^effort: xhigh$/d' skills/brainstorming/SKILL.md +bash -c 'source /dev/stdin <<<"$(sed -n "/Effort tiering (BDR-NEXT)/,/^unset _s _f/p" install-plugins.sh)"' +grep -c '^effort: xhigh' skills/brainstorming/SKILL.md +``` +Expected: `1` (the extracted block re-added the line without running the whole installer). + +- [ ] **Step 5: Suites and commit** + +Run: `make test` +```bash +git add skills/brainstorming/SKILL.md skills/writing-plans/SKILL.md install-plugins.sh lib/tests/effort-routing.test.sh +git commit -m "feat(effort): xhigh on the vendored brainstorming and writing-plans, re-applied at resync" +``` + +--- + +### Task 10: BDR id, CHANGELOG, registries, journal, TODO reconcile + +**Files:** +- Modify: `lib/effort-shift.md`, `lib/model-gate.md`, `lib/tests/effort-routing.test.sh`, `hooks/session-start.sh`, `hooks/statusline.sh`, `install-plugins.sh`, 14 orchestrator files (every `BDR-NEXT` token) +- Modify: `CHANGELOG.md` (`## [Unreleased]` → `### Added`), `.claude/memory/decisions.md`, `.claude/memory/evals.md`, `.claude/memory/journal.md`, `.claude/tasks/TODO.md` + +- [ ] **Step 1: Compute the id and replace the token everywhere** + +```bash +N=$(( $(grep -o -E 'BDR-[0-9]+' .claude/memory/decisions.md | sort -t- -k2 -n | tail -1 | cut -d- -f2) + 1 )) +echo "BDR-$N" +grep -rl 'BDR-NEXT' --include='*.md' --include='*.sh' --include='*.json' . | grep -v '^./docs/superpowers/' | xargs sed -i "s/BDR-NEXT/BDR-$N/g" +grep -rn 'BDR-NEXT' . --include='*.md' --include='*.sh' | grep -v '^./docs/superpowers/' | wc -l +``` +Expected: `0` (the spec and this plan keep the token as history). + +- [ ] **Step 2: CHANGELOG under `## [Unreleased]` → `### Added` (first bullet position)** + +```bash +python3 - <<'PY' +p="CHANGELOG.md"; s=open(p).read() +anchor="## [Unreleased]\n\n### Added\n" +assert s.count(anchor)==1 +entry=("- **Effort tiering (BDR-$N)**: reasoning effort routed per role and per phase. " +"Session default `high`; `effort:` pins on the 20 repo-authored agents; entry level on " +"the 33 user-invoked skills (low → xhigh); five shifter skills `effort-low` … `effort-max` " +"loaded at phase boundaries through `lib/effort-shift.md`, with `max` at the verify-secure " +"caps and ship-feature 4b; `/effort-max` as the turn-scoped relaunch lever; statusline shows " +"the live level; session banner warns when `CLAUDE_CODE_EFFORT_LEVEL` silences the pins; " +"census `lib/tests/effort-routing.test.sh`.\n") +open(p,"w").write(s.replace(anchor, anchor+entry)) +PY +sed -i "s/BDR-\$N/BDR-$N/" CHANGELOG.md +``` + +- [ ] **Step 3: BDR entry (index row after the last row, section at the end), caveman English** + +Index row (columns `| ID | Date | Decision | Status |` — copy the exact header of the table in `decisions.md` and match it): +``` +| BDR- | 2026-09-28 | Effort tiering: session high, agent effort pins (BDR-077 second axis), skill entry levels, five shifter skills for phase shifts, max at loop caps + 4b | accepted | +``` +Section: +``` +## BDR- — Effort tiering: session high, pins, skill levels, phase shifts, max at escalation [accepted] (2026-09-28) +- **Decision**: settings `effortLevel` high; `effort:` pin on 20 repo-authored agents by role (low appliers, medium executors, high judgment on sonnet/opus, xhigh challengers + gates); `effort:` on 33 user-invoked skills = run entry level; `skills/effort-{low,medium,high,xhigh,max}` loaded by orchestrators at phase boundaries per `lib/effort-shift.md` (medium at dispatch, own level before challenge synthesis, low at bookkeeping tail, max at verify-secure caps + ship-feature 4b); STOP texts suggest `/effort-max`; statusline live level; banner warns on `CLAUDE_CODE_EFFORT_LEVEL`; census `lib/tests/effort-routing.test.sh`. +- **Why**: session-wide xhigh burned thinking on bookkeeping; measurement (EVAL-035) put 97 % of thinking in the main loop, so the main-loop lever (skill effort, verified LRN-179) carries the savings; pins = explicitness + future models. +- **Alternatives rejected**: executor pins only (executors think 26 tok/request); escalation-diagnoser agent fable+max (no context, one more agent, main-loop max keeps the failure context); reflection in fable skill-runner children with session medium (loses interactivity/context); settings.json rewrite mid-run (global side effect, LRN-098); `maxEffortLevel` caps (hide a mis-pin the census should fail). +- **Caveats**: shifts inert in `-p`/SDK; a turn-ending prose gate resets to session level (re-assert wired where reflection resumes); one effort per agent file → mode-based agents pin their judgment mode; vendored superpowers patch re-applied by install-plugins STEP 8e. +- **Refs**: spec `docs/superpowers/specs/2026-09-28-effort-tiering-design.md`, plan `docs/superpowers/plans/2026-09-28-effort-tiering.md`, [[LRN-179]], [[EVAL-035]], [[BDR-077]]. +``` +Replace `` by the computed id. Insert the row after the last `| BDR-` row with the same python pattern as Task 4's lock insertion; append the section at the end of the file. + +- [ ] **Step 4: EVAL row + section for the Task 4 A/B (columns `| ID | Date | Output | Action |`)** + +``` +| EVAL-036 | 2026-09-28 | A/B `/reconcile` headless, session high vs skill low (Task 4): requests , output , thinking , ms | keep low on bookkeeping skills; repeat on a reflection skill before touching the medium/high split | +``` +Section with `- **Date**`, `- **Method**` (the Task 4 Step 1 command), `- **Result**` (the two lines), `- **Anomaly**` (anything odd: for example thinking near zero in both runs means effort did not matter for that skill), `- **Action**`. Use the next free EVAL id (`grep -o -E 'EVAL-[0-9]+' .claude/memory/evals.md | sort -t- -k2 -n | tail -1`). + +- [ ] **Step 5: Journal line and TODO reconcile** + +Append under today's heading in `.claude/memory/journal.md` (create the `## 2026-09-28` heading if absent): `- effort tiering shipped on feature/effort-tiering: session high, 20 pins, 33 skill levels, 5 shifters, max at caps + 4b; census green; finish awaits human signal.` +In `.claude/tasks/TODO.md`, tick the four wave checkboxes of the `effort tiering` section. + +- [ ] **Step 6: Suites, then commit code and docs, then the memory surgically** + +Run: `make test && shellcheck *.sh hooks/*.sh lib/*.sh` +```bash +git add CHANGELOG.md lib hooks install-plugins.sh skills agents +git commit -m "docs(effort): BDR-$N id, CHANGELOG entry" +bash lib/memory-commit.sh commit "chore(memory): BDR-$N effort tiering, EVAL A/B, journal, TODO" +git status --short +``` +Expected: clean tree, both commits pushed by the post-commit hook. + +--- + +### Task 11: Keep the transcript audit script (spec §9 tooling) + +**Files:** +- Create: `lib/effort-audit.py` (from the spike's `effort_split2.py`, cleaned: functions ≤ 25 logic lines, 80-char lines, no globals beyond constants) +- Modify: `lib/effort-shift.md` (one line under Mechanics: `Measure with python3 ~/.claude/lib/effort-audit.py [projects-root]`) +- Test: `lib/tests/effort-routing.test.sh` + +- [ ] **Step 1: Lock** + +```bash +python3 - <<'PY' +p="lib/tests/effort-routing.test.sh"; s=open(p).read() +locks="""# ── 11) audit tooling +has "lib/effort-shift.md" 'effort-audit.py' +[ -x "$R/lib/effort-audit.py" ] && ok || ko "lib/effort-audit.py missing or not executable" + +# ── summary""" +open(p,"w").write(s.replace("# ── summary", locks)) +PY +``` + +- [ ] **Step 2: Write the script** + +```bash +cat > lib/effort-audit.py <<'EOF' +#!/usr/bin/env python3 +"""Sum output/thinking/cache tokens per (scope, model, effort) over Claude Code +transcripts. scope = main (session jsonl) | sub (subagents/*.jsonl or +isSidechain records). Read-only. Usage: effort-audit.py [projects-root]""" +import collections +import glob +import json +import os +import sys + +WEIGHTS = {"in": 1.0, "cc": 1.25, "cr": 0.1, "out": 5.0} # relative to input price +FIELDS = ("in", "cc", "cr", "out", "think") + + +def usage_row(usage): + """Map one API usage block to the five counted fields.""" + details = usage.get("output_tokens_details") or {} + return { + "in": usage.get("input_tokens", 0) or 0, + "cc": usage.get("cache_creation_input_tokens", 0) or 0, + "cr": usage.get("cache_read_input_tokens", 0) or 0, + "out": usage.get("output_tokens", 0) or 0, + "think": details.get("thinking_tokens", 0) or 0, + } + + +def scan(path, scope, agg): + """Add every assistant record of one transcript to agg.""" + with open(path, errors="ignore") as handle: + for line in handle: + try: + rec = json.loads(line) + except ValueError: + continue + msg = rec.get("message") or {} + if rec.get("type") != "assistant" or not msg.get("usage"): + continue + sub = scope == "sub" or bool(rec.get("isSidechain")) + key = ("sub" if sub else "main", + str(msg.get("model", "?")).replace("claude-", ""), + str(rec.get("effort") or "?")) + row = usage_row(msg["usage"]) + agg[key]["msgs"] += 1 + for field in FIELDS: + agg[key][field] += row[field] + + +def weighted(counter): + return sum(counter[f] * WEIGHTS[f] for f in WEIGHTS) + + +def report(agg): + """Print the per-key table, then the main/sub split and the thinking share.""" + total = collections.Counter() + for counter in agg.values(): + total.update(counter) + total_w = weighted(total) or 1 + print(f"{'scope':5} {'model':22} {'effort':7} {'msgs':>6} {'think/msg':>9} " + f"{'think_tok':>10} {'out_tok':>10} {'cache_read':>12} {'%wcost':>7}") + for (scope, model, effort), c in sorted(agg.items(), key=lambda kv: -weighted(kv[1])): + per_msg = c["think"] / max(c["msgs"], 1) + print(f"{scope:5} {model:22} {effort:7} {c['msgs']:6d} {per_msg:9.0f} " + f"{c['think']:10d} {c['out']:10d} {c['cr']:12d} {100 * weighted(c) / total_w:6.1f}%") + by_scope = collections.defaultdict(collections.Counter) + for (scope, _, _), c in agg.items(): + by_scope[scope].update(c) + for scope, c in by_scope.items(): + print(f" {scope:5} weighted-cost {100 * weighted(c) / total_w:5.1f}% " + f"thinking {100 * c['think'] / max(total['think'], 1):5.1f}% requests {c['msgs']}") + print(f" thinking = {100 * total['think'] * WEIGHTS['out'] / total_w:.1f}% of weighted cost; " + f"cache reads = {100 * total['cr'] * WEIGHTS['cr'] / total_w:.1f}%") + + +def main(): + root = os.path.expanduser(sys.argv[1] if len(sys.argv) > 1 else "~/.claude/projects") + agg = collections.defaultdict(collections.Counter) + for project in sorted(glob.glob(os.path.join(root, "*"))): + if not os.path.isdir(project): + continue + for path in glob.glob(os.path.join(project, "*.jsonl")): + scan(path, "main", agg) + for path in glob.glob(os.path.join(project, "*", "subagents", "*.jsonl")): + scan(path, "sub", agg) + report(agg) + + +if __name__ == "__main__": + main() +EOF +chmod +x lib/effort-audit.py +python3 lib/effort-audit.py | head -5 +``` +Expected: the table header and the top rows, `main fable-5-1` first. + +- [ ] **Step 3: Pointer in the include, suites, commit** + +```bash +python3 - <<'PY' +p="lib/effort-shift.md"; s=open(p).read() +anchor="## Shifters\n" +assert s.count(anchor)==1 +s=s.replace(anchor, "Measure the split any time: `python3 ~/.claude/lib/effort-audit.py`\n(thinking/output/cache tokens per scope, model and effort).\n\n"+anchor) +open(p,"w").write(s) +PY +make test suite=lib/tests/effort-routing.test.sh +git add lib/effort-audit.py lib/effort-shift.md lib/tests/effort-routing.test.sh +git commit -m "feat(effort): transcript audit script for the thinking/cost split" +``` + +--- + +## Self-review against the spec + +- **§4 D1** → Task 2 (settings, banner, statusline). **D2** → Task 3. **D3** → Tasks 4 and 9. **D4** → Tasks 5, 6, 7, 8. **D5** → Task 2. **§6** every file listed has a task. **§7** every census item has a lock: 1 (Task 3 `fm_has_effort`/`fm_no_effort`), 2 (Task 3), 3 (Tasks 4, 9), 4 (Task 5), 5 (Tasks 6, 7), 6 (Task 1), 7 and 8 run as existing suites in `make test`. **§8** waves = Tasks 1-3 / 4, 9 / 5-8 / 10-11. **§9** → Task 4 Steps 1 and 5, EVAL in Task 10, tooling in Task 11. +- **Placeholders**: `BDR-NEXT` is a defined token with a defined replacement step (Task 10); the three `grep -n -m1` recipes in Task 6 Step 4 and Task 8 Step 2 name the expected match and the exact insertion to make. +- **Names**: `Skill(effort-)`, `lib/effort-shift.md`, `fm_has_effort`, `ins_before`/`ins_after`/`ins_after_para`, `lvl`, `pin` are spelled identically across tasks. +- **Review Focus**: 1 → Task 5 lock `Headless sessions`; 2 → Task 2 Step 6; 3 → Task 6 Step 7 + lock; 4 → Task 3 `fm_no_effort status-reporter`; 5 → Task 9 Steps 3-4 + lock. From bb28ecefa22433172f4fafbef6a67e28418451dc Mon Sep 17 00:00:00 2001 From: bastien Date: Mon, 28 Sep 2026 18:47:03 +0200 Subject: [PATCH 04/23] docs(plan): ins_before_para helper for prose anchors (SDD preflight ruling) --- .../plans/2026-09-28-effort-tiering.md | 20 ++++++++++++++----- 1 file changed, 15 insertions(+), 5 deletions(-) diff --git a/docs/superpowers/plans/2026-09-28-effort-tiering.md b/docs/superpowers/plans/2026-09-28-effort-tiering.md index edfbfe2..bded7b8 100644 --- a/docs/superpowers/plans/2026-09-28-effort-tiering.md +++ b/docs/superpowers/plans/2026-09-28-effort-tiering.md @@ -523,7 +523,7 @@ git commit -m "feat(effort): five shifter skills, lib/effort-shift.md, model-gat **Interfaces:** - Consumes: shifter names and the include path from Task 5. -- Produces: helpers `ins_before`, `ins_after`, `ins_after_para` (local to this task's shell). +- Produces: helpers `ins_before`, `ins_after`, `ins_after_para`, `ins_before_para` (local to this task's shell). `ins_before` is for anchors inside code blocks (a standalone `Agent(` line); the `_para` forms are for anchors inside prose, where a bare insertion would split a sentence. - [ ] **Step 1: Locks (above the summary block)** @@ -565,6 +565,13 @@ j=next(k for k in range(i,len(L)) if L[k].strip()=="") L[j:j]=t.split("\\n"); open(f,"w").write("\n".join(L)) PY } +ins_before_para() { python3 - "$1" "$2" "$3" <<'PY' +import sys; f,a,t=sys.argv[1:]; L=open(f).read().split("\n") +i=next(k for k,l in enumerate(L) if a in l) +j=next(k for k in range(i,-1,-1) if L[k].strip()=="")+1 # first line of the paragraph +L[j:j]=t.split("\\n"); open(f,"w").write("\n".join(L)) +PY +} INC='EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation.' ``` `next(...)` raises `StopIteration` when an anchor is absent: that is the intended failure, fix the anchor rather than the helper. @@ -593,7 +600,7 @@ ins_before skills/web-validate/SKILL.md 'Agent(' "$M" ins_before skills/harden/SKILL.md 'Agent(' "$M" ins_after skills/ship-feature/SKILL.md '## STEP 4 — IMPLEMENT' "First: \`Skill(effort-medium)\` (effort-shift: dispatch span starts)." ins_after skills/init-project/SKILL.md '## STEP 8 — IMPLEMENT' "First: \`Skill(effort-medium)\` (effort-shift: dispatch span starts)." -ins_before skills/onboard/SKILL.md 'Agent(subagent_type="onboarder")' "\`Skill(effort-medium)\` first (effort-shift: dispatch span starts)." +ins_before_para skills/onboard/SKILL.md 'Agent(subagent_type="onboarder")' "\`Skill(effort-medium)\` first (effort-shift: dispatch span starts)." ins_after agents/client-handover-writer.md '## STEP 3 — BASELINE AUDITS' "First: \`Skill(effort-medium)\` (effort-shift: dispatch span starts)." ins_before skills/tour/SKILL.md 'Agent(subagent_type="general-purpose",' "$M" ins_before skills/audit-delta/SKILL.md 'Agent(subagent_type="security-auditor", description="audit-delta security' "$M" @@ -604,9 +611,12 @@ Anchors verified 2026-09-28: `web-validate` and `harden` open their first dispat ```bash for s in feat hotfix bugfix seo geo harden web-validate; do - ins_before "skills/$s/SKILL.md" 'lib/challenge-plan.md' "\`Skill(effort-high)\` first (effort-shift: reflection resumes)."; done + ins_before_para "skills/$s/SKILL.md" 'lib/challenge-plan.md' "\`Skill(effort-high)\` first (effort-shift: reflection resumes)."; done for s in ship-feature init-project onboard code-clean audit-delta; do - ins_before "skills/$s/SKILL.md" 'lib/challenge-plan.md' "\`Skill(effort-xhigh)\` first (effort-shift: reflection resumes)."; done + ins_before_para "skills/$s/SKILL.md" 'lib/challenge-plan.md' "\`Skill(effort-xhigh)\` first (effort-shift: reflection resumes)."; done +``` +The challenge include is referenced mid-sentence in every skill (`… harden it. Run\n\`$HOME/.claude/lib/challenge-plan.md\` with …`), hence the paragraph form. +```bash ``` `seo`, `geo` and `web-validate` dispatch their applier after the challenge (seo l.557, geo l.117, web-validate l.312; the first `Agent(subagent_type="hotfixer")` in each file), so a second medium shift goes there: ```bash @@ -618,7 +628,7 @@ for s in seo geo web-validate; do ins_before "skills/$s/SKILL.md" 'Agent(subagen ```bash for s in feat hotfix bugfix ship-feature init-project; do - ins_before "skills/$s/SKILL.md" 'lib/capitalize-commit.md' "\`Skill(effort-low)\` first (effort-shift: bookkeeping tail).\\n"; done + ins_before_para "skills/$s/SKILL.md" 'lib/capitalize-commit.md' "\`Skill(effort-low)\` first (effort-shift: bookkeeping tail).\\n"; done ``` - [ ] **Step 7: Nested re-assert in feat (commit-change runs at low)** From 5437638437175607644e012f052ce44106c1b096 Mon Sep 17 00:00:00 2001 From: bastien Date: Mon, 28 Sep 2026 18:49:40 +0200 Subject: [PATCH 05/23] test(effort): census suite skeleton with flip-test and settings lock --- lib/tests/effort-routing.test.sh | 34 ++++++++++++++++++++++++++++++++ 1 file changed, 34 insertions(+) create mode 100755 lib/tests/effort-routing.test.sh diff --git a/lib/tests/effort-routing.test.sh b/lib/tests/effort-routing.test.sh new file mode 100755 index 0000000..f3f5504 --- /dev/null +++ b/lib/tests/effort-routing.test.sh @@ -0,0 +1,34 @@ +#!/usr/bin/env bash +# lib/tests/effort-routing.test.sh — census: effort tiering (BDR-NEXT) +# agent pins, skill entry levels, shifter skills, orchestrator wiring, settings. +# shellcheck disable=SC2015 # A && ok || ko is deliberate here: ok/ko never fail, so C never masks a true A +set -u +R="$(cd "$(dirname "$0")/../.." && pwd)" +pass=0; fail=0 +ok() { pass=$((pass+1)); } +ko() { fail=$((fail+1)); printf 'FAIL %s\n' "$1"; } +has() { if grep -qF "$2" "$R/$1"; then ok; else ko "$1 missing: $2"; fi; } +lacks() { if grep -qF "$2" "$R/$1"; then ko "$1 must NOT contain: $2"; else ok; fi; } +# frontmatter = the lines between the first two '---' lines +fm() { awk 'NR==1&&/^---$/{p=1;next} p&&/^---$/{exit} p' "$1"; } +fm_effort() { fm "$1" | grep -E '^effort: (low|medium|high|xhigh|max)$' | head -1 | cut -d' ' -f2; } +fm_has_effort() { + got="$(fm_effort "$R/$1")" + if [ "$got" = "$2" ]; then ok; else ko "$1 frontmatter effort must be '$2', got '${got:-none}'"; fi +} +fm_no_effort() { if fm "$R/$1" | grep -q '^effort:'; then ko "$1 must NOT pin effort"; else ok; fi; } + +# ── flip-test: the frontmatter reader must accept a valid level and reject an invalid one +FIX="$(mktemp -d)"; trap 'rm -rf "$FIX"' EXIT +printf -- '---\nname: good\neffort: xhigh\n---\nbody with effort: low in prose\n' > "$FIX/good.md" +printf -- '---\nname: bad\neffort: turbo\n---\n' > "$FIX/bad.md" +[ "$(fm_effort "$FIX/good.md")" = "xhigh" ] && ok || ko "flip: valid level not read" +[ -z "$(fm_effort "$FIX/bad.md")" ] && ok || ko "flip: invalid level accepted" +[ "$(fm "$FIX/good.md" | grep -c 'prose')" -eq 0 ] && ok || ko "flip: body leaked into frontmatter" + +# ── 1) session default (spec D1) +has "settings.json" '"effortLevel": "high"' + +# ── summary (later tasks insert their locks ABOVE this line) +printf 'effort-routing census: %d pass, %d fail\n' "$pass" "$fail" +[ "$fail" -eq 0 ] From 45ae0d12170cd7a03fe0ea5c0c30131b30ae246f Mon Sep 17 00:00:00 2001 From: bastien Date: Mon, 28 Sep 2026 18:52:36 +0200 Subject: [PATCH 06/23] feat(effort): session default high, env-var warning, live effort in statusline --- hooks/session-start.sh | 7 +++++++ hooks/statusline.sh | 11 ++++++----- lib/tests/effort-routing.test.sh | 4 ++++ settings.json | 2 +- 4 files changed, 18 insertions(+), 6 deletions(-) diff --git a/hooks/session-start.sh b/hooks/session-start.sh index 3ab0237..90ba97c 100644 --- a/hooks/session-start.sh +++ b/hooks/session-start.sh @@ -107,6 +107,12 @@ fi REPO_DIR="${_repo_dir:-}" unset _claude_real _repo_dir +# Effort tiering (BDR-NEXT): this env var beats every skill/agent `effort:` pin. +EFFORT_WARN="" +if [ -n "${CLAUDE_CODE_EFFORT_LEVEL:-}" ]; then + EFFORT_WARN="⚠️ CLAUDE_CODE_EFFORT_LEVEL=${CLAUDE_CODE_EFFORT_LEVEL} set: skill/agent effort pins ignored" +fi + # Detect plan and set passive token budget PLAN=$(detect_plan 2>/dev/null || echo "pro") case "$PLAN" in @@ -253,5 +259,6 @@ unset _remote_ver REPO_DIR echo "│ 💡 /plugin-check before starting a new project │" echo "│ 🩺 make doctor full diagnostic │" echo "└───────────────────────────────────────────────────┘" +[ -n "$EFFORT_WARN" ] && printf '%s\n' "$EFFORT_WARN" echo "" unset TOKEN_WARN diff --git a/hooks/statusline.sh b/hooks/statusline.sh index 5e76c84..c9e2332 100755 --- a/hooks/statusline.sh +++ b/hooks/statusline.sh @@ -33,13 +33,14 @@ if [ -z "$PROFILE" ] || [ "$PROFILE" = "none" ]; then PROFILE="$DEFAULT_PROFILE" fi -# Effort level from settings.json (.effortLevel — set by /effort or manual edit). -# settings.json is the source-of-truth, symlinked into ~/.claude/settings.json. -EFFORT="?" -if [ -f "$REPO/settings.json" ]; then +# Effort level: the live value when the harness exports it (skill/agent +# `effort:` shifts included, BDR-NEXT), else the persisted settings.json key +# (.effortLevel — set by /effort or manual edit; symlinked into ~/.claude). +EFFORT="${CLAUDE_EFFORT:-}" +if [ -z "$EFFORT" ] && [ -f "$REPO/settings.json" ]; then EFFORT=$(jq -r '.effortLevel // "?"' "$REPO/settings.json" 2>/dev/null) - [ -z "$EFFORT" ] && EFFORT="?" fi +[ -z "$EFFORT" ] && EFFORT="?" # Session duration (from total_duration_ms) DURATION_MS=$(echo "$INPUT" | jq -r \ diff --git a/lib/tests/effort-routing.test.sh b/lib/tests/effort-routing.test.sh index f3f5504..1de335f 100755 --- a/lib/tests/effort-routing.test.sh +++ b/lib/tests/effort-routing.test.sh @@ -29,6 +29,10 @@ printf -- '---\nname: bad\neffort: turbo\n---\n' > "$FIX/bad.md" # ── 1) session default (spec D1) has "settings.json" '"effortLevel": "high"' +# ── 2) hooks: env-var warning + live effort in the statusline (spec D1, D5) +has "hooks/session-start.sh" 'CLAUDE_CODE_EFFORT_LEVEL' +has "hooks/statusline.sh" 'CLAUDE_EFFORT' + # ── summary (later tasks insert their locks ABOVE this line) printf 'effort-routing census: %d pass, %d fail\n' "$pass" "$fail" [ "$fail" -eq 0 ] diff --git a/settings.json b/settings.json index d34217e..a008190 100644 --- a/settings.json +++ b/settings.json @@ -444,7 +444,7 @@ } }, "feedbackDrafts": "off", - "effortLevel": "xhigh", + "effortLevel": "high", "remoteControlAtStartup": true, "inputNeededNotifEnabled": true, "skipAutoPermissionPrompt": true, From 9223fda99f9471a0c0865663afa4021929c96138 Mon Sep 17 00:00:00 2001 From: bastien Date: Mon, 28 Sep 2026 18:59:49 +0200 Subject: [PATCH 07/23] feat(effort): pin effort on the 20 repo-authored agents (BDR-077 second axis) --- agents/analyzer.md | 1 + agents/bugfixer.md | 1 + agents/code-cleaner.md | 1 + agents/commit-changer.md | 1 + agents/doc-syncer.md | 1 + agents/feater.md | 1 + agents/geo-analyzer.md | 1 + agents/handover-doc-writer.md | 1 + agents/hotfixer.md | 1 + agents/onboarder.md | 1 + agents/plan-challenger.md | 1 + agents/plugin-advisor.md | 1 + agents/plugin-probe.md | 1 + agents/refactorer.md | 1 + agents/release-executor.md | 1 + agents/scaffolder.md | 2 +- agents/security-auditor.md | 1 + agents/seo-analyzer.md | 1 + agents/validator-analyzer.md | 1 + agents/verifier.md | 1 + lib/tests/effort-routing.test.sh | 8 ++++++++ skills/init-project/SKILL.md | 2 +- 22 files changed, 29 insertions(+), 2 deletions(-) diff --git a/agents/analyzer.md b/agents/analyzer.md index b7f7728..84b7f0a 100644 --- a/agents/analyzer.md +++ b/agents/analyzer.md @@ -3,6 +3,7 @@ name: analyzer description: Analyze code, codebase, or problem before any modification. Produces a factual report without proposing solutions. Use proactively before any refactoring, design, or implementation. tools: Read, Grep, Glob, Bash model: opus +effort: high memory: project --- diff --git a/agents/bugfixer.md b/agents/bugfixer.md index 3bbd971..3094d12 100644 --- a/agents/bugfixer.md +++ b/agents/bugfixer.md @@ -3,6 +3,7 @@ name: bugfixer description: Bug-fix EXECUTOR — dispatched by /bugfix with a closed DIAGNOSIS + FIX PLAN + contract. Applies the fix and a regression test, runs the suite, reports. No investigation, no questions, no commit. tools: Read, Edit, Write, Bash, Grep, Glob model: sonnet +effort: medium --- # BUGFIXER — fix executor diff --git a/agents/code-cleaner.md b/agents/code-cleaner.md index 3f94d21..6c32af2 100644 --- a/agents/code-cleaner.md +++ b/agents/code-cleaner.md @@ -3,6 +3,7 @@ name: code-cleaner description: Cleanup EXECUTOR (PHASE 2) — dispatched by /code-clean with an APPROVED scope. Deletes approved dead code, hands style/structural items to the refactorer, re-audits. Zero behavior change. No audit, no questions, no commit. tools: Read, Edit, Write, Bash, Grep, Glob model: sonnet +effort: medium --- # CODE-CLEANER — cleanup executor (PHASE 2) diff --git a/agents/commit-changer.md b/agents/commit-changer.md index 4b7712d..b74d68b 100644 --- a/agents/commit-changer.md +++ b/agents/commit-changer.md @@ -3,6 +3,7 @@ name: commit-changer description: Retrace-and-commit engine — dispatched by /commit-change. Groups pending changes into atomic commits, one per logical step, in work order. tools: Bash, Read, Grep, Glob model: sonnet +effort: high --- # Git Smart Commit diff --git a/agents/doc-syncer.md b/agents/doc-syncer.md index a042395..08bf745 100644 --- a/agents/doc-syncer.md +++ b/agents/doc-syncer.md @@ -3,6 +3,7 @@ name: doc-syncer description: 'Two-mode public-doc sync agent — MODE: audit (dispatched model="opus" — drift detection, semantic analysis, drafts, PATCH PLAN, read-only) and MODE: patch (sonnet pin — applies the APPROVED plan, oracle-checked, emits CHANGE SUMMARY + PATCHED_FILES). The validation gate lives in the DISPATCHER (BDR-077). Convention-aware (Diátaxis, Keep a Changelog); never touches .claude/.' tools: Read, Write, Edit, Bash, Grep, Glob model: sonnet +effort: high --- # DOC SYNCER diff --git a/agents/feater.md b/agents/feater.md index e834ef6..99e9a92 100644 --- a/agents/feater.md +++ b/agents/feater.md @@ -3,6 +3,7 @@ name: feater description: Small-feature EXECUTOR — dispatched by /feat with a closed plan + contract. Implements to the letter, tests, reports. No planning, no questions, no commit. tools: Read, Edit, Write, Bash, Grep, Glob model: sonnet +effort: medium --- # FEATER — plan executor diff --git a/agents/geo-analyzer.md b/agents/geo-analyzer.md index f23e0e7..22f37d9 100644 --- a/agents/geo-analyzer.md +++ b/agents/geo-analyzer.md @@ -3,6 +3,7 @@ name: geo-analyzer description: GEO audit agent for AI search engines — dispatched by /geo and /seo. Audits AI crawlers, llms.txt, entity signals, Schema.org; emits a fix bundle (dispatcher applies), scored report. Classical SEO → seo-analyzer agent. tools: Read, Edit, Write, Bash, Grep, Glob, WebFetch, WebSearch model: opus +effort: xhigh --- # GEO — Generative Engine Optimization audit, fix & strategy diff --git a/agents/handover-doc-writer.md b/agents/handover-doc-writer.md index 85d7e36..79b63dc 100644 --- a/agents/handover-doc-writer.md +++ b/agents/handover-doc-writer.md @@ -3,6 +3,7 @@ name: handover-doc-writer description: 'Two-mode deliverable writer — MODE: synthesize (dispatched model="opus" — memory+git clustering, 6-chapter synthesis into a run-scoped draft) and MODE: render (sonnet pin — annexes, precheck, deterministic gates, MD + branded HTML/PDF from the draft). Dispatched twice by client-handover with the resolved PACKAGE. No audits, no questions, no dispatch.' tools: Read, Write, Edit, Bash, Grep, Glob, WebSearch, WebFetch model: sonnet +effort: high --- # HANDOVER DOC WRITER diff --git a/agents/hotfixer.md b/agents/hotfixer.md index 5f4eb1e..1aeb4e9 100644 --- a/agents/hotfixer.md +++ b/agents/hotfixer.md @@ -3,6 +3,7 @@ name: hotfixer description: Quick-fix executor — dispatched by /hotfix, which owns the routing and gitflow gate. Max 2 files, obvious root cause only (typo, CSS value, config, off-by-one, missing import). tools: Read, Edit, Write, Bash, Grep, Glob model: sonnet +effort: low --- # HOTFIXER — closed-fix executor / L1 fix-bundle applier diff --git a/agents/onboarder.md b/agents/onboarder.md index 4a8a792..209d1d7 100644 --- a/agents/onboarder.md +++ b/agents/onboarder.md @@ -3,6 +3,7 @@ name: onboarder description: Generate claude-config files (CLAUDE.md, settings.json, .claudeignore, .gitignore safety, .claude/tasks/ + .claude/memory/ + .claude/audits/) for an existing project. Pure config generator — no interview, no audit. Called by /onboard orchestrator. tools: Read, Write, Edit, Bash, Glob, Grep model: sonnet +effort: medium --- # ONBOARDER (config generator) diff --git a/agents/plan-challenger.md b/agents/plan-challenger.md index 50d91fa..58a8426 100644 --- a/agents/plan-challenger.md +++ b/agents/plan-challenger.md @@ -3,6 +3,7 @@ name: plan-challenger description: Fresh independent plan challenger — reads a PLAN file from disk and adversarially attacks it through ONE assigned lens (correctness | robustness | simplicity), then renders structured findings + a verdict. Report-only, never fixes, never implements. Dispatched fresh; blind to the other lenses. tools: Read, Grep, Glob, Bash model: opus +effort: xhigh --- # PLAN-CHALLENGER AGENT diff --git a/agents/plugin-advisor.md b/agents/plugin-advisor.md index 1772299..2237d53 100644 --- a/agents/plugin-advisor.md +++ b/agents/plugin-advisor.md @@ -3,6 +3,7 @@ name: plugin-advisor description: Plugin-fit REASONER — dispatched by lib/plugin-gate.md with a PROBE REPORT (from plugin-probe). Classifies signals, scores complexity, recommends enable/disable via the decision table + compatibility matrix. Report-only. tools: Read, Glob, Grep model: opus +effort: xhigh --- # PLUGIN ADVISOR diff --git a/agents/plugin-probe.md b/agents/plugin-probe.md index 31c4eac..9f427be 100644 --- a/agents/plugin-probe.md +++ b/agents/plugin-probe.md @@ -3,6 +3,7 @@ name: plugin-probe description: Mechanical detection probe — dispatched by lib/plugin-gate.md BEFORE the plugin-advisor reasoner. Runs the CLI/filesystem probes, reports raw facts as a PROBE REPORT. No analysis, no recommendations. tools: Bash, Read, Glob, Grep model: sonnet +effort: low --- # PLUGIN PROBE diff --git a/agents/refactorer.md b/agents/refactorer.md index dbd4145..e5480f6 100644 --- a/agents/refactorer.md +++ b/agents/refactorer.md @@ -3,6 +3,7 @@ name: refactorer description: Refactor existing code without changing external behavior. Applies strict project norms. Use on legacy or non-compliant code. tools: Read, Write, Edit, Grep, Glob, Bash model: sonnet +effort: high --- # REFACTORER diff --git a/agents/release-executor.md b/agents/release-executor.md index 204582f..f33a180 100644 --- a/agents/release-executor.md +++ b/agents/release-executor.md @@ -3,6 +3,7 @@ name: release-executor description: Mechanical release executor — dispatched by /release-candidate for its two spans (prep, finish+tag). Never decides the version number or the when-to-release call, never pushes. tools: Read, Edit, Write, Bash, Grep, Glob model: sonnet +effort: low --- # RELEASE-EXECUTOR — mechanical release spans diff --git a/agents/scaffolder.md b/agents/scaffolder.md index a758ddb..607bde3 100644 --- a/agents/scaffolder.md +++ b/agents/scaffolder.md @@ -3,7 +3,7 @@ name: scaffolder description: Create empty project skeleton. Generates CLAUDE.md, settings, structure, config, empty entry points, installs deps, optional Docker. NO business logic. tools: Read, Write, Edit, Bash, Glob, Grep model: sonnet -effort: high +effort: medium --- # SCAFFOLDER diff --git a/agents/security-auditor.md b/agents/security-auditor.md index 01352d7..5f1b586 100644 --- a/agents/security-auditor.md +++ b/agents/security-auditor.md @@ -3,6 +3,7 @@ name: security-auditor description: 'SAST security gate — runs the pinned semgrep rulesets + the CLAUDE.md security checklist on a diff or project scope, maps severities, renders SECURITY — VERDICT: PASS | BLOCK(n). Blocks HIGH/CRITICAL only, reports the rest. Never fixes code. Fresh dispatch, no iteration history.' tools: Read, Grep, Glob, Bash, Write model: sonnet +effort: xhigh --- # SECURITY-AUDITOR AGENT diff --git a/agents/seo-analyzer.md b/agents/seo-analyzer.md index 7233632..9d5d580 100644 --- a/agents/seo-analyzer.md +++ b/agents/seo-analyzer.md @@ -3,6 +3,7 @@ name: seo-analyzer description: 'Classical SEO audit agent (Google, Bing) — dispatched from /seo. Live audit: Core Web Vitals, on-page, technical, local SEO, legal (FR). Emits a fix bundle (dispatcher applies) + scored report. AI/GEO → geo-analyzer agent.' tools: Read, Edit, Write, Bash, Grep, Glob, WebFetch, WebSearch model: opus +effort: xhigh --- # SEO — Classical Search Engines audit, fix & strategy diff --git a/agents/validator-analyzer.md b/agents/validator-analyzer.md index 57112d0..3156755 100644 --- a/agents/validator-analyzer.md +++ b/agents/validator-analyzer.md @@ -3,6 +3,7 @@ name: validator-analyzer description: Web standards audit agent — W3C HTML validity (validator.nu), W3C CSS validity (jigsaw.w3.org), WCAG 2.1 accessibility (axe-core, pa11y, WAVE). Dispatched from /web-validate. Produces scored .claude/audits/VALIDATE.md report with concrete diffs for auto-fixable issues and user actions for judgment-required fixes. Complementary to /harden (security), /seo (indexability), /geo (AI extraction). tools: Read, Edit, Write, Bash, Grep, Glob, WebFetch model: sonnet +effort: low --- # Validator — W3C + WCAG audit diff --git a/agents/verifier.md b/agents/verifier.md index c9cdc2d..3fb7116 100644 --- a/agents/verifier.md +++ b/agents/verifier.md @@ -3,6 +3,7 @@ name: verifier description: Fresh independent verifier — reads a CONTRACT file from disk and renders a structured verdict (CONFORME / ECARTS / ERROR) on the implemented diff. Report-only, never fixes. Dispatched fresh at every iteration; receives no iteration history. tools: Read, Grep, Glob, Bash model: sonnet +effort: xhigh --- # VERIFIER AGENT diff --git a/lib/tests/effort-routing.test.sh b/lib/tests/effort-routing.test.sh index 1de335f..a34a89b 100755 --- a/lib/tests/effort-routing.test.sh +++ b/lib/tests/effort-routing.test.sh @@ -33,6 +33,14 @@ has "settings.json" '"effortLevel": "high"' has "hooks/session-start.sh" 'CLAUDE_CODE_EFFORT_LEVEL' has "hooks/statusline.sh" 'CLAUDE_EFFORT' +# ── 3) agent pins (spec D2): one effort per agent file, judgment mode wins on mode-based agents +for a in hotfixer release-executor plugin-probe validator-analyzer; do fm_has_effort "agents/$a.md" low; done +for a in feater bugfixer code-cleaner onboarder scaffolder; do fm_has_effort "agents/$a.md" medium; done +for a in refactorer analyzer commit-changer doc-syncer handover-doc-writer; do fm_has_effort "agents/$a.md" high; done +for a in plan-challenger plugin-advisor verifier security-auditor seo-analyzer geo-analyzer; do fm_has_effort "agents/$a.md" xhigh; done +for a in interviewer client-handover-writer status-reporter; do fm_no_effort "agents/$a.md"; done +has "skills/init-project/SKILL.md" 'pin sonnet, effort medium' + # ── summary (later tasks insert their locks ABOVE this line) printf 'effort-routing census: %d pass, %d fail\n' "$pass" "$fail" [ "$fail" -eq 0 ] diff --git a/skills/init-project/SKILL.md b/skills/init-project/SKILL.md index a156cf7..e5d240c 100644 --- a/skills/init-project/SKILL.md +++ b/skills/init-project/SKILL.md @@ -95,7 +95,7 @@ contract, each tagged `[gated ]`. STEP 9's verifier judges against this enriched contract. ## STEP 5 — SCAFFOLD -Dispatch `Agent(subagent_type="scaffolder")` (pin sonnet, effort high — +Dispatch `Agent(subagent_type="scaffolder")` (pin sonnet, effort medium — BDR-077 : le design est CLOS au gate #1, le scaffold est de l'exécution, plus jamais inline sur le modèle de session). Pass IN THE PROMPT (LRN-126 — every field the scaffolder consumes crosses the dispatch): BRIEF (verbatim) From 94189adbc61fc7f183d2ea6b5a7d26a74b21d742 Mon Sep 17 00:00:00 2001 From: bastien Date: Mon, 28 Sep 2026 19:11:21 +0200 Subject: [PATCH 08/23] feat(effort): entry effort level on the 31 user-invoked skills (spec D3) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A/B /reconcile headless — BEFORE requests=18 output=12374 thinking=3135 effort={'high'} duration_ms=96518 / AFTER requests=15 output=9038 thinking=2248 effort={'low'} duration_ms=78410 --- lib/tests/effort-routing.test.sh | 6 ++++++ skills/analyze/SKILL.md | 1 + skills/audit-delta/SKILL.md | 1 + skills/bugfix/SKILL.md | 1 + skills/capitalize/SKILL.md | 1 + skills/client-handover/SKILL.md | 1 + skills/close/SKILL.md | 1 + skills/code-clean/SKILL.md | 1 + skills/commit-change/SKILL.md | 1 + skills/deploy/SKILL.md | 1 + skills/doc/SKILL.md | 1 + skills/feat/SKILL.md | 1 + skills/geo/SKILL.md | 1 + skills/gitflow/SKILL.md | 1 + skills/harden/SKILL.md | 1 + skills/hotfix/SKILL.md | 1 + skills/init-project/SKILL.md | 1 + skills/onboard/SKILL.md | 1 + skills/plugin-check/SKILL.md | 1 + skills/profile/SKILL.md | 1 + skills/prune-memory/SKILL.md | 1 + skills/reconcile/SKILL.md | 1 + skills/refactor/SKILL.md | 1 + skills/release-candidate/SKILL.md | 1 + skills/seo/SKILL.md | 1 + skills/ship-feature/SKILL.md | 1 + skills/status/SKILL.md | 1 + skills/tour/SKILL.md | 1 + skills/web-validate/SKILL.md | 1 + 29 files changed, 34 insertions(+) diff --git a/lib/tests/effort-routing.test.sh b/lib/tests/effort-routing.test.sh index a34a89b..4944d82 100755 --- a/lib/tests/effort-routing.test.sh +++ b/lib/tests/effort-routing.test.sh @@ -41,6 +41,12 @@ for a in plan-challenger plugin-advisor verifier security-auditor seo-analyzer g for a in interviewer client-handover-writer status-reporter; do fm_no_effort "agents/$a.md"; done has "skills/init-project/SKILL.md" 'pin sonnet, effort medium' +# ── 4) skill entry levels (spec D3): the user's invocation sets the run's level +for s in status commit-change release-candidate doc capitalize close reconcile deploy profile plugin-check; do fm_has_effort "skills/$s/SKILL.md" low; done +for s in gitflow prune-memory find-docs; do fm_has_effort "skills/$s/SKILL.md" medium; done +for s in feat hotfix bugfix refactor web-validate harden seo geo; do fm_has_effort "skills/$s/SKILL.md" high; done +for s in ship-feature init-project onboard tour audit-delta analyze code-clean client-handover; do fm_has_effort "skills/$s/SKILL.md" xhigh; done + # ── summary (later tasks insert their locks ABOVE this line) printf 'effort-routing census: %d pass, %d fail\n' "$pass" "$fail" [ "$fail" -eq 0 ] diff --git a/skills/analyze/SKILL.md b/skills/analyze/SKILL.md index db6c353..8e5944b 100644 --- a/skills/analyze/SKILL.md +++ b/skills/analyze/SKILL.md @@ -1,5 +1,6 @@ --- name: analyze +effort: xhigh description: 'Deep factual code analysis (read-only) or DEBUG mode (pass error/stack trace) — no solutions proposed, no file modifications. Triggers: "analyze", "analyse", "how does X work", "comment ça marche", "investigate only", "root cause only, no fix", "pourquoi ce comportement", "debug analysis". Fix wanted → /bugfix or /hotfix instead.' argument-hint: allowed-tools: Read, Grep, Glob, Bash diff --git a/skills/audit-delta/SKILL.md b/skills/audit-delta/SKILL.md index 82f8b43..02ac235 100644 --- a/skills/audit-delta/SKILL.md +++ b/skills/audit-delta/SKILL.md @@ -1,5 +1,6 @@ --- name: audit-delta +effort: xhigh description: | Use when the user wants a recurring code audit scoped to changes since the previous run (full codebase on first run), on selectable axes: diff --git a/skills/bugfix/SKILL.md b/skills/bugfix/SKILL.md index 733b314..7366e07 100644 --- a/skills/bugfix/SKILL.md +++ b/skills/bugfix/SKILL.md @@ -1,5 +1,6 @@ --- name: bugfix +effort: high description: | Structured bug fix with root cause investigation. For bugs where the cause isn't immediately obvious, spans multiple files, or diff --git a/skills/capitalize/SKILL.md b/skills/capitalize/SKILL.md index c637b7c..74b96fb 100644 --- a/skills/capitalize/SKILL.md +++ b/skills/capitalize/SKILL.md @@ -1,5 +1,6 @@ --- name: capitalize +effort: low description: | Use when about to /clear or /compact, or closing a session, with decisions, learnings, blockers, evals, or TODO changes not yet written diff --git a/skills/client-handover/SKILL.md b/skills/client-handover/SKILL.md index 8d07813..73ebfaf 100644 --- a/skills/client-handover/SKILL.md +++ b/skills/client-handover/SKILL.md @@ -1,5 +1,6 @@ --- name: client-handover +effort: xhigh description: | Use when finalizing a project for non-technical client delivery — final audits, live-site validation, branded deliverable (MD + HTML + diff --git a/skills/close/SKILL.md b/skills/close/SKILL.md index 014faef..af68e3b 100644 --- a/skills/close/SKILL.md +++ b/skills/close/SKILL.md @@ -1,5 +1,6 @@ --- name: close +effort: low description: | End-of-session ritual — flush what was decided, learned, and blocked into `.claude/memory/`, reconcile `.claude/tasks/TODO.md`, and log a journal line. diff --git a/skills/code-clean/SKILL.md b/skills/code-clean/SKILL.md index 56fc4ce..30079f2 100644 --- a/skills/code-clean/SKILL.md +++ b/skills/code-clean/SKILL.md @@ -1,5 +1,6 @@ --- name: code-clean +effort: xhigh description: | Full codebase cleanup: dead code, style/norm enforcement, structural issues. Two-phase: read-only audit, then approved fixes only diff --git a/skills/commit-change/SKILL.md b/skills/commit-change/SKILL.md index f2b2579..876b162 100644 --- a/skills/commit-change/SKILL.md +++ b/skills/commit-change/SKILL.md @@ -1,5 +1,6 @@ --- name: commit-change +effort: low description: | Analyze all pending changes (staged, unstaged, untracked) and create atomic commits grouped by logical unit, retracing the work. Any git diff --git a/skills/deploy/SKILL.md b/skills/deploy/SKILL.md index 7413e52..841549a 100644 --- a/skills/deploy/SKILL.md +++ b/skills/deploy/SKILL.md @@ -1,5 +1,6 @@ --- name: deploy +effort: low description: | Use when deploying a project via its per-project runbook — instantiates the delta since last deploy, hands off for out-of-band execution, resumes cold, learns from errors. diff --git a/skills/doc/SKILL.md b/skills/doc/SKILL.md index db4d8e1..4f8d107 100644 --- a/skills/doc/SKILL.md +++ b/skills/doc/SKILL.md @@ -1,5 +1,6 @@ --- name: doc +effort: low description: | Use when documentation may be out of sync with code — features added/removed vs README / INSTALL / DEPLOY / CHANGELOG. Stack-aware diff --git a/skills/feat/SKILL.md b/skills/feat/SKILL.md index 06a49aa..b013c0a 100644 --- a/skills/feat/SKILL.md +++ b/skills/feat/SKILL.md @@ -1,5 +1,6 @@ --- name: feat +effort: high description: | Small feature implementation (1-5 files). Reflection inline (scope, plan, contract — session model), execution dispatched to the diff --git a/skills/geo/SKILL.md b/skills/geo/SKILL.md index f988cac..3ed2600 100644 --- a/skills/geo/SKILL.md +++ b/skills/geo/SKILL.md @@ -1,5 +1,6 @@ --- name: geo +effort: high description: | Use when a web project needs AI-search visibility audit — ChatGPT, Perplexity, Gemini, AI Overviews, Copilot… Standalone GEO; dispatches diff --git a/skills/gitflow/SKILL.md b/skills/gitflow/SKILL.md index fdffdd3..b5df56e 100644 --- a/skills/gitflow/SKILL.md +++ b/skills/gitflow/SKILL.md @@ -1,5 +1,6 @@ --- name: gitflow +effort: medium description: Use when a project needs gitflow branch operations — bootstrapping main+develop, starting a typed branch (feature/bugfix/release/hotfix), or integrating finished work by directed merge — or when an orchestrator must branch or merge under the gitflow model. Use when about to merge any branch into develop or main. --- diff --git a/skills/harden/SKILL.md b/skills/harden/SKILL.md index e867928..5aad845 100644 --- a/skills/harden/SKILL.md +++ b/skills/harden/SKILL.md @@ -1,5 +1,6 @@ --- name: harden +effort: high description: | Web hardening audit — HTTPS/TLS, HSTS, security headers (CSP, X-Frame-Options…), cookie flags, canonical, custom 404, server config diff --git a/skills/hotfix/SKILL.md b/skills/hotfix/SKILL.md index 2e55596..ff11901 100644 --- a/skills/hotfix/SKILL.md +++ b/skills/hotfix/SKILL.md @@ -1,5 +1,6 @@ --- name: hotfix +effort: high description: | Quick fix for superficial bugs: typos, CSS issues, config errors, off-by-one, wrong variable name, missing import, broken link. diff --git a/skills/init-project/SKILL.md b/skills/init-project/SKILL.md index e5d240c..069779a 100644 --- a/skills/init-project/SKILL.md +++ b/skills/init-project/SKILL.md @@ -1,5 +1,6 @@ --- name: init-project +effort: xhigh description: 'Use when initializing a brand-new project from scratch — needs interview, design, scaffold, and TDD implementation. Multi-agent orchestrator: plugin-advisor + interviewer + analyzer + scaffolder with two validation gates. Triggers: "init project", "new project", "start project from scratch", "scaffold project", "init-project".' argument-hint: allowed-tools: Read, Write, Edit, Bash, Grep, Glob, Agent, Skill diff --git a/skills/onboard/SKILL.md b/skills/onboard/SKILL.md index b3ae11c..eba468d 100644 --- a/skills/onboard/SKILL.md +++ b/skills/onboard/SKILL.md @@ -1,5 +1,6 @@ --- name: onboard +effort: xhigh description: 'Use when bringing an existing repo into the claude-config framework — needs archetype detection, config install, full multi-axis audit (debt/SEO/GEO/UI-UX/perf/security/a11y/docs), and prioritized backlog. Multi-agent orchestrator. Do NOT use for repos created via /init-project. Triggers: "onboard", "onboard project", "audit existing repo", "setup existing project".' argument-hint: '[optional hints: "Python FastAPI" | "Next.js monorepo" | "force-archetype:wordpress"]' allowed-tools: Read, Write, Edit, Bash, Glob, Grep, Agent, Skill diff --git a/skills/plugin-check/SKILL.md b/skills/plugin-check/SKILL.md index d81fc6a..6ad5cbd 100644 --- a/skills/plugin-check/SKILL.md +++ b/skills/plugin-check/SKILL.md @@ -1,5 +1,6 @@ --- name: plugin-check +effort: low description: 'Audit active plugins vs project needs. Read-only advisory recommending enable/disable. Triggers: "plugin-check", "quels plugins".' argument-hint: '[ex: "React + FastAPI" or "Rust CLI, no frontend"]' allowed-tools: Read, Bash, Glob, Grep, Agent diff --git a/skills/profile/SKILL.md b/skills/profile/SKILL.md index 1815d51..be40ece 100644 --- a/skills/profile/SKILL.md +++ b/skills/profile/SKILL.md @@ -1,5 +1,6 @@ --- name: profile +effort: low description: | Partition Claude skills by purpose: design, dev, qa, audit, minimal. Toggles symlinks between skills/ and skills-disabled/ to keep only diff --git a/skills/prune-memory/SKILL.md b/skills/prune-memory/SKILL.md index 21ba59d..27f0d10 100644 --- a/skills/prune-memory/SKILL.md +++ b/skills/prune-memory/SKILL.md @@ -1,5 +1,6 @@ --- name: prune-memory +effort: medium description: | Use when .claude/memory/ registries grow too large or noisy — superseded entries verbose, similar entries cluttering, journal stale, caveman style diff --git a/skills/reconcile/SKILL.md b/skills/reconcile/SKILL.md index 2585df1..187bb92 100644 --- a/skills/reconcile/SKILL.md +++ b/skills/reconcile/SKILL.md @@ -1,5 +1,6 @@ --- name: reconcile +effort: low description: Use when you need the REAL open-work state of a project and the TODO or memory registries may be stale — "is the queue empty?", "what's left open?", "qu'est-ce qui reste", before /close, after a break, or when a checkbox/status looks doubtful. Confronts declared status (TODO checkboxes, registry statuses) against real git/fs state and surfaces the gaps. NOT memory curation (that is /prune-memory). --- diff --git a/skills/refactor/SKILL.md b/skills/refactor/SKILL.md index 67e5abd..b3720dd 100644 --- a/skills/refactor/SKILL.md +++ b/skills/refactor/SKILL.md @@ -1,5 +1,6 @@ --- name: refactor +effort: high description: 'Improve code quality without changing behavior — strict norm enforcement, targeted scope (file/module). Full-codebase audit+cleanup → /code-clean. Triggers: "refactor", "clean up code", "normaliser".' argument-hint: allowed-tools: Read, Write, Edit, Grep, Glob, Bash, Agent diff --git a/skills/release-candidate/SKILL.md b/skills/release-candidate/SKILL.md index bafc6fc..7eddd35 100644 --- a/skills/release-candidate/SKILL.md +++ b/skills/release-candidate/SKILL.md @@ -1,5 +1,6 @@ --- name: release-candidate +effort: low description: 'Use when develop is ahead of main and you want to cut a versioned release — finalize version.txt + CHANGELOG, merge develop→main via the gitflow fan-out, tag it, and push. Triggers: "cut a release", "release candidate", "tag a version", "ship develop to main". NOT feature/bugfix integration (that is gitflow finish via /ship-feature) nor a hotfix.' allowed-tools: - Read diff --git a/skills/seo/SKILL.md b/skills/seo/SKILL.md index c7206f5..a4cf440 100644 --- a/skills/seo/SKILL.md +++ b/skills/seo/SKILL.md @@ -1,5 +1,6 @@ --- name: seo +effort: high description: | Use when a web project needs SEO + GEO audit or optimization — classical search (Google, Bing) AND AI search (ChatGPT, Perplexity, AI diff --git a/skills/ship-feature/SKILL.md b/skills/ship-feature/SKILL.md index 17a4233..76bc316 100644 --- a/skills/ship-feature/SKILL.md +++ b/skills/ship-feature/SKILL.md @@ -1,5 +1,6 @@ --- name: ship-feature +effort: xhigh description: 'Use when shipping a new feature end-to-end — needs design brainstorm, planning, TDD implementation with subagents, error recovery, code review, and finish. Multi-agent orchestrator (9-step pipeline). Triggers: "ship feature", "ship-feature", "build and merge", "feature end-to-end", "implement and ship".' argument-hint: allowed-tools: Read, Write, Edit, Bash, Grep, Glob diff --git a/skills/status/SKILL.md b/skills/status/SKILL.md index 04fdf9d..39ce739 100644 --- a/skills/status/SKILL.md +++ b/skills/status/SKILL.md @@ -1,5 +1,6 @@ --- name: status +effort: low description: 'Consolidated project snapshot — plugins + passive token cost, git state, recent commits, GSD v2 milestone progress. Read-only. Run at session start or after a break. Open-work reconciliation (stale TODO vs real git) → /reconcile. Triggers: "status", "sitrep", "where are we", "project state", "after break".' argument-hint: (no arguments needed) allowed-tools: Read, Bash, Glob, Grep, Agent diff --git a/skills/tour/SKILL.md b/skills/tour/SKILL.md index 7cb455f..b19dfc1 100644 --- a/skills/tour/SKILL.md +++ b/skills/tour/SKILL.md @@ -1,5 +1,6 @@ --- name: tour +effort: xhigh description: | Use when the user wants ONE grouped pass over a whole project (or a list of projects) covering all hygiene axes together: code cleanup + diff --git a/skills/web-validate/SKILL.md b/skills/web-validate/SKILL.md index 7168333..1fa34b6 100644 --- a/skills/web-validate/SKILL.md +++ b/skills/web-validate/SKILL.md @@ -1,5 +1,6 @@ --- name: web-validate +effort: high description: | Use when a web project needs W3C HTML/CSS validity or WCAG 2.1 accessibility audit. Dispatches the validator-analyzer agent, strict From bac235cb3342401ac3b604008b7143dcac13dd44 Mon Sep 17 00:00:00 2001 From: bastien Date: Mon, 28 Sep 2026 19:21:15 +0200 Subject: [PATCH 09/23] feat(effort): xhigh on the vendored brainstorming and writing-plans, re-applied at resync --- install-plugins.sh | 11 +++++++++++ lib/tests/effort-routing.test.sh | 9 +++++++++ 2 files changed, 20 insertions(+) diff --git a/install-plugins.sh b/install-plugins.sh index 5e2a0a0..c669cde 100644 --- a/install-plugins.sh +++ b/install-plugins.sh @@ -934,6 +934,17 @@ for _ext_skill in "${EXT_SKILL_NAMES[@]}"; do done echo "" +# Effort tiering (BDR-NEXT): the vendored brainstorming/writing-plans carry an +# effort pin upstream lacks; re-apply after every resync (census lock in +# lib/tests/effort-routing.test.sh alarms if this ever stops working). +for _s in brainstorming writing-plans; do + _f="$(cd "$(dirname "$0")" && pwd)/skills-external/$_s/SKILL.md" + if [ -f "$_f" ] && ! grep -q '^effort:' "$_f"; then + sed -i "0,/^name: $_s\$/s//&\neffort: xhigh/" "$_f" + fi +done +unset _s _f + # ============================================================ # STEP 8.5 — EXTERNAL SKILLS (npx skills add …) # ============================================================ diff --git a/lib/tests/effort-routing.test.sh b/lib/tests/effort-routing.test.sh index 4944d82..4e17b3e 100755 --- a/lib/tests/effort-routing.test.sh +++ b/lib/tests/effort-routing.test.sh @@ -47,6 +47,15 @@ for s in gitflow prune-memory find-docs; do fm_has_effort "skills/$s/SKILL.md" m for s in feat hotfix bugfix refactor web-validate harden seo geo; do fm_has_effort "skills/$s/SKILL.md" high; done for s in ship-feature init-project onboard tour audit-delta analyze code-clean client-handover; do fm_has_effort "skills/$s/SKILL.md" xhigh; done +# ── 9) vendored superpowers carry xhigh (spec D3). The files live in skills-external/ (gitignored, +# machine-owned), so the durable artifact is the install-plugins.sh re-apply; the frontmatter +# check skips VISIBLY when the skill is not vendored yet (fresh clone before make plugin). +for s in brainstorming writing-plans; do + if [ -f "$R/skills-external/$s/SKILL.md" ]; then fm_has_effort "skills-external/$s/SKILL.md" xhigh + else printf 'SKIP skills-external/%s/SKILL.md not vendored yet (run make plugin)\n' "$s"; fi +done +has "install-plugins.sh" 'effort: xhigh' + # ── summary (later tasks insert their locks ABOVE this line) printf 'effort-routing census: %d pass, %d fail\n' "$pass" "$fail" [ "$fail" -eq 0 ] From 3de9d4f85ae1676a0e1c148ee8f890ebc97b0592 Mon Sep 17 00:00:00 2001 From: bastien Date: Mon, 28 Sep 2026 19:23:43 +0200 Subject: [PATCH 10/23] fix(effort): find-docs is ctx7-generated and gitignored, no entry level (30 skills, not 31) --- lib/tests/effort-routing.test.sh | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/lib/tests/effort-routing.test.sh b/lib/tests/effort-routing.test.sh index 4e17b3e..a50db5d 100755 --- a/lib/tests/effort-routing.test.sh +++ b/lib/tests/effort-routing.test.sh @@ -43,7 +43,7 @@ has "skills/init-project/SKILL.md" 'pin sonnet, effort medium' # ── 4) skill entry levels (spec D3): the user's invocation sets the run's level for s in status commit-change release-candidate doc capitalize close reconcile deploy profile plugin-check; do fm_has_effort "skills/$s/SKILL.md" low; done -for s in gitflow prune-memory find-docs; do fm_has_effort "skills/$s/SKILL.md" medium; done +for s in gitflow prune-memory; do fm_has_effort "skills/$s/SKILL.md" medium; done for s in feat hotfix bugfix refactor web-validate harden seo geo; do fm_has_effort "skills/$s/SKILL.md" high; done for s in ship-feature init-project onboard tour audit-delta analyze code-clean client-handover; do fm_has_effort "skills/$s/SKILL.md" xhigh; done From 4a450ea6bc10a0d071ca20453d34abb0fdb7099f Mon Sep 17 00:00:00 2001 From: bastien Date: Mon, 28 Sep 2026 19:32:26 +0200 Subject: [PATCH 11/23] feat(effort): five shifter skills, lib/effort-shift.md, model-gate second axis --- lib/effort-shift.md | 57 ++++++++++++++++++++++++++++++++ lib/model-gate.md | 4 +++ lib/tests/effort-routing.test.sh | 7 ++++ skills/effort-high/SKILL.md | 6 ++++ skills/effort-low/SKILL.md | 6 ++++ skills/effort-max/SKILL.md | 6 ++++ skills/effort-medium/SKILL.md | 6 ++++ skills/effort-xhigh/SKILL.md | 6 ++++ 8 files changed, 98 insertions(+) create mode 100644 lib/effort-shift.md create mode 100644 skills/effort-high/SKILL.md create mode 100644 skills/effort-low/SKILL.md create mode 100644 skills/effort-max/SKILL.md create mode 100644 skills/effort-medium/SKILL.md create mode 100644 skills/effort-xhigh/SKILL.md diff --git a/lib/effort-shift.md b/lib/effort-shift.md new file mode 100644 index 0000000..71be69b --- /dev/null +++ b/lib/effort-shift.md @@ -0,0 +1,57 @@ +# Effort shift — phase-level reasoning effort on the main loop (BDR-NEXT) + +Shared include, companion of `lib/model-gate.md`: the gate fixes WHICH model +reflects, this include fixes HOW HARD each phase thinks. The rungs are the +user's: low (fix a line, run a script) · medium (day-to-day) · high +(refactor, resisting bug) · xhigh (architecture, audit before validation) · +max (stuck error, judged need). + +## Mechanics (verified on Claude Code 2.1.283) + +- A skill's `effort:` frontmatter applies from the moment it loads to the + end of the turn: on the user's `/skill` and on a `Skill(...)` call by + Claude in an interactive session. Last loaded wins, both directions. The + prompt cache survives a shift. +- Dispatched agents run on their own `effort:` pin, never on a shift. + Unpinned agents inherit the level in force at dispatch. +- Headless sessions (`-p`, `claude agents`, SDK) ignore skill-level effort: + the run stays at the session level. `CLAUDE_CODE_EFFORT_LEVEL` beats every + frontmatter; keep it unset (the session banner warns). + +## Shifters + +`Skill(effort-low)` · `Skill(effort-medium)` · `Skill(effort-high)` · +`Skill(effort-xhigh)` · `Skill(effort-max)`. One tool call, one-line body. +Typed by the user, `/effort-max` is a turn-scoped max: the relaunch lever +after a STOP. `ultrathink` only adds an in-context nudge; the API level +does not move. + +## Wiring — per orchestrator + +1. A dispatch span starts (executor, collector, fan-out) → + `Skill(effort-medium)`. +2. Reflection resumes after a dispatch span (challenge synthesis, verdict, + plan revision) → `Skill(effort-)`. Concretely: + the line before every `lib/challenge-plan.md` call. +3. The bookkeeping tail (memory commit, doc commit) → `Skill(effort-low)`. +4. Escalation → `Skill(effort-max)`, then the skill's own level again once + the diagnosis is produced. Automatic points: verify-secure loop caps + (GATE 0 floor, GATE 1 conformity, GATE 2 security) and ship-feature + STEP 4b. Not automatic, by doctrine: the challenge fail-safe (a mute + challenger is an infrastructure failure) and "gone WRONG → STOP" (STOP + precedes any further reasoning); their STOP text names the level + reached and suggests `/effort-max` for the relaunch. + +## Re-assert + +- After any nested `Skill(...)` whose frontmatter carries a different + effort (feat → commit-change), reload the orchestrator's own level. +- After a prose gate that ends the turn, the resumed turn runs at the + session level. If the resumed phase is reflection, its first step is + `Skill(effort-)`; dispatch and orchestration phases need + nothing. + +## Never + +- A shift never inside a dispatched agent: pins rule there. +- Max is for diagnosis, not for retrying the same fix harder. diff --git a/lib/model-gate.md b/lib/model-gate.md index 97aab7b..d75bd86 100644 --- a/lib/model-gate.md +++ b/lib/model-gate.md @@ -45,3 +45,7 @@ site — `model: "fable"` when the child performs reflection/orchestration on the main loop's behalf (skill-runners), otherwise its complexity tier (opus = dispatched judgment, sonnet = execution/collection, haiku = short mechanical probes). + +Effort is the second axis of the same table (BDR-NEXT): every typed agent +carries an `effort:` pin next to `model:`, and the main loop shifts per phase +through `lib/effort-shift.md`. Nothing dispatched inherits either axis. diff --git a/lib/tests/effort-routing.test.sh b/lib/tests/effort-routing.test.sh index a50db5d..abe0bb6 100755 --- a/lib/tests/effort-routing.test.sh +++ b/lib/tests/effort-routing.test.sh @@ -56,6 +56,13 @@ for s in brainstorming writing-plans; do done has "install-plugins.sh" 'effort: xhigh' +# ── 5) shifter skills + include (spec D4) +for l in low medium high xhigh max; do fm_has_effort "skills/effort-$l/SKILL.md" "$l"; has "skills/effort-$l/SKILL.md" "name: effort-$l"; done +has "lib/effort-shift.md" 'Headless sessions' +has "lib/effort-shift.md" 'Skill(effort-max)' +has "lib/effort-shift.md" 'never inside a dispatched agent' +has "lib/model-gate.md" 'lib/effort-shift.md' + # ── summary (later tasks insert their locks ABOVE this line) printf 'effort-routing census: %d pass, %d fail\n' "$pass" "$fail" [ "$fail" -eq 0 ] diff --git a/skills/effort-high/SKILL.md b/skills/effort-high/SKILL.md new file mode 100644 index 0000000..cbeaeaa --- /dev/null +++ b/skills/effort-high/SKILL.md @@ -0,0 +1,6 @@ +--- +name: effort-high +description: Investigation shift. Deeper reasoning for diagnosis, LOCATE, contract drafting, refactor judgement inside feat, hotfix and bugfix runs. +effort: high +--- +Effort shifted to high for the rest of this turn (lib/effort-shift.md). Continue with the caller's next step. diff --git a/skills/effort-low/SKILL.md b/skills/effort-low/SKILL.md new file mode 100644 index 0000000..16da6a4 --- /dev/null +++ b/skills/effort-low/SKILL.md @@ -0,0 +1,6 @@ +--- +name: effort-low +description: Bookkeeping shift. Lowers reasoning to the cheapest level for the rest of the turn: journal lines, memory commits, capitalize, release bookkeeping, status output. +effort: low +--- +Effort shifted to low for the rest of this turn (lib/effort-shift.md). Continue with the caller's next step. diff --git a/skills/effort-max/SKILL.md b/skills/effort-max/SKILL.md new file mode 100644 index 0000000..06de425 --- /dev/null +++ b/skills/effort-max/SKILL.md @@ -0,0 +1,6 @@ +--- +name: effort-max +description: Escalation shift. Maximum reasoning when a verify or security loop hits its cap, a gate fails twice, or error recovery starts in ship-feature. +effort: max +--- +Effort shifted to max for the rest of this turn (lib/effort-shift.md). Continue with the caller's next step. diff --git a/skills/effort-medium/SKILL.md b/skills/effort-medium/SKILL.md new file mode 100644 index 0000000..6115c86 --- /dev/null +++ b/skills/effort-medium/SKILL.md @@ -0,0 +1,6 @@ +--- +name: effort-medium +description: Orchestration shift. Standard reasoning between two dispatches: read a subagent report, pick the next step, relay a gate verdict, route a branch. +effort: medium +--- +Effort shifted to medium for the rest of this turn (lib/effort-shift.md). Continue with the caller's next step. diff --git a/skills/effort-xhigh/SKILL.md b/skills/effort-xhigh/SKILL.md new file mode 100644 index 0000000..6a60abd --- /dev/null +++ b/skills/effort-xhigh/SKILL.md @@ -0,0 +1,6 @@ +--- +name: effort-xhigh +description: Reflection shift. Deep reasoning for brainstorm, planning, challenge synthesis and audit verdicts before a human validation gate. +effort: xhigh +--- +Effort shifted to xhigh for the rest of this turn (lib/effort-shift.md). Continue with the caller's next step. From 3c58160d0c2c5b7874bf94c8a833e7848dcc2638 Mon Sep 17 00:00:00 2001 From: bastien Date: Mon, 28 Sep 2026 19:44:22 +0200 Subject: [PATCH 12/23] feat(effort): wire phase shifts in the 13 orchestrators and the handover writer --- agents/client-handover-writer.md | 2 ++ lib/tests/effort-routing.test.sh | 10 ++++++++++ skills/audit-delta/SKILL.md | 3 +++ skills/bugfix/SKILL.md | 5 +++++ skills/code-clean/SKILL.md | 3 +++ skills/feat/SKILL.md | 6 ++++++ skills/geo/SKILL.md | 4 ++++ skills/harden/SKILL.md | 3 +++ skills/hotfix/SKILL.md | 5 +++++ skills/init-project/SKILL.md | 5 +++++ skills/onboard/SKILL.md | 3 +++ skills/seo/SKILL.md | 4 ++++ skills/ship-feature/SKILL.md | 5 +++++ skills/tour/SKILL.md | 2 ++ skills/web-validate/SKILL.md | 4 ++++ 15 files changed, 64 insertions(+) diff --git a/agents/client-handover-writer.md b/agents/client-handover-writer.md index 6b0af6b..125220b 100644 --- a/agents/client-handover-writer.md +++ b/agents/client-handover-writer.md @@ -225,6 +225,7 @@ Store `DEPLOYED_URL` for STEP 7. If empty, ask user during STEP 6. --- ## STEP 3 — BASELINE AUDITS (parallel) +First: `Skill(effort-medium)` (effort-shift: dispatch span starts). Goal: capture `SCORE_*_BEFORE` so the client doc shows the delta. @@ -261,6 +262,7 @@ the gate. this pipeline (initial audits, fix-loop re-dispatches, commit-change, web-validate) carries `model: "fable"` — the child hosts gated orchestration on the pipeline's behalf; it must never inherit the session model. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation. For web projects, dispatch in **a single message with two parallel Agent calls** (each with `model: "fable"`): diff --git a/lib/tests/effort-routing.test.sh b/lib/tests/effort-routing.test.sh index abe0bb6..c0d5bac 100755 --- a/lib/tests/effort-routing.test.sh +++ b/lib/tests/effort-routing.test.sh @@ -63,6 +63,16 @@ has "lib/effort-shift.md" 'Skill(effort-max)' has "lib/effort-shift.md" 'never inside a dispatched agent' has "lib/model-gate.md" 'lib/effort-shift.md' +# ── 6) orchestrator wiring (spec D4) +for s in feat hotfix bugfix ship-feature init-project onboard tour code-clean seo geo harden web-validate audit-delta; do + has "skills/$s/SKILL.md" 'lib/effort-shift.md'; has "skills/$s/SKILL.md" 'Skill(effort-medium)'; done +has "agents/client-handover-writer.md" 'lib/effort-shift.md'; has "agents/client-handover-writer.md" 'Skill(effort-medium)' +for s in feat hotfix bugfix; do has "skills/$s/SKILL.md" 'Skill(effort-high)'; done +for s in ship-feature init-project onboard code-clean audit-delta; do has "skills/$s/SKILL.md" 'Skill(effort-xhigh)'; done +for s in seo geo harden web-validate; do has "skills/$s/SKILL.md" 'Skill(effort-high)'; done +for s in feat hotfix bugfix ship-feature init-project; do has "skills/$s/SKILL.md" 'Skill(effort-low)'; done +has "skills/feat/SKILL.md" 'effort-shift: nested commit-change' + # ── summary (later tasks insert their locks ABOVE this line) printf 'effort-routing census: %d pass, %d fail\n' "$pass" "$fail" [ "$fail" -eq 0 ] diff --git a/skills/audit-delta/SKILL.md b/skills/audit-delta/SKILL.md index 02ac235..1443c91 100644 --- a/skills/audit-delta/SKILL.md +++ b/skills/audit-delta/SKILL.md @@ -29,6 +29,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation. Audit only what changed since the last run, on the axes the user picks. Per axis: **audit → approval gate → fix → re-verify → marker update**, @@ -169,6 +170,7 @@ Then show the user the same compact table inline. ### 3b-bis. CHALLENGE THE PROPOSALS (before the gate) +`Skill(effort-xhigh)` first (effort-shift: reflection resumes). This axis' findings + proposed fixes are a proposal set worth attacking before the human gate. Persist THIS axis' finding list (not the whole append-only report) to `.claude/tasks/plans/--.md`, then run @@ -254,6 +256,7 @@ Then offer to capitalize (per CLAUDE.md): recurring finding patterns → below on the same delta (the SAST is a deterministic floor, the reasoned pass covers what grep/rules miss): ``` +Skill(effort-medium) # effort-shift: dispatch span starts Agent(subagent_type="security-auditor", description="audit-delta security — semgrep SAST", prompt="MODE: audit\nSCOPE: \nREPORT: .claude/audits/.audit-delta-semgrep.md\nFollow agents/security-auditor.md exactly. Pinned rulesets, no login. Write ONLY to REPORT. End with REPORT_WRITTEN: .") ``` diff --git a/skills/bugfix/SKILL.md b/skills/bugfix/SKILL.md index 7366e07..4dfa490 100644 --- a/skills/bugfix/SKILL.md +++ b/skills/bugfix/SKILL.md @@ -26,6 +26,7 @@ allowed-tools: MODEL GATE (blocking): run `$HOME/.claude/lib/model-gate.md` BEFORE any step below. Verdict `small` → STOP — print the gate's remedy, end the turn, dispatch nothing. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation. ## REQUEST $ARGUMENTS @@ -123,6 +124,7 @@ RISK: bug report left open → one batch of questions, before STEP 3b. The trivial fast-path is not exempt: a 1-line fix with a visible choice still asks. +`Skill(effort-high)` first (effort-shift: reflection resumes). ## STEP 3b — CHALLENGE THE FIX PLAN (before the contract) Unless the fix is the trivial 1-2 line case STEP 3 already fast-paths, the DIAGNOSIS + FIX PLAN is a reflection worth attacking before it hardens into a @@ -157,6 +159,7 @@ branch it's a no-op (commit in place). Never `finish`. Dispatch the executor — sonnet by frontmatter pin, do not override: ``` +Skill(effort-medium) # effort-shift: dispatch span starts Agent(subagent_type="bugfixer") prompt: "CONTRACT: DIAGNOSIS: @@ -278,6 +281,8 @@ A bugfix with an understood root cause is almost always worth one entry: If the bug was trivial and the root cause not transferable → skip with `CAPITALIZE: trivial, skip`. +`Skill(effort-low)` first (effort-shift: bookkeeping tail). + **Then commit the memory** — follow `$HOME/.claude/lib/capitalize-commit.md`: it surgically commits what capitalize just wrote (`.claude/memory` + `.claude/tasks` only, never `git add -A`) as one `chore(memory)` commit, reports the memory-commit diff --git a/skills/code-clean/SKILL.md b/skills/code-clean/SKILL.md index 30079f2..ec1dcea 100644 --- a/skills/code-clean/SKILL.md +++ b/skills/code-clean/SKILL.md @@ -26,6 +26,7 @@ allowed-tools: MODEL GATE (blocking): run `$HOME/.claude/lib/model-gate.md` BEFORE any step below. Verdict `small` → STOP — print the gate's remedy, end the turn, dispatch nothing. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation. ## TARGET $ARGUMENTS @@ -120,6 +121,7 @@ TOTALS: If no issues found: report clean state and stop. +`Skill(effort-xhigh)` first (effort-shift: reflection resumes). ## STEP 3b — CHALLENGE THE SCOPE (before approval) The STEP 3 report is the proposed cleanup scope — worth attacking before the human approves it. It is still inline, so FIRST persist it to @@ -173,6 +175,7 @@ is approved, stop — no dispatch. 2. **Dispatch the executor** — sonnet by frontmatter pin, do not override: ``` +Skill(effort-medium) # effort-shift: dispatch span starts Agent(subagent_type="code-cleaner") prompt: "SCOPE: .claude/audits/CODE-CLEAN-SCOPE.md APPROVED: diff --git a/skills/feat/SKILL.md b/skills/feat/SKILL.md index b013c0a..b39a042 100644 --- a/skills/feat/SKILL.md +++ b/skills/feat/SKILL.md @@ -26,6 +26,7 @@ allowed-tools: MODEL GATE (blocking): run `$HOME/.claude/lib/model-gate.md` BEFORE any step below. Verdict `small` → STOP — print the gate's remedy, end the turn, dispatch nothing. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation. ## REQUEST $ARGUMENTS @@ -122,6 +123,7 @@ request left open → one batch of questions BEFORE dispatching; answers land in the contract's CLARIFICATIONS `[gated]` and in the plan. A choice that surfaces only during execution comes back as `NEED-DECISION` (STEP 3). +`Skill(effort-high)` first (effort-shift: reflection resumes). ## STEP 1b — CHALLENGE THE PLAN (before branching) The STEP 1 plan is a reflection worth attacking before a branch is spent on it. Persist it to `.claude/tasks/plans/--.md`, then run @@ -144,6 +146,7 @@ branch it's a no-op (commit in place). Never `finish`. Dispatch the executor — sonnet by frontmatter pin, do not override: ``` +Skill(effort-medium) # effort-shift: dispatch span starts Agent(subagent_type="feater") prompt: "CONTRACT: PLAN: @@ -200,6 +203,7 @@ test), consider splitting into 2-3 atomic commits grouped by logical unit — or run `/commit-change` on the pending work (it dispatches the commit-changer (propose opus / apply sonnet, BDR-077); never inline-load the bare agent, it is now a propose/apply executor). +Then `Skill(effort-high)` (effort-shift: nested commit-change loaded at low; reload feat's level). Print summary: ``` @@ -250,6 +254,8 @@ Always append a 1-line entry to today's heading in `.claude/memory/journal.md`. If no substantive capture candidate → skip with `CAPITALIZE: nothing to log`. +`Skill(effort-low)` first (effort-shift: bookkeeping tail). + **Then commit the memory** — follow `$HOME/.claude/lib/capitalize-commit.md`: it surgically commits what capitalize just wrote (`.claude/memory` + `.claude/tasks` only, never `git add -A`) as one `chore(memory)` commit, reports the memory-commit diff --git a/skills/geo/SKILL.md b/skills/geo/SKILL.md index 3ed2600..d5b21bd 100644 --- a/skills/geo/SKILL.md +++ b/skills/geo/SKILL.md @@ -29,6 +29,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation. Dispatches the `geo-analyzer` subagent (audit + fix bundle), then applies the bundle from THIS main loop at **L1** — same shape as `/web-validate` @@ -46,6 +47,7 @@ every phase (LRN-126). Clean `.audit/geo-signals-.md` after apply. **A — collect (sonnet):** ``` +Skill(effort-medium) # effort-shift: dispatch span starts Agent(subagent_type="geo-analyzer", model="sonnet") prompt: "MODE: collect RUNID: @@ -83,6 +85,7 @@ Do NOT apply any fix and do NOT dispatch any sub-agent — /geo applies your bundle." ``` +`Skill(effort-high)` first (effort-shift: reflection resumes). ## STEP 1b — CHALLENGE THE FIX BUNDLE (advisory, before apply) The analyzer returned a `## FIX BUNDLE` — worth attacking before any edit lands. **Skip if intervention mode = conservative** (nothing is applied). Else persist the @@ -115,6 +118,7 @@ intent, not header wording: **AUTO** = no-confirmation items (G1–G4/G6); For each AUTO item, dispatch its `applier` at L1, passing the item verbatim: ``` +Skill(effort-medium) # effort-shift: dispatch span starts Agent(subagent_type="hotfixer") # or "feater" per the item's applier prompt: ". diff --git a/skills/harden/SKILL.md b/skills/harden/SKILL.md index 5aad845..7277deb 100644 --- a/skills/harden/SKILL.md +++ b/skills/harden/SKILL.md @@ -29,6 +29,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation. This skill orchestrates a narrow-scope hardening audit: TLS + security headers + redirects + canonical + custom 404 + server configs. It @@ -260,6 +261,7 @@ seo-analyzer will run in parallel. Spawn a single seo-analyzer subagent with an explicit IN/OUT scope list. ``` +Skill(effort-medium) # effort-shift: dispatch span starts Agent( subagent_type="seo-analyzer", description="harden — narrow-scope web hardening audit", @@ -519,6 +521,7 @@ Extract the score and critical-alert count from `.claude/audits/HARDEN.md` for t --- +`Skill(effort-high)` first (effort-shift: reflection resumes). ## STEP 2b — CHALLENGE THE FIX BUNDLE (MODE=fix only, advisory) Skip if MODE=audit (no bundle exists). Else, before the STEP 3 gate, harden the bundle: extract the `## 8. Fix bundle` section from HARDEN.md to diff --git a/skills/hotfix/SKILL.md b/skills/hotfix/SKILL.md index ff11901..e5fa263 100644 --- a/skills/hotfix/SKILL.md +++ b/skills/hotfix/SKILL.md @@ -24,6 +24,7 @@ allowed-tools: MODEL GATE (blocking): run `$HOME/.claude/lib/model-gate.md` BEFORE any step below. Verdict `small` → STOP — print the gate's remedy, end the turn, dispatch nothing. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation. ## REQUEST $ARGUMENTS @@ -93,6 +94,7 @@ point. Run it ONLY when the settled fix touches control flow or behaviour — an off-by-one, a wrong operator/variable, a behaviour-changing config value, or a missing import that alters execution. In doubt → it is probably a `/bugfix`. +`Skill(effort-high)` first (effort-shift: reflection resumes). For a logic fix: persist the STEP 1 located fix (root cause + the exact edit) to `.claude/tasks/plans/--.md`, then run `$HOME/.claude/lib/challenge-plan.md` with `PLAN` = that file, `KIND` = @@ -136,6 +138,7 @@ mentioned: STOP and ask `"working tree dirty: stash and continue, or abort?"`. Dispatch the executor — sonnet by frontmatter pin, do not override: ``` +Skill(effort-medium) # effort-shift: dispatch span starts Agent(subagent_type="hotfixer") prompt: "CONTRACT: LOCATED: @@ -233,6 +236,8 @@ Always append a 1-line entry to today's heading in `.claude/memory/journal.md` ( **Language rule**: the journal line and any proposed BLK/LRN entries are ALWAYS written English AND caveman — fragments, articles dropped, code/IDs/quoted errors verbatim — per CLAUDE.md "Memory registries" (Always English, always caveman). +`Skill(effort-low)` first (effort-shift: bookkeeping tail). + **Then commit the memory** — follow `$HOME/.claude/lib/capitalize-commit.md`: it surgically commits what capitalize just wrote (`.claude/memory` + `.claude/tasks` only, never `git add -A`) as one `chore(memory)` commit, reports the memory-commit diff --git a/skills/init-project/SKILL.md b/skills/init-project/SKILL.md index 069779a..48fb483 100644 --- a/skills/init-project/SKILL.md +++ b/skills/init-project/SKILL.md @@ -14,6 +14,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation. ## REQUEST $ARGUMENTS @@ -183,6 +184,7 @@ implemented on a `feature/*` branch off `develop` (STEP 8). Invoke `writing-plans` (vendored superpowers skill) with BRIEF + skeleton. Granular tasks (2-5 min each), exact file paths, TDD: tests before code. +`Skill(effort-xhigh)` first (effort-shift: reflection resumes). ## STEP 6b — CHALLENGE THE PLAN (before the gate) Before the human sees the implementation plan, harden it. Run `$HOME/.claude/lib/challenge-plan.md` with `PLAN` = the plan STEP 6 wrote under @@ -209,6 +211,7 @@ Approve and start? (yes / request changes) Changes → back to STEP 6. Approved → continue. ## STEP 8 — IMPLEMENT +First: `Skill(effort-medium)` (effort-shift: dispatch span starts). Start the MVP feature branch off develop, then implement on it: ```bash bash "$HOME/.claude/lib/gitflow.sh" start feature mvp @@ -311,6 +314,8 @@ articles dropped, code/IDs/quoted errors verbatim — per CLAUDE.md "Memory registries" (Always English, always caveman). The gate may mirror the user's language; entries must not. +`Skill(effort-low)` first (effort-shift: bookkeeping tail). + **Then commit the memory** — follow `$HOME/.claude/lib/capitalize-commit.md`: it surgically commits the approved founding decisions (`.claude/memory` + `.claude/tasks` only, never `git add -A`) as one `chore(memory)` commit, BEFORE diff --git a/skills/onboard/SKILL.md b/skills/onboard/SKILL.md index eba468d..b89a83b 100644 --- a/skills/onboard/SKILL.md +++ b/skills/onboard/SKILL.md @@ -14,6 +14,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation. ## REQUEST $ARGUMENTS @@ -91,6 +92,7 @@ STOP. La réponse détermine si STEP 1 tourne une fois (A) ou N fois (C) ou avec ## STEP 2 — BASELINE CONFIG (onboarder agent) +`Skill(effort-medium)` first (effort-shift: dispatch span starts). Dispatch `Agent(subagent_type="onboarder")` (pin sonnet — BDR-077 : config templating = exécution, plus jamais inline sur le modèle de session). Un BLOCAGE (clé manquante, CLAUDE.md existant) revient en rapport — l'agent ne @@ -890,6 +892,7 @@ Vérifier que les 4 fichiers `.claude/audits/ONBOARD_REPORT.md`, `.claude/audits --- +`Skill(effort-xhigh)` first (effort-shift: reflection resumes). ## STEP 7b — CHALLENGE THE PROPOSALS (before the human gate) The 4 audit files are on disk; `AUDIT_PROPOSALS.md` is the artifact worth attacking before the human spends a gate on it. Run diff --git a/skills/seo/SKILL.md b/skills/seo/SKILL.md index a4cf440..03f8115 100644 --- a/skills/seo/SKILL.md +++ b/skills/seo/SKILL.md @@ -30,6 +30,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation. This skill orchestrates TWO specialist agents running in parallel, then merges their output into a single `.claude/audits/SEO.md` report. It is the main @@ -324,6 +325,7 @@ templating. **PHASE A — collect (both domains, one message):** ``` +Skill(effort-medium) # effort-shift: dispatch span starts Agent(subagent_type="seo-analyzer", model="sonnet") prompt: """ MODE: collect @@ -507,6 +509,7 @@ write GEO.md/SEO.md — /seo applies your bundle in STEP 1.5 and merges the reports." ``` +`Skill(effort-high)` first (effort-shift: reflection resumes). ## STEP 1b — CHALLENGE THE FIX BUNDLE (advisory, before apply) Both envelopes now carry a `## FIX BUNDLE` — worth attacking before any edit lands. **Skip if intervention mode = conservative** (nothing is applied). Else persist both @@ -555,6 +558,7 @@ The two bundles may touch the same shared template (meta vs JSON-LD). Apply For each AUTO item, dispatch its `applier` at L1, passing the item verbatim: ``` +Skill(effort-medium) # effort-shift: dispatch span starts Agent(subagent_type="hotfixer") # or "feater" per the item's applier prompt: ". diff --git a/skills/ship-feature/SKILL.md b/skills/ship-feature/SKILL.md index 76bc316..04aedbc 100644 --- a/skills/ship-feature/SKILL.md +++ b/skills/ship-feature/SKILL.md @@ -14,6 +14,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation. ## REQUEST $ARGUMENTS @@ -123,6 +124,7 @@ every VISIBLE / PUBLIC NAME / SCOPE choice the plan settles that neither the request nor the STEP 1 brainstorm settled (check the contract's CLARIFICATIONS first) → one batch before STEP 2b; answers append to the contract `[gated]`. +`Skill(effort-xhigh)` first (effort-shift: reflection resumes). ## STEP 2b — CHALLENGE THE PLAN (adversarial, before the gate) Before the human sees the plan, harden it. Run `$HOME/.claude/lib/challenge-plan.md`: - `PLAN` = the plan STEP 2 wrote under `docs/superpowers/plans/` @@ -170,6 +172,7 @@ judges the diff against this ENRICHED contract, not the STEP 0e seed — so a criterion the design introduced is verified, not lost. ## STEP 4 — IMPLEMENT +First: `Skill(effort-medium)` (effort-shift: dispatch span starts). Start the feature branch off develop, then implement on it: ```bash bash "$HOME/.claude/lib/gitflow.sh" start feature @@ -268,6 +271,8 @@ Feature shipped implies at least one design decision worth capturing. Run this B If nothing substantive to log → print `CAPITALIZE: nothing substantive to log` and skip. +`Skill(effort-low)` first (effort-shift: bookkeeping tail). + **Then commit the memory** — follow `$HOME/.claude/lib/capitalize-commit.md`: it surgically commits what capitalize just wrote (`.claude/memory` + `.claude/tasks` only, never `git add -A`) as one `chore(memory)` commit, reports the memory-commit diff --git a/skills/tour/SKILL.md b/skills/tour/SKILL.md index b19dfc1..b2ed14a 100644 --- a/skills/tour/SKILL.md +++ b/skills/tour/SKILL.md @@ -30,6 +30,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation. One pipeline per project: **security → clean → re-verify → reconcile → doc → convergence re-audit**, looping until a full pass applies zero new @@ -86,6 +87,7 @@ Model discipline (the user-fixed invariant behind this mode): Runner dispatch, one per project: ``` +Skill(effort-medium) # effort-shift: dispatch span starts Agent(subagent_type="general-purpose", description="tour runner — ", prompt="Read ~/.claude/skills/tour/SKILL.md and execute STEP 1 → STEP 3 diff --git a/skills/web-validate/SKILL.md b/skills/web-validate/SKILL.md index 1fa34b6..0b886ce 100644 --- a/skills/web-validate/SKILL.md +++ b/skills/web-validate/SKILL.md @@ -28,6 +28,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation. This skill orchestrates a narrow-scope standards audit : @@ -179,6 +180,7 @@ Spawn a single `validator-analyzer` subagent with explicit scope and collected context : ``` +Skill(effort-medium) # effort-shift: dispatch span starts Agent( subagent_type="validator-analyzer", description="validate — W3C HTML + CSS + WCAG audit", @@ -252,6 +254,7 @@ grep -c '^### \[Critique\]' .claude/audits/VALIDATE.md --- +`Skill(effort-high)` first (effort-shift: reflection resumes). ## STEP 2b — CHALLENGE THE FIX BUNDLE (MODE=fix only, advisory) Skip if MODE=audit (no bundle exists). Else, before the STEP 3 gate, harden the bundle: extract the `## 5. Fix bundle` section from VALIDATE.md to @@ -310,6 +313,7 @@ Options : share files: ``` +Skill(effort-medium) # effort-shift: dispatch span starts Agent(subagent_type="hotfixer") prompt: ". From a117e7ed678101c09f731ff79d3216fbe7696e9f Mon Sep 17 00:00:00 2001 From: bastien Date: Mon, 28 Sep 2026 19:55:47 +0200 Subject: [PATCH 13/23] fix(effort): shifts are sent with the step's first tool call (harness pairing rule); challenge shifts under their heading; fence indentation --- agents/client-handover-writer.md | 4 ++-- lib/effort-shift.md | 13 ++++++++++++- lib/tests/effort-routing.test.sh | 5 +++++ skills/audit-delta/SKILL.md | 6 +++--- skills/bugfix/SKILL.md | 8 ++++---- skills/code-clean/SKILL.md | 6 +++--- skills/feat/SKILL.md | 10 +++++----- skills/geo/SKILL.md | 8 ++++---- skills/harden/SKILL.md | 6 +++--- skills/hotfix/SKILL.md | 8 ++++---- skills/init-project/SKILL.md | 8 ++++---- skills/onboard/SKILL.md | 6 +++--- skills/seo/SKILL.md | 8 ++++---- skills/ship-feature/SKILL.md | 8 ++++---- skills/tour/SKILL.md | 4 ++-- skills/web-validate/SKILL.md | 8 ++++---- 16 files changed, 66 insertions(+), 50 deletions(-) diff --git a/agents/client-handover-writer.md b/agents/client-handover-writer.md index 125220b..0e0e204 100644 --- a/agents/client-handover-writer.md +++ b/agents/client-handover-writer.md @@ -225,7 +225,7 @@ Store `DEPLOYED_URL` for STEP 7. If empty, ask user during STEP 6. --- ## STEP 3 — BASELINE AUDITS (parallel) -First: `Skill(effort-medium)` (effort-shift: dispatch span starts). +First: `Skill(effort-medium)` (effort-shift: dispatch span starts; send it in the same message as this step's first dispatch). Goal: capture `SCORE_*_BEFORE` so the client doc shows the delta. @@ -262,7 +262,7 @@ the gate. this pipeline (initial audits, fix-loop re-dispatches, commit-change, web-validate) carries `model: "fable"` — the child hosts gated orchestration on the pipeline's behalf; it must never inherit the session model. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. For web projects, dispatch in **a single message with two parallel Agent calls** (each with `model: "fable"`): diff --git a/lib/effort-shift.md b/lib/effort-shift.md index 71be69b..85d823c 100644 --- a/lib/effort-shift.md +++ b/lib/effort-shift.md @@ -8,6 +8,16 @@ max (stuck error, judged need). ## Mechanics (verified on Claude Code 2.1.283) +- **Pairing rule**: a `Skill(effort-)` call applies its effort only + when the same assistant message carries at least one other tool call + after it; a lone Skill call is a no-op. Send the shift together with the + step's first tool call, shift first. That paired call already runs at the + new level: pair a downward shift with a pinned-agent dispatch or a + Read/Bash, never with a built-in judgment dispatch (`general-purpose`, + `model: "opus"`), which would inherit it. +- Re-loading a shifter already loaded in the conversation re-applies its + effort (the harness only dedupes the skill text), so bounce-back + sequences such as medium → max → medium work. - A skill's `effort:` frontmatter applies from the moment it loads to the end of the turn: on the user's `/skill` and on a `Skill(...)` call by Claude in an interactive session. Last loaded wins, both directions. The @@ -21,7 +31,8 @@ max (stuck error, judged need). ## Shifters `Skill(effort-low)` · `Skill(effort-medium)` · `Skill(effort-high)` · -`Skill(effort-xhigh)` · `Skill(effort-max)`. One tool call, one-line body. +`Skill(effort-xhigh)` · `Skill(effort-max)`. One tool call, one-line body, +always sent with another tool call (Pairing rule). Typed by the user, `/effort-max` is a turn-scoped max: the relaunch lever after a STOP. `ultrathink` only adds an in-context nudge; the API level does not move. diff --git a/lib/tests/effort-routing.test.sh b/lib/tests/effort-routing.test.sh index c0d5bac..6dc7391 100755 --- a/lib/tests/effort-routing.test.sh +++ b/lib/tests/effort-routing.test.sh @@ -73,6 +73,11 @@ for s in seo geo harden web-validate; do has "skills/$s/SKILL.md" 'Skill(effort- for s in feat hotfix bugfix ship-feature init-project; do has "skills/$s/SKILL.md" 'Skill(effort-low)'; done has "skills/feat/SKILL.md" 'effort-shift: nested commit-change' +# ── 6b) pairing rule documented (R11) +has "lib/effort-shift.md" 'lone Skill call is a no-op' +has "lib/effort-shift.md" 're-applies its' +[ "$(grep -c 'a lone Skill call is a no-op' "$R/skills/feat/SKILL.md")" -ge 1 ] && ok || ko "feat INC line must carry the pairing rule" + # ── summary (later tasks insert their locks ABOVE this line) printf 'effort-routing census: %d pass, %d fail\n' "$pass" "$fail" [ "$fail" -eq 0 ] diff --git a/skills/audit-delta/SKILL.md b/skills/audit-delta/SKILL.md index 1443c91..3cbda45 100644 --- a/skills/audit-delta/SKILL.md +++ b/skills/audit-delta/SKILL.md @@ -29,7 +29,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. Audit only what changed since the last run, on the axes the user picks. Per axis: **audit → approval gate → fix → re-verify → marker update**, @@ -170,7 +170,7 @@ Then show the user the same compact table inline. ### 3b-bis. CHALLENGE THE PROPOSALS (before the gate) -`Skill(effort-xhigh)` first (effort-shift: reflection resumes). +`Skill(effort-xhigh)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). This axis' findings + proposed fixes are a proposal set worth attacking before the human gate. Persist THIS axis' finding list (not the whole append-only report) to `.claude/tasks/plans/--.md`, then run @@ -256,7 +256,7 @@ Then offer to capitalize (per CLAUDE.md): recurring finding patterns → below on the same delta (the SAST is a deterministic floor, the reasoned pass covers what grep/rules miss): ``` -Skill(effort-medium) # effort-shift: dispatch span starts + Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message Agent(subagent_type="security-auditor", description="audit-delta security — semgrep SAST", prompt="MODE: audit\nSCOPE: \nREPORT: .claude/audits/.audit-delta-semgrep.md\nFollow agents/security-auditor.md exactly. Pinned rulesets, no login. Write ONLY to REPORT. End with REPORT_WRITTEN: .") ``` diff --git a/skills/bugfix/SKILL.md b/skills/bugfix/SKILL.md index 4dfa490..af186a5 100644 --- a/skills/bugfix/SKILL.md +++ b/skills/bugfix/SKILL.md @@ -26,7 +26,7 @@ allowed-tools: MODEL GATE (blocking): run `$HOME/.claude/lib/model-gate.md` BEFORE any step below. Verdict `small` → STOP — print the gate's remedy, end the turn, dispatch nothing. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. ## REQUEST $ARGUMENTS @@ -124,8 +124,8 @@ RISK: bug report left open → one batch of questions, before STEP 3b. The trivial fast-path is not exempt: a 1-line fix with a visible choice still asks. -`Skill(effort-high)` first (effort-shift: reflection resumes). ## STEP 3b — CHALLENGE THE FIX PLAN (before the contract) +`Skill(effort-high)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). Unless the fix is the trivial 1-2 line case STEP 3 already fast-paths, the DIAGNOSIS + FIX PLAN is a reflection worth attacking before it hardens into a contract. Persist it to `.claude/tasks/plans/--.md`, then run @@ -159,7 +159,7 @@ branch it's a no-op (commit in place). Never `finish`. Dispatch the executor — sonnet by frontmatter pin, do not override: ``` -Skill(effort-medium) # effort-shift: dispatch span starts +Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message Agent(subagent_type="bugfixer") prompt: "CONTRACT: DIAGNOSIS: @@ -281,7 +281,7 @@ A bugfix with an understood root cause is almost always worth one entry: If the bug was trivial and the root cause not transferable → skip with `CAPITALIZE: trivial, skip`. -`Skill(effort-low)` first (effort-shift: bookkeeping tail). +`Skill(effort-low)` first (effort-shift: bookkeeping tail; send it in the same message as the memory-commit command). **Then commit the memory** — follow `$HOME/.claude/lib/capitalize-commit.md`: it surgically commits what capitalize just wrote (`.claude/memory` + `.claude/tasks` diff --git a/skills/code-clean/SKILL.md b/skills/code-clean/SKILL.md index ec1dcea..0821ebb 100644 --- a/skills/code-clean/SKILL.md +++ b/skills/code-clean/SKILL.md @@ -26,7 +26,7 @@ allowed-tools: MODEL GATE (blocking): run `$HOME/.claude/lib/model-gate.md` BEFORE any step below. Verdict `small` → STOP — print the gate's remedy, end the turn, dispatch nothing. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. ## TARGET $ARGUMENTS @@ -121,8 +121,8 @@ TOTALS: If no issues found: report clean state and stop. -`Skill(effort-xhigh)` first (effort-shift: reflection resumes). ## STEP 3b — CHALLENGE THE SCOPE (before approval) +`Skill(effort-xhigh)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). The STEP 3 report is the proposed cleanup scope — worth attacking before the human approves it. It is still inline, so FIRST persist it to `.claude/tasks/plans/--.md` (STEP 3 report format, one item @@ -175,7 +175,7 @@ is approved, stop — no dispatch. 2. **Dispatch the executor** — sonnet by frontmatter pin, do not override: ``` -Skill(effort-medium) # effort-shift: dispatch span starts + Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message Agent(subagent_type="code-cleaner") prompt: "SCOPE: .claude/audits/CODE-CLEAN-SCOPE.md APPROVED: diff --git a/skills/feat/SKILL.md b/skills/feat/SKILL.md index b39a042..a5de32b 100644 --- a/skills/feat/SKILL.md +++ b/skills/feat/SKILL.md @@ -26,7 +26,7 @@ allowed-tools: MODEL GATE (blocking): run `$HOME/.claude/lib/model-gate.md` BEFORE any step below. Verdict `small` → STOP — print the gate's remedy, end the turn, dispatch nothing. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. ## REQUEST $ARGUMENTS @@ -123,8 +123,8 @@ request left open → one batch of questions BEFORE dispatching; answers land in the contract's CLARIFICATIONS `[gated]` and in the plan. A choice that surfaces only during execution comes back as `NEED-DECISION` (STEP 3). -`Skill(effort-high)` first (effort-shift: reflection resumes). ## STEP 1b — CHALLENGE THE PLAN (before branching) +`Skill(effort-high)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). The STEP 1 plan is a reflection worth attacking before a branch is spent on it. Persist it to `.claude/tasks/plans/--.md`, then run `$HOME/.claude/lib/challenge-plan.md` with `PLAN` = that file, `KIND` = `build-plan`, @@ -146,7 +146,7 @@ branch it's a no-op (commit in place). Never `finish`. Dispatch the executor — sonnet by frontmatter pin, do not override: ``` -Skill(effort-medium) # effort-shift: dispatch span starts +Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message Agent(subagent_type="feater") prompt: "CONTRACT: PLAN: @@ -203,7 +203,7 @@ test), consider splitting into 2-3 atomic commits grouped by logical unit — or run `/commit-change` on the pending work (it dispatches the commit-changer (propose opus / apply sonnet, BDR-077); never inline-load the bare agent, it is now a propose/apply executor). -Then `Skill(effort-high)` (effort-shift: nested commit-change loaded at low; reload feat's level). +Then `Skill(effort-high)` (effort-shift: nested commit-change loaded at low; reload feat's level, sent with the next tool call). Print summary: ``` @@ -254,7 +254,7 @@ Always append a 1-line entry to today's heading in `.claude/memory/journal.md`. If no substantive capture candidate → skip with `CAPITALIZE: nothing to log`. -`Skill(effort-low)` first (effort-shift: bookkeeping tail). +`Skill(effort-low)` first (effort-shift: bookkeeping tail; send it in the same message as the memory-commit command). **Then commit the memory** — follow `$HOME/.claude/lib/capitalize-commit.md`: it surgically commits what capitalize just wrote (`.claude/memory` + `.claude/tasks` diff --git a/skills/geo/SKILL.md b/skills/geo/SKILL.md index d5b21bd..c06c8f6 100644 --- a/skills/geo/SKILL.md +++ b/skills/geo/SKILL.md @@ -29,7 +29,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. Dispatches the `geo-analyzer` subagent (audit + fix bundle), then applies the bundle from THIS main loop at **L1** — same shape as `/web-validate` @@ -47,7 +47,7 @@ every phase (LRN-126). Clean `.audit/geo-signals-.md` after apply. **A — collect (sonnet):** ``` -Skill(effort-medium) # effort-shift: dispatch span starts +Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message Agent(subagent_type="geo-analyzer", model="sonnet") prompt: "MODE: collect RUNID: @@ -85,8 +85,8 @@ Do NOT apply any fix and do NOT dispatch any sub-agent — /geo applies your bundle." ``` -`Skill(effort-high)` first (effort-shift: reflection resumes). ## STEP 1b — CHALLENGE THE FIX BUNDLE (advisory, before apply) +`Skill(effort-high)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). The analyzer returned a `## FIX BUNDLE` — worth attacking before any edit lands. **Skip if intervention mode = conservative** (nothing is applied). Else persist the bundle verbatim to `.claude/tasks/plans/--.md`, then run @@ -118,7 +118,7 @@ intent, not header wording: **AUTO** = no-confirmation items (G1–G4/G6); For each AUTO item, dispatch its `applier` at L1, passing the item verbatim: ``` -Skill(effort-medium) # effort-shift: dispatch span starts +Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message Agent(subagent_type="hotfixer") # or "feater" per the item's applier prompt: ". diff --git a/skills/harden/SKILL.md b/skills/harden/SKILL.md index 7277deb..ad3ef56 100644 --- a/skills/harden/SKILL.md +++ b/skills/harden/SKILL.md @@ -29,7 +29,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. This skill orchestrates a narrow-scope hardening audit: TLS + security headers + redirects + canonical + custom 404 + server configs. It @@ -261,7 +261,7 @@ seo-analyzer will run in parallel. Spawn a single seo-analyzer subagent with an explicit IN/OUT scope list. ``` -Skill(effort-medium) # effort-shift: dispatch span starts +Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message Agent( subagent_type="seo-analyzer", description="harden — narrow-scope web hardening audit", @@ -521,8 +521,8 @@ Extract the score and critical-alert count from `.claude/audits/HARDEN.md` for t --- -`Skill(effort-high)` first (effort-shift: reflection resumes). ## STEP 2b — CHALLENGE THE FIX BUNDLE (MODE=fix only, advisory) +`Skill(effort-high)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). Skip if MODE=audit (no bundle exists). Else, before the STEP 3 gate, harden the bundle: extract the `## 8. Fix bundle` section from HARDEN.md to `.claude/tasks/plans/--.md` (a clean, blind-judgeable artifact), then run diff --git a/skills/hotfix/SKILL.md b/skills/hotfix/SKILL.md index e5fa263..f395cd5 100644 --- a/skills/hotfix/SKILL.md +++ b/skills/hotfix/SKILL.md @@ -24,7 +24,7 @@ allowed-tools: MODEL GATE (blocking): run `$HOME/.claude/lib/model-gate.md` BEFORE any step below. Verdict `small` → STOP — print the gate's remedy, end the turn, dispatch nothing. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. ## REQUEST $ARGUMENTS @@ -94,7 +94,7 @@ point. Run it ONLY when the settled fix touches control flow or behaviour — an off-by-one, a wrong operator/variable, a behaviour-changing config value, or a missing import that alters execution. In doubt → it is probably a `/bugfix`. -`Skill(effort-high)` first (effort-shift: reflection resumes). +`Skill(effort-high)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). For a logic fix: persist the STEP 1 located fix (root cause + the exact edit) to `.claude/tasks/plans/--.md`, then run `$HOME/.claude/lib/challenge-plan.md` with `PLAN` = that file, `KIND` = @@ -138,7 +138,7 @@ mentioned: STOP and ask `"working tree dirty: stash and continue, or abort?"`. Dispatch the executor — sonnet by frontmatter pin, do not override: ``` -Skill(effort-medium) # effort-shift: dispatch span starts +Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message Agent(subagent_type="hotfixer") prompt: "CONTRACT: LOCATED: @@ -236,7 +236,7 @@ Always append a 1-line entry to today's heading in `.claude/memory/journal.md` ( **Language rule**: the journal line and any proposed BLK/LRN entries are ALWAYS written English AND caveman — fragments, articles dropped, code/IDs/quoted errors verbatim — per CLAUDE.md "Memory registries" (Always English, always caveman). -`Skill(effort-low)` first (effort-shift: bookkeeping tail). +`Skill(effort-low)` first (effort-shift: bookkeeping tail; send it in the same message as the memory-commit command). **Then commit the memory** — follow `$HOME/.claude/lib/capitalize-commit.md`: it surgically commits what capitalize just wrote (`.claude/memory` + `.claude/tasks` diff --git a/skills/init-project/SKILL.md b/skills/init-project/SKILL.md index 48fb483..66e6c91 100644 --- a/skills/init-project/SKILL.md +++ b/skills/init-project/SKILL.md @@ -14,7 +14,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. ## REQUEST $ARGUMENTS @@ -184,8 +184,8 @@ implemented on a `feature/*` branch off `develop` (STEP 8). Invoke `writing-plans` (vendored superpowers skill) with BRIEF + skeleton. Granular tasks (2-5 min each), exact file paths, TDD: tests before code. -`Skill(effort-xhigh)` first (effort-shift: reflection resumes). ## STEP 6b — CHALLENGE THE PLAN (before the gate) +`Skill(effort-xhigh)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). Before the human sees the implementation plan, harden it. Run `$HOME/.claude/lib/challenge-plan.md` with `PLAN` = the plan STEP 6 wrote under `docs/superpowers/plans/`, `KIND` = `build-plan`, `SCOPE` = the skeleton + task file @@ -211,7 +211,7 @@ Approve and start? (yes / request changes) Changes → back to STEP 6. Approved → continue. ## STEP 8 — IMPLEMENT -First: `Skill(effort-medium)` (effort-shift: dispatch span starts). +First: `Skill(effort-medium)` (effort-shift: dispatch span starts; send it in the same message as this step's first dispatch). Start the MVP feature branch off develop, then implement on it: ```bash bash "$HOME/.claude/lib/gitflow.sh" start feature mvp @@ -314,7 +314,7 @@ articles dropped, code/IDs/quoted errors verbatim — per CLAUDE.md "Memory registries" (Always English, always caveman). The gate may mirror the user's language; entries must not. -`Skill(effort-low)` first (effort-shift: bookkeeping tail). +`Skill(effort-low)` first (effort-shift: bookkeeping tail; send it in the same message as the memory-commit command). **Then commit the memory** — follow `$HOME/.claude/lib/capitalize-commit.md`: it surgically commits the approved founding decisions (`.claude/memory` + diff --git a/skills/onboard/SKILL.md b/skills/onboard/SKILL.md index b89a83b..5d7f526 100644 --- a/skills/onboard/SKILL.md +++ b/skills/onboard/SKILL.md @@ -14,7 +14,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. ## REQUEST $ARGUMENTS @@ -92,7 +92,7 @@ STOP. La réponse détermine si STEP 1 tourne une fois (A) ou N fois (C) ou avec ## STEP 2 — BASELINE CONFIG (onboarder agent) -`Skill(effort-medium)` first (effort-shift: dispatch span starts). +`Skill(effort-medium)` first (effort-shift: dispatch span starts; send it in the same message as this step's first dispatch). Dispatch `Agent(subagent_type="onboarder")` (pin sonnet — BDR-077 : config templating = exécution, plus jamais inline sur le modèle de session). Un BLOCAGE (clé manquante, CLAUDE.md existant) revient en rapport — l'agent ne @@ -892,8 +892,8 @@ Vérifier que les 4 fichiers `.claude/audits/ONBOARD_REPORT.md`, `.claude/audits --- -`Skill(effort-xhigh)` first (effort-shift: reflection resumes). ## STEP 7b — CHALLENGE THE PROPOSALS (before the human gate) +`Skill(effort-xhigh)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). The 4 audit files are on disk; `AUDIT_PROPOSALS.md` is the artifact worth attacking before the human spends a gate on it. Run `$HOME/.claude/lib/challenge-plan.md` with `PLAN` = diff --git a/skills/seo/SKILL.md b/skills/seo/SKILL.md index 03f8115..ebace86 100644 --- a/skills/seo/SKILL.md +++ b/skills/seo/SKILL.md @@ -30,7 +30,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. This skill orchestrates TWO specialist agents running in parallel, then merges their output into a single `.claude/audits/SEO.md` report. It is the main @@ -325,7 +325,7 @@ templating. **PHASE A — collect (both domains, one message):** ``` -Skill(effort-medium) # effort-shift: dispatch span starts +Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message Agent(subagent_type="seo-analyzer", model="sonnet") prompt: """ MODE: collect @@ -509,8 +509,8 @@ write GEO.md/SEO.md — /seo applies your bundle in STEP 1.5 and merges the reports." ``` -`Skill(effort-high)` first (effort-shift: reflection resumes). ## STEP 1b — CHALLENGE THE FIX BUNDLE (advisory, before apply) +`Skill(effort-high)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). Both envelopes now carry a `## FIX BUNDLE` — worth attacking before any edit lands. **Skip if intervention mode = conservative** (nothing is applied). Else persist both bundles (seo + geo, verbatim) to `.claude/tasks/plans/--.md`, then run @@ -558,7 +558,7 @@ The two bundles may touch the same shared template (meta vs JSON-LD). Apply For each AUTO item, dispatch its `applier` at L1, passing the item verbatim: ``` -Skill(effort-medium) # effort-shift: dispatch span starts +Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message Agent(subagent_type="hotfixer") # or "feater" per the item's applier prompt: ". diff --git a/skills/ship-feature/SKILL.md b/skills/ship-feature/SKILL.md index 04aedbc..fb8e3a6 100644 --- a/skills/ship-feature/SKILL.md +++ b/skills/ship-feature/SKILL.md @@ -14,7 +14,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. ## REQUEST $ARGUMENTS @@ -124,8 +124,8 @@ every VISIBLE / PUBLIC NAME / SCOPE choice the plan settles that neither the request nor the STEP 1 brainstorm settled (check the contract's CLARIFICATIONS first) → one batch before STEP 2b; answers append to the contract `[gated]`. -`Skill(effort-xhigh)` first (effort-shift: reflection resumes). ## STEP 2b — CHALLENGE THE PLAN (adversarial, before the gate) +`Skill(effort-xhigh)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). Before the human sees the plan, harden it. Run `$HOME/.claude/lib/challenge-plan.md`: - `PLAN` = the plan STEP 2 wrote under `docs/superpowers/plans/` - `KIND` = `build-plan` @@ -172,7 +172,7 @@ judges the diff against this ENRICHED contract, not the STEP 0e seed — so a criterion the design introduced is verified, not lost. ## STEP 4 — IMPLEMENT -First: `Skill(effort-medium)` (effort-shift: dispatch span starts). +First: `Skill(effort-medium)` (effort-shift: dispatch span starts; send it in the same message as this step's first dispatch). Start the feature branch off develop, then implement on it: ```bash bash "$HOME/.claude/lib/gitflow.sh" start feature @@ -271,7 +271,7 @@ Feature shipped implies at least one design decision worth capturing. Run this B If nothing substantive to log → print `CAPITALIZE: nothing substantive to log` and skip. -`Skill(effort-low)` first (effort-shift: bookkeeping tail). +`Skill(effort-low)` first (effort-shift: bookkeeping tail; send it in the same message as the memory-commit command). **Then commit the memory** — follow `$HOME/.claude/lib/capitalize-commit.md`: it surgically commits what capitalize just wrote (`.claude/memory` + `.claude/tasks` diff --git a/skills/tour/SKILL.md b/skills/tour/SKILL.md index b2ed14a..8894645 100644 --- a/skills/tour/SKILL.md +++ b/skills/tour/SKILL.md @@ -30,7 +30,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. One pipeline per project: **security → clean → re-verify → reconcile → doc → convergence re-audit**, looping until a full pass applies zero new @@ -87,7 +87,7 @@ Model discipline (the user-fixed invariant behind this mode): Runner dispatch, one per project: ``` -Skill(effort-medium) # effort-shift: dispatch span starts +Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message Agent(subagent_type="general-purpose", description="tour runner — ", prompt="Read ~/.claude/skills/tour/SKILL.md and execute STEP 1 → STEP 3 diff --git a/skills/web-validate/SKILL.md b/skills/web-validate/SKILL.md index 0b886ce..cae8c15 100644 --- a/skills/web-validate/SKILL.md +++ b/skills/web-validate/SKILL.md @@ -28,7 +28,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. This skill orchestrates a narrow-scope standards audit : @@ -180,7 +180,7 @@ Spawn a single `validator-analyzer` subagent with explicit scope and collected context : ``` -Skill(effort-medium) # effort-shift: dispatch span starts +Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message Agent( subagent_type="validator-analyzer", description="validate — W3C HTML + CSS + WCAG audit", @@ -254,8 +254,8 @@ grep -c '^### \[Critique\]' .claude/audits/VALIDATE.md --- -`Skill(effort-high)` first (effort-shift: reflection resumes). ## STEP 2b — CHALLENGE THE FIX BUNDLE (MODE=fix only, advisory) +`Skill(effort-high)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). Skip if MODE=audit (no bundle exists). Else, before the STEP 3 gate, harden the bundle: extract the `## 5. Fix bundle` section from VALIDATE.md to `.claude/tasks/plans/--.md` (a clean, blind-judgeable artifact), then run @@ -313,7 +313,7 @@ Options : share files: ``` -Skill(effort-medium) # effort-shift: dispatch span starts + Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message Agent(subagent_type="hotfixer") prompt: ". From 557e4cc3172f03beca6ec9e08b0e9cd39a7ada89 Mon Sep 17 00:00:00 2001 From: bastien Date: Mon, 28 Sep 2026 20:02:56 +0200 Subject: [PATCH 14/23] feat(effort): max at the verify-secure caps and ship-feature 4b; STOP texts suggest /effort-max --- lib/challenge-plan.md | 2 ++ lib/tests/effort-routing.test.sh | 6 ++++++ lib/verify-secure-loop.md | 9 +++++---- skills/ship-feature/SKILL.md | 7 +++++-- 4 files changed, 18 insertions(+), 6 deletions(-) diff --git a/lib/challenge-plan.md b/lib/challenge-plan.md index 91e72ef..66d4cb1 100644 --- a/lib/challenge-plan.md +++ b/lib/challenge-plan.md @@ -59,6 +59,8 @@ silently downgrade the judgment. (The executor gates stay sonnet.) A challenger that returns a malformed/empty verdict, a missing `PROOF`, or dies → retry ONCE with a fresh challenger; a 2nd failure on that lens → STOP and escalate +(the STOP text names the level reached, `$CLAUDE_EFFORT`, and suggests `/effort-max` +for the relaunch; no shift here: a mute challenger is an infrastructure failure) to the human, NAMING the lens. Never carry "plan challenged" into the gate on a silently dropped lens (`verify-secure-loop.md`: "a mute verifier is NEVER a PASS"). diff --git a/lib/tests/effort-routing.test.sh b/lib/tests/effort-routing.test.sh index 6dc7391..536145e 100755 --- a/lib/tests/effort-routing.test.sh +++ b/lib/tests/effort-routing.test.sh @@ -78,6 +78,12 @@ has "lib/effort-shift.md" 'lone Skill call is a no-op' has "lib/effort-shift.md" 're-applies its' [ "$(grep -c 'a lone Skill call is a no-op' "$R/skills/feat/SKILL.md")" -ge 1 ] && ok || ko "feat INC line must carry the pairing rule" +# ── 7) escalation at max (spec D4) +[ "$(grep -c 'Skill(effort-max)' "$R/lib/verify-secure-loop.md")" -eq 3 ] && ok || ko "verify-secure-loop.md must shift to max at its 3 caps" +has "skills/ship-feature/SKILL.md" 'Skill(effort-max)' +has "lib/challenge-plan.md" '/effort-max' +has "lib/verify-secure-loop.md" '/effort-max' + # ── summary (later tasks insert their locks ABOVE this line) printf 'effort-routing census: %d pass, %d fail\n' "$pass" "$fail" [ "$fail" -eq 0 ] diff --git a/lib/verify-secure-loop.md b/lib/verify-secure-loop.md index 5bb4fd1..2540670 100644 --- a/lib/verify-secure-loop.md +++ b/lib/verify-secure-loop.md @@ -35,7 +35,7 @@ single `GATES — VERDICT:` line: - `UNMET(n)` → hand the dev the CONTRACT path + the `NOT-MET` rows verbatim, nothing else; re-run GATE 0. **No verifier is dispatched** — a red build or a red suite is not a judgement call, and paying an LLM to discover it is - waste. **Max 3 floor iterations** → STOP + human escalation with the rows. + waste. **Max 3 floor iterations** → `Skill(effort-max)` (effort-shift: cap reached, diagnose at max before escalating; send it in the same message as the first tool call that gathers the escalation evidence), then STOP + human escalation with the rows. - `ABANDONED(n)` → floor green but a handoff stands. Continue to GATE 1; the verifier surfaces it and its `ABANDONED(n)` verdict routes to the human gate. @@ -74,7 +74,7 @@ Parse its single `VERIFY — VERDICT:` line: lines (NOT-MET / out-of-scope), nothing else: re-dispatch a FRESH executor with those inputs only, never redo the fix by hand. Then re-run GATE 0 and re-dispatch a FRESH verifier. Repeat. - **Max 3 conformity iterations** → STOP + human escalation with the + **Max 3 conformity iterations** → `Skill(effort-max)` (effort-shift: cap reached, diagnose at max before escalating; send it in the same message as the first tool call that gathers the escalation evidence), then STOP + human escalation with the CRITERIA table (the contract-vs-realized diff). - `ABANDONED(n)` → direct human gate, never a dev loop (a dev cannot close what was proven impossible). The human lifts the abandonment or accepts @@ -104,8 +104,9 @@ Parse its single `SECURITY — VERDICT:` line: (re-dispatch a FRESH executor, never fix by hand). Then re-run GATE 0, then **re-verify the REQUEST first** (GATE 1, fresh verifier) — a security fix can drift the behavior — **then re-run GATE 2** (fresh auditor), in that - order. **Max 3 security iterations** → STOP + human escalation with the - BLOCKING table. + order. **Max 3 security iterations** → `Skill(effort-max)` (effort-shift: cap reached, diagnose at max before escalating; send it in the same message as the first tool call that gathers the escalation evidence), then STOP + human escalation with the + BLOCKING table. Every STOP text names the level reached (`$CLAUDE_EFFORT`) + and suggests `/effort-max` for the relaunch. - `DEGRADED` (semgrep absent) → does NOT block on the tool's absence; surface the checklist result + recommend `make plugin`. A DEGRADED run that still BLOCKs (grep-caught secret/injection) blocks like any other. diff --git a/skills/ship-feature/SKILL.md b/skills/ship-feature/SKILL.md index fb8e3a6..693d724 100644 --- a/skills/ship-feature/SKILL.md +++ b/skills/ship-feature/SKILL.md @@ -191,7 +191,8 @@ this loop. ## STEP 4b — ERROR RECOVERY (if STEP 4 fails) If a subagent returns a build error, failing test, or type error: -1. Load `$HOME/.claude/agents/analyzer.md` in DEBUG MODE on the exact error output. +1. `Skill(effort-max)` (effort-shift: error recovery; send it in the same message as the Read of the analyzer file below), then load + `$HOME/.claude/agents/analyzer.md` in DEBUG MODE on the exact error output. Produce: root cause hypotheses (ordered), affected files, what NOT to touch. 2. Present gate: ``` @@ -207,8 +208,10 @@ OPTIONS : C) Abort feature — preserve work done so far ``` 3. Wait for user choice. Do NOT auto-fix. Do NOT proceed without explicit approval. -4. If A → apply minimal fix, re-run STEP 4 for the failed task only. Max 2 retry attempts. +4. On resume the turn is at the session level (effort-shift: turn reset). + If A → `Skill(effort-medium)` sent with the re-dispatch, apply minimal fix, re-run STEP 4 for the failed task only. Max 2 retry attempts. If still failing after 2 → fall back to options B or C. + If B or C → `Skill(effort-xhigh)` first, sent with the next tool call. If B → before skipping: scan remaining task list for tasks that depend on the failed task (look for references to the same file or function in subsequent tasks). If dependents found → present: "Tasks [N, M] depend on the skipped task. From dd9488964b2354c289bf62e968cdc69ecb8d9d93 Mon Sep 17 00:00:00 2001 From: bastien Date: Mon, 28 Sep 2026 20:04:11 +0200 Subject: [PATCH 15/23] feat(effort): re-assert the skill level after prose gates that end the turn --- lib/tests/effort-routing.test.sh | 3 +++ skills/bugfix/SKILL.md | 1 + 2 files changed, 4 insertions(+) diff --git a/lib/tests/effort-routing.test.sh b/lib/tests/effort-routing.test.sh index 536145e..b3ef678 100755 --- a/lib/tests/effort-routing.test.sh +++ b/lib/tests/effort-routing.test.sh @@ -84,6 +84,9 @@ has "skills/ship-feature/SKILL.md" 'Skill(effort-max)' has "lib/challenge-plan.md" '/effort-max' has "lib/verify-secure-loop.md" '/effort-max' +# ── 8) turn-reset re-assert after a prose gate followed by reflection +has "skills/bugfix/SKILL.md" 'effort-shift: turn reset' + # ── summary (later tasks insert their locks ABOVE this line) printf 'effort-routing census: %d pass, %d fail\n' "$pass" "$fail" [ "$fail" -eq 0 ] diff --git a/skills/bugfix/SKILL.md b/skills/bugfix/SKILL.md index af186a5..9e54f09 100644 --- a/skills/bugfix/SKILL.md +++ b/skills/bugfix/SKILL.md @@ -119,6 +119,7 @@ RISK: obvious fix. - If the fix is significant (>10 lines, multiple files, behavior change): wait for user approval. + On resume: `Skill(effort-high)` first, sent with the next tool call (effort-shift: turn reset). - Then run pass B of `$HOME/.claude/lib/contract-interview.md` against the FIX PLAN: every VISIBLE / PUBLIC NAME / SCOPE choice it settles that the bug report left open → one batch of questions, before STEP 3b. The trivial From 1e3339358ce755f49c8c7dcf0bf1f2ac70069a5d Mon Sep 17 00:00:00 2001 From: bastien Date: Mon, 28 Sep 2026 20:14:00 +0200 Subject: [PATCH 16/23] feat(effort): transcript audit script for the thinking/cost split --- lib/effort-audit.py | 88 ++++++++++++++++++++++++++++++++ lib/effort-shift.md | 3 ++ lib/tests/effort-routing.test.sh | 4 ++ 3 files changed, 95 insertions(+) create mode 100755 lib/effort-audit.py diff --git a/lib/effort-audit.py b/lib/effort-audit.py new file mode 100755 index 0000000..2b74d13 --- /dev/null +++ b/lib/effort-audit.py @@ -0,0 +1,88 @@ +#!/usr/bin/env python3 +"""Sum output/thinking/cache tokens per (scope, model, effort) over Claude Code +transcripts. scope = main (session jsonl) | sub (subagents/*.jsonl or +isSidechain records). Read-only. Usage: effort-audit.py [projects-root]""" +import collections +import glob +import json +import os +import sys + +WEIGHTS = {"in": 1.0, "cc": 1.25, "cr": 0.1, "out": 5.0} # relative to input price +FIELDS = ("in", "cc", "cr", "out", "think") + + +def usage_row(usage): + """Map one API usage block to the five counted fields.""" + details = usage.get("output_tokens_details") or {} + return { + "in": usage.get("input_tokens", 0) or 0, + "cc": usage.get("cache_creation_input_tokens", 0) or 0, + "cr": usage.get("cache_read_input_tokens", 0) or 0, + "out": usage.get("output_tokens", 0) or 0, + "think": details.get("thinking_tokens", 0) or 0, + } + + +def scan(path, scope, agg): + """Add every assistant record of one transcript to agg.""" + with open(path, errors="ignore") as handle: + for line in handle: + try: + rec = json.loads(line) + except ValueError: + continue + msg = rec.get("message") or {} + if rec.get("type") != "assistant" or not msg.get("usage"): + continue + sub = scope == "sub" or bool(rec.get("isSidechain")) + key = ("sub" if sub else "main", + str(msg.get("model", "?")).replace("claude-", ""), + str(rec.get("effort") or "?")) + row = usage_row(msg["usage"]) + agg[key]["msgs"] += 1 + for field in FIELDS: + agg[key][field] += row[field] + + +def weighted(counter): + return sum(counter[f] * WEIGHTS[f] for f in WEIGHTS) + + +def report(agg): + """Print the per-key table, then the main/sub split and the thinking share.""" + total = collections.Counter() + for counter in agg.values(): + total.update(counter) + total_w = weighted(total) or 1 + print(f"{'scope':5} {'model':22} {'effort':7} {'msgs':>6} {'think/msg':>9} " + f"{'think_tok':>10} {'out_tok':>10} {'cache_read':>12} {'%wcost':>7}") + for (scope, model, effort), c in sorted(agg.items(), key=lambda kv: -weighted(kv[1])): + per_msg = c["think"] / max(c["msgs"], 1) + print(f"{scope:5} {model:22} {effort:7} {c['msgs']:6d} {per_msg:9.0f} " + f"{c['think']:10d} {c['out']:10d} {c['cr']:12d} {100 * weighted(c) / total_w:6.1f}%") + by_scope = collections.defaultdict(collections.Counter) + for (scope, _, _), c in agg.items(): + by_scope[scope].update(c) + for scope, c in by_scope.items(): + print(f" {scope:5} weighted-cost {100 * weighted(c) / total_w:5.1f}% " + f"thinking {100 * c['think'] / max(total['think'], 1):5.1f}% requests {c['msgs']}") + print(f" thinking = {100 * total['think'] * WEIGHTS['out'] / total_w:.1f}% of weighted cost; " + f"cache reads = {100 * total['cr'] * WEIGHTS['cr'] / total_w:.1f}%") + + +def main(): + root = os.path.expanduser(sys.argv[1] if len(sys.argv) > 1 else "~/.claude/projects") + agg = collections.defaultdict(collections.Counter) + for project in sorted(glob.glob(os.path.join(root, "*"))): + if not os.path.isdir(project): + continue + for path in glob.glob(os.path.join(project, "*.jsonl")): + scan(path, "main", agg) + for path in glob.glob(os.path.join(project, "*", "subagents", "*.jsonl")): + scan(path, "sub", agg) + report(agg) + + +if __name__ == "__main__": + main() diff --git a/lib/effort-shift.md b/lib/effort-shift.md index 85d823c..3d07925 100644 --- a/lib/effort-shift.md +++ b/lib/effort-shift.md @@ -28,6 +28,9 @@ max (stuck error, judged need). the run stays at the session level. `CLAUDE_CODE_EFFORT_LEVEL` beats every frontmatter; keep it unset (the session banner warns). +Measure the split any time: `python3 ~/.claude/lib/effort-audit.py` +(thinking/output/cache tokens per scope, model and effort). + ## Shifters `Skill(effort-low)` · `Skill(effort-medium)` · `Skill(effort-high)` · diff --git a/lib/tests/effort-routing.test.sh b/lib/tests/effort-routing.test.sh index b3ef678..6128221 100755 --- a/lib/tests/effort-routing.test.sh +++ b/lib/tests/effort-routing.test.sh @@ -87,6 +87,10 @@ has "lib/verify-secure-loop.md" '/effort-max' # ── 8) turn-reset re-assert after a prose gate followed by reflection has "skills/bugfix/SKILL.md" 'effort-shift: turn reset' +# ── 11) audit tooling +has "lib/effort-shift.md" 'effort-audit.py' +[ -x "$R/lib/effort-audit.py" ] && ok || ko "lib/effort-audit.py missing or not executable" + # ── summary (later tasks insert their locks ABOVE this line) printf 'effort-routing census: %d pass, %d fail\n' "$pass" "$fail" [ "$fail" -eq 0 ] From 98ef991958032d4301c634254b958b82f06800cc Mon Sep 17 00:00:00 2001 From: bastien Date: Mon, 28 Sep 2026 20:20:05 +0200 Subject: [PATCH 17/23] docs(effort): BDR-107 id, CHANGELOG entry, spec corrected for the rulings (vendored pins, exclusions, pairing rule) --- CHANGELOG.md | 1 + agents/client-handover-writer.md | 2 +- .../specs/2026-09-28-effort-tiering-design.md | 27 ++++++++++++------- hooks/session-start.sh | 2 +- hooks/statusline.sh | 2 +- install-plugins.sh | 2 +- lib/effort-shift.md | 2 +- lib/model-gate.md | 2 +- lib/tests/effort-routing.test.sh | 2 +- skills/audit-delta/SKILL.md | 2 +- skills/bugfix/SKILL.md | 2 +- skills/code-clean/SKILL.md | 2 +- skills/feat/SKILL.md | 2 +- skills/geo/SKILL.md | 2 +- skills/harden/SKILL.md | 2 +- skills/hotfix/SKILL.md | 2 +- skills/init-project/SKILL.md | 2 +- skills/onboard/SKILL.md | 2 +- skills/seo/SKILL.md | 2 +- skills/ship-feature/SKILL.md | 2 +- skills/tour/SKILL.md | 2 +- skills/web-validate/SKILL.md | 2 +- 22 files changed, 39 insertions(+), 29 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 29876a7..867c199 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -7,6 +7,7 @@ Format follows [Keep a Changelog](https://keepachangelog.com/). ## [Unreleased] ### Added +- **Effort tiering (BDR-107)**: reasoning effort routed per role and per phase. Session default `high`; `effort:` pins on the 20 repo-authored agents; entry level on 28 tracked user-invoked skills plus the two vendored superpowers skills (re-applied by `install-plugins.sh` after resync); five shifter skills `effort-low` … `effort-max` loaded at phase boundaries per `lib/effort-shift.md`, always sent with the step's first tool call (a lone Skill call is a no-op on 2.1.283), with `max` at the verify-secure caps and ship-feature 4b; `/effort-max` as the turn-scoped relaunch lever; statusline shows the live level; session banner warns when `CLAUDE_CODE_EFFORT_LEVEL` silences the pins; census `lib/tests/effort-routing.test.sh`; transcript audit `lib/effort-audit.py`. - **Design gate asks the user to sign in to 21st instead of skipping it**: `lib/design-tool-gate.sh` adds a three-state 21st auth predicate (`twentyfirst_auth_state`, honors `TWENTYFIRST_TOKEN`/`API_KEY_21ST` or a diff --git a/agents/client-handover-writer.md b/agents/client-handover-writer.md index 0e0e204..13d32e9 100644 --- a/agents/client-handover-writer.md +++ b/agents/client-handover-writer.md @@ -262,7 +262,7 @@ the gate. this pipeline (initial audits, fix-loop re-dispatches, commit-change, web-validate) carries `model: "fable"` — the child hosts gated orchestration on the pipeline's behalf; it must never inherit the session model. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. For web projects, dispatch in **a single message with two parallel Agent calls** (each with `model: "fable"`): diff --git a/docs/superpowers/specs/2026-09-28-effort-tiering-design.md b/docs/superpowers/specs/2026-09-28-effort-tiering-design.md index 05fab60..79a3419 100644 --- a/docs/superpowers/specs/2026-09-28-effort-tiering-design.md +++ b/docs/superpowers/specs/2026-09-28-effort-tiering-design.md @@ -34,6 +34,8 @@ repo. | Subagent frontmatter `effort:` | Applied to the subagent. Absent → **inherits the session level**. | built-in on sonnet printed `xhigh`; impeccable agent pinned `medium` printed `medium` | | Skill frontmatter `effort:`, user-typed `/skill` | Applied for the **rest of the turn**, AskUserQuestion included. | headless `/effort-probe-low`: every request at `low` | | Skill frontmatter `effort:`, loaded by Claude through the Skill tool, **interactive** session | Applied for the rest of the turn. Last loaded skill wins, up and down. | this session: `xhigh` → probe max → `$CLAUDE_EFFORT=max`, request records `effort=max` → probe xhigh → back to `xhigh` | +| Same, pairing rule | Applies **only when the Skill call shares the assistant message with another tool call after it**; a lone Skill call is a no-op. The paired call already runs at the new level. | this session, 8/8 observations | +| Same, re-load | A shifter already loaded in the conversation re-applies its effort when loaded again (paired); only its text is deduped. | this session | | Same, **headless** (`-p`) | **Not applied** (neither `effort:` nor `model:`). | three `-p` runs, transcript effort unchanged | | Prompt cache on a mid-turn shift | **Preserved** on Fable 5.1: first request at max read 206,996 cached tokens, wrote 1,164. | this session | | Agent tool call site | No `effort` parameter (only `model`). One agent file = one effort. | tool schema | @@ -97,14 +99,16 @@ Applies from the user's invocation for the rest of the turn. | effort | Skills | |---|---| | low | status, commit-change, release-candidate, doc, capitalize, close, reconcile, deploy, profile, plugin-check | -| medium | gitflow, prune-memory, find-docs | +| medium | gitflow, prune-memory | | high | feat, hotfix, bugfix, refactor, web-validate, harden, seo, geo | -| xhigh | ship-feature, init-project, onboard, tour, audit-delta, analyze, code-clean, client-handover, spec, skillify, brainstorming, writing-plans | -| unlisted | session default, by design (gstack and plugin skills are external; graphify is machine-owned) | +| xhigh | ship-feature, init-project, onboard, tour, audit-delta, analyze, code-clean, client-handover, brainstorming, writing-plans | +| unlisted | session default, by design: gstack skills (`spec` and `skillify` are gstack), plugin skills, and machine-generated skills (`graphify`, `find-docs`) | -`brainstorming` and `writing-plans` are vendored superpowers skills: the -one-line patch drifts from upstream at each resync; a census lock (§7) -makes the loss loud. +`brainstorming` and `writing-plans` are vendored superpowers skills living in +`skills-external/` (gitignored, symlinked into `skills/`): the pin is applied +to the real file and never committed; `install-plugins.sh` re-applies it after +every resync, and the census checks it whenever the file is present (visible +SKIP otherwise). A skill loaded by Claude as a sub-step (feat → commit-change) also shifts the level for the rest of the turn (interactive, §2), so orchestrators @@ -121,6 +125,9 @@ mirroring `lib/model-gate.md`: - A shift is a `Skill(effort-)` call on the main loop. Never inside a dispatched agent (agents run on their pin). One tool round-trip, cache-safe (§2). +- **Pairing rule**: the shift is sent in the same assistant message as the + step's first tool call, shift first; a lone Skill call is a no-op (§2). + Re-loading a shifter re-applies its effort. - Orchestrator wiring, three points each: `effort-medium` when the plan is closed and the dispatch phase starts; `effort-low` before the capitalize / journal / doc-commit tail; `effort-max` at an escalation @@ -176,7 +183,7 @@ each subagent's effort (2.1.243). |---|---| | `settings.json` | `effortLevel` → `high` (curated config: read the diff, LRN-098) | | `agents/*.md` (20) | `effort:` line per D2; `skills/init-project/SKILL.md:98` citer | -| `skills/*/SKILL.md` (33) | `effort:` line per D3, including the two vendored superpowers skills | +| `skills/*/SKILL.md` (28 tracked) + `skills-external/{brainstorming,writing-plans}/SKILL.md` (not committed) | `effort:` line per D3; `install-plugins.sh` re-applies the two vendored pins after resync | | `skills/effort-{low,medium,high,xhigh,max}/SKILL.md` | new, frontmatter + one sentence | | `lib/effort-shift.md` | new include: protocol, wiring points, escalation, turn reset, headless note | | `lib/model-gate.md` §4 | one paragraph: effort is the second axis, pointer to the include | @@ -185,6 +192,7 @@ each subagent's effort (2.1.243). | orchestrator SKILL.md (feat, hotfix, bugfix, ship-feature, init-project, onboard, tour, code-clean, seo, geo, harden, web-validate, client-handover, audit-delta) | include line + the three wiring points; ship-feature 4b max | | `hooks/statusline.sh`, `hooks/session-start.sh` | live effort display; env-var warning | | `lib/tests/effort-routing.test.sh` | new census suite (§7) | +| `lib/effort-audit.py` | transcript audit script (§9) | | `CHANGELOG.md`, `.claude/memory/*` | release note; BDR + LRN + EVAL + journal (§8) | ## 7. Tests and census (`make test`) @@ -196,8 +204,9 @@ New suite `lib/tests/effort-routing.test.sh`, `grep -qF` locks in the in its first 10 frontmatter lines, level in the allowed set; the "none" list has no `effort:`. 2. Tier locks per D2 (one `has` per agent). -3. Skill locks per D3 (one `has` per skill), including `brainstorming` and - `writing-plans` (the resync alarm). +3. Skill locks per D3 (one per skill); the two vendored skills are checked + when present (visible SKIP otherwise); `install-plugins.sh` carries the + re-apply block. 4. The five shifter skills exist with the exact `name:` and `effort:`. 5. `lib/effort-shift.md` is included by every orchestrator in the §6 list; `verify-secure-loop.md` and `ship-feature/SKILL.md` contain the diff --git a/hooks/session-start.sh b/hooks/session-start.sh index 90ba97c..ef22f19 100644 --- a/hooks/session-start.sh +++ b/hooks/session-start.sh @@ -107,7 +107,7 @@ fi REPO_DIR="${_repo_dir:-}" unset _claude_real _repo_dir -# Effort tiering (BDR-NEXT): this env var beats every skill/agent `effort:` pin. +# Effort tiering (BDR-107): this env var beats every skill/agent `effort:` pin. EFFORT_WARN="" if [ -n "${CLAUDE_CODE_EFFORT_LEVEL:-}" ]; then EFFORT_WARN="⚠️ CLAUDE_CODE_EFFORT_LEVEL=${CLAUDE_CODE_EFFORT_LEVEL} set: skill/agent effort pins ignored" diff --git a/hooks/statusline.sh b/hooks/statusline.sh index c9e2332..52200b3 100755 --- a/hooks/statusline.sh +++ b/hooks/statusline.sh @@ -34,7 +34,7 @@ if [ -z "$PROFILE" ] || [ "$PROFILE" = "none" ]; then fi # Effort level: the live value when the harness exports it (skill/agent -# `effort:` shifts included, BDR-NEXT), else the persisted settings.json key +# `effort:` shifts included, BDR-107), else the persisted settings.json key # (.effortLevel — set by /effort or manual edit; symlinked into ~/.claude). EFFORT="${CLAUDE_EFFORT:-}" if [ -z "$EFFORT" ] && [ -f "$REPO/settings.json" ]; then diff --git a/install-plugins.sh b/install-plugins.sh index c669cde..935ec9e 100644 --- a/install-plugins.sh +++ b/install-plugins.sh @@ -934,7 +934,7 @@ for _ext_skill in "${EXT_SKILL_NAMES[@]}"; do done echo "" -# Effort tiering (BDR-NEXT): the vendored brainstorming/writing-plans carry an +# Effort tiering (BDR-107): the vendored brainstorming/writing-plans carry an # effort pin upstream lacks; re-apply after every resync (census lock in # lib/tests/effort-routing.test.sh alarms if this ever stops working). for _s in brainstorming writing-plans; do diff --git a/lib/effort-shift.md b/lib/effort-shift.md index 3d07925..502716b 100644 --- a/lib/effort-shift.md +++ b/lib/effort-shift.md @@ -1,4 +1,4 @@ -# Effort shift — phase-level reasoning effort on the main loop (BDR-NEXT) +# Effort shift — phase-level reasoning effort on the main loop (BDR-107) Shared include, companion of `lib/model-gate.md`: the gate fixes WHICH model reflects, this include fixes HOW HARD each phase thinks. The rungs are the diff --git a/lib/model-gate.md b/lib/model-gate.md index d75bd86..22c897f 100644 --- a/lib/model-gate.md +++ b/lib/model-gate.md @@ -46,6 +46,6 @@ the main loop's behalf (skill-runners), otherwise its complexity tier (opus = dispatched judgment, sonnet = execution/collection, haiku = short mechanical probes). -Effort is the second axis of the same table (BDR-NEXT): every typed agent +Effort is the second axis of the same table (BDR-107): every typed agent carries an `effort:` pin next to `model:`, and the main loop shifts per phase through `lib/effort-shift.md`. Nothing dispatched inherits either axis. diff --git a/lib/tests/effort-routing.test.sh b/lib/tests/effort-routing.test.sh index 6128221..5921c31 100755 --- a/lib/tests/effort-routing.test.sh +++ b/lib/tests/effort-routing.test.sh @@ -1,5 +1,5 @@ #!/usr/bin/env bash -# lib/tests/effort-routing.test.sh — census: effort tiering (BDR-NEXT) +# lib/tests/effort-routing.test.sh — census: effort tiering (BDR-107) # agent pins, skill entry levels, shifter skills, orchestrator wiring, settings. # shellcheck disable=SC2015 # A && ok || ko is deliberate here: ok/ko never fail, so C never masks a true A set -u diff --git a/skills/audit-delta/SKILL.md b/skills/audit-delta/SKILL.md index 3cbda45..7b838ae 100644 --- a/skills/audit-delta/SKILL.md +++ b/skills/audit-delta/SKILL.md @@ -29,7 +29,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. Audit only what changed since the last run, on the axes the user picks. Per axis: **audit → approval gate → fix → re-verify → marker update**, diff --git a/skills/bugfix/SKILL.md b/skills/bugfix/SKILL.md index 9e54f09..96869cb 100644 --- a/skills/bugfix/SKILL.md +++ b/skills/bugfix/SKILL.md @@ -26,7 +26,7 @@ allowed-tools: MODEL GATE (blocking): run `$HOME/.claude/lib/model-gate.md` BEFORE any step below. Verdict `small` → STOP — print the gate's remedy, end the turn, dispatch nothing. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. ## REQUEST $ARGUMENTS diff --git a/skills/code-clean/SKILL.md b/skills/code-clean/SKILL.md index 0821ebb..6d9a7b9 100644 --- a/skills/code-clean/SKILL.md +++ b/skills/code-clean/SKILL.md @@ -26,7 +26,7 @@ allowed-tools: MODEL GATE (blocking): run `$HOME/.claude/lib/model-gate.md` BEFORE any step below. Verdict `small` → STOP — print the gate's remedy, end the turn, dispatch nothing. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. ## TARGET $ARGUMENTS diff --git a/skills/feat/SKILL.md b/skills/feat/SKILL.md index a5de32b..4ca234d 100644 --- a/skills/feat/SKILL.md +++ b/skills/feat/SKILL.md @@ -26,7 +26,7 @@ allowed-tools: MODEL GATE (blocking): run `$HOME/.claude/lib/model-gate.md` BEFORE any step below. Verdict `small` → STOP — print the gate's remedy, end the turn, dispatch nothing. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. ## REQUEST $ARGUMENTS diff --git a/skills/geo/SKILL.md b/skills/geo/SKILL.md index c06c8f6..17b3206 100644 --- a/skills/geo/SKILL.md +++ b/skills/geo/SKILL.md @@ -29,7 +29,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. Dispatches the `geo-analyzer` subagent (audit + fix bundle), then applies the bundle from THIS main loop at **L1** — same shape as `/web-validate` diff --git a/skills/harden/SKILL.md b/skills/harden/SKILL.md index ad3ef56..7aad9e6 100644 --- a/skills/harden/SKILL.md +++ b/skills/harden/SKILL.md @@ -29,7 +29,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. This skill orchestrates a narrow-scope hardening audit: TLS + security headers + redirects + canonical + custom 404 + server configs. It diff --git a/skills/hotfix/SKILL.md b/skills/hotfix/SKILL.md index f395cd5..4ae566d 100644 --- a/skills/hotfix/SKILL.md +++ b/skills/hotfix/SKILL.md @@ -24,7 +24,7 @@ allowed-tools: MODEL GATE (blocking): run `$HOME/.claude/lib/model-gate.md` BEFORE any step below. Verdict `small` → STOP — print the gate's remedy, end the turn, dispatch nothing. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. ## REQUEST $ARGUMENTS diff --git a/skills/init-project/SKILL.md b/skills/init-project/SKILL.md index 66e6c91..cf28ef7 100644 --- a/skills/init-project/SKILL.md +++ b/skills/init-project/SKILL.md @@ -14,7 +14,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. ## REQUEST $ARGUMENTS diff --git a/skills/onboard/SKILL.md b/skills/onboard/SKILL.md index 5d7f526..c8d7e79 100644 --- a/skills/onboard/SKILL.md +++ b/skills/onboard/SKILL.md @@ -14,7 +14,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. ## REQUEST $ARGUMENTS diff --git a/skills/seo/SKILL.md b/skills/seo/SKILL.md index ebace86..31cd20a 100644 --- a/skills/seo/SKILL.md +++ b/skills/seo/SKILL.md @@ -30,7 +30,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. This skill orchestrates TWO specialist agents running in parallel, then merges their output into a single `.claude/audits/SEO.md` report. It is the main diff --git a/skills/ship-feature/SKILL.md b/skills/ship-feature/SKILL.md index 693d724..60d5149 100644 --- a/skills/ship-feature/SKILL.md +++ b/skills/ship-feature/SKILL.md @@ -14,7 +14,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. ## REQUEST $ARGUMENTS diff --git a/skills/tour/SKILL.md b/skills/tour/SKILL.md index 8894645..00a04ec 100644 --- a/skills/tour/SKILL.md +++ b/skills/tour/SKILL.md @@ -30,7 +30,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. One pipeline per project: **security → clean → re-verify → reconcile → doc → convergence re-audit**, looping until a full pass applies zero new diff --git a/skills/web-validate/SKILL.md b/skills/web-validate/SKILL.md index cae8c15..e6f93ee 100644 --- a/skills/web-validate/SKILL.md +++ b/skills/web-validate/SKILL.md @@ -28,7 +28,7 @@ Run `$HOME/.claude/lib/model-gate.md`. Reflection here (planning, audit judgment, loop decisions) requires Fable/Opus. Verdict `small` → STOP: the gate prints the remedy; end the turn — no later step, no dispatch. Nominal (big) path is silent. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. This skill orchestrates a narrow-scope standards audit : From 5ed96aa8ed610746cc52065af18a73f28ce2c50d Mon Sep 17 00:00:00 2001 From: bastien Date: Mon, 28 Sep 2026 20:20:10 +0200 Subject: [PATCH 18/23] chore(memory): BDR-107 effort tiering, EVAL-036 A/B, journal, TODO W1-W4 ticked --- .claude/memory/decisions.md | 9 +++++++++ .claude/memory/evals.md | 8 ++++++++ .claude/memory/journal.md | 1 + .claude/tasks/TODO.md | 8 ++++---- 4 files changed, 22 insertions(+), 4 deletions(-) diff --git a/.claude/memory/decisions.md b/.claude/memory/decisions.md index 0410532..707cf57 100644 --- a/.claude/memory/decisions.md +++ b/.claude/memory/decisions.md @@ -128,6 +128,7 @@ rules: | BDR-104 | 2026-09-28 | MengTo motion pack: vendor 5 scroll skills pinned via shared lib/vendor-skills.sh + build personal skill site-motion; 17 skipped | accepted | | BDR-105 | 2026-09-28 | skill-catalog prune: 9 gstack out via GSTACK_REMOVED, full ⊇ every profile, max = everything, brightdata + frontend-design plugin off, security-guidance Stop review off, design gate asks `21st login` and waits | accepted | | BDR-106 | 2026-09-28 | superpowers: 7 wired skills vendored at v6.4.1 via lib/vendor-skills.sh (always_on lock class), plugin + marketplace dropped, citers by bare name, doctrine map for the 4 non-vendored refs | accepted | +| BDR-107 | 2026-09-28 | Effort tiering: session high, effort pins on 20 agents (BDR-077 second axis), entry level on 30 skills, five paired shifter skills, max at loop caps + ship-feature 4b | accepted | --- @@ -1330,3 +1331,11 @@ Branch feature/user-writing-web-rules, UNMERGED (human gate). - **Caveats**: upstream cross-refs to the plugin prefix and the 8 dropped skills remain in the vendored text (a call on a dropped name fails, doctrine map applies); no upstream auto-update (bump the pin deliberately); the harness hot-loaded the 7 bare names in the running session after link.sh, the plugin names leave at restart; `superpowers-marketplace` cache dir may linger empty; other machines: `make plugin` (vendors) + `make link`, then uninstall the cached plugin by hand (CHANGELOG). - **Reference**: 18f8c89 (wiring), ddea411 (citers/docs/settings); contract `2026-09-28-superpowers-vendored-1357` (12 criteria, oracles in `.oracles/`), plan r3 after 3 challengers (simplicity CONCERNS(2), robustness CONCERNS(3), correctness FATAL(5)) + confirmation CONCERNS(1); executors 2/2 DONE first pass; GATE 0 MET, verifier CONFORME 12/12, security PASS; catalog 82 skills, plugin passive cost 670 t (ui-ux-pro-max only). Links [[BDR-105]] [[BDR-102]] [[BDR-104]] [[BDR-065]] [[LRN-178]] [[EVAL-034]]. - **Amendment 2026-09-28 (merge)**: `gitflow finish` → 65665a5, no conflict, pushed, local + origin copies removed; the 7 vendored skills stay linked after the merge. Whole prune (tiers 1 + 2) on develop. + +## BDR-107 — Effort tiering: session high, agent pins, skill entry levels, paired phase shifts, max at escalation [accepted] (2026-09-28) +- **Decision**: settings `effortLevel` high (was xhigh). `effort:` pin on 20 repo-authored agents by role: low appliers (hotfixer, release-executor, plugin-probe, validator-analyzer), medium executors (feater, bugfixer, code-cleaner, onboarder, scaffolder), high judgment (refactorer, analyzer, commit-changer, doc-syncer, handover-doc-writer), xhigh challengers + gates (plan-challenger, plugin-advisor, verifier, security-auditor, seo-analyzer, geo-analyzer); none on interviewer/client-handover-writer (inline-load), status-reporter (haiku), impeccable-* (vendored). `effort:` on 28 tracked user-invoked skills = run entry level (low bookkeeping, medium gitflow/prune-memory, high feat/hotfix/bugfix/refactor/audits-with-fix, xhigh orchestrators) + xhigh on vendored brainstorming/writing-plans (skills-external/, re-applied by install-plugins STEP 8e). Five shifter skills `effort-{low,medium,high,xhigh,max}` loaded by orchestrators per `lib/effort-shift.md`: medium at dispatch span, own level before challenge synthesis, low at bookkeeping tail, max at verify-secure caps (GATE 0/1/2) + ship-feature 4b; re-assert after nested skill / prose gate. STOP texts name `$CLAUDE_EFFORT`, suggest `/effort-max`. statusline shows `$CLAUDE_EFFORT`; banner warns on `CLAUDE_CODE_EFFORT_LEVEL`. Census `lib/tests/effort-routing.test.sh`. Audit script `lib/effort-audit.py`. +- **Why**: session-wide xhigh burned thinking on bookkeeping; EVAL-035: 97 % of thinking in the main loop, sonnet subagents ~26 tok/request → main-loop levers (entry level, shifts) carry the savings; pins = explicitness + future models. A/B `/reconcile` high→low: requests 18→15, output −27 %, thinking −28 %, time −19 % (EVAL-036). +- **Harness facts (2.1.283)**: skill `effort:` applies on user slash invocation and on interactive Skill-tool load; the Skill-tool load applies ONLY when paired with another tool call in the same message (lone call = no-op); re-load re-applies (text deduped); not applied in `-p`/SDK; prompt cache kept across a shift; `CLAUDE_CODE_EFFORT_LEVEL` beats every frontmatter; one effort per agent file, no call-site override; unpinned agents inherit the level in force at dispatch. +- **Alternatives rejected**: executor pins only (they barely think); escalation-diagnoser agent fable+max (no context, one more agent; main-loop max keeps the failure context); reflection in fable skill-runner children with session medium (loses interactivity); settings.json rewrite mid-run (LRN-098 class); `maxEffortLevel` caps (hide a mis-pin the census should fail); pins on machine-generated skills (find-docs: ctx7 regenerates, gitignored) or gstack skills (spec, skillify). +- **Caveats**: shifts inert headless; a prose gate ending the turn resets to session level (re-assert wired in bugfix and ship-feature 4b); mode-based agents pin their judgment mode; a shift paired with a built-in judgment dispatch would downgrade it (pair with Read/Bash instead); `lib/gitflow-test.sh` T16a red on this machine = gitleaks not installed, unrelated. +- **Refs**: spec `docs/superpowers/specs/2026-09-28-effort-tiering-design.md`, plan `docs/superpowers/plans/2026-09-28-effort-tiering.md`, [[LRN-179]], [[EVAL-035]], [[EVAL-036]], [[BDR-077]]. diff --git a/.claude/memory/evals.md b/.claude/memory/evals.md index 07573d0..bfdca25 100644 --- a/.claude/memory/evals.md +++ b/.claude/memory/evals.md @@ -56,6 +56,7 @@ rules: | EVAL-033 | 2026-09-28 | case 7: 2 analyzers + 2 executors + 3 re-dispatches; verifiers caught shape, convention and my wrong count; security caught an env override | brief names the scratchpad path explicitly (3 /tmp leftovers); keep blind verifiers; count claims get an artifact | | EVAL-034 | 2026-09-28 | catalog prune + 21st gate: two challenge rounds each found what r3 missed (nested SKILL.md, fixture cp lists, in-session export); my ledgers failed twice (heredoc CHECKs); 5 executors DONE first pass; verifier gap = tool false positive | keep the confirmation pass on any plan that changed materially; one-line CHECKs; grep fixture cp lists before a `source` | | EVAL-035 | 2026-09-28 | thinking-share measurement, 6 days of transcripts (10,955 requests): thinking = 8 % of weighted spend, 97 % of it in the main loop; sonnet subagents at xhigh think 26 tok/request; cache reads = 53 % | pins = explicitness not savings; main-loop effort + context size are the levers; A/B after rollout | +| EVAL-036 | 2026-09-28 | A/B `/reconcile` headless, session high vs skill entry low: requests 18→15, output 12374→9038 (−27 %), thinking 3135→2248 (−28 %), time 96.5→78.4 s (−19 %), n=1 | keep low on bookkeeping skills; repeat on a reflection skill before touching the medium/high split | --- @@ -338,3 +339,10 @@ Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itse - **Result**: main loop 67 % of weighted spend, 97 % of thinking (Fable 1,430 think-tok/request); sonnet subagents 5,268 requests at xhigh, 26 think-tok/request; thinking = 8 % of spend, all output 16 %, cache reads 53 % (main-loop context ~320 k tok/request). Window 6 days only. Indirect effect of effort (fewer steps → fewer requests) unmeasured. - **Anomaly**: design was framed around executor pins; one script inverted it before any edit. Measure before routing. - **Action**: pins stay (explicitness, future models); main-loop skill effort + phase shifts carry the savings; A/B `/reconcile` high vs xhigh after rollout; context size = bigger lever, separate track. + +## EVAL-036 — A/B `/reconcile` headless: skill entry level low vs session high +- **Date**: 2026-09-28 +- **Method**: Task 4 of the effort-tiering plan; `claude -p "/reconcile" --output-format json --allowedTools Read Grep Glob "Bash(git status:*)" "Bash(git log:*)"` before (session `high`, no frontmatter) and after (`effort: low` on the skill); per-request `usage` summed from the session jsonl. +- **Result**: requests 18→15, output tokens 12374→9038 (−27 %), thinking 3135→2248 (−28 %), duration 96.5 s→78.4 s (−19 %); transcript effort field high→low confirmed. n=1, same repo state. +- **Anomaly**: none; the indirect effect (fewer steps at lower effort) is real, which EVAL-035's static split could not show. +- **Action**: keep low on bookkeeping skills; repeat on a reflection skill (feat) before touching the medium/high split; `lib/effort-audit.py` makes the split measurable any time. diff --git a/.claude/memory/journal.md b/.claude/memory/journal.md index 06ea276..65555bc 100644 --- a/.claude/memory/journal.md +++ b/.claude/memory/journal.md @@ -542,3 +542,4 @@ rules: - User go "merge le tier 2": feature/superpowers-vendored merged into develop via `gitflow finish` → 65665a5, no conflict, pushed, copies removed by the lib. develop == origin/develop, no working branch anywhere. Whole skill-catalog prune (BDR-105 + BDR-106) on develop: catalog 82 skills, plugin passive cost 670 t, no session injection. Open for the user: `21st login`, claude.ai skills off, floor-guard `xit(` hotfix (BLK-023), two /tmp fixture dirs, other machines `make plugin` + `make link` + uninstall the cached plugin. - /hotfix BLK-023 (user: "fais le hotfix du floor-guard"): `skip_kind` substring match → `xit(` ⊂ `exit(`. Fix 0deb559 on bugfix/floor-guard-xit-boundary: bare Jasmine names via `SKIP_IDENT_RE` lookbehind, 4 flip fixtures (12/12). 3 challengers (2 SOLID, robustness CONCERNS(2): fixture line itself flaggable on a test path → waiver comment outside the echo; my criterion-2 live oracle vacuous → dropped — same LRN-173 class, plus I wrote a heredoc CHECK again before catching it, [[LRN-176]]). Hotfixer DONE first pass, oracles MET, security PASS. UNMERGED — human gate. - User go "oui pour le changelog et merge le": CHANGELOG floor-guard entry amended via doc-syncer patch + doc-commit (018dfa3), bugfix/floor-guard-xit-boundary merged into develop via `gitflow finish` → c9f9b40, pushed, copies removed. develop == origin/develop, no working branch anywhere. Day total on develop: skill-catalog prune tiers 1 + 2 (BDR-105, BDR-106), 21st sign-in gate, BLK-023 resolved. +- effort tiering built on feature/effort-tiering (BDR-107): session high, 20 agent pins, 28+2 skill entry levels, 5 paired shifters, max at caps + 4b, census 129+ locks green, A/B −27 % output on /reconcile; finish awaits human signal. diff --git a/.claude/tasks/TODO.md b/.claude/tasks/TODO.md index 4932005..9425917 100644 --- a/.claude/tasks/TODO.md +++ b/.claude/tasks/TODO.md @@ -4,10 +4,10 @@ Spec `docs/superpowers/specs/2026-09-28-effort-tiering-design.md`, plan `docs/superpowers/plans/2026-09-28-effort-tiering.md`. Approved 2026-09-28: session high, A+B+C, max on the main loop at the loop caps + ship-feature 4b, superpowers patch. -- [ ] W1 settings high + banner warning + statusline live level + 20 agent pins + census suite (Tasks 1-3) -- [ ] W2 33 skill entry levels + superpowers xhigh with resync re-apply (Tasks 4, 9) -- [ ] W3 five shifters + lib/effort-shift.md + orchestrator wiring + max at caps/4b + gate audit (Tasks 5-8) -- [ ] W4 BDR id + CHANGELOG + EVAL A/B + journal + audit script (Tasks 10-11) +- [x] W1 settings high + banner warning + statusline live level + 20 agent pins + census suite (Tasks 1-3) +- [x] W2 33 skill entry levels + superpowers xhigh with resync re-apply (Tasks 4, 9) +- [x] W3 five shifters + lib/effort-shift.md + orchestrator wiring + max at caps/4b + gate audit (Tasks 5-8) +- [x] W4 BDR id + CHANGELOG + EVAL A/B + journal + audit script (Tasks 10-11) ## 2026-09-28 — tier 2: vendor 7 superpowers skills, drop the plugin (feature/superpowers-vendored) User go "fais le tier 2" (decision 2026-09-28, batch 1). Contract From a3b479e984f3319f1e100b5a99428fc7032e8110 Mon Sep 17 00:00:00 2001 From: bastien Date: Mon, 28 Sep 2026 20:27:46 +0200 Subject: [PATCH 19/23] fix(effort): reflow effort-audit.py to 80 columns (R12) --- lib/effort-audit.py | 32 +++++++++++++++++++++----------- 1 file changed, 21 insertions(+), 11 deletions(-) diff --git a/lib/effort-audit.py b/lib/effort-audit.py index 2b74d13..bc39494 100755 --- a/lib/effort-audit.py +++ b/lib/effort-audit.py @@ -8,7 +8,8 @@ import json import os import sys -WEIGHTS = {"in": 1.0, "cc": 1.25, "cr": 0.1, "out": 5.0} # relative to input price +# Weights relative to input price. +WEIGHTS = {"in": 1.0, "cc": 1.25, "cr": 0.1, "out": 5.0} FIELDS = ("in", "cc", "cr", "out", "think") @@ -50,36 +51,45 @@ def weighted(counter): def report(agg): - """Print the per-key table, then the main/sub split and the thinking share.""" + """Print the per-key table, then the main/sub split and the thinking + share.""" total = collections.Counter() for counter in agg.values(): total.update(counter) total_w = weighted(total) or 1 print(f"{'scope':5} {'model':22} {'effort':7} {'msgs':>6} {'think/msg':>9} " f"{'think_tok':>10} {'out_tok':>10} {'cache_read':>12} {'%wcost':>7}") - for (scope, model, effort), c in sorted(agg.items(), key=lambda kv: -weighted(kv[1])): + ranked = sorted(agg.items(), key=lambda kv: -weighted(kv[1])) + for (scope, model, effort), c in ranked: per_msg = c["think"] / max(c["msgs"], 1) - print(f"{scope:5} {model:22} {effort:7} {c['msgs']:6d} {per_msg:9.0f} " - f"{c['think']:10d} {c['out']:10d} {c['cr']:12d} {100 * weighted(c) / total_w:6.1f}%") + print(f"{scope:5} {model:22} {effort:7} {c['msgs']:6d} " + f"{per_msg:9.0f} {c['think']:10d} {c['out']:10d} " + f"{c['cr']:12d} {100 * weighted(c) / total_w:6.1f}%") by_scope = collections.defaultdict(collections.Counter) for (scope, _, _), c in agg.items(): by_scope[scope].update(c) for scope, c in by_scope.items(): - print(f" {scope:5} weighted-cost {100 * weighted(c) / total_w:5.1f}% " - f"thinking {100 * c['think'] / max(total['think'], 1):5.1f}% requests {c['msgs']}") - print(f" thinking = {100 * total['think'] * WEIGHTS['out'] / total_w:.1f}% of weighted cost; " - f"cache reads = {100 * total['cr'] * WEIGHTS['cr'] / total_w:.1f}%") + print(f" {scope:5} weighted-cost " + f"{100 * weighted(c) / total_w:5.1f}% thinking " + f"{100 * c['think'] / max(total['think'], 1):5.1f}% " + f"requests {c['msgs']}") + print(f" thinking = " + f"{100 * total['think'] * WEIGHTS['out'] / total_w:.1f}% " + f"of weighted cost; cache reads = " + f"{100 * total['cr'] * WEIGHTS['cr'] / total_w:.1f}%") def main(): - root = os.path.expanduser(sys.argv[1] if len(sys.argv) > 1 else "~/.claude/projects") + root = os.path.expanduser( + sys.argv[1] if len(sys.argv) > 1 else "~/.claude/projects") agg = collections.defaultdict(collections.Counter) for project in sorted(glob.glob(os.path.join(root, "*"))): if not os.path.isdir(project): continue for path in glob.glob(os.path.join(project, "*.jsonl")): scan(path, "main", agg) - for path in glob.glob(os.path.join(project, "*", "subagents", "*.jsonl")): + sub_glob = os.path.join(project, "*", "subagents", "*.jsonl") + for path in glob.glob(sub_glob): scan(path, "sub", agg) report(agg) From 58c3a3e9b7b6d21fe26e2ad4ff08c6e5dc38d692 Mon Sep 17 00:00:00 2001 From: bastien Date: Mon, 28 Sep 2026 20:48:27 +0200 Subject: [PATCH 20/23] fix(effort): re-raise judgment dispatches, planning re-asserts, pairing caveat, dedupe audit script (final review I1-I3) --- agents/client-handover-writer.md | 5 +++-- lib/effort-audit.py | 9 ++++++++- lib/effort-shift.md | 15 ++++++++++++--- lib/model-gate.md | 4 +++- lib/tests/effort-routing.test.sh | 16 ++++++++++++++-- skills/audit-delta/SKILL.md | 2 +- skills/bugfix/SKILL.md | 2 +- skills/code-clean/SKILL.md | 2 +- skills/feat/SKILL.md | 2 +- skills/geo/SKILL.md | 2 +- skills/harden/SKILL.md | 2 +- skills/hotfix/SKILL.md | 2 +- skills/init-project/SKILL.md | 4 +++- skills/onboard/SKILL.md | 3 +-- skills/seo/SKILL.md | 2 +- skills/ship-feature/SKILL.md | 5 ++++- skills/tour/SKILL.md | 1 - skills/web-validate/SKILL.md | 2 +- 18 files changed, 57 insertions(+), 23 deletions(-) diff --git a/agents/client-handover-writer.md b/agents/client-handover-writer.md index 13d32e9..1ba9c0c 100644 --- a/agents/client-handover-writer.md +++ b/agents/client-handover-writer.md @@ -97,6 +97,8 @@ Parse `$ARGUMENTS` for optional flags: --- +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. + ## STEP 1 — PRE-FLIGHT ```bash @@ -225,7 +227,7 @@ Store `DEPLOYED_URL` for STEP 7. If empty, ask user during STEP 6. --- ## STEP 3 — BASELINE AUDITS (parallel) -First: `Skill(effort-medium)` (effort-shift: dispatch span starts; send it in the same message as this step's first dispatch). +First: `Skill(effort-high)` (effort-shift: judgment dispatch; the fable skill-runners are built-ins and inherit the level in force; high is the entry level of the audits they run). Goal: capture `SCORE_*_BEFORE` so the client doc shows the delta. @@ -262,7 +264,6 @@ the gate. this pipeline (initial audits, fix-loop re-dispatches, commit-change, web-validate) carries `model: "fable"` — the child hosts gated orchestration on the pipeline's behalf; it must never inherit the session model. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. For web projects, dispatch in **a single message with two parallel Agent calls** (each with `model: "fable"`): diff --git a/lib/effort-audit.py b/lib/effort-audit.py index bc39494..2d84556 100755 --- a/lib/effort-audit.py +++ b/lib/effort-audit.py @@ -26,7 +26,10 @@ def usage_row(usage): def scan(path, scope, agg): - """Add every assistant record of one transcript to agg.""" + """Add every assistant record of one transcript to agg, once per + message id (the transcript writes one record per content block, + all sharing the same id and usage).""" + seen = set() with open(path, errors="ignore") as handle: for line in handle: try: @@ -36,6 +39,10 @@ def scan(path, scope, agg): msg = rec.get("message") or {} if rec.get("type") != "assistant" or not msg.get("usage"): continue + mid = msg.get("id") + if mid in seen: + continue + seen.add(mid) sub = scope == "sub" or bool(rec.get("isSidechain")) key = ("sub" if sub else "main", str(msg.get("model", "?")).replace("claude-", ""), diff --git a/lib/effort-shift.md b/lib/effort-shift.md index 502716b..38b7ad7 100644 --- a/lib/effort-shift.md +++ b/lib/effort-shift.md @@ -19,9 +19,11 @@ max (stuck error, judged need). effort (the harness only dedupes the skill text), so bounce-back sequences such as medium → max → medium work. - A skill's `effort:` frontmatter applies from the moment it loads to the - end of the turn: on the user's `/skill` and on a `Skill(...)` call by - Claude in an interactive session. Last loaded wins, both directions. The - prompt cache survives a shift. + end of the turn: on the user's `/skill` unconditionally, and on a + `Skill(...)` call by Claude only under the pairing rule above (a skill + Claude loads alone, such as `brainstorming` or `writing-plans`, applies + nothing). Last loaded wins, both directions. The prompt cache survives a + shift. - Dispatched agents run on their own `effort:` pin, never on a shift. Unpinned agents inherit the level in force at dispatch. - Headless sessions (`-p`, `claude agents`, SDK) ignore skill-level effort: @@ -55,6 +57,11 @@ does not move. challenger is an infrastructure failure) and "gone WRONG → STOP" (STOP precedes any further reasoning); their STOP text names the level reached and suggests `/effort-max` for the relaunch. +5. Before any built-in or unpinned dispatch that carries judgment (a + `general-purpose` with `model: "opus"` or `"fable"`, the code reviewer + of requesting-code-review, a skill-runner) → `Skill(effort-)` + paired with that dispatch: built-ins inherit the level in force, and a + medium set earlier in the span would downgrade them. ## Re-assert @@ -69,3 +76,5 @@ does not move. - A shift never inside a dispatched agent: pins rule there. - Max is for diagnosis, not for retrying the same fix harder. +- A medium shift never precedes a judgment dispatch in the same span + without an own-level shift paired with that dispatch. diff --git a/lib/model-gate.md b/lib/model-gate.md index 22c897f..874cd23 100644 --- a/lib/model-gate.md +++ b/lib/model-gate.md @@ -48,4 +48,6 @@ mechanical probes). Effort is the second axis of the same table (BDR-107): every typed agent carries an `effort:` pin next to `model:`, and the main loop shifts per phase -through `lib/effort-shift.md`. Nothing dispatched inherits either axis. +through `lib/effort-shift.md`. No typed agent inherits either axis; +built-ins inherit the effort in force at dispatch, so an orchestrator shifts +before dispatching them (`lib/effort-shift.md`, wiring point 5). diff --git a/lib/tests/effort-routing.test.sh b/lib/tests/effort-routing.test.sh index 5921c31..a6e4003 100755 --- a/lib/tests/effort-routing.test.sh +++ b/lib/tests/effort-routing.test.sh @@ -65,8 +65,11 @@ has "lib/model-gate.md" 'lib/effort-shift.md' # ── 6) orchestrator wiring (spec D4) for s in feat hotfix bugfix ship-feature init-project onboard tour code-clean seo geo harden web-validate audit-delta; do - has "skills/$s/SKILL.md" 'lib/effort-shift.md'; has "skills/$s/SKILL.md" 'Skill(effort-medium)'; done -has "agents/client-handover-writer.md" 'lib/effort-shift.md'; has "agents/client-handover-writer.md" 'Skill(effort-medium)' + has "skills/$s/SKILL.md" 'lib/effort-shift.md'; has "skills/$s/SKILL.md" 'a lone Skill call is a no-op'; done +for s in feat hotfix bugfix ship-feature init-project code-clean seo geo harden web-validate audit-delta; do + has "skills/$s/SKILL.md" 'Skill(effort-medium)'; done +lacks "skills/onboard/SKILL.md" 'Skill(effort-medium)'; lacks "skills/tour/SKILL.md" 'Skill(effort-medium)' +has "agents/client-handover-writer.md" 'lib/effort-shift.md'; lacks "agents/client-handover-writer.md" 'Skill(effort-medium)'; has "agents/client-handover-writer.md" 'Skill(effort-high)' for s in feat hotfix bugfix; do has "skills/$s/SKILL.md" 'Skill(effort-high)'; done for s in ship-feature init-project onboard code-clean audit-delta; do has "skills/$s/SKILL.md" 'Skill(effort-xhigh)'; done for s in seo geo harden web-validate; do has "skills/$s/SKILL.md" 'Skill(effort-high)'; done @@ -91,6 +94,15 @@ has "skills/bugfix/SKILL.md" 'effort-shift: turn reset' has "lib/effort-shift.md" 'effort-audit.py' [ -x "$R/lib/effort-audit.py" ] && ok || ko "lib/effort-audit.py missing or not executable" +# ── 6c) judgment dispatches re-raised, planning re-asserts, stronger locks (final review I1/I2/M5) +for s in ship-feature init-project; do has "skills/$s/SKILL.md" 'effort-shift: judgment dispatch'; has "skills/$s/SKILL.md" 'effort-shift: turn reset'; done +has "agents/client-handover-writer.md" 'effort-shift: judgment dispatch' +has "lib/effort-shift.md" 'Before any built-in or unpinned dispatch' +has "lib/model-gate.md" 'built-ins inherit the effort in force' +has "skills/ship-feature/SKILL.md" 'effort-shift: error recovery' +for s in feat hotfix bugfix seo geo harden web-validate ship-feature init-project onboard code-clean audit-delta; do has "skills/$s/SKILL.md" 'effort-shift: own level before the challenge'; done +has "install-plugins.sh" 'for _s in brainstorming writing-plans; do' + # ── summary (later tasks insert their locks ABOVE this line) printf 'effort-routing census: %d pass, %d fail\n' "$pass" "$fail" [ "$fail" -eq 0 ] diff --git a/skills/audit-delta/SKILL.md b/skills/audit-delta/SKILL.md index 7b838ae..4a13098 100644 --- a/skills/audit-delta/SKILL.md +++ b/skills/audit-delta/SKILL.md @@ -170,7 +170,7 @@ Then show the user the same compact table inline. ### 3b-bis. CHALLENGE THE PROPOSALS (before the gate) -`Skill(effort-xhigh)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). +`Skill(effort-xhigh)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). This axis' findings + proposed fixes are a proposal set worth attacking before the human gate. Persist THIS axis' finding list (not the whole append-only report) to `.claude/tasks/plans/--.md`, then run diff --git a/skills/bugfix/SKILL.md b/skills/bugfix/SKILL.md index 96869cb..84051fd 100644 --- a/skills/bugfix/SKILL.md +++ b/skills/bugfix/SKILL.md @@ -126,7 +126,7 @@ RISK: fast-path is not exempt: a 1-line fix with a visible choice still asks. ## STEP 3b — CHALLENGE THE FIX PLAN (before the contract) -`Skill(effort-high)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). +`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). Unless the fix is the trivial 1-2 line case STEP 3 already fast-paths, the DIAGNOSIS + FIX PLAN is a reflection worth attacking before it hardens into a contract. Persist it to `.claude/tasks/plans/--.md`, then run diff --git a/skills/code-clean/SKILL.md b/skills/code-clean/SKILL.md index 6d9a7b9..72458cb 100644 --- a/skills/code-clean/SKILL.md +++ b/skills/code-clean/SKILL.md @@ -122,7 +122,7 @@ TOTALS: If no issues found: report clean state and stop. ## STEP 3b — CHALLENGE THE SCOPE (before approval) -`Skill(effort-xhigh)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). +`Skill(effort-xhigh)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). The STEP 3 report is the proposed cleanup scope — worth attacking before the human approves it. It is still inline, so FIRST persist it to `.claude/tasks/plans/--.md` (STEP 3 report format, one item diff --git a/skills/feat/SKILL.md b/skills/feat/SKILL.md index 4ca234d..4b10386 100644 --- a/skills/feat/SKILL.md +++ b/skills/feat/SKILL.md @@ -124,7 +124,7 @@ in the contract's CLARIFICATIONS `[gated]` and in the plan. A choice that surfaces only during execution comes back as `NEED-DECISION` (STEP 3). ## STEP 1b — CHALLENGE THE PLAN (before branching) -`Skill(effort-high)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). +`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). The STEP 1 plan is a reflection worth attacking before a branch is spent on it. Persist it to `.claude/tasks/plans/--.md`, then run `$HOME/.claude/lib/challenge-plan.md` with `PLAN` = that file, `KIND` = `build-plan`, diff --git a/skills/geo/SKILL.md b/skills/geo/SKILL.md index 17b3206..ee76ad9 100644 --- a/skills/geo/SKILL.md +++ b/skills/geo/SKILL.md @@ -86,7 +86,7 @@ your bundle." ``` ## STEP 1b — CHALLENGE THE FIX BUNDLE (advisory, before apply) -`Skill(effort-high)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). +`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). The analyzer returned a `## FIX BUNDLE` — worth attacking before any edit lands. **Skip if intervention mode = conservative** (nothing is applied). Else persist the bundle verbatim to `.claude/tasks/plans/--.md`, then run diff --git a/skills/harden/SKILL.md b/skills/harden/SKILL.md index 7aad9e6..d6d53e3 100644 --- a/skills/harden/SKILL.md +++ b/skills/harden/SKILL.md @@ -522,7 +522,7 @@ Extract the score and critical-alert count from `.claude/audits/HARDEN.md` for t --- ## STEP 2b — CHALLENGE THE FIX BUNDLE (MODE=fix only, advisory) -`Skill(effort-high)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). +`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). Skip if MODE=audit (no bundle exists). Else, before the STEP 3 gate, harden the bundle: extract the `## 8. Fix bundle` section from HARDEN.md to `.claude/tasks/plans/--.md` (a clean, blind-judgeable artifact), then run diff --git a/skills/hotfix/SKILL.md b/skills/hotfix/SKILL.md index 4ae566d..61530b1 100644 --- a/skills/hotfix/SKILL.md +++ b/skills/hotfix/SKILL.md @@ -94,7 +94,7 @@ point. Run it ONLY when the settled fix touches control flow or behaviour — an off-by-one, a wrong operator/variable, a behaviour-changing config value, or a missing import that alters execution. In doubt → it is probably a `/bugfix`. -`Skill(effort-high)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). +`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). For a logic fix: persist the STEP 1 located fix (root cause + the exact edit) to `.claude/tasks/plans/--.md`, then run `$HOME/.claude/lib/challenge-plan.md` with `PLAN` = that file, `KIND` = diff --git a/skills/init-project/SKILL.md b/skills/init-project/SKILL.md index cf28ef7..9ee91d6 100644 --- a/skills/init-project/SKILL.md +++ b/skills/init-project/SKILL.md @@ -181,11 +181,12 @@ This is the deterministic scaffold commit owner (closes BLK-010). The MVP is implemented on a `feature/*` branch off `develop` (STEP 8). ## STEP 6 — PLAN +`Skill(effort-xhigh)` first, sent with the next tool call (effort-shift: turn reset; gate #1 ended the turn and the vendored `writing-plans` pin applies only when the user invokes it). Invoke `writing-plans` (vendored superpowers skill) with BRIEF + skeleton. Granular tasks (2-5 min each), exact file paths, TDD: tests before code. ## STEP 6b — CHALLENGE THE PLAN (before the gate) -`Skill(effort-xhigh)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). +`Skill(effort-xhigh)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). Before the human sees the implementation plan, harden it. Run `$HOME/.claude/lib/challenge-plan.md` with `PLAN` = the plan STEP 6 wrote under `docs/superpowers/plans/`, `KIND` = `build-plan`, `SCOPE` = the skeleton + task file @@ -259,6 +260,7 @@ deferred to a later /onboard) and turns the informal analyze into a verdict against the founding contract. Distinct axis from STEP 10 code review ([[LRN-095]]) — both run. +`Skill(effort-xhigh)` first, sent with the review dispatch (effort-shift: judgment dispatch; the reviewer is a built-in and inherits the level in force). ## STEP 10 — CODE REVIEW Invoke `requesting-code-review` (vendored superpowers skill). **Model routing (BDR-077):** the review subagent it dispatches MUST carry `model: "opus"` in the Agent call — diff --git a/skills/onboard/SKILL.md b/skills/onboard/SKILL.md index c8d7e79..bf7bc43 100644 --- a/skills/onboard/SKILL.md +++ b/skills/onboard/SKILL.md @@ -92,7 +92,6 @@ STOP. La réponse détermine si STEP 1 tourne une fois (A) ou N fois (C) ou avec ## STEP 2 — BASELINE CONFIG (onboarder agent) -`Skill(effort-medium)` first (effort-shift: dispatch span starts; send it in the same message as this step's first dispatch). Dispatch `Agent(subagent_type="onboarder")` (pin sonnet — BDR-077 : config templating = exécution, plus jamais inline sur le modèle de session). Un BLOCAGE (clé manquante, CLAUDE.md existant) revient en rapport — l'agent ne @@ -893,7 +892,7 @@ Vérifier que les 4 fichiers `.claude/audits/ONBOARD_REPORT.md`, `.claude/audits --- ## STEP 7b — CHALLENGE THE PROPOSALS (before the human gate) -`Skill(effort-xhigh)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). +`Skill(effort-xhigh)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). The 4 audit files are on disk; `AUDIT_PROPOSALS.md` is the artifact worth attacking before the human spends a gate on it. Run `$HOME/.claude/lib/challenge-plan.md` with `PLAN` = diff --git a/skills/seo/SKILL.md b/skills/seo/SKILL.md index 31cd20a..ea3a82e 100644 --- a/skills/seo/SKILL.md +++ b/skills/seo/SKILL.md @@ -510,7 +510,7 @@ the reports." ``` ## STEP 1b — CHALLENGE THE FIX BUNDLE (advisory, before apply) -`Skill(effort-high)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). +`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). Both envelopes now carry a `## FIX BUNDLE` — worth attacking before any edit lands. **Skip if intervention mode = conservative** (nothing is applied). Else persist both bundles (seo + geo, verbatim) to `.claude/tasks/plans/--.md`, then run diff --git a/skills/ship-feature/SKILL.md b/skills/ship-feature/SKILL.md index 60d5149..941f02c 100644 --- a/skills/ship-feature/SKILL.md +++ b/skills/ship-feature/SKILL.md @@ -114,8 +114,10 @@ Inject ONLY what constrains: the NON-BINDING count does NOT enter the brainstorm (the injection inherits the OUTPUT filter — detail what binds, drop what doesn't). Consumption = INPUT INJECTION (we can't modify the external skill; we control its input). Refine request into validated design via Socratic questioning. Don't proceed until design approved. +Turns after a user reply run at the session level until a tool call is paired with `Skill(effort-xhigh)` (effort-shift: turn reset). ## STEP 2 — PLAN +`Skill(effort-xhigh)` first, sent with the next tool call (effort-shift: turn reset; brainstorm turns after a user reply run at the session level, and the vendored `brainstorming` pin applies only when the user invokes it). Invoke `writing-plans` (vendored superpowers skill) with the validated design AND the 0d digest: every task must be consistent with the in-force constraints; where a task implements or affects one, note the ID inline. Break design into tasks (2-5 min each). Each task: exact file paths, full code, verification steps. @@ -125,7 +127,7 @@ request nor the STEP 1 brainstorm settled (check the contract's CLARIFICATIONS first) → one batch before STEP 2b; answers append to the contract `[gated]`. ## STEP 2b — CHALLENGE THE PLAN (adversarial, before the gate) -`Skill(effort-xhigh)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). +`Skill(effort-xhigh)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). Before the human sees the plan, harden it. Run `$HOME/.claude/lib/challenge-plan.md`: - `PLAN` = the plan STEP 2 wrote under `docs/superpowers/plans/` - `KIND` = `build-plan` @@ -240,6 +242,7 @@ against the contract. It is a DISTINCT axis from STEP 6 code review (contract conformity + security vs. craft/design) — both run, neither subsumes the other ([[LRN-095]]). +`Skill(effort-xhigh)` first, sent with the review dispatch (effort-shift: judgment dispatch; the reviewer is a built-in and inherits the level in force). ## STEP 6 — CODE REVIEW Invoke `requesting-code-review` (vendored superpowers skill). **Model routing (BDR-077):** the review subagent it dispatches MUST carry `model: "opus"` in the Agent call — diff --git a/skills/tour/SKILL.md b/skills/tour/SKILL.md index 00a04ec..3f7ca46 100644 --- a/skills/tour/SKILL.md +++ b/skills/tour/SKILL.md @@ -87,7 +87,6 @@ Model discipline (the user-fixed invariant behind this mode): Runner dispatch, one per project: ``` -Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message Agent(subagent_type="general-purpose", description="tour runner — ", prompt="Read ~/.claude/skills/tour/SKILL.md and execute STEP 1 → STEP 3 diff --git a/skills/web-validate/SKILL.md b/skills/web-validate/SKILL.md index e6f93ee..0901830 100644 --- a/skills/web-validate/SKILL.md +++ b/skills/web-validate/SKILL.md @@ -255,7 +255,7 @@ grep -c '^### \[Critique\]' .claude/audits/VALIDATE.md --- ## STEP 2b — CHALLENGE THE FIX BUNDLE (MODE=fix only, advisory) -`Skill(effort-high)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). +`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). Skip if MODE=audit (no bundle exists). Else, before the STEP 3 gate, harden the bundle: extract the `## 5. Fix bundle` section from VALIDATE.md to `.claude/tasks/plans/--.md` (a clean, blind-judgeable artifact), then run From a70430e683a53c57b1035512f74ffd2add56bb30 Mon Sep 17 00:00:00 2001 From: bastien Date: Mon, 28 Sep 2026 20:48:35 +0200 Subject: [PATCH 21/23] chore(memory): EVAL-037 deduped counts, BDR-107 correction, LRN-180 pairing rule, TODO count --- .claude/memory/decisions.md | 1 + .claude/memory/evals.md | 8 ++++++++ .claude/memory/learnings.md | 5 +++++ .claude/tasks/TODO.md | 2 +- 4 files changed, 15 insertions(+), 1 deletion(-) diff --git a/.claude/memory/decisions.md b/.claude/memory/decisions.md index 707cf57..d78f265 100644 --- a/.claude/memory/decisions.md +++ b/.claude/memory/decisions.md @@ -1339,3 +1339,4 @@ Branch feature/user-writing-web-rules, UNMERGED (human gate). - **Alternatives rejected**: executor pins only (they barely think); escalation-diagnoser agent fable+max (no context, one more agent; main-loop max keeps the failure context); reflection in fable skill-runner children with session medium (loses interactivity); settings.json rewrite mid-run (LRN-098 class); `maxEffortLevel` caps (hide a mis-pin the census should fail); pins on machine-generated skills (find-docs: ctx7 regenerates, gitignored) or gstack skills (spec, skillify). - **Caveats**: shifts inert headless; a prose gate ending the turn resets to session level (re-assert wired in bugfix and ship-feature 4b); mode-based agents pin their judgment mode; a shift paired with a built-in judgment dispatch would downgrade it (pair with Read/Bash instead); `lib/gitflow-test.sh` T16a red on this machine = gitleaks not installed, unrelated. - **Refs**: spec `docs/superpowers/specs/2026-09-28-effort-tiering-design.md`, plan `docs/superpowers/plans/2026-09-28-effort-tiering.md`, [[LRN-179]], [[EVAL-035]], [[EVAL-036]], [[BDR-077]]. +- **Correction (2026-09-28)**: EVAL-035 counted one record per content block (~2.8× on request counts); deduped figures in [[EVAL-037]]: main-loop thinking 99.9% of total thinking (was 96.6%), thinking 5.6% of weighted cost (was 8.4%), sonnet think/request 26→0.2 tok. Conclusions hold, sharper: main loop still carries almost all thinking, executors stay cheap. diff --git a/.claude/memory/evals.md b/.claude/memory/evals.md index bfdca25..2854bf5 100644 --- a/.claude/memory/evals.md +++ b/.claude/memory/evals.md @@ -57,6 +57,7 @@ rules: | EVAL-034 | 2026-09-28 | catalog prune + 21st gate: two challenge rounds each found what r3 missed (nested SKILL.md, fixture cp lists, in-session export); my ledgers failed twice (heredoc CHECKs); 5 executors DONE first pass; verifier gap = tool false positive | keep the confirmation pass on any plan that changed materially; one-line CHECKs; grep fixture cp lists before a `source` | | EVAL-035 | 2026-09-28 | thinking-share measurement, 6 days of transcripts (10,955 requests): thinking = 8 % of weighted spend, 97 % of it in the main loop; sonnet subagents at xhigh think 26 tok/request; cache reads = 53 % | pins = explicitness not savings; main-loop effort + context size are the levers; A/B after rollout | | EVAL-036 | 2026-09-28 | A/B `/reconcile` headless, session high vs skill entry low: requests 18→15, output 12374→9038 (−27 %), thinking 3135→2248 (−28 %), time 96.5→78.4 s (−19 %), n=1 | keep low on bookkeeping skills; repeat on a reflection skill before touching the medium/high split | +| EVAL-037 | 2026-09-28 | correction of EVAL-035/036 counts: transcript records are per content block; deduped by message.id → main-loop thinking share 99.9%, thinking share of weighted cost 5.6%, sonnet think/msg 26→0.2, A/B requests 9→8 | conclusions hold (sharper: main-loop thinking 96.6%→99.9%, weighted-cost thinking corrected 8.4%→5.6%); effort-audit.py dedupes from a3b479e+ | --- @@ -346,3 +347,10 @@ Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itse - **Result**: requests 18→15, output tokens 12374→9038 (−27 %), thinking 3135→2248 (−28 %), duration 96.5 s→78.4 s (−19 %); transcript effort field high→low confirmed. n=1, same repo state. - **Anomaly**: none; the indirect effect (fewer steps at lower effort) is real, which EVAL-035's static split could not show. - **Action**: keep low on bookkeeping skills; repeat on a reflection skill (feat) before touching the medium/high split; `lib/effort-audit.py` makes the split measurable any time. + +## EVAL-037 — correction of EVAL-035/036: one transcript record per content block, deduped by message.id +- **Date**: 2026-09-28 +- **Output checked**: EVAL-035 (8 % thinking / 97 % main loop / 26 tok per sonnet request) and EVAL-036 (requests 18→15), produced by `effort-audit.py` counting every assistant record; final review found duplicates (same `message.id` + identical `usage`, one record per content block, ~2.8× on this repo's last 6 transcripts). +- **Result (deduped)**: main weighted-cost 61.4 %, thinking share 99.9 % (was 96.6 %); sub weighted-cost 38.6 %, thinking share 0.1 %; thinking = 5.6 % of weighted cost (was 8.4 %, inflated by duplicate counting); sonnet think/request 26→0.2 tok (sub, xhigh); A/B `/reconcile` (EVAL-036 rerun, deduped) requests 9→8, output 6129→4706, thinking 1550→1104 — the raw undeduped counts on the same transcripts are 18→15, matching EVAL-036 exactly (the bug, not the finding). +- **Anomaly**: the main-loop-carries-almost-all-thinking split got SHARPER after dedup (96.6→99.9 %), not weaker — duplication was near-uniform across content blocks, so ratios among scopes barely moved; only the absolute request/token counts and the overall thinking-share-of-cost figure were inflated (~2.2-2.8× depending on transcript mix). +- **Action**: `lib/effort-audit.py` dedupes by `message.id` from this commit; cite EVAL-037, not EVAL-035, for the split. diff --git a/.claude/memory/learnings.md b/.claude/memory/learnings.md index b90abfa..3d0fe40 100644 --- a/.claude/memory/learnings.md +++ b/.claude/memory/learnings.md @@ -199,6 +199,7 @@ rules: | LRN-177 | 2026-09-28 | gstack skills hardcode `~/.claude/skills/gstack/` (83 paths: bin, scripts, ETHOS.md, */sections, review/specialists, make-pdf/dist, freeze/bin…); only bin + browse/dist were linked → dead skills and vacuous hooks (exit 127); ./setup plants a global symlink; whole-dir link exposes nested SKILL.md; `apply` is additive, `set` parks | any gstack wiring change, any "gstack skill fails" report | | LRN-178 | 2026-09-28 | a top-level `source` added to a lib breaks every hermetic suite that copies that lib alone into a fixture; grep the `cp` lists before adding one, or source lazily inside the branch that needs it | adding `source` to profile.sh / toggle-external.sh / any lib the suites copy | | LRN-179 | 2026-09-28 | Skill `effort:` frontmatter shifts the MAIN LOOP for the rest of the turn on user slash invocation AND on interactive Skill-tool loads (last loaded wins, both directions, prompt cache kept); NOT applied in `-p`/headless; agent pins always honoured, unpinned agents inherit session | effort tiering; any skill or agent that must think more or less than the session | +| LRN-180 | 2026-09-28 | Skill-tool effort override needs a paired tool call: a lone Skill(effort-*) call is a no-op; a load in the same message as another tool call applies (the paired call already sees it); re-load re-applies (text deduped); skills Claude loads alone (brainstorming, writing-plans) apply nothing | every orchestrator shift; amends LRN-179 | --- @@ -1645,3 +1646,7 @@ Rule: when editing a doctrine file under structure locks, grep the test's lock s ## LRN-179 — skill `effort:` shifts the main loop for the rest of the turn, interactive only - **Context**: effort-tiering spike 2026-09-28, Claude Code 2.1.283, Fable 5.1. Probes = `$CLAUDE_EFFORT` in Bash + transcript `effort` field per request. User-typed `/probe-low` → whole turn `low`. Skill-tool load in interactive session → `max` then `xhigh`, last loaded wins, both directions; first request after the switch read 206,996 cached tokens, wrote 1,164 (cache kept). Three `-p` runs: neither `effort:` nor `model:` skill frontmatter applied via Skill tool. Agent pin honoured (impeccable `medium`), unpinned built-in on sonnet inherited `xhigh`. Docs agent claimed "ultrathink keyword does not exist": wrong, docs = in-context nudge, API effort unchanged. Harness claims get verified against the harness ([[LRN-046]]). - **Apply**: main-loop effort per phase = `Skill(effort-)` on the main loop, never inside a dispatched agent; headless runs stay at session level; keep `CLAUDE_CODE_EFFORT_LEVEL` unset (beats every frontmatter). Spec `docs/superpowers/specs/2026-09-28-effort-tiering-design.md`. + +## LRN-180 — Skill-tool effort override needs a paired tool call; a lone Skill call is a no-op (2.1.283) +- **Context**: effort-tiering smoke. Six lone `Skill(effort-*)` / probe loads left `$CLAUDE_EFFORT` unchanged; every load issued in the same assistant message as another tool call applied, and the paired Bash already saw the new level. Re-loading an already-loaded shifter re-applies (text deduped: "already loaded above"). Final review: `brainstorming` / `writing-plans` loaded alone by ship-feature and init-project → their vendored xhigh pin inert. Amends [[LRN-179]]. +- **Apply**: `Skill(effort-)` always travels with the step's first tool call, shift first; pair a downward shift with a pinned executor or a Read/Bash, never with a built-in judgment dispatch; before any built-in judgment dispatch, pair the own-level shift with it; skills Claude loads alone do not apply their pin → re-assert with a paired shift at the resumed planning step ([[BDR-107]]). diff --git a/.claude/tasks/TODO.md b/.claude/tasks/TODO.md index 9425917..b6c4f18 100644 --- a/.claude/tasks/TODO.md +++ b/.claude/tasks/TODO.md @@ -5,7 +5,7 @@ Spec `docs/superpowers/specs/2026-09-28-effort-tiering-design.md`, plan `docs/superpowers/plans/2026-09-28-effort-tiering.md`. Approved 2026-09-28: session high, A+B+C, max on the main loop at the loop caps + ship-feature 4b, superpowers patch. - [x] W1 settings high + banner warning + statusline live level + 20 agent pins + census suite (Tasks 1-3) -- [x] W2 33 skill entry levels + superpowers xhigh with resync re-apply (Tasks 4, 9) +- [x] W2 28+2 skill entry levels + superpowers xhigh with resync re-apply (Tasks 4, 9) - [x] W3 five shifters + lib/effort-shift.md + orchestrator wiring + max at caps/4b + gate audit (Tasks 5-8) - [x] W4 BDR id + CHANGELOG + EVAL A/B + journal + audit script (Tasks 10-11) From e529801411a8f78644492c1ed9efb8d5af2fd4bb Mon Sep 17 00:00:00 2001 From: bastien Date: Mon, 28 Sep 2026 20:50:26 +0200 Subject: [PATCH 22/23] fix(effort): judgment-dispatch shift under its heading (ship-feature, init-project) --- skills/init-project/SKILL.md | 2 +- skills/ship-feature/SKILL.md | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/skills/init-project/SKILL.md b/skills/init-project/SKILL.md index 9ee91d6..a7bc5ff 100644 --- a/skills/init-project/SKILL.md +++ b/skills/init-project/SKILL.md @@ -260,8 +260,8 @@ deferred to a later /onboard) and turns the informal analyze into a verdict against the founding contract. Distinct axis from STEP 10 code review ([[LRN-095]]) — both run. -`Skill(effort-xhigh)` first, sent with the review dispatch (effort-shift: judgment dispatch; the reviewer is a built-in and inherits the level in force). ## STEP 10 — CODE REVIEW +`Skill(effort-xhigh)` first, sent with the review dispatch (effort-shift: judgment dispatch; the reviewer is a built-in and inherits the level in force). Invoke `requesting-code-review` (vendored superpowers skill). **Model routing (BDR-077):** the review subagent it dispatches MUST carry `model: "opus"` in the Agent call — craft review is dispatched judgment, never inherited from the session. Fix diff --git a/skills/ship-feature/SKILL.md b/skills/ship-feature/SKILL.md index 941f02c..a941f35 100644 --- a/skills/ship-feature/SKILL.md +++ b/skills/ship-feature/SKILL.md @@ -242,8 +242,8 @@ against the contract. It is a DISTINCT axis from STEP 6 code review (contract conformity + security vs. craft/design) — both run, neither subsumes the other ([[LRN-095]]). -`Skill(effort-xhigh)` first, sent with the review dispatch (effort-shift: judgment dispatch; the reviewer is a built-in and inherits the level in force). ## STEP 6 — CODE REVIEW +`Skill(effort-xhigh)` first, sent with the review dispatch (effort-shift: judgment dispatch; the reviewer is a built-in and inherits the level in force). Invoke `requesting-code-review` (vendored superpowers skill). **Model routing (BDR-077):** the review subagent it dispatches MUST carry `model: "opus"` in the Agent call — craft review is dispatched judgment, never inherited from the session. Fix From 5b3ea682b46bc396d218892dd5a8b58fe021f56e Mon Sep 17 00:00:00 2001 From: bastien Date: Mon, 28 Sep 2026 21:21:31 +0200 Subject: [PATCH 23/23] chore: purge transient planning artifacts (BDR-065) --- .../plans/2026-09-28-effort-tiering.md | 1109 ----------------- .../specs/2026-09-28-effort-tiering-design.md | 250 ---- 2 files changed, 1359 deletions(-) delete mode 100644 docs/superpowers/plans/2026-09-28-effort-tiering.md delete mode 100644 docs/superpowers/specs/2026-09-28-effort-tiering-design.md diff --git a/docs/superpowers/plans/2026-09-28-effort-tiering.md b/docs/superpowers/plans/2026-09-28-effort-tiering.md deleted file mode 100644 index bded7b8..0000000 --- a/docs/superpowers/plans/2026-09-28-effort-tiering.md +++ /dev/null @@ -1,1109 +0,0 @@ -# Effort Tiering Implementation Plan - -> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. - -**Goal:** Route reasoning effort per role and per phase (low → max) across the claude-config orchestrators, instead of one session-wide `xhigh`. - -**Architecture:** Effort becomes the second axis of the BDR-077 routing table. Three layers, each a one-line frontmatter mechanism the harness already honours: `effort:` pins on the 20 repo-authored agents (dispatched work), `effort:` on the 33 user-invoked skills (the run's entry level), and five empty "shifter" skills the orchestrators load at phase boundaries (`Skill(effort-)`), including `max` at the loop caps and ship-feature error recovery. A census suite locks every value. - -**Tech Stack:** bash, GNU sed, python3 stdlib, jq, shellcheck, `make test` (suites under `lib/tests/*.test.sh` are auto-discovered). - -**Spec:** `docs/superpowers/specs/2026-09-28-effort-tiering-design.md` (read it first; every decision below is argued there, §2 holds the harness evidence, §3 the measurement). - -## Global Constraints - -- Claude Code ≥ 2.1.267 on the executing machine (skill/agent `effort:` honoured on Fable); the spike ran on 2.1.283. -- `CLAUDE_CODE_EFFORT_LEVEL` must be unset in every session that runs a smoke: it silences every frontmatter override. -- Branch `feature/effort-tiering` (exists, off develop). Every task ends with a commit on it; never `--no-verify`; never commit on develop/main; `gitflow finish` only on the human's signal. -- `make test` green after every task; `shellcheck *.sh hooks/*.sh lib/*.sh` clean after any shell edit. -- Allowed effort values, exactly: `low`, `medium`, `high`, `xhigh`, `max`. `max` never in `settings.json` (harness rejects it there). -- Never edit: `agents/impeccable-*.md`, `skills/graphify/**`, anything under `skills-external/`. The two vendored superpowers files edited (`skills/brainstorming/SKILL.md`, `skills/writing-plans/SKILL.md`) get a resync re-apply in Task 9. -- No new `CLAUDE.md "…"` citations anywhere (doctrine-citers census); the doctrine lives in `lib/`. -- `BDR-NEXT` is a literal token used in lib text and CHANGELOG until Task 10 computes the real id and replaces it. It must not survive Task 10. -- Shell: functions ≤ 25 logic lines, 80-char lines. Registry entries: English, caveman. -- Commit messages: no attribution lines. - -## Review Focus - -1. A run launched headless (`claude -p`, `claude agents`, SDK) never shifts: the include must say so, and the census locks that sentence (Task 5). -2. `CLAUDE_CODE_EFFORT_LEVEL` exported in the user's shell silently disables every pin and shift: the session banner must warn, tested by running the hook with the variable set (Task 2). -3. A nested skill at a different level (feat → commit-change at low) leaves the rest of the run at low: feat must re-assert `Skill(effort-high)` right after, locked by the census (Task 6). -4. An `effort:` on a model without effort support (haiku) is at best ignored: `status-reporter` must carry none, locked (Task 3). -5. A superpowers resync overwrites the vendored frontmatter: the census lock alarms, and install-plugins re-applies it (Task 9). - ---- - -### Task 1: Census suite skeleton with flip-test and the settings lock - -**Files:** -- Create: `lib/tests/effort-routing.test.sh` -- Test: itself (`make test suite=lib/tests/effort-routing.test.sh`) - -**Interfaces:** -- Produces: helpers `has`, `lacks`, `fm`, `fm_effort`, `fm_has_effort `, `fm_no_effort `, counters `pass`/`fail`, final line `effort-routing census: N pass, M fail`. Later tasks append `has`/`fm_has_effort` lines to this file, above the summary block. - -- [ ] **Step 1: Write the suite with the flip-test and one real lock that fails today** - -```bash -cat > lib/tests/effort-routing.test.sh <<'EOF' -#!/usr/bin/env bash -# lib/tests/effort-routing.test.sh — census: effort tiering (BDR-NEXT) -# agent pins, skill entry levels, shifter skills, orchestrator wiring, settings. -set -u -R="$(cd "$(dirname "$0")/../.." && pwd)" -pass=0; fail=0 -ok() { pass=$((pass+1)); } -ko() { fail=$((fail+1)); printf 'FAIL %s\n' "$1"; } -has() { if grep -qF "$2" "$R/$1"; then ok; else ko "$1 missing: $2"; fi; } -lacks() { if grep -qF "$2" "$R/$1"; then ko "$1 must NOT contain: $2"; else ok; fi; } -# frontmatter = the lines between the first two '---' lines -fm() { awk 'NR==1&&/^---$/{p=1;next} p&&/^---$/{exit} p' "$1"; } -fm_effort() { fm "$1" | grep -E '^effort: (low|medium|high|xhigh|max)$' | head -1 | cut -d' ' -f2; } -fm_has_effort() { - got="$(fm_effort "$R/$1")" - if [ "$got" = "$2" ]; then ok; else ko "$1 frontmatter effort must be '$2', got '${got:-none}'"; fi -} -fm_no_effort() { if fm "$R/$1" | grep -q '^effort:'; then ko "$1 must NOT pin effort"; else ok; fi; } - -# ── flip-test: the frontmatter reader must accept a valid level and reject an invalid one -FIX="$(mktemp -d)"; trap 'rm -rf "$FIX"' EXIT -printf -- '---\nname: good\neffort: xhigh\n---\nbody with effort: low in prose\n' > "$FIX/good.md" -printf -- '---\nname: bad\neffort: turbo\n---\n' > "$FIX/bad.md" -[ "$(fm_effort "$FIX/good.md")" = "xhigh" ] && ok || ko "flip: valid level not read" -[ -z "$(fm_effort "$FIX/bad.md")" ] && ok || ko "flip: invalid level accepted" -[ "$(fm "$FIX/good.md" | grep -c 'prose')" -eq 0 ] && ok || ko "flip: body leaked into frontmatter" - -# ── 1) session default (spec D1) -has "settings.json" '"effortLevel": "high"' - -# ── summary (later tasks insert their locks ABOVE this line) -printf 'effort-routing census: %d pass, %d fail\n' "$pass" "$fail" -[ "$fail" -eq 0 ] -EOF -chmod +x lib/tests/effort-routing.test.sh -``` - -- [ ] **Step 2: Run it, expect the flip-test to pass and the settings lock to fail** - -Run: `make test suite=lib/tests/effort-routing.test.sh` -Expected: `FAIL settings.json missing: "effortLevel": "high"` then `effort-routing census: 3 pass, 1 fail`, non-zero exit. - -- [ ] **Step 3: Shellcheck** - -Run: `shellcheck lib/tests/effort-routing.test.sh` -Expected: no output. - -- [ ] **Step 4: Commit** - -```bash -git add lib/tests/effort-routing.test.sh -git commit -m "test(effort): census suite skeleton with flip-test and settings lock" -``` - ---- - -### Task 2: Session default `high`, banner warning, live effort in the statusline - -**Files:** -- Modify: `settings.json` (line with `"effortLevel"`) -- Modify: `hooks/session-start.sh` (after the line `unset _claude_real _repo_dir`, and after the banner's closing box line) -- Modify: `hooks/statusline.sh:36-41` (the `EFFORT=` block) -- Test: `lib/tests/effort-routing.test.sh` - -**Interfaces:** -- Produces: banner line `⚠️ CLAUDE_CODE_EFFORT_LEVEL= set: skill/agent effort pins ignored` when the variable is exported; statusline `effort: `. - -- [ ] **Step 1: Add the two hook locks to the census (above the summary block)** - -```bash -python3 - <<'PY' -p="lib/tests/effort-routing.test.sh"; s=open(p).read() -s=s.replace("# ── summary", """# ── 2) hooks: env-var warning + live effort in the statusline (spec D1, D5) -has "hooks/session-start.sh" 'CLAUDE_CODE_EFFORT_LEVEL' -has "hooks/statusline.sh" 'CLAUDE_EFFORT' - -# ── summary""") -open(p,"w").write(s) -PY -``` - -- [ ] **Step 2: Run the suite, expect 3 failures (settings + two hooks)** - -Run: `make test suite=lib/tests/effort-routing.test.sh` -Expected: three `FAIL` lines, non-zero exit. - -- [ ] **Step 3: Set the session default to `high`** - -```bash -sed -i 's/"effortLevel": "xhigh"/"effortLevel": "high"/' settings.json -git diff settings.json -``` -Expected diff: exactly one changed line. If anything else moved (LRN-098: `/effort` and `/model` rewrite this file), `git checkout settings.json` and redo the sed. - -- [ ] **Step 4: Banner warning in session-start.sh** - -Insert after the line `unset _claude_real _repo_dir`: - -```bash -python3 - <<'PY' -p="hooks/session-start.sh"; L=open(p).read().split("\n") -i=L.index("unset _claude_real _repo_dir") -L[i+1:i+1]=[ -"", -"# Effort tiering (BDR-NEXT): this env var beats every skill/agent `effort:` pin.", -"EFFORT_WARN=\"\"", -"if [ -n \"${CLAUDE_CODE_EFFORT_LEVEL:-}\" ]; then", -" EFFORT_WARN=\"⚠️ CLAUDE_CODE_EFFORT_LEVEL=${CLAUDE_CODE_EFFORT_LEVEL} set: skill/agent effort pins ignored\"", -"fi", -] -open(p,"w").write("\n".join(L)) -PY -grep -n '└' hooks/session-start.sh -``` -Expected: two lines, both plain `echo` statements (an early fix-hint box near l.32, the main banner near l.255). The python below inserts after the last one: - -```bash -python3 - <<'PY' -p="hooks/session-start.sh"; L=open(p).read().split("\n") -i=max(k for k,l in enumerate(L) if "└" in l) -L.insert(i+1, '[ -n "$EFFORT_WARN" ] && printf \'%s\\n\' "$EFFORT_WARN"') -open(p,"w").write("\n".join(L)) -PY -``` - -- [ ] **Step 5: Live effort in the statusline** - -Replace the block from the comment `# Effort level from settings.json` through the `fi` that closes `if [ -f "$REPO/settings.json" ]; then`: - -```bash -python3 - <<'PY' -p="hooks/statusline.sh"; s=open(p).read() -old_start=s.index("# Effort level from settings.json") -old_end=s.index("fi\n", s.index('jq -r \'.effortLevel', old_start))+3 -new='''# Effort level: the live value when the harness exports it (skill/agent -# `effort:` shifts included, BDR-NEXT), else the persisted settings.json key -# (.effortLevel — set by /effort or manual edit; symlinked into ~/.claude). -EFFORT="${CLAUDE_EFFORT:-}" -if [ -z "$EFFORT" ] && [ -f "$REPO/settings.json" ]; then - EFFORT=$(jq -r '.effortLevel // "?"' "$REPO/settings.json" 2>/dev/null) -fi -[ -z "$EFFORT" ] && EFFORT="?" -''' -open(p,"w").write(s[:old_start]+new+s[old_end:]) -PY -sed -n 34,46p hooks/statusline.sh -``` -Expected: the new block, no duplicate `EFFORT=` lines. - -- [ ] **Step 6: Run the hook with the variable set, then the suites** - -```bash -CLAUDE_CODE_EFFORT_LEVEL=medium bash hooks/session-start.sh 2>/dev/null | grep -c 'effort pins ignored' -bash hooks/session-start.sh 2>/dev/null | grep -c 'effort pins ignored' -shellcheck hooks/session-start.sh hooks/statusline.sh -make test suite=lib/tests/effort-routing.test.sh -``` -Expected: `1`, then `0`, shellcheck silent, census all pass. - -- [ ] **Step 7: Full test run and commit** - -Run: `make test` -Expected: every suite green (curated-config-guard accepts a hand edit of settings.json). - -```bash -git add settings.json hooks/session-start.sh hooks/statusline.sh lib/tests/effort-routing.test.sh -git commit -m "feat(effort): session default high, env-var warning, live effort in statusline" -``` - ---- - -### Task 3: Agent effort pins (spec D2) - -**Files:** -- Modify: 20 files `agents/.md` (line 5 is `model: sonnet` or `model: opus` in every one of them; `agents/scaffolder.md` already has `effort: high` on line 6) -- Modify: `skills/init-project/SKILL.md` (the sentence `(pin sonnet, effort high — BDR-077`) -- Test: `lib/tests/effort-routing.test.sh` - -**Interfaces:** -- Produces: `effort: ` on line 6 of each pinned agent. - -- [ ] **Step 1: Add the 23 locks to the census (above the summary block)** - -```bash -python3 - <<'PY' -p="lib/tests/effort-routing.test.sh"; s=open(p).read() -locks="""# ── 3) agent pins (spec D2): one effort per agent file, judgment mode wins on mode-based agents -for a in hotfixer release-executor plugin-probe validator-analyzer; do fm_has_effort "agents/$a.md" low; done -for a in feater bugfixer code-cleaner onboarder scaffolder; do fm_has_effort "agents/$a.md" medium; done -for a in refactorer analyzer commit-changer doc-syncer handover-doc-writer; do fm_has_effort "agents/$a.md" high; done -for a in plan-challenger plugin-advisor verifier security-auditor seo-analyzer geo-analyzer; do fm_has_effort "agents/$a.md" xhigh; done -for a in interviewer client-handover-writer status-reporter; do fm_no_effort "agents/$a.md"; done -has "skills/init-project/SKILL.md" 'pin sonnet, effort medium' - -# ── summary""" -open(p,"w").write(s.replace("# ── summary", locks)) -PY -make test suite=lib/tests/effort-routing.test.sh | tail -1 -``` -Expected: `... 20 fail` (19 pins + the citer; the three `fm_no_effort` pass already). - -- [ ] **Step 2: Apply the pins** - -```bash -pin() { L=$1; shift; for a in "$@"; do - sed -i "0,/^model: \(sonnet\|opus\)$/s//&\neffort: $L/" "agents/$a.md"; done; } -pin low hotfixer release-executor plugin-probe validator-analyzer -pin medium feater bugfixer code-cleaner onboarder -sed -i 's/^effort: high$/effort: medium/' agents/scaffolder.md -pin high refactorer analyzer commit-changer doc-syncer handover-doc-writer -pin xhigh plan-challenger plugin-advisor verifier security-auditor seo-analyzer geo-analyzer -sed -i 's/(pin sonnet, effort high — BDR-077/(pin sonnet, effort medium — BDR-077/' skills/init-project/SKILL.md -grep -c '^effort:' agents/*.md | grep -v ':0' | grep -v impeccable | wc -l -``` -Expected: `20`. - -- [ ] **Step 3: Suites** - -Run: `make test suite=lib/tests/effort-routing.test.sh && make test suite=lib/tests/model-routing.test.sh` -Expected: both green (the model locks read `model: sonnet` on line 5, untouched). - -- [ ] **Step 4: Smoke, planted input, disk-verified** - -From an interactive session in this repo, dispatch: -``` -Agent(subagent_type="release-executor", description="effort pin smoke", - prompt="Diagnostic only, no release work. Run exactly one bash command and report its raw output: echo CLAUDE_EFFORT=$CLAUDE_EFFORT") -``` -Expected report: `CLAUDE_EFFORT=low`. Then read the subagent transcript: -```bash -f=$(ls -t ~/.claude/projects/-home-bchanot-Documents-claude/*/subagents/*.jsonl | head -1) -grep -o '"effort":"[a-z]*"' "$f" | sort | uniq -c -``` -Expected: only `"effort":"low"`. - -- [ ] **Step 5: Commit** - -```bash -git add agents/*.md skills/init-project/SKILL.md lib/tests/effort-routing.test.sh -git commit -m "feat(effort): pin effort on the 20 repo-authored agents (BDR-077 second axis)" -``` - ---- - -### Task 4: Skill entry levels (spec D3) with a before/after measurement - -**Files:** -- Modify: 31 files `skills//SKILL.md` (line 2 is `name: ` in every one of them; the two superpowers files are Task 9) -- Test: `lib/tests/effort-routing.test.sh` - -**Interfaces:** -- Produces: `effort: ` on line 3 of each listed skill; the `lvl` helper reused by Task 9. - -- [ ] **Step 1: Baseline measurement BEFORE the change (spec §9)** - -```bash -S=/tmp/claude-1000/-home-bchanot-Documents-claude/e593bc78-da6b-469b-9d0c-08d1a4aa8373/scratchpad -mkdir -p "$S"; cd ~/Documents/claude -claude -p "/reconcile" --output-format json --allowedTools "Read" "Grep" "Glob" "Bash(git status:*)" "Bash(git log:*)" > "$S/ab-before.json" 2>/dev/null -python3 - "$S/ab-before.json" <<'PY' -import json,sys,glob,os -d=json.load(open(sys.argv[1])); sid=d["session_id"] -P=os.path.expanduser("~/.claude/projects/-home-bchanot-Documents-claude") -n=o=t=0; eff=set() -for line in open(f"{P}/{sid}.jsonl", errors="ignore"): - r=json.loads(line) - if r.get("type")!="assistant": continue - u=r["message"].get("usage") or {}; n+=1; o+=u.get("output_tokens",0) - t+=(u.get("output_tokens_details") or {}).get("thinking_tokens",0); eff.add(r.get("effort")) -print(f"BEFORE requests={n} output={o} thinking={t} effort={eff} duration_ms={d.get('duration_ms')}") -PY -``` -Expected: one line, `effort={'high'}` (session default after Task 2). Keep the line for Task 10. - -- [ ] **Step 2: Add the 31 locks to the census** - -```bash -python3 - <<'PY' -p="lib/tests/effort-routing.test.sh"; s=open(p).read() -locks="""# ── 4) skill entry levels (spec D3): the user's invocation sets the run's level -for s in status commit-change release-candidate doc capitalize close reconcile deploy profile plugin-check; do fm_has_effort "skills/$s/SKILL.md" low; done -for s in gitflow prune-memory find-docs; do fm_has_effort "skills/$s/SKILL.md" medium; done -for s in feat hotfix bugfix refactor web-validate harden seo geo; do fm_has_effort "skills/$s/SKILL.md" high; done -for s in ship-feature init-project onboard tour audit-delta analyze code-clean client-handover spec skillify; do fm_has_effort "skills/$s/SKILL.md" xhigh; done - -# ── summary""" -open(p,"w").write(s.replace("# ── summary", locks)) -PY -make test suite=lib/tests/effort-routing.test.sh | tail -1 -``` -Expected: `... 31 fail`. - -- [ ] **Step 3: Apply the levels** - -```bash -lvl() { L=$1; shift; for s in "$@"; do - sed -i "0,/^name: $s\$/s//&\neffort: $L/" "skills/$s/SKILL.md"; done; } -lvl low status commit-change release-candidate doc capitalize close reconcile deploy profile plugin-check -lvl medium gitflow prune-memory find-docs -lvl high feat hotfix bugfix refactor web-validate harden seo geo -lvl xhigh ship-feature init-project onboard tour audit-delta analyze code-clean client-handover spec skillify -grep -l '^effort:' skills/*/SKILL.md | wc -l -``` -Expected: `31`. - -- [ ] **Step 4: Suites** - -Run: `make test suite=lib/tests/effort-routing.test.sh && make test suite=lib/tests/skill-routing-census.test.sh` -Expected: both green. - -- [ ] **Step 5: Measurement AFTER, same command as Step 1 with `ab-after.json` and the label `AFTER`** - -Expected: `effort={'low'}`. Record both lines in the commit body; Task 10 turns them into the EVAL. - -- [ ] **Step 6: Commit** - -```bash -git add skills/*/SKILL.md lib/tests/effort-routing.test.sh -git commit -m "feat(effort): entry effort level on the 31 user-invoked skills (spec D3)" \ - -m "A/B /reconcile headless — " -``` - ---- - -### Task 5: Shifter skills, `lib/effort-shift.md`, model-gate paragraph (spec D4) - -**Files:** -- Create: `skills/effort-low/SKILL.md`, `skills/effort-medium/SKILL.md`, `skills/effort-high/SKILL.md`, `skills/effort-xhigh/SKILL.md`, `skills/effort-max/SKILL.md` -- Create: `lib/effort-shift.md` -- Modify: `lib/model-gate.md` (end of §4) -- Test: `lib/tests/effort-routing.test.sh`, `lib/tests/skill-routing-census.test.sh` - -**Interfaces:** -- Produces: skill names `effort-low`, `effort-medium`, `effort-high`, `effort-xhigh`, `effort-max`; the include path `$HOME/.claude/lib/effort-shift.md`; the wiring vocabulary Tasks 6-8 insert: `Skill(effort-)` lines with a trailing `# effort-shift: ` comment. - -- [ ] **Step 1: Locks (above the summary block)** - -```bash -python3 - <<'PY' -p="lib/tests/effort-routing.test.sh"; s=open(p).read() -locks="""# ── 5) shifter skills + include (spec D4) -for l in low medium high xhigh max; do fm_has_effort "skills/effort-$l/SKILL.md" "$l"; has "skills/effort-$l/SKILL.md" "name: effort-$l"; done -has "lib/effort-shift.md" 'Headless sessions' -has "lib/effort-shift.md" 'Skill(effort-max)' -has "lib/effort-shift.md" 'never inside a dispatched agent' -has "lib/model-gate.md" 'lib/effort-shift.md' - -# ── summary""" -open(p,"w").write(s.replace("# ── summary", locks)) -PY -``` - -- [ ] **Step 2: Create the five shifters (descriptions pre-validated against the routing census, pairwise ≤ 0.03)** - -```bash -mk() { mkdir -p "skills/effort-$1"; printf '%s\n' '---' "name: effort-$1" "description: $2" "effort: $1" '---' \ - "Effort shifted to $1 for the rest of this turn (lib/effort-shift.md). Continue with the caller's next step." \ - > "skills/effort-$1/SKILL.md"; } -mk low "Bookkeeping shift. Lowers reasoning to the cheapest level for the rest of the turn: journal lines, memory commits, capitalize, release bookkeeping, status output." -mk medium "Orchestration shift. Standard reasoning between two dispatches: read a subagent report, pick the next step, relay a gate verdict, route a branch." -mk high "Investigation shift. Deeper reasoning for diagnosis, LOCATE, contract drafting, refactor judgement inside feat, hotfix and bugfix runs." -mk xhigh "Reflection shift. Deep reasoning for brainstorm, planning, challenge synthesis and audit verdicts before a human validation gate." -mk max "Escalation shift. Maximum reasoning when a verify or security loop hits its cap, a gate fails twice, or error recovery starts in ship-feature." -``` - -- [ ] **Step 3: Write the include** - -```bash -cat > lib/effort-shift.md <<'EOF' -# Effort shift — phase-level reasoning effort on the main loop (BDR-NEXT) - -Shared include, companion of `lib/model-gate.md`: the gate fixes WHICH model -reflects, this include fixes HOW HARD each phase thinks. The rungs are the -user's: low (fix a line, run a script) · medium (day-to-day) · high -(refactor, resisting bug) · xhigh (architecture, audit before validation) · -max (stuck error, judged need). - -## Mechanics (verified on Claude Code 2.1.283) - -- A skill's `effort:` frontmatter applies from the moment it loads to the - end of the turn: on the user's `/skill` and on a `Skill(...)` call by - Claude in an interactive session. Last loaded wins, both directions. The - prompt cache survives a shift. -- Dispatched agents run on their own `effort:` pin, never on a shift. - Unpinned agents inherit the level in force at dispatch. -- Headless sessions (`-p`, `claude agents`, SDK) ignore skill-level effort: - the run stays at the session level. `CLAUDE_CODE_EFFORT_LEVEL` beats every - frontmatter; keep it unset (the session banner warns). - -## Shifters - -`Skill(effort-low)` · `Skill(effort-medium)` · `Skill(effort-high)` · -`Skill(effort-xhigh)` · `Skill(effort-max)`. One tool call, one-line body. -Typed by the user, `/effort-max` is a turn-scoped max: the relaunch lever -after a STOP. `ultrathink` only adds an in-context nudge; the API level -does not move. - -## Wiring — per orchestrator - -1. A dispatch span starts (executor, collector, fan-out) → - `Skill(effort-medium)`. -2. Reflection resumes after a dispatch span (challenge synthesis, verdict, - plan revision) → `Skill(effort-)`. Concretely: - the line before every `lib/challenge-plan.md` call. -3. The bookkeeping tail (memory commit, doc commit) → `Skill(effort-low)`. -4. Escalation → `Skill(effort-max)`, then the skill's own level again once - the diagnosis is produced. Automatic points: verify-secure loop caps - (GATE 0 floor, GATE 1 conformity, GATE 2 security) and ship-feature - STEP 4b. Not automatic, by doctrine: the challenge fail-safe (a mute - challenger is an infrastructure failure) and "gone WRONG → STOP" (STOP - precedes any further reasoning); their STOP text names the level - reached and suggests `/effort-max` for the relaunch. - -## Re-assert - -- After any nested `Skill(...)` whose frontmatter carries a different - effort (feat → commit-change), reload the orchestrator's own level. -- After a prose gate that ends the turn, the resumed turn runs at the - session level. If the resumed phase is reflection, its first step is - `Skill(effort-)`; dispatch and orchestration phases need - nothing. - -## Never - -- A shift never inside a dispatched agent: pins rule there. -- Max is for diagnosis, not for retrying the same fix harder. -EOF -``` - -- [ ] **Step 4: Model-gate paragraph (append to §4)** - -```bash -cat >> lib/model-gate.md <<'EOF' - -Effort is the second axis of the same table (BDR-NEXT): every typed agent -carries an `effort:` pin next to `model:`, and the main loop shifts per phase -through `lib/effort-shift.md`. Nothing dispatched inherits either axis. -EOF -``` - -- [ ] **Step 5: Suites** - -Run: `make test suite=lib/tests/effort-routing.test.sh && make test suite=lib/tests/skill-routing-census.test.sh && make test suite=lib/tests/profile-census.test.sh` -Expected: all green; the routing census prints no `effort-` pair under WARN. - -- [ ] **Step 6: Main-session smoke (interactive session only, cannot be headless)** - -In an interactive session in this repo, after `/reload-skills`: call `Skill(effort-max)`, then Bash `echo $CLAUDE_EFFORT`, then `Skill(effort-xhigh)`, then the echo again. -Expected: `max`, then `xhigh`. Transcript check for the cache: -```bash -f=~/.claude/projects/-home-bchanot-Documents-claude/$(ls -t ~/.claude/projects/-home-bchanot-Documents-claude/ | grep jsonl | head -1) -python3 -c " -import json,sys -for l in open('$f',errors='ignore'): - r=json.loads(l) - if r.get('type')=='assistant': - u=r['message'].get('usage') or {}; print(r.get('effort'), u.get('cache_creation_input_tokens'), u.get('cache_read_input_tokens'))" | tail -6 -``` -Expected: the first `max` row has `cache_creation` in the low thousands and `cache_read` unchanged from the row before (no cache bust). - -- [ ] **Step 7: Commit** - -```bash -git add skills/effort-*/SKILL.md lib/effort-shift.md lib/model-gate.md lib/tests/effort-routing.test.sh -git commit -m "feat(effort): five shifter skills, lib/effort-shift.md, model-gate second axis" -``` - ---- - -### Task 6: Orchestrator wiring (include line, medium at dispatch, own level at challenge, low at the tail, nested re-assert) - -**Files:** -- Modify: `skills/{feat,hotfix,bugfix,ship-feature,init-project,onboard,tour,code-clean,seo,geo,harden,web-validate,audit-delta}/SKILL.md`, `agents/client-handover-writer.md` (the client-handover skill loads this agent inline; its dispatches live there) -- Test: `lib/tests/effort-routing.test.sh` - -**Interfaces:** -- Consumes: shifter names and the include path from Task 5. -- Produces: helpers `ins_before`, `ins_after`, `ins_after_para`, `ins_before_para` (local to this task's shell). `ins_before` is for anchors inside code blocks (a standalone `Agent(` line); the `_para` forms are for anchors inside prose, where a bare insertion would split a sentence. - -- [ ] **Step 1: Locks (above the summary block)** - -```bash -python3 - <<'PY' -p="lib/tests/effort-routing.test.sh"; s=open(p).read() -locks="""# ── 6) orchestrator wiring (spec D4) -for s in feat hotfix bugfix ship-feature init-project onboard tour code-clean seo geo harden web-validate audit-delta; do - has "skills/$s/SKILL.md" 'lib/effort-shift.md'; has "skills/$s/SKILL.md" 'Skill(effort-medium)'; done -has "agents/client-handover-writer.md" 'lib/effort-shift.md'; has "agents/client-handover-writer.md" 'Skill(effort-medium)' -for s in feat hotfix bugfix; do has "skills/$s/SKILL.md" 'Skill(effort-high)'; done -for s in ship-feature init-project onboard code-clean audit-delta; do has "skills/$s/SKILL.md" 'Skill(effort-xhigh)'; done -for s in seo geo harden web-validate; do has "skills/$s/SKILL.md" 'Skill(effort-high)'; done -for s in feat hotfix bugfix ship-feature init-project; do has "skills/$s/SKILL.md" 'Skill(effort-low)'; done -has "skills/feat/SKILL.md" 'effort-shift: nested commit-change' - -# ── summary""" -open(p,"w").write(s.replace("# ── summary", locks)) -PY -``` - -- [ ] **Step 2: Define the three insertion helpers (exact-string anchors, first occurrence)** - -```bash -ins_before() { python3 - "$1" "$2" "$3" <<'PY' -import sys; f,a,t=sys.argv[1:]; L=open(f).read().split("\n") -i=next(k for k,l in enumerate(L) if a in l); L[i:i]=t.split("\\n"); open(f,"w").write("\n".join(L)) -PY -} -ins_after() { python3 - "$1" "$2" "$3" <<'PY' -import sys; f,a,t=sys.argv[1:]; L=open(f).read().split("\n") -i=next(k for k,l in enumerate(L) if a in l); L[i+1:i+1]=t.split("\\n"); open(f,"w").write("\n".join(L)) -PY -} -ins_after_para() { python3 - "$1" "$2" "$3" <<'PY' -import sys; f,a,t=sys.argv[1:]; L=open(f).read().split("\n") -i=next(k for k,l in enumerate(L) if a in l) -j=next(k for k in range(i,len(L)) if L[k].strip()=="") -L[j:j]=t.split("\\n"); open(f,"w").write("\n".join(L)) -PY -} -ins_before_para() { python3 - "$1" "$2" "$3" <<'PY' -import sys; f,a,t=sys.argv[1:]; L=open(f).read().split("\n") -i=next(k for k,l in enumerate(L) if a in l) -j=next(k for k in range(i,-1,-1) if L[k].strip()=="")+1 # first line of the paragraph -L[j:j]=t.split("\\n"); open(f,"w").write("\n".join(L)) -PY -} -INC='EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-NEXT): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation.' -``` -`next(...)` raises `StopIteration` when an anchor is absent: that is the intended failure, fix the anchor rather than the helper. - -- [ ] **Step 3: Include line, after the model-gate paragraph, in all 13 skills and the writer agent** - -```bash -for s in feat hotfix bugfix ship-feature init-project onboard tour code-clean seo geo harden web-validate audit-delta; do - ins_after_para "skills/$s/SKILL.md" 'lib/model-gate.md' "$INC"; done -ins_after_para agents/client-handover-writer.md 'model: "fable"' "$INC" -grep -c 'lib/effort-shift.md' skills/*/SKILL.md agents/client-handover-writer.md | grep -v ':0' | wc -l -``` -Expected: `14`. - -- [ ] **Step 4: Medium at the first executor/collector dispatch (anchors verified in the repo on 2026-09-28)** - -```bash -M='Skill(effort-medium) # effort-shift: dispatch span starts' -ins_before skills/feat/SKILL.md 'Agent(subagent_type="feater")' "$M" -ins_before skills/hotfix/SKILL.md 'Agent(subagent_type="hotfixer")' "$M" -ins_before skills/bugfix/SKILL.md 'Agent(subagent_type="bugfixer")' "$M" -ins_before skills/code-clean/SKILL.md 'Agent(subagent_type="code-cleaner")' "$M" -ins_before skills/seo/SKILL.md 'Agent(subagent_type="seo-analyzer", model="sonnet")' "$M" -ins_before skills/geo/SKILL.md 'Agent(subagent_type="geo-analyzer", model="sonnet")' "$M" -ins_before skills/web-validate/SKILL.md 'Agent(' "$M" -ins_before skills/harden/SKILL.md 'Agent(' "$M" -ins_after skills/ship-feature/SKILL.md '## STEP 4 — IMPLEMENT' "First: \`Skill(effort-medium)\` (effort-shift: dispatch span starts)." -ins_after skills/init-project/SKILL.md '## STEP 8 — IMPLEMENT' "First: \`Skill(effort-medium)\` (effort-shift: dispatch span starts)." -ins_before_para skills/onboard/SKILL.md 'Agent(subagent_type="onboarder")' "\`Skill(effort-medium)\` first (effort-shift: dispatch span starts)." -ins_after agents/client-handover-writer.md '## STEP 3 — BASELINE AUDITS' "First: \`Skill(effort-medium)\` (effort-shift: dispatch span starts)." -ins_before skills/tour/SKILL.md 'Agent(subagent_type="general-purpose",' "$M" -ins_before skills/audit-delta/SKILL.md 'Agent(subagent_type="security-auditor", description="audit-delta security' "$M" -``` -Anchors verified 2026-09-28: `web-validate` and `harden` open their first dispatch with a bare `Agent(` line (l.181 and l.262), the first `Agent(` in each file; `seo` (l.326) and `geo` (l.48) open the collect dispatch with the full call line, unique as first occurrence. `MODE: collect` is not an anchor: it sits inside the prompt string. - -- [ ] **Step 5: Own level before every challenge-plan call (reflection resumes)** - -```bash -for s in feat hotfix bugfix seo geo harden web-validate; do - ins_before_para "skills/$s/SKILL.md" 'lib/challenge-plan.md' "\`Skill(effort-high)\` first (effort-shift: reflection resumes)."; done -for s in ship-feature init-project onboard code-clean audit-delta; do - ins_before_para "skills/$s/SKILL.md" 'lib/challenge-plan.md' "\`Skill(effort-xhigh)\` first (effort-shift: reflection resumes)."; done -``` -The challenge include is referenced mid-sentence in every skill (`… harden it. Run\n\`$HOME/.claude/lib/challenge-plan.md\` with …`), hence the paragraph form. -```bash -``` -`seo`, `geo` and `web-validate` dispatch their applier after the challenge (seo l.557, geo l.117, web-validate l.312; the first `Agent(subagent_type="hotfixer")` in each file), so a second medium shift goes there: -```bash -for s in seo geo web-validate; do ins_before "skills/$s/SKILL.md" 'Agent(subagent_type="hotfixer")' "$M"; done -``` -`harden` applies inline in its STEP 3 (main loop, after the user's confirmation): no applier dispatch, the own-level shift before its challenge line is its last shift. - -- [ ] **Step 6: Low at the bookkeeping tail (the five skills with a memory-commit include)** - -```bash -for s in feat hotfix bugfix ship-feature init-project; do - ins_before_para "skills/$s/SKILL.md" 'lib/capitalize-commit.md' "\`Skill(effort-low)\` first (effort-shift: bookkeeping tail).\\n"; done -``` - -- [ ] **Step 7: Nested re-assert in feat (commit-change runs at low)** - -```bash -grep -n -E 'commit-change' skills/feat/SKILL.md -``` -Expected: two consecutive prose lines near l.199 (`… or run \`/commit-change\` on the pending work (it dispatches the …`). The sentence continues, so append at the end of that paragraph, not after the line: -```bash -ins_after_para skills/feat/SKILL.md '/commit-change' "Then \`Skill(effort-high)\` (effort-shift: nested commit-change loaded at low; reload feat's level)." -``` - -- [ ] **Step 8: Suites, then a read-through of each edited file around the insertions** - -Run: `make test suite=lib/tests/effort-routing.test.sh && make test suite=lib/tests/model-routing.test.sh && make test suite=lib/tests/loops-light.test.sh` -Expected: all green (the model-routing and loops-light locks match single lines that this task never splits). - -```bash -git diff -U1 -- skills agents | grep -E '^\+' | grep -v '^+++' | wc -l -``` -Expected: about 40 added lines, none inside a YAML frontmatter block (every insertion sits below the second `---`). - -- [ ] **Step 9: Commit** - -```bash -git add skills/*/SKILL.md agents/client-handover-writer.md lib/tests/effort-routing.test.sh -git commit -m "feat(effort): wire phase shifts in the 13 orchestrators and the handover writer" -``` - ---- - -### Task 7: Escalation points at max (loop caps, ship-feature 4b) and STOP texts - -**Files:** -- Modify: `lib/verify-secure-loop.md` (the three `**Max 3 … iterations** → STOP + human escalation` sentences, lines 38, 77, 107 on 2026-09-28) -- Modify: `skills/ship-feature/SKILL.md` (STEP 4b, `1. Load \`$HOME/.claude/agents/analyzer.md\` in DEBUG MODE`, and step 4 `If A →` / `If B →`) -- Modify: `lib/challenge-plan.md` (line `retry ONCE with a fresh challenger; a 2nd failure on that lens → STOP and escalate`) -- Test: `lib/tests/effort-routing.test.sh` - -**Interfaces:** -- Consumes: `Skill(effort-max)` and `/effort-max` from Task 5. - -- [ ] **Step 1: Locks** - -```bash -python3 - <<'PY' -p="lib/tests/effort-routing.test.sh"; s=open(p).read() -locks="""# ── 7) escalation at max (spec D4) -[ "$(grep -c 'Skill(effort-max)' "$R/lib/verify-secure-loop.md")" -eq 3 ] && ok || ko "verify-secure-loop.md must shift to max at its 3 caps" -has "skills/ship-feature/SKILL.md" 'Skill(effort-max)' -has "lib/challenge-plan.md" '/effort-max' -has "lib/verify-secure-loop.md" '/effort-max' - -# ── summary""" -open(p,"w").write(s.replace("# ── summary", locks)) -PY -``` - -- [ ] **Step 2: Loop caps** - -```bash -python3 - <<'PY' -p="lib/verify-secure-loop.md"; s=open(p).read() -for cap in ("floor","conformity","security"): - old=f"**Max 3 {cap} iterations** → STOP + human escalation" - new=(f"**Max 3 {cap} iterations** → `Skill(effort-max)` (effort-shift: cap reached, " - f"diagnose at max before escalating), then STOP + human escalation") - assert s.count(old)==1, cap; s=s.replace(old,new) -s=s.replace("STOP + human escalation with the\n BLOCKING table.", - "STOP + human escalation with the\n BLOCKING table. Every STOP text names the level reached (`$CLAUDE_EFFORT`)\n and suggests `/effort-max` for the relaunch.") -open(p,"w").write(s) -PY -grep -c 'Skill(effort-max)' lib/verify-secure-loop.md; grep -c '/effort-max' lib/verify-secure-loop.md -``` -Expected: `3` and `1`. - -- [ ] **Step 3: ship-feature 4b** - -```bash -python3 - <<'PY' -p="skills/ship-feature/SKILL.md"; s=open(p).read() -old="1. Load `$HOME/.claude/agents/analyzer.md` in DEBUG MODE on the exact error output." -assert s.count(old)==1 -s=s.replace(old, "1. `Skill(effort-max)` (effort-shift: error recovery), then load\n `$HOME/.claude/agents/analyzer.md` in DEBUG MODE on the exact error output.") -old2="4. If A → apply minimal fix, re-run STEP 4 for the failed task only." -assert s.count(old2)==1 -s=s.replace(old2, "4. On resume the turn is at the session level (effort-shift: turn reset).\n If A → `Skill(effort-medium)`, apply minimal fix, re-run STEP 4 for the failed task only.") -old3=" If B → before skipping:" -assert s.count(old3)==1 -s=s.replace(old3, " If B or C → `Skill(effort-xhigh)` first.\n If B → before skipping:") -open(p,"w").write(s) -PY -``` - -- [ ] **Step 4: Challenge fail-safe STOP text** - -```bash -python3 - <<'PY' -p="lib/challenge-plan.md"; s=open(p).read() -old="retry ONCE with a fresh challenger; a 2nd failure on that lens → STOP and escalate" -assert s.count(old)==1 -s=s.replace(old, old+"\n(the STOP text names the level reached, `$CLAUDE_EFFORT`, and suggests `/effort-max`\nfor the relaunch; no shift here: a mute challenger is an infrastructure failure)") -open(p,"w").write(s) -PY -``` - -- [ ] **Step 5: Suites and commit** - -Run: `make test` -Expected: green. - -```bash -git add lib/verify-secure-loop.md lib/challenge-plan.md skills/ship-feature/SKILL.md lib/tests/effort-routing.test.sh -git commit -m "feat(effort): max at the verify-secure caps and ship-feature 4b; STOP texts suggest /effort-max" -``` - ---- - -### Task 8: Turn-ending gate audit and re-assert - -**Files:** -- Modify: `skills/bugfix/SKILL.md` (the gate `behavior change): wait for user approval.` before pass B of the contract interview) -- Possibly modify: any other orchestrator where the audit below finds a prose gate followed by reflection -- Test: `lib/tests/effort-routing.test.sh` - -- [ ] **Step 1: Lock** - -```bash -python3 - <<'PY' -p="lib/tests/effort-routing.test.sh"; s=open(p).read() -locks="""# ── 8) turn-reset re-assert after a prose gate followed by reflection -has "skills/bugfix/SKILL.md" 'effort-shift: turn reset' - -# ── summary""" -open(p,"w").write(s.replace("# ── summary", locks)) -PY -``` - -- [ ] **Step 2: Audit every prose gate** - -```bash -grep -n -i -E "end the turn|end your turn|wait for (the )?(user|human)|STOP and wait|wait for user" \ - skills/{feat,hotfix,bugfix,ship-feature,init-project,onboard,tour,code-clean,seo,geo,harden,web-validate,audit-delta}/SKILL.md \ - lib/contract-interview.md lib/challenge-plan.md lib/plugin-gate.md lib/verify-secure-loop.md -``` -Known on 2026-09-28: `bugfix:119` (resume = contract pass B, reflection → re-assert), `ship-feature:205` (handled in Task 7), `ship-feature:14` and `init-project:14` (model-gate STOP, the run ends → nothing). Classify every other hit the same way: model-gate STOP or loop-cap STOP → nothing; resume into dispatch/orchestration → nothing; resume into reflection → re-assert with the skill's own level. - -- [ ] **Step 3: Re-assert in bugfix** - -```bash -python3 - <<'PY' -p="skills/bugfix/SKILL.md"; s=open(p).read() -old=" behavior change): wait for user approval.\n" -assert s.count(old)==1 -s=s.replace(old, old+" On resume: `Skill(effort-high)` first (effort-shift: turn reset).\n") -open(p,"w").write(s) -PY -``` -Apply the same one-line pattern to any other reflection resume found in Step 2, with that skill's level. - -- [ ] **Step 4: Manual verification of the reset itself (interactive, once)** - -In an interactive session: type `/effort-max`, wait for the reply, then send a plain message such as `echo test` and read the transcript: -```bash -f=~/.claude/projects/-home-bchanot-Documents-claude/$(ls -t ~/.claude/projects/-home-bchanot-Documents-claude/ | grep jsonl | head -1) -grep -o '"effort":"[a-z]*"' "$f" | tail -4 -``` -Expected: `max` rows for the first turn, `high` for the second (the session default from Task 2). - -- [ ] **Step 5: Suites and commit** - -Run: `make test` -```bash -git add skills/*/SKILL.md lib/tests/effort-routing.test.sh -git commit -m "feat(effort): re-assert the skill level after prose gates that end the turn" -``` - ---- - -### Task 9: Vendored superpowers patch with resync re-apply - -**Files:** -- Modify: `skills/brainstorming/SKILL.md`, `skills/writing-plans/SKILL.md` (line 2 `name: …`) -- Modify: `install-plugins.sh` (end of the STEP 8e block that vendors the 7 superpowers skills) -- Test: `lib/tests/effort-routing.test.sh` - -- [ ] **Step 1: Locks** - -```bash -python3 - <<'PY' -p="lib/tests/effort-routing.test.sh"; s=open(p).read() -locks="""# ── 9) vendored superpowers carry xhigh; a resync that drops it fails here (spec D3) -for s in brainstorming writing-plans; do fm_has_effort "skills/$s/SKILL.md" xhigh; done -has "install-plugins.sh" 'effort: xhigh' - -# ── summary""" -open(p,"w").write(s.replace("# ── summary", locks)) -PY -``` - -- [ ] **Step 2: Patch the two files (same `lvl` helper as Task 4)** - -```bash -lvl() { L=$1; shift; for s in "$@"; do - sed -i "0,/^name: $s\$/s//&\neffort: $L/" "skills/$s/SKILL.md"; done; } -lvl xhigh brainstorming writing-plans -sed -n 1,4p skills/brainstorming/SKILL.md skills/writing-plans/SKILL.md -``` -Expected: `effort: xhigh` on line 3 of both. - -- [ ] **Step 3: Re-apply after every resync in install-plugins.sh** - -```bash -grep -n -i 'STEP 8e' install-plugins.sh -``` -Expected: three hits on 2026-09-28: a cross-reference comment near l.535, the heading `# ── Step 8e: Agent Skills …` near l.908, and its `echo` near l.915. The block ends where the `# ====` banner of STEP 8.5 begins (near l.937). Insert the re-apply right before that banner: -```bash -python3 - <<'PY' -p="install-plugins.sh"; L=open(p).read().split("\n") -i=next(k for k,l in enumerate(L) if l.startswith("# ── Step 8e:")) -j=next(k for k in range(i+1,len(L)) if L[k].startswith("# ====")) # the STEP 8.5 banner -L[j:j]=[ -"# Effort tiering (BDR-NEXT): the vendored brainstorming/writing-plans carry an", -"# effort pin upstream lacks; re-apply after every resync (census lock in", -"# lib/tests/effort-routing.test.sh alarms if this ever stops working).", -"for _s in brainstorming writing-plans; do", -" _f=\"$(cd \"$(dirname \"$0\")\" && pwd)/skills/$_s/SKILL.md\"", -" if [ -f \"$_f\" ] && ! grep -q '^effort:' \"$_f\"; then", -" sed -i \"0,/^name: $_s\\$/s//&\\neffort: xhigh/\" \"$_f\"", -" fi", -"done", -"unset _s _f", -"", -] -open(p,"w").write("\n".join(L)) -PY -shellcheck install-plugins.sh -``` -Expected: shellcheck silent, and `sed -n '/^unset _s _f/,+2p' install-plugins.sh` shows the blank line then the `# ====` banner of STEP 8.5. - -- [ ] **Step 4: Prove the re-apply works** - -```bash -sed -i '/^effort: xhigh$/d' skills/brainstorming/SKILL.md -bash -c 'source /dev/stdin <<<"$(sed -n "/Effort tiering (BDR-NEXT)/,/^unset _s _f/p" install-plugins.sh)"' -grep -c '^effort: xhigh' skills/brainstorming/SKILL.md -``` -Expected: `1` (the extracted block re-added the line without running the whole installer). - -- [ ] **Step 5: Suites and commit** - -Run: `make test` -```bash -git add skills/brainstorming/SKILL.md skills/writing-plans/SKILL.md install-plugins.sh lib/tests/effort-routing.test.sh -git commit -m "feat(effort): xhigh on the vendored brainstorming and writing-plans, re-applied at resync" -``` - ---- - -### Task 10: BDR id, CHANGELOG, registries, journal, TODO reconcile - -**Files:** -- Modify: `lib/effort-shift.md`, `lib/model-gate.md`, `lib/tests/effort-routing.test.sh`, `hooks/session-start.sh`, `hooks/statusline.sh`, `install-plugins.sh`, 14 orchestrator files (every `BDR-NEXT` token) -- Modify: `CHANGELOG.md` (`## [Unreleased]` → `### Added`), `.claude/memory/decisions.md`, `.claude/memory/evals.md`, `.claude/memory/journal.md`, `.claude/tasks/TODO.md` - -- [ ] **Step 1: Compute the id and replace the token everywhere** - -```bash -N=$(( $(grep -o -E 'BDR-[0-9]+' .claude/memory/decisions.md | sort -t- -k2 -n | tail -1 | cut -d- -f2) + 1 )) -echo "BDR-$N" -grep -rl 'BDR-NEXT' --include='*.md' --include='*.sh' --include='*.json' . | grep -v '^./docs/superpowers/' | xargs sed -i "s/BDR-NEXT/BDR-$N/g" -grep -rn 'BDR-NEXT' . --include='*.md' --include='*.sh' | grep -v '^./docs/superpowers/' | wc -l -``` -Expected: `0` (the spec and this plan keep the token as history). - -- [ ] **Step 2: CHANGELOG under `## [Unreleased]` → `### Added` (first bullet position)** - -```bash -python3 - <<'PY' -p="CHANGELOG.md"; s=open(p).read() -anchor="## [Unreleased]\n\n### Added\n" -assert s.count(anchor)==1 -entry=("- **Effort tiering (BDR-$N)**: reasoning effort routed per role and per phase. " -"Session default `high`; `effort:` pins on the 20 repo-authored agents; entry level on " -"the 33 user-invoked skills (low → xhigh); five shifter skills `effort-low` … `effort-max` " -"loaded at phase boundaries through `lib/effort-shift.md`, with `max` at the verify-secure " -"caps and ship-feature 4b; `/effort-max` as the turn-scoped relaunch lever; statusline shows " -"the live level; session banner warns when `CLAUDE_CODE_EFFORT_LEVEL` silences the pins; " -"census `lib/tests/effort-routing.test.sh`.\n") -open(p,"w").write(s.replace(anchor, anchor+entry)) -PY -sed -i "s/BDR-\$N/BDR-$N/" CHANGELOG.md -``` - -- [ ] **Step 3: BDR entry (index row after the last row, section at the end), caveman English** - -Index row (columns `| ID | Date | Decision | Status |` — copy the exact header of the table in `decisions.md` and match it): -``` -| BDR- | 2026-09-28 | Effort tiering: session high, agent effort pins (BDR-077 second axis), skill entry levels, five shifter skills for phase shifts, max at loop caps + 4b | accepted | -``` -Section: -``` -## BDR- — Effort tiering: session high, pins, skill levels, phase shifts, max at escalation [accepted] (2026-09-28) -- **Decision**: settings `effortLevel` high; `effort:` pin on 20 repo-authored agents by role (low appliers, medium executors, high judgment on sonnet/opus, xhigh challengers + gates); `effort:` on 33 user-invoked skills = run entry level; `skills/effort-{low,medium,high,xhigh,max}` loaded by orchestrators at phase boundaries per `lib/effort-shift.md` (medium at dispatch, own level before challenge synthesis, low at bookkeeping tail, max at verify-secure caps + ship-feature 4b); STOP texts suggest `/effort-max`; statusline live level; banner warns on `CLAUDE_CODE_EFFORT_LEVEL`; census `lib/tests/effort-routing.test.sh`. -- **Why**: session-wide xhigh burned thinking on bookkeeping; measurement (EVAL-035) put 97 % of thinking in the main loop, so the main-loop lever (skill effort, verified LRN-179) carries the savings; pins = explicitness + future models. -- **Alternatives rejected**: executor pins only (executors think 26 tok/request); escalation-diagnoser agent fable+max (no context, one more agent, main-loop max keeps the failure context); reflection in fable skill-runner children with session medium (loses interactivity/context); settings.json rewrite mid-run (global side effect, LRN-098); `maxEffortLevel` caps (hide a mis-pin the census should fail). -- **Caveats**: shifts inert in `-p`/SDK; a turn-ending prose gate resets to session level (re-assert wired where reflection resumes); one effort per agent file → mode-based agents pin their judgment mode; vendored superpowers patch re-applied by install-plugins STEP 8e. -- **Refs**: spec `docs/superpowers/specs/2026-09-28-effort-tiering-design.md`, plan `docs/superpowers/plans/2026-09-28-effort-tiering.md`, [[LRN-179]], [[EVAL-035]], [[BDR-077]]. -``` -Replace `` by the computed id. Insert the row after the last `| BDR-` row with the same python pattern as Task 4's lock insertion; append the section at the end of the file. - -- [ ] **Step 4: EVAL row + section for the Task 4 A/B (columns `| ID | Date | Output | Action |`)** - -``` -| EVAL-036 | 2026-09-28 | A/B `/reconcile` headless, session high vs skill low (Task 4): requests , output , thinking , ms | keep low on bookkeeping skills; repeat on a reflection skill before touching the medium/high split | -``` -Section with `- **Date**`, `- **Method**` (the Task 4 Step 1 command), `- **Result**` (the two lines), `- **Anomaly**` (anything odd: for example thinking near zero in both runs means effort did not matter for that skill), `- **Action**`. Use the next free EVAL id (`grep -o -E 'EVAL-[0-9]+' .claude/memory/evals.md | sort -t- -k2 -n | tail -1`). - -- [ ] **Step 5: Journal line and TODO reconcile** - -Append under today's heading in `.claude/memory/journal.md` (create the `## 2026-09-28` heading if absent): `- effort tiering shipped on feature/effort-tiering: session high, 20 pins, 33 skill levels, 5 shifters, max at caps + 4b; census green; finish awaits human signal.` -In `.claude/tasks/TODO.md`, tick the four wave checkboxes of the `effort tiering` section. - -- [ ] **Step 6: Suites, then commit code and docs, then the memory surgically** - -Run: `make test && shellcheck *.sh hooks/*.sh lib/*.sh` -```bash -git add CHANGELOG.md lib hooks install-plugins.sh skills agents -git commit -m "docs(effort): BDR-$N id, CHANGELOG entry" -bash lib/memory-commit.sh commit "chore(memory): BDR-$N effort tiering, EVAL A/B, journal, TODO" -git status --short -``` -Expected: clean tree, both commits pushed by the post-commit hook. - ---- - -### Task 11: Keep the transcript audit script (spec §9 tooling) - -**Files:** -- Create: `lib/effort-audit.py` (from the spike's `effort_split2.py`, cleaned: functions ≤ 25 logic lines, 80-char lines, no globals beyond constants) -- Modify: `lib/effort-shift.md` (one line under Mechanics: `Measure with python3 ~/.claude/lib/effort-audit.py [projects-root]`) -- Test: `lib/tests/effort-routing.test.sh` - -- [ ] **Step 1: Lock** - -```bash -python3 - <<'PY' -p="lib/tests/effort-routing.test.sh"; s=open(p).read() -locks="""# ── 11) audit tooling -has "lib/effort-shift.md" 'effort-audit.py' -[ -x "$R/lib/effort-audit.py" ] && ok || ko "lib/effort-audit.py missing or not executable" - -# ── summary""" -open(p,"w").write(s.replace("# ── summary", locks)) -PY -``` - -- [ ] **Step 2: Write the script** - -```bash -cat > lib/effort-audit.py <<'EOF' -#!/usr/bin/env python3 -"""Sum output/thinking/cache tokens per (scope, model, effort) over Claude Code -transcripts. scope = main (session jsonl) | sub (subagents/*.jsonl or -isSidechain records). Read-only. Usage: effort-audit.py [projects-root]""" -import collections -import glob -import json -import os -import sys - -WEIGHTS = {"in": 1.0, "cc": 1.25, "cr": 0.1, "out": 5.0} # relative to input price -FIELDS = ("in", "cc", "cr", "out", "think") - - -def usage_row(usage): - """Map one API usage block to the five counted fields.""" - details = usage.get("output_tokens_details") or {} - return { - "in": usage.get("input_tokens", 0) or 0, - "cc": usage.get("cache_creation_input_tokens", 0) or 0, - "cr": usage.get("cache_read_input_tokens", 0) or 0, - "out": usage.get("output_tokens", 0) or 0, - "think": details.get("thinking_tokens", 0) or 0, - } - - -def scan(path, scope, agg): - """Add every assistant record of one transcript to agg.""" - with open(path, errors="ignore") as handle: - for line in handle: - try: - rec = json.loads(line) - except ValueError: - continue - msg = rec.get("message") or {} - if rec.get("type") != "assistant" or not msg.get("usage"): - continue - sub = scope == "sub" or bool(rec.get("isSidechain")) - key = ("sub" if sub else "main", - str(msg.get("model", "?")).replace("claude-", ""), - str(rec.get("effort") or "?")) - row = usage_row(msg["usage"]) - agg[key]["msgs"] += 1 - for field in FIELDS: - agg[key][field] += row[field] - - -def weighted(counter): - return sum(counter[f] * WEIGHTS[f] for f in WEIGHTS) - - -def report(agg): - """Print the per-key table, then the main/sub split and the thinking share.""" - total = collections.Counter() - for counter in agg.values(): - total.update(counter) - total_w = weighted(total) or 1 - print(f"{'scope':5} {'model':22} {'effort':7} {'msgs':>6} {'think/msg':>9} " - f"{'think_tok':>10} {'out_tok':>10} {'cache_read':>12} {'%wcost':>7}") - for (scope, model, effort), c in sorted(agg.items(), key=lambda kv: -weighted(kv[1])): - per_msg = c["think"] / max(c["msgs"], 1) - print(f"{scope:5} {model:22} {effort:7} {c['msgs']:6d} {per_msg:9.0f} " - f"{c['think']:10d} {c['out']:10d} {c['cr']:12d} {100 * weighted(c) / total_w:6.1f}%") - by_scope = collections.defaultdict(collections.Counter) - for (scope, _, _), c in agg.items(): - by_scope[scope].update(c) - for scope, c in by_scope.items(): - print(f" {scope:5} weighted-cost {100 * weighted(c) / total_w:5.1f}% " - f"thinking {100 * c['think'] / max(total['think'], 1):5.1f}% requests {c['msgs']}") - print(f" thinking = {100 * total['think'] * WEIGHTS['out'] / total_w:.1f}% of weighted cost; " - f"cache reads = {100 * total['cr'] * WEIGHTS['cr'] / total_w:.1f}%") - - -def main(): - root = os.path.expanduser(sys.argv[1] if len(sys.argv) > 1 else "~/.claude/projects") - agg = collections.defaultdict(collections.Counter) - for project in sorted(glob.glob(os.path.join(root, "*"))): - if not os.path.isdir(project): - continue - for path in glob.glob(os.path.join(project, "*.jsonl")): - scan(path, "main", agg) - for path in glob.glob(os.path.join(project, "*", "subagents", "*.jsonl")): - scan(path, "sub", agg) - report(agg) - - -if __name__ == "__main__": - main() -EOF -chmod +x lib/effort-audit.py -python3 lib/effort-audit.py | head -5 -``` -Expected: the table header and the top rows, `main fable-5-1` first. - -- [ ] **Step 3: Pointer in the include, suites, commit** - -```bash -python3 - <<'PY' -p="lib/effort-shift.md"; s=open(p).read() -anchor="## Shifters\n" -assert s.count(anchor)==1 -s=s.replace(anchor, "Measure the split any time: `python3 ~/.claude/lib/effort-audit.py`\n(thinking/output/cache tokens per scope, model and effort).\n\n"+anchor) -open(p,"w").write(s) -PY -make test suite=lib/tests/effort-routing.test.sh -git add lib/effort-audit.py lib/effort-shift.md lib/tests/effort-routing.test.sh -git commit -m "feat(effort): transcript audit script for the thinking/cost split" -``` - ---- - -## Self-review against the spec - -- **§4 D1** → Task 2 (settings, banner, statusline). **D2** → Task 3. **D3** → Tasks 4 and 9. **D4** → Tasks 5, 6, 7, 8. **D5** → Task 2. **§6** every file listed has a task. **§7** every census item has a lock: 1 (Task 3 `fm_has_effort`/`fm_no_effort`), 2 (Task 3), 3 (Tasks 4, 9), 4 (Task 5), 5 (Tasks 6, 7), 6 (Task 1), 7 and 8 run as existing suites in `make test`. **§8** waves = Tasks 1-3 / 4, 9 / 5-8 / 10-11. **§9** → Task 4 Steps 1 and 5, EVAL in Task 10, tooling in Task 11. -- **Placeholders**: `BDR-NEXT` is a defined token with a defined replacement step (Task 10); the three `grep -n -m1` recipes in Task 6 Step 4 and Task 8 Step 2 name the expected match and the exact insertion to make. -- **Names**: `Skill(effort-)`, `lib/effort-shift.md`, `fm_has_effort`, `ins_before`/`ins_after`/`ins_after_para`, `lvl`, `pin` are spelled identically across tasks. -- **Review Focus**: 1 → Task 5 lock `Headless sessions`; 2 → Task 2 Step 6; 3 → Task 6 Step 7 + lock; 4 → Task 3 `fm_no_effort status-reporter`; 5 → Task 9 Steps 3-4 + lock. diff --git a/docs/superpowers/specs/2026-09-28-effort-tiering-design.md b/docs/superpowers/specs/2026-09-28-effort-tiering-design.md deleted file mode 100644 index 79a3419..0000000 --- a/docs/superpowers/specs/2026-09-28-effort-tiering-design.md +++ /dev/null @@ -1,250 +0,0 @@ -# Effort tiering — design - -Date: 2026-09-28 · Branch: `feature/effort-tiering` · Status: draft for review - -## 1. Intent - -Adapt the reasoning effort along a development run, not hold the whole -session at `xhigh`. The user's five-rung scale is the contract: - -| Rung | User definition | Examples | -|---|---|---| -| low | fix a line, rename a file, run a script | journal, commit, release bookkeeping | -| medium | day-to-day work | implement a closed plan, orchestrate between dispatches | -| high | a refactor, a bug that resists | investigation, diagnosis, contract drafting | -| xhigh | architecture, audit before validation | brainstorm, plan, challenge synthesis, gates | -| max | a stuck error, an error that cannot be recovered, or judged need | loop caps, error recovery | - -Automatic wherever the harness allows it. Where it does not, the user gets a -one-keystroke lever, never a silent default. - -Effort is a second axis on the BDR-077 routing table: BDR-077 fixed WHICH -MODEL runs each role and forbade inherit; this design fixes HOW HARD it -thinks, with the same no-inherit principle. - -## 2. What the harness allows (verified on Claude Code 2.1.283, 2026-09-28) - -Sources: code.claude.com/docs (model-config, skills, sub-agents, hooks), -the CHANGELOG (2.1.120, 2.1.149, 2.1.267, 2.1.280) and live probes in this -repo. - -| Mechanism | Verified behaviour | Evidence | -|---|---|---| -| Session level | Resolution order: `CLAUDE_CODE_EFFORT_LEVEL` env > `--effort` / `/effort` > settings (`modelSettings` per model, else top-level `effortLevel`) > model default (`high` on Fable 5.1). `max` is session-only, never persisted. `/effort auto` clears the per-model saved level only; a top-level `effortLevel` still applies. | docs | -| Subagent frontmatter `effort:` | Applied to the subagent. Absent → **inherits the session level**. | built-in on sonnet printed `xhigh`; impeccable agent pinned `medium` printed `medium` | -| Skill frontmatter `effort:`, user-typed `/skill` | Applied for the **rest of the turn**, AskUserQuestion included. | headless `/effort-probe-low`: every request at `low` | -| Skill frontmatter `effort:`, loaded by Claude through the Skill tool, **interactive** session | Applied for the rest of the turn. Last loaded skill wins, up and down. | this session: `xhigh` → probe max → `$CLAUDE_EFFORT=max`, request records `effort=max` → probe xhigh → back to `xhigh` | -| Same, pairing rule | Applies **only when the Skill call shares the assistant message with another tool call after it**; a lone Skill call is a no-op. The paired call already runs at the new level. | this session, 8/8 observations | -| Same, re-load | A shifter already loaded in the conversation re-applies its effort when loaded again (paired); only its text is deduped. | this session | -| Same, **headless** (`-p`) | **Not applied** (neither `effort:` nor `model:`). | three `-p` runs, transcript effort unchanged | -| Prompt cache on a mid-turn shift | **Preserved** on Fable 5.1: first request at max read 206,996 cached tokens, wrote 1,164. | this session | -| Agent tool call site | No `effort` parameter (only `model`). One agent file = one effort. | tool schema | -| Hooks | Read `$CLAUDE_EFFORT` / `effort.level`; **cannot change** the level. | docs | -| `ultrathink` keyword | In-context nudge only; the effort sent to the API is unchanged. | docs | -| Env var | `CLAUDE_CODE_EFFORT_LEVEL` beats every frontmatter override. Unset on this machine. | docs + `env` | - -## 3. What the numbers say (6 days of local transcripts, all projects, 10,955 requests) - -Weights relative to input price: output ×5, cache read ×0.1, cache write ×1.25. - -| Item | Share | -|---|---| -| Cache reads (context re-read per request) | 53 % of weighted spend | -| All output tokens | 16 % | -| of which thinking | 8 % | -| Thinking located in the main loop | 97 % of thinking | -| Mean thinking per request: Fable main loop / sonnet subagent at xhigh | 1,430 / 26 tokens | -| Mean cached context per main-loop request | ~320 k tokens | - -Consequences. Executors barely think even at xhigh: pinning them is about -explicitness and future models (Opus 5.5 "thinks more per turn at a given -level"), not savings today. The direct lever of effort is single-digit -percent; the indirect lever (fewer steps at lower effort → fewer requests → -fewer cache reads) is unmeasured and gets an A/B in §9. The dominant cost is -main-loop context size, out of scope here (see `/capitalize`, `/clear`). - -## 4. Decisions - -### D1. Session default `high` -`settings.json` `effortLevel`: `xhigh` → `high`, explicit rather than -deleted: the statusline reads the key, and LRN-139 wants a visible value to -sweep at every model bump. Interactive chat outside a skill runs at the -model default; the user raises with `/effort xhigh` (session) or the new -`/effort-max` shifter (turn, see D4). `CLAUDE_CODE_EFFORT_LEVEL` must stay -unset (it would silence every override below); the session-start banner -warns if it is set. - -### D2. Agent pins (approach A) — repo-authored agents only - -| effort | Agents | -|---|---| -| low | hotfixer, release-executor, plugin-probe, validator-analyzer | -| medium | feater, bugfixer, code-cleaner, onboarder, scaffolder (was `high`; citer `skills/init-project/SKILL.md:98` updated) | -| high | refactorer, analyzer, commit-changer, doc-syncer, handover-doc-writer | -| xhigh | plan-challenger, plugin-advisor, verifier, security-auditor, seo-analyzer, geo-analyzer | -| none | interviewer, client-handover-writer (inline-load only, a pin would be inert and misleading, BDR-076 precedent); status-reporter (haiku, no effort support); `impeccable-*` (vendored) | - -Rules. One effort per agent file, so a mode-based agent (BDR-077) pins the -level of its **judgment** mode and its mechanical modes over-tier: the -fail-safe direction, and free on sonnet per §3. Built-ins (Explore, -general-purpose, Plan) cannot be pinned at the call site and inherit the -main loop's current level; Explore on Fable thinks ~1 token per request, -so no wrapper agent is created. Verifier and security-auditor sit at xhigh -by the user's own definition ("audit before validation"); on sonnet the -cost difference is nil. - -### D3. Skill frontmatter effort (approach B) — the run's entry level -Applies from the user's invocation for the rest of the turn. - -| effort | Skills | -|---|---| -| low | status, commit-change, release-candidate, doc, capitalize, close, reconcile, deploy, profile, plugin-check | -| medium | gitflow, prune-memory | -| high | feat, hotfix, bugfix, refactor, web-validate, harden, seo, geo | -| xhigh | ship-feature, init-project, onboard, tour, audit-delta, analyze, code-clean, client-handover, brainstorming, writing-plans | -| unlisted | session default, by design: gstack skills (`spec` and `skillify` are gstack), plugin skills, and machine-generated skills (`graphify`, `find-docs`) | - -`brainstorming` and `writing-plans` are vendored superpowers skills living in -`skills-external/` (gitignored, symlinked into `skills/`): the pin is applied -to the real file and never committed; `install-plugins.sh` re-applies it after -every resync, and the census checks it whenever the file is present (visible -SKIP otherwise). - -A skill loaded by Claude as a sub-step (feat → commit-change) also shifts -the level for the rest of the turn (interactive, §2), so orchestrators -re-assert their own level after any nested Skill call whose level differs -(D4 protocol). - -### D4. Phase shifts inside a run (approach C) -Five one-line skills, no body beyond a sentence, user-invocable: -`effort-low`, `effort-medium`, `effort-high`, `effort-xhigh`, `effort-max`. -Descriptions as pre-validated against the routing census (pairwise -similarity ≤ 0.03). Protocol in a shared include `lib/effort-shift.md`, -mirroring `lib/model-gate.md`: - -- A shift is a `Skill(effort-)` call on the main loop. Never inside a - dispatched agent (agents run on their pin). One tool round-trip, - cache-safe (§2). -- **Pairing rule**: the shift is sent in the same assistant message as the - step's first tool call, shift first; a lone Skill call is a no-op (§2). - Re-loading a shifter re-applies its effort. -- Orchestrator wiring, three points each: `effort-medium` when the plan is - closed and the dispatch phase starts; `effort-low` before the - capitalize / journal / doc-commit tail; `effort-max` at an escalation - point, then the skill's own level again once the diagnosis is produced. -- Re-assert the skill's own level after any nested `Skill(...)` call whose - frontmatter carries a different effort (D3): the nested level would - otherwise hold for the rest of the turn. -- **Escalation points (automatic max)**: verify-secure loop GATE 1 cap - (3 conformity rounds) and GATE 2 cap (3 security rounds), before the - human-escalation table is composed; ship-feature STEP 4b, so the - inline analyzer DEBUG read runs at max. Full conversation context is the - asset here; a fresh diagnoser agent was considered and dropped (YAGNI: - no context, one more agent, same effort). -- **Not automatic, by doctrine**: the challenge fail-safe (a mute - challenger is an infrastructure failure, not a reasoning problem) and - the "gone WRONG → STOP" rule (STOP precedes any further reasoning). Both - STOP messages name the level reached and suggest `/effort-max` for the - relaunch: a turn-scoped max the user gets by typing one command. -- **Turn reset**: a prose gate that ends the turn (model-gate STOP, loop - cap STOP, and the four prose gates found in bugfix, ship-feature ×2, - init-project) drops the resumed turn to the session level. The plan - audits each such gate: if the resumed phase is reflection, the resume - step re-asserts with `Skill(effort-xhigh)`; if it is dispatch or - orchestration, session `high` is adequate and nothing is added. -- **Headless limitation**: `-p`, `claude agents` and SDK sessions ignore - skill-level effort (§2); runs there stay at the session level. Documented - in the include, no mitigation. - -### D5. Visibility -`hooks/statusline.sh` shows `$CLAUDE_EFFORT` when set (the live level, -shifts included) and falls back to the settings key. `/tasks` already shows -each subagent's effort (2.1.243). - -## 5. Alternatives rejected - -- **Keep xhigh, pin executors only**: executors think ~26 tokens per - request; the burn is in the main loop (§3). -- **Escalation diagnoser agent (`model: fable`, `effort: max`)**: chosen - before the interactive probe proved C viable; dropped because the - main-loop shift keeps the full failure context and adds no agent. -- **Move reflection into `model: fable` skill-runner children with - `effort: xhigh`, session at medium**: loses conversation context and - interactivity (BDR-077 retention criteria), heavy re-architecture for a - lever C delivers in five one-line files. -- **Rewrite settings.json mid-run to shift effort**: global side effect on - every session, LRN-098 drift class, fights the harness. -- **`maxEffortLevel` cap on sonnet**: pins already bound each agent; a cap - would hide a mis-pin instead of failing it in the census. - -## 6. Files touched - -| Area | Change | -|---|---| -| `settings.json` | `effortLevel` → `high` (curated config: read the diff, LRN-098) | -| `agents/*.md` (20) | `effort:` line per D2; `skills/init-project/SKILL.md:98` citer | -| `skills/*/SKILL.md` (28 tracked) + `skills-external/{brainstorming,writing-plans}/SKILL.md` (not committed) | `effort:` line per D3; `install-plugins.sh` re-applies the two vendored pins after resync | -| `skills/effort-{low,medium,high,xhigh,max}/SKILL.md` | new, frontmatter + one sentence | -| `lib/effort-shift.md` | new include: protocol, wiring points, escalation, turn reset, headless note | -| `lib/model-gate.md` §4 | one paragraph: effort is the second axis, pointer to the include | -| `lib/verify-secure-loop.md` | `Skill(effort-max)` before each cap's human-escalation table; STOP text names the level | -| `lib/challenge-plan.md` | STOP text names the level, suggests `/effort-max` | -| orchestrator SKILL.md (feat, hotfix, bugfix, ship-feature, init-project, onboard, tour, code-clean, seo, geo, harden, web-validate, client-handover, audit-delta) | include line + the three wiring points; ship-feature 4b max | -| `hooks/statusline.sh`, `hooks/session-start.sh` | live effort display; env-var warning | -| `lib/tests/effort-routing.test.sh` | new census suite (§7) | -| `lib/effort-audit.py` | transcript audit script (§9) | -| `CHANGELOG.md`, `.claude/memory/*` | release note; BDR + LRN + EVAL + journal (§8) | - -## 7. Tests and census (`make test`) - -New suite `lib/tests/effort-routing.test.sh`, `grep -qF` locks in the -`model-routing.test.sh` style, flip-tested first (BDR-100): - -1. Every repo-authored agent outside the "none" list has `effort: ` - in its first 10 frontmatter lines, level in the allowed set; the "none" - list has no `effort:`. -2. Tier locks per D2 (one `has` per agent). -3. Skill locks per D3 (one per skill); the two vendored skills are checked - when present (visible SKIP otherwise); `install-plugins.sh` carries the - re-apply block. -4. The five shifter skills exist with the exact `name:` and `effort:`. -5. `lib/effort-shift.md` is included by every orchestrator in the §6 list; - `verify-secure-loop.md` and `ship-feature/SKILL.md` contain the - `Skill(effort-max)` lock. -6. `settings.json` `effortLevel` is `high`. -7. `skill-routing-census` stays green with the five new descriptions - (pre-validated). -8. `doctrine-citers` stays green: no new `CLAUDE.md "…"` citation; the - doctrine lives in `lib/`. - -Per-wave smoke, planted input, disk-verified (BDR-077 precedent): -W1 dispatch a pinned agent that echoes `$CLAUDE_EFFORT`; W2 invoke `/status` -and read `effort=low` in the transcript; W3 run a skill through a shift -and read the request sequence; W4 statusline shows the live level. - -## 8. Rollout - -Four waves on `feature/effort-tiering`, one commit each, smoke as merge -gate, human signal for `gitflow finish`: - -- W1 settings + agent pins + test suite + model-gate paragraph. -- W2 skill frontmatter (D3) + superpowers patch. -- W3 shifter skills + `lib/effort-shift.md` + orchestrator wiring + - escalation points + turn-reset audit. -- W4 statusline + banner warning + CHANGELOG + registries. - -Registries: BDR (effort tiering, this spec's decisions and rejected -alternatives), LRN (skill effort applies on user invocation and on -interactive Skill-tool loads, not in `-p`; shifts are cache-safe), EVAL -(the §3 measurement and its method), journal line. - -## 9. Measurement after rollout - -A/B on a repeatable skill run (`/reconcile` on this repo, session `high` -vs `xhigh`): requests, output tokens, thinking tokens, wall time, from the -transcript. Records whether the indirect lever exists. Goes to EVAL. - -## 10. Out of scope - -Main-loop context size (the 53 %), gstack and plugin skills, the -`impeccable-*` agents, `graphify` (machine-owned), headless sessions.