- lib/effort-pins.txt (map) + lib/effort-pins.sh (idempotent re-apply) replace the hardcoded brainstorming/writing-plans loop; called after the last vendoring step of install-plugins.sh AND update-all.sh (the resync dropped the pins until the next make plugin) - design stack high uniform (last loaded wins), superpowers, agent-skills, 21st pack pinned from the map; skills-perso low, pdf-translate medium, site-motion high - doctrine: design stack loads paired with the first Read; one level per stack (CLAUDE.global.md, lib/effort-shift.md) - lib/effort-audit.py prints thinking coverage per scope (sub-agent records carry no thinking count on ~94 % of requests) - census map-driven + fixture suite lib/tests/effort-pins.test.sh; docs README/USAGE/CHANGELOG; contract + TODO plan
86 lines
4.4 KiB
Markdown
86 lines
4.4 KiB
Markdown
# Effort shift — phase-level reasoning effort on the main loop (BDR-107)
|
|
|
|
Shared include, companion of `lib/model-gate.md`: the gate fixes WHICH model
|
|
reflects, this include fixes HOW HARD each phase thinks. The rungs are the
|
|
user's: low (fix a line, run a script) · medium (day-to-day) · high
|
|
(refactor, resisting bug) · xhigh (architecture, audit before validation) ·
|
|
max (stuck error, judged need).
|
|
|
|
## Mechanics (verified on Claude Code 2.1.283)
|
|
|
|
- **Pairing rule**: a `Skill(effort-<level>)` call applies its effort only
|
|
when the same assistant message carries at least one other tool call
|
|
after it; a lone Skill call is a no-op. Send the shift together with the
|
|
step's first tool call, shift first. That paired call already runs at the
|
|
new level: pair a downward shift with a pinned-agent dispatch or a
|
|
Read/Bash, never with a built-in judgment dispatch (`general-purpose`,
|
|
`model: "opus"`), which would inherit it.
|
|
- Re-loading a shifter already loaded in the conversation re-applies its
|
|
effort (the harness only dedupes the skill text), so bounce-back
|
|
sequences such as medium → max → medium work.
|
|
- A skill's `effort:` frontmatter applies from the moment it loads to the
|
|
end of the turn: on the user's `/skill` unconditionally, and on a
|
|
`Skill(...)` call by Claude only under the pairing rule above (a skill
|
|
Claude loads alone, such as `brainstorming` or `writing-plans`, applies
|
|
nothing). Last loaded wins, both directions. The prompt cache survives a
|
|
shift.
|
|
- **Stacked skills share one level**: skills that load together in one
|
|
build (the design stack) all pin the same level, since the last loaded
|
|
wins. Vendored externals get their level from `lib/effort-pins.txt`,
|
|
re-applied by `lib/effort-pins.sh` after every vendoring step; repo
|
|
skills carry it in their frontmatter.
|
|
- Dispatched agents run on their own `effort:` pin, never on a shift.
|
|
Unpinned agents inherit the level in force at dispatch.
|
|
- Headless sessions (`-p`, `claude agents`, SDK) ignore skill-level effort:
|
|
the run stays at the session level. `CLAUDE_CODE_EFFORT_LEVEL` beats every
|
|
frontmatter; keep it unset (the session banner warns).
|
|
|
|
Measure the split any time: `python3 ~/.claude/lib/effort-audit.py`
|
|
(thinking/output/cache tokens per scope, model and effort).
|
|
|
|
## Shifters
|
|
|
|
`Skill(effort-low)` · `Skill(effort-medium)` · `Skill(effort-high)` ·
|
|
`Skill(effort-xhigh)` · `Skill(effort-max)`. One tool call, one-line body,
|
|
always sent with another tool call (Pairing rule).
|
|
Typed by the user, `/effort-max` is a turn-scoped max: the relaunch lever
|
|
after a STOP. `ultrathink` only adds an in-context nudge; the API level
|
|
does not move.
|
|
|
|
## Wiring — per orchestrator
|
|
|
|
1. A dispatch span starts (executor, collector, fan-out) →
|
|
`Skill(effort-medium)`.
|
|
2. Reflection resumes after a dispatch span (challenge synthesis, verdict,
|
|
plan revision) → `Skill(effort-<the skill's own level>)`. Concretely:
|
|
the line before every `lib/challenge-plan.md` call.
|
|
3. The bookkeeping tail (memory commit, doc commit) → `Skill(effort-low)`.
|
|
4. Escalation → `Skill(effort-max)`, then the skill's own level again once
|
|
the diagnosis is produced. Automatic points: verify-secure loop caps
|
|
(GATE 0 floor, GATE 1 conformity, GATE 2 security) and ship-feature
|
|
STEP 4b. Not automatic, by doctrine: the challenge fail-safe (a mute
|
|
challenger is an infrastructure failure) and "gone WRONG → STOP" (STOP
|
|
precedes any further reasoning); their STOP text names the level
|
|
reached and suggests `/effort-max` for the relaunch.
|
|
5. Before any built-in or unpinned dispatch that carries judgment (a
|
|
`general-purpose` with `model: "opus"` or `"fable"`, the code reviewer
|
|
of requesting-code-review, a skill-runner) → `Skill(effort-<own level>)`
|
|
paired with that dispatch: built-ins inherit the level in force, and a
|
|
medium set earlier in the span would downgrade them.
|
|
|
|
## Re-assert
|
|
|
|
- After any nested `Skill(...)` whose frontmatter carries a different
|
|
effort (feat → commit-change), reload the orchestrator's own level.
|
|
- After a prose gate that ends the turn, the resumed turn runs at the
|
|
session level. If the resumed phase is reflection, its first step is
|
|
`Skill(effort-<own level>)`; dispatch and orchestration phases need
|
|
nothing.
|
|
|
|
## Never
|
|
|
|
- A shift never inside a dispatched agent: pins rule there.
|
|
- Max is for diagnosis, not for retrying the same fix harder.
|
|
- A medium shift never precedes a judgment dispatch in the same span
|
|
without an own-level shift paired with that dispatch.
|