fix(effort): re-raise judgment dispatches, planning re-asserts, pairing caveat, dedupe audit script (final review I1-I3)

This commit is contained in:
bastien
2026-09-28 20:48:27 +02:00
parent a3b479e984
commit 58c3a3e9b7
18 changed files with 57 additions and 23 deletions
+3 -2
View File
@@ -97,6 +97,8 @@ Parse `$ARGUMENTS` for optional flags:
---
EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op.
## STEP 1 — PRE-FLIGHT
```bash
@@ -225,7 +227,7 @@ Store `DEPLOYED_URL` for STEP 7. If empty, ask user during STEP 6.
---
## STEP 3 — BASELINE AUDITS (parallel)
First: `Skill(effort-medium)` (effort-shift: dispatch span starts; send it in the same message as this step's first dispatch).
First: `Skill(effort-high)` (effort-shift: judgment dispatch; the fable skill-runners are built-ins and inherit the level in force; high is the entry level of the audits they run).
Goal: capture `SCORE_*_BEFORE` so the client doc shows the delta.
@@ -262,7 +264,6 @@ the gate.
this pipeline (initial audits, fix-loop re-dispatches, commit-change,
web-validate) carries `model: "fable"` — the child hosts gated orchestration
on the pipeline's behalf; it must never inherit the session model.
EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op.
For web projects, dispatch in **a single message with two parallel Agent calls** (each with `model: "fable"`):
+8 -1
View File
@@ -26,7 +26,10 @@ def usage_row(usage):
def scan(path, scope, agg):
"""Add every assistant record of one transcript to agg."""
"""Add every assistant record of one transcript to agg, once per
message id (the transcript writes one record per content block,
all sharing the same id and usage)."""
seen = set()
with open(path, errors="ignore") as handle:
for line in handle:
try:
@@ -36,6 +39,10 @@ def scan(path, scope, agg):
msg = rec.get("message") or {}
if rec.get("type") != "assistant" or not msg.get("usage"):
continue
mid = msg.get("id")
if mid in seen:
continue
seen.add(mid)
sub = scope == "sub" or bool(rec.get("isSidechain"))
key = ("sub" if sub else "main",
str(msg.get("model", "?")).replace("claude-", ""),
+12 -3
View File
@@ -19,9 +19,11 @@ max (stuck error, judged need).
effort (the harness only dedupes the skill text), so bounce-back
sequences such as medium → max → medium work.
- A skill's `effort:` frontmatter applies from the moment it loads to the
end of the turn: on the user's `/skill` and on a `Skill(...)` call by
Claude in an interactive session. Last loaded wins, both directions. The
prompt cache survives a shift.
end of the turn: on the user's `/skill` unconditionally, and on a
`Skill(...)` call by Claude only under the pairing rule above (a skill
Claude loads alone, such as `brainstorming` or `writing-plans`, applies
nothing). Last loaded wins, both directions. The prompt cache survives a
shift.
- Dispatched agents run on their own `effort:` pin, never on a shift.
Unpinned agents inherit the level in force at dispatch.
- Headless sessions (`-p`, `claude agents`, SDK) ignore skill-level effort:
@@ -55,6 +57,11 @@ does not move.
challenger is an infrastructure failure) and "gone WRONG → STOP" (STOP
precedes any further reasoning); their STOP text names the level
reached and suggests `/effort-max` for the relaunch.
5. Before any built-in or unpinned dispatch that carries judgment (a
`general-purpose` with `model: "opus"` or `"fable"`, the code reviewer
of requesting-code-review, a skill-runner) → `Skill(effort-<own level>)`
paired with that dispatch: built-ins inherit the level in force, and a
medium set earlier in the span would downgrade them.
## Re-assert
@@ -69,3 +76,5 @@ does not move.
- A shift never inside a dispatched agent: pins rule there.
- Max is for diagnosis, not for retrying the same fix harder.
- A medium shift never precedes a judgment dispatch in the same span
without an own-level shift paired with that dispatch.
+3 -1
View File
@@ -48,4 +48,6 @@ mechanical probes).
Effort is the second axis of the same table (BDR-107): every typed agent
carries an `effort:` pin next to `model:`, and the main loop shifts per phase
through `lib/effort-shift.md`. Nothing dispatched inherits either axis.
through `lib/effort-shift.md`. No typed agent inherits either axis;
built-ins inherit the effort in force at dispatch, so an orchestrator shifts
before dispatching them (`lib/effort-shift.md`, wiring point 5).
+14 -2
View File
@@ -65,8 +65,11 @@ has "lib/model-gate.md" 'lib/effort-shift.md'
# ── 6) orchestrator wiring (spec D4)
for s in feat hotfix bugfix ship-feature init-project onboard tour code-clean seo geo harden web-validate audit-delta; do
has "skills/$s/SKILL.md" 'lib/effort-shift.md'; has "skills/$s/SKILL.md" 'Skill(effort-medium)'; done
has "agents/client-handover-writer.md" 'lib/effort-shift.md'; has "agents/client-handover-writer.md" 'Skill(effort-medium)'
has "skills/$s/SKILL.md" 'lib/effort-shift.md'; has "skills/$s/SKILL.md" 'a lone Skill call is a no-op'; done
for s in feat hotfix bugfix ship-feature init-project code-clean seo geo harden web-validate audit-delta; do
has "skills/$s/SKILL.md" 'Skill(effort-medium)'; done
lacks "skills/onboard/SKILL.md" 'Skill(effort-medium)'; lacks "skills/tour/SKILL.md" 'Skill(effort-medium)'
has "agents/client-handover-writer.md" 'lib/effort-shift.md'; lacks "agents/client-handover-writer.md" 'Skill(effort-medium)'; has "agents/client-handover-writer.md" 'Skill(effort-high)'
for s in feat hotfix bugfix; do has "skills/$s/SKILL.md" 'Skill(effort-high)'; done
for s in ship-feature init-project onboard code-clean audit-delta; do has "skills/$s/SKILL.md" 'Skill(effort-xhigh)'; done
for s in seo geo harden web-validate; do has "skills/$s/SKILL.md" 'Skill(effort-high)'; done
@@ -91,6 +94,15 @@ has "skills/bugfix/SKILL.md" 'effort-shift: turn reset'
has "lib/effort-shift.md" 'effort-audit.py'
[ -x "$R/lib/effort-audit.py" ] && ok || ko "lib/effort-audit.py missing or not executable"
# ── 6c) judgment dispatches re-raised, planning re-asserts, stronger locks (final review I1/I2/M5)
for s in ship-feature init-project; do has "skills/$s/SKILL.md" 'effort-shift: judgment dispatch'; has "skills/$s/SKILL.md" 'effort-shift: turn reset'; done
has "agents/client-handover-writer.md" 'effort-shift: judgment dispatch'
has "lib/effort-shift.md" 'Before any built-in or unpinned dispatch'
has "lib/model-gate.md" 'built-ins inherit the effort in force'
has "skills/ship-feature/SKILL.md" 'effort-shift: error recovery'
for s in feat hotfix bugfix seo geo harden web-validate ship-feature init-project onboard code-clean audit-delta; do has "skills/$s/SKILL.md" 'effort-shift: own level before the challenge'; done
has "install-plugins.sh" 'for _s in brainstorming writing-plans; do'
# ── summary (later tasks insert their locks ABOVE this line)
printf 'effort-routing census: %d pass, %d fail\n' "$pass" "$fail"
[ "$fail" -eq 0 ]
+1 -1
View File
@@ -170,7 +170,7 @@ Then show the user the same compact table inline.
### 3b-bis. CHALLENGE THE PROPOSALS (before the gate)
`Skill(effort-xhigh)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch).
`Skill(effort-xhigh)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch).
This axis' findings + proposed fixes are a proposal set worth attacking before
the human gate. Persist THIS axis' finding list (not the whole append-only
report) to `.claude/tasks/plans/<date>-<axis>-<HHMM>.md`, then run
+1 -1
View File
@@ -126,7 +126,7 @@ RISK: <low/medium — what could go wrong>
fast-path is not exempt: a 1-line fix with a visible choice still asks.
## STEP 3b — CHALLENGE THE FIX PLAN (before the contract)
`Skill(effort-high)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch).
`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch).
Unless the fix is the trivial 1-2 line case STEP 3 already fast-paths, the
DIAGNOSIS + FIX PLAN is a reflection worth attacking before it hardens into a
contract. Persist it to `.claude/tasks/plans/<date>-<slug>-<HHMM>.md`, then run
+1 -1
View File
@@ -122,7 +122,7 @@ TOTALS: <N blocking, N warn, N info>
If no issues found: report clean state and stop.
## STEP 3b — CHALLENGE THE SCOPE (before approval)
`Skill(effort-xhigh)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch).
`Skill(effort-xhigh)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch).
The STEP 3 report is the proposed cleanup scope — worth attacking before the
human approves it. It is still inline, so FIRST persist it to
`.claude/tasks/plans/<date>-<slug>-<HHMM>.md` (STEP 3 report format, one item
+1 -1
View File
@@ -124,7 +124,7 @@ in the contract's CLARIFICATIONS `[gated]` and in the plan. A choice that
surfaces only during execution comes back as `NEED-DECISION` (STEP 3).
## STEP 1b — CHALLENGE THE PLAN (before branching)
`Skill(effort-high)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch).
`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch).
The STEP 1 plan is a reflection worth attacking before a branch is spent on it.
Persist it to `.claude/tasks/plans/<date>-<slug>-<HHMM>.md`, then run
`$HOME/.claude/lib/challenge-plan.md` with `PLAN` = that file, `KIND` = `build-plan`,
+1 -1
View File
@@ -86,7 +86,7 @@ your bundle."
```
## STEP 1b — CHALLENGE THE FIX BUNDLE (advisory, before apply)
`Skill(effort-high)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch).
`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch).
The analyzer returned a `## FIX BUNDLE` — worth attacking before any edit lands.
**Skip if intervention mode = conservative** (nothing is applied). Else persist the
bundle verbatim to `.claude/tasks/plans/<date>-<slug>-<HHMM>.md`, then run
+1 -1
View File
@@ -522,7 +522,7 @@ Extract the score and critical-alert count from `.claude/audits/HARDEN.md` for t
---
## STEP 2b — CHALLENGE THE FIX BUNDLE (MODE=fix only, advisory)
`Skill(effort-high)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch).
`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch).
Skip if MODE=audit (no bundle exists). Else, before the STEP 3 gate, harden the bundle:
extract the `## 8. Fix bundle` section from HARDEN.md to
`.claude/tasks/plans/<date>-<slug>-<HHMM>.md` (a clean, blind-judgeable artifact), then run
+1 -1
View File
@@ -94,7 +94,7 @@ point. Run it ONLY when the settled fix touches control flow or behaviour — an
off-by-one, a wrong operator/variable, a behaviour-changing config value, or a
missing import that alters execution. In doubt → it is probably a `/bugfix`.
`Skill(effort-high)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch).
`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch).
For a logic fix: persist the STEP 1 located fix (root cause + the exact edit) to
`.claude/tasks/plans/<date>-<slug>-<HHMM>.md`, then run
`$HOME/.claude/lib/challenge-plan.md` with `PLAN` = that file, `KIND` =
+3 -1
View File
@@ -181,11 +181,12 @@ This is the deterministic scaffold commit owner (closes BLK-010). The MVP is
implemented on a `feature/*` branch off `develop` (STEP 8).
## STEP 6 — PLAN
`Skill(effort-xhigh)` first, sent with the next tool call (effort-shift: turn reset; gate #1 ended the turn and the vendored `writing-plans` pin applies only when the user invokes it).
Invoke `writing-plans` (vendored superpowers skill) with BRIEF + skeleton.
Granular tasks (2-5 min each), exact file paths, TDD: tests before code.
## STEP 6b — CHALLENGE THE PLAN (before the gate)
`Skill(effort-xhigh)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch).
`Skill(effort-xhigh)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch).
Before the human sees the implementation plan, harden it. Run
`$HOME/.claude/lib/challenge-plan.md` with `PLAN` = the plan STEP 6 wrote under
`docs/superpowers/plans/`, `KIND` = `build-plan`, `SCOPE` = the skeleton + task file
@@ -259,6 +260,7 @@ deferred to a later /onboard) and turns the informal analyze into a verdict
against the founding contract. Distinct axis from STEP 10 code review
([[LRN-095]]) — both run.
`Skill(effort-xhigh)` first, sent with the review dispatch (effort-shift: judgment dispatch; the reviewer is a built-in and inherits the level in force).
## STEP 10 — CODE REVIEW
Invoke `requesting-code-review` (vendored superpowers skill). **Model routing (BDR-077):** the
review subagent it dispatches MUST carry `model: "opus"` in the Agent call —
+1 -2
View File
@@ -92,7 +92,6 @@ STOP. La réponse détermine si STEP 1 tourne une fois (A) ou N fois (C) ou avec
## STEP 2 — BASELINE CONFIG (onboarder agent)
`Skill(effort-medium)` first (effort-shift: dispatch span starts; send it in the same message as this step's first dispatch).
Dispatch `Agent(subagent_type="onboarder")` (pin sonnet — BDR-077 : config
templating = exécution, plus jamais inline sur le modèle de session). Un
BLOCAGE (clé manquante, CLAUDE.md existant) revient en rapport — l'agent ne
@@ -893,7 +892,7 @@ Vérifier que les 4 fichiers `.claude/audits/ONBOARD_REPORT.md`, `.claude/audits
---
## STEP 7b — CHALLENGE THE PROPOSALS (before the human gate)
`Skill(effort-xhigh)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch).
`Skill(effort-xhigh)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch).
The 4 audit files are on disk; `AUDIT_PROPOSALS.md` is the artifact worth
attacking before the human spends a gate on it. Run
`$HOME/.claude/lib/challenge-plan.md` with `PLAN` =
+1 -1
View File
@@ -510,7 +510,7 @@ the reports."
```
## STEP 1b — CHALLENGE THE FIX BUNDLE (advisory, before apply)
`Skill(effort-high)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch).
`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch).
Both envelopes now carry a `## FIX BUNDLE` — worth attacking before any edit lands.
**Skip if intervention mode = conservative** (nothing is applied). Else persist both
bundles (seo + geo, verbatim) to `.claude/tasks/plans/<date>-<slug>-<HHMM>.md`, then run
+4 -1
View File
@@ -114,8 +114,10 @@ Inject ONLY what constrains: the NON-BINDING count does NOT enter the brainstorm
(the injection inherits the OUTPUT filter — detail what binds, drop what doesn't).
Consumption = INPUT INJECTION (we can't modify the external skill; we control its input).
Refine request into validated design via Socratic questioning. Don't proceed until design approved.
Turns after a user reply run at the session level until a tool call is paired with `Skill(effort-xhigh)` (effort-shift: turn reset).
## STEP 2 — PLAN
`Skill(effort-xhigh)` first, sent with the next tool call (effort-shift: turn reset; brainstorm turns after a user reply run at the session level, and the vendored `brainstorming` pin applies only when the user invokes it).
Invoke `writing-plans` (vendored superpowers skill) with the validated design AND the 0d digest: every task
must be consistent with the in-force constraints; where a task implements or affects one,
note the ID inline. Break design into tasks (2-5 min each). Each task: exact file paths, full code, verification steps.
@@ -125,7 +127,7 @@ request nor the STEP 1 brainstorm settled (check the contract's CLARIFICATIONS
first) → one batch before STEP 2b; answers append to the contract `[gated]`.
## STEP 2b — CHALLENGE THE PLAN (adversarial, before the gate)
`Skill(effort-xhigh)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch).
`Skill(effort-xhigh)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch).
Before the human sees the plan, harden it. Run `$HOME/.claude/lib/challenge-plan.md`:
- `PLAN` = the plan STEP 2 wrote under `docs/superpowers/plans/`
- `KIND` = `build-plan`
@@ -240,6 +242,7 @@ against the contract. It is a DISTINCT axis from STEP 6 code review (contract
conformity + security vs. craft/design) — both run, neither subsumes the
other ([[LRN-095]]).
`Skill(effort-xhigh)` first, sent with the review dispatch (effort-shift: judgment dispatch; the reviewer is a built-in and inherits the level in force).
## STEP 6 — CODE REVIEW
Invoke `requesting-code-review` (vendored superpowers skill). **Model routing (BDR-077):** the
review subagent it dispatches MUST carry `model: "opus"` in the Agent call —
-1
View File
@@ -87,7 +87,6 @@ Model discipline (the user-fixed invariant behind this mode):
Runner dispatch, one per project:
```
Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message
Agent(subagent_type="general-purpose",
description="tour runner — <project basename>",
prompt="Read ~/.claude/skills/tour/SKILL.md and execute STEP 1 → STEP 3
+1 -1
View File
@@ -255,7 +255,7 @@ grep -c '^### \[Critique\]' .claude/audits/VALIDATE.md
---
## STEP 2b — CHALLENGE THE FIX BUNDLE (MODE=fix only, advisory)
`Skill(effort-high)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch).
`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch).
Skip if MODE=audit (no bundle exists). Else, before the STEP 3 gate, harden the bundle:
extract the `## 5. Fix bundle` section from VALIDATE.md to
`.claude/tasks/plans/<date>-<slug>-<HHMM>.md` (a clean, blind-judgeable artifact), then run