From 58c3a3e9b7b6d21fe26e2ad4ff08c6e5dc38d692 Mon Sep 17 00:00:00 2001 From: bastien Date: Mon, 28 Sep 2026 20:48:27 +0200 Subject: [PATCH] fix(effort): re-raise judgment dispatches, planning re-asserts, pairing caveat, dedupe audit script (final review I1-I3) --- agents/client-handover-writer.md | 5 +++-- lib/effort-audit.py | 9 ++++++++- lib/effort-shift.md | 15 ++++++++++++--- lib/model-gate.md | 4 +++- lib/tests/effort-routing.test.sh | 16 ++++++++++++++-- skills/audit-delta/SKILL.md | 2 +- skills/bugfix/SKILL.md | 2 +- skills/code-clean/SKILL.md | 2 +- skills/feat/SKILL.md | 2 +- skills/geo/SKILL.md | 2 +- skills/harden/SKILL.md | 2 +- skills/hotfix/SKILL.md | 2 +- skills/init-project/SKILL.md | 4 +++- skills/onboard/SKILL.md | 3 +-- skills/seo/SKILL.md | 2 +- skills/ship-feature/SKILL.md | 5 ++++- skills/tour/SKILL.md | 1 - skills/web-validate/SKILL.md | 2 +- 18 files changed, 57 insertions(+), 23 deletions(-) diff --git a/agents/client-handover-writer.md b/agents/client-handover-writer.md index 13d32e9..1ba9c0c 100644 --- a/agents/client-handover-writer.md +++ b/agents/client-handover-writer.md @@ -97,6 +97,8 @@ Parse `$ARGUMENTS` for optional flags: --- +EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. + ## STEP 1 — PRE-FLIGHT ```bash @@ -225,7 +227,7 @@ Store `DEPLOYED_URL` for STEP 7. If empty, ask user during STEP 6. --- ## STEP 3 — BASELINE AUDITS (parallel) -First: `Skill(effort-medium)` (effort-shift: dispatch span starts; send it in the same message as this step's first dispatch). +First: `Skill(effort-high)` (effort-shift: judgment dispatch; the fable skill-runners are built-ins and inherit the level in force; high is the entry level of the audits they run). Goal: capture `SCORE_*_BEFORE` so the client doc shows the delta. @@ -262,7 +264,6 @@ the gate. this pipeline (initial audits, fix-loop re-dispatches, commit-change, web-validate) carries `model: "fable"` — the child hosts gated orchestration on the pipeline's behalf; it must never inherit the session model. -EFFORT SHIFTS: follow `$HOME/.claude/lib/effort-shift.md` (BDR-107): medium when a dispatch span starts, own level before challenge synthesis, low at the bookkeeping tail, max at escalation; every shift goes in the same message as the step's first tool call, a lone Skill call is a no-op. For web projects, dispatch in **a single message with two parallel Agent calls** (each with `model: "fable"`): diff --git a/lib/effort-audit.py b/lib/effort-audit.py index bc39494..2d84556 100755 --- a/lib/effort-audit.py +++ b/lib/effort-audit.py @@ -26,7 +26,10 @@ def usage_row(usage): def scan(path, scope, agg): - """Add every assistant record of one transcript to agg.""" + """Add every assistant record of one transcript to agg, once per + message id (the transcript writes one record per content block, + all sharing the same id and usage).""" + seen = set() with open(path, errors="ignore") as handle: for line in handle: try: @@ -36,6 +39,10 @@ def scan(path, scope, agg): msg = rec.get("message") or {} if rec.get("type") != "assistant" or not msg.get("usage"): continue + mid = msg.get("id") + if mid in seen: + continue + seen.add(mid) sub = scope == "sub" or bool(rec.get("isSidechain")) key = ("sub" if sub else "main", str(msg.get("model", "?")).replace("claude-", ""), diff --git a/lib/effort-shift.md b/lib/effort-shift.md index 502716b..38b7ad7 100644 --- a/lib/effort-shift.md +++ b/lib/effort-shift.md @@ -19,9 +19,11 @@ max (stuck error, judged need). effort (the harness only dedupes the skill text), so bounce-back sequences such as medium → max → medium work. - A skill's `effort:` frontmatter applies from the moment it loads to the - end of the turn: on the user's `/skill` and on a `Skill(...)` call by - Claude in an interactive session. Last loaded wins, both directions. The - prompt cache survives a shift. + end of the turn: on the user's `/skill` unconditionally, and on a + `Skill(...)` call by Claude only under the pairing rule above (a skill + Claude loads alone, such as `brainstorming` or `writing-plans`, applies + nothing). Last loaded wins, both directions. The prompt cache survives a + shift. - Dispatched agents run on their own `effort:` pin, never on a shift. Unpinned agents inherit the level in force at dispatch. - Headless sessions (`-p`, `claude agents`, SDK) ignore skill-level effort: @@ -55,6 +57,11 @@ does not move. challenger is an infrastructure failure) and "gone WRONG → STOP" (STOP precedes any further reasoning); their STOP text names the level reached and suggests `/effort-max` for the relaunch. +5. Before any built-in or unpinned dispatch that carries judgment (a + `general-purpose` with `model: "opus"` or `"fable"`, the code reviewer + of requesting-code-review, a skill-runner) → `Skill(effort-)` + paired with that dispatch: built-ins inherit the level in force, and a + medium set earlier in the span would downgrade them. ## Re-assert @@ -69,3 +76,5 @@ does not move. - A shift never inside a dispatched agent: pins rule there. - Max is for diagnosis, not for retrying the same fix harder. +- A medium shift never precedes a judgment dispatch in the same span + without an own-level shift paired with that dispatch. diff --git a/lib/model-gate.md b/lib/model-gate.md index 22c897f..874cd23 100644 --- a/lib/model-gate.md +++ b/lib/model-gate.md @@ -48,4 +48,6 @@ mechanical probes). Effort is the second axis of the same table (BDR-107): every typed agent carries an `effort:` pin next to `model:`, and the main loop shifts per phase -through `lib/effort-shift.md`. Nothing dispatched inherits either axis. +through `lib/effort-shift.md`. No typed agent inherits either axis; +built-ins inherit the effort in force at dispatch, so an orchestrator shifts +before dispatching them (`lib/effort-shift.md`, wiring point 5). diff --git a/lib/tests/effort-routing.test.sh b/lib/tests/effort-routing.test.sh index 5921c31..a6e4003 100755 --- a/lib/tests/effort-routing.test.sh +++ b/lib/tests/effort-routing.test.sh @@ -65,8 +65,11 @@ has "lib/model-gate.md" 'lib/effort-shift.md' # ── 6) orchestrator wiring (spec D4) for s in feat hotfix bugfix ship-feature init-project onboard tour code-clean seo geo harden web-validate audit-delta; do - has "skills/$s/SKILL.md" 'lib/effort-shift.md'; has "skills/$s/SKILL.md" 'Skill(effort-medium)'; done -has "agents/client-handover-writer.md" 'lib/effort-shift.md'; has "agents/client-handover-writer.md" 'Skill(effort-medium)' + has "skills/$s/SKILL.md" 'lib/effort-shift.md'; has "skills/$s/SKILL.md" 'a lone Skill call is a no-op'; done +for s in feat hotfix bugfix ship-feature init-project code-clean seo geo harden web-validate audit-delta; do + has "skills/$s/SKILL.md" 'Skill(effort-medium)'; done +lacks "skills/onboard/SKILL.md" 'Skill(effort-medium)'; lacks "skills/tour/SKILL.md" 'Skill(effort-medium)' +has "agents/client-handover-writer.md" 'lib/effort-shift.md'; lacks "agents/client-handover-writer.md" 'Skill(effort-medium)'; has "agents/client-handover-writer.md" 'Skill(effort-high)' for s in feat hotfix bugfix; do has "skills/$s/SKILL.md" 'Skill(effort-high)'; done for s in ship-feature init-project onboard code-clean audit-delta; do has "skills/$s/SKILL.md" 'Skill(effort-xhigh)'; done for s in seo geo harden web-validate; do has "skills/$s/SKILL.md" 'Skill(effort-high)'; done @@ -91,6 +94,15 @@ has "skills/bugfix/SKILL.md" 'effort-shift: turn reset' has "lib/effort-shift.md" 'effort-audit.py' [ -x "$R/lib/effort-audit.py" ] && ok || ko "lib/effort-audit.py missing or not executable" +# ── 6c) judgment dispatches re-raised, planning re-asserts, stronger locks (final review I1/I2/M5) +for s in ship-feature init-project; do has "skills/$s/SKILL.md" 'effort-shift: judgment dispatch'; has "skills/$s/SKILL.md" 'effort-shift: turn reset'; done +has "agents/client-handover-writer.md" 'effort-shift: judgment dispatch' +has "lib/effort-shift.md" 'Before any built-in or unpinned dispatch' +has "lib/model-gate.md" 'built-ins inherit the effort in force' +has "skills/ship-feature/SKILL.md" 'effort-shift: error recovery' +for s in feat hotfix bugfix seo geo harden web-validate ship-feature init-project onboard code-clean audit-delta; do has "skills/$s/SKILL.md" 'effort-shift: own level before the challenge'; done +has "install-plugins.sh" 'for _s in brainstorming writing-plans; do' + # ── summary (later tasks insert their locks ABOVE this line) printf 'effort-routing census: %d pass, %d fail\n' "$pass" "$fail" [ "$fail" -eq 0 ] diff --git a/skills/audit-delta/SKILL.md b/skills/audit-delta/SKILL.md index 7b838ae..4a13098 100644 --- a/skills/audit-delta/SKILL.md +++ b/skills/audit-delta/SKILL.md @@ -170,7 +170,7 @@ Then show the user the same compact table inline. ### 3b-bis. CHALLENGE THE PROPOSALS (before the gate) -`Skill(effort-xhigh)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). +`Skill(effort-xhigh)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). This axis' findings + proposed fixes are a proposal set worth attacking before the human gate. Persist THIS axis' finding list (not the whole append-only report) to `.claude/tasks/plans/--.md`, then run diff --git a/skills/bugfix/SKILL.md b/skills/bugfix/SKILL.md index 96869cb..84051fd 100644 --- a/skills/bugfix/SKILL.md +++ b/skills/bugfix/SKILL.md @@ -126,7 +126,7 @@ RISK: fast-path is not exempt: a 1-line fix with a visible choice still asks. ## STEP 3b — CHALLENGE THE FIX PLAN (before the contract) -`Skill(effort-high)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). +`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). Unless the fix is the trivial 1-2 line case STEP 3 already fast-paths, the DIAGNOSIS + FIX PLAN is a reflection worth attacking before it hardens into a contract. Persist it to `.claude/tasks/plans/--.md`, then run diff --git a/skills/code-clean/SKILL.md b/skills/code-clean/SKILL.md index 6d9a7b9..72458cb 100644 --- a/skills/code-clean/SKILL.md +++ b/skills/code-clean/SKILL.md @@ -122,7 +122,7 @@ TOTALS: If no issues found: report clean state and stop. ## STEP 3b — CHALLENGE THE SCOPE (before approval) -`Skill(effort-xhigh)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). +`Skill(effort-xhigh)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). The STEP 3 report is the proposed cleanup scope — worth attacking before the human approves it. It is still inline, so FIRST persist it to `.claude/tasks/plans/--.md` (STEP 3 report format, one item diff --git a/skills/feat/SKILL.md b/skills/feat/SKILL.md index 4ca234d..4b10386 100644 --- a/skills/feat/SKILL.md +++ b/skills/feat/SKILL.md @@ -124,7 +124,7 @@ in the contract's CLARIFICATIONS `[gated]` and in the plan. A choice that surfaces only during execution comes back as `NEED-DECISION` (STEP 3). ## STEP 1b — CHALLENGE THE PLAN (before branching) -`Skill(effort-high)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). +`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). The STEP 1 plan is a reflection worth attacking before a branch is spent on it. Persist it to `.claude/tasks/plans/--.md`, then run `$HOME/.claude/lib/challenge-plan.md` with `PLAN` = that file, `KIND` = `build-plan`, diff --git a/skills/geo/SKILL.md b/skills/geo/SKILL.md index 17b3206..ee76ad9 100644 --- a/skills/geo/SKILL.md +++ b/skills/geo/SKILL.md @@ -86,7 +86,7 @@ your bundle." ``` ## STEP 1b — CHALLENGE THE FIX BUNDLE (advisory, before apply) -`Skill(effort-high)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). +`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). The analyzer returned a `## FIX BUNDLE` — worth attacking before any edit lands. **Skip if intervention mode = conservative** (nothing is applied). Else persist the bundle verbatim to `.claude/tasks/plans/--.md`, then run diff --git a/skills/harden/SKILL.md b/skills/harden/SKILL.md index 7aad9e6..d6d53e3 100644 --- a/skills/harden/SKILL.md +++ b/skills/harden/SKILL.md @@ -522,7 +522,7 @@ Extract the score and critical-alert count from `.claude/audits/HARDEN.md` for t --- ## STEP 2b — CHALLENGE THE FIX BUNDLE (MODE=fix only, advisory) -`Skill(effort-high)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). +`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). Skip if MODE=audit (no bundle exists). Else, before the STEP 3 gate, harden the bundle: extract the `## 8. Fix bundle` section from HARDEN.md to `.claude/tasks/plans/--.md` (a clean, blind-judgeable artifact), then run diff --git a/skills/hotfix/SKILL.md b/skills/hotfix/SKILL.md index 4ae566d..61530b1 100644 --- a/skills/hotfix/SKILL.md +++ b/skills/hotfix/SKILL.md @@ -94,7 +94,7 @@ point. Run it ONLY when the settled fix touches control flow or behaviour — an off-by-one, a wrong operator/variable, a behaviour-changing config value, or a missing import that alters execution. In doubt → it is probably a `/bugfix`. -`Skill(effort-high)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). +`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). For a logic fix: persist the STEP 1 located fix (root cause + the exact edit) to `.claude/tasks/plans/--.md`, then run `$HOME/.claude/lib/challenge-plan.md` with `PLAN` = that file, `KIND` = diff --git a/skills/init-project/SKILL.md b/skills/init-project/SKILL.md index cf28ef7..9ee91d6 100644 --- a/skills/init-project/SKILL.md +++ b/skills/init-project/SKILL.md @@ -181,11 +181,12 @@ This is the deterministic scaffold commit owner (closes BLK-010). The MVP is implemented on a `feature/*` branch off `develop` (STEP 8). ## STEP 6 — PLAN +`Skill(effort-xhigh)` first, sent with the next tool call (effort-shift: turn reset; gate #1 ended the turn and the vendored `writing-plans` pin applies only when the user invokes it). Invoke `writing-plans` (vendored superpowers skill) with BRIEF + skeleton. Granular tasks (2-5 min each), exact file paths, TDD: tests before code. ## STEP 6b — CHALLENGE THE PLAN (before the gate) -`Skill(effort-xhigh)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). +`Skill(effort-xhigh)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). Before the human sees the implementation plan, harden it. Run `$HOME/.claude/lib/challenge-plan.md` with `PLAN` = the plan STEP 6 wrote under `docs/superpowers/plans/`, `KIND` = `build-plan`, `SCOPE` = the skeleton + task file @@ -259,6 +260,7 @@ deferred to a later /onboard) and turns the informal analyze into a verdict against the founding contract. Distinct axis from STEP 10 code review ([[LRN-095]]) — both run. +`Skill(effort-xhigh)` first, sent with the review dispatch (effort-shift: judgment dispatch; the reviewer is a built-in and inherits the level in force). ## STEP 10 — CODE REVIEW Invoke `requesting-code-review` (vendored superpowers skill). **Model routing (BDR-077):** the review subagent it dispatches MUST carry `model: "opus"` in the Agent call — diff --git a/skills/onboard/SKILL.md b/skills/onboard/SKILL.md index c8d7e79..bf7bc43 100644 --- a/skills/onboard/SKILL.md +++ b/skills/onboard/SKILL.md @@ -92,7 +92,6 @@ STOP. La réponse détermine si STEP 1 tourne une fois (A) ou N fois (C) ou avec ## STEP 2 — BASELINE CONFIG (onboarder agent) -`Skill(effort-medium)` first (effort-shift: dispatch span starts; send it in the same message as this step's first dispatch). Dispatch `Agent(subagent_type="onboarder")` (pin sonnet — BDR-077 : config templating = exécution, plus jamais inline sur le modèle de session). Un BLOCAGE (clé manquante, CLAUDE.md existant) revient en rapport — l'agent ne @@ -893,7 +892,7 @@ Vérifier que les 4 fichiers `.claude/audits/ONBOARD_REPORT.md`, `.claude/audits --- ## STEP 7b — CHALLENGE THE PROPOSALS (before the human gate) -`Skill(effort-xhigh)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). +`Skill(effort-xhigh)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). The 4 audit files are on disk; `AUDIT_PROPOSALS.md` is the artifact worth attacking before the human spends a gate on it. Run `$HOME/.claude/lib/challenge-plan.md` with `PLAN` = diff --git a/skills/seo/SKILL.md b/skills/seo/SKILL.md index 31cd20a..ea3a82e 100644 --- a/skills/seo/SKILL.md +++ b/skills/seo/SKILL.md @@ -510,7 +510,7 @@ the reports." ``` ## STEP 1b — CHALLENGE THE FIX BUNDLE (advisory, before apply) -`Skill(effort-high)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). +`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). Both envelopes now carry a `## FIX BUNDLE` — worth attacking before any edit lands. **Skip if intervention mode = conservative** (nothing is applied). Else persist both bundles (seo + geo, verbatim) to `.claude/tasks/plans/--.md`, then run diff --git a/skills/ship-feature/SKILL.md b/skills/ship-feature/SKILL.md index 60d5149..941f02c 100644 --- a/skills/ship-feature/SKILL.md +++ b/skills/ship-feature/SKILL.md @@ -114,8 +114,10 @@ Inject ONLY what constrains: the NON-BINDING count does NOT enter the brainstorm (the injection inherits the OUTPUT filter — detail what binds, drop what doesn't). Consumption = INPUT INJECTION (we can't modify the external skill; we control its input). Refine request into validated design via Socratic questioning. Don't proceed until design approved. +Turns after a user reply run at the session level until a tool call is paired with `Skill(effort-xhigh)` (effort-shift: turn reset). ## STEP 2 — PLAN +`Skill(effort-xhigh)` first, sent with the next tool call (effort-shift: turn reset; brainstorm turns after a user reply run at the session level, and the vendored `brainstorming` pin applies only when the user invokes it). Invoke `writing-plans` (vendored superpowers skill) with the validated design AND the 0d digest: every task must be consistent with the in-force constraints; where a task implements or affects one, note the ID inline. Break design into tasks (2-5 min each). Each task: exact file paths, full code, verification steps. @@ -125,7 +127,7 @@ request nor the STEP 1 brainstorm settled (check the contract's CLARIFICATIONS first) → one batch before STEP 2b; answers append to the contract `[gated]`. ## STEP 2b — CHALLENGE THE PLAN (adversarial, before the gate) -`Skill(effort-xhigh)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). +`Skill(effort-xhigh)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). Before the human sees the plan, harden it. Run `$HOME/.claude/lib/challenge-plan.md`: - `PLAN` = the plan STEP 2 wrote under `docs/superpowers/plans/` - `KIND` = `build-plan` @@ -240,6 +242,7 @@ against the contract. It is a DISTINCT axis from STEP 6 code review (contract conformity + security vs. craft/design) — both run, neither subsumes the other ([[LRN-095]]). +`Skill(effort-xhigh)` first, sent with the review dispatch (effort-shift: judgment dispatch; the reviewer is a built-in and inherits the level in force). ## STEP 6 — CODE REVIEW Invoke `requesting-code-review` (vendored superpowers skill). **Model routing (BDR-077):** the review subagent it dispatches MUST carry `model: "opus"` in the Agent call — diff --git a/skills/tour/SKILL.md b/skills/tour/SKILL.md index 00a04ec..3f7ca46 100644 --- a/skills/tour/SKILL.md +++ b/skills/tour/SKILL.md @@ -87,7 +87,6 @@ Model discipline (the user-fixed invariant behind this mode): Runner dispatch, one per project: ``` -Skill(effort-medium) # effort-shift: dispatch span starts; send with the Agent call below in ONE message Agent(subagent_type="general-purpose", description="tour runner — ", prompt="Read ~/.claude/skills/tour/SKILL.md and execute STEP 1 → STEP 3 diff --git a/skills/web-validate/SKILL.md b/skills/web-validate/SKILL.md index e6f93ee..0901830 100644 --- a/skills/web-validate/SKILL.md +++ b/skills/web-validate/SKILL.md @@ -255,7 +255,7 @@ grep -c '^### \[Critique\]' .claude/audits/VALIDATE.md --- ## STEP 2b — CHALLENGE THE FIX BUNDLE (MODE=fix only, advisory) -`Skill(effort-high)` first (effort-shift: reflection resumes; send it in the same message as the challenger dispatch). +`Skill(effort-high)` first (effort-shift: own level before the challenge; send it in the same message as the challenger dispatch). Skip if MODE=audit (no bundle exists). Else, before the STEP 3 gate, harden the bundle: extract the `## 5. Fix bundle` section from VALIDATE.md to `.claude/tasks/plans/--.md` (a clean, blind-judgeable artifact), then run