fix(effort): re-raise judgment dispatches, planning re-asserts, pairing caveat, dedupe audit script (final review I1-I3)

This commit is contained in:
bastien
2026-09-28 20:48:27 +02:00
parent a3b479e984
commit 58c3a3e9b7
18 changed files with 57 additions and 23 deletions
+8 -1
View File
@@ -26,7 +26,10 @@ def usage_row(usage):
def scan(path, scope, agg):
"""Add every assistant record of one transcript to agg."""
"""Add every assistant record of one transcript to agg, once per
message id (the transcript writes one record per content block,
all sharing the same id and usage)."""
seen = set()
with open(path, errors="ignore") as handle:
for line in handle:
try:
@@ -36,6 +39,10 @@ def scan(path, scope, agg):
msg = rec.get("message") or {}
if rec.get("type") != "assistant" or not msg.get("usage"):
continue
mid = msg.get("id")
if mid in seen:
continue
seen.add(mid)
sub = scope == "sub" or bool(rec.get("isSidechain"))
key = ("sub" if sub else "main",
str(msg.get("model", "?")).replace("claude-", ""),
+12 -3
View File
@@ -19,9 +19,11 @@ max (stuck error, judged need).
effort (the harness only dedupes the skill text), so bounce-back
sequences such as medium → max → medium work.
- A skill's `effort:` frontmatter applies from the moment it loads to the
end of the turn: on the user's `/skill` and on a `Skill(...)` call by
Claude in an interactive session. Last loaded wins, both directions. The
prompt cache survives a shift.
end of the turn: on the user's `/skill` unconditionally, and on a
`Skill(...)` call by Claude only under the pairing rule above (a skill
Claude loads alone, such as `brainstorming` or `writing-plans`, applies
nothing). Last loaded wins, both directions. The prompt cache survives a
shift.
- Dispatched agents run on their own `effort:` pin, never on a shift.
Unpinned agents inherit the level in force at dispatch.
- Headless sessions (`-p`, `claude agents`, SDK) ignore skill-level effort:
@@ -55,6 +57,11 @@ does not move.
challenger is an infrastructure failure) and "gone WRONG → STOP" (STOP
precedes any further reasoning); their STOP text names the level
reached and suggests `/effort-max` for the relaunch.
5. Before any built-in or unpinned dispatch that carries judgment (a
`general-purpose` with `model: "opus"` or `"fable"`, the code reviewer
of requesting-code-review, a skill-runner) → `Skill(effort-<own level>)`
paired with that dispatch: built-ins inherit the level in force, and a
medium set earlier in the span would downgrade them.
## Re-assert
@@ -69,3 +76,5 @@ does not move.
- A shift never inside a dispatched agent: pins rule there.
- Max is for diagnosis, not for retrying the same fix harder.
- A medium shift never precedes a judgment dispatch in the same span
without an own-level shift paired with that dispatch.
+3 -1
View File
@@ -48,4 +48,6 @@ mechanical probes).
Effort is the second axis of the same table (BDR-107): every typed agent
carries an `effort:` pin next to `model:`, and the main loop shifts per phase
through `lib/effort-shift.md`. Nothing dispatched inherits either axis.
through `lib/effort-shift.md`. No typed agent inherits either axis;
built-ins inherit the effort in force at dispatch, so an orchestrator shifts
before dispatching them (`lib/effort-shift.md`, wiring point 5).
+14 -2
View File
@@ -65,8 +65,11 @@ has "lib/model-gate.md" 'lib/effort-shift.md'
# ── 6) orchestrator wiring (spec D4)
for s in feat hotfix bugfix ship-feature init-project onboard tour code-clean seo geo harden web-validate audit-delta; do
has "skills/$s/SKILL.md" 'lib/effort-shift.md'; has "skills/$s/SKILL.md" 'Skill(effort-medium)'; done
has "agents/client-handover-writer.md" 'lib/effort-shift.md'; has "agents/client-handover-writer.md" 'Skill(effort-medium)'
has "skills/$s/SKILL.md" 'lib/effort-shift.md'; has "skills/$s/SKILL.md" 'a lone Skill call is a no-op'; done
for s in feat hotfix bugfix ship-feature init-project code-clean seo geo harden web-validate audit-delta; do
has "skills/$s/SKILL.md" 'Skill(effort-medium)'; done
lacks "skills/onboard/SKILL.md" 'Skill(effort-medium)'; lacks "skills/tour/SKILL.md" 'Skill(effort-medium)'
has "agents/client-handover-writer.md" 'lib/effort-shift.md'; lacks "agents/client-handover-writer.md" 'Skill(effort-medium)'; has "agents/client-handover-writer.md" 'Skill(effort-high)'
for s in feat hotfix bugfix; do has "skills/$s/SKILL.md" 'Skill(effort-high)'; done
for s in ship-feature init-project onboard code-clean audit-delta; do has "skills/$s/SKILL.md" 'Skill(effort-xhigh)'; done
for s in seo geo harden web-validate; do has "skills/$s/SKILL.md" 'Skill(effort-high)'; done
@@ -91,6 +94,15 @@ has "skills/bugfix/SKILL.md" 'effort-shift: turn reset'
has "lib/effort-shift.md" 'effort-audit.py'
[ -x "$R/lib/effort-audit.py" ] && ok || ko "lib/effort-audit.py missing or not executable"
# ── 6c) judgment dispatches re-raised, planning re-asserts, stronger locks (final review I1/I2/M5)
for s in ship-feature init-project; do has "skills/$s/SKILL.md" 'effort-shift: judgment dispatch'; has "skills/$s/SKILL.md" 'effort-shift: turn reset'; done
has "agents/client-handover-writer.md" 'effort-shift: judgment dispatch'
has "lib/effort-shift.md" 'Before any built-in or unpinned dispatch'
has "lib/model-gate.md" 'built-ins inherit the effort in force'
has "skills/ship-feature/SKILL.md" 'effort-shift: error recovery'
for s in feat hotfix bugfix seo geo harden web-validate ship-feature init-project onboard code-clean audit-delta; do has "skills/$s/SKILL.md" 'effort-shift: own level before the challenge'; done
has "install-plugins.sh" 'for _s in brainstorming writing-plans; do'
# ── summary (later tasks insert their locks ABOVE this line)
printf 'effort-routing census: %d pass, %d fail\n' "$pass" "$fail"
[ "$fail" -eq 0 ]