Files
claude/.claude/tasks/contracts/2026-09-29-effort-round-1315.md
T
bastien cd3a745857 fix(effort): resync re-applies the pins after the 21st pack refresh, order locked
The 21st pack refresh (update-all 7.4) rewrites every 21st-* SKILL.md after
the superpowers refresh; the re-apply now sits after it, the census locks
the order in both scripts. Contract: criterion 3 anchor, shellcheck
directive authorized, tracked design-motion-principles copy gated.
2026-09-29 15:09:45 +02:00

7.7 KiB
Raw Blame History

CONTRACT — effort-round

  • date: 2026-09-29 | flow: feat by hand (feature/* off develop) | branch: feature/effort-round
  • status: active

REQUEST (verbatim — IMMUTABLE)

en se basant sur le meme tableau que la derniere fois [low: corriger une ligne, renommer un fichier, lancer un script · medium: le travail courant · high: un refactor, un bug qui resiste · xhigh: architecture, audit avant validation · max: quand une erreur coince, une erreur ne se rattrape pas, ou qu'on juge avoir besoin de beaucoup de reflexion], quand on a pin les orchestrateurs et leur sous agent a des efforts, j'aimerais que tu fasse une ronde de tout les skill et que tu mete un niveau d'effort en plus du model pin. D'ailleurs les model pin, c'est du par exemple Sonnet ou du Sonnet 5.5 (version du model pinned) ? Car il faudrait utiliser les versions qui vont bien avec la tache qu'ils ont a acomplir. [answered: model pins stay tier aliases; the latest version of a tier is also the cheapest or same-priced, the quality/price trade-off is tier × effort]

CLARIFICATIONS

  • User choices 2026-09-29 (AskUserQuestion): design stack high uniform; hotfix stays high; every vendored external of the proposed table gets a pin; README/USAGE docs in the same branch.
  • Round result: 30 existing entry levels hold against the table; 3 repo skills had none (skills-perso low, pdf-translate medium, site-motion high); vendored externals get theirs from lib/effort-pins.txt re-applied by lib/effort-pins.sh; impeccable, graphify, find-docs, gstack, darwin-skill and the five shifters stay unpinned (machine-owned, BDR-107).
  • Defect found in passing, fixed here: update-all.sh re-fetched the vendored skills but never re-applied the pins (lost until the next make plugin).
  • Defect found in passing, surfaced not fixed: ~94 % of sub-agent usage records carry no output_tokens_details, so lib/effort-audit.py read zero thinking on sub-agents; the script now prints coverage and a CAVEAT; EVAL-037's "executors stay cheap" is a measurement gap (registry correction pending user approval).
  • lib/tests/effort-routing.test.sh line 4 widens its shellcheck directive from SC2015 to SC2015,SC2016: the new has … '$REPO' locks are literal source text, the $REPO must NOT expand (authorized; a test file, informational). [verifier 2026-09-29 gap 3]
  • Frontmatter placement of the inserted effort: line (after name:, else before the closing ---) has no harness effect; locked by the fixture suite only.

ACCEPTANCE CRITERIA

  1. Map + helper: lib/effort-pins.sh inserts, keeps, replaces (frontmatter only), skips a missing skill, is idempotent, rejects a bad level / traversal name / three-field line before writing, parses the real map. CHECK: out=$(make test suite=lib/tests/effort-pins.test.sh 2>&1); echo "$out" | grep -q 'effort-pins: [0-9]* pass, 0 fail' || { echo "$out" | grep FAIL; exit 1; }; echo PINS_GREEN EXPECT: PINS_GREEN EVIDENCE: MET exit=0 marker-found :: PINS_GREEN
  2. Census: the effort-routing suite is green and locks the three new repo levels, the map-driven vendored check, the design-stack single level, both re-apply call sites and the doctrine pointer. CHECK: out=$(make test suite=lib/tests/effort-routing.test.sh 2>&1); echo "$out" | grep -q 'census: [0-9]* pass, 0 fail' || { echo "$out" | grep FAIL; exit 1; }; for k in skills-perso pdf-translate site-motion effort-pins.txt 'stack_levels' 'apply_effort_pins'; do grep -q "$k" lib/tests/effort-routing.test.sh || { echo "census lacks $k"; exit 1; }; done; echo CENSUS_GREEN EXPECT: CENSUS_GREEN EVIDENCE: MET exit=0 marker-found :: CENSUS_GREEN
  3. Re-apply wired after the LAST vendoring step of both scripts, hardcoded loop gone: in install-plugins.sh the call follows the 21st pack staging block; in update-all.sh it follows the 21st pack refresh (§7.4, the last step that rewrites a SKILL.md), which itself follows the superpowers refresh. [verifier 2026-09-29: the first placement sat after the superpowers refresh only, the 21st refresh ran later and dropped seven pins] CHECK: a=$(grep -n 'apply_effort_pins "$REPO"' install-plugins.sh | cut -d: -f1); b=$(grep -n 'rm -rf "$TFD_STAGE"' install-plugins.sh | tail -1 | cut -d: -f1); c=$(grep -n 'apply_effort_pins "$REPO"' update-all.sh | cut -d: -f1); d=$(grep -n 'skills-external/$_tfd_name' update-all.sh | tail -1 | cut -d: -f1); e=$(grep -n 'vendor_pinned_skills superpowers refresh' update-all.sh | cut -d: -f1); [ "$(echo "$a" | wc -l)" -eq 1 ] && [ "$a" -gt "$b" ] && [ "$(echo "$c" | wc -l)" -eq 1 ] && [ -n "$d" ] && [ "$c" -gt "$d" ] && [ "$c" -gt "$e" ] && ! grep -q 'for _s in brainstorming writing-plans' install-plugins.sh && bash -n install-plugins.sh && bash -n update-all.sh && echo RESYNC_OK EXPECT: RESYNC_OK EVIDENCE: MET exit=0 marker-found :: RESYNC_OK
  4. Live tree: every map entry whose skill is vendored on this machine carries that level in its frontmatter (idempotent re-run applies 0). CHECK: out=$(bash lib/effort-pins.sh 2>&1) && echo "$out" | grep -q ' 0 applied, [0-9]* already at level' && echo LIVE_AT_LEVEL EXPECT: LIVE_AT_LEVEL EVIDENCE: MET exit=0 marker-found :: LIVE_AT_LEVEL
  5. Audit script: compiles, runs on a fixture with two records lacking output_tokens_details and one carrying it, reports 33 % coverage for that scope and the CAVEAT line (below 50 %). CHECK: python3 -m py_compile lib/effort-audit.py && D=$(mktemp -d) && mkdir -p "$D/p" && printf '%s\n%s\n%s\n' '{"type":"assistant","message":{"id":"m1","model":"claude-sonnet-5-5","usage":{"input_tokens":1,"output_tokens":10}}}' '{"type":"assistant","message":{"id":"m2","model":"claude-sonnet-5-5","usage":{"input_tokens":1,"output_tokens":10}}}' '{"type":"assistant","message":{"id":"m3","model":"claude-sonnet-5-5","usage":{"input_tokens":1,"output_tokens":10,"output_tokens_details":{"thinking_tokens":4}}}}' > "$D/p/s.jsonl" && out=$(python3 lib/effort-audit.py "$D") && echo "$out" | grep -q 'thinking counted on 33% of them' && echo "$out" | grep -q 'CAVEAT: main' && echo AUDIT_OK EXPECT: AUDIT_OK EVIDENCE: MET exit=0 marker-found :: AUDIT_OK
  6. Doctrine + docs: CLAUDE.global.md ≤ 320 lines with the paired-load line; README "## Effort routing"; USAGE "### Niveau d'effort"; CHANGELOG Added + Fixed entries; doctrine-citers census green. CHECK: [ "$(wc -l < CLAUDE.global.md)" -le 320 ] && grep -q 'a lone Skill call applies no effort' CLAUDE.global.md && grep -q '^## Effort routing' README.md && grep -q "^### Niveau d'effort" USAGE.md && grep -q 'Effort round (BDR-108)' CHANGELOG.md && grep -q 'never re-applied the effort pins' CHANGELOG.md && make test suite=lib/tests/doctrine-citers.test.sh 2>&1 | grep -q 'FAIL=0' && echo DOCS_OK EXPECT: DOCS_OK EVIDENCE: MET exit=0 marker-found :: DOCS_OK
  7. Health stack: shellcheck clean on the touched shell files and the Health Stack set; no-vacuous-locks green. CHECK: shellcheck lib/effort-pins.sh lib/tests/effort-pins.test.sh lib/tests/effort-routing.test.sh install-plugins.sh update-all.sh .sh hooks/.sh lib/*.sh && make test suite=lib/tests/no-vacuous-locks.test.sh >/dev/null 2>&1 && echo LINT_OK EXPECT: LINT_OK EVIDENCE: MET exit=0 marker-found :: LINT_OK

FILE SCOPE

  • lib/effort-pins.txt, lib/effort-pins.sh (new); lib/tests/effort-pins.test.sh (new); lib/tests/effort-routing.test.sh
  • install-plugins.sh, update-all.sh (re-apply call), lib/effort-audit.py (coverage)
  • skills/skills-perso/SKILL.md, skills/pdf-translate/SKILL.md, skills/site-motion/SKILL.md (effort line)
  • lib/effort-shift.md, CLAUDE.global.md (doctrine), README.md, USAGE.md, CHANGELOG.md
  • .claude/tasks/TODO.md, .claude/tasks/contracts/ (this file)
  • skills-external/design-motion-principles/SKILL.md [gated 2026-09-29] — the only vendored external tracked in git; its copy carries the effort: high line the resync re-applies (user choice: gate, not untrack)