- lib/effort-pins.txt (map) + lib/effort-pins.sh (idempotent re-apply) replace the hardcoded brainstorming/writing-plans loop; called after the last vendoring step of install-plugins.sh AND update-all.sh (the resync dropped the pins until the next make plugin) - design stack high uniform (last loaded wins), superpowers, agent-skills, 21st pack pinned from the map; skills-perso low, pdf-translate medium, site-motion high - doctrine: design stack loads paired with the first Read; one level per stack (CLAUDE.global.md, lib/effort-shift.md) - lib/effort-audit.py prints thinking coverage per scope (sub-agent records carry no thinking count on ~94 % of requests) - census map-driven + fixture suite lib/tests/effort-pins.test.sh; docs README/USAGE/CHANGELOG; contract + TODO plan
52 lines
6.9 KiB
Markdown
52 lines
6.9 KiB
Markdown
# CONTRACT — effort-round
|
||
- date: 2026-09-29 | flow: feat by hand (feature/* off develop) | branch: feature/effort-round
|
||
- status: active
|
||
|
||
## REQUEST (verbatim — IMMUTABLE)
|
||
> en se basant sur le meme tableau que la derniere fois [low: corriger une ligne, renommer un fichier, lancer un script · medium: le travail courant · high: un refactor, un bug qui resiste · xhigh: architecture, audit avant validation · max: quand une erreur coince, une erreur ne se rattrape pas, ou qu'on juge avoir besoin de beaucoup de reflexion], quand on a pin les orchestrateurs et leur sous agent a des efforts, j'aimerais que tu fasse une ronde de tout les skill et que tu mete un niveau d'effort en plus du model pin. D'ailleurs les model pin, c'est du par exemple Sonnet ou du Sonnet 5.5 (version du model pinned) ? Car il faudrait utiliser les versions qui vont bien avec la tache qu'ils ont a acomplir.
|
||
> [answered: model pins stay tier aliases; the latest version of a tier is also the cheapest or same-priced, the quality/price trade-off is tier × effort]
|
||
|
||
## CLARIFICATIONS
|
||
- User choices 2026-09-29 (AskUserQuestion): design stack high uniform; hotfix stays high; every vendored external of the proposed table gets a pin; README/USAGE docs in the same branch.
|
||
- Round result: 30 existing entry levels hold against the table; 3 repo skills had none (skills-perso low, pdf-translate medium, site-motion high); vendored externals get theirs from `lib/effort-pins.txt` re-applied by `lib/effort-pins.sh`; impeccable, graphify, find-docs, gstack, darwin-skill and the five shifters stay unpinned (machine-owned, BDR-107).
|
||
- Defect found in passing, fixed here: `update-all.sh` re-fetched the vendored skills but never re-applied the pins (lost until the next `make plugin`).
|
||
- Defect found in passing, surfaced not fixed: ~94 % of sub-agent usage records carry no `output_tokens_details`, so `lib/effort-audit.py` read zero thinking on sub-agents; the script now prints coverage and a CAVEAT; EVAL-037's "executors stay cheap" is a measurement gap (registry correction pending user approval).
|
||
- Frontmatter placement of the inserted `effort:` line (after `name:`, else before the closing `---`) has no harness effect; locked by the fixture suite only.
|
||
|
||
## ACCEPTANCE CRITERIA
|
||
1. Map + helper: `lib/effort-pins.sh` inserts, keeps, replaces (frontmatter only), skips a missing skill, is idempotent, rejects a bad level / traversal name / three-field line before writing, parses the real map.
|
||
CHECK: out=$(make test suite=lib/tests/effort-pins.test.sh 2>&1); echo "$out" | grep -q 'effort-pins: [0-9]* pass, 0 fail' || { echo "$out" | grep FAIL; exit 1; }; echo PINS_GREEN
|
||
EXPECT: PINS_GREEN
|
||
EVIDENCE: MET exit=0 marker-found :: PINS_GREEN
|
||
2. Census: the effort-routing suite is green and locks the three new repo levels, the map-driven vendored check, the design-stack single level, both re-apply call sites and the doctrine pointer.
|
||
CHECK: out=$(make test suite=lib/tests/effort-routing.test.sh 2>&1); echo "$out" | grep -q 'census: [0-9]* pass, 0 fail' || { echo "$out" | grep FAIL; exit 1; }; for k in skills-perso pdf-translate site-motion effort-pins.txt 'stack_levels' 'apply_effort_pins'; do grep -q "$k" lib/tests/effort-routing.test.sh || { echo "census lacks $k"; exit 1; }; done; echo CENSUS_GREEN
|
||
EXPECT: CENSUS_GREEN
|
||
EVIDENCE: MET exit=0 marker-found :: CENSUS_GREEN
|
||
3. Re-apply wired after the LAST vendoring step of both scripts, hardcoded loop gone: in install-plugins.sh the call follows the 21st pack staging block; in update-all.sh it follows the superpowers refresh.
|
||
CHECK: a=$(grep -n 'apply_effort_pins "$REPO"' install-plugins.sh | cut -d: -f1); b=$(grep -n 'rm -rf "$TFD_STAGE"' install-plugins.sh | tail -1 | cut -d: -f1); c=$(grep -n 'apply_effort_pins "$REPO"' update-all.sh | cut -d: -f1); d=$(grep -n 'vendor_pinned_skills superpowers refresh' update-all.sh | cut -d: -f1); [ "$(echo "$a" | wc -l)" -eq 1 ] && [ "$a" -gt "$b" ] && [ "$(echo "$c" | wc -l)" -eq 1 ] && [ "$c" -gt "$d" ] && ! grep -q 'for _s in brainstorming writing-plans' install-plugins.sh && bash -n install-plugins.sh && bash -n update-all.sh && echo RESYNC_OK
|
||
EXPECT: RESYNC_OK
|
||
EVIDENCE: MET exit=0 marker-found :: RESYNC_OK
|
||
4. Live tree: every map entry whose skill is vendored on this machine carries that level in its frontmatter (idempotent re-run applies 0).
|
||
CHECK: out=$(bash lib/effort-pins.sh 2>&1) && echo "$out" | grep -q ' 0 applied, [0-9]* already at level' && echo LIVE_AT_LEVEL
|
||
EXPECT: LIVE_AT_LEVEL
|
||
EVIDENCE: MET exit=0 marker-found :: LIVE_AT_LEVEL
|
||
5. Audit script: compiles, runs on a fixture with two records lacking `output_tokens_details` and one carrying it, reports 33 % coverage for that scope and the CAVEAT line (below 50 %).
|
||
CHECK: python3 -m py_compile lib/effort-audit.py && D=$(mktemp -d) && mkdir -p "$D/p" && printf '%s\n%s\n%s\n' '{"type":"assistant","message":{"id":"m1","model":"claude-sonnet-5-5","usage":{"input_tokens":1,"output_tokens":10}}}' '{"type":"assistant","message":{"id":"m2","model":"claude-sonnet-5-5","usage":{"input_tokens":1,"output_tokens":10}}}' '{"type":"assistant","message":{"id":"m3","model":"claude-sonnet-5-5","usage":{"input_tokens":1,"output_tokens":10,"output_tokens_details":{"thinking_tokens":4}}}}' > "$D/p/s.jsonl" && out=$(python3 lib/effort-audit.py "$D") && echo "$out" | grep -q 'thinking counted on 33% of them' && echo "$out" | grep -q 'CAVEAT: main' && echo AUDIT_OK
|
||
EXPECT: AUDIT_OK
|
||
EVIDENCE: MET exit=0 marker-found :: AUDIT_OK
|
||
6. Doctrine + docs: CLAUDE.global.md ≤ 320 lines with the paired-load line; README "## Effort routing"; USAGE "### Niveau d'effort"; CHANGELOG Added + Fixed entries; doctrine-citers census green.
|
||
CHECK: [ "$(wc -l < CLAUDE.global.md)" -le 320 ] && grep -q 'a lone Skill call applies no effort' CLAUDE.global.md && grep -q '^## Effort routing' README.md && grep -q "^### Niveau d'effort" USAGE.md && grep -q 'Effort round (BDR-108)' CHANGELOG.md && grep -q 'never re-applied the effort pins' CHANGELOG.md && make test suite=lib/tests/doctrine-citers.test.sh 2>&1 | grep -q 'FAIL=0' && echo DOCS_OK
|
||
EXPECT: DOCS_OK
|
||
EVIDENCE: MET exit=0 marker-found :: DOCS_OK
|
||
7. Health stack: shellcheck clean on the touched shell files and the Health Stack set; no-vacuous-locks green.
|
||
CHECK: shellcheck lib/effort-pins.sh lib/tests/effort-pins.test.sh lib/tests/effort-routing.test.sh install-plugins.sh update-all.sh *.sh hooks/*.sh lib/*.sh && make test suite=lib/tests/no-vacuous-locks.test.sh >/dev/null 2>&1 && echo LINT_OK
|
||
EXPECT: LINT_OK
|
||
EVIDENCE: MET exit=0 marker-found :: LINT_OK
|
||
|
||
## FILE SCOPE
|
||
- lib/effort-pins.txt, lib/effort-pins.sh (new); lib/tests/effort-pins.test.sh (new); lib/tests/effort-routing.test.sh
|
||
- install-plugins.sh, update-all.sh (re-apply call), lib/effort-audit.py (coverage)
|
||
- skills/skills-perso/SKILL.md, skills/pdf-translate/SKILL.md, skills/site-motion/SKILL.md (effort line)
|
||
- lib/effort-shift.md, CLAUDE.global.md (doctrine), README.md, USAGE.md, CHANGELOG.md
|
||
- .claude/tasks/TODO.md, .claude/tasks/contracts/ (this file)
|