chore(memory): BDR-107 effort tiering, EVAL-036 A/B, journal, TODO W1-W4 ticked

This commit is contained in:
bastien
2026-09-28 20:20:10 +02:00
parent 98ef991958
commit 5ed96aa8ed
4 changed files with 22 additions and 4 deletions
+9
View File
@@ -128,6 +128,7 @@ rules:
| BDR-104 | 2026-09-28 | MengTo motion pack: vendor 5 scroll skills pinned via shared lib/vendor-skills.sh + build personal skill site-motion; 17 skipped | accepted |
| BDR-105 | 2026-09-28 | skill-catalog prune: 9 gstack out via GSTACK_REMOVED, full ⊇ every profile, max = everything, brightdata + frontend-design plugin off, security-guidance Stop review off, design gate asks `21st login` and waits | accepted |
| BDR-106 | 2026-09-28 | superpowers: 7 wired skills vendored at v6.4.1 via lib/vendor-skills.sh (always_on lock class), plugin + marketplace dropped, citers by bare name, doctrine map for the 4 non-vendored refs | accepted |
| BDR-107 | 2026-09-28 | Effort tiering: session high, effort pins on 20 agents (BDR-077 second axis), entry level on 30 skills, five paired shifter skills, max at loop caps + ship-feature 4b | accepted |
---
@@ -1330,3 +1331,11 @@ Branch feature/user-writing-web-rules, UNMERGED (human gate).
- **Caveats**: upstream cross-refs to the plugin prefix and the 8 dropped skills remain in the vendored text (a call on a dropped name fails, doctrine map applies); no upstream auto-update (bump the pin deliberately); the harness hot-loaded the 7 bare names in the running session after link.sh, the plugin names leave at restart; `superpowers-marketplace` cache dir may linger empty; other machines: `make plugin` (vendors) + `make link`, then uninstall the cached plugin by hand (CHANGELOG).
- **Reference**: 18f8c89 (wiring), ddea411 (citers/docs/settings); contract `2026-09-28-superpowers-vendored-1357` (12 criteria, oracles in `.oracles/`), plan r3 after 3 challengers (simplicity CONCERNS(2), robustness CONCERNS(3), correctness FATAL(5)) + confirmation CONCERNS(1); executors 2/2 DONE first pass; GATE 0 MET, verifier CONFORME 12/12, security PASS; catalog 82 skills, plugin passive cost 670 t (ui-ux-pro-max only). Links [[BDR-105]] [[BDR-102]] [[BDR-104]] [[BDR-065]] [[LRN-178]] [[EVAL-034]].
- **Amendment 2026-09-28 (merge)**: `gitflow finish` → 65665a5, no conflict, pushed, local + origin copies removed; the 7 vendored skills stay linked after the merge. Whole prune (tiers 1 + 2) on develop.
## BDR-107 — Effort tiering: session high, agent pins, skill entry levels, paired phase shifts, max at escalation [accepted] (2026-09-28)
- **Decision**: settings `effortLevel` high (was xhigh). `effort:` pin on 20 repo-authored agents by role: low appliers (hotfixer, release-executor, plugin-probe, validator-analyzer), medium executors (feater, bugfixer, code-cleaner, onboarder, scaffolder), high judgment (refactorer, analyzer, commit-changer, doc-syncer, handover-doc-writer), xhigh challengers + gates (plan-challenger, plugin-advisor, verifier, security-auditor, seo-analyzer, geo-analyzer); none on interviewer/client-handover-writer (inline-load), status-reporter (haiku), impeccable-* (vendored). `effort:` on 28 tracked user-invoked skills = run entry level (low bookkeeping, medium gitflow/prune-memory, high feat/hotfix/bugfix/refactor/audits-with-fix, xhigh orchestrators) + xhigh on vendored brainstorming/writing-plans (skills-external/, re-applied by install-plugins STEP 8e). Five shifter skills `effort-{low,medium,high,xhigh,max}` loaded by orchestrators per `lib/effort-shift.md`: medium at dispatch span, own level before challenge synthesis, low at bookkeeping tail, max at verify-secure caps (GATE 0/1/2) + ship-feature 4b; re-assert after nested skill / prose gate. STOP texts name `$CLAUDE_EFFORT`, suggest `/effort-max`. statusline shows `$CLAUDE_EFFORT`; banner warns on `CLAUDE_CODE_EFFORT_LEVEL`. Census `lib/tests/effort-routing.test.sh`. Audit script `lib/effort-audit.py`.
- **Why**: session-wide xhigh burned thinking on bookkeeping; EVAL-035: 97 % of thinking in the main loop, sonnet subagents ~26 tok/request → main-loop levers (entry level, shifts) carry the savings; pins = explicitness + future models. A/B `/reconcile` high→low: requests 18→15, output −27 %, thinking −28 %, time −19 % (EVAL-036).
- **Harness facts (2.1.283)**: skill `effort:` applies on user slash invocation and on interactive Skill-tool load; the Skill-tool load applies ONLY when paired with another tool call in the same message (lone call = no-op); re-load re-applies (text deduped); not applied in `-p`/SDK; prompt cache kept across a shift; `CLAUDE_CODE_EFFORT_LEVEL` beats every frontmatter; one effort per agent file, no call-site override; unpinned agents inherit the level in force at dispatch.
- **Alternatives rejected**: executor pins only (they barely think); escalation-diagnoser agent fable+max (no context, one more agent; main-loop max keeps the failure context); reflection in fable skill-runner children with session medium (loses interactivity); settings.json rewrite mid-run (LRN-098 class); `maxEffortLevel` caps (hide a mis-pin the census should fail); pins on machine-generated skills (find-docs: ctx7 regenerates, gitignored) or gstack skills (spec, skillify).
- **Caveats**: shifts inert headless; a prose gate ending the turn resets to session level (re-assert wired in bugfix and ship-feature 4b); mode-based agents pin their judgment mode; a shift paired with a built-in judgment dispatch would downgrade it (pair with Read/Bash instead); `lib/gitflow-test.sh` T16a red on this machine = gitleaks not installed, unrelated.
- **Refs**: spec `docs/superpowers/specs/2026-09-28-effort-tiering-design.md`, plan `docs/superpowers/plans/2026-09-28-effort-tiering.md`, [[LRN-179]], [[EVAL-035]], [[EVAL-036]], [[BDR-077]].