From 5ed96aa8ed610746cc52065af18a73f28ce2c50d Mon Sep 17 00:00:00 2001 From: bastien Date: Mon, 28 Sep 2026 20:20:10 +0200 Subject: [PATCH] chore(memory): BDR-107 effort tiering, EVAL-036 A/B, journal, TODO W1-W4 ticked --- .claude/memory/decisions.md | 9 +++++++++ .claude/memory/evals.md | 8 ++++++++ .claude/memory/journal.md | 1 + .claude/tasks/TODO.md | 8 ++++---- 4 files changed, 22 insertions(+), 4 deletions(-) diff --git a/.claude/memory/decisions.md b/.claude/memory/decisions.md index 0410532..707cf57 100644 --- a/.claude/memory/decisions.md +++ b/.claude/memory/decisions.md @@ -128,6 +128,7 @@ rules: | BDR-104 | 2026-09-28 | MengTo motion pack: vendor 5 scroll skills pinned via shared lib/vendor-skills.sh + build personal skill site-motion; 17 skipped | accepted | | BDR-105 | 2026-09-28 | skill-catalog prune: 9 gstack out via GSTACK_REMOVED, full ⊇ every profile, max = everything, brightdata + frontend-design plugin off, security-guidance Stop review off, design gate asks `21st login` and waits | accepted | | BDR-106 | 2026-09-28 | superpowers: 7 wired skills vendored at v6.4.1 via lib/vendor-skills.sh (always_on lock class), plugin + marketplace dropped, citers by bare name, doctrine map for the 4 non-vendored refs | accepted | +| BDR-107 | 2026-09-28 | Effort tiering: session high, effort pins on 20 agents (BDR-077 second axis), entry level on 30 skills, five paired shifter skills, max at loop caps + ship-feature 4b | accepted | --- @@ -1330,3 +1331,11 @@ Branch feature/user-writing-web-rules, UNMERGED (human gate). - **Caveats**: upstream cross-refs to the plugin prefix and the 8 dropped skills remain in the vendored text (a call on a dropped name fails, doctrine map applies); no upstream auto-update (bump the pin deliberately); the harness hot-loaded the 7 bare names in the running session after link.sh, the plugin names leave at restart; `superpowers-marketplace` cache dir may linger empty; other machines: `make plugin` (vendors) + `make link`, then uninstall the cached plugin by hand (CHANGELOG). - **Reference**: 18f8c89 (wiring), ddea411 (citers/docs/settings); contract `2026-09-28-superpowers-vendored-1357` (12 criteria, oracles in `.oracles/`), plan r3 after 3 challengers (simplicity CONCERNS(2), robustness CONCERNS(3), correctness FATAL(5)) + confirmation CONCERNS(1); executors 2/2 DONE first pass; GATE 0 MET, verifier CONFORME 12/12, security PASS; catalog 82 skills, plugin passive cost 670 t (ui-ux-pro-max only). Links [[BDR-105]] [[BDR-102]] [[BDR-104]] [[BDR-065]] [[LRN-178]] [[EVAL-034]]. - **Amendment 2026-09-28 (merge)**: `gitflow finish` → 65665a5, no conflict, pushed, local + origin copies removed; the 7 vendored skills stay linked after the merge. Whole prune (tiers 1 + 2) on develop. + +## BDR-107 — Effort tiering: session high, agent pins, skill entry levels, paired phase shifts, max at escalation [accepted] (2026-09-28) +- **Decision**: settings `effortLevel` high (was xhigh). `effort:` pin on 20 repo-authored agents by role: low appliers (hotfixer, release-executor, plugin-probe, validator-analyzer), medium executors (feater, bugfixer, code-cleaner, onboarder, scaffolder), high judgment (refactorer, analyzer, commit-changer, doc-syncer, handover-doc-writer), xhigh challengers + gates (plan-challenger, plugin-advisor, verifier, security-auditor, seo-analyzer, geo-analyzer); none on interviewer/client-handover-writer (inline-load), status-reporter (haiku), impeccable-* (vendored). `effort:` on 28 tracked user-invoked skills = run entry level (low bookkeeping, medium gitflow/prune-memory, high feat/hotfix/bugfix/refactor/audits-with-fix, xhigh orchestrators) + xhigh on vendored brainstorming/writing-plans (skills-external/, re-applied by install-plugins STEP 8e). Five shifter skills `effort-{low,medium,high,xhigh,max}` loaded by orchestrators per `lib/effort-shift.md`: medium at dispatch span, own level before challenge synthesis, low at bookkeeping tail, max at verify-secure caps (GATE 0/1/2) + ship-feature 4b; re-assert after nested skill / prose gate. STOP texts name `$CLAUDE_EFFORT`, suggest `/effort-max`. statusline shows `$CLAUDE_EFFORT`; banner warns on `CLAUDE_CODE_EFFORT_LEVEL`. Census `lib/tests/effort-routing.test.sh`. Audit script `lib/effort-audit.py`. +- **Why**: session-wide xhigh burned thinking on bookkeeping; EVAL-035: 97 % of thinking in the main loop, sonnet subagents ~26 tok/request → main-loop levers (entry level, shifts) carry the savings; pins = explicitness + future models. A/B `/reconcile` high→low: requests 18→15, output −27 %, thinking −28 %, time −19 % (EVAL-036). +- **Harness facts (2.1.283)**: skill `effort:` applies on user slash invocation and on interactive Skill-tool load; the Skill-tool load applies ONLY when paired with another tool call in the same message (lone call = no-op); re-load re-applies (text deduped); not applied in `-p`/SDK; prompt cache kept across a shift; `CLAUDE_CODE_EFFORT_LEVEL` beats every frontmatter; one effort per agent file, no call-site override; unpinned agents inherit the level in force at dispatch. +- **Alternatives rejected**: executor pins only (they barely think); escalation-diagnoser agent fable+max (no context, one more agent; main-loop max keeps the failure context); reflection in fable skill-runner children with session medium (loses interactivity); settings.json rewrite mid-run (LRN-098 class); `maxEffortLevel` caps (hide a mis-pin the census should fail); pins on machine-generated skills (find-docs: ctx7 regenerates, gitignored) or gstack skills (spec, skillify). +- **Caveats**: shifts inert headless; a prose gate ending the turn resets to session level (re-assert wired in bugfix and ship-feature 4b); mode-based agents pin their judgment mode; a shift paired with a built-in judgment dispatch would downgrade it (pair with Read/Bash instead); `lib/gitflow-test.sh` T16a red on this machine = gitleaks not installed, unrelated. +- **Refs**: spec `docs/superpowers/specs/2026-09-28-effort-tiering-design.md`, plan `docs/superpowers/plans/2026-09-28-effort-tiering.md`, [[LRN-179]], [[EVAL-035]], [[EVAL-036]], [[BDR-077]]. diff --git a/.claude/memory/evals.md b/.claude/memory/evals.md index 07573d0..bfdca25 100644 --- a/.claude/memory/evals.md +++ b/.claude/memory/evals.md @@ -56,6 +56,7 @@ rules: | EVAL-033 | 2026-09-28 | case 7: 2 analyzers + 2 executors + 3 re-dispatches; verifiers caught shape, convention and my wrong count; security caught an env override | brief names the scratchpad path explicitly (3 /tmp leftovers); keep blind verifiers; count claims get an artifact | | EVAL-034 | 2026-09-28 | catalog prune + 21st gate: two challenge rounds each found what r3 missed (nested SKILL.md, fixture cp lists, in-session export); my ledgers failed twice (heredoc CHECKs); 5 executors DONE first pass; verifier gap = tool false positive | keep the confirmation pass on any plan that changed materially; one-line CHECKs; grep fixture cp lists before a `source` | | EVAL-035 | 2026-09-28 | thinking-share measurement, 6 days of transcripts (10,955 requests): thinking = 8 % of weighted spend, 97 % of it in the main loop; sonnet subagents at xhigh think 26 tok/request; cache reads = 53 % | pins = explicitness not savings; main-loop effort + context size are the levers; A/B after rollout | +| EVAL-036 | 2026-09-28 | A/B `/reconcile` headless, session high vs skill entry low: requests 18→15, output 12374→9038 (−27 %), thinking 3135→2248 (−28 %), time 96.5→78.4 s (−19 %), n=1 | keep low on bookkeeping skills; repeat on a reflection skill before touching the medium/high split | --- @@ -338,3 +339,10 @@ Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itse - **Result**: main loop 67 % of weighted spend, 97 % of thinking (Fable 1,430 think-tok/request); sonnet subagents 5,268 requests at xhigh, 26 think-tok/request; thinking = 8 % of spend, all output 16 %, cache reads 53 % (main-loop context ~320 k tok/request). Window 6 days only. Indirect effect of effort (fewer steps → fewer requests) unmeasured. - **Anomaly**: design was framed around executor pins; one script inverted it before any edit. Measure before routing. - **Action**: pins stay (explicitness, future models); main-loop skill effort + phase shifts carry the savings; A/B `/reconcile` high vs xhigh after rollout; context size = bigger lever, separate track. + +## EVAL-036 — A/B `/reconcile` headless: skill entry level low vs session high +- **Date**: 2026-09-28 +- **Method**: Task 4 of the effort-tiering plan; `claude -p "/reconcile" --output-format json --allowedTools Read Grep Glob "Bash(git status:*)" "Bash(git log:*)"` before (session `high`, no frontmatter) and after (`effort: low` on the skill); per-request `usage` summed from the session jsonl. +- **Result**: requests 18→15, output tokens 12374→9038 (−27 %), thinking 3135→2248 (−28 %), duration 96.5 s→78.4 s (−19 %); transcript effort field high→low confirmed. n=1, same repo state. +- **Anomaly**: none; the indirect effect (fewer steps at lower effort) is real, which EVAL-035's static split could not show. +- **Action**: keep low on bookkeeping skills; repeat on a reflection skill (feat) before touching the medium/high split; `lib/effort-audit.py` makes the split measurable any time. diff --git a/.claude/memory/journal.md b/.claude/memory/journal.md index 06ea276..65555bc 100644 --- a/.claude/memory/journal.md +++ b/.claude/memory/journal.md @@ -542,3 +542,4 @@ rules: - User go "merge le tier 2": feature/superpowers-vendored merged into develop via `gitflow finish` → 65665a5, no conflict, pushed, copies removed by the lib. develop == origin/develop, no working branch anywhere. Whole skill-catalog prune (BDR-105 + BDR-106) on develop: catalog 82 skills, plugin passive cost 670 t, no session injection. Open for the user: `21st login`, claude.ai skills off, floor-guard `xit(` hotfix (BLK-023), two /tmp fixture dirs, other machines `make plugin` + `make link` + uninstall the cached plugin. - /hotfix BLK-023 (user: "fais le hotfix du floor-guard"): `skip_kind` substring match → `xit(` ⊂ `exit(`. Fix 0deb559 on bugfix/floor-guard-xit-boundary: bare Jasmine names via `SKIP_IDENT_RE` lookbehind, 4 flip fixtures (12/12). 3 challengers (2 SOLID, robustness CONCERNS(2): fixture line itself flaggable on a test path → waiver comment outside the echo; my criterion-2 live oracle vacuous → dropped — same LRN-173 class, plus I wrote a heredoc CHECK again before catching it, [[LRN-176]]). Hotfixer DONE first pass, oracles MET, security PASS. UNMERGED — human gate. - User go "oui pour le changelog et merge le": CHANGELOG floor-guard entry amended via doc-syncer patch + doc-commit (018dfa3), bugfix/floor-guard-xit-boundary merged into develop via `gitflow finish` → c9f9b40, pushed, copies removed. develop == origin/develop, no working branch anywhere. Day total on develop: skill-catalog prune tiers 1 + 2 (BDR-105, BDR-106), 21st sign-in gate, BLK-023 resolved. +- effort tiering built on feature/effort-tiering (BDR-107): session high, 20 agent pins, 28+2 skill entry levels, 5 paired shifters, max at caps + 4b, census 129+ locks green, A/B −27 % output on /reconcile; finish awaits human signal. diff --git a/.claude/tasks/TODO.md b/.claude/tasks/TODO.md index 4932005..9425917 100644 --- a/.claude/tasks/TODO.md +++ b/.claude/tasks/TODO.md @@ -4,10 +4,10 @@ Spec `docs/superpowers/specs/2026-09-28-effort-tiering-design.md`, plan `docs/superpowers/plans/2026-09-28-effort-tiering.md`. Approved 2026-09-28: session high, A+B+C, max on the main loop at the loop caps + ship-feature 4b, superpowers patch. -- [ ] W1 settings high + banner warning + statusline live level + 20 agent pins + census suite (Tasks 1-3) -- [ ] W2 33 skill entry levels + superpowers xhigh with resync re-apply (Tasks 4, 9) -- [ ] W3 five shifters + lib/effort-shift.md + orchestrator wiring + max at caps/4b + gate audit (Tasks 5-8) -- [ ] W4 BDR id + CHANGELOG + EVAL A/B + journal + audit script (Tasks 10-11) +- [x] W1 settings high + banner warning + statusline live level + 20 agent pins + census suite (Tasks 1-3) +- [x] W2 33 skill entry levels + superpowers xhigh with resync re-apply (Tasks 4, 9) +- [x] W3 five shifters + lib/effort-shift.md + orchestrator wiring + max at caps/4b + gate audit (Tasks 5-8) +- [x] W4 BDR id + CHANGELOG + EVAL A/B + journal + audit script (Tasks 10-11) ## 2026-09-28 — tier 2: vendor 7 superpowers skills, drop the plugin (feature/superpowers-vendored) User go "fais le tier 2" (decision 2026-09-28, batch 1). Contract