chore(memory): BDR-108 effort round, LRN-181/182, BLK-024 resync pins, EVAL-038 sub-agent thinking unmeasured, journal, TODO parked LOW

This commit is contained in:
bastien
2026-09-29 15:41:25 +02:00
parent bb46ee22eb
commit 54a93eabe2
6 changed files with 41 additions and 1 deletions
+7
View File
@@ -43,6 +43,7 @@ rules:
| BLK-021 | 2026-09-22 | Bash tool dead mid-session ("every command exits 1"): /tmp usrquota blown by a dead session's probe HOMEs — 2… | open |
| BLK-022 | 2026-09-22 | `hooks/guard-bash.sh` withheld by the safety classifier; executable spec shipped instead — 2026-09-22 | open |
| BLK-023 | 2026-09-28 | floor-guard SKIP pattern `xit(` (Jasmine) matches any `exit(` in python/JS test helpers → false ECARTS; workaround: no `exit(` in inline python, bash derives rc from output — 2026-09-28 | resolved |
| BLK-024 | 2026-09-29 | update-all.sh re-fetched vendored skills but never re-applied the effort pins (lost until next `make plugin`); my first fix placed the re-apply BEFORE the late 21st refresh — rtk-truncated grep read as complete — 2026-09-29 | resolved |
---
@@ -268,3 +269,9 @@ rules:
- **Real cause**: `lib/floor-guard.sh` SKIP_SUBSTRINGS holds the bare fragment `'xit('` to catch Jasmine's `xit(…)`; `skip_kind()` is a plain substring match, so `sys.exit(`, `SystemExit(`, `process.exit(` all hit.
- **Solution**: workaround applied — the inline python prints violations only, the bash wrapper derives the return code from the captured output (no `exit(` anywhere). Root fix pending: word-bound the pattern (`(^|[^a-zA-Z_.])xit\(`) or match `xit(` only in JS/TS test files; hotfix-sized.
- **Status**: resolved 2026-09-28 — hotfix 0deb559 (bugfix/floor-guard-xit-boundary): the four bare Jasmine identifiers moved into `SKIP_IDENT_RE` with lookbehind `(?<![A-Za-z0-9_.])`, dotted/decorator forms stay substrings; fixtures SKIP_EXIT_CLEAN (RED before, GREEN after) + xit/fit/fdescribe flags. Residual `shortcut:` in the guard: `def fit(` / `function xit(` still match, `xit (` / `xit.each(` still do not (as before). Links [[BDR-105]], [[BDR-102]] (floor-guard origin), [[EVAL-034]].
## BLK-024 — resync dropped the vendored effort pins, twice — 2026-09-29
- **Friction**: [[BDR-107]] re-applied brainstorming/writing-plans xhigh only in install-plugins.sh STEP 8e; update-all.sh §7.3 re-fetches at the same commit → SKILL.md overwritten, `effort:` gone until the next `make plugin`. Latent since 2026-09-28.
- **Real cause (second instance)**: my re-apply call landed after the superpowers refresh; update-all.sh §7.4 (21st pack) runs LATER and `rm -rf` + `mv` every 21st-* SKILL.md. My grep of update-all.sh was truncated by rtk ("+28 more hidden") and I read the partial listing as the whole file. Fresh verifier caught it (ECARTS).
- **Solution**: `lib/effort-pins.txt` + `lib/effort-pins.sh` called ONCE after the LAST vendoring step of both scripts; census locks the order by line number (`ln_last`). Rule: a truncated tool listing is not a census; re-run without the pager or grep the anchor directly.
- **Status**: resolved 2026-09-29 (feature/effort-round, [[BDR-108]]).
+8
View File
@@ -129,6 +129,7 @@ rules:
| BDR-105 | 2026-09-28 | skill-catalog prune: 9 gstack out via GSTACK_REMOVED, full ⊇ every profile, max = everything, brightdata + frontend-design plugin off, security-guidance Stop review off, design gate asks `21st login` and waits | accepted |
| BDR-106 | 2026-09-28 | superpowers: 7 wired skills vendored at v6.4.1 via lib/vendor-skills.sh (always_on lock class), plugin + marketplace dropped, citers by bare name, doctrine map for the 4 non-vendored refs | accepted |
| BDR-107 | 2026-09-28 | Effort tiering: session high, effort pins on 20 agents (BDR-077 second axis), entry level on 30 skills, five paired shifter skills, max at loop caps + ship-feature 4b | accepted |
| BDR-108 | 2026-09-29 | Effort round: level on every skill next to its model pin (3 repo + 25 vendored via `lib/effort-pins.txt` re-applied after the LAST vendoring step of install + resync), design stack ONE level (high), model pins stay tier aliases: quality/price trade-off = tier × effort, never version | accepted |
---
@@ -1340,3 +1341,10 @@ Branch feature/user-writing-web-rules, UNMERGED (human gate).
- **Caveats**: shifts inert headless; a prose gate ending the turn resets to session level (re-assert wired in bugfix and ship-feature 4b); mode-based agents pin their judgment mode; a shift paired with a built-in judgment dispatch would downgrade it (pair with Read/Bash instead); `lib/gitflow-test.sh` T16a red on this machine = gitleaks not installed, unrelated.
- **Refs**: spec `docs/superpowers/specs/2026-09-28-effort-tiering-design.md`, plan `docs/superpowers/plans/2026-09-28-effort-tiering.md`, [[LRN-179]], [[EVAL-035]], [[EVAL-036]], [[BDR-077]].
- **Correction (2026-09-28)**: EVAL-035 counted one record per content block (~2.8× on request counts); deduped figures in [[EVAL-037]]: main-loop thinking 99.9% of total thinking (was 96.6%), thinking 5.6% of weighted cost (was 8.4%), sonnet think/request 26→0.2 tok. Conclusions hold, sharper: main loop still carries almost all thinking, executors stay cheap.
## BDR-108 — Effort round: every skill carries a level next to its model pin; model pins stay tier aliases [accepted] (2026-09-29)
- **Decision**: 3 repo skills pinned (skills-perso low, pdf-translate medium, site-motion high). 25 vendored externals (superpowers 7, agent-skills 3, design stack 9, 21st pack 6) get level from `lib/effort-pins.txt`, applied by `lib/effort-pins.sh` after LAST vendoring step of install-plugins.sh (21st pack, STEP 8.7) AND update-all.sh (§7.4). Design stack = ONE level, high. hotfix stays high. Model pins stay aliases (`sonnet` `opus` `haiku` `fable`). Doctrine: design stack loads paired with first Read; census order-locked by line number; `lib/effort-audit.py` prints thinking coverage.
- **Why**: user rungs (low fix-a-line · medium day-to-day · high refactor/resisting bug · xhigh architecture/audit · max stuck). Latest version of each tier = cheapest or same price (Sonnet 5.5 = Sonnet 5, Opus 5.5 < Opus 5, Haiku 4.5 alone, Fable 5.1 = Fable 5) → no version arbitration, only tier × effort. Stacked skills: last loaded wins → two levels in a stack = effort depends on load order. Aliases track generation free (transcripts: `sonnet` → sonnet-5 then sonnet-5-5).
- **Alternatives rejected**: full model IDs in frontmatter (maintenance, Agent-tool call site enum cannot pin a version, older gen never cheaper); hotfix → medium (no A/B on a reflection skill yet, [[EVAL-036]]); design stack medium; untrack design-motion-principles (only tracked external, gated in contract instead); pins on gstack / impeccable / graphify / find-docs / darwin (machine-owned, [[BDR-107]]).
- **Gates**: GATE 0 MET; verifier ECARTS(3): resync re-apply sat BEFORE the 21st refresh (real, fixed by fresh executor), tracked file out of scope (gated), shellcheck directive unauthorized (clarified) → CONFORME 7/7; security PASS + 4 LOW hardened (criteria 8-9); `make test` 44 suites rc 0.
- **Refs**: contract `.claude/tasks/contracts/2026-09-29-effort-round-1315.md`, [[BDR-107]], [[BDR-077]], [[LRN-181]], [[LRN-182]], [[BLK-024]], [[EVAL-038]].
+8
View File
@@ -58,6 +58,7 @@ rules:
| EVAL-035 | 2026-09-28 | thinking-share measurement, 6 days of transcripts (10,955 requests): thinking = 8 % of weighted spend, 97 % of it in the main loop; sonnet subagents at xhigh think 26 tok/request; cache reads = 53 % | pins = explicitness not savings; main-loop effort + context size are the levers; A/B after rollout |
| EVAL-036 | 2026-09-28 | A/B `/reconcile` headless, session high vs skill entry low: requests 18→15, output 12374→9038 (−27 %), thinking 3135→2248 (−28 %), time 96.5→78.4 s (−19 %), n=1 | keep low on bookkeeping skills; repeat on a reflection skill before touching the medium/high split |
| EVAL-037 | 2026-09-28 | correction of EVAL-035/036 counts: transcript records are per content block; deduped by message.id → main-loop thinking share 99.9%, thinking share of weighted cost 5.6%, sonnet think/msg 26→0.2, A/B requests 9→8 | conclusions hold (sharper: main-loop thinking 96.6%→99.9%, weighted-cost thinking corrected 8.4%→5.6%); effort-audit.py dedupes from a3b479e+ |
| EVAL-038 | 2026-09-29 | correction of EVAL-037: 94 % of sub-agent usage records carry no `output_tokens_details` (Fable subs at xhigh read 0 thinking, impossible with always-on thinking) → sub-agent thinking UNMEASURED, not ≈0; main loop 100 % counted; weighted-cost split (61/39) still holds | `effort-audit.py` prints coverage + CAVEAT; cite the cost split only; agent effort pins stay unmeasured; a tier move on a price argument = judgment, not figure |
---
@@ -354,3 +355,10 @@ Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itse
- **Result (deduped)**: main weighted-cost 61.4 %, thinking share 99.9 % (was 96.6 %); sub weighted-cost 38.6 %, thinking share 0.1 %; thinking = 5.6 % of weighted cost (was 8.4 %, inflated by duplicate counting); sonnet think/request 26→0.2 tok (sub, xhigh); A/B `/reconcile` (EVAL-036 rerun, deduped) requests 9→8, output 6129→4706, thinking 1550→1104 — the raw undeduped counts on the same transcripts are 18→15, matching EVAL-036 exactly (the bug, not the finding).
- **Anomaly**: the main-loop-carries-almost-all-thinking split got SHARPER after dedup (96.6→99.9 %), not weaker — duplication was near-uniform across content blocks, so ratios among scopes barely moved; only the absolute request/token counts and the overall thinking-share-of-cost figure were inflated (~2.2-2.8× depending on transcript mix).
- **Action**: `lib/effort-audit.py` dedupes by `message.id` from this commit; cite EVAL-037, not EVAL-035, for the split.
## EVAL-038 — correction of EVAL-037: sub-agent thinking is unmeasured, not ≈0
- **Date**: 2026-09-29
- **Output checked**: [[EVAL-037]] "sonnet think/request 26→0.2 tok, executors stay cheap; main loop carries 99.9 % of thinking".
- **Method**: scan of the last 400 transcripts, dedup by message.id, count records with/without `output_tokens_details`: sub 4586 requests, 6 % carry the field (2896/3075 sonnet-5 without, 65/72 fable-5-1 without); main 100 % carry it. A Fable 5.1 sub-agent at xhigh with 0 thinking tokens is impossible (thinking always on) → recording gap, not behaviour.
- **Anomaly**: "main loop = 99.9 % of thinking" is a coverage artefact. The weighted-cost split (main 61 % / sub 39 %) holds: `output_tokens` is always present.
- **Action**: `lib/effort-audit.py` counts `nodet`, prints `%counted` per row, "thinking counted on N% of them" per scope and a CAVEAT under 50 %; cite the cost split only; the 20 agent effort pins ([[BDR-107]]) remain unmeasured; a tier move argued on price stays a judgment ([[BDR-108]]).
+4
View File
@@ -545,3 +545,7 @@ rules:
- effort tiering built on feature/effort-tiering (BDR-107): session high, 20 agent pins, 28+2 skill entry levels, 5 paired shifters, max at caps + 4b, census 129+ locks green, A/B −27 % output on /reconcile; finish awaits human signal.
- User go "ok merge le": feature/effort-tiering merged into develop (94ede35) via gitflow finish, spec + plan purged (BDR-065), branch removed local + origin; leftovers for the user: .claude/skills/effort-probe-* and .superpowers/sdd/ scratch (deletes refused), gitleaks install (T16a), statusline visual check.
- From dotfiles repo (config): commit blocked, pre-commit ran `gitleaks git --staged`, Ubuntu apt gitleaks 8.16 has no `git` subcmd → exit 1 read as leak, every commit blocked. bugfix/gitleaks-protect-fallback 347073a: generator probes `gitleaks git --help`, falls back `protect --staged`; hooks regenerated; T16c symlink farm /usr/bin minus gitleaks (short PATH no longer hid an apt binary). make test rc 0, 170/0. User go "merge les deux": merged into develop 55b77e3 via gitflow finish, branch removed local + origin. Learning captured in config repo LRN-013.
## 2026-09-29
- Effort round on feature/effort-round ([[BDR-108]]): user table re-applied to all 88 linked skills; 30 existing levels hold, 3 repo skills pinned, 25 vendored externals pinned from `lib/effort-pins.txt` via `lib/effort-pins.sh` after the LAST vendoring step of install + resync (resync had dropped the BDR-107 pins, [[BLK-024]]); design stack ONE level high ([[LRN-181]]); model pins stay aliases, trade-off = tier × effort ([[LRN-182]]). Verifier ECARTS(3) caught my re-apply placed before the late 21st refresh (rtk-truncated grep read as complete) → fresh executor, CONFORME 7/7 then 9/9 after the 4-LOW hardening; security PASS ×2; `make test` 44 suites rc 0. Sub-agent thinking found unmeasured, not ≈0 ([[EVAL-038]]). 5 residual LOW parked in TODO. UNMERGED — human gate.
+10
View File
@@ -200,6 +200,8 @@ rules:
| LRN-178 | 2026-09-28 | a top-level `source` added to a lib breaks every hermetic suite that copies that lib alone into a fixture; grep the `cp` lists before adding one, or source lazily inside the branch that needs it | adding `source` to profile.sh / toggle-external.sh / any lib the suites copy |
| LRN-179 | 2026-09-28 | Skill `effort:` frontmatter shifts the MAIN LOOP for the rest of the turn on user slash invocation AND on interactive Skill-tool loads (last loaded wins, both directions, prompt cache kept); NOT applied in `-p`/headless; agent pins always honoured, unpinned agents inherit session | effort tiering; any skill or agent that must think more or less than the session |
| LRN-180 | 2026-09-28 | Skill-tool effort override needs a paired tool call: a lone Skill(effort-*) call is a no-op; a load in the same message as another tool call applies (the paired call already sees it); re-load re-applies (text deduped); skills Claude loads alone (brainstorming, writing-plans) apply nothing | every orchestrator shift; amends LRN-179 |
| LRN-181 | 2026-09-29 | Stacked skills share ONE effort level: skill `effort:` = last loaded wins, so a stack loaded in one build (design toolchain) with two levels gets an effort that depends on load order; a skill Claude loads alone applies nothing (LRN-180) | one level per stack in `lib/effort-pins.txt`; load the stack paired with the first Read; copy the stack level when vendoring a new design skill |
| LRN-182 | 2026-09-29 | Effort/thinking baselines are generation-bound and aliases move silently: `sonnet` resolved sonnet-5 then sonnet-5-5 mid-period, Sonnet 5.5 recalibrated its effort levels; EVAL-036 measured one generation | re-run `lib/effort-audit.py` after an alias moves; cite the generation in any effort measurement; never pin a version for it (older gen never cheaper) |
---
@@ -1650,3 +1652,11 @@ Rule: when editing a doctrine file under structure locks, grep the test's lock s
## LRN-180 — Skill-tool effort override needs a paired tool call; a lone Skill call is a no-op (2.1.283)
- **Context**: effort-tiering smoke. Six lone `Skill(effort-*)` / probe loads left `$CLAUDE_EFFORT` unchanged; every load issued in the same assistant message as another tool call applied, and the paired Bash already saw the new level. Re-loading an already-loaded shifter re-applies (text deduped: "already loaded above"). Final review: `brainstorming` / `writing-plans` loaded alone by ship-feature and init-project → their vendored xhigh pin inert. Amends [[LRN-179]].
- **Apply**: `Skill(effort-<level>)` always travels with the step's first tool call, shift first; pair a downward shift with a pinned executor or a Read/Bash, never with a built-in judgment dispatch; before any built-in judgment dispatch, pair the own-level shift with it; skills Claude loads alone do not apply their pin → re-assert with a paired shift at the resumed planning step ([[BDR-107]]).
## LRN-181 — Stacked skills share one effort level; a lone load applies none
- **Context**: design toolchain loads 5-8 skills in one build. Skill `effort:` frontmatter = last loaded wins, both directions ([[LRN-179]]). Two levels inside the stack → effort depends on load order, invisible. Plus [[LRN-180]]: a Skill call Claude issues alone is a no-op.
- **Apply**: one level per stack (`lib/effort-pins.txt` design section, census `stack_levels` lock, site-motion frontmatter matches); doctrine "load the stack paired with the first Read of the target file"; new vendored design skill → copy the stack level. [[BDR-108]]
## LRN-182 — Effort baselines are generation-bound; model aliases move silently
- **Context**: transcripts of the last weeks show `sonnet` → claude-sonnet-5 (3069 msgs) then claude-sonnet-5-5 (recent), `opus` → opus-5 then opus-5-5, `fable` → fable-5 then fable-5-1. API reference: Sonnet 5.5 recalibrated effort levels ("start at medium for agentic coding"). [[EVAL-036]] A/B ran on one generation.
- **Apply**: after an alias moves (new model in a tier) re-run `python3 lib/effort-audit.py` and re-read the pins; write the generation next to any effort figure; keep aliases (latest = cheapest or same price, never pin a version for a measurement). [[BDR-108]]