From a70430e683a53c57b1035512f74ffd2add56bb30 Mon Sep 17 00:00:00 2001 From: bastien Date: Mon, 28 Sep 2026 20:48:35 +0200 Subject: [PATCH] chore(memory): EVAL-037 deduped counts, BDR-107 correction, LRN-180 pairing rule, TODO count --- .claude/memory/decisions.md | 1 + .claude/memory/evals.md | 8 ++++++++ .claude/memory/learnings.md | 5 +++++ .claude/tasks/TODO.md | 2 +- 4 files changed, 15 insertions(+), 1 deletion(-) diff --git a/.claude/memory/decisions.md b/.claude/memory/decisions.md index 707cf57..d78f265 100644 --- a/.claude/memory/decisions.md +++ b/.claude/memory/decisions.md @@ -1339,3 +1339,4 @@ Branch feature/user-writing-web-rules, UNMERGED (human gate). - **Alternatives rejected**: executor pins only (they barely think); escalation-diagnoser agent fable+max (no context, one more agent; main-loop max keeps the failure context); reflection in fable skill-runner children with session medium (loses interactivity); settings.json rewrite mid-run (LRN-098 class); `maxEffortLevel` caps (hide a mis-pin the census should fail); pins on machine-generated skills (find-docs: ctx7 regenerates, gitignored) or gstack skills (spec, skillify). - **Caveats**: shifts inert headless; a prose gate ending the turn resets to session level (re-assert wired in bugfix and ship-feature 4b); mode-based agents pin their judgment mode; a shift paired with a built-in judgment dispatch would downgrade it (pair with Read/Bash instead); `lib/gitflow-test.sh` T16a red on this machine = gitleaks not installed, unrelated. - **Refs**: spec `docs/superpowers/specs/2026-09-28-effort-tiering-design.md`, plan `docs/superpowers/plans/2026-09-28-effort-tiering.md`, [[LRN-179]], [[EVAL-035]], [[EVAL-036]], [[BDR-077]]. +- **Correction (2026-09-28)**: EVAL-035 counted one record per content block (~2.8× on request counts); deduped figures in [[EVAL-037]]: main-loop thinking 99.9% of total thinking (was 96.6%), thinking 5.6% of weighted cost (was 8.4%), sonnet think/request 26→0.2 tok. Conclusions hold, sharper: main loop still carries almost all thinking, executors stay cheap. diff --git a/.claude/memory/evals.md b/.claude/memory/evals.md index bfdca25..2854bf5 100644 --- a/.claude/memory/evals.md +++ b/.claude/memory/evals.md @@ -57,6 +57,7 @@ rules: | EVAL-034 | 2026-09-28 | catalog prune + 21st gate: two challenge rounds each found what r3 missed (nested SKILL.md, fixture cp lists, in-session export); my ledgers failed twice (heredoc CHECKs); 5 executors DONE first pass; verifier gap = tool false positive | keep the confirmation pass on any plan that changed materially; one-line CHECKs; grep fixture cp lists before a `source` | | EVAL-035 | 2026-09-28 | thinking-share measurement, 6 days of transcripts (10,955 requests): thinking = 8 % of weighted spend, 97 % of it in the main loop; sonnet subagents at xhigh think 26 tok/request; cache reads = 53 % | pins = explicitness not savings; main-loop effort + context size are the levers; A/B after rollout | | EVAL-036 | 2026-09-28 | A/B `/reconcile` headless, session high vs skill entry low: requests 18→15, output 12374→9038 (−27 %), thinking 3135→2248 (−28 %), time 96.5→78.4 s (−19 %), n=1 | keep low on bookkeeping skills; repeat on a reflection skill before touching the medium/high split | +| EVAL-037 | 2026-09-28 | correction of EVAL-035/036 counts: transcript records are per content block; deduped by message.id → main-loop thinking share 99.9%, thinking share of weighted cost 5.6%, sonnet think/msg 26→0.2, A/B requests 9→8 | conclusions hold (sharper: main-loop thinking 96.6%→99.9%, weighted-cost thinking corrected 8.4%→5.6%); effort-audit.py dedupes from a3b479e+ | --- @@ -346,3 +347,10 @@ Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itse - **Result**: requests 18→15, output tokens 12374→9038 (−27 %), thinking 3135→2248 (−28 %), duration 96.5 s→78.4 s (−19 %); transcript effort field high→low confirmed. n=1, same repo state. - **Anomaly**: none; the indirect effect (fewer steps at lower effort) is real, which EVAL-035's static split could not show. - **Action**: keep low on bookkeeping skills; repeat on a reflection skill (feat) before touching the medium/high split; `lib/effort-audit.py` makes the split measurable any time. + +## EVAL-037 — correction of EVAL-035/036: one transcript record per content block, deduped by message.id +- **Date**: 2026-09-28 +- **Output checked**: EVAL-035 (8 % thinking / 97 % main loop / 26 tok per sonnet request) and EVAL-036 (requests 18→15), produced by `effort-audit.py` counting every assistant record; final review found duplicates (same `message.id` + identical `usage`, one record per content block, ~2.8× on this repo's last 6 transcripts). +- **Result (deduped)**: main weighted-cost 61.4 %, thinking share 99.9 % (was 96.6 %); sub weighted-cost 38.6 %, thinking share 0.1 %; thinking = 5.6 % of weighted cost (was 8.4 %, inflated by duplicate counting); sonnet think/request 26→0.2 tok (sub, xhigh); A/B `/reconcile` (EVAL-036 rerun, deduped) requests 9→8, output 6129→4706, thinking 1550→1104 — the raw undeduped counts on the same transcripts are 18→15, matching EVAL-036 exactly (the bug, not the finding). +- **Anomaly**: the main-loop-carries-almost-all-thinking split got SHARPER after dedup (96.6→99.9 %), not weaker — duplication was near-uniform across content blocks, so ratios among scopes barely moved; only the absolute request/token counts and the overall thinking-share-of-cost figure were inflated (~2.2-2.8× depending on transcript mix). +- **Action**: `lib/effort-audit.py` dedupes by `message.id` from this commit; cite EVAL-037, not EVAL-035, for the split. diff --git a/.claude/memory/learnings.md b/.claude/memory/learnings.md index b90abfa..3d0fe40 100644 --- a/.claude/memory/learnings.md +++ b/.claude/memory/learnings.md @@ -199,6 +199,7 @@ rules: | LRN-177 | 2026-09-28 | gstack skills hardcode `~/.claude/skills/gstack/` (83 paths: bin, scripts, ETHOS.md, */sections, review/specialists, make-pdf/dist, freeze/bin…); only bin + browse/dist were linked → dead skills and vacuous hooks (exit 127); ./setup plants a global symlink; whole-dir link exposes nested SKILL.md; `apply` is additive, `set` parks | any gstack wiring change, any "gstack skill fails" report | | LRN-178 | 2026-09-28 | a top-level `source` added to a lib breaks every hermetic suite that copies that lib alone into a fixture; grep the `cp` lists before adding one, or source lazily inside the branch that needs it | adding `source` to profile.sh / toggle-external.sh / any lib the suites copy | | LRN-179 | 2026-09-28 | Skill `effort:` frontmatter shifts the MAIN LOOP for the rest of the turn on user slash invocation AND on interactive Skill-tool loads (last loaded wins, both directions, prompt cache kept); NOT applied in `-p`/headless; agent pins always honoured, unpinned agents inherit session | effort tiering; any skill or agent that must think more or less than the session | +| LRN-180 | 2026-09-28 | Skill-tool effort override needs a paired tool call: a lone Skill(effort-*) call is a no-op; a load in the same message as another tool call applies (the paired call already sees it); re-load re-applies (text deduped); skills Claude loads alone (brainstorming, writing-plans) apply nothing | every orchestrator shift; amends LRN-179 | --- @@ -1645,3 +1646,7 @@ Rule: when editing a doctrine file under structure locks, grep the test's lock s ## LRN-179 — skill `effort:` shifts the main loop for the rest of the turn, interactive only - **Context**: effort-tiering spike 2026-09-28, Claude Code 2.1.283, Fable 5.1. Probes = `$CLAUDE_EFFORT` in Bash + transcript `effort` field per request. User-typed `/probe-low` → whole turn `low`. Skill-tool load in interactive session → `max` then `xhigh`, last loaded wins, both directions; first request after the switch read 206,996 cached tokens, wrote 1,164 (cache kept). Three `-p` runs: neither `effort:` nor `model:` skill frontmatter applied via Skill tool. Agent pin honoured (impeccable `medium`), unpinned built-in on sonnet inherited `xhigh`. Docs agent claimed "ultrathink keyword does not exist": wrong, docs = in-context nudge, API effort unchanged. Harness claims get verified against the harness ([[LRN-046]]). - **Apply**: main-loop effort per phase = `Skill(effort-)` on the main loop, never inside a dispatched agent; headless runs stay at session level; keep `CLAUDE_CODE_EFFORT_LEVEL` unset (beats every frontmatter). Spec `docs/superpowers/specs/2026-09-28-effort-tiering-design.md`. + +## LRN-180 — Skill-tool effort override needs a paired tool call; a lone Skill call is a no-op (2.1.283) +- **Context**: effort-tiering smoke. Six lone `Skill(effort-*)` / probe loads left `$CLAUDE_EFFORT` unchanged; every load issued in the same assistant message as another tool call applied, and the paired Bash already saw the new level. Re-loading an already-loaded shifter re-applies (text deduped: "already loaded above"). Final review: `brainstorming` / `writing-plans` loaded alone by ship-feature and init-project → their vendored xhigh pin inert. Amends [[LRN-179]]. +- **Apply**: `Skill(effort-)` always travels with the step's first tool call, shift first; pair a downward shift with a pinned executor or a Read/Bash, never with a built-in judgment dispatch; before any built-in judgment dispatch, pair the own-level shift with it; skills Claude loads alone do not apply their pin → re-assert with a paired shift at the resumed planning step ([[BDR-107]]). diff --git a/.claude/tasks/TODO.md b/.claude/tasks/TODO.md index 9425917..b6c4f18 100644 --- a/.claude/tasks/TODO.md +++ b/.claude/tasks/TODO.md @@ -5,7 +5,7 @@ Spec `docs/superpowers/specs/2026-09-28-effort-tiering-design.md`, plan `docs/superpowers/plans/2026-09-28-effort-tiering.md`. Approved 2026-09-28: session high, A+B+C, max on the main loop at the loop caps + ship-feature 4b, superpowers patch. - [x] W1 settings high + banner warning + statusline live level + 20 agent pins + census suite (Tasks 1-3) -- [x] W2 33 skill entry levels + superpowers xhigh with resync re-apply (Tasks 4, 9) +- [x] W2 28+2 skill entry levels + superpowers xhigh with resync re-apply (Tasks 4, 9) - [x] W3 five shifters + lib/effort-shift.md + orchestrator wiring + max at caps/4b + gate audit (Tasks 5-8) - [x] W4 BDR id + CHANGELOG + EVAL A/B + journal + audit script (Tasks 10-11)