From 854b74e9a47b5eef0cdecdec8332be151d7b6d12 Mon Sep 17 00:00:00 2001 From: bastien Date: Mon, 28 Sep 2026 18:24:34 +0200 Subject: [PATCH] =?UTF-8?q?chore(memory):=20LRN-179=20+=20EVAL-035=20?= =?UTF-8?q?=E2=80=94=20effort=20spike=20facts,=20thinking-share=20measurem?= =?UTF-8?q?ent?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- .claude/memory/evals.md | 8 ++++++++ .claude/memory/learnings.md | 5 +++++ 2 files changed, 13 insertions(+) diff --git a/.claude/memory/evals.md b/.claude/memory/evals.md index 509f95d..07573d0 100644 --- a/.claude/memory/evals.md +++ b/.claude/memory/evals.md @@ -55,6 +55,7 @@ rules: | EVAL-032 | 2026-09-27 | 4 parallel feater executors, one tree, gate loop: verifier caught a vacuous test, security caught a partial-write; my oracles wrong twice | keep same-tree parallel dispatch with disjoint FILE SCOPE + orchestrator-owned shared files; blind verifier stays; measure oracles on precedents | | EVAL-033 | 2026-09-28 | case 7: 2 analyzers + 2 executors + 3 re-dispatches; verifiers caught shape, convention and my wrong count; security caught an env override | brief names the scratchpad path explicitly (3 /tmp leftovers); keep blind verifiers; count claims get an artifact | | EVAL-034 | 2026-09-28 | catalog prune + 21st gate: two challenge rounds each found what r3 missed (nested SKILL.md, fixture cp lists, in-session export); my ledgers failed twice (heredoc CHECKs); 5 executors DONE first pass; verifier gap = tool false positive | keep the confirmation pass on any plan that changed materially; one-line CHECKs; grep fixture cp lists before a `source` | +| EVAL-035 | 2026-09-28 | thinking-share measurement, 6 days of transcripts (10,955 requests): thinking = 8 % of weighted spend, 97 % of it in the main loop; sonnet subagents at xhigh think 26 tok/request; cache reads = 53 % | pins = explicitness not savings; main-loop effort + context size are the levers; A/B after rollout | --- @@ -330,3 +331,10 @@ Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itse - **Result**: prune — challengers closed 8 MAJOR at r3, the confirmation pass still found 1 BLOCKER (nested SKILL.md in browser-skills/openclaw/node_modules) + 3 MAJOR (setup's global symlink, update-all 3rd copy, fixture cp lists); executors 4/4 DONE first pass; GATE 0 UNMET(4) = my heredoc CHECKs ([[LRN-176]]); verifier ECARTS(1) = floor-guard false positive ([[BLK-023]]), CONFORME at iteration 2; security PASS. 21st gate — three lenses: my shared-helper reflex = BLOCKER ×2 ([[LRN-178]]), my `export TWENTYFIRST_TOKEN` remedy = MAJOR (env does not persist); confirmation pass pinned the diagnostic format; executor DONE first pass, CONFORME 7/7, PASS. - **Anomalies**: (1) both times the confirmation pass found real defects after "all MAJOR closed" → r3 is not a stopping point; (2) every gate failure of the day was mine (ledger format, tool pattern), none the executors'; (3) verifier and challengers each re-ran the live oracles themselves (link.sh, `set full`, the gate) — cheap, decisive; (4) the user's rule ("full ⊇ every profile") arrived at pass B and inverted a settled plan step: pass B before challenge is the right order. - **Action**: keep the single confirmation pass mandatory when a plan changed materially; contract CHECKs one line, files under `.oracles/`; grep fixture `cp` lists before any new `source`; run the live oracle once by hand before dispatching the verifier. + +## EVAL-035 — effort burn measured, premise corrected: subagents don't think, the main loop does +- **Date**: 2026-09-28 +- **Output checked**: my hypothesis "executors inherit xhigh → that is the burn" vs `effort_split2.py` (scratchpad) over `~/.claude/projects/*`: main jsonl + `*/subagents/*.jsonl`, `isSidechain` split; weights output ×5, cache read ×0.1, cache write ×1.25. +- **Result**: main loop 67 % of weighted spend, 97 % of thinking (Fable 1,430 think-tok/request); sonnet subagents 5,268 requests at xhigh, 26 think-tok/request; thinking = 8 % of spend, all output 16 %, cache reads 53 % (main-loop context ~320 k tok/request). Window 6 days only. Indirect effect of effort (fewer steps → fewer requests) unmeasured. +- **Anomaly**: design was framed around executor pins; one script inverted it before any edit. Measure before routing. +- **Action**: pins stay (explicitness, future models); main-loop skill effort + phase shifts carry the savings; A/B `/reconcile` high vs xhigh after rollout; context size = bigger lever, separate track. diff --git a/.claude/memory/learnings.md b/.claude/memory/learnings.md index 269bfcf..b90abfa 100644 --- a/.claude/memory/learnings.md +++ b/.claude/memory/learnings.md @@ -198,6 +198,7 @@ rules: | LRN-176 | 2026-09-28 | gates.sh `CHECK:` is single-line: a heredoc body reads as prose, the oracle runs `python3 -` on empty stdin and lands NOT-MET "marker absent", never ERROR; multi-line oracle → `.oracles/*.py` | writing contract oracles longer than one line | | LRN-177 | 2026-09-28 | gstack skills hardcode `~/.claude/skills/gstack/` (83 paths: bin, scripts, ETHOS.md, */sections, review/specialists, make-pdf/dist, freeze/bin…); only bin + browse/dist were linked → dead skills and vacuous hooks (exit 127); ./setup plants a global symlink; whole-dir link exposes nested SKILL.md; `apply` is additive, `set` parks | any gstack wiring change, any "gstack skill fails" report | | LRN-178 | 2026-09-28 | a top-level `source` added to a lib breaks every hermetic suite that copies that lib alone into a fixture; grep the `cp` lists before adding one, or source lazily inside the branch that needs it | adding `source` to profile.sh / toggle-external.sh / any lib the suites copy | +| LRN-179 | 2026-09-28 | Skill `effort:` frontmatter shifts the MAIN LOOP for the rest of the turn on user slash invocation AND on interactive Skill-tool loads (last loaded wins, both directions, prompt cache kept); NOT applied in `-p`/headless; agent pins always honoured, unpinned agents inherit session | effort tiering; any skill or agent that must think more or less than the session | --- @@ -1640,3 +1641,7 @@ Rule: when editing a doctrine file under structure locks, grep the test's lock s ## LRN-178 — before a new top-level `source`, grep the fixture `cp` lists - **Context**: twice in one day. E1b's `source gstack-removed.sh` in profile.sh/toggle-external.sh needed a `cp` line in three suites (profile-default, profile-set-managed, toggle-external-repo-resolution) — caught by the confirmation challenger, fixed in scope. My 21st helper plan would have added a second top-level `source` to toggle-external.sh with no fixture update → four suites red under `set -euo pipefail`; two challengers flagged it as BLOCKER, the helper was dropped. - **Apply**: `grep -n "cp .*lib/" lib/tests/*.sh` before adding a `source` to a lib; either widen every fixture copy in the same change or source lazily inside the one branch that needs it. Prefer the inline predicate when only one caller needs the new semantics ([[BDR-105]]). + +## LRN-179 — skill `effort:` shifts the main loop for the rest of the turn, interactive only +- **Context**: effort-tiering spike 2026-09-28, Claude Code 2.1.283, Fable 5.1. Probes = `$CLAUDE_EFFORT` in Bash + transcript `effort` field per request. User-typed `/probe-low` → whole turn `low`. Skill-tool load in interactive session → `max` then `xhigh`, last loaded wins, both directions; first request after the switch read 206,996 cached tokens, wrote 1,164 (cache kept). Three `-p` runs: neither `effort:` nor `model:` skill frontmatter applied via Skill tool. Agent pin honoured (impeccable `medium`), unpinned built-in on sonnet inherited `xhigh`. Docs agent claimed "ultrathink keyword does not exist": wrong, docs = in-context nudge, API effort unchanged. Harness claims get verified against the harness ([[LRN-046]]). +- **Apply**: main-loop effort per phase = `Skill(effort-)` on the main loop, never inside a dispatched agent; headless runs stay at session level; keep `CLAUDE_CODE_EFFORT_LEVEL` unset (beats every frontmatter). Spec `docs/superpowers/specs/2026-09-28-effort-tiering-design.md`.