chore(memory): LRN-179 + EVAL-035 — effort spike facts, thinking-share measurement
This commit is contained in:
@@ -55,6 +55,7 @@ rules:
|
||||
| EVAL-032 | 2026-09-27 | 4 parallel feater executors, one tree, gate loop: verifier caught a vacuous test, security caught a partial-write; my oracles wrong twice | keep same-tree parallel dispatch with disjoint FILE SCOPE + orchestrator-owned shared files; blind verifier stays; measure oracles on precedents |
|
||||
| EVAL-033 | 2026-09-28 | case 7: 2 analyzers + 2 executors + 3 re-dispatches; verifiers caught shape, convention and my wrong count; security caught an env override | brief names the scratchpad path explicitly (3 /tmp leftovers); keep blind verifiers; count claims get an artifact |
|
||||
| EVAL-034 | 2026-09-28 | catalog prune + 21st gate: two challenge rounds each found what r3 missed (nested SKILL.md, fixture cp lists, in-session export); my ledgers failed twice (heredoc CHECKs); 5 executors DONE first pass; verifier gap = tool false positive | keep the confirmation pass on any plan that changed materially; one-line CHECKs; grep fixture cp lists before a `source` |
|
||||
| EVAL-035 | 2026-09-28 | thinking-share measurement, 6 days of transcripts (10,955 requests): thinking = 8 % of weighted spend, 97 % of it in the main loop; sonnet subagents at xhigh think 26 tok/request; cache reads = 53 % | pins = explicitness not savings; main-loop effort + context size are the levers; A/B after rollout |
|
||||
|
||||
---
|
||||
|
||||
@@ -330,3 +331,10 @@ Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itse
|
||||
- **Result**: prune — challengers closed 8 MAJOR at r3, the confirmation pass still found 1 BLOCKER (nested SKILL.md in browser-skills/openclaw/node_modules) + 3 MAJOR (setup's global symlink, update-all 3rd copy, fixture cp lists); executors 4/4 DONE first pass; GATE 0 UNMET(4) = my heredoc CHECKs ([[LRN-176]]); verifier ECARTS(1) = floor-guard false positive ([[BLK-023]]), CONFORME at iteration 2; security PASS. 21st gate — three lenses: my shared-helper reflex = BLOCKER ×2 ([[LRN-178]]), my `export TWENTYFIRST_TOKEN` remedy = MAJOR (env does not persist); confirmation pass pinned the diagnostic format; executor DONE first pass, CONFORME 7/7, PASS.
|
||||
- **Anomalies**: (1) both times the confirmation pass found real defects after "all MAJOR closed" → r3 is not a stopping point; (2) every gate failure of the day was mine (ledger format, tool pattern), none the executors'; (3) verifier and challengers each re-ran the live oracles themselves (link.sh, `set full`, the gate) — cheap, decisive; (4) the user's rule ("full ⊇ every profile") arrived at pass B and inverted a settled plan step: pass B before challenge is the right order.
|
||||
- **Action**: keep the single confirmation pass mandatory when a plan changed materially; contract CHECKs one line, files under `.oracles/`; grep fixture `cp` lists before any new `source`; run the live oracle once by hand before dispatching the verifier.
|
||||
|
||||
## EVAL-035 — effort burn measured, premise corrected: subagents don't think, the main loop does
|
||||
- **Date**: 2026-09-28
|
||||
- **Output checked**: my hypothesis "executors inherit xhigh → that is the burn" vs `effort_split2.py` (scratchpad) over `~/.claude/projects/*`: main jsonl + `*/subagents/*.jsonl`, `isSidechain` split; weights output ×5, cache read ×0.1, cache write ×1.25.
|
||||
- **Result**: main loop 67 % of weighted spend, 97 % of thinking (Fable 1,430 think-tok/request); sonnet subagents 5,268 requests at xhigh, 26 think-tok/request; thinking = 8 % of spend, all output 16 %, cache reads 53 % (main-loop context ~320 k tok/request). Window 6 days only. Indirect effect of effort (fewer steps → fewer requests) unmeasured.
|
||||
- **Anomaly**: design was framed around executor pins; one script inverted it before any edit. Measure before routing.
|
||||
- **Action**: pins stay (explicitness, future models); main-loop skill effort + phase shifts carry the savings; A/B `/reconcile` high vs xhigh after rollout; context size = bigger lever, separate track.
|
||||
|
||||
Reference in New Issue
Block a user