chore(memory): BDR-108 effort round, LRN-181/182, BLK-024 resync pins, EVAL-038 sub-agent thinking unmeasured, journal, TODO parked LOW

This commit is contained in:
bastien
2026-09-29 15:41:25 +02:00
parent bb46ee22eb
commit 54a93eabe2
6 changed files with 41 additions and 1 deletions
+8
View File
@@ -58,6 +58,7 @@ rules:
| EVAL-035 | 2026-09-28 | thinking-share measurement, 6 days of transcripts (10,955 requests): thinking = 8 % of weighted spend, 97 % of it in the main loop; sonnet subagents at xhigh think 26 tok/request; cache reads = 53 % | pins = explicitness not savings; main-loop effort + context size are the levers; A/B after rollout |
| EVAL-036 | 2026-09-28 | A/B `/reconcile` headless, session high vs skill entry low: requests 18→15, output 12374→9038 (−27 %), thinking 3135→2248 (−28 %), time 96.5→78.4 s (−19 %), n=1 | keep low on bookkeeping skills; repeat on a reflection skill before touching the medium/high split |
| EVAL-037 | 2026-09-28 | correction of EVAL-035/036 counts: transcript records are per content block; deduped by message.id → main-loop thinking share 99.9%, thinking share of weighted cost 5.6%, sonnet think/msg 26→0.2, A/B requests 9→8 | conclusions hold (sharper: main-loop thinking 96.6%→99.9%, weighted-cost thinking corrected 8.4%→5.6%); effort-audit.py dedupes from a3b479e+ |
| EVAL-038 | 2026-09-29 | correction of EVAL-037: 94 % of sub-agent usage records carry no `output_tokens_details` (Fable subs at xhigh read 0 thinking, impossible with always-on thinking) → sub-agent thinking UNMEASURED, not ≈0; main loop 100 % counted; weighted-cost split (61/39) still holds | `effort-audit.py` prints coverage + CAVEAT; cite the cost split only; agent effort pins stay unmeasured; a tier move on a price argument = judgment, not figure |
---
@@ -354,3 +355,10 @@ Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itse
- **Result (deduped)**: main weighted-cost 61.4 %, thinking share 99.9 % (was 96.6 %); sub weighted-cost 38.6 %, thinking share 0.1 %; thinking = 5.6 % of weighted cost (was 8.4 %, inflated by duplicate counting); sonnet think/request 26→0.2 tok (sub, xhigh); A/B `/reconcile` (EVAL-036 rerun, deduped) requests 9→8, output 6129→4706, thinking 1550→1104 — the raw undeduped counts on the same transcripts are 18→15, matching EVAL-036 exactly (the bug, not the finding).
- **Anomaly**: the main-loop-carries-almost-all-thinking split got SHARPER after dedup (96.6→99.9 %), not weaker — duplication was near-uniform across content blocks, so ratios among scopes barely moved; only the absolute request/token counts and the overall thinking-share-of-cost figure were inflated (~2.2-2.8× depending on transcript mix).
- **Action**: `lib/effort-audit.py` dedupes by `message.id` from this commit; cite EVAL-037, not EVAL-035, for the split.
## EVAL-038 — correction of EVAL-037: sub-agent thinking is unmeasured, not ≈0
- **Date**: 2026-09-29
- **Output checked**: [[EVAL-037]] "sonnet think/request 26→0.2 tok, executors stay cheap; main loop carries 99.9 % of thinking".
- **Method**: scan of the last 400 transcripts, dedup by message.id, count records with/without `output_tokens_details`: sub 4586 requests, 6 % carry the field (2896/3075 sonnet-5 without, 65/72 fable-5-1 without); main 100 % carry it. A Fable 5.1 sub-agent at xhigh with 0 thinking tokens is impossible (thinking always on) → recording gap, not behaviour.
- **Anomaly**: "main loop = 99.9 % of thinking" is a coverage artefact. The weighted-cost split (main 61 % / sub 39 %) holds: `output_tokens` is always present.
- **Action**: `lib/effort-audit.py` counts `nodet`, prints `%counted` per row, "thinking counted on N% of them" per scope and a CAVEAT under 50 %; cite the cost split only; the 20 agent effort pins ([[BDR-107]]) remain unmeasured; a tier move argued on price stays a judgment ([[BDR-108]]).