From 54a93eabe2ae4057d4b47eb18545e451659a6376 Mon Sep 17 00:00:00 2001 From: bastien Date: Tue, 29 Sep 2026 15:41:25 +0200 Subject: [PATCH] chore(memory): BDR-108 effort round, LRN-181/182, BLK-024 resync pins, EVAL-038 sub-agent thinking unmeasured, journal, TODO parked LOW --- .claude/memory/blockers.md | 7 +++++++ .claude/memory/decisions.md | 8 ++++++++ .claude/memory/evals.md | 8 ++++++++ .claude/memory/journal.md | 4 ++++ .claude/memory/learnings.md | 10 ++++++++++ .claude/tasks/TODO.md | 5 ++++- 6 files changed, 41 insertions(+), 1 deletion(-) diff --git a/.claude/memory/blockers.md b/.claude/memory/blockers.md index b2322c5..c762f06 100644 --- a/.claude/memory/blockers.md +++ b/.claude/memory/blockers.md @@ -43,6 +43,7 @@ rules: | BLK-021 | 2026-09-22 | Bash tool dead mid-session ("every command exits 1"): /tmp usrquota blown by a dead session's probe HOMEs — 2… | open | | BLK-022 | 2026-09-22 | `hooks/guard-bash.sh` withheld by the safety classifier; executable spec shipped instead — 2026-09-22 | open | | BLK-023 | 2026-09-28 | floor-guard SKIP pattern `xit(` (Jasmine) matches any `exit(` in python/JS test helpers → false ECARTS; workaround: no `exit(` in inline python, bash derives rc from output — 2026-09-28 | resolved | +| BLK-024 | 2026-09-29 | update-all.sh re-fetched vendored skills but never re-applied the effort pins (lost until next `make plugin`); my first fix placed the re-apply BEFORE the late 21st refresh — rtk-truncated grep read as complete — 2026-09-29 | resolved | --- @@ -268,3 +269,9 @@ rules: - **Real cause**: `lib/floor-guard.sh` SKIP_SUBSTRINGS holds the bare fragment `'xit('` to catch Jasmine's `xit(…)`; `skip_kind()` is a plain substring match, so `sys.exit(`, `SystemExit(`, `process.exit(` all hit. - **Solution**: workaround applied — the inline python prints violations only, the bash wrapper derives the return code from the captured output (no `exit(` anywhere). Root fix pending: word-bound the pattern (`(^|[^a-zA-Z_.])xit\(`) or match `xit(` only in JS/TS test files; hotfix-sized. - **Status**: resolved 2026-09-28 — hotfix 0deb559 (bugfix/floor-guard-xit-boundary): the four bare Jasmine identifiers moved into `SKIP_IDENT_RE` with lookbehind `(?)` always travels with the step's first tool call, shift first; pair a downward shift with a pinned executor or a Read/Bash, never with a built-in judgment dispatch; before any built-in judgment dispatch, pair the own-level shift with it; skills Claude loads alone do not apply their pin → re-assert with a paired shift at the resumed planning step ([[BDR-107]]). + +## LRN-181 — Stacked skills share one effort level; a lone load applies none +- **Context**: design toolchain loads 5-8 skills in one build. Skill `effort:` frontmatter = last loaded wins, both directions ([[LRN-179]]). Two levels inside the stack → effort depends on load order, invisible. Plus [[LRN-180]]: a Skill call Claude issues alone is a no-op. +- **Apply**: one level per stack (`lib/effort-pins.txt` design section, census `stack_levels` lock, site-motion frontmatter matches); doctrine "load the stack paired with the first Read of the target file"; new vendored design skill → copy the stack level. [[BDR-108]] + +## LRN-182 — Effort baselines are generation-bound; model aliases move silently +- **Context**: transcripts of the last weeks show `sonnet` → claude-sonnet-5 (3069 msgs) then claude-sonnet-5-5 (recent), `opus` → opus-5 then opus-5-5, `fable` → fable-5 then fable-5-1. API reference: Sonnet 5.5 recalibrated effort levels ("start at medium for agentic coding"). [[EVAL-036]] A/B ran on one generation. +- **Apply**: after an alias moves (new model in a tier) re-run `python3 lib/effort-audit.py` and re-read the pins; write the generation next to any effort figure; keep aliases (latest = cheapest or same price, never pin a version for a measurement). [[BDR-108]] diff --git a/.claude/tasks/TODO.md b/.claude/tasks/TODO.md index a43b26e..6ef2144 100644 --- a/.claude/tasks/TODO.md +++ b/.claude/tasks/TODO.md @@ -17,7 +17,10 @@ quality/price trade-off is tier × effort, never version. - [x] S5 doctrine: Design work paired load + one level per stack (CLAUDE.global.md, lib/effort-shift.md) - [x] S6 docs: README effort section, USAGE niveau d'effort, CHANGELOG - [x] S7 contract + GATE 0 + fresh verifier + security gate, make test, shellcheck — GATE 0 MET, verifier ECARTS(3) → executor moved the resync re-apply after the 21st refresh (real gap), scope gated, directive authorized → CONFORME 7/7; security PASS (4 LOW on the helper, see journal); make test 44 suites rc 0 -- [ ] S8 registries (BDR-108, LRN, BLK resync, EVAL correction) after user approval; journal +- [x] S8 registries BDR-108, LRN-181, LRN-182, BLK-024, EVAL-038 (user go) + journal +- [x] S9 hardening of lib/effort-pins.sh (4 LOW, user go): fresh executor, T11-T14, verifier CONFORME 9/9, security PASS +- [ ] parked LOW (security re-gate 2026-09-29, none exploitable): no RETURN trap on the mktemp sibling (SIGINT during awk leaves `SKILL.md.XXXXXX`); T13 never reaches the post-write re-read branch (CRLF opener fails `_effort_pin_closed` first, fixture with LF delimiters + CRLF `name:` line would); T14 fails under root (chmod ignored); `WORK="$(mktemp -d)"` unguarded in the suite (`|| exit 1`); install-plugins.sh `err()` uses `echo -e` on the rejected map line +- UNMERGED — human gate ("merge it") ## 2026-09-28 — effort tiering: session high, agent pins, skill levels, phase shifts (feature/effort-tiering) Spec `docs/superpowers/specs/2026-09-28-effort-tiering-design.md`, plan