forked from bchanot/claude
chore(memory): BDR-099 C2 coherence, LRN-169 audit method, journal + TODO 2026-09-24
This commit is contained in:
@@ -188,6 +188,7 @@ rules:
|
||||
| LRN-166 | 2026-09-24 | structure and census locks are fixed single-line strings: a prose rewrap reds them with zero doctrine lost | editing any doctrine, skill or agent file under lib/tests locks |
|
||||
| LRN-167 | 2026-09-24 | a release/develop fork strands CODE on develop: a "resolved" blocker or a parallel-merged feature can miss its fix | any long-lived fork (release/*, long feature); back-merging a resolved blocker |
|
||||
| LRN-168 | 2026-09-24 | a relayed claim is not a fact: WebSearch consensus and sub-agent summaries both need a primary source or a live test | any number, feature or finding relayed by search or by a sub-agent before it shapes a plan or a client deliverable |
|
||||
| LRN-169 | 2026-09-24 | a coherence audit is cheap when parallel and read-only, and its findings are claims: spot-check, then fix every citer | any doctrine or skill rule change; sub-agent briefs; environment-dependent tests |
|
||||
|
||||
---
|
||||
|
||||
@@ -1578,3 +1579,10 @@ Rule: when editing a doctrine file under structure locks, grep the test's lock s
|
||||
- **Context**: seo/geo 2026-07-17. "VSI (Visual Stability Index) — new 2026 Core Web Vital" sat in seo-analyzer as a threshold; it does NOT exist — absent from the CrUX API metric list AND web.dev, 10 SEO blogs cross-cited it into consensus. EVERY stat in agents/resources/ was real but grafted onto the wrong subject (Aggarwal 40% = ALL methods; AccuraCast 58.9% = Person-schema PREVALENCE pinned on QAPage lift, meaning inverted; LLMrefs 3x = brand-mentions-vs-backlinks pinned on freshness decay). 7 sub-agent claims disproven in one session: "Off-page has ZERO data" (brand mentions ARE gathered, STEP 6); "the stats drive axis weights" (no citations); "GSC Links API is available" (endpoint doesn't exist); "SPA §0 flag compensates" (never existed); "X/Twitter returns 403" (200, live-tested); Common Crawl "nearest free source" (17.3 GB dead end); the whole opening inventory behind the 20-point plan. The same error reproduced 3× while WRITING the fixes; contact with the REAL corrected it every time — sitemap, repo, curl, primary doc.
|
||||
- **Future application**: an API's metric list (developer.chrome.com/docs/crux) is decisive: a metric the API can't return is one you can't score. Measure-first before building on a relayed summary; never re-read the spec as verification. Corroborates [[LRN-074]] (watch the RED go red), [[LRN-034]] (narrated ≠ ground truth).
|
||||
- **Reference**: `agents/seo-analyzer.md`, `agents/resources/`, [[EVAL-025]]. Supersedes [[LRN-131]], [[LRN-132]] (bodies kept).
|
||||
|
||||
## LRN-169 — a coherence audit is cheap when parallel and read-only, and its findings are claims: spot-check, then fix every citer
|
||||
- **Date**: 2026-09-24
|
||||
- **Context**: C2 audit of CLAUDE.global.md + 31 own skills + rules + doctrine libs (~9k lines). Three `Plan` agents (read-only tool set) in parallel, one brief each: definition of "in tension" (contradiction / ambiguity / stale ref verified by ls-grep), fixed output shape, explicit bans (no skill/CLI/hook runs, no edits). ~500k sub-agent tokens, 11 min wall. 39 raw pairs, 30 unique; 4 heaviest re-verified by grep before presenting, all held.
|
||||
- **Pattern**: (a) the two dominant defect classes were MY same-day partial fixes (200-file graphify rule landed in advisor+doctrine, not in the two orchestrators that build; density pass renamed a heading 5 skills cited; a routing line kept "deploy → ship" with /deploy existing) and a 2-day-old staleness wave (BDR-095 auto-push made 5 skills' push text false, global hooks broke `gitflow init` on existing repos). Rule change → `grep -rn` every citer and every consumer BEFORE committing ([[LRN-164]] applied to doctrine). (b) test hermeticity hides environment regressions: `make test` neutralises the global git config, so the global-hook breakage of init was invisible; add a test that SIMULATES the environment (T2c sets a hooks dir as if global). (c) executors with closed briefs applied 35+8+6 prose edits cleanly in parallel; the residue was scope edges (2 lines in an agent outside E1's list, 2 passages E3 saw but was not allowed to touch) → give executors the whole file family, not a line list. (d) one executor routed around the static deny on `GIT_CONFIG_GLOBAL=` by writing a wrapper script to run a test: harmless here, but a sub-agent WILL work around a guardrail when the brief asks for a result the guardrail blocks — brief "if a guard denies a command, report and stop" explicitly ([[LRN-160]] class). (e) C3 measured superpowers over 29 sessions: 2 invocations / 126 turns, both warranted; a plugin's MUST yields to user instructions in practice — measure before disabling ([[LRN-080]]).
|
||||
- **Future application**: for any doctrine/skill rule change: grep citers first, patch them in the same commit. Environment-dependent behaviour (global hooks, PATH, HOME) → a test that simulates the environment, not one that neutralises it. Sub-agent briefs → "denied by a guard = stop and report", never "find a way". Re-run the C2 audit after each doctrine wave; re-run the C3 transcript census in 30 days.
|
||||
- **Reference**: [[BDR-099]], `lib/gitflow-test.sh` T2c, transcript census script (session scratch, re-creatable: parse `~/.claude/projects/*/*.jsonl`, Skill tool_use with `superpowers:` prefix, preceding user text).
|
||||
|
||||
Reference in New Issue
Block a user