diff --git a/.claude/memory/decisions.md b/.claude/memory/decisions.md index c4a99fa..ad77cd2 100644 --- a/.claude/memory/decisions.md +++ b/.claude/memory/decisions.md @@ -1076,3 +1076,6 @@ Old routing "Bug → investigate (bugfix if gstack off)" + gstack ON by default ### BDR-081 — Config recalibrated for Claude 5 family (Opus 5 dispatch tier) [accepted] (2026-07-30) Opus 5 (released 2026-07-24) now backs every `model: opus` pin (BDR-076/077) + any `/model opus` session. Research (official migration guide + web + registries): Opus 5 OVER-delegates (inverts LRN-030 Opus 4.8 trait that CLAUDE.global.md:43-47 compensated), self-verifies (explicit verify instructions → over-verification, "removing them reduces wasted tokens with no loss in quality"), literal following (conservative-reporting clauses depress recall; MUST/CRITICAL over-triggers), scope expansion = named regression, written deliverables +30-40%. Claude Code injects Opus-5-only anti-delegation prompt sections (heron_brook + subagent_steer_delegation, issue #80988, server-gated, no opt-out) — prose caps would triple-stack. Shipped: delegation block → model-neutral WHEN-guidance + explicit gates carve-out (verifier/security/challenge still dispatch as written); "staff engineer" self-check bar dropped; finish-whole-task clause folded into Deviations (gone-WRONG→STOP still wins); deliverable-length rule; design hook `\bux\b` dropped (`\bui\b` KEPT — 0 FP, 1 logged TP, lock-tested); plan-challenger grounded-doubt→[MINOR] in-place reword (grammar byte-identical). Plan challenged by 3 blind Opus 5 plan-challengers: correctness CONCERNS(4) / robustness FATAL(5, BLOCKER: all surfaces symlink-deployed LIVE — gates fire post-deployment) / simplicity CONCERNS(4); every fix adopted as prescribed (scratch-validation before live hook write, minimal diffs, ux-only, MINOR-routing). Alternatives rejected: leave as-is (nudge actively counter-productive); hard spawn caps in prose (harness injects one); confidence axis on challenger grammar (consumer unwired); dropping \bui\b (no evidence). NOT touched: verify-secure-loop + fresh gates (harness architecture BDR-049/050, ≠ model self-check prose); Security/Architecture sections (BDR-021); settings effortLevel xhigh (user pref — Opus 5 carry-over trap → LRN-139); superpowers plugin wording (external upstream). Plan+synthesis: .claude/tasks/plans/2026-07-30-opus5-config-tuning-1238.md. Branch feature/opus5-config-tuning, unmerged (human gate). + +### BDR-082 — seo/geo analyzers de-prescribed for Opus 5 (C1) [accepted] (2026-08-02) +BDR-081 N5 follow-on, user-directed apparatus (plan+3-lens challenge+census+dogfood). Method: audience×mode-range invariant — dedup ONLY verbatim same-audience (spec rule / bundle-item payload / phase-local caveat) same-mode-range repeats; cross-mode + agent↔dispatcher twins stay (standalone paths need them). Census-FIRST: lib/tests/seo-geo-contract.test.sh 71 locks (verdict grammar, sentinels, ALL STEP headers incl. interiors, item fields, score labels, envelope keys), flip-proven 7 mutations→7 FAILs, committed BEFORE reword. Shipped: self-output verification removed (":970 run twice"→conditional integrity guard; ":1217"→single-shot-scoped), 2 pre-BDR-061 vestigials fixed, caps softened (P0-rule/MANDATORY/ALWAYS→plain content rules), 2 essays compressed, checklist :1309→routing map rows verbatim (challenger caught it = routing table, NOT self-check), true same-range dups only (seo Handoff+landing-page blocks; geo ZERO — all claimed pairs distinct on inspection). FROZEN: guard-first url-guard orderings, :550 denominator-before-sampling (ordering IS the honesty mechanism), R2/NAP/COVERAGE/citation invariants, external-freshness checks (world drift ≠ self-verification). Deltas: seo 1528→1503 l ("P0 rule" 2→0, ALWAYS 1→0, MUST 5→4, NEVER 9→9 = class-B bans kept); geo 1106→1107 (MANDATORY 1→0, MUST 4→3). Plan challenged correctness FATAL / robustness FATAL(3 BLOCKER) / simplicity CONCERNS + confirmation FATAL(9) — every BLOCKER closed by named change (§5bis record). Dogfood before/after on frozen zenquality copy: judge-replay on frozen signals (zero collect variance) + templates + fresh collects + e2e judge + 42/42 assert battery BOTH sets + blind reader "interchangeable; all deltas = presentation variance both directions OR after MORE spec-conformant". Alternatives rejected: keyword dedup (challengers proved audience/range-blind — most annex "twins" were distinct obligations), FULL/aggressive dogfood (billing gate killed nested CLI; left as user option), banner/shape locks (LLM-convention layers wobble — lock strings only). Evidence: .audit/dogfood-baseline/ (18 artifacts + DOGFOOD-VERDICT.md), plan .claude/tasks/plans/2026-07-30-seo-geo-deprescription-1402.md. Branch feature/seo-geo-deprescription, UNMERGED (human gate). diff --git a/.claude/memory/journal.md b/.claude/memory/journal.md index d489fdb..e690acd 100644 --- a/.claude/memory/journal.md +++ b/.claude/memory/journal.md @@ -430,3 +430,6 @@ rules: ## 2026-07-30 - User: Opus 5 "needs more freedom" → analyse config + adapt. Research 3-agent (registries / config audit / web) + official migration guide: over-delegation (inverts LRN-030), over-verification, literal following, scope expansion, #80988 injections. Plan challenged 3 blind Opus 5 plan-challengers — robustness FATAL (BLOCKER: symlink-live deployment), all fixes adopted. Shipped: CLAUDE.global.md recalibrated (delegation when-guidance, staff-bar dropped, finish-whole-task, deliverable-length; 308/320), design hook \bux\b dropped flip-tested (22/0), plan-challenger grounded-doubt→[MINOR] (44/0). BDR-081 + LRN-139. feature/opus5-config-tuning, UNMERGED. + +## 2026-08-02 +- C1 seo/geo de-prescription EXECUTED end-to-end: census-first 71 locks flip-proven → reword under audience×range invariant (adafa35/c7646a9) → controlled dogfood (judge-replay frozen signals + templates + fresh collects + e2e + blind reader) → 42/42 both sets, zero contract regression, recall improved. Plan survived 4 challenge passes (2 FATAL + confirmation FATAL(9), all closed by name). BDR-082 + LRN-140. Nested-CLI dogfood died on monthly spend limit → inline pipeline (canonical /seo shape). feature/seo-geo-deprescription UNMERGED (human gate). Chantiers C2-C4 pending. diff --git a/.claude/memory/learnings.md b/.claude/memory/learnings.md index 1c27fa5..6a91f43 100644 --- a/.claude/memory/learnings.md +++ b/.claude/memory/learnings.md @@ -1361,3 +1361,10 @@ rules: - **Opus 5 traps found**: (a) Claude Code injects Opus-5-only anti-delegation prompt sections (heron_brook + subagent_steer_delegation, issue #80988; server-gated, no opt-out, absent from transcripts) — own prose stacks on top blindly; (b) NO model-default effort hold on Opus 5 — persisted effortLevel (xhigh, settings.json) silently carries over, against "start high, sweep low/medium"; run /effort sweep per model; (c) effort does NOT shorten visible output/deliverables — only prose length rules do (+30-40% docs). - **future application**: at every model-generation bump, grep config for trait-compensating language ("counters model tendency…", "default to X") and re-verify the premise; prefer WHEN-guidance (conditions where X pays) over directional nudges — survives inversions unchanged. - **link**: [[LRN-030]] [[BDR-081]]. + +## LRN-140 — de-prescription findings: dedup evaporates, self-verify is default, recall survives (2026-08-02) +- **pattern 1 — inventory dedup counts lie**: line-level inspection killed most "duplicate" pairs (seo 9 families→2 real merges; geo 7→0). Twins differ by AUDIENCE (bundle-item payload read by fresh applier vs spec rule) or MODE-RANGE (collect/judge/template/RULES) or are distinct obligations sharing a keyword (30/70 ×3 = three different rules). Dedup rule that survives: verbatim + same-audience + same-range ONLY. +- **pattern 2 — Opus 5 self-verifies unprompted**: "run it twice" instruction REMOVED → after-judge still ran score engine twice, identical output. Removing verify-prose does not remove the behavior; its value = no compounding, no contradiction burn. Confirms BDR-081 E3 mechanism, refines the payoff claim. +- **pattern 3 — de-prescription does NOT depress recall**: reworded collect caught  -encoded phone AT COLLECT (baseline collect missed it); reworded judge found new RGPD finding + self-caught false positive + corrected collect coverage claim 21/21→20/21. Integrity/honesty invariants (kept class B) carry the discipline, not the caps. +- **pattern 4 — lock strings, never shapes**: LLM-convention output layers (banners, fences, table columns, section order) wobble run-to-run in BOTH directions — baseline itself deviated from spec where after conformed (§0 ENTRIES, BUNDLE-before-SCORING). Stable contract = census-locked literal strings; anything unlocked drifts and MUST be tolerated by consumers (tier recognition "by intent" is the right pattern). +- **link**: [[BDR-082]] [[BDR-081]] [[LRN-139]] [[LRN-113]]. diff --git a/.claude/tasks/TODO.md b/.claude/tasks/TODO.md index 4de4c90..e60b70f 100644 --- a/.claude/tasks/TODO.md +++ b/.claude/tasks/TODO.md @@ -29,11 +29,18 @@ adopted as prescribed (plan §5bis, v2 items below). ## 2026-07-30 — Claude 5 follow-on chantiers (user directive, checkpoint between each) Order fixed, one branch per chantier, no merge without per-chantier signal. -- [ ] C1 dé-prescription seo-analyzer.md + geo-analyzer.md (opus pins → Opus 5): - separate machine-parsed contracts (fix-bundles, ownership matrices, - output formats — verbatim) from process choreography (MUST/MANDATORY on - "how" → when-guidance). Dedicated plan + 3-lens challenge, census tests - same commits, real /seo dogfood before/after (zero format regression). +- [x] C1 dé-prescription seo-analyzer.md + geo-analyzer.md — DONE 2026-08-02. + Census-first 71 locks flip-proven (9681b46) → rewords under + audience×range invariant (adafa35 seo, c7646a9 geo) → controlled + before/after dogfood: judge-replay on frozen signals + templates + + fresh collects + e2e judge + blind reader = 42/42 both sets, zero + contract regression, recall improved. Plan challenged 4 passes + (FATAL/FATAL/CONCERNS + confirmation FATAL(9), all closed by name). + BDR-082 + LRN-140. Evidence .audit/dogfood-baseline/ (19 artifacts). + Branch feature/seo-geo-deprescription UNMERGED — human gate. + Residual for gate: §6bis dynamically-unverified list (FULL branches, + apply path — census-locked statically); FULL/aggressive dry-run = user + option; nested-CLI dogfood blocked by monthly spend limit (inline used). - [ ] C2 self-contradiction audit CLAUDE.global.md + own skills: list rule pairs in tension, propose resolution per pair, apply after user OK. /doctor as assistant, not authority. diff --git a/CHANGELOG.md b/CHANGELOG.md index e8769b4..ad138db 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -7,6 +7,23 @@ Format follows [Keep a Changelog](https://keepachangelog.com/). ## [Unreleased] ### Changed +- **seo-analyzer + geo-analyzer de-prescribed for Opus 5 (BDR-082)** — + process choreography converted to when-guidance under an + audience×mode-range invariant; self-output verification demands removed + (the score-engine "run it twice" became a conditional integrity guard); + two pre-BDR-061 vestigial rules fixed; P0/MANDATORY/ALWAYS caps softened + to plain content rules. Machine contract byte-frozen and locked by the + new `lib/tests/seo-geo-contract.test.sh` census (71 locks, flip-proven); + proven by a controlled before/after `/seo` dogfood — judge replay on + frozen signals, 42/42 presence assertions on both runs, blind structural + reader: interchangeable, recall improved. + +### Added +- **`lib/tests/seo-geo-contract.test.sh`** — census locking the seo/geo + agent ⇄ dispatcher machine contract: judge verdict grammar, FIX BUNDLE + + READY-TO-APPLY sentinel, signals handoff, every STEP header (interiors + included), bundle item fields, score labels, scoring blocks, envelope + keys (46→71 assertions across the C1 chantier). - **Global instruction layer recalibrated for the Claude 5 family (BDR-081)** — delegation block is now model-neutral when-guidance (the Opus 4.8 under-delegation counter inverted on Opus 5, which over-delegates and gets