diff --git a/.claude/audits/DARWIN-2026-08-26-card.png b/.claude/audits/DARWIN-2026-08-26-card.png new file mode 100644 index 0000000..6d43293 Binary files /dev/null and b/.claude/audits/DARWIN-2026-08-26-card.png differ diff --git a/.claude/audits/DARWIN-2026-08-26.md b/.claude/audits/DARWIN-2026-08-26.md new file mode 100644 index 0000000..c07c13e --- /dev/null +++ b/.claude/audits/DARWIN-2026-08-26.md @@ -0,0 +1,90 @@ +# Darwin run 2026-08-25/26: fresh baseline + threshold optimization + bug pass + +Branch `feature/darwin-optimize-20260825`, 26 commits, 39 files, +299/-142. +Log: `~/.agents/skills/darwin-skill/results.tsv` (fresh, the May file was wiped +by the 2026-06-23 reinstall). Method: darwin v2.1. Absolute scores served as +triage only; every keep/revert decision came from a paired same-judge majority +(3 judges per round, before/after read in one call). + +## Scope + +54 units: 31 personal skill-systems (SKILL.md + dispatched agents judged +together, per EVAL-004) and 23 agents. Excluded: gstack/external symlinks +(BDR-015/043, LRN-070), darwin-skill itself (BDR-058 pin), and find-docs, +newly identified as machine-owned ctx7 output (gitignored, installer-written). + +## Baseline (7 blind judges, dims scored 1-10, totals recomputed main-thread per LRN-018) + +Mean 83.4 (skills 83.5, agents 83.3). Best: deploy, release-candidate, +release-executor (90.4). Worst: skills-perso 63.5. All dim8 rows marked +dry_run by design; live execution happened later, inside the paired rounds. +13 units scored below the user-set threshold of 80. + +## Phase 2: threshold loop, 13/13 units, 0 reverts + +Every round was validated by 3 paired judges (neutral, skeptic, realism). +All verdicts 3-0 better. + +| Unit (baseline) | Round(s) | What changed | +|---|---|---| +| skills-perso (63.5) | d8 | Detection rebuilt on the link.sh convention: symlink = external, real dir = personal, gitignored = machine-generated. Live result 8/31 to 31/31, zero false positives | +| interviewer (70.9) | d3, d9 | Failure-mode table (vague, "you decide", contradiction, partial, balloon) + 2-round budget; DO-NOT list | +| onboarder (71.5) | d8 | BRIEF contract split REQUIRED/OPTIONAL; null enrichment becomes TODO placeholders; STOP kept for required keys and unresolved monorepo. Kills the guaranteed first-dispatch bounce vs /onboard STEP 2 | +| pdf-translate (72.3) | d3/d8 | 8-row failure table: deps, >30 pages gate, zero-output, illisible markers, design-html/browse fallbacks, QA cap 2, stale workdir | +| refactor (75.6) + refactorer (76.8) | d4/d3 | No-tests STOP gate + GO-WITHOUT-TESTS arbitration in the dispatcher; mid-run test-failure revert protocol; code-cleaner inline carve-out | +| profile (77.3) | d3 | 6-row failure table, every row fact-checked against profile.sh (rc=1 paths, partial toggle, split plugin leg, BLK-006 contradiction); fixture de-drift | +| plugin-probe (78.5) + plugin-advisor (77.5) | d8 | FRAMEWORK-DEPS now exact dep@version (preact false-hit killed, fallback actually fires; the old `\|\| true` silently emitted nothing and tripped the advisor's fail-closed path on non-Node projects); frontend/fast-libs derivable; PLAN echoed-or-unknown, invention removed | +| analyze (77.7) + analyzer (78.0) | d1, d2 | Bilingual triggers + fix-wanted disambiguator; TASKS ordered, each step mapped to its OUTPUT section | +| status-reporter (78.0) | d5 x2 | Fabrication-forcing token field replaced, then restored producibly from doctor.sh constants (a skeptic judge found the source); dead ROADMAP row rewritten post-ADR-013 | +| gitflow (78.4) | d3 | 7-row failure table keyed to lib return codes; rc=4 conflict resume empirically verified; human merge gate untouched | + +## Bug pass: verified defects in above-threshold units, 8 commits, all kept 3-0 + +- hotfix: `git restore .` on every failure branch wiped tolerated in-progress + user edits. Now: `git stash create` pre-flight snapshot + file-scoped + restore + fresh-dispatch-only security gate. Two skeptic residuals amended + (RULES bullet, FILE(S) new-file marker). +- init-project: allowed-tools lacked Agent and Skill while every step + dispatches. commit-change: conflict grep now covers all 7 unmerged codes. + tour: --report-only no longer commits (could land on develop). +- harden: severity rule now defers to the calibrated guide; the late SSL Labs + grade has an assigned actor. +- plan-challenger: ERROR joined the load-bearing verdict grammar. +- handover writers: stale chapter refs corrected (glossary/tone to §6, + cross-links and THRESHOLD-OVERRIDE to §5); STEP 14.5 verification deferred + post-write; anchor gate ordered into STEP 16. +- security-auditor: /hotfix no-verifier carve-out documented. close: STEP 5C + enumerated, --no-push passthrough added. +- prune-memory: false "v1-untested" note replaced by the real tests/ state. + code-clean: executor attribution corrected (code-cleaner, refactorer inline). +- Fixtures de-drifted: plugin-check (PLUGIN CHECK block, real plugin names), + onboard (nextjs-app-router). + +`make test` green (0 RED, rc=0) after one census rewrap: a locked phrase had +been line-wrapped and the single-line grep lock caught it. + +## Residual findings, logged not fixed + +- analyze triggers: "how does X work" brushes graphify's territory; graphify's + graph-exists routing still wins. +- pdf-translate: pdfinfo row assumes poppler (fitz also has page count); "GB" + slightly overstated near the 30-page gate. +- web-validate: .validate-cache mkdir lives in a skipped STEP 0 + (self-recoverable); axis budgets 35/25/40 never reconciled with the base-100 + deduction table. seo/geo minor wording items. verifier/doc-syncer/audit-delta + restatement redundancy (cosmetic). handover-doc-writer STEP 14.5 umbrella + line still says "BEFORE STEP 15" while the inner note overrides it. +- bugfix trivial-fast-path boundary loosely defined; feat prompt-3 expectation + vs full gate pipeline. + +## Methodology notes + +- v2.1 paired majority produced 36 unit-round verdicts and 24 batch verdicts, + all better, 0 reverts, 0 ties. The May-2026 run under absolute-delta scoring + had reverted 2 edits on judge noise; this run had no such event. +- Judges live-executed wherever the artifact was executable (skills-perso + detection, profile.sh probes, plugin grep on scratch manifests, doctor.sh + grep, git merge no-op resume). Behavior outranked prose in 5 units. +- Two grep-exit-masking bugs surfaced (a `head` pipe swallowing the fallback's + trigger), one in the probe being fixed, one in this run's own test harness. + The pattern is worth a learning entry. diff --git a/.claude/memory/blockers.md b/.claude/memory/blockers.md index 77b4478..21dfa9c 100644 --- a/.claude/memory/blockers.md +++ b/.claude/memory/blockers.md @@ -215,3 +215,30 @@ rules: - **Solution** (workaround): dispatcher ran `gitflow.sh finish` + tag inline after its own human gate — where the signal is real. Release completed clean (main `648bc6e`, tag v1.3.1). - **Status**: open. Candidate fixes: (a) quote gate evidence verbatim in span prompt — untested vs classifier; (b) move finish+tag span permanently inline in /release-candidate — keeps prep span dispatched, costs the sonnet pin on ~5 mechanical commands, cheap; (c) permission rule allowing subagent `gitflow.sh finish` — weakens the guard, refused. Decide at next release. - **Reference**: skill `release-candidate` STEP 5. Pattern adjacent [[LRN-089]] (ambient-state/context assumptions across boundaries). Journal 2026-07-20. + +## BLK-019 — notify-attention bell silent, toast OK (VS Code client default) — 2026-09-01 +- **Friction**: hook fired, Windows toast arrived, native bell never audible. User heard only Windows toast sound. Looked like half-broken hook. +- **Real cause**: not hook. Toast proves full `terminalSequence` reached terminal, `\a\a` sits at head of that same string → BEL emitted. VS Code defaults `accessibility.signals.terminalBell` to `"auto"` = sound OFF unless screen reader active. +- **Solution**: `"accessibility.signals.terminalBell": { "sound": "on" }` in CLIENT-side user settings.json (`c:/Users//AppData/Roaming/Code/User/`). Unreachable from remote: real SSH remote, not WSL (no `/mnt/c`, `/proc/version` no Microsoft). User applied, retest → both channels OK. +- **Status**: resolved (per-client-machine, not repo-portable). +- **Reference**: `~/.claude/hooks/notify-attention.sh` header already documented the setting; never applied. New client machine → bell mute again while toast works. Silent-degradation class [[LRN-047]]. + +## BLK-020 — notify-attention: both channels dead on one VS Code client — 2026-09-02 +- **Friction**: client-side prereqs applied on Windows box (ext `wenbopan.vscode-terminal-osc-notifier` + `accessibility.signals.terminalBell` sound:on), window reloaded. AskUserQuestion → nothing. `idle_prompt` 90s wait → nothing. Direct write `\a\a` + OSC 777 to claude own pty (`/dev/pts/2`) → nothing. Second client machine, same SSH server, same hook, same registries → both channels OK. +- **Server side cleared**: hook dry-run emits `BELx2 + OSC 777 + ST` correctly, `jq` present, matcher covers `idle_prompt`, ext NOT wrongly installed remote-side. Not a hook bug — same class as [[BLK-019]] (client default silently degrades). +- **Real cause**: unresolved. Facts: claude runs under `dtach -c ~/.dtach/claude-190012`; claude fd1 = `/dev/pts/2` (inner pty, dtach master side), REAL VS Code terminal = `/dev/pts/1` held by dtach client pid 742794. `VSCODE_SHELL_INTEGRATION` unset this terminal; ext marketplace doc requires shell integration ON. BUT other working session (`claude-154323`) also runs under dtach → dtach alone insufficient explanation, weight shifts back to client-side. +- **Probes run**: direct write to `/dev/pts/1` (real VS Code pty, chain alive: bash pts/1 → dtach client 742794 S+ → master → claude pts/2) → no bell, no toast. Visible-marker injection both paths → user saw neither, BUT inconclusive: claude TUI repaints, injected text clobbered next frame. Only BEL is repaint-proof, and BEL stays silent. +- **Client settings verified by user**: settings.json path correct (no VS Code profile indirection), `terminalBell` sound on, ext installed + enabled local side. VS Code recent (server dirs 2026-08), so ≥ 1.93 ext requirement met. +- **Next probe**: user opens FRESH VS Code integrated terminal (no dtach, no claude TUI) and runs `printf '\a\a\033]777;notify;Test;hello\033\\'`. Isolates client renderer from claude/dtach path. Beep+toast there → fault in claude/dtach path; nothing → client-side, diff against working machine. +- **Fresh-terminal probe (decisive)**: user ran `printf '\a\a\033]777;notify;Test;hello\033\\'` in NEW VS Code terminal → toast OK, bell still silent. Splits one symptom into TWO independent faults. +- **Fault A (toast in claude session)**: ext parses only terminals created AFTER its activation. Claude terminal pts/1 born 19:00, ext installed later same day → that terminal never hooked. Fix: restart claude in fresh terminal, or re-attach existing dtach session from one (`dtach -a ~/.dtach/`; dtach broadcasts to multiple clients, no session loss). NOT a dtach filtering bug — earlier hypothesis wrong. +- **Fault B (bell)**: silent even in fresh terminal where toast works → not terminal path, VS Code audio side. Toast sound = Windows notification (works); bell = VS Code process audio (mute). Suspects: signal volume option, Windows volume mixer entry for Code, output device. Probe: palette `Help: List Signal Sounds` → Terminal Bell plays preview or not. +- **Fault A RESOLVED (verified 2026-09-02)**: re-attached session from fresh terminal (`dtach -a ~/.dtach/claude-190012`, new client pts/3). Both sends toasted — one through session path (pts/2, dtach broadcast), one direct. Rule: ext hooks only terminals born AFTER its activation → install ext, THEN start/re-attach claude session. dtach broadcast means zero session loss. +- **Fault B still open**: bell silent on every path. New signal: toasts arrive but user reports NO sound at all, while [[BLK-019]] machine got audible Windows toast sound. Both audio channels dead + both visual channels fine → common factor is client audio output, not terminal stream. Suspects ranked: Windows per-app notification sound off for Code, system/app volume mixer mute, wrong output device, `accessibility.signalOptions.volume` 0. +- **Fault B ROOT CAUSE ISOLATED (2026-09-02)**: palette `Help: List Signal Sounds` → Terminal Bell preview plays NO sound, while Windows toast sound IS audible. Preview bypasses terminal, BEL, hook, dtach, ext entirely → VS Code renderer audio itself mute on this box. Toast sound emitted by Windows shell, not by Code → explains why one audio channel works and other does not. +- **Fix candidates (client, ranked)**: (1) Windows per-app volume mixer — Code muted/0, or per-app OUTPUT DEVICE pointing at disconnected device (mixer only lists app after it attempts playback → hit preview first, then open mixer); (2) VS Code `accessibility.signalOptions.volume` = 0 kills all signals; (3) compare both against working machine. +- **Pragmatic out**: toast already carries audible Windows sound → attention signal functional without bell. Bell is redundant channel, not blocker. +- **Fault B RESOLVED (2026-09-03)**: cause = Windows per-app volume mixer, Code entry at 0. Toast audible throughout because Windows shell emits that sound, not Code → masked a plain app-volume mute. User set volume up → bell audible. +- **Status**: resolved (A: ext hooks only terminals born after activation → install ext THEN start/re-attach session; B: Code app volume 0 in Windows mixer). +- **Lesson**: two independent client faults presented as one symptom ("nothing works"). Splitting probe = run signal in FRESH terminal + play VS Code's own sound preview. Preview bypasses terminal/BEL/hook/dtach/ext → isolates renderer audio in one step. Do that FIRST next time, before any server-side archaeology. +- **Reference**: [[BLK-019]] bell-only variant (resolved differently — setting alone insufficient here), [[LRN-145]] terminalSequence-not-/dev/tty pattern. Silent-degradation class [[LRN-047]]. diff --git a/.claude/memory/decisions.md b/.claude/memory/decisions.md index 15c51ba..f1c0bd5 100644 --- a/.claude/memory/decisions.md +++ b/.claude/memory/decisions.md @@ -92,6 +92,11 @@ rules: | BDR-072 | 2026-07-17 | SPA: honest refuse (On-page N/A, not zero), no headless browser (R2 over R1) | accepted | | BDR-073 | 2026-07-17 | Scoring: LLM judges findings+severity, engine does the arithmetic (deterministic /20) | accepted | | BDR-080 | 2026-07-21 | Bug routing inverted: /bugfix primary, /investigate explicit-only | accepted | +| BDR-083 | 2026-08-24 | Contract gates: deterministic floor (GATE 0) under the fresh verifier | accepted | +| BDR-084 | 2026-08-24 | /tour multi-project: parallel runners (LRN-083 derogation, bounded), runner inherits session model | accepted | +| BDR-085 | 2026-08-25 | User permanent rules: writing-style always-on in rules/, web build+security path-scoped | accepted | +| BDR-086 | 2026-08-26 | darwin: threshold gates full loops; verified defects fixed regardless of unit score (paired-validated, batched checkpoint) | accepted | +| BDR-087 | 2026-09-03 | Stop hook = attention signal only, never control flow; one script for Notification + Stop | accepted | --- @@ -1073,3 +1078,48 @@ Audit (user ask "profile toggles externals both ways?"): ASYMMETRIC. Enable side ### BDR-080 — bug routing inverted: /bugfix primary, /investigate explicit-only [accepted] (2026-07-21) Old routing "Bug → investigate (bugfix if gstack off)" + gstack ON by default → every bug took path bypassing own quality pipeline (gitflow aiguillage, contract, fresh verifier + security gates, doc-sync, `.claude/memory` registries) — /bugfix relegated to near-never fallback. Skill comparison: same core doctrine (root-cause iron law, hypothesis loop, regression test, 3-strike stop, >5-files alert) but incompatible wrappers — investigate monolithic (same context investigates+fixes+verifies, ~1075-line SKILL.md w/ gstack preamble/telemetry/onboarding, capitalizes to `~/.gstack` learnings.jsonl framework never reads at session start); bugfix orchestrator (reflection inline, sonnet bugfixer executor, fresh gates — BDR-066, LRN-083). Composition rejected: skills superpose in context, don't compose — invoking investigate inside bugfix = two full workflows, two completion protocols, two memory systems loaded at once. Decision: CLAUDE.global.md routing line inverted — bugfix primary; investigate ONLY on explicit ask for gstack ecosystem (cross-project learnings, /freeze scope lock, long no-commit investigation). Alternatives rejected: keep investigate primary (bypasses framework), embed investigate inside bugfix (context conflict, dual memory). Known drift noted at write time: Index table rows BDR-074..079 missing (pre-existing, /prune-memory scope). + +### BDR-081 — Config recalibrated for Claude 5 family (Opus 5 dispatch tier) [accepted] (2026-07-30) +Opus 5 (released 2026-07-24) now backs every `model: opus` pin (BDR-076/077) + any `/model opus` session. Research (official migration guide + web + registries): Opus 5 OVER-delegates (inverts LRN-030 Opus 4.8 trait that CLAUDE.global.md:43-47 compensated), self-verifies (explicit verify instructions → over-verification, "removing them reduces wasted tokens with no loss in quality"), literal following (conservative-reporting clauses depress recall; MUST/CRITICAL over-triggers), scope expansion = named regression, written deliverables +30-40%. Claude Code injects Opus-5-only anti-delegation prompt sections (heron_brook + subagent_steer_delegation, issue #80988, server-gated, no opt-out) — prose caps would triple-stack. Shipped: delegation block → model-neutral WHEN-guidance + explicit gates carve-out (verifier/security/challenge still dispatch as written); "staff engineer" self-check bar dropped; finish-whole-task clause folded into Deviations (gone-WRONG→STOP still wins); deliverable-length rule; design hook `\bux\b` dropped (`\bui\b` KEPT — 0 FP, 1 logged TP, lock-tested); plan-challenger grounded-doubt→[MINOR] in-place reword (grammar byte-identical). Plan challenged by 3 blind Opus 5 plan-challengers: correctness CONCERNS(4) / robustness FATAL(5, BLOCKER: all surfaces symlink-deployed LIVE — gates fire post-deployment) / simplicity CONCERNS(4); every fix adopted as prescribed (scratch-validation before live hook write, minimal diffs, ux-only, MINOR-routing). Alternatives rejected: leave as-is (nudge actively counter-productive); hard spawn caps in prose (harness injects one); confidence axis on challenger grammar (consumer unwired); dropping \bui\b (no evidence). NOT touched: verify-secure-loop + fresh gates (harness architecture BDR-049/050, ≠ model self-check prose); Security/Architecture sections (BDR-021); settings effortLevel xhigh (user pref — Opus 5 carry-over trap → LRN-139); superpowers plugin wording (external upstream). Plan+synthesis: .claude/tasks/plans/2026-07-30-opus5-config-tuning-1238.md. Branch feature/opus5-config-tuning, unmerged (human gate). + +### BDR-082 — seo/geo analyzers de-prescribed for Opus 5 (C1) [accepted] (2026-08-02) +BDR-081 N5 follow-on, user-directed apparatus (plan+3-lens challenge+census+dogfood). Method: audience×mode-range invariant — dedup ONLY verbatim same-audience (spec rule / bundle-item payload / phase-local caveat) same-mode-range repeats; cross-mode + agent↔dispatcher twins stay (standalone paths need them). Census-FIRST: lib/tests/seo-geo-contract.test.sh 71 locks (verdict grammar, sentinels, ALL STEP headers incl. interiors, item fields, score labels, envelope keys), flip-proven 7 mutations→7 FAILs, committed BEFORE reword. Shipped: self-output verification removed (":970 run twice"→conditional integrity guard; ":1217"→single-shot-scoped), 2 pre-BDR-061 vestigials fixed, caps softened (P0-rule/MANDATORY/ALWAYS→plain content rules), 2 essays compressed, checklist :1309→routing map rows verbatim (challenger caught it = routing table, NOT self-check), true same-range dups only (seo Handoff+landing-page blocks; geo ZERO — all claimed pairs distinct on inspection). FROZEN: guard-first url-guard orderings, :550 denominator-before-sampling (ordering IS the honesty mechanism), R2/NAP/COVERAGE/citation invariants, external-freshness checks (world drift ≠ self-verification). Deltas: seo 1528→1503 l ("P0 rule" 2→0, ALWAYS 1→0, MUST 5→4, NEVER 9→9 = class-B bans kept); geo 1106→1107 (MANDATORY 1→0, MUST 4→3). Plan challenged correctness FATAL / robustness FATAL(3 BLOCKER) / simplicity CONCERNS + confirmation FATAL(9) — every BLOCKER closed by named change (§5bis record). Dogfood before/after on frozen zenquality copy: judge-replay on frozen signals (zero collect variance) + templates + fresh collects + e2e judge + 42/42 assert battery BOTH sets + blind reader "interchangeable; all deltas = presentation variance both directions OR after MORE spec-conformant". Alternatives rejected: keyword dedup (challengers proved audience/range-blind — most annex "twins" were distinct obligations), FULL/aggressive dogfood (billing gate killed nested CLI; left as user option), banner/shape locks (LLM-convention layers wobble — lock strings only). Evidence: .audit/dogfood-baseline/ (18 artifacts + DOGFOOD-VERDICT.md), plan .claude/tasks/plans/2026-07-30-seo-geo-deprescription-1402.md. Branch feature/seo-geo-deprescription, UNMERGED (human gate). + +### BDR-083 — contract gates: deterministic floor (GATE 0) under the verifier [accepted] (2026-08-24) +User asked what to take from `unlazy` skill (Leonxlnx/unlazy 2.1.0, MIT). Verdict on its verification ARCHITECTURE: teaches nothing we lack — contract + fresh blind verifier + bounded loops + order invariant already shipped (BDR-049/050/066, LRN-083). Real gap found elsewhere: between executor and GATE 1, NO deterministic floor. GATE 1 = LLM dispatch; verifier's mandatory `PROOF:` line = a line the verifier WRITES — nothing structurally stops it being produced without executing anything (LRN-048 demands a pass prove it looked; the proof is self-reported prose). Decision: import unlazy's gate ledger INTO the existing contract, never alongside it. Palier 2, user-chosen over doctrine-only / defer. +TAKEN: criterion carries an oracle (indented `CHECK:` cmd + `EXPECT:` success-only marker + `EVIDENCE:` slot); fail-closed = exit 0 AND marker (a nonzero process never passes because its error text carries the token); evidence persisted INTO the contract → the fresh verifier reads fact, not the executor's report; `ABANDON: ` = impossible criterion never deleted, blocks CONFORME, routes to human gate (new verdict token `ABANDONED(n)` — distinct routing from ECARTS ⇒ distinct token, not a sub-line to re-derive); 4 gate-authoring rules (observe the named artifact / success-only marker / positive control before any absence check / recompute supplied numbers, never copy one into EXPECT); 4-pass executor discipline (feater full; bugfixer narrowed to fix+test under "keep the fix minimal", pass 3 = negative control proving the regression test fails without the fix). +REFUSED + why: Stop hook `decision:"block"` — contradicts "STOP + human escalation", "gone WRONG → STOP re-plan", "merge only on explicit human signal"; a hook FORCING continuation is the inverse of our gates; its 6-block release either traps the session or gives up; each block = an agent continuation = real tokens. Approval store `~/.unlazy/approved` (binds ledger+cmd+CWD+shell+timeout+platform+full PATH) — exists to execute ledgers INHERITED from untrusted repos; our contracts are authored by our own orchestrator in our own repo ⇒ biggest chunk of their 28k checker closes zero threat here. `.unlazy//` tree (PLAN+GATES+gates/+status.log+session+hook-state+locks/) — a 4th bookkeeping tree beside .claude/tasks/{contracts,plans} + memory/ + audits/. `tree N` Depth-Tree effort arithmetic — disowned by unlazy's OWN research/validation-protocol.md (v1 six-run figures unreproducible), while the repo DESCRIPTION still advertises the retracted claim. Node checker (28k .mjs + 54k .mjs tests) — lib stack is 100% bash, Health Stack = `shellcheck *.sh hooks/*.sh lib/*.sh` would cover none of it. `OWNS:` ownership leases — deferred (Palier 3): our parallel dispatches (seo/geo, 3 plan-challengers) are read-only, the write-collision problem does not exist yet. +Shipped: lib/gates.sh (~250 l bash; `status` never executes and never writes · `run` ALWAYS re-executes every runnable criterion — trusting written evidence is the failure being closed, so there is no incremental mode to get wrong; rc 0 MET / 2 UNMET|malformed / 3 ABANDONED; parse fails closed on partial oracle, duplicate id, unindented attribute, runnable-without-EVIDENCE, and executes nothing when the ledger is malformed). GATE 0 in lib/verify-secure-loop.md (red floor → executor re-dispatch with the NOT-MET rows, NO verifier spawned; own 3-iteration budget, separate from conformity; malformed ledger fixed in the main loop, never dispatched to a dev). Order invariant now GATE 0→1→2 on every re-loop. lib/contract-interview.md: ORACLES section + template + ABANDONMENT lifecycle + per-flow oracle weight. agents/verifier.md: oracle-consumption rules — a red or unrun oracle is NEVER overridden by reading code; a MET oracle proves the ORACLE, not the English sentence ⇒ vacuous oracle = NOT-MET, the one judgement no command can make; verifier may re-run a CHECK but never edits the contract. lib/tests/gates.test.sh 64 assertions (sentinel-proved non-execution, with its own positive control asserted first). +Alternatives rejected: Palier 1 doctrine-only (CHECK:/EXPECT: become decorative without an executant); port the Node checker (stack break, shellcheck-blind); fold ABANDONED into ECARTS (would send a dev to fix the impossible and eat the 3-iteration budget); `status` revalidating old evidence (that trust is the failure being closed). +Branch feature/contract-gates, UNMERGED (human gate). `make test` rc 0, shellcheck clean, e2e verified on a real contract in the documented template. + +### BDR-084 — /tour multi-project: parallel runners, bounded LRN-083 derogation [accepted] (2026-08-24) +User asked whether agent parallelism on independent tasks is ACTIVE. Measured first (LRN-080): (a) mechanics — nested probe, 1 dispatched orchestrator fanned 3 sub-agents, execution windows all overlap, 9.1s vs ~18s sequential ⇒ nested parallel dispatch WORKS; (b) doctrine — already prescribed at 3 layers (harness "single message" injection; /seo, challenge-plan, /cso, graphify explicit same-message mandates; graphify even anti-sequential wording); remaining serializations all MOTIVATED (audit-delta crash-resilience documented, verify-secure-loop order invariant); (c) behavior — probe orchestrator batched spontaneously without being told "parallel" (N=1), this session fanned 8+7 agents/message during the RED. Conclusion: nothing to add globally — a CLAUDE.md "parallelize" line would duplicate-stack the harness injection (BDR-081 anti-pattern). +ONE real sequential-but-independent candidate: /tour multi-project (independent repos, one by one, no documented reason). User gate: option "tout paralléliser" chosen over report-only-only and no-change, WITH the model invariant "orchestrateur garde le modèle orchestrateur; skills/agents suivent leurs orchestrateurs définis". +Decision: STEP 0 routes (1 project = inline unchanged; ≥2 = STEP 0b fan-out). One general-purpose runner per project, ALL in ONE message, dispatched with NO model override — inherits the session model (model-gate already validated big; a runner carries tour's reflection: fix decisions, convergence). Inside a runner every agent keeps its defined tier (security-auditor sonnet, Phase B opus, doc-syncer sonnet two-mode). Dead/mute runner ⇒ explicit `RUNNER FAILED` summary row (mute is never a pass). Capitalize offer stays MAIN LOOP ONLY (registries = shared state). +LRN-083 derogation, bounded: per-project fix loop + convergence now run INSIDE the dispatched runner. Bounded because nothing a runner decides touches shared state — independent repos, per-repo chore branches, branches stay UNMERGED for human review exactly as inline (report-as-approval-gate design unchanged). Precedent: client-handover-writer already a dispatched orchestrator running parallel audit loops (BDR-077). +Alternatives rejected: report-only-only parallel (my recommendation — user overrode: full parallel wanted); one sub-orchestrator agent .md file (drift risk vs SKILL.md, the runner reads the skill from disk instead — client-handover→/seo precedent); pinning the runner (would put tour reflection on an executor tier — inverts BDR-076); global CLAUDE.md parallelism line (duplicate of harness injection). Census §12: 6 locks (fan-out present, no-pin, single-message, capitalize main-loop, RUNNER FAILED, no pinned runner), flip-tested. Branch feature/tour-parallel, UNMERGED (human gate). + +### BDR-085 — user permanent rules: writing-style always-on in rules/, web rules path-scoped [accepted] (2026-08-25) +User supplied 4-block permanent rule text (writing / website / code security / self-check), asked: coverage check, conflict check, integrate. Coverage verdict: security CORE (parameterized queries, input validation, env-var secrets, AuthN/AuthZ split + default deny, no stack traces, fail closed, least privilege) ALREADY in CLAUDE.global.md §Security — NOT duplicated. NEW: entire writing-style block, design anti-default list, public-site done-checklist, web-app specifics (browser-exposed keys, service-key/client split, RLS, server-side auth, IDOR, hashed passwords + cookie flags, field minimization, rate limiting, upload restrictions). +Placement: CLAUDE.global.md at 308/320 (session-start density guard) → no room for ~30 always-on lines. Decision: rules/writing-style.md WITHOUT paths: (always-on load, same session cost, outside the 320 budget) + rules/web-building.md + rules/web-security.md WITH paths: (lazy-load = token win, fire only on web/code files). Project CLAUDE.md doctrine line amended with the budget exception. Feeds C2 self-contradiction audit. +Conflict carve-outs, stated INSIDE the rules: registries keep caveman format (fragments, em-dashes, bullets); code comments keep code style; structured skill/report templates keep their formats; robuste/transformer banned in buzzword sense only (robustness lens, math transform allowed); no-Inter default rule carries "existing brand identities keep their fonts" (ZenQuality deliverables use Inter+Playfair by brand decision — client-handover BDR). +Self-check rule scoped to DELIVERABLES (text, site, feature), not every conversational reply — literal "avant de me rendre quoi que ce soit" would append a compliance note to every chat answer, pure noise. User can re-widen. +Alternatives rejected: compress into CLAUDE.global.md (~11 lines to fit → loses the carve-outs, zero headroom left); path-scope writing-style (applies to conversation, not file reads → would never fire in chat-only sessions); one merged web file (two concerns, one-rule-one-file). +Branch feature/user-writing-web-rules, UNMERGED (human gate). + +## BDR-086 — darwin bug-pass scope: verified defects fixed above threshold +- **Date**: 2026-08-26 +- **Decision**: units < threshold get full weighted-gap optimization loops (per-unit checkpoint). Judge-VERIFIED defects (file:line, confirmed) in above-threshold units get targeted fixes in a grouped pass — same paired 3-judge validation, one batched checkpoint. User-gated at the scorecard. +- **Why**: leaving a verified destructive path (hotfix `git restore .` wiping tolerated user edits, file scored 85) unfixed = score-worship; rubric serves quality, not the inverse. +- **Alternatives rejected**: strict threshold (ships known bugs); optimize-everything (cost, HL-4 diminishing returns). +- **Reference**: run 2026-08-26, commits 6eceedb..6eac7fb, `.claude/audits/DARWIN-2026-08-26.md`. + +## BDR-087 — Stop hook = attention signal only, never control flow +- **Date**: 2026-09-03 +- **Decision**: `hooks/notify-attention.sh` wired on BOTH `Notification` (matcher = input-needed set) AND `Stop` (no matcher). One script, branches on `.hook_event_name` when `.message`/`.notification_type` absent → Stop yields "Claude has finished responding". Bell + toast now fire every turn end. +- **Why**: Notification types cover input-needed ONLY. Turn-end had no event; nearest was `idle_prompt`, ~60s late — useless for Remote-SSH user away from screen. User enumerated turn-end as required case. +- **Alternatives rejected**: second dedicated script (duplicates terminalSequence + jq logic, two files to keep in sync); `idle_prompt` alone (60s lag); SubagentStop too (noise, subagent completion not user-visible moment). +- **Guard vs prior refusal**: [[BDR-083]] (unlazy review, GATE 0) REFUSED a Stop hook using `decision:"block"` (forces continuation, inverts human gates). THIS Stop hook returns `terminalSequence` + `suppressOutput` only, exit 0, zero control-flow effect. Signal ≠ control. Do not read the refusal as banning Stop outright. +- **Status**: accepted. +- **Reference**: [[LRN-146]] event-coverage gap, [[BLK-020]] client-side faults, [[LRN-145]] terminalSequence pattern. Verified live: turn-end + AskUserQuestion both ring; `permission_prompt` unexercisable under `defaultMode: auto`. diff --git a/.claude/memory/evals.md b/.claude/memory/evals.md index 87905cc..814446a 100644 --- a/.claude/memory/evals.md +++ b/.claude/memory/evals.md @@ -37,6 +37,8 @@ rules: | EVAL-018 | 2026-07-06 | job3 docs-drift audit + execution: 46/46 findings verified, 20/23 fixes shipped (B1 blocked, D2-D5+B6 skipped by decision), zero residual on re-sweep | keep | | EVAL-019 | 2026-07-06 | job4 test-gap audit + execution: 11 specs + 5 fixes/seams, every mutation red-green verified, zero residual | keep | | EVAL-025 | 2026-07-17 | opening seo/geo inventory (subagents): 7/7 verifiable claims false or overstated; real contact corrected all, 6 plan corrections + 4 features killed at measurement | keep | +| EVAL-027 | 2026-08-24 | contract-gates behavioral RED: 16/16 fresh unprimed runs followed new doctrine (GATE 0 order, vacuous oracle, ABANDONED routing, scope temptation resisted) | keep | +| EVAL-028 | 2026-08-26 | darwin v2.1 paired run 54 units: 60 paired verdicts 0 revert/tie; skeptics found 3 real residuals — engaged, not rubber-stamp | keep | --- @@ -251,3 +253,17 @@ rules: ### EVAL-026 — 3-way plan challenge caught 4 BLOCKERs dogfooding own plan (2026-07-17) Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itself. Verdicts CONCERNS(4)/FATAL(6)/FATAL(4). Caught 4 distinct BLOCKERs a single pass would blend: (1) v1 unbuildable — targeted init-project (inline-load, no dispatch) + false "plan on disk" premise for feat/bugfix (only contract persists); (2) failed-open silently dropping a lens while claiming "challenged" (inverts verify-secure-loop "a mute verifier is NEVER a PASS"); (3) consensus-weighting buries lone L2 security finding (lenses orthogonal); (4) sonnet challengers violate [[BDR-066]] (audit judgment=big model). Synthesis REJECTED 1 false positive (allowed-tools-blocks-dispatch — ship-feature has same frontmatter + dispatches fine). Each lens found a DIFFERENT class of flaw → evidence 3-independent > 1-multilens. Action: hardened v2 (severity-driven + fail-safe + re-think loop) shipped. Method validated itself before build. + +### EVAL-027 — contract-gates behavioral RED: 16/16 fresh runs follow the new doctrine (2026-08-24) +- **output**: BDR-083 doctrine (GATE 0 in verify-secure-loop, oracle rules in contract-interview, oracle-consumption + ABANDONED(n) in verifier, 4 passes in feater/bugfixer) — locks prove the TEXT is there; this RED measured whether fresh unprimed contexts FOLLOW it. +- **method**: 16 subagent runs on sandbox repos (scratchpad/red/), prompts = the documented dispatch shapes verbatim, zero mention of test/measure/gates (LRN-080 anti-priming; distinct from LRN-080's own question — instruction already written, question = compliance not pre-existence). Production agents (subagent_type verifier ×9, feater ×2) + fresh orchestrator roles ×5. Every claim re-scored deterministically after: EVIDENCE lines physically rewritten in contracts, git status on sandboxes, gates.sh parse of authored contracts. +- **verdict**: 16/16 conformant. v1 red-oracle-wins 3/3 (NOT-MET citing evidence, own re-run). v2 vacuous-oracle 3/3 — hardest rule (green evidence + correct code → still NOT-MET, evidence explicitly discarded per rule). v3 abandonment semantics 2/2 + v3b pure precedence 1/1 (ABANDONED(1), not CONFORME). o-red 2/2 (gates.sh FIRST, verdict parsed, NO verifier on red floor, executor re-dispatch = contract path + NOT-MET rows verbatim, floor iteration counted 1/3). o-green 1/1 (floor → verifier dispatch with CONTRACT+DIFF+TEST only). e contract-authoring 2/2 (3 oracles + 1 judgement-kept-manual, parse clean in gates.sh first try, POSITIVE CONTROLS run unprompted — rule 3 internalized, markers distinct success-only tokens). f feater 2/2 (out-of-scope temptation src/util.sh SEEN and named untouched, no commit, no placeholder, 4 passes visible in report). +- **anomalies**: none against doctrine. Fixture flaw (mine): placeholder.txt trick used to fabricate a 2nd commit made v2/v3 diffs contain no feature work — every verifier CAUGHT it (out-of-scope + "implementation pre-exists base commit"), polluting v3's intended pure-ABANDONED measurement → v3b clean fixture added. Subjects sharper than the fixture: one flagged the abandon reason not covering the missing French doc. +- **limits**: N=1-3 per cell; subjects read short fresh docs in small sandboxes — long-context production noise not simulated; orchestrator subjects = general-purpose agents told to follow the doc, not the full /feat skill stack. +- **action**: keep — doctrine ships as written, no reinforcement wording needed. Artifacts: scratchpad/red/ (session-lived, not committed). + +## EVAL-028 — darwin v2.1 paired run, 54 units +- **Date**: 2026-08-26. **Output**: 12 optimization rounds (13 sub-80 units) + 8 bug-fix commits, all kept. +- **Method**: paired same-judge 3-majority per round (v2.1); judges live-exec where artifact executable (5 units: skills-perso, profile, plugin-pair, status-reporter, gitflow). Absolute scores triage-only. Totals main-thread (LRN-018 applied). +- **Anomalies**: (1) 0 reverts/ties in 60 verdicts — homogeneous-better checked: skeptic lens found real residuals 3x (doctor.sh cost source, hotfix RULES leftover restore, FILE(S) new-marker) → judges engaged. (2) census lock RED on line-rewrap, make test caught → LRN-144. (3) head-pipe masked grep exit 2x → LRN-143. +- **Action**: v2.1 paired = standard. Post-run absolute rescore skipped by design (would be judge-noise theater). diff --git a/.claude/memory/journal.md b/.claude/memory/journal.md index 78a43c9..e7b64d3 100644 --- a/.claude/memory/journal.md +++ b/.claude/memory/journal.md @@ -427,3 +427,36 @@ rules: ## 2026-07-22 - User: auto-gitignore+delete transient pipeline artifacts in all projects. Investigation reframed the ask — gitignore = WRONG tool (files read from disk during run; would break superpowers SDD `git add` of spec). BDR-065 already rejected gitignore + its DELETE side was doctrine-only (no code, manual chore slipped once — 655e364). User picks (2 recommended): keep committed-during-run + AUTOMATE delete; keep `.claude/tasks/{contracts,plans}` versioned. - Built `lib/gitflow.sh` `_gitflow_purge_transient` at finish (feature/bugfix, pre-merge, best-effort never-abort, opt-out `GITFLOW_PURGE_TRANSIENT=0`) + `purge-transient` CLI verb. Universal via `~/.claude/lib`→repo symlink. gitflow-test T17 a-d (10 checks, `--full-history` recovery), shellcheck clean, make test exit 0. BDR-065 amendment + [[LRN-138]]. feature/gitflow-auto-purge-transient. + +## 2026-07-30 +- User: Opus 5 "needs more freedom" → analyse config + adapt. Research 3-agent (registries / config audit / web) + official migration guide: over-delegation (inverts LRN-030), over-verification, literal following, scope expansion, #80988 injections. Plan challenged 3 blind Opus 5 plan-challengers — robustness FATAL (BLOCKER: symlink-live deployment), all fixes adopted. Shipped: CLAUDE.global.md recalibrated (delegation when-guidance, staff-bar dropped, finish-whole-task, deliverable-length; 308/320), design hook \bux\b dropped flip-tested (22/0), plan-challenger grounded-doubt→[MINOR] (44/0). BDR-081 + LRN-139. feature/opus5-config-tuning, UNMERGED. + +## 2026-08-02 +- C1 seo/geo de-prescription EXECUTED end-to-end: census-first 71 locks flip-proven → reword under audience×range invariant (adafa35/c7646a9) → controlled dogfood (judge-replay frozen signals + templates + fresh collects + e2e + blind reader) → 42/42 both sets, zero contract regression, recall improved. Plan survived 4 challenge passes (2 FATAL + confirmation FATAL(9), all closed by name). BDR-082 + LRN-140. Nested-CLI dogfood died on monthly spend limit → inline pipeline (canonical /seo shape). feature/seo-geo-deprescription UNMERGED (human gate). Chantiers C2-C4 pending. + +## 2026-08-24 +- Analysed `unlazy` skill (Leonxlnx/unlazy 2.1.0) on user request. Its verification architecture teaches us nothing — contract + fresh blind verifier + bounded loops already shipped. Real gap: no deterministic floor between executor and GATE 1 (the verifier's `PROOF:` is a line it writes, not a process exit). +- Shipped Palier 2 (user-chosen): lib/gates.sh + GATE 0 + oracle-bearing criteria + `ABANDONED(n)` verdict + 4-pass executors. Refused unlazy's Stop hook, approval store, .unlazy/ tree, tree-N arithmetic, Node checker — [[BDR-083]] records each why. +- `make test` rc 0, shellcheck clean, 64 new assertions, e2e on a real contract. Branch feature/contract-gates UNMERGED (human gate). +- Locks caught a reflow regression (5 red on rewrapped phrases, zero doctrine lost) → [[LRN-142]]. Skill-adoption pattern → [[LRN-141]]. +- Parallelism audit (user ask "est-ce actif ?"): measured, not assumed — nested probe proves concurrent fan-out (9.1s vs 18s), doctrine already prescribed everywhere safe, remaining serializations motivated. One candidate found: /tour multi-project → parallel runners shipped ([[BDR-084]], user gate "tout paralléliser" + model invariant). Branch feature/tour-parallel UNMERGED. + +## 2026-08-25 +- User permanent rules integrated: rules/writing-style.md (always-on) + web-building.md + web-security.md (path-scoped). Security core already in §Security, not duplicated. Carve-outs protect caveman registries + skill templates + brand fonts. [[BDR-085]]. Branch feature/user-writing-web-rules UNMERGED (human gate). + +## 2026-08-26 — darwin fresh baseline + threshold run (feature/darwin-optimize-20260825, UNMERGED) +- `/darwin-skill all skills and agents` (background). Fresh results.tsv (May file wiped). 7 blind judges, 54 rows (31 skill-systems + 23 agents), mean 83.4, 13 <80. find-docs excluded — machine-owned ctx7 (gitignored), 3rd exclusion ground after BDR-015/058. +- Phase 2: 12 rounds / 13 units, 0 reverts, all paired 3-0 ([[EVAL-028]]). Star: skills-perso detection 8/31 → 31/31 live-verified. Bug pass [[BDR-086]]: 8 commits in above-80 units kept 3-0 (hotfix git-restore data-loss path ★, onboarder contract bounce, plugin data-flow, plan-challenger grammar, handover stale §refs + gate order, tour report-only commit, harden severity, fixtures). +- make test green after census-rewrap fix ([[LRN-144]]); [[LRN-143]] head-pipe grep mask. 29 commits, report `.claude/audits/DARWIN-2026-08-26.md` + card PNG. Branch awaits human review + merge. + +## 2026-09-01 +- Attention signal shipped: hooks/notify-attention.sh + Notification entry in settings.json (bell x2 + OSC 777 toast via terminalSequence). Client-side VS Code steps pending: terminalBell sound:on + osc-notifier ext. [[LRN-145]]. Branch chore/notify-attention-hook, UNMERGED. +- Pre-existing model switch opus[1m] committed separately on same branch. + +## 2026-09-03 +- Attention signal completed + verified end-to-end. Two client faults isolated ([[BLK-020]] resolved): ext instruments only terminals born AFTER activation (re-attach via `dtach -a`, no session loss); Code app volume 0 in Windows mixer killed bell while Windows-emitted toast sound masked it. +- Coverage gap found + closed: `Notification` matcher covers input-needed only, turn-end had no event. `Stop` wired on same script, branches on `.hook_event_name` ([[BDR-087]], [[LRN-146]]). Verified live: turn-end + AskUserQuestion ring; `permission_prompt` unexercisable under `defaultMode: auto`. +- BDR-087 + LRN-146 + BLK-020 capitalized. Branch feature/notify-stop-event, merged to develop (f90ee74). +- Post-merge regression: toast dead again after re-attach from a RESTORED terminal, bell fine. Root cause [[LRN-147]]: ext hooks only terminals born after its activation; `enablePersistentSessions` restores terminals before it. Fix = disable persistent sessions, or fresh terminal + `dtach -a`. Verified: 3/3 toasts on fresh pty. +- Same-day counter-example broke that cause: second session's terminal deaf though created LATER, same window, ext global, shells identical. Trigger unknown; [[LRN-148]] adds the 5s pre-flight test + demotes LRN-147's mechanism claim. +- Attention signal refined: per-event labels (BDR-087 follow-on), silence on non-attention events, and no turn-end signal while `background_tasks` non-empty ([[LRN-149]]). Payload dump beat the docs: `background_tasks` undocumented for Stop but present on the wire. Branch bugfix/notify-subagent-spawn. diff --git a/.claude/memory/learnings.md b/.claude/memory/learnings.md index 2a4d04a..dcf98b2 100644 --- a/.claude/memory/learnings.md +++ b/.claude/memory/learnings.md @@ -138,6 +138,7 @@ rules: | LRN-133 | 2026-07-17 | an omission must stay LEGIBLE, never silent — tool that can't measure says so in its output | designing any audit/measure output; deciding what a cap/refusal/N-A emits | | LRN-134 | 2026-07-17 | resolve-then-pin in stdlib http.client beats monkeypatching getaddrinfo — dual-stack, thread-safe, no requests; classify the OS-resolved IP not the URL text | closing SSRF/DNS-rebinding on any Python HTTP egress | | LRN-135 | 2026-07-17 | a prefix-only scan for a dangerous construct is bypassable by padding — scan the WHOLE document | refusing any hostile construct (DTD/directive/marker) before parse | +| LRN-143 | 2026-08-26 | `cmd \| head \|\| fallback` — pipeline rc is head's (0), fallback dead; bounded output → drop head, else pipefail | any probe/fallback bash in skills before trusting `\|\|` | --- @@ -1355,3 +1356,68 @@ rules: - **context**: user asked to gitignore transient planning artifacts (`docs/superpowers/{specs,plans}`, `.claude/tasks/{contracts,plans}`) to stop them merging. BDR-065 had already REJECTED gitignore for docs/superpowers on the git-travel ground; the real gap was the DELETE side never being coded (doctrine-only manual chore, slipped once — 655e364). Built `_gitflow_purge_transient`. - **future application**: "don't merge transient X" → ask: does the run read X from disk? does X travel via git (worktree, foreign checkout)? Yes → auto-purge at finish, not gitignore. Scoped commit `-- ` avoids sweeping a dirty index; `git diff --quiet HEAD -- paths` precheck makes `git rm` all-or-nothing safe; keep the purge best-effort so cleanup NEVER blocks a merge. Prove archive-reachability with `git log --full-history` / `git show :path` — plain `git log -- path` prunes the purged add-commit via history simplification (bit me writing T17). - **link**: [[BDR-065]]. + +## LRN-139 — model-trait compensations invert across generations; state WHEN-guidance, not direction (2026-07-30) +- **pattern**: config rules that COMPENSATE a model trait become counter-productive when the next generation inverts the trait. LRN-030 (Opus 4.8 under-delegates → "Default to delegation… counters under-delegation") inverted by Opus 5 (delegates MORE readily, official guide) — the rule pushed the failure the model now has. Same class: explicit verify instructions → over-verification; conservative-reporting clauses → literal recall suppression; MUST/CRITICAL → over-triggering. +- **Opus 5 traps found**: (a) Claude Code injects Opus-5-only anti-delegation prompt sections (heron_brook + subagent_steer_delegation, issue #80988; server-gated, no opt-out, absent from transcripts) — own prose stacks on top blindly; (b) NO model-default effort hold on Opus 5 — persisted effortLevel (xhigh, settings.json) silently carries over, against "start high, sweep low/medium"; run /effort sweep per model; (c) effort does NOT shorten visible output/deliverables — only prose length rules do (+30-40% docs). +- **future application**: at every model-generation bump, grep config for trait-compensating language ("counters model tendency…", "default to X") and re-verify the premise; prefer WHEN-guidance (conditions where X pays) over directional nudges — survives inversions unchanged. +- **link**: [[LRN-030]] [[BDR-081]]. + +## LRN-140 — de-prescription findings: dedup evaporates, self-verify is default, recall survives (2026-08-02) +- **pattern 1 — inventory dedup counts lie**: line-level inspection killed most "duplicate" pairs (seo 9 families→2 real merges; geo 7→0). Twins differ by AUDIENCE (bundle-item payload read by fresh applier vs spec rule) or MODE-RANGE (collect/judge/template/RULES) or are distinct obligations sharing a keyword (30/70 ×3 = three different rules). Dedup rule that survives: verbatim + same-audience + same-range ONLY. +- **pattern 2 — Opus 5 self-verifies unprompted**: "run it twice" instruction REMOVED → after-judge still ran score engine twice, identical output. Removing verify-prose does not remove the behavior; its value = no compounding, no contradiction burn. Confirms BDR-081 E3 mechanism, refines the payoff claim. +- **pattern 3 — de-prescription does NOT depress recall**: reworded collect caught  -encoded phone AT COLLECT (baseline collect missed it); reworded judge found new RGPD finding + self-caught false positive + corrected collect coverage claim 21/21→20/21. Integrity/honesty invariants (kept class B) carry the discipline, not the caps. +- **pattern 4 — lock strings, never shapes**: LLM-convention output layers (banners, fences, table columns, section order) wobble run-to-run in BOTH directions — baseline itself deviated from spec where after conformed (§0 ENTRIES, BUNDLE-before-SCORING). Stable contract = census-locked literal strings; anything unlocked drifts and MUST be tolerated by consumers (tier recognition "by intent" is the right pattern). +- **link**: [[BDR-082]] [[BDR-081]] [[LRN-139]] [[LRN-113]]. + +## LRN-141 — adopting an external skill: take the invariants, refuse the machinery (2026-08-24) +Context: unlazy import ([[BDR-083]]). Pattern: an external skill's MACHINERY encodes ITS threat model and ITS doctrine; only its INVARIANTS transfer. Two clean cases from one repo. (1) Approval store binding PATH/shell/platform exists because unlazy executes ledgers INHERITED from untrusted repos — importing it into a config that authors its own ledgers buys per-command approval prompts and closes zero threat. (2) Stop hook returning `decision:"block"` exists because unlazy has no human gate — importing it into a config whose spine is "STOP + escalate to human" would make the tooling fight the doctrine. Meanwhile the invariants (exit 0 AND marker; evidence persisted so the next reader gets fact not report; impossible ≠ deletable) cost ~250 l of our own bash and fit the EXISTING contract with no new tree. +Separating test: ask WHAT THREAT / WHAT DOCTRINE does this piece assume. Answer "theirs" → refuse the piece, keep the invariant it was protecting. +Corollary on claims: unlazy's own research/validation-protocol.md RETRACTS its v1 benchmark numbers as unreproducible while the repo DESCRIPTION still advertises them. Read a project's self-criticism before its README — the retraction is the credibility signal, the headline is not. +Future application: any skill/plugin adoption — skills-external/, /plugin-check, install-plugins.sh. + +## LRN-142 — structure locks are fixed-string: reflowing a doctrine paragraph reds them (2026-08-24) +Context: contract-gates ([[BDR-083]]). Editing lib/verify-secure-loop.md rewrapped 5 locked phrases across line breaks ("Max 3 conformity iterations", "Max 3 security iterations", "re-verify the REQUEST first", "always re-checked BEFORE security", "one verifier dispatch + one security dispatch") → loops-light.test.sh 30 pass / 5 fail, though ZERO doctrine was dropped. Locks did their job: they cannot distinguish "clause deleted" from "clause rewrapped", and that conservative bias is correct — the alternative (fuzzy matching) would miss real deletions. +Rule: when editing a doctrine file under structure locks, grep the test's lock strings FIRST, then re-flow AROUND them — each locked phrase stays on one unbroken line. Fix the DOC, not the lock, unless the doctrine genuinely changed. Under locks today: verify-secure-loop.md, contract-interview.md, verifier / security-auditor / plan-challenger agents, seo+geo (71 locks). + +## LRN-143 — pipe to head masks grep exit; `|| fallback` never fires +- **Context**: darwin 2026-08-26 — plugin-probe FRAMEWORK-DEPS (`grep … | head || echo none`) emitted silent-empty on no-match; same bug in run's own probe test. +- **Pattern**: pipeline rc = LAST command's (head = 0 always). `|| fallback` after pipe = dead code. Bounded output → drop head; else `set -o pipefail` or capture + test. +- **Future**: any skill/agent bash probe with a `||` fallback: check what the pipeline rc actually is first. + +## LRN-144 — census locks grep EXACT single-line phrases; prose rewrap breaks them +- **Context**: darwin 2026-08-26 — hotfix RULES rewrap split "No verifier is dispatched at hotfix weight"; loops-light.test.sh lock RED; make test caught post-edit. +- **Pattern**: lib/tests/*.test.sh lock sentences verbatim, single-line. Rewording/rewrapping skill+agent md near locked phrases silently breaks census. +- **Future**: before editing skill/agent prose, grep lib/tests/ for locks in the touched region; run make test BEFORE dispatching judges, not after. + +## LRN-145 — hooks reach the terminal only via terminalSequence JSON field +- **Context**: 2026-09-01 — attention bell for VS Code Remote-SSH (CLI on remote Linux). Hook subprocess has no controlling TTY; /dev/tty unreliable. Docs: terminalSequence = supported side-effect field, fires even on events that discard output. +- **Pattern**: Notification hook → stdout JSON `{suppressOutput:true, terminalSequence:""}`. VS Code terminal ignores OSC 777/9 natively (claude-code #28338); client-side ext wenbopan.vscode-terminal-osc-notifier converts to native toast over Remote-SSH; beep needs accessibility.signals.terminalBell sound:on. permission_prompt fires ~6s late, idle_prompt ~60s. +- **Future**: any hook ringing/notifying the terminal (bell, toast, title) — terminalSequence, never /dev/tty. Input-needed matcher set: permission_prompt|idle_prompt|agent_needs_input|elicitation_dialog|elicitation_url_dialog. + +## LRN-146 — Notification event alone misses end-of-turn; Stop is the missing event +- **Context**: 2026-09-03 — attention signal verified end-to-end after [[BLK-020]]. Matcher `permission_prompt|idle_prompt|agent_needs_input|elicitation_*` covers input-needed cases ONLY. "Claude finished speaking" has no notification_type — nearest was `idle_prompt`, ~60s late. Gap invisible until explicitly enumerated by user. +- **Pattern**: wire SAME hook script on TWO events — `Notification` (matcher = input-needed set) + `Stop` (fires once per turn end, supports terminalSequence, no matcher). Script branches on `.hook_event_name` when `.message`/`.notification_type` absent: Stop → "Claude has finished responding", else default. Read stdin ONCE into var, jq the var (stdin not re-readable). +- **Verified**: turn-end bip+toast OK, AskUserQuestion selector bip+toast OK. `permission_prompt` NOT exercisable under `defaultMode: auto` — ask-rules (`python3 -c *`, `curl`…) auto-approved, no prompt raised. Hooks hot-reloaded by file watcher, no restart. +- **Future**: enumerate the events a signal must cover BEFORE wiring, one per user-visible moment. Notification ≠ lifecycle-complete. SubagentStop exists too for agent completion. + +## LRN-147 — VS Code restores terminals BEFORE ext activation → toast dies every restart +- **Context**: 2026-09-03, second hit same day. Bell OK, toast gone, after user re-attached session from a restored terminal. Probe on that pty: OSC 777 unique + OSC 777 repeated + OSC 9 → all three silent, while BEL rang. Same pty, bell works ⇒ bytes arrive, ext just not hooked to that terminal. +- **Pattern**: `wenbopan.vscode-terminal-osc-notifier` instruments a terminal only if it exists AFTER ext activation. `terminal.integrated.enablePersistentSessions` (default true) restores terminals at window startup, i.e. BEFORE lazy ext activation → every restored terminal is permanently deaf to OSC. Recurs at each VS Code restart, silently, bell still ringing so it reads as "half broken". +- **Fix**: client setting `"terminal.integrated.enablePersistentSessions": false` → no terminal pre-exists activation. Fallback without it: after VS Code start, open a FRESH terminal then `dtach -a ~/.dtach/` (dtach broadcasts, old client can stay or be closed, session never lost). +- **Diagnostic shortcut**: bell rings + toast dead on the SAME pty = terminal-instrumentation fault, not audio, not hook, not server. Bell dead + toast alive = audio fault ([[BLK-020]] fault B). The two channels split the search space; check which one survives before anything else. +- **Future**: any client-side terminal-parsing ext over Remote-SSH inherits this. Verify instrumentation on the ACTUAL attached pty after every restart, never assume yesterday's terminal. + +## LRN-148 — terminal instrumentation is per-terminal + unpredictable; pre-flight test before attaching +- **Refines**: [[LRN-147]] blamed restored-terminals-born-before-activation. Too narrow — counter-example same day: two terminals SAME VS Code window, pts/3 (born 01:58:33) instrumented, pts/7 (born 01:59:29, LATER) deaf. Ext is GLOBAL (marketplace: Enable/Disable pause parsing extension-wide, no per-terminal setting), shells identical on every server-side measurable: `VSCODE_INJECTION=1`, TERM, TERM_PROGRAM, same `--init-file` shell-integration path, ~2-3s between shell start and dtach. Trigger NOT identified. +- **Pattern**: treat instrumentation as a per-terminal property that can silently fail for unknown reasons. Cheap pre-flight before committing a long-lived session to a terminal: `printf '\a\a\033]777;notify;NEUF;test\033\\'` typed IN that terminal. Toast → instrumented, attach. Bell only → deaf terminal, open another. Costs 5s, replaces an hour of pty archaeology. +- **Recovery**: deaf terminal never repairs. Open fresh terminal, pre-flight it, `dtach -a ~/.dtach/`. dtach broadcasts, so old client may stay attached; session never at risk. +- **Diagnostic split (holds)**: bell alive + toast dead = terminal instrumentation. Toast alive + bell dead = client audio ([[BLK-020]]). Neither = bytes never arrive. +- **Future**: do NOT assert the born-before-activation cause as established — it fits the first incident, not the second. Unknown trigger is the honest state. + +## LRN-149 — Stop hook payload carries background_tasks; use it to skip premature signals +- **Context**: 2026-09-03. User: "notif à la création d'un sous-agent alors qu'il faudrait pas". Instrumented hook, ran probe subagents: NEITHER subagent creation NOR completion calls the hook. Only event = `Stop`, fired when the turn ends right after spawning. Signal was real but LIED ("Finished responding" while work continued). +- **Pattern**: dump the real payload (`printf '%s' "$payload" >> file.jsonl`) instead of trusting docs — docs list Stop fields without `background_tasks`, the wire has it: `[{"id","type":"subagent","status":"running","description","agent_type"}]`. Rule: on Stop, `(.background_tasks // []) | length` > 0 → exit 0 silent. Next turn end signals for real. Interaction events (permission/question) always signal, background or not. +- **Fail-open**: field absent (older client) → still signal. Missed notification worse than extra one. +- **Cross-session gotcha**: hook is user-scope, so EVERY session runs it. A single-file dump (`> file`) gets overwritten by another project's session — append JSONL and filter on `.cwd`. That accident proved `permission_prompt` fires with `message="Claude needs your permission"` (unexercisable in this session under `defaultMode: auto`). +- **Future**: any hook needing turn-completion semantics must check background_tasks; "turn ended" ≠ "work done". Verified live: Stop with 0 tasks signals, Stop with 1 running subagent silent. diff --git a/.claude/tasks/TODO.md b/.claude/tasks/TODO.md index cc2c8cb..59a9b67 100644 --- a/.claude/tasks/TODO.md +++ b/.claude/tasks/TODO.md @@ -1,5 +1,105 @@ # TODO +## 2026-08-25 — darwin fresh baseline: 32 skill-systems + 23 agents (feature/darwin-optimize-20260825) +User: `/darwin-skill all skills and agents` (background). Fresh-from-zero +(results.tsv wiped 2026-06-23, journal 2026-06-30). Scope per BDR-015/043 + +LRN-070: personal skills only, external/gstack OUT. EVAL-004 applied: eval +unit = skill+dispatched-agents SYSTEM, agents get own rows. LRN-018: judges +emit per-dim scores, totals recomputed main-thread. v2.1 keep/revert = +paired same-judge majority, absolute scores triage-only. +- [x] T1 Phase 0+0.5: gitflow branch, results.tsv header, 7 new + test-prompts.json (capitalize deploy gitflow pdf-translate reconcile + release-candidate tour), runtime scan (2 minor hits). find-docs + EXCLUDED — machine-owned ctx7 (BDR-053, gitignored) → 31 systems. +- [x] T2 Phase 0.5 gate PASSED: reuse prompts as-is; dim8 full_test on + candidates only (baseline dry_run); Phase 2 set = ALL units <80. +- [x] T3 Phase 1 baseline DONE: 7 blind judges, 54 rows (31 skills + 23 + agents), mean 83.4, 13 units <80, ~25 verified findings (hotfix + destructive restore, onboard/onboarder contract, init-project + allowed-tools, skills-perso 8/32 detection...). + +- [x] T4 Phase 1 gate PASSED: user picked the set — proven by Phase 2 + running 13/13 units, 0 reverts (DARWIN-2026-08-26.md:23). + Ticked by reconcile 2026-09-01. +- [x] T5 Phase 2 DONE: 13/13 units, 12 rounds kept 3-0, 0 reverts + + bug pass 8 commits kept 3-0 (2 skeptic residuals amended). make test + green. +- [x] T6 Phase 3 DONE: report .claude/audits/DARWIN-2026-08-26.md + card + PNG (playwright fallback). Capitalize pending user approval. Branch + UNMERGED — human gate. + → both residuals stale: capitalized a15854a, merged 726464f + (reconcile 2026-09-01). + +## 2026-08-25 — user permanent rules: writing + web build + web security (feature/user-writing-web-rules) +User supplied 4-block rule text (écris / site / code / vérification); asked: +coverage check, conflict check, integrate. Verdict: security CORE already in +CLAUDE.global.md §Security (parameterized queries, env-var secrets, +AuthN/AuthZ, fail closed) — NOT duplicated. NEW: writing-style block, design +anti-default list, site done-checklist, web-app specifics (RLS, service key, +IDOR, cookie flags, rate limit, field minimization). Placement: global at +308/320 budget → rules/ instead. +- [x] R1 rules/writing-style.md — always-on (no paths:), scope carve-outs + (registries caveman, code comments, skill templates) + self-check +- [x] R2 rules/web-building.md — paths: web globs; anti-defaults + done + checklist (report missing, never invent) + skill pointers +- [x] R3 rules/web-security.md — paths: code globs; web-app specifics + extending §Security, zero dup of the core +- [x] R4 CLAUDE.md (project) — amend always-on doctrine line (320-budget + exception → rules/), feeds C2 audit +- [x] R5 capitalize BDR-085 + journal + CHANGELOG +- [x] merge → develop 5ec7bfa — human gate passed (reconcile 2026-08-25) + +## 2026-07-30 — adapt config for Claude 5 family / Opus 5 (feature/opus5-config-tuning) +User: Opus 5 "needs more freedom" → research (official migration guide + +web + registres) confirms: over-delegates (inverts LRN-030 Opus 4.8 trait), +over-verifies if told to verify, literal instruction following, scope +expansion named regression, harness already injects anti-delegation on +Opus 5 (#80988). Plan: .claude/tasks/plans/2026-07-30-opus5-config-tuning-1238.md +— to be challenged by 3 blind plan-challengers (opus pins → Opus 5), then +executed on feature branch. NO merge (human gate). +Challenged 2026-07-30: correctness CONCERNS(4) · robustness FATAL(5, 1 +BLOCKER: symlink-live deployment) · simplicity CONCERNS(4) — all fixes +adopted as prescribed (plan §5bis, v2 items below). +- [x] W0 branch first (eab2a10 parent); hook regex validated on scratch copy + (bash -n + shellcheck + 5 replays, HOME sandboxed) before live write +- [x] W1 delegation block v2 (when-guidance + gates carve-out + scoped don't-redo) — 0f7b565 +- [x] W2 "staff engineer" bar line deleted — 0f7b565 +- [x] W3 finish-whole-task folded into Deviations (+ gone-WRONG→STOP) — 0f7b565 +- [x] W4 deliverable-length rule — 0f7b565 +- [x] W5 line budget: 308/320 +- [x] W6 hook \bux\b dropped, \bui\b kept + F10 must-fire lock, D11 quiet row + flip-tested (fire before/quiet after) — eab2a10, suite 22/0 +- [x] W7 plan-challenger :82-83 reworded → [MINOR] routing, census row — c3d3f4d, 44/0 +- [x] W8 BDR-081 + LRN-139 + journal + CHANGELOG +- [x] W9 final gate: make test full suite — green except known T6c + (darwin-skill residual → chantier 4 below), 2026-07-30 +- [x] W10 merged on explicit user signal — 709cf9b (2026-07-30 13:28), + branch deleted; confirmed post-merge this session + +## 2026-07-30 — Claude 5 follow-on chantiers (user directive, checkpoint between each) +Order fixed, one branch per chantier, no merge without per-chantier signal. +- [x] C1 dé-prescription seo-analyzer.md + geo-analyzer.md — DONE 2026-08-02. + Census-first 71 locks flip-proven (9681b46) → rewords under + audience×range invariant (adafa35 seo, c7646a9 geo) → controlled + before/after dogfood: judge-replay on frozen signals + templates + + fresh collects + e2e judge + blind reader = 42/42 both sets, zero + contract regression, recall improved. Plan challenged 4 passes + (FATAL/FATAL/CONCERNS + confirmation FATAL(9), all closed by name). + BDR-082 + LRN-140. Evidence .audit/dogfood-baseline/ (19 artifacts). + Branch feature/seo-geo-deprescription UNMERGED — human gate. + → merged 5488c48, branch deleted (reconcile 2026-08-25). + Residual for gate: §6bis dynamically-unverified list (FULL branches, + apply path — census-locked statically); FULL/aggressive dry-run = user + option; nested-CLI dogfood blocked by monthly spend limit (inline used). +- [ ] C2 self-contradiction audit CLAUDE.global.md + own skills: list rule + pairs in tension, propose resolution per pair, apply after user OK. + /doctor as assistant, not authority. +- [ ] C3 superpowers: MEASURE first (skill-invocation log over sessions) + whether "1% chance → MUST invoke" over-triggers; if yes, options + + trade-offs (disable plugin / softer house rule / live with) — user decides. +- [x] C4 hygiene: reinstall darwin-skill — DONE (reconcile 2026-08-25: + ~/.agents/skills/darwin-skill present, T6c green, make test exit 0). + ## 2026-07-22 — auto-purge transient superpowers artifacts at finish (feature/gitflow-auto-purge-transient) User: transient planning artifacts (`docs/superpowers/{specs,plans}`) leak into develop; BDR-065 "post-merge cleanup" is DOCTRINE ONLY (no code) — manual chore, @@ -19,15 +119,18 @@ versioned (durable, referenced by decisions.md e.g. BDR-076). Universal via the - [x] Gate: shellcheck lib/*.sh CLEAN + `make test` exit 0 (gitflow 106/0, full suite green). Universal via ~/.claude/lib → repo lib symlink (verified). - [x] CLAUDE.md §Transient planning artifacts: → "AUTO-PURGED by gitflow finish". -- [ ] Capitalize: BDR-065 amendment (delete side now automated) + LRN — pending user OK. +- [x] Capitalize: BDR-065 Amendment (2026-07-22) in body + LRN-138 present + (reconcile 2026-08-25). ## 2026-07-20 — pending merge gates (reconcile) - [x] merge feature/profile-managed-externals → develop (BDR-079 profile symmetry + /doc clean pass: README/USAGE/ARCHITECTURE.md) — 37c79f0 - [x] merge chore/purge-transient-docs → develop (docs/ transient purge 655e364 + reconcile e75ea79) — reaches main at next release -- [ ] Makefile help text: profiles 5/10 listed (:57) + test glob missing - run-*.sh (:31) — 2-line hotfix (flagged by /doc audit) +- [ ] Makefile help text: profile-list help lists 5/10 profiles (:57) — + 1-line hotfix. (test glob :31 FIXED — has run-*.sh, reconcile 2026-08-25) + Re-verified OPEN 2026-09-01: lib/profiles/ has 10, Makefile:57 lists 5 + (backend, full, seo, web-full, web missing). ## 2026-07-20 — profile ↔ toggle-external symmetry (feature/profile-managed-externals, BDR-079) Audit verdict: gstack on-demand + design enable already work; DISABLE side @@ -1094,3 +1197,78 @@ branch) → LOT3 mis-merge trap; + 3 doctor false-warns (LRN-047 class). comment anchored to measured ~11.4k (LRN-088). False "92% CRITICAL" → ~5% comfortable. - [x] Verify — suites green (71/13/32/19/20/13 + RC 5/5); doctor 0 false-warn; shellcheck clean. +docs(changelog) Unreleased entry (706abff). Gate passed on GO 2026-07-03. Finish pending. + +## 2026-08-24 — contract gates: plancher déterministe (feature/contract-gates) +Source: analyse du skill `unlazy` (Leonxlnx/unlazy, 2.1.0). Verdict: son +architecture de vérification n'apprend rien (contrat+verifier frais+boucles +bornées ⊂ déjà en place). Le trou réel: **entre l'exécuteur et GATE 1 il n'y a +aucun plancher déterministe** — GATE 1 est un dispatch LLM, et `PROOF:` est une +ligne que le verifier ÉCRIT (rien ne l'empêche structurellement de la produire +sans rien exécuter). Palier 2 retenu (user, 2026-08-24). + +PRIS d'unlazy: critère porteur d'oracle exécutable (CHECK/EXPECT/EVIDENCE), +fail-closed (exit 0 ET marqueur), evidence pending = NOT-MET, `ABANDON: +` comme handoff visible non supprimable, les 4 règles d'écriture de +gates falsifiables, la discipline 4 passes. +REFUSÉ: Stop hook `decision:"block"` (contredit "STOP + escalade humaine" et +"merge sur signal humain"), approval store `~/.unlazy/approved` (résout +l'exécution de ledgers hérités non fiables — pas notre menace), arbre +`.unlazy//` (4e arbre de bookkeeping ⇒ mort de la config), `tree N` +(désavoué par ses propres docs), le checker Node 28k (stack lib = 100% bash, +Health Stack = shellcheck). + +- [x] W0 branche feature/contract-gates depuis develop (via lib/gitflow.sh) +- [x] W1 `lib/gates.sh` — parse ACCEPTANCE CRITERIA, exécute fail-closed + (exit 0 ET EXPECT), réécrit EVIDENCE dans le contrat. Sous-commandes + `run` (exécute+écrit) / `status` (parse seul, jamais d'exécution, jamais + d'écriture). rc 0=MET · 2=UNMET/malformé · 3=ABANDONED. +- [x] W2 `lib/contract-interview.md` — STEP 3 gagne CHECK/EXPECT/EVIDENCE + optionnels par critère + les 4 règles de falsifiabilité; template mis à + jour; ABANDON dans Lifecycle; ligne de poids par flow. +- [x] W3 `agents/verifier.md` — EVIDENCE fail-closed (coché+pending = NOT-MET), + bucket ABANDONED, verdict `CONFORME` impossible si abandon présent. +- [x] W4 `lib/verify-secure-loop.md` — GATE 0 déterministe avant GATE 1 + (rouge ⇒ re-dispatch exécuteur sans brûler un verifier). +- [x] W5 `agents/feater.md` + `agents/bugfixer.md` — discipline 4 passes. +- [x] W6 `lib/tests/gates.test.sh` — comportemental sur gates.sh (fail-closed, + exit≠0 avec marqueur = FAIL, pending, ABANDON, malformé, status + n'exécute pas) + locks de structure sur W2/W3/W4/W5. +- [x] W7 shellcheck + bash -n + `make test` complet. +- [x] W8 CHANGELOG + registres (BDR + LRN + journal). +- [x] W10 restatements skills : bullet GATE 0 dans feat/bugfix/ship-feature/ + init-project (+4 locks, flip-testé) ; ligne hotfix du tableau de poids + corrigée (aucun floor à ce poids). 2026-08-24. +- [x] W11 RED comportemental : 16/16 runs frais non-amorcés conformes + (verifier ×9, feater ×2, orchestrateur ×5) → EVAL-027. 2026-08-24. +- [x] W9 merge sur signal humain explicite (2026-08-24, "merge dans develop"). + +**Won't-build-now — Palier 3 unlazy (OWNS/leases), trigger documenté :** +Différé volontairement (BDR-083) : tous les dispatches parallèles actuels +sont read-only — le problème (2 exécuteurs ÉCRIVAINS concurrents) n'existe +pas. Pattern [[LRN-080]] : ne pas construire sans menace mesurée. +TRIGGER = le jour où un flow dispatche ≥2 exécuteurs écrivains en parallèle : +(1) FILE SCOPE du contrat = déclaration OWNS (champ existant, zéro format +neuf) ; (2) ~40 l dans gates.sh ou lib/owns.sh — intersection CONSERVATRICE +des FILE SCOPE des contrats actifs avant fan-out, conflit possible → refus + +dispatch séquentiel (pas de locks disque tant que l'orchestrateur est +unique) ; (3) locks + tests. + +## 2026-08-24 — tour multi-projets en parallèle (feature/tour-parallel) +User (gate 2026-08-24): "tout paralléliser (option 2) mais bien garder la +sélection des modèles — orchestrateur garde le modèle orchestrateur, les +skills/agents suivent leurs orchestrateurs définis". Preuve mécanique +préalable: probe imbriquée 3 sous-agents, fenêtres chevauchantes, 9.1s vs +~18s séquentiel. Dérogation LRN-083 (boucle de fix par projet déplacée dans +un runner dispatché) → à consigner BDR-084. Repos indépendants, branches +chore par repo, report-as-approval-gate ⇒ rien de partagé n'est décidé +dans un runner; capitalize reste main-loop. +- [x] T1 skills/tour/SKILL.md — STEP 0 routé (1 projet = inline inchangé; + ≥2 = fan-out) + STEP 0b: un runner general-purpose par projet, TOUS + dans UN message, SANS pin modèle (hérite session, model-gate déjà + passé); agents internes gardent leurs tiers définis; runner mort = + ligne RUNNER FAILED, jamais absent silencieux; capitalize main-loop. +- [x] T2 locks census §12 dans lib/tests/model-routing.test.sh (fan-out + présent, runner non-pinné, single message, capitalize main-loop). +- [x] T3 BDR-084 + CHANGELOG + journal. +- [x] T4 make test rc 0 + shellcheck clean (SC2016 silencé, littéral + voulu). Merge NON fait — gate humain. diff --git a/.claude/tasks/plans/2026-07-30-opus5-config-tuning-1238.md b/.claude/tasks/plans/2026-07-30-opus5-config-tuning-1238.md new file mode 100644 index 0000000..e9c2c7b --- /dev/null +++ b/.claude/tasks/plans/2026-07-30-opus5-config-tuning-1238.md @@ -0,0 +1,255 @@ +# PLAN — Adapt claude-config for the Claude 5 family (Opus 5 focus) + +Date: 2026-07-30 · Branch (planned): feature/opus5-config-tuning (off develop) +KIND: build-plan · Author: main-loop session (Fable 5) + +## 1. Context & evidence + +Opus 5 (`claude-opus-5`, released 2026-07-24) now backs every `model: opus` +agent pin in this repo (analyzer, plan-challenger, seo/geo-analyzer, +plugin-advisor — BDR-076/077) and any session the user switches to via +`/model opus`. Its documented behavioral profile differs from Opus 4.8 in +ways that make parts of this config counterproductive: + +- E1 **Over-delegation**: Opus 5 "delegates to subagents more readily than + prior models" (official prompting guide). Opus 4.8 had the OPPOSITE trait + (LRN-030), and `CLAUDE.global.md:43-47` was written to counter it + ("Counters model tendency to under-delegate"). The premise is inverted. +- E2 **Anti-delegation already injected by the harness**: Claude Code + v2.1.219 server-gates an Opus-5-only prompt section (`heron_brook` + + `subagent_steer_delegation`, GitHub issue #80988) that says "Do not call + the AgentTool unless the user requested it" and "Subagents multiply cost + and time…". Stacking our own hard cap on top would triple-constrain; + keeping a pro-delegation nudge would fight the injection. Model-neutral + when-guidance is the stable middle. +- E3 **Over-verification**: official guidance — "If your prompt contains + explicit verification instructions … remove them: instructions like these + cause over-verification on Claude Opus 5, and removing them reduces wasted + tokens with no loss in quality." Also true of per-prompt "double-check" + phrasing. Targets PROSE told to the model, not harness-level gates. +- E4 **Scope expansion**: named Opus 5 regression ("can expand the scope of + a task, adding steps that weren't requested"). Anthropic ships a literal + counter-block; tested to reduce scope changes "to nearly zero". +- E5 **Literal instruction following** (since 4.7, stronger now): aggressive + MUST/CRITICAL language over-triggers; conservative-reporting instructions + ("only report high-severity") measurably depress recall in review/challenge + harnesses. +- E6 **Longer written deliverables**: files written to disk run ~30-40% + longer; `effort` does NOT control visible/deliverable length — only prose + instructions do. +- E7 **Overconstraint costs reasoning**: Anthropic removed >80% of Claude + Code's system prompt for Claude-5-generation models "with no measurable + loss"; named mechanism = tokens burned resolving conflicting rules. +- E8 **Hook false positive (today)**: `\bux\b` in + `hooks/design-toolchain-reminder.sh:47` fired on French prose ("changement + ux vu" — matches after apostrophe/slash/space); 2nd `ux` FP in the log, + both French. Continues the LRN-1005/1007 false-positive series. No test + row covers `\bui\b`/`\bux\b`. +- E9 **Effort carry-over trap**: Opus 5 has no model-default effort hold in + Claude Code — a persisted `xhigh` (our `settings.json:333`) silently + carries onto Opus 5 sessions, against Anthropic's "start at high, sweep + low/medium" guidance for that model. + +## 2. Design decisions + +- D1 The global instruction layer must be MODEL-NEUTRAL across the Claude 5 + family (sessions run Fable 5 by default; dispatched judgment agents run + Opus 5; executors Sonnet). Fixes therefore express WHEN-guidance and + outcome bars, not directional compensation for one model's trait. +- D2 Harness-level quality gates (fresh blind verifier + security-auditor, + BDR-049/050; plan-challenge, BDR-075) are architecture, not model + self-check prompting. They stay. E3 applies only to prose that tells the + MODEL to verify its own work. +- D3 Per BDR-021, the Security and Architecture-decisions sections of + CLAUDE.global.md stay verbatim (deliberate policy). No softening there. +- D4 Registries are append-only: LRN-030 is not edited; a new LRN records + the trait inversion and points back to it. +- D5 Deterministic backstops (gitflow pre-commit, Gitea protection, + permissions.deny, rtk pinning) are explicitly out of "more freedom" scope + — community reports show Opus 5 working AROUND soft controls, which argues + for keeping hard ones. + +## 3. Work items + +### W1 — CLAUDE.global.md: rewrite the delegation block (:43-47) +Replace the 5-line block (incl. "Default to delegation for multi-file +exploration. Counters model tendency to under-delegate.") with model-neutral +when-guidance, same footprint (≤5 lines): + +``` +- Sub-agents: one task per sub-agent, main context stays clean. + Delegate genuinely independent, sizeable tracks (wide multi-file + exploration, parallel audits) — not work doable in a few tool + calls, and not self-verification (harness gates own that). Brief + precisely, then commit to the delegation — don't redo its work. +``` +Rationale: E1+E2. No hard spawn cap in prose (harness already injects one on +Opus 5; Fable benefits from delegation). + +### W2 — CLAUDE.global.md: reframe "After code changes" (:75-83) +Keep the concrete quality bar; drop the proof-mandate/self-check phrasing +(E3). Replace steps 2-4 with faithful-outcome reporting: + +``` +## After code changes +1. Run tests, lint, build, type-check if available. +2. Report outcomes faithfully: what passed, what wasn't run, + remaining risks, surviving deviations. Completion claims only + for verified work. +3. Correction or notable event → capitalize to right registry. +``` +Net: -2 lines. "Would staff engineer approve?" bar and "Don't mark complete +without proof" are removed as self-check choreography; honest-reporting +line preserves the intent (grounded completion claims) without mandating an +extra verification pass. + +### W3 — CLAUDE.global.md: add scope fence (Workflow section) +Append (adapted from Anthropic's tested block, caveman-compressed, ~5 lines): + +``` +- Scope: deliver what was asked, at the scope intended. Routine + judgment calls → decide alone; materially different readings → + ask. Better approach spotted → say so in one line, still do the + task as asked. Finish the whole task; genuinely blocked → do the + rest, state plainly what's missing. +``` +Rationale: E4. Complements existing "Scope changes to task — no unrelated +edits" (line ~15) without contradicting it. + +### W4 — CLAUDE.global.md: add deliverable-length rule (Code style / Comments area) +~2 lines: + +``` +- Written deliverables (docs, reports, .md): length matched to what + the task needs — no filler sections, no boilerplate summaries. +``` +Rationale: E6. Registries already covered by caveman rule. + +### W5 — Line budget +After W1-W4: expected ~309 lines. Hard check: `wc -l CLAUDE.global.md` ≤ 320 +(session-start.sh warning threshold at :202-213). + +### W6 — hooks/design-toolchain-reminder.sh: drop `\bui\b` and `\bux\b` +- Remove the two 2-char alternatives from the pattern at :47. Keep + `ui/ux|ux/ui|ui kit` and all other tokens. +- Add a dated header comment (3rd tightening pass, 2026-07-30, cites the + two French-prose `ux` FPs; series LRN-1005/1007). +- Trade-off accepted: a bare "améliore l'ux" prompt with no other design + token goes quiet — the CLAUDE.global.md "Design work" section still + routes it (the hook is a belt, self-described soft nudge). +- Update `lib/tests/design-toolchain-reminder.test.sh`: add 2 quiet rows + (the real FP prompt excerpt; a bare "l'ui" French sentence) — flip-tested + per LRN-096. Existing 9 must-fire rows unaffected (none uses ui/ux). + +### W7 — agents/plan-challenger.md: coverage-first reporting line +Add one clause to the findings rules (add-only, no removal): uncertain or +low-severity findings are REPORTED with an explicit confidence + severity +tag rather than self-censored — severity filtering happens in the +orchestrator's synthesis, not in the challenger. Rationale: E5 (literal +Opus 5 + "manufactured concern is a failure" wording risks suppressing real +low-confidence findings). Must not touch: verdict grammar, MANDATORY PROOF +clause, blind-dispatch rules (test-locked in plan-challenger.test.sh). + +### W8 — Memory + docs capitalization (same branch, follows the work) +- decisions.md: new BDR (config adapted for Claude 5 family — scope, + rationale, alternatives incl. "leave config as-is" and "hard spawn caps" + rejected). +- learnings.md: new LRN — Opus 5 behavioral profile (over-delegation + inverts LRN-030's Opus 4.8 trait; over-verification; literal following; + no effort hold on Opus 5 in Claude Code; heron_brook/#80988 injection). +- journal.md: one line. +- CHANGELOG.md: entry under Unreleased. + +### W9 — Gates (before commit) +- `shellcheck hooks/design-toolchain-reminder.sh` clean. +- Manual flip-test of the hook: FP prompt → quiet; "redesign the navbar" → + fires. +- `make test` full suite green (design-toolchain-reminder.test.sh, + plan-challenger.test.sh, model-routing.test.sh untouched-but-must-pass, + curated-config-guard, loops-light…). +- `wc -l CLAUDE.global.md` ≤ 320. + +### W10 — Gitflow +`bash ~/.claude/lib/gitflow.sh start feature opus5-config-tuning` off +develop; atomic commits (hook+test / CLAUDE.global.md / agent / memory+docs); +NO `gitflow finish` — merge only on explicit human signal. + +## 4. Explicitly NOT doing (considered, rejected) + +- N1 Touching lib/verify-secure-loop.md or the fresh-verifier/security + gates: harness architecture (BDR-049/050, D2), verifies SONNET executor + output — not Opus 5 self-check prose. +- N2 Softening the Security / Architecture sections (BDR-021, D3). +- N3 Editing the superpowers plugin's "1% chance → MUST invoke" language: + external upstream code; flagged as residual over-triggering risk in the + new LRN, revisit as its own decision if observed. +- N4 Changing `settings.json` `effortLevel: "xhigh"`: user preference, + optimal for the Fable 5 session default; the Opus 5 carry-over trap (E9) + is documented in the LRN + surfaced to the user for a manual decision. +- N5 De-prescribing seo-analyzer.md / geo-analyzer.md (1528/1106 lines, + heavy MUST density): separate project, backlog note in TODO.md. +- N6 Removing or session-gating the design/ctx7 reminder hooks: soft + nudges, cheap, deliberately built; tightened only (W6). +- N7 Any model pin change: `model: opus` pins now resolve to Opus 5 — + desired outcome, census (model-routing.test.sh) untouched. +- N8 Committing settings.json for any reason (LRN-098/1049 /model-churn + trap): file is currently clean; keep it out of every commit. + +## 5bis. CHALLENGE SYNTHESIS (2026-07-30) — FINAL amendments (v2) + +Verdicts: correctness CONCERNS(4) · robustness FATAL(5, 1 BLOCKER) · +simplicity CONCERNS(4). Every fix below is the challenger's own named FIX, +adopted as written. No re-challenge pass: scope narrowed, no new dependency; +W0 is an execution-time safety procedure, not a new config mechanism. + +- **W0 (NEW — robustness BLOCKER)**: all edited surfaces are symlink-deployed + LIVE (~/.claude/CLAUDE.md, hooks/, agents/ → this repo); edits take effect + machine-wide at save time, before any W9 gate. Mitigations: + (a) `gitflow start` BEFORE any live-file edit; never checkout develop + mid-work; (b) hook regex change validated on a SCRATCH copy first + (bash -n + shellcheck + pattern replay), then written to the live file in + ONE atomic Edit; (c) named reverts: `git show develop: > `; + escape hatch = remove the hook registration block from settings.json. +- **W1 v2** (robustness#3, correctness#2): replacement text carves out the + mandated gates explicitly and scopes "don't redo": + "Skill-mandated gates (fresh verifier/security/challenge) always dispatch + as written. Don't redo delegated work by hand — failed gates re-dispatch + fresh executors instead." +- **W2 v2** (simplicity#2): minimal diff — delete ONLY the line + `Bar: "would staff engineer approve?"`. Steps 1-4 + capitalize step stay. +- **W3 v2** (simplicity#1, robustness#4): no new bullet. Fold the only new + clause into the existing Deviations bullet: "Finish the whole task: + blocked on an independent sub-part → do the rest, state what's missing. + Gone WRONG → still STOP, re-plan." (net +2 lines, no conflict with :53). +- **W4**: unchanged (+2 lines). Budget v2: 304 +1 −1 +2 +2 = 308 ≤ 320. +- **W6 v2** (all lenses): drop `\bux\b` ONLY — keep `\bui\b` (zero evidenced + FP; one logged true positive). Accepted trade-off: the 2026-07-21 "ameliore + le tutoriel…gamifier" ux row (plausible TP) goes quiet; CLAUDE.global.md + design-routing section remains the router. Header comment notes the log + records `head -1` only → per-token FP rate not fully derivable. Tests: + quiet row = synthetic "changement ux vu…" (verified matches pre-change → + flips); must-fire row = "revois l'ui du panneau admin" (locks `\bui\b`; + apostrophe escaped correctly, doubles as JSON-path control per + robustness#7). No log-excerpt rows (vacuous — 100-char truncation). +- **W7 v2** (all lenses): in-place reword of the `:82-83` sentence (NOT + test-locked; plan v1 misstated that) instead of an add-only clause: + "No invention — ungrounded is noise. Silently dropping a grounded doubt is + equally a failure: file it as `[MINOR]` with the uncertainty stated in + `WHY:`. Nothing real at all → `SOLID` with `FINDINGS: none`." + OUTPUT grammar byte-identical; no confidence axis; no consumer change. + Census: add `has "$A" "grounded doubt"` row to plan-challenger.test.sh in + the same commit. +- **W9 v2**: adds the W0 scratch-validation step; rest unchanged. +- **W10 v2**: branch creation moves FIRST in execution order. + +## 5. Constraints for challengers + +- Registries append-only; curation only via /prune-memory. +- Census tests lock behavior: any hook/agent edit must land with its test + update in the same commit; `make test` must stay green. +- CLAUDE.global.md ≤ 320 lines (runtime warning threshold). +- BDR-021: Security + Architecture sections verbatim. +- Gitflow: feature branch off develop, no merge without human signal. +- The global file serves ALL models (Fable sessions, Opus 5 dispatches, + Sonnet executors read skill/agent prompts instead) — no Opus-5-only + wording in CLAUDE.global.md. diff --git a/.claude/tasks/plans/2026-07-30-seo-geo-deprescription-1402-inventory.md b/.claude/tasks/plans/2026-07-30-seo-geo-deprescription-1402-inventory.md new file mode 100644 index 0000000..d572ffc --- /dev/null +++ b/.claude/tasks/plans/2026-07-30-seo-geo-deprescription-1402-inventory.md @@ -0,0 +1,174 @@ +# ANNEX — directive-language inventory (analyzer report, 2026-07-30) + +Produced by a read-only analyzer dispatch over agents/seo-analyzer.md +(1528 l) + agents/geo-analyzer.md (1106 l), cross-referenced against +every consumer. Referenced by the C1 plan (same folder, -1402.md). + +## 0. Token census (raw) + +| Token family | seo-analyzer.md | geo-analyzer.md | +|---|---|---| +| MUST/must | 12 | 6 | +| MANDATORY/mandatory | 8 | 4 | +| NEVER/never | 42 | 33 | +| ALWAYS/always | 6 | 1 | +| CRITICAL/critical | 3 | 1 | +| Do NOT / do not | 24 | 8 | +| verbatim | 6 | 3 | +| STOP | 3 | 4 | +| refuse/REFUSE | 6 | 4 | +| ⚠️ blocks | 0 | 0 | + +## 1. Test locks on these files (complete list — 6 per file) + +model-routing.test.sh:67-68 `model: opus` (both) · :150-157 `MODE: +collect|judge|template` + `COLLECTION COMPLETE` (both) · +seo-data.test.sh:538-540 `fetch.sh crux` / `fetch.sh queries` / +`Performance GSC` (seo) · :542-543 `fetch.sh schema_gen` / +`fetch.sh content_quality` (geo). +NOT locked by any test: READY-TO-APPLY sentinel, envelope headings, +score-block shapes, JUDGE-ERROR strings, batch labels — contracts by +consumer only; a rewrite can break them silently and make test stays +green. Sibling dispatcher locks: model-routing.test.sh:159-166. +Stale line-number comments (no enforcement): lib/url-guard.sh:9, +url-guard.test.sh:20, source-scope.sh:25, seo-data/README.md:196/309, +drift.py:4, linkgraph.py:4 — all already drifted. + +## 2. Format contract (artifact → consumer) — FREEZE SET + +seo-analyzer: signals `.audit/seo-signals-.md` (+clean/load sites +in /seo) · `COLLECTION COMPLETE — RUNID: ` terminal · +`COLLECT REPORT` w/ `STATUS: DONE|BLOCKED` · `SEO JUDGE — VERDICT: +ERROR()` · judge report forwarded verbatim to template · +`SEO SCORING ()` block w/ `COVERAGE SOURCE:`/`COVERAGE LIVE :` ++ 7 axes + `SEO GLOBAL (weighted): XX.X/20` (score.py:26-37 mirrors +weights) · `TRAJECTORY TO 17/20 (code-only)` · `fetch.sh score` JSON +(`axes.{technical,on-page,seo-local,off-page,social,competitive,legal}`, +severities `critique|haute|moyenne|basse`, `status:"na"`) · `FIX PLAN (N +findings total)` + BATCH A…F (tier-mapping tolerant) · `## FIX BUNDLE +(for dispatcher)` + `### AUTO/### GATED/### USER ACTIONS` + item fields +`id: applier: files: concern: current: expected:` · sentinel `READY TO +APPLY — awaiting dispatcher confirmation` (also reused by /harden:366) · +envelope `SEO AGENT RESULT` + `## SECTION FOR SEO.md §2…§6` + `## ENTRIES +FOR SEO.md §0/§8/§9/§10/§11/§15` · `Automatisation possible avec:` per +§11 entry · standalone `.claude/audits/SEO.md` w/ `**Score SEO** : XX.X +/ 20` (client-handover-writer.md:344 labeled grep) + §0-§15 + Historique. +geo-analyzer: same families with GEO names; envelope `GEO AGENT RESULT` ++ `## SECTION FOR SEO.md §7` (7.1-7.6); `**Score GEO** : XX.X / 20` +(handover parses it only inside SEO.md, allow_fallback=no); G1-G7 +batches (G1-G4/G6 AUTO · G5 GATED · G7 USER). Both: STEP NUMBERS are +addressed by dispatchers (seo: 2-5/6-11/12-14; geo: 0-5/6-12/13-15; +also depth-matrix.md:17-19,37) — renumbering re-points dispatch prompts. +Engine interfaces: fetch.sh verbs {crux,queries,inspect,cannibal, +sitemap,rendercheck,linkgraph,score,schema_gen,content_quality} · +url-guard.sh host|url · source-scope.sh findargs|list · resources/*.md. + +## 3-4. Site classification counts + +| | seo | geo | total | +|---|---|---|---| +| A machine-parsed contract | ~52 | ~41 | ~93 (12 test-locked) | +| B safety/policy invariant | ~30 | ~31 | ~61 | +| C process choreography | ~21 | ~12 | ~33 | +| D other/domain-fact | ~20 | ~13 | ~33 | + +### Class C sites — seo-analyzer.md (rewrite targets) +:61 "First action." · :143-148 CMS-detect-before-edit ordering · +:208-210 "keep the two consistent" (runtime cross-file reconcile) · +:508 "run this BEFORE anything else in STEP 5" (ordering; the refusal +rule itself is B) · :550-553 "Record the denominator BEFORE sampling" +(ordering; honesty rule is B) · :602-604 "Sanity-check the grouping +before trusting it" (self-verify) · :606-618 sampling-method essay · +:661-680 C1a 20-line rationale (rule itself is B at :1493-1501) · +:875 per-item method · :970-971 "Run it twice on the same file before +publishing" (exact BDR-081 over-verification class) · :1147 "AUTO items +are a commitment, not a suggestion." · :1149-1157 P0 CMS-plugin-first +mandate · :1159-1162 P0 Bing mandate (dup of geo :777-786) · :1217 "Do +not proceed to STEP 12 until this plan is printed." · :1260-1261 + +:1342-1350 + :1502-1503 landing-page rule ×3 · :1309-1320 bundle +completeness checklist (10 checkboxes self-audit) · :1504 "Preserve +existing valid SEO." · :1522-1523 WebSearch-on-FULL extra-verify · +:1525-1526 "Transparency. Every automated change logged" (VESTIGIAL — +agent applies nothing, pre-BDR-061). + +### Class C sites — geo-analyzer.md +:48 "copy these patterns" · :124 "First action." + :127-139 ask-block +(unreachable when dispatched) · :230 conditional skip · :262-269 + +:873 + :1063-1065 PERMISSIVE default ×3 · :360 ordering · :394 "20-50 +real customer questions (P0)" · :777-786 MANDATORY AI-index submission +(dup of seo :1159-1162) · :811 "Consolidate EVERY finding" · :823 +"Print the plan before STEP 13" · :1102-1103 WebSearch extra-verify · +:1106 "Every automated change logged in §14" (VESTIGIAL; §15 log is +dispatcher's per :959). + +### Class B anchors (keep obligation, dedup emphasis) +CWD/TARGET MISMATCH twins (seo :117-126 ≈ geo :173-181) · url-guard +mandatory (seo :287-291 ≈ geo :273-277) · R2 refuse-to-score (seo +:519-548, geo :548-557; BDR-072) · NAP direction rule (seo :801-812, +geo :1073-1087; LRN-032-zenquality) · COVERAGE mandatory (seo +:1110-1130, geo :725-729; LRN-133) · never-apply/L1 (BDR-061; LRN-105 +named-ban) · C1a build-output ban · no-invented-content/DGCCRF · +"Compute the scores, do not feel them (I7)" (BDR-073) · §14 mandatory +disclosure lines (backlinks BDR-071, security headers I4) · honest +llms.txt framing · cite-sources (LRN-131). + +## 6. Duplication map (sweep ALL twins — LRN-113) + +seo internal: never-apply ×4 (:1227-1234, :1352-1357, :1468-1472, +:1527-1528) · landing-page ×3 (:1260, :1342, :1502) · shared-file +discipline ×2 (:1254, :1486) · bundle self-containment ×2 (:1249, +:1473) · COVERAGE ×4 (:438, :1000, :1096, :1110) · security-headers- +not-scored ×3 (:281, :977, :994) · 30/70 ×3 (:397, :614, :1165) · +sentinel-verbatim ×3 (:1301, :1304, :1397). +geo internal: PERMISSIVE ×3 · never-apply ×4 (:826-832, :842-848, +:1031-1034, :1104-1105) · tier-mapping ×2 (:824, :850) · +content_quality-advisory ×2 (:584, :622) · shared-file ×2 (:858, +:1047) · llms-honest ×2 (:337, :1066) · cite-sources ×2 (:17, :1089). +Cross-agent twins (stay twins — both files dispatch standalone): +CWD block · url-guard block · MODE DETECTION · MODE BOUNDARY · R2 · +COVERAGE · NAP rule · RULES section skeleton · C1a · automation rule · +Bing/AI-index action · CDN/WAF check. +Agent↔dispatcher duplication (stays — dispatch prompt is per-run +context, agent spec serves standalone/no-MODE paths): NAP ×4 total · +shared-file ×7 · security-headers ×5 · domain split · weights 80/20- +75/25 · Historique · never-re-derive (test-locked dispatcher side). + +## 7. Contradictions / ambiguities found + +1. seo :1525-1526 + geo :1106 vestigial "automated change logged" + (agent applies nothing; geo :959 says dispatcher fills §15). +2. Ask-the-user blocks unreachable in dispatched path (seo :64-75, + :88-112; geo :127-139, :153-168); /geo:41 states it outright. +3. Collect boundary wording: agents "STEP 0-5 ONLY" vs /seo "STEP 2-5 + only (context replaces STEP 0-1)" — works by prompt override. +4. geo judge does live work (sameAs curls :477-492, web_search) unlike + pure-judgment seo judge — asymmetric split, by design. +5. /harden imposes its own output contract (HARDEN.md, /100) the agent + spec never acknowledges; keys on "NARROW-SCOPE" in dispatch prompt. +6. "LRN-032" cite is ambiguous in THIS repo (local LRN-032 = different + lesson; the NAP lesson is zenquality's registry) — keep the + "zenquality" qualifier wherever cited. +7. geo :376-377 uncited FAQ-citation-rate claim vs geo :1089-1097 + cite-sources rule (LRN-131 failure class). +8. Score-label parse fragility: client-handover extract_score fallback + greps FIRST X/20 in file — losing the `Score SEO` label would + silently read `TRAJECTORY TO 17/20` as 17.0. (Latent, downstream.) +9. GEO scoring has no deterministic engine (score.py covers SEO axes + only) — BDR-073 binds only half the pair. + +## 8. Binding memory (from the analyzer's read-before) + +IN FORCE: BDR-081 (premise) · LRN-139 (when-guidance shape) · BDR-061 +(bundle+sentinel decision) · BDR-077 (mode split, fail-closed, locks +survive) · BDR-073 (deterministic scoring) · BDR-072 (R2 refuse) · +BDR-071 (off-page ceiling + §14 line) · BDR-010/LRN-011 (labeled +scores gate) · LRN-133 (omission legible) · LRN-131/132/EVAL-025 +(WebSearch ≠ verification) · LRN-105 (named ban stays explicit) · +LRN-080/088 (measure before delete → dogfood) · LRN-113 (sweep whole +surface) · LRN-093 (no vacuous locks; single-line anchors) · +LRN-126/137 (mode split carries data paths) · BLK-017 (Bing deferred). + +## 9. Open questions → dispatcher decisions (see plan §4b) + +Q1 freeze scope · Q2 census extension · Q3 dedup strategy · +Q4 vestigial lines · Q5 /harden //onboard reconciliation. diff --git a/.claude/tasks/plans/2026-07-30-seo-geo-deprescription-1402.md b/.claude/tasks/plans/2026-07-30-seo-geo-deprescription-1402.md new file mode 100644 index 0000000..2664fc2 --- /dev/null +++ b/.claude/tasks/plans/2026-07-30-seo-geo-deprescription-1402.md @@ -0,0 +1,342 @@ +# PLAN v2 — De-prescribe seo-analyzer.md + geo-analyzer.md for Opus 5 + +Date: 2026-07-30 · Branch: feature/seo-geo-deprescription (off develop, started) +KIND: build-plan · Author: main-loop session (Fable 5) +Parent decision: BDR-081 N5 (deferred as separate project) · Method: LRN-139 +v2: revised after the 3-lens challenge (§5bis) — every BLOCKER closed by a +named change; one confirmation challenger pass follows before execution. + +## 1. Context & evidence (v2 — sizing corrected per simplicity#1) + +Both agents are opus-pinned (BDR-076) → every judge phase runs Opus 5. +BDR-081 profile applies: literal following, over-verification when told +to verify, conflicting/duplicated rules burn reasoning tokens. These are +the LONGEST agent files in the repo (1528 + 1106 l) with real downstream +parsers — NOT the densest (measured: ~4.5 directive hits/100 l, ranks +20th/22nd; security-auditor is 17/100). What this pass buys, honestly: +(a) removal of self-output-verification demands (the one pattern the +baseline dogfood caught live: the judge reported "run twice, identical +output" — seo:970 firing), (b) removal of vestigial pre-BDR-061 lines +and 2 real contradictions, (c) small same-audience/same-range dedup, +(d) caps→when-guidance on choreography. The verification apparatus +(census + 3-lens challenge + before/after dogfood) is USER-DIRECTED for +this chantier, not derived from the density premise. + +## 2. Contract surface (v2 — split per correctness#5) + +### 2a. Machine-parsed (named non-LLM consumer: test, script, or literal +grep in a dispatcher step) — byte-frozen +- `model: opus`, `MODE: collect|judge|template`, `COLLECTION COMPLETE` + (model-routing.test.sh:67-68,150-157). +- `fetch.sh crux|queries` + `Performance GSC` (seo), `fetch.sh + schema_gen|content_quality` (geo) (seo-data.test.sh:538-543). +- `SEO|GEO JUDGE — VERDICT: ERROR(` — dispatcher ERROR CONTRACT + fail-closes on it (skills/seo:316-318, skills/geo:65-68). +- `## FIX BUNDLE` + sentinel `READY TO APPLY — awaiting dispatcher + confirmation` — apply step keys on it (skills/seo:524, skills/geo:101; + reused by /harden:366). +- `.audit/-signals-.md` names + fail-closed load. +- STEP numbering: dispatchers address ranges literally (seo 2-5/6-11/ + 12-14; geo 0-5/6-12/13-15; depth-matrix:17-19,37). +- `**Score SEO** : XX.X / 20` / `**Score GEO** : XX.X / 20` labels — + client-handover-writer.md:344-345 labeled grep (BDR-010/LRN-011); + losing the SEO label silently falls back to first-X/20-in-file. +- Bundle item fields `id: applier: files: current: expected:` — pasted + verbatim into hotfixer/feater at L1; `applier: bash` run in-loop. +- url-guard call sites: seo-analyzer.md:287-295, geo-analyzer.md:273-280 + (NOT ":257" as v1 said — robustness#5) + sitemap-URL guard seo:573-582. +- `NARROW-SCOPE` keying of the I4 carve-out (seo:981-983) — /harden's + dispatch prompt relies on it. + +### 2b. LLM-convention contracts (no code consumer; the dispatcher LLM +merges by these shapes) — locked in the census, still frozen +`SEO|GEO AGENT RESULT` envelopes · `## SECTION FOR SEO.md §N` · +`## ENTRIES FOR SEO.md` · `SEO|GEO SCORING (` blocks + `COVERAGE +SOURCE`/`COVERAGE LIVE` lines + `GLOBAL (weighted)` · `TRAJECTORY TO +17/20 (code-only)` · `FIX PLAN (` (seo) · batch labels A-F / G1-G7 +(tier recognition tolerant, labels nominal) · `COLLECT REPORT` + +`STATUS: DONE|BLOCKED` · `Automatisation possible avec:` · §0-§15 +report skeleton + Historique. CROSS-AGENT NOTES emit-instruction lives +in /seo's dispatch prompts (dispatcher-side lock only). + +## 3. Class B invariants — obligation kept, single strongest statement; +security ORDERINGS byte-frozen (robustness#5/#7) + +- Guard-first orderings, frozen verbatim: seo:287-291 / geo:273-277 + ("Guard the domain before it reaches a shell… Run the guard FIRST… + never 'clean up' the value and retry") + seo:573-582 URL loop. +- seo:550 "Record the denominator BEFORE sampling" — the ordering IS + the honesty mechanism (a post-hoc denominator is self-serving); + frozen; only surrounding prose may compress. +- NAP direction rule (LRN-032-zenquality — keep the qualifier, the bare + ID is ambiguous in this repo), R2 refuse-to-score (BDR-072), COVERAGE + obligations (LRN-133 — note :436-439 is a DISTINCT index-reach + obligation, not a repeat), §14 mandatory disclosure lines (BDR-071 + backlinks verbatim line, I4 security-headers), never-apply/L1 + (BDR-061; LRN-105 named ban), C1a build-output ban, no-invented- + content/DGCCRF, deterministic scoring (BDR-073), fail-closed judge, + shared-file Edit-not-Write discipline, honest llms.txt framing, + cite-sources (LRN-131). +- External-freshness checks are NOT self-verification (robustness#6): + seo:1522-1523 + geo:1102-1103 verify a DRIFTING WORLD feeding an + AUTO-tier robots.txt edit — kept, reworded as when-guidance ("crawler + lists shift; cross-check before emitting G1 from the dated resource"). + +## 4. Work items v2 + +- P0 SEQUENCING + LIVE-TREE EXPOSURE (robustness#4, conf#2/#3/#4/#9): + agents/ resolves through ~/.claude symlinks to the WORKING TREE — + edits are live between Edit calls, before any commit. Rules: + (1) the FULL baseline completes before the first agent edit — + signals + judge reports + TEMPLATE envelopes + merged SEO.md + + HUMAN-ACTIONS.md (conf#2: without frozen template artifacts the + template-range edits would have no differential and P0 makes one + unobtainable later); + (2) all baseline artifacts copied to the DURABLE, gitignored + `.audit/dogfood-baseline/` in this repo before the first edit + (conf#9: the session scratchpad dies with the session/reboot; + LRN-124: .audit/** is never committed); + (3) freeze window: no /seo //geo //harden //onboard AND no + /client-handover (spawns /seo — conf#3) nor any skill transitively + dispatching either analyzer, in ANY project, until the after-dogfood + verdict; + (4) aborts (conf#4): mid-reword interrupt or after-dogfood failure → + `git checkout HEAD -- agents/seo-analyzer.md agents/geo-analyzer.md` + (in-flight revert, index-safe); `git checkout develop -- agents/…` + is reserved for a WHOLE-BRANCH abandon; after an abort the named + exit is either (a) fix + re-run the after-dogfood, or (b) present + the static evidence (census + git diff review) to the human who may + accept or abandon at the gate — no open-ended reverted state. +- P1 CENSUS (commit 1, test-only, green pre-reword — compatible with + §7's same-commit rule: it locks EXISTING state and changes no agent + file; reword commits carry any census DELTA): DONE in working tree — + lib/tests/seo-geo-contract.test.sh 54/0, shellcheck clean, real + flip-test run: 7 scratch mutations → 7 FAILs (not "by construction" — + robustness#10). File-qualified locks (correctness#4): `FIX PLAN (` + + `applier: bash` + `Score SEO` seo-only; `Score GEO` geo-only. + Incidental locks dropped (CROSS-AGENT NOTE agent-side, bare + `applier:`). Item fields locked both files. v3 (conf#5): EVERY + `## STEP n —` header locked, interiors included (seo 0-14, geo 0-15) + — census now 71/0. Freeze mechanism for the + ~40 A-sites the census does not cover: reviewed `git diff -U0 + agents/*.md` on each reword commit (simplicity#4). +- P2 REWORD seo-analyzer.md (commit 2): + (a) Self-OUTPUT verification, v3 (conf#1/#8 — neither is deleted + outright): :970-971 "run it twice" → when-guidance integrity + guard ("if the findings JSON changed after scoring, re-run and + explain the move" — score.py is deterministic, so a moving + output means mutated findings: anti-score-shopping, BDR-073; + the unconditional double-run the baseline judge burned goes + away, the guard stays); :1217 "Do not proceed until printed" → + when-guidance scoped to the single-shot path ("single-shot runs + print the FIX PLAN before STEP 12 serializes it" — MODE: judge + stops at 11, but /harden //onboard execute the whole file, + conf#1). The completeness checklist :1309-1320 is NOT deleted: + its routing rows (stock-photo→GATED(E), compression→AUTO(bash) + or §11, aggregateRating→AUTO(hotfixer), structural→GATED(D)…) + are unique routing content (robustness#3) — reshape into a plain + mapping table, drop only the checkbox self-audit framing. + (b) DELETE vestigial :1525-1526 (contradicts BDR-061; Q4). + (c) DEDUP under the invariant (correctness#1 + robustness#1): only + VERBATIM same-AUDIENCE (spec rule / bundle-item payload / + phase-local caveat) same-MODE-RANGE (collect 0-5 / judge 6-11 / + template 12-14 / RULES=global) repeats merge. Expected survivors + per family listed at execution in the commit message; honest + net: never-apply 4→3 (RULES pair merges; template-range + statements stay), sentinel-verbatim reminders 3→2, landing-page + 3→2 (payload instance :1260 + one spec statement; :1342 vs + :1502 merge), bundle-self-containment 2→1 (same range). + NOT deduped (v1 was wrong — distinct rules or cross-range): + COVERAGE ×4, 30/70 ×3, security-headers ×3, shared-file + discipline (payload vs spec audiences). + (d) SOFTEN caps/orderings to when-guidance, keeping semantics: + :61, :508 (gate stays before on-page scoring; emphasis drops), + :875, :1147, :1149-1157 CMS-plugin-first folded together with + :143-148 into ONE statement (correctness#3 — two strengths of + one rule otherwise), :1159-1162 Bing (content rule kept, caps + drop; FULL-only → statically verified), essays :606-618 + + :661-680 compressed keeping the rule + LRN citations; :602-604 + kept as a when-guidance failure detector ("families ≈ URLs → + the heuristic broke — say so"), not deleted (robustness#8). + (e) Dispositions completing the C-list (correctness#3): :208-210 → + static pointer ("the CDN/WAF twin check lives in geo STEP 4"); + :1504 KEEP as-is (one-line scope guard). +- P3 REWORD geo-analyzer.md (commit 3), same invariant: + PERMISSIVE ×3: ALL survive (collect/template/RULES ranges; + :873 is the item-level default guarding an unconfirmed AUTO + robots.txt edit — named survivor, robustness#9). never-apply 4→3 + (RULES pair merges). tier-mapping :824/:850 BOTH stay (judge vs + template ranges). content_quality-advisory 2→1 (same range). + shared-file 2× stays (payload vs spec). llms-honest 2× stays + (collect vs RULES). cite-sources 2× stays (:17 guards the header + stats specifically). :1106 vestigial → reworded to the truth + (dispatcher fills the log — matches :959; Q4). :777-786 caps → + plain content rule (FULL-only). :1102-1103 → freshness + when-guidance (kept — §3). :394 quantity softened ("substantial, + real customer questions"). :48 softened. :124-139 ask-block KEPT + (standalone path). Orderings :360/:811/:823 softened. :376-377 + uncited claim → honest framing (no invented source). +- P4 DOGFOOD AFTER (v3 — ordered by decisiveness, conf#7): fresh copy + of zenquality-frozen; phases in this order so a mid-run death still + leaves the decisive evidence (billing class already realised once): + (ii-first) judges fed the FROZEN baseline signals + (.audit/dogfood-baseline/) → judge reports vs frozen baseline judge + reports, ZERO collect variance — the decisive Opus-judge-prose + differential; (iii) templates on those judge reports → envelopes, + compared against the frozen BASELINE envelopes — the template + verdict anchors on ENVELOPES only (SEO.md/HUMAN-ACTIONS.md are + dispatcher-merged by this authoring session, non-attributable — + conf#10); (i-last) fresh collects, same pre-answered context → + (a) shape check of signals/COLLECT REPORT vs baseline, (b) + FIELD-LEVEL diff of the fresh signals vs baseline signals (record + blocks, COVERAGE counts, denominators — a shape-valid file with a + dropped field must be caught, conf#6), and (c) ONE end-to-end seo + judge on the FRESH signals (the domain with the most collect-range + edits) so the reworded collect→judge handoff runs at least once. + Comparison mechanical-first: presence-assertion script (named home: + `.audit/dogfood-baseline/assert-after.sh`, session-reproducible, + never committed — conf#11) + a FRESH reader agent diffing + before/after WITHOUT this plan in context (correctness#7); the + authoring session only arbitrates its report. If the after-run dies: + P0(4) abort + named exit applies; no merge request meanwhile. +- P5 GATES: make test full suite (census + model-routing + seo-data + + no-vacuous-locks) · shellcheck on touched .sh · per-RANGE grep sweep + for every deduped family (asserts the named survivor lines exist in + their ranges — mode-blind ≥1× sweep is insufficient, robustness#1) · + MEASURED deltas recorded (simplicity#7): wc -l + directive-token + census (annex §0 grep set) per file, before/after, into the BDR. + (v1's manual MODE/STEP sweep dropped — the census asserts it, + simplicity#5.) +- P6 CAPITALIZE: BDR (decision, invariant, deltas, alternatives), LRN + (audience×range dedup invariant — reusable), journal, CHANGELOG. + TODO C1 checked. NO merge (human gate). Checkpoint report includes + the DYNAMICALLY-UNVERIFIED list (§6bis). + +## 4b. Dispatcher decisions (v2) + +- Q1 freeze scope: all §2a byte-frozen + §2b frozen via census; the + remaining unlocked A-prose freeze = per-commit git diff review. +- Q2 census: done (P1), flip-proven. +- Q3 dedup: WITHIN-file, same-AUDIENCE, same-MODE-RANGE, verbatim + repeats only. Cross-agent + agent↔dispatcher twins stay. (Mechanism + note correcting robustness#1's premise: every dispatch loads the FULL + agent file; the risk is ATTENTIONAL — a literal-following model told + "run STEP 13-15" deprioritizes guidance scoped to another step's + body — not access. Same fix either way.) +- Q4 vestigial: seo :1525-1526 DELETE; geo :1106 REWORD to + dispatcher-owns-log (correctness#6 resolved). +- Q5 /harden //onboard: out of scope (N6); their dispatch-prompt + contracts are untouched by agent-file rewording; `NARROW-SCOPE` + keying frozen (§2a). + +## 5. Dogfood protocol (v2) + +Baseline (DONE for collect+judge SEO; geo judge in flight at v2 time): +frozen zenquality copy (no .env), inline pipeline (canonical /seo shape +— the nested-CLI attempt died on the CLI monthly spend limit, recorded), +absolute PROJECT ROOT in every dispatch, `/seo local conservative`, +STEP 0 pre-answered, NAP = NAP-KIT.md (user-confirmed 2026-07-10). +Baseline artifacts frozen under the DURABLE `.audit/dogfood-baseline/` +(gitignored, never committed — conf#9): signals ×2, judge reports ×2, +template ENVELOPES ×2, merged SEO.md, HUMAN-ACTIONS.md (conf#2 — the +template phase runs to completion BEFORE the first agent edit). +After-run per P4. LIMITS stated honestly +(robustness#2): conservative never enters STEP 1b/1.5 (no applier parses +an item this run — the item-field contract is census-locked statically); +LOCAL never executes STEP 3-4/6-7 FULL branches (Bing/AI-index emission +text, live checks — the FULL-only conditionals were exercised and +correctly declined in the baseline judge). These stay on the +§6bis unverified list for the human gate; a FULL/aggressive dry-run is +an OPTION the user may order at checkpoint, not part of this plan. + +## 5bis. CHALLENGE SYNTHESIS (2026-07-30) + +Verdicts: correctness FATAL(3) [1 BLOCKER, 2 MAJOR, 4 MINOR] · +robustness FATAL(9) [3 BLOCKER, 6 MAJOR, 2 MINOR] · simplicity +CONCERNS(3) [3 MAJOR, 4 MINOR]. All three lenses returned. Every +BLOCKER closed by a named v2 change: +- correctness#1 (audience-blind dedup) + robustness#1 (mode-blind + dedup) → §4b Q3 invariant + P2(c)/P3 rewritten + P5 per-range sweep. +- robustness#2 (dogfood can't reach riskiest edits) → §5 honest limits + + §6bis unverified list + P2(d)/P3 minimal-diff on FULL-only sites + + static census cover; FULL/aggressive run offered to the human, not + silently added (billing exposure robustness#11). +- robustness#3 (routing table misfiled as self-check) → P2(a) keeps + routing rows verbatim. +Majors adopted: R4 live-tree abort path (P0) · R5 url-guard anchors +corrected + security orderings frozen (§2a/§3) · R6 external-freshness +kept (§3) · R7 :550 frozen (§3) · R8 :602 kept as detector (P2(d)) · +R9 :873 named survivor (P3) · C2 folded into R2's resolution · C3 full +dispositions (P2(d)/(e), P3) · S1 §1 rewritten · S2 controlled +judge-replay (P4) · S3 mechanical presence script (P4). Minors adopted: +C4 file-qualified locks · C5 §2 split · C6 three inconsistencies +resolved (P1 note, Q4, N1 marker) · C7 fresh-reader diff · S4 diff- +review freeze · S5 sweep dropped · S6+R10 census corrected+flip-proven · +S7 measured deltas. Rejected/scoped: S1's apparatus-shrinking (the +apparatus is user-directed); R1's access premise corrected to +attentional (fix adopted unchanged). + +CONFIRMATION PASS (robustness lens, v2 → v3): FATAL(9) — 2 BLOCKER + +7 MAJOR/MINOR, all targeting the v2 amendments as asked. Closed by +name: conf#1 no-MODE single-shot → §6bis + P2(a) :1217 scoped-softened +· conf#2 missing baseline template artifacts → P0(1) full-baseline +precondition · conf#3 /client-handover freeze → P0(3) · conf#4 abort +HEAD-vs-develop + named exit → P0(4) · conf#5 interior STEP locks → +census extended to all headers (71/0) · conf#6 collect→judge seam → +P4(i) field-diff + one end-to-end seo judge on fresh signals · conf#7 +decisiveness order → P4 reordered (ii)→(iii)→(i) · conf#8 :970 +anti-score-shopping → when-guidance reword, not deletion · conf#9 +volatile baseline → durable .audit/dogfood-baseline/ · conf#10 +dispatcher-owned artifacts → envelope-anchored template verdict · +conf#11 script home named. Challenge budget exhausted (1 re-pass max): +residual risk goes to the human gate with this record. + +## 6. Explicitly NOT doing + +- N1 No dispatcher (SKILL.md) edits. +- N2 No scoring-weight, axis, or depth-matrix changes. +- N3 No model-pin changes (BDR-076). +- N4 No weakening of class-B invariants (§3 hardened in v2: security + orderings byte-frozen). +- N5 No new modes, no pipeline reshaping (BDR-077). +- N6 No /harden //onboard contract reconciliation (annex §7.5). +- N7 No collect-boundary wording fix (works by prompt override). +- N8 No cross-agent shared-resource consolidation. +- N9 No deterministic GEO score engine (annex §7.9). +- N10 No FULL/aggressive dogfood in this plan (user option at gate). + +## 6bis. Dynamically-unverified edit surface (for the human gate) + +Sites edited by P2/P3 that no dogfood run executes: FULL-branch content +(seo :1159-1162 Bing emission, geo :777-786 AI-index emission, both +freshness when-guidances), apply-path parsing (STEP 1b/1.5 — item +pasted into appliers; covered statically by census item-field locks + +frozen bundle templates), STEP 6-7 external-presence prose, and the +no-MODE single-shot path (conf#1: /harden and /onboard dispatch the +agents without a MODE line — "all steps in sequence" — so the whole +reworded body drives those runs; every never-apply and ordering +statement that path relies on keeps a surviving instance, and :1217 +is softened-scoped to it, never deleted). Mitigation: minimal diffs +there (caps→plain only), census locks, git-diff review. + +## 4c. Backlog surfaced (not this branch) + +- Score-label fallback fragility in client-handover-writer.md (can read + `TRAJECTORY TO 17/20` as 17.0 if the label vanishes) — annex §7.8. +- Stale lib/ line-number comments pointing at agent lines (annex §1). +- Baseline judge's gate observation: /client-handover 17/20 gate passes + with an open `critique` finding — "open critique = independent + blocker" is worth its own decision. + +## 7. Constraints for challengers + +- Registries append-only; census green throughout; reword commits keep + 54/0 + model-routing + seo-data locks green. +- Agent files symlink-live INCLUDING between Edit calls (P0 abort path). +- §2a byte-identical; §2b frozen; STEP numbering preserved; §3 security + orderings verbatim. +- Dedup only same-audience + same-mode-range verbatim repeats; named + survivors per family in commit messages; P5 per-range sweep. +- The judge phase is Opus 5; collect/template Sonnet — literal + following applies to all (E5 "since 4.7"). +- Baseline artifacts frozen before first edit; after-run design per P4. diff --git a/.gitignore b/.gitignore index e834484..d884dcf 100644 --- a/.gitignore +++ b/.gitignore @@ -142,6 +142,12 @@ desktop.ini # an update. The source is always re-synced, so no offline copy is needed. skills-external/frontend-design/ +# Emil Design Eng — machine-owned copy curl'd from emilkowalski/skill by +# install-plugins.sh (Step 8, when absent) and re-fetched on every update-all.sh +# run. Not vendored: tracking it produced a repo diff each time upstream shipped +# an edit. The source is always re-fetched, so no offline copy is needed. +skills-external/emil-design-eng/ + # Impeccable — machine-owned dist produced by `npx impeccable skills install` # (install-plugins.sh Step 8d, update-all.sh), pinned in plugins.lock.json. # Not vendored: the installer owns the layout and rewrites it on update diff --git a/CHANGELOG.md b/CHANGELOG.md index e7df2b7..edb9a9b 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -6,6 +6,130 @@ Format follows [Keep a Changelog](https://keepachangelog.com/). ## [Unreleased] +## [1.5.0] — 2026-09-13 + +### Added +- **Attention signals on the terminal (BDR-087)** — new + `hooks/notify-attention.sh`, wired on `Notification` (input-needed + matcher) and on `Stop` (no matcher). Returns a double BEL plus an + OSC 777 toast through the `terminalSequence` JSON field, since hooks + have no controlling TTY. Signal only: `suppressOutput`, exit 0, zero + control-flow effect, which is what separates it from the `decision: + "block"` Stop hook [[BDR-083]] refused. Each event reaches the toast + as a readable label instead of a snake_case type; events needing no + attention (`agent_completed`, `auth_success`) exit silently; a turn + that ends with `background_tasks` still running stays quiet and + signals at the real end. Client-side prerequisites over Remote-SSH + are documented in the script header ([[BLK-020]]): VS Code + `accessibility.signals.terminalBell` for the beep, an OSC notifier + extension for the Windows toast. +- **User permanent rules (BDR-085)** — three new rules/ files from the + user's rule text: `writing-style.md` (always-on: em-dash ban, no slop + vocabulary, no hedging chains, deliverable self-check), + `web-building.md` (path-scoped: design anti-defaults + public-site done + checklist), `web-security.md` (path-scoped: RLS, service-key/client + split, IDOR, cookie flags, rate limiting — extends §Security, no dup). + Project CLAUDE.md rules/ doctrine gains the 320-budget exception. +- **/tour multi-project parallel fan-out (BDR-084)** — two or more + project paths now dispatch one runner per repo in a single message + (independent working trees, nothing collides) instead of processing + them one by one. The runner inherits the session model (no pin — it + carries tour's reflection); every agent inside keeps its defined tier. + A dead runner surfaces as an explicit `RUNNER FAILED` summary row; the + gated capitalize offer stays in the main loop. Bounded LRN-083 + derogation recorded in BDR-084. Census §12: 6 locks, flip-tested. + Mechanics proven first: nested probe, 3 overlapping agent windows, + 9.1s vs ~18s sequential. +- **Contract gates — deterministic floor under the fresh verifier (BDR-083)** — + an acceptance criterion can now carry an oracle (`CHECK:` command + + `EXPECT:` success-only marker + `EVIDENCE:` slot). `lib/gates.sh run + ` executes them fail-closed — MET requires exit 0 **and** the + marker — and writes the outcome back into the contract, so the fresh + verifier reads evidence as fact instead of trusting the executor's report. + New `GATE 0` in `lib/verify-secure-loop.md` runs the floor before any + verifier is dispatched: a red build no longer costs an LLM dispatch to + discover. `ABANDON: ` makes an impossible criterion a visible + handoff that blocks `CONFORME` and routes to the human gate (new verifier + verdict `ABANDONED(n)`). `feater` and `bugfixer` gain a four-pass + completion discipline, scoped so it can never widen the contract. + Adapted from the `unlazy` skill (Leonxlnx/unlazy, MIT); its Stop hook, + approval store, `.unlazy/` tree, depth-tree arithmetic and Node checker + were deliberately refused — see BDR-083 for each reason. + The four orchestrator skills (`feat`, `bugfix`, `ship-feature`, + `init-project`) restate the GATE 0 bullet ahead of GATE 1 (locked); + hotfix explicitly runs no floor. Behavioral RED: 16/16 fresh unprimed + runs followed the new doctrine (EVAL-027). + 64 new assertions in `lib/tests/gates.test.sh`. +- **`lib/tests/seo-geo-contract.test.sh`** — census locking the seo/geo + agent ⇄ dispatcher machine contract: judge verdict grammar, FIX BUNDLE + + READY-TO-APPLY sentinel, signals handoff, every STEP header (interiors + included), bundle item fields, score labels, scoring blocks, envelope + keys (46→71 assertions across the C1 chantier). + +### Changed +- **Skill and agent quality campaign, 54 units (BDR-086)** — full darwin + v2.1 pass over the 31 personal skill-systems and 23 agents, excluding + the gstack/external symlinks and machine-owned units. Fresh baseline + mean 83.4; the 13 units under the user-set threshold of 80 were + optimized to completion, and verified defects in above-threshold units + were fixed in a grouped pass rather than left to ship because the score + was good enough. Every round was validated by a paired 3-judge majority + reading before and after in one call: 36 unit-round verdicts, 24 batch + verdicts, all better, zero reverts. Full report and residual findings: + `.claude/audits/DARWIN-2026-08-26.md`. +- **seo-analyzer + geo-analyzer de-prescribed for Opus 5 (BDR-082)** — + process choreography converted to when-guidance under an + audience×mode-range invariant; self-output verification demands removed + (the score-engine "run it twice" became a conditional integrity guard); + two pre-BDR-061 vestigial rules fixed; P0/MANDATORY/ALWAYS caps softened + to plain content rules. Machine contract byte-frozen and locked by the + new `lib/tests/seo-geo-contract.test.sh` census (71 locks, flip-proven); + proven by a controlled before/after `/seo` dogfood — judge replay on + frozen signals, 42/42 presence assertions on both runs, blind structural + reader: interchangeable, recall improved. +- **Global instruction layer recalibrated for the Claude 5 family (BDR-081)** — + delegation block is now model-neutral when-guidance (the Opus 4.8 + under-delegation counter inverted on Opus 5, which over-delegates and gets + an injected harness cap); "staff engineer" self-check bar dropped (Opus 5 + over-verification trigger); finish-whole-task clause added to Deviations; + written-deliverable length rule added. 308/320 lines. +- **Default session model is now `opus[1m]`** (was `claude-fable-5[1m]`). +- **`skills-external/emil-design-eng/` untracked** — the file is curl'd + from upstream by `install-plugins.sh` when absent and re-fetched by + every `update-all.sh` run, so tracking it produced a repo diff on each + upstream edit. Same category as `frontend-design/` and `impeccable/`, + already ignored on that rationale; a fresh clone re-fetches it. + `design-motion-principles/` has the same overwrite behaviour but no + bootstrap clone yet, so it stays tracked until that gap closes. + +### Fixed +- **hotfix wiped tolerated in-progress edits on its revert path** — every + failure branch ran `git restore .`, destroying user edits the run had + tolerated. Now a `git stash create` pre-flight snapshot plus a + file-scoped restore, and the security gate is fresh-dispatch only. +- **skills-perso listed 8 of 31 personal skills** — detection rebuilt on + the `link.sh` symlink convention (symlink = external, real dir = + personal, gitignored = machine-generated). Live result 31/31, no false + positives. +- **plan-challenger** — `ERROR` joined the load-bearing verdict grammar + (STEP 1 emitted it, the parser enum omitted it); grounded-but-uncertain + findings now file as `[MINOR]` with the uncertainty stated, instead of + being self-censored (Opus 5 follows conservative-reporting clauses + literally). +- **design-toolchain hook** — dropped `\bux\b` (2 French-prose false + positives; 3rd tightening pass, series LRN-1005/1007); `\bui\b` kept and + locked by a must-fire test row. +- **Agent and skill defects found by the campaign's judges** — + `init-project` allowed-tools lacked `Agent` and `Skill` while every step + dispatches; `commit-change` conflict grep now covers all 7 unmerged + codes; `tour --report-only` no longer commits; `harden` severity defers + to the calibrated guide and the late SSL Labs grade has an assigned + actor; handover writers' stale chapter refs corrected and the anchor + gate ordered; `security-auditor` documents the hotfix no-verifier + carve-out; `close` enumerates STEP 5C and passes `--no-push` through; + `prune-memory` drops a false "v1-untested" note; `code-clean` attributes + its executor correctly; plugin-check and onboard fixtures de-drifted. + ## [1.4.0] — 2026-07-22 ### Added diff --git a/CLAUDE.global.md b/CLAUDE.global.md index 3aa4c2d..3ab8178 100644 --- a/CLAUDE.global.md +++ b/CLAUDE.global.md @@ -22,6 +22,8 @@ Apply unless repo-specific instructions override. - Document intent, not mechanics. Use project doc style (docstring, JSDoc…). - Explicit, consistent, meaningful names. Straight control flow, no hidden side effects. +- Written deliverables (docs, reports, .md): length matched to what + the task needs — no filler sections, no boilerplate summaries. ## Refactoring - Priority: safety → readability → consistency. @@ -40,11 +42,12 @@ Apply unless repo-specific instructions override. - Confirm before implementing only when real trade-offs exist (multiple valid approaches, breaking change, destructive action) — else proceed. - Minimal changes unless broader refactor requested. State trade-offs. -- Sub-agents keep main context clean — one task per sub-agent. - More compute on hard problems. Task fans out across independent - items (many files, parallel searches, multi-point checks) → delegate - to sub-agents, don't iterate serially. Default to delegation for - multi-file exploration. Counters model tendency to under-delegate. +- Sub-agents: one task per sub-agent, main context stays clean. + Delegate genuinely independent, sizeable tracks (wide multi-file + exploration, parallel audits) — not work doable in a few tool + calls. Skill-mandated gates (fresh verifier/security/challenge) + always dispatch as written. Don't redo delegated work by hand — + failed gates re-dispatch fresh executors instead. - One question upfront if needed — don't interrupt mid-task. *Exception: skill-mandated gates and checkpoints (orchestrator validation gates, approval gates, darwin checkpoints) always fire.* @@ -53,6 +56,8 @@ Apply unless repo-specific instructions override. - Something goes wrong → STOP, re-plan. Never push through. - Deviations: minor or clearly justified → do, explain after. Significant or shaky justification → ask before deviating. + Finish the whole task: blocked on an independent sub-part → do + the rest, state what's missing. Gone WRONG → still STOP, re-plan. - Root causes only. No temp fixes. Never assume — verify paths, APIs, variables before use. @@ -77,7 +82,6 @@ Apply unless repo-specific instructions override. 2. Report what verified, what not. 3. List remaining risks, surviving deviations. 4. Don't mark complete without proof it works. - Bar: "would staff engineer approve?" 5. Correction or notable event → capitalize to right registry (see "Memory registries"). diff --git a/CLAUDE.md b/CLAUDE.md index a7b5941..94091bb 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -17,7 +17,9 @@ A rule WITH `paths:` YAML frontmatter (glob list) loads lazily — only when Claude reads a file matching a glob; a rule WITHOUT it loads at session start, same cost as the global memory. Extract from CLAUDE.global.md only what can be path-scoped (the token win) or what is generated; always-on -doctrine stays in CLAUDE.global.md. `paths:` globs match against the +doctrine stays in CLAUDE.global.md. Exception: a standalone user-authored +rule set that would bust the 320-line density budget may live here WITHOUT +`paths:` (always-on load) — writing-style.md (BDR-085). `paths:` globs match against the CURRENT project's tree — a broad glob (e.g. `rules/**`) can fire in foreign projects; keep rule bodies tiny. Docs: https://code.claude.com/docs/en/memory.md#path-specific-rules diff --git a/agents/analyzer.md b/agents/analyzer.md index 135d405..7eccc81 100644 --- a/agents/analyzer.md +++ b/agents/analyzer.md @@ -25,14 +25,13 @@ Produce a clear analysis without proposing solutions. --- -## TASKS +## TASKS (in order — each step feeds the OUTPUT section named) -- Identify relevant parts of the codebase -- Understand current behavior -- List dependencies -- Highlight constraints -- Detect risks -- Identify ambiguities +1. **Locate** — find the relevant parts of the codebase (Glob/Grep from the target) → file list +2. **Understand** — read them; describe current behavior as-is → CONTEXT, KEY COMPONENTS +3. **Map dependencies** — imports, call sites, data flow in/out → KEY COMPONENTS roles +4. **Constrain** — invariants, contracts, conventions the code obeys → CONSTRAINTS +5. **Assess** — risks with probability, then ambiguities → RISKS, OPEN QUESTIONS --- diff --git a/agents/bugfixer.md b/agents/bugfixer.md index c1771ab..f24cd33 100644 --- a/agents/bugfixer.md +++ b/agents/bugfixer.md @@ -46,6 +46,24 @@ Every choice was made in the plan or is a NEED-DECISION to report. security/verifier dispatch, editing `.claude/**` or memory registries, user questions (you cannot ask — report instead), attribution trailers of any kind. +## FOUR PASSES — over the fix and its test, nothing else + +Loop these until a full pass finds nothing. They apply to the fix and the +regression test ONLY — "keep the fix minimal" above still governs. They make +the minimal fix COMPLETE; they never widen it. + +1. **Complete.** The ROOT CAUSE named in DIAGNOSIS is closed, not just the + reported symptom. No placeholder, no deferred remainder. +2. **Expert reread.** Does the fix hold for the neighbouring inputs and error + paths that reach the same root cause, or only for the one case reported? +3. **Negative control.** Confirm the regression test actually FAILS without + the fix — stash it, run the test, restore. A test that passes both ways + proves nothing, and a green suite then certifies nothing. +4. **Polish.** Naming and comments on what you touched. Nothing else. + +A pass that wants a file outside the contract FILE SCOPE is a +`NEED-DECISION`, not a pass. + ## OUTPUT — end with exactly this report (your final message) ``` diff --git a/agents/client-handover-writer.md b/agents/client-handover-writer.md index 6a71e36..c8883d1 100644 --- a/agents/client-handover-writer.md +++ b/agents/client-handover-writer.md @@ -732,7 +732,7 @@ write `.claude/audits/THRESHOLD-OVERRIDE.md` documenting: - Top 3 unresolved issues per axis - User's stated reason -This file is referenced in §4 of the client doc ("Ce qui vous reste à faire") +This file is referenced in §5 of the client doc ("Ce qui vous reste à faire") so the client knows what's still below the bar. If `ALL_PASS = false`: diff --git a/agents/feater.md b/agents/feater.md index 6d43318..1346f59 100644 --- a/agents/feater.md +++ b/agents/feater.md @@ -57,6 +57,25 @@ report below is optional on this path (the dispatcher needs the edit applied editing `.claude/**` or memory registries, user questions (you cannot ask — report instead), attribution trailers of any kind. +## FOUR PASSES — before you report DONE + +Do not stop at the first version that runs. Loop these until a full pass +finds nothing: + +1. **Complete.** The whole deliverable the plan names is implemented. No + placeholder, no TODO, no deferred remainder you plan to mention in NOTES. +2. **Expert reread.** Read it as someone who owns this codebase. Where you + took the cheap version of a part, replace it with the one the plan asked + for. +3. **Defect hunt.** Correctness, error paths, integration with the callers + you did NOT touch, portability. Fix what you find. +4. **Polish.** Low-cost only: naming, comment density, dead code you + introduced. + +Every pass stays inside the plan and the contract FILE SCOPE. A pass that +wants to leave either is a `NEED-DECISION`, not a pass — these passes make +the requested work COMPLETE, they never widen it. + ## OUTPUT — end with exactly this report (your final message) ``` diff --git a/agents/geo-analyzer.md b/agents/geo-analyzer.md index 8c63520..f23e0e7 100644 --- a/agents/geo-analyzer.md +++ b/agents/geo-analyzer.md @@ -45,7 +45,7 @@ This anchors the agent's output so the user can compare audits over time. effort : weight: <1-5> ``` -Worked examples (1 per axis, copy these patterns when reporting): +Worked examples (1 per axis — the reporting shape to match): ``` [HIGH] [ai-crawlers] GPTBot blocked in robots.txt @@ -391,7 +391,7 @@ Emit finding: FAQ PAGE : present at | absent FAQ SCHEMA : FAQPage (collection) | QAPage (single Q) | none Q&A COUNT : | not applicable -RECOMMENDATION : CREATE /faq with 20-50 real customer questions (P0 for GEO) | ADD schema to existing page | OK +RECOMMENDATION : CREATE /faq with real customer questions (typically dozens — high GEO priority) | ADD schema to existing page | OK ``` If absent and site is informational/service/B2B → emit as MEDIUM-term @@ -774,9 +774,8 @@ High-impact, low-effort. For each: - Expected impact (high/medium/low) - AUTO (bundled in STEP 13, applied by the dispatcher) or USER (documented in §11 of SEO.md) -**MANDATORY user action — AI index submission**: every FULL audit -MUST emit these 3 user actions (they are the entry points for AI -search engines into your site): +**AI index submission** (FULL audits — emit these 3 user actions; +they are the entry points for AI search engines into the site): 1. **Bing Webmaster Tools** — submit + verify sitemap. Critical because ChatGPT Search, Copilot, DuckDuckGo index through Bing. @@ -808,7 +807,8 @@ Additionally, if business is local: **Apple Business Connect** ## STEP 12 — TRIAGE FIX BATCHES `[both]` -Consolidate EVERY finding from STEPs 4-9 into structured batches. +Consolidate the findings from STEPs 4-9 into structured batches — +every finding lands in exactly one batch. | Batch | Agent | Scope | Confirmation | |---|---|---|---| @@ -820,7 +820,8 @@ Consolidate EVERY finding from STEPs 4-9 into structured batches. | **G6 — Entity @id + sameAs wiring** | `feater` | JSON-LD graph restructure | No | | **G7 — User actions** | documented in §11 | Wikidata, KP, monitoring | N/A | -Print the plan before STEP 13, then map into the bundle tiers: +Single-shot runs (no MODE line) print this plan before STEP 13 +serializes it; `MODE: judge` simply ends at STEP 12. Tier mapping: G1–G4/G6 → AUTO, G5 → GATED, G7 → USER ACTIONS. **Apply-vs-report is the DISPATCHER's call, not yours.** You ALWAYS emit @@ -1101,6 +1102,6 @@ PROCHAINE ETAPE : `automation-catalog.md`. No exceptions. - **WebSearch on FULL audits** to cross-check crawler list + tool landscape before emitting — these shift quickly. -- **Dispatcher verifies.** Build pass + invalid-JSON-LD revert happen in - the dispatcher after it applies the bundle — never in this agent. -- **Transparency.** Every automated change logged in §14. +- **Dispatcher verifies.** Build pass, invalid-JSON-LD revert and the + applied-change log (SEO.md §15) happen in the dispatcher after it + applies the bundle — never in this agent. diff --git a/agents/handover-doc-writer.md b/agents/handover-doc-writer.md index 298881c..0575dd2 100644 --- a/agents/handover-doc-writer.md +++ b/agents/handover-doc-writer.md @@ -424,7 +424,7 @@ Wrong — has date prefix: ### 6.3 Glossaire (optionnel) -[Include only if at least 4 of the terms below appear in chapter 4. +[Include only if at least 4 of the terms below appear in chapter 6. Format: term — one-line plain-language definition. Sort alphabetically. This is the ONLY place internal tooling names may be mentioned by their internal label, and only when explaining what they correspond @@ -465,7 +465,7 @@ des audits de santé. Pour toute question, contactez [contact].* 1. Address the client directly ("votre site", "vous pouvez"). 2. Chapters 1–3: replace every tech term with a user-facing equivalent. 3. No abbreviations the client wouldn't use (HTTPS yes, CSP no — unless - in chapter 4 with definition). + in chapter 6 with definition). 4. Concrete numbers > adjectives. 5. Short paragraphs. Bullet lists for things you can count. 6. **Score deltas explained in plain words**. Never just dump numbers. @@ -535,9 +535,9 @@ The chapter must include: 8. **Outils gratuits pour vérifier votre présence**. -Cross-link this chapter from §4 (owner responsibilities — "Ce qui vous +Cross-link this chapter from §5 (owner responsibilities — "Ce qui vous reste à faire"). Items in this §7 annex that are recurring belong in -§4's cadence checklist (Mensuel / Trimestriel / Annuel). +§5's cadence checklist (Mensuel / Trimestriel / Annuel). --- @@ -624,7 +624,9 @@ checkbox: (`LANG=en`: "Items already checked have been validated.") -### Verification +### Verification (deferred — run right AFTER STEP 15 writes `$OUTPUT_MD`; +the pre-checks themselves are applied to the in-memory body here, the +file does not exist yet) ```bash # At least one pre-check expected for any project with real history. @@ -687,6 +689,9 @@ awk '/^## 1\./{flag=1} /^## 6\./{flag=0} flag' "$OUTPUT" \ **Anchor-resolution gate** (clickable section refs work). ```bash +# ORDER: run this gate in STEP 16, immediately AFTER the HTML render — +# $OUTPUT_HTML does not exist yet at STEP 15. A broken anchor found here +# loops back to fix the markdown ref, then re-render. grep -oE '\]\(#[a-z0-9-]+\)' "$OUTPUT_MD" | tr -d ']()#' | sort -u > /tmp/refs.txt grep -oE 'id="[^"]+"' "$OUTPUT_HTML" | sed 's/id="//;s/"//' | sort -u > /tmp/ids.txt comm -23 /tmp/refs.txt /tmp/ids.txt diff --git a/agents/hotfixer.md b/agents/hotfixer.md index c531917..89843fe 100644 --- a/agents/hotfixer.md +++ b/agents/hotfixer.md @@ -75,7 +75,7 @@ the edit applied + self-verified, not the report grammar). ``` HOTFIX-EXEC REPORT STATUS : DONE | BLOCKED -FILE(S) : +FILE(S) : FIX : SMOKE : NOTES : diff --git a/agents/interviewer.md b/agents/interviewer.md index bbddf47..0321318 100644 --- a/agents/interviewer.md +++ b/agents/interviewer.md @@ -14,6 +14,17 @@ Gather context. Produce complete PROJECT BRIEF as single source of truth. - If the initial prompt already provides name + purpose + stack + features + architecture → skip questions and generate the BRIEF directly. - Otherwise ask only what's genuinely missing, in a single structured block. - After answers: produce BRIEF. One follow-up allowed if answer is ambiguous. +- Hard budget: 2 question rounds total (initial block + one follow-up). The BRIEF ships after round 2 no matter what — gaps become OPEN DECISIONS, never a third round. + +## FAILURE MODES + +| Trigger | First response | If still unresolved | +|---|---|---| +| Answer vague/ambiguous | One targeted follow-up on that item only | Record item in OPEN DECISIONS with the safest reading, marked `(assumed)` — never invent a confident value | +| "I don't know / you decide" | Propose ONE concrete default + why, ask yes/no | Take the default, mark `(assumed)`, list in OPEN DECISIONS | +| Contradictory answers (e.g. embedded runtime + managed cloud DB) | Name the contradiction, ask which side wins | Put BOTH options in OPEN DECISIONS; do not silently pick one | +| Partial answer to the block | Re-ask ONLY the missing items in the follow-up round | Missing fields → `none stated` + OPEN DECISIONS entry | +| Feature list balloons (>10) | Keep the 10 the user ranks first as V1 | Overflow goes to OUT OF SCOPE with a `(deferred by budget)` tag | ## QUESTIONS (skip answered ones) @@ -60,3 +71,12 @@ OPEN DECISIONS: ``` Stop after BRIEF. Orchestrator handles next step. + +## DO NOT + +- Design, architect, or implement anything — the BRIEF is the entire deliverable. +- Recommend a stack/framework unless the user asks or a FAILURE MODES default applies. +- Re-ask a question the initial prompt or a previous answer already covered. +- Exceed the 2-round budget, whatever is still missing. +- Fill any BRIEF field with an invented value — `(assumed)` + OPEN DECISIONS is the only path for gaps. +- Editorialize on the user's choices (no "great choice", no unsolicited warnings — one factual flag in OPEN DECISIONS if a choice conflicts with a stated constraint). diff --git a/agents/onboarder.md b/agents/onboarder.md index b214c6d..21163d3 100644 --- a/agents/onboarder.md +++ b/agents/onboarder.md @@ -12,33 +12,40 @@ Generate the baseline claude-config files in a project directory. No interview, --- -## INPUTS REQUIRED (passed by orchestrator) +## INPUTS (passed by orchestrator) 1. `PROJECT_ROOT` — absolute path where files should be written -2. `BRIEF` — dict with keys filled by orchestrator STEP 1-3: +2. `BRIEF` — dict. Two tiers: + +**REQUIRED (STOP if missing — the orchestrator's STEP 2 minimal brief always carries these):** - `archetype` (e.g., "nextjs-app-router", "wordpress", "dotfiles-meta") - - `archetype_category` (cms | static | framework | api | cli | library | mobile | meta) - `project_name` - `stack` (language/framework/versions) - `purpose` (1-3 sentences) - `build_cmd`, `test_cmd`, `lint_cmd` (or "N/A") - - `folder_tree` (max 2 levels) - - `architecture_notes` - - `conventions` - - `exceptions_to_global_rules` - - `key_deps` (list with one-line purpose each) - - `workflow_notes` - - `is_monorepo` (bool) + `packages` list if true - - `monorepo_mode` ("A" | "B:" | "C") — only if is_monorepo -If any key is missing, PRINT what's missing and STOP. Do NOT invent values. +**OPTIONAL enrichment (normally `null` on first dispatch — the interview fills them at STEP 3, AFTER this agent runs):** + - `archetype_category` (cms | static | framework | api | cli | library | mobile | meta — derive from `archetype` when null) + - `folder_tree`, `architecture_notes`, `conventions`, + `exceptions_to_global_rules`, `key_deps`, `workflow_notes` + - `is_monorepo` (bool) + `packages` + `monorepo_mode` ("A" | "B:" | "C") + +Contract: +- A REQUIRED key missing → PRINT what's missing and STOP. Do NOT invent values. +- An OPTIONAL key null/missing → generate the DRAFT anyway: the matching + CLAUDE.md section gets the placeholder ``, + never an invented value. List every placeholder in OUTPUT. +- EXCEPTION — unresolved monorepo: workspace markers present in the tree + (`pnpm-workspace.yaml`, `workspaces` in package.json, `apps/`+`packages/`) + but `monorepo_mode` null → STOP. Path resolution is ambiguous; the + orchestrator's STEP 1b gate must arbitrate first. --- ## PHASE 1 — GENERATE CLAUDE.md Read `~/.claude/templates/project-CLAUDE.md` as base. -Fill sections from BRIEF. Preserve global CLAUDE.md compatibility (this file extends, doesn't override silently). +Fill sections from BRIEF; null enrichment keys become their `` placeholder. Preserve global CLAUDE.md compatibility (this file extends, doesn't override silently). Write to `${PROJECT_ROOT}/CLAUDE.md`. @@ -149,6 +156,7 @@ FILES WRITTEN: ✅ .claude/memory/evals.md (created | unchanged) ✅ .claude/audits/ (created | unchanged) [✅ ROADMAP.md] (if generate_roadmap) +PLACEHOLDERS : ``` --- @@ -158,4 +166,4 @@ FILES WRITTEN: - NO audit (handled downstream by orchestrator). - NO destructive writes: never overwrite CLAUDE.md if it exists without asking (print path + STOP, let orchestrator decide). - Respect monorepo mode: path resolution depends on `monorepo_mode` in BRIEF. -- If any BRIEF key is missing, STOP and report — do not guess. +- If a REQUIRED BRIEF key is missing (or monorepo unresolved), STOP and report — do not guess. Null OPTIONAL keys are normal on first dispatch: placeholder, don't stop. diff --git a/agents/plan-challenger.md b/agents/plan-challenger.md index e5c0fda..0e0a1c1 100644 --- a/agents/plan-challenger.md +++ b/agents/plan-challenger.md @@ -63,7 +63,7 @@ Ground EVERY finding in the plan text (quote the section) or the real code ## OUTPUT (exact format — machine-parsed by the orchestrator) ``` -CHALLENGE — LENS: — VERDICT: SOLID | CONCERNS(n) | FATAL(n) +CHALLENGE — LENS: — VERDICT: SOLID | CONCERNS(n) | FATAL(n) | ERROR() PLAN: FINDINGS: 1. [BLOCKER] — WHY: — FIX: @@ -79,13 +79,17 @@ PROOF: read files, inspected , checked plan §<…> - Report-only. Never edit, write, or implement — naming the flaw precisely is the whole job. -- No invention. If your lens finds nothing real, return `SOLID` with - `FINDINGS: none` — a manufactured concern is a failure, not diligence. +- No invention — ungrounded is noise. Silently dropping a grounded doubt is + equally a failure: file it as `[MINOR]` with the uncertainty stated in + `WHY:`. Nothing real at all → `SOLID` with `FINDINGS: none`. - `PROOF` is MANDATORY. A verdict without a `PROOF` line is a structural failure the orchestrator discards. - Stay in your lens. A finding outside it belongs to another challenger. - The verdict grammar is load-bearing: exactly one - `CHALLENGE — LENS: … — VERDICT:` line, spelled as above. + `CHALLENGE — LENS: … — VERDICT:` line, spelled as above. `ERROR()` + (STEP 1's missing/unreadable-plan verdict) is part of the grammar: it + carries only the `PLAN:` line — no FINDINGS, no PROOF — and the + orchestrator treats it as a dispatcher-side failure, not a challenge result. ## ORCHESTRATOR PROTOCOL (consumer contract — wiring reference) diff --git a/agents/plugin-advisor.md b/agents/plugin-advisor.md index 7d9141e..a7cd0df 100644 --- a/agents/plugin-advisor.md +++ b/agents/plugin-advisor.md @@ -25,6 +25,18 @@ field. PROBE REPORT missing or a field absent → emit `PLUGIN CHECK — VERDICT: ERROR(probe report missing/invalid: )` and STOP. Fail closed: no recommendations over invented detection. +`FRAMEWORK-DEPS` carries exact `"dep": "version"` pairs (or +`framework-deps-none`). Derive signal classes from those names + versions: +`frontend` = react/react-dom/vue/nuxt/svelte/astro/next present; +`fast-libs` = next, react ≥18 (version prefix), prisma/@prisma/client, +supabase/@supabase/supabase-js, drizzle-orm, expo. Never re-scan the +manifest to make this split. + +`REQUEST` MAY carry `PLAN: Max|Pro|Free` from the dispatcher. Echo it in +the output. Absent → output `PLAN: unknown (not provided)` and SKIP the +plan-budget WARN (absolute COST ESTIMATE still reported). Never assume a +plan. + --- ## PHASE 2 — ANALYZE @@ -86,7 +98,7 @@ ACTIVE: [plugin — status, one line each] PROFILE: [active skill profile — name + match%, or "custom"] SIGNALS: [detected signals] COMPLEXITY: % — -PLAN: (budget: ~t passive tokens) +PLAN: (budget: ~t | n/a) COST ESTIMATE: ~Xt passive tokens (all active plugins combined) RECOMMENDATIONS: @@ -315,7 +327,7 @@ or by applying a profile that lists it (e.g. `apply web` to restore - Active toggle plugins not needed for this task (dead passive cost) - Multi-session feature + `gsd` CLI not installed → `npm install -g gsd-pi` -- Total passive cost > 50% of plan budget (Pro: ~5500t, Max: ~10000t, Free: ~2500t) +- Total passive cost > 50% of plan budget (Pro: ~5500t, Max: ~10000t, Free: ~2500t) — only when PLAN was provided; PLAN unknown → skip this WARN - **Next.js/React 18+/Prisma/Supabase detected + context7 not configured** → Risk: Claude may generate code using outdated APIs (App Router changes frequently) → Fix: `npm install -g ctx7 && ctx7 setup --claude` diff --git a/agents/plugin-probe.md b/agents/plugin-probe.md index 16f646d..31c4eac 100644 --- a/agents/plugin-probe.md +++ b/agents/plugin-probe.md @@ -34,7 +34,8 @@ command -v rtk &>/dev/null && rtk --version 2>/dev/null | head -1 || echo "rtk-n # Project signals (run from project root) ls package.json pyproject.toml Cargo.toml go.mod 2>/dev/null | head -5 -grep -rl "next\|react\|vue\|prisma\|supabase" package.json 2>/dev/null | head -3 || true +# Exact-key dep match with versions ("react": won't match "preact":) +grep -ohE '"(next|react|react-dom|vue|nuxt|svelte|astro|prisma|@prisma/client|@supabase/supabase-js|supabase|drizzle-orm|expo)"[[:space:]]*:[[:space:]]*"[^"]*"' package.json 2>/dev/null || echo "framework-deps-none" find . -name "*.tsx" -o -name "*.jsx" 2>/dev/null | head -3 | wc -l find . -name "docker-compose*" -o -name "Dockerfile" 2>/dev/null | head -3 | wc -l @@ -72,7 +73,7 @@ EXTERNAL : PROFILE : CLIS : ctx7= gsd= rtk= MANIFESTS : -FRAMEWORK-DEPS: +FRAMEWORK-DEPS: TSX-JSX-COUNT : DOCKER-COUNT : ANIM : eligibility= installed= diff --git a/agents/refactorer.md b/agents/refactorer.md index 067256e..dceb723 100644 --- a/agents/refactorer.md +++ b/agents/refactorer.md @@ -19,9 +19,18 @@ Improve code without ever changing its external behavior. 1. Analyze the target — list ALL violations 2. Produce the report BEFORE touching anything -3. Check that tests exist (if not — report before modifying) +3. Check that tests exist covering the target. + 🛑 **STOP — no tests**: emit the PRE-REPORT with `TESTS PRESENT: no` and + end WITHOUT editing. Zero-behavioral-regression is unverifiable without + tests; the dispatcher arbitrates. Proceed on a no-test target ONLY when + the dispatch prompt carries the explicit token `GO-WITHOUT-TESTS`. + (Inline-load inside code-cleaner: the orchestrator's APPROVED scope is + that token — note `TESTS PRESENT: no` in the output, don't stop.) 4. Refactor function by function -5. Verify tests pass after each modification +5. Run the tests after each modification. + Test fails → revert THAT modification, record it under + `VIOLATIONS NOT FIXED` (reason: "test regression on refactor"), continue + with the next violation. Never leave the suite red between steps. --- @@ -60,6 +69,7 @@ TESTS PRESENT: yes / no - Zero behavioral regression - Existing tests must pass +- No tests on the target → PRE-REPORT + STOP (unless dispatched with `GO-WITHOUT-TESTS`) - Do not modify business logic under the guise of refactoring - Do not refactor unrelated parts diff --git a/agents/security-auditor.md b/agents/security-auditor.md index a7d22b5..4ec655d 100644 --- a/agents/security-auditor.md +++ b/agents/security-auditor.md @@ -147,7 +147,9 @@ In audit mode, ALSO write this same block (plus per-finding detail) to ## ORCHESTRATOR PROTOCOL (consumer contract — wiring reference) - The security gate runs AFTER the request-conformity verdict is CONFORME - (verifier), never before. + (verifier), never before — EXCEPT under /hotfix, which by design runs no + verifier: there the gate fires directly on the smoke-passed diff (its + one-attempt model reverts on BLOCK instead of looping). - Dispatch a FRESH auditor each iteration — no context reuse. Input = mode + scope + (report) + (context), nothing else. - Parse the `SECURITY — VERDICT:` line: diff --git a/agents/seo-analyzer.md b/agents/seo-analyzer.md index 9c09f16..7233632 100644 --- a/agents/seo-analyzer.md +++ b/agents/seo-analyzer.md @@ -58,8 +58,8 @@ STEP 1-2 business/tech context is consumed by all later steps). ## STEP 0 — AUDIT DEPTH -**First action.** If a parent skill (`/seo` dispatcher) passed depth -in $ARGUMENTS, use it. Otherwise: +If a parent skill (`/seo` dispatcher) passed depth in $ARGUMENTS, use +it. Otherwise: ``` SEO AUDIT DEPTH — choose one: @@ -141,11 +141,10 @@ Record rendering: **SSR / SSG / SPA / hybrid / ISR**. ### CMS detection + SEO plugin presence (plugin-first strategy) -Before proposing any manual edit, detect if the site runs on a CMS -and whether a SEO plugin is already handling the heavy lifting. If a -CMS is detected WITHOUT a SEO plugin, the highest-priority quick win -is to install the appropriate plugin — editing theme files manually -is a last resort and creates maintenance debt. +Detect whether the site runs on a CMS and whether a SEO plugin is +already handling the heavy lifting; record the signals. The +plugin-first ranking policy (CMS without plugin → installation is the +top quick win) lives in STEP 10. ```bash # WordPress signals @@ -206,8 +205,7 @@ topology — TLS terminated upstream, the origin sees plain HTTP plus `/harden` reuses this agent for its entire config-hardening axis, so a wrong topology call scores a client's server config against a file that never ran. -geo-analyzer STEP 4 already carries the matching CDN/WAF-override check — -keep the two consistent. +(The same CDN/WAF-override check lives in geo-analyzer STEP 4.) ```bash # Server / hosting @@ -505,7 +503,7 @@ Fetch rendered HTML. Extract and analyze: ## STEP 5 — ON-PAGE AUDIT `[both]` -### Rendering gate — run this BEFORE anything else in STEP 5 (R2) +### Rendering gate (R2) — it gates every on-page check below ```bash bash ~/.claude/lib/seo-data/fetch.sh rendercheck --url "https://$DOMAIN/" @@ -599,9 +597,9 @@ doorway-page risk — the exact thing the 30/70 rule exists to catch — is invisible. Group by shared parent AND by shared slug prefix; if ≥3 URLs share a prefix of 2+ hyphen tokens, that is a family whatever the depth. -Sanity-check the grouping before trusting it: a site whose sitemap yields -almost as many families as URLs has probably defeated your heuristic, not -proved it has no templates. +A sitemap that yields almost as many families as URLs has probably +defeated the heuristic, not proved the site has no templates — say so +instead of trusting the grouping. **Sample by finding class, because the classes need opposite samples:** @@ -611,10 +609,9 @@ proved it has no templates. | **Duplication / 30-70 / cannibalisation** | **≥3 from the LARGEST family** | invisible with one page each. You cannot tell whether 25 city pages are 70% unique by reading one of them. | | Per-page content (title/description length, H1 wording) | spread across families + GSC position 4-10 quick wins | these vary per page even from one template. | -"One per template" is right for code and **wrong for the 30/70 rule** — a -rule this spec mandates in §9. Sampling one page per family makes that check -structurally impossible, so take the third page of the biggest family even -though it is "the same template". +The split is deliberate: one-per-family alone makes the §9 30/70 check +structurally impossible — hence ≥3 pages from the biggest family, even +though they share a template. An un-sampled family is an un-audited family. Name the ones you skipped. @@ -658,16 +655,11 @@ mapfile -t FEXCL < <(bash ~/.claude/lib/source-scope.sh findargs) find . "${FEXCL[@]}" -type f \( -iname "*.jpg" -o -iname "*.jpeg" -o -iname "*.png" -o -iname "*.gif" \) -printf "%s %p\n" 2>/dev/null | sort -rn | head -20 ``` -**Why the guard, and why `find` specifically (C1a).** `grep` and `find` -disagree about this repo and you use both. Claude Code routes `grep` through -ugrep with `--ignore-files`, so it honours `.gitignore` and never descends -into a gitignored `dist/`. `find` honours nothing. Measured on a real Astro -repo: this command returned **92 images, 45 of them under `dist/`** — every -asset twice, source and generated copy, byte-identical. So "top 20 by size" -was ~10 real images dressed as 20, and a batch-C item -(`cwebp -q 80 -o .webp`) could target `dist/og-image.png`, whose -`.webp` the dispatcher's own `npm run build` then erases. The fix lands, -verification passes, nothing survives. +**Why the guard, and why `find` specifically (C1a).** Claude Code routes +`grep` through ugrep with `--ignore-files` (honours `.gitignore`); `find` +honours nothing. Measured on a real Astro repo: without the guard this +command returned 92 images, 45 under `dist/` — and a batch-C item built +on that targets an artifact the dispatcher's own `npm run build` erases. `FEXCL` MUST be consumed as a quoted array. `find . $FEXCL …` lets the shell glob `*/dist/*` against the CWD and hand the matches to find as search paths @@ -967,8 +959,9 @@ disagree, and `/client-handover` gates on 17/20. **N/A is not a zero** and the engine will not let it behave like one. - `status: "error"` → malformed findings. Fix them; never fall back to eyeballing a number. -- Run it twice on the same file before publishing. If the output moved, your - findings moved, and that is the thing to explain. +- The engine is deterministic: if you modified the findings JSON after + scoring, re-run and explain the move — a shifted score means shifted + findings, never engine noise. **Technical axis note:** CWV scored on CrUX field data (75th percentile, real users, from STEP 4) when available; otherwise lab PageSpeed @@ -1144,22 +1137,19 @@ For each: - Expected impact (high / medium / low) - AUTO (bundled in STEP 12, applied by the dispatcher) or USER (in SEO.md §11, with automation options) -AUTO items are a commitment, not a suggestion. +**CMS plugin first**: a CMS detected in STEP 2 without a SEO plugin +makes plugin installation the top quick win — +RankMath/Yoast/SEOPress (WordPress), Yoast SEO (Drupal), SEO Suite +Ultimate (Magento), Plug in SEO (Shopify) deliver meta + sitemap + +OG + breadcrumbs + JSON-LD in ~15 min of admin UI, where hand-editing +theme files first creates duplication, conflicts, and maintenance +debt. See `~/.claude/agents/resources/automation-catalog.md` CMS +plugins section for the exact install path per CMS. -**P0 rule — CMS plugin first**: if STEP 2 detected a CMS without a -SEO plugin, the FIRST quick win MUST be plugin installation. Reason: -installing RankMath/Yoast/SEOPress (WordPress), Yoast SEO (Drupal), -SEO Suite Ultimate (Magento), Plug in SEO (Shopify) takes ~15 min -via admin UI and delivers meta + sitemap + OG + breadcrumbs + JSON-LD -in one shot. Editing theme files by hand before this creates -duplication, conflicts, and maintenance debt. See -`~/.claude/agents/resources/automation-catalog.md` CMS plugins -section for the exact install path per CMS. - -**P0 rule — Bing Webmaster Tools**: on FULL audit, ALWAYS emit -"Submit site to Bing Webmaster Tools" as a user action — ChatGPT -Search uses the Bing index, so this is also a GEO signal. See -automation-catalog.md for IndexNow + Bing. +**Bing Webmaster Tools** (FULL audits): emit "Submit site to Bing +Webmaster Tools" as a user action — ChatGPT Search uses the Bing +index, so this is also a GEO signal. See automation-catalog.md for +IndexNow + Bing. ### Medium term (1-3 months) City/service pages (30/70 rule: 30% shared, 70% unique per city), @@ -1214,7 +1204,8 @@ BATCH F — USER ACTIONS (N items, documented in SEO.md §11 with automation cat ... ``` -Do not proceed to STEP 12 until this plan is printed. +Single-shot runs (no MODE line) print this plan before STEP 12 +serializes it; `MODE: judge` simply ends here. --- @@ -1306,18 +1297,19 @@ as the last line of the bundle — the dispatcher keys its apply step on it. Do NOT run any post-fix verification (build/lint, NAP consistency); the dispatcher does that after it applies. Your job ends at the sentinel. -### Bundle completeness checklist (did every finding reach the bundle?) +### Finding-class → tier routing (complete map: every finding lands in +exactly one tier; §11 mirrors USER ACTIONS) -- [ ] Meta/title/OG/canonical → AUTO (hotfixer) -- [ ] JSON-LD LocalBusiness/Organization → AUTO (hotfixer/feater) — detailed GEO schema → geo-analyzer -- [ ] Image alt/dimensions → AUTO (hotfixer); compression → AUTO (bash) or §11 if tools absent -- [ ] robots.txt / sitemap.xml → AUTO (hotfixer) — AI-bot directives → geo-analyzer -- [ ] .htaccess security headers, image/video sitemap, hreflang → AUTO (feater) -- [ ] Legal pages, CMP, footer links → AUTO (feater) -- [ ] Heading hierarchy, noindex on technical pages → AUTO (hotfixer) -- [ ] Unverifiable aggregateRating removal → AUTO (hotfixer); stock-photo testimonials → GATED (E) -- [ ] Structural / new pages → GATED (D) -- [ ] Video transcripts, GMB, directories → USER ACTIONS (§11) +- Meta/title/OG/canonical → AUTO (hotfixer) +- JSON-LD LocalBusiness/Organization → AUTO (hotfixer/feater) — detailed GEO schema → geo-analyzer +- Image alt/dimensions → AUTO (hotfixer); compression → AUTO (bash) or §11 if tools absent +- robots.txt / sitemap.xml → AUTO (hotfixer) — AI-bot directives → geo-analyzer +- .htaccess security headers, image/video sitemap, hreflang → AUTO (feater) +- Legal pages, CMP, footer links → AUTO (feater) +- Heading hierarchy, noindex on technical pages → AUTO (hotfixer) +- Unverifiable aggregateRating removal → AUTO (hotfixer); stock-photo testimonials → GATED (E) +- Structural / new pages → GATED (D) +- Video transcripts, GMB, directories → USER ACTIONS (§11) ### Framework-specific notes @@ -1339,23 +1331,6 @@ Carry the relevant note into each bundle item so the applier honors it: - **Ghost** — Native SEO strong (meta + OG + JSON-LD out of box). Usually no plugin needed; handle gaps via `default.hbs` edits. - **Wix / Squarespace / Webflow (hosted CMS)** — No theme file access. ALL SEO changes happen in the admin UI: meta, alt, sitemap, redirects, JSON-LD (partial). Agent emits detailed USER action list per panel to touch — cannot auto-apply anything. -### Landing page rule - -Zero visible change on landing/homepage except: -- Meta tags (invisible) -- Footer links (discreet) -- JSON-LD (invisible) -- Image fixes: compression, alt, dimensions (invisible or quasi) - -Anything else → batch D (confirmation). - -### Handoff to dispatcher - -Post-fix verification (build/lint, NAP consistency across JSON-LD / -visible / GMB, revert-on-break) and the §15 change log are the -DISPATCHER's responsibility, AFTER it applies the bundle at L1. You -emitted the bundle terminated by the sentinel — stop here. - --- ## STEP 13 — OUTPUT `[both]` @@ -1519,10 +1494,10 @@ PROCHAINE ETAPE : ### Process - **Every user action lists automation.** Mandatory from `~/.claude/agents/resources/automation-catalog.md`. -- **WebSearch on FULL** to validate tool landscape + cross-check - competitor state before emitting. +- **WebSearch on FULL when naming drifting externals** — tool + landscapes and competitor state shift; cross-check before a + recommendation names them. - **Iterative SEO.md.** Preserve Historique section. -- **Transparency.** Every automated change logged with file, change, - reason. -- **Dispatcher verifies.** Build/lint pass + revert-on-break happen in - the dispatcher after it applies the bundle — never in this agent. +- **Dispatcher verifies.** Build/lint pass, revert-on-break and the §15 + change log happen in the dispatcher after it applies the bundle — + never in this agent. diff --git a/agents/status-reporter.md b/agents/status-reporter.md index 33439dd..98f6081 100644 --- a/agents/status-reporter.md +++ b/agents/status-reporter.md @@ -1,6 +1,6 @@ --- name: status-reporter -description: Read-only project-status engine — dispatched by /status. Collects plugins, token budget, git state, build/tests, GSD milestone into one snapshot. +description: Read-only project-status engine — dispatched by /status. Collects plugin roster + passive-cost estimate (doctor.sh constants), git state, build/tests, GSD milestone into one snapshot. tools: Read, Bash, Glob, Grep model: haiku --- @@ -23,8 +23,12 @@ cat ~/.claude/lib/../version.txt 2>/dev/null || echo "unknown" # lib symlink re command -v rtk &>/dev/null && echo "rtk: installed" || echo "rtk: missing" command -v gsd &>/dev/null && gsd --version 2>/dev/null | head -1 || echo "gsd: not installed" -# Token estimate (passive) -# (approximate from known plugin costs) +# Passive token cost — source of truth: doctor.sh's constants block +# (PLUGIN_TOKENS + per detect_* line). Read it, sum ONLY the plugins +# found active above. Never invent a number outside these constants. +grep -E 'PLUGIN_TOKENS \+ [0-9]+' "$(readlink -f "$HOME/.claude/lib")/../doctor.sh" 2>/dev/null +# grep empty (doctor.sh missing/moved) → report the plugin count only and +# defer cost to /plugin-check. ``` Check `~/.claude/plugins/cache` for active marketplace plugins. @@ -134,7 +138,7 @@ PROJECT STATUS CONFIG Version : v - Plugins ON: (~t passive) + Plugins ON: (~t passive — doctor.sh constants; full audit → /plugin-check) GSD v2 : installed / not installed PROJECT @@ -174,7 +178,7 @@ The report is best-effort: a single failing data source must not abort the whole |---|---| | Permission denied on `git` (sandbox/CI without `.git` access) | Mark `Branch: N/A (permission denied)`, `Uncommitted: N/A`, `RECENT COMMITS: N/A`. Continue to PROJECT/GSD sections. | | Permission denied on `~/.claude/plugins/cache` or `~/.claude.json` | Mark `Plugins ON: unknown (cannot read cache)`. Continue. | -| `.gsd/ROADMAP.md` exists but unparseable (malformed checkboxes, encoding issue) | Mark `Progress: N/A (ROADMAP.md unreadable)`, do NOT abort the section — still print `Status: initialized` and `Milestone: N/A`. | +| gsd CLI snapshot fails or `.gsd/` state unreadable (`gsd.db`, `STATE.md`, per-milestone `-ROADMAP.md` — post-ADR-013 layout) | Mark `Progress: N/A (gsd state unreadable)`, do NOT abort the section — still print `Status: initialized` and `Milestone: N/A`. | | `package.json` / `pyproject.toml` parse error | Mark `Tests: N/A (manifest parse error)`. Continue. | | `python3` not available in PATH | Skip the python parsing fallbacks; rely on log files + bash-only checks. Mark Tests as `unknown` if no log found. | | All sections fail | Print a minimal envelope with each section showing `N/A (data source unavailable)` and a one-line `DIAGNOSTIC: ` footer. Exit code 0 (status reporter never blocks). | diff --git a/agents/verifier.md b/agents/verifier.md index f6fd9bf..ba6a7ae 100644 --- a/agents/verifier.md +++ b/agents/verifier.md @@ -48,6 +48,25 @@ Rules: read the diff AND enough surrounding code to judge behavior; run criterion. Never mark `MET` from naming, comments, or plausibility — only from behavior you observed or code you read. +### Criteria carrying an oracle (`CHECK:` / `EXPECT:` / `EVIDENCE:`) + +`lib/gates.sh run` already executed these and wrote the outcome over the +`EVIDENCE:` line. Read it from the contract and treat it as fact: + +- `EVIDENCE: NOT-MET …` or `EVIDENCE: pending` → the criterion is `NOT-MET`. + Reading the code NEVER overrides a red or unrun oracle. Cite the evidence + line as your evidence. +- `EVIDENCE: MET …` → the declared command passed. That is the strongest + evidence available for that criterion — but it proves the ORACLE, not the + English sentence. Read the `CHECK:` and confirm it observes the artifact + the criterion names. A vacuous oracle (`1. invoices reconcile` + + `CHECK: echo ok`) is `NOT-MET` — reason `vacuous oracle`, quoting the + command. That judgement is yours alone; no command can make it. + +You may re-run a `CHECK:` yourself to settle a doubt (Bash is read-only, and +these commands are observation). You may NOT edit the contract — an evidence +line you disagree with is reported, never rewritten. + ## STEP 3 — SCOPE CHECK List the files actually touched (`git diff --name-only` over `DIFF`). @@ -58,19 +77,30 @@ only enters the contract through a human micro-gate. ## STEP 4 — VERDICT -`CONFORME` ⇔ ALL criteria `MET` AND zero out-of-scope files. -Anything else is `ECARTS(n)` where n = count(NOT-MET) + count(UNVERIFIABLE) -+ count(out-of-scope files). +Read the contract's `ABANDON:` lines. An abandoned criterion is `ABANDONED` +— never `MET`, never counted as a gap the dev can close. + +Precedence, first match wins — fix what is fixable before escalating what +is not: + +1. `ERROR()` — the contract is missing or unreadable. +2. `ECARTS(n)` — n = count(NOT-MET) + count(UNVERIFIABLE) + count(out-of-scope + files). Surface any abandonment in the same report. +3. `ABANDONED(n)` — zero gaps remain, but n abandonments stand. This is NOT + a pass and NOT a dev loop: it routes straight to the human gate. +4. `CONFORME` — ALL criteria `MET`, zero out-of-scope files, zero + abandonments. ## OUTPUT (exact format — machine-parsed by the orchestrator) ``` -VERIFY — VERDICT: CONFORME | ECARTS(n) | ERROR() +VERIFY — VERDICT: CONFORME | ECARTS(n) | ABANDONED(n) | ERROR() CONTRACT: CRITERIA: - 1. — MET — + 1. — MET — 2. — NOT-MET — expected <…> / actual <…> — 3. — UNVERIFIABLE — + 4. — ABANDONED — SCOPE: in-scope files; out-of-scope: PROOF: read files, ran , checked / criteria ``` @@ -82,6 +112,8 @@ PROOF: read files, ran , checked / criteria - `UNVERIFIABLE` ≠ `MET`. A criterion you did not check is `UNVERIFIABLE`, never silently dropped: the checked count in `PROOF` must equal the contract's criteria count. +- `ABANDONED` ≠ `MET`. An abandonment is a visible handoff, never a pass — + report it verbatim even when everything else is green. - `PROOF` is MANDATORY. A `CONFORME` without a `PROOF` line is invalid — the orchestrator discards it as a structural failure (LRN-048: a pass must prove it looked). @@ -103,6 +135,9 @@ loop, never here): with the CRITERIA table (the contract-vs-realized diff). - Remaining `UNVERIFIABLE` while everything else is MET → direct human gate (a dev cannot fix unverifiability). + - `ABANDONED(n)` → direct human gate, never a dev loop. The human either + lifts the abandonment (the criterion was fixable after all) or accepts + the partial delivery; the run is never reported as fully complete. - Structural failure (`ERROR(…)`, missing/duplicated VERDICT line, unparsable output, agent crash, `CONFORME` without `PROOF`) → retry ONCE with a fresh verifier; a 2nd structural failure → human diff --git a/hooks/design-toolchain-reminder.sh b/hooks/design-toolchain-reminder.sh index e39b76a..68547eb 100755 --- a/hooks/design-toolchain-reminder.sh +++ b/hooks/design-toolchain-reminder.sh @@ -44,7 +44,11 @@ lc="$(printf '%s' "$prompt" | tr '[:upper:]' '[:lower:]')" # "design system", "redesign", "front-?end design". dashboard -> \bdashboard\b # so a filename like ecc_dashboard.py no longer matches while "admin dashboard" # still does. animation kept (rarely non-UI). -pattern='redesign|refonte|refont|ui/ux|ux/ui|\bui\b|\bux\b|ui kit|design system|design-system|front-?end design|\bnavbar\b|\bsidebar\b|\bmodal\b|\bbouton\b|\bbutton\b|formulaire|\bhero\b|\bheader\b|\bfooter\b|dropdown|tooltip|\bbadge\b|\bchart\b|graphique|accordion|carousel|\bslider\b|landing|\bdashboard\b|homepage|home page|\baccueil\b|\bécran\b|\becran\b|portfolio|maquette|mockup|wireframe|prototype|\bjoli\b|\bjolie\b|\bbeau\b|\bbelle\b|esth[eé]tique|aesthetic|\bvisuel\b|\bvisual\b|embellir|fignol|peaufin|polish|styliser|styling|stylesheet|\bskin\b|charte graphique|\bbrand\b|branding|\blogo\b|favicon|ic[oô]ne|\bicon\b|\bcss\b|tailwind|shadcn|couleur|gradient|d[eé]grad[eé]|\bombre\b|spacing|espacement|\bmarge\b|\bpadding\b|\bmargin\b|\bradius\b|arrondi|\bhover\b|dark mode|light mode|typograph|\bfont\b|\bfonts\b|font pairing|\bpolice\b|animation|\bmotion\b|micro-interaction|keyframe|glassmorph|neumorph|claymorph|skeuomorph|brutalis|bento|minimalis|responsive|figma' +# Tightened 2026-07-30 (3rd pass): dropped \bux\b — bare "ux" matched inside +# French prose ("changement ux vu…"; 2 logged FPs, both FR). \bui\b KEPT +# (zero logged FP, one logged true positive). NB: the log records only the +# FIRST match per fire (head -1), so per-token FP rates aren't derivable. +pattern='redesign|refonte|refont|ui/ux|ux/ui|\bui\b|ui kit|design system|design-system|front-?end design|\bnavbar\b|\bsidebar\b|\bmodal\b|\bbouton\b|\bbutton\b|formulaire|\bhero\b|\bheader\b|\bfooter\b|dropdown|tooltip|\bbadge\b|\bchart\b|graphique|accordion|carousel|\bslider\b|landing|\bdashboard\b|homepage|home page|\baccueil\b|\bécran\b|\becran\b|portfolio|maquette|mockup|wireframe|prototype|\bjoli\b|\bjolie\b|\bbeau\b|\bbelle\b|esth[eé]tique|aesthetic|\bvisuel\b|\bvisual\b|embellir|fignol|peaufin|polish|styliser|styling|stylesheet|\bskin\b|charte graphique|\bbrand\b|branding|\blogo\b|favicon|ic[oô]ne|\bicon\b|\bcss\b|tailwind|shadcn|couleur|gradient|d[eé]grad[eé]|\bombre\b|spacing|espacement|\bmarge\b|\bpadding\b|\bmargin\b|\bradius\b|arrondi|\bhover\b|dark mode|light mode|typograph|\bfont\b|\bfonts\b|font pairing|\bpolice\b|animation|\bmotion\b|micro-interaction|keyframe|glassmorph|neumorph|claymorph|skeuomorph|brutalis|bento|minimalis|responsive|figma' if printf '%s' "$lc" | grep -Eq "$pattern"; then # Counter: log the fire (time, matched token, excerpt) — best-effort, never blocks. diff --git a/hooks/notify-attention.sh b/hooks/notify-attention.sh new file mode 100755 index 0000000..a6113cc --- /dev/null +++ b/hooks/notify-attention.sh @@ -0,0 +1,67 @@ +#!/usr/bin/env bash +# Notification + Stop hook — signal the user through the terminal when +# Claude needs input (permission, question, idle wait) or has finished +# responding. Each case gets its own readable label so the toast says +# which one fired. +# +# Runs on the remote (Linux); the only channel that crosses SSH into the +# VS Code client is the terminal stream. Hooks have no controlling TTY, +# so the sequence goes through the supported `terminalSequence` JSON +# output field and Claude Code writes it to the terminal: +# - BEL x2 (double beep) -> sound, needs VS Code setting +# accessibility.signals.terminalBell { "sound": "on" } AND a non-zero +# volume for Code in the Windows volume mixer (BLK-020). +# - OSC 777 notify -> Windows toast via the client-side extension +# "Terminal Notification" (wenbopan.vscode-terminal-osc-notifier). +# A terminal can be deaf to OSC while the bell still rings; test it +# before attaching a session to it (LRN-148). +# Both are invisible no-ops in terminals that ignore them. +set -u + +payload=$(cat 2>/dev/null) +read_field() { + printf '%s' "$payload" | jq -r "$1 // empty" 2>/dev/null \ + | tr -d '\000-\037' | cut -c1-160 +} + +# How many background tasks are still running as the hook fires. +background_count() { + count=$(printf '%s' "$payload" | jq -r '(.background_tasks // []) | length' 2>/dev/null) + case "$count" in ''|*[!0-9]*) echo 0 ;; *) echo "$count" ;; esac +} + +event=$(read_field '.notification_type') +[ -n "$event" ] || event=$(read_field '.hook_event_name') + + + +case "$event" in + # Turn end while a subagent still runs is not the real end: stay silent, + # the next turn end will signal once the work is actually done. + Stop) [ "$(background_count)" -eq 0 ] || exit 0 + label="Finished responding" ;; + permission_prompt) label="Needs your permission" ;; + agent_needs_input) label="Asks you a question" ;; + idle_prompt) label="Waiting for you" ;; + elicitation_dialog|elicitation_url_dialog) label="Needs your input" ;; + # anything else (agent_completed, auth_success, quota_*) stays silent: + # signal only for turn end and moments needing the user. + *) exit 0 ;; +esac + +detail=$(read_field '.message') +if [ -n "$detail" ]; then + # Claude Code's own wording often restates the label ("Claude needs your + # permission"). Append it only when it actually adds something. + short=$(printf '%s' "$detail" | tr '[:upper:]' '[:lower:]' | sed 's/^claude //') + case "$(printf '%s' "$label" | tr '[:upper:]' '[:lower:]')" in + *"$short"*) : ;; + *) label="${label}: ${detail}" ;; + esac +fi + +bell=$(printf '\a') +esc=$(printf '\033') +seq="${bell}${bell}${esc}]777;notify;Claude Code;${label}${esc}\\" +jq -cn --arg seq "$seq" '{suppressOutput: true, terminalSequence: $seq}' +exit 0 diff --git a/lib/contract-interview.md b/lib/contract-interview.md index 9dcc1c4..47256b7 100644 --- a/lib/contract-interview.md +++ b/lib/contract-interview.md @@ -35,6 +35,38 @@ ask what the repo can answer — verify paths/APIs/behavior yourself first. this conversation. - FILE SCOPE: paths/zones expected to change, or `repo-wide — `. +### ORACLES — a criterion a command can decide carries one + +Give such a criterion an indented `CHECK:` (the command), `EXPECT:` (a +success-only marker), and `EVIDENCE: pending`. +`bash ~/.claude/lib/gates.sh run ` executes it fail-closed — MET +requires exit 0 **AND** the marker — and writes the result back over the +`EVIDENCE:` line. That persisted evidence is what the fresh verifier reads +as fact instead of trusting the executor's report (GATE 0 in +`lib/verify-secure-loop.md`). + +Both attributes or neither. `CHECK:` without `EXPECT:` is a parse error, not +a manual criterion — the runner refuses the whole ledger. Leave a criterion +oracle-free when no command can decide it; the verifier judges those. + +Four authoring rules — a gate that cannot fail proves nothing: + +1. **Observe the named artifact.** The check reads the file, service, or + measurement the criterion's own words name — never a proxy for it. + `1. invoices reconcile` + `CHECK: echo ok` is valid and worthless. +2. **Success-only marker.** The script runs every assertion, exits nonzero + on any failure, and prints the `EXPECT:` string only after all pass. +3. **Positive control before any absence check.** Run the same logic against + a fixture known to trip it and confirm it fails. A missing file, a wrong + path, and a broken pattern all look exactly like valid absence. +4. **Recompute supplied numbers.** Never copy a figure from the request into + `EXPECT:` — the script derives it from source and prints its own marker. + A number that is its own proof proves nothing. + +`CHECK:` is shell code run with our privileges. It is safe only because we +author it in our own repo — never build one out of externally-supplied text +(a scraped URL, a client string); route those through `lib/url-guard.sh`. + ## STEP 4 — WRITE TO DISK (immediately, before any next step) Path: `.claude/tasks/contracts/--.md` @@ -57,8 +89,13 @@ Q: / A: (or: none — request complete) ## ACCEPTANCE CRITERIA -1. -2. +1. + CHECK: + EXPECT: + EVIDENCE: pending +2. + +(ABANDON: — only for a criterion proven impossible) ## FILE SCOPE @@ -78,6 +115,13 @@ Print one line to the user, then continue the flow: this micro-gate: human approves → FILE SCOPE gains the entry `[gated]`; human declines → the dev removes the edit. Without this gate the dev justifies everything and scope constrains nothing. +- **ABANDONMENT**: a criterion proven impossible within the authorized task + is NEVER deleted and never quietly downgraded. Keep it, append + `ABANDON: ` under the criteria, and name it + in the final report. An abandonment is a visible handoff, not a pass: the + verifier cannot return `CONFORME` while one stands, and the run cannot be + described as fully complete. This is the structural half of the house rule + "blocked on an independent sub-part → do the rest, state what's missing". - **Deep re-scope** (the request itself changes): NEW contract file with `supersedes: ` in its header — never a rewrite of the old one. - **Aborted run**: delete the contract file, or commit it with @@ -96,6 +140,15 @@ Print one line to the user, then continue the flow: | init-project | Full. The interviewer's PROJECT BRIEF pours into the contract (V1 features → criteria). | | onboard | Audit-scope contract (interview answers → what to audit, which axes). | +Oracles follow the same proportion. hotfix: none — that flow runs no floor +(and no verifier); the hotfixer runs build/tests itself. feat / bugfix: the +suite criterion at minimum, and for bugfix the regression test the DIAGNOSIS +names — its `CHECK:` runs that test alone, so a green result means the +reproduction actually flipped. +ship-feature / init-project: build, suite, and every criterion a command can +settle. onboard: audit criteria are mostly judgement — leave them oracle-free +rather than invent a check that cannot fail. + ## Hand-off rule Downstream consumers (plan step, dev subagents, verifier) receive the diff --git a/lib/gates.sh b/lib/gates.sh new file mode 100644 index 0000000..36b5404 --- /dev/null +++ b/lib/gates.sh @@ -0,0 +1,323 @@ +#!/usr/bin/env bash +# Deterministic floor under GATE 1: execute the acceptance criteria that the +# contract itself declares as oracles, fail-closed, and persist the evidence +# INTO the contract file. +# +# bash ~/.claude/lib/gates.sh status # parse only, never runs +# bash ~/.claude/lib/gates.sh run # execute + write evidence +# +# rc 0 = MET every runnable criterion passed, no abandonment standing +# 2 = UNMET a runnable criterion failed, or the ledger is malformed +# 3 = ABANDONED runnable criteria all passed, an abandonment still stands +# +# WHY: GATE 1 (lib/verify-secure-loop.md) is an LLM dispatch, and the +# verifier's mandatory `PROOF:` line is a line the verifier WRITES — nothing +# structurally stops it from being produced without anything being executed. +# This runs what the contract declares BEFORE a verifier is ever spawned: a +# red floor sends the executor back for free. Adapted from the `unlazy` skill +# (Leonxlnx/unlazy) — its gate ledger, minus the machinery we do not need. +# +# `run` always re-executes every runnable criterion, including ones already +# recorded MET. Trusting written evidence is exactly the failure this closes, +# so there is no incremental mode to get it wrong with. +# +# TRUST BOUNDARY: `CHECK:` is shell code, run with this process's privileges +# and environment. That is safe here only because the contract is authored by +# our own orchestrator in our own repo — which is why there is no approval +# store (we never execute ledgers inherited from a foreign repo). NEVER build +# a `CHECK:` out of externally-supplied text; route such values through +# lib/url-guard.sh first. +set -uo pipefail + +TIMEOUT="${GATES_TIMEOUT:-120}" +EVIDENCE_CAP=140 + +# Module-level parse tables, index-aligned. Bash has no record type; threading +# eight parallel arrays through every call would cost more readability than +# the explicit data flow buys. +_ID=(); _TEXT=(); _CHECK=(); _EXPECT=(); _EVLINE=(); _EVTEXT=() +_STATUS=(); _EVID=() +_ABANDON_ID=(); _ABANDON_WHY=() +_ERRORS=() +_CUR=-1 + +_die() { printf 'GATES — VERDICT: ERROR(%s)\n' "$1"; exit 2; } +_err() { _ERRORS+=("$1"); } + +_trim() { + local s="$1" + s="${s#"${s%%[![:space:]]*}"}" + printf '%s' "${s%"${s##*[![:space:]]}"}" +} + +# ── parse ─────────────────────────────────────────────────────────────────── + +_new_crit() { # _new_crit + local i + for ((i = 0; i < ${#_ID[@]}; i++)); do + if [ "${_ID[i]}" = "$1" ]; then + _err "duplicate criterion id: $1" + # Orphan what follows instead of aliasing it onto the previous + # criterion, which would hand one gate another gate's oracle. + _CUR=-1 + return 0 + fi + done + _ID+=("$1"); _TEXT+=("$2") + _CHECK+=(""); _EXPECT+=(""); _EVLINE+=("0"); _EVTEXT+=("") + _CUR=$((${#_ID[@]} - 1)) +} + +_set_attr() { # _set_attr + if [ "$_CUR" -lt 0 ]; then + _err "$1 at line $3 belongs to no criterion" + return 0 + fi + case "$1" in + CHECK) _CHECK[_CUR]="$2" ;; + EXPECT) _EXPECT[_CUR]="$2" ;; + EVIDENCE) _EVLINE[_CUR]="$3"; _EVTEXT[_CUR]="$2" ;; + esac +} + +# An UNINDENTED attribute is diagnosed, never absorbed: silently ignoring it +# would demote a runnable criterion to a manual one, which is the one parse +# bug that turns this checker into a rubber stamp. +_absorb() { # _absorb + local body + if [[ "$1" =~ ^([0-9]+)\.[[:space:]]+(.*)$ ]]; then + _new_crit "${BASH_REMATCH[1]}" "${BASH_REMATCH[2]}" + elif [[ "$1" =~ ^ABANDON:[[:space:]]*([0-9]+)?[[:space:]]*(.*)$ ]]; then + _ABANDON_ID+=("${BASH_REMATCH[1]}"); _ABANDON_WHY+=("${BASH_REMATCH[2]}") + elif [[ "$1" =~ ^(CHECK|EXPECT|EVIDENCE): ]]; then + _err "unindented ${BASH_REMATCH[1]}: at line $2" + elif [[ "$1" =~ ^[[:space:]]+(CHECK|EXPECT|EVIDENCE):(.*)$ ]]; then + body="$(_trim "${BASH_REMATCH[2]}")" + _set_attr "${BASH_REMATCH[1]}" "$body" "$2" + fi +} + +_parse() { # _parse + local line n=0 fence=0 inblock=0 + while IFS= read -r line || [ -n "$line" ]; do + n=$((n + 1)) + case "$line" in '```'*) fence=$((1 - fence)); continue ;; esac + [ "$fence" -eq 1 ] && continue + case "$line" in + '## ACCEPTANCE CRITERIA'*) inblock=1; continue ;; + '## '*) inblock=0; continue ;; + esac + [ "$inblock" -eq 1 ] && _absorb "$line" "$n" + done < "$1" +} + +# ── validation ────────────────────────────────────────────────────────────── + +_validate_oracles() { + local i + for ((i = 0; i < ${#_ID[@]}; i++)); do + if [ -n "${_CHECK[i]}" ] && [ -z "${_EXPECT[i]}" ]; then + _err "criterion ${_ID[i]}: CHECK without EXPECT (partial oracle)" + elif [ -z "${_CHECK[i]}" ] && [ -n "${_EXPECT[i]}" ]; then + _err "criterion ${_ID[i]}: EXPECT without CHECK (partial oracle)" + elif [ -n "${_CHECK[i]}" ] && [ "${_EVLINE[i]}" = "0" ]; then + _err "criterion ${_ID[i]}: runnable but has no EVIDENCE: line" + fi + done +} + +_validate_abandons() { + local i j found + for ((i = 0; i < ${#_ABANDON_ID[@]}; i++)); do + found=0 + for ((j = 0; j < ${#_ID[@]}; j++)); do + [ "${_ID[j]}" = "${_ABANDON_ID[i]}" ] && found=1 + done + [ "$found" -eq 1 ] || + _err "ABANDON names unknown criterion: '${_ABANDON_ID[i]}'" + [ -n "$(_trim "${_ABANDON_WHY[i]}")" ] || + _err "ABANDON ${_ABANDON_ID[i]}: blank reason (a handoff needs one)" + done +} + +_is_abandoned() { # _is_abandoned + local i + for ((i = 0; i < ${#_ABANDON_ID[@]}; i++)); do + [ "${_ABANDON_ID[i]}" = "$1" ] && return 0 + done + return 1 +} + +# ── execution ─────────────────────────────────────────────────────────────── + +# One line, capped, newlines flattened: the smallest output that proves the +# outcome. Full logs stay in the terminal, never in the contract. +_decisive() { # _decisive + local flat + flat="$(printf '%s' "$1" | tr '\n\r\t' ' ' | tr -s ' ')" + flat="$(_trim "$flat")" + if [ "${#flat}" -gt "$EVIDENCE_CAP" ]; then + printf '%s…' "${flat:0:$EVIDENCE_CAP}" + else + printf '%s' "$flat" + fi +} + +# Fail-closed: exit 0 AND the marker. A nonzero process never passes because +# its error text happens to contain the expected token. +_run_one() { # _run_one + local i="$1" out rc + out="$(timeout "$TIMEOUT" bash -c "${_CHECK[i]}" 2>&1)" + rc=$? + _STATUS[i]="NOT-MET" + if [ "$rc" -eq 124 ]; then + _EVID[i]="NOT-MET timeout=${TIMEOUT}s" + elif [ "$rc" -ne 0 ]; then + _EVID[i]="NOT-MET exit=$rc (nonzero) :: $(_decisive "$out")" + elif [[ "$out" != *"${_EXPECT[i]}"* ]]; then + _EVID[i]="NOT-MET exit=0 marker-absent :: $(_decisive "$out")" + else + _STATUS[i]="MET" + _EVID[i]="MET exit=0 marker-found :: $(_decisive "$out")" + fi +} + +_run_all() { + local i + for ((i = 0; i < ${#_ID[@]}; i++)); do + _STATUS[i]=""; _EVID[i]="" + [ -n "${_CHECK[i]}" ] && _run_one "$i" + done +} + +_evline_owner() { # _evline_owner — echoes idx, or nothing + local i + for ((i = 0; i < ${#_ID[@]}; i++)); do + if [ "${_EVLINE[i]}" = "$1" ] && [ -n "${_EVID[i]}" ]; then + printf '%s' "$i" + return 0 + fi + done +} + +# Rewrites only the EVIDENCE lines of criteria that actually ran; every other +# byte of the contract is copied through, indentation included. +_write_back() { # _write_back + local tmp line n=0 idx + tmp="$(mktemp)" || _die "mktemp failed" + while IFS= read -r line || [ -n "$line" ]; do + n=$((n + 1)) + idx="$(_evline_owner "$n")" + if [ -n "$idx" ]; then + printf '%s%s\n' "${line%%[![:space:]]*}" "EVIDENCE: ${_EVID[idx]}" + else + printf '%s\n' "$line" + fi + done < "$1" > "$tmp" + cat "$tmp" > "$1" && rm -f "$tmp" +} + +# ── report ────────────────────────────────────────────────────────────────── + +# A recorded `pending`, or a criterion that never ran, is PENDING — never MET. +# `status` reports what the file says; it does not revalidate old evidence. +_row_state() { # _row_state + local i="$1" + _is_abandoned "${_ID[i]}" && { printf 'ABANDONED'; return 0; } + [ -z "${_CHECK[i]}" ] && { printf 'MANUAL'; return 0; } + [ -n "${_STATUS[i]:-}" ] && { printf '%s' "${_STATUS[i]}"; return 0; } + case "${_EVTEXT[i]}" in + MET' '*) printf 'MET-RECORDED' ;; + *) printf 'PENDING' ;; + esac +} + +_report_rows() { + local i state + for ((i = 0; i < ${#_ID[@]}; i++)); do + state="$(_row_state "$i")" + printf ' %-3s %-13s %s\n' "${_ID[i]}" "$state" "${_TEXT[i]}" + done +} + +_report_abandons() { + local i + for ((i = 0; i < ${#_ABANDON_ID[@]}; i++)); do + printf ' ABANDONED %s — %s\n' "${_ABANDON_ID[i]}" "${_ABANDON_WHY[i]}" + done +} + +_count_state() { # _count_state + local i n=0 + for ((i = 0; i < ${#_ID[@]}; i++)); do + [ "$(_row_state "$i")" = "$1" ] && n=$((n + 1)) + done + printf '%s' "$n" +} + +_verdict() { # _verdict — prints the line, returns the rc + local unmet pending abandoned + if [ "${#_ERRORS[@]}" -gt 0 ]; then + printf 'GATES — VERDICT: ERROR(%s)\n' "${#_ERRORS[@]}" + return 2 + fi + unmet="$(_count_state NOT-MET)" + pending="$(_count_state PENDING)" + abandoned="$(_count_state ABANDONED)" + [ "$unmet" -gt 0 ] && + { printf 'GATES — VERDICT: UNMET(%s)\n' "$unmet"; return 2; } + if [ "$1" = "status" ] && [ "$pending" -gt 0 ]; then + printf 'GATES — VERDICT: PENDING(%s)\n' "$pending" + return 2 + fi + [ "$abandoned" -gt 0 ] && + { printf 'GATES — VERDICT: ABANDONED(%s)\n' "$abandoned"; return 3; } + printf 'GATES — VERDICT: MET\n' + return 0 +} + +_report() { # _report + local rc + printf 'GATES — %s (%s)\n' "$2" "$1" + _report_rows + _report_abandons + [ "${#_ERRORS[@]}" -gt 0 ] && printf ' ERROR %s\n' "${_ERRORS[@]}" + printf 'RUNNABLE: %s of %s criteria; timeout %ss\n' \ + "$(_runnable_count)" "${#_ID[@]}" "$TIMEOUT" + _verdict "$1" + rc=$? + return "$rc" +} + +_runnable_count() { + local i n=0 + for ((i = 0; i < ${#_ID[@]}; i++)); do + [ -n "${_CHECK[i]}" ] && n=$((n + 1)) + done + printf '%s' "$n" +} + +# ── entry point ───────────────────────────────────────────────────────────── + +main() { # main + local mode="$1" file="$2" + [ -r "$file" ] || _die "contract unreadable: $file" + _parse "$file" + [ "${#_ID[@]}" -gt 0 ] || + _die "no numbered criteria under ## ACCEPTANCE CRITERIA" + _validate_oracles + _validate_abandons + if [ "$mode" = "run" ] && [ "${#_ERRORS[@]}" -eq 0 ]; then + _run_all + _write_back "$file" + fi + _report "$mode" "$file" +} + +case "${1:-}" in + status|run) + [ $# -eq 2 ] || _die "usage: gates.sh {status|run} " + main "$1" "$2" + ;; + *) _die "usage: gates.sh {status|run} " ;; +esac diff --git a/lib/tests/contract-verifier.test.sh b/lib/tests/contract-verifier.test.sh index af7b99f..ce77347 100644 --- a/lib/tests/contract-verifier.test.sh +++ b/lib/tests/contract-verifier.test.sh @@ -67,7 +67,7 @@ fi tr_ "frontmatter name" "$AGT" "^name: verifier$" tr_ "tools read-only set" "$AGT" "^tools: Read, Grep, Glob, Bash$" tn "no write-capable tools" "$AGT" "^tools:.*(Edit|Write|NotebookEdit)" -tf "verdict grammar" "$AGT" "VERIFY — VERDICT: CONFORME | ECARTS(n) | ERROR()" +tf "verdict grammar" "$AGT" "VERIFY — VERDICT: CONFORME | ECARTS(n) | ABANDONED(n) | ERROR()" tf "blind — no iteration history" "$AGT" "NEVER receive iteration history" tf "blind — complete every time" "$AGT" "every verification is complete and blind" tf "unverifiable is not met" "$AGT" "\`UNVERIFIABLE\` ≠ \`MET\`" diff --git a/lib/tests/design-toolchain-reminder.test.sh b/lib/tests/design-toolchain-reminder.test.sh index 959882f..53977cb 100644 --- a/lib/tests/design-toolchain-reminder.test.sh +++ b/lib/tests/design-toolchain-reminder.test.sh @@ -22,6 +22,7 @@ check D8-dash-file "$(fire 'ecc_dashboard.py')" quiet # --- Harness-generated inputs must be QUIET even with UI tokens --- check D9-tasknotif "$(fire ' x add css header fonts')" quiet check D10-notif-file "$(fire ' design-motion-principles keyframe done')" quiet +check D11-bare-ux "$(fire 'changement ux vu de tes trouvailles')" quiet # --- Real UI signals must still FIRE --- check F1-button "$(fire 'add a button')" fire @@ -33,6 +34,7 @@ check F6-frontdesign "$(fire 'frontend design work')" fire check F7-admin-dash "$(fire 'admin dashboard screen')" fire check F8-animation "$(fire 'add an animation')" fire check F9-designsys "$(fire 'our design system')" fire +check F10-bare-ui "$(fire 'revois l'\''ui du panneau admin')" fire # --- Fire is logged (time + token + excerpt) --- tmp="$(mktemp -d)" diff --git a/lib/tests/gates.test.sh b/lib/tests/gates.test.sh new file mode 100644 index 0000000..76d5fa7 --- /dev/null +++ b/lib/tests/gates.test.sh @@ -0,0 +1,317 @@ +#!/usr/bin/env bash +# ============================================================ +# lib/gates.sh — behavioural tests + structure locks for the +# deterministic floor (GATE 0, lib/verify-secure-loop.md). +# +# Fail-closed is the entire point of this runner, so every +# "looks green but must not pass" case is asserted explicitly: +# nonzero exit carrying the marker, marker absent, timeout, +# unindented attribute silently demoting a gate to manual. +# Non-execution is proved with a sentinel file, and the +# sentinel's own positive control is asserted first — an +# absence check that was never able to fire proves nothing. +# ============================================================ +set -uo pipefail + +REPO="$(cd "$(dirname "$0")/../.." && pwd)" +GATES="$REPO/lib/gates.sh" +WORK="$(mktemp -d)" +trap 'rm -rf "$WORK"' EXIT +PASS=0; FAIL=0; N=0 +LAST="" + +ok() { echo " PASS $1"; PASS=$((PASS + 1)); } +bad() { echo " FAIL $1 — $2"; FAIL=$((FAIL + 1)); } + +# gate