From 0e7f171405f7adeea70eb8ea4061a298a165b105 Mon Sep 17 00:00:00 2001 From: Bastien Chanot Date: Mon, 6 Jul 2026 12:09:30 +0200 Subject: [PATCH] =?UTF-8?q?job2:=20capitalize=20=E2=80=94=20journal=202026?= =?UTF-8?q?-07-06=20+=20EVAL-017=20(audit=20shipped,=20verify-pass=20anoma?= =?UTF-8?q?lies)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- .claude/memory/evals.md | 9 +++++++++ .claude/memory/journal.md | 7 +++++++ 2 files changed, 16 insertions(+) diff --git a/.claude/memory/evals.md b/.claude/memory/evals.md index 22ea9ff..298bc60 100644 --- a/.claude/memory/evals.md +++ b/.claude/memory/evals.md @@ -160,3 +160,12 @@ rules: - **method**: real prod deploy (VPS). Independent live proof post-mark: curl bchanot.fr → 200 + nosniff + X-Frame-Options + CSP + HSTS + versionless server — tour SEC-2 fixed end-to-end, tour→prod loop closed. - **anomalies**: (1) NOT exercised: cold cross-session resume + STEP 4 learn (0 incidents) — natural test at next deploy/failure. (2) UX gap, user feedback: compound `ssh host "cd … && …"` one-liners ≠ wanted session style (one command per line), and the checklist lived only on disk — skill patched same day (step=block grammar, shape rule, hand-back prints NEXT.sh inline; template + bchanot-cv runbook restyled). Re-dogfood at next deploy. - **action**: keep. Two-moment contract works in-session; disk artifacts coherent throughout. + +## EVAL-017 — job2 audit: fresh-context verify pass caught 3 explorer false claims + +- **Date**: 2026-07-06 +- **output**: `.audit/job2-report.md` — 17 findings, 26 diffs, execution prompt. 4 explorers (skills/agents/hooks+lib/registry x-ref) + 1 docs agent (claude-code-guide), then 3 fresh verifiers re-checked all 17 findings + 9 registry quotes from list+paths only. +- **method**: verifiers blind to auditor reasoning. Mid-run session-limit kill all 3 → resumed from transcript via SendMessage, all completed. +- **result**: 15/17 REPRODUCED, 2 PARTIALLY (wording only: F3 "exactly 4"→4-of-54; F14 soft precondition existed). 0 discarded. Registry quotes 9/9 verbatim. Exact char counts 100% match (4840 total agents). +- **anomalies**: 3 explorer false claims, ALL about harness semantics not file content: (1) agents-explorer — `Agent` tool "non-canonical" + `memory:`/`effort:` frontmatter "invalid": wrong, all documented; (2) skills-explorer — skills/gstack/ "stray orphan": refuted by link.sh:54-57 deliberate plumbing; (3) guide agent — `[1m]` model suffix "invalid ANSI": refuted, /model writes it itself. File-content claims (counts, quotes, refs): zero errors. +- **action**: harness-semantics claims from explorers ALWAYS cross-check vs docs/live evidence; file-content claims reliable after one verify pass. diff --git a/.claude/memory/journal.md b/.claude/memory/journal.md index c21a95d..016fc4e 100644 --- a/.claude/memory/journal.md +++ b/.claude/memory/journal.md @@ -331,3 +331,10 @@ rules: - Built /tour skill (grouped sweep clean+security+reconcile+doc, auto, 1..N projects, convergence loop bounded 3×) via writing-skills TDD + skill-creator guidance: RED 6 gaps → GREEN 6/6 closed disk-verified → REFACTOR 2 holes (scratch self-block, BREAKING tag). [[BDR-052]] [[LRN-099]] [[LRN-100]] [[EVAL-014]]. Merged feature/tour-skill → develop + release/1.0.0 on user GO. settings.json /model side-effect reverted (Opus 4.8 1M default restored, attribution backstop kept). - /deploy first real run (bchanot-cv): bootstrap→mark full cycle, live-proven (full security-header stack live — tour→prod closed, tag deploy/2026-07-05). Skill patched post-run on user UX feedback: session-style NEXT.sh (one command per line) + hand-back prints the checklist inline ([[EVAL-016]]); template + generated runbook restyled. impeccable chain + Node 24 baseline shipped develop+RC, pushed. settings.json: +inputNeededNotifEnabled committed (layout unchanged). - /deploy pass 2 (user feedback live): checklist DISPLAY-ONLY — NEXT.sh file eliminated (throwaway artifact, PENDING+runbook regenerate anywhere), hand-back ends the turn with the checklist as final text (a print above AskUserQuestion never reached the user, [[LRN-102]]). Skill+template+CHANGELOG patched; legacy NEXT.sh removed from bchanot-cv; deploy run 2 (residuals b24c58b) re-handed-back inline. + +## 2026-07-06 + +- job1 fixes merged develop (`c6d5e03`): CLAUDE.md gitflow density pass, F14 hook pointer-only, line-count guard, [[LRN-103]]. +- job2 config-smell audit shipped read-only: `.audit/job2-report.md` — surface skills/agents/hooks/plugins/settings(.local), 17 findings (3 RISK perms, 6 DRIFT, 2 BLOAT, 3 OVERLAP, 2 DEAD, 1 struct), 26 diffs base c6d5e03, 0 decision-conflicts, all fresh-context verified [[EVAL-017]]. Live catch: design hook fired on audit's own task-notifications (14/20 recent fires). +- Brief premise corrected: Edit/Bash(hooks/*.sh) permission rule NEVER existed — was config-protection case arm (:37) + job1 sentinel bypasses. Phase-0 UNREFERENCED metrics 100% broken (grep -q kills -l). +- User GO full execution incl. 3 RISK: cp/mv→ask, find -exec deny mirror, settings.local prune (python3 -, rtk git *). F9 fable default committed (user re-chose via /model), F16 gitflow-migrate.sh removed (git-recoverable), F8/find-docs skip (generator-owned). Executor = Sonnet subagent on chore/job2-fixes, NO finish.