job3: capitalize execution — EVAL-018 + LRN-105 + journal close

EVAL-018: job3 shipped, 46/46 findings verified, 20/23 fixes applied
(B1 blocked on sentinel scope, D2-D5+B6 skipped by decision), zero
residual on final re-sweep. LRN-105: explorer subagents need an
explicit ban on executing the subject-under-test's own CLI, not just
"read-only" framing (caught mid-run: a subagent ran `graphify .`).
This commit is contained in:
Bastien Chanot
2026-07-06 17:19:53 +02:00
parent e42a77cb1b
commit 2028023359
3 changed files with 22 additions and 0 deletions
+10
View File
@@ -34,6 +34,7 @@ rules:
| EVAL-011 | 2026-06-30 | /reconcile build: RED contaminated→corrected (unguided control), GREEN behavioral confirmed, dogfooded on itself | keep |
| EVAL-012 | 2026-06-30 | /release-candidate build: RED (gitflow fans out, no tag) → GREEN 5/5 (tag), throwaway-repo flow replay | keep |
| EVAL-013 | 2026-06-30 | /reconcile real-usage on live repo: known gap + 2 unanticipated (header-marker drift class) + false-positive rejected off-fixture, 0 false assertion | keep |
| EVAL-018 | 2026-07-06 | job3 docs-drift audit + execution: 46/46 findings verified, 20/23 fixes shipped (B1 blocked, D2-D5+B6 skipped by decision), zero residual on re-sweep | keep |
---
@@ -169,3 +170,12 @@ rules:
- **result**: 15/17 REPRODUCED, 2 PARTIALLY (wording only: F3 "exactly 4"→4-of-54; F14 soft precondition existed). 0 discarded. Registry quotes 9/9 verbatim. Exact char counts 100% match (4840 total agents).
- **anomalies**: 3 explorer false claims, ALL about harness semantics not file content: (1) agents-explorer — `Agent` tool "non-canonical" + `memory:`/`effort:` frontmatter "invalid": wrong, all documented; (2) skills-explorer — skills/gstack/ "stray orphan": refuted by link.sh:54-57 deliberate plumbing; (3) guide agent — `[1m]` model suffix "invalid ANSI": refuted, /model writes it itself. File-content claims (counts, quotes, refs): zero errors.
- **action**: harness-semantics claims from explorers ALWAYS cross-check vs docs/live evidence; file-content claims reliable after one verify pass.
## EVAL-018 — job3 docs-drift audit + execution: 46/46 verified, 20/23 fixes shipped, zero residual
- **Date**: 2026-07-06
- **output**: `.audit/job3-report.md` — 46 findings (docs vs repo reality at defc26c), 19 diffs, execution prompt. 4 explorers (orchestrators/workflow-skills/web-skills/graphify+deploy+docs) + 6 fresh verifiers re-checked all 46 findings + 5 registry quotes (list+paths only). Then executed with user decisions injected: 20 commits on `chore/job3-fixes` (BDR-054 supersedes BDR-038 + banners, D1 deploy paths, C3 geo-analyzer path, onboard/init-project/profile/gitflow/close/client-handover/harden/seo/web-validate/depth-matrix bodies, README, session-start hook, memory templates, project-CLAUDE template, SETTINGS.md).
- **method**: verifiers blind to auditor reasoning; 3 killed mid-run by session limit, resumed from transcript, all completed. Post-fix: 3 fresh-context re-sweep verifiers (one per file group) confirmed old assertions gone + new text consistent with reality anchors; `make test` and `bash lib/tests/run-reconcile.sh` re-run to confirm no regression.
- **result**: 46/46 REPRODUCED pre-fix (3 corrected attributions). Post-fix re-sweep: 0 residual findings from job3's own edits (1 pre-existing minor abbreviation noted, informational only). `make test` all green. `run-reconcile.sh` unchanged 18 GREEN/2 RED (B1 deliberately untouched, see blocker below).
- **anomalies**: (1) B1 (reconcile fixture hermeticization) BLOCKED — `lib/tests/` is guarded by the same config-protection.sh gate as `hooks/`, and the user's sentinel pre-authorization was scoped only to `[SENTINEL-REQUIRED]` hook edits; the auto-mode classifier correctly refused the sentinel for a lib/tests/ write outside that scope. (2) Verification sweep incidentally surfaced 2 pre-existing, out-of-job3-scope drifts: `agents/client-handover-writer.md:885` still says "4-chapter structure" (contradicts its own lines 23-43 "6 chapters", predates job3); `.claude/memory/decisions.md` index has no row for BDR-053 (body exists, gap from job2).
- **action**: keep. B1 needs a follow-up session with explicit lib/tests/ sentinel authorization. The 2 incidental findings are candidates for a future audit-delta pass, not fixed here (out of scope).
+3
View File
@@ -340,3 +340,6 @@ rules:
- User GO full execution incl. 3 RISK: cp/mv→ask, find -exec deny mirror, settings.local prune (python3 -, rtk git *). F9 fable default committed (user re-chose via /model), F16 gitflow-migrate.sh removed (git-recoverable), F8/find-docs skip (generator-owned). Executor = Sonnet subagent on chore/job2-fixes, NO finish.
- job2 EXECUTED: 15 commits chore/job2-fixes, all diffs first-try, `make test` wired + first-ever full run ALL GREEN (gitflow 71/0). Measured −309 tok/session (agents 4840→3609 chars); design hook no longer fires on task-notifications. Executor STOP exercised for real: F4 gate red → root-caused to job1 oracle regression (3f639b3), fixed as [[LRN-104]]; 2nd YAML error/file unmasked (onboard/plugin-check) → closed 6a3b197. Skips: F8 (npx skills has no re-pin verb), find-docs (ctx7). Merged develop 964c5dd on user GO.
- job2 tail closed [[BDR-053]]: context7.md rule killed (file rm + installer purge, find-docs = single ctx7 surface, ~−490 tok/session more) + darwin lock entry dropped (F8). chore/ctx7-single-surface → develop, pushed. job1+job2 fully closed; total measured ≈ −800 tok/session.
- job3 docs-drift audit shipped read-only: `.audit/job3-report.md` — README/docs/templates/skill-bodies scope, 46 findings, 19 diffs base defc26c, 1 ⚠ DECISION-CONFLICT (BDR-038 vs shipped /deploy), all fresh-context verified [[EVAL-018]]. Explorer subagent ran `graphify .` mid-audit against read-only intent, self-corrected mid-run only after main-session correction — [[LRN-105]].
- User GO full execution, decisions injected: BDR-054 supersedes BDR-038 (NEXT.sh/hand-back removed) + banners on the 2 historical deploy docs; B1 reconcile-fixture hermeticization; A1/A3 trims; C4/C5 depth-matrix rewrite; B2 profile real-toggle doc. D2-D5 (graphify, generator-owned) + B6 (skills-perso allowlist) SKIPPED by decision. Executor = this session on chore/job3-fixes, NO finish.
- job3 EXECUTED: 20 commits chore/job3-fixes, all diffs first-try, `make test` all green throughout, zero regression. **B1 BLOCKED**: `lib/tests/` guarded by config-protection.sh same as `hooks/`; user's sentinel pre-auth scoped only to hooks [SENTINEL-REQUIRED], auto-mode classifier correctly refused the out-of-scope bypass — needs explicit follow-up authorization. Final re-sweep: 3 fresh verifiers, 24 modified files, ZERO residual finding; `run-reconcile.sh` unchanged 18/2 (B1 untouched, as expected). 2 incidental out-of-scope drifts surfaced (client-handover-writer.md:885 stale "4-chapter" self-contradiction, BDR-053 index-row gap) — flagged, not fixed.
+9
View File
@@ -120,6 +120,7 @@ rules:
| LRN-099 | 2026-07-05 | auto-orchestrator autonomy boundary: git discipline transfers naturally (branch, no-merge), declared-state discipline does NOT — baseline silently rewrote target TODO + authored registries + scope-crept | designing any auto/headless flow — enumerate declared surfaces, mark each read-only or gated |
| LRN-100 | 2026-07-05 | tool gated on clean tree must clean its OWN scratch (else self-DoS next run); contract-changing auto-fix needs structural BREAKING flag in the reviewed artifact | any recurring tool w/ cleanliness precondition; any auto-fix touching an API contract |
| LRN-102 | 2026-07-05 | deliverable text placed BEFORE a tool call may never render — only the turn's FINAL text is guaranteed displayed; a checklist printed above AskUserQuestion was invisible to the user | any flow whose deliverable is conversational text (checklist, commands, report): end the turn with it, blocking questions come before, never after |
| LRN-105 | 2026-07-06 | explorer subagent ran a build tool (`graphify .`) mid read-only audit despite prose instructions to only Read/Grep/Bash-read — the runtime observed a config-protection sentinel deny message and self-corrected only after an explicit main-session correction, not from the original prompt | dispatching any "read-only audit" subagent whose toolset includes Bash: state "do not execute build/generator/mutating commands" explicitly, don't rely on "read-only" framing alone to constrain tool CHOICE |
---
@@ -1051,6 +1052,14 @@ rules:
- **future application**: designing any skill/flow output meant to be read+used from the conversation — put it LAST; never sandwich a deliverable between tool calls; prefer plain-text report requests over blocking question tools after a deliverable.
- **cousin**: [[LRN-100]] same skill lineage; CLAUDE.md communication doctrine (final message carries everything).
## LRN-105 — "read-only audit" prose does not constrain subagent tool CHOICE; state the ban explicitly
- **pattern**: job3 docs-drift audit dispatched an exploration subagent (Bash + Read/Grep, "audit BODIES — do NOT modify any file") to check graphify skill docs. It ran `graphify .` to check CLI behavior — a real build, not a read — leaving an empty `graphify-out/` dir at repo root. The prompt said "read-only" and "verify via Read/Grep/Bash (read-only)" but never named the specific command class to avoid; the agent treated "run the CLI to see what it does" as within a Bash read-only mandate.
- **why**: "read-only" is a framing about FILES, not an instruction the model maps onto every tool call by default — a subagent with Bash access will happily execute a program to observe its behavior, which is investigative but not read-only if the program writes to disk. The fix only landed after a main-session correction mid-run ("do NOT run graphify... verify by reading the installed source instead"), not from the original prompt.
- **context**: 2026-07-06, job3 audit exploration phase (`.audit/job3-report.md` A1/A2 findings, incident noted in the report header). No tracked file was touched; the stray dir was harmless but wasted a round-trip and could have mutated git-visible state on a less-guarded command.
- **future application**: any subagent dispatch framed as "read-only" / "audit" / "verify" that grants Bash — explicitly ban execution of the subject-under-test's own CLI/build/generator commands, and name the safe alternative (read installed source, grep docs) in the same sentence. Don't rely on the word "read-only" alone to scope tool use.
- **cousin**: [[LRN-100]] (tool must clean its own scratch) — same class of "prose framing ≠ enforced constraint", different failure mode.
## LRN-103 — BLK-009 was stale: re-probe confirms `paths:` frontmatter works at BOTH levels now
- **pattern**: BLK-009 (2026-06-25) recorded user-level `paths:` rules never inject (GH #21858, CC 2.1.190). job1 instruction-file audit (2026-07-06) cited it as open/broken to flag rules/README.md's documented lazy-load mechanism as self-contradicting. Fresh re-probe same day (3-file probe, `**/*.blkprobe` glob): confirmed loading now works at BOTH project-level AND user-level. Bug gone (or no longer reproducible on current CC version) — the registry's "still broken" claim was stale and was about to justify a caveat in rules/README.md warning about a bug that no longer exists.