diff --git a/.claude/memory/decisions.md b/.claude/memory/decisions.md index fb3b918..d785cab 100644 --- a/.claude/memory/decisions.md +++ b/.claude/memory/decisions.md @@ -73,6 +73,7 @@ rules: | BDR-049 | 2026-07-03 | verifier = fresh + blind (no iteration history) + disk-contract + PROOF-or-fail; mute ≠ PASS; scope enrichment via human micro-gate | accepted | | BDR-050 | 2026-07-03 | universal pipeline (contract→dev inline→fresh verify→fresh security, loops bounded 3× in main loop) with per-flow weighting; hotfix failure = revert not loop | accepted | | BDR-051 | 2026-07-04 | contract enrich-at-gate: the contract grows ONLY at a human micro-gate ([gated] marker); the verifier judges the ENRICHED contract, not the seed | accepted | +| BDR-052 | 2026-07-05 | /tour auto mode = branch-as-gate: no mid-run approval gates; unmerged chore branch + per-project TOUR.md = deferred human gate; reconcile report-only; loop bounded 3× | accepted | --- @@ -825,3 +826,11 @@ rules: - **Rationale**: the raw request underspecifies (a one-line "add validation" hides the schema-rejection requirement the design surfaces). If the verifier judged only the seed, every design decision would be unverified. Gating the growth keeps the contract honest (no silent scope creep) AND complete (design criteria are verified). The only flow where the contract is mutable mid-run — bounded to gate moments. - **Alternatives rejected**: freeze the contract at creation (design criteria unverified — the seed is too thin); let the dev enrich (the [[BDR-049]] failure mode — dev justifies everything, scope constrains nothing); a second contract per design (loses the single-reference property). - **Reference**: ship-feature STEP 0e+3, init-project STEP 1+4, feature/verify-loops `1c69de2`. Behavioral GREEN: a `[gated 2026-07-04]` design criterion (reject unknown config keys) was read + judged NOT-MET by a fresh verifier across 3 rounds (dogfood). Builds on [[BDR-049]] [[BDR-050]]. + +## BDR-052 — /tour auto mode: branch-as-gate, declared state read-only + +- **Date**: 2026-07-05 +- **Decision**: /tour (grouped sweep clean+security+reconcile+doc, 1..N projects) runs auto, NO mid-run approval gates. Compensations: (1) fixes on `chore/tour-` via gitflow lib, skill NEVER finish/merge/push — unmerged branch + per-project append-only `.claude/audits/TOUR.md` = the human gate, deferred not deleted; (2) reconcile phase REPORT-ONLY even in auto — target TODO + registries read-only, gaps = `suggested` rows applied later via /reconcile; (3) convergence loop bounded 3× ([[LRN-083]]), residuals reported honestly; (4) security floor = security-auditor (pinned semgrep, [[LRN-047]] BLOCK HIGH/CRITICAL) every iteration + cso posture once (gstack ON); CRITICAL/HIGH contract-changing fix applied but tagged **BREAKING** in report+summary; (5) dirty tree / no develop / no lib → report-only, never stash, never hand-branch. +- **Rationale**: mid-run gates defeat the skill's point (hands-off grouped sweep, user away). Auto-checking TODO reproduces the exact lie /reconcile catches — RED-proven, baseline did it. Branch+report = same approval semantics as audit-delta's 3c gate, moved after the fact where a headless run can afford it. +- **Alternatives rejected**: per-phase AskUserQuestion gates (audit-delta model — blocks headless); one consolidated pre-fix gate (still blocks); auto-edit TODO on oracle proof (inference ≠ approval); plain-branch fallback on non-gitflow repos (violates lib-only doctrine → report-only instead). +- **Reference**: skills/tour/SKILL.md + CLAUDE.md routing (feature/tour-skill `73e6a1c`). TDD trail [[LRN-099]] [[LRN-100]] [[EVAL-014]]. diff --git a/.claude/memory/evals.md b/.claude/memory/evals.md index 1443d81..4eb7e26 100644 --- a/.claude/memory/evals.md +++ b/.claude/memory/evals.md @@ -144,3 +144,11 @@ rules: - **method**: real run, no fixture. Per declared item, oracle vs git/fs: `oracle_path_present` (SKILL.md d3d6ced), `oracle_msg_committed`, `oracle_merge_done` (3 branches merged+deleted), tag v4.0.0 + version.txt. `blk_open` → 3 external (BLK-001/003/009, no drift). `deferrals` (marked) + `contradiction_candidates`. Measurable: 1 primary gap (/release-candidate QUEUED-but-done, oracle-proven) + 3 secondary (header-marker drift) found · 1 false positive rejected · 0 false gap asserted. - **anomalies**: none wrong. 2 capabilities PROVEN that [[EVAL-011]] did NOT: (a) finds UNANTICIPATED gaps — the 3 `[branch X]` headers = a header-marker drift CLASS beyond checkbox drift, not designed-for, caught anyway (merge_done=YES + no local branch). Coverage wider than spec. (b) rejects FALSE POSITIVE on REAL data — `--help` candidate (BDR-001 title ⇄ TODO L134) surfaced as CANDIDATE not verdict; review → both WON'T-BUILD, aligned, not contradiction. Recursive coherence holds OFF-fixture. Design note: NO merge-time header-update hook — merge does merge, /reconcile = periodic catch (separation kept, finding 1). - **action**: keep. Real-world value proven — known gap + 2 unknown + false-positive rejected, zero false assertion. + +## EVAL-014 — /tour GREEN run: 6/6 RED gaps closed, disk-verified; re-verify caught agent's own regression + +- **Date**: 2026-07-05 +- **output**: GREEN subagent run w/ skill on fresh seeded fixture: 3 iterations CONVERGED, 7 commits on `chore/tour-2026-07-04` (unmerged), TOUR.md 18 findings (SEC×6 / CLN×5 / REC×2 / DOC×2 / INF×2), functional suite 8/8 PASS, semgrep PASS(0) final. RED baseline same fixture = 6 gaps ([[LRN-099]]). +- **method**: main session verified ON DISK, not from agent summary: TODO zero-diff vs develop ✓, no target `.claude/memory/` created ✓, per-iteration semgrep report files present ✓, TOUR.md committed ✓, zero scope creep (no .gitignore) ✓, main/develop untouched + branch unmerged ✓, 3-iteration bound held ✓. +- **anomalies**: (1) scratch semgrep files untracked → tree dirty at end, would self-block next run — patched STEP 3.2 [[LRN-100]]; (2) SEC-2 API-BREAKING fix (new required header) unflagged — patched template BREAKING tag; (3) positive: it2 re-verify caught regression of agent's OWN fix (`compare_digest(str)` raises on non-ASCII → 500 not 403), fixed + functionally proven it3 — re-verify loop has real teeth. +- **action**: keep (skill shipped). REFACTOR additions not re-run through 3rd full pass — re-test at first real use ([[LRN-100]]). diff --git a/.claude/memory/journal.md b/.claude/memory/journal.md index 8a86e0a..97c41aa 100644 --- a/.claude/memory/journal.md +++ b/.claude/memory/journal.md @@ -325,3 +325,7 @@ rules: - Merged verify-loops chantier + default-model chore into develop (user pushed). Cut release/4.1.0 (prep + RC gate 8/8 green) — awaiting GO. - rules/ dir built + symlinked via link.sh (feature/rules-dir `06391a6`): real feature verified (paths-scoped lazy rules); context7.md machine-owned → gitignored (find-docs pattern). "contexts dir" request REFUSED — feature doesn't exist (official docs via claude-code-guide); intent already covered by agents/skills. [[LRN-097]]. + +## 2026-07-05 + +- Built /tour skill (grouped sweep clean+security+reconcile+doc, auto, 1..N projects, convergence loop bounded 3×) via writing-skills TDD + skill-creator guidance: RED 6 gaps → GREEN 6/6 closed disk-verified → REFACTOR 2 holes (scratch self-block, BREAKING tag). [[BDR-052]] [[LRN-099]] [[LRN-100]] [[EVAL-014]]. Merged feature/tour-skill → develop + release/1.0.0 on user GO. settings.json /model side-effect reverted (Opus 4.8 1M default restored, attribution backstop kept). diff --git a/.claude/memory/learnings.md b/.claude/memory/learnings.md index 1c7007d..5380088 100644 --- a/.claude/memory/learnings.md +++ b/.claude/memory/learnings.md @@ -117,6 +117,8 @@ rules: | LRN-095 | 2026-07-03 | orthogonal gates don't contaminate — a conformity verifier must PASS correct-but-insecure code (security is a separate gate's job); proven live (CONFORME on a feature carrying a SQLi); fusing the two degrades each | designing multi-dimension review/verify/audit gates | | LRN-096 | 2026-07-04 | a backstop/guard is code — reliable ONLY after a flip-test proves it CAN fail; an unproven guard replacing an advisory = a vacuous guard (LRN-048 applied to guards); flip-test mandatory at guard creation | building any deterministic guard/lint/backstop | | LRN-097 | 2026-07-04 | community blog pattern ≠ official feature — "contexts dir" doesn't exist in Claude Code; verify feature against official docs (claude-code-guide) BEFORE building infra; the intent was already covered by real mechanisms (agents/skills/rules) | any "add support for X" request naming a Claude Code feature | +| LRN-099 | 2026-07-05 | auto-orchestrator autonomy boundary: git discipline transfers naturally (branch, no-merge), declared-state discipline does NOT — baseline silently rewrote target TODO + authored registries + scope-crept | designing any auto/headless flow — enumerate declared surfaces, mark each read-only or gated | +| LRN-100 | 2026-07-05 | tool gated on clean tree must clean its OWN scratch (else self-DoS next run); contract-changing auto-fix needs structural BREAKING flag in the reviewed artifact | any recurring tool w/ cleanliness precondition; any auto-fix touching an API contract | --- @@ -1023,3 +1025,19 @@ rules: - **context**: 2026-07-04 rules-dir chantier. `rules/` (real feature, verified: paths-scoped lazy loading) was built; `contexts/` (nonexistent) was refused with the doc citation. - **future application**: "add support for X" where X is a Claude Code/tool feature — claude-code-guide first, build second. Same discipline for any tool: feature existence is a fact to verify, not assume. - **cousin**: [[LRN-086]] provenance discipline; [[LRN-046]] verify before trust; CLAUDE.md "Never assume — verify". + +## LRN-099 — Auto-orchestrator autonomy boundary: working branch YES, declared/shared state NO + +- **pattern**: /tour RED baseline (no skill, pressure "injoignable, reboucle jusqu'à propre"): git discipline held NATURALLY (gitflow lib branch, no merge w/o signal, atomic commits — doctrine survived into subagent) BUT state-write discipline failed across the board: target TODO silently rewritten (boxes checked, restructured), BDR/journal entries authored autonomously, unrequested bootstrap (.gitignore + registries "bonus hygiene"). Plus: security = ad-hoc grep+ruff (no semgrep floor), findings only in final chat msg (no reviewable artifact), loop unbounded (converged pass 2 by luck). +- **why**: model generalizes commit discipline from doctrine; "declared state = someone's approval surface" NOT in its prior — such writes look helpful. Auto-flow skills must lock declared-state writes explicitly (read-only rules, report-only phases), not just git verbs. +- **context**: 2026-07-04 /tour TDD, seeded fixture (vuln + dead code + lying TODO + stale README). 6 gaps → 6 counters in SKILL.md; GREEN closed all, disk-verified. +- **future application**: designing any auto/headless flow — enumerate SHARED/DECLARED surfaces (TODO, registries, human-facing docs, config), mark each read-only or gated. Never assume git discipline implies state discipline. +- **cousin**: [[LRN-083]] bounded loops in main loop; /reconcile principle (inferred checkbox = the lie). + +## LRN-100 — Clean-tree-gated tools must clean own scratch (self-DoS); breaking auto-fix needs structural flag + +- **pattern**: /tour GREEN left 4 untracked scratch files (`.tour-semgrep*.md`) → tree dirty at end → NEXT run hits own "dirty tree → report-only" precondition = self-block. Same run: HIGH security fix adding required auth header = API-BREAKING, reported plain "fixed" — branch diff doesn't shout contract change. +- **why**: preconditions designed against user WIP also fire on the tool's own residue → scratch cleanup = explicit end-of-run step. Both = omission failures → structural counters (template slot: STEP 3.2 cleanup, BREAKING tag in report template + summary), NOT prohibition prose (writing-skills "match form to failure"). +- **context**: 2026-07-04 /tour GREEN on fixture; both patched at REFACTOR (SKILL.md STEP 3). Additions template-structural, NOT re-run through 3rd full pass (cost) — re-test first real use. +- **future application**: any recurring tool gated on repo cleanliness → audit what IT leaves behind; any auto-applied fix changing a contract → structural BREAKING flag in the human-reviewed artifact. +- **cousin**: [[LRN-099]] same chantier; [[LRN-071]] swallowed-failure class (silent residue ≈ masked state). diff --git a/.claude/tasks/TODO.md b/.claude/tasks/TODO.md index 562cc68..e3915db 100644 --- a/.claude/tasks/TODO.md +++ b/.claude/tasks/TODO.md @@ -1,5 +1,40 @@ # TODO +## 2026-07-04 — skill /tour (tir groupé multi-projets, feature/tour-skill) +Goal: 1 orchestrateur = clean-code + sécurité (security-auditor/semgrep [+cso si +gstack ON]) + reconcile + doc, mode auto, sur 1..N projets. Boucle de convergence +(fixes peuvent invalider l'audit précédent) BORNÉE 3× (LRN-083). Build via +superpowers:writing-skills (TDD, pattern audit-delta/reconcile) + guidance +skill-creator (structure, description trigger-pushy). +Design verrouillé : +- auto = fixes committés sur `chore/tour-` par repo (gitflow lib), JAMAIS + finish/merge (signal humain only). Tree sale ou pas de develop → report-only. +- ordre par repo : sécurité → clean → re-verify (checks projet, fail=revert + fail-closed) → reconcile (REPORT-ONLY, jamais d'auto-coche TODO) → doc + (mode silencieux doc-syncer) → re-audit convergence. +- convergence = 1 passe complète à zéro finding nouveau + checks verts ; + sinon re-boucle, max 3 itérations, résidus rapportés honnêtement. +- rapport `.claude/audits/TOUR.md` par repo + synthèse inline multi-repos. +- registres : offre capitalize gatée en fin, jamais silencieux. +- [x] RED : fixture repo → baseline SANS skill. 6 gaps : TODO cible ré-écrit + silencieusement ; registres écrits de façon autonome ; sécu = grep ad-hoc + sans semgrep ; zéro rapport persistant ; scope creep (.gitignore + + registres bootstrap) ; boucle sans borne déclarée. (Bien fait : branche + gitflow via lib, pas de merge, commits atomiques, convergence passe 2.) +- [x] GREEN : skills/tour/SKILL.md — run avec skill sur fixture-green, + 6/6 gaps fermés VÉRIFIÉS sur disque (TODO zero-diff, 0 registre, + semgrep chaque itération, TOUR.md committé 18 findings, 0 scope + creep, 3 it. bornées convergées, chore branch non mergée) +- [x] REFACTOR : 2 trous du GREEN patchés (scratch semgrep non trackés → + auto-blocage du prochain run, STEP 3.2 cleanup ; fix sécu cassant + non signalé → tag BREAKING structurel dans template). Additions + template-structurelles NON re-testées par un 3e run complet (coût) — + re-test au premier usage réel. +- [x] Routage CLAUDE.md (ligne « Grouped all-axes sweep → tour ») +- [x] Commit branche + capitalize (BDR-052, LRN-099/100, EVAL-014, journal) +- [x] GO user 2026-07-05 : merge develop + release/1.0.0 ; settings.json + restauré (Opus 4.8 1M défaut, backstop attribution conservé) + ## 2026-07-03 — verify loops + semgrep gate + contract (chantier orchestrateurs) Archi validée au gate (session 2026-07-03). Cible : contract sur DISQUE dès création (fichier de run, pattern DIAGNOSIS) + verifier frais (verdict structuré diff --git a/CHANGELOG.md b/CHANGELOG.md index 1d73b6d..e690e41 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -6,6 +6,9 @@ Format follows [Keep a Changelog](https://keepachangelog.com/). ## [Unreleased] +### Added +- `/tour` skill — grouped all-axes sweep over one or several projects: security (pinned-semgrep `security-auditor` agent + `/cso` posture when gstack is ON) → cleanup → re-verify → reconcile (report-only, never edits the target TODO/registries) → doc sync, looping until a full pass applies zero fixes (bounded at 3 iterations). Fixes land on a `chore/tour-` branch the skill never merges; each project gets an append-only `.claude/audits/TOUR.md` report with BREAKING tags on contract-changing security fixes. Built TDD (superpowers:writing-skills): baseline run showed silent TODO rewrites, autonomous registry writes, grep-as-security-pass, no persistent report, scope creep and an unbounded loop — each countered and verified on a seeded fixture. + ### Fixed - `gitflow_finish` ignored its ` ` arguments and always merged the checked-out branch — naming a different branch silently merged the wrong one. The arguments are now an optional safety assertion: if given and not equal to the current branch, `finish` refuses with a clear error instead of merging. No-argument calls (the only real caller) are unchanged. - `doctor.sh` false-warnings removed (a check that cries wolf is one you learn to ignore): `cargo` absence no longer claims "RTK unavailable" (RTK ships as a prebuilt binary); `check_symlink` no longer flags files reached through directory-level symlinks (e.g. `hooks/session-start.sh`); the GStack check counts the per-skill symlinks instead of a `skills/gstack` link that `link.sh` deliberately removes; the token-budget estimate is measured against the ~200k context window instead of a mis-framed "~11k session budget" that produced a false "92% CRITICAL". diff --git a/CLAUDE.md b/CLAUDE.md index f94825a..f17c452 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -269,6 +269,8 @@ only the non-obvious cases: gstack fallbacks, disambiguation, cryptic names. - Cut a release / tag a version (develop ahead of main) → release-candidate - Docs post-ship → document-release (doc if gstack off); stale-doc audit → doc - Audit of changes since last run → audit-delta +- Grouped all-axes sweep (clean+security+reconcile+doc, "tir groupé", + tour of one or more projects, fix + loop until clean) → tour - Open-work inventory / "queue empty?" / stale TODO vs real git → reconcile - Design / UI (build, system, audit, polish) → see "Design work" below - Architecture review → plan-eng-review diff --git a/settings.json b/settings.json index ecbf04d..3e9e71d 100644 --- a/settings.json +++ b/settings.json @@ -226,6 +226,11 @@ "additionalDirectories": [] }, "model": "claude-opus-4-8[1m]", + "attribution": { + "commit": "", + "pr": "", + "sessionUrl": false + }, "hooks": { "SessionStart": [ { diff --git a/skills/tour/SKILL.md b/skills/tour/SKILL.md new file mode 100644 index 0000000..1d96f9f --- /dev/null +++ b/skills/tour/SKILL.md @@ -0,0 +1,263 @@ +--- +name: tour +description: | + Use when the user wants ONE grouped pass over a whole project (or a + list of projects) covering all hygiene axes together: code cleanup + + security (semgrep/cso) + TODO-vs-reality check + doc sync, auto-fixing + and re-auditing until a clean pass. Use it whenever the user asks for + a "tour" of their projects, a grouped/combined audit-and-fix, or a + periodic all-axes sweep — even without naming the axes. + NOT one axis alone (/code-clean, /cso, /audit-delta, /reconcile, /doc), + one bug (/hotfix, /bugfix), dashboard (/health), branch diff (/review). + Triggers: "tour", "tir groupé", "grand ménage", "fais un tour sur les + projets", "sweep", "full pass", "vérifie et corrige tout", "passe + tout au propre". +argument-hint: "[project paths… — blank = current repo] [--report-only]" +allowed-tools: + - Read + - Edit + - Write + - Bash + - Grep + - Glob + - Agent + - AskUserQuestion +--- + +# /tour — grouped multi-axis sweep (clean + security + reconcile + doc) + +One pipeline per project: **security → clean → re-verify → reconcile → +doc → convergence re-audit**, looping until a full pass applies zero new +fixes. Auto mode by design: fixes are committed on a dedicated +`chore/tour-` branch that this skill **never merges** — the branch +plus its report IS the approval gate, reviewed by the human afterwards. + +Core principle: **autonomy on the working branch, never on shared +state.** The skill may edit code freely on its own branch; it may NOT +silently rewrite declared state (target TODO, memory registries) or +integrate anything (merge/finish/push). + +## When NOT to use + +| Situation | Skill | +|-----------|-------| +| One axis only (cleanup / security / TODO / doc) | `/code-clean`, `/cso`, `/reconcile`, `/doc` | +| Recurring single-axis audit scoped to the delta | `/audit-delta` | +| One obvious bug | `/hotfix`, `/bugfix` | +| Quality dashboard, no fixes | `/health` | +| Review a branch/PR diff | `/review`, `/code-review` | +| All axes, fix, loop to clean, 1..N projects | **this skill** | + +## STEP 0 — ARGS & PROJECT LIST + +- Paths in `$ARGUMENTS` → project list, processed **sequentially** in + the given order. No paths → current repo only. +- `--report-only` → run every audit, apply NO fix, write reports only. +- A failure in one project never aborts the tour: record it in that + project's report section and move to the next. + +## STEP 1 — PRECONDITIONS (per project) + +All git commands use `git -C `. Check, in order: + +1. Is a git repository → else SKIP (recorded, not an error). +2. Working tree **clean** (`git status --porcelain` empty) → else this + project runs **report-only**: never mix the user's WIP with tour + fixes, never stash someone else's work. +3. `develop` exists and `~/.claude/lib/gitflow.sh` is available → start + the working branch via the lib, never by hand: + `bash ~/.claude/lib/gitflow.sh start chore tour-YYYY-MM-DD` + (append `-2`, `-3`… if the branch already exists). Missing develop + or lib → **report-only** + suggest `gitflow init` in the report. +4. Detect project checks once (tests, lint, build, type-check — from + package.json/Makefile/CLAUDE.md). Record what exists; "none found" + is itself a report line. + +Report file: `.claude/audits/TOUR.md` in the target project (create +`.claude/audits/` if absent). Append-only — never rewrite past runs. + +## STEP 2 — ITERATION LOOP (max 3 per project) + +Each iteration runs phases A→D in fixed order. **Convergence** = one +full iteration that applies **zero fixes** and finds **zero new +findings** with project checks green. Converged → STEP 3. Not converged +after 3 iterations → STOP, residuals stay `open` in the report, say so +honestly in the summary. Never loop past 3. + +### Phase A — SECURITY (deterministic floor first) + +1. Dispatch the SAST gate (fresh every iteration): + ``` + Agent(subagent_type="security-auditor", description="tour security — semgrep SAST", + prompt="MODE: audit\nSCOPE: project (full tree, respect .gitignore)\nPROJECT: \nREPORT: .claude/audits/.tour-semgrep.md\nFollow agents/security-auditor.md exactly. Pinned rulesets, no login. Write ONLY to REPORT. End with REPORT_WRITTEN: .") + ``` + semgrep ABSENT → DEGRADED (checklist only) is surfaced in the + report, not a silent downgrade and not a blocker. +2. gstack ON (`/cso` available) → **iteration 1 only**, dispatch a cso + posture audit (deps CVE, OWASP) in audit mode; fold its findings in. +3. Fix policy (skip in `--report-only`): CRITICAL/HIGH → fix now. + MEDIUM/LOW → fix only if local and behavior-preserving, else leave + `open`. Every fix minimal, CLAUDE.md security defaults apply. + A CRITICAL/HIGH fix that changes the API contract (new required + header/param, changed status codes, moved paths) is still applied — + but its report row and the global summary line carry a **BREAKING** + tag, so the human review cannot miss it. +4. Commit scoped: `git add ` (never `-A`), + `fix(security): …`. + +### Phase B — CLEAN + +1. Dispatch a read-only cleanup audit (code-cleaner agent if available, + else analyzer/general): dead code, unused imports/exports, + commented-out blocks, stale flags, norm violations. Findings as + `id | file:line | finding | proposed fix`. +2. Apply **behavior-preserving** fixes only. A finding that would change + behavior is a bug, not cleanup → log to + `.claude/audits/BUGS-FOUND.md`, leave the code alone. +3. Commit scoped: `chore(clean): …`. + +### Phase C — RE-VERIFY (after any fix) + +1. Run the project checks found in STEP 1. Lint alone is NOT + verification when tests/build exist. +2. Fresh read-only subagent re-audits the files modified this + iteration (same axis prompts). Pass = approved findings resolved AND + zero new findings introduced. +3. Fail → fix → recheck, max 3 attempts inside the iteration; still + failing → **revert this phase's commits** (fail closed), findings + back to `open`, recorded in the report. + +### Phase D — RECONCILE + DOC + +1. **Reconcile — REPORT-ONLY, always, even in auto mode.** Confront + declared state (target TODO checkboxes, registry statuses) against + real state (git log, files, branches) — reuse `lib/reconcile.sh` + oracles when available. Every gap goes in the report as + `declared X | real Y | suggested edit`. **Never check a box, never + restructure, never edit the target project's TODO.md or + `.claude/memory/`** — an inferred checkbox is exactly the lie + /reconcile exists to catch. The human applies suggestions via + `/reconcile` later. +2. **Doc sync** — dispatch doc-syncer in AUTOMATIC (silent) mode: + public docs only (README, INSTALL, USAGE, CHANGELOG…), never + `.claude/**`, never CLAUDE.md. Commit its `PATCHED_FILES:` via + `bash ~/.claude/lib/doc-commit.sh` when available, else a scoped + `docs: …` commit of exactly those paths. + +### End of iteration + +Fixes were applied (any phase) OR new findings appeared → run another +iteration (fixes can invalidate earlier audits — that is the point of +the loop). Otherwise → converged. + +## STEP 3 — REPORT, CLEANUP & SUMMARY (per project, then global) + +Append to `.claude/audits/TOUR.md`, then close the run in this exact +order: + +1. Write the run section (template below). +2. **Delete the scratch audit files** this run created + (`.claude/audits/.tour-semgrep*` and similar) — their content is + folded into TOUR.md. A tree left dirty here forces the NEXT tour + into report-only: the skill must not self-block. +3. Commit the report as the run's final commit (`docs(tour): report`). +4. Confirm `git status --porcelain` is clean (runtime junk the sandbox + cannot delete, e.g. `__pycache__/`, becomes a report residual line). + +```markdown +## Tour 2026-07-04 — branch chore/tour-2026-07-04 — 2 iterations — CONVERGED +| ID | Axis | File | Sev | Finding | Status | +|----|------|------|-----|---------|--------| +| SEC-1 | security | app.py:17 | high | shell=True + concat | fixed | +| SEC-2 | security | app.py:14 | high | no authz on POST /backup | fixed — **BREAKING**: new required X-Backup-Token header | +| CLN-1 | clean | utils.py:9 | - | dead legacy_md5 | fixed | +| REC-1 | reconcile | TODO.md | - | "/health" unchecked, shipped 2d92696 | suggested | +| DOC-1 | doc | README.md | - | phantom /status endpoint | fixed | +Checks: pytest PASS, ruff PASS. Residuals: none. Commits: 5. BREAKING: 1 (SEC-2). +``` + +Global summary inline, one line per project (append `BREAKING: n` to +any project line whose fixes changed an API contract): + +``` +TOUR COMPLETE — 2026-07-04 + ~/proj/api : CONVERGED (2 it.) — 3 fixed, 1 suggested | chore/tour-2026-07-04, 4 commits + ~/proj/site : NOT CONVERGED (3 it.) — 2 open residuals | chore/tour-2026-07-04, 6 commits + ~/proj/lib : report-only (dirty tree) | no branch + Branches left UNMERGED — review each, then `gitflow finish` on your GO. + Reconcile suggestions pending — apply via /reconcile. +``` + +Then offer to capitalize (gated, per CLAUDE.md): recurring cross-project +patterns → learnings, tour verdict → evals. Never write registries +without that approval — neither this repo's nor any target project's. + +## Rules + +- Branch via the gitflow lib; **never `gitflow finish`, never merge, + never push** — no exceptions, "the tour is green" is not a signal. +- Scoped pathspecs only; `git add -A` is forbidden. +- Target TODO.md and target `.claude/memory/` are READ-ONLY. Reconcile + produces suggestions, not edits. +- Only the four axes. No unrequested bootstrap (.gitignore, registries, + configs, features) — infrastructure gaps are report lines, not work. +- Security floor = security-auditor (pinned semgrep + checklist). An + ad-hoc grep is never "the security pass". +- Max 3 iterations per project; max 3 fix attempts per re-verify. + Residuals are reported, not silently retried forever. +- Dirty tree / no develop / no gitflow lib → report-only, stated in the + report. Never stash, never branch by hand. +- Reports append-only. One report per project, in that project. + +## Common mistakes + +| Mistake | Fix | +|---------|-----| +| Checking TODO boxes "obviously done" during reconcile | Report-only. Suggested edits, human applies. | +| Writing BDR/LRN/journal entries in the target project | Registries only via the gated capitalize offer, end of tour. | +| grep/ruff pass = security done | security-auditor agent (pinned semgrep) is the floor, every iteration. | +| Findings live only in the final chat message | TOUR.md is what the human reviews before merging. Write it. | +| "Bonus hygiene" (.gitignore, templates, bootstrap) | Out of scope. Report line, not work. | +| Loop "until clean" with no bound | Max 3 iterations, then honest residuals. | +| Merging/finishing because everything is green | Green ≠ GO. Branch stays; human merges. | +| Stashing a dirty tree to proceed | Report-only for that project. | +| One TOUR.md for all projects in the config repo | Each project gets its own `.claude/audits/TOUR.md`. | +| Fixing a behavior-changing "cleanup" finding | That is a bug → BUGS-FOUND.md, untouched code. | +| Scratch audit files left untracked at the end | Delete them in STEP 3.2 — a dirty tree self-blocks the next tour. | +| Contract-changing security fix reported as plain "fixed" | Tag **BREAKING** in the row AND the summary line. | + +## Red flags — STOP + +- About to `Edit` a target project's TODO.md or `.claude/memory/*`. +- About to run `gitflow finish`, `git merge`, or `git push`. +- About to `git add -A` or commit on `main`/`develop`. +- Starting iteration 4, or "just one more loop, it's almost clean". +- Security phase done without the security-auditor agent and without a + DEGRADED notice in the report. +- Creating any file the audit did not require (.gitignore, templates). +- Ending a project's run with `git status --porcelain` non-empty and no + residual line explaining every leftover path. + +## TDD note (skill itself) + +Baseline-tested per superpowers:writing-skills (2026-07-04, seeded +fixture, no skill): the agent branched correctly via gitflow and did not +merge, BUT (1) silently rewrote the target TODO (checked boxes, +restructured) during "reconcile"; (2) authored BDR/journal registry +entries autonomously; (3) ran security as ad-hoc grep + ruff — no +semgrep, no pinned rulesets; (4) left findings only in its final chat +message — no persistent report to review before merge; (5) bootstrapped +unrequested .gitignore + memory registries ("bonus hygiene"); +(6) looped without a stated bound (converged at pass 2 by luck). Phase +D.1, the registry rule, Phase A.1, STEP 3, the scope rule and the +3-iteration bound counter each observed failure. + +GREEN run (same day, fresh fixture, skill followed): all six gaps +closed — TODO zero-diff, no registry writes, semgrep every iteration, +TOUR.md committed, no scope creep, converged in 3 bounded iterations on +an unmerged chore branch. Two new holes surfaced and patched +(REFACTOR): scratch semgrep files left untracked (would self-block the +next run — STEP 3.2) and a contract-changing security fix not flagged +(BREAKING tag in template). The REFACTOR additions are +template-structural and were not re-run through a third full fixture +pass — re-test on first real use.