New lib/tests/fixtures/decisions-snapshot.md (neutral name, LRN-077
style): carries a --help token (so reconcile_contradiction_candidates
still surfaces the BDR-001 ⇄ --help-chantier candidate against
todo-snapshot.md), a "one-line ticket" line, and representative
OUT-OF-SCOPE/DEFERRED/follow-up context. T3 and T5 in run-reconcile.sh
now read this fixture instead of the LIVE $MEM/decisions.md; deleted
the $MEM variable definition and its stale comment. Closes J4-10
(FIXTURE-DRIFT): T3/T5 were the last live-registry reads in this
suite (T2 was fixed in job3-B1) — any legitimate prune/reword of the
real decisions.md would have reded the suite for a reason unrelated
to the reconcile engine itself.
grep -c '$MEM' lib/tests/run-reconcile.sh == 0 (verified).
GREEN: real repo, 26/26 passed (all 4 T3 markers + T5 candidate found
via the fixture).
Red demo (per spec — no code mutation, this is a fixture-substitution
spec): lean scratch copy, pointed T3's decisions-arg at /dev/null
transiently → "one-line ticket" (the only marker living solely in the
decisions-side fixture, not in todo-snapshot.md) goes missing, RED;
the other 3 markers stay green (satisfied by todo-snapshot.md alone).
Proves the assertions actually read the fixture rather than passing
vacuously.
New T7 block in lib/tests/run-reconcile.sh (+6 assertions, 20→26):
a throwaway git repo under mktemp with a LOCAL BARE origin drives the
3 previously-unexercised oracles live: tree_clean (dirty→rc≠0, clean→
rc0), pushed (pushed to origin FIRST so origin/main exists — else
rev-list is vacuously empty — then rc0 when synced, rc≠0 once 1 ahead),
msg_committed (rc0 for a present commit message, rc≠0 for an absent
one). Closes J4-12 (DEGRADED, prerequisite of SPEC-09/10): these 3
oracles backed report-only /reconcile output with zero test coverage —
a silent inversion would mis-report open-work state.
Mutation (lean scratch copy — only lib/reconcile.sh + lib/tests/
run-reconcile.sh + its fixtures + .claude/memory/decisions.md, not the
whole repo/.git, per the /tmp-exhaustion lesson from SPEC-01/02/04):
inverted tree_clean's rc (`-z` → `-n` on the porcelain-status check;
the report's literal "--quiet → negated" wording doesn't match this
function's actual `[ -z ... ]` shape, so applied the equivalent
semantic inversion). RED: both T7a assertions fail (dirty reads as
clean and vice versa); T7b/T7c stay green, confirming the mutation is
localized. (T6a/b/c red in the lean copy too, expected — no real git
history / skills dir there — unrelated to the mutation.)
GREEN: real repo unmutated, 26/26 passed (T7 included).
New T15 block in lib/gitflow-test.sh (+7 assertions, 83→90): fresh git
init sandbox with NO identity (GIT_CONFIG_GLOBAL=/dev/null
GIT_CONFIG_SYSTEM=/dev/null, git 2.53 supports the override) →
gitflow_init must return rc 1 AND leave zero mutation: no develop
branch, unborn HEAD, hooksPath unset, nothing staged, no .gitignore/
.githooks written. Closes J4-06 (WEAK): every test repo up to now set
an identity first, so this precheck never fired.
Mutation (lean scratch copy — only lib/gitflow.sh + lib/gitflow-test.sh
+ templates/gitignore/standard.gitignore, not the whole repo/.git, to
avoid repeating the /tmp exhaustion from the SPEC-01/02/04 full-repo
copies): deleted the identity precheck (gitflow.sh:178-179). RED: 3/7
T15 assertions fail — "nothing staged", "no .gitignore written", "no
.githooks written" — while rc stays 1 and HEAD stays unborn (git itself
still refuses the identity-less commit). This is the half-applied-init
failure mode named in the finding (BLK-012 class): same exit code, but
now via a partial mutation instead of a clean upfront refusal — exactly
why the spec pins zero-mutation checks beyond rc alone.
GREEN: real repo unmutated, 90/90 passed (T15 included).
New T14 block in lib/gitflow-test.sh (+3 assertions, 80→83), direct
.githooks/pre-commit invocation (T10-style): T14a mixed code+.claude
staged together on main → BLOCKED (whitelist must not let code ride
along .claude/). T14b MERGE_HEAD present + code staged on main →
exit 0 (conflict-resolution commit exemption, gitflow.sh:222). T14c
hook installed+activated BEFORE the first commit (gitflow_install_hook,
not gitflow_init's deferred activation) → root commit still succeeds
(gitflow.sh:221). Closes J4-05 (WEAK): these 3 exemption paths were
untested — a whitelist regression, or the root/merge exemptions
breaking, would have been silent.
Mutations (scratch copy, applied via Bash/sed — not Edit/Write, avoids
tripping config-protection's path-suffix guard on lib/gitflow.sh for a
throwaway file that's never committed), one at a time, each reverted
before the next:
- T14c: deleted the root-commit guard (gitflow.sh:221,
`git rev-parse --verify -q HEAD ... || exit 0`) → T14c reds alone.
- T14b: deleted the MERGE_HEAD guard (gitflow.sh:222) → T14b reds alone.
- T14a: report's candidate mutation ("remove grep -v '^\.claude/'")
self-corrects (still blocks mixed, via the inverted over-blocking
direction — doesn't red). Used the pinned alternative instead:
`head -1` → `head -0` in the whitelist check (gitflow.sh:230),
neutering the non-empty test so every protected-branch commit is
wrongly allowed. T14a reds, plus (expected, same root cause) the
pre-existing T3 "block direct code on main" and T10 DRIFT(main)/
DRIFT(develop) also red — consistent with a whitelist regression
of this shape being a broad, not narrow, break.
GREEN: real repo unmutated, 83/83 passed (T14a/b/c included).
New T13 block in lib/gitflow-test.sh (+9 assertions, 71→80):
T13a release finish → main gets the commit, develop gets it via
merge-back, release branch deleted. T13b two open releases + a
finished hotfix → hotfix commit present in BOTH release branches.
T13c bugfix finish → develop only, main untouched, branch deleted.
Closes J4-02 (CRITICAL): a half-landed release (main-only or
develop-only) or a mis-based bugfix finish was invisible to the
only test suite that exercises gitflow_finish's fan-out.
Mutation (scratch copy, applied via Bash/perl — not Edit/Write, so
config-protection's path-suffix guard on lib/gitflow.sh isn't
tripped for a throwaway file that's never committed): deleted the
develop merge-back line in gitflow_finish's release arm
(gitflow.sh:122-125). RED: T13a fails 3/3 (rc 5 — _gitflow_delete
refuses because develop never got the merge, so the branch isn't
fully merged; develop missing the commit; branch not deleted).
GREEN: real repo unmutated, 80/80 passed (T13a/b/c included).
Makefile test target now loops lib/tests/run-*.sh in addition to
*.test.sh + gitflow-test.sh, special-casing run-release-candidate.sh
with RC_WORK=$(mktemp -d) RC_TAG=1. Closes J4-01 (CRITICAL): the 5
run-*.sh suites (memory-commit 13, doc-commit 32, doc-shape 19,
reconcile 20, release 5/5) were excluded from the repo's only
aggregate gate.
Mutation (scratch copy, never the working tree): dropped the
`-- "${changed[@]}"` pathspec from lib/memory-commit.sh:86's commit
call. RED: run-deterministic.sh T2 fails (pre-staged dangling code
gets embarked instead of staying staged), make test exits 2.
GREEN: real repo unmutated, make test exits 0, all suites incl. the
5 previously-excluded ones (RESULT: 13/32/19 passed, 20 GREEN, 5
GREEN RC_TAG=1).
skills/reconcile/SKILL.md:53 claimed "20/20, shellcheck clean" but the
suite read .claude/memory/blockers.md live, so closing BLK-009
(d1e7423) turned T2b/T2c red for a correct reason unrelated to the
engine. Froze a post-BLK-009 snapshot (lib/tests/fixtures/
blockers-snapshot.md) and pointed T2 at it instead of the live file —
same approach the other T1/T3/T4 fixtures already use. Updated T2b/T2c
expectations to match (BLK-009 resolved, open = {001,003}). Suite is
back to 20/20 GREEN, shellcheck clean, matching the skill's own claim.
EVAL-018: job3 shipped, 46/46 findings verified, 20/23 fixes applied
(B1 blocked on sentinel scope, D2-D5+B6 skipped by decision), zero
residual on final re-sweep. LRN-105: explorer subagents need an
explicit ban on executing the subject-under-test's own CLI, not just
"read-only" framing (caught mid-run: a subagent ran `graphify .`).
Both sections never matched any component (envelope §1-§9 vs the real
§0-§15 structure; 8-axis integer weights vs the agents' 7/4 and 6/5
percentage-weight scoring). Kept what this file legitimately owns: the
depth-decision matrix and the sibling-skill dedup rules.
- A1: delete STEP 5d (graphify --mode quick doesn't exist in the CLI; the
command always failed, masked by `|| true` — STEP 10's full pass already
covers the graph)
- A2: STEP 10 full-pass build flag --output -> --out
- A8: pipeline is 12 steps (STEP 0-11), not 11; header template unchanged
(N/11 correctly denotes the max index of a 0-indexed 12-step sequence)
- A2: graphify build flag --output -> --out
- A4: ROADMAP xref points to the real /onboard add gsd path, not a nonexistent STEP 9 decision
- A5: exact gitflow init commit message (matches lib/gitflow.sh:163)
- A6: bare skill names (design-review, browse) — no gstack: namespace exists
- A7: eval pattern for recommend_anim_install_cmd (the function only echoes; must eval its output)
BDR-038 recorded NEXT.sh file + AskUserQuestion hand-back as the /deploy
design; 52f6678 removed both (LRN-102: pre-tool-call text may never render)
with no superseding decision. BDR-054 regularizes it. One-line banners on
docs/plans/2026-06-27-deploy-skill.md and docs/specs/2026-06-27-deploy-skill-design.md
point to the shipped behavior; historical body left untouched.
Live failure (run 2): the checklist printed above AskUserQuestion never
reached the user. Fix is structural: the checklist is never written to a
file (throwaway — PENDING.json + live runbook regenerate it in any
session) and every hand-back/re-display ends the turn with the full
checklist as the FINAL text, no tool call after it. Cold resume without a
report regenerates + re-displays. Artifacts 5 -> 4 files; bootstrap
gitignore step drops NEXT.sh; mistakes/red-flags updated (no tool call
after the print, no file 'for reference').
Generator-owned files updated by the out-of-band 'make plugin' run
(SKILL.md, extraction-spec, query reference, version marker). Committed
deliberately after diff review; CHANGELOG Unreleased notes the bump.
First-real-run UX feedback (EVAL-016): one command per line as typed in an
interactive session (ssh opens the box, following lines run on it, local
steps flagged), never folded ssh compounds; the hand-back prints the full
checklist in the conversation (and every re-hand-back reprints it). Step
defined as a block (header + command lines to next blank line), @delta
governs the block. Template restyled to match.
Fresh installs and too-old hosts now get Node 24 (NodeSource setup_24.x /
brew node@24). GSD v2 (>=22) still satisfied. Next 'make plugin' on a
Node-22 host upgrades in place and unlocks impeccable Step 8d.
Complementary to frontend-design (kept: build-time aesthetic direction).
impeccable adds what the chain lacked: 45 deterministic anti-slop rules
(npx impeccable detect, exit 0/2, --json) — the design counterpart of the
semgrep gate — plus 23 design verbs under one /impeccable skill and
persistent per-project design context.
- plugins.lock.json: CLI pinned 3.2.0 (rules update = audit output change
on unchanged code, LRN-077 class); skill dist = its own release track
- install-plugins.sh Step 8d: staged npx install (tmpdir) -> moved to
skills-external/impeccable (machine-owned, gitignored, ctx7 pattern);
never writes through the ~/.claude/skills symlink into the tracked tree
- update-all.sh: pin-honored refresh, Node<24 or failure -> dist kept
- Node >= 24 required (host at 22): steps skip gracefully, activation
deferred to a deliberate Node bump
- link.sh EXTERNAL_SKILLS, profiles (design/web/web-full/full),
plugin-advisor, CLAUDE.md design routing, design-gate, README, CHANGELOG
- NOT in design GATE-BLOCK yet: promotion after first dogfood
Claude Code loads modular rule files from ~/.claude/rules/ (user scope,
recursive, markdown, optional paths: frontmatter for lazy path-scoped
loading — stable, symlink-supported). The repo had no rules/ at all, so
the ctx7 setup had created ~/.claude/rules as a REAL directory outside
version control — invisible to the repo, unreproducible on a new machine.
- rules/README.md — doctrine: one rule per file; paths:-scoped extraction
is the token win, always-on doctrine stays in CLAUDE.md
- link.sh — rules added to the symlinked-dirs loop
- .gitignore — rules/context7.md ignored (machine-owned: `ctx7 setup`
(re)writes it, same treatment as skills/find-docs/)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XpphkdTosUzokBDNG7PToS
Pin the default model to claude-opus-4-8[1m] in settings.json (the tracked
config deployed to ~/.claude). Previously no model key was committed, so the
default resolved to the tier default; a working-tree pin to claude-fable-5[1m]
was never committed. Takes effect at next session start.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XpphkdTosUzokBDNG7PToS
ship-feature: STEP 0e CONTRACT (request verbatim) → ENRICHED at the STEP 3
validation gate (design criteria appended [gated <date>], the human
micro-gate) → STEP 5 VERIFY+SECURE judges the branch against the ENRICHED
contract via the shared include. Distinct axis from STEP 6 code review, both
run (LRN-095).
init-project: contract seeded from the PROJECT BRIEF (V1 features → criteria)
→ ENRICHED at VALIDATION GATE #1 → STEP 9 VERIFY+SECURE. Adds the security
gate init-project previously lacked (was deferred to a later /onboard).
onboard: explicit NO verify-loop — it produces an audit report, not a change
to verify against a request; contract is scope-only, security-auditor runs
MODE audit (report-only), never a gate. Documented to prevent a misplaced
symmetry loop (BDR-050: dev pipeline != audit).
lib/tests/no-vacuous-locks.test.sh: deterministic backstop for LRN-093 (2nd
recurrence in this chantier → the advisory alone did not hold). Refuses a
literal \n in any grep/tf/tr_/tn pattern across lib/tests/*.test.sh;
flip-tested against a synthetic offender so the guard proves it bites.
lib/tests/loops-heavy.test.sh: 18 structure locks green.
Behavioral dogfood (both vigilance points, real): (1) enrichment — a fresh
verifier reads and checks a [gated] design criterion (ECARTS naming it
precisely); (2) escalation — 3 consecutive ECARTS on the same criterion →
orchestrator STOPs at the max-3 bound + presents the CONTRACT-vs-REALIZED
table, no 4th loop, no commit. First real exercise of the infinite-loop guard.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XpphkdTosUzokBDNG7PToS
lib/verify-secure-loop.md: shared main-loop include. GATE 1 fresh verifier
(blind, contract from disk) → CONFORME straight to GATE 2, ECARTS loop max
3; GATE 2 fresh security-auditor (MODE gate) → PASS to commit, BLOCK loop
max 3 with re-verify-request-FIRST order invariant. Mute agent never a PASS.
feater.md: STEP 0.7 CONTRACT (proportional, silent on a clear feature) +
STEP 3 VERIFY+SECURE via the include. Nominal = one verifier + one security
dispatch; the loop only costs when it loops.
bugfixer.md: STEP 3.5 CONTRACT fed by the DIAGNOSIS (bug report verbatim +
reproduced-then-gone + regression test criteria) + STEP 5 fresh gates via
the include. Renumbered STEP 5 sub-steps (gates before the commit gate).
hotfixer.md: STEP 1.7 CONTRACT (silent autofill, zero questions) + STEP 3
security gate whose FAILURE REVERTS (git restore to pre-flight SHA + escalate
to /bugfix), never loops — the 1-attempt model preserved. No fresh verifier
at hotfix weight (the smoke-check verifies the trivial contract). Adds the
Agent tool to hotfixer.md + hotfix/SKILL.md for the security dispatch.
lib/tests/loops-light.test.sh: 27 structure locks green, shellcheck clean.
Behavioral pipeline dogfood on a fixture (feat adding a feature WITH a SQLi):
GATE1 CONFORME (feature present, SQLi not a conformity gap — orthogonal
gates) → GATE2 BLOCK(1) (checklist caught the %-interp SQLi semgrep's taint
rules missed) → [fix to parameterized] → re-verify CONFORME (order invariant,
feature intact) → re-scan PASS. Loop converges to green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XpphkdTosUzokBDNG7PToS
agents/security-auditor.md: fresh read-only-on-code SAST gate. Pinned
rulesets p/security-audit + p/secrets + p/owasp-top-ten (owasp REQUIRED —
measured: the 2-ruleset baseline misses SQLi + path-traversal entirely on
realistic Flask code), never --config auto, never auto login (BDR-048).
Severity map: secrets ERROR → CRITICAL, other ERROR → HIGH (block),
WARNING/INFO → reported. gate mode (diff, no Write) vs audit mode (Write
only to REPORT, rule-locked). DEGRADED (semgrep absent) still runs the
checklist and still blocks — never a vacuous pass (LRN-048). Anti-gaming:
a new un-gated nosemgrep suppression is BLOCKING. PROOF mandatory, mute
auditor never a PASS, blind (no iteration history), blocks HIGH/CRITICAL
only (LRN-047).
Grafts: onboard STEP 6 L3a dispatches it in audit mode (report
.onboard-audit/semgrep.md) in BOTH gstack branches — complement to cso
(cso is a gstack submodule, unmodifiable); synthesis picks it up via the
existing .onboard-audit/ sweep. audit-delta security axis runs the SAST
pass first, folds findings into the existing gate/fix/re-verify flow.
lib/tests/security-auditor.test.sh: 28 structure locks green, shellcheck
clean. Behavioral dogfood (fresh agents on a planted fixture):
BLOCK(9) on the vuln commit (2 secrets→CRITICAL, semgrep+checklist
complementarity — checklist caught the 6 semgrep missed off-context);
BLOCK(1) on a new nosemgrep suppression (understood semgrep's 0 was the
mask); DEGRADED → BLOCK(7) on grep-detectable secrets with semgrep hidden.
FP measured on real repos (faunosteo, game): owasp adds only hygiene
findings, contained by diff-scoping.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XpphkdTosUzokBDNG7PToS
lib/contract-interview.md: mandatory upstream passage for all orchestrators.
Verbatim REQUEST (immutable), proportional questions (complete request =
zero, max 3 one batch), testable criteria + file scope, written to disk
immediately (.claude/tasks/contracts/<date>-<slug>-<HHMM>.md — a context-only
contract dies at compaction). Lifecycle: enrichment only at human gates
([gated] marker, scope micro-gate), supersedes for re-scope, aborted runs
deleted or committed status:aborted — never left dirty. Hand-off = path,
not restatement.
agents/verifier.md: fresh read-only verifier. Reads the contract from disk,
renders VERIFY — VERDICT: CONFORME | ECARTS(n) | ERROR. Blind: never
receives iteration history. PROOF line mandatory (LRN-048), UNVERIFIABLE
never MET, mute verifier never a PASS (retry once fresh, 2nd structural
failure = human escalation). Orchestrator protocol documented in-file
(max 3 iterations, re-verify request before security).
lib/tests/contract-verifier.test.sh: 31 deterministic structure locks on
the load-bearing doctrine clauses — green, shellcheck clean.
Behavioral dogfood (2 fresh subagents on a planted fixture): gap case →
ECARTS(2) exactly as planted (NOT-MET located + out-of-scope flagged);
conform case with injected fake iteration history → CONFORME, noise
ignored, real python spot-check as evidence. Both outputs parse-clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XpphkdTosUzokBDNG7PToS
Step 7.5 in install-plugins.sh: pipx install semgrep==<pin> behind a
command -v guard (LRN-085 pattern), version echo on skip, login is
Pro-rules-only guidance — never run automatically (ctx7 pattern).
Step 6.2 in update-all.sh: pin-honored update that displays the version
jump (cur → pin) before pipx install --force; latest only when unpinned.
plugins.lock.json: semgrep pinned 1.168.0 — semgrep is a BLOCKING gate,
a silent upgrade means new BLOCKs on unchanged code (gsd-pin pattern).
Dogfooded via extracted real blocks: fresh install, idempotent re-run,
pin-match skip, jump display + clean warn on bogus pin. Rulesets
p/security-audit + p/secrets fetch anonymously (no login) and detect.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XpphkdTosUzokBDNG7PToS