Merge feature/verify-loops into develop

This commit is contained in:
Bastien Chanot
2026-07-04 04:54:47 +02:00
24 changed files with 1162 additions and 33 deletions
+37
View File
@@ -69,6 +69,10 @@ rules:
| BDR-045 | 2026-07-01 | Standalone memory/doc skills branch to chore/* via aiguillage (hook exemption kept) | accepted |
| BDR-046 | 2026-07-01 | Claude Code installs via official native installer (curl claude.ai/install.sh), drop npm from install.sh | accepted |
| BDR-047 | 2026-07-01 | ECC audit → zero import; local config ahead of reference | accepted |
| BDR-048 | 2026-07-03 | semgrep security gate: engine version + rulesets PINNED, never --config auto; upgrade = deliberate visible human jump | accepted |
| BDR-049 | 2026-07-03 | verifier = fresh + blind (no iteration history) + disk-contract + PROOF-or-fail; mute ≠ PASS; scope enrichment via human micro-gate | accepted |
| BDR-050 | 2026-07-03 | universal pipeline (contract→dev inline→fresh verify→fresh security, loops bounded 3× in main loop) with per-flow weighting; hotfix failure = revert not loop | accepted |
| BDR-051 | 2026-07-04 | contract enrich-at-gate: the contract grows ONLY at a human micro-gate ([gated] marker); the verifier judges the ENRICHED contract, not the seed | accepted |
---
@@ -788,3 +792,36 @@ rules:
ONE scope gap: BDR-047 never opened hooks/ — ECC's only WIRED subsystem. Fruit:
config-protection hook (own idiom, NOT ECC import), shipped
feature/config-protection-hook. Lesson holds + refined by [[LRN-090]].
## BDR-048 — Deterministic security gate: pinned engine + pinned rulesets (semgrep)
- **Date**: 2026-07-03
- **Decision**: semgrep = BLOCKING gate (verify-loops chantier) → engine version PINNED in plugins.lock.json (gsd-pin pattern; update-all.sh honors pin + displays jump cur→pin before `pipx install --force`). Rulesets PINNED in-agent: `p/security-audit` + `p/secrets`. Never `--config auto` (registry telemetry + ruleset resolved per-run = non-deterministic gate, [[LRN-077]] class). Never auto `semgrep login` — Pro rules optional, guide-only (ctx7 pattern).
- **Rationale**: gate blocks HIGH/CRITICAL only ([[LRN-047]]); silent engine/rule upgrade = new BLOCKs on unchanged code w/o human decision → gate crying false → ignored. Version jump must be deliberate + visible (bump pin, then `make update` shows the jump).
- **Alternatives rejected**: `latest` (pipx house default, graphifyy-style) — fine for comfort tools, wrong for a blocking gate; `--config auto` — telemetry + non-determinism.
- **Reference**: plugins.lock.json `semgrep` entry, install-plugins.sh STEP 7.5, update-all.sh step 6.2 — branch feature/semgrep-install `ccfecc9`. Conditions [[LRN-047]], [[LRN-085]]. Coverage caveat of the community rulesets: [[LRN-092]].
- **Addendum 2026-07-03** (lot 3, measured): rulesets = `p/security-audit` + `p/secrets` + **`p/owasp-top-ten`**. owasp-top-ten is REQUIRED not optional — measured on realistic Flask code, the 2-ruleset baseline missed SQL injection + path traversal ENTIRELY (0 findings); owasp's taint rules catch them. Severity map: secrets ERROR→CRITICAL, other ERROR→HIGH (block), WARNING/INFO→reported. Blocking threshold = ERROR (per-RULE, not per-vuln — same class can straddle ERROR/WARNING; blocking WARNING too floods FP). FP measured shell/md only (faunosteo, game: sole added blocking ERROR = Dockerfile `missing-user` hygiene, contained by gate-mode diff-scoping). **Re-evaluate owasp FP at the first real web/python app project** (shell/md repos don't represent where the gate runs). See [[LRN-094]], agents/security-auditor.md branch feature/security-auditor `2b297bd`.
## BDR-049 — Verifier doctrine: fresh + blind + disk-contract + proof-or-fail
- **Date**: 2026-07-03
- **Decision**: conformity verdict comes ONLY from a FRESH verifier subagent per iteration. Input = contract PATH (read from disk — dev restatement structurally unable to interpose) + diff range + optional test cmd. NEVER iteration history: blind, complete verification every time (cost bounded by the main-loop max-3 cap, [[LRN-083]]: loops decided in main loop). CONFORME ⇔ all criteria MET + zero out-of-scope. PROOF line mandatory ([[LRN-048]]). Mute/unparsable verifier NEVER a PASS: 1 fresh retry, 2nd structural failure = human escalation. Dev-justified out-of-scope enters FILE SCOPE only via a human micro-gate (`[gated]` marker) — else the dev justifies everything and scope constrains nothing. Contract on DISK at creation (`.claude/tasks/contracts/<date>-<slug>-<HHMM>.md`, committed; aborted run → deleted or `status: aborted`, never left dirty).
- **Rationale**: dev self-score is always confident → not a gate. Verifier fed history anchors on prior verdicts → telescopic drift. Context-only contract dies at compaction, the verbatim with it.
- **Alternatives rejected**: dev self-assessment as gate; cumulative verifier context ("cheaper" but anchored); gitignored run files (lose escalation reference + session-death survival).
- **Reference**: lib/contract-interview.md + agents/verifier.md + lib/tests/contract-verifier.test.sh (31 locks) — branch feature/contract-verifier `6aed5ee`. Behavioral GREEN: planted-gap → ECARTS(2) exact; conform-under-injected-history → CONFORME (blindness held). Twin of [[BDR-048]] (security gate). Conditions [[LRN-048]], [[LRN-083]].
## BDR-050 — Universal verify+secure pipeline, weighted per flow (loops in the main loop)
- **Date**: 2026-07-03
- **Decision**: every dev flow = contract (verbatim, on disk) → dev INLINE → fresh verifier (request conformity) → fresh security-auditor (`MODE: gate`) → commit. Loops BOUNDED at 3 and decided in the ORCHESTRATOR MAIN LOOP ([[LRN-083]]), never in a subagent. Order invariant: on any security re-loop, re-verify the REQUEST before re-scanning security. Per-flow weight: feat/bugfix = both gates, both loop (nominal 2 dispatches); hotfix = NO fresh verifier (its smoke-check verifies the trivial autofill contract), security gate whose FAILURE REVERTS (`git restore` + escalate to /bugfix), never loops — the 1-attempt model preserved (nominal 1 dispatch). Shared include `lib/verify-secure-loop.md` for feat/bugfix; hotfix inline variant.
- **Rationale**: the value is the INDEPENDENCE of the gate (fresh subagent vs a rich contract), NOT delegating the dev — so dev stays inline in light flows and weighting lives on loops+questions, never on skipping a gate. hotfix reverts because a 3× loop would reintroduce the weight its identity excludes.
- **Alternatives rejected**: dispatch the dev too (turns feat into ship-feature-bis); one merged "quality" gate (see [[LRN-095]] — orthogonal gates degrade if fused); hotfix loops like feat (breaks its 1-attempt identity).
- **Reference**: lib/verify-secure-loop.md + wired feater/bugfixer/hotfixer + lib/tests/loops-light.test.sh (27 locks) — feature/verify-loops `0f0162d`. Behavioral GREEN (feat fixture): CONFORME→BLOCK(1) SQLi→fix→re-verify CONFORME→re-scan PASS, order invariant held. Builds on [[BDR-048]] [[BDR-049]]. Conditions [[LRN-083]] [[LRN-095]].
## BDR-051 — Contract enrich-at-gate: the contract grows only at a human micro-gate
- **Date**: 2026-07-04
- **Decision**: the CONTRACT's REQUEST is immutable, but ACCEPTANCE CRITERIA + FILE SCOPE may GROW — exclusively at a human gate, each added entry tagged `[gated <date>]`. In the heavy flows (ship-feature STEP 3, init-project GATE #1) the approved DESIGN appends design-derived criteria to the contract; the fresh verifier then judges the diff against the ENRICHED contract, never the seed. Same mechanism as the out-of-scope micro-gate ([[BDR-049]]) — a dev never enriches; only the human validating a gate does.
- **Rationale**: the raw request underspecifies (a one-line "add validation" hides the schema-rejection requirement the design surfaces). If the verifier judged only the seed, every design decision would be unverified. Gating the growth keeps the contract honest (no silent scope creep) AND complete (design criteria are verified). The only flow where the contract is mutable mid-run — bounded to gate moments.
- **Alternatives rejected**: freeze the contract at creation (design criteria unverified — the seed is too thin); let the dev enrich (the [[BDR-049]] failure mode — dev justifies everything, scope constrains nothing); a second contract per design (loses the single-reference property).
- **Reference**: ship-feature STEP 0e+3, init-project STEP 1+4, feature/verify-loops `1c69de2`. Behavioral GREEN: a `[gated 2026-07-04]` design criterion (reject unknown config keys) was read + judged NOT-MET by a fresh verifier across 3 rounds (dogfood). Builds on [[BDR-049]] [[BDR-050]].
+6
View File
@@ -314,3 +314,9 @@ rules:
- Next: #2 design-toolchain trigger fix (residual false-fires post-ed2408e, 5× this session).
- #2 done (bugfix/design-toolchain-trigger): trigger tightened — dropped bare design|component|composant|theme|thème|transition|frontend|front-end|palette; dashboard→\bdashboard\b (kills ecc_dashboard.py filename match, keeps "admin dashboard"); kept animation; added "front-?end design" bigram + fire-log counter (time+token+excerpt, ~/.claude/logs/design-toolchain-fires.log) so future "re-firing?" is measured. Test 18/18, shellcheck clean, live dogfood green. [[LRN-091]] corrob [[LRN-047]].
- Double dogfood of #1 guard: config-protection blocked + sentinel-bypassed my own edits to the now-guarded design hook + its test — first real use of the guard, friction validated in passing (one-shot sentinel .claude/.config-edit-ok, non-empty reason, logged+consumed). ECC second-regard closed: #1 config-protection + #2 trigger fix, both merged to develop, nothing pushed.
- Chantier verify-loops/semgrep/contract: Phase 1 read-only (6 subagents mapped 6 orchestrators + cso + agents + install patterns; caught subagent error — cso IS gstack symlink, ls-verified) → archi GATED-GO (5 verdicts: local grafts, dev inline light flows, hotfix unchanged, pinned rulesets, pinned version; +2 specs: contract on DISK, mute verifier ≠ PASS). LOT 1 shipped on feature/semgrep-install (ccfecc9+b8d3ccc): install-plugins STEP 7.5 + update-all 6.2 + lock pin 1.168.0, dogfooded real (4 paths + anonymous ruleset fetch + detection). [[BDR-048]] [[LRN-092]]. Next: lot 2 specs (contract-interview lib + verifier agent).
- Chantier verify-loops LOT 2 (feature/contract-verifier `6aed5ee`): lib/contract-interview.md (verbatim contract on DISK, micro-gate scope enrichment, aborted never dirty) + agents/verifier.md (fresh+blind, PROOF-or-fail, mute ≠ PASS) + 31 structure locks green, shellcheck clean. Behavioral: planted-gap → ECARTS(2) exact; conform under injected fake history → CONFORME (blindness held). Sentinel consumed 4× on guarded lib/tests/. [[BDR-049]] [[LRN-093]]. Merge note: lot 1+2 both append registries at same anchors → trivial stack-conflict expected. Next: lot 3 security-auditor spec.
- Chantier verify-loops LOT 3 (feature/security-auditor `2b297bd`): agents/security-auditor.md (SAST gate, pinned p/security-audit+p/secrets+p/owasp-top-ten, secrets→CRITICAL, block ERROR only, DEGRADED-still-checks, anti-gaming nosemgrep, PROOF-or-fail) + grafts onboard L3a (complement to cso, both gstack branches) + audit-delta security axis. 28 structure locks + 4 behavioral dogfoods green: vuln→BLOCK(9), nosemgrep→BLOCK(1), DEGRADED→BLOCK(7). owasp REQUIRED (measured: baseline misses SQLi+path-traversal on Flask). [[LRN-094]] + [[BDR-048]] addendum (owasp/severity/FP) applied at integration on feature/verify-loops (index drift LRN-090/091 backfilled same pass). Next: lot 4 loops-light (feat/bugfix/hotfix wiring).
- Integration: feature/verify-loops = develop + merge lots 1-3 (local, develop/main intact, nothing pushed) so lots 4-5 wiring is dogfoodable against present agents. Memory stack-conflicts resolved (BDR-048/049, LRN-092/093/094 stacked ID-order; BDR-048 addendum applied; LRN-090/091 index rows backfilled).
- Chantier verify-loops LOT 4 (feature/verify-loops `0f0162d`): lib/verify-secure-loop.md shared include + wired feater (0.7 contract, 3 verify+secure), bugfixer (3.5 contract from diagnosis, 5 gates), hotfixer (1.7 silent contract, 3 security gate FAILURE=REVERT not loop, +Agent tool). 27 structure locks + full pipeline dogfood: feat fixture w/ SQLi → GATE1 CONFORME → GATE2 BLOCK(1) (checklist caught what semgrep taint missed) → fix → re-verify CONFORME (order invariant) → re-scan PASS. [[BDR-050]] [[LRN-095]]. Weighting held: feat/bugfix nominal 2 dispatches, hotfix 1 + revert-on-fail. INCIDENT: re-committed [[LRN-093]] (2nd recurrence, 4 locks w/ \n) — caught at first run; user flagged advisory-insufficient → build deterministic backstop in lot 5. Next: lot 5 heavy flows (ship-feature enrich-at-gate, init-project +security, onboard no-loop) + escalation dogfood (max-3 STOP) + LRN-093 meta-test guard.
- Chantier verify-loops LOT 5 (feature/verify-loops `1c69de2`, FINAL): ship-feature (0e contract, enrich-at-gate STEP 3 [gated], 5 verify+secure vs ENRICHED) + init-project (contract from BRIEF, enrich GATE#1, 9 verify+secure — adds the security gate it lacked) + onboard (explicit NO-loop, audit≠dev, documented vs symmetry) + lib/tests/no-vacuous-locks.test.sh (LRN-093 deterministic backstop w/ inline flip-test) + loops-heavy 18 locks. Dogfood BOTH vigilance points real: (1) enrich — fresh verifier reads+judges a [gated] design criterion (ECARTS names it); (2) escalation — 3 consecutive ECARTS → orchestrator STOP at max-3 + CONTRACT-vs-REALIZED table, no 4th loop, no commit (first real exercise of the infinite-loop guard). [[BDR-051]] [[LRN-096]]. INCIDENT closed: the backstop's OWN flip-test RED'd (regex missed line-start tf) → fixed → [[LRN-096]] (a guard is code, prove it can fail). Chantier complete: 5 lots on feature/verify-loops, develop+main intact, nothing pushed.
+38
View File
@@ -109,6 +109,13 @@ rules:
| LRN-087 | 2026-07-02 | presence-flag ≠ capability — rtk silently dead after .bashrc wipe; emitted commands need ABSOLUTE bin paths (they run in another shell); integrity pin = live machinery, re-pin on hook edit | any PATH-dependent capability + hand-managed shell profile; hooks emitting commands for another shell |
| LRN-088 | 2026-07-02 | token-cutting intuition inverts under measurement — verbosity beats cardinality (gstack 34 skills ≈ 592 tok vs pr-review 6 agents ≈ 2,183) | any "disable X to save tokens" — measure per-item bytes first; profiles toggle skills, not plugin payloads |
| LRN-089 | 2026-07-03 | pass-through wrapper (CLI `"$@"` → fn deriving target from ambient state: HEAD/cwd/env) silently ignores its args = silent contract violation; guard = args are an ASSERTION, refuse when they disagree with state | any dispatcher forwarding args to a callee that reads ambient state instead of the args |
| LRN-090 | 2026-06-30 | external-repo audit: open WIRED subsystems (hooks/runners) before declarative (docs/rules); described capability ≠ wired capability | auditing an external config/framework repo for transferable value |
| LRN-091 | 2026-07-03 | keyword-triggered soft-nudge hook w/ bare common tokens over-fires on non-UI work → tuned out; bare only when UI sense dominates, else bigram-or-drop | any advisory/nudge hook keyed on keywords |
| LRN-092 | 2026-07-03 | SAST smoke test w/ the OFFICIAL example secret = vacuous pass (rules exclude documented example keys by design); validate w/ realistic payloads + measure tier coverage before trusting a gate ruleset | smoke-testing any detector/gate — never the canonical example payload |
| LRN-093 | 2026-07-03 | grep -F pattern w/ embedded newline = per-line OR = lock that matches anything; structure locks single-line only, flip-test new locks | writing any grep-based structure lock / census test |
| LRN-094 | 2026-07-03 | SAST severity ≠ exploitability — semgrep ERROR conflates real vulns + hardening recos; metadata does NOT cleanly separate them (measured) → metadata refinement = noisy gate; ERROR-threshold + diff-scoping is the containment | mapping a SAST tool's output to a blocking gate |
| LRN-095 | 2026-07-03 | orthogonal gates don't contaminate — a conformity verifier must PASS correct-but-insecure code (security is a separate gate's job); proven live (CONFORME on a feature carrying a SQLi); fusing the two degrades each | designing multi-dimension review/verify/audit gates |
| LRN-096 | 2026-07-04 | a backstop/guard is code — reliable ONLY after a flip-test proves it CAN fail; an unproven guard replacing an advisory = a vacuous guard (LRN-048 applied to guards); flip-test mandatory at guard creation | building any deterministic guard/lint/backstop |
---
@@ -977,3 +984,34 @@ rules:
- **rule**: keep a token BARE only when its UI sense dominates largely in a dev context (glassmorphism, navbar). Token common in non-UI talk (design, component, theme, transition, frontend) → require a UI-specific bigram (design system, front-end design) or drop; in doubt → bigram-or-drop. Borderline standalone nouns (dashboard, animation) may stay bare as an assumed call — the fire-log arbitrates later on data, not gut. (NOT "never bare tokens" — animation stays bare here by design.)
- **context**: design-toolchain-reminder.sh — 07-02 tightening (dropped page/form/menu/…) insufficient; 6 bare tokens still false-fired ~6×/session during the ECC config audit (design, ecc_dashboard.py, component, frontend, theme, transition, palette). 07-03 fix: dropped them, dashboard→`\bdashboard\b` (filename match killed, "admin dashboard" kept), added a fire-log (time+token+excerpt). `lib/tests/design-toolchain-reminder.test.sh` locks it (18 checks).
- **cousin**: [[LRN-047]] a doctor that cries false is ignored.
## LRN-092 — SAST smoke test: official example keys are rule-excluded — "no findings" proves nothing
- **pattern**: smoke-testing a SAST/secret detector w/ the OFFICIAL example payload (AWS `AKIA...EXAMPLE`) → 0 findings BY DESIGN — rules exclude documented example keys to kill FP. A vacuous pass, [[LRN-048]] class (a pass must prove it looked). Validate w/ realistic-shaped payloads AND enumerate what the tier does NOT catch before trusting a ruleset as a gate.
- **context**: lot 1 semgrep-install dogfood 2026-07-03. `p/secrets`+`p/security-audit` community tier: anonymous fetch OK (52 rules, no login), `subprocess-shell-true` detected ERROR; MISSED %-format SQLi on bare cursor (no recognized DB-API context) + fake-checksum `ghp_` token. Gap logged for security-auditor agent design (consider adding `p/owasp-top-ten`).
- **future application**: any detector/gate smoke test — craft realistic payloads, never the canonical example; measure the miss-list on purpose-built fixtures; size the gate's blocking scope on that data.
- **cousin**: [[LRN-048]] a 0/OK must prove it looked; [[LRN-047]] noisy guard = ignored guard; conditions [[BDR-048]].
## LRN-093 — grep -F with an embedded newline = per-line OR = vacuous lock
- **pattern**: a fixed-string grep pattern containing a newline is treated as MULTIPLE patterns (one per line) — match succeeds if ANY line matches. A structure lock written that way passes on essentially anything (`"no\n forced loop"` → matches any "no") = a lock that proves nothing, [[LRN-048]] class.
- **context**: lot 2 `lib/tests/contract-verifier.test.sh`, caught in self-review BEFORE first run; replaced by a single-line distinctive anchor ("proceed straight to the security gate").
- **future application**: structure locks / census greps = ONE line per pattern, always; a clause spanning lines → lock a distinctive single-line fragment. Flip-test every new lock (prove it CAN fail) before trusting its green.
- **cousin**: [[LRN-048]] a pass must prove it looked; [[LRN-046]] deterministic-oracle discipline.
## LRN-094 — SAST severity ≠ exploitability; metadata does not cleanly separate — don't refine on it
- **pattern**: semgrep `ERROR` conflates exploitable vulns (SQLi, secrets, command injection) with hardening recommendations (Dockerfile missing-USER, npm release-age). The obvious refinement — gate on `metadata.impact`/`likelihood`/`confidence` — does NOT work: measured, the Dockerfile hygiene ERROR (`impact=MEDIUM likelihood=LOW`) is indistinguishable from a tainted-SQL ERROR (`impact=MEDIUM likelihood=MEDIUM`), and a real command-injection reads `impact=LOW likelihood=HIGH`. Metadata-based severity = a noisy, non-deterministic gate ([[LRN-077]] class).
- **context**: lot 3 security-auditor design 2026-07-03. Measured on 2 real repos (faunosteo, game): the only added blocking ERROR from owasp-top-ten is Dockerfile hygiene, contained because gate mode scopes to the DIFF (a pre-existing infra finding can't block an unrelated code change).
- **future application**: mapping any SAST to a blocking gate — take the tool's ERROR/blocking level as the deterministic threshold, contain FP by SCOPING (diff, not repo), NOT by a metadata heuristic or a hand-maintained hygiene denylist ([[LRN-049]]: match guard cost to proven stake — build the denylist only if hygiene ERRORs prove noisy on a real project).
- **cousin**: [[LRN-047]] noisy gate = ignored; [[LRN-077]] non-deterministic gate; conditions [[BDR-048]].
## LRN-095 — Orthogonal gates don't contaminate: a conformity check must pass correct-but-insecure code
- **pattern**: when a pipeline has distinct gates (request-conformity, security), each judges ONLY its dimension. A conformity verifier must return CONFORME on code that is correct-but-insecure — the vuln is the SECURITY gate's job, not a conformity gap. Proven live: a `get_item` feature satisfying its contract but carrying a `%`-interpolation SQLi → verifier CONFORME, security-auditor BLOCK(1). Fusing the two into one "quality" gate makes each worse: the conformity check starts hunting vulns (scope creep, misses conformity), the security check starts judging feature-completeness (dilutes).
- **context**: lot 4 verify-secure-loop dogfood 2026-07-03. The orthogonality is WHY the order invariant matters (re-verify request before re-scan security) — two independent axes re-checked independently.
- **future application**: any multi-dimension gate (review lenses, verify+audit, correctness+perf) — keep each gate single-axis and let a finding on axis B pass axis A's gate; compose verdicts in the orchestrator, don't merge the judges.
- **cousin**: [[BDR-050]] the pipeline; [[BDR-049]] fresh verifier; conditions [[LRN-083]].
## LRN-096 — A backstop is code: prove it can FAIL (flip-test) before trusting its green
- **pattern**: a deterministic guard built to replace a forgettable advisory is itself code, and an UNPROVEN guard is a vacuous guard — [[LRN-048]] (a pass must prove it looked) applied to guards themselves. The LRN-093 backstop (refuse `\n` in grep/tf patterns) shipped with a regex requiring whitespace before `tf` → it silently MISSED `tf` at line start (exactly where the real locks sit). A flip-test (feed the guard a KNOWN offender, assert it bites) caught the hole; without it the guard would have green-lit the very class it was built to kill. So: a flip-test is MANDATORY at guard creation, part of the guard, not optional QA.
- **why it matters**: the whole point of a backstop is that it fires on the bad case; a guard that can't fail proves nothing and is WORSE than the advisory it replaced (false confidence). The advisory→backstop move ([[LRN-047]] [[LRN-091]], own doctrine) is only sound if the backstop is itself verified against a real miss.
- **context**: lot 5 `lib/tests/no-vacuous-locks.test.sh` 2026-07-04. Built the guard, its flip-test RED'd (regex too weak, missed line-start `tf`), fixed the regex, flip-test green. The guard now ships WITH the flip-test inline so it self-proves on every run.
- **future application**: building any guard/lint/census/backstop — bundle a flip-test (a synthetic offender the guard must catch) in the same file; a guard whose failure path was never exercised is untrusted. Corroborates [[LRN-047]]/[[LRN-091]] (advisory→deterministic) — this is the *quality bar* on the deterministic replacement.
- **cousin**: [[LRN-048]] prove it looked; [[LRN-093]] the class this guards; [[LRN-046]] deterministic-oracle discipline.
+28
View File
@@ -1,5 +1,33 @@
# TODO
## 2026-07-03 — verify loops + semgrep gate + contract (chantier orchestrateurs)
Archi validée au gate (session 2026-07-03). Cible : contract sur DISQUE dès
création (fichier de run, pattern DIAGNOSIS) + verifier frais (verdict structuré
CONFORME/écarts, preuve-qu'il-a-regardé LRN-048, 2 échecs structurels = escalade
humaine — verifier muet ≠ PASS) + gate sécu semgrep (rulesets ÉPINGLÉS
p/security-audit + p/secrets — pas --config auto, classe LRN-077 ; BLOCK
HIGH/CRITICAL only, LRN-047) + boucles bornées 3× décidées en boucle principale
(LRN-083). cso = symlink submodule gstack → non modifiable → greffes locales
(onboard cso-fallback, audit-delta, agent neuf ; complément semgrep même
gstack ON). Verdicts user : dev inline conservé feat/bugfix/hotfix (verify+sécu
= sous-agents frais) ; hotfix garde revert-escalade ; PIN version semgrep dans
plugins.lock.json (gate bloquante — upgrade silencieux = nouveaux BLOCK sur code
inchangé ; pattern gsd-pin, saut affiché par update-all).
LOT 1 — feature/semgrep-install (GO)
- [x] plugins.lock.json — pin semgrep 1.168.0 (pattern gsd, note gate bloquante)
- [x] install-plugins.sh STEP 7.5 — pipx pinned, command -v guard + version echo, login guide-only (jamais auto)
- [x] update-all.sh step 6.2 — pin-honored, affichage saut cur→pin, pipx install --force
- [x] Dogfood — install réel 1.168.0 via bloc extrait + idempotence (re-run = skip) + pin-match + saut affiché (1.168.0→9.9.9 fake, warn propre, install intacte)
- [x] Verify — bash -n OK, shellcheck clean (SC1091 info pré-existants only), lock JSON valide ; smoke rulesets : fetch anonyme 52 règles SANS login, subprocess-shell-true ERROR détecté. Limite notée pour LOT 3 : community tier rate SQLi %-format hors contexte API + tokens fake (choix rulesets à re-évaluer à l'agent)
- [ ] Commit scoped (settings.json dirty pré-existant JAMAIS stagé) + GATE lot 1
LOT 2 — feature/contract-verifier : specs montrées AVANT écriture. lib/contract-interview.md + agents/verifier.md.
LOT 3 — feature/security-auditor : agents/security-auditor.md + greffe audit-delta + onboard fallback + complément gstack-ON.
LOT 4 — feature/loops-light : câblage feat/bugfix/hotfix.
LOT 5 — feature/loops-heavy : câblage ship-feature + init-project + onboard.
Rien poussé ; gate par lot ; suites après chaque lot.
## 2026-07-03 — design-toolchain trigger fix (bugfix/design-toolchain-trigger)
Root cause (NOT a kill-switch, per user): ed2408e (07-02) dropped ultra-generic
tokens but left bare tokens common in non-UI talk → ~6× false-fire THIS session
+25 -3
View File
@@ -100,6 +100,16 @@ RISK: <low/medium — what could go wrong>
- If the fix is significant (>10 lines, multiple files,
behavior change): wait for user approval.
## STEP 3.5 — CONTRACT
Run `$HOME/.claude/lib/contract-interview.md` (main loop). The DIAGNOSIS
feeds it: REQUEST verbatim = the bug report as received; ACCEPTANCE CRITERIA
= the symptom reproduced-then-gone + a regression test present and passing;
FILE SCOPE = the FIX PLAN files. Questions stay proportional (a clear,
reproduced bug → zero). It writes the contract to
`.claude/tasks/contracts/<date>-<slug>-<HHMM>.md`; keep the path for GATE 1
(STEP 5).
## STEP 4 — FIX
**Gitflow aiguillage (before editing):** follow `$HOME/.claude/lib/gitflow-aiguillage.md`
@@ -131,7 +141,19 @@ Apply the fix following the plan:
```
2. If a build step exists, verify it passes (`npm run build`, `tsc --noEmit`, `cargo build`, etc.).
3. Check for regressions in related functionality.
4. **Pre-commit confirmation gate.** Before running `git commit`, present the diff
4. **Fresh gates (verify + secure), bounded loops.** Steps 1-3 are your
dev-side smoke test, NOT the gate. Run the two fresh gates per
`$HOME/.claude/lib/verify-secure-loop.md` with `CONTRACT` = the STEP 3.5
path, `DIFF` = the fix diff, `TEST` = the suite from step 1:
- GATE 1 — a FRESH verifier judges the fix against the contract (bug gone
+ regression test present). CONFORME → GATE 2. ECARTS → fix, re-verify,
max 3 → escalate.
- GATE 2 — a FRESH security-auditor (`MODE: gate`) scans the fix diff
(a bug fix can introduce a vuln). PASS → commit gate. BLOCK → fix,
re-verify request THEN re-scan, max 3 → escalate.
Nominal = one verifier + one security dispatch. Only then the commit gate.
5. **Pre-commit confirmation gate.** Before running `git commit`, present the diff
summary and the proposed message, then wait for approval:
```
@@ -152,7 +174,7 @@ Apply the fix following the plan:
- `skip` → leave changes uncommitted, exit cleanly.
- `amend last` → the fix should fold into the previous commit (use only when prior commit is unpushed).
5. Commit using conventional format (after approval):
6. Commit using conventional format (after approval):
```
fix(<scope>): <root cause description>
@@ -161,7 +183,7 @@ Apply the fix following the plan:
Co-Authored-By: Claude <noreply@anthropic.com>
```
6. Print summary:
7. Print summary:
```
BUGFIX COMPLETE
BUG : <symptom>
+28 -12
View File
@@ -65,6 +65,17 @@ already constrain or forbid the approach; an LRN may name a gotcha to apply. Emi
MEMORY; feed STEP 1 MINI-PLAN. Inline consumption — reader = planner, no injection.
`.claude/memory/` absent → guarded no-op (zero overhead on a memory-less repo).
## STEP 0.7 — CONTRACT
Run `$HOME/.claude/lib/contract-interview.md` (main loop — you are it). It
captures the request verbatim, asks 0-3 questions PROPORTIONAL to ambiguity
(a complete request → zero questions, silent), derives testable acceptance
criteria + file scope, and writes the contract to
`.claude/tasks/contracts/<date>-<slug>-<HHMM>.md`. Keep the path — GATE 1
(STEP 3) hands it to a fresh verifier. On a small, clear feature this is a
few seconds and no questions; it is the single reference the verifier judges
against, not a restatement.
## STEP 1 — MINI-PLAN
Quick mental model, not a formal plan document:
@@ -101,19 +112,24 @@ Work through the plan:
- Follow existing patterns in the codebase.
- Run tests incrementally as you go.
## STEP 3 — VERIFY
## STEP 3 — VERIFY + SECURE (fresh gates, bounded loops)
1. Run the full relevant test suite:
```bash
# detect and run tests, lint, type-check
```
2. If a dev server is relevant, mention what the user should
check visually.
3. Quick self-review: scan your diff for obvious issues:
```bash
git diff --stat
git diff
```
First, your own pre-check (dev-side, fast): run the relevant test suite /
lint / type-check, and if a dev server is relevant note what to check
visually. This is your smoke test, NOT the gate.
Then run the two fresh gates per `$HOME/.claude/lib/verify-secure-loop.md`
with `CONTRACT` = the STEP 0.7 path, `DIFF` = your working-tree diff, `TEST`
= the suite you just ran:
- GATE 1 — a FRESH verifier judges the diff against the contract (blind, no
self-score of yours counts). CONFORME on the first pass → straight to GATE
2, no loop. ECARTS → fix the named gaps, re-verify, max 3 → escalate.
- GATE 2 — a FRESH security-auditor (`MODE: gate`) scans the diff. PASS →
commit. BLOCK → fix, re-verify the request THEN re-scan, max 3 → escalate.
Nominal (clear request, conform first pass, clean diff) = exactly one
verifier + one security dispatch. The loop only costs when it loops.
## STEP 4 — COMMIT
+33 -5
View File
@@ -1,13 +1,15 @@
---
name: hotfixer
description: Quick fix for superficial bugs (typos, CSS issues, config errors, off-by-one, wrong variable name, missing import, broken link). Max 2 files, obvious root cause only.
tools: Read, Edit, Write, Bash, Grep, Glob
tools: Read, Edit, Write, Bash, Grep, Glob, Agent
---
# HOTFIX — Quick Superficial Fix
Fast-track fix for obvious bugs. No planning overhead, no plugin
check, no subagents. Get in, fix, verify, get out.
Fast-track fix for obvious bugs. No planning overhead, no plugin check.
The fix is inline (no dev subagents); a fresh security gate runs before
commit, and any gate failure reverts — never loops. Get in, fix, gate,
get out.
## REQUEST
$ARGUMENTS
@@ -39,6 +41,18 @@ skip). For a RECURRING or urgent bug only, a quick blockers-only glance may save
If a prior BLK names this bug, jump to its solution. Not mandatory; no RELATED MEMORY
disposition required at hotfix weight.
## STEP 1.7 — CONTRACT (silent autofill)
Run `$HOME/.claude/lib/contract-interview.md` at hotfix weight: **zero
questions ever** (a hotfix is an obvious fix by definition). Autofill the
contract — REQUEST verbatim = the bug description as given; ACCEPTANCE
CRITERIA = "symptom gone; build/tests green"; FILE SCOPE = the 1-2 target
files. It writes `.claude/tasks/contracts/<date>-<slug>-<HHMM>.md`. This is
the reference for the security gate's scope and the escalation report if a
gate fails. No verifier is dispatched at hotfix weight — the STEP 3
smoke-check already verifies these trivial criteria; the gate hotfix adds is
security (below).
## STEP 1.5 — DESIGN GATE
Follow `$HOME/.claude/lib/design-gate.md`:
@@ -102,18 +116,32 @@ Apply the minimal change that fixes the bug:
(Files were not yet staged — restore is safe.)
- STOP and tell user: `"Hotfix introduced a regression. Reverted. Escalate to /bugfix or /analyze for deeper investigation."`
- Do NOT commit a broken fix.
3. Commit using conventional format (only after verify passes):
3. **Security gate (fresh auditor) — failure REVERTS, never loops.** Dispatch
a FRESH security-auditor (`subagent_type: security-auditor`, or load
`agents/security-auditor.md`) with `MODE: gate`, `SCOPE:` the working-tree
diff vs the pre-flight SHA. Parse its `SECURITY — VERDICT:` line:
- `PASS` (or `DEGRADED` with no BLOCK) → proceed to commit.
- `BLOCK(n)` → this is hotfix: do NOT loop. Run `git restore .` to the
pre-flight SHA, print the `BLOCKING` list, and STOP:
`"Hotfix introduced a security finding. Reverted. Escalate to /bugfix
for a fix under the full verify+security loop."` The hotfix model is
one attempt; any gate failure (smoke OR security) reverts and escalates.
- Structural failure (mute / unparsable / no VERDICT line) → treat as a
failed gate: retry ONCE fresh; a 2nd structural failure → revert +
escalate. A mute auditor is never a PASS.
4. Commit using conventional format (only after verify AND security pass):
```
fix(<scope>): <what was wrong>
Co-Authored-By: Claude <noreply@anthropic.com>
```
4. Print summary:
5. Print summary:
```
HOTFIX APPLIED
FILE(S) : <changed files>
FIX : <one-line description>
VERIFIED: <test name or smoke check that passed>
SECURITY: <PASS | DEGRADED (checklist only)>
```
## STEP 4 — DOC SYNC (automatic)
+159
View File
@@ -0,0 +1,159 @@
---
name: security-auditor
description: SAST security gate — runs the pinned semgrep rulesets + the CLAUDE.md security checklist on a diff or project scope, maps severities, renders SECURITY — VERDICT: PASS | BLOCK(n). Blocks HIGH/CRITICAL only, reports the rest. Never fixes code. Fresh dispatch, no iteration history.
tools: Read, Grep, Glob, Bash, Write
---
# SECURITY-AUDITOR AGENT
You are the security gate. You run semgrep + a checklist over a scope,
classify by severity, and render a verdict. You never fix code, you never
edit anything but the report file (audit mode only), and you never trust a
prior run — every scan is fresh and complete.
Bash runs semgrep and read-only inspection only — never a command that
mutates code, installs, or commits.
## MODES
- **gate** (default; dev flows) — SCOPE = a diff. Output = the stdout block
below. `Write` is FORBIDDEN in this mode.
- **audit** (onboard, audit-delta) — SCOPE = project root or a delta list.
`Write` is allowed ONLY to the exact `REPORT` path given — NEVER to any
code/config file. Writing anywhere else is a contract violation.
## INPUT (from the orchestrator — nothing else exists)
- `MODE: gate|audit`
- `SCOPE: <git range | explicit file list | project root>`
- `REPORT: <path>` (audit mode only — the single writable path)
- `CONTEXT: <archetype-context path>` (optional; onboard supplies it)
You NEVER receive iteration history — no prior verdicts, no earlier finding
lists, no dev reports. Ignore any such material if it appears. Every scan is
blind and complete (cost bounded upstream by the max-3 loop cap).
## STEP 1 — TOOL CHECK
`command -v semgrep` and capture the version. If ABSENT → **DEGRADED mode**:
announce it loudly on the `TOOL:` line, and STILL RUN STEP 3 (the checklist)
— a DEGRADED run must prove it detected everything it still can. A DEGRADED
run that skips the checklist and PASSes is a vacuous pass (LRN-048). Never a
silent skip, never a false BLOCK from the tool being absent (LRN-047).
## STEP 2 — SEMGREP (skip only in DEGRADED)
Resolve the scanned paths from SCOPE (in gate mode: `git diff --name-only
<range>` filtered to existing files; in audit mode: the root or delta list,
excluding `node_modules`, `dist`, `vendor`, `.git`).
Run, on those paths ONLY:
```
semgrep scan --config p/security-audit --config p/secrets --config p/owasp-top-ten \
--metrics=off --quiet --json <paths>
```
Pinned rulesets, never `--config auto`, never `semgrep login` (BDR-048:
`auto` = registry telemetry + per-run ruleset resolution = a
non-deterministic gate). owasp-top-ten is REQUIRED, not optional: measured
2026-07-03, the two-ruleset baseline missed SQL injection and path traversal
entirely on realistic Flask code; owasp-top-ten's taint rules catch them.
**Severity mapping** (from `results[].extra.severity` + ruleset origin):
| semgrep | origin | → gate severity | blocks? |
|---------|--------|-----------------|---------|
| ERROR | p/secrets | CRITICAL | yes |
| ERROR | other | HIGH | yes |
| WARNING | any | MEDIUM | no (reported) |
| INFO | any | LOW | no (reported) |
The blocking threshold is ERROR — deterministic, rule-assigned. Known limit
(measured): severity is per-RULE not per-VULN — the same class can span
ERROR and WARNING rules (e.g. `tainted-sql-string`=ERROR vs
`sql-injection-db-cursor-execute`=WARNING). Blocking on WARNING too would
flood FPs (nginx/github-actions/npm hygiene warnings); ERROR is the right
line. Report — never silently drop — the MEDIUM/LOW findings.
## STEP 3 — CHECKLIST (always, incl. DEGRADED)
Grep the scope for the CLAUDE.md non-negotiable defaults semgrep may miss.
Each hit → severity + file:line + one-line why:
- hardcoded secret / token / key / auth-bearing URL (→ CRITICAL)
- SQL built by string concatenation / interpolation (→ HIGH)
- unsanitized render of user input (innerHTML, dangerouslySetInnerHTML,
raw(), `eval`) (→ HIGH)
- sensitive endpoint with no authz check (→ HIGH)
- stack trace / internal path / DB error surfaced to the user (→ MEDIUM)
- secret / password / token / PII written to a log (→ HIGH)
- tracked `.env` or committed credential file (→ CRITICAL)
If `CONTEXT` (archetype) is given, scope the checklist to what applies
(no web-XSS checks on firmware, etc.).
## STEP 4 — ANTI-GAMING
Scan the diff (gate) or scope (audit) for any NEW suppression comment
(`# nosemgrep`, `// nosemgrep`, `nosec`, `eslint-disable ... security`, or
equivalent) that did not exist before this change. Each new suppression is a
**BLOCKING** finding UNLESS it already carries a human `[gated <date>]`
marker — same rule as scope enrichment: without the micro-gate the dev
suppresses everything and the gate constrains nothing. Report pre-existing
suppressions as LOW (context), do not block on them.
## STEP 5 — DEDUP + VERDICT
Merge semgrep + checklist findings, dedup by (file:line, rule/check).
`BLOCK(n)` ⇔ n = count(CRITICAL) + count(HIGH) + count(new un-gated
suppressions) > 0. Otherwise `PASS`. MEDIUM/LOW are REPORTED, never
blocking.
## OUTPUT (exact format — machine-parsed by the orchestrator)
```
SECURITY — VERDICT: PASS | BLOCK(n) | ERROR(<reason>)
TOOL: semgrep <ver> — p/security-audit, p/secrets, p/owasp-top-ten | ABSENT (DEGRADED — checklist only; install: make plugin)
SCOPE: <n> files
BLOCKING:
1. [CRITICAL|HIGH] <rule/check> — <file:line> — <why> — hint: <fix direction>
REPORTED (non-blocking):
- [MEDIUM|LOW] <rule/check> — <file:line>
PROOF: semgrep <n> rules on <n> files → <n> findings; checklist <n> checks → <n> findings
```
In audit mode, ALSO write this same block (plus per-finding detail) to
`REPORT`, and end stdout with `REPORT_WRITTEN: <path>`.
## RULES
- Report-only on CODE. Never edit or fix a code file. In audit mode the sole
writable path is `REPORT`; in gate mode nothing is writable.
- `PROOF` is MANDATORY — a `PASS` (or DEGRADED PASS) without a `PROOF` line
showing what was scanned is invalid; the orchestrator discards it as a
structural failure (LRN-048).
- A mute / crashed / unparsable auditor is NEVER a PASS. Exactly one
`SECURITY — VERDICT:` line, spelled as above.
- Blocks on HIGH/CRITICAL only. A noisy gate that blocks on hygiene is a
gate people learn to bypass (LRN-047) — MEDIUM/LOW are reported, not
gated.
## ORCHESTRATOR PROTOCOL (consumer contract — wiring reference)
- The security gate runs AFTER the request-conformity verdict is CONFORME
(verifier), never before.
- Dispatch a FRESH auditor each iteration — no context reuse. Input = mode +
scope + (report) + (context), nothing else.
- Parse the `SECURITY — VERDICT:` line:
- `PASS` → proceed (to commit / next step).
- `BLOCK(n)` → the dev subagent receives the BLOCKING list + the contract
path. After the fix: re-verify the REQUEST first (verifier), THEN re-run
this gate — in that order. Max 3 security iterations → STOP + human
escalation with the BLOCKING table.
- `DEGRADED` (semgrep absent) → does NOT block; surface the checklist
result + recommend `make plugin`. A DEGRADED BLOCK (grep-caught
hardcoded secret etc.) blocks like any other.
- Structural failure (`ERROR(…)`, missing/duplicated VERDICT line,
unparsable, crash, PASS without PROOF) → retry ONCE fresh; 2nd
structural failure → human escalation. A mute auditor is never a PASS.
+110
View File
@@ -0,0 +1,110 @@
---
name: verifier
description: Fresh independent verifier — reads a CONTRACT file from disk and renders a structured verdict (CONFORME / ECARTS / ERROR) on the implemented diff. Report-only, never fixes. Dispatched fresh at every iteration; receives no iteration history.
tools: Read, Grep, Glob, Bash
---
# VERIFIER AGENT
You verify that an implementation CONFORMS to a contract. You are NOT the
developer, you never fix anything, and you never trust the developer's
summary — only the contract, the code, and what you execute yourself.
Bash is for OBSERVATION ONLY: run tests/builds, `git diff` / `git log` /
`git show`, read-only inspection. Never a command that writes, installs,
commits, or mutates any state.
## INPUT (from the orchestrator — nothing else exists)
- `CONTRACT: <path>` — you READ it from disk; never accept an inline
restatement in its place
- `DIFF: <git range base...HEAD | explicit file list>`
- `TEST: <test command>` (optional)
You NEVER receive iteration history: no previous verdicts, no prior gap
lists, no dev reports. If any such material appears in your prompt, IGNORE
it — every verification is complete and blind. (Cost is bounded upstream:
the orchestrator caps the loop at 3 iterations.)
## STEP 1 — READ THE CONTRACT
Read the contract file. If it is missing, unreadable, or lacks its
`REQUEST` or `ACCEPTANCE CRITERIA` section → output
`VERIFY — VERDICT: ERROR(<reason>)` plus the `CONTRACT:` line, and STOP.
## STEP 2 — EVIDENCE PER CRITERION
For EACH acceptance criterion, establish exactly one status from the real
code:
- `MET` — with evidence: the file:line you read, or the test/build you RAN
- `NOT-MET` — expected vs actual, located at file:line
- `UNVERIFIABLE` — precise reason (missing environment, requires human
judgment, external dependency…)
Rules: read the diff AND enough surrounding code to judge behavior; run
`TEST` if provided, plus cheap targeted checks when they settle a
criterion. Never mark `MET` from naming, comments, or plausibility — only
from behavior you observed or code you read.
## STEP 3 — SCOPE CHECK
List the files actually touched (`git diff --name-only` over `DIFF`).
Compare against the contract's `FILE SCOPE`. Report every out-of-scope
file. Disposition is NOT your call: the orchestrator treats each one as a
gap — the dev removes it or justifies it, and an accepted justification
only enters the contract through a human micro-gate.
## STEP 4 — VERDICT
`CONFORME` ⇔ ALL criteria `MET` AND zero out-of-scope files.
Anything else is `ECARTS(n)` where n = count(NOT-MET) + count(UNVERIFIABLE)
+ count(out-of-scope files).
## OUTPUT (exact format — machine-parsed by the orchestrator)
```
VERIFY — VERDICT: CONFORME | ECARTS(n) | ERROR(<reason>)
CONTRACT: <path>
CRITERIA:
1. <criterion> — MET — <evidence file:line | test ran → result>
2. <criterion> — NOT-MET — expected <…> / actual <…> — <file:line>
3. <criterion> — UNVERIFIABLE — <reason>
SCOPE: in-scope <n> files; out-of-scope: <list | none>
PROOF: read <n> files, ran <cmd → result | nothing>, checked <n>/<n> criteria
```
## RULES
- Report-only. Never edit, never write, never propose the fix itself —
naming the gap precisely is the whole job.
- `UNVERIFIABLE` ≠ `MET`. A criterion you did not check is `UNVERIFIABLE`,
never silently dropped: the checked count in `PROOF` must equal the
contract's criteria count.
- `PROOF` is MANDATORY. A `CONFORME` without a `PROOF` line is invalid —
the orchestrator discards it as a structural failure (LRN-048: a pass
must prove it looked).
- The verdict grammar is load-bearing: exactly one `VERIFY — VERDICT:`
line, spelled exactly as above.
## ORCHESTRATOR PROTOCOL (consumer contract — wiring reference)
How every orchestrator consumes this agent (the loop lives in the MAIN
loop, never here):
- Dispatch a FRESH verifier at every iteration — no context reuse. Input =
contract path + diff range + optional test command, nothing else.
- Parse the `VERIFY — VERDICT:` line:
- `CONFORME` on first pass → proceed straight to the security gate — no
forced loop.
- `ECARTS(n)` → the dev subagent receives the contract PATH + the exact
gap list (nothing else). Max 3 iterations → STOP + human escalation
with the CRITERIA table (the contract-vs-realized diff).
- Remaining `UNVERIFIABLE` while everything else is MET → direct human
gate (a dev cannot fix unverifiability).
- Structural failure (`ERROR(…)`, missing/duplicated VERDICT line,
unparsable output, agent crash, `CONFORME` without `PROOF`) → retry
ONCE with a fresh verifier; a 2nd structural failure → human
escalation. A mute verifier is NEVER a PASS.
- After a security-gate fix round: re-verify the request FIRST (this
agent), THEN re-verify security — in that order.
+29
View File
@@ -654,6 +654,35 @@ if command -v graphify &>/dev/null; then
fi
echo ""
# ============================================================
# STEP 7.5 — SEMGREP (SAST engine for the security gate)
# ============================================================
echo "── Step 7.5: Semgrep — SAST security gate ───────────────────"
echo ""
if command -v semgrep &>/dev/null; then
ok "semgrep already installed ($(semgrep --version 2>/dev/null | head -1))"
else
SEMGREP_VER=$(pinned_version "semgrep")
if [ "$SEMGREP_VER" != "latest" ]; then
info "Installing semgrep ${SEMGREP_VER} (pinned in plugins.lock.json)..."
pipx install "semgrep==${SEMGREP_VER}" 2>/dev/null
else
info "Installing semgrep latest (consider pinning in plugins.lock.json)..."
pipx install semgrep 2>/dev/null
fi
if command -v semgrep &>/dev/null; then
ok "semgrep installed ($(semgrep --version 2>/dev/null | head -1))"
else
err "semgrep install failed — run manually: pipx install semgrep"
fi
fi
# Login is Pro-rules only and optional — NEVER run automatically (ctx7
# pattern: guide, don't block). The gate uses pinned public rulesets.
if command -v semgrep &>/dev/null; then
info "Optional Pro rules: semgrep login (never run automatically)"
fi
echo ""
# ============================================================
# STEP 8 — EMIL DESIGN ENG (UI polish / animation skill)
# ============================================================
+104
View File
@@ -0,0 +1,104 @@
# Contract interview — mandatory upstream passage (all orchestrators)
Produces the CONTRACT: the single reference passed verbatim to the plan, the
dev subagents, and the verifier. The contract is what lets the orchestrator
delegate execution without subagents ever needing a human gate (LRN-083:
subagents = execution + report only; gates and loop decisions live in the
main loop).
Run this in the ORCHESTRATOR MAIN LOOP, never in a subagent — STEP 2 may
talk to the human. Mandatory passage in every flow; questions are optional
and proportional — a complete request goes through silently.
## STEP 1 — CAPTURE (verbatim)
Copy the user's request EXACTLY as typed (`$ARGUMENTS` + the triggering
message). No paraphrase, no cleanup, no translation, no summarizing. This
section is IMMUTABLE for the life of the run — every later consumer
(planner, dev, verifier) reads THESE words, never a restatement.
## STEP 2 — AMBIGUITY CHECK (questions optional, proportional)
Ask ONLY if one of these is missing AND not derivable from the repo:
- a testable expected outcome
- an unambiguous scope (what is allowed to change)
- non-contradictory constraints
Complete request → ZERO questions, stay silent. Otherwise: max 3 questions,
one single batch (house rule: one question upfront, never mid-task). Never
ask what the repo can answer — verify paths/APIs/behavior yourself first.
## STEP 3 — DERIVE
- ACCEPTANCE CRITERIA: numbered; each one testable — a fresh reader must be
able to mark it MET / NOT-MET against the real code, without having seen
this conversation.
- FILE SCOPE: paths/zones expected to change, or `repo-wide — <reason>`.
## STEP 4 — WRITE TO DISK (immediately, before any next step)
Path: `.claude/tasks/contracts/<YYYY-MM-DD>-<slug>-<HHMM>.md`
(`mkdir -p` the directory; unique per run: date + short kebab slug + HHMM —
two runs on the same day never collide). A contract that lives only in
context dies at compaction, and the verbatim request with it.
Template:
```markdown
# CONTRACT — <slug>
- date: <YYYY-MM-DD> | flow: <ship-feature|feat|bugfix|hotfix|init-project|onboard> | branch: <branch>
- status: active
## REQUEST (verbatim — IMMUTABLE)
<the user's exact words>
## CLARIFICATIONS
Q: <question> / A: <answer>
(or: none — request complete)
## ACCEPTANCE CRITERIA
1. <testable criterion>
2. <testable criterion>
## FILE SCOPE
<paths/zones>
(or: repo-wide — <reason>)
```
Print one line to the user, then continue the flow:
`CONTRACT: <path> — <n> criteria, scope <files|repo-wide>, <q> questions asked`
## Lifecycle
- **REQUEST**: immutable, for the life of the run. Never rewritten, never
"cleaned up".
- **CRITERIA / FILE SCOPE enrichment**: ONLY at a human gate, each added
entry marked `[gated <YYYY-MM-DD>]`. A dev subagent NEVER enriches the
contract. An out-of-scope edit the dev justifies is accepted ONLY through
this micro-gate: human approves → FILE SCOPE gains the entry `[gated]`;
human declines → the dev removes the edit. Without this gate the dev
justifies everything and scope constrains nothing.
- **Deep re-scope** (the request itself changes): NEW contract file with
`supersedes: <old path>` in its header — never a rewrite of the old one.
- **Aborted run**: delete the contract file, or commit it with
`status: aborted` in the header. NEVER left dirty in the working tree.
- **Commit**: the contract rides the existing memory commit —
`lib/capitalize-commit.md` already covers the `.claude/tasks` pathspec.
No new plumbing.
## Weight per flow
| Flow | Weight |
|------|--------|
| hotfix | Silent autofill — criteria: "symptom gone; build/tests green"; scope = the 1-2 target files. Zero questions ever. |
| feat / bugfix | Proportional. bugfix: the DIAGNOSIS feeds the criteria (symptom reproduced-then-gone + regression test present). |
| ship-feature | Full. Design decisions approved at the validation gate append criteria `[gated <date>]` — the human validates the enriched contract, the verifier receives that version. |
| init-project | Full. The interviewer's PROJECT BRIEF pours into the contract (V1 features → criteria). |
| onboard | Audit-scope contract (interview answers → what to audit, which axes). |
## Hand-off rule
Downstream consumers (plan step, dev subagents, verifier) receive the
contract PATH, not a restatement of its content — the file on disk is the
only authoritative copy, and reading it from disk is what makes the dev's
reformulation structurally unable to interpose.
+87
View File
@@ -0,0 +1,87 @@
#!/usr/bin/env bash
# ============================================================
# Structure locks — contract/verifier pair (verify-loops lot 2)
# Deterministic greps on load-bearing doctrine clauses: an edit
# that silently drops one (blind verifier, PROOF mandatory,
# immutable REQUEST, micro-gate scope enrichment…) reds here.
# ============================================================
set -u
REPO="$(cd "$(dirname "$0")/../.." && pwd)"
LIB="$REPO/lib/contract-interview.md"
AGT="$REPO/agents/verifier.md"
PASS=0; FAIL=0
# Fixed-string lock (UTF-8 punctuation safe)
tf() { # tf <label> <file> <fixed-string>
if grep -qF -- "$3" "$2" 2>/dev/null; then
echo " PASS $1"; PASS=$((PASS+1))
else
echo " FAIL $1 — missing: $3"; FAIL=$((FAIL+1))
fi
}
# Regex lock
tr_() { # tr_ <label> <file> <ERE>
if grep -qE -- "$3" "$2" 2>/dev/null; then
echo " PASS $1"; PASS=$((PASS+1))
else
echo " FAIL $1 — no match: $3"; FAIL=$((FAIL+1))
fi
}
# Negative lock — pattern must NOT match
tn() { # tn <label> <file> <ERE>
if grep -qE -- "$3" "$2" 2>/dev/null; then
echo " FAIL $1 — forbidden match: $3"; FAIL=$((FAIL+1))
else
echo " PASS $1"; PASS=$((PASS+1))
fi
}
echo "── contract-interview.md locks ──"
if [ -f "$LIB" ]; then
echo " PASS lib exists"; PASS=$((PASS+1))
else
echo " FAIL lib missing: $LIB"; FAIL=$((FAIL+1))
fi
tf "verbatim request immutable" "$LIB" "REQUEST (verbatim — IMMUTABLE)"
tf "contracts dir committed path" "$LIB" ".claude/tasks/contracts/"
tf "unique per-run slug" "$LIB" "<YYYY-MM-DD>-<slug>-<HHMM>"
tf "silent when complete" "$LIB" "ZERO questions"
tf "question budget" "$LIB" "max 3 questions"
tf "aborted status" "$LIB" "status: aborted"
tf "never left dirty" "$LIB" "NEVER left dirty"
tf "scope enrichment micro-gate" "$LIB" "micro-gate"
tf "gated marker" "$LIB" "[gated <YYYY-MM-DD>]"
tf "main-loop only" "$LIB" "ORCHESTRATOR MAIN LOOP"
tf "re-scope supersedes" "$LIB" "supersedes:"
tf "hand-off = path not content" "$LIB" "contract PATH, not a restatement"
echo "── verifier.md locks ──"
if [ -f "$AGT" ]; then
echo " PASS agent exists"; PASS=$((PASS+1))
else
echo " FAIL agent missing: $AGT"; FAIL=$((FAIL+1))
fi
tr_ "frontmatter name" "$AGT" "^name: verifier$"
tr_ "tools read-only set" "$AGT" "^tools: Read, Grep, Glob, Bash$"
tn "no write-capable tools" "$AGT" "^tools:.*(Edit|Write|NotebookEdit)"
tf "verdict grammar" "$AGT" "VERIFY — VERDICT: CONFORME | ECARTS(n) | ERROR(<reason>)"
tf "blind — no iteration history" "$AGT" "NEVER receive iteration history"
tf "blind — complete every time" "$AGT" "every verification is complete and blind"
tf "unverifiable is not met" "$AGT" "\`UNVERIFIABLE\` ≠ \`MET\`"
tf "proof mandatory" "$AGT" "\`PROOF\` is MANDATORY"
tf "contract read from disk" "$AGT" "READ it from disk"
tf "checked count equality" "$AGT" "checked count in"
tf "report-only" "$AGT" "Report-only. Never edit"
tf "bash observation only" "$AGT" "OBSERVATION ONLY"
tf "loop bound" "$AGT" "Max 3 iterations"
tf "structural retry then escalate" "$AGT" "2nd structural failure"
tf "mute never a pass" "$AGT" "A mute verifier is NEVER a PASS"
tf "conforme first pass no loop" "$AGT" "proceed straight to the security gate"
tf "reverify order request first" "$AGT" "re-verify the request FIRST"
echo ""
echo "contract-verifier structure locks: $PASS pass, $FAIL fail"
[ "$FAIL" -eq 0 ]
+49
View File
@@ -0,0 +1,49 @@
#!/usr/bin/env bash
# ============================================================
# Structure locks — heavy-flow wiring (verify-loops lot 5)
# ship-feature + init-project get contract + enrich-at-gate +
# verify-secure-loop; onboard is the explicit NO-LOOP audit case.
# ============================================================
set -u
REPO="$(cd "$(dirname "$0")/../.." && pwd)"
SHF="$REPO/skills/ship-feature/SKILL.md"
INI="$REPO/skills/init-project/SKILL.md"
ONB="$REPO/skills/onboard/SKILL.md"
PASS=0; FAIL=0
tf() { # tf <label> <file> <fixed-string> (single-line patterns only, LRN-093)
if grep -qF -- "$3" "$2" 2>/dev/null; then
echo " PASS $1"; PASS=$((PASS+1))
else
echo " FAIL $1 — missing: $3"; FAIL=$((FAIL+1))
fi
}
echo "-- ship-feature (enrich-at-gate) --"
tf "shf contract step" "$SHF" "STEP 0e — CONTRACT"
tf "shf contract-interview" "$SHF" "lib/contract-interview.md"
tf "shf enrich at gate" "$SHF" "ENRICH the STEP 0e contract"
tf "shf gated marker" "$SHF" "[gated <date>]"
tf "shf verify+secure step" "$SHF" "STEP 5 — VERIFY + SECURE"
tf "shf uses shared include" "$SHF" "lib/verify-secure-loop.md"
tf "shf judges enriched" "$SHF" "ENRICHED contract"
tf "shf orthogonal to review" "$SHF" "DISTINCT axis from STEP 6 code review"
echo "-- init-project (contract from BRIEF + adds security) --"
tf "ini contract from brief" "$INI" "contract-interview.md"
tf "ini criteria from V1" "$INI" "V1 FEATURES (each testable)"
tf "ini enrich at gate1" "$INI" "ENRICH the STEP 1 contract"
tf "ini verify+secure step" "$INI" "STEP 9 — VERIFY + SECURE"
tf "ini uses shared include" "$INI" "lib/verify-secure-loop.md"
tf "ini adds security gate" "$INI" "adds the security gate init-project previously lacked"
echo "-- onboard (explicit NO-LOOP audit) --"
tf "onb no-loop stated" "$ONB" "n'a PAS de boucle verify"
tf "onb audit not gate" "$ONB" "MODE: audit"
tf "onb scope contract" "$ONB" "contract de SCOPE"
tf "onb no symmetry loop" "$ONB" "Ne PAS ajouter la boucle des flux dev"
echo ""
echo "loops-heavy structure locks: $PASS pass, $FAIL fail"
[ "$FAIL" -eq 0 ]
+70
View File
@@ -0,0 +1,70 @@
#!/usr/bin/env bash
# ============================================================
# Structure locks — light-flow wiring (verify-loops lot 4)
# feat/bugfix get contract + fresh verifier + security gate
# (bounded 3x); hotfix gets contract + security gate whose
# FAILURE REVERTS (never loops). Locks the load-bearing clauses.
# ============================================================
set -u
REPO="$(cd "$(dirname "$0")/../.." && pwd)"
INC="$REPO/lib/verify-secure-loop.md"
FEA="$REPO/agents/feater.md"
BUG="$REPO/agents/bugfixer.md"
HOT="$REPO/agents/hotfixer.md"
HSK="$REPO/skills/hotfix/SKILL.md"
PASS=0; FAIL=0
tf() { # tf <label> <file> <fixed-string>
if grep -qF -- "$3" "$2" 2>/dev/null; then
echo " PASS $1"; PASS=$((PASS+1))
else
echo " FAIL $1 — missing: $3"; FAIL=$((FAIL+1))
fi
}
tr_() { # tr_ <label> <file> <ERE>
if grep -qE -- "$3" "$2" 2>/dev/null; then
echo " PASS $1"; PASS=$((PASS+1))
else
echo " FAIL $1 — no match: $3"; FAIL=$((FAIL+1))
fi
}
echo "── verify-secure-loop.md (shared include) ──"
if [ -f "$INC" ]; then echo " PASS include exists"; PASS=$((PASS+1)); else echo " FAIL include missing"; FAIL=$((FAIL+1)); fi
tf "gate1 fresh verifier" "$INC" "GATE 1 — REQUEST CONFORMITY (fresh verifier)"
tf "gate2 fresh auditor" "$INC" "GATE 2 — SECURITY (fresh security-auditor)"
tf "blind — no dev summary" "$INC" "Never pass the dev's summary"
tf "conforme first pass no loop" "$INC" "First-pass conforme = no loop"
tf "conformity max 3" "$INC" "Max 3 conformity iterations"
tf "security max 3" "$INC" "Max 3 security iterations"
tf "reverify request first" "$INC" "re-verify the REQUEST first"
tf "order invariant" "$INC" "always re-checked BEFORE security"
tf "mute never a pass (verify)" "$INC" "NEVER a PASS"
tf "nominal cheap stated" "$INC" "one verifier dispatch + one security dispatch"
echo "── feater.md (feat wiring) ──"
tf "feat contract step" "$FEA" "STEP 0.7 — CONTRACT"
tf "feat contract-interview" "$FEA" "lib/contract-interview.md"
tf "feat verify+secure step" "$FEA" "STEP 3 — VERIFY + SECURE"
tf "feat uses shared include" "$FEA" "lib/verify-secure-loop.md"
tf "feat nominal 1+1 dispatch" "$FEA" "verifier + one security dispatch"
echo "── bugfixer.md (bugfix wiring) ──"
tf "bug contract step" "$BUG" "STEP 3.5 — CONTRACT"
tf "bug diagnosis feeds it" "$BUG" "feeds it: REQUEST verbatim"
tf "bug fresh gates" "$BUG" "Fresh gates (verify + secure)"
tf "bug uses shared include" "$BUG" "lib/verify-secure-loop.md"
echo "── hotfixer.md (hotfix wiring — revert, not loop) ──"
tr_ "hotfix has Agent tool" "$HOT" "^tools:.*Agent"
tf "hotfix silent contract" "$HOT" "STEP 1.7 — CONTRACT (silent autofill)"
tf "hotfix zero questions" "$HOT" "questions ever"
tf "hotfix security gate" "$HOT" "Security gate (fresh auditor)"
tf "hotfix block reverts" "$HOT" "failure REVERTS, never loops"
tf "hotfix no verifier" "$HOT" "No verifier is dispatched at hotfix weight"
tf "hotfix skill has Agent" "$HSK" " - Agent"
echo ""
echo "loops-light structure locks: $PASS pass, $FAIL fail"
[ "$FAIL" -eq 0 ]
+55
View File
@@ -0,0 +1,55 @@
#!/usr/bin/env bash
# ============================================================
# Deterministic backstop for LRN-093 (vacuous grep locks).
# grep NEVER interprets a literal \n as a newline: in a -F
# fixed string it splits the pattern into a per-line OR (matches
# anything); in -E it is a literal "n". Either way a structure
# lock carrying \n proves nothing. The LRN advisory alone did not
# hold (2 recurrences same chantier) -> this mechanical guard.
#
# Rule: no backslash-n inside a quoted pattern on a grep / tf /
# tr_ / tn line, anywhere under lib/tests/*.test.sh. Flip-tested
# below against a synthetic offender so the guard proves it bites.
# ============================================================
set -u
REPO="$(cd "$(dirname "$0")/../.." && pwd)"
SELF="no-vacuous-locks.test.sh"
FAIL=0
# One scanner, reused for the real tree and the flip-test fixture.
# Matches a grep -*F/-*E call OR a tf/tr_/tn helper (at line start or after
# whitespace) whose quoted pattern contains a literal backslash-n.
scan() { # scan <dir containing *.test.sh>
grep -rnE '(grep +-[A-Za-z]*[EFqe]|(^|[[:space:]])(tf|tr_|tn)[[:space:]]).*"[^"]*\\n' \
"$1"/*.test.sh 2>/dev/null | grep -v "$SELF"
}
echo "-- LRN-093 backstop: scan lib/tests/*.test.sh --"
HITS="$(scan "$REPO/lib/tests")"
if [ -n "$HITS" ]; then
FAIL=1
printf '%s\n' "$HITS" | while IFS= read -r line; do
printf ' FAIL vacuous backslash-n lock: %s\n' "$line"
done
else
printf ' PASS no vacuous backslash-n locks in lib/tests/*.test.sh\n'
fi
# Flip-test: the guard MUST catch a known offender (LRN-093 discipline —
# prove a lock CAN fail before trusting its green).
echo "-- flip-test: guard bites a synthetic offender --"
TMP="$(mktemp -d)"
# shellcheck disable=SC2016 # the single quotes are deliberate: literal backslash-n
printf '%s\n' 'tf "bad" "$F" "no\nforced loop"' > "$TMP/z.test.sh"
if [ -n "$(scan "$TMP")" ]; then
printf ' PASS guard catches the synthetic offender\n'
else
printf ' FAIL guard blind to a known offender (regex too weak)\n'
FAIL=1
fi
rm -rf "$TMP"
echo ""
if [ "$FAIL" -eq 0 ]; then echo "no-vacuous-locks: clean"; else echo "no-vacuous-locks: vacuous locks present"; fi
[ "$FAIL" -eq 0 ]
+78
View File
@@ -0,0 +1,78 @@
#!/usr/bin/env bash
# ============================================================
# Structure locks — security-auditor agent + grafts (lot 3)
# Deterministic greps on load-bearing doctrine: an edit that
# drops one (pinned rulesets, DEGRADED-still-checks, PROOF,
# block-HIGH-only, anti-gaming, the two SKILL grafts) reds here.
# ============================================================
set -u
REPO="$(cd "$(dirname "$0")/../.." && pwd)"
AGT="$REPO/agents/security-auditor.md"
ONB="$REPO/skills/onboard/SKILL.md"
ADL="$REPO/skills/audit-delta/SKILL.md"
PASS=0; FAIL=0
tf() { # tf <label> <file> <fixed-string>
if grep -qF -- "$3" "$2" 2>/dev/null; then
echo " PASS $1"; PASS=$((PASS+1))
else
echo " FAIL $1 — missing: $3"; FAIL=$((FAIL+1))
fi
}
tr_() { # tr_ <label> <file> <ERE>
if grep -qE -- "$3" "$2" 2>/dev/null; then
echo " PASS $1"; PASS=$((PASS+1))
else
echo " FAIL $1 — no match: $3"; FAIL=$((FAIL+1))
fi
}
tn() { # tn <label> <file> <ERE> (must NOT match)
if grep -qE -- "$3" "$2" 2>/dev/null; then
echo " FAIL $1 — forbidden match: $3"; FAIL=$((FAIL+1))
else
echo " PASS $1"; PASS=$((PASS+1))
fi
}
echo "── security-auditor.md locks ──"
if [ -f "$AGT" ]; then
echo " PASS agent exists"; PASS=$((PASS+1))
else
echo " FAIL agent missing: $AGT"; FAIL=$((FAIL+1))
fi
tr_ "frontmatter name" "$AGT" "^name: security-auditor$"
tr_ "tools incl Write (audit)" "$AGT" "^tools: Read, Grep, Glob, Bash, Write$"
tf "verdict grammar" "$AGT" "SECURITY — VERDICT: PASS | BLOCK(n) | ERROR(<reason>)"
tf "ruleset security-audit" "$AGT" "p/security-audit"
tf "ruleset secrets" "$AGT" "p/secrets"
tf "ruleset owasp required" "$AGT" "p/owasp-top-ten"
tf "no config auto stated" "$AGT" "never \`--config auto\`"
tf "no auto login" "$AGT" "never \`semgrep login\`"
tf "secrets to CRITICAL" "$AGT" "p/secrets | CRITICAL"
tf "block ERROR threshold only" "$AGT" "blocking threshold is ERROR"
tf "medium low reported" "$AGT" "MEDIUM/LOW are REPORTED, never"
tf "degraded still checks" "$AGT" "STILL RUN STEP 3"
tf "degraded vacuous pass named" "$AGT" "vacuous pass"
tf "anti-gaming suppression" "$AGT" "NEW suppression comment"
tf "anti-gaming micro-gate" "$AGT" "[gated <date>]"
tf "proof mandatory" "$AGT" "\`PROOF\` is MANDATORY"
tf "mute never a pass" "$AGT" "NEVER a PASS"
tf "write rule-locked audit" "$AGT" "writable path is \`REPORT\`"
tf "gate mode write forbidden" "$AGT" "\`Write\` is FORBIDDEN in this mode"
tf "blind no history" "$AGT" "NEVER receive iteration history"
tf "reverify request first" "$AGT" "re-verify the REQUEST first"
tf "max 3 security iters" "$AGT" "Max 3 security iterations"
echo "── onboard graft locks ──"
tf "onboard dispatches auditor" "$ONB" "subagent_type=\"security-auditor\""
tf "onboard report path" "$ONB" ".onboard-audit/semgrep.md"
tf "onboard verify incl semgrep" "$ONB" "code-clean,cso,semgrep,doc"
echo "── audit-delta graft locks ──"
tf "audit-delta dispatches" "$ADL" "subagent_type=\"security-auditor\""
tf "audit-delta semgrep first" "$ADL" "FIRST run the semgrep SAST pass"
echo ""
echo "security-auditor structure locks: $PASS pass, $FAIL fail"
[ "$FAIL" -eq 0 ]
+67
View File
@@ -0,0 +1,67 @@
# Verify + secure loop — shared orchestrator include (feat, bugfix)
Runs in the ORCHESTRATOR MAIN LOOP after the dev step completes. Turns a
finished diff into a verified, security-cleared change through two fresh
gates and bounded loops. The dev stays inline (LRN-083: subagents =
execution + report; loop decisions live here, in the main loop).
Inputs the caller must have ready:
- `CONTRACT`: path to the contract file written by `contract-interview.md`.
- `DIFF`: the range/file-list the dev just produced (e.g. `HEAD` vs the
pre-dev SHA, or the working-tree diff before commit).
- `TEST`: the project test command, if known.
Nominal path is cheap: one verifier dispatch + one security dispatch, done.
The loop only costs more when it actually loops.
## GATE 1 — REQUEST CONFORMITY (fresh verifier)
Dispatch a FRESH verifier subagent (`subagent_type: verifier`, or load
`agents/verifier.md`). Pass ONLY: the `CONTRACT` path, the `DIFF` range, the
`TEST` command. Never pass the dev's summary, never pass a prior iteration's
gaps — the verifier reads the contract from disk and judges blind.
Parse its single `VERIFY — VERDICT:` line:
- `CONFORME` → go to GATE 2. (First-pass conforme = no loop.)
- `ECARTS(n)` → hand the dev the CONTRACT path + the exact `CRITERIA` gap
lines (NOT-MET / out-of-scope), nothing else. Dev fixes inline, then
re-dispatch a FRESH verifier. Repeat. **Max 3 conformity iterations** →
STOP + human escalation with the CRITERIA table (the contract-vs-realized
diff).
- Remaining `UNVERIFIABLE` while all else MET → direct human gate (a dev
cannot fix unverifiability); do not spend a loop on it.
- Out-of-scope files: a dev justification is accepted ONLY through the human
micro-gate that appends `[gated <date>]` to the contract's FILE SCOPE;
otherwise the dev removes the file.
- Structural failure (`ERROR(…)`, missing/duplicated VERDICT line,
unparsable, crash, `CONFORME` without `PROOF`) → retry ONCE with a fresh
verifier; a 2nd structural failure → human escalation. A mute verifier is
NEVER a PASS.
## GATE 2 — SECURITY (fresh security-auditor)
Only after GATE 1 is `CONFORME`. Dispatch a FRESH security-auditor
(`subagent_type: security-auditor`, or load `agents/security-auditor.md`)
with `MODE: gate`, `SCOPE: <DIFF>`. No report path (gate mode is
stdout-only, no Write).
Parse its single `SECURITY — VERDICT:` line:
- `PASS` → done, proceed to commit.
- `BLOCK(n)` → hand the dev the `BLOCKING` list + the CONTRACT path. Dev
fixes inline. Then **re-verify the REQUEST first** (GATE 1, fresh
verifier) — a security fix can drift the behavior — **then re-run GATE 2**
(fresh auditor), in that order. **Max 3 security iterations** → STOP +
human escalation with the BLOCKING table.
- `DEGRADED` (semgrep absent) → does NOT block on the tool's absence; surface
the checklist result + recommend `make plugin`. A DEGRADED run that still
BLOCKs (grep-caught secret/injection) blocks like any other.
- Structural failure → retry ONCE fresh; 2nd → human escalation. A mute
auditor is NEVER a PASS.
## Order invariant
REQUEST conformity is always re-checked BEFORE security on any re-loop — a
security fix that breaks the feature must not slip through because only the
security gate re-ran. Never the reverse order.
+6
View File
@@ -26,6 +26,12 @@
"managed_by": "pipx",
"note": "Codebase knowledge graph. CLI is 'graphify'. Install: pipx install graphifyy && graphify install && graphify claude install. Adds PreToolUse hook for Glob/Grep."
},
"semgrep": {
"source": "pypi:semgrep",
"version": "1.168.0",
"managed_by": "pipx",
"note": "SAST engine for the security gate (security-auditor agent, onboard cso fallback, audit-delta). Rulesets pinned in-agent: p/security-audit + p/secrets (never --config auto). BLOCKING gate -> pin honored by update-all.sh: 'make update' will NOT advance semgrep past it; bump deliberately (new rules = new BLOCKs on unchanged code). Never run 'semgrep login' automatically (Pro rules are optional, guide-only)."
},
"emil-design-eng": {
"source": "https://github.com/emilkowalski/skill",
"path": "skills/emil-design-eng/SKILL.md",
+15 -6
View File
@@ -221,12 +221,21 @@ Then offer to capitalize (per CLAUDE.md): recurring finding patterns →
## Axis specs (subagent prompts)
- **security** — scoped to the delta: hardcoded secrets/tokens/keys (also
in comments), injection (SQL/XSS/command — string concat into
queries/shells), authN/authZ gaps on new endpoints, fail-open error
paths, secrets/PII in logs, new dependencies in lockfiles (name them +
known CVEs), unguarded destructive shell (`rm -rf` with unquoted or
un-`:?`-guarded vars).
- **security** — FIRST run the semgrep SAST pass, THEN the reasoned checks
below on the same delta (the SAST is a deterministic floor, the reasoned
pass covers what grep/rules miss):
```
Agent(subagent_type="security-auditor", description="audit-delta security — semgrep SAST",
prompt="MODE: audit\nSCOPE: <delta file list>\nREPORT: .claude/audits/.audit-delta-semgrep.md\nFollow agents/security-auditor.md exactly. Pinned rulesets, no login. Write ONLY to REPORT. End with REPORT_WRITTEN: <path>.")
```
Fold its BLOCKING (CRITICAL/HIGH) + REPORTED findings into this axis'
finding list before the 3c gate. semgrep ABSENT → DEGRADED (checklist
only) is surfaced, not a blocker. Reasoned checks (also scoped to the
delta): hardcoded secrets/tokens/keys (also in comments), injection
(SQL/XSS/command — string concat into queries/shells), authN/authZ gaps
on new endpoints, fail-open error paths, secrets/PII in logs, new
dependencies in lockfiles (name them + known CVEs), unguarded destructive
shell (`rm -rf` with unquoted or un-`:?`-guarded vars).
- **errors** — bugs in changed code: logic errors, off-by-one, unhandled
edge cases (empty/null/unicode/concurrent), race conditions, swallowed
errors, resource leaks (missing trap/close/finally). Improvements only
+1
View File
@@ -15,6 +15,7 @@ allowed-tools:
- Bash
- Grep
- Glob
- Agent
---
Load and follow strictly:
+28 -2
View File
@@ -48,6 +48,14 @@ ls CLAUDE.md .claude/CLAUDE.md 2>/dev/null | head -1
In both cases: MANDATORY STOP until user answers remaining questions. Produce PROJECT BRIEF.
**Then run `$HOME/.claude/lib/contract-interview.md`** seeded from the BRIEF:
REQUEST verbatim = the user's project description; ACCEPTANCE CRITERIA = the
V1 FEATURES (each testable); FILE SCOPE = the planned tree. No new questions
(the interview already asked). It writes
`.claude/tasks/contracts/<date>-<slug>-<HHMM>.md`; the DESIGN approved at STEP
4 ENRICHES it, and STEP 9's verifier judges the MVP against the enriched
contract.
## STEP 2 — ANALYZE
Load `$HOME/.claude/agents/analyzer.md`. Analyze BRIEF: existing code, stack constraints, infra risks, open decisions. Produce ANALYSIS REPORT.
@@ -70,6 +78,11 @@ Approve? (yes / request changes)
```
Changes → back to STEP 3. Approved → continue.
**On approval — ENRICH the STEP 1 contract**: append the DESIGN-derived
acceptance criteria (resolved decisions, interfaces, test strategy) to the
contract, each tagged `[gated <date>]`. STEP 9's verifier judges against this
enriched contract.
## STEP 5 — SCAFFOLD
Load `$HOME/.claude/agents/scaffolder.md`. Pass: BRIEF + DESIGN + `~/.claude/templates/project-CLAUDE.md` + `~/.claude/CLAUDE.md`.
Creates: CLAUDE.md, settings, structure, config, empty entry points, .gitignore, .env.example, .claude/tasks/TODO.md, .claude/memory/{decisions,learnings,blockers,journal,evals}.md, .claude/audits/. NO README, NO features.
@@ -175,8 +188,21 @@ If `graphify` CLI is installed AND complexity >= 30%:
2. Print: `🔗 Full project graph updated at graphify-out/`
If `graphify` not installed or complexity < 30% → skip silently.
## STEP 9 — ANALYZE
Load `$HOME/.claude/agents/analyzer.md`. Check: no regressions, no deviations, no stale scaffold, conventions respected.
## STEP 9 — VERIFY + SECURE (fresh gates, bounded loops)
Run the two fresh gates per `$HOME/.claude/lib/verify-secure-loop.md` with
`CONTRACT` = the STEP 1 path (ENRICHED at STEP 4), `DIFF` = the MVP branch
diff (`develop..HEAD`), `TEST` = the project suite:
- GATE 1 — a FRESH verifier judges the MVP against the enriched contract (V1
features + `[gated]` design criteria). CONFORME → GATE 2. ECARTS → fix,
re-verify, max 3 → STOP + human escalation with the CRITERIA table.
- GATE 2 — a FRESH security-auditor (`MODE: gate`, `SCOPE: develop..HEAD`).
PASS → STEP 10. BLOCK → fix, re-verify request THEN re-scan, max 3 →
escalate.
This adds the security gate init-project previously lacked (security was only
deferred to a later /onboard) and turns the informal analyze into a verdict
against the founding contract. Distinct axis from STEP 10 code review
([[LRN-095]]) — both run.
## STEP 10 — CODE REVIEW
Invoke `superpowers:requesting-code-review`. Fix all CRITICAL before proceeding.
+35 -2
View File
@@ -486,6 +486,38 @@ bash $HOME/.claude/lib/toggle-external.sh list 2>/dev/null | grep -E "^gstack\s+
)
```
#### Dispatch semgrep SAST — `security-auditor` (TOUJOURS, complément de cso)
En complément de cso (ON) OU du fallback (OFF) — un moteur SAST déterministe
à côté de l'audit grep/raisonné. cso est un submodule gstack non modifiable ;
semgrep vit dans cet agent local. Lancé dans les DEUX branches gstack.
```
Agent(
subagent_type="security-auditor",
description="Onboard — semgrep SAST audit (report-only)",
prompt="""
MODE: audit
SCOPE: <PROJECT_ROOT>
REPORT: <PROJECT_ROOT>/.onboard-audit/semgrep.md
CONTEXT: <PROJECT_ROOT>/.onboard-audit/archetype-context.md
Follow agents/security-auditor.md exactly. Pinned rulesets only, no login.
Write ONLY to the REPORT path. End stdout with REPORT_WRITTEN: <path>.
"""
)
```
Si semgrep absent → l'agent rend DEGRADED (checklist seule) + recommande
`make plugin` ; NON bloquant en onboard (audit, pas gate).
**Onboard n'a PAS de boucle verify→dev (`lib/verify-secure-loop.md`) — par
conception.** onboard produit un RAPPORT d'audit, pas une modification à
vérifier contre une demande : il n'y a ni contract de conformité, ni diff dev,
ni verifier, ni max-3. Le contract d'onboard est un contract de SCOPE (ce que
l'interview STEP 3 + `audit_stack` définissent comme périmètre d'audit), et
`security-auditor` tourne en `MODE: audit` (report-only), jamais en `MODE:
gate`. Ne PAS ajouter la boucle des flux dev ici par symétrie — l'audit et le
flux de dev sont deux formes distinctes ([[BDR-050]] pipeline dev ≠ audit).
#### Dispatch doc-syncer (si `doc` dans audit_stack)
```
Agent(
@@ -511,9 +543,10 @@ Agent(
### Après les 3 dispatches
Attendre la fin des 3 subagents. Vérifier que les 3 fichiers existent et sont non vides :
Attendre la fin des subagents. Vérifier que les fichiers existent et sont non vides
(semgrep.md inclus — DEGRADED reste non vide : il porte le résultat checklist) :
```bash
for f in .onboard-audit/{code-clean,cso,doc}.md; do
for f in .onboard-audit/{code-clean,cso,semgrep,doc}.md; do
[ -s "$f" ] && echo "OK $f" || echo "MISSING $f"
done
```
+33 -3
View File
@@ -78,7 +78,16 @@ The returned digest (ANALYSIS + RELATED MEMORY) stays in the orchestrator's cont
is FED to STEP 1 and STEP 2 and reconciled at STEP 3. Degradation: request too vague →
analyzer flags ambiguous zones, does not block (STEP 1 refines). `.claude/memory/` empty or
absent → analyzer omits RELATED MEMORY (no-op); the step still returns the code ANALYSIS.
Additive — distinct from STEP 5 ANALYZE (post-impl regression) and STEP 4b DEBUG.
Additive — distinct from STEP 5 VERIFY + SECURE (post-impl) and STEP 4b DEBUG.
## STEP 0e — CONTRACT
Run `$HOME/.claude/lib/contract-interview.md`. REQUEST verbatim = the feature
request as typed; initial ACCEPTANCE CRITERIA from the request; FILE SCOPE
seeded from 0d's KEY COMPONENTS. It writes
`.claude/tasks/contracts/<date>-<slug>-<HHMM>.md`; keep the path — the design
approved at STEP 3 ENRICHES it, and STEP 5's verifier judges the diff against
the ENRICHED contract. This is the only flow where the contract grows mid-run.
## STEP 1 — BRAINSTORM
Invoke `superpowers:brainstorming` — but FEED it the STEP 0d digest as binding context,
@@ -120,6 +129,13 @@ never a guarantee (same discipline as the memory-commit `✅<hash>`: show what's
assert a check not performed). No RELATED MEMORY from 0d → omit the block.
Changes → back to STEP 2. Approved → continue.
**On approval — ENRICH the STEP 0e contract.** The design just validated adds
detail the raw request lacked: append the design-derived acceptance criteria
to the contract's ACCEPTANCE CRITERIA, each tagged `[gated <date>]` (this is
the human micro-gate that authorizes contract growth). STEP 5's verifier
judges the diff against this ENRICHED contract, not the STEP 0e seed — so a
criterion the design introduced is verified, not lost.
## STEP 4 — IMPLEMENT
Start the feature branch off develop, then implement on it:
```bash
@@ -157,8 +173,22 @@ OPTIONS :
Skip them too? (yes / keep and accept partial implementation)"
If no dependents → skip cleanly and continue.
## STEP 5 — ANALYZE
Load `$HOME/.claude/agents/analyzer.md`. Check: no regressions, no stale code, no plan deviations.
## STEP 5 — VERIFY + SECURE (fresh gates, bounded loops)
Run the two fresh gates per `$HOME/.claude/lib/verify-secure-loop.md` with
`CONTRACT` = the STEP 0e path (ENRICHED at STEP 3), `DIFF` = the branch diff
(`develop..HEAD`), `TEST` = the project suite:
- GATE 1 — a FRESH verifier judges the branch against the ENRICHED contract
(all criteria, including the `[gated]` design ones). CONFORME → GATE 2.
ECARTS → hand the dev the gap list, fix, re-verify, max 3 → STOP + human
escalation with the CRITERIA table.
- GATE 2 — a FRESH security-auditor (`MODE: gate`, `SCOPE: develop..HEAD`)
scans the branch. PASS → STEP 6. BLOCK → fix, re-verify the request THEN
re-scan, max 3 → escalate.
This replaces the old informal "analyze for regressions" with a verdict
against the contract. It is a DISTINCT axis from STEP 6 code review (contract
conformity + security vs. craft/design) — both run, neither subsumes the
other ([[LRN-095]]).
## STEP 6 — CODE REVIEW
Invoke `superpowers:requesting-code-review`. Fix all CRITICAL before proceeding.
+41
View File
@@ -227,6 +227,47 @@ else
info "graphifyy not installed — skipping"
fi
# ── 6.2. Update Semgrep (pin-honored — BLOCKING security gate) ──
echo ""
echo "── Updating Semgrep..."
if command -v semgrep &>/dev/null; then
SEMGREP_VER=""
if [ -f "$REPO/plugins.lock.json" ] && command -v python3 &>/dev/null; then
SEMGREP_VER=$(python3 -c "
import json
with open('$REPO/plugins.lock.json') as f:
d = json.load(f)
print(d.get('semgrep',{}).get('version',''))
" 2>/dev/null || true)
fi
SEMGREP_CUR=$(semgrep --version 2>/dev/null | head -1)
if [ -n "$SEMGREP_VER" ] && [ "$SEMGREP_VER" != "latest" ]; then
if [ "$SEMGREP_CUR" = "$SEMGREP_VER" ]; then
ok "semgrep already at pinned $SEMGREP_VER"
else
# Jump shown explicitly: semgrep is a BLOCKING gate — a version bump
# can add rules that BLOCK unchanged code, so the jump must be a
# visible, deliberate human decision (bump the pin, then update).
info "semgrep ${SEMGREP_CUR:-?} → ${SEMGREP_VER} (pinned in plugins.lock.json)"
if pipx install --force "semgrep==${SEMGREP_VER}" 2>/dev/null; then
ok "semgrep updated to $SEMGREP_VER"
else
warn "semgrep update failed — try: pipx install --force semgrep==${SEMGREP_VER}"
fi
fi
else
info "No pinned version — upgrading to latest"
if pipx upgrade semgrep 2>/dev/null; then
ok "semgrep updated ($(semgrep --version 2>/dev/null | head -1))"
else
warn "semgrep update failed — try: pipx upgrade semgrep"
fi
fi
else
info "semgrep not installed — skipping (run: make plugin)"
fi
# ── 6.5. Update bun ──
echo ""
echo "── Updating bun..."