diff --git a/.claude/memory/decisions.md b/.claude/memory/decisions.md index 6e1cfae..fb3b918 100644 --- a/.claude/memory/decisions.md +++ b/.claude/memory/decisions.md @@ -69,6 +69,10 @@ rules: | BDR-045 | 2026-07-01 | Standalone memory/doc skills branch to chore/* via aiguillage (hook exemption kept) | accepted | | BDR-046 | 2026-07-01 | Claude Code installs via official native installer (curl claude.ai/install.sh), drop npm from install.sh | accepted | | BDR-047 | 2026-07-01 | ECC audit → zero import; local config ahead of reference | accepted | +| BDR-048 | 2026-07-03 | semgrep security gate: engine version + rulesets PINNED, never --config auto; upgrade = deliberate visible human jump | accepted | +| BDR-049 | 2026-07-03 | verifier = fresh + blind (no iteration history) + disk-contract + PROOF-or-fail; mute ≠ PASS; scope enrichment via human micro-gate | accepted | +| BDR-050 | 2026-07-03 | universal pipeline (contract→dev inline→fresh verify→fresh security, loops bounded 3× in main loop) with per-flow weighting; hotfix failure = revert not loop | accepted | +| BDR-051 | 2026-07-04 | contract enrich-at-gate: the contract grows ONLY at a human micro-gate ([gated] marker); the verifier judges the ENRICHED contract, not the seed | accepted | --- @@ -788,3 +792,36 @@ rules: ONE scope gap: BDR-047 never opened hooks/ — ECC's only WIRED subsystem. Fruit: config-protection hook (own idiom, NOT ECC import), shipped feature/config-protection-hook. Lesson holds + refined by [[LRN-090]]. + +## BDR-048 — Deterministic security gate: pinned engine + pinned rulesets (semgrep) + +- **Date**: 2026-07-03 +- **Decision**: semgrep = BLOCKING gate (verify-loops chantier) → engine version PINNED in plugins.lock.json (gsd-pin pattern; update-all.sh honors pin + displays jump cur→pin before `pipx install --force`). Rulesets PINNED in-agent: `p/security-audit` + `p/secrets`. Never `--config auto` (registry telemetry + ruleset resolved per-run = non-deterministic gate, [[LRN-077]] class). Never auto `semgrep login` — Pro rules optional, guide-only (ctx7 pattern). +- **Rationale**: gate blocks HIGH/CRITICAL only ([[LRN-047]]); silent engine/rule upgrade = new BLOCKs on unchanged code w/o human decision → gate crying false → ignored. Version jump must be deliberate + visible (bump pin, then `make update` shows the jump). +- **Alternatives rejected**: `latest` (pipx house default, graphifyy-style) — fine for comfort tools, wrong for a blocking gate; `--config auto` — telemetry + non-determinism. +- **Reference**: plugins.lock.json `semgrep` entry, install-plugins.sh STEP 7.5, update-all.sh step 6.2 — branch feature/semgrep-install `ccfecc9`. Conditions [[LRN-047]], [[LRN-085]]. Coverage caveat of the community rulesets: [[LRN-092]]. +- **Addendum 2026-07-03** (lot 3, measured): rulesets = `p/security-audit` + `p/secrets` + **`p/owasp-top-ten`**. owasp-top-ten is REQUIRED not optional — measured on realistic Flask code, the 2-ruleset baseline missed SQL injection + path traversal ENTIRELY (0 findings); owasp's taint rules catch them. Severity map: secrets ERROR→CRITICAL, other ERROR→HIGH (block), WARNING/INFO→reported. Blocking threshold = ERROR (per-RULE, not per-vuln — same class can straddle ERROR/WARNING; blocking WARNING too floods FP). FP measured shell/md only (faunosteo, game: sole added blocking ERROR = Dockerfile `missing-user` hygiene, contained by gate-mode diff-scoping). **Re-evaluate owasp FP at the first real web/python app project** (shell/md repos don't represent where the gate runs). See [[LRN-094]], agents/security-auditor.md branch feature/security-auditor `2b297bd`. + +## BDR-049 — Verifier doctrine: fresh + blind + disk-contract + proof-or-fail + +- **Date**: 2026-07-03 +- **Decision**: conformity verdict comes ONLY from a FRESH verifier subagent per iteration. Input = contract PATH (read from disk — dev restatement structurally unable to interpose) + diff range + optional test cmd. NEVER iteration history: blind, complete verification every time (cost bounded by the main-loop max-3 cap, [[LRN-083]]: loops decided in main loop). CONFORME ⇔ all criteria MET + zero out-of-scope. PROOF line mandatory ([[LRN-048]]). Mute/unparsable verifier NEVER a PASS: 1 fresh retry, 2nd structural failure = human escalation. Dev-justified out-of-scope enters FILE SCOPE only via a human micro-gate (`[gated]` marker) — else the dev justifies everything and scope constrains nothing. Contract on DISK at creation (`.claude/tasks/contracts/--.md`, committed; aborted run → deleted or `status: aborted`, never left dirty). +- **Rationale**: dev self-score is always confident → not a gate. Verifier fed history anchors on prior verdicts → telescopic drift. Context-only contract dies at compaction, the verbatim with it. +- **Alternatives rejected**: dev self-assessment as gate; cumulative verifier context ("cheaper" but anchored); gitignored run files (lose escalation reference + session-death survival). +- **Reference**: lib/contract-interview.md + agents/verifier.md + lib/tests/contract-verifier.test.sh (31 locks) — branch feature/contract-verifier `6aed5ee`. Behavioral GREEN: planted-gap → ECARTS(2) exact; conform-under-injected-history → CONFORME (blindness held). Twin of [[BDR-048]] (security gate). Conditions [[LRN-048]], [[LRN-083]]. + +## BDR-050 — Universal verify+secure pipeline, weighted per flow (loops in the main loop) + +- **Date**: 2026-07-03 +- **Decision**: every dev flow = contract (verbatim, on disk) → dev INLINE → fresh verifier (request conformity) → fresh security-auditor (`MODE: gate`) → commit. Loops BOUNDED at 3 and decided in the ORCHESTRATOR MAIN LOOP ([[LRN-083]]), never in a subagent. Order invariant: on any security re-loop, re-verify the REQUEST before re-scanning security. Per-flow weight: feat/bugfix = both gates, both loop (nominal 2 dispatches); hotfix = NO fresh verifier (its smoke-check verifies the trivial autofill contract), security gate whose FAILURE REVERTS (`git restore` + escalate to /bugfix), never loops — the 1-attempt model preserved (nominal 1 dispatch). Shared include `lib/verify-secure-loop.md` for feat/bugfix; hotfix inline variant. +- **Rationale**: the value is the INDEPENDENCE of the gate (fresh subagent vs a rich contract), NOT delegating the dev — so dev stays inline in light flows and weighting lives on loops+questions, never on skipping a gate. hotfix reverts because a 3× loop would reintroduce the weight its identity excludes. +- **Alternatives rejected**: dispatch the dev too (turns feat into ship-feature-bis); one merged "quality" gate (see [[LRN-095]] — orthogonal gates degrade if fused); hotfix loops like feat (breaks its 1-attempt identity). +- **Reference**: lib/verify-secure-loop.md + wired feater/bugfixer/hotfixer + lib/tests/loops-light.test.sh (27 locks) — feature/verify-loops `0f0162d`. Behavioral GREEN (feat fixture): CONFORME→BLOCK(1) SQLi→fix→re-verify CONFORME→re-scan PASS, order invariant held. Builds on [[BDR-048]] [[BDR-049]]. Conditions [[LRN-083]] [[LRN-095]]. + +## BDR-051 — Contract enrich-at-gate: the contract grows only at a human micro-gate + +- **Date**: 2026-07-04 +- **Decision**: the CONTRACT's REQUEST is immutable, but ACCEPTANCE CRITERIA + FILE SCOPE may GROW — exclusively at a human gate, each added entry tagged `[gated ]`. In the heavy flows (ship-feature STEP 3, init-project GATE #1) the approved DESIGN appends design-derived criteria to the contract; the fresh verifier then judges the diff against the ENRICHED contract, never the seed. Same mechanism as the out-of-scope micro-gate ([[BDR-049]]) — a dev never enriches; only the human validating a gate does. +- **Rationale**: the raw request underspecifies (a one-line "add validation" hides the schema-rejection requirement the design surfaces). If the verifier judged only the seed, every design decision would be unverified. Gating the growth keeps the contract honest (no silent scope creep) AND complete (design criteria are verified). The only flow where the contract is mutable mid-run — bounded to gate moments. +- **Alternatives rejected**: freeze the contract at creation (design criteria unverified — the seed is too thin); let the dev enrich (the [[BDR-049]] failure mode — dev justifies everything, scope constrains nothing); a second contract per design (loses the single-reference property). +- **Reference**: ship-feature STEP 0e+3, init-project STEP 1+4, feature/verify-loops `1c69de2`. Behavioral GREEN: a `[gated 2026-07-04]` design criterion (reject unknown config keys) was read + judged NOT-MET by a fresh verifier across 3 rounds (dogfood). Builds on [[BDR-049]] [[BDR-050]]. diff --git a/.claude/memory/journal.md b/.claude/memory/journal.md index d91825b..27303e3 100644 --- a/.claude/memory/journal.md +++ b/.claude/memory/journal.md @@ -314,3 +314,9 @@ rules: - Next: #2 design-toolchain trigger fix (residual false-fires post-ed2408e, 5× this session). - #2 done (bugfix/design-toolchain-trigger): trigger tightened — dropped bare design|component|composant|theme|thème|transition|frontend|front-end|palette; dashboard→\bdashboard\b (kills ecc_dashboard.py filename match, keeps "admin dashboard"); kept animation; added "front-?end design" bigram + fire-log counter (time+token+excerpt, ~/.claude/logs/design-toolchain-fires.log) so future "re-firing?" is measured. Test 18/18, shellcheck clean, live dogfood green. [[LRN-091]] corrob [[LRN-047]]. - Double dogfood of #1 guard: config-protection blocked + sentinel-bypassed my own edits to the now-guarded design hook + its test — first real use of the guard, friction validated in passing (one-shot sentinel .claude/.config-edit-ok, non-empty reason, logged+consumed). ECC second-regard closed: #1 config-protection + #2 trigger fix, both merged to develop, nothing pushed. +- Chantier verify-loops/semgrep/contract: Phase 1 read-only (6 subagents mapped 6 orchestrators + cso + agents + install patterns; caught subagent error — cso IS gstack symlink, ls-verified) → archi GATED-GO (5 verdicts: local grafts, dev inline light flows, hotfix unchanged, pinned rulesets, pinned version; +2 specs: contract on DISK, mute verifier ≠ PASS). LOT 1 shipped on feature/semgrep-install (ccfecc9+b8d3ccc): install-plugins STEP 7.5 + update-all 6.2 + lock pin 1.168.0, dogfooded real (4 paths + anonymous ruleset fetch + detection). [[BDR-048]] [[LRN-092]]. Next: lot 2 specs (contract-interview lib + verifier agent). +- Chantier verify-loops LOT 2 (feature/contract-verifier `6aed5ee`): lib/contract-interview.md (verbatim contract on DISK, micro-gate scope enrichment, aborted never dirty) + agents/verifier.md (fresh+blind, PROOF-or-fail, mute ≠ PASS) + 31 structure locks green, shellcheck clean. Behavioral: planted-gap → ECARTS(2) exact; conform under injected fake history → CONFORME (blindness held). Sentinel consumed 4× on guarded lib/tests/. [[BDR-049]] [[LRN-093]]. Merge note: lot 1+2 both append registries at same anchors → trivial stack-conflict expected. Next: lot 3 security-auditor spec. +- Chantier verify-loops LOT 3 (feature/security-auditor `2b297bd`): agents/security-auditor.md (SAST gate, pinned p/security-audit+p/secrets+p/owasp-top-ten, secrets→CRITICAL, block ERROR only, DEGRADED-still-checks, anti-gaming nosemgrep, PROOF-or-fail) + grafts onboard L3a (complement to cso, both gstack branches) + audit-delta security axis. 28 structure locks + 4 behavioral dogfoods green: vuln→BLOCK(9), nosemgrep→BLOCK(1), DEGRADED→BLOCK(7). owasp REQUIRED (measured: baseline misses SQLi+path-traversal on Flask). [[LRN-094]] + [[BDR-048]] addendum (owasp/severity/FP) applied at integration on feature/verify-loops (index drift LRN-090/091 backfilled same pass). Next: lot 4 loops-light (feat/bugfix/hotfix wiring). +- Integration: feature/verify-loops = develop + merge lots 1-3 (local, develop/main intact, nothing pushed) so lots 4-5 wiring is dogfoodable against present agents. Memory stack-conflicts resolved (BDR-048/049, LRN-092/093/094 stacked ID-order; BDR-048 addendum applied; LRN-090/091 index rows backfilled). +- Chantier verify-loops LOT 4 (feature/verify-loops `0f0162d`): lib/verify-secure-loop.md shared include + wired feater (0.7 contract, 3 verify+secure), bugfixer (3.5 contract from diagnosis, 5 gates), hotfixer (1.7 silent contract, 3 security gate FAILURE=REVERT not loop, +Agent tool). 27 structure locks + full pipeline dogfood: feat fixture w/ SQLi → GATE1 CONFORME → GATE2 BLOCK(1) (checklist caught what semgrep taint missed) → fix → re-verify CONFORME (order invariant) → re-scan PASS. [[BDR-050]] [[LRN-095]]. Weighting held: feat/bugfix nominal 2 dispatches, hotfix 1 + revert-on-fail. INCIDENT: re-committed [[LRN-093]] (2nd recurrence, 4 locks w/ \n) — caught at first run; user flagged advisory-insufficient → build deterministic backstop in lot 5. Next: lot 5 heavy flows (ship-feature enrich-at-gate, init-project +security, onboard no-loop) + escalation dogfood (max-3 STOP) + LRN-093 meta-test guard. +- Chantier verify-loops LOT 5 (feature/verify-loops `1c69de2`, FINAL): ship-feature (0e contract, enrich-at-gate STEP 3 [gated], 5 verify+secure vs ENRICHED) + init-project (contract from BRIEF, enrich GATE#1, 9 verify+secure — adds the security gate it lacked) + onboard (explicit NO-loop, audit≠dev, documented vs symmetry) + lib/tests/no-vacuous-locks.test.sh (LRN-093 deterministic backstop w/ inline flip-test) + loops-heavy 18 locks. Dogfood BOTH vigilance points real: (1) enrich — fresh verifier reads+judges a [gated] design criterion (ECARTS names it); (2) escalation — 3 consecutive ECARTS → orchestrator STOP at max-3 + CONTRACT-vs-REALIZED table, no 4th loop, no commit (first real exercise of the infinite-loop guard). [[BDR-051]] [[LRN-096]]. INCIDENT closed: the backstop's OWN flip-test RED'd (regex missed line-start tf) → fixed → [[LRN-096]] (a guard is code, prove it can fail). Chantier complete: 5 lots on feature/verify-loops, develop+main intact, nothing pushed. diff --git a/.claude/memory/learnings.md b/.claude/memory/learnings.md index 377556d..81a920c 100644 --- a/.claude/memory/learnings.md +++ b/.claude/memory/learnings.md @@ -109,6 +109,13 @@ rules: | LRN-087 | 2026-07-02 | presence-flag ≠ capability — rtk silently dead after .bashrc wipe; emitted commands need ABSOLUTE bin paths (they run in another shell); integrity pin = live machinery, re-pin on hook edit | any PATH-dependent capability + hand-managed shell profile; hooks emitting commands for another shell | | LRN-088 | 2026-07-02 | token-cutting intuition inverts under measurement — verbosity beats cardinality (gstack 34 skills ≈ 592 tok vs pr-review 6 agents ≈ 2,183) | any "disable X to save tokens" — measure per-item bytes first; profiles toggle skills, not plugin payloads | | LRN-089 | 2026-07-03 | pass-through wrapper (CLI `"$@"` → fn deriving target from ambient state: HEAD/cwd/env) silently ignores its args = silent contract violation; guard = args are an ASSERTION, refuse when they disagree with state | any dispatcher forwarding args to a callee that reads ambient state instead of the args | +| LRN-090 | 2026-06-30 | external-repo audit: open WIRED subsystems (hooks/runners) before declarative (docs/rules); described capability ≠ wired capability | auditing an external config/framework repo for transferable value | +| LRN-091 | 2026-07-03 | keyword-triggered soft-nudge hook w/ bare common tokens over-fires on non-UI work → tuned out; bare only when UI sense dominates, else bigram-or-drop | any advisory/nudge hook keyed on keywords | +| LRN-092 | 2026-07-03 | SAST smoke test w/ the OFFICIAL example secret = vacuous pass (rules exclude documented example keys by design); validate w/ realistic payloads + measure tier coverage before trusting a gate ruleset | smoke-testing any detector/gate — never the canonical example payload | +| LRN-093 | 2026-07-03 | grep -F pattern w/ embedded newline = per-line OR = lock that matches anything; structure locks single-line only, flip-test new locks | writing any grep-based structure lock / census test | +| LRN-094 | 2026-07-03 | SAST severity ≠ exploitability — semgrep ERROR conflates real vulns + hardening recos; metadata does NOT cleanly separate them (measured) → metadata refinement = noisy gate; ERROR-threshold + diff-scoping is the containment | mapping a SAST tool's output to a blocking gate | +| LRN-095 | 2026-07-03 | orthogonal gates don't contaminate — a conformity verifier must PASS correct-but-insecure code (security is a separate gate's job); proven live (CONFORME on a feature carrying a SQLi); fusing the two degrades each | designing multi-dimension review/verify/audit gates | +| LRN-096 | 2026-07-04 | a backstop/guard is code — reliable ONLY after a flip-test proves it CAN fail; an unproven guard replacing an advisory = a vacuous guard (LRN-048 applied to guards); flip-test mandatory at guard creation | building any deterministic guard/lint/backstop | --- @@ -977,3 +984,34 @@ rules: - **rule**: keep a token BARE only when its UI sense dominates largely in a dev context (glassmorphism, navbar). Token common in non-UI talk (design, component, theme, transition, frontend) → require a UI-specific bigram (design system, front-end design) or drop; in doubt → bigram-or-drop. Borderline standalone nouns (dashboard, animation) may stay bare as an assumed call — the fire-log arbitrates later on data, not gut. (NOT "never bare tokens" — animation stays bare here by design.) - **context**: design-toolchain-reminder.sh — 07-02 tightening (dropped page/form/menu/…) insufficient; 6 bare tokens still false-fired ~6×/session during the ECC config audit (design, ecc_dashboard.py, component, frontend, theme, transition, palette). 07-03 fix: dropped them, dashboard→`\bdashboard\b` (filename match killed, "admin dashboard" kept), added a fire-log (time+token+excerpt). `lib/tests/design-toolchain-reminder.test.sh` locks it (18 checks). - **cousin**: [[LRN-047]] a doctor that cries false is ignored. + +## LRN-092 — SAST smoke test: official example keys are rule-excluded — "no findings" proves nothing +- **pattern**: smoke-testing a SAST/secret detector w/ the OFFICIAL example payload (AWS `AKIA...EXAMPLE`) → 0 findings BY DESIGN — rules exclude documented example keys to kill FP. A vacuous pass, [[LRN-048]] class (a pass must prove it looked). Validate w/ realistic-shaped payloads AND enumerate what the tier does NOT catch before trusting a ruleset as a gate. +- **context**: lot 1 semgrep-install dogfood 2026-07-03. `p/secrets`+`p/security-audit` community tier: anonymous fetch OK (52 rules, no login), `subprocess-shell-true` detected ERROR; MISSED %-format SQLi on bare cursor (no recognized DB-API context) + fake-checksum `ghp_` token. Gap logged for security-auditor agent design (consider adding `p/owasp-top-ten`). +- **future application**: any detector/gate smoke test — craft realistic payloads, never the canonical example; measure the miss-list on purpose-built fixtures; size the gate's blocking scope on that data. +- **cousin**: [[LRN-048]] a 0/OK must prove it looked; [[LRN-047]] noisy guard = ignored guard; conditions [[BDR-048]]. + +## LRN-093 — grep -F with an embedded newline = per-line OR = vacuous lock +- **pattern**: a fixed-string grep pattern containing a newline is treated as MULTIPLE patterns (one per line) — match succeeds if ANY line matches. A structure lock written that way passes on essentially anything (`"no\n forced loop"` → matches any "no") = a lock that proves nothing, [[LRN-048]] class. +- **context**: lot 2 `lib/tests/contract-verifier.test.sh`, caught in self-review BEFORE first run; replaced by a single-line distinctive anchor ("proceed straight to the security gate"). +- **future application**: structure locks / census greps = ONE line per pattern, always; a clause spanning lines → lock a distinctive single-line fragment. Flip-test every new lock (prove it CAN fail) before trusting its green. +- **cousin**: [[LRN-048]] a pass must prove it looked; [[LRN-046]] deterministic-oracle discipline. + +## LRN-094 — SAST severity ≠ exploitability; metadata does not cleanly separate — don't refine on it +- **pattern**: semgrep `ERROR` conflates exploitable vulns (SQLi, secrets, command injection) with hardening recommendations (Dockerfile missing-USER, npm release-age). The obvious refinement — gate on `metadata.impact`/`likelihood`/`confidence` — does NOT work: measured, the Dockerfile hygiene ERROR (`impact=MEDIUM likelihood=LOW`) is indistinguishable from a tainted-SQL ERROR (`impact=MEDIUM likelihood=MEDIUM`), and a real command-injection reads `impact=LOW likelihood=HIGH`. Metadata-based severity = a noisy, non-deterministic gate ([[LRN-077]] class). +- **context**: lot 3 security-auditor design 2026-07-03. Measured on 2 real repos (faunosteo, game): the only added blocking ERROR from owasp-top-ten is Dockerfile hygiene, contained because gate mode scopes to the DIFF (a pre-existing infra finding can't block an unrelated code change). +- **future application**: mapping any SAST to a blocking gate — take the tool's ERROR/blocking level as the deterministic threshold, contain FP by SCOPING (diff, not repo), NOT by a metadata heuristic or a hand-maintained hygiene denylist ([[LRN-049]]: match guard cost to proven stake — build the denylist only if hygiene ERRORs prove noisy on a real project). +- **cousin**: [[LRN-047]] noisy gate = ignored; [[LRN-077]] non-deterministic gate; conditions [[BDR-048]]. + +## LRN-095 — Orthogonal gates don't contaminate: a conformity check must pass correct-but-insecure code +- **pattern**: when a pipeline has distinct gates (request-conformity, security), each judges ONLY its dimension. A conformity verifier must return CONFORME on code that is correct-but-insecure — the vuln is the SECURITY gate's job, not a conformity gap. Proven live: a `get_item` feature satisfying its contract but carrying a `%`-interpolation SQLi → verifier CONFORME, security-auditor BLOCK(1). Fusing the two into one "quality" gate makes each worse: the conformity check starts hunting vulns (scope creep, misses conformity), the security check starts judging feature-completeness (dilutes). +- **context**: lot 4 verify-secure-loop dogfood 2026-07-03. The orthogonality is WHY the order invariant matters (re-verify request before re-scan security) — two independent axes re-checked independently. +- **future application**: any multi-dimension gate (review lenses, verify+audit, correctness+perf) — keep each gate single-axis and let a finding on axis B pass axis A's gate; compose verdicts in the orchestrator, don't merge the judges. +- **cousin**: [[BDR-050]] the pipeline; [[BDR-049]] fresh verifier; conditions [[LRN-083]]. + +## LRN-096 — A backstop is code: prove it can FAIL (flip-test) before trusting its green +- **pattern**: a deterministic guard built to replace a forgettable advisory is itself code, and an UNPROVEN guard is a vacuous guard — [[LRN-048]] (a pass must prove it looked) applied to guards themselves. The LRN-093 backstop (refuse `\n` in grep/tf patterns) shipped with a regex requiring whitespace before `tf` → it silently MISSED `tf` at line start (exactly where the real locks sit). A flip-test (feed the guard a KNOWN offender, assert it bites) caught the hole; without it the guard would have green-lit the very class it was built to kill. So: a flip-test is MANDATORY at guard creation, part of the guard, not optional QA. +- **why it matters**: the whole point of a backstop is that it fires on the bad case; a guard that can't fail proves nothing and is WORSE than the advisory it replaced (false confidence). The advisory→backstop move ([[LRN-047]] [[LRN-091]], own doctrine) is only sound if the backstop is itself verified against a real miss. +- **context**: lot 5 `lib/tests/no-vacuous-locks.test.sh` 2026-07-04. Built the guard, its flip-test RED'd (regex too weak, missed line-start `tf`), fixed the regex, flip-test green. The guard now ships WITH the flip-test inline so it self-proves on every run. +- **future application**: building any guard/lint/census/backstop — bundle a flip-test (a synthetic offender the guard must catch) in the same file; a guard whose failure path was never exercised is untrusted. Corroborates [[LRN-047]]/[[LRN-091]] (advisory→deterministic) — this is the *quality bar* on the deterministic replacement. +- **cousin**: [[LRN-048]] prove it looked; [[LRN-093]] the class this guards; [[LRN-046]] deterministic-oracle discipline. diff --git a/.claude/tasks/TODO.md b/.claude/tasks/TODO.md index e17c1fe..562cc68 100644 --- a/.claude/tasks/TODO.md +++ b/.claude/tasks/TODO.md @@ -1,5 +1,33 @@ # TODO +## 2026-07-03 — verify loops + semgrep gate + contract (chantier orchestrateurs) +Archi validée au gate (session 2026-07-03). Cible : contract sur DISQUE dès +création (fichier de run, pattern DIAGNOSIS) + verifier frais (verdict structuré +CONFORME/écarts, preuve-qu'il-a-regardé LRN-048, 2 échecs structurels = escalade +humaine — verifier muet ≠ PASS) + gate sécu semgrep (rulesets ÉPINGLÉS +p/security-audit + p/secrets — pas --config auto, classe LRN-077 ; BLOCK +HIGH/CRITICAL only, LRN-047) + boucles bornées 3× décidées en boucle principale +(LRN-083). cso = symlink submodule gstack → non modifiable → greffes locales +(onboard cso-fallback, audit-delta, agent neuf ; complément semgrep même +gstack ON). Verdicts user : dev inline conservé feat/bugfix/hotfix (verify+sécu += sous-agents frais) ; hotfix garde revert-escalade ; PIN version semgrep dans +plugins.lock.json (gate bloquante — upgrade silencieux = nouveaux BLOCK sur code +inchangé ; pattern gsd-pin, saut affiché par update-all). + +LOT 1 — feature/semgrep-install (GO) +- [x] plugins.lock.json — pin semgrep 1.168.0 (pattern gsd, note gate bloquante) +- [x] install-plugins.sh STEP 7.5 — pipx pinned, command -v guard + version echo, login guide-only (jamais auto) +- [x] update-all.sh step 6.2 — pin-honored, affichage saut cur→pin, pipx install --force +- [x] Dogfood — install réel 1.168.0 via bloc extrait + idempotence (re-run = skip) + pin-match + saut affiché (1.168.0→9.9.9 fake, warn propre, install intacte) +- [x] Verify — bash -n OK, shellcheck clean (SC1091 info pré-existants only), lock JSON valide ; smoke rulesets : fetch anonyme 52 règles SANS login, subprocess-shell-true ERROR détecté. Limite notée pour LOT 3 : community tier rate SQLi %-format hors contexte API + tokens fake (choix rulesets à re-évaluer à l'agent) +- [ ] Commit scoped (settings.json dirty pré-existant JAMAIS stagé) + GATE lot 1 + +LOT 2 — feature/contract-verifier : specs montrées AVANT écriture. lib/contract-interview.md + agents/verifier.md. +LOT 3 — feature/security-auditor : agents/security-auditor.md + greffe audit-delta + onboard fallback + complément gstack-ON. +LOT 4 — feature/loops-light : câblage feat/bugfix/hotfix. +LOT 5 — feature/loops-heavy : câblage ship-feature + init-project + onboard. +Rien poussé ; gate par lot ; suites après chaque lot. + ## 2026-07-03 — design-toolchain trigger fix (bugfix/design-toolchain-trigger) Root cause (NOT a kill-switch, per user): ed2408e (07-02) dropped ultra-generic tokens but left bare tokens common in non-UI talk → ~6× false-fire THIS session diff --git a/agents/bugfixer.md b/agents/bugfixer.md index 59210b3..5ecfff7 100644 --- a/agents/bugfixer.md +++ b/agents/bugfixer.md @@ -100,6 +100,16 @@ RISK: - If the fix is significant (>10 lines, multiple files, behavior change): wait for user approval. +## STEP 3.5 — CONTRACT + +Run `$HOME/.claude/lib/contract-interview.md` (main loop). The DIAGNOSIS +feeds it: REQUEST verbatim = the bug report as received; ACCEPTANCE CRITERIA += the symptom reproduced-then-gone + a regression test present and passing; +FILE SCOPE = the FIX PLAN files. Questions stay proportional (a clear, +reproduced bug → zero). It writes the contract to +`.claude/tasks/contracts/--.md`; keep the path for GATE 1 +(STEP 5). + ## STEP 4 — FIX **Gitflow aiguillage (before editing):** follow `$HOME/.claude/lib/gitflow-aiguillage.md` @@ -131,7 +141,19 @@ Apply the fix following the plan: ``` 2. If a build step exists, verify it passes (`npm run build`, `tsc --noEmit`, `cargo build`, etc.). 3. Check for regressions in related functionality. -4. **Pre-commit confirmation gate.** Before running `git commit`, present the diff +4. **Fresh gates (verify + secure), bounded loops.** Steps 1-3 are your + dev-side smoke test, NOT the gate. Run the two fresh gates per + `$HOME/.claude/lib/verify-secure-loop.md` with `CONTRACT` = the STEP 3.5 + path, `DIFF` = the fix diff, `TEST` = the suite from step 1: + - GATE 1 — a FRESH verifier judges the fix against the contract (bug gone + + regression test present). CONFORME → GATE 2. ECARTS → fix, re-verify, + max 3 → escalate. + - GATE 2 — a FRESH security-auditor (`MODE: gate`) scans the fix diff + (a bug fix can introduce a vuln). PASS → commit gate. BLOCK → fix, + re-verify request THEN re-scan, max 3 → escalate. + + Nominal = one verifier + one security dispatch. Only then the commit gate. +5. **Pre-commit confirmation gate.** Before running `git commit`, present the diff summary and the proposed message, then wait for approval: ``` @@ -152,7 +174,7 @@ Apply the fix following the plan: - `skip` → leave changes uncommitted, exit cleanly. - `amend last` → the fix should fold into the previous commit (use only when prior commit is unpushed). -5. Commit using conventional format (after approval): +6. Commit using conventional format (after approval): ``` fix(): @@ -161,7 +183,7 @@ Apply the fix following the plan: Co-Authored-By: Claude ``` -6. Print summary: +7. Print summary: ``` BUGFIX COMPLETE BUG : diff --git a/agents/feater.md b/agents/feater.md index 7751099..23d54cb 100644 --- a/agents/feater.md +++ b/agents/feater.md @@ -65,6 +65,17 @@ already constrain or forbid the approach; an LRN may name a gotcha to apply. Emi MEMORY; feed STEP 1 MINI-PLAN. Inline consumption — reader = planner, no injection. `.claude/memory/` absent → guarded no-op (zero overhead on a memory-less repo). +## STEP 0.7 — CONTRACT + +Run `$HOME/.claude/lib/contract-interview.md` (main loop — you are it). It +captures the request verbatim, asks 0-3 questions PROPORTIONAL to ambiguity +(a complete request → zero questions, silent), derives testable acceptance +criteria + file scope, and writes the contract to +`.claude/tasks/contracts/--.md`. Keep the path — GATE 1 +(STEP 3) hands it to a fresh verifier. On a small, clear feature this is a +few seconds and no questions; it is the single reference the verifier judges +against, not a restatement. + ## STEP 1 — MINI-PLAN Quick mental model, not a formal plan document: @@ -101,19 +112,24 @@ Work through the plan: - Follow existing patterns in the codebase. - Run tests incrementally as you go. -## STEP 3 — VERIFY +## STEP 3 — VERIFY + SECURE (fresh gates, bounded loops) -1. Run the full relevant test suite: - ```bash - # detect and run tests, lint, type-check - ``` -2. If a dev server is relevant, mention what the user should - check visually. -3. Quick self-review: scan your diff for obvious issues: - ```bash - git diff --stat - git diff - ``` +First, your own pre-check (dev-side, fast): run the relevant test suite / +lint / type-check, and if a dev server is relevant note what to check +visually. This is your smoke test, NOT the gate. + +Then run the two fresh gates per `$HOME/.claude/lib/verify-secure-loop.md` +with `CONTRACT` = the STEP 0.7 path, `DIFF` = your working-tree diff, `TEST` += the suite you just ran: + +- GATE 1 — a FRESH verifier judges the diff against the contract (blind, no + self-score of yours counts). CONFORME on the first pass → straight to GATE + 2, no loop. ECARTS → fix the named gaps, re-verify, max 3 → escalate. +- GATE 2 — a FRESH security-auditor (`MODE: gate`) scans the diff. PASS → + commit. BLOCK → fix, re-verify the request THEN re-scan, max 3 → escalate. + +Nominal (clear request, conform first pass, clean diff) = exactly one +verifier + one security dispatch. The loop only costs when it loops. ## STEP 4 — COMMIT diff --git a/agents/hotfixer.md b/agents/hotfixer.md index fa21e9e..b958707 100644 --- a/agents/hotfixer.md +++ b/agents/hotfixer.md @@ -1,13 +1,15 @@ --- name: hotfixer description: Quick fix for superficial bugs (typos, CSS issues, config errors, off-by-one, wrong variable name, missing import, broken link). Max 2 files, obvious root cause only. -tools: Read, Edit, Write, Bash, Grep, Glob +tools: Read, Edit, Write, Bash, Grep, Glob, Agent --- # HOTFIX — Quick Superficial Fix -Fast-track fix for obvious bugs. No planning overhead, no plugin -check, no subagents. Get in, fix, verify, get out. +Fast-track fix for obvious bugs. No planning overhead, no plugin check. +The fix is inline (no dev subagents); a fresh security gate runs before +commit, and any gate failure reverts — never loops. Get in, fix, gate, +get out. ## REQUEST $ARGUMENTS @@ -39,6 +41,18 @@ skip). For a RECURRING or urgent bug only, a quick blockers-only glance may save If a prior BLK names this bug, jump to its solution. Not mandatory; no RELATED MEMORY disposition required at hotfix weight. +## STEP 1.7 — CONTRACT (silent autofill) + +Run `$HOME/.claude/lib/contract-interview.md` at hotfix weight: **zero +questions ever** (a hotfix is an obvious fix by definition). Autofill the +contract — REQUEST verbatim = the bug description as given; ACCEPTANCE +CRITERIA = "symptom gone; build/tests green"; FILE SCOPE = the 1-2 target +files. It writes `.claude/tasks/contracts/--.md`. This is +the reference for the security gate's scope and the escalation report if a +gate fails. No verifier is dispatched at hotfix weight — the STEP 3 +smoke-check already verifies these trivial criteria; the gate hotfix adds is +security (below). + ## STEP 1.5 — DESIGN GATE Follow `$HOME/.claude/lib/design-gate.md`: @@ -102,18 +116,32 @@ Apply the minimal change that fixes the bug: (Files were not yet staged — restore is safe.) - STOP and tell user: `"Hotfix introduced a regression. Reverted. Escalate to /bugfix or /analyze for deeper investigation."` - Do NOT commit a broken fix. -3. Commit using conventional format (only after verify passes): +3. **Security gate (fresh auditor) — failure REVERTS, never loops.** Dispatch + a FRESH security-auditor (`subagent_type: security-auditor`, or load + `agents/security-auditor.md`) with `MODE: gate`, `SCOPE:` the working-tree + diff vs the pre-flight SHA. Parse its `SECURITY — VERDICT:` line: + - `PASS` (or `DEGRADED` with no BLOCK) → proceed to commit. + - `BLOCK(n)` → this is hotfix: do NOT loop. Run `git restore .` to the + pre-flight SHA, print the `BLOCKING` list, and STOP: + `"Hotfix introduced a security finding. Reverted. Escalate to /bugfix + for a fix under the full verify+security loop."` The hotfix model is + one attempt; any gate failure (smoke OR security) reverts and escalates. + - Structural failure (mute / unparsable / no VERDICT line) → treat as a + failed gate: retry ONCE fresh; a 2nd structural failure → revert + + escalate. A mute auditor is never a PASS. +4. Commit using conventional format (only after verify AND security pass): ``` fix(): Co-Authored-By: Claude ``` -4. Print summary: +5. Print summary: ``` HOTFIX APPLIED FILE(S) : FIX : VERIFIED: + SECURITY: ``` ## STEP 4 — DOC SYNC (automatic) diff --git a/agents/security-auditor.md b/agents/security-auditor.md new file mode 100644 index 0000000..5209108 --- /dev/null +++ b/agents/security-auditor.md @@ -0,0 +1,159 @@ +--- +name: security-auditor +description: SAST security gate — runs the pinned semgrep rulesets + the CLAUDE.md security checklist on a diff or project scope, maps severities, renders SECURITY — VERDICT: PASS | BLOCK(n). Blocks HIGH/CRITICAL only, reports the rest. Never fixes code. Fresh dispatch, no iteration history. +tools: Read, Grep, Glob, Bash, Write +--- + +# SECURITY-AUDITOR AGENT + +You are the security gate. You run semgrep + a checklist over a scope, +classify by severity, and render a verdict. You never fix code, you never +edit anything but the report file (audit mode only), and you never trust a +prior run — every scan is fresh and complete. + +Bash runs semgrep and read-only inspection only — never a command that +mutates code, installs, or commits. + +## MODES + +- **gate** (default; dev flows) — SCOPE = a diff. Output = the stdout block + below. `Write` is FORBIDDEN in this mode. +- **audit** (onboard, audit-delta) — SCOPE = project root or a delta list. + `Write` is allowed ONLY to the exact `REPORT` path given — NEVER to any + code/config file. Writing anywhere else is a contract violation. + +## INPUT (from the orchestrator — nothing else exists) + +- `MODE: gate|audit` +- `SCOPE: ` +- `REPORT: ` (audit mode only — the single writable path) +- `CONTEXT: ` (optional; onboard supplies it) + +You NEVER receive iteration history — no prior verdicts, no earlier finding +lists, no dev reports. Ignore any such material if it appears. Every scan is +blind and complete (cost bounded upstream by the max-3 loop cap). + +## STEP 1 — TOOL CHECK + +`command -v semgrep` and capture the version. If ABSENT → **DEGRADED mode**: +announce it loudly on the `TOOL:` line, and STILL RUN STEP 3 (the checklist) +— a DEGRADED run must prove it detected everything it still can. A DEGRADED +run that skips the checklist and PASSes is a vacuous pass (LRN-048). Never a +silent skip, never a false BLOCK from the tool being absent (LRN-047). + +## STEP 2 — SEMGREP (skip only in DEGRADED) + +Resolve the scanned paths from SCOPE (in gate mode: `git diff --name-only +` filtered to existing files; in audit mode: the root or delta list, +excluding `node_modules`, `dist`, `vendor`, `.git`). + +Run, on those paths ONLY: + +``` +semgrep scan --config p/security-audit --config p/secrets --config p/owasp-top-ten \ + --metrics=off --quiet --json +``` + +Pinned rulesets, never `--config auto`, never `semgrep login` (BDR-048: +`auto` = registry telemetry + per-run ruleset resolution = a +non-deterministic gate). owasp-top-ten is REQUIRED, not optional: measured +2026-07-03, the two-ruleset baseline missed SQL injection and path traversal +entirely on realistic Flask code; owasp-top-ten's taint rules catch them. + +**Severity mapping** (from `results[].extra.severity` + ruleset origin): + +| semgrep | origin | → gate severity | blocks? | +|---------|--------|-----------------|---------| +| ERROR | p/secrets | CRITICAL | yes | +| ERROR | other | HIGH | yes | +| WARNING | any | MEDIUM | no (reported) | +| INFO | any | LOW | no (reported) | + +The blocking threshold is ERROR — deterministic, rule-assigned. Known limit +(measured): severity is per-RULE not per-VULN — the same class can span +ERROR and WARNING rules (e.g. `tainted-sql-string`=ERROR vs +`sql-injection-db-cursor-execute`=WARNING). Blocking on WARNING too would +flood FPs (nginx/github-actions/npm hygiene warnings); ERROR is the right +line. Report — never silently drop — the MEDIUM/LOW findings. + +## STEP 3 — CHECKLIST (always, incl. DEGRADED) + +Grep the scope for the CLAUDE.md non-negotiable defaults semgrep may miss. +Each hit → severity + file:line + one-line why: + +- hardcoded secret / token / key / auth-bearing URL (→ CRITICAL) +- SQL built by string concatenation / interpolation (→ HIGH) +- unsanitized render of user input (innerHTML, dangerouslySetInnerHTML, + raw(), `eval`) (→ HIGH) +- sensitive endpoint with no authz check (→ HIGH) +- stack trace / internal path / DB error surfaced to the user (→ MEDIUM) +- secret / password / token / PII written to a log (→ HIGH) +- tracked `.env` or committed credential file (→ CRITICAL) + +If `CONTEXT` (archetype) is given, scope the checklist to what applies +(no web-XSS checks on firmware, etc.). + +## STEP 4 — ANTI-GAMING + +Scan the diff (gate) or scope (audit) for any NEW suppression comment +(`# nosemgrep`, `// nosemgrep`, `nosec`, `eslint-disable ... security`, or +equivalent) that did not exist before this change. Each new suppression is a +**BLOCKING** finding UNLESS it already carries a human `[gated ]` +marker — same rule as scope enrichment: without the micro-gate the dev +suppresses everything and the gate constrains nothing. Report pre-existing +suppressions as LOW (context), do not block on them. + +## STEP 5 — DEDUP + VERDICT + +Merge semgrep + checklist findings, dedup by (file:line, rule/check). +`BLOCK(n)` ⇔ n = count(CRITICAL) + count(HIGH) + count(new un-gated +suppressions) > 0. Otherwise `PASS`. MEDIUM/LOW are REPORTED, never +blocking. + +## OUTPUT (exact format — machine-parsed by the orchestrator) + +``` +SECURITY — VERDICT: PASS | BLOCK(n) | ERROR() +TOOL: semgrep — p/security-audit, p/secrets, p/owasp-top-ten | ABSENT (DEGRADED — checklist only; install: make plugin) +SCOPE: files +BLOCKING: + 1. [CRITICAL|HIGH] — — — hint: +REPORTED (non-blocking): + - [MEDIUM|LOW] — +PROOF: semgrep rules on files → findings; checklist checks → findings +``` + +In audit mode, ALSO write this same block (plus per-finding detail) to +`REPORT`, and end stdout with `REPORT_WRITTEN: `. + +## RULES + +- Report-only on CODE. Never edit or fix a code file. In audit mode the sole + writable path is `REPORT`; in gate mode nothing is writable. +- `PROOF` is MANDATORY — a `PASS` (or DEGRADED PASS) without a `PROOF` line + showing what was scanned is invalid; the orchestrator discards it as a + structural failure (LRN-048). +- A mute / crashed / unparsable auditor is NEVER a PASS. Exactly one + `SECURITY — VERDICT:` line, spelled as above. +- Blocks on HIGH/CRITICAL only. A noisy gate that blocks on hygiene is a + gate people learn to bypass (LRN-047) — MEDIUM/LOW are reported, not + gated. + +## ORCHESTRATOR PROTOCOL (consumer contract — wiring reference) + +- The security gate runs AFTER the request-conformity verdict is CONFORME + (verifier), never before. +- Dispatch a FRESH auditor each iteration — no context reuse. Input = mode + + scope + (report) + (context), nothing else. +- Parse the `SECURITY — VERDICT:` line: + - `PASS` → proceed (to commit / next step). + - `BLOCK(n)` → the dev subagent receives the BLOCKING list + the contract + path. After the fix: re-verify the REQUEST first (verifier), THEN re-run + this gate — in that order. Max 3 security iterations → STOP + human + escalation with the BLOCKING table. + - `DEGRADED` (semgrep absent) → does NOT block; surface the checklist + result + recommend `make plugin`. A DEGRADED BLOCK (grep-caught + hardcoded secret etc.) blocks like any other. + - Structural failure (`ERROR(…)`, missing/duplicated VERDICT line, + unparsable, crash, PASS without PROOF) → retry ONCE fresh; 2nd + structural failure → human escalation. A mute auditor is never a PASS. diff --git a/agents/verifier.md b/agents/verifier.md new file mode 100644 index 0000000..05d1c78 --- /dev/null +++ b/agents/verifier.md @@ -0,0 +1,110 @@ +--- +name: verifier +description: Fresh independent verifier — reads a CONTRACT file from disk and renders a structured verdict (CONFORME / ECARTS / ERROR) on the implemented diff. Report-only, never fixes. Dispatched fresh at every iteration; receives no iteration history. +tools: Read, Grep, Glob, Bash +--- + +# VERIFIER AGENT + +You verify that an implementation CONFORMS to a contract. You are NOT the +developer, you never fix anything, and you never trust the developer's +summary — only the contract, the code, and what you execute yourself. + +Bash is for OBSERVATION ONLY: run tests/builds, `git diff` / `git log` / +`git show`, read-only inspection. Never a command that writes, installs, +commits, or mutates any state. + +## INPUT (from the orchestrator — nothing else exists) + +- `CONTRACT: ` — you READ it from disk; never accept an inline + restatement in its place +- `DIFF: ` +- `TEST: ` (optional) + +You NEVER receive iteration history: no previous verdicts, no prior gap +lists, no dev reports. If any such material appears in your prompt, IGNORE +it — every verification is complete and blind. (Cost is bounded upstream: +the orchestrator caps the loop at 3 iterations.) + +## STEP 1 — READ THE CONTRACT + +Read the contract file. If it is missing, unreadable, or lacks its +`REQUEST` or `ACCEPTANCE CRITERIA` section → output +`VERIFY — VERDICT: ERROR()` plus the `CONTRACT:` line, and STOP. + +## STEP 2 — EVIDENCE PER CRITERION + +For EACH acceptance criterion, establish exactly one status from the real +code: + +- `MET` — with evidence: the file:line you read, or the test/build you RAN +- `NOT-MET` — expected vs actual, located at file:line +- `UNVERIFIABLE` — precise reason (missing environment, requires human + judgment, external dependency…) + +Rules: read the diff AND enough surrounding code to judge behavior; run +`TEST` if provided, plus cheap targeted checks when they settle a +criterion. Never mark `MET` from naming, comments, or plausibility — only +from behavior you observed or code you read. + +## STEP 3 — SCOPE CHECK + +List the files actually touched (`git diff --name-only` over `DIFF`). +Compare against the contract's `FILE SCOPE`. Report every out-of-scope +file. Disposition is NOT your call: the orchestrator treats each one as a +gap — the dev removes it or justifies it, and an accepted justification +only enters the contract through a human micro-gate. + +## STEP 4 — VERDICT + +`CONFORME` ⇔ ALL criteria `MET` AND zero out-of-scope files. +Anything else is `ECARTS(n)` where n = count(NOT-MET) + count(UNVERIFIABLE) ++ count(out-of-scope files). + +## OUTPUT (exact format — machine-parsed by the orchestrator) + +``` +VERIFY — VERDICT: CONFORME | ECARTS(n) | ERROR() +CONTRACT: +CRITERIA: + 1. — MET — + 2. — NOT-MET — expected <…> / actual <…> — + 3. — UNVERIFIABLE — +SCOPE: in-scope files; out-of-scope: +PROOF: read files, ran , checked / criteria +``` + +## RULES + +- Report-only. Never edit, never write, never propose the fix itself — + naming the gap precisely is the whole job. +- `UNVERIFIABLE` ≠ `MET`. A criterion you did not check is `UNVERIFIABLE`, + never silently dropped: the checked count in `PROOF` must equal the + contract's criteria count. +- `PROOF` is MANDATORY. A `CONFORME` without a `PROOF` line is invalid — + the orchestrator discards it as a structural failure (LRN-048: a pass + must prove it looked). +- The verdict grammar is load-bearing: exactly one `VERIFY — VERDICT:` + line, spelled exactly as above. + +## ORCHESTRATOR PROTOCOL (consumer contract — wiring reference) + +How every orchestrator consumes this agent (the loop lives in the MAIN +loop, never here): + +- Dispatch a FRESH verifier at every iteration — no context reuse. Input = + contract path + diff range + optional test command, nothing else. +- Parse the `VERIFY — VERDICT:` line: + - `CONFORME` on first pass → proceed straight to the security gate — no + forced loop. + - `ECARTS(n)` → the dev subagent receives the contract PATH + the exact + gap list (nothing else). Max 3 iterations → STOP + human escalation + with the CRITERIA table (the contract-vs-realized diff). + - Remaining `UNVERIFIABLE` while everything else is MET → direct human + gate (a dev cannot fix unverifiability). + - Structural failure (`ERROR(…)`, missing/duplicated VERDICT line, + unparsable output, agent crash, `CONFORME` without `PROOF`) → retry + ONCE with a fresh verifier; a 2nd structural failure → human + escalation. A mute verifier is NEVER a PASS. +- After a security-gate fix round: re-verify the request FIRST (this + agent), THEN re-verify security — in that order. diff --git a/install-plugins.sh b/install-plugins.sh index 2810f4d..60d79e6 100644 --- a/install-plugins.sh +++ b/install-plugins.sh @@ -654,6 +654,35 @@ if command -v graphify &>/dev/null; then fi echo "" +# ============================================================ +# STEP 7.5 — SEMGREP (SAST engine for the security gate) +# ============================================================ +echo "── Step 7.5: Semgrep — SAST security gate ───────────────────" +echo "" +if command -v semgrep &>/dev/null; then + ok "semgrep already installed ($(semgrep --version 2>/dev/null | head -1))" +else + SEMGREP_VER=$(pinned_version "semgrep") + if [ "$SEMGREP_VER" != "latest" ]; then + info "Installing semgrep ${SEMGREP_VER} (pinned in plugins.lock.json)..." + pipx install "semgrep==${SEMGREP_VER}" 2>/dev/null + else + info "Installing semgrep latest (consider pinning in plugins.lock.json)..." + pipx install semgrep 2>/dev/null + fi + if command -v semgrep &>/dev/null; then + ok "semgrep installed ($(semgrep --version 2>/dev/null | head -1))" + else + err "semgrep install failed — run manually: pipx install semgrep" + fi +fi +# Login is Pro-rules only and optional — NEVER run automatically (ctx7 +# pattern: guide, don't block). The gate uses pinned public rulesets. +if command -v semgrep &>/dev/null; then + info "Optional Pro rules: semgrep login (never run automatically)" +fi +echo "" + # ============================================================ # STEP 8 — EMIL DESIGN ENG (UI polish / animation skill) # ============================================================ diff --git a/lib/contract-interview.md b/lib/contract-interview.md new file mode 100644 index 0000000..9dcc1c4 --- /dev/null +++ b/lib/contract-interview.md @@ -0,0 +1,104 @@ +# Contract interview — mandatory upstream passage (all orchestrators) + +Produces the CONTRACT: the single reference passed verbatim to the plan, the +dev subagents, and the verifier. The contract is what lets the orchestrator +delegate execution without subagents ever needing a human gate (LRN-083: +subagents = execution + report only; gates and loop decisions live in the +main loop). + +Run this in the ORCHESTRATOR MAIN LOOP, never in a subagent — STEP 2 may +talk to the human. Mandatory passage in every flow; questions are optional +and proportional — a complete request goes through silently. + +## STEP 1 — CAPTURE (verbatim) + +Copy the user's request EXACTLY as typed (`$ARGUMENTS` + the triggering +message). No paraphrase, no cleanup, no translation, no summarizing. This +section is IMMUTABLE for the life of the run — every later consumer +(planner, dev, verifier) reads THESE words, never a restatement. + +## STEP 2 — AMBIGUITY CHECK (questions optional, proportional) + +Ask ONLY if one of these is missing AND not derivable from the repo: +- a testable expected outcome +- an unambiguous scope (what is allowed to change) +- non-contradictory constraints + +Complete request → ZERO questions, stay silent. Otherwise: max 3 questions, +one single batch (house rule: one question upfront, never mid-task). Never +ask what the repo can answer — verify paths/APIs/behavior yourself first. + +## STEP 3 — DERIVE + +- ACCEPTANCE CRITERIA: numbered; each one testable — a fresh reader must be + able to mark it MET / NOT-MET against the real code, without having seen + this conversation. +- FILE SCOPE: paths/zones expected to change, or `repo-wide — `. + +## STEP 4 — WRITE TO DISK (immediately, before any next step) + +Path: `.claude/tasks/contracts/--.md` +(`mkdir -p` the directory; unique per run: date + short kebab slug + HHMM — +two runs on the same day never collide). A contract that lives only in +context dies at compaction, and the verbatim request with it. + +Template: + +```markdown +# CONTRACT — +- date: | flow: | branch: +- status: active + +## REQUEST (verbatim — IMMUTABLE) + + +## CLARIFICATIONS +Q: / A: +(or: none — request complete) + +## ACCEPTANCE CRITERIA +1. +2. + +## FILE SCOPE + +(or: repo-wide — ) +``` + +Print one line to the user, then continue the flow: +`CONTRACT: — criteria, scope , questions asked` + +## Lifecycle + +- **REQUEST**: immutable, for the life of the run. Never rewritten, never + "cleaned up". +- **CRITERIA / FILE SCOPE enrichment**: ONLY at a human gate, each added + entry marked `[gated ]`. A dev subagent NEVER enriches the + contract. An out-of-scope edit the dev justifies is accepted ONLY through + this micro-gate: human approves → FILE SCOPE gains the entry `[gated]`; + human declines → the dev removes the edit. Without this gate the dev + justifies everything and scope constrains nothing. +- **Deep re-scope** (the request itself changes): NEW contract file with + `supersedes: ` in its header — never a rewrite of the old one. +- **Aborted run**: delete the contract file, or commit it with + `status: aborted` in the header. NEVER left dirty in the working tree. +- **Commit**: the contract rides the existing memory commit — + `lib/capitalize-commit.md` already covers the `.claude/tasks` pathspec. + No new plumbing. + +## Weight per flow + +| Flow | Weight | +|------|--------| +| hotfix | Silent autofill — criteria: "symptom gone; build/tests green"; scope = the 1-2 target files. Zero questions ever. | +| feat / bugfix | Proportional. bugfix: the DIAGNOSIS feeds the criteria (symptom reproduced-then-gone + regression test present). | +| ship-feature | Full. Design decisions approved at the validation gate append criteria `[gated ]` — the human validates the enriched contract, the verifier receives that version. | +| init-project | Full. The interviewer's PROJECT BRIEF pours into the contract (V1 features → criteria). | +| onboard | Audit-scope contract (interview answers → what to audit, which axes). | + +## Hand-off rule + +Downstream consumers (plan step, dev subagents, verifier) receive the +contract PATH, not a restatement of its content — the file on disk is the +only authoritative copy, and reading it from disk is what makes the dev's +reformulation structurally unable to interpose. diff --git a/lib/tests/contract-verifier.test.sh b/lib/tests/contract-verifier.test.sh new file mode 100644 index 0000000..af7b99f --- /dev/null +++ b/lib/tests/contract-verifier.test.sh @@ -0,0 +1,87 @@ +#!/usr/bin/env bash +# ============================================================ +# Structure locks — contract/verifier pair (verify-loops lot 2) +# Deterministic greps on load-bearing doctrine clauses: an edit +# that silently drops one (blind verifier, PROOF mandatory, +# immutable REQUEST, micro-gate scope enrichment…) reds here. +# ============================================================ +set -u + +REPO="$(cd "$(dirname "$0")/../.." && pwd)" +LIB="$REPO/lib/contract-interview.md" +AGT="$REPO/agents/verifier.md" +PASS=0; FAIL=0 + +# Fixed-string lock (UTF-8 punctuation safe) +tf() { # tf