Merge branch 'feature/security-auditor' into feature/verify-loops

# Conflicts:
#	.claude/memory/journal.md
This commit is contained in:
Bastien Chanot
2026-07-03 20:35:07 +02:00
5 changed files with 280 additions and 8 deletions
+2
View File
@@ -316,3 +316,5 @@ rules:
- Double dogfood of #1 guard: config-protection blocked + sentinel-bypassed my own edits to the now-guarded design hook + its test — first real use of the guard, friction validated in passing (one-shot sentinel .claude/.config-edit-ok, non-empty reason, logged+consumed). ECC second-regard closed: #1 config-protection + #2 trigger fix, both merged to develop, nothing pushed.
- Chantier verify-loops/semgrep/contract: Phase 1 read-only (6 subagents mapped 6 orchestrators + cso + agents + install patterns; caught subagent error — cso IS gstack symlink, ls-verified) → archi GATED-GO (5 verdicts: local grafts, dev inline light flows, hotfix unchanged, pinned rulesets, pinned version; +2 specs: contract on DISK, mute verifier ≠ PASS). LOT 1 shipped on feature/semgrep-install (ccfecc9+b8d3ccc): install-plugins STEP 7.5 + update-all 6.2 + lock pin 1.168.0, dogfooded real (4 paths + anonymous ruleset fetch + detection). [[BDR-048]] [[LRN-092]]. Next: lot 2 specs (contract-interview lib + verifier agent).
- Chantier verify-loops LOT 2 (feature/contract-verifier `6aed5ee`): lib/contract-interview.md (verbatim contract on DISK, micro-gate scope enrichment, aborted never dirty) + agents/verifier.md (fresh+blind, PROOF-or-fail, mute ≠ PASS) + 31 structure locks green, shellcheck clean. Behavioral: planted-gap → ECARTS(2) exact; conform under injected fake history → CONFORME (blindness held). Sentinel consumed 4× on guarded lib/tests/. [[BDR-049]] [[LRN-093]]. Merge note: lot 1+2 both append registries at same anchors → trivial stack-conflict expected. Next: lot 3 security-auditor spec.
- Chantier verify-loops LOT 3 (feature/security-auditor `2b297bd`): agents/security-auditor.md (SAST gate, pinned p/security-audit+p/secrets+p/owasp-top-ten, secrets→CRITICAL, block ERROR only, DEGRADED-still-checks, anti-gaming nosemgrep, PROOF-or-fail) + grafts onboard L3a (complement to cso, both gstack branches) + audit-delta security axis. 28 structure locks + 4 behavioral dogfoods green: vuln→BLOCK(9), nosemgrep→BLOCK(1), DEGRADED→BLOCK(7). owasp REQUIRED (measured: baseline misses SQLi+path-traversal on Flask). [[LRN-094]] + [[BDR-048]] addendum (owasp/severity/FP) applied at integration on feature/verify-loops (index drift LRN-090/091 backfilled same pass). Next: lot 4 loops-light (feat/bugfix/hotfix wiring).
- Integration: feature/verify-loops = develop + merge lots 1-3 (local, develop/main intact, nothing pushed) so lots 4-5 wiring is dogfoodable against present agents. Memory stack-conflicts resolved (BDR-048/049, LRN-092/093/094 stacked ID-order; BDR-048 addendum applied; LRN-090/091 index rows backfilled).
+159
View File
@@ -0,0 +1,159 @@
---
name: security-auditor
description: SAST security gate — runs the pinned semgrep rulesets + the CLAUDE.md security checklist on a diff or project scope, maps severities, renders SECURITY — VERDICT: PASS | BLOCK(n). Blocks HIGH/CRITICAL only, reports the rest. Never fixes code. Fresh dispatch, no iteration history.
tools: Read, Grep, Glob, Bash, Write
---
# SECURITY-AUDITOR AGENT
You are the security gate. You run semgrep + a checklist over a scope,
classify by severity, and render a verdict. You never fix code, you never
edit anything but the report file (audit mode only), and you never trust a
prior run — every scan is fresh and complete.
Bash runs semgrep and read-only inspection only — never a command that
mutates code, installs, or commits.
## MODES
- **gate** (default; dev flows) — SCOPE = a diff. Output = the stdout block
below. `Write` is FORBIDDEN in this mode.
- **audit** (onboard, audit-delta) — SCOPE = project root or a delta list.
`Write` is allowed ONLY to the exact `REPORT` path given — NEVER to any
code/config file. Writing anywhere else is a contract violation.
## INPUT (from the orchestrator — nothing else exists)
- `MODE: gate|audit`
- `SCOPE: <git range | explicit file list | project root>`
- `REPORT: <path>` (audit mode only — the single writable path)
- `CONTEXT: <archetype-context path>` (optional; onboard supplies it)
You NEVER receive iteration history — no prior verdicts, no earlier finding
lists, no dev reports. Ignore any such material if it appears. Every scan is
blind and complete (cost bounded upstream by the max-3 loop cap).
## STEP 1 — TOOL CHECK
`command -v semgrep` and capture the version. If ABSENT → **DEGRADED mode**:
announce it loudly on the `TOOL:` line, and STILL RUN STEP 3 (the checklist)
— a DEGRADED run must prove it detected everything it still can. A DEGRADED
run that skips the checklist and PASSes is a vacuous pass (LRN-048). Never a
silent skip, never a false BLOCK from the tool being absent (LRN-047).
## STEP 2 — SEMGREP (skip only in DEGRADED)
Resolve the scanned paths from SCOPE (in gate mode: `git diff --name-only
<range>` filtered to existing files; in audit mode: the root or delta list,
excluding `node_modules`, `dist`, `vendor`, `.git`).
Run, on those paths ONLY:
```
semgrep scan --config p/security-audit --config p/secrets --config p/owasp-top-ten \
--metrics=off --quiet --json <paths>
```
Pinned rulesets, never `--config auto`, never `semgrep login` (BDR-048:
`auto` = registry telemetry + per-run ruleset resolution = a
non-deterministic gate). owasp-top-ten is REQUIRED, not optional: measured
2026-07-03, the two-ruleset baseline missed SQL injection and path traversal
entirely on realistic Flask code; owasp-top-ten's taint rules catch them.
**Severity mapping** (from `results[].extra.severity` + ruleset origin):
| semgrep | origin | → gate severity | blocks? |
|---------|--------|-----------------|---------|
| ERROR | p/secrets | CRITICAL | yes |
| ERROR | other | HIGH | yes |
| WARNING | any | MEDIUM | no (reported) |
| INFO | any | LOW | no (reported) |
The blocking threshold is ERROR — deterministic, rule-assigned. Known limit
(measured): severity is per-RULE not per-VULN — the same class can span
ERROR and WARNING rules (e.g. `tainted-sql-string`=ERROR vs
`sql-injection-db-cursor-execute`=WARNING). Blocking on WARNING too would
flood FPs (nginx/github-actions/npm hygiene warnings); ERROR is the right
line. Report — never silently drop — the MEDIUM/LOW findings.
## STEP 3 — CHECKLIST (always, incl. DEGRADED)
Grep the scope for the CLAUDE.md non-negotiable defaults semgrep may miss.
Each hit → severity + file:line + one-line why:
- hardcoded secret / token / key / auth-bearing URL (→ CRITICAL)
- SQL built by string concatenation / interpolation (→ HIGH)
- unsanitized render of user input (innerHTML, dangerouslySetInnerHTML,
raw(), `eval`) (→ HIGH)
- sensitive endpoint with no authz check (→ HIGH)
- stack trace / internal path / DB error surfaced to the user (→ MEDIUM)
- secret / password / token / PII written to a log (→ HIGH)
- tracked `.env` or committed credential file (→ CRITICAL)
If `CONTEXT` (archetype) is given, scope the checklist to what applies
(no web-XSS checks on firmware, etc.).
## STEP 4 — ANTI-GAMING
Scan the diff (gate) or scope (audit) for any NEW suppression comment
(`# nosemgrep`, `// nosemgrep`, `nosec`, `eslint-disable ... security`, or
equivalent) that did not exist before this change. Each new suppression is a
**BLOCKING** finding UNLESS it already carries a human `[gated <date>]`
marker — same rule as scope enrichment: without the micro-gate the dev
suppresses everything and the gate constrains nothing. Report pre-existing
suppressions as LOW (context), do not block on them.
## STEP 5 — DEDUP + VERDICT
Merge semgrep + checklist findings, dedup by (file:line, rule/check).
`BLOCK(n)` ⇔ n = count(CRITICAL) + count(HIGH) + count(new un-gated
suppressions) > 0. Otherwise `PASS`. MEDIUM/LOW are REPORTED, never
blocking.
## OUTPUT (exact format — machine-parsed by the orchestrator)
```
SECURITY — VERDICT: PASS | BLOCK(n) | ERROR(<reason>)
TOOL: semgrep <ver> — p/security-audit, p/secrets, p/owasp-top-ten | ABSENT (DEGRADED — checklist only; install: make plugin)
SCOPE: <n> files
BLOCKING:
1. [CRITICAL|HIGH] <rule/check> — <file:line> — <why> — hint: <fix direction>
REPORTED (non-blocking):
- [MEDIUM|LOW] <rule/check> — <file:line>
PROOF: semgrep <n> rules on <n> files → <n> findings; checklist <n> checks → <n> findings
```
In audit mode, ALSO write this same block (plus per-finding detail) to
`REPORT`, and end stdout with `REPORT_WRITTEN: <path>`.
## RULES
- Report-only on CODE. Never edit or fix a code file. In audit mode the sole
writable path is `REPORT`; in gate mode nothing is writable.
- `PROOF` is MANDATORY — a `PASS` (or DEGRADED PASS) without a `PROOF` line
showing what was scanned is invalid; the orchestrator discards it as a
structural failure (LRN-048).
- A mute / crashed / unparsable auditor is NEVER a PASS. Exactly one
`SECURITY — VERDICT:` line, spelled as above.
- Blocks on HIGH/CRITICAL only. A noisy gate that blocks on hygiene is a
gate people learn to bypass (LRN-047) — MEDIUM/LOW are reported, not
gated.
## ORCHESTRATOR PROTOCOL (consumer contract — wiring reference)
- The security gate runs AFTER the request-conformity verdict is CONFORME
(verifier), never before.
- Dispatch a FRESH auditor each iteration — no context reuse. Input = mode +
scope + (report) + (context), nothing else.
- Parse the `SECURITY — VERDICT:` line:
- `PASS` → proceed (to commit / next step).
- `BLOCK(n)` → the dev subagent receives the BLOCKING list + the contract
path. After the fix: re-verify the REQUEST first (verifier), THEN re-run
this gate — in that order. Max 3 security iterations → STOP + human
escalation with the BLOCKING table.
- `DEGRADED` (semgrep absent) → does NOT block; surface the checklist
result + recommend `make plugin`. A DEGRADED BLOCK (grep-caught
hardcoded secret etc.) blocks like any other.
- Structural failure (`ERROR(…)`, missing/duplicated VERDICT line,
unparsable, crash, PASS without PROOF) → retry ONCE fresh; 2nd
structural failure → human escalation. A mute auditor is never a PASS.
+78
View File
@@ -0,0 +1,78 @@
#!/usr/bin/env bash
# ============================================================
# Structure locks — security-auditor agent + grafts (lot 3)
# Deterministic greps on load-bearing doctrine: an edit that
# drops one (pinned rulesets, DEGRADED-still-checks, PROOF,
# block-HIGH-only, anti-gaming, the two SKILL grafts) reds here.
# ============================================================
set -u
REPO="$(cd "$(dirname "$0")/../.." && pwd)"
AGT="$REPO/agents/security-auditor.md"
ONB="$REPO/skills/onboard/SKILL.md"
ADL="$REPO/skills/audit-delta/SKILL.md"
PASS=0; FAIL=0
tf() { # tf <label> <file> <fixed-string>
if grep -qF -- "$3" "$2" 2>/dev/null; then
echo " PASS $1"; PASS=$((PASS+1))
else
echo " FAIL $1 — missing: $3"; FAIL=$((FAIL+1))
fi
}
tr_() { # tr_ <label> <file> <ERE>
if grep -qE -- "$3" "$2" 2>/dev/null; then
echo " PASS $1"; PASS=$((PASS+1))
else
echo " FAIL $1 — no match: $3"; FAIL=$((FAIL+1))
fi
}
tn() { # tn <label> <file> <ERE> (must NOT match)
if grep -qE -- "$3" "$2" 2>/dev/null; then
echo " FAIL $1 — forbidden match: $3"; FAIL=$((FAIL+1))
else
echo " PASS $1"; PASS=$((PASS+1))
fi
}
echo "── security-auditor.md locks ──"
if [ -f "$AGT" ]; then
echo " PASS agent exists"; PASS=$((PASS+1))
else
echo " FAIL agent missing: $AGT"; FAIL=$((FAIL+1))
fi
tr_ "frontmatter name" "$AGT" "^name: security-auditor$"
tr_ "tools incl Write (audit)" "$AGT" "^tools: Read, Grep, Glob, Bash, Write$"
tf "verdict grammar" "$AGT" "SECURITY — VERDICT: PASS | BLOCK(n) | ERROR(<reason>)"
tf "ruleset security-audit" "$AGT" "p/security-audit"
tf "ruleset secrets" "$AGT" "p/secrets"
tf "ruleset owasp required" "$AGT" "p/owasp-top-ten"
tf "no config auto stated" "$AGT" "never \`--config auto\`"
tf "no auto login" "$AGT" "never \`semgrep login\`"
tf "secrets to CRITICAL" "$AGT" "p/secrets | CRITICAL"
tf "block ERROR threshold only" "$AGT" "blocking threshold is ERROR"
tf "medium low reported" "$AGT" "MEDIUM/LOW are REPORTED, never"
tf "degraded still checks" "$AGT" "STILL RUN STEP 3"
tf "degraded vacuous pass named" "$AGT" "vacuous pass"
tf "anti-gaming suppression" "$AGT" "NEW suppression comment"
tf "anti-gaming micro-gate" "$AGT" "[gated <date>]"
tf "proof mandatory" "$AGT" "\`PROOF\` is MANDATORY"
tf "mute never a pass" "$AGT" "NEVER a PASS"
tf "write rule-locked audit" "$AGT" "writable path is \`REPORT\`"
tf "gate mode write forbidden" "$AGT" "\`Write\` is FORBIDDEN in this mode"
tf "blind no history" "$AGT" "NEVER receive iteration history"
tf "reverify request first" "$AGT" "re-verify the REQUEST first"
tf "max 3 security iters" "$AGT" "Max 3 security iterations"
echo "── onboard graft locks ──"
tf "onboard dispatches auditor" "$ONB" "subagent_type=\"security-auditor\""
tf "onboard report path" "$ONB" ".onboard-audit/semgrep.md"
tf "onboard verify incl semgrep" "$ONB" "code-clean,cso,semgrep,doc"
echo "── audit-delta graft locks ──"
tf "audit-delta dispatches" "$ADL" "subagent_type=\"security-auditor\""
tf "audit-delta semgrep first" "$ADL" "FIRST run the semgrep SAST pass"
echo ""
echo "security-auditor structure locks: $PASS pass, $FAIL fail"
[ "$FAIL" -eq 0 ]
+15 -6
View File
@@ -221,12 +221,21 @@ Then offer to capitalize (per CLAUDE.md): recurring finding patterns →
## Axis specs (subagent prompts)
- **security** — scoped to the delta: hardcoded secrets/tokens/keys (also
in comments), injection (SQL/XSS/command — string concat into
queries/shells), authN/authZ gaps on new endpoints, fail-open error
paths, secrets/PII in logs, new dependencies in lockfiles (name them +
known CVEs), unguarded destructive shell (`rm -rf` with unquoted or
un-`:?`-guarded vars).
- **security** — FIRST run the semgrep SAST pass, THEN the reasoned checks
below on the same delta (the SAST is a deterministic floor, the reasoned
pass covers what grep/rules miss):
```
Agent(subagent_type="security-auditor", description="audit-delta security — semgrep SAST",
prompt="MODE: audit\nSCOPE: <delta file list>\nREPORT: .claude/audits/.audit-delta-semgrep.md\nFollow agents/security-auditor.md exactly. Pinned rulesets, no login. Write ONLY to REPORT. End with REPORT_WRITTEN: <path>.")
```
Fold its BLOCKING (CRITICAL/HIGH) + REPORTED findings into this axis'
finding list before the 3c gate. semgrep ABSENT → DEGRADED (checklist
only) is surfaced, not a blocker. Reasoned checks (also scoped to the
delta): hardcoded secrets/tokens/keys (also in comments), injection
(SQL/XSS/command — string concat into queries/shells), authN/authZ gaps
on new endpoints, fail-open error paths, secrets/PII in logs, new
dependencies in lockfiles (name them + known CVEs), unguarded destructive
shell (`rm -rf` with unquoted or un-`:?`-guarded vars).
- **errors** — bugs in changed code: logic errors, off-by-one, unhandled
edge cases (empty/null/unicode/concurrent), race conditions, swallowed
errors, resource leaks (missing trap/close/finally). Improvements only
+26 -2
View File
@@ -486,6 +486,29 @@ bash $HOME/.claude/lib/toggle-external.sh list 2>/dev/null | grep -E "^gstack\s+
)
```
#### Dispatch semgrep SAST — `security-auditor` (TOUJOURS, complément de cso)
En complément de cso (ON) OU du fallback (OFF) — un moteur SAST déterministe
à côté de l'audit grep/raisonné. cso est un submodule gstack non modifiable ;
semgrep vit dans cet agent local. Lancé dans les DEUX branches gstack.
```
Agent(
subagent_type="security-auditor",
description="Onboard — semgrep SAST audit (report-only)",
prompt="""
MODE: audit
SCOPE: <PROJECT_ROOT>
REPORT: <PROJECT_ROOT>/.onboard-audit/semgrep.md
CONTEXT: <PROJECT_ROOT>/.onboard-audit/archetype-context.md
Follow agents/security-auditor.md exactly. Pinned rulesets only, no login.
Write ONLY to the REPORT path. End stdout with REPORT_WRITTEN: <path>.
"""
)
```
Si semgrep absent → l'agent rend DEGRADED (checklist seule) + recommande
`make plugin` ; NON bloquant en onboard (audit, pas gate).
#### Dispatch doc-syncer (si `doc` dans audit_stack)
```
Agent(
@@ -511,9 +534,10 @@ Agent(
### Après les 3 dispatches
Attendre la fin des 3 subagents. Vérifier que les 3 fichiers existent et sont non vides :
Attendre la fin des subagents. Vérifier que les fichiers existent et sont non vides
(semgrep.md inclus — DEGRADED reste non vide : il porte le résultat checklist) :
```bash
for f in .onboard-audit/{code-clean,cso,doc}.md; do
for f in .onboard-audit/{code-clean,cso,semgrep,doc}.md; do
[ -s "$f" ] && echo "OK $f" || echo "MISSING $f"
done
```