Merge branch 'feature/contract-verifier' into feature/verify-loops
# Conflicts: # .claude/memory/decisions.md # .claude/memory/journal.md # .claude/memory/learnings.md
This commit is contained in:
@@ -70,6 +70,7 @@ rules:
|
||||
| BDR-046 | 2026-07-01 | Claude Code installs via official native installer (curl claude.ai/install.sh), drop npm from install.sh | accepted |
|
||||
| BDR-047 | 2026-07-01 | ECC audit → zero import; local config ahead of reference | accepted |
|
||||
| BDR-048 | 2026-07-03 | semgrep security gate: engine version + rulesets PINNED, never --config auto; upgrade = deliberate visible human jump | accepted |
|
||||
| BDR-049 | 2026-07-03 | verifier = fresh + blind (no iteration history) + disk-contract + PROOF-or-fail; mute ≠ PASS; scope enrichment via human micro-gate | accepted |
|
||||
|
||||
---
|
||||
|
||||
@@ -797,3 +798,12 @@ rules:
|
||||
- **Rationale**: gate blocks HIGH/CRITICAL only ([[LRN-047]]); silent engine/rule upgrade = new BLOCKs on unchanged code w/o human decision → gate crying false → ignored. Version jump must be deliberate + visible (bump pin, then `make update` shows the jump).
|
||||
- **Alternatives rejected**: `latest` (pipx house default, graphifyy-style) — fine for comfort tools, wrong for a blocking gate; `--config auto` — telemetry + non-determinism.
|
||||
- **Reference**: plugins.lock.json `semgrep` entry, install-plugins.sh STEP 7.5, update-all.sh step 6.2 — branch feature/semgrep-install `ccfecc9`. Conditions [[LRN-047]], [[LRN-085]]. Coverage caveat of the community rulesets: [[LRN-092]].
|
||||
- **Addendum 2026-07-03** (lot 3, measured): rulesets = `p/security-audit` + `p/secrets` + **`p/owasp-top-ten`**. owasp-top-ten is REQUIRED not optional — measured on realistic Flask code, the 2-ruleset baseline missed SQL injection + path traversal ENTIRELY (0 findings); owasp's taint rules catch them. Severity map: secrets ERROR→CRITICAL, other ERROR→HIGH (block), WARNING/INFO→reported. Blocking threshold = ERROR (per-RULE, not per-vuln — same class can straddle ERROR/WARNING; blocking WARNING too floods FP). FP measured shell/md only (faunosteo, game: sole added blocking ERROR = Dockerfile `missing-user` hygiene, contained by gate-mode diff-scoping). **Re-evaluate owasp FP at the first real web/python app project** (shell/md repos don't represent where the gate runs). See [[LRN-094]], agents/security-auditor.md branch feature/security-auditor `2b297bd`.
|
||||
|
||||
## BDR-049 — Verifier doctrine: fresh + blind + disk-contract + proof-or-fail
|
||||
|
||||
- **Date**: 2026-07-03
|
||||
- **Decision**: conformity verdict comes ONLY from a FRESH verifier subagent per iteration. Input = contract PATH (read from disk — dev restatement structurally unable to interpose) + diff range + optional test cmd. NEVER iteration history: blind, complete verification every time (cost bounded by the main-loop max-3 cap, [[LRN-083]]: loops decided in main loop). CONFORME ⇔ all criteria MET + zero out-of-scope. PROOF line mandatory ([[LRN-048]]). Mute/unparsable verifier NEVER a PASS: 1 fresh retry, 2nd structural failure = human escalation. Dev-justified out-of-scope enters FILE SCOPE only via a human micro-gate (`[gated]` marker) — else the dev justifies everything and scope constrains nothing. Contract on DISK at creation (`.claude/tasks/contracts/<date>-<slug>-<HHMM>.md`, committed; aborted run → deleted or `status: aborted`, never left dirty).
|
||||
- **Rationale**: dev self-score is always confident → not a gate. Verifier fed history anchors on prior verdicts → telescopic drift. Context-only contract dies at compaction, the verbatim with it.
|
||||
- **Alternatives rejected**: dev self-assessment as gate; cumulative verifier context ("cheaper" but anchored); gitignored run files (lose escalation reference + session-death survival).
|
||||
- **Reference**: lib/contract-interview.md + agents/verifier.md + lib/tests/contract-verifier.test.sh (31 locks) — branch feature/contract-verifier `6aed5ee`. Behavioral GREEN: planted-gap → ECARTS(2) exact; conform-under-injected-history → CONFORME (blindness held). Twin of [[BDR-048]] (security gate). Conditions [[LRN-048]], [[LRN-083]].
|
||||
|
||||
@@ -315,3 +315,4 @@ rules:
|
||||
- #2 done (bugfix/design-toolchain-trigger): trigger tightened — dropped bare design|component|composant|theme|thème|transition|frontend|front-end|palette; dashboard→\bdashboard\b (kills ecc_dashboard.py filename match, keeps "admin dashboard"); kept animation; added "front-?end design" bigram + fire-log counter (time+token+excerpt, ~/.claude/logs/design-toolchain-fires.log) so future "re-firing?" is measured. Test 18/18, shellcheck clean, live dogfood green. [[LRN-091]] corrob [[LRN-047]].
|
||||
- Double dogfood of #1 guard: config-protection blocked + sentinel-bypassed my own edits to the now-guarded design hook + its test — first real use of the guard, friction validated in passing (one-shot sentinel .claude/.config-edit-ok, non-empty reason, logged+consumed). ECC second-regard closed: #1 config-protection + #2 trigger fix, both merged to develop, nothing pushed.
|
||||
- Chantier verify-loops/semgrep/contract: Phase 1 read-only (6 subagents mapped 6 orchestrators + cso + agents + install patterns; caught subagent error — cso IS gstack symlink, ls-verified) → archi GATED-GO (5 verdicts: local grafts, dev inline light flows, hotfix unchanged, pinned rulesets, pinned version; +2 specs: contract on DISK, mute verifier ≠ PASS). LOT 1 shipped on feature/semgrep-install (ccfecc9+b8d3ccc): install-plugins STEP 7.5 + update-all 6.2 + lock pin 1.168.0, dogfooded real (4 paths + anonymous ruleset fetch + detection). [[BDR-048]] [[LRN-092]]. Next: lot 2 specs (contract-interview lib + verifier agent).
|
||||
- Chantier verify-loops LOT 2 (feature/contract-verifier `6aed5ee`): lib/contract-interview.md (verbatim contract on DISK, micro-gate scope enrichment, aborted never dirty) + agents/verifier.md (fresh+blind, PROOF-or-fail, mute ≠ PASS) + 31 structure locks green, shellcheck clean. Behavioral: planted-gap → ECARTS(2) exact; conform under injected fake history → CONFORME (blindness held). Sentinel consumed 4× on guarded lib/tests/. [[BDR-049]] [[LRN-093]]. Merge note: lot 1+2 both append registries at same anchors → trivial stack-conflict expected. Next: lot 3 security-auditor spec.
|
||||
|
||||
@@ -109,7 +109,11 @@ rules:
|
||||
| LRN-087 | 2026-07-02 | presence-flag ≠ capability — rtk silently dead after .bashrc wipe; emitted commands need ABSOLUTE bin paths (they run in another shell); integrity pin = live machinery, re-pin on hook edit | any PATH-dependent capability + hand-managed shell profile; hooks emitting commands for another shell |
|
||||
| LRN-088 | 2026-07-02 | token-cutting intuition inverts under measurement — verbosity beats cardinality (gstack 34 skills ≈ 592 tok vs pr-review 6 agents ≈ 2,183) | any "disable X to save tokens" — measure per-item bytes first; profiles toggle skills, not plugin payloads |
|
||||
| LRN-089 | 2026-07-03 | pass-through wrapper (CLI `"$@"` → fn deriving target from ambient state: HEAD/cwd/env) silently ignores its args = silent contract violation; guard = args are an ASSERTION, refuse when they disagree with state | any dispatcher forwarding args to a callee that reads ambient state instead of the args |
|
||||
| LRN-090 | 2026-06-30 | external-repo audit: open WIRED subsystems (hooks/runners) before declarative (docs/rules); described capability ≠ wired capability | auditing an external config/framework repo for transferable value |
|
||||
| LRN-091 | 2026-07-03 | keyword-triggered soft-nudge hook w/ bare common tokens over-fires on non-UI work → tuned out; bare only when UI sense dominates, else bigram-or-drop | any advisory/nudge hook keyed on keywords |
|
||||
| LRN-092 | 2026-07-03 | SAST smoke test w/ the OFFICIAL example secret = vacuous pass (rules exclude documented example keys by design); validate w/ realistic payloads + measure tier coverage before trusting a gate ruleset | smoke-testing any detector/gate — never the canonical example payload |
|
||||
| LRN-093 | 2026-07-03 | grep -F pattern w/ embedded newline = per-line OR = lock that matches anything; structure locks single-line only, flip-test new locks | writing any grep-based structure lock / census test |
|
||||
| LRN-094 | 2026-07-03 | SAST severity ≠ exploitability — semgrep ERROR conflates real vulns + hardening recos; metadata does NOT cleanly separate them (measured) → metadata refinement = noisy gate; ERROR-threshold + diff-scoping is the containment | mapping a SAST tool's output to a blocking gate |
|
||||
|
||||
---
|
||||
|
||||
@@ -984,3 +988,15 @@ rules:
|
||||
- **context**: lot 1 semgrep-install dogfood 2026-07-03. `p/secrets`+`p/security-audit` community tier: anonymous fetch OK (52 rules, no login), `subprocess-shell-true` detected ERROR; MISSED %-format SQLi on bare cursor (no recognized DB-API context) + fake-checksum `ghp_` token. Gap logged for security-auditor agent design (consider adding `p/owasp-top-ten`).
|
||||
- **future application**: any detector/gate smoke test — craft realistic payloads, never the canonical example; measure the miss-list on purpose-built fixtures; size the gate's blocking scope on that data.
|
||||
- **cousin**: [[LRN-048]] a 0/OK must prove it looked; [[LRN-047]] noisy guard = ignored guard; conditions [[BDR-048]].
|
||||
|
||||
## LRN-093 — grep -F with an embedded newline = per-line OR = vacuous lock
|
||||
- **pattern**: a fixed-string grep pattern containing a newline is treated as MULTIPLE patterns (one per line) — match succeeds if ANY line matches. A structure lock written that way passes on essentially anything (`"no\n forced loop"` → matches any "no") = a lock that proves nothing, [[LRN-048]] class.
|
||||
- **context**: lot 2 `lib/tests/contract-verifier.test.sh`, caught in self-review BEFORE first run; replaced by a single-line distinctive anchor ("proceed straight to the security gate").
|
||||
- **future application**: structure locks / census greps = ONE line per pattern, always; a clause spanning lines → lock a distinctive single-line fragment. Flip-test every new lock (prove it CAN fail) before trusting its green.
|
||||
- **cousin**: [[LRN-048]] a pass must prove it looked; [[LRN-046]] deterministic-oracle discipline.
|
||||
|
||||
## LRN-094 — SAST severity ≠ exploitability; metadata does not cleanly separate — don't refine on it
|
||||
- **pattern**: semgrep `ERROR` conflates exploitable vulns (SQLi, secrets, command injection) with hardening recommendations (Dockerfile missing-USER, npm release-age). The obvious refinement — gate on `metadata.impact`/`likelihood`/`confidence` — does NOT work: measured, the Dockerfile hygiene ERROR (`impact=MEDIUM likelihood=LOW`) is indistinguishable from a tainted-SQL ERROR (`impact=MEDIUM likelihood=MEDIUM`), and a real command-injection reads `impact=LOW likelihood=HIGH`. Metadata-based severity = a noisy, non-deterministic gate ([[LRN-077]] class).
|
||||
- **context**: lot 3 security-auditor design 2026-07-03. Measured on 2 real repos (faunosteo, game): the only added blocking ERROR from owasp-top-ten is Dockerfile hygiene, contained because gate mode scopes to the DIFF (a pre-existing infra finding can't block an unrelated code change).
|
||||
- **future application**: mapping any SAST to a blocking gate — take the tool's ERROR/blocking level as the deterministic threshold, contain FP by SCOPING (diff, not repo), NOT by a metadata heuristic or a hand-maintained hygiene denylist ([[LRN-049]]: match guard cost to proven stake — build the denylist only if hygiene ERRORs prove noisy on a real project).
|
||||
- **cousin**: [[LRN-047]] noisy gate = ignored; [[LRN-077]] non-deterministic gate; conditions [[BDR-048]].
|
||||
|
||||
Reference in New Issue
Block a user