Merge feature/contract-gates into develop

This commit is contained in:
Bastien Chanot
2026-08-24 13:44:06 +02:00
20 changed files with 965 additions and 23 deletions
+9
View File
@@ -92,6 +92,7 @@ rules:
| BDR-072 | 2026-07-17 | SPA: honest refuse (On-page N/A, not zero), no headless browser (R2 over R1) | accepted | | BDR-072 | 2026-07-17 | SPA: honest refuse (On-page N/A, not zero), no headless browser (R2 over R1) | accepted |
| BDR-073 | 2026-07-17 | Scoring: LLM judges findings+severity, engine does the arithmetic (deterministic /20) | accepted | | BDR-073 | 2026-07-17 | Scoring: LLM judges findings+severity, engine does the arithmetic (deterministic /20) | accepted |
| BDR-080 | 2026-07-21 | Bug routing inverted: /bugfix primary, /investigate explicit-only | accepted | | BDR-080 | 2026-07-21 | Bug routing inverted: /bugfix primary, /investigate explicit-only | accepted |
| BDR-083 | 2026-08-24 | Contract gates: deterministic floor (GATE 0) under the fresh verifier | accepted |
--- ---
@@ -1079,3 +1080,11 @@ Opus 5 (released 2026-07-24) now backs every `model: opus` pin (BDR-076/077) + a
### BDR-082 — seo/geo analyzers de-prescribed for Opus 5 (C1) [accepted] (2026-08-02) ### BDR-082 — seo/geo analyzers de-prescribed for Opus 5 (C1) [accepted] (2026-08-02)
BDR-081 N5 follow-on, user-directed apparatus (plan+3-lens challenge+census+dogfood). Method: audience×mode-range invariant — dedup ONLY verbatim same-audience (spec rule / bundle-item payload / phase-local caveat) same-mode-range repeats; cross-mode + agent↔dispatcher twins stay (standalone paths need them). Census-FIRST: lib/tests/seo-geo-contract.test.sh 71 locks (verdict grammar, sentinels, ALL STEP headers incl. interiors, item fields, score labels, envelope keys), flip-proven 7 mutations→7 FAILs, committed BEFORE reword. Shipped: self-output verification removed (":970 run twice"→conditional integrity guard; ":1217"→single-shot-scoped), 2 pre-BDR-061 vestigials fixed, caps softened (P0-rule/MANDATORY/ALWAYS→plain content rules), 2 essays compressed, checklist :1309→routing map rows verbatim (challenger caught it = routing table, NOT self-check), true same-range dups only (seo Handoff+landing-page blocks; geo ZERO — all claimed pairs distinct on inspection). FROZEN: guard-first url-guard orderings, :550 denominator-before-sampling (ordering IS the honesty mechanism), R2/NAP/COVERAGE/citation invariants, external-freshness checks (world drift ≠ self-verification). Deltas: seo 1528→1503 l ("P0 rule" 2→0, ALWAYS 1→0, MUST 5→4, NEVER 9→9 = class-B bans kept); geo 1106→1107 (MANDATORY 1→0, MUST 4→3). Plan challenged correctness FATAL / robustness FATAL(3 BLOCKER) / simplicity CONCERNS + confirmation FATAL(9) — every BLOCKER closed by named change (§5bis record). Dogfood before/after on frozen zenquality copy: judge-replay on frozen signals (zero collect variance) + templates + fresh collects + e2e judge + 42/42 assert battery BOTH sets + blind reader "interchangeable; all deltas = presentation variance both directions OR after MORE spec-conformant". Alternatives rejected: keyword dedup (challengers proved audience/range-blind — most annex "twins" were distinct obligations), FULL/aggressive dogfood (billing gate killed nested CLI; left as user option), banner/shape locks (LLM-convention layers wobble — lock strings only). Evidence: .audit/dogfood-baseline/ (18 artifacts + DOGFOOD-VERDICT.md), plan .claude/tasks/plans/2026-07-30-seo-geo-deprescription-1402.md. Branch feature/seo-geo-deprescription, UNMERGED (human gate). BDR-081 N5 follow-on, user-directed apparatus (plan+3-lens challenge+census+dogfood). Method: audience×mode-range invariant — dedup ONLY verbatim same-audience (spec rule / bundle-item payload / phase-local caveat) same-mode-range repeats; cross-mode + agent↔dispatcher twins stay (standalone paths need them). Census-FIRST: lib/tests/seo-geo-contract.test.sh 71 locks (verdict grammar, sentinels, ALL STEP headers incl. interiors, item fields, score labels, envelope keys), flip-proven 7 mutations→7 FAILs, committed BEFORE reword. Shipped: self-output verification removed (":970 run twice"→conditional integrity guard; ":1217"→single-shot-scoped), 2 pre-BDR-061 vestigials fixed, caps softened (P0-rule/MANDATORY/ALWAYS→plain content rules), 2 essays compressed, checklist :1309→routing map rows verbatim (challenger caught it = routing table, NOT self-check), true same-range dups only (seo Handoff+landing-page blocks; geo ZERO — all claimed pairs distinct on inspection). FROZEN: guard-first url-guard orderings, :550 denominator-before-sampling (ordering IS the honesty mechanism), R2/NAP/COVERAGE/citation invariants, external-freshness checks (world drift ≠ self-verification). Deltas: seo 1528→1503 l ("P0 rule" 2→0, ALWAYS 1→0, MUST 5→4, NEVER 9→9 = class-B bans kept); geo 1106→1107 (MANDATORY 1→0, MUST 4→3). Plan challenged correctness FATAL / robustness FATAL(3 BLOCKER) / simplicity CONCERNS + confirmation FATAL(9) — every BLOCKER closed by named change (§5bis record). Dogfood before/after on frozen zenquality copy: judge-replay on frozen signals (zero collect variance) + templates + fresh collects + e2e judge + 42/42 assert battery BOTH sets + blind reader "interchangeable; all deltas = presentation variance both directions OR after MORE spec-conformant". Alternatives rejected: keyword dedup (challengers proved audience/range-blind — most annex "twins" were distinct obligations), FULL/aggressive dogfood (billing gate killed nested CLI; left as user option), banner/shape locks (LLM-convention layers wobble — lock strings only). Evidence: .audit/dogfood-baseline/ (18 artifacts + DOGFOOD-VERDICT.md), plan .claude/tasks/plans/2026-07-30-seo-geo-deprescription-1402.md. Branch feature/seo-geo-deprescription, UNMERGED (human gate).
### BDR-083 — contract gates: deterministic floor (GATE 0) under the verifier [accepted] (2026-08-24)
User asked what to take from `unlazy` skill (Leonxlnx/unlazy 2.1.0, MIT). Verdict on its verification ARCHITECTURE: teaches nothing we lack — contract + fresh blind verifier + bounded loops + order invariant already shipped (BDR-049/050/066, LRN-083). Real gap found elsewhere: between executor and GATE 1, NO deterministic floor. GATE 1 = LLM dispatch; verifier's mandatory `PROOF:` line = a line the verifier WRITES — nothing structurally stops it being produced without executing anything (LRN-048 demands a pass prove it looked; the proof is self-reported prose). Decision: import unlazy's gate ledger INTO the existing contract, never alongside it. Palier 2, user-chosen over doctrine-only / defer.
TAKEN: criterion carries an oracle (indented `CHECK:` cmd + `EXPECT:` success-only marker + `EVIDENCE:` slot); fail-closed = exit 0 AND marker (a nonzero process never passes because its error text carries the token); evidence persisted INTO the contract → the fresh verifier reads fact, not the executor's report; `ABANDON: <id> <non-blank reason>` = impossible criterion never deleted, blocks CONFORME, routes to human gate (new verdict token `ABANDONED(n)` — distinct routing from ECARTS ⇒ distinct token, not a sub-line to re-derive); 4 gate-authoring rules (observe the named artifact / success-only marker / positive control before any absence check / recompute supplied numbers, never copy one into EXPECT); 4-pass executor discipline (feater full; bugfixer narrowed to fix+test under "keep the fix minimal", pass 3 = negative control proving the regression test fails without the fix).
REFUSED + why: Stop hook `decision:"block"` — contradicts "STOP + human escalation", "gone WRONG → STOP re-plan", "merge only on explicit human signal"; a hook FORCING continuation is the inverse of our gates; its 6-block release either traps the session or gives up; each block = an agent continuation = real tokens. Approval store `~/.unlazy/approved` (binds ledger+cmd+CWD+shell+timeout+platform+full PATH) — exists to execute ledgers INHERITED from untrusted repos; our contracts are authored by our own orchestrator in our own repo ⇒ biggest chunk of their 28k checker closes zero threat here. `.unlazy/<scope>/` tree (PLAN+GATES+gates/+status.log+session+hook-state+locks/) — a 4th bookkeeping tree beside .claude/tasks/{contracts,plans} + memory/ + audits/. `tree N` Depth-Tree effort arithmetic — disowned by unlazy's OWN research/validation-protocol.md (v1 six-run figures unreproducible), while the repo DESCRIPTION still advertises the retracted claim. Node checker (28k .mjs + 54k .mjs tests) — lib stack is 100% bash, Health Stack = `shellcheck *.sh hooks/*.sh lib/*.sh` would cover none of it. `OWNS:` ownership leases — deferred (Palier 3): our parallel dispatches (seo/geo, 3 plan-challengers) are read-only, the write-collision problem does not exist yet.
Shipped: lib/gates.sh (~250 l bash; `status` never executes and never writes · `run` ALWAYS re-executes every runnable criterion — trusting written evidence is the failure being closed, so there is no incremental mode to get wrong; rc 0 MET / 2 UNMET|malformed / 3 ABANDONED; parse fails closed on partial oracle, duplicate id, unindented attribute, runnable-without-EVIDENCE, and executes nothing when the ledger is malformed). GATE 0 in lib/verify-secure-loop.md (red floor → executor re-dispatch with the NOT-MET rows, NO verifier spawned; own 3-iteration budget, separate from conformity; malformed ledger fixed in the main loop, never dispatched to a dev). Order invariant now GATE 0→1→2 on every re-loop. lib/contract-interview.md: ORACLES section + template + ABANDONMENT lifecycle + per-flow oracle weight. agents/verifier.md: oracle-consumption rules — a red or unrun oracle is NEVER overridden by reading code; a MET oracle proves the ORACLE, not the English sentence ⇒ vacuous oracle = NOT-MET, the one judgement no command can make; verifier may re-run a CHECK but never edits the contract. lib/tests/gates.test.sh 64 assertions (sentinel-proved non-execution, with its own positive control asserted first).
Alternatives rejected: Palier 1 doctrine-only (CHECK:/EXPECT: become decorative without an executant); port the Node checker (stack break, shellcheck-blind); fold ABANDONED into ECARTS (would send a dev to fix the impossible and eat the 3-iteration budget); `status` revalidating old evidence (that trust is the failure being closed).
Branch feature/contract-gates, UNMERGED (human gate). `make test` rc 0, shellcheck clean, e2e verified on a real contract in the documented template.
+9
View File
@@ -37,6 +37,7 @@ rules:
| EVAL-018 | 2026-07-06 | job3 docs-drift audit + execution: 46/46 findings verified, 20/23 fixes shipped (B1 blocked, D2-D5+B6 skipped by decision), zero residual on re-sweep | keep | | EVAL-018 | 2026-07-06 | job3 docs-drift audit + execution: 46/46 findings verified, 20/23 fixes shipped (B1 blocked, D2-D5+B6 skipped by decision), zero residual on re-sweep | keep |
| EVAL-019 | 2026-07-06 | job4 test-gap audit + execution: 11 specs + 5 fixes/seams, every mutation red-green verified, zero residual | keep | | EVAL-019 | 2026-07-06 | job4 test-gap audit + execution: 11 specs + 5 fixes/seams, every mutation red-green verified, zero residual | keep |
| EVAL-025 | 2026-07-17 | opening seo/geo inventory (subagents): 7/7 verifiable claims false or overstated; real contact corrected all, 6 plan corrections + 4 features killed at measurement | keep | | EVAL-025 | 2026-07-17 | opening seo/geo inventory (subagents): 7/7 verifiable claims false or overstated; real contact corrected all, 6 plan corrections + 4 features killed at measurement | keep |
| EVAL-027 | 2026-08-24 | contract-gates behavioral RED: 16/16 fresh unprimed runs followed new doctrine (GATE 0 order, vacuous oracle, ABANDONED routing, scope temptation resisted) | keep |
--- ---
@@ -251,3 +252,11 @@ rules:
### EVAL-026 — 3-way plan challenge caught 4 BLOCKERs dogfooding own plan (2026-07-17) ### EVAL-026 — 3-way plan challenge caught 4 BLOCKERs dogfooding own plan (2026-07-17)
Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itself. Verdicts CONCERNS(4)/FATAL(6)/FATAL(4). Caught 4 distinct BLOCKERs a single pass would blend: (1) v1 unbuildable — targeted init-project (inline-load, no dispatch) + false "plan on disk" premise for feat/bugfix (only contract persists); (2) failed-open silently dropping a lens while claiming "challenged" (inverts verify-secure-loop "a mute verifier is NEVER a PASS"); (3) consensus-weighting buries lone L2 security finding (lenses orthogonal); (4) sonnet challengers violate [[BDR-066]] (audit judgment=big model). Synthesis REJECTED 1 false positive (allowed-tools-blocks-dispatch — ship-feature has same frontmatter + dispatches fine). Each lens found a DIFFERENT class of flaw → evidence 3-independent > 1-multilens. Action: hardened v2 (severity-driven + fail-safe + re-think loop) shipped. Method validated itself before build. Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itself. Verdicts CONCERNS(4)/FATAL(6)/FATAL(4). Caught 4 distinct BLOCKERs a single pass would blend: (1) v1 unbuildable — targeted init-project (inline-load, no dispatch) + false "plan on disk" premise for feat/bugfix (only contract persists); (2) failed-open silently dropping a lens while claiming "challenged" (inverts verify-secure-loop "a mute verifier is NEVER a PASS"); (3) consensus-weighting buries lone L2 security finding (lenses orthogonal); (4) sonnet challengers violate [[BDR-066]] (audit judgment=big model). Synthesis REJECTED 1 false positive (allowed-tools-blocks-dispatch — ship-feature has same frontmatter + dispatches fine). Each lens found a DIFFERENT class of flaw → evidence 3-independent > 1-multilens. Action: hardened v2 (severity-driven + fail-safe + re-think loop) shipped. Method validated itself before build.
### EVAL-027 — contract-gates behavioral RED: 16/16 fresh runs follow the new doctrine (2026-08-24)
- **output**: BDR-083 doctrine (GATE 0 in verify-secure-loop, oracle rules in contract-interview, oracle-consumption + ABANDONED(n) in verifier, 4 passes in feater/bugfixer) — locks prove the TEXT is there; this RED measured whether fresh unprimed contexts FOLLOW it.
- **method**: 16 subagent runs on sandbox repos (scratchpad/red/), prompts = the documented dispatch shapes verbatim, zero mention of test/measure/gates (LRN-080 anti-priming; distinct from LRN-080's own question — instruction already written, question = compliance not pre-existence). Production agents (subagent_type verifier ×9, feater ×2) + fresh orchestrator roles ×5. Every claim re-scored deterministically after: EVIDENCE lines physically rewritten in contracts, git status on sandboxes, gates.sh parse of authored contracts.
- **verdict**: 16/16 conformant. v1 red-oracle-wins 3/3 (NOT-MET citing evidence, own re-run). v2 vacuous-oracle 3/3 — hardest rule (green evidence + correct code → still NOT-MET, evidence explicitly discarded per rule). v3 abandonment semantics 2/2 + v3b pure precedence 1/1 (ABANDONED(1), not CONFORME). o-red 2/2 (gates.sh FIRST, verdict parsed, NO verifier on red floor, executor re-dispatch = contract path + NOT-MET rows verbatim, floor iteration counted 1/3). o-green 1/1 (floor → verifier dispatch with CONTRACT+DIFF+TEST only). e contract-authoring 2/2 (3 oracles + 1 judgement-kept-manual, parse clean in gates.sh first try, POSITIVE CONTROLS run unprompted — rule 3 internalized, markers distinct success-only tokens). f feater 2/2 (out-of-scope temptation src/util.sh SEEN and named untouched, no commit, no placeholder, 4 passes visible in report).
- **anomalies**: none against doctrine. Fixture flaw (mine): placeholder.txt trick used to fabricate a 2nd commit made v2/v3 diffs contain no feature work — every verifier CAUGHT it (out-of-scope + "implementation pre-exists base commit"), polluting v3's intended pure-ABANDONED measurement → v3b clean fixture added. Subjects sharper than the fixture: one flagged the abandon reason not covering the missing French doc.
- **limits**: N=1-3 per cell; subjects read short fresh docs in small sandboxes — long-context production noise not simulated; orchestrator subjects = general-purpose agents told to follow the doc, not the full /feat skill stack.
- **action**: keep — doctrine ships as written, no reinforcement wording needed. Artifacts: scratchpad/red/ (session-lived, not committed).
+6
View File
@@ -433,3 +433,9 @@ rules:
## 2026-08-02 ## 2026-08-02
- C1 seo/geo de-prescription EXECUTED end-to-end: census-first 71 locks flip-proven → reword under audience×range invariant (adafa35/c7646a9) → controlled dogfood (judge-replay frozen signals + templates + fresh collects + e2e + blind reader) → 42/42 both sets, zero contract regression, recall improved. Plan survived 4 challenge passes (2 FATAL + confirmation FATAL(9), all closed by name). BDR-082 + LRN-140. Nested-CLI dogfood died on monthly spend limit → inline pipeline (canonical /seo shape). feature/seo-geo-deprescription UNMERGED (human gate). Chantiers C2-C4 pending. - C1 seo/geo de-prescription EXECUTED end-to-end: census-first 71 locks flip-proven → reword under audience×range invariant (adafa35/c7646a9) → controlled dogfood (judge-replay frozen signals + templates + fresh collects + e2e + blind reader) → 42/42 both sets, zero contract regression, recall improved. Plan survived 4 challenge passes (2 FATAL + confirmation FATAL(9), all closed by name). BDR-082 + LRN-140. Nested-CLI dogfood died on monthly spend limit → inline pipeline (canonical /seo shape). feature/seo-geo-deprescription UNMERGED (human gate). Chantiers C2-C4 pending.
## 2026-08-24
- Analysed `unlazy` skill (Leonxlnx/unlazy 2.1.0) on user request. Its verification architecture teaches us nothing — contract + fresh blind verifier + bounded loops already shipped. Real gap: no deterministic floor between executor and GATE 1 (the verifier's `PROOF:` is a line it writes, not a process exit).
- Shipped Palier 2 (user-chosen): lib/gates.sh + GATE 0 + oracle-bearing criteria + `ABANDONED(n)` verdict + 4-pass executors. Refused unlazy's Stop hook, approval store, .unlazy/ tree, tree-N arithmetic, Node checker — [[BDR-083]] records each why.
- `make test` rc 0, shellcheck clean, 64 new assertions, e2e on a real contract. Branch feature/contract-gates UNMERGED (human gate).
- Locks caught a reflow regression (5 red on rewrapped phrases, zero doctrine lost) → [[LRN-142]]. Skill-adoption pattern → [[LRN-141]].
+10
View File
@@ -1368,3 +1368,13 @@ rules:
- **pattern 3 — de-prescription does NOT depress recall**: reworded collect caught &nbsp;-encoded phone AT COLLECT (baseline collect missed it); reworded judge found new RGPD finding + self-caught false positive + corrected collect coverage claim 21/21→20/21. Integrity/honesty invariants (kept class B) carry the discipline, not the caps. - **pattern 3 — de-prescription does NOT depress recall**: reworded collect caught &nbsp;-encoded phone AT COLLECT (baseline collect missed it); reworded judge found new RGPD finding + self-caught false positive + corrected collect coverage claim 21/21→20/21. Integrity/honesty invariants (kept class B) carry the discipline, not the caps.
- **pattern 4 — lock strings, never shapes**: LLM-convention output layers (banners, fences, table columns, section order) wobble run-to-run in BOTH directions — baseline itself deviated from spec where after conformed (§0 ENTRIES, BUNDLE-before-SCORING). Stable contract = census-locked literal strings; anything unlocked drifts and MUST be tolerated by consumers (tier recognition "by intent" is the right pattern). - **pattern 4 — lock strings, never shapes**: LLM-convention output layers (banners, fences, table columns, section order) wobble run-to-run in BOTH directions — baseline itself deviated from spec where after conformed (§0 ENTRIES, BUNDLE-before-SCORING). Stable contract = census-locked literal strings; anything unlocked drifts and MUST be tolerated by consumers (tier recognition "by intent" is the right pattern).
- **link**: [[BDR-082]] [[BDR-081]] [[LRN-139]] [[LRN-113]]. - **link**: [[BDR-082]] [[BDR-081]] [[LRN-139]] [[LRN-113]].
## LRN-141 — adopting an external skill: take the invariants, refuse the machinery (2026-08-24)
Context: unlazy import ([[BDR-083]]). Pattern: an external skill's MACHINERY encodes ITS threat model and ITS doctrine; only its INVARIANTS transfer. Two clean cases from one repo. (1) Approval store binding PATH/shell/platform exists because unlazy executes ledgers INHERITED from untrusted repos — importing it into a config that authors its own ledgers buys per-command approval prompts and closes zero threat. (2) Stop hook returning `decision:"block"` exists because unlazy has no human gate — importing it into a config whose spine is "STOP + escalate to human" would make the tooling fight the doctrine. Meanwhile the invariants (exit 0 AND marker; evidence persisted so the next reader gets fact not report; impossible ≠ deletable) cost ~250 l of our own bash and fit the EXISTING contract with no new tree.
Separating test: ask WHAT THREAT / WHAT DOCTRINE does this piece assume. Answer "theirs" → refuse the piece, keep the invariant it was protecting.
Corollary on claims: unlazy's own research/validation-protocol.md RETRACTS its v1 benchmark numbers as unreproducible while the repo DESCRIPTION still advertises them. Read a project's self-criticism before its README — the retraction is the credibility signal, the headline is not.
Future application: any skill/plugin adoption — skills-external/, /plugin-check, install-plugins.sh.
## LRN-142 — structure locks are fixed-string: reflowing a doctrine paragraph reds them (2026-08-24)
Context: contract-gates ([[BDR-083]]). Editing lib/verify-secure-loop.md rewrapped 5 locked phrases across line breaks ("Max 3 conformity iterations", "Max 3 security iterations", "re-verify the REQUEST first", "always re-checked BEFORE security", "one verifier dispatch + one security dispatch") → loops-light.test.sh 30 pass / 5 fail, though ZERO doctrine was dropped. Locks did their job: they cannot distinguish "clause deleted" from "clause rewrapped", and that conservative bias is correct — the alternative (fuzzy matching) would miss real deletions.
Rule: when editing a doctrine file under structure locks, grep the test's lock strings FIRST, then re-flow AROUND them — each locked phrase stays on one unbroken line. Fix the DOC, not the lock, unless the doctrine genuinely changed. Under locks today: verify-secure-loop.md, contract-interview.md, verifier / security-auditor / plan-challenger agents, seo+geo (71 locks).
+55
View File
@@ -1144,3 +1144,58 @@ branch) → LOT3 mis-merge trap; + 3 doctor false-warns (LRN-047 class).
comment anchored to measured ~11.4k (LRN-088). False "92% CRITICAL" → ~5% comfortable. comment anchored to measured ~11.4k (LRN-088). False "92% CRITICAL" → ~5% comfortable.
- [x] Verify — suites green (71/13/32/19/20/13 + RC 5/5); doctor 0 false-warn; shellcheck clean. - [x] Verify — suites green (71/13/32/19/20/13 + RC 5/5); doctor 0 false-warn; shellcheck clean.
+docs(changelog) Unreleased entry (706abff). Gate passed on GO 2026-07-03. Finish pending. +docs(changelog) Unreleased entry (706abff). Gate passed on GO 2026-07-03. Finish pending.
## 2026-08-24 — contract gates: plancher déterministe (feature/contract-gates)
Source: analyse du skill `unlazy` (Leonxlnx/unlazy, 2.1.0). Verdict: son
architecture de vérification n'apprend rien (contrat+verifier frais+boucles
bornées ⊂ déjà en place). Le trou réel: **entre l'exécuteur et GATE 1 il n'y a
aucun plancher déterministe** — GATE 1 est un dispatch LLM, et `PROOF:` est une
ligne que le verifier ÉCRIT (rien ne l'empêche structurellement de la produire
sans rien exécuter). Palier 2 retenu (user, 2026-08-24).
PRIS d'unlazy: critère porteur d'oracle exécutable (CHECK/EXPECT/EVIDENCE),
fail-closed (exit 0 ET marqueur), evidence pending = NOT-MET, `ABANDON: <id>
<raison>` comme handoff visible non supprimable, les 4 règles d'écriture de
gates falsifiables, la discipline 4 passes.
REFUSÉ: Stop hook `decision:"block"` (contredit "STOP + escalade humaine" et
"merge sur signal humain"), approval store `~/.unlazy/approved` (résout
l'exécution de ledgers hérités non fiables — pas notre menace), arbre
`.unlazy/<scope>/` (4e arbre de bookkeeping ⇒ mort de la config), `tree N`
(désavoué par ses propres docs), le checker Node 28k (stack lib = 100% bash,
Health Stack = shellcheck).
- [x] W0 branche feature/contract-gates depuis develop (via lib/gitflow.sh)
- [x] W1 `lib/gates.sh` — parse ACCEPTANCE CRITERIA, exécute fail-closed
(exit 0 ET EXPECT), réécrit EVIDENCE dans le contrat. Sous-commandes
`run` (exécute+écrit) / `status` (parse seul, jamais d'exécution, jamais
d'écriture). rc 0=MET · 2=UNMET/malformé · 3=ABANDONED.
- [x] W2 `lib/contract-interview.md` — STEP 3 gagne CHECK/EXPECT/EVIDENCE
optionnels par critère + les 4 règles de falsifiabilité; template mis à
jour; ABANDON dans Lifecycle; ligne de poids par flow.
- [x] W3 `agents/verifier.md` — EVIDENCE fail-closed (coché+pending = NOT-MET),
bucket ABANDONED, verdict `CONFORME` impossible si abandon présent.
- [x] W4 `lib/verify-secure-loop.md` — GATE 0 déterministe avant GATE 1
(rouge ⇒ re-dispatch exécuteur sans brûler un verifier).
- [x] W5 `agents/feater.md` + `agents/bugfixer.md` — discipline 4 passes.
- [x] W6 `lib/tests/gates.test.sh` — comportemental sur gates.sh (fail-closed,
exit≠0 avec marqueur = FAIL, pending, ABANDON, malformé, status
n'exécute pas) + locks de structure sur W2/W3/W4/W5.
- [x] W7 shellcheck + bash -n + `make test` complet.
- [x] W8 CHANGELOG + registres (BDR + LRN + journal).
- [x] W10 restatements skills : bullet GATE 0 dans feat/bugfix/ship-feature/
init-project (+4 locks, flip-testé) ; ligne hotfix du tableau de poids
corrigée (aucun floor à ce poids). 2026-08-24.
- [x] W11 RED comportemental : 16/16 runs frais non-amorcés conformes
(verifier ×9, feater ×2, orchestrateur ×5) → EVAL-027. 2026-08-24.
- [x] W9 merge sur signal humain explicite (2026-08-24, "merge dans develop").
**Won't-build-now — Palier 3 unlazy (OWNS/leases), trigger documenté :**
Différé volontairement (BDR-083) : tous les dispatches parallèles actuels
sont read-only — le problème (2 exécuteurs ÉCRIVAINS concurrents) n'existe
pas. Pattern [[LRN-080]] : ne pas construire sans menace mesurée.
TRIGGER = le jour où un flow dispatche ≥2 exécuteurs écrivains en parallèle :
(1) FILE SCOPE du contrat = déclaration OWNS (champ existant, zéro format
neuf) ; (2) ~40 l dans gates.sh ou lib/owns.sh — intersection CONSERVATRICE
des FILE SCOPE des contrats actifs avant fan-out, conflit possible → refus +
dispatch séquentiel (pas de locks disque tant que l'orchestrateur est
unique) ; (3) locks + tests.
+22
View File
@@ -6,6 +6,28 @@ Format follows [Keep a Changelog](https://keepachangelog.com/).
## [Unreleased] ## [Unreleased]
### Added
- **Contract gates — deterministic floor under the fresh verifier (BDR-083)** —
an acceptance criterion can now carry an oracle (`CHECK:` command +
`EXPECT:` success-only marker + `EVIDENCE:` slot). `lib/gates.sh run
<contract>` executes them fail-closed — MET requires exit 0 **and** the
marker — and writes the outcome back into the contract, so the fresh
verifier reads evidence as fact instead of trusting the executor's report.
New `GATE 0` in `lib/verify-secure-loop.md` runs the floor before any
verifier is dispatched: a red build no longer costs an LLM dispatch to
discover. `ABANDON: <id> <reason>` makes an impossible criterion a visible
handoff that blocks `CONFORME` and routes to the human gate (new verifier
verdict `ABANDONED(n)`). `feater` and `bugfixer` gain a four-pass
completion discipline, scoped so it can never widen the contract.
Adapted from the `unlazy` skill (Leonxlnx/unlazy, MIT); its Stop hook,
approval store, `.unlazy/` tree, depth-tree arithmetic and Node checker
were deliberately refused — see BDR-083 for each reason.
The four orchestrator skills (`feat`, `bugfix`, `ship-feature`,
`init-project`) restate the GATE 0 bullet ahead of GATE 1 (locked);
hotfix explicitly runs no floor. Behavioral RED: 16/16 fresh unprimed
runs followed the new doctrine (EVAL-027).
64 new assertions in `lib/tests/gates.test.sh`.
### Changed ### Changed
- **seo-analyzer + geo-analyzer de-prescribed for Opus 5 (BDR-082)** — - **seo-analyzer + geo-analyzer de-prescribed for Opus 5 (BDR-082)** —
process choreography converted to when-guidance under an process choreography converted to when-guidance under an
+18
View File
@@ -46,6 +46,24 @@ Every choice was made in the plan or is a NEED-DECISION to report.
security/verifier dispatch, editing `.claude/**` or memory registries, user security/verifier dispatch, editing `.claude/**` or memory registries, user
questions (you cannot ask — report instead), attribution trailers of any kind. questions (you cannot ask — report instead), attribution trailers of any kind.
## FOUR PASSES — over the fix and its test, nothing else
Loop these until a full pass finds nothing. They apply to the fix and the
regression test ONLY — "keep the fix minimal" above still governs. They make
the minimal fix COMPLETE; they never widen it.
1. **Complete.** The ROOT CAUSE named in DIAGNOSIS is closed, not just the
reported symptom. No placeholder, no deferred remainder.
2. **Expert reread.** Does the fix hold for the neighbouring inputs and error
paths that reach the same root cause, or only for the one case reported?
3. **Negative control.** Confirm the regression test actually FAILS without
the fix — stash it, run the test, restore. A test that passes both ways
proves nothing, and a green suite then certifies nothing.
4. **Polish.** Naming and comments on what you touched. Nothing else.
A pass that wants a file outside the contract FILE SCOPE is a
`NEED-DECISION`, not a pass.
## OUTPUT — end with exactly this report (your final message) ## OUTPUT — end with exactly this report (your final message)
``` ```
+19
View File
@@ -57,6 +57,25 @@ report below is optional on this path (the dispatcher needs the edit applied
editing `.claude/**` or memory registries, user questions (you cannot editing `.claude/**` or memory registries, user questions (you cannot
ask — report instead), attribution trailers of any kind. ask — report instead), attribution trailers of any kind.
## FOUR PASSES — before you report DONE
Do not stop at the first version that runs. Loop these until a full pass
finds nothing:
1. **Complete.** The whole deliverable the plan names is implemented. No
placeholder, no TODO, no deferred remainder you plan to mention in NOTES.
2. **Expert reread.** Read it as someone who owns this codebase. Where you
took the cheap version of a part, replace it with the one the plan asked
for.
3. **Defect hunt.** Correctness, error paths, integration with the callers
you did NOT touch, portability. Fix what you find.
4. **Polish.** Low-cost only: naming, comment density, dead code you
introduced.
Every pass stays inside the plan and the contract FILE SCOPE. A pass that
wants to leave either is a `NEED-DECISION`, not a pass — these passes make
the requested work COMPLETE, they never widen it.
## OUTPUT — end with exactly this report (your final message) ## OUTPUT — end with exactly this report (your final message)
``` ```
+40 -5
View File
@@ -48,6 +48,25 @@ Rules: read the diff AND enough surrounding code to judge behavior; run
criterion. Never mark `MET` from naming, comments, or plausibility — only criterion. Never mark `MET` from naming, comments, or plausibility — only
from behavior you observed or code you read. from behavior you observed or code you read.
### Criteria carrying an oracle (`CHECK:` / `EXPECT:` / `EVIDENCE:`)
`lib/gates.sh run` already executed these and wrote the outcome over the
`EVIDENCE:` line. Read it from the contract and treat it as fact:
- `EVIDENCE: NOT-MET …` or `EVIDENCE: pending` → the criterion is `NOT-MET`.
Reading the code NEVER overrides a red or unrun oracle. Cite the evidence
line as your evidence.
- `EVIDENCE: MET …` → the declared command passed. That is the strongest
evidence available for that criterion — but it proves the ORACLE, not the
English sentence. Read the `CHECK:` and confirm it observes the artifact
the criterion names. A vacuous oracle (`1. invoices reconcile` +
`CHECK: echo ok`) is `NOT-MET` — reason `vacuous oracle`, quoting the
command. That judgement is yours alone; no command can make it.
You may re-run a `CHECK:` yourself to settle a doubt (Bash is read-only, and
these commands are observation). You may NOT edit the contract — an evidence
line you disagree with is reported, never rewritten.
## STEP 3 — SCOPE CHECK ## STEP 3 — SCOPE CHECK
List the files actually touched (`git diff --name-only` over `DIFF`). List the files actually touched (`git diff --name-only` over `DIFF`).
@@ -58,19 +77,30 @@ only enters the contract through a human micro-gate.
## STEP 4 — VERDICT ## STEP 4 — VERDICT
`CONFORME` ⇔ ALL criteria `MET` AND zero out-of-scope files. Read the contract's `ABANDON:` lines. An abandoned criterion is `ABANDONED`
Anything else is `ECARTS(n)` where n = count(NOT-MET) + count(UNVERIFIABLE) — never `MET`, never counted as a gap the dev can close.
+ count(out-of-scope files).
Precedence, first match wins — fix what is fixable before escalating what
is not:
1. `ERROR(<reason>)` — the contract is missing or unreadable.
2. `ECARTS(n)` — n = count(NOT-MET) + count(UNVERIFIABLE) + count(out-of-scope
files). Surface any abandonment in the same report.
3. `ABANDONED(n)` — zero gaps remain, but n abandonments stand. This is NOT
a pass and NOT a dev loop: it routes straight to the human gate.
4. `CONFORME` — ALL criteria `MET`, zero out-of-scope files, zero
abandonments.
## OUTPUT (exact format — machine-parsed by the orchestrator) ## OUTPUT (exact format — machine-parsed by the orchestrator)
``` ```
VERIFY — VERDICT: CONFORME | ECARTS(n) | ERROR(<reason>) VERIFY — VERDICT: CONFORME | ECARTS(n) | ABANDONED(n) | ERROR(<reason>)
CONTRACT: <path> CONTRACT: <path>
CRITERIA: CRITERIA:
1. <criterion> — MET — <evidence file:line | test ran → result> 1. <criterion> — MET — <EVIDENCE line | file:line | test ran → result>
2. <criterion> — NOT-MET — expected <…> / actual <…> — <file:line> 2. <criterion> — NOT-MET — expected <…> / actual <…> — <file:line>
3. <criterion> — UNVERIFIABLE — <reason> 3. <criterion> — UNVERIFIABLE — <reason>
4. <criterion> — ABANDONED — <the reason recorded in the contract>
SCOPE: in-scope <n> files; out-of-scope: <list | none> SCOPE: in-scope <n> files; out-of-scope: <list | none>
PROOF: read <n> files, ran <cmd → result | nothing>, checked <n>/<n> criteria PROOF: read <n> files, ran <cmd → result | nothing>, checked <n>/<n> criteria
``` ```
@@ -82,6 +112,8 @@ PROOF: read <n> files, ran <cmd → result | nothing>, checked <n>/<n> criteria
- `UNVERIFIABLE` ≠ `MET`. A criterion you did not check is `UNVERIFIABLE`, - `UNVERIFIABLE` ≠ `MET`. A criterion you did not check is `UNVERIFIABLE`,
never silently dropped: the checked count in `PROOF` must equal the never silently dropped: the checked count in `PROOF` must equal the
contract's criteria count. contract's criteria count.
- `ABANDONED` ≠ `MET`. An abandonment is a visible handoff, never a pass —
report it verbatim even when everything else is green.
- `PROOF` is MANDATORY. A `CONFORME` without a `PROOF` line is invalid — - `PROOF` is MANDATORY. A `CONFORME` without a `PROOF` line is invalid —
the orchestrator discards it as a structural failure (LRN-048: a pass the orchestrator discards it as a structural failure (LRN-048: a pass
must prove it looked). must prove it looked).
@@ -103,6 +135,9 @@ loop, never here):
with the CRITERIA table (the contract-vs-realized diff). with the CRITERIA table (the contract-vs-realized diff).
- Remaining `UNVERIFIABLE` while everything else is MET → direct human - Remaining `UNVERIFIABLE` while everything else is MET → direct human
gate (a dev cannot fix unverifiability). gate (a dev cannot fix unverifiability).
- `ABANDONED(n)` → direct human gate, never a dev loop. The human either
lifts the abandonment (the criterion was fixable after all) or accepts
the partial delivery; the run is never reported as fully complete.
- Structural failure (`ERROR(…)`, missing/duplicated VERDICT line, - Structural failure (`ERROR(…)`, missing/duplicated VERDICT line,
unparsable output, agent crash, `CONFORME` without `PROOF`) → retry unparsable output, agent crash, `CONFORME` without `PROOF`) → retry
ONCE with a fresh verifier; a 2nd structural failure → human ONCE with a fresh verifier; a 2nd structural failure → human
+55 -2
View File
@@ -35,6 +35,38 @@ ask what the repo can answer — verify paths/APIs/behavior yourself first.
this conversation. this conversation.
- FILE SCOPE: paths/zones expected to change, or `repo-wide — <reason>`. - FILE SCOPE: paths/zones expected to change, or `repo-wide — <reason>`.
### ORACLES — a criterion a command can decide carries one
Give such a criterion an indented `CHECK:` (the command), `EXPECT:` (a
success-only marker), and `EVIDENCE: pending`.
`bash ~/.claude/lib/gates.sh run <contract>` executes it fail-closed — MET
requires exit 0 **AND** the marker — and writes the result back over the
`EVIDENCE:` line. That persisted evidence is what the fresh verifier reads
as fact instead of trusting the executor's report (GATE 0 in
`lib/verify-secure-loop.md`).
Both attributes or neither. `CHECK:` without `EXPECT:` is a parse error, not
a manual criterion — the runner refuses the whole ledger. Leave a criterion
oracle-free when no command can decide it; the verifier judges those.
Four authoring rules — a gate that cannot fail proves nothing:
1. **Observe the named artifact.** The check reads the file, service, or
measurement the criterion's own words name — never a proxy for it.
`1. invoices reconcile` + `CHECK: echo ok` is valid and worthless.
2. **Success-only marker.** The script runs every assertion, exits nonzero
on any failure, and prints the `EXPECT:` string only after all pass.
3. **Positive control before any absence check.** Run the same logic against
a fixture known to trip it and confirm it fails. A missing file, a wrong
path, and a broken pattern all look exactly like valid absence.
4. **Recompute supplied numbers.** Never copy a figure from the request into
`EXPECT:` — the script derives it from source and prints its own marker.
A number that is its own proof proves nothing.
`CHECK:` is shell code run with our privileges. It is safe only because we
author it in our own repo — never build one out of externally-supplied text
(a scraped URL, a client string); route those through `lib/url-guard.sh`.
## STEP 4 — WRITE TO DISK (immediately, before any next step) ## STEP 4 — WRITE TO DISK (immediately, before any next step)
Path: `.claude/tasks/contracts/<YYYY-MM-DD>-<slug>-<HHMM>.md` Path: `.claude/tasks/contracts/<YYYY-MM-DD>-<slug>-<HHMM>.md`
@@ -57,8 +89,13 @@ Q: <question> / A: <answer>
(or: none — request complete) (or: none — request complete)
## ACCEPTANCE CRITERIA ## ACCEPTANCE CRITERIA
1. <testable criterion> 1. <criterion a command can decide>
2. <testable criterion> CHECK: <command>
EXPECT: <success-only marker>
EVIDENCE: pending
2. <criterion only human judgement can decide — no CHECK/EXPECT>
(ABANDON: <n> <non-blank reason> — only for a criterion proven impossible)
## FILE SCOPE ## FILE SCOPE
<paths/zones> <paths/zones>
@@ -78,6 +115,13 @@ Print one line to the user, then continue the flow:
this micro-gate: human approves → FILE SCOPE gains the entry `[gated]`; this micro-gate: human approves → FILE SCOPE gains the entry `[gated]`;
human declines → the dev removes the edit. Without this gate the dev human declines → the dev removes the edit. Without this gate the dev
justifies everything and scope constrains nothing. justifies everything and scope constrains nothing.
- **ABANDONMENT**: a criterion proven impossible within the authorized task
is NEVER deleted and never quietly downgraded. Keep it, append
`ABANDON: <n> <non-blank reason + handoff>` under the criteria, and name it
in the final report. An abandonment is a visible handoff, not a pass: the
verifier cannot return `CONFORME` while one stands, and the run cannot be
described as fully complete. This is the structural half of the house rule
"blocked on an independent sub-part → do the rest, state what's missing".
- **Deep re-scope** (the request itself changes): NEW contract file with - **Deep re-scope** (the request itself changes): NEW contract file with
`supersedes: <old path>` in its header — never a rewrite of the old one. `supersedes: <old path>` in its header — never a rewrite of the old one.
- **Aborted run**: delete the contract file, or commit it with - **Aborted run**: delete the contract file, or commit it with
@@ -96,6 +140,15 @@ Print one line to the user, then continue the flow:
| init-project | Full. The interviewer's PROJECT BRIEF pours into the contract (V1 features → criteria). | | init-project | Full. The interviewer's PROJECT BRIEF pours into the contract (V1 features → criteria). |
| onboard | Audit-scope contract (interview answers → what to audit, which axes). | | onboard | Audit-scope contract (interview answers → what to audit, which axes). |
Oracles follow the same proportion. hotfix: none — that flow runs no floor
(and no verifier); the hotfixer runs build/tests itself. feat / bugfix: the
suite criterion at minimum, and for bugfix the regression test the DIAGNOSIS
names — its `CHECK:` runs that test alone, so a green result means the
reproduction actually flipped.
ship-feature / init-project: build, suite, and every criterion a command can
settle. onboard: audit criteria are mostly judgement — leave them oracle-free
rather than invent a check that cannot fail.
## Hand-off rule ## Hand-off rule
Downstream consumers (plan step, dev subagents, verifier) receive the Downstream consumers (plan step, dev subagents, verifier) receive the
+323
View File
@@ -0,0 +1,323 @@
#!/usr/bin/env bash
# Deterministic floor under GATE 1: execute the acceptance criteria that the
# contract itself declares as oracles, fail-closed, and persist the evidence
# INTO the contract file.
#
# bash ~/.claude/lib/gates.sh status <contract> # parse only, never runs
# bash ~/.claude/lib/gates.sh run <contract> # execute + write evidence
#
# rc 0 = MET every runnable criterion passed, no abandonment standing
# 2 = UNMET a runnable criterion failed, or the ledger is malformed
# 3 = ABANDONED runnable criteria all passed, an abandonment still stands
#
# WHY: GATE 1 (lib/verify-secure-loop.md) is an LLM dispatch, and the
# verifier's mandatory `PROOF:` line is a line the verifier WRITES — nothing
# structurally stops it from being produced without anything being executed.
# This runs what the contract declares BEFORE a verifier is ever spawned: a
# red floor sends the executor back for free. Adapted from the `unlazy` skill
# (Leonxlnx/unlazy) — its gate ledger, minus the machinery we do not need.
#
# `run` always re-executes every runnable criterion, including ones already
# recorded MET. Trusting written evidence is exactly the failure this closes,
# so there is no incremental mode to get it wrong with.
#
# TRUST BOUNDARY: `CHECK:` is shell code, run with this process's privileges
# and environment. That is safe here only because the contract is authored by
# our own orchestrator in our own repo — which is why there is no approval
# store (we never execute ledgers inherited from a foreign repo). NEVER build
# a `CHECK:` out of externally-supplied text; route such values through
# lib/url-guard.sh first.
set -uo pipefail
TIMEOUT="${GATES_TIMEOUT:-120}"
EVIDENCE_CAP=140
# Module-level parse tables, index-aligned. Bash has no record type; threading
# eight parallel arrays through every call would cost more readability than
# the explicit data flow buys.
_ID=(); _TEXT=(); _CHECK=(); _EXPECT=(); _EVLINE=(); _EVTEXT=()
_STATUS=(); _EVID=()
_ABANDON_ID=(); _ABANDON_WHY=()
_ERRORS=()
_CUR=-1
_die() { printf 'GATES — VERDICT: ERROR(%s)\n' "$1"; exit 2; }
_err() { _ERRORS+=("$1"); }
_trim() {
local s="$1"
s="${s#"${s%%[![:space:]]*}"}"
printf '%s' "${s%"${s##*[![:space:]]}"}"
}
# ── parse ───────────────────────────────────────────────────────────────────
_new_crit() { # _new_crit <id> <text>
local i
for ((i = 0; i < ${#_ID[@]}; i++)); do
if [ "${_ID[i]}" = "$1" ]; then
_err "duplicate criterion id: $1"
# Orphan what follows instead of aliasing it onto the previous
# criterion, which would hand one gate another gate's oracle.
_CUR=-1
return 0
fi
done
_ID+=("$1"); _TEXT+=("$2")
_CHECK+=(""); _EXPECT+=(""); _EVLINE+=("0"); _EVTEXT+=("")
_CUR=$((${#_ID[@]} - 1))
}
_set_attr() { # _set_attr <CHECK|EXPECT|EVIDENCE> <value> <lineno>
if [ "$_CUR" -lt 0 ]; then
_err "$1 at line $3 belongs to no criterion"
return 0
fi
case "$1" in
CHECK) _CHECK[_CUR]="$2" ;;
EXPECT) _EXPECT[_CUR]="$2" ;;
EVIDENCE) _EVLINE[_CUR]="$3"; _EVTEXT[_CUR]="$2" ;;
esac
}
# An UNINDENTED attribute is diagnosed, never absorbed: silently ignoring it
# would demote a runnable criterion to a manual one, which is the one parse
# bug that turns this checker into a rubber stamp.
_absorb() { # _absorb <raw-line> <lineno>
local body
if [[ "$1" =~ ^([0-9]+)\.[[:space:]]+(.*)$ ]]; then
_new_crit "${BASH_REMATCH[1]}" "${BASH_REMATCH[2]}"
elif [[ "$1" =~ ^ABANDON:[[:space:]]*([0-9]+)?[[:space:]]*(.*)$ ]]; then
_ABANDON_ID+=("${BASH_REMATCH[1]}"); _ABANDON_WHY+=("${BASH_REMATCH[2]}")
elif [[ "$1" =~ ^(CHECK|EXPECT|EVIDENCE): ]]; then
_err "unindented ${BASH_REMATCH[1]}: at line $2"
elif [[ "$1" =~ ^[[:space:]]+(CHECK|EXPECT|EVIDENCE):(.*)$ ]]; then
body="$(_trim "${BASH_REMATCH[2]}")"
_set_attr "${BASH_REMATCH[1]}" "$body" "$2"
fi
}
_parse() { # _parse <file>
local line n=0 fence=0 inblock=0
while IFS= read -r line || [ -n "$line" ]; do
n=$((n + 1))
case "$line" in '```'*) fence=$((1 - fence)); continue ;; esac
[ "$fence" -eq 1 ] && continue
case "$line" in
'## ACCEPTANCE CRITERIA'*) inblock=1; continue ;;
'## '*) inblock=0; continue ;;
esac
[ "$inblock" -eq 1 ] && _absorb "$line" "$n"
done < "$1"
}
# ── validation ──────────────────────────────────────────────────────────────
_validate_oracles() {
local i
for ((i = 0; i < ${#_ID[@]}; i++)); do
if [ -n "${_CHECK[i]}" ] && [ -z "${_EXPECT[i]}" ]; then
_err "criterion ${_ID[i]}: CHECK without EXPECT (partial oracle)"
elif [ -z "${_CHECK[i]}" ] && [ -n "${_EXPECT[i]}" ]; then
_err "criterion ${_ID[i]}: EXPECT without CHECK (partial oracle)"
elif [ -n "${_CHECK[i]}" ] && [ "${_EVLINE[i]}" = "0" ]; then
_err "criterion ${_ID[i]}: runnable but has no EVIDENCE: line"
fi
done
}
_validate_abandons() {
local i j found
for ((i = 0; i < ${#_ABANDON_ID[@]}; i++)); do
found=0
for ((j = 0; j < ${#_ID[@]}; j++)); do
[ "${_ID[j]}" = "${_ABANDON_ID[i]}" ] && found=1
done
[ "$found" -eq 1 ] ||
_err "ABANDON names unknown criterion: '${_ABANDON_ID[i]}'"
[ -n "$(_trim "${_ABANDON_WHY[i]}")" ] ||
_err "ABANDON ${_ABANDON_ID[i]}: blank reason (a handoff needs one)"
done
}
_is_abandoned() { # _is_abandoned <criterion-id>
local i
for ((i = 0; i < ${#_ABANDON_ID[@]}; i++)); do
[ "${_ABANDON_ID[i]}" = "$1" ] && return 0
done
return 1
}
# ── execution ───────────────────────────────────────────────────────────────
# One line, capped, newlines flattened: the smallest output that proves the
# outcome. Full logs stay in the terminal, never in the contract.
_decisive() { # _decisive <combined-output>
local flat
flat="$(printf '%s' "$1" | tr '\n\r\t' ' ' | tr -s ' ')"
flat="$(_trim "$flat")"
if [ "${#flat}" -gt "$EVIDENCE_CAP" ]; then
printf '%s…' "${flat:0:$EVIDENCE_CAP}"
else
printf '%s' "$flat"
fi
}
# Fail-closed: exit 0 AND the marker. A nonzero process never passes because
# its error text happens to contain the expected token.
_run_one() { # _run_one <idx>
local i="$1" out rc
out="$(timeout "$TIMEOUT" bash -c "${_CHECK[i]}" 2>&1)"
rc=$?
_STATUS[i]="NOT-MET"
if [ "$rc" -eq 124 ]; then
_EVID[i]="NOT-MET timeout=${TIMEOUT}s"
elif [ "$rc" -ne 0 ]; then
_EVID[i]="NOT-MET exit=$rc (nonzero) :: $(_decisive "$out")"
elif [[ "$out" != *"${_EXPECT[i]}"* ]]; then
_EVID[i]="NOT-MET exit=0 marker-absent :: $(_decisive "$out")"
else
_STATUS[i]="MET"
_EVID[i]="MET exit=0 marker-found :: $(_decisive "$out")"
fi
}
_run_all() {
local i
for ((i = 0; i < ${#_ID[@]}; i++)); do
_STATUS[i]=""; _EVID[i]=""
[ -n "${_CHECK[i]}" ] && _run_one "$i"
done
}
_evline_owner() { # _evline_owner <lineno> — echoes idx, or nothing
local i
for ((i = 0; i < ${#_ID[@]}; i++)); do
if [ "${_EVLINE[i]}" = "$1" ] && [ -n "${_EVID[i]}" ]; then
printf '%s' "$i"
return 0
fi
done
}
# Rewrites only the EVIDENCE lines of criteria that actually ran; every other
# byte of the contract is copied through, indentation included.
_write_back() { # _write_back <file>
local tmp line n=0 idx
tmp="$(mktemp)" || _die "mktemp failed"
while IFS= read -r line || [ -n "$line" ]; do
n=$((n + 1))
idx="$(_evline_owner "$n")"
if [ -n "$idx" ]; then
printf '%s%s\n' "${line%%[![:space:]]*}" "EVIDENCE: ${_EVID[idx]}"
else
printf '%s\n' "$line"
fi
done < "$1" > "$tmp"
cat "$tmp" > "$1" && rm -f "$tmp"
}
# ── report ──────────────────────────────────────────────────────────────────
# A recorded `pending`, or a criterion that never ran, is PENDING — never MET.
# `status` reports what the file says; it does not revalidate old evidence.
_row_state() { # _row_state <idx>
local i="$1"
_is_abandoned "${_ID[i]}" && { printf 'ABANDONED'; return 0; }
[ -z "${_CHECK[i]}" ] && { printf 'MANUAL'; return 0; }
[ -n "${_STATUS[i]:-}" ] && { printf '%s' "${_STATUS[i]}"; return 0; }
case "${_EVTEXT[i]}" in
MET' '*) printf 'MET-RECORDED' ;;
*) printf 'PENDING' ;;
esac
}
_report_rows() {
local i state
for ((i = 0; i < ${#_ID[@]}; i++)); do
state="$(_row_state "$i")"
printf ' %-3s %-13s %s\n' "${_ID[i]}" "$state" "${_TEXT[i]}"
done
}
_report_abandons() {
local i
for ((i = 0; i < ${#_ABANDON_ID[@]}; i++)); do
printf ' ABANDONED %s — %s\n' "${_ABANDON_ID[i]}" "${_ABANDON_WHY[i]}"
done
}
_count_state() { # _count_state <state>
local i n=0
for ((i = 0; i < ${#_ID[@]}; i++)); do
[ "$(_row_state "$i")" = "$1" ] && n=$((n + 1))
done
printf '%s' "$n"
}
_verdict() { # _verdict <mode> — prints the line, returns the rc
local unmet pending abandoned
if [ "${#_ERRORS[@]}" -gt 0 ]; then
printf 'GATES — VERDICT: ERROR(%s)\n' "${#_ERRORS[@]}"
return 2
fi
unmet="$(_count_state NOT-MET)"
pending="$(_count_state PENDING)"
abandoned="$(_count_state ABANDONED)"
[ "$unmet" -gt 0 ] &&
{ printf 'GATES — VERDICT: UNMET(%s)\n' "$unmet"; return 2; }
if [ "$1" = "status" ] && [ "$pending" -gt 0 ]; then
printf 'GATES — VERDICT: PENDING(%s)\n' "$pending"
return 2
fi
[ "$abandoned" -gt 0 ] &&
{ printf 'GATES — VERDICT: ABANDONED(%s)\n' "$abandoned"; return 3; }
printf 'GATES — VERDICT: MET\n'
return 0
}
_report() { # _report <mode> <file>
local rc
printf 'GATES — %s (%s)\n' "$2" "$1"
_report_rows
_report_abandons
[ "${#_ERRORS[@]}" -gt 0 ] && printf ' ERROR %s\n' "${_ERRORS[@]}"
printf 'RUNNABLE: %s of %s criteria; timeout %ss\n' \
"$(_runnable_count)" "${#_ID[@]}" "$TIMEOUT"
_verdict "$1"
rc=$?
return "$rc"
}
_runnable_count() {
local i n=0
for ((i = 0; i < ${#_ID[@]}; i++)); do
[ -n "${_CHECK[i]}" ] && n=$((n + 1))
done
printf '%s' "$n"
}
# ── entry point ─────────────────────────────────────────────────────────────
main() { # main <status|run> <contract>
local mode="$1" file="$2"
[ -r "$file" ] || _die "contract unreadable: $file"
_parse "$file"
[ "${#_ID[@]}" -gt 0 ] ||
_die "no numbered criteria under ## ACCEPTANCE CRITERIA"
_validate_oracles
_validate_abandons
if [ "$mode" = "run" ] && [ "${#_ERRORS[@]}" -eq 0 ]; then
_run_all
_write_back "$file"
fi
_report "$mode" "$file"
}
case "${1:-}" in
status|run)
[ $# -eq 2 ] || _die "usage: gates.sh {status|run} <contract-path>"
main "$1" "$2"
;;
*) _die "usage: gates.sh {status|run} <contract-path>" ;;
esac
+1 -1
View File
@@ -67,7 +67,7 @@ fi
tr_ "frontmatter name" "$AGT" "^name: verifier$" tr_ "frontmatter name" "$AGT" "^name: verifier$"
tr_ "tools read-only set" "$AGT" "^tools: Read, Grep, Glob, Bash$" tr_ "tools read-only set" "$AGT" "^tools: Read, Grep, Glob, Bash$"
tn "no write-capable tools" "$AGT" "^tools:.*(Edit|Write|NotebookEdit)" tn "no write-capable tools" "$AGT" "^tools:.*(Edit|Write|NotebookEdit)"
tf "verdict grammar" "$AGT" "VERIFY — VERDICT: CONFORME | ECARTS(n) | ERROR(<reason>)" tf "verdict grammar" "$AGT" "VERIFY — VERDICT: CONFORME | ECARTS(n) | ABANDONED(n) | ERROR(<reason>)"
tf "blind — no iteration history" "$AGT" "NEVER receive iteration history" tf "blind — no iteration history" "$AGT" "NEVER receive iteration history"
tf "blind — complete every time" "$AGT" "every verification is complete and blind" tf "blind — complete every time" "$AGT" "every verification is complete and blind"
tf "unverifiable is not met" "$AGT" "\`UNVERIFIABLE\` ≠ \`MET\`" tf "unverifiable is not met" "$AGT" "\`UNVERIFIABLE\` ≠ \`MET\`"
+317
View File
@@ -0,0 +1,317 @@
#!/usr/bin/env bash
# ============================================================
# lib/gates.sh — behavioural tests + structure locks for the
# deterministic floor (GATE 0, lib/verify-secure-loop.md).
#
# Fail-closed is the entire point of this runner, so every
# "looks green but must not pass" case is asserted explicitly:
# nonzero exit carrying the marker, marker absent, timeout,
# unindented attribute silently demoting a gate to manual.
# Non-execution is proved with a sentinel file, and the
# sentinel's own positive control is asserted first — an
# absence check that was never able to fire proves nothing.
# ============================================================
set -uo pipefail
REPO="$(cd "$(dirname "$0")/../.." && pwd)"
GATES="$REPO/lib/gates.sh"
WORK="$(mktemp -d)"
trap 'rm -rf "$WORK"' EXIT
PASS=0; FAIL=0; N=0
LAST=""
ok() { echo " PASS $1"; PASS=$((PASS + 1)); }
bad() { echo " FAIL $1 — $2"; FAIL=$((FAIL + 1)); }
# gate <label> <mode> <expected-verdict> <expected-rc> <<< fixture-on-stdin
gate() {
local label="$1" mode="$2" want="$3" wantrc="$4" out rc
N=$((N + 1)); LAST="$WORK/c$N.md"
cat > "$LAST"
out="$(GATES_TIMEOUT="${GATES_TIMEOUT:-120}" \
bash "$GATES" "$mode" "$LAST" 2>&1)"
rc=$?
if [[ "$out" == *"$want"* ]] && [ "$rc" -eq "$wantrc" ]; then
ok "$label"
else
bad "$label" "want '$want' rc=$wantrc, got rc=$rc"
printf '%s\n' "$out" | sed 's/^/ /'
fi
}
has() {
if grep -qF -- "$2" "$LAST"; then ok "$1"; else bad "$1" "missing: $2"; fi
}
exists() {
if [ -e "$1" ]; then ok "$2"; else bad "$2" "sentinel absent: $1"; fi
}
absent() {
if [ -e "$1" ]; then bad "$2" "sentinel created: $1"; else ok "$2"; fi
}
echo "── fail-closed execution ──"
gate "exit 0 + marker = MET" run "GATES — VERDICT: MET" 0 <<'EOF'
## ACCEPTANCE CRITERIA
1. green
CHECK: echo "MARKER-OK"
EXPECT: MARKER-OK
EVIDENCE: pending
EOF
has "evidence written back" "EVIDENCE: MET exit=0 marker-found"
# The case a naive checker gets wrong: the marker IS in the output, but the
# process failed. Substring matching alone would certify a broken build.
gate "nonzero exit + marker = UNMET" run "GATES — VERDICT: UNMET(1)" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. lies
CHECK: echo "MARKER-OK"; exit 7
EXPECT: MARKER-OK
EVIDENCE: pending
EOF
has "nonzero recorded honestly" "NOT-MET exit=7 (nonzero)"
gate "exit 0 + no marker = UNMET" run "GATES — VERDICT: UNMET(1)" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. silent success is not success
CHECK: echo "something else"
EXPECT: MARKER-OK
EVIDENCE: pending
EOF
has "marker-absent recorded" "NOT-MET exit=0 marker-absent"
GATES_TIMEOUT=1 gate "timeout = UNMET" run "GATES — VERDICT: UNMET(1)" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. hangs
CHECK: sleep 5
EXPECT: never
EVIDENCE: pending
EOF
has "timeout recorded" "NOT-MET timeout=1s"
gate "manual-only contract passes through" run "RUNNABLE: 0 of 2" 0 <<'EOF'
## ACCEPTANCE CRITERIA
1. a human reads the copy
2. the design matches the brief
EOF
echo "── non-execution (sentinel), positive control first ──"
# Positive control: prove the sentinel mechanism can fire at all.
gate "sentinel fires when a CHECK runs" run "GATES — VERDICT: MET" 0 <<EOF
## ACCEPTANCE CRITERIA
1. control
CHECK: touch "$WORK/fired"; echo "M"
EXPECT: M
EVIDENCE: pending
EOF
exists "$WORK/fired" "positive control: sentinel created"
gate "status never executes" status "GATES — VERDICT: PENDING(1)" 2 <<EOF
## ACCEPTANCE CRITERIA
1. must not run
CHECK: touch "$WORK/status-ran"; echo "M"
EXPECT: M
EVIDENCE: pending
EOF
absent "$WORK/status-ran" "status did not execute"
has "status did not write evidence" "EVIDENCE: pending"
gate "fenced example is not a gate" run "RUNNABLE: 1 of 1" 0 <<EOF
## ACCEPTANCE CRITERIA
1. real
CHECK: echo "R"
EXPECT: R
EVIDENCE: pending
\`\`\`markdown
2. documentation example, invisible to the parser
CHECK: touch "$WORK/fenced-ran"; echo "nope"
EXPECT: nope
EVIDENCE: pending
\`\`\`
EOF
absent "$WORK/fenced-ran" "fenced CHECK never executed"
gate "malformed ledger executes nothing" run "GATES — VERDICT: ERROR" 2 <<EOF
## ACCEPTANCE CRITERIA
1. would run if the ledger parsed
CHECK: touch "$WORK/malformed-ran"; echo "M"
EXPECT: M
EVIDENCE: pending
2. partial oracle poisons the whole ledger
CHECK: echo "x"
EVIDENCE: pending
EOF
absent "$WORK/malformed-ran" "malformed ledger did not execute"
has "malformed ledger not written" "EVIDENCE: pending"
echo "── parse strictness ──"
gate "CHECK without EXPECT" run "CHECK without EXPECT" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. partial
CHECK: echo x
EVIDENCE: pending
EOF
gate "EXPECT without CHECK" run "EXPECT without CHECK" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. partial
EXPECT: x
EVIDENCE: pending
EOF
# An unindented CHECK must be diagnosed, never absorbed: silently ignoring it
# demotes a runnable criterion to a manual one — the one parse bug that turns
# this runner into a rubber stamp.
gate "unindented attribute is diagnosed" run "unindented CHECK:" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. sneaky
CHECK: echo x
EXPECT: x
EVIDENCE: pending
EOF
gate "runnable without EVIDENCE line" run "has no EVIDENCE: line" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. no ledger slot
CHECK: echo x
EXPECT: x
EOF
gate "duplicate criterion id" run "duplicate criterion id: 1" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. first
EVIDENCE: pending
1. second
EVIDENCE: pending
EOF
# After a rejected duplicate the following attributes must be orphaned, not
# aliased onto the previous criterion — that would hand one gate another's
# oracle and let a stale EVIDENCE line satisfy it.
gate "duplicate orphans what follows" run "belongs to no criterion" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. real
CHECK: echo x
EXPECT: x
EVIDENCE: pending
1. duplicate
CHECK: echo y
EXPECT: y
EVIDENCE: pending
EOF
gate "no numbered criteria" run "no numbered criteria" 2 <<'EOF'
## ACCEPTANCE CRITERIA
nothing numbered here
EOF
echo "── abandonment ──"
gate "valid abandonment = rc 3" run "GATES — VERDICT: ABANDONED(1)" 3 <<'EOF'
## ACCEPTANCE CRITERIA
1. green
CHECK: echo "M"
EXPECT: M
EVIDENCE: pending
2. impossible
EVIDENCE: pending
ABANDON: 2 upstream API offline; handoff recorded in BLK-099
EOF
gate "blank abandonment reason" run "blank reason" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. green
EVIDENCE: pending
ABANDON: 1
EOF
gate "abandonment naming nothing" run "unknown criterion" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. green
EVIDENCE: pending
ABANDON: 9 names a criterion that does not exist
EOF
echo "── usage ──"
# usage <label> <expected-substring> <argv...>
usage() {
local label="$1" want="$2" out rc; shift 2
out="$(bash "$GATES" "$@" 2>&1)"; rc=$?
if [ "$rc" -eq 2 ] && [[ "$out" == *"$want"* ]]; then
ok "$label"
else
bad "$label" "rc=$rc out=$out"
fi
}
usage "no args = ERROR rc 2" "usage:"
usage "missing contract = ERROR rc 2" "contract unreadable" run "$WORK/nope.md"
usage "unknown mode refused" "usage:" frobnicate "$WORK/c1.md"
# ── structure locks on the doctrine this runner is wired into ───────────────
CI="$REPO/lib/contract-interview.md"
VS="$REPO/lib/verify-secure-loop.md"
AGT="$REPO/agents/verifier.md"
FE="$REPO/agents/feater.md"
BF="$REPO/agents/bugfixer.md"
lock() { # lock <label> <file> <fixed-string>
if grep -qF -- "$3" "$2" 2>/dev/null; then
ok "$1"
else
bad "$1" "missing: $3"
fi
}
echo "── contract-interview.md oracle doctrine ──"
lock "oracle section" "$CI" "### ORACLES"
lock "runner named" "$CI" "lib/gates.sh run <contract>"
lock "fail-closed spelled out" "$CI" "exit 0 **AND** the marker"
lock "both or neither" "$CI" "Both attributes or neither"
lock "rule observe artifact" "$CI" "Observe the named artifact"
lock "rule success-only" "$CI" "Success-only marker"
lock "rule positive control" "$CI" "Positive control before any absence check"
lock "rule recompute numbers" "$CI" "Recompute supplied numbers"
lock "shell trust boundary" "$CI" "url-guard.sh"
lock "template carries oracle" "$CI" "EXPECT: <success-only marker>"
lock "abandonment lifecycle" "$CI" "**ABANDONMENT**"
lock "abandonment not deleted" "$CI" "NEVER deleted"
echo "── verify-secure-loop.md GATE 0 ──"
lock "gate 0 exists" "$VS" "## GATE 0 — DETERMINISTIC FLOOR"
lock "gate 0 no dispatch" "$VS" "**No verifier is dispatched**"
lock "gate 0 loop bound" "$VS" "Max 3 floor iterations"
lock "gate 0 budget separate" "$VS" "not eat the conformity budget"
lock "malformed = main loop" "$VS" "never dispatch a dev for it"
lock "order invariant" "$VS" "GATE 0 → GATE 1 → GATE 2"
echo "── verifier.md oracle + abandonment ──"
lock "verdict grammar" "$AGT" \
"VERIFY — VERDICT: CONFORME | ECARTS(n) | ABANDONED(n) | ERROR(<reason>)"
lock "red oracle wins" "$AGT" "NEVER overrides a red or unrun oracle"
lock "vacuous oracle caught" "$AGT" "vacuous oracle"
lock "oracle != english" "$AGT" "proves the ORACLE, not the"
lock "never edits contract" "$AGT" "reported, never rewritten"
lock "abandoned is not met" "$AGT" "\`ABANDONED\` ≠ \`MET\`"
lock "abandoned routes human" "$AGT" "direct human gate, never a dev loop"
echo "── executor four passes ──"
lock "feater passes" "$FE" "## FOUR PASSES"
lock "feater no placeholder" "$FE" "no deferred remainder you plan"
lock "feater never widens" "$FE" "they never widen it"
lock "bugfixer passes" "$BF" "## FOUR PASSES"
lock "bugfixer stays minimal" "$BF" "keep the fix minimal"
lock "bugfixer neg control" "$BF" "**Negative control.**"
lock "bugfixer test must fail" "$BF" "A test that passes both ways"
echo ""
echo "gates: $PASS pass, $FAIL fail"
[ "$FAIL" -eq 0 ]
+2
View File
@@ -27,6 +27,7 @@ tf "shf enrich at gate" "$SHF" "ENRICH the STEP 0e contract"
tf "shf gated marker" "$SHF" "[gated <date>]" tf "shf gated marker" "$SHF" "[gated <date>]"
tf "shf verify+secure step" "$SHF" "STEP 5 — VERIFY + SECURE" tf "shf verify+secure step" "$SHF" "STEP 5 — VERIFY + SECURE"
tf "shf uses shared include" "$SHF" "lib/verify-secure-loop.md" tf "shf uses shared include" "$SHF" "lib/verify-secure-loop.md"
tf "shf gate0 floor" "$SHF" "GATE 0 — deterministic floor"
tf "shf judges enriched" "$SHF" "ENRICHED contract" tf "shf judges enriched" "$SHF" "ENRICHED contract"
tf "shf orthogonal to review" "$SHF" "DISTINCT axis from STEP 6 code review" tf "shf orthogonal to review" "$SHF" "DISTINCT axis from STEP 6 code review"
@@ -36,6 +37,7 @@ tf "ini criteria from V1" "$INI" "V1 FEATURES (each testable)"
tf "ini enrich at gate1" "$INI" "ENRICH the STEP 1 contract" tf "ini enrich at gate1" "$INI" "ENRICH the STEP 1 contract"
tf "ini verify+secure step" "$INI" "STEP 9 — VERIFY + SECURE" tf "ini verify+secure step" "$INI" "STEP 9 — VERIFY + SECURE"
tf "ini uses shared include" "$INI" "lib/verify-secure-loop.md" tf "ini uses shared include" "$INI" "lib/verify-secure-loop.md"
tf "ini gate0 floor" "$INI" "GATE 0 — deterministic floor"
tf "ini adds security gate" "$INI" "adds the security gate init-project previously lacked" tf "ini adds security gate" "$INI" "adds the security gate init-project previously lacked"
echo "-- onboard (explicit NO-LOOP audit) --" echo "-- onboard (explicit NO-LOOP audit) --"
+2
View File
@@ -57,6 +57,7 @@ tf "feat contract step" "$FSK" "STEP 0.7 — CONTRACT"
tf "feat contract-interview" "$FSK" "lib/contract-interview.md" tf "feat contract-interview" "$FSK" "lib/contract-interview.md"
tf "feat verify+secure step" "$FSK" "STEP 4 — VERIFY + SECURE" tf "feat verify+secure step" "$FSK" "STEP 4 — VERIFY + SECURE"
tf "feat uses shared include" "$FSK" "lib/verify-secure-loop.md" tf "feat uses shared include" "$FSK" "lib/verify-secure-loop.md"
tf "feat gate0 floor" "$FSK" "GATE 0 — deterministic floor"
tf "feat nominal 1+1 dispatch" "$FSK" "verifier + one security dispatch" tf "feat nominal 1+1 dispatch" "$FSK" "verifier + one security dispatch"
tf "feat dispatches feater" "$FSK" 'subagent_type="feater"' tf "feat dispatches feater" "$FSK" 'subagent_type="feater"'
@@ -65,6 +66,7 @@ tf "bug contract step" "$BSK" "STEP 3.5 — CONTRACT"
tf "bug diagnosis feeds it" "$BSK" "feeds it: REQUEST verbatim" tf "bug diagnosis feeds it" "$BSK" "feeds it: REQUEST verbatim"
tf "bug fresh gates" "$BSK" "the two fresh gates per" tf "bug fresh gates" "$BSK" "the two fresh gates per"
tf "bug uses shared include" "$BSK" "lib/verify-secure-loop.md" tf "bug uses shared include" "$BSK" "lib/verify-secure-loop.md"
tf "bug gate0 floor" "$BSK" "GATE 0 — deterministic floor"
tf "bug dispatches bugfixer" "$BSK" 'subagent_type="bugfixer"' tf "bug dispatches bugfixer" "$BSK" 'subagent_type="bugfixer"'
echo "── agents/bugfixer.md (bugfix executor — sonnet, no Agent) ──" echo "── agents/bugfixer.md (bugfix executor — sonnet, no Agent) ──"
+54 -12
View File
@@ -13,8 +13,40 @@ Inputs the caller must have ready:
pre-dev SHA, or the working-tree diff before commit). pre-dev SHA, or the working-tree diff before commit).
- `TEST`: the project test command, if known. - `TEST`: the project test command, if known.
Nominal path is cheap: one verifier dispatch + one security dispatch, done. Nominal path is cheap — a free floor run, then
The loop only costs more when it actually loops. one verifier dispatch + one security dispatch, done. The loop only costs
more when it actually loops.
## GATE 0 — DETERMINISTIC FLOOR (no dispatch, no model)
Before spending a verifier dispatch, execute the oracles the contract itself
declares:
```bash
bash ~/.claude/lib/gates.sh run "$CONTRACT"
```
It runs every `CHECK:` fail-closed (MET requires exit 0 AND the `EXPECT:`
marker) and writes the outcome back over each `EVIDENCE:` line. Parse its
single `GATES — VERDICT:` line:
- `MET` → floor green, go to GATE 1. An all-manual contract lands here too
(`RUNNABLE: 0 of n`) and passes straight through.
- `UNMET(n)` → hand the dev the CONTRACT path + the `NOT-MET` rows verbatim,
nothing else; re-run GATE 0. **No verifier is dispatched** — a red build or
a red suite is not a judgement call, and paying an LLM to discover it is
waste. **Max 3 floor iterations** → STOP + human escalation with the rows.
- `ABANDONED(n)` → floor green but a handoff stands. Continue to GATE 1; the
verifier surfaces it and its `ABANDONED(n)` verdict routes to the human
gate.
- `ERROR(n)` → the ledger is malformed (partial oracle, duplicate id,
unindented attribute, runnable criterion with no `EVIDENCE:` line). The
contract is the ORCHESTRATOR's own artifact — fix it here in the main loop,
never dispatch a dev for it.
Floor iterations are counted separately from GATE 1's: a cheap loop here does
not eat the conformity budget. GATE 0 also runs unchanged after every
security fix round, before re-verifying the request.
## GATE 1 — REQUEST CONFORMITY (fresh verifier) ## GATE 1 — REQUEST CONFORMITY (fresh verifier)
@@ -29,9 +61,13 @@ Parse its single `VERIFY — VERDICT:` line:
- `ECARTS(n)` → hand the dev the CONTRACT path + the exact `CRITERIA` gap - `ECARTS(n)` → hand the dev the CONTRACT path + the exact `CRITERIA` gap
lines (NOT-MET / out-of-scope), nothing else. Inline dev fixes in place; lines (NOT-MET / out-of-scope), nothing else. Inline dev fixes in place;
a dispatched dev is re-dispatched FRESH with those inputs only. Then a dispatched dev is re-dispatched FRESH with those inputs only. Then
re-dispatch a FRESH verifier. Repeat. **Max 3 conformity iterations** → re-run GATE 0 and re-dispatch a FRESH verifier. Repeat.
STOP + human escalation with the CRITERIA table (the contract-vs-realized **Max 3 conformity iterations** → STOP + human escalation with the
diff). CRITERIA table (the contract-vs-realized diff).
- `ABANDONED(n)` → direct human gate, never a dev loop (a dev cannot close
what was proven impossible). The human lifts the abandonment or accepts
the partial delivery; either way the run is never reported as fully
complete, and the abandonment is named in the final report.
- Remaining `UNVERIFIABLE` while all else MET → direct human gate (a dev - Remaining `UNVERIFIABLE` while all else MET → direct human gate (a dev
cannot fix unverifiability); do not spend a loop on it. cannot fix unverifiability); do not spend a loop on it.
- Out-of-scope files: a dev justification is accepted ONLY through the human - Out-of-scope files: a dev justification is accepted ONLY through the human
@@ -53,10 +89,11 @@ Parse its single `SECURITY — VERDICT:` line:
- `PASS` → done, proceed to commit. - `PASS` → done, proceed to commit.
- `BLOCK(n)` → hand the dev the `BLOCKING` list + the CONTRACT path (inline - `BLOCK(n)` → hand the dev the `BLOCKING` list + the CONTRACT path (inline
fix, or FRESH executor re-dispatch). Then **re-verify the REQUEST first** (GATE 1, fresh fix, or FRESH executor re-dispatch). Then re-run GATE 0, then
verifier) — a security fix can drift the behavior — **then re-run GATE 2** **re-verify the REQUEST first** (GATE 1, fresh verifier) — a security fix
(fresh auditor), in that order. **Max 3 security iterations** → STOP + can drift the behavior — **then re-run GATE 2** (fresh auditor), in that
human escalation with the BLOCKING table. order. **Max 3 security iterations** → STOP + human escalation with the
BLOCKING table.
- `DEGRADED` (semgrep absent) → does NOT block on the tool's absence; surface - `DEGRADED` (semgrep absent) → does NOT block on the tool's absence; surface
the checklist result + recommend `make plugin`. A DEGRADED run that still the checklist result + recommend `make plugin`. A DEGRADED run that still
BLOCKs (grep-caught secret/injection) blocks like any other. BLOCKs (grep-caught secret/injection) blocks like any other.
@@ -65,6 +102,11 @@ Parse its single `SECURITY — VERDICT:` line:
## Order invariant ## Order invariant
REQUEST conformity is always re-checked BEFORE security on any re-loop — a Every re-loop replays the gates in order: **GATE 0 → GATE 1 → GATE 2**,
security fix that breaks the feature must not slip through because only the never a subset and never reversed.
security gate re-ran. Never the reverse order.
The floor runs first because it is free, and because a red build makes the
verifier's verdict meaningless. REQUEST conformity is
always re-checked BEFORE security on any re-loop — a security fix that breaks
the feature must not slip through because only the security gate re-ran.
Never the reverse order.
+6 -1
View File
@@ -172,6 +172,11 @@ Parse the `BUGFIX-EXEC REPORT`:
1. Run the two fresh gates per `$HOME/.claude/lib/verify-secure-loop.md` with 1. Run the two fresh gates per `$HOME/.claude/lib/verify-secure-loop.md` with
`CONTRACT` = the STEP 3.5 path, `DIFF` = the executor's working-tree diff, `CONTRACT` = the STEP 3.5 path, `DIFF` = the executor's working-tree diff,
`TEST` = the suite named in its report: `TEST` = the suite named in its report:
- GATE 0 — deterministic floor, no dispatch: `bash ~/.claude/lib/gates.sh
run "$CONTRACT"` executes the criteria's declared oracles fail-closed
(the regression-test criterion included). UNMET → re-dispatch a FRESH
bugfixer with the NOT-MET rows verbatim — no verifier is spent on a red
floor; own budget, max 3 → escalate. MET → GATE 1.
- GATE 1 — a FRESH verifier judges the fix against the contract (bug gone - GATE 1 — a FRESH verifier judges the fix against the contract (bug gone
+ regression test present). CONFORME on the first pass → straight to + regression test present). CONFORME on the first pass → straight to
GATE 2, no loop. ECARTS → the "dev" of the loop is the dispatched GATE 2, no loop. ECARTS → the "dev" of the loop is the dispatched
@@ -183,7 +188,7 @@ Parse the `BUGFIX-EXEC REPORT`:
path; re-verify the request THEN re-scan, max 3 → escalate. path; re-verify the request THEN re-scan, max 3 → escalate.
Loop decisions stay HERE, in the main loop (LRN-083). Nominal = one Loop decisions stay HERE, in the main loop (LRN-083). Nominal = one
executor + one verifier + one security dispatch. executor + a free floor run + one verifier + one security dispatch.
2. **Pre-commit confirmation gate.** Before running `git commit`, present the diff 2. **Pre-commit confirmation gate.** Before running `git commit`, present the diff
summary and the proposed message, then wait for approval: summary and the proposed message, then wait for approval:
+7 -2
View File
@@ -161,6 +161,11 @@ Run the two fresh gates per `$HOME/.claude/lib/verify-secure-loop.md` with
`CONTRACT` = the STEP 0.7 path, `DIFF` = the working-tree diff the executor `CONTRACT` = the STEP 0.7 path, `DIFF` = the working-tree diff the executor
produced, `TEST` = the suite named in its report: produced, `TEST` = the suite named in its report:
- GATE 0 — deterministic floor, no dispatch: `bash ~/.claude/lib/gates.sh
run "$CONTRACT"` executes the criteria's declared oracles fail-closed.
UNMET → re-dispatch a FRESH feater with the NOT-MET rows verbatim — no
verifier is spent on a red floor; own budget, max 3 → escalate.
MET (an all-manual contract too) → GATE 1.
- GATE 1 — a FRESH verifier judges the diff against the contract (blind). - GATE 1 — a FRESH verifier judges the diff against the contract (blind).
CONFORME on the first pass → straight to GATE 2, no loop. ECARTS → the CONFORME on the first pass → straight to GATE 2, no loop. ECARTS → the
"dev" of the loop is the dispatched executor: re-dispatch a FRESH feater "dev" of the loop is the dispatched executor: re-dispatch a FRESH feater
@@ -171,8 +176,8 @@ produced, `TEST` = the suite named in its report:
CONTRACT path; re-verify the request THEN re-scan, max 3 → escalate. CONTRACT path; re-verify the request THEN re-scan, max 3 → escalate.
Loop decisions stay HERE, in the main loop (LRN-083). Nominal (clear Loop decisions stay HERE, in the main loop (LRN-083). Nominal (clear
request, conform first pass, clean diff) = one executor + one request, conform first pass, clean diff) = one executor + a free floor
verifier + one security dispatch. run + one verifier + one security dispatch.
## STEP 5 — COMMIT ## STEP 5 — COMMIT
+5
View File
@@ -227,6 +227,11 @@ If `graphify` not installed or complexity < 30% → skip silently.
Run the two fresh gates per `$HOME/.claude/lib/verify-secure-loop.md` with Run the two fresh gates per `$HOME/.claude/lib/verify-secure-loop.md` with
`CONTRACT` = the STEP 1 path (ENRICHED at STEP 4), `DIFF` = the MVP branch `CONTRACT` = the STEP 1 path (ENRICHED at STEP 4), `DIFF` = the MVP branch
diff (`develop..HEAD`), `TEST` = the project suite: diff (`develop..HEAD`), `TEST` = the project suite:
- GATE 0 — deterministic floor, no dispatch: `bash ~/.claude/lib/gates.sh
run "$CONTRACT"` executes the criteria's declared oracles fail-closed.
UNMET → hand the dev with the NOT-MET rows verbatim — no
verifier is spent on a red floor; own budget, max 3 → escalate.
MET (an all-manual contract too) → GATE 1.
- GATE 1 — a FRESH verifier judges the MVP against the enriched contract (V1 - GATE 1 — a FRESH verifier judges the MVP against the enriched contract (V1
features + `[gated]` design criteria). CONFORME → GATE 2. ECARTS → fix, features + `[gated]` design criteria). CONFORME → GATE 2. ECARTS → fix,
re-verify, max 3 → STOP + human escalation with the CRITERIA table. re-verify, max 3 → STOP + human escalation with the CRITERIA table.
+5
View File
@@ -210,6 +210,11 @@ OPTIONS :
Run the two fresh gates per `$HOME/.claude/lib/verify-secure-loop.md` with Run the two fresh gates per `$HOME/.claude/lib/verify-secure-loop.md` with
`CONTRACT` = the STEP 0e path (ENRICHED at STEP 3), `DIFF` = the branch diff `CONTRACT` = the STEP 0e path (ENRICHED at STEP 3), `DIFF` = the branch diff
(`develop..HEAD`), `TEST` = the project suite: (`develop..HEAD`), `TEST` = the project suite:
- GATE 0 — deterministic floor, no dispatch: `bash ~/.claude/lib/gates.sh
run "$CONTRACT"` executes the criteria's declared oracles fail-closed.
UNMET → hand the dev with the NOT-MET rows verbatim — no
verifier is spent on a red floor; own budget, max 3 → escalate.
MET (an all-manual contract too) → GATE 1.
- GATE 1 — a FRESH verifier judges the branch against the ENRICHED contract - GATE 1 — a FRESH verifier judges the branch against the ENRICHED contract
(all criteria, including the `[gated]` design ones). CONFORME → GATE 2. (all criteria, including the `[gated]` design ones). CONFORME → GATE 2.
ECARTS → hand the dev the gap list, fix, re-verify, max 3 → STOP + human ECARTS → hand the dev the gap list, fix, re-verify, max 3 → STOP + human