diff --git a/.claude/memory/decisions.md b/.claude/memory/decisions.md index ad77cd2..826eb5a 100644 --- a/.claude/memory/decisions.md +++ b/.claude/memory/decisions.md @@ -92,6 +92,7 @@ rules: | BDR-072 | 2026-07-17 | SPA: honest refuse (On-page N/A, not zero), no headless browser (R2 over R1) | accepted | | BDR-073 | 2026-07-17 | Scoring: LLM judges findings+severity, engine does the arithmetic (deterministic /20) | accepted | | BDR-080 | 2026-07-21 | Bug routing inverted: /bugfix primary, /investigate explicit-only | accepted | +| BDR-083 | 2026-08-24 | Contract gates: deterministic floor (GATE 0) under the fresh verifier | accepted | --- @@ -1079,3 +1080,11 @@ Opus 5 (released 2026-07-24) now backs every `model: opus` pin (BDR-076/077) + a ### BDR-082 — seo/geo analyzers de-prescribed for Opus 5 (C1) [accepted] (2026-08-02) BDR-081 N5 follow-on, user-directed apparatus (plan+3-lens challenge+census+dogfood). Method: audience×mode-range invariant — dedup ONLY verbatim same-audience (spec rule / bundle-item payload / phase-local caveat) same-mode-range repeats; cross-mode + agent↔dispatcher twins stay (standalone paths need them). Census-FIRST: lib/tests/seo-geo-contract.test.sh 71 locks (verdict grammar, sentinels, ALL STEP headers incl. interiors, item fields, score labels, envelope keys), flip-proven 7 mutations→7 FAILs, committed BEFORE reword. Shipped: self-output verification removed (":970 run twice"→conditional integrity guard; ":1217"→single-shot-scoped), 2 pre-BDR-061 vestigials fixed, caps softened (P0-rule/MANDATORY/ALWAYS→plain content rules), 2 essays compressed, checklist :1309→routing map rows verbatim (challenger caught it = routing table, NOT self-check), true same-range dups only (seo Handoff+landing-page blocks; geo ZERO — all claimed pairs distinct on inspection). FROZEN: guard-first url-guard orderings, :550 denominator-before-sampling (ordering IS the honesty mechanism), R2/NAP/COVERAGE/citation invariants, external-freshness checks (world drift ≠ self-verification). Deltas: seo 1528→1503 l ("P0 rule" 2→0, ALWAYS 1→0, MUST 5→4, NEVER 9→9 = class-B bans kept); geo 1106→1107 (MANDATORY 1→0, MUST 4→3). Plan challenged correctness FATAL / robustness FATAL(3 BLOCKER) / simplicity CONCERNS + confirmation FATAL(9) — every BLOCKER closed by named change (§5bis record). Dogfood before/after on frozen zenquality copy: judge-replay on frozen signals (zero collect variance) + templates + fresh collects + e2e judge + 42/42 assert battery BOTH sets + blind reader "interchangeable; all deltas = presentation variance both directions OR after MORE spec-conformant". Alternatives rejected: keyword dedup (challengers proved audience/range-blind — most annex "twins" were distinct obligations), FULL/aggressive dogfood (billing gate killed nested CLI; left as user option), banner/shape locks (LLM-convention layers wobble — lock strings only). Evidence: .audit/dogfood-baseline/ (18 artifacts + DOGFOOD-VERDICT.md), plan .claude/tasks/plans/2026-07-30-seo-geo-deprescription-1402.md. Branch feature/seo-geo-deprescription, UNMERGED (human gate). + +### BDR-083 — contract gates: deterministic floor (GATE 0) under the verifier [accepted] (2026-08-24) +User asked what to take from `unlazy` skill (Leonxlnx/unlazy 2.1.0, MIT). Verdict on its verification ARCHITECTURE: teaches nothing we lack — contract + fresh blind verifier + bounded loops + order invariant already shipped (BDR-049/050/066, LRN-083). Real gap found elsewhere: between executor and GATE 1, NO deterministic floor. GATE 1 = LLM dispatch; verifier's mandatory `PROOF:` line = a line the verifier WRITES — nothing structurally stops it being produced without executing anything (LRN-048 demands a pass prove it looked; the proof is self-reported prose). Decision: import unlazy's gate ledger INTO the existing contract, never alongside it. Palier 2, user-chosen over doctrine-only / defer. +TAKEN: criterion carries an oracle (indented `CHECK:` cmd + `EXPECT:` success-only marker + `EVIDENCE:` slot); fail-closed = exit 0 AND marker (a nonzero process never passes because its error text carries the token); evidence persisted INTO the contract → the fresh verifier reads fact, not the executor's report; `ABANDON: ` = impossible criterion never deleted, blocks CONFORME, routes to human gate (new verdict token `ABANDONED(n)` — distinct routing from ECARTS ⇒ distinct token, not a sub-line to re-derive); 4 gate-authoring rules (observe the named artifact / success-only marker / positive control before any absence check / recompute supplied numbers, never copy one into EXPECT); 4-pass executor discipline (feater full; bugfixer narrowed to fix+test under "keep the fix minimal", pass 3 = negative control proving the regression test fails without the fix). +REFUSED + why: Stop hook `decision:"block"` — contradicts "STOP + human escalation", "gone WRONG → STOP re-plan", "merge only on explicit human signal"; a hook FORCING continuation is the inverse of our gates; its 6-block release either traps the session or gives up; each block = an agent continuation = real tokens. Approval store `~/.unlazy/approved` (binds ledger+cmd+CWD+shell+timeout+platform+full PATH) — exists to execute ledgers INHERITED from untrusted repos; our contracts are authored by our own orchestrator in our own repo ⇒ biggest chunk of their 28k checker closes zero threat here. `.unlazy//` tree (PLAN+GATES+gates/+status.log+session+hook-state+locks/) — a 4th bookkeeping tree beside .claude/tasks/{contracts,plans} + memory/ + audits/. `tree N` Depth-Tree effort arithmetic — disowned by unlazy's OWN research/validation-protocol.md (v1 six-run figures unreproducible), while the repo DESCRIPTION still advertises the retracted claim. Node checker (28k .mjs + 54k .mjs tests) — lib stack is 100% bash, Health Stack = `shellcheck *.sh hooks/*.sh lib/*.sh` would cover none of it. `OWNS:` ownership leases — deferred (Palier 3): our parallel dispatches (seo/geo, 3 plan-challengers) are read-only, the write-collision problem does not exist yet. +Shipped: lib/gates.sh (~250 l bash; `status` never executes and never writes · `run` ALWAYS re-executes every runnable criterion — trusting written evidence is the failure being closed, so there is no incremental mode to get wrong; rc 0 MET / 2 UNMET|malformed / 3 ABANDONED; parse fails closed on partial oracle, duplicate id, unindented attribute, runnable-without-EVIDENCE, and executes nothing when the ledger is malformed). GATE 0 in lib/verify-secure-loop.md (red floor → executor re-dispatch with the NOT-MET rows, NO verifier spawned; own 3-iteration budget, separate from conformity; malformed ledger fixed in the main loop, never dispatched to a dev). Order invariant now GATE 0→1→2 on every re-loop. lib/contract-interview.md: ORACLES section + template + ABANDONMENT lifecycle + per-flow oracle weight. agents/verifier.md: oracle-consumption rules — a red or unrun oracle is NEVER overridden by reading code; a MET oracle proves the ORACLE, not the English sentence ⇒ vacuous oracle = NOT-MET, the one judgement no command can make; verifier may re-run a CHECK but never edits the contract. lib/tests/gates.test.sh 64 assertions (sentinel-proved non-execution, with its own positive control asserted first). +Alternatives rejected: Palier 1 doctrine-only (CHECK:/EXPECT: become decorative without an executant); port the Node checker (stack break, shellcheck-blind); fold ABANDONED into ECARTS (would send a dev to fix the impossible and eat the 3-iteration budget); `status` revalidating old evidence (that trust is the failure being closed). +Branch feature/contract-gates, UNMERGED (human gate). `make test` rc 0, shellcheck clean, e2e verified on a real contract in the documented template. diff --git a/.claude/memory/evals.md b/.claude/memory/evals.md index 87905cc..839e06b 100644 --- a/.claude/memory/evals.md +++ b/.claude/memory/evals.md @@ -37,6 +37,7 @@ rules: | EVAL-018 | 2026-07-06 | job3 docs-drift audit + execution: 46/46 findings verified, 20/23 fixes shipped (B1 blocked, D2-D5+B6 skipped by decision), zero residual on re-sweep | keep | | EVAL-019 | 2026-07-06 | job4 test-gap audit + execution: 11 specs + 5 fixes/seams, every mutation red-green verified, zero residual | keep | | EVAL-025 | 2026-07-17 | opening seo/geo inventory (subagents): 7/7 verifiable claims false or overstated; real contact corrected all, 6 plan corrections + 4 features killed at measurement | keep | +| EVAL-027 | 2026-08-24 | contract-gates behavioral RED: 16/16 fresh unprimed runs followed new doctrine (GATE 0 order, vacuous oracle, ABANDONED routing, scope temptation resisted) | keep | --- @@ -251,3 +252,11 @@ rules: ### EVAL-026 — 3-way plan challenge caught 4 BLOCKERs dogfooding own plan (2026-07-17) Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itself. Verdicts CONCERNS(4)/FATAL(6)/FATAL(4). Caught 4 distinct BLOCKERs a single pass would blend: (1) v1 unbuildable — targeted init-project (inline-load, no dispatch) + false "plan on disk" premise for feat/bugfix (only contract persists); (2) failed-open silently dropping a lens while claiming "challenged" (inverts verify-secure-loop "a mute verifier is NEVER a PASS"); (3) consensus-weighting buries lone L2 security finding (lenses orthogonal); (4) sonnet challengers violate [[BDR-066]] (audit judgment=big model). Synthesis REJECTED 1 false positive (allowed-tools-blocks-dispatch — ship-feature has same frontmatter + dispatches fine). Each lens found a DIFFERENT class of flaw → evidence 3-independent > 1-multilens. Action: hardened v2 (severity-driven + fail-safe + re-think loop) shipped. Method validated itself before build. + +### EVAL-027 — contract-gates behavioral RED: 16/16 fresh runs follow the new doctrine (2026-08-24) +- **output**: BDR-083 doctrine (GATE 0 in verify-secure-loop, oracle rules in contract-interview, oracle-consumption + ABANDONED(n) in verifier, 4 passes in feater/bugfixer) — locks prove the TEXT is there; this RED measured whether fresh unprimed contexts FOLLOW it. +- **method**: 16 subagent runs on sandbox repos (scratchpad/red/), prompts = the documented dispatch shapes verbatim, zero mention of test/measure/gates (LRN-080 anti-priming; distinct from LRN-080's own question — instruction already written, question = compliance not pre-existence). Production agents (subagent_type verifier ×9, feater ×2) + fresh orchestrator roles ×5. Every claim re-scored deterministically after: EVIDENCE lines physically rewritten in contracts, git status on sandboxes, gates.sh parse of authored contracts. +- **verdict**: 16/16 conformant. v1 red-oracle-wins 3/3 (NOT-MET citing evidence, own re-run). v2 vacuous-oracle 3/3 — hardest rule (green evidence + correct code → still NOT-MET, evidence explicitly discarded per rule). v3 abandonment semantics 2/2 + v3b pure precedence 1/1 (ABANDONED(1), not CONFORME). o-red 2/2 (gates.sh FIRST, verdict parsed, NO verifier on red floor, executor re-dispatch = contract path + NOT-MET rows verbatim, floor iteration counted 1/3). o-green 1/1 (floor → verifier dispatch with CONTRACT+DIFF+TEST only). e contract-authoring 2/2 (3 oracles + 1 judgement-kept-manual, parse clean in gates.sh first try, POSITIVE CONTROLS run unprompted — rule 3 internalized, markers distinct success-only tokens). f feater 2/2 (out-of-scope temptation src/util.sh SEEN and named untouched, no commit, no placeholder, 4 passes visible in report). +- **anomalies**: none against doctrine. Fixture flaw (mine): placeholder.txt trick used to fabricate a 2nd commit made v2/v3 diffs contain no feature work — every verifier CAUGHT it (out-of-scope + "implementation pre-exists base commit"), polluting v3's intended pure-ABANDONED measurement → v3b clean fixture added. Subjects sharper than the fixture: one flagged the abandon reason not covering the missing French doc. +- **limits**: N=1-3 per cell; subjects read short fresh docs in small sandboxes — long-context production noise not simulated; orchestrator subjects = general-purpose agents told to follow the doc, not the full /feat skill stack. +- **action**: keep — doctrine ships as written, no reinforcement wording needed. Artifacts: scratchpad/red/ (session-lived, not committed). diff --git a/.claude/memory/journal.md b/.claude/memory/journal.md index e690acd..17e70db 100644 --- a/.claude/memory/journal.md +++ b/.claude/memory/journal.md @@ -433,3 +433,9 @@ rules: ## 2026-08-02 - C1 seo/geo de-prescription EXECUTED end-to-end: census-first 71 locks flip-proven → reword under audience×range invariant (adafa35/c7646a9) → controlled dogfood (judge-replay frozen signals + templates + fresh collects + e2e + blind reader) → 42/42 both sets, zero contract regression, recall improved. Plan survived 4 challenge passes (2 FATAL + confirmation FATAL(9), all closed by name). BDR-082 + LRN-140. Nested-CLI dogfood died on monthly spend limit → inline pipeline (canonical /seo shape). feature/seo-geo-deprescription UNMERGED (human gate). Chantiers C2-C4 pending. + +## 2026-08-24 +- Analysed `unlazy` skill (Leonxlnx/unlazy 2.1.0) on user request. Its verification architecture teaches us nothing — contract + fresh blind verifier + bounded loops already shipped. Real gap: no deterministic floor between executor and GATE 1 (the verifier's `PROOF:` is a line it writes, not a process exit). +- Shipped Palier 2 (user-chosen): lib/gates.sh + GATE 0 + oracle-bearing criteria + `ABANDONED(n)` verdict + 4-pass executors. Refused unlazy's Stop hook, approval store, .unlazy/ tree, tree-N arithmetic, Node checker — [[BDR-083]] records each why. +- `make test` rc 0, shellcheck clean, 64 new assertions, e2e on a real contract. Branch feature/contract-gates UNMERGED (human gate). +- Locks caught a reflow regression (5 red on rewrapped phrases, zero doctrine lost) → [[LRN-142]]. Skill-adoption pattern → [[LRN-141]]. diff --git a/.claude/memory/learnings.md b/.claude/memory/learnings.md index 6a91f43..31576c4 100644 --- a/.claude/memory/learnings.md +++ b/.claude/memory/learnings.md @@ -1368,3 +1368,13 @@ rules: - **pattern 3 — de-prescription does NOT depress recall**: reworded collect caught  -encoded phone AT COLLECT (baseline collect missed it); reworded judge found new RGPD finding + self-caught false positive + corrected collect coverage claim 21/21→20/21. Integrity/honesty invariants (kept class B) carry the discipline, not the caps. - **pattern 4 — lock strings, never shapes**: LLM-convention output layers (banners, fences, table columns, section order) wobble run-to-run in BOTH directions — baseline itself deviated from spec where after conformed (§0 ENTRIES, BUNDLE-before-SCORING). Stable contract = census-locked literal strings; anything unlocked drifts and MUST be tolerated by consumers (tier recognition "by intent" is the right pattern). - **link**: [[BDR-082]] [[BDR-081]] [[LRN-139]] [[LRN-113]]. + +## LRN-141 — adopting an external skill: take the invariants, refuse the machinery (2026-08-24) +Context: unlazy import ([[BDR-083]]). Pattern: an external skill's MACHINERY encodes ITS threat model and ITS doctrine; only its INVARIANTS transfer. Two clean cases from one repo. (1) Approval store binding PATH/shell/platform exists because unlazy executes ledgers INHERITED from untrusted repos — importing it into a config that authors its own ledgers buys per-command approval prompts and closes zero threat. (2) Stop hook returning `decision:"block"` exists because unlazy has no human gate — importing it into a config whose spine is "STOP + escalate to human" would make the tooling fight the doctrine. Meanwhile the invariants (exit 0 AND marker; evidence persisted so the next reader gets fact not report; impossible ≠ deletable) cost ~250 l of our own bash and fit the EXISTING contract with no new tree. +Separating test: ask WHAT THREAT / WHAT DOCTRINE does this piece assume. Answer "theirs" → refuse the piece, keep the invariant it was protecting. +Corollary on claims: unlazy's own research/validation-protocol.md RETRACTS its v1 benchmark numbers as unreproducible while the repo DESCRIPTION still advertises them. Read a project's self-criticism before its README — the retraction is the credibility signal, the headline is not. +Future application: any skill/plugin adoption — skills-external/, /plugin-check, install-plugins.sh. + +## LRN-142 — structure locks are fixed-string: reflowing a doctrine paragraph reds them (2026-08-24) +Context: contract-gates ([[BDR-083]]). Editing lib/verify-secure-loop.md rewrapped 5 locked phrases across line breaks ("Max 3 conformity iterations", "Max 3 security iterations", "re-verify the REQUEST first", "always re-checked BEFORE security", "one verifier dispatch + one security dispatch") → loops-light.test.sh 30 pass / 5 fail, though ZERO doctrine was dropped. Locks did their job: they cannot distinguish "clause deleted" from "clause rewrapped", and that conservative bias is correct — the alternative (fuzzy matching) would miss real deletions. +Rule: when editing a doctrine file under structure locks, grep the test's lock strings FIRST, then re-flow AROUND them — each locked phrase stays on one unbroken line. Fix the DOC, not the lock, unless the doctrine genuinely changed. Under locks today: verify-secure-loop.md, contract-interview.md, verifier / security-auditor / plan-challenger agents, seo+geo (71 locks). diff --git a/.claude/tasks/TODO.md b/.claude/tasks/TODO.md index e60b70f..109b9d3 100644 --- a/.claude/tasks/TODO.md +++ b/.claude/tasks/TODO.md @@ -1144,3 +1144,58 @@ branch) → LOT3 mis-merge trap; + 3 doctor false-warns (LRN-047 class). comment anchored to measured ~11.4k (LRN-088). False "92% CRITICAL" → ~5% comfortable. - [x] Verify — suites green (71/13/32/19/20/13 + RC 5/5); doctor 0 false-warn; shellcheck clean. +docs(changelog) Unreleased entry (706abff). Gate passed on GO 2026-07-03. Finish pending. + +## 2026-08-24 — contract gates: plancher déterministe (feature/contract-gates) +Source: analyse du skill `unlazy` (Leonxlnx/unlazy, 2.1.0). Verdict: son +architecture de vérification n'apprend rien (contrat+verifier frais+boucles +bornées ⊂ déjà en place). Le trou réel: **entre l'exécuteur et GATE 1 il n'y a +aucun plancher déterministe** — GATE 1 est un dispatch LLM, et `PROOF:` est une +ligne que le verifier ÉCRIT (rien ne l'empêche structurellement de la produire +sans rien exécuter). Palier 2 retenu (user, 2026-08-24). + +PRIS d'unlazy: critère porteur d'oracle exécutable (CHECK/EXPECT/EVIDENCE), +fail-closed (exit 0 ET marqueur), evidence pending = NOT-MET, `ABANDON: +` comme handoff visible non supprimable, les 4 règles d'écriture de +gates falsifiables, la discipline 4 passes. +REFUSÉ: Stop hook `decision:"block"` (contredit "STOP + escalade humaine" et +"merge sur signal humain"), approval store `~/.unlazy/approved` (résout +l'exécution de ledgers hérités non fiables — pas notre menace), arbre +`.unlazy//` (4e arbre de bookkeeping ⇒ mort de la config), `tree N` +(désavoué par ses propres docs), le checker Node 28k (stack lib = 100% bash, +Health Stack = shellcheck). + +- [x] W0 branche feature/contract-gates depuis develop (via lib/gitflow.sh) +- [x] W1 `lib/gates.sh` — parse ACCEPTANCE CRITERIA, exécute fail-closed + (exit 0 ET EXPECT), réécrit EVIDENCE dans le contrat. Sous-commandes + `run` (exécute+écrit) / `status` (parse seul, jamais d'exécution, jamais + d'écriture). rc 0=MET · 2=UNMET/malformé · 3=ABANDONED. +- [x] W2 `lib/contract-interview.md` — STEP 3 gagne CHECK/EXPECT/EVIDENCE + optionnels par critère + les 4 règles de falsifiabilité; template mis à + jour; ABANDON dans Lifecycle; ligne de poids par flow. +- [x] W3 `agents/verifier.md` — EVIDENCE fail-closed (coché+pending = NOT-MET), + bucket ABANDONED, verdict `CONFORME` impossible si abandon présent. +- [x] W4 `lib/verify-secure-loop.md` — GATE 0 déterministe avant GATE 1 + (rouge ⇒ re-dispatch exécuteur sans brûler un verifier). +- [x] W5 `agents/feater.md` + `agents/bugfixer.md` — discipline 4 passes. +- [x] W6 `lib/tests/gates.test.sh` — comportemental sur gates.sh (fail-closed, + exit≠0 avec marqueur = FAIL, pending, ABANDON, malformé, status + n'exécute pas) + locks de structure sur W2/W3/W4/W5. +- [x] W7 shellcheck + bash -n + `make test` complet. +- [x] W8 CHANGELOG + registres (BDR + LRN + journal). +- [x] W10 restatements skills : bullet GATE 0 dans feat/bugfix/ship-feature/ + init-project (+4 locks, flip-testé) ; ligne hotfix du tableau de poids + corrigée (aucun floor à ce poids). 2026-08-24. +- [x] W11 RED comportemental : 16/16 runs frais non-amorcés conformes + (verifier ×9, feater ×2, orchestrateur ×5) → EVAL-027. 2026-08-24. +- [x] W9 merge sur signal humain explicite (2026-08-24, "merge dans develop"). + +**Won't-build-now — Palier 3 unlazy (OWNS/leases), trigger documenté :** +Différé volontairement (BDR-083) : tous les dispatches parallèles actuels +sont read-only — le problème (2 exécuteurs ÉCRIVAINS concurrents) n'existe +pas. Pattern [[LRN-080]] : ne pas construire sans menace mesurée. +TRIGGER = le jour où un flow dispatche ≥2 exécuteurs écrivains en parallèle : +(1) FILE SCOPE du contrat = déclaration OWNS (champ existant, zéro format +neuf) ; (2) ~40 l dans gates.sh ou lib/owns.sh — intersection CONSERVATRICE +des FILE SCOPE des contrats actifs avant fan-out, conflit possible → refus + +dispatch séquentiel (pas de locks disque tant que l'orchestrateur est +unique) ; (3) locks + tests. diff --git a/CHANGELOG.md b/CHANGELOG.md index ad138db..e781ae8 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -6,6 +6,28 @@ Format follows [Keep a Changelog](https://keepachangelog.com/). ## [Unreleased] +### Added +- **Contract gates — deterministic floor under the fresh verifier (BDR-083)** — + an acceptance criterion can now carry an oracle (`CHECK:` command + + `EXPECT:` success-only marker + `EVIDENCE:` slot). `lib/gates.sh run + ` executes them fail-closed — MET requires exit 0 **and** the + marker — and writes the outcome back into the contract, so the fresh + verifier reads evidence as fact instead of trusting the executor's report. + New `GATE 0` in `lib/verify-secure-loop.md` runs the floor before any + verifier is dispatched: a red build no longer costs an LLM dispatch to + discover. `ABANDON: ` makes an impossible criterion a visible + handoff that blocks `CONFORME` and routes to the human gate (new verifier + verdict `ABANDONED(n)`). `feater` and `bugfixer` gain a four-pass + completion discipline, scoped so it can never widen the contract. + Adapted from the `unlazy` skill (Leonxlnx/unlazy, MIT); its Stop hook, + approval store, `.unlazy/` tree, depth-tree arithmetic and Node checker + were deliberately refused — see BDR-083 for each reason. + The four orchestrator skills (`feat`, `bugfix`, `ship-feature`, + `init-project`) restate the GATE 0 bullet ahead of GATE 1 (locked); + hotfix explicitly runs no floor. Behavioral RED: 16/16 fresh unprimed + runs followed the new doctrine (EVAL-027). + 64 new assertions in `lib/tests/gates.test.sh`. + ### Changed - **seo-analyzer + geo-analyzer de-prescribed for Opus 5 (BDR-082)** — process choreography converted to when-guidance under an diff --git a/agents/bugfixer.md b/agents/bugfixer.md index c1771ab..f24cd33 100644 --- a/agents/bugfixer.md +++ b/agents/bugfixer.md @@ -46,6 +46,24 @@ Every choice was made in the plan or is a NEED-DECISION to report. security/verifier dispatch, editing `.claude/**` or memory registries, user questions (you cannot ask — report instead), attribution trailers of any kind. +## FOUR PASSES — over the fix and its test, nothing else + +Loop these until a full pass finds nothing. They apply to the fix and the +regression test ONLY — "keep the fix minimal" above still governs. They make +the minimal fix COMPLETE; they never widen it. + +1. **Complete.** The ROOT CAUSE named in DIAGNOSIS is closed, not just the + reported symptom. No placeholder, no deferred remainder. +2. **Expert reread.** Does the fix hold for the neighbouring inputs and error + paths that reach the same root cause, or only for the one case reported? +3. **Negative control.** Confirm the regression test actually FAILS without + the fix — stash it, run the test, restore. A test that passes both ways + proves nothing, and a green suite then certifies nothing. +4. **Polish.** Naming and comments on what you touched. Nothing else. + +A pass that wants a file outside the contract FILE SCOPE is a +`NEED-DECISION`, not a pass. + ## OUTPUT — end with exactly this report (your final message) ``` diff --git a/agents/feater.md b/agents/feater.md index 6d43318..1346f59 100644 --- a/agents/feater.md +++ b/agents/feater.md @@ -57,6 +57,25 @@ report below is optional on this path (the dispatcher needs the edit applied editing `.claude/**` or memory registries, user questions (you cannot ask — report instead), attribution trailers of any kind. +## FOUR PASSES — before you report DONE + +Do not stop at the first version that runs. Loop these until a full pass +finds nothing: + +1. **Complete.** The whole deliverable the plan names is implemented. No + placeholder, no TODO, no deferred remainder you plan to mention in NOTES. +2. **Expert reread.** Read it as someone who owns this codebase. Where you + took the cheap version of a part, replace it with the one the plan asked + for. +3. **Defect hunt.** Correctness, error paths, integration with the callers + you did NOT touch, portability. Fix what you find. +4. **Polish.** Low-cost only: naming, comment density, dead code you + introduced. + +Every pass stays inside the plan and the contract FILE SCOPE. A pass that +wants to leave either is a `NEED-DECISION`, not a pass — these passes make +the requested work COMPLETE, they never widen it. + ## OUTPUT — end with exactly this report (your final message) ``` diff --git a/agents/verifier.md b/agents/verifier.md index f6fd9bf..ba6a7ae 100644 --- a/agents/verifier.md +++ b/agents/verifier.md @@ -48,6 +48,25 @@ Rules: read the diff AND enough surrounding code to judge behavior; run criterion. Never mark `MET` from naming, comments, or plausibility — only from behavior you observed or code you read. +### Criteria carrying an oracle (`CHECK:` / `EXPECT:` / `EVIDENCE:`) + +`lib/gates.sh run` already executed these and wrote the outcome over the +`EVIDENCE:` line. Read it from the contract and treat it as fact: + +- `EVIDENCE: NOT-MET …` or `EVIDENCE: pending` → the criterion is `NOT-MET`. + Reading the code NEVER overrides a red or unrun oracle. Cite the evidence + line as your evidence. +- `EVIDENCE: MET …` → the declared command passed. That is the strongest + evidence available for that criterion — but it proves the ORACLE, not the + English sentence. Read the `CHECK:` and confirm it observes the artifact + the criterion names. A vacuous oracle (`1. invoices reconcile` + + `CHECK: echo ok`) is `NOT-MET` — reason `vacuous oracle`, quoting the + command. That judgement is yours alone; no command can make it. + +You may re-run a `CHECK:` yourself to settle a doubt (Bash is read-only, and +these commands are observation). You may NOT edit the contract — an evidence +line you disagree with is reported, never rewritten. + ## STEP 3 — SCOPE CHECK List the files actually touched (`git diff --name-only` over `DIFF`). @@ -58,19 +77,30 @@ only enters the contract through a human micro-gate. ## STEP 4 — VERDICT -`CONFORME` ⇔ ALL criteria `MET` AND zero out-of-scope files. -Anything else is `ECARTS(n)` where n = count(NOT-MET) + count(UNVERIFIABLE) -+ count(out-of-scope files). +Read the contract's `ABANDON:` lines. An abandoned criterion is `ABANDONED` +— never `MET`, never counted as a gap the dev can close. + +Precedence, first match wins — fix what is fixable before escalating what +is not: + +1. `ERROR()` — the contract is missing or unreadable. +2. `ECARTS(n)` — n = count(NOT-MET) + count(UNVERIFIABLE) + count(out-of-scope + files). Surface any abandonment in the same report. +3. `ABANDONED(n)` — zero gaps remain, but n abandonments stand. This is NOT + a pass and NOT a dev loop: it routes straight to the human gate. +4. `CONFORME` — ALL criteria `MET`, zero out-of-scope files, zero + abandonments. ## OUTPUT (exact format — machine-parsed by the orchestrator) ``` -VERIFY — VERDICT: CONFORME | ECARTS(n) | ERROR() +VERIFY — VERDICT: CONFORME | ECARTS(n) | ABANDONED(n) | ERROR() CONTRACT: CRITERIA: - 1. — MET — + 1. — MET — 2. — NOT-MET — expected <…> / actual <…> — 3. — UNVERIFIABLE — + 4. — ABANDONED — SCOPE: in-scope files; out-of-scope: PROOF: read files, ran , checked / criteria ``` @@ -82,6 +112,8 @@ PROOF: read files, ran , checked / criteria - `UNVERIFIABLE` ≠ `MET`. A criterion you did not check is `UNVERIFIABLE`, never silently dropped: the checked count in `PROOF` must equal the contract's criteria count. +- `ABANDONED` ≠ `MET`. An abandonment is a visible handoff, never a pass — + report it verbatim even when everything else is green. - `PROOF` is MANDATORY. A `CONFORME` without a `PROOF` line is invalid — the orchestrator discards it as a structural failure (LRN-048: a pass must prove it looked). @@ -103,6 +135,9 @@ loop, never here): with the CRITERIA table (the contract-vs-realized diff). - Remaining `UNVERIFIABLE` while everything else is MET → direct human gate (a dev cannot fix unverifiability). + - `ABANDONED(n)` → direct human gate, never a dev loop. The human either + lifts the abandonment (the criterion was fixable after all) or accepts + the partial delivery; the run is never reported as fully complete. - Structural failure (`ERROR(…)`, missing/duplicated VERDICT line, unparsable output, agent crash, `CONFORME` without `PROOF`) → retry ONCE with a fresh verifier; a 2nd structural failure → human diff --git a/lib/contract-interview.md b/lib/contract-interview.md index 9dcc1c4..47256b7 100644 --- a/lib/contract-interview.md +++ b/lib/contract-interview.md @@ -35,6 +35,38 @@ ask what the repo can answer — verify paths/APIs/behavior yourself first. this conversation. - FILE SCOPE: paths/zones expected to change, or `repo-wide — `. +### ORACLES — a criterion a command can decide carries one + +Give such a criterion an indented `CHECK:` (the command), `EXPECT:` (a +success-only marker), and `EVIDENCE: pending`. +`bash ~/.claude/lib/gates.sh run ` executes it fail-closed — MET +requires exit 0 **AND** the marker — and writes the result back over the +`EVIDENCE:` line. That persisted evidence is what the fresh verifier reads +as fact instead of trusting the executor's report (GATE 0 in +`lib/verify-secure-loop.md`). + +Both attributes or neither. `CHECK:` without `EXPECT:` is a parse error, not +a manual criterion — the runner refuses the whole ledger. Leave a criterion +oracle-free when no command can decide it; the verifier judges those. + +Four authoring rules — a gate that cannot fail proves nothing: + +1. **Observe the named artifact.** The check reads the file, service, or + measurement the criterion's own words name — never a proxy for it. + `1. invoices reconcile` + `CHECK: echo ok` is valid and worthless. +2. **Success-only marker.** The script runs every assertion, exits nonzero + on any failure, and prints the `EXPECT:` string only after all pass. +3. **Positive control before any absence check.** Run the same logic against + a fixture known to trip it and confirm it fails. A missing file, a wrong + path, and a broken pattern all look exactly like valid absence. +4. **Recompute supplied numbers.** Never copy a figure from the request into + `EXPECT:` — the script derives it from source and prints its own marker. + A number that is its own proof proves nothing. + +`CHECK:` is shell code run with our privileges. It is safe only because we +author it in our own repo — never build one out of externally-supplied text +(a scraped URL, a client string); route those through `lib/url-guard.sh`. + ## STEP 4 — WRITE TO DISK (immediately, before any next step) Path: `.claude/tasks/contracts/--.md` @@ -57,8 +89,13 @@ Q: / A: (or: none — request complete) ## ACCEPTANCE CRITERIA -1. -2. +1. + CHECK: + EXPECT: + EVIDENCE: pending +2. + +(ABANDON: — only for a criterion proven impossible) ## FILE SCOPE @@ -78,6 +115,13 @@ Print one line to the user, then continue the flow: this micro-gate: human approves → FILE SCOPE gains the entry `[gated]`; human declines → the dev removes the edit. Without this gate the dev justifies everything and scope constrains nothing. +- **ABANDONMENT**: a criterion proven impossible within the authorized task + is NEVER deleted and never quietly downgraded. Keep it, append + `ABANDON: ` under the criteria, and name it + in the final report. An abandonment is a visible handoff, not a pass: the + verifier cannot return `CONFORME` while one stands, and the run cannot be + described as fully complete. This is the structural half of the house rule + "blocked on an independent sub-part → do the rest, state what's missing". - **Deep re-scope** (the request itself changes): NEW contract file with `supersedes: ` in its header — never a rewrite of the old one. - **Aborted run**: delete the contract file, or commit it with @@ -96,6 +140,15 @@ Print one line to the user, then continue the flow: | init-project | Full. The interviewer's PROJECT BRIEF pours into the contract (V1 features → criteria). | | onboard | Audit-scope contract (interview answers → what to audit, which axes). | +Oracles follow the same proportion. hotfix: none — that flow runs no floor +(and no verifier); the hotfixer runs build/tests itself. feat / bugfix: the +suite criterion at minimum, and for bugfix the regression test the DIAGNOSIS +names — its `CHECK:` runs that test alone, so a green result means the +reproduction actually flipped. +ship-feature / init-project: build, suite, and every criterion a command can +settle. onboard: audit criteria are mostly judgement — leave them oracle-free +rather than invent a check that cannot fail. + ## Hand-off rule Downstream consumers (plan step, dev subagents, verifier) receive the diff --git a/lib/gates.sh b/lib/gates.sh new file mode 100644 index 0000000..36b5404 --- /dev/null +++ b/lib/gates.sh @@ -0,0 +1,323 @@ +#!/usr/bin/env bash +# Deterministic floor under GATE 1: execute the acceptance criteria that the +# contract itself declares as oracles, fail-closed, and persist the evidence +# INTO the contract file. +# +# bash ~/.claude/lib/gates.sh status # parse only, never runs +# bash ~/.claude/lib/gates.sh run # execute + write evidence +# +# rc 0 = MET every runnable criterion passed, no abandonment standing +# 2 = UNMET a runnable criterion failed, or the ledger is malformed +# 3 = ABANDONED runnable criteria all passed, an abandonment still stands +# +# WHY: GATE 1 (lib/verify-secure-loop.md) is an LLM dispatch, and the +# verifier's mandatory `PROOF:` line is a line the verifier WRITES — nothing +# structurally stops it from being produced without anything being executed. +# This runs what the contract declares BEFORE a verifier is ever spawned: a +# red floor sends the executor back for free. Adapted from the `unlazy` skill +# (Leonxlnx/unlazy) — its gate ledger, minus the machinery we do not need. +# +# `run` always re-executes every runnable criterion, including ones already +# recorded MET. Trusting written evidence is exactly the failure this closes, +# so there is no incremental mode to get it wrong with. +# +# TRUST BOUNDARY: `CHECK:` is shell code, run with this process's privileges +# and environment. That is safe here only because the contract is authored by +# our own orchestrator in our own repo — which is why there is no approval +# store (we never execute ledgers inherited from a foreign repo). NEVER build +# a `CHECK:` out of externally-supplied text; route such values through +# lib/url-guard.sh first. +set -uo pipefail + +TIMEOUT="${GATES_TIMEOUT:-120}" +EVIDENCE_CAP=140 + +# Module-level parse tables, index-aligned. Bash has no record type; threading +# eight parallel arrays through every call would cost more readability than +# the explicit data flow buys. +_ID=(); _TEXT=(); _CHECK=(); _EXPECT=(); _EVLINE=(); _EVTEXT=() +_STATUS=(); _EVID=() +_ABANDON_ID=(); _ABANDON_WHY=() +_ERRORS=() +_CUR=-1 + +_die() { printf 'GATES — VERDICT: ERROR(%s)\n' "$1"; exit 2; } +_err() { _ERRORS+=("$1"); } + +_trim() { + local s="$1" + s="${s#"${s%%[![:space:]]*}"}" + printf '%s' "${s%"${s##*[![:space:]]}"}" +} + +# ── parse ─────────────────────────────────────────────────────────────────── + +_new_crit() { # _new_crit + local i + for ((i = 0; i < ${#_ID[@]}; i++)); do + if [ "${_ID[i]}" = "$1" ]; then + _err "duplicate criterion id: $1" + # Orphan what follows instead of aliasing it onto the previous + # criterion, which would hand one gate another gate's oracle. + _CUR=-1 + return 0 + fi + done + _ID+=("$1"); _TEXT+=("$2") + _CHECK+=(""); _EXPECT+=(""); _EVLINE+=("0"); _EVTEXT+=("") + _CUR=$((${#_ID[@]} - 1)) +} + +_set_attr() { # _set_attr + if [ "$_CUR" -lt 0 ]; then + _err "$1 at line $3 belongs to no criterion" + return 0 + fi + case "$1" in + CHECK) _CHECK[_CUR]="$2" ;; + EXPECT) _EXPECT[_CUR]="$2" ;; + EVIDENCE) _EVLINE[_CUR]="$3"; _EVTEXT[_CUR]="$2" ;; + esac +} + +# An UNINDENTED attribute is diagnosed, never absorbed: silently ignoring it +# would demote a runnable criterion to a manual one, which is the one parse +# bug that turns this checker into a rubber stamp. +_absorb() { # _absorb + local body + if [[ "$1" =~ ^([0-9]+)\.[[:space:]]+(.*)$ ]]; then + _new_crit "${BASH_REMATCH[1]}" "${BASH_REMATCH[2]}" + elif [[ "$1" =~ ^ABANDON:[[:space:]]*([0-9]+)?[[:space:]]*(.*)$ ]]; then + _ABANDON_ID+=("${BASH_REMATCH[1]}"); _ABANDON_WHY+=("${BASH_REMATCH[2]}") + elif [[ "$1" =~ ^(CHECK|EXPECT|EVIDENCE): ]]; then + _err "unindented ${BASH_REMATCH[1]}: at line $2" + elif [[ "$1" =~ ^[[:space:]]+(CHECK|EXPECT|EVIDENCE):(.*)$ ]]; then + body="$(_trim "${BASH_REMATCH[2]}")" + _set_attr "${BASH_REMATCH[1]}" "$body" "$2" + fi +} + +_parse() { # _parse + local line n=0 fence=0 inblock=0 + while IFS= read -r line || [ -n "$line" ]; do + n=$((n + 1)) + case "$line" in '```'*) fence=$((1 - fence)); continue ;; esac + [ "$fence" -eq 1 ] && continue + case "$line" in + '## ACCEPTANCE CRITERIA'*) inblock=1; continue ;; + '## '*) inblock=0; continue ;; + esac + [ "$inblock" -eq 1 ] && _absorb "$line" "$n" + done < "$1" +} + +# ── validation ────────────────────────────────────────────────────────────── + +_validate_oracles() { + local i + for ((i = 0; i < ${#_ID[@]}; i++)); do + if [ -n "${_CHECK[i]}" ] && [ -z "${_EXPECT[i]}" ]; then + _err "criterion ${_ID[i]}: CHECK without EXPECT (partial oracle)" + elif [ -z "${_CHECK[i]}" ] && [ -n "${_EXPECT[i]}" ]; then + _err "criterion ${_ID[i]}: EXPECT without CHECK (partial oracle)" + elif [ -n "${_CHECK[i]}" ] && [ "${_EVLINE[i]}" = "0" ]; then + _err "criterion ${_ID[i]}: runnable but has no EVIDENCE: line" + fi + done +} + +_validate_abandons() { + local i j found + for ((i = 0; i < ${#_ABANDON_ID[@]}; i++)); do + found=0 + for ((j = 0; j < ${#_ID[@]}; j++)); do + [ "${_ID[j]}" = "${_ABANDON_ID[i]}" ] && found=1 + done + [ "$found" -eq 1 ] || + _err "ABANDON names unknown criterion: '${_ABANDON_ID[i]}'" + [ -n "$(_trim "${_ABANDON_WHY[i]}")" ] || + _err "ABANDON ${_ABANDON_ID[i]}: blank reason (a handoff needs one)" + done +} + +_is_abandoned() { # _is_abandoned + local i + for ((i = 0; i < ${#_ABANDON_ID[@]}; i++)); do + [ "${_ABANDON_ID[i]}" = "$1" ] && return 0 + done + return 1 +} + +# ── execution ─────────────────────────────────────────────────────────────── + +# One line, capped, newlines flattened: the smallest output that proves the +# outcome. Full logs stay in the terminal, never in the contract. +_decisive() { # _decisive + local flat + flat="$(printf '%s' "$1" | tr '\n\r\t' ' ' | tr -s ' ')" + flat="$(_trim "$flat")" + if [ "${#flat}" -gt "$EVIDENCE_CAP" ]; then + printf '%s…' "${flat:0:$EVIDENCE_CAP}" + else + printf '%s' "$flat" + fi +} + +# Fail-closed: exit 0 AND the marker. A nonzero process never passes because +# its error text happens to contain the expected token. +_run_one() { # _run_one + local i="$1" out rc + out="$(timeout "$TIMEOUT" bash -c "${_CHECK[i]}" 2>&1)" + rc=$? + _STATUS[i]="NOT-MET" + if [ "$rc" -eq 124 ]; then + _EVID[i]="NOT-MET timeout=${TIMEOUT}s" + elif [ "$rc" -ne 0 ]; then + _EVID[i]="NOT-MET exit=$rc (nonzero) :: $(_decisive "$out")" + elif [[ "$out" != *"${_EXPECT[i]}"* ]]; then + _EVID[i]="NOT-MET exit=0 marker-absent :: $(_decisive "$out")" + else + _STATUS[i]="MET" + _EVID[i]="MET exit=0 marker-found :: $(_decisive "$out")" + fi +} + +_run_all() { + local i + for ((i = 0; i < ${#_ID[@]}; i++)); do + _STATUS[i]=""; _EVID[i]="" + [ -n "${_CHECK[i]}" ] && _run_one "$i" + done +} + +_evline_owner() { # _evline_owner — echoes idx, or nothing + local i + for ((i = 0; i < ${#_ID[@]}; i++)); do + if [ "${_EVLINE[i]}" = "$1" ] && [ -n "${_EVID[i]}" ]; then + printf '%s' "$i" + return 0 + fi + done +} + +# Rewrites only the EVIDENCE lines of criteria that actually ran; every other +# byte of the contract is copied through, indentation included. +_write_back() { # _write_back + local tmp line n=0 idx + tmp="$(mktemp)" || _die "mktemp failed" + while IFS= read -r line || [ -n "$line" ]; do + n=$((n + 1)) + idx="$(_evline_owner "$n")" + if [ -n "$idx" ]; then + printf '%s%s\n' "${line%%[![:space:]]*}" "EVIDENCE: ${_EVID[idx]}" + else + printf '%s\n' "$line" + fi + done < "$1" > "$tmp" + cat "$tmp" > "$1" && rm -f "$tmp" +} + +# ── report ────────────────────────────────────────────────────────────────── + +# A recorded `pending`, or a criterion that never ran, is PENDING — never MET. +# `status` reports what the file says; it does not revalidate old evidence. +_row_state() { # _row_state + local i="$1" + _is_abandoned "${_ID[i]}" && { printf 'ABANDONED'; return 0; } + [ -z "${_CHECK[i]}" ] && { printf 'MANUAL'; return 0; } + [ -n "${_STATUS[i]:-}" ] && { printf '%s' "${_STATUS[i]}"; return 0; } + case "${_EVTEXT[i]}" in + MET' '*) printf 'MET-RECORDED' ;; + *) printf 'PENDING' ;; + esac +} + +_report_rows() { + local i state + for ((i = 0; i < ${#_ID[@]}; i++)); do + state="$(_row_state "$i")" + printf ' %-3s %-13s %s\n' "${_ID[i]}" "$state" "${_TEXT[i]}" + done +} + +_report_abandons() { + local i + for ((i = 0; i < ${#_ABANDON_ID[@]}; i++)); do + printf ' ABANDONED %s — %s\n' "${_ABANDON_ID[i]}" "${_ABANDON_WHY[i]}" + done +} + +_count_state() { # _count_state + local i n=0 + for ((i = 0; i < ${#_ID[@]}; i++)); do + [ "$(_row_state "$i")" = "$1" ] && n=$((n + 1)) + done + printf '%s' "$n" +} + +_verdict() { # _verdict — prints the line, returns the rc + local unmet pending abandoned + if [ "${#_ERRORS[@]}" -gt 0 ]; then + printf 'GATES — VERDICT: ERROR(%s)\n' "${#_ERRORS[@]}" + return 2 + fi + unmet="$(_count_state NOT-MET)" + pending="$(_count_state PENDING)" + abandoned="$(_count_state ABANDONED)" + [ "$unmet" -gt 0 ] && + { printf 'GATES — VERDICT: UNMET(%s)\n' "$unmet"; return 2; } + if [ "$1" = "status" ] && [ "$pending" -gt 0 ]; then + printf 'GATES — VERDICT: PENDING(%s)\n' "$pending" + return 2 + fi + [ "$abandoned" -gt 0 ] && + { printf 'GATES — VERDICT: ABANDONED(%s)\n' "$abandoned"; return 3; } + printf 'GATES — VERDICT: MET\n' + return 0 +} + +_report() { # _report + local rc + printf 'GATES — %s (%s)\n' "$2" "$1" + _report_rows + _report_abandons + [ "${#_ERRORS[@]}" -gt 0 ] && printf ' ERROR %s\n' "${_ERRORS[@]}" + printf 'RUNNABLE: %s of %s criteria; timeout %ss\n' \ + "$(_runnable_count)" "${#_ID[@]}" "$TIMEOUT" + _verdict "$1" + rc=$? + return "$rc" +} + +_runnable_count() { + local i n=0 + for ((i = 0; i < ${#_ID[@]}; i++)); do + [ -n "${_CHECK[i]}" ] && n=$((n + 1)) + done + printf '%s' "$n" +} + +# ── entry point ───────────────────────────────────────────────────────────── + +main() { # main + local mode="$1" file="$2" + [ -r "$file" ] || _die "contract unreadable: $file" + _parse "$file" + [ "${#_ID[@]}" -gt 0 ] || + _die "no numbered criteria under ## ACCEPTANCE CRITERIA" + _validate_oracles + _validate_abandons + if [ "$mode" = "run" ] && [ "${#_ERRORS[@]}" -eq 0 ]; then + _run_all + _write_back "$file" + fi + _report "$mode" "$file" +} + +case "${1:-}" in + status|run) + [ $# -eq 2 ] || _die "usage: gates.sh {status|run} " + main "$1" "$2" + ;; + *) _die "usage: gates.sh {status|run} " ;; +esac diff --git a/lib/tests/contract-verifier.test.sh b/lib/tests/contract-verifier.test.sh index af7b99f..ce77347 100644 --- a/lib/tests/contract-verifier.test.sh +++ b/lib/tests/contract-verifier.test.sh @@ -67,7 +67,7 @@ fi tr_ "frontmatter name" "$AGT" "^name: verifier$" tr_ "tools read-only set" "$AGT" "^tools: Read, Grep, Glob, Bash$" tn "no write-capable tools" "$AGT" "^tools:.*(Edit|Write|NotebookEdit)" -tf "verdict grammar" "$AGT" "VERIFY — VERDICT: CONFORME | ECARTS(n) | ERROR()" +tf "verdict grammar" "$AGT" "VERIFY — VERDICT: CONFORME | ECARTS(n) | ABANDONED(n) | ERROR()" tf "blind — no iteration history" "$AGT" "NEVER receive iteration history" tf "blind — complete every time" "$AGT" "every verification is complete and blind" tf "unverifiable is not met" "$AGT" "\`UNVERIFIABLE\` ≠ \`MET\`" diff --git a/lib/tests/gates.test.sh b/lib/tests/gates.test.sh new file mode 100644 index 0000000..76d5fa7 --- /dev/null +++ b/lib/tests/gates.test.sh @@ -0,0 +1,317 @@ +#!/usr/bin/env bash +# ============================================================ +# lib/gates.sh — behavioural tests + structure locks for the +# deterministic floor (GATE 0, lib/verify-secure-loop.md). +# +# Fail-closed is the entire point of this runner, so every +# "looks green but must not pass" case is asserted explicitly: +# nonzero exit carrying the marker, marker absent, timeout, +# unindented attribute silently demoting a gate to manual. +# Non-execution is proved with a sentinel file, and the +# sentinel's own positive control is asserted first — an +# absence check that was never able to fire proves nothing. +# ============================================================ +set -uo pipefail + +REPO="$(cd "$(dirname "$0")/../.." && pwd)" +GATES="$REPO/lib/gates.sh" +WORK="$(mktemp -d)" +trap 'rm -rf "$WORK"' EXIT +PASS=0; FAIL=0; N=0 +LAST="" + +ok() { echo " PASS $1"; PASS=$((PASS + 1)); } +bad() { echo " FAIL $1 — $2"; FAIL=$((FAIL + 1)); } + +# gate