feat(gates): deterministic floor (GATE 0) under the fresh verifier

GATE 1 is an LLM dispatch and the verifier's mandatory PROOF: line is a line
the verifier writes — nothing structurally stops it being produced without
anything being executed. Nothing deterministic sat between the executor and
that dispatch.

An acceptance criterion can now carry an oracle: indented CHECK: (command),
EXPECT: (success-only marker), EVIDENCE: (slot). lib/gates.sh runs them
fail-closed — MET requires exit 0 AND the marker, so a nonzero process never
passes on its error text carrying the token — and writes the outcome back
into the contract, so the fresh verifier reads evidence as fact rather than
trusting the executor's report.

GATE 0 runs that floor before any verifier is dispatched; a red build sends
the executor back for free, on its own iteration budget. ABANDON: <id>
<reason> turns an impossible criterion into a visible handoff that blocks
CONFORME and routes to the human gate, via the new ABANDONED(n) verdict —
a distinct token because it routes distinctly, never a dev loop. feater and
bugfixer gain a four-pass completion discipline, scoped so a pass can never
widen the contract.

The runner's parse fails closed on partial oracles, duplicate ids,
unindented attributes and runnable criteria with no EVIDENCE: line, and
executes nothing at all when the ledger is malformed. status never executes
and never writes; run always re-executes, since trusting written evidence is
the failure being closed.

Adapted from the unlazy skill (Leonxlnx/unlazy, MIT). Its Stop hook,
approval store, .unlazy/ tree, depth-tree arithmetic and Node checker were
deliberately refused — BDR-083 records each reason.

64 assertions in lib/tests/gates.test.sh, non-execution proved by sentinel
with its own positive control asserted first.
This commit is contained in:
Bastien Chanot
2026-08-24 13:12:38 +02:00
parent 5488c4870f
commit 63310467ca
8 changed files with 826 additions and 20 deletions
+18
View File
@@ -46,6 +46,24 @@ Every choice was made in the plan or is a NEED-DECISION to report.
security/verifier dispatch, editing `.claude/**` or memory registries, user
questions (you cannot ask — report instead), attribution trailers of any kind.
## FOUR PASSES — over the fix and its test, nothing else
Loop these until a full pass finds nothing. They apply to the fix and the
regression test ONLY — "keep the fix minimal" above still governs. They make
the minimal fix COMPLETE; they never widen it.
1. **Complete.** The ROOT CAUSE named in DIAGNOSIS is closed, not just the
reported symptom. No placeholder, no deferred remainder.
2. **Expert reread.** Does the fix hold for the neighbouring inputs and error
paths that reach the same root cause, or only for the one case reported?
3. **Negative control.** Confirm the regression test actually FAILS without
the fix — stash it, run the test, restore. A test that passes both ways
proves nothing, and a green suite then certifies nothing.
4. **Polish.** Naming and comments on what you touched. Nothing else.
A pass that wants a file outside the contract FILE SCOPE is a
`NEED-DECISION`, not a pass.
## OUTPUT — end with exactly this report (your final message)
```
+19
View File
@@ -57,6 +57,25 @@ report below is optional on this path (the dispatcher needs the edit applied
editing `.claude/**` or memory registries, user questions (you cannot
ask — report instead), attribution trailers of any kind.
## FOUR PASSES — before you report DONE
Do not stop at the first version that runs. Loop these until a full pass
finds nothing:
1. **Complete.** The whole deliverable the plan names is implemented. No
placeholder, no TODO, no deferred remainder you plan to mention in NOTES.
2. **Expert reread.** Read it as someone who owns this codebase. Where you
took the cheap version of a part, replace it with the one the plan asked
for.
3. **Defect hunt.** Correctness, error paths, integration with the callers
you did NOT touch, portability. Fix what you find.
4. **Polish.** Low-cost only: naming, comment density, dead code you
introduced.
Every pass stays inside the plan and the contract FILE SCOPE. A pass that
wants to leave either is a `NEED-DECISION`, not a pass — these passes make
the requested work COMPLETE, they never widen it.
## OUTPUT — end with exactly this report (your final message)
```
+40 -5
View File
@@ -48,6 +48,25 @@ Rules: read the diff AND enough surrounding code to judge behavior; run
criterion. Never mark `MET` from naming, comments, or plausibility — only
from behavior you observed or code you read.
### Criteria carrying an oracle (`CHECK:` / `EXPECT:` / `EVIDENCE:`)
`lib/gates.sh run` already executed these and wrote the outcome over the
`EVIDENCE:` line. Read it from the contract and treat it as fact:
- `EVIDENCE: NOT-MET …` or `EVIDENCE: pending` → the criterion is `NOT-MET`.
Reading the code NEVER overrides a red or unrun oracle. Cite the evidence
line as your evidence.
- `EVIDENCE: MET …` → the declared command passed. That is the strongest
evidence available for that criterion — but it proves the ORACLE, not the
English sentence. Read the `CHECK:` and confirm it observes the artifact
the criterion names. A vacuous oracle (`1. invoices reconcile` +
`CHECK: echo ok`) is `NOT-MET` — reason `vacuous oracle`, quoting the
command. That judgement is yours alone; no command can make it.
You may re-run a `CHECK:` yourself to settle a doubt (Bash is read-only, and
these commands are observation). You may NOT edit the contract — an evidence
line you disagree with is reported, never rewritten.
## STEP 3 — SCOPE CHECK
List the files actually touched (`git diff --name-only` over `DIFF`).
@@ -58,19 +77,30 @@ only enters the contract through a human micro-gate.
## STEP 4 — VERDICT
`CONFORME` ⇔ ALL criteria `MET` AND zero out-of-scope files.
Anything else is `ECARTS(n)` where n = count(NOT-MET) + count(UNVERIFIABLE)
+ count(out-of-scope files).
Read the contract's `ABANDON:` lines. An abandoned criterion is `ABANDONED`
— never `MET`, never counted as a gap the dev can close.
Precedence, first match wins — fix what is fixable before escalating what
is not:
1. `ERROR(<reason>)` — the contract is missing or unreadable.
2. `ECARTS(n)` — n = count(NOT-MET) + count(UNVERIFIABLE) + count(out-of-scope
files). Surface any abandonment in the same report.
3. `ABANDONED(n)` — zero gaps remain, but n abandonments stand. This is NOT
a pass and NOT a dev loop: it routes straight to the human gate.
4. `CONFORME` — ALL criteria `MET`, zero out-of-scope files, zero
abandonments.
## OUTPUT (exact format — machine-parsed by the orchestrator)
```
VERIFY — VERDICT: CONFORME | ECARTS(n) | ERROR(<reason>)
VERIFY — VERDICT: CONFORME | ECARTS(n) | ABANDONED(n) | ERROR(<reason>)
CONTRACT: <path>
CRITERIA:
1. <criterion> — MET — <evidence file:line | test ran → result>
1. <criterion> — MET — <EVIDENCE line | file:line | test ran → result>
2. <criterion> — NOT-MET — expected <…> / actual <…> — <file:line>
3. <criterion> — UNVERIFIABLE — <reason>
4. <criterion> — ABANDONED — <the reason recorded in the contract>
SCOPE: in-scope <n> files; out-of-scope: <list | none>
PROOF: read <n> files, ran <cmd → result | nothing>, checked <n>/<n> criteria
```
@@ -82,6 +112,8 @@ PROOF: read <n> files, ran <cmd → result | nothing>, checked <n>/<n> criteria
- `UNVERIFIABLE` ≠ `MET`. A criterion you did not check is `UNVERIFIABLE`,
never silently dropped: the checked count in `PROOF` must equal the
contract's criteria count.
- `ABANDONED` ≠ `MET`. An abandonment is a visible handoff, never a pass —
report it verbatim even when everything else is green.
- `PROOF` is MANDATORY. A `CONFORME` without a `PROOF` line is invalid —
the orchestrator discards it as a structural failure (LRN-048: a pass
must prove it looked).
@@ -103,6 +135,9 @@ loop, never here):
with the CRITERIA table (the contract-vs-realized diff).
- Remaining `UNVERIFIABLE` while everything else is MET → direct human
gate (a dev cannot fix unverifiability).
- `ABANDONED(n)` → direct human gate, never a dev loop. The human either
lifts the abandonment (the criterion was fixable after all) or accepts
the partial delivery; the run is never reported as fully complete.
- Structural failure (`ERROR(…)`, missing/duplicated VERDICT line,
unparsable output, agent crash, `CONFORME` without `PROOF`) → retry
ONCE with a fresh verifier; a 2nd structural failure → human
+54 -2
View File
@@ -35,6 +35,38 @@ ask what the repo can answer — verify paths/APIs/behavior yourself first.
this conversation.
- FILE SCOPE: paths/zones expected to change, or `repo-wide — <reason>`.
### ORACLES — a criterion a command can decide carries one
Give such a criterion an indented `CHECK:` (the command), `EXPECT:` (a
success-only marker), and `EVIDENCE: pending`.
`bash ~/.claude/lib/gates.sh run <contract>` executes it fail-closed — MET
requires exit 0 **AND** the marker — and writes the result back over the
`EVIDENCE:` line. That persisted evidence is what the fresh verifier reads
as fact instead of trusting the executor's report (GATE 0 in
`lib/verify-secure-loop.md`).
Both attributes or neither. `CHECK:` without `EXPECT:` is a parse error, not
a manual criterion — the runner refuses the whole ledger. Leave a criterion
oracle-free when no command can decide it; the verifier judges those.
Four authoring rules — a gate that cannot fail proves nothing:
1. **Observe the named artifact.** The check reads the file, service, or
measurement the criterion's own words name — never a proxy for it.
`1. invoices reconcile` + `CHECK: echo ok` is valid and worthless.
2. **Success-only marker.** The script runs every assertion, exits nonzero
on any failure, and prints the `EXPECT:` string only after all pass.
3. **Positive control before any absence check.** Run the same logic against
a fixture known to trip it and confirm it fails. A missing file, a wrong
path, and a broken pattern all look exactly like valid absence.
4. **Recompute supplied numbers.** Never copy a figure from the request into
`EXPECT:` — the script derives it from source and prints its own marker.
A number that is its own proof proves nothing.
`CHECK:` is shell code run with our privileges. It is safe only because we
author it in our own repo — never build one out of externally-supplied text
(a scraped URL, a client string); route those through `lib/url-guard.sh`.
## STEP 4 — WRITE TO DISK (immediately, before any next step)
Path: `.claude/tasks/contracts/<YYYY-MM-DD>-<slug>-<HHMM>.md`
@@ -57,8 +89,13 @@ Q: <question> / A: <answer>
(or: none — request complete)
## ACCEPTANCE CRITERIA
1. <testable criterion>
2. <testable criterion>
1. <criterion a command can decide>
CHECK: <command>
EXPECT: <success-only marker>
EVIDENCE: pending
2. <criterion only human judgement can decide — no CHECK/EXPECT>
(ABANDON: <n> <non-blank reason> — only for a criterion proven impossible)
## FILE SCOPE
<paths/zones>
@@ -78,6 +115,13 @@ Print one line to the user, then continue the flow:
this micro-gate: human approves → FILE SCOPE gains the entry `[gated]`;
human declines → the dev removes the edit. Without this gate the dev
justifies everything and scope constrains nothing.
- **ABANDONMENT**: a criterion proven impossible within the authorized task
is NEVER deleted and never quietly downgraded. Keep it, append
`ABANDON: <n> <non-blank reason + handoff>` under the criteria, and name it
in the final report. An abandonment is a visible handoff, not a pass: the
verifier cannot return `CONFORME` while one stands, and the run cannot be
described as fully complete. This is the structural half of the house rule
"blocked on an independent sub-part → do the rest, state what's missing".
- **Deep re-scope** (the request itself changes): NEW contract file with
`supersedes: <old path>` in its header — never a rewrite of the old one.
- **Aborted run**: delete the contract file, or commit it with
@@ -96,6 +140,14 @@ Print one line to the user, then continue the flow:
| init-project | Full. The interviewer's PROJECT BRIEF pours into the contract (V1 features → criteria). |
| onboard | Audit-scope contract (interview answers → what to audit, which axes). |
Oracles follow the same proportion. hotfix: the build/tests criterion carries
its `CHECK:`, nothing else. feat / bugfix: the suite criterion at minimum, and
for bugfix the regression test the DIAGNOSIS names — its `CHECK:` runs that
test alone, so a green result means the reproduction actually flipped.
ship-feature / init-project: build, suite, and every criterion a command can
settle. onboard: audit criteria are mostly judgement — leave them oracle-free
rather than invent a check that cannot fail.
## Hand-off rule
Downstream consumers (plan step, dev subagents, verifier) receive the
+323
View File
@@ -0,0 +1,323 @@
#!/usr/bin/env bash
# Deterministic floor under GATE 1: execute the acceptance criteria that the
# contract itself declares as oracles, fail-closed, and persist the evidence
# INTO the contract file.
#
# bash ~/.claude/lib/gates.sh status <contract> # parse only, never runs
# bash ~/.claude/lib/gates.sh run <contract> # execute + write evidence
#
# rc 0 = MET every runnable criterion passed, no abandonment standing
# 2 = UNMET a runnable criterion failed, or the ledger is malformed
# 3 = ABANDONED runnable criteria all passed, an abandonment still stands
#
# WHY: GATE 1 (lib/verify-secure-loop.md) is an LLM dispatch, and the
# verifier's mandatory `PROOF:` line is a line the verifier WRITES — nothing
# structurally stops it from being produced without anything being executed.
# This runs what the contract declares BEFORE a verifier is ever spawned: a
# red floor sends the executor back for free. Adapted from the `unlazy` skill
# (Leonxlnx/unlazy) — its gate ledger, minus the machinery we do not need.
#
# `run` always re-executes every runnable criterion, including ones already
# recorded MET. Trusting written evidence is exactly the failure this closes,
# so there is no incremental mode to get it wrong with.
#
# TRUST BOUNDARY: `CHECK:` is shell code, run with this process's privileges
# and environment. That is safe here only because the contract is authored by
# our own orchestrator in our own repo — which is why there is no approval
# store (we never execute ledgers inherited from a foreign repo). NEVER build
# a `CHECK:` out of externally-supplied text; route such values through
# lib/url-guard.sh first.
set -uo pipefail
TIMEOUT="${GATES_TIMEOUT:-120}"
EVIDENCE_CAP=140
# Module-level parse tables, index-aligned. Bash has no record type; threading
# eight parallel arrays through every call would cost more readability than
# the explicit data flow buys.
_ID=(); _TEXT=(); _CHECK=(); _EXPECT=(); _EVLINE=(); _EVTEXT=()
_STATUS=(); _EVID=()
_ABANDON_ID=(); _ABANDON_WHY=()
_ERRORS=()
_CUR=-1
_die() { printf 'GATES — VERDICT: ERROR(%s)\n' "$1"; exit 2; }
_err() { _ERRORS+=("$1"); }
_trim() {
local s="$1"
s="${s#"${s%%[![:space:]]*}"}"
printf '%s' "${s%"${s##*[![:space:]]}"}"
}
# ── parse ───────────────────────────────────────────────────────────────────
_new_crit() { # _new_crit <id> <text>
local i
for ((i = 0; i < ${#_ID[@]}; i++)); do
if [ "${_ID[i]}" = "$1" ]; then
_err "duplicate criterion id: $1"
# Orphan what follows instead of aliasing it onto the previous
# criterion, which would hand one gate another gate's oracle.
_CUR=-1
return 0
fi
done
_ID+=("$1"); _TEXT+=("$2")
_CHECK+=(""); _EXPECT+=(""); _EVLINE+=("0"); _EVTEXT+=("")
_CUR=$((${#_ID[@]} - 1))
}
_set_attr() { # _set_attr <CHECK|EXPECT|EVIDENCE> <value> <lineno>
if [ "$_CUR" -lt 0 ]; then
_err "$1 at line $3 belongs to no criterion"
return 0
fi
case "$1" in
CHECK) _CHECK[_CUR]="$2" ;;
EXPECT) _EXPECT[_CUR]="$2" ;;
EVIDENCE) _EVLINE[_CUR]="$3"; _EVTEXT[_CUR]="$2" ;;
esac
}
# An UNINDENTED attribute is diagnosed, never absorbed: silently ignoring it
# would demote a runnable criterion to a manual one, which is the one parse
# bug that turns this checker into a rubber stamp.
_absorb() { # _absorb <raw-line> <lineno>
local body
if [[ "$1" =~ ^([0-9]+)\.[[:space:]]+(.*)$ ]]; then
_new_crit "${BASH_REMATCH[1]}" "${BASH_REMATCH[2]}"
elif [[ "$1" =~ ^ABANDON:[[:space:]]*([0-9]+)?[[:space:]]*(.*)$ ]]; then
_ABANDON_ID+=("${BASH_REMATCH[1]}"); _ABANDON_WHY+=("${BASH_REMATCH[2]}")
elif [[ "$1" =~ ^(CHECK|EXPECT|EVIDENCE): ]]; then
_err "unindented ${BASH_REMATCH[1]}: at line $2"
elif [[ "$1" =~ ^[[:space:]]+(CHECK|EXPECT|EVIDENCE):(.*)$ ]]; then
body="$(_trim "${BASH_REMATCH[2]}")"
_set_attr "${BASH_REMATCH[1]}" "$body" "$2"
fi
}
_parse() { # _parse <file>
local line n=0 fence=0 inblock=0
while IFS= read -r line || [ -n "$line" ]; do
n=$((n + 1))
case "$line" in '```'*) fence=$((1 - fence)); continue ;; esac
[ "$fence" -eq 1 ] && continue
case "$line" in
'## ACCEPTANCE CRITERIA'*) inblock=1; continue ;;
'## '*) inblock=0; continue ;;
esac
[ "$inblock" -eq 1 ] && _absorb "$line" "$n"
done < "$1"
}
# ── validation ──────────────────────────────────────────────────────────────
_validate_oracles() {
local i
for ((i = 0; i < ${#_ID[@]}; i++)); do
if [ -n "${_CHECK[i]}" ] && [ -z "${_EXPECT[i]}" ]; then
_err "criterion ${_ID[i]}: CHECK without EXPECT (partial oracle)"
elif [ -z "${_CHECK[i]}" ] && [ -n "${_EXPECT[i]}" ]; then
_err "criterion ${_ID[i]}: EXPECT without CHECK (partial oracle)"
elif [ -n "${_CHECK[i]}" ] && [ "${_EVLINE[i]}" = "0" ]; then
_err "criterion ${_ID[i]}: runnable but has no EVIDENCE: line"
fi
done
}
_validate_abandons() {
local i j found
for ((i = 0; i < ${#_ABANDON_ID[@]}; i++)); do
found=0
for ((j = 0; j < ${#_ID[@]}; j++)); do
[ "${_ID[j]}" = "${_ABANDON_ID[i]}" ] && found=1
done
[ "$found" -eq 1 ] ||
_err "ABANDON names unknown criterion: '${_ABANDON_ID[i]}'"
[ -n "$(_trim "${_ABANDON_WHY[i]}")" ] ||
_err "ABANDON ${_ABANDON_ID[i]}: blank reason (a handoff needs one)"
done
}
_is_abandoned() { # _is_abandoned <criterion-id>
local i
for ((i = 0; i < ${#_ABANDON_ID[@]}; i++)); do
[ "${_ABANDON_ID[i]}" = "$1" ] && return 0
done
return 1
}
# ── execution ───────────────────────────────────────────────────────────────
# One line, capped, newlines flattened: the smallest output that proves the
# outcome. Full logs stay in the terminal, never in the contract.
_decisive() { # _decisive <combined-output>
local flat
flat="$(printf '%s' "$1" | tr '\n\r\t' ' ' | tr -s ' ')"
flat="$(_trim "$flat")"
if [ "${#flat}" -gt "$EVIDENCE_CAP" ]; then
printf '%s…' "${flat:0:$EVIDENCE_CAP}"
else
printf '%s' "$flat"
fi
}
# Fail-closed: exit 0 AND the marker. A nonzero process never passes because
# its error text happens to contain the expected token.
_run_one() { # _run_one <idx>
local i="$1" out rc
out="$(timeout "$TIMEOUT" bash -c "${_CHECK[i]}" 2>&1)"
rc=$?
_STATUS[i]="NOT-MET"
if [ "$rc" -eq 124 ]; then
_EVID[i]="NOT-MET timeout=${TIMEOUT}s"
elif [ "$rc" -ne 0 ]; then
_EVID[i]="NOT-MET exit=$rc (nonzero) :: $(_decisive "$out")"
elif [[ "$out" != *"${_EXPECT[i]}"* ]]; then
_EVID[i]="NOT-MET exit=0 marker-absent :: $(_decisive "$out")"
else
_STATUS[i]="MET"
_EVID[i]="MET exit=0 marker-found :: $(_decisive "$out")"
fi
}
_run_all() {
local i
for ((i = 0; i < ${#_ID[@]}; i++)); do
_STATUS[i]=""; _EVID[i]=""
[ -n "${_CHECK[i]}" ] && _run_one "$i"
done
}
_evline_owner() { # _evline_owner <lineno> — echoes idx, or nothing
local i
for ((i = 0; i < ${#_ID[@]}; i++)); do
if [ "${_EVLINE[i]}" = "$1" ] && [ -n "${_EVID[i]}" ]; then
printf '%s' "$i"
return 0
fi
done
}
# Rewrites only the EVIDENCE lines of criteria that actually ran; every other
# byte of the contract is copied through, indentation included.
_write_back() { # _write_back <file>
local tmp line n=0 idx
tmp="$(mktemp)" || _die "mktemp failed"
while IFS= read -r line || [ -n "$line" ]; do
n=$((n + 1))
idx="$(_evline_owner "$n")"
if [ -n "$idx" ]; then
printf '%s%s\n' "${line%%[![:space:]]*}" "EVIDENCE: ${_EVID[idx]}"
else
printf '%s\n' "$line"
fi
done < "$1" > "$tmp"
cat "$tmp" > "$1" && rm -f "$tmp"
}
# ── report ──────────────────────────────────────────────────────────────────
# A recorded `pending`, or a criterion that never ran, is PENDING — never MET.
# `status` reports what the file says; it does not revalidate old evidence.
_row_state() { # _row_state <idx>
local i="$1"
_is_abandoned "${_ID[i]}" && { printf 'ABANDONED'; return 0; }
[ -z "${_CHECK[i]}" ] && { printf 'MANUAL'; return 0; }
[ -n "${_STATUS[i]:-}" ] && { printf '%s' "${_STATUS[i]}"; return 0; }
case "${_EVTEXT[i]}" in
MET' '*) printf 'MET-RECORDED' ;;
*) printf 'PENDING' ;;
esac
}
_report_rows() {
local i state
for ((i = 0; i < ${#_ID[@]}; i++)); do
state="$(_row_state "$i")"
printf ' %-3s %-13s %s\n' "${_ID[i]}" "$state" "${_TEXT[i]}"
done
}
_report_abandons() {
local i
for ((i = 0; i < ${#_ABANDON_ID[@]}; i++)); do
printf ' ABANDONED %s — %s\n' "${_ABANDON_ID[i]}" "${_ABANDON_WHY[i]}"
done
}
_count_state() { # _count_state <state>
local i n=0
for ((i = 0; i < ${#_ID[@]}; i++)); do
[ "$(_row_state "$i")" = "$1" ] && n=$((n + 1))
done
printf '%s' "$n"
}
_verdict() { # _verdict <mode> — prints the line, returns the rc
local unmet pending abandoned
if [ "${#_ERRORS[@]}" -gt 0 ]; then
printf 'GATES — VERDICT: ERROR(%s)\n' "${#_ERRORS[@]}"
return 2
fi
unmet="$(_count_state NOT-MET)"
pending="$(_count_state PENDING)"
abandoned="$(_count_state ABANDONED)"
[ "$unmet" -gt 0 ] &&
{ printf 'GATES — VERDICT: UNMET(%s)\n' "$unmet"; return 2; }
if [ "$1" = "status" ] && [ "$pending" -gt 0 ]; then
printf 'GATES — VERDICT: PENDING(%s)\n' "$pending"
return 2
fi
[ "$abandoned" -gt 0 ] &&
{ printf 'GATES — VERDICT: ABANDONED(%s)\n' "$abandoned"; return 3; }
printf 'GATES — VERDICT: MET\n'
return 0
}
_report() { # _report <mode> <file>
local rc
printf 'GATES — %s (%s)\n' "$2" "$1"
_report_rows
_report_abandons
[ "${#_ERRORS[@]}" -gt 0 ] && printf ' ERROR %s\n' "${_ERRORS[@]}"
printf 'RUNNABLE: %s of %s criteria; timeout %ss\n' \
"$(_runnable_count)" "${#_ID[@]}" "$TIMEOUT"
_verdict "$1"
rc=$?
return "$rc"
}
_runnable_count() {
local i n=0
for ((i = 0; i < ${#_ID[@]}; i++)); do
[ -n "${_CHECK[i]}" ] && n=$((n + 1))
done
printf '%s' "$n"
}
# ── entry point ─────────────────────────────────────────────────────────────
main() { # main <status|run> <contract>
local mode="$1" file="$2"
[ -r "$file" ] || _die "contract unreadable: $file"
_parse "$file"
[ "${#_ID[@]}" -gt 0 ] ||
_die "no numbered criteria under ## ACCEPTANCE CRITERIA"
_validate_oracles
_validate_abandons
if [ "$mode" = "run" ] && [ "${#_ERRORS[@]}" -eq 0 ]; then
_run_all
_write_back "$file"
fi
_report "$mode" "$file"
}
case "${1:-}" in
status|run)
[ $# -eq 2 ] || _die "usage: gates.sh {status|run} <contract-path>"
main "$1" "$2"
;;
*) _die "usage: gates.sh {status|run} <contract-path>" ;;
esac
+1 -1
View File
@@ -67,7 +67,7 @@ fi
tr_ "frontmatter name" "$AGT" "^name: verifier$"
tr_ "tools read-only set" "$AGT" "^tools: Read, Grep, Glob, Bash$"
tn "no write-capable tools" "$AGT" "^tools:.*(Edit|Write|NotebookEdit)"
tf "verdict grammar" "$AGT" "VERIFY — VERDICT: CONFORME | ECARTS(n) | ERROR(<reason>)"
tf "verdict grammar" "$AGT" "VERIFY — VERDICT: CONFORME | ECARTS(n) | ABANDONED(n) | ERROR(<reason>)"
tf "blind — no iteration history" "$AGT" "NEVER receive iteration history"
tf "blind — complete every time" "$AGT" "every verification is complete and blind"
tf "unverifiable is not met" "$AGT" "\`UNVERIFIABLE\` ≠ \`MET\`"
+317
View File
@@ -0,0 +1,317 @@
#!/usr/bin/env bash
# ============================================================
# lib/gates.sh — behavioural tests + structure locks for the
# deterministic floor (GATE 0, lib/verify-secure-loop.md).
#
# Fail-closed is the entire point of this runner, so every
# "looks green but must not pass" case is asserted explicitly:
# nonzero exit carrying the marker, marker absent, timeout,
# unindented attribute silently demoting a gate to manual.
# Non-execution is proved with a sentinel file, and the
# sentinel's own positive control is asserted first — an
# absence check that was never able to fire proves nothing.
# ============================================================
set -uo pipefail
REPO="$(cd "$(dirname "$0")/../.." && pwd)"
GATES="$REPO/lib/gates.sh"
WORK="$(mktemp -d)"
trap 'rm -rf "$WORK"' EXIT
PASS=0; FAIL=0; N=0
LAST=""
ok() { echo " PASS $1"; PASS=$((PASS + 1)); }
bad() { echo " FAIL $1 — $2"; FAIL=$((FAIL + 1)); }
# gate <label> <mode> <expected-verdict> <expected-rc> <<< fixture-on-stdin
gate() {
local label="$1" mode="$2" want="$3" wantrc="$4" out rc
N=$((N + 1)); LAST="$WORK/c$N.md"
cat > "$LAST"
out="$(GATES_TIMEOUT="${GATES_TIMEOUT:-120}" \
bash "$GATES" "$mode" "$LAST" 2>&1)"
rc=$?
if [[ "$out" == *"$want"* ]] && [ "$rc" -eq "$wantrc" ]; then
ok "$label"
else
bad "$label" "want '$want' rc=$wantrc, got rc=$rc"
printf '%s\n' "$out" | sed 's/^/ /'
fi
}
has() {
if grep -qF -- "$2" "$LAST"; then ok "$1"; else bad "$1" "missing: $2"; fi
}
exists() {
if [ -e "$1" ]; then ok "$2"; else bad "$2" "sentinel absent: $1"; fi
}
absent() {
if [ -e "$1" ]; then bad "$2" "sentinel created: $1"; else ok "$2"; fi
}
echo "── fail-closed execution ──"
gate "exit 0 + marker = MET" run "GATES — VERDICT: MET" 0 <<'EOF'
## ACCEPTANCE CRITERIA
1. green
CHECK: echo "MARKER-OK"
EXPECT: MARKER-OK
EVIDENCE: pending
EOF
has "evidence written back" "EVIDENCE: MET exit=0 marker-found"
# The case a naive checker gets wrong: the marker IS in the output, but the
# process failed. Substring matching alone would certify a broken build.
gate "nonzero exit + marker = UNMET" run "GATES — VERDICT: UNMET(1)" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. lies
CHECK: echo "MARKER-OK"; exit 7
EXPECT: MARKER-OK
EVIDENCE: pending
EOF
has "nonzero recorded honestly" "NOT-MET exit=7 (nonzero)"
gate "exit 0 + no marker = UNMET" run "GATES — VERDICT: UNMET(1)" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. silent success is not success
CHECK: echo "something else"
EXPECT: MARKER-OK
EVIDENCE: pending
EOF
has "marker-absent recorded" "NOT-MET exit=0 marker-absent"
GATES_TIMEOUT=1 gate "timeout = UNMET" run "GATES — VERDICT: UNMET(1)" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. hangs
CHECK: sleep 5
EXPECT: never
EVIDENCE: pending
EOF
has "timeout recorded" "NOT-MET timeout=1s"
gate "manual-only contract passes through" run "RUNNABLE: 0 of 2" 0 <<'EOF'
## ACCEPTANCE CRITERIA
1. a human reads the copy
2. the design matches the brief
EOF
echo "── non-execution (sentinel), positive control first ──"
# Positive control: prove the sentinel mechanism can fire at all.
gate "sentinel fires when a CHECK runs" run "GATES — VERDICT: MET" 0 <<EOF
## ACCEPTANCE CRITERIA
1. control
CHECK: touch "$WORK/fired"; echo "M"
EXPECT: M
EVIDENCE: pending
EOF
exists "$WORK/fired" "positive control: sentinel created"
gate "status never executes" status "GATES — VERDICT: PENDING(1)" 2 <<EOF
## ACCEPTANCE CRITERIA
1. must not run
CHECK: touch "$WORK/status-ran"; echo "M"
EXPECT: M
EVIDENCE: pending
EOF
absent "$WORK/status-ran" "status did not execute"
has "status did not write evidence" "EVIDENCE: pending"
gate "fenced example is not a gate" run "RUNNABLE: 1 of 1" 0 <<EOF
## ACCEPTANCE CRITERIA
1. real
CHECK: echo "R"
EXPECT: R
EVIDENCE: pending
\`\`\`markdown
2. documentation example, invisible to the parser
CHECK: touch "$WORK/fenced-ran"; echo "nope"
EXPECT: nope
EVIDENCE: pending
\`\`\`
EOF
absent "$WORK/fenced-ran" "fenced CHECK never executed"
gate "malformed ledger executes nothing" run "GATES — VERDICT: ERROR" 2 <<EOF
## ACCEPTANCE CRITERIA
1. would run if the ledger parsed
CHECK: touch "$WORK/malformed-ran"; echo "M"
EXPECT: M
EVIDENCE: pending
2. partial oracle poisons the whole ledger
CHECK: echo "x"
EVIDENCE: pending
EOF
absent "$WORK/malformed-ran" "malformed ledger did not execute"
has "malformed ledger not written" "EVIDENCE: pending"
echo "── parse strictness ──"
gate "CHECK without EXPECT" run "CHECK without EXPECT" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. partial
CHECK: echo x
EVIDENCE: pending
EOF
gate "EXPECT without CHECK" run "EXPECT without CHECK" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. partial
EXPECT: x
EVIDENCE: pending
EOF
# An unindented CHECK must be diagnosed, never absorbed: silently ignoring it
# demotes a runnable criterion to a manual one — the one parse bug that turns
# this runner into a rubber stamp.
gate "unindented attribute is diagnosed" run "unindented CHECK:" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. sneaky
CHECK: echo x
EXPECT: x
EVIDENCE: pending
EOF
gate "runnable without EVIDENCE line" run "has no EVIDENCE: line" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. no ledger slot
CHECK: echo x
EXPECT: x
EOF
gate "duplicate criterion id" run "duplicate criterion id: 1" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. first
EVIDENCE: pending
1. second
EVIDENCE: pending
EOF
# After a rejected duplicate the following attributes must be orphaned, not
# aliased onto the previous criterion — that would hand one gate another's
# oracle and let a stale EVIDENCE line satisfy it.
gate "duplicate orphans what follows" run "belongs to no criterion" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. real
CHECK: echo x
EXPECT: x
EVIDENCE: pending
1. duplicate
CHECK: echo y
EXPECT: y
EVIDENCE: pending
EOF
gate "no numbered criteria" run "no numbered criteria" 2 <<'EOF'
## ACCEPTANCE CRITERIA
nothing numbered here
EOF
echo "── abandonment ──"
gate "valid abandonment = rc 3" run "GATES — VERDICT: ABANDONED(1)" 3 <<'EOF'
## ACCEPTANCE CRITERIA
1. green
CHECK: echo "M"
EXPECT: M
EVIDENCE: pending
2. impossible
EVIDENCE: pending
ABANDON: 2 upstream API offline; handoff recorded in BLK-099
EOF
gate "blank abandonment reason" run "blank reason" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. green
EVIDENCE: pending
ABANDON: 1
EOF
gate "abandonment naming nothing" run "unknown criterion" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. green
EVIDENCE: pending
ABANDON: 9 names a criterion that does not exist
EOF
echo "── usage ──"
# usage <label> <expected-substring> <argv...>
usage() {
local label="$1" want="$2" out rc; shift 2
out="$(bash "$GATES" "$@" 2>&1)"; rc=$?
if [ "$rc" -eq 2 ] && [[ "$out" == *"$want"* ]]; then
ok "$label"
else
bad "$label" "rc=$rc out=$out"
fi
}
usage "no args = ERROR rc 2" "usage:"
usage "missing contract = ERROR rc 2" "contract unreadable" run "$WORK/nope.md"
usage "unknown mode refused" "usage:" frobnicate "$WORK/c1.md"
# ── structure locks on the doctrine this runner is wired into ───────────────
CI="$REPO/lib/contract-interview.md"
VS="$REPO/lib/verify-secure-loop.md"
AGT="$REPO/agents/verifier.md"
FE="$REPO/agents/feater.md"
BF="$REPO/agents/bugfixer.md"
lock() { # lock <label> <file> <fixed-string>
if grep -qF -- "$3" "$2" 2>/dev/null; then
ok "$1"
else
bad "$1" "missing: $3"
fi
}
echo "── contract-interview.md oracle doctrine ──"
lock "oracle section" "$CI" "### ORACLES"
lock "runner named" "$CI" "lib/gates.sh run <contract>"
lock "fail-closed spelled out" "$CI" "exit 0 **AND** the marker"
lock "both or neither" "$CI" "Both attributes or neither"
lock "rule observe artifact" "$CI" "Observe the named artifact"
lock "rule success-only" "$CI" "Success-only marker"
lock "rule positive control" "$CI" "Positive control before any absence check"
lock "rule recompute numbers" "$CI" "Recompute supplied numbers"
lock "shell trust boundary" "$CI" "url-guard.sh"
lock "template carries oracle" "$CI" "EXPECT: <success-only marker>"
lock "abandonment lifecycle" "$CI" "**ABANDONMENT**"
lock "abandonment not deleted" "$CI" "NEVER deleted"
echo "── verify-secure-loop.md GATE 0 ──"
lock "gate 0 exists" "$VS" "## GATE 0 — DETERMINISTIC FLOOR"
lock "gate 0 no dispatch" "$VS" "**No verifier is dispatched**"
lock "gate 0 loop bound" "$VS" "Max 3 floor iterations"
lock "gate 0 budget separate" "$VS" "not eat the conformity budget"
lock "malformed = main loop" "$VS" "never dispatch a dev for it"
lock "order invariant" "$VS" "GATE 0 → GATE 1 → GATE 2"
echo "── verifier.md oracle + abandonment ──"
lock "verdict grammar" "$AGT" \
"VERIFY — VERDICT: CONFORME | ECARTS(n) | ABANDONED(n) | ERROR(<reason>)"
lock "red oracle wins" "$AGT" "NEVER overrides a red or unrun oracle"
lock "vacuous oracle caught" "$AGT" "vacuous oracle"
lock "oracle != english" "$AGT" "proves the ORACLE, not the"
lock "never edits contract" "$AGT" "reported, never rewritten"
lock "abandoned is not met" "$AGT" "\`ABANDONED\` ≠ \`MET\`"
lock "abandoned routes human" "$AGT" "direct human gate, never a dev loop"
echo "── executor four passes ──"
lock "feater passes" "$FE" "## FOUR PASSES"
lock "feater no placeholder" "$FE" "no deferred remainder you plan"
lock "feater never widens" "$FE" "they never widen it"
lock "bugfixer passes" "$BF" "## FOUR PASSES"
lock "bugfixer stays minimal" "$BF" "keep the fix minimal"
lock "bugfixer neg control" "$BF" "**Negative control.**"
lock "bugfixer test must fail" "$BF" "A test that passes both ways"
echo ""
echo "gates: $PASS pass, $FAIL fail"
[ "$FAIL" -eq 0 ]
+54 -12
View File
@@ -13,8 +13,40 @@ Inputs the caller must have ready:
pre-dev SHA, or the working-tree diff before commit).
- `TEST`: the project test command, if known.
Nominal path is cheap: one verifier dispatch + one security dispatch, done.
The loop only costs more when it actually loops.
Nominal path is cheap — a free floor run, then
one verifier dispatch + one security dispatch, done. The loop only costs
more when it actually loops.
## GATE 0 — DETERMINISTIC FLOOR (no dispatch, no model)
Before spending a verifier dispatch, execute the oracles the contract itself
declares:
```bash
bash ~/.claude/lib/gates.sh run "$CONTRACT"
```
It runs every `CHECK:` fail-closed (MET requires exit 0 AND the `EXPECT:`
marker) and writes the outcome back over each `EVIDENCE:` line. Parse its
single `GATES — VERDICT:` line:
- `MET` → floor green, go to GATE 1. An all-manual contract lands here too
(`RUNNABLE: 0 of n`) and passes straight through.
- `UNMET(n)` → hand the dev the CONTRACT path + the `NOT-MET` rows verbatim,
nothing else; re-run GATE 0. **No verifier is dispatched** — a red build or
a red suite is not a judgement call, and paying an LLM to discover it is
waste. **Max 3 floor iterations** → STOP + human escalation with the rows.
- `ABANDONED(n)` → floor green but a handoff stands. Continue to GATE 1; the
verifier surfaces it and its `ABANDONED(n)` verdict routes to the human
gate.
- `ERROR(n)` → the ledger is malformed (partial oracle, duplicate id,
unindented attribute, runnable criterion with no `EVIDENCE:` line). The
contract is the ORCHESTRATOR's own artifact — fix it here in the main loop,
never dispatch a dev for it.
Floor iterations are counted separately from GATE 1's: a cheap loop here does
not eat the conformity budget. GATE 0 also runs unchanged after every
security fix round, before re-verifying the request.
## GATE 1 — REQUEST CONFORMITY (fresh verifier)
@@ -29,9 +61,13 @@ Parse its single `VERIFY — VERDICT:` line:
- `ECARTS(n)` → hand the dev the CONTRACT path + the exact `CRITERIA` gap
lines (NOT-MET / out-of-scope), nothing else. Inline dev fixes in place;
a dispatched dev is re-dispatched FRESH with those inputs only. Then
re-dispatch a FRESH verifier. Repeat. **Max 3 conformity iterations** →
STOP + human escalation with the CRITERIA table (the contract-vs-realized
diff).
re-run GATE 0 and re-dispatch a FRESH verifier. Repeat.
**Max 3 conformity iterations** → STOP + human escalation with the
CRITERIA table (the contract-vs-realized diff).
- `ABANDONED(n)` → direct human gate, never a dev loop (a dev cannot close
what was proven impossible). The human lifts the abandonment or accepts
the partial delivery; either way the run is never reported as fully
complete, and the abandonment is named in the final report.
- Remaining `UNVERIFIABLE` while all else MET → direct human gate (a dev
cannot fix unverifiability); do not spend a loop on it.
- Out-of-scope files: a dev justification is accepted ONLY through the human
@@ -53,10 +89,11 @@ Parse its single `SECURITY — VERDICT:` line:
- `PASS` → done, proceed to commit.
- `BLOCK(n)` → hand the dev the `BLOCKING` list + the CONTRACT path (inline
fix, or FRESH executor re-dispatch). Then **re-verify the REQUEST first** (GATE 1, fresh
verifier) — a security fix can drift the behavior — **then re-run GATE 2**
(fresh auditor), in that order. **Max 3 security iterations** → STOP +
human escalation with the BLOCKING table.
fix, or FRESH executor re-dispatch). Then re-run GATE 0, then
**re-verify the REQUEST first** (GATE 1, fresh verifier) — a security fix
can drift the behavior — **then re-run GATE 2** (fresh auditor), in that
order. **Max 3 security iterations** → STOP + human escalation with the
BLOCKING table.
- `DEGRADED` (semgrep absent) → does NOT block on the tool's absence; surface
the checklist result + recommend `make plugin`. A DEGRADED run that still
BLOCKs (grep-caught secret/injection) blocks like any other.
@@ -65,6 +102,11 @@ Parse its single `SECURITY — VERDICT:` line:
## Order invariant
REQUEST conformity is always re-checked BEFORE security on any re-loop — a
security fix that breaks the feature must not slip through because only the
security gate re-ran. Never the reverse order.
Every re-loop replays the gates in order: **GATE 0 → GATE 1 → GATE 2**,
never a subset and never reversed.
The floor runs first because it is free, and because a red build makes the
verifier's verdict meaningless. REQUEST conformity is
always re-checked BEFORE security on any re-loop — a security fix that breaks
the feature must not slip through because only the security gate re-ran.
Never the reverse order.