Merge bugfix/make-test-names-red-suites into develop
This commit is contained in:
@@ -61,6 +61,7 @@ rules:
|
|||||||
| EVAL-038 | 2026-09-29 | correction of EVAL-037: 94 % of sub-agent usage records carry no `output_tokens_details` (Fable subs at xhigh read 0 thinking, impossible with always-on thinking) → sub-agent thinking UNMEASURED, not ≈0; main loop 100 % counted; weighted-cost split (61/39) still holds | `effort-audit.py` prints coverage + CAVEAT; cite the cost split only; agent effort pins stay unmeasured; a tier move on a price argument = judgment, not figure |
|
| EVAL-038 | 2026-09-29 | correction of EVAL-037: 94 % of sub-agent usage records carry no `output_tokens_details` (Fable subs at xhigh read 0 thinking, impossible with always-on thinking) → sub-agent thinking UNMEASURED, not ≈0; main loop 100 % counted; weighted-cost split (61/39) still holds | `effort-audit.py` prints coverage + CAVEAT; cite the cost split only; agent effort pins stay unmeasured; a tier move on a price argument = judgment, not figure |
|
||||||
| EVAL-039 | 2026-09-30 | ship-feature run higgsfield-pack: plan dry-run in scratch → 0 executor failure on 7 tasks; challenge found 7 MAJOR I missed; floor-guard caught 2 shellcheck suppressions of mine; final review found README/code gap | keep |
|
| EVAL-039 | 2026-09-30 | ship-feature run higgsfield-pack: plan dry-run in scratch → 0 executor failure on 7 tasks; challenge found 7 MAJOR I missed; floor-guard caught 2 shellcheck suppressions of mine; final review found README/code gap | keep |
|
||||||
| EVAL-040 | 2026-10-08 | model-router w1a plan: 3 challengers + 1 confirmation found 2 BLOCKER + 14 MAJOR on a plan judged closed; executor then passed every gate first time | keep the round, never dispatch a mod plan without it |
|
| EVAL-040 | 2026-10-08 | model-router w1a plan: 3 challengers + 1 confirmation found 2 BLOCKER + 14 MAJOR on a plan judged closed; executor then passed every gate first time | keep the round, never dispatch a mod plan without it |
|
||||||
|
| EVAL-041 | 2026-10-09 | model-router W1-C plan: 3 lenses FATAL (4 BLOCKER + 20 MAJOR) then 2 confirmations each FATAL with a NEW BLOCKER in my own revision; executor DONE first pass, 3 short text/hardening rounds | one confirmation is not enough when a revision removes a whole mechanism; the plan carried the risk, the code almost none |
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -379,3 +380,10 @@ Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itse
|
|||||||
- **Anomaly**: r1 carried 2 BLOCKER (explicit Agent `model` overridden at every step; Skill bridge answer shape refused by the output schema → doublon kept) + 10 MAJOR, all invisible to me: spike levers carried over as design (`agentsNext`, per-step model rewrite), a table copying 21 pins = third source of truth, writes on main from sub-agent loops. Confirmation found 4 more MAJOR (explicit params vs in-agent writes, `e.wait`, vacuous test assertions, haiku effort). Executor then DONE first pass, verifier CONFORME 6/6, security PASS: the plan was the whole risk.
|
- **Anomaly**: r1 carried 2 BLOCKER (explicit Agent `model` overridden at every step; Skill bridge answer shape refused by the output schema → doublon kept) + 10 MAJOR, all invisible to me: spike levers carried over as design (`agentsNext`, per-step model rewrite), a table copying 21 pins = third source of truth, writes on main from sub-agent loops. Confirmation found 4 more MAJOR (explicit params vs in-agent writes, `e.wait`, vacuous test assertions, haiku effort). Executor then DONE first pass, verifier CONFORME 6/6, security PASS: the plan was the whole risk.
|
||||||
- **Action**: a mod plan always goes through the full round + confirmation; test assertions must read the one line that carries the value; spike code is a FACT source, never a design source ([[BDR-115]], [[LRN-205]], [[LRN-206]]).
|
- **Action**: a mod plan always goes through the full round + confirmation; test assertions must read the one line that carries the value; spike code is a FACT source, never a design source ([[BDR-115]], [[LRN-205]], [[LRN-206]]).
|
||||||
|
|
||||||
|
## EVAL-041 — model-router W1-C: the plan was the whole risk, two confirmations were needed
|
||||||
|
- **Date**: 2026-10-09
|
||||||
|
- **Output checked**: plan `.claude/tasks/plans/2026-10-09-model-router-tiers-1237.md` r1 → r4 (absolute tiers, breaker, derived phases), written with every engine fact in hand.
|
||||||
|
- **Method**: 3 blind opus challengers (simplicity FATAL(6), correctness FATAL(11), robustness FATAL(11)) → r2; confirmation FATAL(8) with a NEW BLOCKER introduced by r2 (per-step engine-fallback detection climbing the backoff) → r3; second confirmation FATAL(4) with a NEW BLOCKER introduced by r3 (`lastPlan` reset vs kept) → r4; executor DONE first pass; verifier ECARTS ×2 on texts (3 gaps) + hardening (4 items) → CONFORME; security PASS.
|
||||||
|
- **Anomaly**: 6 BLOCKER + 29 MAJOR over four revisions, each confirmation found a flaw my own fix had introduced; the doctrine cap (one confirmation) would have shipped r2 with a 5-hour false outage. The executor never needed a re-dispatch for logic: all later rounds were text truthfulness and hardening.
|
||||||
|
- **Action**: when a revision REMOVES or REPLACES a mechanism, re-challenge once more (state the deviation); keep plan sections additive with an explicit precedence line (r4 > r3 > r2) so executors and verifiers read one law; name superseded clauses of prior contracts in the Disposition. Links [[EVAL-040]], [[LRN-207]], [[BDR-115]].
|
||||||
|
|
||||||
|
|||||||
@@ -588,3 +588,4 @@ rules:
|
|||||||
- model-router W1-C adaptive tiers landed (d0fa100): plan r1→r4 through 3 lenses (all FATAL: 4 BLOCKER + 20 MAJOR) + 2 confirmations (1 BLOCKER each, in my own r2 then r3) → deviation from the one-confirmation cap, stated. Design: absolute tiers, StopFailure-kind breaker + PostModelSwitch auto, fallback chain, main upgrade under a 200k cap (fails closed), sticky turnModel, derived orchestrate (background dispatches), prompt default rules with skip rules; classifier deferred. feater DONE → 2 gap rounds (texts) → hardening (leaveDown gates, cap fail-closed, one-way prefix, log key) → verifier CONFORME, security PASS (LOW only). 58 tests. Lesson: I sent iteration history in a security brief; the auditor contract forbids it (blind scan) — scope only next time. Live checks pending after /reload-plugins (R16/T8). Next: user reload, live test, wave 2.
|
- model-router W1-C adaptive tiers landed (d0fa100): plan r1→r4 through 3 lenses (all FATAL: 4 BLOCKER + 20 MAJOR) + 2 confirmations (1 BLOCKER each, in my own r2 then r3) → deviation from the one-confirmation cap, stated. Design: absolute tiers, StopFailure-kind breaker + PostModelSwitch auto, fallback chain, main upgrade under a 200k cap (fails closed), sticky turnModel, derived orchestrate (background dispatches), prompt default rules with skip rules; classifier deferred. feater DONE → 2 gap rounds (texts) → hardening (leaveDown gates, cap fail-closed, one-way prefix, log key) → verifier CONFORME, security PASS (LOW only). 58 tests. Lesson: I sent iteration history in a security brief; the auditor contract forbids it (blind scan) — scope only next time. Live checks pending after /reload-plugins (R16/T8). Next: user reload, live test, wave 2.
|
||||||
- W1-C live checks after the user's /reload-plugins (skills-dir copy, 12 hooks): derived orchestrate real (main high→medium during a background Explore→high after), Explore sonnet/medium, route plan → xhigh, Skill(effort-low) bridge → low, both confirmed in engine records; /route show resolves 10 phases to full ids, down none; steps carry bare ids (no [1m]) → suffix carry inert here. Breaker/auto/StopFailure order wait for a real incident. Branch feature/model-router-mod: 22 commits ahead, unpushed (manual), merge = human signal. Wave 2 queued.
|
- W1-C live checks after the user's /reload-plugins (skills-dir copy, 12 hooks): derived orchestrate real (main high→medium during a background Explore→high after), Explore sonnet/medium, route plan → xhigh, Skill(effort-low) bridge → low, both confirmed in engine records; /route show resolves 10 phases to full ids, down none; steps carry bare ids (no [1m]) → suffix carry inert here. Breaker/auto/StopFailure order wait for a real incident. Branch feature/model-router-mod: 22 commits ahead, unpushed (manual), merge = human signal. Wave 2 queued.
|
||||||
- User go 'ok merge': full make test green (except env red design-tool-gate), shellcheck clean → gitflow finish feature model-router-mod → develop abbdf79 (22 commits: waves 0, 1-A, 1-B1, 1-B2, 1-C + docs + registries). Manual mode: develop NOT pushed, the user publishes by hand from the terminal. Branch removed locally. User will /clear before wave 2.
|
- User go 'ok merge': full make test green (except env red design-tool-gate), shellcheck clean → gitflow finish feature model-router-mod → develop abbdf79 (22 commits: waves 0, 1-A, 1-B1, 1-B2, 1-C + docs + registries). Manual mode: develop NOT pushed, the user publishes by hand from the terminal. Branch removed locally. User will /clear before wave 2.
|
||||||
|
- /hotfix make-test-names-red-suites (user: '9 min pour un merge?'): measured from transcript timestamps, the merge took <20 s; 454 s went to make test run TWICE (full + per-suite sweep to name the red suite, because the aggregate rc is silent). Fix: Makefile test prints FAIL <suite> + summary, rc unchanged (GNU make returns 2 on a failed recipe; my first oracle expected 1). hotfixer DONE, GATE 0 MET 2/2, security PASS. Branch bugfix/make-test-names-red-suites UNMERGED (human signal). Method note: a mods/-only diff needs only the mods suite + doctrine census, the full run once before merge.
|
||||||
|
|||||||
@@ -1765,3 +1765,15 @@ Rule: when editing a doctrine file under structure locks, grep the test's lock s
|
|||||||
- **Context**: model-router tests. Kit `$` is `EngineCall<E> = (e: Args<E>)`: `command.run` needs `origin` + `presentation`, `prompt.submit` needs `wait` + `origin`, `agent.spawn` needs `tool_use_id, description, provider, parentModel, background, fork`; the test file is type-checked with the hooks (tsc include). `session.start` does NOT fire at load → every test boots with a bottom `on('session.start')` + `$.session.start({ cwd, surface: null, isInteractive: false })`. A hook calling `next` hits the kit's bottom which throws unless the test registered one (`on('agent.spawn', ($, e) => ({ model: e.model, agentId: 'a1' }))`). `$.turn.step` returns a stream: drain with `for await` then await `.result` (awaiting `.result` alone runs no hook). No fs/network/process: defaults path only. Assert on the ONE line that carries the value (a `show()` listing every phase always contains every id and level).
|
- **Context**: model-router tests. Kit `$` is `EngineCall<E> = (e: Args<E>)`: `command.run` needs `origin` + `presentation`, `prompt.submit` needs `wait` + `origin`, `agent.spawn` needs `tool_use_id, description, provider, parentModel, background, fork`; the test file is type-checked with the hooks (tsc include). `session.start` does NOT fire at load → every test boots with a bottom `on('session.start')` + `$.session.start({ cwd, surface: null, isInteractive: false })`. A hook calling `next` hits the kit's bottom which throws unless the test registered one (`on('agent.spawn', ($, e) => ({ model: e.model, agentId: 'a1' }))`). `$.turn.step` returns a stream: drain with `for await` then await `.result` (awaiting `.result` alone runs no hook). No fs/network/process: defaults path only. Assert on the ONE line that carries the value (a `show()` listing every phase always contains every id and level).
|
||||||
- **Apply**: write the boot helper first, type every input from the declarations, never relax a test to dodge a type. Links [[BDR-115]].
|
- **Apply**: write the boot helper first, type every input from the declarations, never relax a test to dodge a type. Links [[BDR-115]].
|
||||||
|
|
||||||
|
## LRN-207 — Model availability = typed API errors (`classic.StopFailure` rate_limit|overloaded|billing_error|model_not_found) + `PostModelSwitch auto`; never the turn-end reason, never `rateLimits`; sticky state and breaker target = two fields
|
||||||
|
- **Context**: model-router W1-C. `turn.complete reason: error` covers context-limit and network errors → a breaker fed by it marked fable down 15 min on a context overflow (challenge r1). `rateLimits` kinds are account windows (five_hour, seven_day, spend_limit), not per model. A per-step "engine fallback detection" (`e.model` ≠ `$.session.model()`) re-marked the session model at EVERY step → backoff climbed to the 5 h cap in one turn (confirmation r2). One field used both as sticky cur and as breaker target contradicted itself (reset at turn end vs kept).
|
||||||
|
- **Apply**: feed a breaker only from typed error kinds that name unavailability; mark per EPISODE (idempotent while down), backoff 15→300 min, `model_not_found` until reload; a user `/model` clears, `/clear` keeps (account-wide). Keep `turnModel` (sticky, reset per turn) and `lastPlan` (breaker target, kept) separate; unrouted steps pass `e.model` verbatim so the engine's own fallback is respected. Links [[BDR-115]], [[LRN-204]].
|
||||||
|
|
||||||
|
## LRN-208 — Security-auditor brief = scope only; iteration history and claimed closures violate the blind-scan contract
|
||||||
|
- **Context**: W1-C gate 2026-10-09. I sent "previous PASS with 6 MEDIUM, closed: …" in the brief. Auditor: "The auditor contract says that is never sent and must be ignored … Please do not send it next time" (it scanned blind anyway). Same rule as the verifier (lib/verify-secure-loop.md: fresh, no history).
|
||||||
|
- **Apply**: SCOPE + a neutral CONTEXT of what the code does; never prior verdicts, closures or accepted residuals. Residuals live in TODO, the auditor rediscovers them (that is the point). Links [[LRN-083]].
|
||||||
|
|
||||||
|
## LRN-209 — Scope the test run to the diff; one full `make test` before the merge; `make test` now names the red suite; GNU make rc 2 on a failed recipe
|
||||||
|
- **Context**: 2026-10-09 the pre-merge check took 454 s = full `make test` (48 suites) + a per-suite re-run to NAME the red one (aggregate rc, permanent env red design-tool-gate). Three more full passes earlier that day for mods/-only diffs. Hotfix efdd491: the recipe prints `FAIL <suite>` + `all suites green` / `<n> suite(s) red: …`. My oracle expected rc 1; GNU make returns 2 when a recipe line fails.
|
||||||
|
- **Apply**: a diff confined to one component runs that component's suite + the doctrine census; the full suite runs ONCE before `gitflow finish`; read the FAIL lines, never re-run per suite. Oracles on `make` test `[ $rc -ne 0 ]`, not `-eq 1`. Links [[LRN-173]].
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,24 @@
|
|||||||
|
# CONTRACT — make-test-names-red-suites
|
||||||
|
- date: 2026-10-09 | flow: hotfix | branch: bugfix/make-test-names-red-suites
|
||||||
|
- status: active
|
||||||
|
|
||||||
|
## REQUEST (verbatim — IMMUTABLE)
|
||||||
|
Skill args: "Makefile `test` target: print `FAIL <suite>` for every red suite and a final summary line (`<n> suite(s) red: <names>` or `all suites green`) so a full `make test` names the failing suites itself; exit code unchanged (1 on any red). Today only `== <suite>` headers print and the aggregate rc forces a second per-suite run to find the red one."
|
||||||
|
User (fr): "c'est quand meme long 9 min pour faire un merge non ? … ou est le bottlneck ?" → measured: the pre-merge `make test` (454 s) was a full run PLUS a per-suite re-run to name the red suite; "oui vas y fait le maintenant".
|
||||||
|
|
||||||
|
## CLARIFICATIONS
|
||||||
|
Q: wording / A: given by the request: `FAIL <suite>` per red suite right after it runs, then one summary line `<n> suite(s) red: <names>` or `all suites green`. [user]
|
||||||
|
Q: exit code / A: unchanged: 1 when any suite is red, 0 otherwise. [user]
|
||||||
|
|
||||||
|
## ACCEPTANCE CRITERIA
|
||||||
|
1. Symptom gone: a run with one red suite prints `FAIL <that suite>` and `1 suite(s) red: <that suite>` and exits non-zero (GNU make reports a failed recipe as 2); a run with only green suites prints `all suites green` and exits 0. Checked on a two-suite fixture through `make test suite="<green> <red>"`-style invocations (the `suite` variable already accepts a list).
|
||||||
|
CHECK: cd /Users/b.chanot/Documents/claude && W=$(mktemp -d) && printf '#!/usr/bin/env bash\nexit 0\n' > "$W/green.test.sh" && printf '#!/usr/bin/env bash\nexit 1\n' > "$W/red.test.sh" && out=$(make test suite="$W/green.test.sh $W/red.test.sh" 2>&1); rc=$?; [ $rc -ne 0 ] && echo "$out" | grep -q "^FAIL $W/red.test.sh" && echo "$out" | grep -q "1 suite(s) red: $W/red.test.sh" && out2=$(make test suite="$W/green.test.sh" 2>&1); rc2=$?; [ $rc2 -eq 0 ] && echo "$out2" | grep -q "all suites green" && echo SUMMARY-OK
|
||||||
|
EXPECT: SUMMARY-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: SUMMARY-OK
|
||||||
|
2. Build/tests green: the Makefile still runs the real suites (`make test suite=lib/tests/mods.test.sh` exits 0 and prints `all suites green`); `make -n test` parses.
|
||||||
|
CHECK: cd /Users/b.chanot/Documents/claude && make -n test >/dev/null && out=$(make test suite=lib/tests/mods.test.sh 2>&1); rc=$?; [ $rc -eq 0 ] && echo "$out" | grep -q "all suites green" && echo REAL-SUITE-OK
|
||||||
|
EXPECT: REAL-SUITE-OK
|
||||||
|
EVIDENCE: MET exit=0 marker-found :: REAL-SUITE-OK
|
||||||
|
|
||||||
|
## FILE SCOPE
|
||||||
|
Makefile
|
||||||
@@ -34,12 +34,15 @@ test: ## Run deterministic tests hermetically (one: make test suite=lib/tests/x.
|
|||||||
@# fire inside the throwaway repos the suites build. The export lives
|
@# fire inside the throwaway repos the suites build. The export lives
|
||||||
@# HERE so nobody has to type the (denied) env-prefix form by hand.
|
@# HERE so nobody has to type the (denied) env-prefix form by hand.
|
||||||
@export GIT_CONFIG_GLOBAL=/dev/null GIT_CONFIG_SYSTEM=/dev/null; \
|
@export GIT_CONFIG_GLOBAL=/dev/null GIT_CONFIG_SYSTEM=/dev/null; \
|
||||||
fail=0; for t in $(or $(suite),$(SUITES)); do \
|
fail=0; red=""; for t in $(or $(suite),$(SUITES)); do \
|
||||||
echo "== $$t"; \
|
echo "== $$t"; \
|
||||||
case "$$(basename "$$t")" in \
|
case "$$(basename "$$t")" in \
|
||||||
run-release-candidate.sh) RC_WORK=$$(mktemp -d) RC_TAG=1 bash "$$t" || fail=1 ;; \
|
run-release-candidate.sh) RC_WORK=$$(mktemp -d) RC_TAG=1 bash "$$t" || { fail=1; red="$$red $$t"; echo "FAIL $$t"; } ;; \
|
||||||
*) bash "$$t" || fail=1 ;; \
|
*) bash "$$t" || { fail=1; red="$$red $$t"; echo "FAIL $$t"; } ;; \
|
||||||
esac; done; exit $$fail
|
esac; done; \
|
||||||
|
if [ $$fail -eq 0 ]; then echo "all suites green"; \
|
||||||
|
else echo "$$(echo $$red | wc -w | tr -d ' ') suite(s) red:$$red"; fi; \
|
||||||
|
exit $$fail
|
||||||
|
|
||||||
scan-secrets: ## Gitleaks sweep: this repo's history + ~/.claude. Extra repos: make scan-secrets repos="path1 path2"
|
scan-secrets: ## Gitleaks sweep: this repo's history + ~/.claude. Extra repos: make scan-secrets repos="path1 path2"
|
||||||
@command -v gitleaks >/dev/null 2>&1 || { echo "gitleaks not installed — https://github.com/gitleaks/gitleaks"; exit 1; }
|
@command -v gitleaks >/dev/null 2>&1 || { echo "gitleaks not installed — https://github.com/gitleaks/gitleaks"; exit 1; }
|
||||||
|
|||||||
Reference in New Issue
Block a user