chore(memory): case 2 registries, contracts, CHANGELOG — BDR-102 LRN-172 LRN-173 EVAL-032

This commit is contained in:
bastien
2026-09-27 20:18:41 +02:00
parent 1a8e6decdb
commit 740138337c
10 changed files with 273 additions and 0 deletions
+10
View File
@@ -123,6 +123,7 @@ rules:
| BDR-099 | 2026-09-24 | C2 coherence: 30 doctrine/skill tensions resolved, doctrine wins, BDR-068 kept as the written exception | accepted |
| BDR-100 | 2026-09-24 | Guardrail evasion and partial rule changes get mechanisms, not lessons: refusal ends the attempt, citers census in make test | accepted |
| BDR-101 | 2026-09-25 | `full` = default profile: no selection ⇒ full in force, `reset` applies it, install applies it | accepted |
| BDR-102 | 2026-09-27 | agent-skills: no plugin, vendor 3 skills + build floor-guard + routing census + rest-api rule | accepted |
---
@@ -1270,3 +1271,12 @@ Branch feature/user-writing-web-rules, UNMERGED (human gate).
- **Alternatives rejected**: label-only reset (lies about plugins/externals); additive reset (`gstack on` + `apply full`, state ⊇ full, label ambiguous); keep `none` sentinel + statusline `?` (request unmet); install display-only (fresh machine ≠ full, 21st pack parked); public `profile.sh default` verb for installer (reset already IS "go to default"); one shared cache parser (statusline must not spawn profile.sh → 3 copies kept, each commented).
- **Caveats**: `make plugin` re-run re-applies selected profile → manual layering (`gstack on` over `dev`) trimmed back. Plugin legs of Step 11 install-immutable ([[BDR-028]] EXIT guard; committed enabledPlugins already match full). LOW security note: cache content not charset-checked before path use (pre-existing in `read_profile`) → follow-up.
- **Reference**: commits e196328 (residue scrub), 0d035fc (profile), 1bbdad0 (install); contract/plan `2026-09-25-default-profile-full-1254`; `lib/tests/profile-default.test.sh` 29 checks. Links [[BDR-017]] [[BDR-018]] [[BDR-079]] [[BDR-093]] [[LRN-170]] [[EVAL-031]].
## BDR-102 — agent-skills: no plugin, vendor 3 skills + build floor-guard + routing census + rest-api rule
- **Date**: 2026-09-27
- **Status**: accepted, feature/agent-skills-borrow, UNMERGED (human gate)
- **Decision**: addyosmani/agent-skills (99.4k stars, 25 skills) NOT installed as plugin. Borrowed 4 things, user go: (1) `observability-and-instrumentation`, `deprecation-and-migration`, `ci-cd-and-automation` vendored emil-way at pinned commit 2686b620 (lock `agent-skills`, install Step 8e tmp+mv, update-all 7.3, link.sh, toggle-external per-name, profiles full/backend/dev); (2) `lib/floor-guard.sh` diff-scoped bar-weakening detector, verifier STEP 3 mandatory, waiver `floor-guard: allow <reason>`; (3) `lib/tests/skill-routing-census.test.sh` TF-IDF description-collision census, WARN 0.50 / FAIL 0.75; (4) `rules/rest-api.md` path-scoped, distilled from api-and-interface-design minus the one-version rule.
- **Why**: 20/25 skills already covered (superpowers, personal skills, gstack, built-ins). Real gaps grep-verified: observability (archetype question only), migration (one strangler line), CI build (ship/cso only detect), bar-weakening guard (security-auditor flags nosemgrep only), collision census (routing section exists because of collisions; darwin scores quality not collisions). Upstream comparison doc: never stack two skill routers.
- **Alternatives rejected**: plugin install (1.8k tok/session for 20 % novelty, `/spec` `/review` `/ship` collide with gstack, `/code-simplify` vs `/simplify`, trunk-based git + one-version API contradict doctrine, second router); vendor api-and-interface-design whole (one-version rule vs § Web APIs → distilled); grouped `agent-skills` toggle pack (no shared installer, per-name like emil); web-full profile for the trio (design-class, allowlist wins over "lists bugfix"); Tier 2 prompt ranking (follow-up).
- **Caveats**: vendored prompts = third-party content loaded into sessions, the pin is the review point (security-auditor scanned: benign, no hidden Unicode); floor-guard waiver is self-service, WAIVED informational → security MEDIUM, design decision pending with user (require CLARIFICATIONS ack outside test fixtures?); floor-guard prints raw diff snippets a verifier reads (LOW, framing follow-up); update-all 7.3 failure branch leaves `.tmp` like emil (parity, not fixed); `make link` after merge to symlink the trio (user).
- **Reference**: commits d28c45e (trio), 2b25cb4 (floor-guard), 409db51 (census), 1a8e6de (rest-api); contracts `.claude/tasks/contracts/2026-09-27-{agent-skills-vendor,floor-guard,skill-routing-census,rest-api-rule}-1525.md`; gates MET ×4, verifiers CONFORME ×4 (2 re-dispatches), security PASS ×2. Links [[BDR-100]] [[LRN-172]] [[LRN-173]] [[EVAL-032]]. Case 1 of the same review: ladder in doctrine, feature/yagni-ladder 9315c6c.
+8
View File
@@ -52,6 +52,7 @@ rules:
| EVAL-029 | 2026-09-15 | 4-agent plan challenge: 6 BLOCKER; 3 of 3 confirmation-pass BLOCKERs came from the fixes themselves; caught a false 654 MB orphan claim | keep |
| EVAL-030 | 2026-09-24 | 2026-09-24 self-audit: two regressions and one guardrail bypass came from my own process, not from the tools | BDR-100 mechanisms shipped; re-run census at next doctrine wave |
| EVAL-031 | 2026-09-25 | /feat run for BDR-101: challenge round earned its cost, two blockers sat in my own premises | keep challenge round on state-detection plans; check live state before planning; pin grep in oracles |
| EVAL-032 | 2026-09-27 | 4 parallel feater executors, one tree, gate loop: verifier caught a vacuous test, security caught a partial-write; my oracles wrong twice | keep same-tree parallel dispatch with disjoint FILE SCOPE + orchestrator-owned shared files; blind verifier stays; measure oracles on precedents |
---
@@ -306,3 +307,10 @@ Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itse
- **Method**: challenge lib (3 lenses + 1 confirmation), gates.sh floor, fresh verifier, fresh security-auditor, full `make test` (236 green + 2 pre-existing T16a).
- **Anomalies**: (1) plan asserted "all gstack enabled" without one `ls skills/`; banner said gstack OFF ([[LRN-170]]). (2) two sub-agents hit same grep-shim quirk ([[LRN-171]]). (3) `git add .env.example` denied (`git add .env*` glob, [[BDR-069]] collateral): edit left unstaged for user, not routed around; edit itself went through python script while `Edit(**/.env.*)` denied — surfaced to user. (4) UserPromptSubmit design hook fired on "design skills" (false positive, no UI work).
- **Action**: keep challenge round for any plan touching state detection; check live state before planning; pin grep in oracles. Links [[BDR-101]].
## EVAL-032 — 4 parallel feater executors, one tree, gate loop (case 2 of the 6-repo review)
- **Date**: 2026-09-27
- **Method**: 4 contracts, 4 feater executors dispatched in one turn on the same working tree (disjoint FILE SCOPE, CHANGELOG reserved to the orchestrator), gates.sh run per contract, fresh verifier per contract, security-auditor on the whole diff then on the re-touched files, full `make test`.
- **Result**: 4/4 CONFORME after 2 re-dispatches; security PASS ×2. Verifier A3 caught a vacuous distinct-pair test the executor had self-justified ([[LRN-172]]). Security caught first-download without tmp+mv (partial file accepted forever) + python source splicing → fixed by fresh executor, re-verified, re-audited. Executors never touched each other's files; A4 verifier counted the orchestrator-reserved CHANGELOG as ECARTS(1), correct by contract wording.
- **Anomalies**: (1) my oracles wrong twice ([[LRN-173]]); (2) gates.sh ERROR(3) on first run, `EVIDENCE: pending` missing; (3) 3 guardrail denials on sub-agents (`export GIT_CONFIG_GLOBAL` inline ×2 incl. a verifier, `rm -rf /tmp/tmp.AAyJzvufO6` executor cleanup), all reported, none evaded — [[BDR-100]] live; (4) security-auditor miscounted the sha as 41 chars (it is 40) — verify sub-agent claims before acting; (5) `make test` rc 1 from the 2 pre-existing T16a, my first grep filter hid the totals.
- **Action**: keep the pattern; add `EVIDENCE: pending` to the contract skeleton; floor-guard waiver policy → user decision; snippet framing for LLM-consumed output → follow-up.
+2
View File
@@ -518,3 +518,5 @@ rules:
## 2026-09-27
- bugfix/gitignore-diagram-allowlist merged into develop on user go, `gitflow finish` → facd26d, pushed, copies removed by the lib. Day 2026-09-25 lot fully on develop: BDR-101 default profile, full +4 gstack skills, gitignore allowlist. No working branch anywhere; `skills/diagram` ignored.
- 6-repo review case 2 (agent-skills 99.4k stars): plugin rejected (1.8k tok/session, `/spec` `/review` `/ship` collide with gstack, trunk-based git + one-version API vs doctrine, second router, upstream says never stack routers). User go on 4 borrows → [[BDR-102]]: trio vendored emil-way at pinned 2686b620 (d28c45e), `lib/floor-guard.sh` + verifier STEP 3 (2b25cb4), `lib/tests/skill-routing-census.test.sh` 120 skills max 0.52 (409db51), `rules/rest-api.md` (1a8e6de). 4 feater executors in parallel, same tree; gates MET ×4, verifiers CONFORME ×4 after 2 re-dispatches (A3 vacuous N=2 fixture [[LRN-172]]; A1 tmp+mv + argv from security), security PASS ×2. My oracles wrong twice ([[LRN-173]]), [[EVAL-032]]. `make test` 35 suites green minus 2 pre-existing T16a (gitleaks), shellcheck clean. feature/agent-skills-borrow UNMERGED. Guardrails fired 3× on sub-agents, none evaded; `/tmp/tmp.AAyJzvufO6` scratch dir left for the user (rm -rf refused, correctly).
- Case 3 (ui-skills 9.2k): 7 own skills + registry of 36 third-party + 47-lesson site playbook (React components, not agent files). Verdict given: extend rules/web-building.md with ~12 stack-agnostic micro-rules, install nothing (CLI/MCP = curl of raw SKILL.md, third router, baseline-ui stack mandates vs Astro doctrine). Awaiting user; case 4 reticle material prefetched.
+11
View File
@@ -191,6 +191,8 @@ rules:
| LRN-169 | 2026-09-24 | a coherence audit is cheap when parallel and read-only, and its findings are claims: spot-check, then fix every citer | any doctrine or skill rule change; sub-agent briefs; environment-dependent tests |
| LRN-170 | 2026-09-25 | "count == 0" ≠ "all on" when default state is "nothing installed": verify a fast-path premise on the live tree | status/current/detect commands, installer "is X applied?" checks, plan premises copied from stale comments |
| LRN-171 | 2026-09-25 | sub-agent sandbox: grep shim returns EMPTY inside `$(...)` for patterns holding literal `$VAR` — oracles pin `command grep` | contract CHECK lines, hermetic test greps, hooks parsing grep output |
| LRN-172 | 2026-09-27 | TF-IDF cosine on a 2-doc corpus is identically 0: similarity self-tests need N ≥ 4, a same-corpus positive control and a sensitivity re-run | fixtures for any corpus-normalised statistic (idf, z-score, ranking), "distinct pair passes" tests |
| LRN-173 | 2026-09-27 | contract oracles written from memory failed twice: run the CHECK on the precedent files first, census greps via `git grep` (tracked only), `EVIDENCE: pending` mandatory for gates.sh | contract CHECK lines, precedent-mirroring criteria, gates.sh ledgers |
---
@@ -1604,3 +1606,12 @@ Rule: when editing a doctrine file under structure locks, grep the test's lock s
- **Pattern**: oracles + test assertions portable across main shell / sub-agent shells pin `/usr/bin/grep` or `command grep`; avoid `$VAR` literals mid-pattern (`-F` or `--`).
- **Where applicable**: contract `CHECK:` lines run by executors/verifiers; hermetic test greps; any hook riding on grep output.
- **Reference**: contract `2026-09-25-default-profile-full-1254` criterion 12; verifier + executor reports 2026-09-25. Links [[LRN-074]] [[BDR-101]].
## LRN-172 — TF-IDF cosine on a 2-doc corpus is identically 0: similarity self-tests need N ≥ 4, a same-corpus positive control and a sensitivity re-run
- **Context**: A3 executor wrote a "distinct pair passes" fixture as a bare 2-doc corpus, documented the 0.00 as "the point being proven". Fresh verifier proved a near-duplicate pair also scores 0.00 at N=2: idf = log(N/df) = 0 for shared terms, unique terms never meet. Executor self-report "all markers printed" was true; the test was vacuous anyway.
- **Fix shape**: one N=4 corpus: near-dup control 0.90 → FAIL, distinct pair 0.00 → silent, then swap one doc for a near-copy → 0.62 WARN appears. Marker printed only when all three hold.
- **Apply**: any fixture for a corpus-normalised statistic asserts both directions in one corpus; "passes on a trivial corpus" proves nothing; markers prove the oracle, not the intent → keep the blind verifier ([[BDR-102]] [[EVAL-032]]).
## LRN-173 — contract oracles written from memory failed twice: run the CHECK on the precedent files first, census greps via `git grep`, `EVIDENCE: pending` mandatory
- **Context**: same run, two orchestrator oracle bugs. (1) "≤ 80 chars" over the whole file: rules/web-building.md line 2 (`paths:` frontmatter) is already 110 chars, so the criterion contradicted the precedent it named; executor returned NEED-DECISION instead of bending. (2) emil-citers census with `grep -rl` hit gitignored `install-*.log` at the repo root; `git grep -l` (tracked only) is the right census tool. (3) gates.sh `run` errors `runnable but has no EVIDENCE: line` unless each criterion carries `EVIDENCE: pending`.
- **Apply**: before shipping a CHECK, run it against the files it claims to mirror; census oracles = `git grep`; contract skeleton carries `EVIDENCE: pending` per criterion (check /feat's template writes it). Oracle fixes are orchestrator-owned, never a re-dispatch ([[BDR-102]]).