forked from bchanot/claude
chore(memory): case 7 registries, contracts, CHANGELOG — BDR-104 LRN-174 EVAL-033
This commit is contained in:
@@ -125,6 +125,7 @@ rules:
|
||||
| BDR-101 | 2026-09-25 | `full` = default profile: no selection ⇒ full in force, `reset` applies it, install applies it | accepted |
|
||||
| BDR-102 | 2026-09-27 | agent-skills: no plugin, vendor 3 skills + build floor-guard + routing census + rest-api rule | accepted |
|
||||
| BDR-103 | 2026-09-27 | 6-repo review: 5 verdicts, 3 criteria (grep-verified coverage, per-session cost, doctrine conflict); stars decided nothing | accepted |
|
||||
| BDR-104 | 2026-09-28 | MengTo motion pack: vendor 5 scroll skills pinned via shared lib/vendor-skills.sh + build personal skill site-motion; 17 skipped | accepted |
|
||||
|
||||
---
|
||||
|
||||
@@ -1297,3 +1298,12 @@ Branch feature/user-writing-web-rules, UNMERGED (human gate).
|
||||
- **Alternatives rejected**: install-then-prune (sunk cost, [[BDR-047]] ECC lesson); one bulk verdict (user wanted one case per turn, each with a build-vs-install call); building reticle's engine (1 286 server files).
|
||||
- **Caveats**: borrowed prompts change upstream with no diff, the pin is the review point; do not re-audit these six expecting more; mengto motion pack from the ui-skills registry is the open follow-up if site-level choreography (GSAP/ScrollTrigger, WebGL hero, masked reveals) proves thin locally.
|
||||
- **Reference**: branches feature/yagni-ladder (9315c6c, de7371d), feature/agent-skills-borrow (d28c45e 2b25cb4 409db51 1a8e6de 7401383 6617889 d71f3a7), feature/web-building-microrules (a2e654d 5a27372), chore/six-repo-review-notes (da35cde 197225a). Links [[BDR-102]] [[LRN-172]] [[LRN-173]] [[EVAL-032]] [[BDR-047]] [[BDR-006]].
|
||||
|
||||
## BDR-104 — MengTo motion pack: vendor 5 scroll skills via a shared helper + build `site-motion`; 17 skipped
|
||||
- **Date**: 2026-09-28
|
||||
- **Status**: accepted, feature/mengto-site-motion, UNMERGED (human gate)
|
||||
- **Decision**: (1) `lib/vendor-skills.sh` = one `vendor_pinned_skills <lock-key> [refresh]` for every curl-vendored upstream (agent-skills moved onto it; lock `skills` as list or dict of file lists; python argv lock read; `..`/charset guard; `VENDOR_BASE_URL` only as `file://`; tmp+mv; refresh skips never-installed skills, update-all convention). (2) Vendored at a965851: scroll-world-storytelling, build-threejs-scroll-worlds (+5 refs incl. scroll-conductor.js), scroll-scrubbed-visual-sequence, scroll-scrubbed-word-reveal, scroll-progress-timeline; text only, never demo/agents/binaries; design/web/web-full/full profiles. (3) `skills/site-motion` personal: invariants of the other 17 (gates, engine choice, Lenis sync, Astro ClientRouter lifecycle, numbered recipes, pitfalls); routed in CLAUDE.global.md Build UI chain + design-gate (not GATE-BLOCK).
|
||||
- **Why**: user asked whether the UI profiles carry motion knowledge for lively sites. Census: component polish deep (emil 27 KB, motion cookbook, impeccable animate), site choreography thin as workflows; ui-ux-pro-max CSVs cover terms but as search rows ([[LRN-174]]). Two analyzers read 22 skills: 5 clean workflows/recipes absent locally; 17 covered, generic, buggy (reduced-motion `clearProps`, gate before `registerPlugin`, no-JS opacity 0, refresh-rate `phi`) or Codex/Xcode machinery → distil per [[LRN-141]]. Registry sample (11) ≠ catalog (88): always list the tree.
|
||||
- **Alternatives rejected**: vendor all 22 (bugs + duplicates + 17×~100 tok descriptions); distil only (loses the two deep workflows whose value is their full text); second inline curl loop (duplication; helper instead); `VENDOR_BASE_URL` gated by a companion var (file:// prefix check suffices); grouped toggle pack (per-name like emil).
|
||||
- **Caveats**: vendored prompts change upstream with no diff, pin = review point; GSAP-first content vs [[BDR-005]] `motion` default, site-motion states the allowance; LOW hardening open: `re.match` `$` accepts a trailing newline, `source`/`path`/`sha` lock fields not charset-checked; `make link` + `bash lib/profile.sh apply full` after merge (user); sub-agent leftovers `/tmp/mengto-verify` (user removes).
|
||||
- **Reference**: commits 2a1ad17 (helper + vendoring), ba14b5e (site-motion); contracts `.claude/tasks/contracts/2026-09-27-{mengto-vendor,site-motion-skill}-0002.md`; gates MET, verifiers CONFORME after 3 re-dispatches (frontmatter shape, refresh convention, security env override + traversal), security PASS ×2, `make test` 36 suites green minus 2 pre-existing T16a. Links [[BDR-103]] [[BDR-102]] [[LRN-141]] [[LRN-174]] [[EVAL-033]].
|
||||
|
||||
@@ -53,6 +53,7 @@ rules:
|
||||
| EVAL-030 | 2026-09-24 | 2026-09-24 self-audit: two regressions and one guardrail bypass came from my own process, not from the tools | BDR-100 mechanisms shipped; re-run census at next doctrine wave |
|
||||
| EVAL-031 | 2026-09-25 | /feat run for BDR-101: challenge round earned its cost, two blockers sat in my own premises | keep challenge round on state-detection plans; check live state before planning; pin grep in oracles |
|
||||
| EVAL-032 | 2026-09-27 | 4 parallel feater executors, one tree, gate loop: verifier caught a vacuous test, security caught a partial-write; my oracles wrong twice | keep same-tree parallel dispatch with disjoint FILE SCOPE + orchestrator-owned shared files; blind verifier stays; measure oracles on precedents |
|
||||
| EVAL-033 | 2026-09-28 | case 7: 2 analyzers + 2 executors + 3 re-dispatches; verifiers caught shape, convention and my wrong count; security caught an env override | brief names the scratchpad path explicitly (3 /tmp leftovers); keep blind verifiers; count claims get an artifact |
|
||||
|
||||
---
|
||||
|
||||
@@ -314,3 +315,10 @@ Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itse
|
||||
- **Result**: 4/4 CONFORME after 2 re-dispatches; security PASS ×2. Verifier A3 caught a vacuous distinct-pair test the executor had self-justified ([[LRN-172]]). Security caught first-download without tmp+mv (partial file accepted forever) + python source splicing → fixed by fresh executor, re-verified, re-audited. Executors never touched each other's files; A4 verifier counted the orchestrator-reserved CHANGELOG as ECARTS(1), correct by contract wording.
|
||||
- **Anomalies**: (1) my oracles wrong twice ([[LRN-173]]); (2) gates.sh ERROR(3) on first run, `EVIDENCE: pending` missing; (3) 3 guardrail denials on sub-agents (`export GIT_CONFIG_GLOBAL` inline ×2 incl. a verifier, `rm -rf /tmp/tmp.AAyJzvufO6` executor cleanup), all reported, none evaded — [[BDR-100]] live; (4) security-auditor miscounted the sha as 41 chars (it is 40) — verify sub-agent claims before acting; (5) `make test` rc 1 from the 2 pre-existing T16a, my first grep filter hid the totals.
|
||||
- **Action**: keep the pattern; add `EVIDENCE: pending` to the contract skeleton; floor-guard waiver policy → user decision; snippet framing for LLM-consumed output → follow-up.
|
||||
|
||||
## EVAL-033 — case 7 execution: analyzers first, then two executors, three re-dispatches
|
||||
- **Date**: 2026-09-28
|
||||
- **Method**: two read-only analyzers (22 skills, fixed per-skill format, grep overlap against local assets) → verdict → two contracts → two feater executors in parallel (disjoint scopes, profiles owned by one, CHANGELOG by me) → gates.sh → fresh verifiers → security ×2 → full `make test`.
|
||||
- **Result**: both CONFORME after 3 re-dispatches: frontmatter shape (executor mirrored external peers instead of the named personal ones), update-all refresh convention (executor's "additive" install broke "not installed — skipping"), security MEDIUM env override + LOW traversal. Verifier also doubted my CHANGELOG "sixteen skipped" → it was seventeen. Byte-for-byte fidelity of 14 files confirmed twice.
|
||||
- **Anomalies**: (1) three sub-agents wrote to `/tmp` outside the scratchpad then could not `rm -rf` (refused, correctly) — the brief must name the scratchpad path; (2) executors' self-justified deviations were plausible each time and wrong twice → blind verifier stays mandatory; (3) analyzer reports at ~120-180 words per skill were the right grain, two of them fit my context; (4) `make test` rc 1 is still the 2 pre-existing T16a, my filter now shows totals.
|
||||
- **Action**: brief template line "scratch only under <scratchpad>"; count claims in CHANGELOG/journal cite the list they count; keep the analyzer-first pattern for any pack > 5 skills.
|
||||
|
||||
@@ -528,3 +528,6 @@ rules:
|
||||
- User go "merge le tout": the four review branches merged into develop via `gitflow finish` → b3597eb (yagni-ladder), 04cb057 (agent-skills-borrow), 68fcdaf (web-building-microrules), 39d5b15 (six-repo-review-notes); 7 registry conflicts (TODO ×3, journal ×3, CHANGELOG, decisions ×2) resolved by a scratch resolver keeping both sides in order (TODO/CHANGELOG incoming first, registries HEAD first), merge commits by hand, finish re-run removed local + origin copies. develop == origin/develop, no review branch left. Post-merge: 0 conflict markers, BDR-101→103 in order, `make test` 35 suites green minus 2 pre-existing T16a, shellcheck clean. BDR-103 written on user go. Open for the user: `make link` + `bash lib/profile.sh apply full` (trio symlinks), `rm -rf /tmp/tmp.AAyJzvufO6`.
|
||||
- User asked whether the UI profiles already carry motion-design knowledge for lively modern sites. Census answer: micro-interaction + component polish deep (emil 27 KB, motion cookbook 14 sections incl. scroll-driven, impeccable animate + detect); site-level choreography thin (GSAP/ScrollTrigger storytelling, Lenis, WebGL hero, masked reveals, marquee as workflows) and Astro View Transitions at zero mentions despite Astro-first. Candidate: mengto motion pack from the ui-skills registry, same three criteria; user decides.
|
||||
- Correction to the motion census above: ui-ux-pro-max's data CSVs (motion.csv 17 rows: GSAP reveal/pin/scrub, SplitText, parallax, magnetic; stacks/threejs.csv 53 rows; stacks/astro.csv rows 28-31 ViewTransitions: ClientRouter, `transition:name`, no-JS fallback; landing.csv scrollytelling) cover what I called absent. My grep skipped the plugin's data files. They are search-DB rows reached through the skill's search tool, not build workflows. Case 7 (mengto pack, 22 skills read by two analyzers): verdict pending user decision.
|
||||
|
||||
## 2026-09-28
|
||||
- Case 7 (MengTo motion pack), user go "l'hybride" → [[BDR-104]]: `lib/vendor-skills.sh` shared helper (agent-skills moved onto it), 5 scroll skills vendored at a965851 (2a1ad17), `skills/site-motion` personal skill + routing (ba14b5e). Two analyzers read 22 skills first; 17 skipped (bugs, duplicates, covered, Codex/Xcode machinery). Gates MET, verifiers CONFORME after 3 re-dispatches (frontmatter shape, update-all refresh convention, security env override + traversal), security PASS ×2, `make test` 36 suites green minus 2 pre-existing T16a, shellcheck clean. [[LRN-174]] [[EVAL-033]]. feature/mengto-site-motion UNMERGED. Open for the user: `make link` + `bash lib/profile.sh apply full`, `rm -rf /tmp/mengto-verify`, LOW hardening (regex trailing newline, source/path/sha charset).
|
||||
|
||||
@@ -193,6 +193,7 @@ rules:
|
||||
| LRN-171 | 2026-09-25 | sub-agent sandbox: grep shim returns EMPTY inside `$(...)` for patterns holding literal `$VAR` — oracles pin `command grep` | contract CHECK lines, hermetic test greps, hooks parsing grep output |
|
||||
| LRN-172 | 2026-09-27 | TF-IDF cosine on a 2-doc corpus is identically 0: similarity self-tests need N ≥ 4, a same-corpus positive control and a sensitivity re-run | fixtures for any corpus-normalised statistic (idf, z-score, ranking), "distinct pair passes" tests |
|
||||
| LRN-173 | 2026-09-27 | contract oracles written from memory failed twice: run the CHECK on the precedent files first, census greps via `git grep` (tracked only), `EVIDENCE: pending` mandatory for gates.sh | contract CHECK lines, precedent-mirroring criteria, gates.sh ledgers |
|
||||
| LRN-174 | 2026-09-28 | a coverage census must grep plugin DATA files (CSV/JSON search DBs), not only SKILL.md prose; and a registry sample is not the upstream catalog, list the tree | before claiming a gap in installed skills; before scoping an external-repo evaluation |
|
||||
|
||||
---
|
||||
|
||||
@@ -1615,3 +1616,7 @@ Rule: when editing a doctrine file under structure locks, grep the test's lock s
|
||||
## LRN-173 — contract oracles written from memory failed twice: run the CHECK on the precedent files first, census greps via `git grep`, `EVIDENCE: pending` mandatory
|
||||
- **Context**: same run, two orchestrator oracle bugs. (1) "≤ 80 chars" over the whole file: rules/web-building.md line 2 (`paths:` frontmatter) is already 110 chars, so the criterion contradicted the precedent it named; executor returned NEED-DECISION instead of bending. (2) emil-citers census with `grep -rl` hit gitignored `install-*.log` at the repo root; `git grep -l` (tracked only) is the right census tool. (3) gates.sh `run` errors `runnable but has no EVIDENCE: line` unless each criterion carries `EVIDENCE: pending`.
|
||||
- **Apply**: before shipping a CHECK, run it against the files it claims to mirror; census oracles = `git grep`; contract skeleton carries `EVIDENCE: pending` per criterion (check /feat's template writes it). Oracle fixes are orchestrator-owned, never a re-dispatch ([[BDR-102]]).
|
||||
|
||||
## LRN-174 — a coverage census greps plugin data files too; a registry sample is not the catalog
|
||||
- **Context**: I told the user Astro View Transitions had zero local mentions. ui-ux-pro-max's `data/stacks/astro.csv` rows 28-31 carry ClientRouter, `transition:name`, no-JS fallback; `motion.csv` carries GSAP pin/scrub, SplitText, parallax. My grep covered SKILL.md prose and archetypes, not the plugin's CSV search DB. Same day: the ui-skills registry showed 11 MengTo skills; the repo tree has 88 web-design skills, and the substantive ones were outside the sample.
|
||||
- **Apply**: census = `grep -rl` over the plugin cache including data dirs, then say "row in a search DB" vs "workflow"; evaluating an upstream = `git/trees?recursive=1` first, sample never. Correct the user the moment the miss is found ([[BDR-104]]).
|
||||
|
||||
Reference in New Issue
Block a user