chore(memory): case 2 registries, contracts, CHANGELOG — BDR-102 LRN-172 LRN-173 EVAL-032
This commit is contained in:
@@ -123,6 +123,7 @@ rules:
|
||||
| BDR-099 | 2026-09-24 | C2 coherence: 30 doctrine/skill tensions resolved, doctrine wins, BDR-068 kept as the written exception | accepted |
|
||||
| BDR-100 | 2026-09-24 | Guardrail evasion and partial rule changes get mechanisms, not lessons: refusal ends the attempt, citers census in make test | accepted |
|
||||
| BDR-101 | 2026-09-25 | `full` = default profile: no selection ⇒ full in force, `reset` applies it, install applies it | accepted |
|
||||
| BDR-102 | 2026-09-27 | agent-skills: no plugin, vendor 3 skills + build floor-guard + routing census + rest-api rule | accepted |
|
||||
|
||||
---
|
||||
|
||||
@@ -1270,3 +1271,12 @@ Branch feature/user-writing-web-rules, UNMERGED (human gate).
|
||||
- **Alternatives rejected**: label-only reset (lies about plugins/externals); additive reset (`gstack on` + `apply full`, state ⊇ full, label ambiguous); keep `none` sentinel + statusline `?` (request unmet); install display-only (fresh machine ≠ full, 21st pack parked); public `profile.sh default` verb for installer (reset already IS "go to default"); one shared cache parser (statusline must not spawn profile.sh → 3 copies kept, each commented).
|
||||
- **Caveats**: `make plugin` re-run re-applies selected profile → manual layering (`gstack on` over `dev`) trimmed back. Plugin legs of Step 11 install-immutable ([[BDR-028]] EXIT guard; committed enabledPlugins already match full). LOW security note: cache content not charset-checked before path use (pre-existing in `read_profile`) → follow-up.
|
||||
- **Reference**: commits e196328 (residue scrub), 0d035fc (profile), 1bbdad0 (install); contract/plan `2026-09-25-default-profile-full-1254`; `lib/tests/profile-default.test.sh` 29 checks. Links [[BDR-017]] [[BDR-018]] [[BDR-079]] [[BDR-093]] [[LRN-170]] [[EVAL-031]].
|
||||
|
||||
## BDR-102 — agent-skills: no plugin, vendor 3 skills + build floor-guard + routing census + rest-api rule
|
||||
- **Date**: 2026-09-27
|
||||
- **Status**: accepted, feature/agent-skills-borrow, UNMERGED (human gate)
|
||||
- **Decision**: addyosmani/agent-skills (99.4k stars, 25 skills) NOT installed as plugin. Borrowed 4 things, user go: (1) `observability-and-instrumentation`, `deprecation-and-migration`, `ci-cd-and-automation` vendored emil-way at pinned commit 2686b620 (lock `agent-skills`, install Step 8e tmp+mv, update-all 7.3, link.sh, toggle-external per-name, profiles full/backend/dev); (2) `lib/floor-guard.sh` diff-scoped bar-weakening detector, verifier STEP 3 mandatory, waiver `floor-guard: allow <reason>`; (3) `lib/tests/skill-routing-census.test.sh` TF-IDF description-collision census, WARN 0.50 / FAIL 0.75; (4) `rules/rest-api.md` path-scoped, distilled from api-and-interface-design minus the one-version rule.
|
||||
- **Why**: 20/25 skills already covered (superpowers, personal skills, gstack, built-ins). Real gaps grep-verified: observability (archetype question only), migration (one strangler line), CI build (ship/cso only detect), bar-weakening guard (security-auditor flags nosemgrep only), collision census (routing section exists because of collisions; darwin scores quality not collisions). Upstream comparison doc: never stack two skill routers.
|
||||
- **Alternatives rejected**: plugin install (1.8k tok/session for 20 % novelty, `/spec` `/review` `/ship` collide with gstack, `/code-simplify` vs `/simplify`, trunk-based git + one-version API contradict doctrine, second router); vendor api-and-interface-design whole (one-version rule vs § Web APIs → distilled); grouped `agent-skills` toggle pack (no shared installer, per-name like emil); web-full profile for the trio (design-class, allowlist wins over "lists bugfix"); Tier 2 prompt ranking (follow-up).
|
||||
- **Caveats**: vendored prompts = third-party content loaded into sessions, the pin is the review point (security-auditor scanned: benign, no hidden Unicode); floor-guard waiver is self-service, WAIVED informational → security MEDIUM, design decision pending with user (require CLARIFICATIONS ack outside test fixtures?); floor-guard prints raw diff snippets a verifier reads (LOW, framing follow-up); update-all 7.3 failure branch leaves `.tmp` like emil (parity, not fixed); `make link` after merge to symlink the trio (user).
|
||||
- **Reference**: commits d28c45e (trio), 2b25cb4 (floor-guard), 409db51 (census), 1a8e6de (rest-api); contracts `.claude/tasks/contracts/2026-09-27-{agent-skills-vendor,floor-guard,skill-routing-census,rest-api-rule}-1525.md`; gates MET ×4, verifiers CONFORME ×4 (2 re-dispatches), security PASS ×2. Links [[BDR-100]] [[LRN-172]] [[LRN-173]] [[EVAL-032]]. Case 1 of the same review: ladder in doctrine, feature/yagni-ladder 9315c6c.
|
||||
|
||||
@@ -52,6 +52,7 @@ rules:
|
||||
| EVAL-029 | 2026-09-15 | 4-agent plan challenge: 6 BLOCKER; 3 of 3 confirmation-pass BLOCKERs came from the fixes themselves; caught a false 654 MB orphan claim | keep |
|
||||
| EVAL-030 | 2026-09-24 | 2026-09-24 self-audit: two regressions and one guardrail bypass came from my own process, not from the tools | BDR-100 mechanisms shipped; re-run census at next doctrine wave |
|
||||
| EVAL-031 | 2026-09-25 | /feat run for BDR-101: challenge round earned its cost, two blockers sat in my own premises | keep challenge round on state-detection plans; check live state before planning; pin grep in oracles |
|
||||
| EVAL-032 | 2026-09-27 | 4 parallel feater executors, one tree, gate loop: verifier caught a vacuous test, security caught a partial-write; my oracles wrong twice | keep same-tree parallel dispatch with disjoint FILE SCOPE + orchestrator-owned shared files; blind verifier stays; measure oracles on precedents |
|
||||
|
||||
---
|
||||
|
||||
@@ -306,3 +307,10 @@ Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itse
|
||||
- **Method**: challenge lib (3 lenses + 1 confirmation), gates.sh floor, fresh verifier, fresh security-auditor, full `make test` (236 green + 2 pre-existing T16a).
|
||||
- **Anomalies**: (1) plan asserted "all gstack enabled" without one `ls skills/`; banner said gstack OFF ([[LRN-170]]). (2) two sub-agents hit same grep-shim quirk ([[LRN-171]]). (3) `git add .env.example` denied (`git add .env*` glob, [[BDR-069]] collateral): edit left unstaged for user, not routed around; edit itself went through python script while `Edit(**/.env.*)` denied — surfaced to user. (4) UserPromptSubmit design hook fired on "design skills" (false positive, no UI work).
|
||||
- **Action**: keep challenge round for any plan touching state detection; check live state before planning; pin grep in oracles. Links [[BDR-101]].
|
||||
|
||||
## EVAL-032 — 4 parallel feater executors, one tree, gate loop (case 2 of the 6-repo review)
|
||||
- **Date**: 2026-09-27
|
||||
- **Method**: 4 contracts, 4 feater executors dispatched in one turn on the same working tree (disjoint FILE SCOPE, CHANGELOG reserved to the orchestrator), gates.sh run per contract, fresh verifier per contract, security-auditor on the whole diff then on the re-touched files, full `make test`.
|
||||
- **Result**: 4/4 CONFORME after 2 re-dispatches; security PASS ×2. Verifier A3 caught a vacuous distinct-pair test the executor had self-justified ([[LRN-172]]). Security caught first-download without tmp+mv (partial file accepted forever) + python source splicing → fixed by fresh executor, re-verified, re-audited. Executors never touched each other's files; A4 verifier counted the orchestrator-reserved CHANGELOG as ECARTS(1), correct by contract wording.
|
||||
- **Anomalies**: (1) my oracles wrong twice ([[LRN-173]]); (2) gates.sh ERROR(3) on first run, `EVIDENCE: pending` missing; (3) 3 guardrail denials on sub-agents (`export GIT_CONFIG_GLOBAL` inline ×2 incl. a verifier, `rm -rf /tmp/tmp.AAyJzvufO6` executor cleanup), all reported, none evaded — [[BDR-100]] live; (4) security-auditor miscounted the sha as 41 chars (it is 40) — verify sub-agent claims before acting; (5) `make test` rc 1 from the 2 pre-existing T16a, my first grep filter hid the totals.
|
||||
- **Action**: keep the pattern; add `EVIDENCE: pending` to the contract skeleton; floor-guard waiver policy → user decision; snippet framing for LLM-consumed output → follow-up.
|
||||
|
||||
@@ -518,3 +518,5 @@ rules:
|
||||
|
||||
## 2026-09-27
|
||||
- bugfix/gitignore-diagram-allowlist merged into develop on user go, `gitflow finish` → facd26d, pushed, copies removed by the lib. Day 2026-09-25 lot fully on develop: BDR-101 default profile, full +4 gstack skills, gitignore allowlist. No working branch anywhere; `skills/diagram` ignored.
|
||||
- 6-repo review case 2 (agent-skills 99.4k stars): plugin rejected (1.8k tok/session, `/spec` `/review` `/ship` collide with gstack, trunk-based git + one-version API vs doctrine, second router, upstream says never stack routers). User go on 4 borrows → [[BDR-102]]: trio vendored emil-way at pinned 2686b620 (d28c45e), `lib/floor-guard.sh` + verifier STEP 3 (2b25cb4), `lib/tests/skill-routing-census.test.sh` 120 skills max 0.52 (409db51), `rules/rest-api.md` (1a8e6de). 4 feater executors in parallel, same tree; gates MET ×4, verifiers CONFORME ×4 after 2 re-dispatches (A3 vacuous N=2 fixture [[LRN-172]]; A1 tmp+mv + argv from security), security PASS ×2. My oracles wrong twice ([[LRN-173]]), [[EVAL-032]]. `make test` 35 suites green minus 2 pre-existing T16a (gitleaks), shellcheck clean. feature/agent-skills-borrow UNMERGED. Guardrails fired 3× on sub-agents, none evaded; `/tmp/tmp.AAyJzvufO6` scratch dir left for the user (rm -rf refused, correctly).
|
||||
- Case 3 (ui-skills 9.2k): 7 own skills + registry of 36 third-party + 47-lesson site playbook (React components, not agent files). Verdict given: extend rules/web-building.md with ~12 stack-agnostic micro-rules, install nothing (CLI/MCP = curl of raw SKILL.md, third router, baseline-ui stack mandates vs Astro doctrine). Awaiting user; case 4 reticle material prefetched.
|
||||
|
||||
@@ -191,6 +191,8 @@ rules:
|
||||
| LRN-169 | 2026-09-24 | a coherence audit is cheap when parallel and read-only, and its findings are claims: spot-check, then fix every citer | any doctrine or skill rule change; sub-agent briefs; environment-dependent tests |
|
||||
| LRN-170 | 2026-09-25 | "count == 0" ≠ "all on" when default state is "nothing installed": verify a fast-path premise on the live tree | status/current/detect commands, installer "is X applied?" checks, plan premises copied from stale comments |
|
||||
| LRN-171 | 2026-09-25 | sub-agent sandbox: grep shim returns EMPTY inside `$(...)` for patterns holding literal `$VAR` — oracles pin `command grep` | contract CHECK lines, hermetic test greps, hooks parsing grep output |
|
||||
| LRN-172 | 2026-09-27 | TF-IDF cosine on a 2-doc corpus is identically 0: similarity self-tests need N ≥ 4, a same-corpus positive control and a sensitivity re-run | fixtures for any corpus-normalised statistic (idf, z-score, ranking), "distinct pair passes" tests |
|
||||
| LRN-173 | 2026-09-27 | contract oracles written from memory failed twice: run the CHECK on the precedent files first, census greps via `git grep` (tracked only), `EVIDENCE: pending` mandatory for gates.sh | contract CHECK lines, precedent-mirroring criteria, gates.sh ledgers |
|
||||
|
||||
---
|
||||
|
||||
@@ -1604,3 +1606,12 @@ Rule: when editing a doctrine file under structure locks, grep the test's lock s
|
||||
- **Pattern**: oracles + test assertions portable across main shell / sub-agent shells pin `/usr/bin/grep` or `command grep`; avoid `$VAR` literals mid-pattern (`-F` or `--`).
|
||||
- **Where applicable**: contract `CHECK:` lines run by executors/verifiers; hermetic test greps; any hook riding on grep output.
|
||||
- **Reference**: contract `2026-09-25-default-profile-full-1254` criterion 12; verifier + executor reports 2026-09-25. Links [[LRN-074]] [[BDR-101]].
|
||||
|
||||
## LRN-172 — TF-IDF cosine on a 2-doc corpus is identically 0: similarity self-tests need N ≥ 4, a same-corpus positive control and a sensitivity re-run
|
||||
- **Context**: A3 executor wrote a "distinct pair passes" fixture as a bare 2-doc corpus, documented the 0.00 as "the point being proven". Fresh verifier proved a near-duplicate pair also scores 0.00 at N=2: idf = log(N/df) = 0 for shared terms, unique terms never meet. Executor self-report "all markers printed" was true; the test was vacuous anyway.
|
||||
- **Fix shape**: one N=4 corpus: near-dup control 0.90 → FAIL, distinct pair 0.00 → silent, then swap one doc for a near-copy → 0.62 WARN appears. Marker printed only when all three hold.
|
||||
- **Apply**: any fixture for a corpus-normalised statistic asserts both directions in one corpus; "passes on a trivial corpus" proves nothing; markers prove the oracle, not the intent → keep the blind verifier ([[BDR-102]] [[EVAL-032]]).
|
||||
|
||||
## LRN-173 — contract oracles written from memory failed twice: run the CHECK on the precedent files first, census greps via `git grep`, `EVIDENCE: pending` mandatory
|
||||
- **Context**: same run, two orchestrator oracle bugs. (1) "≤ 80 chars" over the whole file: rules/web-building.md line 2 (`paths:` frontmatter) is already 110 chars, so the criterion contradicted the precedent it named; executor returned NEED-DECISION instead of bending. (2) emil-citers census with `grep -rl` hit gitignored `install-*.log` at the repo root; `git grep -l` (tracked only) is the right census tool. (3) gates.sh `run` errors `runnable but has no EVIDENCE: line` unless each criterion carries `EVIDENCE: pending`.
|
||||
- **Apply**: before shipping a CHECK, run it against the files it claims to mirror; census oracles = `git grep`; contract skeleton carries `EVIDENCE: pending` per criterion (check /feat's template writes it). Oracle fixes are orchestrator-owned, never a re-dispatch ([[BDR-102]]).
|
||||
|
||||
@@ -1,5 +1,31 @@
|
||||
# TODO
|
||||
|
||||
## 2026-09-27 — case 2 of the 6-repo review: borrow from agent-skills (feature/agent-skills-borrow)
|
||||
User go "ok pour les 4" after the analysis: plugin rejected (1.8k tok/session for
|
||||
20 % novelty, /spec /review /ship collide with gstack, trunk-based git and the
|
||||
one-version API rule contradict the doctrine, second router). Four independent
|
||||
chantiers, one contract each under `.claude/tasks/contracts/2026-09-27-*`,
|
||||
dispatched to feater executors; gates replayed by the orchestrator (gates.sh →
|
||||
fresh verifier → fresh security-auditor). Case 1 lives on feature/yagni-ladder.
|
||||
- [x] A1 d28c45e vendor observability-and-instrumentation, deprecation-and-migration,
|
||||
ci-cd-and-automation (emil precedent, pinned commit 2686b620) — agent-skills-vendor
|
||||
- [x] A2 2b25cb4 `lib/floor-guard.sh` diff-scoped bar-weakening detector + verifier step
|
||||
+ suite — floor-guard
|
||||
- [x] A3 409db51 `lib/tests/skill-routing-census.test.sh` description-collision census
|
||||
(measured: 120 skills, max 0.52 careful~guard, 0 >= 0.75) — skill-routing-census
|
||||
- [x] A4 1a8e6de `rules/rest-api.md` path-scoped rule from api-and-interface-design,
|
||||
one-version rule dropped — rest-api-rule
|
||||
- [x] A5 gates MET ×4, verifiers CONFORME ×4 (2 re-dispatches), security PASS ×2,
|
||||
make test 35 suites green minus 2 pre-existing T16a, shellcheck clean;
|
||||
BDR-102 LRN-172 LRN-173 EVAL-032. UNMERGED — human gate.
|
||||
Follow-up (not started): per-skill positive/negative prompt ranking (Tier 2
|
||||
second half); `make link` after merge to symlink the 3 skills (user runs it);
|
||||
floor-guard waiver policy (security MEDIUM: self-service `floor-guard: allow`
|
||||
→ require a CLARIFICATIONS ack outside test fixtures?) — user decision;
|
||||
frame floor-guard snippets as data in the verifier step (LOW); `/tmp/tmp.AAyJzvufO6`
|
||||
scratch dir from an executor proof, `rm -rf` refused → user removes; contract
|
||||
skeleton must carry `EVIDENCE: pending` (LRN-173, check /feat's template).
|
||||
|
||||
## 2026-09-25 — full profile +4 gstack web/doc skills (bugfix/full-profile-web-doc-skills)
|
||||
User go after the "why is gstack off under full?" answer (it was not: unapplied
|
||||
default). `scrape`, `skillify`, `diagram`, `make-pdf` join full.profile; the rest
|
||||
|
||||
@@ -0,0 +1,62 @@
|
||||
# CONTRACT — agent-skills-vendor
|
||||
- date: 2026-09-27 | flow: feat (ad-hoc dispatch, /feat gates replayed by the orchestrator) | branch: feature/agent-skills-borrow
|
||||
- status: active
|
||||
|
||||
## REQUEST (verbatim — IMMUTABLE)
|
||||
> Vendor three skills from addyosmani/agent-skills as machine-owned copies, the emil-design-eng way: `observability-and-instrumentation`, `deprecation-and-migration`, `ci-cd-and-automation`. Source pinned to commit `2686b620fc1fed2e8f60c704839c766b8594c6b6` (main, 2026-09-26) in `plugins.lock.json`; the scripts read the pin from the lock, never hardcode it. Each lands in `skills-external/<name>/SKILL.md` (gitignored, curl'd by install-plugins.sh, refreshed by update-all.sh at the pinned commit), symlinked by link.sh, registered wherever emil-design-eng is registered when the semantics apply (toggle-external registry, profiles that carry the dev skills, tests that enumerate externals, CHANGELOG). User go 2026-09-27 ("ok pour les 4", case 2 of the 6-repo review).
|
||||
|
||||
## CLARIFICATIONS
|
||||
- Profiles: add the three to `full.profile` and to every profile that lists `bugfix` (dev-class); never to design/web profiles.
|
||||
- Not mirrored on purpose (design-only citers): lib/design-gate.md, lib/tests/fixtures/registry-index-drift.md, lib/profiles/{design,web,web-full}.profile, agents/plugin-advisor.md, agents/plugin-probe.md, CLAUDE.global.md.
|
||||
- Upstream SKILL.md copied byte-for-byte: no edits, no rewrite of internal links (dangling cross-skill mentions accepted).
|
||||
- The executor materializes the three files with the same curl the install step uses (network read allowed); it never runs `link.sh`, `make link`, `make plugin` or `update-all.sh` (they touch `~/.claude`).
|
||||
- Lock entry shape: `"agent-skills": {"source": "https://github.com/addyosmani/agent-skills", "commit": "<sha>", "skills": [...], "managed_by": "curl", "note": "..."}`. A helper reading it may follow the `pinned_version` pattern of install-plugins.sh.
|
||||
- The three raw URLs: `https://raw.githubusercontent.com/addyosmani/agent-skills/<sha>/skills/<name>/SKILL.md` (no extra files under these three skills upstream).
|
||||
|
||||
## ACCEPTANCE CRITERIA
|
||||
1. Three vendored files present, frontmatter name = dir name.
|
||||
CHECK: ok=1; for s in observability-and-instrumentation deprecation-and-migration ci-cd-and-automation; do f="skills-external/$s/SKILL.md"; [ -f "$f" ] && grep -q "^name: $s\$" "$f" || { echo "bad $s"; ok=0; }; done; [ "$ok" -eq 1 ] && echo VENDORED
|
||||
EXPECT: VENDORED
|
||||
EVIDENCE: MET exit=0 marker-found :: VENDORED
|
||||
2. Pin recorded in the lock with the three names.
|
||||
CHECK: python3 -c 'import json; d=json.load(open("plugins.lock.json"))["agent-skills"]; assert d["commit"]=="2686b620fc1fed2e8f60c704839c766b8594c6b6", d; assert set(d["skills"])=={"observability-and-instrumentation","deprecation-and-migration","ci-cd-and-automation"}, d; print("PINNED")'
|
||||
EXPECT: PINNED
|
||||
EVIDENCE: MET exit=0 marker-found :: PINNED
|
||||
3. Install and update steps exist and are lock-driven (no hardcoded sha).
|
||||
CHECK: grep -q "agent-skills" install-plugins.sh && grep -q "agent-skills" update-all.sh && ! grep -q "2686b620" install-plugins.sh update-all.sh link.sh lib/toggle-external.sh && echo LOCK_DRIVEN
|
||||
EXPECT: LOCK_DRIVEN
|
||||
EVIDENCE: MET exit=0 marker-found :: LOCK_DRIVEN
|
||||
4. Both the copy and the symlink path are gitignored.
|
||||
CHECK: ok=1; for s in observability-and-instrumentation deprecation-and-migration ci-cd-and-automation; do git check-ignore -q "skills-external/$s" && git check-ignore -q "skills/$s" || { echo "not ignored $s"; ok=0; }; done; [ "$ok" -eq 1 ] && echo IGNORED
|
||||
EXPECT: IGNORED
|
||||
EVIDENCE: MET exit=0 marker-found :: IGNORED
|
||||
5. link.sh symlinks them (EXTERNAL_SKILLS list).
|
||||
CHECK: ok=1; for s in observability-and-instrumentation deprecation-and-migration ci-cd-and-automation; do grep -q "$s" link.sh || { echo "not linked $s"; ok=0; }; done; [ "$ok" -eq 1 ] && echo LINK_LISTED
|
||||
EXPECT: LINK_LISTED
|
||||
EVIDENCE: MET exit=0 marker-found :: LINK_LISTED
|
||||
6. Emil citers census mirrored (tracked files only: the ignored install-*.log files at the root also name emil), except the design-only allowlist.
|
||||
CHECK: allow=" lib/design-gate.md lib/tests/fixtures/registry-index-drift.md lib/profiles/design.profile lib/profiles/web.profile lib/profiles/web-full.profile agents/plugin-advisor.md agents/plugin-probe.md CLAUDE.global.md "; miss=0; for f in $(git grep -l emil-design-eng -- . ':!skills-external' ':!skills' ':!.claude'); do grep -q observability-and-instrumentation "$f" && continue; case "$allow" in *" $f "*) ;; *) echo "not mirrored: $f"; miss=1 ;; esac; done; [ "$miss" -eq 0 ] && echo CENSUS_MIRRORED
|
||||
EXPECT: CENSUS_MIRRORED
|
||||
EVIDENCE: MET exit=0 marker-found :: CENSUS_MIRRORED
|
||||
7. Profile, toggle-external and doctrine-citers suites green.
|
||||
CHECK: out=$(make test suite="lib/tests/profile-default.test.sh lib/tests/profile-set-managed.test.sh lib/tests/toggle-external-repo-resolution.test.sh lib/tests/doctrine-citers.test.sh" 2>&1); echo "$out" | grep -qE "FAIL=[1-9]" && { echo "$out" | tail -20; exit 1; }; echo SUITES_GREEN
|
||||
EXPECT: SUITES_GREEN
|
||||
EVIDENCE: MET exit=0 marker-found :: SUITES_GREEN
|
||||
8. shellcheck clean on the touched scripts.
|
||||
CHECK: shellcheck install-plugins.sh update-all.sh link.sh lib/toggle-external.sh lib/profile.sh && echo SHELLCHECK_OK
|
||||
EXPECT: SHELLCHECK_OK
|
||||
EVIDENCE: MET exit=0 marker-found :: SHELLCHECK_OK
|
||||
|
||||
## FILE SCOPE
|
||||
- plugins.lock.json, install-plugins.sh, update-all.sh, link.sh, .gitignore
|
||||
- lib/toggle-external.sh, lib/profile.sh (only if the externals registry lives there), lib/profiles/full.profile + dev-class profiles
|
||||
- lib/tests/profile-default.test.sh, lib/tests/profile-set-managed.test.sh, lib/tests/toggle-external-repo-resolution.test.sh (only if they enumerate externals)
|
||||
- CHANGELOG.md (Unreleased entry)
|
||||
- skills-external/<name>/SKILL.md x3 (materialized, gitignored)
|
||||
|
||||
## PLAN
|
||||
1. `grep -rn emil-design-eng` census (files listed in criterion 6) → mirror file by file, same comment density.
|
||||
2. Lock entry; install-plugins.sh new step next to Step 8 ("agent-skills — 3 skills, pinned commit") reading the sha from the lock (python3/jq like the existing helpers), curl each raw URL into skills-external/<name>/SKILL.md, skip when present, `err` with the manual command on failure; update-all.sh step re-curls at the pinned sha (tmp + mv, emil precedent).
|
||||
3. link.sh EXTERNAL_SKILLS += 3; .gitignore: `skills/<name>` in the symlink allowlist + `skills-external/<name>/` with a short comment naming the source (emil precedent).
|
||||
4. toggle-external registry + profiles + tests that enumerate externals.
|
||||
5. Materialize the three files with the curl; run criteria 1-8; report the emil citers you deliberately did not mirror and why.
|
||||
@@ -0,0 +1,49 @@
|
||||
# CONTRACT — floor-guard
|
||||
- date: 2026-09-27 | flow: feat (ad-hoc dispatch, /feat gates replayed by the orchestrator) | branch: feature/agent-skills-borrow
|
||||
- status: active
|
||||
|
||||
## REQUEST (verbatim — IMMUTABLE)
|
||||
> Build `lib/floor-guard.sh`, a diff-scoped deterministic detector of a quietly weakened quality bar, adapted from addyosmani/agent-skills `constraint-driven-development` (floor guard) to this repo's gate model. Usage `bash ~/.claude/lib/floor-guard.sh <base-ref> [-- <pathspec>...]` over `git diff <base-ref>` (working tree included). Kinds: SUPPRESS (added checker silencing: `@ts-ignore`, `@ts-expect-error` without a trailing reason, `eslint-disable*`, `# noqa`, `# type: ignore`, `nosemgrep`, `nosec`, `shellcheck disable`), SKIP (added `.skip(`, `.only(`, `xit(`, `xdescribe(`, `fit(`, `fdescribe(`, `it.todo(`, `@pytest.mark.skip`, `@unittest.skip`, `t.Skip(` in test files), DELETED_TEST (deleted file whose path matches a test pattern), ASSERT_DROP (a test file whose assertion-line count decreases: `expect(`, `assert`, `should`, `.toBe`), STUB (added `not implemented`, `NotImplementedError`, empty `catch` block, `except: pass`), THRESHOLD_DOWN (a numeric value decreased on the same key in coverage/quality config files: `jest.config*`, `vitest.config*`, `.nycrc*`, `codecov*`, `sonar-project.properties`, `lighthouserc*`, `CONSTRAINTS.md`). Waiver: an added line carrying `floor-guard: allow <reason>` is printed as WAIVED and not counted. Output: one `FLOOR <KIND> <file>:<line> <snippet>` per finding, then `FLOOR GUARD: clean` (rc 0) or `FLOOR GUARD: <n> finding(s), <m> waived` (rc 2); rc 3 on usage error. Wire it as a mandatory verifier step (agents/verifier.md) and document it in lib/verify-secure-loop.md GATE 1; hermetic suite `lib/tests/floor-guard.test.sh`. User go 2026-09-27 ("ok pour les 4", case 2 item 2).
|
||||
|
||||
## CLARIFICATIONS
|
||||
- Language: bash entry point; the diff parsing may live in an embedded python3 heredoc (precedent `lib/tests/run-review-guards.sh`). Functions <= 25 logic lines, 80-char lines.
|
||||
- File classes: test file = path contains `test`, `spec`, `__tests__`, or matches `*.test.*`, `*.spec.*`, `*_test.go`, `*_test.py`, `test_*.py`; config file = the names listed under THRESHOLD_DOWN. SKIP and ASSERT_DROP apply to test files only; SUPPRESS and STUB to any file; THRESHOLD_DOWN to config files only.
|
||||
- The guard's own pattern table contains the trigger strings: those source lines carry `# floor-guard: allow pattern table` so the guard stays clean on itself (this also exercises the waiver path for real).
|
||||
- Verifier step: `bash ~/.claude/lib/floor-guard.sh <base>` where base = the branch's gitflow base (develop; main for hotfix/release); rc 2 → verdict ECARTS listing each FLOOR line, unless the contract's CLARIFICATIONS explicitly authorize that exact weakening (quote it in the verdict).
|
||||
- Suite: throwaway repos (`make test` exports GIT_CONFIG_GLOBAL=/dev/null); one RED fixture per kind, one WAIVED fixture, one CLEAN fixture, each printed as `PASS <KIND>`; summary `PASS=n FAIL=m` like the other suites.
|
||||
- Self-detection: the suite file itself contains the trigger strings; the self-run criterion excludes it by pathspec.
|
||||
- Out of scope: a pre-commit hook, running it on the repo's own diff in `make test`, parsers beyond the regexes above.
|
||||
|
||||
## ACCEPTANCE CRITERIA
|
||||
1. Suite green, every kind flip-tested.
|
||||
CHECK: out=$(make test suite=lib/tests/floor-guard.test.sh 2>&1); for k in SUPPRESS SKIP DELETED_TEST ASSERT_DROP STUB THRESHOLD_DOWN WAIVED CLEAN; do echo "$out" | grep -q "PASS $k" || { echo "missing PASS $k"; echo "$out" | tail -15; exit 1; }; done; echo "$out" | grep -qE "FAIL=[1-9]" && exit 1; echo KINDS_GREEN
|
||||
EXPECT: KINDS_GREEN
|
||||
EVIDENCE: MET exit=0 marker-found :: KINDS_GREEN
|
||||
2. Usage error is rc 3.
|
||||
CHECK: bash lib/floor-guard.sh >/dev/null 2>&1; [ $? -eq 3 ] && echo RC_USAGE
|
||||
EXPECT: RC_USAGE
|
||||
EVIDENCE: MET exit=0 marker-found :: RC_USAGE
|
||||
3. Self-run clean on this branch (suite file excluded by pathspec), waivers visible.
|
||||
CHECK: out=$(bash lib/floor-guard.sh develop -- . ':!lib/tests/floor-guard.test.sh' 2>&1); rc=$?; echo "$out" | tail -3; [ $rc -eq 0 ] && echo "$out" | grep -q WAIVED && echo SELF_CLEAN
|
||||
EXPECT: SELF_CLEAN
|
||||
EVIDENCE: MET exit=0 marker-found :: WAIVED STUB lib/floor-guard.sh:105 'not implemented', # floor-guard: allow pattern table WAIVED STUB lib/floor-guard.sh:106 'NotImplementedE…
|
||||
4. Verifier wired, loop documented.
|
||||
CHECK: grep -q "floor-guard.sh" agents/verifier.md && grep -q "floor-guard" lib/verify-secure-loop.md && echo WIRED
|
||||
EXPECT: WIRED
|
||||
EVIDENCE: MET exit=0 marker-found :: WIRED
|
||||
5. shellcheck and doctrine-citers clean.
|
||||
CHECK: shellcheck lib/floor-guard.sh lib/tests/floor-guard.test.sh && out=$(make test suite=lib/tests/doctrine-citers.test.sh 2>&1) && ! echo "$out" | grep -qE "FAIL=[1-9]" && echo LINT_OK
|
||||
EXPECT: LINT_OK
|
||||
EVIDENCE: MET exit=0 marker-found :: LINT_OK
|
||||
|
||||
## FILE SCOPE
|
||||
- lib/floor-guard.sh (new), lib/tests/floor-guard.test.sh (new)
|
||||
- agents/verifier.md (one mandatory step), lib/verify-secure-loop.md (one paragraph under GATE 1)
|
||||
- CHANGELOG.md (Unreleased entry)
|
||||
|
||||
## PLAN
|
||||
1. Script: arg parsing (base, optional `--` pathspec), `git diff --unified=0 <base> -- <pathspec>` plus `git diff --diff-filter=D --name-only <base> -- <pathspec>`, rc contract, header comment stating WHY (BDR-100 class: deterministic floor under an LLM gate).
|
||||
2. Python parser on stdin: walk hunks, classify added lines by kind and file class, count assertion lines removed vs added per test file, compare numeric values on identical keys in config files (`key: 80` → `key: 60`, JSON or YAML-ish), detect waivers.
|
||||
3. Output lines + summary + rc.
|
||||
4. Suite: helper `mk_repo` (git init -q, base commit with a test file holding 3 assertions, a `vitest.config.ts` with `lines: 80`, a source file), one fixture per kind → assert rc 2 and the FLOOR line; WAIVED → rc 0 and a WAIVED line; CLEAN → rc 0.
|
||||
5. verifier.md step + verify-secure-loop.md paragraph + CHANGELOG; run criteria 1-5.
|
||||
@@ -0,0 +1,34 @@
|
||||
# CONTRACT — rest-api-rule
|
||||
- date: 2026-09-27 | flow: feat (ad-hoc dispatch, /feat gates replayed by the orchestrator) | branch: feature/agent-skills-borrow
|
||||
- status: active
|
||||
|
||||
## REQUEST (verbatim — IMMUTABLE)
|
||||
> Write `rules/rest-api.md`, a path-scoped user rule distilled from addyosmani/agent-skills `api-and-interface-design` (commit 2686b620fc1fed2e8f60c704839c766b8594c6b6), the way `rules/web-building.md` is written: `paths:` frontmatter, English, <= 45 lines, 80-char lines. Keep: contract-first order (typed interface → schemas with server-generated fields apart → error codes → validate at boundaries only); one error envelope `{ error: { code, message, details? } }` with the HTTP map 400 invalid / 401 auth / 403 forbidden / 404 missing / 409 conflict / 422 semantic / 500 server; every list endpoint paginated (`page`, `pageSize`, `totalItems`, `totalPages`), filters as query params; idempotency (key derived from intent, atomic claim via unique constraint, payload guard, explicit in-flight duplicate policy 409 / wait / 202, retention beyond the longest retry path incl. dead-letter); naming (plural nouns, camelCase params and fields, UPPER_SNAKE enums, is/has/can booleans); one Hyrum's law line. Drop the upstream one-version rule: versioning points to the CLAUDE.md heading "Web APIs — always versioned". User go 2026-09-27 ("ok pour les 4", case 2 item 4).
|
||||
|
||||
## CLARIFICATIONS
|
||||
- Globs: `["**/api/**", "**/routes/**", "**/controllers/**", "**/*.route.*", "**/*.controller.*", "**/openapi.*", "**/*.openapi.*"]`.
|
||||
- Fetch the upstream text with curl at the pinned commit to distill from; never vendor it.
|
||||
- Cite the doctrine as `CLAUDE.md § Web APIs — always versioned` (exact heading, so the doctrine-citers census resolves it).
|
||||
- No routing line in CLAUDE.global.md, no README change, rules/README.md unchanged; CHANGELOG Unreleased entry.
|
||||
|
||||
## ACCEPTANCE CRITERIA
|
||||
1. Shape: frontmatter, paths JSON list, <= 45 lines, body lines <= 80 chars (the one-line `paths:` frontmatter is exempt, repo precedent: web-building.md / web-security.md line 2).
|
||||
CHECK: f=rules/rest-api.md; [ -f "$f" ] && [ "$(head -1 "$f")" = "---" ] && python3 -c 'import re,json; s=open("rules/rest-api.md").read(); m=re.search(r"^paths: (.*)$",s,re.M); assert len(json.loads(m.group(1)))>=5' && [ "$(wc -l < "$f")" -le 45 ] && ! awk 'NR>3 && length>80' "$f" | grep -q . && echo RULE_SHAPE
|
||||
EXPECT: RULE_SHAPE
|
||||
EVIDENCE: MET exit=0 marker-found :: RULE_SHAPE
|
||||
2. Content present, one-version rule absent.
|
||||
CHECK: f=rules/rest-api.md; ok=1; for k in "code" "message" "409" "422" "pageSize" "totalItems" "idempoten" "unique" "camelCase" "UPPER_SNAKE" "Hyrum" "always versioned"; do grep -qi -- "$k" "$f" || { echo "missing $k"; ok=0; }; done; grep -qi "extend rather than fork" "$f" && { echo "one-version rule present"; ok=0; }; [ "$ok" -eq 1 ] && echo RULE_CONTENT
|
||||
EXPECT: RULE_CONTENT
|
||||
EVIDENCE: MET exit=0 marker-found :: RULE_CONTENT
|
||||
3. Doctrine citation resolves.
|
||||
CHECK: out=$(make test suite=lib/tests/doctrine-citers.test.sh 2>&1); echo "$out" | grep -qE "FAIL=[1-9]" && { echo "$out" | tail -10; exit 1; }; echo CITERS_OK
|
||||
EXPECT: CITERS_OK
|
||||
EVIDENCE: MET exit=0 marker-found :: CITERS_OK
|
||||
|
||||
## FILE SCOPE
|
||||
- rules/rest-api.md (new), CHANGELOG.md (Unreleased entry)
|
||||
|
||||
## PLAN
|
||||
1. curl the upstream SKILL.md at the pin into the scratch dir; read it.
|
||||
2. Write the rule: title "REST API — contract, errors, lists, idempotency"; sections Contract first · Errors · Lists · Idempotency · Naming · Versioning (one line pointing to the doctrine heading).
|
||||
3. CHANGELOG line; run criteria 1-3.
|
||||
@@ -0,0 +1,42 @@
|
||||
# CONTRACT — skill-routing-census
|
||||
- date: 2026-09-27 | flow: feat (ad-hoc dispatch, /feat gates replayed by the orchestrator) | branch: feature/agent-skills-borrow
|
||||
- status: active
|
||||
|
||||
## REQUEST (verbatim — IMMUTABLE)
|
||||
> Build `lib/tests/skill-routing-census.test.sh`: a deterministic census of skill-description collisions across the live catalog, adapted from addyosmani/agent-skills evals Tier 2. Catalog = every `SKILL.md` under `~/.claude/skills/*/` (symlinks resolved) plus plugin skills under `~/.claude/plugins/cache/*/*/*/skills/*/SKILL.md` and `~/.claude/plugins/cache/*/*/*/.claude/skills/*/SKILL.md` (roots overridable via `SKILL_ROUTING_ROOTS`, colon-separated dirs, for fixtures). Extract `description` (scalar or `|`/`>` block), tokenize (lowercase, `[a-z][a-z0-9-]+`, stopwords, suffix stemming s/es/ed/ing), TF-IDF cosine over all pairs. Print `skills with description: N`, the top 10 pairs as `0.52 a ~ b`, WARN lines for pairs >= `SKILL_ROUTING_WARN` (default 0.50), FAIL for pairs >= `SKILL_ROUTING_FAIL` (default 0.75). Self-test on fixtures: a near-duplicate pair must FAIL (`FIXTURE_COLLISION_DETECTED`), a distinct pair must pass (`FIXTURE_DISTINCT_OK`). Measured today: 120 skills, max 0.52 (careful ~ guard), 0 pairs >= 0.75 → green with one WARN. User go 2026-09-27 ("ok pour les 4", case 2 item 3).
|
||||
|
||||
## CLARIFICATIONS
|
||||
- python3 embedded in the bash suite is allowed (precedent run-review-guards.sh). If the python body exceeds ~120 lines, put it in `lib/skill-routing-census.py` and keep the suite as the wrapper.
|
||||
- Reference implementation (read it, reuse the logic, harden the block-description parsing): `/tmp/claude-1000/-home-bchanot-Documents-claude/977f1703-f01d-497a-b794-5b69fafcd35f/scratchpad/census.py`.
|
||||
- Thresholds stay at the upstream defaults; no allowlist file; a WARN never fails the suite.
|
||||
- Dedup by skill directory name: first path wins.
|
||||
- Out of scope (follow-up): positive/negative prompt ranking per skill.
|
||||
- The live census is machine-dependent by design (it audits this machine's catalog); the fixture self-test is the hermetic part. On a machine with an empty catalog the live pass prints `skills with description: 0` and passes.
|
||||
|
||||
## ACCEPTANCE CRITERIA
|
||||
1. Suite green on the live catalog, catalog non-trivial here.
|
||||
CHECK: out=$(make test suite=lib/tests/skill-routing-census.test.sh 2>&1); echo "$out" | grep -qE "FAIL=[1-9]" && { echo "$out" | tail -15; exit 1; }; echo "$out" | grep -qE "skills with description: *[0-9]{2,}" && echo LIVE_GREEN
|
||||
EXPECT: LIVE_GREEN
|
||||
EVIDENCE: MET exit=0 marker-found :: LIVE_GREEN
|
||||
2. Fixture flip: collision detected, distinct pair passes.
|
||||
CHECK: out=$(make test suite=lib/tests/skill-routing-census.test.sh 2>&1); echo "$out" | grep -q "FIXTURE_COLLISION_DETECTED" && echo "$out" | grep -q "FIXTURE_DISTINCT_OK" && echo FLIP_TESTED
|
||||
EXPECT: FLIP_TESTED
|
||||
EVIDENCE: MET exit=0 marker-found :: FLIP_TESTED
|
||||
3. Report shape: top pairs with two-decimal scores.
|
||||
CHECK: out=$(make test suite=lib/tests/skill-routing-census.test.sh 2>&1); echo "$out" | grep -qE "^0\.[0-9]{2} [a-z0-9:_-]+ ~ [a-z0-9:_-]+" && echo REPORT_SHAPE
|
||||
EXPECT: REPORT_SHAPE
|
||||
EVIDENCE: MET exit=0 marker-found :: REPORT_SHAPE
|
||||
4. shellcheck clean.
|
||||
CHECK: shellcheck lib/tests/skill-routing-census.test.sh && echo SHELLCHECK_OK
|
||||
EXPECT: SHELLCHECK_OK
|
||||
EVIDENCE: MET exit=0 marker-found :: SHELLCHECK_OK
|
||||
|
||||
## FILE SCOPE
|
||||
- lib/tests/skill-routing-census.test.sh (new); optional lib/skill-routing-census.py (new)
|
||||
- CHANGELOG.md (Unreleased entry)
|
||||
|
||||
## PLAN
|
||||
1. Roots: default globs; `SKILL_ROUTING_ROOTS` override; dedup by dir name.
|
||||
2. Description extraction: scalar `description: text`; block `description: |` or `>` → join the indented continuation lines.
|
||||
3. Tokens / TF-IDF / cosine as in census.py: stopword set, stem(), (1+log tf)·log(N/df), L2 norm, cosine over combinations.
|
||||
4. Report + thresholds + rc. Suite: run live (a FAIL line → suite RED), then two fixture dirs under mktemp: A/B near-duplicate descriptions → expect a FAIL line → print FIXTURE_COLLISION_DETECTED; C/D distinct → expect rc 0 → print FIXTURE_DISTINCT_OK. `PASS=n FAIL=m` summary like the other suites.
|
||||
@@ -7,6 +7,35 @@ Format follows [Keep a Changelog](https://keepachangelog.com/).
|
||||
## [Unreleased]
|
||||
|
||||
### Added
|
||||
- **Agent Skills trio** (`observability-and-instrumentation`,
|
||||
`deprecation-and-migration`, `ci-cd-and-automation`) — vendored from
|
||||
addyosmani/agent-skills at a pinned commit (`agent-skills` entry in
|
||||
plugins.lock.json), the emil-design-eng way: curl'd into
|
||||
`skills-external/<name>/SKILL.md` by install-plugins.sh Step 8e, refreshed
|
||||
by update-all.sh at the same commit, symlinked by link.sh, registered in
|
||||
`lib/toggle-external.sh`, `lib/profile.sh` and the `full`/`backend`/`dev`
|
||||
profiles. Case 2 of the 6-repo review: the plugin itself was rejected
|
||||
(1.8k tokens per session for 20 % novelty, `/spec` `/review` `/ship`
|
||||
collide with gstack, trunk-based git and the one-version API rule
|
||||
contradict the doctrine, a second skill router).
|
||||
- **`lib/floor-guard.sh`** — diff-scoped deterministic detector of a quietly
|
||||
weakened quality bar (new lint/type suppressions, skipped or deleted
|
||||
tests, dropped assertions, stubs, lowered coverage thresholds), with a
|
||||
`floor-guard: allow <reason>` waiver, rc 0/2/3. Mandatory verifier
|
||||
STEP 3 (`agents/verifier.md`), documented under GATE 1 of
|
||||
`lib/verify-secure-loop.md`. Suite `lib/tests/floor-guard.test.sh`: 6
|
||||
kinds plus a WAIVED and a CLEAN fixture, each flip-tested. Adapted from agent-skills `constraint-driven-development`.
|
||||
- **`lib/tests/skill-routing-census.test.sh`** (+ `lib/skill-routing-census.py`)
|
||||
— TF-IDF cosine census of skill-description collisions across the live
|
||||
catalog (routing ambiguity, not naming): top 10 pairs, WARN ≥ 0.50,
|
||||
FAIL ≥ 0.75, fixture flip-test. Baseline 2026-09-27: 120 skills, max 0.52
|
||||
(`careful` ~ `guard`). Adapted from agent-skills evals Tier 2.
|
||||
- **`rules/rest-api.md`** — path-scoped REST rule distilled from agent-skills
|
||||
`api-and-interface-design`: contract-first order, one error envelope +
|
||||
HTTP map, paginated lists, idempotency (key from intent, atomic claim,
|
||||
payload guard, duplicate policy, retention), naming, Hyrum's law;
|
||||
versioning points to `CLAUDE.md § Web APIs — always versioned` instead of
|
||||
the upstream one-version rule.
|
||||
- **`make test suite=<file>`** runs one suite hermetically; the
|
||||
`GIT_CONFIG_GLOBAL=/dev/null` export lives in the Makefile so nobody types
|
||||
the denied env-prefix form by hand (the reason an executor wrote a wrapper
|
||||
|
||||
Reference in New Issue
Block a user