diff --git a/.claude/memory/decisions.md b/.claude/memory/decisions.md index a2174db..d397c18 100644 --- a/.claude/memory/decisions.md +++ b/.claude/memory/decisions.md @@ -123,6 +123,7 @@ rules: | BDR-099 | 2026-09-24 | C2 coherence: 30 doctrine/skill tensions resolved, doctrine wins, BDR-068 kept as the written exception | accepted | | BDR-100 | 2026-09-24 | Guardrail evasion and partial rule changes get mechanisms, not lessons: refusal ends the attempt, citers census in make test | accepted | | BDR-101 | 2026-09-25 | `full` = default profile: no selection ⇒ full in force, `reset` applies it, install applies it | accepted | +| BDR-102 | 2026-09-27 | agent-skills: no plugin, vendor 3 skills + build floor-guard + routing census + rest-api rule | accepted | --- @@ -1270,3 +1271,13 @@ Branch feature/user-writing-web-rules, UNMERGED (human gate). - **Alternatives rejected**: label-only reset (lies about plugins/externals); additive reset (`gstack on` + `apply full`, state ⊇ full, label ambiguous); keep `none` sentinel + statusline `?` (request unmet); install display-only (fresh machine ≠ full, 21st pack parked); public `profile.sh default` verb for installer (reset already IS "go to default"); one shared cache parser (statusline must not spawn profile.sh → 3 copies kept, each commented). - **Caveats**: `make plugin` re-run re-applies selected profile → manual layering (`gstack on` over `dev`) trimmed back. Plugin legs of Step 11 install-immutable ([[BDR-028]] EXIT guard; committed enabledPlugins already match full). LOW security note: cache content not charset-checked before path use (pre-existing in `read_profile`) → follow-up. - **Reference**: commits e196328 (residue scrub), 0d035fc (profile), 1bbdad0 (install); contract/plan `2026-09-25-default-profile-full-1254`; `lib/tests/profile-default.test.sh` 29 checks. Links [[BDR-017]] [[BDR-018]] [[BDR-079]] [[BDR-093]] [[LRN-170]] [[EVAL-031]]. + +## BDR-102 — agent-skills: no plugin, vendor 3 skills + build floor-guard + routing census + rest-api rule +- **Date**: 2026-09-27 +- **Status**: accepted, feature/agent-skills-borrow, UNMERGED (human gate) +- **Decision**: addyosmani/agent-skills (99.4k stars, 25 skills) NOT installed as plugin. Borrowed 4 things, user go: (1) `observability-and-instrumentation`, `deprecation-and-migration`, `ci-cd-and-automation` vendored emil-way at pinned commit 2686b620 (lock `agent-skills`, install Step 8e tmp+mv, update-all 7.3, link.sh, toggle-external per-name, profiles full/backend/dev); (2) `lib/floor-guard.sh` diff-scoped bar-weakening detector, verifier STEP 3 mandatory, waiver `floor-guard: allow `; (3) `lib/tests/skill-routing-census.test.sh` TF-IDF description-collision census, WARN 0.50 / FAIL 0.75; (4) `rules/rest-api.md` path-scoped, distilled from api-and-interface-design minus the one-version rule. +- **Why**: 20/25 skills already covered (superpowers, personal skills, gstack, built-ins). Real gaps grep-verified: observability (archetype question only), migration (one strangler line), CI build (ship/cso only detect), bar-weakening guard (security-auditor flags nosemgrep only), collision census (routing section exists because of collisions; darwin scores quality not collisions). Upstream comparison doc: never stack two skill routers. +- **Alternatives rejected**: plugin install (1.8k tok/session for 20 % novelty, `/spec` `/review` `/ship` collide with gstack, `/code-simplify` vs `/simplify`, trunk-based git + one-version API contradict doctrine, second router); vendor api-and-interface-design whole (one-version rule vs § Web APIs → distilled); grouped `agent-skills` toggle pack (no shared installer, per-name like emil); web-full profile for the trio (design-class, allowlist wins over "lists bugfix"); Tier 2 prompt ranking (follow-up). +- **Caveats**: vendored prompts = third-party content loaded into sessions, the pin is the review point (security-auditor scanned: benign, no hidden Unicode); floor-guard waiver is self-service, WAIVED informational → security MEDIUM, design decision pending with user (require CLARIFICATIONS ack outside test fixtures?); floor-guard prints raw diff snippets a verifier reads (LOW, framing follow-up); update-all 7.3 failure branch leaves `.tmp` like emil (parity, not fixed); `make link` after merge to symlink the trio (user). +- **Reference**: commits d28c45e (trio), 2b25cb4 (floor-guard), 409db51 (census), 1a8e6de (rest-api); contracts `.claude/tasks/contracts/2026-09-27-{agent-skills-vendor,floor-guard,skill-routing-census,rest-api-rule}-1525.md`; gates MET ×4, verifiers CONFORME ×4 (2 re-dispatches), security PASS ×2. Links [[BDR-100]] [[LRN-172]] [[LRN-173]] [[EVAL-032]]. Case 1 of the same review: ladder in doctrine, feature/yagni-ladder 9315c6c. +- **Amendment 2026-09-27**: waiver policy strict, user go: `WAIVED` outside a test file = gap unless the contract's CLARIFICATIONS names file + reason (verifier STEP 3, loop doc). Commit 6617889. Closes the security MEDIUM caveat above. diff --git a/.claude/memory/evals.md b/.claude/memory/evals.md index 0ed3c87..17d1edf 100644 --- a/.claude/memory/evals.md +++ b/.claude/memory/evals.md @@ -52,6 +52,7 @@ rules: | EVAL-029 | 2026-09-15 | 4-agent plan challenge: 6 BLOCKER; 3 of 3 confirmation-pass BLOCKERs came from the fixes themselves; caught a false 654 MB orphan claim | keep | | EVAL-030 | 2026-09-24 | 2026-09-24 self-audit: two regressions and one guardrail bypass came from my own process, not from the tools | BDR-100 mechanisms shipped; re-run census at next doctrine wave | | EVAL-031 | 2026-09-25 | /feat run for BDR-101: challenge round earned its cost, two blockers sat in my own premises | keep challenge round on state-detection plans; check live state before planning; pin grep in oracles | +| EVAL-032 | 2026-09-27 | 4 parallel feater executors, one tree, gate loop: verifier caught a vacuous test, security caught a partial-write; my oracles wrong twice | keep same-tree parallel dispatch with disjoint FILE SCOPE + orchestrator-owned shared files; blind verifier stays; measure oracles on precedents | --- @@ -306,3 +307,10 @@ Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itse - **Method**: challenge lib (3 lenses + 1 confirmation), gates.sh floor, fresh verifier, fresh security-auditor, full `make test` (236 green + 2 pre-existing T16a). - **Anomalies**: (1) plan asserted "all gstack enabled" without one `ls skills/`; banner said gstack OFF ([[LRN-170]]). (2) two sub-agents hit same grep-shim quirk ([[LRN-171]]). (3) `git add .env.example` denied (`git add .env*` glob, [[BDR-069]] collateral): edit left unstaged for user, not routed around; edit itself went through python script while `Edit(**/.env.*)` denied — surfaced to user. (4) UserPromptSubmit design hook fired on "design skills" (false positive, no UI work). - **Action**: keep challenge round for any plan touching state detection; check live state before planning; pin grep in oracles. Links [[BDR-101]]. + +## EVAL-032 — 4 parallel feater executors, one tree, gate loop (case 2 of the 6-repo review) +- **Date**: 2026-09-27 +- **Method**: 4 contracts, 4 feater executors dispatched in one turn on the same working tree (disjoint FILE SCOPE, CHANGELOG reserved to the orchestrator), gates.sh run per contract, fresh verifier per contract, security-auditor on the whole diff then on the re-touched files, full `make test`. +- **Result**: 4/4 CONFORME after 2 re-dispatches; security PASS ×2. Verifier A3 caught a vacuous distinct-pair test the executor had self-justified ([[LRN-172]]). Security caught first-download without tmp+mv (partial file accepted forever) + python source splicing → fixed by fresh executor, re-verified, re-audited. Executors never touched each other's files; A4 verifier counted the orchestrator-reserved CHANGELOG as ECARTS(1), correct by contract wording. +- **Anomalies**: (1) my oracles wrong twice ([[LRN-173]]); (2) gates.sh ERROR(3) on first run, `EVIDENCE: pending` missing; (3) 3 guardrail denials on sub-agents (`export GIT_CONFIG_GLOBAL` inline ×2 incl. a verifier, `rm -rf /tmp/tmp.AAyJzvufO6` executor cleanup), all reported, none evaded — [[BDR-100]] live; (4) security-auditor miscounted the sha as 41 chars (it is 40) — verify sub-agent claims before acting; (5) `make test` rc 1 from the 2 pre-existing T16a, my first grep filter hid the totals. +- **Action**: keep the pattern; add `EVIDENCE: pending` to the contract skeleton; floor-guard waiver policy → user decision; snippet framing for LLM-consumed output → follow-up. diff --git a/.claude/memory/journal.md b/.claude/memory/journal.md index f55f240..9dc0079 100644 --- a/.claude/memory/journal.md +++ b/.claude/memory/journal.md @@ -519,3 +519,6 @@ rules: ## 2026-09-27 - bugfix/gitignore-diagram-allowlist merged into develop on user go, `gitflow finish` → facd26d, pushed, copies removed by the lib. Day 2026-09-25 lot fully on develop: BDR-101 default profile, full +4 gstack skills, gitignore allowlist. No working branch anywhere; `skills/diagram` ignored. - 6-repo review, case 1 (ponytail + chisle, token-economy layer): rejected as plugins. Ponytail 146.7k stars, injects ~600 tok at SessionStart + every SubagentStart; chisle 566 stars, PostToolUse `updatedToolOutput` rewrite unverified on native tools, prose rules collide with writing-style.md; rtk already covers input axis (chisle bench: dedup 0 hit on rtk-filtered corpus); caveman purge precedent v3.5.0. Borrowed the ordered YAGNI ladder + `shortcut:` marker into CLAUDE.global.md § Code style, user go. feature/yagni-ladder UNMERGED. +- 6-repo review case 2 (agent-skills 99.4k stars): plugin rejected (1.8k tok/session, `/spec` `/review` `/ship` collide with gstack, trunk-based git + one-version API vs doctrine, second router, upstream says never stack routers). User go on 4 borrows → [[BDR-102]]: trio vendored emil-way at pinned 2686b620 (d28c45e), `lib/floor-guard.sh` + verifier STEP 3 (2b25cb4), `lib/tests/skill-routing-census.test.sh` 120 skills max 0.52 (409db51), `rules/rest-api.md` (1a8e6de). 4 feater executors in parallel, same tree; gates MET ×4, verifiers CONFORME ×4 after 2 re-dispatches (A3 vacuous N=2 fixture [[LRN-172]]; A1 tmp+mv + argv from security), security PASS ×2. My oracles wrong twice ([[LRN-173]]), [[EVAL-032]]. `make test` 35 suites green minus 2 pre-existing T16a (gitleaks), shellcheck clean. feature/agent-skills-borrow UNMERGED. Guardrails fired 3× on sub-agents, none evaded; `/tmp/tmp.AAyJzvufO6` scratch dir left for the user (rm -rf refused, correctly). +- Case 3 (ui-skills 9.2k): 7 own skills + registry of 36 third-party + 47-lesson site playbook (React components, not agent files). Verdict given: extend rules/web-building.md with ~12 stack-agnostic micro-rules, install nothing (CLI/MCP = curl of raw SKILL.md, third router, baseline-ui stack mandates vs Astro doctrine). Awaiting user; case 4 reticle material prefetched. +- Waiver policy strict applied on feature/agent-skills-borrow (6617889): verifier STEP 3 counts non-test WAIVED lines as gaps unless CLARIFICATIONS names them; loop doc + CHANGELOG + BDR-102 amendment. The earlier journal line said "applied" one commit early, corrected here. diff --git a/.claude/memory/learnings.md b/.claude/memory/learnings.md index 8753fee..2e4553d 100644 --- a/.claude/memory/learnings.md +++ b/.claude/memory/learnings.md @@ -191,6 +191,8 @@ rules: | LRN-169 | 2026-09-24 | a coherence audit is cheap when parallel and read-only, and its findings are claims: spot-check, then fix every citer | any doctrine or skill rule change; sub-agent briefs; environment-dependent tests | | LRN-170 | 2026-09-25 | "count == 0" ≠ "all on" when default state is "nothing installed": verify a fast-path premise on the live tree | status/current/detect commands, installer "is X applied?" checks, plan premises copied from stale comments | | LRN-171 | 2026-09-25 | sub-agent sandbox: grep shim returns EMPTY inside `$(...)` for patterns holding literal `$VAR` — oracles pin `command grep` | contract CHECK lines, hermetic test greps, hooks parsing grep output | +| LRN-172 | 2026-09-27 | TF-IDF cosine on a 2-doc corpus is identically 0: similarity self-tests need N ≥ 4, a same-corpus positive control and a sensitivity re-run | fixtures for any corpus-normalised statistic (idf, z-score, ranking), "distinct pair passes" tests | +| LRN-173 | 2026-09-27 | contract oracles written from memory failed twice: run the CHECK on the precedent files first, census greps via `git grep` (tracked only), `EVIDENCE: pending` mandatory for gates.sh | contract CHECK lines, precedent-mirroring criteria, gates.sh ledgers | --- @@ -1604,3 +1606,12 @@ Rule: when editing a doctrine file under structure locks, grep the test's lock s - **Pattern**: oracles + test assertions portable across main shell / sub-agent shells pin `/usr/bin/grep` or `command grep`; avoid `$VAR` literals mid-pattern (`-F` or `--`). - **Where applicable**: contract `CHECK:` lines run by executors/verifiers; hermetic test greps; any hook riding on grep output. - **Reference**: contract `2026-09-25-default-profile-full-1254` criterion 12; verifier + executor reports 2026-09-25. Links [[LRN-074]] [[BDR-101]]. + +## LRN-172 — TF-IDF cosine on a 2-doc corpus is identically 0: similarity self-tests need N ≥ 4, a same-corpus positive control and a sensitivity re-run +- **Context**: A3 executor wrote a "distinct pair passes" fixture as a bare 2-doc corpus, documented the 0.00 as "the point being proven". Fresh verifier proved a near-duplicate pair also scores 0.00 at N=2: idf = log(N/df) = 0 for shared terms, unique terms never meet. Executor self-report "all markers printed" was true; the test was vacuous anyway. +- **Fix shape**: one N=4 corpus: near-dup control 0.90 → FAIL, distinct pair 0.00 → silent, then swap one doc for a near-copy → 0.62 WARN appears. Marker printed only when all three hold. +- **Apply**: any fixture for a corpus-normalised statistic asserts both directions in one corpus; "passes on a trivial corpus" proves nothing; markers prove the oracle, not the intent → keep the blind verifier ([[BDR-102]] [[EVAL-032]]). + +## LRN-173 — contract oracles written from memory failed twice: run the CHECK on the precedent files first, census greps via `git grep`, `EVIDENCE: pending` mandatory +- **Context**: same run, two orchestrator oracle bugs. (1) "≤ 80 chars" over the whole file: rules/web-building.md line 2 (`paths:` frontmatter) is already 110 chars, so the criterion contradicted the precedent it named; executor returned NEED-DECISION instead of bending. (2) emil-citers census with `grep -rl` hit gitignored `install-*.log` at the repo root; `git grep -l` (tracked only) is the right census tool. (3) gates.sh `run` errors `runnable but has no EVIDENCE: line` unless each criterion carries `EVIDENCE: pending`. +- **Apply**: before shipping a CHECK, run it against the files it claims to mirror; census oracles = `git grep`; contract skeleton carries `EVIDENCE: pending` per criterion (check /feat's template writes it). Oracle fixes are orchestrator-owned, never a re-dispatch ([[BDR-102]]). diff --git a/.claude/tasks/TODO.md b/.claude/tasks/TODO.md index 9da1211..0351b3a 100644 --- a/.claude/tasks/TODO.md +++ b/.claude/tasks/TODO.md @@ -1,5 +1,30 @@ # TODO +## 2026-09-27 — case 2 of the 6-repo review: borrow from agent-skills (feature/agent-skills-borrow) +User go "ok pour les 4" after the analysis: plugin rejected (1.8k tok/session for +20 % novelty, /spec /review /ship collide with gstack, trunk-based git and the +one-version API rule contradict the doctrine, second router). Four independent +chantiers, one contract each under `.claude/tasks/contracts/2026-09-27-*`, +dispatched to feater executors; gates replayed by the orchestrator (gates.sh → +fresh verifier → fresh security-auditor). Case 1 lives on feature/yagni-ladder. +- [x] A1 d28c45e vendor observability-and-instrumentation, deprecation-and-migration, + ci-cd-and-automation (emil precedent, pinned commit 2686b620) — agent-skills-vendor +- [x] A2 2b25cb4 `lib/floor-guard.sh` diff-scoped bar-weakening detector + verifier step + + suite — floor-guard +- [x] A3 409db51 `lib/tests/skill-routing-census.test.sh` description-collision census + (measured: 120 skills, max 0.52 careful~guard, 0 >= 0.75) — skill-routing-census +- [x] A4 1a8e6de `rules/rest-api.md` path-scoped rule from api-and-interface-design, + one-version rule dropped — rest-api-rule +- [x] A5 gates MET ×4, verifiers CONFORME ×4 (2 re-dispatches), security PASS ×2, + make test 35 suites green minus 2 pre-existing T16a, shellcheck clean; + BDR-102 LRN-172 LRN-173 EVAL-032. UNMERGED — human gate. +Follow-up (not started): per-skill positive/negative prompt ranking (Tier 2 +second half); `make link` after merge to symlink the 3 skills (user runs it); +floor-guard waiver policy: user chose strict (CLARIFICATIONS ack outside +test files, else gap) → applied in agents/verifier.md STEP 3 + loop doc; +frame floor-guard snippets as data in the verifier step (LOW); `/tmp/tmp.AAyJzvufO6` +scratch dir from an executor proof, `rm -rf` refused → user removes; contract +skeleton must carry `EVIDENCE: pending` (LRN-173, check /feat's template). ## 2026-09-27 — YAGNI ladder + shortcut marker in doctrine (feature/yagni-ladder) Case 1 of the 6-repo review (ponytail, chisle). Both rejected as plugins: per-turn and per-subagent injection, prose rules colliding with writing-style.md, caveman diff --git a/.claude/tasks/contracts/2026-09-27-agent-skills-vendor-1525.md b/.claude/tasks/contracts/2026-09-27-agent-skills-vendor-1525.md new file mode 100644 index 0000000..b3294a9 --- /dev/null +++ b/.claude/tasks/contracts/2026-09-27-agent-skills-vendor-1525.md @@ -0,0 +1,62 @@ +# CONTRACT — agent-skills-vendor +- date: 2026-09-27 | flow: feat (ad-hoc dispatch, /feat gates replayed by the orchestrator) | branch: feature/agent-skills-borrow +- status: active + +## REQUEST (verbatim — IMMUTABLE) +> Vendor three skills from addyosmani/agent-skills as machine-owned copies, the emil-design-eng way: `observability-and-instrumentation`, `deprecation-and-migration`, `ci-cd-and-automation`. Source pinned to commit `2686b620fc1fed2e8f60c704839c766b8594c6b6` (main, 2026-09-26) in `plugins.lock.json`; the scripts read the pin from the lock, never hardcode it. Each lands in `skills-external//SKILL.md` (gitignored, curl'd by install-plugins.sh, refreshed by update-all.sh at the pinned commit), symlinked by link.sh, registered wherever emil-design-eng is registered when the semantics apply (toggle-external registry, profiles that carry the dev skills, tests that enumerate externals, CHANGELOG). User go 2026-09-27 ("ok pour les 4", case 2 of the 6-repo review). + +## CLARIFICATIONS +- Profiles: add the three to `full.profile` and to every profile that lists `bugfix` (dev-class); never to design/web profiles. +- Not mirrored on purpose (design-only citers): lib/design-gate.md, lib/tests/fixtures/registry-index-drift.md, lib/profiles/{design,web,web-full}.profile, agents/plugin-advisor.md, agents/plugin-probe.md, CLAUDE.global.md. +- Upstream SKILL.md copied byte-for-byte: no edits, no rewrite of internal links (dangling cross-skill mentions accepted). +- The executor materializes the three files with the same curl the install step uses (network read allowed); it never runs `link.sh`, `make link`, `make plugin` or `update-all.sh` (they touch `~/.claude`). +- Lock entry shape: `"agent-skills": {"source": "https://github.com/addyosmani/agent-skills", "commit": "", "skills": [...], "managed_by": "curl", "note": "..."}`. A helper reading it may follow the `pinned_version` pattern of install-plugins.sh. +- The three raw URLs: `https://raw.githubusercontent.com/addyosmani/agent-skills//skills//SKILL.md` (no extra files under these three skills upstream). + +## ACCEPTANCE CRITERIA +1. Three vendored files present, frontmatter name = dir name. + CHECK: ok=1; for s in observability-and-instrumentation deprecation-and-migration ci-cd-and-automation; do f="skills-external/$s/SKILL.md"; [ -f "$f" ] && grep -q "^name: $s\$" "$f" || { echo "bad $s"; ok=0; }; done; [ "$ok" -eq 1 ] && echo VENDORED + EXPECT: VENDORED + EVIDENCE: MET exit=0 marker-found :: VENDORED +2. Pin recorded in the lock with the three names. + CHECK: python3 -c 'import json; d=json.load(open("plugins.lock.json"))["agent-skills"]; assert d["commit"]=="2686b620fc1fed2e8f60c704839c766b8594c6b6", d; assert set(d["skills"])=={"observability-and-instrumentation","deprecation-and-migration","ci-cd-and-automation"}, d; print("PINNED")' + EXPECT: PINNED + EVIDENCE: MET exit=0 marker-found :: PINNED +3. Install and update steps exist and are lock-driven (no hardcoded sha). + CHECK: grep -q "agent-skills" install-plugins.sh && grep -q "agent-skills" update-all.sh && ! grep -q "2686b620" install-plugins.sh update-all.sh link.sh lib/toggle-external.sh && echo LOCK_DRIVEN + EXPECT: LOCK_DRIVEN + EVIDENCE: MET exit=0 marker-found :: LOCK_DRIVEN +4. Both the copy and the symlink path are gitignored. + CHECK: ok=1; for s in observability-and-instrumentation deprecation-and-migration ci-cd-and-automation; do git check-ignore -q "skills-external/$s" && git check-ignore -q "skills/$s" || { echo "not ignored $s"; ok=0; }; done; [ "$ok" -eq 1 ] && echo IGNORED + EXPECT: IGNORED + EVIDENCE: MET exit=0 marker-found :: IGNORED +5. link.sh symlinks them (EXTERNAL_SKILLS list). + CHECK: ok=1; for s in observability-and-instrumentation deprecation-and-migration ci-cd-and-automation; do grep -q "$s" link.sh || { echo "not linked $s"; ok=0; }; done; [ "$ok" -eq 1 ] && echo LINK_LISTED + EXPECT: LINK_LISTED + EVIDENCE: MET exit=0 marker-found :: LINK_LISTED +6. Emil citers census mirrored (tracked files only: the ignored install-*.log files at the root also name emil), except the design-only allowlist. + CHECK: allow=" lib/design-gate.md lib/tests/fixtures/registry-index-drift.md lib/profiles/design.profile lib/profiles/web.profile lib/profiles/web-full.profile agents/plugin-advisor.md agents/plugin-probe.md CLAUDE.global.md "; miss=0; for f in $(git grep -l emil-design-eng -- . ':!skills-external' ':!skills' ':!.claude'); do grep -q observability-and-instrumentation "$f" && continue; case "$allow" in *" $f "*) ;; *) echo "not mirrored: $f"; miss=1 ;; esac; done; [ "$miss" -eq 0 ] && echo CENSUS_MIRRORED + EXPECT: CENSUS_MIRRORED + EVIDENCE: MET exit=0 marker-found :: CENSUS_MIRRORED +7. Profile, toggle-external and doctrine-citers suites green. + CHECK: out=$(make test suite="lib/tests/profile-default.test.sh lib/tests/profile-set-managed.test.sh lib/tests/toggle-external-repo-resolution.test.sh lib/tests/doctrine-citers.test.sh" 2>&1); echo "$out" | grep -qE "FAIL=[1-9]" && { echo "$out" | tail -20; exit 1; }; echo SUITES_GREEN + EXPECT: SUITES_GREEN + EVIDENCE: MET exit=0 marker-found :: SUITES_GREEN +8. shellcheck clean on the touched scripts. + CHECK: shellcheck install-plugins.sh update-all.sh link.sh lib/toggle-external.sh lib/profile.sh && echo SHELLCHECK_OK + EXPECT: SHELLCHECK_OK + EVIDENCE: MET exit=0 marker-found :: SHELLCHECK_OK + +## FILE SCOPE +- plugins.lock.json, install-plugins.sh, update-all.sh, link.sh, .gitignore +- lib/toggle-external.sh, lib/profile.sh (only if the externals registry lives there), lib/profiles/full.profile + dev-class profiles +- lib/tests/profile-default.test.sh, lib/tests/profile-set-managed.test.sh, lib/tests/toggle-external-repo-resolution.test.sh (only if they enumerate externals) +- CHANGELOG.md (Unreleased entry) +- skills-external//SKILL.md x3 (materialized, gitignored) + +## PLAN +1. `grep -rn emil-design-eng` census (files listed in criterion 6) → mirror file by file, same comment density. +2. Lock entry; install-plugins.sh new step next to Step 8 ("agent-skills — 3 skills, pinned commit") reading the sha from the lock (python3/jq like the existing helpers), curl each raw URL into skills-external//SKILL.md, skip when present, `err` with the manual command on failure; update-all.sh step re-curls at the pinned sha (tmp + mv, emil precedent). +3. link.sh EXTERNAL_SKILLS += 3; .gitignore: `skills/` in the symlink allowlist + `skills-external//` with a short comment naming the source (emil precedent). +4. toggle-external registry + profiles + tests that enumerate externals. +5. Materialize the three files with the curl; run criteria 1-8; report the emil citers you deliberately did not mirror and why. diff --git a/.claude/tasks/contracts/2026-09-27-floor-guard-1525.md b/.claude/tasks/contracts/2026-09-27-floor-guard-1525.md new file mode 100644 index 0000000..38113c5 --- /dev/null +++ b/.claude/tasks/contracts/2026-09-27-floor-guard-1525.md @@ -0,0 +1,49 @@ +# CONTRACT — floor-guard +- date: 2026-09-27 | flow: feat (ad-hoc dispatch, /feat gates replayed by the orchestrator) | branch: feature/agent-skills-borrow +- status: active + +## REQUEST (verbatim — IMMUTABLE) +> Build `lib/floor-guard.sh`, a diff-scoped deterministic detector of a quietly weakened quality bar, adapted from addyosmani/agent-skills `constraint-driven-development` (floor guard) to this repo's gate model. Usage `bash ~/.claude/lib/floor-guard.sh [-- ...]` over `git diff ` (working tree included). Kinds: SUPPRESS (added checker silencing: `@ts-ignore`, `@ts-expect-error` without a trailing reason, `eslint-disable*`, `# noqa`, `# type: ignore`, `nosemgrep`, `nosec`, `shellcheck disable`), SKIP (added `.skip(`, `.only(`, `xit(`, `xdescribe(`, `fit(`, `fdescribe(`, `it.todo(`, `@pytest.mark.skip`, `@unittest.skip`, `t.Skip(` in test files), DELETED_TEST (deleted file whose path matches a test pattern), ASSERT_DROP (a test file whose assertion-line count decreases: `expect(`, `assert`, `should`, `.toBe`), STUB (added `not implemented`, `NotImplementedError`, empty `catch` block, `except: pass`), THRESHOLD_DOWN (a numeric value decreased on the same key in coverage/quality config files: `jest.config*`, `vitest.config*`, `.nycrc*`, `codecov*`, `sonar-project.properties`, `lighthouserc*`, `CONSTRAINTS.md`). Waiver: an added line carrying `floor-guard: allow ` is printed as WAIVED and not counted. Output: one `FLOOR : ` per finding, then `FLOOR GUARD: clean` (rc 0) or `FLOOR GUARD: finding(s), waived` (rc 2); rc 3 on usage error. Wire it as a mandatory verifier step (agents/verifier.md) and document it in lib/verify-secure-loop.md GATE 1; hermetic suite `lib/tests/floor-guard.test.sh`. User go 2026-09-27 ("ok pour les 4", case 2 item 2). + +## CLARIFICATIONS +- Language: bash entry point; the diff parsing may live in an embedded python3 heredoc (precedent `lib/tests/run-review-guards.sh`). Functions <= 25 logic lines, 80-char lines. +- File classes: test file = path contains `test`, `spec`, `__tests__`, or matches `*.test.*`, `*.spec.*`, `*_test.go`, `*_test.py`, `test_*.py`; config file = the names listed under THRESHOLD_DOWN. SKIP and ASSERT_DROP apply to test files only; SUPPRESS and STUB to any file; THRESHOLD_DOWN to config files only. +- The guard's own pattern table contains the trigger strings: those source lines carry `# floor-guard: allow pattern table` so the guard stays clean on itself (this also exercises the waiver path for real). +- Verifier step: `bash ~/.claude/lib/floor-guard.sh ` where base = the branch's gitflow base (develop; main for hotfix/release); rc 2 → verdict ECARTS listing each FLOOR line, unless the contract's CLARIFICATIONS explicitly authorize that exact weakening (quote it in the verdict). +- Suite: throwaway repos (`make test` exports GIT_CONFIG_GLOBAL=/dev/null); one RED fixture per kind, one WAIVED fixture, one CLEAN fixture, each printed as `PASS `; summary `PASS=n FAIL=m` like the other suites. +- Self-detection: the suite file itself contains the trigger strings; the self-run criterion excludes it by pathspec. +- Out of scope: a pre-commit hook, running it on the repo's own diff in `make test`, parsers beyond the regexes above. + +## ACCEPTANCE CRITERIA +1. Suite green, every kind flip-tested. + CHECK: out=$(make test suite=lib/tests/floor-guard.test.sh 2>&1); for k in SUPPRESS SKIP DELETED_TEST ASSERT_DROP STUB THRESHOLD_DOWN WAIVED CLEAN; do echo "$out" | grep -q "PASS $k" || { echo "missing PASS $k"; echo "$out" | tail -15; exit 1; }; done; echo "$out" | grep -qE "FAIL=[1-9]" && exit 1; echo KINDS_GREEN + EXPECT: KINDS_GREEN + EVIDENCE: MET exit=0 marker-found :: KINDS_GREEN +2. Usage error is rc 3. + CHECK: bash lib/floor-guard.sh >/dev/null 2>&1; [ $? -eq 3 ] && echo RC_USAGE + EXPECT: RC_USAGE + EVIDENCE: MET exit=0 marker-found :: RC_USAGE +3. Self-run clean on this branch (suite file excluded by pathspec), waivers visible. + CHECK: out=$(bash lib/floor-guard.sh develop -- . ':!lib/tests/floor-guard.test.sh' 2>&1); rc=$?; echo "$out" | tail -3; [ $rc -eq 0 ] && echo "$out" | grep -q WAIVED && echo SELF_CLEAN + EXPECT: SELF_CLEAN + EVIDENCE: MET exit=0 marker-found :: WAIVED STUB lib/floor-guard.sh:105 'not implemented', # floor-guard: allow pattern table WAIVED STUB lib/floor-guard.sh:106 'NotImplementedE… +4. Verifier wired, loop documented. + CHECK: grep -q "floor-guard.sh" agents/verifier.md && grep -q "floor-guard" lib/verify-secure-loop.md && echo WIRED + EXPECT: WIRED + EVIDENCE: MET exit=0 marker-found :: WIRED +5. shellcheck and doctrine-citers clean. + CHECK: shellcheck lib/floor-guard.sh lib/tests/floor-guard.test.sh && out=$(make test suite=lib/tests/doctrine-citers.test.sh 2>&1) && ! echo "$out" | grep -qE "FAIL=[1-9]" && echo LINT_OK + EXPECT: LINT_OK + EVIDENCE: MET exit=0 marker-found :: LINT_OK + +## FILE SCOPE +- lib/floor-guard.sh (new), lib/tests/floor-guard.test.sh (new) +- agents/verifier.md (one mandatory step), lib/verify-secure-loop.md (one paragraph under GATE 1) +- CHANGELOG.md (Unreleased entry) + +## PLAN +1. Script: arg parsing (base, optional `--` pathspec), `git diff --unified=0 -- ` plus `git diff --diff-filter=D --name-only -- `, rc contract, header comment stating WHY (BDR-100 class: deterministic floor under an LLM gate). +2. Python parser on stdin: walk hunks, classify added lines by kind and file class, count assertion lines removed vs added per test file, compare numeric values on identical keys in config files (`key: 80` → `key: 60`, JSON or YAML-ish), detect waivers. +3. Output lines + summary + rc. +4. Suite: helper `mk_repo` (git init -q, base commit with a test file holding 3 assertions, a `vitest.config.ts` with `lines: 80`, a source file), one fixture per kind → assert rc 2 and the FLOOR line; WAIVED → rc 0 and a WAIVED line; CLEAN → rc 0. +5. verifier.md step + verify-secure-loop.md paragraph + CHANGELOG; run criteria 1-5. diff --git a/.claude/tasks/contracts/2026-09-27-rest-api-rule-1525.md b/.claude/tasks/contracts/2026-09-27-rest-api-rule-1525.md new file mode 100644 index 0000000..050c6e0 --- /dev/null +++ b/.claude/tasks/contracts/2026-09-27-rest-api-rule-1525.md @@ -0,0 +1,34 @@ +# CONTRACT — rest-api-rule +- date: 2026-09-27 | flow: feat (ad-hoc dispatch, /feat gates replayed by the orchestrator) | branch: feature/agent-skills-borrow +- status: active + +## REQUEST (verbatim — IMMUTABLE) +> Write `rules/rest-api.md`, a path-scoped user rule distilled from addyosmani/agent-skills `api-and-interface-design` (commit 2686b620fc1fed2e8f60c704839c766b8594c6b6), the way `rules/web-building.md` is written: `paths:` frontmatter, English, <= 45 lines, 80-char lines. Keep: contract-first order (typed interface → schemas with server-generated fields apart → error codes → validate at boundaries only); one error envelope `{ error: { code, message, details? } }` with the HTTP map 400 invalid / 401 auth / 403 forbidden / 404 missing / 409 conflict / 422 semantic / 500 server; every list endpoint paginated (`page`, `pageSize`, `totalItems`, `totalPages`), filters as query params; idempotency (key derived from intent, atomic claim via unique constraint, payload guard, explicit in-flight duplicate policy 409 / wait / 202, retention beyond the longest retry path incl. dead-letter); naming (plural nouns, camelCase params and fields, UPPER_SNAKE enums, is/has/can booleans); one Hyrum's law line. Drop the upstream one-version rule: versioning points to the CLAUDE.md heading "Web APIs — always versioned". User go 2026-09-27 ("ok pour les 4", case 2 item 4). + +## CLARIFICATIONS +- Globs: `["**/api/**", "**/routes/**", "**/controllers/**", "**/*.route.*", "**/*.controller.*", "**/openapi.*", "**/*.openapi.*"]`. +- Fetch the upstream text with curl at the pinned commit to distill from; never vendor it. +- Cite the doctrine as `CLAUDE.md § Web APIs — always versioned` (exact heading, so the doctrine-citers census resolves it). +- No routing line in CLAUDE.global.md, no README change, rules/README.md unchanged; CHANGELOG Unreleased entry. + +## ACCEPTANCE CRITERIA +1. Shape: frontmatter, paths JSON list, <= 45 lines, body lines <= 80 chars (the one-line `paths:` frontmatter is exempt, repo precedent: web-building.md / web-security.md line 2). + CHECK: f=rules/rest-api.md; [ -f "$f" ] && [ "$(head -1 "$f")" = "---" ] && python3 -c 'import re,json; s=open("rules/rest-api.md").read(); m=re.search(r"^paths: (.*)$",s,re.M); assert len(json.loads(m.group(1)))>=5' && [ "$(wc -l < "$f")" -le 45 ] && ! awk 'NR>3 && length>80' "$f" | grep -q . && echo RULE_SHAPE + EXPECT: RULE_SHAPE + EVIDENCE: MET exit=0 marker-found :: RULE_SHAPE +2. Content present, one-version rule absent. + CHECK: f=rules/rest-api.md; ok=1; for k in "code" "message" "409" "422" "pageSize" "totalItems" "idempoten" "unique" "camelCase" "UPPER_SNAKE" "Hyrum" "always versioned"; do grep -qi -- "$k" "$f" || { echo "missing $k"; ok=0; }; done; grep -qi "extend rather than fork" "$f" && { echo "one-version rule present"; ok=0; }; [ "$ok" -eq 1 ] && echo RULE_CONTENT + EXPECT: RULE_CONTENT + EVIDENCE: MET exit=0 marker-found :: RULE_CONTENT +3. Doctrine citation resolves. + CHECK: out=$(make test suite=lib/tests/doctrine-citers.test.sh 2>&1); echo "$out" | grep -qE "FAIL=[1-9]" && { echo "$out" | tail -10; exit 1; }; echo CITERS_OK + EXPECT: CITERS_OK + EVIDENCE: MET exit=0 marker-found :: CITERS_OK + +## FILE SCOPE +- rules/rest-api.md (new), CHANGELOG.md (Unreleased entry) + +## PLAN +1. curl the upstream SKILL.md at the pin into the scratch dir; read it. +2. Write the rule: title "REST API — contract, errors, lists, idempotency"; sections Contract first · Errors · Lists · Idempotency · Naming · Versioning (one line pointing to the doctrine heading). +3. CHANGELOG line; run criteria 1-3. diff --git a/.claude/tasks/contracts/2026-09-27-skill-routing-census-1525.md b/.claude/tasks/contracts/2026-09-27-skill-routing-census-1525.md new file mode 100644 index 0000000..16ea968 --- /dev/null +++ b/.claude/tasks/contracts/2026-09-27-skill-routing-census-1525.md @@ -0,0 +1,42 @@ +# CONTRACT — skill-routing-census +- date: 2026-09-27 | flow: feat (ad-hoc dispatch, /feat gates replayed by the orchestrator) | branch: feature/agent-skills-borrow +- status: active + +## REQUEST (verbatim — IMMUTABLE) +> Build `lib/tests/skill-routing-census.test.sh`: a deterministic census of skill-description collisions across the live catalog, adapted from addyosmani/agent-skills evals Tier 2. Catalog = every `SKILL.md` under `~/.claude/skills/*/` (symlinks resolved) plus plugin skills under `~/.claude/plugins/cache/*/*/*/skills/*/SKILL.md` and `~/.claude/plugins/cache/*/*/*/.claude/skills/*/SKILL.md` (roots overridable via `SKILL_ROUTING_ROOTS`, colon-separated dirs, for fixtures). Extract `description` (scalar or `|`/`>` block), tokenize (lowercase, `[a-z][a-z0-9-]+`, stopwords, suffix stemming s/es/ed/ing), TF-IDF cosine over all pairs. Print `skills with description: N`, the top 10 pairs as `0.52 a ~ b`, WARN lines for pairs >= `SKILL_ROUTING_WARN` (default 0.50), FAIL for pairs >= `SKILL_ROUTING_FAIL` (default 0.75). Self-test on fixtures: a near-duplicate pair must FAIL (`FIXTURE_COLLISION_DETECTED`), a distinct pair must pass (`FIXTURE_DISTINCT_OK`). Measured today: 120 skills, max 0.52 (careful ~ guard), 0 pairs >= 0.75 → green with one WARN. User go 2026-09-27 ("ok pour les 4", case 2 item 3). + +## CLARIFICATIONS +- python3 embedded in the bash suite is allowed (precedent run-review-guards.sh). If the python body exceeds ~120 lines, put it in `lib/skill-routing-census.py` and keep the suite as the wrapper. +- Reference implementation (read it, reuse the logic, harden the block-description parsing): `/tmp/claude-1000/-home-bchanot-Documents-claude/977f1703-f01d-497a-b794-5b69fafcd35f/scratchpad/census.py`. +- Thresholds stay at the upstream defaults; no allowlist file; a WARN never fails the suite. +- Dedup by skill directory name: first path wins. +- Out of scope (follow-up): positive/negative prompt ranking per skill. +- The live census is machine-dependent by design (it audits this machine's catalog); the fixture self-test is the hermetic part. On a machine with an empty catalog the live pass prints `skills with description: 0` and passes. + +## ACCEPTANCE CRITERIA +1. Suite green on the live catalog, catalog non-trivial here. + CHECK: out=$(make test suite=lib/tests/skill-routing-census.test.sh 2>&1); echo "$out" | grep -qE "FAIL=[1-9]" && { echo "$out" | tail -15; exit 1; }; echo "$out" | grep -qE "skills with description: *[0-9]{2,}" && echo LIVE_GREEN + EXPECT: LIVE_GREEN + EVIDENCE: MET exit=0 marker-found :: LIVE_GREEN +2. Fixture flip: collision detected, distinct pair passes. + CHECK: out=$(make test suite=lib/tests/skill-routing-census.test.sh 2>&1); echo "$out" | grep -q "FIXTURE_COLLISION_DETECTED" && echo "$out" | grep -q "FIXTURE_DISTINCT_OK" && echo FLIP_TESTED + EXPECT: FLIP_TESTED + EVIDENCE: MET exit=0 marker-found :: FLIP_TESTED +3. Report shape: top pairs with two-decimal scores. + CHECK: out=$(make test suite=lib/tests/skill-routing-census.test.sh 2>&1); echo "$out" | grep -qE "^0\.[0-9]{2} [a-z0-9:_-]+ ~ [a-z0-9:_-]+" && echo REPORT_SHAPE + EXPECT: REPORT_SHAPE + EVIDENCE: MET exit=0 marker-found :: REPORT_SHAPE +4. shellcheck clean. + CHECK: shellcheck lib/tests/skill-routing-census.test.sh && echo SHELLCHECK_OK + EXPECT: SHELLCHECK_OK + EVIDENCE: MET exit=0 marker-found :: SHELLCHECK_OK + +## FILE SCOPE +- lib/tests/skill-routing-census.test.sh (new); optional lib/skill-routing-census.py (new) +- CHANGELOG.md (Unreleased entry) + +## PLAN +1. Roots: default globs; `SKILL_ROUTING_ROOTS` override; dedup by dir name. +2. Description extraction: scalar `description: text`; block `description: |` or `>` → join the indented continuation lines. +3. Tokens / TF-IDF / cosine as in census.py: stopword set, stem(), (1+log tf)·log(N/df), L2 norm, cosine over combinations. +4. Report + thresholds + rc. Suite: run live (a FAIL line → suite RED), then two fixture dirs under mktemp: A/B near-duplicate descriptions → expect a FAIL line → print FIXTURE_COLLISION_DETECTED; C/D distinct → expect rc 0 → print FIXTURE_DISTINCT_OK. `PASS=n FAIL=m` summary like the other suites. diff --git a/.gitignore b/.gitignore index 4c0dbdd..a8b9bf3 100644 --- a/.gitignore +++ b/.gitignore @@ -65,6 +65,9 @@ skills/ios-sync skills/design-motion-principles skills/emil-design-eng skills/frontend-design +skills/ci-cd-and-automation +skills/deprecation-and-migration +skills/observability-and-instrumentation # Impeccable — NOT a symlink: `impeccable skills install --scope=global` # writes the skill dir (and its ~15 MB engine binary) straight in through the @@ -177,6 +180,15 @@ skills-external/frontend-design/ # an edit. The source is always re-fetched, so no offline copy is needed. skills-external/emil-design-eng/ +# Agent Skills trio (addyosmani/agent-skills) — machine-owned, curl'd at the +# commit pinned in plugins.lock.json ("agent-skills" entry) by +# install-plugins.sh Step 8e (when absent) and re-fetched at the SAME commit +# by update-all.sh. Not vendored: this is a pin, not a tracked snapshot — +# bump the commit deliberately to pick up an upstream edit. +skills-external/observability-and-instrumentation/ +skills-external/deprecation-and-migration/ +skills-external/ci-cd-and-automation/ + # 21st.dev skill pack — machine-owned: `21st skills install` output, staged by # install-plugins.sh Step 8.7 (the installer refuses to write through the # ~/.claude/skills symlink, so it runs under a throwaway HOME and the skills diff --git a/CHANGELOG.md b/CHANGELOG.md index 397c391..ff20dfa 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -7,6 +7,38 @@ Format follows [Keep a Changelog](https://keepachangelog.com/). ## [Unreleased] ### Added +- **Agent Skills trio** (`observability-and-instrumentation`, + `deprecation-and-migration`, `ci-cd-and-automation`) — vendored from + addyosmani/agent-skills at a pinned commit (`agent-skills` entry in + plugins.lock.json), the emil-design-eng way: curl'd into + `skills-external//SKILL.md` by install-plugins.sh Step 8e, refreshed + by update-all.sh at the same commit, symlinked by link.sh, registered in + `lib/toggle-external.sh`, `lib/profile.sh` and the `full`/`backend`/`dev` + profiles. Case 2 of the 6-repo review: the plugin itself was rejected + (1.8k tokens per session for 20 % novelty, `/spec` `/review` `/ship` + collide with gstack, trunk-based git and the one-version API rule + contradict the doctrine, a second skill router). +- **`lib/floor-guard.sh`** — diff-scoped deterministic detector of a quietly + weakened quality bar (new lint/type suppressions, skipped or deleted + tests, dropped assertions, stubs, lowered coverage thresholds), with a + `floor-guard: allow ` waiver, rc 0/2/3. Mandatory verifier + STEP 3 (`agents/verifier.md`), documented under GATE 1 of + `lib/verify-secure-loop.md`. Suite `lib/tests/floor-guard.test.sh`: 6 + kinds plus a WAIVED and a CLEAN fixture, each flip-tested. Waivers + outside test files count as gaps unless the contract's CLARIFICATIONS + names them (security-gate MEDIUM, user chose strict). Adapted from + agent-skills `constraint-driven-development`. +- **`lib/tests/skill-routing-census.test.sh`** (+ `lib/skill-routing-census.py`) + — TF-IDF cosine census of skill-description collisions across the live + catalog (routing ambiguity, not naming): top 10 pairs, WARN ≥ 0.50, + FAIL ≥ 0.75, fixture flip-test. Baseline 2026-09-27: 120 skills, max 0.52 + (`careful` ~ `guard`). Adapted from agent-skills evals Tier 2. +- **`rules/rest-api.md`** — path-scoped REST rule distilled from agent-skills + `api-and-interface-design`: contract-first order, one error envelope + + HTTP map, paginated lists, idempotency (key from intent, atomic claim, + payload guard, duplicate policy, retention), naming, Hyrum's law; + versioning points to `CLAUDE.md § Web APIs — always versioned` instead of + the upstream one-version rule. - **`make test suite=`** runs one suite hermetically; the `GIT_CONFIG_GLOBAL=/dev/null` export lives in the Makefile so nobody types the denied env-prefix form by hand (the reason an executor wrote a wrapper diff --git a/agents/verifier.md b/agents/verifier.md index 4e902f1..c9cdc2d 100644 --- a/agents/verifier.md +++ b/agents/verifier.md @@ -71,7 +71,33 @@ You may re-run a `CHECK:` yourself to settle a doubt (Bash is read-only, and these commands are observation). You may NOT edit the contract — an evidence line you disagree with is reported, never rewritten. -## STEP 3 — SCOPE CHECK +## STEP 3 — FLOOR GUARD (mandatory, deterministic) + +Run the floor guard over the diff before rendering any verdict — a red or +skipped run here is a structural gap, never a judgment call: + +```bash +bash ~/.claude/lib/floor-guard.sh -- ... +``` + +`` = the branch's gitflow base (develop; main for a hotfix/release). +Parse the single `FLOOR GUARD:` line: + +- `clean` (rc 0) → no unwaived finding; still apply the WAIVED rule below. +- ` finding(s), waived` (rc 2) → each `FLOOR : + ` line is a gap for STEP 5's `ECARTS` count, UNLESS the + contract's `CLARIFICATIONS` explicitly authorizes that exact weakening — + quote the authorizing sentence in the verdict instead of counting it as a + gap. +- `WAIVED :` lines (either rc): on a test file (path + holds `test`, `spec` or `__tests__`) they are informational. Anywhere + else the waiver is self-service by construction, so it is a gap UNLESS + the contract's `CLARIFICATIONS` names that file and the reason — quote + it. The tool prints, the contract authorizes, the verifier counts. +- rc 3 (usage error) → a structural failure like a missing contract: retry + once (base ref or pathspec likely wrong), a second failure escalates. + +## STEP 4 — SCOPE CHECK List the files actually touched (`git diff --name-only` over `DIFF`). Compare against the contract's `FILE SCOPE`. Report every out-of-scope @@ -79,7 +105,7 @@ file. Disposition is NOT your call: the orchestrator treats each one as a gap — the dev removes it or justifies it, and an accepted justification only enters the contract through a human micro-gate. -## STEP 4 — VERDICT +## STEP 5 — VERDICT Read the contract's `ABANDON:` lines. An abandoned criterion is `ABANDONED` — never `MET`, never counted as a gap the dev can close. @@ -89,11 +115,12 @@ is not: 1. `ERROR()` — the contract is missing or unreadable. 2. `ECARTS(n)` — n = count(NOT-MET) + count(UNVERIFIABLE) + count(out-of-scope - files). Surface any abandonment in the same report. + files) + count(unauthorized FLOOR findings from STEP 3). Surface any + abandonment in the same report. 3. `ABANDONED(n)` — zero gaps remain, but n abandonments stand. This is NOT a pass and NOT a dev loop: it routes straight to the human gate. 4. `CONFORME` — ALL criteria `MET`, zero out-of-scope files, zero - abandonments. + unauthorized FLOOR findings, zero abandonments. ## OUTPUT (exact format — machine-parsed by the orchestrator) @@ -106,6 +133,8 @@ CRITERIA: 3. — UNVERIFIABLE — 4. — ABANDONED — SCOPE: in-scope files; out-of-scope: +FLOOR: clean | finding(s) ( waived) — PROOF: read files, ran , checked / criteria ``` @@ -123,6 +152,9 @@ PROOF: read files, ran , checked / criteria - `PROOF` is MANDATORY. A `CONFORME` without a `PROOF` line is invalid — the orchestrator discards it as a structural failure (LRN-048: a pass must prove it looked). +- STEP 3's floor-guard run is MANDATORY, every dispatch. A `CONFORME` or + `ECARTS` without a `FLOOR` line is a structural failure just like a + missing `PROOF` — the run was skipped, not the diff clean. - The verdict grammar is load-bearing: exactly one `VERIFY — VERDICT:` line, spelled exactly as above. @@ -145,8 +177,8 @@ loop, never here): lifts the abandonment (the criterion was fixable after all) or accepts the partial delivery; the run is never reported as fully complete. - Structural failure (`ERROR(…)`, missing/duplicated VERDICT line, - unparsable output, agent crash, `CONFORME` without `PROOF`) → retry - ONCE with a fresh verifier; a 2nd structural failure → human - escalation. A mute verifier is NEVER a PASS. + unparsable output, agent crash, `CONFORME` without `PROOF` or without + `FLOOR`) → retry ONCE with a fresh verifier; a 2nd structural failure → + human escalation. A mute verifier is NEVER a PASS. - After a security-gate fix round: re-verify the request FIRST (this agent), THEN re-verify security — in that order. diff --git a/install-plugins.sh b/install-plugins.sh index 59b730d..f126ff8 100644 --- a/install-plugins.sh +++ b/install-plugins.sh @@ -90,6 +90,22 @@ print(v) fi } +# Read a pinned commit sha from plugins.lock.json (agent-skills style entries +# — no "version", a "commit" field instead). Prints the sha, or "" if the +# entry or the field is absent. +# Usage: pinned_commit "agent-skills" → prints the commit sha or "" +pinned_commit() { + local key="$1" + if [ -f "$REPO/plugins.lock.json" ] && command -v python3 &>/dev/null; then + python3 -c " +import json, sys +with open(sys.argv[1]) as f: + d = json.load(f) +print(d.get(sys.argv[2], {}).get('commit', '')) +" "$REPO/plugins.lock.json" "$key" 2>/dev/null || true + fi +} + # ============================================================ # DETECT OS # ============================================================ @@ -899,6 +915,43 @@ else fi echo "" +# ── Step 8e: Agent Skills (addyosmani/agent-skills, pinned commit) ── +# Three dev-lifecycle skills vendored the emil-design-eng way (curl → +# skills-external//SKILL.md, symlinked by link.sh) but COMMIT-pinned +# instead of tracking main: the sha lives in plugins.lock.json ("agent-skills" +# entry), never hardcoded here. +echo "── Step 8e: Agent Skills (addyosmani/agent-skills) ─────────" +echo "" +AGENT_SKILLS_NAMES=(observability-and-instrumentation deprecation-and-migration ci-cd-and-automation) +AGENT_SKILLS_SHA=$(pinned_commit "agent-skills") +if [ -z "$AGENT_SKILLS_SHA" ]; then + err "agent-skills: no commit pinned in plugins.lock.json — add an \"agent-skills\" entry with a \"commit\" field" +else + for _as_skill in "${AGENT_SKILLS_NAMES[@]}"; do + _as_dir="$REPO/skills-external/$_as_skill" + _as_url="https://raw.githubusercontent.com/addyosmani/agent-skills/$AGENT_SKILLS_SHA/skills/$_as_skill/SKILL.md" + mkdir -p "$_as_dir" + if [ -f "$_as_dir/SKILL.md" ]; then + ok "$_as_skill already downloaded" + else + info "Downloading SKILL.md from addyosmani/agent-skills ($_as_skill)..." + if curl -fsSL "$_as_url" -o "$_as_dir/SKILL.md.tmp" \ + && mv "$_as_dir/SKILL.md.tmp" "$_as_dir/SKILL.md"; then + ok "$_as_skill installed" + else + rm -f "$_as_dir/SKILL.md.tmp" + err "$_as_skill download failed — try: curl -fsSL $_as_url -o $_as_dir/SKILL.md" + fi + fi + if [ -L "$HOME/.claude/skills/$_as_skill" ]; then + ok "$_as_skill symlink OK" + else + info "Symlinking $_as_skill — will be created by link.sh" + fi + done +fi +echo "" + # ============================================================ # STEP 8.5 — EXTERNAL SKILLS (npx skills add …) # ============================================================ @@ -1186,6 +1239,7 @@ echo " 🔄 emil-design-eng — UI polish, animations, component craft (c echo " 🔄 frontend-design — distinctive frontend interfaces, anti-AI-slop (anthropic-agent-skills)" echo " 🔄 impeccable — /impeccable design verbs + 45-rule deterministic detector (npx impeccable detect)" echo " 🔄 design-motion-principles — motion/animation design, 3-designer lens (kylezantos)" +echo " 🔄 agent-skills trio — observability-and-instrumentation, deprecation-and-migration, ci-cd-and-automation (curl → symlink, pinned commit)" echo " 🔄 darwin-skill — autonomous skill optimizer (npx skills, ~/.agents/skills/)" echo " 🔄 21st skill pack — 21st.dev CLI skills; design ones follow the profile (full by default), publishing ones on demand (toggle: lib/toggle-external.sh enable 21st)" echo "" @@ -1194,6 +1248,7 @@ echo " GStack skills symlinked individually into ~/.claude/skills/ (→ submodu echo " Emil Design Eng at: ~/.claude/skills/emil-design-eng/ (symlink → skills-external)" echo " Frontend Design at: ~/.claude/skills/frontend-design/ (symlink → skills-external)" echo " Design Motion Principles at: ~/.claude/skills/design-motion-principles/ (symlink → skills-external)" +echo " Agent Skills trio at: ~/.claude/skills/{observability-and-instrumentation,deprecation-and-migration,ci-cd-and-automation}/ (symlink → skills-external)" echo " npx skills at: ~/.agents/skills/ (symlinked into ~/.claude/skills/)" echo "" echo " → Restart Claude Code — plugins load automatically" diff --git a/lib/floor-guard.sh b/lib/floor-guard.sh new file mode 100755 index 0000000..a573fe0 --- /dev/null +++ b/lib/floor-guard.sh @@ -0,0 +1,297 @@ +#!/usr/bin/env bash +# lib/floor-guard.sh — diff-scoped detector of a quietly weakened quality bar. +# +# bash ~/.claude/lib/floor-guard.sh [-- ...] +# +# rc 0 = clean no floor finding in the diff +# 2 = finding(s), waived +# 3 = usage error (missing , or it does not resolve to a commit) +# +# WHY (BDR-100 class): "no weakened check in this diff" is exactly the kind +# of judgment an LLM verifier can miss, or be talked past one line at a time +# — a single added TS-ignore comment, a skipped test, a dropped assertion, a +# coverage threshold shaved by one point. This makes that judgment +# deterministic: grep the diff for the known ways a change quietly lowers +# the bar, same floor doctrine as gates.sh (contract oracles) and +# doctrine-citers.test.sh (citation census) — a mechanism, not a lesson. +# +# Adapted from addyosmani/agent-skills constraint-driven-development's +# "floor guard" to this repo's own gate model: `git diff` instead of a +# staged-diff assumption, wired into agents/verifier.md STEP 3 rather than a +# pre-commit hook. +# +# Scope: `git diff ` — working tree included (uncommitted changes +# count) — restricted to when given. Every ADDED line is +# classified into one of six kinds (full pattern tables below): +# SUPPRESS a checker-silencing comment added, any file +# SKIP a test disabled or isolated, test files only +# DELETED_TEST a whole test file removed +# ASSERT_DROP a test file's assertion-line count went down +# STUB a not-implemented marker added, any file +# THRESHOLD_DOWN a numeric value lowered on the same key, config files only +# +# Waiver: an added line also carrying `floor-guard: allow ` prints as +# WAIVED and does not count toward the finding total or the rc. +set -uo pipefail + +_usage() { + echo "usage: floor-guard.sh [-- ...]" >&2 + exit 3 +} + +[ $# -ge 1 ] || _usage +BASE_REF="$1"; shift +PATHSPEC=() +if [ $# -gt 0 ]; then + [ "$1" = "--" ] || _usage + shift + PATHSPEC=("$@") +fi +git rev-parse --verify -q "${BASE_REF}^{commit}" >/dev/null 2>&1 || _usage + +TMPDIFF="$(mktemp)" || { echo "floor-guard: mktemp failed" >&2; exit 3; } +trap 'rm -f "$TMPDIFF"' EXIT + +git diff --unified=0 "$BASE_REF" -- "${PATHSPEC[@]}" > "$TMPDIFF" 2>/dev/null +FLOOR_DELETED_FILES="$(git diff --diff-filter=D --name-only \ + "$BASE_REF" -- "${PATHSPEC[@]}" 2>/dev/null)" +export FLOOR_DELETED_FILES + +# "working tree included" means brand-new, still-untracked files too: plain +# `git diff ` never shows them (git only diffs what it already tracks), +# so a file added on this branch and never `git add`-ed would be invisible +# to every kind below. --no-index against /dev/null emits the same unified +# format as the tracked diff above (diff --git / +++ b/path / @@ hunks), +# so the parser needs no separate code path for it. +while IFS= read -r f; do + [ -n "$f" ] || continue + git diff --no-index --unified=0 -- /dev/null "$f" >> "$TMPDIFF" 2>/dev/null +done < <(git ls-files --others --exclude-standard -- "${PATHSPEC[@]}" 2>/dev/null) + +python3 - "$TMPDIFF" <<'PY' +import fnmatch +import os +import re +import sys + +# file classes (CLARIFICATIONS): "path contains test/spec/__tests__" is a +# superset of the explicit globs (*.test.*, *.spec.*, *_test.go, *_test.py, +# test_*.py all contain one of these substrings themselves), so one check +# covers all five. +TEST_SUBSTRINGS = ('test', 'spec', '__tests__') +CONFIG_GLOBS = ('jest.config*', 'vitest.config*', '.nycrc*', 'codecov*', + 'sonar-project.properties', 'lighthouserc*', 'CONSTRAINTS.md') + +# ── pattern tables — the trigger strings themselves, waived on this file's +# own diff so the guard stays clean on itself (also exercises the waiver +# path for real) ──────────────────────────────────────────────────────────── +SUPPRESS_SUBSTRINGS = ( + '@ts-ignore', # floor-guard: allow pattern table + 'eslint-disable', # floor-guard: allow pattern table + '# noqa', # floor-guard: allow pattern table + '# type: ignore', # floor-guard: allow pattern table + 'nosemgrep', # floor-guard: allow pattern table + 'nosec', # floor-guard: allow pattern table + 'shellcheck disable', # floor-guard: allow pattern table +) +TS_EXPECT_ERROR = '@ts-expect-error' # floor-guard: allow pattern table + +SKIP_SUBSTRINGS = ( + '.skip(', '.only(', 'xit(', 'xdescribe(', 'fit(', 'fdescribe(', + 'it.todo(', '@pytest.mark.skip', '@unittest.skip', 't.Skip(', +) + +STUB_SUBSTRINGS = ( + 'not implemented', # floor-guard: allow pattern table + 'NotImplementedError', # floor-guard: allow pattern table +) +EMPTY_CATCH_RE = re.compile(r'catch\s*\([^)]*\)\s*\{\s*\}') +BARE_EXCEPT_RE = re.compile(r'except\b[^:\n]*:\s*pass\b') + +ASSERT_SUBSTRINGS = ('expect(', 'assert', 'should', '.toBe') + +KEYVAL_RE = re.compile(r'["\']?([A-Za-z0-9_.\-]+)["\']?\s*[:=]\s*(-?\d+(?:\.\d+)?)') +WAIVER_RE = re.compile(r'floor-guard:\s*allow\s+(\S.*)$') +HUNK_RE = re.compile(r'^@@ -(\d+)(?:,\d+)? \+(\d+)(?:,\d+)? @@') + + +def is_test_file(path): + return any(s in path for s in TEST_SUBSTRINGS) + + +def is_config_file(path): + base = os.path.basename(path) + return any(fnmatch.fnmatch(base, g) for g in CONFIG_GLOBS) + + +def is_waived(text): + return bool(WAIVER_RE.search(text)) + + +def strip_prefix(raw): + if raw == '/dev/null': + return raw + return raw[2:] if raw[:2] in ('a/', 'b/') else raw + + +def _new_file_entry(files): + entry = {'old_path': None, 'new_path': None, 'adds': [], 'dels': [], + 'first_new': None} + files.append(entry) + return entry + + +def parse_diff(lines): # → list of per-file entries (adds/dels + paths) + files, cur = [], None + old_no = new_no = 0 + for raw in lines: + if raw.startswith('diff --git '): + cur = _new_file_entry(files) + elif raw.startswith('--- '): + cur['old_path'] = strip_prefix(raw[4:]) + elif raw.startswith('+++ '): + cur['new_path'] = strip_prefix(raw[4:]) + elif raw.startswith('@@ '): + m = HUNK_RE.match(raw) + if m: + old_no, new_no = int(m.group(1)), int(m.group(2)) + if cur['first_new'] is None: + cur['first_new'] = new_no + elif raw.startswith('+') and not raw.startswith('+++'): + cur['adds'].append((new_no, raw[1:])); new_no += 1 + elif raw.startswith('-') and not raw.startswith('---'): + cur['dels'].append((old_no, raw[1:])); old_no += 1 + return files + + +def effective_path(entry): + if entry['new_path'] not in (None, '/dev/null'): + return entry['new_path'] + return entry['old_path'] + + +def suppress_kind(text): + if TS_EXPECT_ERROR in text: + after = text.split(TS_EXPECT_ERROR, 1)[1].strip() + return None if after else 'SUPPRESS' + return 'SUPPRESS' if any(p in text for p in SUPPRESS_SUBSTRINGS) else None + + +def stub_kind(text): + if any(p in text for p in STUB_SUBSTRINGS): + return 'STUB' + if EMPTY_CATCH_RE.search(text) or BARE_EXCEPT_RE.search(text): + return 'STUB' + return None + + +def skip_kind(text): + return 'SKIP' if any(p in text for p in SKIP_SUBSTRINGS) else None + + +def line_findings(path, lineno, text, test_file): + out, waived = [], is_waived(text) + for kindfn in (suppress_kind, stub_kind): + kind = kindfn(text) + if kind: + out.append((kind, path, lineno, text, waived)) + if test_file: + kind = skip_kind(text) + if kind: + out.append((kind, path, lineno, text, waived)) + return out + + +def _is_assertion(text): + return any(p in text for p in ASSERT_SUBSTRINGS) + + +def assert_drop_finding(path, entry): + added = sum(1 for _, t in entry['adds'] if _is_assertion(t)) + removed = sum(1 for _, t in entry['dels'] if _is_assertion(t)) + if removed <= added: + return None + lineno = entry['adds'][0][0] if entry['adds'] else (entry['first_new'] or 1) + waived = any(is_waived(t) for _, t in entry['adds']) + snippet = 'assertion lines %d -> %d' % (removed, added) + return ('ASSERT_DROP', path, lineno, snippet, waived) + + +def extract_kv(lines): # → {key: (lineno, value, raw text)} last-wins + kv = {} + for lineno, text in lines: + m = KEYVAL_RE.search(text) + if m: + kv[m.group(1)] = (lineno, float(m.group(2)), text) + return kv + + +def threshold_down_findings(path, entry): + removed_kv = extract_kv(entry['dels']) + added_kv = extract_kv(entry['adds']) + out = [] + for key, (lineno, new_val, text) in added_kv.items(): + old = removed_kv.get(key) + if old and new_val < old[1]: + out.append(('THRESHOLD_DOWN', path, lineno, text, is_waived(text))) + return out + + +def classify_file(entry, deleted_paths): + path = effective_path(entry) + if path is None: + return [] # pure rename/mode-change: no --- / +++ header, no content diff + test_file = is_test_file(path) + if path in deleted_paths and test_file: + return [('DELETED_TEST', path, 1, path, False)] + findings = [] + for lineno, text in entry['adds']: + findings += line_findings(path, lineno, text, test_file) + if test_file: + dropped = assert_drop_finding(path, entry) + if dropped: + findings.append(dropped) + if is_config_file(path): + findings += threshold_down_findings(path, entry) + return findings + + +def load_deleted_paths(): + raw = os.environ.get('FLOOR_DELETED_FILES', '') + return {p for p in raw.splitlines() if p} + + +def snippet_of(text): + return text.strip()[:100] + + +def emit(findings): + ordered = sorted(findings, key=lambda f: (f[1], f[2], f[0])) + n_found = n_waived = 0 + for kind, path, lineno, text, waived in ordered: + tag = 'WAIVED' if waived else 'FLOOR' + print('%s %s %s:%d %s' % (tag, kind, path, lineno, snippet_of(text))) + n_waived += 1 if waived else 0 + n_found += 0 if waived else 1 + if n_found: + print('FLOOR GUARD: %d finding(s), %d waived' % (n_found, n_waived)) + return 2 + print('FLOOR GUARD: clean') + return 0 + + +def main(): + with open(sys.argv[1], 'r', errors='replace') as fh: + lines = fh.read().split('\n') + deleted = load_deleted_paths() + findings = [] + for entry in parse_diff(lines): + findings += classify_file(entry, deleted) + return emit(findings) + + +if __name__ == '__main__': + sys.exit(main()) +PY +rc=$? +exit "$rc" diff --git a/lib/profile.sh b/lib/profile.sh index dd831ca..a88292f 100755 --- a/lib/profile.sh +++ b/lib/profile.sh @@ -79,6 +79,9 @@ MANAGED_EXTERNALS=( 21st-ui-review 21st-cli-use 21st-ai + observability-and-instrumentation + deprecation-and-migration + ci-cd-and-automation ) # MCP servers that are toggle-managed by `set`, both ways (enable AND @@ -756,7 +759,9 @@ NOTE: "set" toggles the MANAGED items automatically, both ways: plugins (ui-ux-pro-max, plugin-dev, pr-review-toolkit), external packs (emil-design-eng, frontend-design, design-motion-principles, impeccable, - the five 21st design skills). Anything outside those allowlists stays + the five 21st design skills, the agent-skills trio + observability-and-instrumentation/deprecation-and-migration/ + ci-cd-and-automation). Anything outside those allowlists stays advisory — run "claude plugin enable|disable" or "bash lib/toggle-external.sh enable|disable " yourself. EOF diff --git a/lib/profiles/backend.profile b/lib/profiles/backend.profile index 7b7c2d7..c200109 100644 --- a/lib/profiles/backend.profile +++ b/lib/profiles/backend.profile @@ -13,6 +13,11 @@ code-clean personal commit-change personal analyze personal +# Dev-lifecycle skills (agent-skills trio) +observability-and-instrumentation external +deprecation-and-migration external +ci-cd-and-automation external + # Ship + review + land ship review diff --git a/lib/profiles/dev.profile b/lib/profiles/dev.profile index 7e2fe91..3415189 100644 --- a/lib/profiles/dev.profile +++ b/lib/profiles/dev.profile @@ -19,6 +19,11 @@ refactor personal code-clean personal commit-change personal +# Dev-lifecycle skills (agent-skills trio) +observability-and-instrumentation external +deprecation-and-migration external +ci-cd-and-automation external + # Session hygiene context-save land-and-deploy diff --git a/lib/profiles/full.profile b/lib/profiles/full.profile index 5ecb2f3..942a1d2 100644 --- a/lib/profiles/full.profile +++ b/lib/profiles/full.profile @@ -83,6 +83,9 @@ emil-design-eng external frontend-design external design-motion-principles external impeccable external +observability-and-instrumentation external +deprecation-and-migration external +ci-cd-and-automation external ui-ux-pro-max plugin@ui-ux-pro-max-skill # pr-review-toolkit REMOVED from full (audit 2026-07-02 #12): heaviest # single plugin cost (~2.2k tokens of agent descriptions/session), useful diff --git a/lib/skill-routing-census.py b/lib/skill-routing-census.py new file mode 100644 index 0000000..89cb70f --- /dev/null +++ b/lib/skill-routing-census.py @@ -0,0 +1,213 @@ +#!/usr/bin/env python3 +"""Deterministic census of skill-description collisions (TF-IDF cosine). + +Adapted from addyosmani/agent-skills evals Tier 2. Reads every SKILL.md +under the live catalog — `~/.claude/skills/*/` plus plugin skills under +`~/.claude/plugins/cache/*/*/*/skills/*/` and the nested +`.../.claude/skills/*/` layout (override with SKILL_ROUTING_ROOTS, +colon-separated dirs, for fixtures) — extracts the `description` +frontmatter field (scalar or `|`/`>` block), tokenizes it, and scores +every pair by TF-IDF cosine. Two skills with near-identical descriptions +route the same prompt to both: a routing collision, not a naming clash. + +Report: `skills with description: N`, the top 10 pairs, then a WARN line +per pair >= SKILL_ROUTING_WARN (default 0.50) and a FAIL line per pair +>= SKILL_ROUTING_FAIL (default 0.75). Exit 1 iff any pair reached FAIL. +""" +import collections +import glob +import itertools +import math +import os +import re +import sys + +DEFAULT_ROOTS = ["~/.claude/skills", "~/.claude/plugins/cache"] +ROOT_SUFFIXES = [ + "*/SKILL.md", + "*/*/*/skills/*/SKILL.md", + "*/*/*/.claude/skills/*/SKILL.md", +] +STOPWORDS = set( + "a an the and or of to in on for with use when you your this that is are " + "be it as by from at into not no if then use using used uses skill " + "skills user users project projects file files code work".split() +) +BLOCK_MARKERS = ("|", ">", "|-", ">-", "|+", ">+") +WARN_DEFAULT = 0.50 +FAIL_DEFAULT = 0.75 + + +def resolve_roots(): + """Root dirs: SKILL_ROUTING_ROOTS override, else the live catalog.""" + override = os.environ.get("SKILL_ROUTING_ROOTS") + return override.split(":") if override else DEFAULT_ROOTS + + +def discover_paths(roots): + """SKILL.md paths across all roots/suffixes, sorted for determinism.""" + paths = [] + for root in roots: + expanded = os.path.expanduser(root) + for suffix in ROOT_SUFFIXES: + paths.extend(sorted(glob.glob(os.path.join(expanded, suffix)))) + return paths + + +def dedup_by_skill_name(paths): + """name (skill dir) -> path; first path wins (symlinks pre-resolved).""" + by_name = {} + for path in paths: + name = os.path.basename(os.path.dirname(path)) + by_name.setdefault(name, path) + return by_name + + +def _frontmatter(text): + """Raw text between the two leading `---` delimiters, or None.""" + match = re.search(r"^---\r?\n(.*?)\r?\n---", text, re.S) + return match.group(1) if match else None + + +def _strip_quotes(value): + if len(value) >= 2 and value[0] == value[-1] and value[0] in "'\"": + return value[1:-1] + return value + + +def _block_text(lines, start): + """Join a YAML block-scalar's indented continuation lines. + + A blank line inside a block scalar does NOT end it (YAML compares + indentation only against non-blank lines) — only a line back at the + frontmatter's column-0 (the next key) does. + """ + collected = [] + for line in lines[start:]: + if line.strip() == "" or line.startswith((" ", "\t")): + collected.append(line.strip()) + else: + break + return " ".join(part for part in collected if part) + + +def extract_description(path): + """The `description:` frontmatter value (scalar or block), or None.""" + try: + with open(path, encoding="utf-8", errors="ignore") as handle: + text = handle.read() + except OSError: + return None + fm = _frontmatter(text) + if fm is None: + return None + lines = fm.split("\n") + for i, line in enumerate(lines): + match = re.match(r"^description:\s*(.*)$", line) + if not match: + continue + value = match.group(1).strip() + if value in BLOCK_MARKERS: + return _block_text(lines, i + 1) or None + return _strip_quotes(value) or None + return None + + +def stem(word): + """Strip a trailing s/es/ed/ing suffix from a long-enough word.""" + for suffix in ("ing", "ed", "es", "s"): + if len(word) > 4 and word.endswith(suffix): + return word[: -len(suffix)] + return word + + +def tokenize(text): + """Lowercase, keep [a-z][a-z0-9-]+ words, drop stopwords/short, stem.""" + words = re.findall(r"[a-z][a-z0-9-]+", text.lower()) + return [stem(w) for w in words if w not in STOPWORDS and len(w) > 2] + + +def build_docs(paths_by_name): + """name -> Counter(tokens), for every skill with a non-empty description.""" + docs = {} + for name, path in paths_by_name.items(): + desc = extract_description(path) + if not desc: + continue + tokens = tokenize(desc) + if tokens: + docs[name] = collections.Counter(tokens) + return docs + + +def _document_frequencies(docs): + df = collections.Counter() + for counter in docs.values(): + for token in counter: + df[token] += 1 + return df + + +def _tfidf_vector(counter, doc_count, df): + """(1 + log tf) * log(N/df), L2-normalized.""" + weights = { + token: (1 + math.log(n)) * math.log(doc_count / df[token]) + for token, n in counter.items() + } + norm = math.sqrt(sum(w * w for w in weights.values())) or 1 + return {token: w / norm for token, w in weights.items()} + + +def tfidf_vectors(docs): + """name -> {token: TF-IDF weight}, L2-normalized.""" + doc_count = len(docs) + df = _document_frequencies(docs) + return {name: _tfidf_vector(c, doc_count, df) for name, c in docs.items()} + + +def cosine(vec_a, vec_b): + return sum(weight * vec_b.get(token, 0) for token, weight in vec_a.items()) + + +def score_pairs(vectors): + """[(score, a, b), ...] over every pair, sorted by score descending.""" + pairs = [ + (cosine(vectors[a], vectors[b]), a, b) + for a, b in itertools.combinations(sorted(vectors), 2) + ] + pairs.sort(reverse=True) + return pairs + + +def _thresholds(): + warn = float(os.environ.get("SKILL_ROUTING_WARN", WARN_DEFAULT)) + fail = float(os.environ.get("SKILL_ROUTING_FAIL", FAIL_DEFAULT)) + return warn, fail + + +def report(doc_count, pairs): + """Print the census report; return True iff a pair reached FAIL.""" + warn, fail = _thresholds() + print(f"skills with description: {doc_count}") + for score, a, b in pairs[:10]: + print(f"{score:.2f} {a} ~ {b}") + failed = False + for score, a, b in pairs: + if score >= fail: + print(f"FAIL {score:.2f} {a} ~ {b}") + failed = True + elif score >= warn: + print(f"WARN {score:.2f} {a} ~ {b}") + return failed + + +def main(): + by_name = dedup_by_skill_name(discover_paths(resolve_roots())) + docs = build_docs(by_name) + vectors = tfidf_vectors(docs) + failed = report(len(docs), score_pairs(vectors)) + return 1 if failed else 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/lib/tests/floor-guard.test.sh b/lib/tests/floor-guard.test.sh new file mode 100644 index 0000000..423dde4 --- /dev/null +++ b/lib/tests/floor-guard.test.sh @@ -0,0 +1,95 @@ +#!/usr/bin/env bash +# lib/tests/floor-guard.test.sh — flip-tests for lib/floor-guard.sh: one RED +# fixture per KIND, one WAIVED fixture, one CLEAN fixture. Each fixture is a +# fresh throwaway repo under $WORK (`make test` exports +# GIT_CONFIG_GLOBAL=/dev/null; core.hooksPath is also pinned per-repo so a +# machine-wide hook never fires here). This file itself carries the trigger +# strings for every kind — the self-run criterion excludes it by pathspec. +set -u +ROOT="$(cd "$(dirname "$0")/../.." && pwd)" +LIB="$ROOT/lib/floor-guard.sh" +WORK="$(mktemp -d)"; trap 'rm -rf "$WORK"' EXIT +pass=0; fail=0 + +# check_kind +check_kind() { + local kind="$1" rc="$2" want_rc="$3" out="$4" want_sub="$5" + if [ "$rc" = "$want_rc" ] && printf '%s\n' "$out" | grep -qF -- "$want_sub"; then + pass=$((pass+1)); echo "PASS $kind" + else + fail=$((fail+1)) + printf 'FAIL %s: rc=%s (want %s), out:\n%s\n' \ + "$kind" "$rc" "$want_rc" "$(printf '%s\n' "$out" | tail -5)" + fi +} + +# mk_repo → path to a fresh throwaway repo: one test file (3 +# assertions), one vitest.config.ts (coverage.lines: 80), one source file. +mk_repo() { + local d="$WORK/$1" + mkdir -p "$d/src" + git init -q "$d" + git -C "$d" config user.email t@t + git -C "$d" config user.name t + git -C "$d" config core.hooksPath /dev/null + printf 'expect(1).toBe(1);\nexpect(2).toBe(2);\nexpect(3).toBe(3);\n' \ + > "$d/sample.test.js" + printf 'export default {\n coverage: {\n lines: 80,\n },\n};\n' \ + > "$d/vitest.config.ts" + printf 'function add(a, b) {\n return a + b;\n}\n' > "$d/src/index.js" + git -C "$d" add -A + git -C "$d" commit -q -m base + echo "$d" +} + +# ── SUPPRESS ────────────────────────────────────────────────────────────── +d=$(mk_repo suppress); base=$(git -C "$d" rev-parse HEAD) +echo '// eslint-disable-next-line no-console' >> "$d/src/index.js" +out=$(cd "$d" && bash "$LIB" "$base" 2>&1); rc=$? +check_kind SUPPRESS "$rc" 2 "$out" 'FLOOR SUPPRESS' + +# ── SKIP ────────────────────────────────────────────────────────────────── +d=$(mk_repo skip); base=$(git -C "$d" rev-parse HEAD) +echo "it.skip('later', () => {});" >> "$d/sample.test.js" +out=$(cd "$d" && bash "$LIB" "$base" 2>&1); rc=$? +check_kind SKIP "$rc" 2 "$out" 'FLOOR SKIP' + +# ── DELETED_TEST ────────────────────────────────────────────────────────── +d=$(mk_repo deleted); base=$(git -C "$d" rev-parse HEAD) +rm "$d/sample.test.js" +out=$(cd "$d" && bash "$LIB" "$base" 2>&1); rc=$? +check_kind DELETED_TEST "$rc" 2 "$out" 'FLOOR DELETED_TEST' + +# ── ASSERT_DROP ─────────────────────────────────────────────────────────── +d=$(mk_repo assertdrop); base=$(git -C "$d" rev-parse HEAD) +printf 'expect(1).toBe(1);\n' > "$d/sample.test.js" +out=$(cd "$d" && bash "$LIB" "$base" 2>&1); rc=$? +check_kind ASSERT_DROP "$rc" 2 "$out" 'FLOOR ASSERT_DROP' + +# ── STUB ────────────────────────────────────────────────────────────────── +d=$(mk_repo stub); base=$(git -C "$d" rev-parse HEAD) +echo "function todo() { throw new Error('not implemented'); }" >> "$d/src/index.js" +out=$(cd "$d" && bash "$LIB" "$base" 2>&1); rc=$? +check_kind STUB "$rc" 2 "$out" 'FLOOR STUB' + +# ── THRESHOLD_DOWN ──────────────────────────────────────────────────────── +d=$(mk_repo threshold); base=$(git -C "$d" rev-parse HEAD) +sed -i 's/lines: 80/lines: 60/' "$d/vitest.config.ts" +out=$(cd "$d" && bash "$LIB" "$base" 2>&1); rc=$? +check_kind THRESHOLD_DOWN "$rc" 2 "$out" 'FLOOR THRESHOLD_DOWN' + +# ── WAIVED ──────────────────────────────────────────────────────────────── +d=$(mk_repo waived); base=$(git -C "$d" rev-parse HEAD) +echo '// eslint-disable-next-line no-console -- floor-guard: allow legacy shim' \ + >> "$d/src/index.js" +out=$(cd "$d" && bash "$LIB" "$base" 2>&1); rc=$? +check_kind WAIVED "$rc" 0 "$out" 'WAIVED' + +# ── CLEAN ───────────────────────────────────────────────────────────────── +d=$(mk_repo clean); base=$(git -C "$d" rev-parse HEAD) +echo '// helper' >> "$d/src/index.js" +out=$(cd "$d" && bash "$LIB" "$base" 2>&1); rc=$? +check_kind CLEAN "$rc" 0 "$out" 'FLOOR GUARD: clean' + +echo "PASS=$pass FAIL=$fail" +[ "$fail" -eq 0 ] diff --git a/lib/tests/profile-default.test.sh b/lib/tests/profile-default.test.sh index ffeae75..ff7a8b5 100644 --- a/lib/tests/profile-default.test.sh +++ b/lib/tests/profile-default.test.sh @@ -18,7 +18,8 @@ check_not() { case "$2" in *"$3"*) fail=$((fail+1)); FX="$(mktemp -d)"; trap 'rm -rf "$FX"' EXIT mkdir -p "$FX/skills" "$FX/skills-disabled" "$FX/lib/profiles" "$FX/bin" \ - "$FX/hooks" "$FX/skills-external/emil-design-eng" + "$FX/hooks" "$FX/skills-external/emil-design-eng" \ + "$FX/skills-external/observability-and-instrumentation" for g in gs-a gs-b gs-c; do mkdir -p "$FX/skills-external/gstack/$g" touch "$FX/skills-external/gstack/$g/SKILL.md" @@ -29,7 +30,8 @@ cp "$ROOT/hooks/statusline.sh" "$FX/hooks/" cat > "$FX/lib/profiles/full.profile" <<'EOF' gs-a gs-b -emil-design-eng external +emil-design-eng external +observability-and-instrumentation external EOF cat > "$FX/lib/profiles/otherish.profile" <<'EOF' gs-c @@ -89,6 +91,7 @@ check T5-gsa-on "$([ -e "$FX/skills/gs-a" ] && echo on || echo off)" on check T5-gsb-on "$([ -e "$FX/skills/gs-b" ] && echo on || echo off)" on check T5-gsc-off "$([ -e "$FX/skills/gs-c" ] && echo on || echo off)" off check T5-emil-on "$([ -e "$FX/skills/emil-design-eng" ] && echo on || echo off)" on +check T5-obs-on "$([ -e "$FX/skills/observability-and-instrumentation" ] && echo on || echo off)" on out="$(run current)" check T5-first-word "$(first_word "$out")" full check_has T5-match "$out" "100% match" diff --git a/lib/tests/profile-set-managed.test.sh b/lib/tests/profile-set-managed.test.sh index 8a4ba37..755749b 100644 --- a/lib/tests/profile-set-managed.test.sh +++ b/lib/tests/profile-set-managed.test.sh @@ -13,7 +13,8 @@ check() { if [ "$2" = "$3" ]; then pass=$((pass+1)); else fail=$((fail+1)); FX="$(mktemp -d)"; trap 'rm -rf "$FX"' EXIT mkdir -p "$FX/skills" "$FX/skills-disabled" "$FX/lib/profiles" "$FX/bin" \ "$FX/skills-external/emil-design-eng" "$FX/skills-external/other-ext" \ - "$FX/skills-external/21st-ui-build" + "$FX/skills-external/21st-ui-build" \ + "$FX/skills-external/observability-and-instrumentation" for g in gs-a gs-b gs-c; do mkdir -p "$FX/skills-external/gstack/$g" touch "$FX/skills-external/gstack/$g/SKILL.md" @@ -26,8 +27,9 @@ ln -s "$FX/skills-external/other-ext" "$FX/skills/other-ext" cat > "$FX/lib/profiles/designish.profile" <<'EOF' gs-a gs-b -emil-design-eng external -21st-ui-build external +emil-design-eng external +21st-ui-build external +observability-and-instrumentation external EOF cat > "$FX/lib/profiles/backendish.profile" <<'EOF' gs-c @@ -53,6 +55,7 @@ check T2-gsb-on "$([ -e "$FX/skills/gs-b" ] && echo on || echo off)" on check T3-gsc-off "$([ -e "$FX/skills/gs-c" ] && echo on || echo off)" off check T4-emil-src "$([ -L "$FX/skills/emil-design-eng" ] && echo on || echo off)" on check T5-21st-src "$([ -L "$FX/skills/21st-ui-build" ] && echo on || echo off)" on +check T5b-obs-src "$([ -L "$FX/skills/observability-and-instrumentation" ] && echo on || echo off)" on check T6-no-mcp "$(grep -c '^mcp ' "$FX/claude-calls.log" || true)" 0 # --- set backendish: managed leftovers parked/unregistered --- @@ -63,6 +66,8 @@ check T9-emil-off "$([ -e "$FX/skills/emil-design-eng" ] && echo on || echo off check T10-emil-park "$([ -e "$FX/skills-disabled/emil-design-eng" ] && echo p || echo n)" p check T11-21st-off "$([ -e "$FX/skills/21st-ui-build" ] && echo on || echo off)" off check T12-21st-park "$([ -e "$FX/skills-disabled/21st-ui-build" ] && echo p || echo n)" p +check T12b-obs-off "$([ -e "$FX/skills/observability-and-instrumentation" ] && echo on || echo off)" off +check T12c-obs-park "$([ -e "$FX/skills-disabled/observability-and-instrumentation" ] && echo p || echo n)" p check T13-other-untouched "$([ -e "$FX/skills/other-ext" ] && echo on || echo off)" on # --- back to designish: parked external restored (not re-sourced) --- @@ -70,6 +75,8 @@ run set designish >/dev/null 2>&1 check T14-emil-back "$([ -e "$FX/skills/emil-design-eng" ] && echo on || echo off)" on check T15-park-gone "$([ -e "$FX/skills-disabled/emil-design-eng" ] && echo p || echo n)" n check T16-21st-back "$([ -e "$FX/skills/21st-ui-build" ] && echo on || echo off)" on +check T16b-obs-back "$([ -e "$FX/skills/observability-and-instrumentation" ] && echo on || echo off)" on +check T16c-obs-park-gone "$([ -e "$FX/skills-disabled/observability-and-instrumentation" ] && echo p || echo n)" n check T17-no-mcp-ever "$(grep -c '^mcp ' "$FX/claude-calls.log" || true)" 0 printf 'PASS=%s FAIL=%s\n' "$pass" "$fail"; [ "$fail" -eq 0 ] diff --git a/lib/tests/skill-routing-census.test.sh b/lib/tests/skill-routing-census.test.sh new file mode 100755 index 0000000..5fe1c41 --- /dev/null +++ b/lib/tests/skill-routing-census.test.sh @@ -0,0 +1,113 @@ +#!/usr/bin/env bash +# lib/tests/skill-routing-census.test.sh — TF-IDF cosine census of +# skill-description collisions across the live catalog (routing ambiguity), +# adapted from addyosmani/agent-skills evals Tier 2. Two skills whose +# descriptions read alike route the same prompt to both — a routing +# collision, not a cosmetic naming clash. +# +# Fixture flip-test math note: with only 2 documents, corpus-wide IDF zeroes +# EVERY term's contribution — a term unique to one doc (df=1) is absent from +# the other vector and can't enter the dot product, a term shared by both +# (df=2) gets idf=log(N/df)=log(1)=0. Cosine is trivially 0 for any 2-doc +# corpus, near-duplicate or not — a bare 2-doc "distinct pair" fixture would +# pass for that reason alone, not because the pair is actually distinct. Both +# fixtures below ride in a 4-doc corpus so IDF carries real signal. The +# distinct-pair fixture also carries its own same-corpus positive control (a +# near-dup pair that must FAIL) and a sensitivity re-run where the "distinct" +# partner is swapped for a near-copy of its counterpart, to prove the WARN/ +# FAIL marker actually fires when the pair collides — not just that it stays +# silent for reasons unrelated to distinctness. +set -u +ROOT="$(cd "$(dirname "$0")/../.." && pwd)" +CENSUS="$ROOT/lib/skill-routing-census.py" +pass=0; fail=0 +check() { if [ "$2" = "$3" ]; then pass=$((pass+1)); else fail=$((fail+1)); + printf 'FAIL %s: got[%s] want[%s]\n' "$1" "$2" "$3"; fi; } + +WORK="$(mktemp -d)"; trap 'rm -rf "$WORK"' EXIT +TRANSLATE_DESC='description: "Translate a PDF, keeping layout and images."' +SECAUDIT_DESC="description: 'Run a security audit for secrets, CVEs, OWASP.'" + +# _mkskill ... — a fixture skill/SKILL.md +_mkskill() { + local dir="$1" name="$2"; shift 2 + mkdir -p "$dir/$name" + { echo "---"; echo "name: $name"; printf '%s\n' "$@"; echo "---"; } \ + > "$dir/$name/SKILL.md" +} + +# _mkcontrol — same-corpus positive control: skill-e/skill-f, a +# near-duplicate deploy-runbook pair that must FAIL wherever it rides. +_mkcontrol() { + local dir="$1" + _mkskill "$dir" skill-e "description: |" \ + " Use when deploying a project via its per-project runbook to ship the" \ + " application to production servers safely with rollback support and" \ + " health checks after each release." + _mkskill "$dir" skill-f "description: >-" \ + " Use when deploying a project via its per-project deployment runbook to" \ + " ship the application to production servers safely with rollback" \ + " support and health checks after each release." +} + +# ── live run: this machine's real catalog (a FAIL line here → suite RED) ── +live_out=$(python3 "$CENSUS" 2>&1); live_rc=$? +printf '%s\n' "$live_out" +check T1-live-run-no-collision "$live_rc" 0 +has_count=$(printf '%s\n' "$live_out" \ + | grep -qE 'skills with description: [0-9]{2,}' && echo yes) +check T2-live-catalog-non-trivial "$has_count" yes + +# ── fixture AB: near-duplicate pair + 2 unrelated fillers (N=4, real signal) ── +AB="$WORK/ab" +_mkskill "$AB" skill-a "description: |" \ + " Use when deploying a project via its per-project runbook to ship the" \ + " application to production servers safely with rollback support and" \ + " health checks after each release." +_mkskill "$AB" skill-b "description: >-" \ + " Use when deploying a project via its per-project deployment runbook to" \ + " ship the application to production servers safely with rollback" \ + " support and health checks after each release." +_mkskill "$AB" filler-translate "$TRANSLATE_DESC" +_mkskill "$AB" filler-secaudit "$SECAUDIT_DESC" +out_ab=$(SKILL_ROUTING_ROOTS="$AB" python3 "$CENSUS" 2>&1) +fail_line=$(printf '%s\n' "$out_ab" \ + | grep -qE '^FAIL 0\.[0-9]{2} skill-a ~ skill-b$' && echo yes) +check T3-collision-fail-line "$fail_line" yes +[ "$fail_line" = yes ] && echo FIXTURE_COLLISION_DETECTED + +# ── fixture CD: distinct pair + same-corpus positive control (N=4) ──────── +# skill-e/skill-f (positive control) must FAIL; skill-c/skill-d (the pair +# under test) must stay silent (no WARN/FAIL line naming them). +CD="$WORK/cd" +_mkcontrol "$CD" +_mkskill "$CD" skill-c "$TRANSLATE_DESC" +_mkskill "$CD" skill-d "$SECAUDIT_DESC" +out_cd=$(SKILL_ROUTING_ROOTS="$CD" python3 "$CENSUS" 2>&1) +control_fail=$(printf '%s\n' "$out_cd" \ + | grep -qE '^FAIL 0\.[0-9]{2} skill-e ~ skill-f$' && echo yes) +check T4-positive-control-fail "$control_fail" yes +distinct_marker=$(printf '%s\n' "$out_cd" \ + | grep -E '^(WARN|FAIL) 0\.[0-9]{2} skill-c ~ skill-d$') +distinct_silent=$([ -z "$distinct_marker" ] && echo yes || echo no) +check T5-distinct-pair-silent "$distinct_silent" yes + +# ── sensitivity re-run: skill-d -> near-copy of skill-c ──────────────────── +# Same corpus, but skill-d's description now near-duplicates skill-c's: the +# marker that stayed silent above must appear here, proving T5 wasn't silent +# by construction (e.g. a broken pattern or a threshold nothing can cross). +CD2="$WORK/cd-nearcopy" +_mkcontrol "$CD2" +_mkskill "$CD2" skill-c "$TRANSLATE_DESC" +_mkskill "$CD2" skill-d \ + 'description: "Translate a PDF document, keeping the layout and images intact."' +out_cd2=$(SKILL_ROUTING_ROOTS="$CD2" python3 "$CENSUS" 2>&1) +sensitivity_marker=$(printf '%s\n' "$out_cd2" \ + | grep -qE '^(WARN|FAIL) 0\.[0-9]{2} skill-c ~ skill-d$' && echo yes) +check T6-sensitivity-marker-appears "$sensitivity_marker" yes + +[ "$control_fail" = yes ] && [ "$distinct_silent" = yes ] \ + && [ "$sensitivity_marker" = yes ] && echo FIXTURE_DISTINCT_OK + +echo "PASS=$pass FAIL=$fail" +[ "$fail" -eq 0 ] diff --git a/lib/tests/toggle-external-repo-resolution.test.sh b/lib/tests/toggle-external-repo-resolution.test.sh index 8c54109..7df0c68 100644 --- a/lib/tests/toggle-external-repo-resolution.test.sh +++ b/lib/tests/toggle-external-repo-resolution.test.sh @@ -14,6 +14,7 @@ check() { if [ "$2" = "$3" ]; then pass=$((pass+1)); else fail=$((fail+1)); SANDBOX="$(mktemp -d)" mkdir -p "$SANDBOX/repo/lib" "$SANDBOX/repo/skills-external/emil-design-eng" \ + "$SANDBOX/repo/skills-external/observability-and-instrumentation" \ "$SANDBOX/repo/skills" "$SANDBOX/home/.claude" cp "$HELPER_SRC" "$SANDBOX/repo/lib/toggle-external.sh" # mark emil-design-eng ENABLED in the real (physical) repo tree @@ -24,5 +25,11 @@ ln -s "$SANDBOX/repo/lib" "$SANDBOX/home/.claude/lib" out="$(bash "$SANDBOX/home/.claude/lib/toggle-external.sh" status emil-design-eng)" check T1-repo-resolves-through-symlink "$out" enabled +# Same case arm, generalized to $tool for the agent-skills trio (this PR) — +# left unlinked, so it must resolve through the symlink as "disabled", not +# "missing" (which would mean REPO fell back to the wrong tree again). +out="$(bash "$SANDBOX/home/.claude/lib/toggle-external.sh" status observability-and-instrumentation)" +check T2-generalized-tool-resolves-through-symlink "$out" disabled + rm -rf "$SANDBOX" printf 'PASS=%s FAIL=%s\n' "$pass" "$fail"; [ "$fail" -eq 0 ] diff --git a/lib/toggle-external.sh b/lib/toggle-external.sh index f51be4b..5c3d4c8 100755 --- a/lib/toggle-external.sh +++ b/lib/toggle-external.sh @@ -21,6 +21,9 @@ # emil-design-eng — single symlink → skills-external/emil-design-eng # darwin-skill — single symlink → ~/.agents/skills/darwin-skill # 21st — 21st.dev skill pack (needs the `21st` CLI + login) +# observability-and-instrumentation, deprecation-and-migration, +# ci-cd-and-automation — the agent-skills trio, same single-symlink shape +# as emil-design-eng (commit-pinned instead of main-branch tracking) # # For fine-grained activation (only design skills, only qa skills, only # audit skills, etc.) instead of all-or-nothing gstack toggling, use: @@ -40,7 +43,8 @@ warn() { echo -e "${YELLOW}⚠${NC} $1"; } err() { echo -e "${RED}✗${NC} $1"; } # All non-plugin tools this script can toggle. -MANAGED_TOOLS=(gstack emil-design-eng darwin-skill 21st) +MANAGED_TOOLS=(gstack emil-design-eng darwin-skill 21st + observability-and-instrumentation deprecation-and-migration ci-cd-and-automation) # Prints the skill names that belong to the "21st" pack. Source of truth: # skills-external/21st-* — the `21st skills install` run in install-plugins.sh @@ -76,9 +80,9 @@ status_tool() { done < <(gstack_skills) echo "disabled" ;; - emil-design-eng) - [ -d "$REPO/skills-external/emil-design-eng" ] || { echo "missing"; return; } - [ -e "$SKILLS_DIR/emil-design-eng" ] && echo "enabled" || echo "disabled" + emil-design-eng|observability-and-instrumentation|deprecation-and-migration|ci-cd-and-automation) + [ -d "$REPO/skills-external/$tool" ] || { echo "missing"; return; } + [ -e "$SKILLS_DIR/$tool" ] && echo "enabled" || echo "disabled" ;; darwin-skill) [ -d "$HOME/.agents/skills/$tool" ] || { echo "missing"; return; } @@ -115,7 +119,7 @@ disable_tool() { done < <(gstack_skills) ok "gstack disabled ($moved symlinks moved)" ;; - emil-design-eng|darwin-skill) + emil-design-eng|darwin-skill|observability-and-instrumentation|deprecation-and-migration|ci-cd-and-automation) if [ -e "$SKILLS_DIR/$tool" ]; then rm -rf "${DISABLED_DIR:?}/${tool:?}" mv "$SKILLS_DIR/$tool" "$DISABLED_DIR/$tool" @@ -165,11 +169,11 @@ enable_tool() { ok "gstack enabled ($moved symlinks restored)" fi ;; - emil-design-eng|darwin-skill) + emil-design-eng|darwin-skill|observability-and-instrumentation|deprecation-and-migration|ci-cd-and-automation) local src case "$tool" in - emil-design-eng) src="$REPO/skills-external/$tool" ;; darwin-skill) src="$HOME/.agents/skills/$tool" ;; + *) src="$REPO/skills-external/$tool" ;; esac if [ -e "$DISABLED_DIR/$tool" ]; then rm -rf "${SKILLS_DIR:?}/${tool:?}" diff --git a/lib/verify-secure-loop.md b/lib/verify-secure-loop.md index cd37cbe..5bb4fd1 100644 --- a/lib/verify-secure-loop.md +++ b/lib/verify-secure-loop.md @@ -55,6 +55,18 @@ Dispatch a FRESH verifier subagent (`subagent_type: verifier`, or load `TEST` command. Never pass the dev's summary, never pass a prior iteration's gaps — the verifier reads the contract from disk and judges blind. +The verifier's STEP 3 (`agents/verifier.md`) runs `lib/floor-guard.sh` +against the diff before it renders any verdict — a deterministic, +diff-scoped check for a quietly weakened quality bar (a suppressed +lint/type check, a skipped or deleted test, a dropped assertion, a lowered +coverage threshold) that an LLM verdict alone can miss or be talked past +one line at a time. Its findings fold straight into that same verifier's +`ECARTS` count unless the contract's `CLARIFICATIONS` explicitly authorizes +the exact weakening; there is no separate gate and no extra dispatch, it +rides this GATE 1 call. A `floor-guard: allow` waiver outside a test file +is a finding too unless the contract's `CLARIFICATIONS` names it: the +waiver is self-service, the contract is human-gated (BDR-102 amendment). + Parse its single `VERIFY — VERDICT:` line: - `CONFORME` → go to GATE 2. (First-pass conforme = no loop.) @@ -74,9 +86,9 @@ Parse its single `VERIFY — VERDICT:` line: micro-gate that appends `[gated ]` to the contract's FILE SCOPE; otherwise the dev removes the file. - Structural failure (`ERROR(…)`, missing/duplicated VERDICT line, - unparsable, crash, `CONFORME` without `PROOF`) → retry ONCE with a fresh - verifier; a 2nd structural failure → human escalation. A mute verifier is - NEVER a PASS. + unparsable, crash, `CONFORME` without `PROOF` or without `FLOOR`) → retry + ONCE with a fresh verifier; a 2nd structural failure → human escalation. + A mute verifier is NEVER a PASS. ## GATE 2 — SECURITY (fresh security-auditor) diff --git a/link.sh b/link.sh index d260640..f9c19d4 100644 --- a/link.sh +++ b/link.sh @@ -93,7 +93,8 @@ fi # impeccable is NOT here: its installer writes the skill straight into # skills/ (and its agents into agents/) at --scope=global, so there is no # skills-external/ copy to symlink. See install-plugins.sh Step 8d. -EXTERNAL_SKILLS=(emil-design-eng frontend-design design-motion-principles) +EXTERNAL_SKILLS=(emil-design-eng frontend-design design-motion-principles + observability-and-instrumentation deprecation-and-migration ci-cd-and-automation) for _ext_skill in "${EXTERNAL_SKILLS[@]}"; do if [ -d "$REPO/skills-external/$_ext_skill" ]; then if [ -L "$CLAUDE/skills/$_ext_skill" ] && [ "$(readlink "$CLAUDE/skills/$_ext_skill")" = "$REPO/skills-external/$_ext_skill" ]; then diff --git a/plugins.lock.json b/plugins.lock.json index db20c6f..60d223c 100644 --- a/plugins.lock.json +++ b/plugins.lock.json @@ -43,6 +43,13 @@ "managed_by": "curl", "note": "Emil Kowalski's design engineering skill — UI polish, animations, component craft. Machine-owned: curl'd to skills-external/emil-design-eng/ (gitignored, re-fetched by update-all.sh), symlinked by link.sh." }, + "agent-skills": { + "source": "https://github.com/addyosmani/agent-skills", + "commit": "2686b620fc1fed2e8f60c704839c766b8594c6b6", + "skills": ["observability-and-instrumentation", "deprecation-and-migration", "ci-cd-and-automation"], + "managed_by": "curl", + "note": "Three dev-lifecycle skills from addyosmani/agent-skills, vendored the emil-design-eng way but COMMIT-pinned (not main-branch tracking): each lands in skills-external//SKILL.md (gitignored, symlinked by link.sh), install-plugins.sh Step 8e curls all three at this commit when absent, update-all.sh re-fetches at the SAME commit on every run (a pin, not an auto-advance). Bump the commit deliberately to pick up upstream changes; the scripts read it from here, never hardcode it." + }, "impeccable": { "source": "npm:impeccable", "version": "4.1.0", diff --git a/rules/rest-api.md b/rules/rest-api.md new file mode 100644 index 0000000..e971f83 --- /dev/null +++ b/rules/rest-api.md @@ -0,0 +1,43 @@ +--- +paths: ["**/api/**", "**/routes/**", "**/controllers/**", "**/*.route.*", "**/*.controller.*", "**/openapi.*", "**/*.openapi.*"] +--- + +# REST API — contract, errors, lists, idempotency + +## Contract first +Order: typed input/output → schemas (server-generated fields like id, +createdAt apart from client input) → error codes → implementation. +Validate at boundaries only (route handlers, external responses, env +loading); trust internal code and your own database reads. + +## Errors +One envelope everywhere: `{ error: { code, message, details? } }`. +HTTP map: 400 invalid · 401 auth · 403 forbidden · 404 missing · +409 conflict · 422 semantic · 500 server (never leak internals). +Never mix throw / null / envelope styles across endpoints. + +## Lists +Every list endpoint paginated: `page`, `pageSize`, `totalItems`, +`totalPages`. Filters as query params (`?status=x&createdAfter=…`), +never in the body. + +## Idempotency +Key from intent, not attempt (`charge:v1:${orderId}`, never +`randomUUID()` or a timestamp). Claim atomically via a unique +constraint — check-then-insert is a race, not a guard. Same key, +different payload: fail loudly, never replay the first response. +Pick the in-flight-duplicate policy on purpose: 409 reject, bounded +wait, or 202 + status URL. Retention outlives the longest retry +path, dead-letter replay included. + +## Naming +Plural nouns for endpoints (`/api/tasks`), no verbs. camelCase query +params and response fields. UPPER_SNAKE enum values. Boolean fields +prefixed is/has/can. + +## Hyrum's law +Every observable behavior — undocumented quirks, error text, timing — +becomes a de facto contract once someone depends on it. + +## Versioning +CLAUDE.md § Web APIs — always versioned. Not repeated here. diff --git a/update-all.sh b/update-all.sh index 8432110..f1384fc 100644 --- a/update-all.sh +++ b/update-all.sh @@ -379,6 +379,38 @@ else info "design-motion-principles not installed — skipping" fi +# ── 7.3. Update Agent Skills (addyosmani/agent-skills, pinned commit) ── +echo "" +echo "── Updating Agent Skills (addyosmani/agent-skills)..." +AGENT_SKILLS_SHA="" +if [ -f "$REPO/plugins.lock.json" ] && command -v python3 &>/dev/null; then + AGENT_SKILLS_SHA=$(python3 -c " +import json, sys +with open(sys.argv[1]) as f: + d = json.load(f) +print(d.get(sys.argv[2], {}).get('commit', '')) +" "$REPO/plugins.lock.json" "agent-skills" 2>/dev/null || true) +fi +AGENT_SKILLS_NAMES=(observability-and-instrumentation deprecation-and-migration ci-cd-and-automation) +if [ -z "$AGENT_SKILLS_SHA" ]; then + warn "agent-skills: no commit pinned in plugins.lock.json — skipping" +else + for _as_skill in "${AGENT_SKILLS_NAMES[@]}"; do + _as_dir="$REPO/skills-external/$_as_skill" + if [ ! -d "$_as_dir" ]; then + info "$_as_skill not installed — skipping (run: make plugin)" + continue + fi + _as_url="https://raw.githubusercontent.com/addyosmani/agent-skills/$AGENT_SKILLS_SHA/skills/$_as_skill/SKILL.md" + if curl -fsSL "$_as_url" -o "$_as_dir/SKILL.md.tmp" \ + && mv "$_as_dir/SKILL.md.tmp" "$_as_dir/SKILL.md"; then + ok "$_as_skill re-fetched at pinned commit" + else + warn "$_as_skill update failed" + fi + done +fi + # ── Impeccable (design detector + skill + subagents) ── # Global scope: the installer writes through the ~/.claude/{skills,agents} # symlinks straight into this repo (install-plugins.sh Step 8d explains why