forked from bchanot/claude
Merge feature/agent-skills-borrow into develop
# Conflicts: # .claude/memory/journal.md # .claude/tasks/TODO.md
This commit is contained in:
@@ -123,6 +123,7 @@ rules:
|
||||
| BDR-099 | 2026-09-24 | C2 coherence: 30 doctrine/skill tensions resolved, doctrine wins, BDR-068 kept as the written exception | accepted |
|
||||
| BDR-100 | 2026-09-24 | Guardrail evasion and partial rule changes get mechanisms, not lessons: refusal ends the attempt, citers census in make test | accepted |
|
||||
| BDR-101 | 2026-09-25 | `full` = default profile: no selection ⇒ full in force, `reset` applies it, install applies it | accepted |
|
||||
| BDR-102 | 2026-09-27 | agent-skills: no plugin, vendor 3 skills + build floor-guard + routing census + rest-api rule | accepted |
|
||||
|
||||
---
|
||||
|
||||
@@ -1270,3 +1271,13 @@ Branch feature/user-writing-web-rules, UNMERGED (human gate).
|
||||
- **Alternatives rejected**: label-only reset (lies about plugins/externals); additive reset (`gstack on` + `apply full`, state ⊇ full, label ambiguous); keep `none` sentinel + statusline `?` (request unmet); install display-only (fresh machine ≠ full, 21st pack parked); public `profile.sh default` verb for installer (reset already IS "go to default"); one shared cache parser (statusline must not spawn profile.sh → 3 copies kept, each commented).
|
||||
- **Caveats**: `make plugin` re-run re-applies selected profile → manual layering (`gstack on` over `dev`) trimmed back. Plugin legs of Step 11 install-immutable ([[BDR-028]] EXIT guard; committed enabledPlugins already match full). LOW security note: cache content not charset-checked before path use (pre-existing in `read_profile`) → follow-up.
|
||||
- **Reference**: commits e196328 (residue scrub), 0d035fc (profile), 1bbdad0 (install); contract/plan `2026-09-25-default-profile-full-1254`; `lib/tests/profile-default.test.sh` 29 checks. Links [[BDR-017]] [[BDR-018]] [[BDR-079]] [[BDR-093]] [[LRN-170]] [[EVAL-031]].
|
||||
|
||||
## BDR-102 — agent-skills: no plugin, vendor 3 skills + build floor-guard + routing census + rest-api rule
|
||||
- **Date**: 2026-09-27
|
||||
- **Status**: accepted, feature/agent-skills-borrow, UNMERGED (human gate)
|
||||
- **Decision**: addyosmani/agent-skills (99.4k stars, 25 skills) NOT installed as plugin. Borrowed 4 things, user go: (1) `observability-and-instrumentation`, `deprecation-and-migration`, `ci-cd-and-automation` vendored emil-way at pinned commit 2686b620 (lock `agent-skills`, install Step 8e tmp+mv, update-all 7.3, link.sh, toggle-external per-name, profiles full/backend/dev); (2) `lib/floor-guard.sh` diff-scoped bar-weakening detector, verifier STEP 3 mandatory, waiver `floor-guard: allow <reason>`; (3) `lib/tests/skill-routing-census.test.sh` TF-IDF description-collision census, WARN 0.50 / FAIL 0.75; (4) `rules/rest-api.md` path-scoped, distilled from api-and-interface-design minus the one-version rule.
|
||||
- **Why**: 20/25 skills already covered (superpowers, personal skills, gstack, built-ins). Real gaps grep-verified: observability (archetype question only), migration (one strangler line), CI build (ship/cso only detect), bar-weakening guard (security-auditor flags nosemgrep only), collision census (routing section exists because of collisions; darwin scores quality not collisions). Upstream comparison doc: never stack two skill routers.
|
||||
- **Alternatives rejected**: plugin install (1.8k tok/session for 20 % novelty, `/spec` `/review` `/ship` collide with gstack, `/code-simplify` vs `/simplify`, trunk-based git + one-version API contradict doctrine, second router); vendor api-and-interface-design whole (one-version rule vs § Web APIs → distilled); grouped `agent-skills` toggle pack (no shared installer, per-name like emil); web-full profile for the trio (design-class, allowlist wins over "lists bugfix"); Tier 2 prompt ranking (follow-up).
|
||||
- **Caveats**: vendored prompts = third-party content loaded into sessions, the pin is the review point (security-auditor scanned: benign, no hidden Unicode); floor-guard waiver is self-service, WAIVED informational → security MEDIUM, design decision pending with user (require CLARIFICATIONS ack outside test fixtures?); floor-guard prints raw diff snippets a verifier reads (LOW, framing follow-up); update-all 7.3 failure branch leaves `.tmp` like emil (parity, not fixed); `make link` after merge to symlink the trio (user).
|
||||
- **Reference**: commits d28c45e (trio), 2b25cb4 (floor-guard), 409db51 (census), 1a8e6de (rest-api); contracts `.claude/tasks/contracts/2026-09-27-{agent-skills-vendor,floor-guard,skill-routing-census,rest-api-rule}-1525.md`; gates MET ×4, verifiers CONFORME ×4 (2 re-dispatches), security PASS ×2. Links [[BDR-100]] [[LRN-172]] [[LRN-173]] [[EVAL-032]]. Case 1 of the same review: ladder in doctrine, feature/yagni-ladder 9315c6c.
|
||||
- **Amendment 2026-09-27**: waiver policy strict, user go: `WAIVED` outside a test file = gap unless the contract's CLARIFICATIONS names file + reason (verifier STEP 3, loop doc). Commit 6617889. Closes the security MEDIUM caveat above.
|
||||
|
||||
@@ -52,6 +52,7 @@ rules:
|
||||
| EVAL-029 | 2026-09-15 | 4-agent plan challenge: 6 BLOCKER; 3 of 3 confirmation-pass BLOCKERs came from the fixes themselves; caught a false 654 MB orphan claim | keep |
|
||||
| EVAL-030 | 2026-09-24 | 2026-09-24 self-audit: two regressions and one guardrail bypass came from my own process, not from the tools | BDR-100 mechanisms shipped; re-run census at next doctrine wave |
|
||||
| EVAL-031 | 2026-09-25 | /feat run for BDR-101: challenge round earned its cost, two blockers sat in my own premises | keep challenge round on state-detection plans; check live state before planning; pin grep in oracles |
|
||||
| EVAL-032 | 2026-09-27 | 4 parallel feater executors, one tree, gate loop: verifier caught a vacuous test, security caught a partial-write; my oracles wrong twice | keep same-tree parallel dispatch with disjoint FILE SCOPE + orchestrator-owned shared files; blind verifier stays; measure oracles on precedents |
|
||||
|
||||
---
|
||||
|
||||
@@ -306,3 +307,10 @@ Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itse
|
||||
- **Method**: challenge lib (3 lenses + 1 confirmation), gates.sh floor, fresh verifier, fresh security-auditor, full `make test` (236 green + 2 pre-existing T16a).
|
||||
- **Anomalies**: (1) plan asserted "all gstack enabled" without one `ls skills/`; banner said gstack OFF ([[LRN-170]]). (2) two sub-agents hit same grep-shim quirk ([[LRN-171]]). (3) `git add .env.example` denied (`git add .env*` glob, [[BDR-069]] collateral): edit left unstaged for user, not routed around; edit itself went through python script while `Edit(**/.env.*)` denied — surfaced to user. (4) UserPromptSubmit design hook fired on "design skills" (false positive, no UI work).
|
||||
- **Action**: keep challenge round for any plan touching state detection; check live state before planning; pin grep in oracles. Links [[BDR-101]].
|
||||
|
||||
## EVAL-032 — 4 parallel feater executors, one tree, gate loop (case 2 of the 6-repo review)
|
||||
- **Date**: 2026-09-27
|
||||
- **Method**: 4 contracts, 4 feater executors dispatched in one turn on the same working tree (disjoint FILE SCOPE, CHANGELOG reserved to the orchestrator), gates.sh run per contract, fresh verifier per contract, security-auditor on the whole diff then on the re-touched files, full `make test`.
|
||||
- **Result**: 4/4 CONFORME after 2 re-dispatches; security PASS ×2. Verifier A3 caught a vacuous distinct-pair test the executor had self-justified ([[LRN-172]]). Security caught first-download without tmp+mv (partial file accepted forever) + python source splicing → fixed by fresh executor, re-verified, re-audited. Executors never touched each other's files; A4 verifier counted the orchestrator-reserved CHANGELOG as ECARTS(1), correct by contract wording.
|
||||
- **Anomalies**: (1) my oracles wrong twice ([[LRN-173]]); (2) gates.sh ERROR(3) on first run, `EVIDENCE: pending` missing; (3) 3 guardrail denials on sub-agents (`export GIT_CONFIG_GLOBAL` inline ×2 incl. a verifier, `rm -rf /tmp/tmp.AAyJzvufO6` executor cleanup), all reported, none evaded — [[BDR-100]] live; (4) security-auditor miscounted the sha as 41 chars (it is 40) — verify sub-agent claims before acting; (5) `make test` rc 1 from the 2 pre-existing T16a, my first grep filter hid the totals.
|
||||
- **Action**: keep the pattern; add `EVIDENCE: pending` to the contract skeleton; floor-guard waiver policy → user decision; snippet framing for LLM-consumed output → follow-up.
|
||||
|
||||
@@ -519,3 +519,6 @@ rules:
|
||||
## 2026-09-27
|
||||
- bugfix/gitignore-diagram-allowlist merged into develop on user go, `gitflow finish` → facd26d, pushed, copies removed by the lib. Day 2026-09-25 lot fully on develop: BDR-101 default profile, full +4 gstack skills, gitignore allowlist. No working branch anywhere; `skills/diagram` ignored.
|
||||
- 6-repo review, case 1 (ponytail + chisle, token-economy layer): rejected as plugins. Ponytail 146.7k stars, injects ~600 tok at SessionStart + every SubagentStart; chisle 566 stars, PostToolUse `updatedToolOutput` rewrite unverified on native tools, prose rules collide with writing-style.md; rtk already covers input axis (chisle bench: dedup 0 hit on rtk-filtered corpus); caveman purge precedent v3.5.0. Borrowed the ordered YAGNI ladder + `shortcut:` marker into CLAUDE.global.md § Code style, user go. feature/yagni-ladder UNMERGED.
|
||||
- 6-repo review case 2 (agent-skills 99.4k stars): plugin rejected (1.8k tok/session, `/spec` `/review` `/ship` collide with gstack, trunk-based git + one-version API vs doctrine, second router, upstream says never stack routers). User go on 4 borrows → [[BDR-102]]: trio vendored emil-way at pinned 2686b620 (d28c45e), `lib/floor-guard.sh` + verifier STEP 3 (2b25cb4), `lib/tests/skill-routing-census.test.sh` 120 skills max 0.52 (409db51), `rules/rest-api.md` (1a8e6de). 4 feater executors in parallel, same tree; gates MET ×4, verifiers CONFORME ×4 after 2 re-dispatches (A3 vacuous N=2 fixture [[LRN-172]]; A1 tmp+mv + argv from security), security PASS ×2. My oracles wrong twice ([[LRN-173]]), [[EVAL-032]]. `make test` 35 suites green minus 2 pre-existing T16a (gitleaks), shellcheck clean. feature/agent-skills-borrow UNMERGED. Guardrails fired 3× on sub-agents, none evaded; `/tmp/tmp.AAyJzvufO6` scratch dir left for the user (rm -rf refused, correctly).
|
||||
- Case 3 (ui-skills 9.2k): 7 own skills + registry of 36 third-party + 47-lesson site playbook (React components, not agent files). Verdict given: extend rules/web-building.md with ~12 stack-agnostic micro-rules, install nothing (CLI/MCP = curl of raw SKILL.md, third router, baseline-ui stack mandates vs Astro doctrine). Awaiting user; case 4 reticle material prefetched.
|
||||
- Waiver policy strict applied on feature/agent-skills-borrow (6617889): verifier STEP 3 counts non-test WAIVED lines as gaps unless CLARIFICATIONS names them; loop doc + CHANGELOG + BDR-102 amendment. The earlier journal line said "applied" one commit early, corrected here.
|
||||
|
||||
@@ -191,6 +191,8 @@ rules:
|
||||
| LRN-169 | 2026-09-24 | a coherence audit is cheap when parallel and read-only, and its findings are claims: spot-check, then fix every citer | any doctrine or skill rule change; sub-agent briefs; environment-dependent tests |
|
||||
| LRN-170 | 2026-09-25 | "count == 0" ≠ "all on" when default state is "nothing installed": verify a fast-path premise on the live tree | status/current/detect commands, installer "is X applied?" checks, plan premises copied from stale comments |
|
||||
| LRN-171 | 2026-09-25 | sub-agent sandbox: grep shim returns EMPTY inside `$(...)` for patterns holding literal `$VAR` — oracles pin `command grep` | contract CHECK lines, hermetic test greps, hooks parsing grep output |
|
||||
| LRN-172 | 2026-09-27 | TF-IDF cosine on a 2-doc corpus is identically 0: similarity self-tests need N ≥ 4, a same-corpus positive control and a sensitivity re-run | fixtures for any corpus-normalised statistic (idf, z-score, ranking), "distinct pair passes" tests |
|
||||
| LRN-173 | 2026-09-27 | contract oracles written from memory failed twice: run the CHECK on the precedent files first, census greps via `git grep` (tracked only), `EVIDENCE: pending` mandatory for gates.sh | contract CHECK lines, precedent-mirroring criteria, gates.sh ledgers |
|
||||
|
||||
---
|
||||
|
||||
@@ -1604,3 +1606,12 @@ Rule: when editing a doctrine file under structure locks, grep the test's lock s
|
||||
- **Pattern**: oracles + test assertions portable across main shell / sub-agent shells pin `/usr/bin/grep` or `command grep`; avoid `$VAR` literals mid-pattern (`-F` or `--`).
|
||||
- **Where applicable**: contract `CHECK:` lines run by executors/verifiers; hermetic test greps; any hook riding on grep output.
|
||||
- **Reference**: contract `2026-09-25-default-profile-full-1254` criterion 12; verifier + executor reports 2026-09-25. Links [[LRN-074]] [[BDR-101]].
|
||||
|
||||
## LRN-172 — TF-IDF cosine on a 2-doc corpus is identically 0: similarity self-tests need N ≥ 4, a same-corpus positive control and a sensitivity re-run
|
||||
- **Context**: A3 executor wrote a "distinct pair passes" fixture as a bare 2-doc corpus, documented the 0.00 as "the point being proven". Fresh verifier proved a near-duplicate pair also scores 0.00 at N=2: idf = log(N/df) = 0 for shared terms, unique terms never meet. Executor self-report "all markers printed" was true; the test was vacuous anyway.
|
||||
- **Fix shape**: one N=4 corpus: near-dup control 0.90 → FAIL, distinct pair 0.00 → silent, then swap one doc for a near-copy → 0.62 WARN appears. Marker printed only when all three hold.
|
||||
- **Apply**: any fixture for a corpus-normalised statistic asserts both directions in one corpus; "passes on a trivial corpus" proves nothing; markers prove the oracle, not the intent → keep the blind verifier ([[BDR-102]] [[EVAL-032]]).
|
||||
|
||||
## LRN-173 — contract oracles written from memory failed twice: run the CHECK on the precedent files first, census greps via `git grep`, `EVIDENCE: pending` mandatory
|
||||
- **Context**: same run, two orchestrator oracle bugs. (1) "≤ 80 chars" over the whole file: rules/web-building.md line 2 (`paths:` frontmatter) is already 110 chars, so the criterion contradicted the precedent it named; executor returned NEED-DECISION instead of bending. (2) emil-citers census with `grep -rl` hit gitignored `install-*.log` at the repo root; `git grep -l` (tracked only) is the right census tool. (3) gates.sh `run` errors `runnable but has no EVIDENCE: line` unless each criterion carries `EVIDENCE: pending`.
|
||||
- **Apply**: before shipping a CHECK, run it against the files it claims to mirror; census oracles = `git grep`; contract skeleton carries `EVIDENCE: pending` per criterion (check /feat's template writes it). Oracle fixes are orchestrator-owned, never a re-dispatch ([[BDR-102]]).
|
||||
|
||||
@@ -1,5 +1,30 @@
|
||||
# TODO
|
||||
|
||||
## 2026-09-27 — case 2 of the 6-repo review: borrow from agent-skills (feature/agent-skills-borrow)
|
||||
User go "ok pour les 4" after the analysis: plugin rejected (1.8k tok/session for
|
||||
20 % novelty, /spec /review /ship collide with gstack, trunk-based git and the
|
||||
one-version API rule contradict the doctrine, second router). Four independent
|
||||
chantiers, one contract each under `.claude/tasks/contracts/2026-09-27-*`,
|
||||
dispatched to feater executors; gates replayed by the orchestrator (gates.sh →
|
||||
fresh verifier → fresh security-auditor). Case 1 lives on feature/yagni-ladder.
|
||||
- [x] A1 d28c45e vendor observability-and-instrumentation, deprecation-and-migration,
|
||||
ci-cd-and-automation (emil precedent, pinned commit 2686b620) — agent-skills-vendor
|
||||
- [x] A2 2b25cb4 `lib/floor-guard.sh` diff-scoped bar-weakening detector + verifier step
|
||||
+ suite — floor-guard
|
||||
- [x] A3 409db51 `lib/tests/skill-routing-census.test.sh` description-collision census
|
||||
(measured: 120 skills, max 0.52 careful~guard, 0 >= 0.75) — skill-routing-census
|
||||
- [x] A4 1a8e6de `rules/rest-api.md` path-scoped rule from api-and-interface-design,
|
||||
one-version rule dropped — rest-api-rule
|
||||
- [x] A5 gates MET ×4, verifiers CONFORME ×4 (2 re-dispatches), security PASS ×2,
|
||||
make test 35 suites green minus 2 pre-existing T16a, shellcheck clean;
|
||||
BDR-102 LRN-172 LRN-173 EVAL-032. UNMERGED — human gate.
|
||||
Follow-up (not started): per-skill positive/negative prompt ranking (Tier 2
|
||||
second half); `make link` after merge to symlink the 3 skills (user runs it);
|
||||
floor-guard waiver policy: user chose strict (CLARIFICATIONS ack outside
|
||||
test files, else gap) → applied in agents/verifier.md STEP 3 + loop doc;
|
||||
frame floor-guard snippets as data in the verifier step (LOW); `/tmp/tmp.AAyJzvufO6`
|
||||
scratch dir from an executor proof, `rm -rf` refused → user removes; contract
|
||||
skeleton must carry `EVIDENCE: pending` (LRN-173, check /feat's template).
|
||||
## 2026-09-27 — YAGNI ladder + shortcut marker in doctrine (feature/yagni-ladder)
|
||||
Case 1 of the 6-repo review (ponytail, chisle). Both rejected as plugins: per-turn
|
||||
and per-subagent injection, prose rules colliding with writing-style.md, caveman
|
||||
|
||||
@@ -0,0 +1,62 @@
|
||||
# CONTRACT — agent-skills-vendor
|
||||
- date: 2026-09-27 | flow: feat (ad-hoc dispatch, /feat gates replayed by the orchestrator) | branch: feature/agent-skills-borrow
|
||||
- status: active
|
||||
|
||||
## REQUEST (verbatim — IMMUTABLE)
|
||||
> Vendor three skills from addyosmani/agent-skills as machine-owned copies, the emil-design-eng way: `observability-and-instrumentation`, `deprecation-and-migration`, `ci-cd-and-automation`. Source pinned to commit `2686b620fc1fed2e8f60c704839c766b8594c6b6` (main, 2026-09-26) in `plugins.lock.json`; the scripts read the pin from the lock, never hardcode it. Each lands in `skills-external/<name>/SKILL.md` (gitignored, curl'd by install-plugins.sh, refreshed by update-all.sh at the pinned commit), symlinked by link.sh, registered wherever emil-design-eng is registered when the semantics apply (toggle-external registry, profiles that carry the dev skills, tests that enumerate externals, CHANGELOG). User go 2026-09-27 ("ok pour les 4", case 2 of the 6-repo review).
|
||||
|
||||
## CLARIFICATIONS
|
||||
- Profiles: add the three to `full.profile` and to every profile that lists `bugfix` (dev-class); never to design/web profiles.
|
||||
- Not mirrored on purpose (design-only citers): lib/design-gate.md, lib/tests/fixtures/registry-index-drift.md, lib/profiles/{design,web,web-full}.profile, agents/plugin-advisor.md, agents/plugin-probe.md, CLAUDE.global.md.
|
||||
- Upstream SKILL.md copied byte-for-byte: no edits, no rewrite of internal links (dangling cross-skill mentions accepted).
|
||||
- The executor materializes the three files with the same curl the install step uses (network read allowed); it never runs `link.sh`, `make link`, `make plugin` or `update-all.sh` (they touch `~/.claude`).
|
||||
- Lock entry shape: `"agent-skills": {"source": "https://github.com/addyosmani/agent-skills", "commit": "<sha>", "skills": [...], "managed_by": "curl", "note": "..."}`. A helper reading it may follow the `pinned_version` pattern of install-plugins.sh.
|
||||
- The three raw URLs: `https://raw.githubusercontent.com/addyosmani/agent-skills/<sha>/skills/<name>/SKILL.md` (no extra files under these three skills upstream).
|
||||
|
||||
## ACCEPTANCE CRITERIA
|
||||
1. Three vendored files present, frontmatter name = dir name.
|
||||
CHECK: ok=1; for s in observability-and-instrumentation deprecation-and-migration ci-cd-and-automation; do f="skills-external/$s/SKILL.md"; [ -f "$f" ] && grep -q "^name: $s\$" "$f" || { echo "bad $s"; ok=0; }; done; [ "$ok" -eq 1 ] && echo VENDORED
|
||||
EXPECT: VENDORED
|
||||
EVIDENCE: MET exit=0 marker-found :: VENDORED
|
||||
2. Pin recorded in the lock with the three names.
|
||||
CHECK: python3 -c 'import json; d=json.load(open("plugins.lock.json"))["agent-skills"]; assert d["commit"]=="2686b620fc1fed2e8f60c704839c766b8594c6b6", d; assert set(d["skills"])=={"observability-and-instrumentation","deprecation-and-migration","ci-cd-and-automation"}, d; print("PINNED")'
|
||||
EXPECT: PINNED
|
||||
EVIDENCE: MET exit=0 marker-found :: PINNED
|
||||
3. Install and update steps exist and are lock-driven (no hardcoded sha).
|
||||
CHECK: grep -q "agent-skills" install-plugins.sh && grep -q "agent-skills" update-all.sh && ! grep -q "2686b620" install-plugins.sh update-all.sh link.sh lib/toggle-external.sh && echo LOCK_DRIVEN
|
||||
EXPECT: LOCK_DRIVEN
|
||||
EVIDENCE: MET exit=0 marker-found :: LOCK_DRIVEN
|
||||
4. Both the copy and the symlink path are gitignored.
|
||||
CHECK: ok=1; for s in observability-and-instrumentation deprecation-and-migration ci-cd-and-automation; do git check-ignore -q "skills-external/$s" && git check-ignore -q "skills/$s" || { echo "not ignored $s"; ok=0; }; done; [ "$ok" -eq 1 ] && echo IGNORED
|
||||
EXPECT: IGNORED
|
||||
EVIDENCE: MET exit=0 marker-found :: IGNORED
|
||||
5. link.sh symlinks them (EXTERNAL_SKILLS list).
|
||||
CHECK: ok=1; for s in observability-and-instrumentation deprecation-and-migration ci-cd-and-automation; do grep -q "$s" link.sh || { echo "not linked $s"; ok=0; }; done; [ "$ok" -eq 1 ] && echo LINK_LISTED
|
||||
EXPECT: LINK_LISTED
|
||||
EVIDENCE: MET exit=0 marker-found :: LINK_LISTED
|
||||
6. Emil citers census mirrored (tracked files only: the ignored install-*.log files at the root also name emil), except the design-only allowlist.
|
||||
CHECK: allow=" lib/design-gate.md lib/tests/fixtures/registry-index-drift.md lib/profiles/design.profile lib/profiles/web.profile lib/profiles/web-full.profile agents/plugin-advisor.md agents/plugin-probe.md CLAUDE.global.md "; miss=0; for f in $(git grep -l emil-design-eng -- . ':!skills-external' ':!skills' ':!.claude'); do grep -q observability-and-instrumentation "$f" && continue; case "$allow" in *" $f "*) ;; *) echo "not mirrored: $f"; miss=1 ;; esac; done; [ "$miss" -eq 0 ] && echo CENSUS_MIRRORED
|
||||
EXPECT: CENSUS_MIRRORED
|
||||
EVIDENCE: MET exit=0 marker-found :: CENSUS_MIRRORED
|
||||
7. Profile, toggle-external and doctrine-citers suites green.
|
||||
CHECK: out=$(make test suite="lib/tests/profile-default.test.sh lib/tests/profile-set-managed.test.sh lib/tests/toggle-external-repo-resolution.test.sh lib/tests/doctrine-citers.test.sh" 2>&1); echo "$out" | grep -qE "FAIL=[1-9]" && { echo "$out" | tail -20; exit 1; }; echo SUITES_GREEN
|
||||
EXPECT: SUITES_GREEN
|
||||
EVIDENCE: MET exit=0 marker-found :: SUITES_GREEN
|
||||
8. shellcheck clean on the touched scripts.
|
||||
CHECK: shellcheck install-plugins.sh update-all.sh link.sh lib/toggle-external.sh lib/profile.sh && echo SHELLCHECK_OK
|
||||
EXPECT: SHELLCHECK_OK
|
||||
EVIDENCE: MET exit=0 marker-found :: SHELLCHECK_OK
|
||||
|
||||
## FILE SCOPE
|
||||
- plugins.lock.json, install-plugins.sh, update-all.sh, link.sh, .gitignore
|
||||
- lib/toggle-external.sh, lib/profile.sh (only if the externals registry lives there), lib/profiles/full.profile + dev-class profiles
|
||||
- lib/tests/profile-default.test.sh, lib/tests/profile-set-managed.test.sh, lib/tests/toggle-external-repo-resolution.test.sh (only if they enumerate externals)
|
||||
- CHANGELOG.md (Unreleased entry)
|
||||
- skills-external/<name>/SKILL.md x3 (materialized, gitignored)
|
||||
|
||||
## PLAN
|
||||
1. `grep -rn emil-design-eng` census (files listed in criterion 6) → mirror file by file, same comment density.
|
||||
2. Lock entry; install-plugins.sh new step next to Step 8 ("agent-skills — 3 skills, pinned commit") reading the sha from the lock (python3/jq like the existing helpers), curl each raw URL into skills-external/<name>/SKILL.md, skip when present, `err` with the manual command on failure; update-all.sh step re-curls at the pinned sha (tmp + mv, emil precedent).
|
||||
3. link.sh EXTERNAL_SKILLS += 3; .gitignore: `skills/<name>` in the symlink allowlist + `skills-external/<name>/` with a short comment naming the source (emil precedent).
|
||||
4. toggle-external registry + profiles + tests that enumerate externals.
|
||||
5. Materialize the three files with the curl; run criteria 1-8; report the emil citers you deliberately did not mirror and why.
|
||||
@@ -0,0 +1,49 @@
|
||||
# CONTRACT — floor-guard
|
||||
- date: 2026-09-27 | flow: feat (ad-hoc dispatch, /feat gates replayed by the orchestrator) | branch: feature/agent-skills-borrow
|
||||
- status: active
|
||||
|
||||
## REQUEST (verbatim — IMMUTABLE)
|
||||
> Build `lib/floor-guard.sh`, a diff-scoped deterministic detector of a quietly weakened quality bar, adapted from addyosmani/agent-skills `constraint-driven-development` (floor guard) to this repo's gate model. Usage `bash ~/.claude/lib/floor-guard.sh <base-ref> [-- <pathspec>...]` over `git diff <base-ref>` (working tree included). Kinds: SUPPRESS (added checker silencing: `@ts-ignore`, `@ts-expect-error` without a trailing reason, `eslint-disable*`, `# noqa`, `# type: ignore`, `nosemgrep`, `nosec`, `shellcheck disable`), SKIP (added `.skip(`, `.only(`, `xit(`, `xdescribe(`, `fit(`, `fdescribe(`, `it.todo(`, `@pytest.mark.skip`, `@unittest.skip`, `t.Skip(` in test files), DELETED_TEST (deleted file whose path matches a test pattern), ASSERT_DROP (a test file whose assertion-line count decreases: `expect(`, `assert`, `should`, `.toBe`), STUB (added `not implemented`, `NotImplementedError`, empty `catch` block, `except: pass`), THRESHOLD_DOWN (a numeric value decreased on the same key in coverage/quality config files: `jest.config*`, `vitest.config*`, `.nycrc*`, `codecov*`, `sonar-project.properties`, `lighthouserc*`, `CONSTRAINTS.md`). Waiver: an added line carrying `floor-guard: allow <reason>` is printed as WAIVED and not counted. Output: one `FLOOR <KIND> <file>:<line> <snippet>` per finding, then `FLOOR GUARD: clean` (rc 0) or `FLOOR GUARD: <n> finding(s), <m> waived` (rc 2); rc 3 on usage error. Wire it as a mandatory verifier step (agents/verifier.md) and document it in lib/verify-secure-loop.md GATE 1; hermetic suite `lib/tests/floor-guard.test.sh`. User go 2026-09-27 ("ok pour les 4", case 2 item 2).
|
||||
|
||||
## CLARIFICATIONS
|
||||
- Language: bash entry point; the diff parsing may live in an embedded python3 heredoc (precedent `lib/tests/run-review-guards.sh`). Functions <= 25 logic lines, 80-char lines.
|
||||
- File classes: test file = path contains `test`, `spec`, `__tests__`, or matches `*.test.*`, `*.spec.*`, `*_test.go`, `*_test.py`, `test_*.py`; config file = the names listed under THRESHOLD_DOWN. SKIP and ASSERT_DROP apply to test files only; SUPPRESS and STUB to any file; THRESHOLD_DOWN to config files only.
|
||||
- The guard's own pattern table contains the trigger strings: those source lines carry `# floor-guard: allow pattern table` so the guard stays clean on itself (this also exercises the waiver path for real).
|
||||
- Verifier step: `bash ~/.claude/lib/floor-guard.sh <base>` where base = the branch's gitflow base (develop; main for hotfix/release); rc 2 → verdict ECARTS listing each FLOOR line, unless the contract's CLARIFICATIONS explicitly authorize that exact weakening (quote it in the verdict).
|
||||
- Suite: throwaway repos (`make test` exports GIT_CONFIG_GLOBAL=/dev/null); one RED fixture per kind, one WAIVED fixture, one CLEAN fixture, each printed as `PASS <KIND>`; summary `PASS=n FAIL=m` like the other suites.
|
||||
- Self-detection: the suite file itself contains the trigger strings; the self-run criterion excludes it by pathspec.
|
||||
- Out of scope: a pre-commit hook, running it on the repo's own diff in `make test`, parsers beyond the regexes above.
|
||||
|
||||
## ACCEPTANCE CRITERIA
|
||||
1. Suite green, every kind flip-tested.
|
||||
CHECK: out=$(make test suite=lib/tests/floor-guard.test.sh 2>&1); for k in SUPPRESS SKIP DELETED_TEST ASSERT_DROP STUB THRESHOLD_DOWN WAIVED CLEAN; do echo "$out" | grep -q "PASS $k" || { echo "missing PASS $k"; echo "$out" | tail -15; exit 1; }; done; echo "$out" | grep -qE "FAIL=[1-9]" && exit 1; echo KINDS_GREEN
|
||||
EXPECT: KINDS_GREEN
|
||||
EVIDENCE: MET exit=0 marker-found :: KINDS_GREEN
|
||||
2. Usage error is rc 3.
|
||||
CHECK: bash lib/floor-guard.sh >/dev/null 2>&1; [ $? -eq 3 ] && echo RC_USAGE
|
||||
EXPECT: RC_USAGE
|
||||
EVIDENCE: MET exit=0 marker-found :: RC_USAGE
|
||||
3. Self-run clean on this branch (suite file excluded by pathspec), waivers visible.
|
||||
CHECK: out=$(bash lib/floor-guard.sh develop -- . ':!lib/tests/floor-guard.test.sh' 2>&1); rc=$?; echo "$out" | tail -3; [ $rc -eq 0 ] && echo "$out" | grep -q WAIVED && echo SELF_CLEAN
|
||||
EXPECT: SELF_CLEAN
|
||||
EVIDENCE: MET exit=0 marker-found :: WAIVED STUB lib/floor-guard.sh:105 'not implemented', # floor-guard: allow pattern table WAIVED STUB lib/floor-guard.sh:106 'NotImplementedE…
|
||||
4. Verifier wired, loop documented.
|
||||
CHECK: grep -q "floor-guard.sh" agents/verifier.md && grep -q "floor-guard" lib/verify-secure-loop.md && echo WIRED
|
||||
EXPECT: WIRED
|
||||
EVIDENCE: MET exit=0 marker-found :: WIRED
|
||||
5. shellcheck and doctrine-citers clean.
|
||||
CHECK: shellcheck lib/floor-guard.sh lib/tests/floor-guard.test.sh && out=$(make test suite=lib/tests/doctrine-citers.test.sh 2>&1) && ! echo "$out" | grep -qE "FAIL=[1-9]" && echo LINT_OK
|
||||
EXPECT: LINT_OK
|
||||
EVIDENCE: MET exit=0 marker-found :: LINT_OK
|
||||
|
||||
## FILE SCOPE
|
||||
- lib/floor-guard.sh (new), lib/tests/floor-guard.test.sh (new)
|
||||
- agents/verifier.md (one mandatory step), lib/verify-secure-loop.md (one paragraph under GATE 1)
|
||||
- CHANGELOG.md (Unreleased entry)
|
||||
|
||||
## PLAN
|
||||
1. Script: arg parsing (base, optional `--` pathspec), `git diff --unified=0 <base> -- <pathspec>` plus `git diff --diff-filter=D --name-only <base> -- <pathspec>`, rc contract, header comment stating WHY (BDR-100 class: deterministic floor under an LLM gate).
|
||||
2. Python parser on stdin: walk hunks, classify added lines by kind and file class, count assertion lines removed vs added per test file, compare numeric values on identical keys in config files (`key: 80` → `key: 60`, JSON or YAML-ish), detect waivers.
|
||||
3. Output lines + summary + rc.
|
||||
4. Suite: helper `mk_repo` (git init -q, base commit with a test file holding 3 assertions, a `vitest.config.ts` with `lines: 80`, a source file), one fixture per kind → assert rc 2 and the FLOOR line; WAIVED → rc 0 and a WAIVED line; CLEAN → rc 0.
|
||||
5. verifier.md step + verify-secure-loop.md paragraph + CHANGELOG; run criteria 1-5.
|
||||
@@ -0,0 +1,34 @@
|
||||
# CONTRACT — rest-api-rule
|
||||
- date: 2026-09-27 | flow: feat (ad-hoc dispatch, /feat gates replayed by the orchestrator) | branch: feature/agent-skills-borrow
|
||||
- status: active
|
||||
|
||||
## REQUEST (verbatim — IMMUTABLE)
|
||||
> Write `rules/rest-api.md`, a path-scoped user rule distilled from addyosmani/agent-skills `api-and-interface-design` (commit 2686b620fc1fed2e8f60c704839c766b8594c6b6), the way `rules/web-building.md` is written: `paths:` frontmatter, English, <= 45 lines, 80-char lines. Keep: contract-first order (typed interface → schemas with server-generated fields apart → error codes → validate at boundaries only); one error envelope `{ error: { code, message, details? } }` with the HTTP map 400 invalid / 401 auth / 403 forbidden / 404 missing / 409 conflict / 422 semantic / 500 server; every list endpoint paginated (`page`, `pageSize`, `totalItems`, `totalPages`), filters as query params; idempotency (key derived from intent, atomic claim via unique constraint, payload guard, explicit in-flight duplicate policy 409 / wait / 202, retention beyond the longest retry path incl. dead-letter); naming (plural nouns, camelCase params and fields, UPPER_SNAKE enums, is/has/can booleans); one Hyrum's law line. Drop the upstream one-version rule: versioning points to the CLAUDE.md heading "Web APIs — always versioned". User go 2026-09-27 ("ok pour les 4", case 2 item 4).
|
||||
|
||||
## CLARIFICATIONS
|
||||
- Globs: `["**/api/**", "**/routes/**", "**/controllers/**", "**/*.route.*", "**/*.controller.*", "**/openapi.*", "**/*.openapi.*"]`.
|
||||
- Fetch the upstream text with curl at the pinned commit to distill from; never vendor it.
|
||||
- Cite the doctrine as `CLAUDE.md § Web APIs — always versioned` (exact heading, so the doctrine-citers census resolves it).
|
||||
- No routing line in CLAUDE.global.md, no README change, rules/README.md unchanged; CHANGELOG Unreleased entry.
|
||||
|
||||
## ACCEPTANCE CRITERIA
|
||||
1. Shape: frontmatter, paths JSON list, <= 45 lines, body lines <= 80 chars (the one-line `paths:` frontmatter is exempt, repo precedent: web-building.md / web-security.md line 2).
|
||||
CHECK: f=rules/rest-api.md; [ -f "$f" ] && [ "$(head -1 "$f")" = "---" ] && python3 -c 'import re,json; s=open("rules/rest-api.md").read(); m=re.search(r"^paths: (.*)$",s,re.M); assert len(json.loads(m.group(1)))>=5' && [ "$(wc -l < "$f")" -le 45 ] && ! awk 'NR>3 && length>80' "$f" | grep -q . && echo RULE_SHAPE
|
||||
EXPECT: RULE_SHAPE
|
||||
EVIDENCE: MET exit=0 marker-found :: RULE_SHAPE
|
||||
2. Content present, one-version rule absent.
|
||||
CHECK: f=rules/rest-api.md; ok=1; for k in "code" "message" "409" "422" "pageSize" "totalItems" "idempoten" "unique" "camelCase" "UPPER_SNAKE" "Hyrum" "always versioned"; do grep -qi -- "$k" "$f" || { echo "missing $k"; ok=0; }; done; grep -qi "extend rather than fork" "$f" && { echo "one-version rule present"; ok=0; }; [ "$ok" -eq 1 ] && echo RULE_CONTENT
|
||||
EXPECT: RULE_CONTENT
|
||||
EVIDENCE: MET exit=0 marker-found :: RULE_CONTENT
|
||||
3. Doctrine citation resolves.
|
||||
CHECK: out=$(make test suite=lib/tests/doctrine-citers.test.sh 2>&1); echo "$out" | grep -qE "FAIL=[1-9]" && { echo "$out" | tail -10; exit 1; }; echo CITERS_OK
|
||||
EXPECT: CITERS_OK
|
||||
EVIDENCE: MET exit=0 marker-found :: CITERS_OK
|
||||
|
||||
## FILE SCOPE
|
||||
- rules/rest-api.md (new), CHANGELOG.md (Unreleased entry)
|
||||
|
||||
## PLAN
|
||||
1. curl the upstream SKILL.md at the pin into the scratch dir; read it.
|
||||
2. Write the rule: title "REST API — contract, errors, lists, idempotency"; sections Contract first · Errors · Lists · Idempotency · Naming · Versioning (one line pointing to the doctrine heading).
|
||||
3. CHANGELOG line; run criteria 1-3.
|
||||
@@ -0,0 +1,42 @@
|
||||
# CONTRACT — skill-routing-census
|
||||
- date: 2026-09-27 | flow: feat (ad-hoc dispatch, /feat gates replayed by the orchestrator) | branch: feature/agent-skills-borrow
|
||||
- status: active
|
||||
|
||||
## REQUEST (verbatim — IMMUTABLE)
|
||||
> Build `lib/tests/skill-routing-census.test.sh`: a deterministic census of skill-description collisions across the live catalog, adapted from addyosmani/agent-skills evals Tier 2. Catalog = every `SKILL.md` under `~/.claude/skills/*/` (symlinks resolved) plus plugin skills under `~/.claude/plugins/cache/*/*/*/skills/*/SKILL.md` and `~/.claude/plugins/cache/*/*/*/.claude/skills/*/SKILL.md` (roots overridable via `SKILL_ROUTING_ROOTS`, colon-separated dirs, for fixtures). Extract `description` (scalar or `|`/`>` block), tokenize (lowercase, `[a-z][a-z0-9-]+`, stopwords, suffix stemming s/es/ed/ing), TF-IDF cosine over all pairs. Print `skills with description: N`, the top 10 pairs as `0.52 a ~ b`, WARN lines for pairs >= `SKILL_ROUTING_WARN` (default 0.50), FAIL for pairs >= `SKILL_ROUTING_FAIL` (default 0.75). Self-test on fixtures: a near-duplicate pair must FAIL (`FIXTURE_COLLISION_DETECTED`), a distinct pair must pass (`FIXTURE_DISTINCT_OK`). Measured today: 120 skills, max 0.52 (careful ~ guard), 0 pairs >= 0.75 → green with one WARN. User go 2026-09-27 ("ok pour les 4", case 2 item 3).
|
||||
|
||||
## CLARIFICATIONS
|
||||
- python3 embedded in the bash suite is allowed (precedent run-review-guards.sh). If the python body exceeds ~120 lines, put it in `lib/skill-routing-census.py` and keep the suite as the wrapper.
|
||||
- Reference implementation (read it, reuse the logic, harden the block-description parsing): `/tmp/claude-1000/-home-bchanot-Documents-claude/977f1703-f01d-497a-b794-5b69fafcd35f/scratchpad/census.py`.
|
||||
- Thresholds stay at the upstream defaults; no allowlist file; a WARN never fails the suite.
|
||||
- Dedup by skill directory name: first path wins.
|
||||
- Out of scope (follow-up): positive/negative prompt ranking per skill.
|
||||
- The live census is machine-dependent by design (it audits this machine's catalog); the fixture self-test is the hermetic part. On a machine with an empty catalog the live pass prints `skills with description: 0` and passes.
|
||||
|
||||
## ACCEPTANCE CRITERIA
|
||||
1. Suite green on the live catalog, catalog non-trivial here.
|
||||
CHECK: out=$(make test suite=lib/tests/skill-routing-census.test.sh 2>&1); echo "$out" | grep -qE "FAIL=[1-9]" && { echo "$out" | tail -15; exit 1; }; echo "$out" | grep -qE "skills with description: *[0-9]{2,}" && echo LIVE_GREEN
|
||||
EXPECT: LIVE_GREEN
|
||||
EVIDENCE: MET exit=0 marker-found :: LIVE_GREEN
|
||||
2. Fixture flip: collision detected, distinct pair passes.
|
||||
CHECK: out=$(make test suite=lib/tests/skill-routing-census.test.sh 2>&1); echo "$out" | grep -q "FIXTURE_COLLISION_DETECTED" && echo "$out" | grep -q "FIXTURE_DISTINCT_OK" && echo FLIP_TESTED
|
||||
EXPECT: FLIP_TESTED
|
||||
EVIDENCE: MET exit=0 marker-found :: FLIP_TESTED
|
||||
3. Report shape: top pairs with two-decimal scores.
|
||||
CHECK: out=$(make test suite=lib/tests/skill-routing-census.test.sh 2>&1); echo "$out" | grep -qE "^0\.[0-9]{2} [a-z0-9:_-]+ ~ [a-z0-9:_-]+" && echo REPORT_SHAPE
|
||||
EXPECT: REPORT_SHAPE
|
||||
EVIDENCE: MET exit=0 marker-found :: REPORT_SHAPE
|
||||
4. shellcheck clean.
|
||||
CHECK: shellcheck lib/tests/skill-routing-census.test.sh && echo SHELLCHECK_OK
|
||||
EXPECT: SHELLCHECK_OK
|
||||
EVIDENCE: MET exit=0 marker-found :: SHELLCHECK_OK
|
||||
|
||||
## FILE SCOPE
|
||||
- lib/tests/skill-routing-census.test.sh (new); optional lib/skill-routing-census.py (new)
|
||||
- CHANGELOG.md (Unreleased entry)
|
||||
|
||||
## PLAN
|
||||
1. Roots: default globs; `SKILL_ROUTING_ROOTS` override; dedup by dir name.
|
||||
2. Description extraction: scalar `description: text`; block `description: |` or `>` → join the indented continuation lines.
|
||||
3. Tokens / TF-IDF / cosine as in census.py: stopword set, stem(), (1+log tf)·log(N/df), L2 norm, cosine over combinations.
|
||||
4. Report + thresholds + rc. Suite: run live (a FAIL line → suite RED), then two fixture dirs under mktemp: A/B near-duplicate descriptions → expect a FAIL line → print FIXTURE_COLLISION_DETECTED; C/D distinct → expect rc 0 → print FIXTURE_DISTINCT_OK. `PASS=n FAIL=m` summary like the other suites.
|
||||
+12
@@ -65,6 +65,9 @@ skills/ios-sync
|
||||
skills/design-motion-principles
|
||||
skills/emil-design-eng
|
||||
skills/frontend-design
|
||||
skills/ci-cd-and-automation
|
||||
skills/deprecation-and-migration
|
||||
skills/observability-and-instrumentation
|
||||
|
||||
# Impeccable — NOT a symlink: `impeccable skills install --scope=global`
|
||||
# writes the skill dir (and its ~15 MB engine binary) straight in through the
|
||||
@@ -177,6 +180,15 @@ skills-external/frontend-design/
|
||||
# an edit. The source is always re-fetched, so no offline copy is needed.
|
||||
skills-external/emil-design-eng/
|
||||
|
||||
# Agent Skills trio (addyosmani/agent-skills) — machine-owned, curl'd at the
|
||||
# commit pinned in plugins.lock.json ("agent-skills" entry) by
|
||||
# install-plugins.sh Step 8e (when absent) and re-fetched at the SAME commit
|
||||
# by update-all.sh. Not vendored: this is a pin, not a tracked snapshot —
|
||||
# bump the commit deliberately to pick up an upstream edit.
|
||||
skills-external/observability-and-instrumentation/
|
||||
skills-external/deprecation-and-migration/
|
||||
skills-external/ci-cd-and-automation/
|
||||
|
||||
# 21st.dev skill pack — machine-owned: `21st skills install` output, staged by
|
||||
# install-plugins.sh Step 8.7 (the installer refuses to write through the
|
||||
# ~/.claude/skills symlink, so it runs under a throwaway HOME and the skills
|
||||
|
||||
@@ -7,6 +7,38 @@ Format follows [Keep a Changelog](https://keepachangelog.com/).
|
||||
## [Unreleased]
|
||||
|
||||
### Added
|
||||
- **Agent Skills trio** (`observability-and-instrumentation`,
|
||||
`deprecation-and-migration`, `ci-cd-and-automation`) — vendored from
|
||||
addyosmani/agent-skills at a pinned commit (`agent-skills` entry in
|
||||
plugins.lock.json), the emil-design-eng way: curl'd into
|
||||
`skills-external/<name>/SKILL.md` by install-plugins.sh Step 8e, refreshed
|
||||
by update-all.sh at the same commit, symlinked by link.sh, registered in
|
||||
`lib/toggle-external.sh`, `lib/profile.sh` and the `full`/`backend`/`dev`
|
||||
profiles. Case 2 of the 6-repo review: the plugin itself was rejected
|
||||
(1.8k tokens per session for 20 % novelty, `/spec` `/review` `/ship`
|
||||
collide with gstack, trunk-based git and the one-version API rule
|
||||
contradict the doctrine, a second skill router).
|
||||
- **`lib/floor-guard.sh`** — diff-scoped deterministic detector of a quietly
|
||||
weakened quality bar (new lint/type suppressions, skipped or deleted
|
||||
tests, dropped assertions, stubs, lowered coverage thresholds), with a
|
||||
`floor-guard: allow <reason>` waiver, rc 0/2/3. Mandatory verifier
|
||||
STEP 3 (`agents/verifier.md`), documented under GATE 1 of
|
||||
`lib/verify-secure-loop.md`. Suite `lib/tests/floor-guard.test.sh`: 6
|
||||
kinds plus a WAIVED and a CLEAN fixture, each flip-tested. Waivers
|
||||
outside test files count as gaps unless the contract's CLARIFICATIONS
|
||||
names them (security-gate MEDIUM, user chose strict). Adapted from
|
||||
agent-skills `constraint-driven-development`.
|
||||
- **`lib/tests/skill-routing-census.test.sh`** (+ `lib/skill-routing-census.py`)
|
||||
— TF-IDF cosine census of skill-description collisions across the live
|
||||
catalog (routing ambiguity, not naming): top 10 pairs, WARN ≥ 0.50,
|
||||
FAIL ≥ 0.75, fixture flip-test. Baseline 2026-09-27: 120 skills, max 0.52
|
||||
(`careful` ~ `guard`). Adapted from agent-skills evals Tier 2.
|
||||
- **`rules/rest-api.md`** — path-scoped REST rule distilled from agent-skills
|
||||
`api-and-interface-design`: contract-first order, one error envelope +
|
||||
HTTP map, paginated lists, idempotency (key from intent, atomic claim,
|
||||
payload guard, duplicate policy, retention), naming, Hyrum's law;
|
||||
versioning points to `CLAUDE.md § Web APIs — always versioned` instead of
|
||||
the upstream one-version rule.
|
||||
- **`make test suite=<file>`** runs one suite hermetically; the
|
||||
`GIT_CONFIG_GLOBAL=/dev/null` export lives in the Makefile so nobody types
|
||||
the denied env-prefix form by hand (the reason an executor wrote a wrapper
|
||||
|
||||
+39
-7
@@ -71,7 +71,33 @@ You may re-run a `CHECK:` yourself to settle a doubt (Bash is read-only, and
|
||||
these commands are observation). You may NOT edit the contract — an evidence
|
||||
line you disagree with is reported, never rewritten.
|
||||
|
||||
## STEP 3 — SCOPE CHECK
|
||||
## STEP 3 — FLOOR GUARD (mandatory, deterministic)
|
||||
|
||||
Run the floor guard over the diff before rendering any verdict — a red or
|
||||
skipped run here is a structural gap, never a judgment call:
|
||||
|
||||
```bash
|
||||
bash ~/.claude/lib/floor-guard.sh <base> -- <pathspec>...
|
||||
```
|
||||
|
||||
`<base>` = the branch's gitflow base (develop; main for a hotfix/release).
|
||||
Parse the single `FLOOR GUARD:` line:
|
||||
|
||||
- `clean` (rc 0) → no unwaived finding; still apply the WAIVED rule below.
|
||||
- `<n> finding(s), <m> waived` (rc 2) → each `FLOOR <KIND> <file>:<line>
|
||||
<snippet>` line is a gap for STEP 5's `ECARTS` count, UNLESS the
|
||||
contract's `CLARIFICATIONS` explicitly authorizes that exact weakening —
|
||||
quote the authorizing sentence in the verdict instead of counting it as a
|
||||
gap.
|
||||
- `WAIVED <KIND> <file>:<line>` lines (either rc): on a test file (path
|
||||
holds `test`, `spec` or `__tests__`) they are informational. Anywhere
|
||||
else the waiver is self-service by construction, so it is a gap UNLESS
|
||||
the contract's `CLARIFICATIONS` names that file and the reason — quote
|
||||
it. The tool prints, the contract authorizes, the verifier counts.
|
||||
- rc 3 (usage error) → a structural failure like a missing contract: retry
|
||||
once (base ref or pathspec likely wrong), a second failure escalates.
|
||||
|
||||
## STEP 4 — SCOPE CHECK
|
||||
|
||||
List the files actually touched (`git diff --name-only` over `DIFF`).
|
||||
Compare against the contract's `FILE SCOPE`. Report every out-of-scope
|
||||
@@ -79,7 +105,7 @@ file. Disposition is NOT your call: the orchestrator treats each one as a
|
||||
gap — the dev removes it or justifies it, and an accepted justification
|
||||
only enters the contract through a human micro-gate.
|
||||
|
||||
## STEP 4 — VERDICT
|
||||
## STEP 5 — VERDICT
|
||||
|
||||
Read the contract's `ABANDON:` lines. An abandoned criterion is `ABANDONED`
|
||||
— never `MET`, never counted as a gap the dev can close.
|
||||
@@ -89,11 +115,12 @@ is not:
|
||||
|
||||
1. `ERROR(<reason>)` — the contract is missing or unreadable.
|
||||
2. `ECARTS(n)` — n = count(NOT-MET) + count(UNVERIFIABLE) + count(out-of-scope
|
||||
files). Surface any abandonment in the same report.
|
||||
files) + count(unauthorized FLOOR findings from STEP 3). Surface any
|
||||
abandonment in the same report.
|
||||
3. `ABANDONED(n)` — zero gaps remain, but n abandonments stand. This is NOT
|
||||
a pass and NOT a dev loop: it routes straight to the human gate.
|
||||
4. `CONFORME` — ALL criteria `MET`, zero out-of-scope files, zero
|
||||
abandonments.
|
||||
unauthorized FLOOR findings, zero abandonments.
|
||||
|
||||
## OUTPUT (exact format — machine-parsed by the orchestrator)
|
||||
|
||||
@@ -106,6 +133,8 @@ CRITERIA:
|
||||
3. <criterion> — UNVERIFIABLE — <reason>
|
||||
4. <criterion> — ABANDONED — <the reason recorded in the contract>
|
||||
SCOPE: in-scope <n> files; out-of-scope: <list | none>
|
||||
FLOOR: clean | <n> finding(s) (<m> waived) — <FLOOR lines, or the
|
||||
CLARIFICATIONS sentence that authorizes each one | none>
|
||||
PROOF: read <n> files, ran <cmd → result | nothing>, checked <n>/<n> criteria
|
||||
```
|
||||
|
||||
@@ -123,6 +152,9 @@ PROOF: read <n> files, ran <cmd → result | nothing>, checked <n>/<n> criteria
|
||||
- `PROOF` is MANDATORY. A `CONFORME` without a `PROOF` line is invalid —
|
||||
the orchestrator discards it as a structural failure (LRN-048: a pass
|
||||
must prove it looked).
|
||||
- STEP 3's floor-guard run is MANDATORY, every dispatch. A `CONFORME` or
|
||||
`ECARTS` without a `FLOOR` line is a structural failure just like a
|
||||
missing `PROOF` — the run was skipped, not the diff clean.
|
||||
- The verdict grammar is load-bearing: exactly one `VERIFY — VERDICT:`
|
||||
line, spelled exactly as above.
|
||||
|
||||
@@ -145,8 +177,8 @@ loop, never here):
|
||||
lifts the abandonment (the criterion was fixable after all) or accepts
|
||||
the partial delivery; the run is never reported as fully complete.
|
||||
- Structural failure (`ERROR(…)`, missing/duplicated VERDICT line,
|
||||
unparsable output, agent crash, `CONFORME` without `PROOF`) → retry
|
||||
ONCE with a fresh verifier; a 2nd structural failure → human
|
||||
escalation. A mute verifier is NEVER a PASS.
|
||||
unparsable output, agent crash, `CONFORME` without `PROOF` or without
|
||||
`FLOOR`) → retry ONCE with a fresh verifier; a 2nd structural failure →
|
||||
human escalation. A mute verifier is NEVER a PASS.
|
||||
- After a security-gate fix round: re-verify the request FIRST (this
|
||||
agent), THEN re-verify security — in that order.
|
||||
|
||||
@@ -90,6 +90,22 @@ print(v)
|
||||
fi
|
||||
}
|
||||
|
||||
# Read a pinned commit sha from plugins.lock.json (agent-skills style entries
|
||||
# — no "version", a "commit" field instead). Prints the sha, or "" if the
|
||||
# entry or the field is absent.
|
||||
# Usage: pinned_commit "agent-skills" → prints the commit sha or ""
|
||||
pinned_commit() {
|
||||
local key="$1"
|
||||
if [ -f "$REPO/plugins.lock.json" ] && command -v python3 &>/dev/null; then
|
||||
python3 -c "
|
||||
import json, sys
|
||||
with open(sys.argv[1]) as f:
|
||||
d = json.load(f)
|
||||
print(d.get(sys.argv[2], {}).get('commit', ''))
|
||||
" "$REPO/plugins.lock.json" "$key" 2>/dev/null || true
|
||||
fi
|
||||
}
|
||||
|
||||
# ============================================================
|
||||
# DETECT OS
|
||||
# ============================================================
|
||||
@@ -899,6 +915,43 @@ else
|
||||
fi
|
||||
echo ""
|
||||
|
||||
# ── Step 8e: Agent Skills (addyosmani/agent-skills, pinned commit) ──
|
||||
# Three dev-lifecycle skills vendored the emil-design-eng way (curl →
|
||||
# skills-external/<name>/SKILL.md, symlinked by link.sh) but COMMIT-pinned
|
||||
# instead of tracking main: the sha lives in plugins.lock.json ("agent-skills"
|
||||
# entry), never hardcoded here.
|
||||
echo "── Step 8e: Agent Skills (addyosmani/agent-skills) ─────────"
|
||||
echo ""
|
||||
AGENT_SKILLS_NAMES=(observability-and-instrumentation deprecation-and-migration ci-cd-and-automation)
|
||||
AGENT_SKILLS_SHA=$(pinned_commit "agent-skills")
|
||||
if [ -z "$AGENT_SKILLS_SHA" ]; then
|
||||
err "agent-skills: no commit pinned in plugins.lock.json — add an \"agent-skills\" entry with a \"commit\" field"
|
||||
else
|
||||
for _as_skill in "${AGENT_SKILLS_NAMES[@]}"; do
|
||||
_as_dir="$REPO/skills-external/$_as_skill"
|
||||
_as_url="https://raw.githubusercontent.com/addyosmani/agent-skills/$AGENT_SKILLS_SHA/skills/$_as_skill/SKILL.md"
|
||||
mkdir -p "$_as_dir"
|
||||
if [ -f "$_as_dir/SKILL.md" ]; then
|
||||
ok "$_as_skill already downloaded"
|
||||
else
|
||||
info "Downloading SKILL.md from addyosmani/agent-skills ($_as_skill)..."
|
||||
if curl -fsSL "$_as_url" -o "$_as_dir/SKILL.md.tmp" \
|
||||
&& mv "$_as_dir/SKILL.md.tmp" "$_as_dir/SKILL.md"; then
|
||||
ok "$_as_skill installed"
|
||||
else
|
||||
rm -f "$_as_dir/SKILL.md.tmp"
|
||||
err "$_as_skill download failed — try: curl -fsSL $_as_url -o $_as_dir/SKILL.md"
|
||||
fi
|
||||
fi
|
||||
if [ -L "$HOME/.claude/skills/$_as_skill" ]; then
|
||||
ok "$_as_skill symlink OK"
|
||||
else
|
||||
info "Symlinking $_as_skill — will be created by link.sh"
|
||||
fi
|
||||
done
|
||||
fi
|
||||
echo ""
|
||||
|
||||
# ============================================================
|
||||
# STEP 8.5 — EXTERNAL SKILLS (npx skills add …)
|
||||
# ============================================================
|
||||
@@ -1186,6 +1239,7 @@ echo " 🔄 emil-design-eng — UI polish, animations, component craft (c
|
||||
echo " 🔄 frontend-design — distinctive frontend interfaces, anti-AI-slop (anthropic-agent-skills)"
|
||||
echo " 🔄 impeccable — /impeccable design verbs + 45-rule deterministic detector (npx impeccable detect)"
|
||||
echo " 🔄 design-motion-principles — motion/animation design, 3-designer lens (kylezantos)"
|
||||
echo " 🔄 agent-skills trio — observability-and-instrumentation, deprecation-and-migration, ci-cd-and-automation (curl → symlink, pinned commit)"
|
||||
echo " 🔄 darwin-skill — autonomous skill optimizer (npx skills, ~/.agents/skills/)"
|
||||
echo " 🔄 21st skill pack — 21st.dev CLI skills; design ones follow the profile (full by default), publishing ones on demand (toggle: lib/toggle-external.sh enable 21st)"
|
||||
echo ""
|
||||
@@ -1194,6 +1248,7 @@ echo " GStack skills symlinked individually into ~/.claude/skills/ (→ submodu
|
||||
echo " Emil Design Eng at: ~/.claude/skills/emil-design-eng/ (symlink → skills-external)"
|
||||
echo " Frontend Design at: ~/.claude/skills/frontend-design/ (symlink → skills-external)"
|
||||
echo " Design Motion Principles at: ~/.claude/skills/design-motion-principles/ (symlink → skills-external)"
|
||||
echo " Agent Skills trio at: ~/.claude/skills/{observability-and-instrumentation,deprecation-and-migration,ci-cd-and-automation}/ (symlink → skills-external)"
|
||||
echo " npx skills at: ~/.agents/skills/ (symlinked into ~/.claude/skills/)"
|
||||
echo ""
|
||||
echo " → Restart Claude Code — plugins load automatically"
|
||||
|
||||
Executable
+297
@@ -0,0 +1,297 @@
|
||||
#!/usr/bin/env bash
|
||||
# lib/floor-guard.sh — diff-scoped detector of a quietly weakened quality bar.
|
||||
#
|
||||
# bash ~/.claude/lib/floor-guard.sh <base-ref> [-- <pathspec>...]
|
||||
#
|
||||
# rc 0 = clean no floor finding in the diff
|
||||
# 2 = <n> finding(s), <m> waived
|
||||
# 3 = usage error (missing <base-ref>, or it does not resolve to a commit)
|
||||
#
|
||||
# WHY (BDR-100 class): "no weakened check in this diff" is exactly the kind
|
||||
# of judgment an LLM verifier can miss, or be talked past one line at a time
|
||||
# — a single added TS-ignore comment, a skipped test, a dropped assertion, a
|
||||
# coverage threshold shaved by one point. This makes that judgment
|
||||
# deterministic: grep the diff for the known ways a change quietly lowers
|
||||
# the bar, same floor doctrine as gates.sh (contract oracles) and
|
||||
# doctrine-citers.test.sh (citation census) — a mechanism, not a lesson.
|
||||
#
|
||||
# Adapted from addyosmani/agent-skills constraint-driven-development's
|
||||
# "floor guard" to this repo's own gate model: `git diff` instead of a
|
||||
# staged-diff assumption, wired into agents/verifier.md STEP 3 rather than a
|
||||
# pre-commit hook.
|
||||
#
|
||||
# Scope: `git diff <base-ref>` — working tree included (uncommitted changes
|
||||
# count) — restricted to <pathspec> when given. Every ADDED line is
|
||||
# classified into one of six kinds (full pattern tables below):
|
||||
# SUPPRESS a checker-silencing comment added, any file
|
||||
# SKIP a test disabled or isolated, test files only
|
||||
# DELETED_TEST a whole test file removed
|
||||
# ASSERT_DROP a test file's assertion-line count went down
|
||||
# STUB a not-implemented marker added, any file
|
||||
# THRESHOLD_DOWN a numeric value lowered on the same key, config files only
|
||||
#
|
||||
# Waiver: an added line also carrying `floor-guard: allow <reason>` prints as
|
||||
# WAIVED and does not count toward the finding total or the rc.
|
||||
set -uo pipefail
|
||||
|
||||
_usage() {
|
||||
echo "usage: floor-guard.sh <base-ref> [-- <pathspec>...]" >&2
|
||||
exit 3
|
||||
}
|
||||
|
||||
[ $# -ge 1 ] || _usage
|
||||
BASE_REF="$1"; shift
|
||||
PATHSPEC=()
|
||||
if [ $# -gt 0 ]; then
|
||||
[ "$1" = "--" ] || _usage
|
||||
shift
|
||||
PATHSPEC=("$@")
|
||||
fi
|
||||
git rev-parse --verify -q "${BASE_REF}^{commit}" >/dev/null 2>&1 || _usage
|
||||
|
||||
TMPDIFF="$(mktemp)" || { echo "floor-guard: mktemp failed" >&2; exit 3; }
|
||||
trap 'rm -f "$TMPDIFF"' EXIT
|
||||
|
||||
git diff --unified=0 "$BASE_REF" -- "${PATHSPEC[@]}" > "$TMPDIFF" 2>/dev/null
|
||||
FLOOR_DELETED_FILES="$(git diff --diff-filter=D --name-only \
|
||||
"$BASE_REF" -- "${PATHSPEC[@]}" 2>/dev/null)"
|
||||
export FLOOR_DELETED_FILES
|
||||
|
||||
# "working tree included" means brand-new, still-untracked files too: plain
|
||||
# `git diff <ref>` never shows them (git only diffs what it already tracks),
|
||||
# so a file added on this branch and never `git add`-ed would be invisible
|
||||
# to every kind below. --no-index against /dev/null emits the same unified
|
||||
# format as the tracked diff above (diff --git / +++ b/path / @@ hunks),
|
||||
# so the parser needs no separate code path for it.
|
||||
while IFS= read -r f; do
|
||||
[ -n "$f" ] || continue
|
||||
git diff --no-index --unified=0 -- /dev/null "$f" >> "$TMPDIFF" 2>/dev/null
|
||||
done < <(git ls-files --others --exclude-standard -- "${PATHSPEC[@]}" 2>/dev/null)
|
||||
|
||||
python3 - "$TMPDIFF" <<'PY'
|
||||
import fnmatch
|
||||
import os
|
||||
import re
|
||||
import sys
|
||||
|
||||
# file classes (CLARIFICATIONS): "path contains test/spec/__tests__" is a
|
||||
# superset of the explicit globs (*.test.*, *.spec.*, *_test.go, *_test.py,
|
||||
# test_*.py all contain one of these substrings themselves), so one check
|
||||
# covers all five.
|
||||
TEST_SUBSTRINGS = ('test', 'spec', '__tests__')
|
||||
CONFIG_GLOBS = ('jest.config*', 'vitest.config*', '.nycrc*', 'codecov*',
|
||||
'sonar-project.properties', 'lighthouserc*', 'CONSTRAINTS.md')
|
||||
|
||||
# ── pattern tables — the trigger strings themselves, waived on this file's
|
||||
# own diff so the guard stays clean on itself (also exercises the waiver
|
||||
# path for real) ────────────────────────────────────────────────────────────
|
||||
SUPPRESS_SUBSTRINGS = (
|
||||
'@ts-ignore', # floor-guard: allow pattern table
|
||||
'eslint-disable', # floor-guard: allow pattern table
|
||||
'# noqa', # floor-guard: allow pattern table
|
||||
'# type: ignore', # floor-guard: allow pattern table
|
||||
'nosemgrep', # floor-guard: allow pattern table
|
||||
'nosec', # floor-guard: allow pattern table
|
||||
'shellcheck disable', # floor-guard: allow pattern table
|
||||
)
|
||||
TS_EXPECT_ERROR = '@ts-expect-error' # floor-guard: allow pattern table
|
||||
|
||||
SKIP_SUBSTRINGS = (
|
||||
'.skip(', '.only(', 'xit(', 'xdescribe(', 'fit(', 'fdescribe(',
|
||||
'it.todo(', '@pytest.mark.skip', '@unittest.skip', 't.Skip(',
|
||||
)
|
||||
|
||||
STUB_SUBSTRINGS = (
|
||||
'not implemented', # floor-guard: allow pattern table
|
||||
'NotImplementedError', # floor-guard: allow pattern table
|
||||
)
|
||||
EMPTY_CATCH_RE = re.compile(r'catch\s*\([^)]*\)\s*\{\s*\}')
|
||||
BARE_EXCEPT_RE = re.compile(r'except\b[^:\n]*:\s*pass\b')
|
||||
|
||||
ASSERT_SUBSTRINGS = ('expect(', 'assert', 'should', '.toBe')
|
||||
|
||||
KEYVAL_RE = re.compile(r'["\']?([A-Za-z0-9_.\-]+)["\']?\s*[:=]\s*(-?\d+(?:\.\d+)?)')
|
||||
WAIVER_RE = re.compile(r'floor-guard:\s*allow\s+(\S.*)$')
|
||||
HUNK_RE = re.compile(r'^@@ -(\d+)(?:,\d+)? \+(\d+)(?:,\d+)? @@')
|
||||
|
||||
|
||||
def is_test_file(path):
|
||||
return any(s in path for s in TEST_SUBSTRINGS)
|
||||
|
||||
|
||||
def is_config_file(path):
|
||||
base = os.path.basename(path)
|
||||
return any(fnmatch.fnmatch(base, g) for g in CONFIG_GLOBS)
|
||||
|
||||
|
||||
def is_waived(text):
|
||||
return bool(WAIVER_RE.search(text))
|
||||
|
||||
|
||||
def strip_prefix(raw):
|
||||
if raw == '/dev/null':
|
||||
return raw
|
||||
return raw[2:] if raw[:2] in ('a/', 'b/') else raw
|
||||
|
||||
|
||||
def _new_file_entry(files):
|
||||
entry = {'old_path': None, 'new_path': None, 'adds': [], 'dels': [],
|
||||
'first_new': None}
|
||||
files.append(entry)
|
||||
return entry
|
||||
|
||||
|
||||
def parse_diff(lines): # → list of per-file entries (adds/dels + paths)
|
||||
files, cur = [], None
|
||||
old_no = new_no = 0
|
||||
for raw in lines:
|
||||
if raw.startswith('diff --git '):
|
||||
cur = _new_file_entry(files)
|
||||
elif raw.startswith('--- '):
|
||||
cur['old_path'] = strip_prefix(raw[4:])
|
||||
elif raw.startswith('+++ '):
|
||||
cur['new_path'] = strip_prefix(raw[4:])
|
||||
elif raw.startswith('@@ '):
|
||||
m = HUNK_RE.match(raw)
|
||||
if m:
|
||||
old_no, new_no = int(m.group(1)), int(m.group(2))
|
||||
if cur['first_new'] is None:
|
||||
cur['first_new'] = new_no
|
||||
elif raw.startswith('+') and not raw.startswith('+++'):
|
||||
cur['adds'].append((new_no, raw[1:])); new_no += 1
|
||||
elif raw.startswith('-') and not raw.startswith('---'):
|
||||
cur['dels'].append((old_no, raw[1:])); old_no += 1
|
||||
return files
|
||||
|
||||
|
||||
def effective_path(entry):
|
||||
if entry['new_path'] not in (None, '/dev/null'):
|
||||
return entry['new_path']
|
||||
return entry['old_path']
|
||||
|
||||
|
||||
def suppress_kind(text):
|
||||
if TS_EXPECT_ERROR in text:
|
||||
after = text.split(TS_EXPECT_ERROR, 1)[1].strip()
|
||||
return None if after else 'SUPPRESS'
|
||||
return 'SUPPRESS' if any(p in text for p in SUPPRESS_SUBSTRINGS) else None
|
||||
|
||||
|
||||
def stub_kind(text):
|
||||
if any(p in text for p in STUB_SUBSTRINGS):
|
||||
return 'STUB'
|
||||
if EMPTY_CATCH_RE.search(text) or BARE_EXCEPT_RE.search(text):
|
||||
return 'STUB'
|
||||
return None
|
||||
|
||||
|
||||
def skip_kind(text):
|
||||
return 'SKIP' if any(p in text for p in SKIP_SUBSTRINGS) else None
|
||||
|
||||
|
||||
def line_findings(path, lineno, text, test_file):
|
||||
out, waived = [], is_waived(text)
|
||||
for kindfn in (suppress_kind, stub_kind):
|
||||
kind = kindfn(text)
|
||||
if kind:
|
||||
out.append((kind, path, lineno, text, waived))
|
||||
if test_file:
|
||||
kind = skip_kind(text)
|
||||
if kind:
|
||||
out.append((kind, path, lineno, text, waived))
|
||||
return out
|
||||
|
||||
|
||||
def _is_assertion(text):
|
||||
return any(p in text for p in ASSERT_SUBSTRINGS)
|
||||
|
||||
|
||||
def assert_drop_finding(path, entry):
|
||||
added = sum(1 for _, t in entry['adds'] if _is_assertion(t))
|
||||
removed = sum(1 for _, t in entry['dels'] if _is_assertion(t))
|
||||
if removed <= added:
|
||||
return None
|
||||
lineno = entry['adds'][0][0] if entry['adds'] else (entry['first_new'] or 1)
|
||||
waived = any(is_waived(t) for _, t in entry['adds'])
|
||||
snippet = 'assertion lines %d -> %d' % (removed, added)
|
||||
return ('ASSERT_DROP', path, lineno, snippet, waived)
|
||||
|
||||
|
||||
def extract_kv(lines): # → {key: (lineno, value, raw text)} last-wins
|
||||
kv = {}
|
||||
for lineno, text in lines:
|
||||
m = KEYVAL_RE.search(text)
|
||||
if m:
|
||||
kv[m.group(1)] = (lineno, float(m.group(2)), text)
|
||||
return kv
|
||||
|
||||
|
||||
def threshold_down_findings(path, entry):
|
||||
removed_kv = extract_kv(entry['dels'])
|
||||
added_kv = extract_kv(entry['adds'])
|
||||
out = []
|
||||
for key, (lineno, new_val, text) in added_kv.items():
|
||||
old = removed_kv.get(key)
|
||||
if old and new_val < old[1]:
|
||||
out.append(('THRESHOLD_DOWN', path, lineno, text, is_waived(text)))
|
||||
return out
|
||||
|
||||
|
||||
def classify_file(entry, deleted_paths):
|
||||
path = effective_path(entry)
|
||||
if path is None:
|
||||
return [] # pure rename/mode-change: no --- / +++ header, no content diff
|
||||
test_file = is_test_file(path)
|
||||
if path in deleted_paths and test_file:
|
||||
return [('DELETED_TEST', path, 1, path, False)]
|
||||
findings = []
|
||||
for lineno, text in entry['adds']:
|
||||
findings += line_findings(path, lineno, text, test_file)
|
||||
if test_file:
|
||||
dropped = assert_drop_finding(path, entry)
|
||||
if dropped:
|
||||
findings.append(dropped)
|
||||
if is_config_file(path):
|
||||
findings += threshold_down_findings(path, entry)
|
||||
return findings
|
||||
|
||||
|
||||
def load_deleted_paths():
|
||||
raw = os.environ.get('FLOOR_DELETED_FILES', '')
|
||||
return {p for p in raw.splitlines() if p}
|
||||
|
||||
|
||||
def snippet_of(text):
|
||||
return text.strip()[:100]
|
||||
|
||||
|
||||
def emit(findings):
|
||||
ordered = sorted(findings, key=lambda f: (f[1], f[2], f[0]))
|
||||
n_found = n_waived = 0
|
||||
for kind, path, lineno, text, waived in ordered:
|
||||
tag = 'WAIVED' if waived else 'FLOOR'
|
||||
print('%s %s %s:%d %s' % (tag, kind, path, lineno, snippet_of(text)))
|
||||
n_waived += 1 if waived else 0
|
||||
n_found += 0 if waived else 1
|
||||
if n_found:
|
||||
print('FLOOR GUARD: %d finding(s), %d waived' % (n_found, n_waived))
|
||||
return 2
|
||||
print('FLOOR GUARD: clean')
|
||||
return 0
|
||||
|
||||
|
||||
def main():
|
||||
with open(sys.argv[1], 'r', errors='replace') as fh:
|
||||
lines = fh.read().split('\n')
|
||||
deleted = load_deleted_paths()
|
||||
findings = []
|
||||
for entry in parse_diff(lines):
|
||||
findings += classify_file(entry, deleted)
|
||||
return emit(findings)
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
sys.exit(main())
|
||||
PY
|
||||
rc=$?
|
||||
exit "$rc"
|
||||
+6
-1
@@ -79,6 +79,9 @@ MANAGED_EXTERNALS=(
|
||||
21st-ui-review
|
||||
21st-cli-use
|
||||
21st-ai
|
||||
observability-and-instrumentation
|
||||
deprecation-and-migration
|
||||
ci-cd-and-automation
|
||||
)
|
||||
|
||||
# MCP servers that are toggle-managed by `set`, both ways (enable AND
|
||||
@@ -756,7 +759,9 @@ NOTE:
|
||||
"set" toggles the MANAGED items automatically, both ways: plugins
|
||||
(ui-ux-pro-max, plugin-dev, pr-review-toolkit), external packs
|
||||
(emil-design-eng, frontend-design, design-motion-principles, impeccable,
|
||||
the five 21st design skills). Anything outside those allowlists stays
|
||||
the five 21st design skills, the agent-skills trio
|
||||
observability-and-instrumentation/deprecation-and-migration/
|
||||
ci-cd-and-automation). Anything outside those allowlists stays
|
||||
advisory — run "claude plugin enable|disable" or
|
||||
"bash lib/toggle-external.sh enable|disable <tool>" yourself.
|
||||
EOF
|
||||
|
||||
@@ -13,6 +13,11 @@ code-clean personal
|
||||
commit-change personal
|
||||
analyze personal
|
||||
|
||||
# Dev-lifecycle skills (agent-skills trio)
|
||||
observability-and-instrumentation external
|
||||
deprecation-and-migration external
|
||||
ci-cd-and-automation external
|
||||
|
||||
# Ship + review + land
|
||||
ship
|
||||
review
|
||||
|
||||
@@ -19,6 +19,11 @@ refactor personal
|
||||
code-clean personal
|
||||
commit-change personal
|
||||
|
||||
# Dev-lifecycle skills (agent-skills trio)
|
||||
observability-and-instrumentation external
|
||||
deprecation-and-migration external
|
||||
ci-cd-and-automation external
|
||||
|
||||
# Session hygiene
|
||||
context-save
|
||||
land-and-deploy
|
||||
|
||||
@@ -83,6 +83,9 @@ emil-design-eng external
|
||||
frontend-design external
|
||||
design-motion-principles external
|
||||
impeccable external
|
||||
observability-and-instrumentation external
|
||||
deprecation-and-migration external
|
||||
ci-cd-and-automation external
|
||||
ui-ux-pro-max plugin@ui-ux-pro-max-skill
|
||||
# pr-review-toolkit REMOVED from full (audit 2026-07-02 #12): heaviest
|
||||
# single plugin cost (~2.2k tokens of agent descriptions/session), useful
|
||||
|
||||
@@ -0,0 +1,213 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Deterministic census of skill-description collisions (TF-IDF cosine).
|
||||
|
||||
Adapted from addyosmani/agent-skills evals Tier 2. Reads every SKILL.md
|
||||
under the live catalog — `~/.claude/skills/*/` plus plugin skills under
|
||||
`~/.claude/plugins/cache/*/*/*/skills/*/` and the nested
|
||||
`.../.claude/skills/*/` layout (override with SKILL_ROUTING_ROOTS,
|
||||
colon-separated dirs, for fixtures) — extracts the `description`
|
||||
frontmatter field (scalar or `|`/`>` block), tokenizes it, and scores
|
||||
every pair by TF-IDF cosine. Two skills with near-identical descriptions
|
||||
route the same prompt to both: a routing collision, not a naming clash.
|
||||
|
||||
Report: `skills with description: N`, the top 10 pairs, then a WARN line
|
||||
per pair >= SKILL_ROUTING_WARN (default 0.50) and a FAIL line per pair
|
||||
>= SKILL_ROUTING_FAIL (default 0.75). Exit 1 iff any pair reached FAIL.
|
||||
"""
|
||||
import collections
|
||||
import glob
|
||||
import itertools
|
||||
import math
|
||||
import os
|
||||
import re
|
||||
import sys
|
||||
|
||||
DEFAULT_ROOTS = ["~/.claude/skills", "~/.claude/plugins/cache"]
|
||||
ROOT_SUFFIXES = [
|
||||
"*/SKILL.md",
|
||||
"*/*/*/skills/*/SKILL.md",
|
||||
"*/*/*/.claude/skills/*/SKILL.md",
|
||||
]
|
||||
STOPWORDS = set(
|
||||
"a an the and or of to in on for with use when you your this that is are "
|
||||
"be it as by from at into not no if then use using used uses skill "
|
||||
"skills user users project projects file files code work".split()
|
||||
)
|
||||
BLOCK_MARKERS = ("|", ">", "|-", ">-", "|+", ">+")
|
||||
WARN_DEFAULT = 0.50
|
||||
FAIL_DEFAULT = 0.75
|
||||
|
||||
|
||||
def resolve_roots():
|
||||
"""Root dirs: SKILL_ROUTING_ROOTS override, else the live catalog."""
|
||||
override = os.environ.get("SKILL_ROUTING_ROOTS")
|
||||
return override.split(":") if override else DEFAULT_ROOTS
|
||||
|
||||
|
||||
def discover_paths(roots):
|
||||
"""SKILL.md paths across all roots/suffixes, sorted for determinism."""
|
||||
paths = []
|
||||
for root in roots:
|
||||
expanded = os.path.expanduser(root)
|
||||
for suffix in ROOT_SUFFIXES:
|
||||
paths.extend(sorted(glob.glob(os.path.join(expanded, suffix))))
|
||||
return paths
|
||||
|
||||
|
||||
def dedup_by_skill_name(paths):
|
||||
"""name (skill dir) -> path; first path wins (symlinks pre-resolved)."""
|
||||
by_name = {}
|
||||
for path in paths:
|
||||
name = os.path.basename(os.path.dirname(path))
|
||||
by_name.setdefault(name, path)
|
||||
return by_name
|
||||
|
||||
|
||||
def _frontmatter(text):
|
||||
"""Raw text between the two leading `---` delimiters, or None."""
|
||||
match = re.search(r"^---\r?\n(.*?)\r?\n---", text, re.S)
|
||||
return match.group(1) if match else None
|
||||
|
||||
|
||||
def _strip_quotes(value):
|
||||
if len(value) >= 2 and value[0] == value[-1] and value[0] in "'\"":
|
||||
return value[1:-1]
|
||||
return value
|
||||
|
||||
|
||||
def _block_text(lines, start):
|
||||
"""Join a YAML block-scalar's indented continuation lines.
|
||||
|
||||
A blank line inside a block scalar does NOT end it (YAML compares
|
||||
indentation only against non-blank lines) — only a line back at the
|
||||
frontmatter's column-0 (the next key) does.
|
||||
"""
|
||||
collected = []
|
||||
for line in lines[start:]:
|
||||
if line.strip() == "" or line.startswith((" ", "\t")):
|
||||
collected.append(line.strip())
|
||||
else:
|
||||
break
|
||||
return " ".join(part for part in collected if part)
|
||||
|
||||
|
||||
def extract_description(path):
|
||||
"""The `description:` frontmatter value (scalar or block), or None."""
|
||||
try:
|
||||
with open(path, encoding="utf-8", errors="ignore") as handle:
|
||||
text = handle.read()
|
||||
except OSError:
|
||||
return None
|
||||
fm = _frontmatter(text)
|
||||
if fm is None:
|
||||
return None
|
||||
lines = fm.split("\n")
|
||||
for i, line in enumerate(lines):
|
||||
match = re.match(r"^description:\s*(.*)$", line)
|
||||
if not match:
|
||||
continue
|
||||
value = match.group(1).strip()
|
||||
if value in BLOCK_MARKERS:
|
||||
return _block_text(lines, i + 1) or None
|
||||
return _strip_quotes(value) or None
|
||||
return None
|
||||
|
||||
|
||||
def stem(word):
|
||||
"""Strip a trailing s/es/ed/ing suffix from a long-enough word."""
|
||||
for suffix in ("ing", "ed", "es", "s"):
|
||||
if len(word) > 4 and word.endswith(suffix):
|
||||
return word[: -len(suffix)]
|
||||
return word
|
||||
|
||||
|
||||
def tokenize(text):
|
||||
"""Lowercase, keep [a-z][a-z0-9-]+ words, drop stopwords/short, stem."""
|
||||
words = re.findall(r"[a-z][a-z0-9-]+", text.lower())
|
||||
return [stem(w) for w in words if w not in STOPWORDS and len(w) > 2]
|
||||
|
||||
|
||||
def build_docs(paths_by_name):
|
||||
"""name -> Counter(tokens), for every skill with a non-empty description."""
|
||||
docs = {}
|
||||
for name, path in paths_by_name.items():
|
||||
desc = extract_description(path)
|
||||
if not desc:
|
||||
continue
|
||||
tokens = tokenize(desc)
|
||||
if tokens:
|
||||
docs[name] = collections.Counter(tokens)
|
||||
return docs
|
||||
|
||||
|
||||
def _document_frequencies(docs):
|
||||
df = collections.Counter()
|
||||
for counter in docs.values():
|
||||
for token in counter:
|
||||
df[token] += 1
|
||||
return df
|
||||
|
||||
|
||||
def _tfidf_vector(counter, doc_count, df):
|
||||
"""(1 + log tf) * log(N/df), L2-normalized."""
|
||||
weights = {
|
||||
token: (1 + math.log(n)) * math.log(doc_count / df[token])
|
||||
for token, n in counter.items()
|
||||
}
|
||||
norm = math.sqrt(sum(w * w for w in weights.values())) or 1
|
||||
return {token: w / norm for token, w in weights.items()}
|
||||
|
||||
|
||||
def tfidf_vectors(docs):
|
||||
"""name -> {token: TF-IDF weight}, L2-normalized."""
|
||||
doc_count = len(docs)
|
||||
df = _document_frequencies(docs)
|
||||
return {name: _tfidf_vector(c, doc_count, df) for name, c in docs.items()}
|
||||
|
||||
|
||||
def cosine(vec_a, vec_b):
|
||||
return sum(weight * vec_b.get(token, 0) for token, weight in vec_a.items())
|
||||
|
||||
|
||||
def score_pairs(vectors):
|
||||
"""[(score, a, b), ...] over every pair, sorted by score descending."""
|
||||
pairs = [
|
||||
(cosine(vectors[a], vectors[b]), a, b)
|
||||
for a, b in itertools.combinations(sorted(vectors), 2)
|
||||
]
|
||||
pairs.sort(reverse=True)
|
||||
return pairs
|
||||
|
||||
|
||||
def _thresholds():
|
||||
warn = float(os.environ.get("SKILL_ROUTING_WARN", WARN_DEFAULT))
|
||||
fail = float(os.environ.get("SKILL_ROUTING_FAIL", FAIL_DEFAULT))
|
||||
return warn, fail
|
||||
|
||||
|
||||
def report(doc_count, pairs):
|
||||
"""Print the census report; return True iff a pair reached FAIL."""
|
||||
warn, fail = _thresholds()
|
||||
print(f"skills with description: {doc_count}")
|
||||
for score, a, b in pairs[:10]:
|
||||
print(f"{score:.2f} {a} ~ {b}")
|
||||
failed = False
|
||||
for score, a, b in pairs:
|
||||
if score >= fail:
|
||||
print(f"FAIL {score:.2f} {a} ~ {b}")
|
||||
failed = True
|
||||
elif score >= warn:
|
||||
print(f"WARN {score:.2f} {a} ~ {b}")
|
||||
return failed
|
||||
|
||||
|
||||
def main():
|
||||
by_name = dedup_by_skill_name(discover_paths(resolve_roots()))
|
||||
docs = build_docs(by_name)
|
||||
vectors = tfidf_vectors(docs)
|
||||
failed = report(len(docs), score_pairs(vectors))
|
||||
return 1 if failed else 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -0,0 +1,95 @@
|
||||
#!/usr/bin/env bash
|
||||
# lib/tests/floor-guard.test.sh — flip-tests for lib/floor-guard.sh: one RED
|
||||
# fixture per KIND, one WAIVED fixture, one CLEAN fixture. Each fixture is a
|
||||
# fresh throwaway repo under $WORK (`make test` exports
|
||||
# GIT_CONFIG_GLOBAL=/dev/null; core.hooksPath is also pinned per-repo so a
|
||||
# machine-wide hook never fires here). This file itself carries the trigger
|
||||
# strings for every kind — the self-run criterion excludes it by pathspec.
|
||||
set -u
|
||||
ROOT="$(cd "$(dirname "$0")/../.." && pwd)"
|
||||
LIB="$ROOT/lib/floor-guard.sh"
|
||||
WORK="$(mktemp -d)"; trap 'rm -rf "$WORK"' EXIT
|
||||
pass=0; fail=0
|
||||
|
||||
# check_kind <KIND> <rc> <want_rc> <out> <want_substr>
|
||||
check_kind() {
|
||||
local kind="$1" rc="$2" want_rc="$3" out="$4" want_sub="$5"
|
||||
if [ "$rc" = "$want_rc" ] && printf '%s\n' "$out" | grep -qF -- "$want_sub"; then
|
||||
pass=$((pass+1)); echo "PASS $kind"
|
||||
else
|
||||
fail=$((fail+1))
|
||||
printf 'FAIL %s: rc=%s (want %s), out:\n%s\n' \
|
||||
"$kind" "$rc" "$want_rc" "$(printf '%s\n' "$out" | tail -5)"
|
||||
fi
|
||||
}
|
||||
|
||||
# mk_repo <name> → path to a fresh throwaway repo: one test file (3
|
||||
# assertions), one vitest.config.ts (coverage.lines: 80), one source file.
|
||||
mk_repo() {
|
||||
local d="$WORK/$1"
|
||||
mkdir -p "$d/src"
|
||||
git init -q "$d"
|
||||
git -C "$d" config user.email t@t
|
||||
git -C "$d" config user.name t
|
||||
git -C "$d" config core.hooksPath /dev/null
|
||||
printf 'expect(1).toBe(1);\nexpect(2).toBe(2);\nexpect(3).toBe(3);\n' \
|
||||
> "$d/sample.test.js"
|
||||
printf 'export default {\n coverage: {\n lines: 80,\n },\n};\n' \
|
||||
> "$d/vitest.config.ts"
|
||||
printf 'function add(a, b) {\n return a + b;\n}\n' > "$d/src/index.js"
|
||||
git -C "$d" add -A
|
||||
git -C "$d" commit -q -m base
|
||||
echo "$d"
|
||||
}
|
||||
|
||||
# ── SUPPRESS ──────────────────────────────────────────────────────────────
|
||||
d=$(mk_repo suppress); base=$(git -C "$d" rev-parse HEAD)
|
||||
echo '// eslint-disable-next-line no-console' >> "$d/src/index.js"
|
||||
out=$(cd "$d" && bash "$LIB" "$base" 2>&1); rc=$?
|
||||
check_kind SUPPRESS "$rc" 2 "$out" 'FLOOR SUPPRESS'
|
||||
|
||||
# ── SKIP ──────────────────────────────────────────────────────────────────
|
||||
d=$(mk_repo skip); base=$(git -C "$d" rev-parse HEAD)
|
||||
echo "it.skip('later', () => {});" >> "$d/sample.test.js"
|
||||
out=$(cd "$d" && bash "$LIB" "$base" 2>&1); rc=$?
|
||||
check_kind SKIP "$rc" 2 "$out" 'FLOOR SKIP'
|
||||
|
||||
# ── DELETED_TEST ──────────────────────────────────────────────────────────
|
||||
d=$(mk_repo deleted); base=$(git -C "$d" rev-parse HEAD)
|
||||
rm "$d/sample.test.js"
|
||||
out=$(cd "$d" && bash "$LIB" "$base" 2>&1); rc=$?
|
||||
check_kind DELETED_TEST "$rc" 2 "$out" 'FLOOR DELETED_TEST'
|
||||
|
||||
# ── ASSERT_DROP ───────────────────────────────────────────────────────────
|
||||
d=$(mk_repo assertdrop); base=$(git -C "$d" rev-parse HEAD)
|
||||
printf 'expect(1).toBe(1);\n' > "$d/sample.test.js"
|
||||
out=$(cd "$d" && bash "$LIB" "$base" 2>&1); rc=$?
|
||||
check_kind ASSERT_DROP "$rc" 2 "$out" 'FLOOR ASSERT_DROP'
|
||||
|
||||
# ── STUB ──────────────────────────────────────────────────────────────────
|
||||
d=$(mk_repo stub); base=$(git -C "$d" rev-parse HEAD)
|
||||
echo "function todo() { throw new Error('not implemented'); }" >> "$d/src/index.js"
|
||||
out=$(cd "$d" && bash "$LIB" "$base" 2>&1); rc=$?
|
||||
check_kind STUB "$rc" 2 "$out" 'FLOOR STUB'
|
||||
|
||||
# ── THRESHOLD_DOWN ────────────────────────────────────────────────────────
|
||||
d=$(mk_repo threshold); base=$(git -C "$d" rev-parse HEAD)
|
||||
sed -i 's/lines: 80/lines: 60/' "$d/vitest.config.ts"
|
||||
out=$(cd "$d" && bash "$LIB" "$base" 2>&1); rc=$?
|
||||
check_kind THRESHOLD_DOWN "$rc" 2 "$out" 'FLOOR THRESHOLD_DOWN'
|
||||
|
||||
# ── WAIVED ────────────────────────────────────────────────────────────────
|
||||
d=$(mk_repo waived); base=$(git -C "$d" rev-parse HEAD)
|
||||
echo '// eslint-disable-next-line no-console -- floor-guard: allow legacy shim' \
|
||||
>> "$d/src/index.js"
|
||||
out=$(cd "$d" && bash "$LIB" "$base" 2>&1); rc=$?
|
||||
check_kind WAIVED "$rc" 0 "$out" 'WAIVED'
|
||||
|
||||
# ── CLEAN ─────────────────────────────────────────────────────────────────
|
||||
d=$(mk_repo clean); base=$(git -C "$d" rev-parse HEAD)
|
||||
echo '// helper' >> "$d/src/index.js"
|
||||
out=$(cd "$d" && bash "$LIB" "$base" 2>&1); rc=$?
|
||||
check_kind CLEAN "$rc" 0 "$out" 'FLOOR GUARD: clean'
|
||||
|
||||
echo "PASS=$pass FAIL=$fail"
|
||||
[ "$fail" -eq 0 ]
|
||||
@@ -18,7 +18,8 @@ check_not() { case "$2" in *"$3"*) fail=$((fail+1));
|
||||
|
||||
FX="$(mktemp -d)"; trap 'rm -rf "$FX"' EXIT
|
||||
mkdir -p "$FX/skills" "$FX/skills-disabled" "$FX/lib/profiles" "$FX/bin" \
|
||||
"$FX/hooks" "$FX/skills-external/emil-design-eng"
|
||||
"$FX/hooks" "$FX/skills-external/emil-design-eng" \
|
||||
"$FX/skills-external/observability-and-instrumentation"
|
||||
for g in gs-a gs-b gs-c; do
|
||||
mkdir -p "$FX/skills-external/gstack/$g"
|
||||
touch "$FX/skills-external/gstack/$g/SKILL.md"
|
||||
@@ -29,7 +30,8 @@ cp "$ROOT/hooks/statusline.sh" "$FX/hooks/"
|
||||
cat > "$FX/lib/profiles/full.profile" <<'EOF'
|
||||
gs-a
|
||||
gs-b
|
||||
emil-design-eng external
|
||||
emil-design-eng external
|
||||
observability-and-instrumentation external
|
||||
EOF
|
||||
cat > "$FX/lib/profiles/otherish.profile" <<'EOF'
|
||||
gs-c
|
||||
@@ -89,6 +91,7 @@ check T5-gsa-on "$([ -e "$FX/skills/gs-a" ] && echo on || echo off)" on
|
||||
check T5-gsb-on "$([ -e "$FX/skills/gs-b" ] && echo on || echo off)" on
|
||||
check T5-gsc-off "$([ -e "$FX/skills/gs-c" ] && echo on || echo off)" off
|
||||
check T5-emil-on "$([ -e "$FX/skills/emil-design-eng" ] && echo on || echo off)" on
|
||||
check T5-obs-on "$([ -e "$FX/skills/observability-and-instrumentation" ] && echo on || echo off)" on
|
||||
out="$(run current)"
|
||||
check T5-first-word "$(first_word "$out")" full
|
||||
check_has T5-match "$out" "100% match"
|
||||
|
||||
@@ -13,7 +13,8 @@ check() { if [ "$2" = "$3" ]; then pass=$((pass+1)); else fail=$((fail+1));
|
||||
FX="$(mktemp -d)"; trap 'rm -rf "$FX"' EXIT
|
||||
mkdir -p "$FX/skills" "$FX/skills-disabled" "$FX/lib/profiles" "$FX/bin" \
|
||||
"$FX/skills-external/emil-design-eng" "$FX/skills-external/other-ext" \
|
||||
"$FX/skills-external/21st-ui-build"
|
||||
"$FX/skills-external/21st-ui-build" \
|
||||
"$FX/skills-external/observability-and-instrumentation"
|
||||
for g in gs-a gs-b gs-c; do
|
||||
mkdir -p "$FX/skills-external/gstack/$g"
|
||||
touch "$FX/skills-external/gstack/$g/SKILL.md"
|
||||
@@ -26,8 +27,9 @@ ln -s "$FX/skills-external/other-ext" "$FX/skills/other-ext"
|
||||
cat > "$FX/lib/profiles/designish.profile" <<'EOF'
|
||||
gs-a
|
||||
gs-b
|
||||
emil-design-eng external
|
||||
21st-ui-build external
|
||||
emil-design-eng external
|
||||
21st-ui-build external
|
||||
observability-and-instrumentation external
|
||||
EOF
|
||||
cat > "$FX/lib/profiles/backendish.profile" <<'EOF'
|
||||
gs-c
|
||||
@@ -53,6 +55,7 @@ check T2-gsb-on "$([ -e "$FX/skills/gs-b" ] && echo on || echo off)" on
|
||||
check T3-gsc-off "$([ -e "$FX/skills/gs-c" ] && echo on || echo off)" off
|
||||
check T4-emil-src "$([ -L "$FX/skills/emil-design-eng" ] && echo on || echo off)" on
|
||||
check T5-21st-src "$([ -L "$FX/skills/21st-ui-build" ] && echo on || echo off)" on
|
||||
check T5b-obs-src "$([ -L "$FX/skills/observability-and-instrumentation" ] && echo on || echo off)" on
|
||||
check T6-no-mcp "$(grep -c '^mcp ' "$FX/claude-calls.log" || true)" 0
|
||||
|
||||
# --- set backendish: managed leftovers parked/unregistered ---
|
||||
@@ -63,6 +66,8 @@ check T9-emil-off "$([ -e "$FX/skills/emil-design-eng" ] && echo on || echo off
|
||||
check T10-emil-park "$([ -e "$FX/skills-disabled/emil-design-eng" ] && echo p || echo n)" p
|
||||
check T11-21st-off "$([ -e "$FX/skills/21st-ui-build" ] && echo on || echo off)" off
|
||||
check T12-21st-park "$([ -e "$FX/skills-disabled/21st-ui-build" ] && echo p || echo n)" p
|
||||
check T12b-obs-off "$([ -e "$FX/skills/observability-and-instrumentation" ] && echo on || echo off)" off
|
||||
check T12c-obs-park "$([ -e "$FX/skills-disabled/observability-and-instrumentation" ] && echo p || echo n)" p
|
||||
check T13-other-untouched "$([ -e "$FX/skills/other-ext" ] && echo on || echo off)" on
|
||||
|
||||
# --- back to designish: parked external restored (not re-sourced) ---
|
||||
@@ -70,6 +75,8 @@ run set designish >/dev/null 2>&1
|
||||
check T14-emil-back "$([ -e "$FX/skills/emil-design-eng" ] && echo on || echo off)" on
|
||||
check T15-park-gone "$([ -e "$FX/skills-disabled/emil-design-eng" ] && echo p || echo n)" n
|
||||
check T16-21st-back "$([ -e "$FX/skills/21st-ui-build" ] && echo on || echo off)" on
|
||||
check T16b-obs-back "$([ -e "$FX/skills/observability-and-instrumentation" ] && echo on || echo off)" on
|
||||
check T16c-obs-park-gone "$([ -e "$FX/skills-disabled/observability-and-instrumentation" ] && echo p || echo n)" n
|
||||
check T17-no-mcp-ever "$(grep -c '^mcp ' "$FX/claude-calls.log" || true)" 0
|
||||
|
||||
printf 'PASS=%s FAIL=%s\n' "$pass" "$fail"; [ "$fail" -eq 0 ]
|
||||
|
||||
Executable
+113
@@ -0,0 +1,113 @@
|
||||
#!/usr/bin/env bash
|
||||
# lib/tests/skill-routing-census.test.sh — TF-IDF cosine census of
|
||||
# skill-description collisions across the live catalog (routing ambiguity),
|
||||
# adapted from addyosmani/agent-skills evals Tier 2. Two skills whose
|
||||
# descriptions read alike route the same prompt to both — a routing
|
||||
# collision, not a cosmetic naming clash.
|
||||
#
|
||||
# Fixture flip-test math note: with only 2 documents, corpus-wide IDF zeroes
|
||||
# EVERY term's contribution — a term unique to one doc (df=1) is absent from
|
||||
# the other vector and can't enter the dot product, a term shared by both
|
||||
# (df=2) gets idf=log(N/df)=log(1)=0. Cosine is trivially 0 for any 2-doc
|
||||
# corpus, near-duplicate or not — a bare 2-doc "distinct pair" fixture would
|
||||
# pass for that reason alone, not because the pair is actually distinct. Both
|
||||
# fixtures below ride in a 4-doc corpus so IDF carries real signal. The
|
||||
# distinct-pair fixture also carries its own same-corpus positive control (a
|
||||
# near-dup pair that must FAIL) and a sensitivity re-run where the "distinct"
|
||||
# partner is swapped for a near-copy of its counterpart, to prove the WARN/
|
||||
# FAIL marker actually fires when the pair collides — not just that it stays
|
||||
# silent for reasons unrelated to distinctness.
|
||||
set -u
|
||||
ROOT="$(cd "$(dirname "$0")/../.." && pwd)"
|
||||
CENSUS="$ROOT/lib/skill-routing-census.py"
|
||||
pass=0; fail=0
|
||||
check() { if [ "$2" = "$3" ]; then pass=$((pass+1)); else fail=$((fail+1));
|
||||
printf 'FAIL %s: got[%s] want[%s]\n' "$1" "$2" "$3"; fi; }
|
||||
|
||||
WORK="$(mktemp -d)"; trap 'rm -rf "$WORK"' EXIT
|
||||
TRANSLATE_DESC='description: "Translate a PDF, keeping layout and images."'
|
||||
SECAUDIT_DESC="description: 'Run a security audit for secrets, CVEs, OWASP.'"
|
||||
|
||||
# _mkskill <root> <name> <frontmatter-line>... — a fixture skill/SKILL.md
|
||||
_mkskill() {
|
||||
local dir="$1" name="$2"; shift 2
|
||||
mkdir -p "$dir/$name"
|
||||
{ echo "---"; echo "name: $name"; printf '%s\n' "$@"; echo "---"; } \
|
||||
> "$dir/$name/SKILL.md"
|
||||
}
|
||||
|
||||
# _mkcontrol <root> — same-corpus positive control: skill-e/skill-f, a
|
||||
# near-duplicate deploy-runbook pair that must FAIL wherever it rides.
|
||||
_mkcontrol() {
|
||||
local dir="$1"
|
||||
_mkskill "$dir" skill-e "description: |" \
|
||||
" Use when deploying a project via its per-project runbook to ship the" \
|
||||
" application to production servers safely with rollback support and" \
|
||||
" health checks after each release."
|
||||
_mkskill "$dir" skill-f "description: >-" \
|
||||
" Use when deploying a project via its per-project deployment runbook to" \
|
||||
" ship the application to production servers safely with rollback" \
|
||||
" support and health checks after each release."
|
||||
}
|
||||
|
||||
# ── live run: this machine's real catalog (a FAIL line here → suite RED) ──
|
||||
live_out=$(python3 "$CENSUS" 2>&1); live_rc=$?
|
||||
printf '%s\n' "$live_out"
|
||||
check T1-live-run-no-collision "$live_rc" 0
|
||||
has_count=$(printf '%s\n' "$live_out" \
|
||||
| grep -qE 'skills with description: [0-9]{2,}' && echo yes)
|
||||
check T2-live-catalog-non-trivial "$has_count" yes
|
||||
|
||||
# ── fixture AB: near-duplicate pair + 2 unrelated fillers (N=4, real signal) ──
|
||||
AB="$WORK/ab"
|
||||
_mkskill "$AB" skill-a "description: |" \
|
||||
" Use when deploying a project via its per-project runbook to ship the" \
|
||||
" application to production servers safely with rollback support and" \
|
||||
" health checks after each release."
|
||||
_mkskill "$AB" skill-b "description: >-" \
|
||||
" Use when deploying a project via its per-project deployment runbook to" \
|
||||
" ship the application to production servers safely with rollback" \
|
||||
" support and health checks after each release."
|
||||
_mkskill "$AB" filler-translate "$TRANSLATE_DESC"
|
||||
_mkskill "$AB" filler-secaudit "$SECAUDIT_DESC"
|
||||
out_ab=$(SKILL_ROUTING_ROOTS="$AB" python3 "$CENSUS" 2>&1)
|
||||
fail_line=$(printf '%s\n' "$out_ab" \
|
||||
| grep -qE '^FAIL 0\.[0-9]{2} skill-a ~ skill-b$' && echo yes)
|
||||
check T3-collision-fail-line "$fail_line" yes
|
||||
[ "$fail_line" = yes ] && echo FIXTURE_COLLISION_DETECTED
|
||||
|
||||
# ── fixture CD: distinct pair + same-corpus positive control (N=4) ────────
|
||||
# skill-e/skill-f (positive control) must FAIL; skill-c/skill-d (the pair
|
||||
# under test) must stay silent (no WARN/FAIL line naming them).
|
||||
CD="$WORK/cd"
|
||||
_mkcontrol "$CD"
|
||||
_mkskill "$CD" skill-c "$TRANSLATE_DESC"
|
||||
_mkskill "$CD" skill-d "$SECAUDIT_DESC"
|
||||
out_cd=$(SKILL_ROUTING_ROOTS="$CD" python3 "$CENSUS" 2>&1)
|
||||
control_fail=$(printf '%s\n' "$out_cd" \
|
||||
| grep -qE '^FAIL 0\.[0-9]{2} skill-e ~ skill-f$' && echo yes)
|
||||
check T4-positive-control-fail "$control_fail" yes
|
||||
distinct_marker=$(printf '%s\n' "$out_cd" \
|
||||
| grep -E '^(WARN|FAIL) 0\.[0-9]{2} skill-c ~ skill-d$')
|
||||
distinct_silent=$([ -z "$distinct_marker" ] && echo yes || echo no)
|
||||
check T5-distinct-pair-silent "$distinct_silent" yes
|
||||
|
||||
# ── sensitivity re-run: skill-d -> near-copy of skill-c ────────────────────
|
||||
# Same corpus, but skill-d's description now near-duplicates skill-c's: the
|
||||
# marker that stayed silent above must appear here, proving T5 wasn't silent
|
||||
# by construction (e.g. a broken pattern or a threshold nothing can cross).
|
||||
CD2="$WORK/cd-nearcopy"
|
||||
_mkcontrol "$CD2"
|
||||
_mkskill "$CD2" skill-c "$TRANSLATE_DESC"
|
||||
_mkskill "$CD2" skill-d \
|
||||
'description: "Translate a PDF document, keeping the layout and images intact."'
|
||||
out_cd2=$(SKILL_ROUTING_ROOTS="$CD2" python3 "$CENSUS" 2>&1)
|
||||
sensitivity_marker=$(printf '%s\n' "$out_cd2" \
|
||||
| grep -qE '^(WARN|FAIL) 0\.[0-9]{2} skill-c ~ skill-d$' && echo yes)
|
||||
check T6-sensitivity-marker-appears "$sensitivity_marker" yes
|
||||
|
||||
[ "$control_fail" = yes ] && [ "$distinct_silent" = yes ] \
|
||||
&& [ "$sensitivity_marker" = yes ] && echo FIXTURE_DISTINCT_OK
|
||||
|
||||
echo "PASS=$pass FAIL=$fail"
|
||||
[ "$fail" -eq 0 ]
|
||||
@@ -14,6 +14,7 @@ check() { if [ "$2" = "$3" ]; then pass=$((pass+1)); else fail=$((fail+1));
|
||||
|
||||
SANDBOX="$(mktemp -d)"
|
||||
mkdir -p "$SANDBOX/repo/lib" "$SANDBOX/repo/skills-external/emil-design-eng" \
|
||||
"$SANDBOX/repo/skills-external/observability-and-instrumentation" \
|
||||
"$SANDBOX/repo/skills" "$SANDBOX/home/.claude"
|
||||
cp "$HELPER_SRC" "$SANDBOX/repo/lib/toggle-external.sh"
|
||||
# mark emil-design-eng ENABLED in the real (physical) repo tree
|
||||
@@ -24,5 +25,11 @@ ln -s "$SANDBOX/repo/lib" "$SANDBOX/home/.claude/lib"
|
||||
out="$(bash "$SANDBOX/home/.claude/lib/toggle-external.sh" status emil-design-eng)"
|
||||
check T1-repo-resolves-through-symlink "$out" enabled
|
||||
|
||||
# Same case arm, generalized to $tool for the agent-skills trio (this PR) —
|
||||
# left unlinked, so it must resolve through the symlink as "disabled", not
|
||||
# "missing" (which would mean REPO fell back to the wrong tree again).
|
||||
out="$(bash "$SANDBOX/home/.claude/lib/toggle-external.sh" status observability-and-instrumentation)"
|
||||
check T2-generalized-tool-resolves-through-symlink "$out" disabled
|
||||
|
||||
rm -rf "$SANDBOX"
|
||||
printf 'PASS=%s FAIL=%s\n' "$pass" "$fail"; [ "$fail" -eq 0 ]
|
||||
|
||||
+11
-7
@@ -21,6 +21,9 @@
|
||||
# emil-design-eng — single symlink → skills-external/emil-design-eng
|
||||
# darwin-skill — single symlink → ~/.agents/skills/darwin-skill
|
||||
# 21st — 21st.dev skill pack (needs the `21st` CLI + login)
|
||||
# observability-and-instrumentation, deprecation-and-migration,
|
||||
# ci-cd-and-automation — the agent-skills trio, same single-symlink shape
|
||||
# as emil-design-eng (commit-pinned instead of main-branch tracking)
|
||||
#
|
||||
# For fine-grained activation (only design skills, only qa skills, only
|
||||
# audit skills, etc.) instead of all-or-nothing gstack toggling, use:
|
||||
@@ -40,7 +43,8 @@ warn() { echo -e "${YELLOW}⚠${NC} $1"; }
|
||||
err() { echo -e "${RED}✗${NC} $1"; }
|
||||
|
||||
# All non-plugin tools this script can toggle.
|
||||
MANAGED_TOOLS=(gstack emil-design-eng darwin-skill 21st)
|
||||
MANAGED_TOOLS=(gstack emil-design-eng darwin-skill 21st
|
||||
observability-and-instrumentation deprecation-and-migration ci-cd-and-automation)
|
||||
|
||||
# Prints the skill names that belong to the "21st" pack. Source of truth:
|
||||
# skills-external/21st-* — the `21st skills install` run in install-plugins.sh
|
||||
@@ -76,9 +80,9 @@ status_tool() {
|
||||
done < <(gstack_skills)
|
||||
echo "disabled"
|
||||
;;
|
||||
emil-design-eng)
|
||||
[ -d "$REPO/skills-external/emil-design-eng" ] || { echo "missing"; return; }
|
||||
[ -e "$SKILLS_DIR/emil-design-eng" ] && echo "enabled" || echo "disabled"
|
||||
emil-design-eng|observability-and-instrumentation|deprecation-and-migration|ci-cd-and-automation)
|
||||
[ -d "$REPO/skills-external/$tool" ] || { echo "missing"; return; }
|
||||
[ -e "$SKILLS_DIR/$tool" ] && echo "enabled" || echo "disabled"
|
||||
;;
|
||||
darwin-skill)
|
||||
[ -d "$HOME/.agents/skills/$tool" ] || { echo "missing"; return; }
|
||||
@@ -115,7 +119,7 @@ disable_tool() {
|
||||
done < <(gstack_skills)
|
||||
ok "gstack disabled ($moved symlinks moved)"
|
||||
;;
|
||||
emil-design-eng|darwin-skill)
|
||||
emil-design-eng|darwin-skill|observability-and-instrumentation|deprecation-and-migration|ci-cd-and-automation)
|
||||
if [ -e "$SKILLS_DIR/$tool" ]; then
|
||||
rm -rf "${DISABLED_DIR:?}/${tool:?}"
|
||||
mv "$SKILLS_DIR/$tool" "$DISABLED_DIR/$tool"
|
||||
@@ -165,11 +169,11 @@ enable_tool() {
|
||||
ok "gstack enabled ($moved symlinks restored)"
|
||||
fi
|
||||
;;
|
||||
emil-design-eng|darwin-skill)
|
||||
emil-design-eng|darwin-skill|observability-and-instrumentation|deprecation-and-migration|ci-cd-and-automation)
|
||||
local src
|
||||
case "$tool" in
|
||||
emil-design-eng) src="$REPO/skills-external/$tool" ;;
|
||||
darwin-skill) src="$HOME/.agents/skills/$tool" ;;
|
||||
*) src="$REPO/skills-external/$tool" ;;
|
||||
esac
|
||||
if [ -e "$DISABLED_DIR/$tool" ]; then
|
||||
rm -rf "${SKILLS_DIR:?}/${tool:?}"
|
||||
|
||||
@@ -55,6 +55,18 @@ Dispatch a FRESH verifier subagent (`subagent_type: verifier`, or load
|
||||
`TEST` command. Never pass the dev's summary, never pass a prior iteration's
|
||||
gaps — the verifier reads the contract from disk and judges blind.
|
||||
|
||||
The verifier's STEP 3 (`agents/verifier.md`) runs `lib/floor-guard.sh`
|
||||
against the diff before it renders any verdict — a deterministic,
|
||||
diff-scoped check for a quietly weakened quality bar (a suppressed
|
||||
lint/type check, a skipped or deleted test, a dropped assertion, a lowered
|
||||
coverage threshold) that an LLM verdict alone can miss or be talked past
|
||||
one line at a time. Its findings fold straight into that same verifier's
|
||||
`ECARTS` count unless the contract's `CLARIFICATIONS` explicitly authorizes
|
||||
the exact weakening; there is no separate gate and no extra dispatch, it
|
||||
rides this GATE 1 call. A `floor-guard: allow` waiver outside a test file
|
||||
is a finding too unless the contract's `CLARIFICATIONS` names it: the
|
||||
waiver is self-service, the contract is human-gated (BDR-102 amendment).
|
||||
|
||||
Parse its single `VERIFY — VERDICT:` line:
|
||||
|
||||
- `CONFORME` → go to GATE 2. (First-pass conforme = no loop.)
|
||||
@@ -74,9 +86,9 @@ Parse its single `VERIFY — VERDICT:` line:
|
||||
micro-gate that appends `[gated <date>]` to the contract's FILE SCOPE;
|
||||
otherwise the dev removes the file.
|
||||
- Structural failure (`ERROR(…)`, missing/duplicated VERDICT line,
|
||||
unparsable, crash, `CONFORME` without `PROOF`) → retry ONCE with a fresh
|
||||
verifier; a 2nd structural failure → human escalation. A mute verifier is
|
||||
NEVER a PASS.
|
||||
unparsable, crash, `CONFORME` without `PROOF` or without `FLOOR`) → retry
|
||||
ONCE with a fresh verifier; a 2nd structural failure → human escalation.
|
||||
A mute verifier is NEVER a PASS.
|
||||
|
||||
## GATE 2 — SECURITY (fresh security-auditor)
|
||||
|
||||
|
||||
@@ -93,7 +93,8 @@ fi
|
||||
# impeccable is NOT here: its installer writes the skill straight into
|
||||
# skills/ (and its agents into agents/) at --scope=global, so there is no
|
||||
# skills-external/ copy to symlink. See install-plugins.sh Step 8d.
|
||||
EXTERNAL_SKILLS=(emil-design-eng frontend-design design-motion-principles)
|
||||
EXTERNAL_SKILLS=(emil-design-eng frontend-design design-motion-principles
|
||||
observability-and-instrumentation deprecation-and-migration ci-cd-and-automation)
|
||||
for _ext_skill in "${EXTERNAL_SKILLS[@]}"; do
|
||||
if [ -d "$REPO/skills-external/$_ext_skill" ]; then
|
||||
if [ -L "$CLAUDE/skills/$_ext_skill" ] && [ "$(readlink "$CLAUDE/skills/$_ext_skill")" = "$REPO/skills-external/$_ext_skill" ]; then
|
||||
|
||||
@@ -43,6 +43,13 @@
|
||||
"managed_by": "curl",
|
||||
"note": "Emil Kowalski's design engineering skill — UI polish, animations, component craft. Machine-owned: curl'd to skills-external/emil-design-eng/ (gitignored, re-fetched by update-all.sh), symlinked by link.sh."
|
||||
},
|
||||
"agent-skills": {
|
||||
"source": "https://github.com/addyosmani/agent-skills",
|
||||
"commit": "2686b620fc1fed2e8f60c704839c766b8594c6b6",
|
||||
"skills": ["observability-and-instrumentation", "deprecation-and-migration", "ci-cd-and-automation"],
|
||||
"managed_by": "curl",
|
||||
"note": "Three dev-lifecycle skills from addyosmani/agent-skills, vendored the emil-design-eng way but COMMIT-pinned (not main-branch tracking): each lands in skills-external/<name>/SKILL.md (gitignored, symlinked by link.sh), install-plugins.sh Step 8e curls all three at this commit when absent, update-all.sh re-fetches at the SAME commit on every run (a pin, not an auto-advance). Bump the commit deliberately to pick up upstream changes; the scripts read it from here, never hardcode it."
|
||||
},
|
||||
"impeccable": {
|
||||
"source": "npm:impeccable",
|
||||
"version": "4.1.0",
|
||||
|
||||
@@ -0,0 +1,43 @@
|
||||
---
|
||||
paths: ["**/api/**", "**/routes/**", "**/controllers/**", "**/*.route.*", "**/*.controller.*", "**/openapi.*", "**/*.openapi.*"]
|
||||
---
|
||||
|
||||
# REST API — contract, errors, lists, idempotency
|
||||
|
||||
## Contract first
|
||||
Order: typed input/output → schemas (server-generated fields like id,
|
||||
createdAt apart from client input) → error codes → implementation.
|
||||
Validate at boundaries only (route handlers, external responses, env
|
||||
loading); trust internal code and your own database reads.
|
||||
|
||||
## Errors
|
||||
One envelope everywhere: `{ error: { code, message, details? } }`.
|
||||
HTTP map: 400 invalid · 401 auth · 403 forbidden · 404 missing ·
|
||||
409 conflict · 422 semantic · 500 server (never leak internals).
|
||||
Never mix throw / null / envelope styles across endpoints.
|
||||
|
||||
## Lists
|
||||
Every list endpoint paginated: `page`, `pageSize`, `totalItems`,
|
||||
`totalPages`. Filters as query params (`?status=x&createdAfter=…`),
|
||||
never in the body.
|
||||
|
||||
## Idempotency
|
||||
Key from intent, not attempt (`charge:v1:${orderId}`, never
|
||||
`randomUUID()` or a timestamp). Claim atomically via a unique
|
||||
constraint — check-then-insert is a race, not a guard. Same key,
|
||||
different payload: fail loudly, never replay the first response.
|
||||
Pick the in-flight-duplicate policy on purpose: 409 reject, bounded
|
||||
wait, or 202 + status URL. Retention outlives the longest retry
|
||||
path, dead-letter replay included.
|
||||
|
||||
## Naming
|
||||
Plural nouns for endpoints (`/api/tasks`), no verbs. camelCase query
|
||||
params and response fields. UPPER_SNAKE enum values. Boolean fields
|
||||
prefixed is/has/can.
|
||||
|
||||
## Hyrum's law
|
||||
Every observable behavior — undocumented quirks, error text, timing —
|
||||
becomes a de facto contract once someone depends on it.
|
||||
|
||||
## Versioning
|
||||
CLAUDE.md § Web APIs — always versioned. Not repeated here.
|
||||
@@ -379,6 +379,38 @@ else
|
||||
info "design-motion-principles not installed — skipping"
|
||||
fi
|
||||
|
||||
# ── 7.3. Update Agent Skills (addyosmani/agent-skills, pinned commit) ──
|
||||
echo ""
|
||||
echo "── Updating Agent Skills (addyosmani/agent-skills)..."
|
||||
AGENT_SKILLS_SHA=""
|
||||
if [ -f "$REPO/plugins.lock.json" ] && command -v python3 &>/dev/null; then
|
||||
AGENT_SKILLS_SHA=$(python3 -c "
|
||||
import json, sys
|
||||
with open(sys.argv[1]) as f:
|
||||
d = json.load(f)
|
||||
print(d.get(sys.argv[2], {}).get('commit', ''))
|
||||
" "$REPO/plugins.lock.json" "agent-skills" 2>/dev/null || true)
|
||||
fi
|
||||
AGENT_SKILLS_NAMES=(observability-and-instrumentation deprecation-and-migration ci-cd-and-automation)
|
||||
if [ -z "$AGENT_SKILLS_SHA" ]; then
|
||||
warn "agent-skills: no commit pinned in plugins.lock.json — skipping"
|
||||
else
|
||||
for _as_skill in "${AGENT_SKILLS_NAMES[@]}"; do
|
||||
_as_dir="$REPO/skills-external/$_as_skill"
|
||||
if [ ! -d "$_as_dir" ]; then
|
||||
info "$_as_skill not installed — skipping (run: make plugin)"
|
||||
continue
|
||||
fi
|
||||
_as_url="https://raw.githubusercontent.com/addyosmani/agent-skills/$AGENT_SKILLS_SHA/skills/$_as_skill/SKILL.md"
|
||||
if curl -fsSL "$_as_url" -o "$_as_dir/SKILL.md.tmp" \
|
||||
&& mv "$_as_dir/SKILL.md.tmp" "$_as_dir/SKILL.md"; then
|
||||
ok "$_as_skill re-fetched at pinned commit"
|
||||
else
|
||||
warn "$_as_skill update failed"
|
||||
fi
|
||||
done
|
||||
fi
|
||||
|
||||
# ── Impeccable (design detector + skill + subagents) ──
|
||||
# Global scope: the installer writes through the ~/.claude/{skills,agents}
|
||||
# symlinks straight into this repo (install-plugins.sh Step 8d explains why
|
||||
|
||||
Reference in New Issue
Block a user