Merge feature/darwin-optimize-20260825 into develop
This commit is contained in:
Binary file not shown.
|
After Width: | Height: | Size: 254 KiB |
@@ -0,0 +1,90 @@
|
|||||||
|
# Darwin run 2026-08-25/26: fresh baseline + threshold optimization + bug pass
|
||||||
|
|
||||||
|
Branch `feature/darwin-optimize-20260825`, 26 commits, 39 files, +299/-142.
|
||||||
|
Log: `~/.agents/skills/darwin-skill/results.tsv` (fresh, the May file was wiped
|
||||||
|
by the 2026-06-23 reinstall). Method: darwin v2.1. Absolute scores served as
|
||||||
|
triage only; every keep/revert decision came from a paired same-judge majority
|
||||||
|
(3 judges per round, before/after read in one call).
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
54 units: 31 personal skill-systems (SKILL.md + dispatched agents judged
|
||||||
|
together, per EVAL-004) and 23 agents. Excluded: gstack/external symlinks
|
||||||
|
(BDR-015/043, LRN-070), darwin-skill itself (BDR-058 pin), and find-docs,
|
||||||
|
newly identified as machine-owned ctx7 output (gitignored, installer-written).
|
||||||
|
|
||||||
|
## Baseline (7 blind judges, dims scored 1-10, totals recomputed main-thread per LRN-018)
|
||||||
|
|
||||||
|
Mean 83.4 (skills 83.5, agents 83.3). Best: deploy, release-candidate,
|
||||||
|
release-executor (90.4). Worst: skills-perso 63.5. All dim8 rows marked
|
||||||
|
dry_run by design; live execution happened later, inside the paired rounds.
|
||||||
|
13 units scored below the user-set threshold of 80.
|
||||||
|
|
||||||
|
## Phase 2: threshold loop, 13/13 units, 0 reverts
|
||||||
|
|
||||||
|
Every round was validated by 3 paired judges (neutral, skeptic, realism).
|
||||||
|
All verdicts 3-0 better.
|
||||||
|
|
||||||
|
| Unit (baseline) | Round(s) | What changed |
|
||||||
|
|---|---|---|
|
||||||
|
| skills-perso (63.5) | d8 | Detection rebuilt on the link.sh convention: symlink = external, real dir = personal, gitignored = machine-generated. Live result 8/31 to 31/31, zero false positives |
|
||||||
|
| interviewer (70.9) | d3, d9 | Failure-mode table (vague, "you decide", contradiction, partial, balloon) + 2-round budget; DO-NOT list |
|
||||||
|
| onboarder (71.5) | d8 | BRIEF contract split REQUIRED/OPTIONAL; null enrichment becomes TODO placeholders; STOP kept for required keys and unresolved monorepo. Kills the guaranteed first-dispatch bounce vs /onboard STEP 2 |
|
||||||
|
| pdf-translate (72.3) | d3/d8 | 8-row failure table: deps, >30 pages gate, zero-output, illisible markers, design-html/browse fallbacks, QA cap 2, stale workdir |
|
||||||
|
| refactor (75.6) + refactorer (76.8) | d4/d3 | No-tests STOP gate + GO-WITHOUT-TESTS arbitration in the dispatcher; mid-run test-failure revert protocol; code-cleaner inline carve-out |
|
||||||
|
| profile (77.3) | d3 | 6-row failure table, every row fact-checked against profile.sh (rc=1 paths, partial toggle, split plugin leg, BLK-006 contradiction); fixture de-drift |
|
||||||
|
| plugin-probe (78.5) + plugin-advisor (77.5) | d8 | FRAMEWORK-DEPS now exact dep@version (preact false-hit killed, fallback actually fires; the old `\|\| true` silently emitted nothing and tripped the advisor's fail-closed path on non-Node projects); frontend/fast-libs derivable; PLAN echoed-or-unknown, invention removed |
|
||||||
|
| analyze (77.7) + analyzer (78.0) | d1, d2 | Bilingual triggers + fix-wanted disambiguator; TASKS ordered, each step mapped to its OUTPUT section |
|
||||||
|
| status-reporter (78.0) | d5 x2 | Fabrication-forcing token field replaced, then restored producibly from doctor.sh constants (a skeptic judge found the source); dead ROADMAP row rewritten post-ADR-013 |
|
||||||
|
| gitflow (78.4) | d3 | 7-row failure table keyed to lib return codes; rc=4 conflict resume empirically verified; human merge gate untouched |
|
||||||
|
|
||||||
|
## Bug pass: verified defects in above-threshold units, 8 commits, all kept 3-0
|
||||||
|
|
||||||
|
- hotfix: `git restore .` on every failure branch wiped tolerated in-progress
|
||||||
|
user edits. Now: `git stash create` pre-flight snapshot + file-scoped
|
||||||
|
restore + fresh-dispatch-only security gate. Two skeptic residuals amended
|
||||||
|
(RULES bullet, FILE(S) new-file marker).
|
||||||
|
- init-project: allowed-tools lacked Agent and Skill while every step
|
||||||
|
dispatches. commit-change: conflict grep now covers all 7 unmerged codes.
|
||||||
|
tour: --report-only no longer commits (could land on develop).
|
||||||
|
- harden: severity rule now defers to the calibrated guide; the late SSL Labs
|
||||||
|
grade has an assigned actor.
|
||||||
|
- plan-challenger: ERROR joined the load-bearing verdict grammar.
|
||||||
|
- handover writers: stale chapter refs corrected (glossary/tone to §6,
|
||||||
|
cross-links and THRESHOLD-OVERRIDE to §5); STEP 14.5 verification deferred
|
||||||
|
post-write; anchor gate ordered into STEP 16.
|
||||||
|
- security-auditor: /hotfix no-verifier carve-out documented. close: STEP 5C
|
||||||
|
enumerated, --no-push passthrough added.
|
||||||
|
- prune-memory: false "v1-untested" note replaced by the real tests/ state.
|
||||||
|
code-clean: executor attribution corrected (code-cleaner, refactorer inline).
|
||||||
|
- Fixtures de-drifted: plugin-check (PLUGIN CHECK block, real plugin names),
|
||||||
|
onboard (nextjs-app-router).
|
||||||
|
|
||||||
|
`make test` green (0 RED, rc=0) after one census rewrap: a locked phrase had
|
||||||
|
been line-wrapped and the single-line grep lock caught it.
|
||||||
|
|
||||||
|
## Residual findings, logged not fixed
|
||||||
|
|
||||||
|
- analyze triggers: "how does X work" brushes graphify's territory; graphify's
|
||||||
|
graph-exists routing still wins.
|
||||||
|
- pdf-translate: pdfinfo row assumes poppler (fitz also has page count); "GB"
|
||||||
|
slightly overstated near the 30-page gate.
|
||||||
|
- web-validate: .validate-cache mkdir lives in a skipped STEP 0
|
||||||
|
(self-recoverable); axis budgets 35/25/40 never reconciled with the base-100
|
||||||
|
deduction table. seo/geo minor wording items. verifier/doc-syncer/audit-delta
|
||||||
|
restatement redundancy (cosmetic). handover-doc-writer STEP 14.5 umbrella
|
||||||
|
line still says "BEFORE STEP 15" while the inner note overrides it.
|
||||||
|
- bugfix trivial-fast-path boundary loosely defined; feat prompt-3 expectation
|
||||||
|
vs full gate pipeline.
|
||||||
|
|
||||||
|
## Methodology notes
|
||||||
|
|
||||||
|
- v2.1 paired majority produced 36 unit-round verdicts and 24 batch verdicts,
|
||||||
|
all better, 0 reverts, 0 ties. The May-2026 run under absolute-delta scoring
|
||||||
|
had reverted 2 edits on judge noise; this run had no such event.
|
||||||
|
- Judges live-executed wherever the artifact was executable (skills-perso
|
||||||
|
detection, profile.sh probes, plugin grep on scratch manifests, doctor.sh
|
||||||
|
grep, git merge no-op resume). Behavior outranked prose in 5 units.
|
||||||
|
- Two grep-exit-masking bugs surfaced (a `head` pipe swallowing the fallback's
|
||||||
|
trigger), one in the probe being fixed, one in this run's own test harness.
|
||||||
|
The pattern is worth a learning entry.
|
||||||
@@ -95,6 +95,7 @@ rules:
|
|||||||
| BDR-083 | 2026-08-24 | Contract gates: deterministic floor (GATE 0) under the fresh verifier | accepted |
|
| BDR-083 | 2026-08-24 | Contract gates: deterministic floor (GATE 0) under the fresh verifier | accepted |
|
||||||
| BDR-084 | 2026-08-24 | /tour multi-project: parallel runners (LRN-083 derogation, bounded), runner inherits session model | accepted |
|
| BDR-084 | 2026-08-24 | /tour multi-project: parallel runners (LRN-083 derogation, bounded), runner inherits session model | accepted |
|
||||||
| BDR-085 | 2026-08-25 | User permanent rules: writing-style always-on in rules/, web build+security path-scoped | accepted |
|
| BDR-085 | 2026-08-25 | User permanent rules: writing-style always-on in rules/, web build+security path-scoped | accepted |
|
||||||
|
| BDR-086 | 2026-08-26 | darwin: threshold gates full loops; verified defects fixed regardless of unit score (paired-validated, batched checkpoint) | accepted |
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -1105,3 +1106,10 @@ Conflict carve-outs, stated INSIDE the rules: registries keep caveman format (fr
|
|||||||
Self-check rule scoped to DELIVERABLES (text, site, feature), not every conversational reply — literal "avant de me rendre quoi que ce soit" would append a compliance note to every chat answer, pure noise. User can re-widen.
|
Self-check rule scoped to DELIVERABLES (text, site, feature), not every conversational reply — literal "avant de me rendre quoi que ce soit" would append a compliance note to every chat answer, pure noise. User can re-widen.
|
||||||
Alternatives rejected: compress into CLAUDE.global.md (~11 lines to fit → loses the carve-outs, zero headroom left); path-scope writing-style (applies to conversation, not file reads → would never fire in chat-only sessions); one merged web file (two concerns, one-rule-one-file).
|
Alternatives rejected: compress into CLAUDE.global.md (~11 lines to fit → loses the carve-outs, zero headroom left); path-scope writing-style (applies to conversation, not file reads → would never fire in chat-only sessions); one merged web file (two concerns, one-rule-one-file).
|
||||||
Branch feature/user-writing-web-rules, UNMERGED (human gate).
|
Branch feature/user-writing-web-rules, UNMERGED (human gate).
|
||||||
|
|
||||||
|
## BDR-086 — darwin bug-pass scope: verified defects fixed above threshold
|
||||||
|
- **Date**: 2026-08-26
|
||||||
|
- **Decision**: units < threshold get full weighted-gap optimization loops (per-unit checkpoint). Judge-VERIFIED defects (file:line, confirmed) in above-threshold units get targeted fixes in a grouped pass — same paired 3-judge validation, one batched checkpoint. User-gated at the scorecard.
|
||||||
|
- **Why**: leaving a verified destructive path (hotfix `git restore .` wiping tolerated user edits, file scored 85) unfixed = score-worship; rubric serves quality, not the inverse.
|
||||||
|
- **Alternatives rejected**: strict threshold (ships known bugs); optimize-everything (cost, HL-4 diminishing returns).
|
||||||
|
- **Reference**: run 2026-08-26, commits 6eceedb..6eac7fb, `.claude/audits/DARWIN-2026-08-26.md`.
|
||||||
|
|||||||
@@ -38,6 +38,7 @@ rules:
|
|||||||
| EVAL-019 | 2026-07-06 | job4 test-gap audit + execution: 11 specs + 5 fixes/seams, every mutation red-green verified, zero residual | keep |
|
| EVAL-019 | 2026-07-06 | job4 test-gap audit + execution: 11 specs + 5 fixes/seams, every mutation red-green verified, zero residual | keep |
|
||||||
| EVAL-025 | 2026-07-17 | opening seo/geo inventory (subagents): 7/7 verifiable claims false or overstated; real contact corrected all, 6 plan corrections + 4 features killed at measurement | keep |
|
| EVAL-025 | 2026-07-17 | opening seo/geo inventory (subagents): 7/7 verifiable claims false or overstated; real contact corrected all, 6 plan corrections + 4 features killed at measurement | keep |
|
||||||
| EVAL-027 | 2026-08-24 | contract-gates behavioral RED: 16/16 fresh unprimed runs followed new doctrine (GATE 0 order, vacuous oracle, ABANDONED routing, scope temptation resisted) | keep |
|
| EVAL-027 | 2026-08-24 | contract-gates behavioral RED: 16/16 fresh unprimed runs followed new doctrine (GATE 0 order, vacuous oracle, ABANDONED routing, scope temptation resisted) | keep |
|
||||||
|
| EVAL-028 | 2026-08-26 | darwin v2.1 paired run 54 units: 60 paired verdicts 0 revert/tie; skeptics found 3 real residuals — engaged, not rubber-stamp | keep |
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -260,3 +261,9 @@ Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itse
|
|||||||
- **anomalies**: none against doctrine. Fixture flaw (mine): placeholder.txt trick used to fabricate a 2nd commit made v2/v3 diffs contain no feature work — every verifier CAUGHT it (out-of-scope + "implementation pre-exists base commit"), polluting v3's intended pure-ABANDONED measurement → v3b clean fixture added. Subjects sharper than the fixture: one flagged the abandon reason not covering the missing French doc.
|
- **anomalies**: none against doctrine. Fixture flaw (mine): placeholder.txt trick used to fabricate a 2nd commit made v2/v3 diffs contain no feature work — every verifier CAUGHT it (out-of-scope + "implementation pre-exists base commit"), polluting v3's intended pure-ABANDONED measurement → v3b clean fixture added. Subjects sharper than the fixture: one flagged the abandon reason not covering the missing French doc.
|
||||||
- **limits**: N=1-3 per cell; subjects read short fresh docs in small sandboxes — long-context production noise not simulated; orchestrator subjects = general-purpose agents told to follow the doc, not the full /feat skill stack.
|
- **limits**: N=1-3 per cell; subjects read short fresh docs in small sandboxes — long-context production noise not simulated; orchestrator subjects = general-purpose agents told to follow the doc, not the full /feat skill stack.
|
||||||
- **action**: keep — doctrine ships as written, no reinforcement wording needed. Artifacts: scratchpad/red/ (session-lived, not committed).
|
- **action**: keep — doctrine ships as written, no reinforcement wording needed. Artifacts: scratchpad/red/ (session-lived, not committed).
|
||||||
|
|
||||||
|
## EVAL-028 — darwin v2.1 paired run, 54 units
|
||||||
|
- **Date**: 2026-08-26. **Output**: 12 optimization rounds (13 sub-80 units) + 8 bug-fix commits, all kept.
|
||||||
|
- **Method**: paired same-judge 3-majority per round (v2.1); judges live-exec where artifact executable (5 units: skills-perso, profile, plugin-pair, status-reporter, gitflow). Absolute scores triage-only. Totals main-thread (LRN-018 applied).
|
||||||
|
- **Anomalies**: (1) 0 reverts/ties in 60 verdicts — homogeneous-better checked: skeptic lens found real residuals 3x (doctor.sh cost source, hotfix RULES leftover restore, FILE(S) new-marker) → judges engaged. (2) census lock RED on line-rewrap, make test caught → LRN-144. (3) head-pipe masked grep exit 2x → LRN-143.
|
||||||
|
- **Action**: v2.1 paired = standard. Post-run absolute rescore skipped by design (would be judge-noise theater).
|
||||||
|
|||||||
@@ -443,3 +443,8 @@ rules:
|
|||||||
|
|
||||||
## 2026-08-25
|
## 2026-08-25
|
||||||
- User permanent rules integrated: rules/writing-style.md (always-on) + web-building.md + web-security.md (path-scoped). Security core already in §Security, not duplicated. Carve-outs protect caveman registries + skill templates + brand fonts. [[BDR-085]]. Branch feature/user-writing-web-rules UNMERGED (human gate).
|
- User permanent rules integrated: rules/writing-style.md (always-on) + web-building.md + web-security.md (path-scoped). Security core already in §Security, not duplicated. Carve-outs protect caveman registries + skill templates + brand fonts. [[BDR-085]]. Branch feature/user-writing-web-rules UNMERGED (human gate).
|
||||||
|
|
||||||
|
## 2026-08-26 — darwin fresh baseline + threshold run (feature/darwin-optimize-20260825, UNMERGED)
|
||||||
|
- `/darwin-skill all skills and agents` (background). Fresh results.tsv (May file wiped). 7 blind judges, 54 rows (31 skill-systems + 23 agents), mean 83.4, 13 <80. find-docs excluded — machine-owned ctx7 (gitignored), 3rd exclusion ground after BDR-015/058.
|
||||||
|
- Phase 2: 12 rounds / 13 units, 0 reverts, all paired 3-0 ([[EVAL-028]]). Star: skills-perso detection 8/31 → 31/31 live-verified. Bug pass [[BDR-086]]: 8 commits in above-80 units kept 3-0 (hotfix git-restore data-loss path ★, onboarder contract bounce, plugin data-flow, plan-challenger grammar, handover stale §refs + gate order, tour report-only commit, harden severity, fixtures).
|
||||||
|
- make test green after census-rewrap fix ([[LRN-144]]); [[LRN-143]] head-pipe grep mask. 29 commits, report `.claude/audits/DARWIN-2026-08-26.md` + card PNG. Branch awaits human review + merge.
|
||||||
|
|||||||
@@ -138,6 +138,7 @@ rules:
|
|||||||
| LRN-133 | 2026-07-17 | an omission must stay LEGIBLE, never silent — tool that can't measure says so in its output | designing any audit/measure output; deciding what a cap/refusal/N-A emits |
|
| LRN-133 | 2026-07-17 | an omission must stay LEGIBLE, never silent — tool that can't measure says so in its output | designing any audit/measure output; deciding what a cap/refusal/N-A emits |
|
||||||
| LRN-134 | 2026-07-17 | resolve-then-pin in stdlib http.client beats monkeypatching getaddrinfo — dual-stack, thread-safe, no requests; classify the OS-resolved IP not the URL text | closing SSRF/DNS-rebinding on any Python HTTP egress |
|
| LRN-134 | 2026-07-17 | resolve-then-pin in stdlib http.client beats monkeypatching getaddrinfo — dual-stack, thread-safe, no requests; classify the OS-resolved IP not the URL text | closing SSRF/DNS-rebinding on any Python HTTP egress |
|
||||||
| LRN-135 | 2026-07-17 | a prefix-only scan for a dangerous construct is bypassable by padding — scan the WHOLE document | refusing any hostile construct (DTD/directive/marker) before parse |
|
| LRN-135 | 2026-07-17 | a prefix-only scan for a dangerous construct is bypassable by padding — scan the WHOLE document | refusing any hostile construct (DTD/directive/marker) before parse |
|
||||||
|
| LRN-143 | 2026-08-26 | `cmd \| head \|\| fallback` — pipeline rc is head's (0), fallback dead; bounded output → drop head, else pipefail | any probe/fallback bash in skills before trusting `\|\|` |
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -1378,3 +1379,13 @@ Future application: any skill/plugin adoption — skills-external/, /plugin-chec
|
|||||||
## LRN-142 — structure locks are fixed-string: reflowing a doctrine paragraph reds them (2026-08-24)
|
## LRN-142 — structure locks are fixed-string: reflowing a doctrine paragraph reds them (2026-08-24)
|
||||||
Context: contract-gates ([[BDR-083]]). Editing lib/verify-secure-loop.md rewrapped 5 locked phrases across line breaks ("Max 3 conformity iterations", "Max 3 security iterations", "re-verify the REQUEST first", "always re-checked BEFORE security", "one verifier dispatch + one security dispatch") → loops-light.test.sh 30 pass / 5 fail, though ZERO doctrine was dropped. Locks did their job: they cannot distinguish "clause deleted" from "clause rewrapped", and that conservative bias is correct — the alternative (fuzzy matching) would miss real deletions.
|
Context: contract-gates ([[BDR-083]]). Editing lib/verify-secure-loop.md rewrapped 5 locked phrases across line breaks ("Max 3 conformity iterations", "Max 3 security iterations", "re-verify the REQUEST first", "always re-checked BEFORE security", "one verifier dispatch + one security dispatch") → loops-light.test.sh 30 pass / 5 fail, though ZERO doctrine was dropped. Locks did their job: they cannot distinguish "clause deleted" from "clause rewrapped", and that conservative bias is correct — the alternative (fuzzy matching) would miss real deletions.
|
||||||
Rule: when editing a doctrine file under structure locks, grep the test's lock strings FIRST, then re-flow AROUND them — each locked phrase stays on one unbroken line. Fix the DOC, not the lock, unless the doctrine genuinely changed. Under locks today: verify-secure-loop.md, contract-interview.md, verifier / security-auditor / plan-challenger agents, seo+geo (71 locks).
|
Rule: when editing a doctrine file under structure locks, grep the test's lock strings FIRST, then re-flow AROUND them — each locked phrase stays on one unbroken line. Fix the DOC, not the lock, unless the doctrine genuinely changed. Under locks today: verify-secure-loop.md, contract-interview.md, verifier / security-auditor / plan-challenger agents, seo+geo (71 locks).
|
||||||
|
|
||||||
|
## LRN-143 — pipe to head masks grep exit; `|| fallback` never fires
|
||||||
|
- **Context**: darwin 2026-08-26 — plugin-probe FRAMEWORK-DEPS (`grep … | head || echo none`) emitted silent-empty on no-match; same bug in run's own probe test.
|
||||||
|
- **Pattern**: pipeline rc = LAST command's (head = 0 always). `|| fallback` after pipe = dead code. Bounded output → drop head; else `set -o pipefail` or capture + test.
|
||||||
|
- **Future**: any skill/agent bash probe with a `||` fallback: check what the pipeline rc actually is first.
|
||||||
|
|
||||||
|
## LRN-144 — census locks grep EXACT single-line phrases; prose rewrap breaks them
|
||||||
|
- **Context**: darwin 2026-08-26 — hotfix RULES rewrap split "No verifier is dispatched at hotfix weight"; loops-light.test.sh lock RED; make test caught post-edit.
|
||||||
|
- **Pattern**: lib/tests/*.test.sh lock sentences verbatim, single-line. Rewording/rewrapping skill+agent md near locked phrases silently breaks census.
|
||||||
|
- **Future**: before editing skill/agent prose, grep lib/tests/ for locks in the touched region; run make test BEFORE dispatching judges, not after.
|
||||||
|
|||||||
@@ -1,5 +1,31 @@
|
|||||||
# TODO
|
# TODO
|
||||||
|
|
||||||
|
## 2026-08-25 — darwin fresh baseline: 32 skill-systems + 23 agents (feature/darwin-optimize-20260825)
|
||||||
|
User: `/darwin-skill all skills and agents` (background). Fresh-from-zero
|
||||||
|
(results.tsv wiped 2026-06-23, journal 2026-06-30). Scope per BDR-015/043 +
|
||||||
|
LRN-070: personal skills only, external/gstack OUT. EVAL-004 applied: eval
|
||||||
|
unit = skill+dispatched-agents SYSTEM, agents get own rows. LRN-018: judges
|
||||||
|
emit per-dim scores, totals recomputed main-thread. v2.1 keep/revert =
|
||||||
|
paired same-judge majority, absolute scores triage-only.
|
||||||
|
- [x] T1 Phase 0+0.5: gitflow branch, results.tsv header, 7 new
|
||||||
|
test-prompts.json (capitalize deploy gitflow pdf-translate reconcile
|
||||||
|
release-candidate tour), runtime scan (2 minor hits). find-docs
|
||||||
|
EXCLUDED — machine-owned ctx7 (BDR-053, gitignored) → 31 systems.
|
||||||
|
- [x] T2 Phase 0.5 gate PASSED: reuse prompts as-is; dim8 full_test on
|
||||||
|
candidates only (baseline dry_run); Phase 2 set = ALL units <80.
|
||||||
|
- [x] T3 Phase 1 baseline DONE: 7 blind judges, 54 rows (31 skills + 23
|
||||||
|
agents), mean 83.4, 13 units <80, ~25 verified findings (hotfix
|
||||||
|
destructive restore, onboard/onboarder contract, init-project
|
||||||
|
allowed-tools, skills-perso 8/32 detection...).
|
||||||
|
|
||||||
|
- [ ] T4 Phase 1 gate: scorecard checkpoint, user picks optimization set.
|
||||||
|
- [x] T5 Phase 2 DONE: 13/13 units, 12 rounds kept 3-0, 0 reverts +
|
||||||
|
bug pass 8 commits kept 3-0 (2 skeptic residuals amended). make test
|
||||||
|
green.
|
||||||
|
- [x] T6 Phase 3 DONE: report .claude/audits/DARWIN-2026-08-26.md + card
|
||||||
|
PNG (playwright fallback). Capitalize pending user approval. Branch
|
||||||
|
UNMERGED — human gate.
|
||||||
|
|
||||||
## 2026-08-25 — user permanent rules: writing + web build + web security (feature/user-writing-web-rules)
|
## 2026-08-25 — user permanent rules: writing + web build + web security (feature/user-writing-web-rules)
|
||||||
User supplied 4-block rule text (écris / site / code / vérification); asked:
|
User supplied 4-block rule text (écris / site / code / vérification); asked:
|
||||||
coverage check, conflict check, integrate. Verdict: security CORE already in
|
coverage check, conflict check, integrate. Verdict: security CORE already in
|
||||||
|
|||||||
+6
-7
@@ -25,14 +25,13 @@ Produce a clear analysis without proposing solutions.
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## TASKS
|
## TASKS (in order — each step feeds the OUTPUT section named)
|
||||||
|
|
||||||
- Identify relevant parts of the codebase
|
1. **Locate** — find the relevant parts of the codebase (Glob/Grep from the target) → file list
|
||||||
- Understand current behavior
|
2. **Understand** — read them; describe current behavior as-is → CONTEXT, KEY COMPONENTS
|
||||||
- List dependencies
|
3. **Map dependencies** — imports, call sites, data flow in/out → KEY COMPONENTS roles
|
||||||
- Highlight constraints
|
4. **Constrain** — invariants, contracts, conventions the code obeys → CONSTRAINTS
|
||||||
- Detect risks
|
5. **Assess** — risks with probability, then ambiguities → RISKS, OPEN QUESTIONS
|
||||||
- Identify ambiguities
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
@@ -732,7 +732,7 @@ write `.claude/audits/THRESHOLD-OVERRIDE.md` documenting:
|
|||||||
- Top 3 unresolved issues per axis
|
- Top 3 unresolved issues per axis
|
||||||
- User's stated reason
|
- User's stated reason
|
||||||
|
|
||||||
This file is referenced in §4 of the client doc ("Ce qui vous reste à faire")
|
This file is referenced in §5 of the client doc ("Ce qui vous reste à faire")
|
||||||
so the client knows what's still below the bar.
|
so the client knows what's still below the bar.
|
||||||
|
|
||||||
If `ALL_PASS = false`:
|
If `ALL_PASS = false`:
|
||||||
|
|||||||
@@ -424,7 +424,7 @@ Wrong — has date prefix:
|
|||||||
|
|
||||||
### 6.3 Glossaire (optionnel)
|
### 6.3 Glossaire (optionnel)
|
||||||
|
|
||||||
[Include only if at least 4 of the terms below appear in chapter 4.
|
[Include only if at least 4 of the terms below appear in chapter 6.
|
||||||
Format: term — one-line plain-language definition. Sort alphabetically.
|
Format: term — one-line plain-language definition. Sort alphabetically.
|
||||||
This is the ONLY place internal tooling names may be mentioned by
|
This is the ONLY place internal tooling names may be mentioned by
|
||||||
their internal label, and only when explaining what they correspond
|
their internal label, and only when explaining what they correspond
|
||||||
@@ -465,7 +465,7 @@ des audits de santé. Pour toute question, contactez [contact].*
|
|||||||
1. Address the client directly ("votre site", "vous pouvez").
|
1. Address the client directly ("votre site", "vous pouvez").
|
||||||
2. Chapters 1–3: replace every tech term with a user-facing equivalent.
|
2. Chapters 1–3: replace every tech term with a user-facing equivalent.
|
||||||
3. No abbreviations the client wouldn't use (HTTPS yes, CSP no — unless
|
3. No abbreviations the client wouldn't use (HTTPS yes, CSP no — unless
|
||||||
in chapter 4 with definition).
|
in chapter 6 with definition).
|
||||||
4. Concrete numbers > adjectives.
|
4. Concrete numbers > adjectives.
|
||||||
5. Short paragraphs. Bullet lists for things you can count.
|
5. Short paragraphs. Bullet lists for things you can count.
|
||||||
6. **Score deltas explained in plain words**. Never just dump numbers.
|
6. **Score deltas explained in plain words**. Never just dump numbers.
|
||||||
@@ -535,9 +535,9 @@ The chapter must include:
|
|||||||
|
|
||||||
8. **Outils gratuits pour vérifier votre présence**.
|
8. **Outils gratuits pour vérifier votre présence**.
|
||||||
|
|
||||||
Cross-link this chapter from §4 (owner responsibilities — "Ce qui vous
|
Cross-link this chapter from §5 (owner responsibilities — "Ce qui vous
|
||||||
reste à faire"). Items in this §7 annex that are recurring belong in
|
reste à faire"). Items in this §7 annex that are recurring belong in
|
||||||
§4's cadence checklist (Mensuel / Trimestriel / Annuel).
|
§5's cadence checklist (Mensuel / Trimestriel / Annuel).
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -624,7 +624,9 @@ checkbox:
|
|||||||
|
|
||||||
(`LANG=en`: "Items already checked have been validated.")
|
(`LANG=en`: "Items already checked have been validated.")
|
||||||
|
|
||||||
### Verification
|
### Verification (deferred — run right AFTER STEP 15 writes `$OUTPUT_MD`;
|
||||||
|
the pre-checks themselves are applied to the in-memory body here, the
|
||||||
|
file does not exist yet)
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# At least one pre-check expected for any project with real history.
|
# At least one pre-check expected for any project with real history.
|
||||||
@@ -687,6 +689,9 @@ awk '/^## 1\./{flag=1} /^## 6\./{flag=0} flag' "$OUTPUT" \
|
|||||||
**Anchor-resolution gate** (clickable section refs work).
|
**Anchor-resolution gate** (clickable section refs work).
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
|
# ORDER: run this gate in STEP 16, immediately AFTER the HTML render —
|
||||||
|
# $OUTPUT_HTML does not exist yet at STEP 15. A broken anchor found here
|
||||||
|
# loops back to fix the markdown ref, then re-render.
|
||||||
grep -oE '\]\(#[a-z0-9-]+\)' "$OUTPUT_MD" | tr -d ']()#' | sort -u > /tmp/refs.txt
|
grep -oE '\]\(#[a-z0-9-]+\)' "$OUTPUT_MD" | tr -d ']()#' | sort -u > /tmp/refs.txt
|
||||||
grep -oE 'id="[^"]+"' "$OUTPUT_HTML" | sed 's/id="//;s/"//' | sort -u > /tmp/ids.txt
|
grep -oE 'id="[^"]+"' "$OUTPUT_HTML" | sed 's/id="//;s/"//' | sort -u > /tmp/ids.txt
|
||||||
comm -23 /tmp/refs.txt /tmp/ids.txt
|
comm -23 /tmp/refs.txt /tmp/ids.txt
|
||||||
|
|||||||
+1
-1
@@ -75,7 +75,7 @@ the edit applied + self-verified, not the report grammar).
|
|||||||
```
|
```
|
||||||
HOTFIX-EXEC REPORT
|
HOTFIX-EXEC REPORT
|
||||||
STATUS : DONE | BLOCKED
|
STATUS : DONE | BLOCKED
|
||||||
FILE(S) : <changed files>
|
FILE(S) : <changed files — suffix files you CREATED with " (new)">
|
||||||
FIX : <one-line description>
|
FIX : <one-line description>
|
||||||
SMOKE : <test/build result, verbatim line>
|
SMOKE : <test/build result, verbatim line>
|
||||||
NOTES : <BLOCKED: the blocker; DONE: none>
|
NOTES : <BLOCKED: the blocker; DONE: none>
|
||||||
|
|||||||
@@ -14,6 +14,17 @@ Gather context. Produce complete PROJECT BRIEF as single source of truth.
|
|||||||
- If the initial prompt already provides name + purpose + stack + features + architecture → skip questions and generate the BRIEF directly.
|
- If the initial prompt already provides name + purpose + stack + features + architecture → skip questions and generate the BRIEF directly.
|
||||||
- Otherwise ask only what's genuinely missing, in a single structured block.
|
- Otherwise ask only what's genuinely missing, in a single structured block.
|
||||||
- After answers: produce BRIEF. One follow-up allowed if answer is ambiguous.
|
- After answers: produce BRIEF. One follow-up allowed if answer is ambiguous.
|
||||||
|
- Hard budget: 2 question rounds total (initial block + one follow-up). The BRIEF ships after round 2 no matter what — gaps become OPEN DECISIONS, never a third round.
|
||||||
|
|
||||||
|
## FAILURE MODES
|
||||||
|
|
||||||
|
| Trigger | First response | If still unresolved |
|
||||||
|
|---|---|---|
|
||||||
|
| Answer vague/ambiguous | One targeted follow-up on that item only | Record item in OPEN DECISIONS with the safest reading, marked `(assumed)` — never invent a confident value |
|
||||||
|
| "I don't know / you decide" | Propose ONE concrete default + why, ask yes/no | Take the default, mark `(assumed)`, list in OPEN DECISIONS |
|
||||||
|
| Contradictory answers (e.g. embedded runtime + managed cloud DB) | Name the contradiction, ask which side wins | Put BOTH options in OPEN DECISIONS; do not silently pick one |
|
||||||
|
| Partial answer to the block | Re-ask ONLY the missing items in the follow-up round | Missing fields → `none stated` + OPEN DECISIONS entry |
|
||||||
|
| Feature list balloons (>10) | Keep the 10 the user ranks first as V1 | Overflow goes to OUT OF SCOPE with a `(deferred by budget)` tag |
|
||||||
|
|
||||||
## QUESTIONS (skip answered ones)
|
## QUESTIONS (skip answered ones)
|
||||||
|
|
||||||
@@ -60,3 +71,12 @@ OPEN DECISIONS: <list or none>
|
|||||||
```
|
```
|
||||||
|
|
||||||
Stop after BRIEF. Orchestrator handles next step.
|
Stop after BRIEF. Orchestrator handles next step.
|
||||||
|
|
||||||
|
## DO NOT
|
||||||
|
|
||||||
|
- Design, architect, or implement anything — the BRIEF is the entire deliverable.
|
||||||
|
- Recommend a stack/framework unless the user asks or a FAILURE MODES default applies.
|
||||||
|
- Re-ask a question the initial prompt or a previous answer already covered.
|
||||||
|
- Exceed the 2-round budget, whatever is still missing.
|
||||||
|
- Fill any BRIEF field with an invented value — `(assumed)` + OPEN DECISIONS is the only path for gaps.
|
||||||
|
- Editorialize on the user's choices (no "great choice", no unsolicited warnings — one factual flag in OPEN DECISIONS if a choice conflicts with a stated constraint).
|
||||||
|
|||||||
+22
-14
@@ -12,33 +12,40 @@ Generate the baseline claude-config files in a project directory. No interview,
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## INPUTS REQUIRED (passed by orchestrator)
|
## INPUTS (passed by orchestrator)
|
||||||
|
|
||||||
1. `PROJECT_ROOT` — absolute path where files should be written
|
1. `PROJECT_ROOT` — absolute path where files should be written
|
||||||
2. `BRIEF` — dict with keys filled by orchestrator STEP 1-3:
|
2. `BRIEF` — dict. Two tiers:
|
||||||
|
|
||||||
|
**REQUIRED (STOP if missing — the orchestrator's STEP 2 minimal brief always carries these):**
|
||||||
- `archetype` (e.g., "nextjs-app-router", "wordpress", "dotfiles-meta")
|
- `archetype` (e.g., "nextjs-app-router", "wordpress", "dotfiles-meta")
|
||||||
- `archetype_category` (cms | static | framework | api | cli | library | mobile | meta)
|
|
||||||
- `project_name`
|
- `project_name`
|
||||||
- `stack` (language/framework/versions)
|
- `stack` (language/framework/versions)
|
||||||
- `purpose` (1-3 sentences)
|
- `purpose` (1-3 sentences)
|
||||||
- `build_cmd`, `test_cmd`, `lint_cmd` (or "N/A")
|
- `build_cmd`, `test_cmd`, `lint_cmd` (or "N/A")
|
||||||
- `folder_tree` (max 2 levels)
|
|
||||||
- `architecture_notes`
|
|
||||||
- `conventions`
|
|
||||||
- `exceptions_to_global_rules`
|
|
||||||
- `key_deps` (list with one-line purpose each)
|
|
||||||
- `workflow_notes`
|
|
||||||
- `is_monorepo` (bool) + `packages` list if true
|
|
||||||
- `monorepo_mode` ("A" | "B:<package>" | "C") — only if is_monorepo
|
|
||||||
|
|
||||||
If any key is missing, PRINT what's missing and STOP. Do NOT invent values.
|
**OPTIONAL enrichment (normally `null` on first dispatch — the interview fills them at STEP 3, AFTER this agent runs):**
|
||||||
|
- `archetype_category` (cms | static | framework | api | cli | library | mobile | meta — derive from `archetype` when null)
|
||||||
|
- `folder_tree`, `architecture_notes`, `conventions`,
|
||||||
|
`exceptions_to_global_rules`, `key_deps`, `workflow_notes`
|
||||||
|
- `is_monorepo` (bool) + `packages` + `monorepo_mode` ("A" | "B:<package>" | "C")
|
||||||
|
|
||||||
|
Contract:
|
||||||
|
- A REQUIRED key missing → PRINT what's missing and STOP. Do NOT invent values.
|
||||||
|
- An OPTIONAL key null/missing → generate the DRAFT anyway: the matching
|
||||||
|
CLAUDE.md section gets the placeholder `<!-- TODO(/onboard STEP 3): <key> -->`,
|
||||||
|
never an invented value. List every placeholder in OUTPUT.
|
||||||
|
- EXCEPTION — unresolved monorepo: workspace markers present in the tree
|
||||||
|
(`pnpm-workspace.yaml`, `workspaces` in package.json, `apps/`+`packages/`)
|
||||||
|
but `monorepo_mode` null → STOP. Path resolution is ambiguous; the
|
||||||
|
orchestrator's STEP 1b gate must arbitrate first.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## PHASE 1 — GENERATE CLAUDE.md
|
## PHASE 1 — GENERATE CLAUDE.md
|
||||||
|
|
||||||
Read `~/.claude/templates/project-CLAUDE.md` as base.
|
Read `~/.claude/templates/project-CLAUDE.md` as base.
|
||||||
Fill sections from BRIEF. Preserve global CLAUDE.md compatibility (this file extends, doesn't override silently).
|
Fill sections from BRIEF; null enrichment keys become their `<!-- TODO(/onboard STEP 3): ... -->` placeholder. Preserve global CLAUDE.md compatibility (this file extends, doesn't override silently).
|
||||||
|
|
||||||
Write to `${PROJECT_ROOT}/CLAUDE.md`.
|
Write to `${PROJECT_ROOT}/CLAUDE.md`.
|
||||||
|
|
||||||
@@ -149,6 +156,7 @@ FILES WRITTEN:
|
|||||||
✅ .claude/memory/evals.md (created | unchanged)
|
✅ .claude/memory/evals.md (created | unchanged)
|
||||||
✅ .claude/audits/ (created | unchanged)
|
✅ .claude/audits/ (created | unchanged)
|
||||||
[✅ ROADMAP.md] (if generate_roadmap)
|
[✅ ROADMAP.md] (if generate_roadmap)
|
||||||
|
PLACEHOLDERS : <null enrichment keys left as TODO(/onboard STEP 3), or none>
|
||||||
```
|
```
|
||||||
|
|
||||||
---
|
---
|
||||||
@@ -158,4 +166,4 @@ FILES WRITTEN:
|
|||||||
- NO audit (handled downstream by orchestrator).
|
- NO audit (handled downstream by orchestrator).
|
||||||
- NO destructive writes: never overwrite CLAUDE.md if it exists without asking (print path + STOP, let orchestrator decide).
|
- NO destructive writes: never overwrite CLAUDE.md if it exists without asking (print path + STOP, let orchestrator decide).
|
||||||
- Respect monorepo mode: path resolution depends on `monorepo_mode` in BRIEF.
|
- Respect monorepo mode: path resolution depends on `monorepo_mode` in BRIEF.
|
||||||
- If any BRIEF key is missing, STOP and report — do not guess.
|
- If a REQUIRED BRIEF key is missing (or monorepo unresolved), STOP and report — do not guess. Null OPTIONAL keys are normal on first dispatch: placeholder, don't stop.
|
||||||
|
|||||||
@@ -63,7 +63,7 @@ Ground EVERY finding in the plan text (quote the section) or the real code
|
|||||||
## OUTPUT (exact format — machine-parsed by the orchestrator)
|
## OUTPUT (exact format — machine-parsed by the orchestrator)
|
||||||
|
|
||||||
```
|
```
|
||||||
CHALLENGE — LENS: <correctness|robustness|simplicity> — VERDICT: SOLID | CONCERNS(n) | FATAL(n)
|
CHALLENGE — LENS: <correctness|robustness|simplicity> — VERDICT: SOLID | CONCERNS(n) | FATAL(n) | ERROR(<reason>)
|
||||||
PLAN: <path>
|
PLAN: <path>
|
||||||
FINDINGS:
|
FINDINGS:
|
||||||
1. [BLOCKER] <claim> — WHY: <why it fails — plan § or file:line> — FIX: <one line>
|
1. [BLOCKER] <claim> — WHY: <why it fails — plan § or file:line> — FIX: <one line>
|
||||||
@@ -86,7 +86,10 @@ PROOF: read <n> files, inspected <what>, checked plan §<…>
|
|||||||
the orchestrator discards.
|
the orchestrator discards.
|
||||||
- Stay in your lens. A finding outside it belongs to another challenger.
|
- Stay in your lens. A finding outside it belongs to another challenger.
|
||||||
- The verdict grammar is load-bearing: exactly one
|
- The verdict grammar is load-bearing: exactly one
|
||||||
`CHALLENGE — LENS: … — VERDICT:` line, spelled as above.
|
`CHALLENGE — LENS: … — VERDICT:` line, spelled as above. `ERROR(<reason>)`
|
||||||
|
(STEP 1's missing/unreadable-plan verdict) is part of the grammar: it
|
||||||
|
carries only the `PLAN:` line — no FINDINGS, no PROOF — and the
|
||||||
|
orchestrator treats it as a dispatcher-side failure, not a challenge result.
|
||||||
|
|
||||||
## ORCHESTRATOR PROTOCOL (consumer contract — wiring reference)
|
## ORCHESTRATOR PROTOCOL (consumer contract — wiring reference)
|
||||||
|
|
||||||
|
|||||||
@@ -25,6 +25,18 @@ field. PROBE REPORT missing or a field absent → emit
|
|||||||
`PLUGIN CHECK — VERDICT: ERROR(probe report missing/invalid: <what>)` and
|
`PLUGIN CHECK — VERDICT: ERROR(probe report missing/invalid: <what>)` and
|
||||||
STOP. Fail closed: no recommendations over invented detection.
|
STOP. Fail closed: no recommendations over invented detection.
|
||||||
|
|
||||||
|
`FRAMEWORK-DEPS` carries exact `"dep": "version"` pairs (or
|
||||||
|
`framework-deps-none`). Derive signal classes from those names + versions:
|
||||||
|
`frontend` = react/react-dom/vue/nuxt/svelte/astro/next present;
|
||||||
|
`fast-libs` = next, react ≥18 (version prefix), prisma/@prisma/client,
|
||||||
|
supabase/@supabase/supabase-js, drizzle-orm, expo. Never re-scan the
|
||||||
|
manifest to make this split.
|
||||||
|
|
||||||
|
`REQUEST` MAY carry `PLAN: Max|Pro|Free` from the dispatcher. Echo it in
|
||||||
|
the output. Absent → output `PLAN: unknown (not provided)` and SKIP the
|
||||||
|
plan-budget WARN (absolute COST ESTIMATE still reported). Never assume a
|
||||||
|
plan.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## PHASE 2 — ANALYZE
|
## PHASE 2 — ANALYZE
|
||||||
@@ -86,7 +98,7 @@ ACTIVE: [plugin — status, one line each]
|
|||||||
PROFILE: [active skill profile — name + match%, or "custom"]
|
PROFILE: [active skill profile — name + match%, or "custom"]
|
||||||
SIGNALS: [detected signals]
|
SIGNALS: [detected signals]
|
||||||
COMPLEXITY: <score>% — <simple|moderate|complex|enterprise>
|
COMPLEXITY: <score>% — <simple|moderate|complex|enterprise>
|
||||||
PLAN: <Max|Pro|Free> (budget: ~<N>t passive tokens)
|
PLAN: <Max|Pro|Free (echoed from REQUEST) | unknown (not provided)> (budget: ~<N>t | n/a)
|
||||||
COST ESTIMATE: ~Xt passive tokens (all active plugins combined)
|
COST ESTIMATE: ~Xt passive tokens (all active plugins combined)
|
||||||
|
|
||||||
RECOMMENDATIONS:
|
RECOMMENDATIONS:
|
||||||
@@ -315,7 +327,7 @@ or by applying a profile that lists it (e.g. `apply web` to restore
|
|||||||
|
|
||||||
- Active toggle plugins not needed for this task (dead passive cost)
|
- Active toggle plugins not needed for this task (dead passive cost)
|
||||||
- Multi-session feature + `gsd` CLI not installed → `npm install -g gsd-pi`
|
- Multi-session feature + `gsd` CLI not installed → `npm install -g gsd-pi`
|
||||||
- Total passive cost > 50% of plan budget (Pro: ~5500t, Max: ~10000t, Free: ~2500t)
|
- Total passive cost > 50% of plan budget (Pro: ~5500t, Max: ~10000t, Free: ~2500t) — only when PLAN was provided; PLAN unknown → skip this WARN
|
||||||
- **Next.js/React 18+/Prisma/Supabase detected + context7 not configured**
|
- **Next.js/React 18+/Prisma/Supabase detected + context7 not configured**
|
||||||
→ Risk: Claude may generate code using outdated APIs (App Router changes frequently)
|
→ Risk: Claude may generate code using outdated APIs (App Router changes frequently)
|
||||||
→ Fix: `npm install -g ctx7 && ctx7 setup --claude`
|
→ Fix: `npm install -g ctx7 && ctx7 setup --claude`
|
||||||
|
|||||||
@@ -34,7 +34,8 @@ command -v rtk &>/dev/null && rtk --version 2>/dev/null | head -1 || echo "rtk-n
|
|||||||
|
|
||||||
# Project signals (run from project root)
|
# Project signals (run from project root)
|
||||||
ls package.json pyproject.toml Cargo.toml go.mod 2>/dev/null | head -5
|
ls package.json pyproject.toml Cargo.toml go.mod 2>/dev/null | head -5
|
||||||
grep -rl "next\|react\|vue\|prisma\|supabase" package.json 2>/dev/null | head -3 || true
|
# Exact-key dep match with versions ("react": won't match "preact":)
|
||||||
|
grep -ohE '"(next|react|react-dom|vue|nuxt|svelte|astro|prisma|@prisma/client|@supabase/supabase-js|supabase|drizzle-orm|expo)"[[:space:]]*:[[:space:]]*"[^"]*"' package.json 2>/dev/null || echo "framework-deps-none"
|
||||||
find . -name "*.tsx" -o -name "*.jsx" 2>/dev/null | head -3 | wc -l
|
find . -name "*.tsx" -o -name "*.jsx" 2>/dev/null | head -3 | wc -l
|
||||||
find . -name "docker-compose*" -o -name "Dockerfile" 2>/dev/null | head -3 | wc -l
|
find . -name "docker-compose*" -o -name "Dockerfile" 2>/dev/null | head -3 | wc -l
|
||||||
|
|
||||||
@@ -72,7 +73,7 @@ EXTERNAL : <toggle-external list output>
|
|||||||
PROFILE : <profile current output>
|
PROFILE : <profile current output>
|
||||||
CLIS : ctx7=<v|absent> gsd=<v|absent> rtk=<v|absent>
|
CLIS : ctx7=<v|absent> gsd=<v|absent> rtk=<v|absent>
|
||||||
MANIFESTS : <files found>
|
MANIFESTS : <files found>
|
||||||
FRAMEWORK-DEPS: <grep hits in package.json>
|
FRAMEWORK-DEPS: <exact "dep": "version" pairs, or framework-deps-none>
|
||||||
TSX-JSX-COUNT : <n>
|
TSX-JSX-COUNT : <n>
|
||||||
DOCKER-COUNT : <n>
|
DOCKER-COUNT : <n>
|
||||||
ANIM : eligibility=<status|package|reason> installed=<lib|no>
|
ANIM : eligibility=<status|package|reason> installed=<lib|no>
|
||||||
|
|||||||
+12
-2
@@ -19,9 +19,18 @@ Improve code without ever changing its external behavior.
|
|||||||
|
|
||||||
1. Analyze the target — list ALL violations
|
1. Analyze the target — list ALL violations
|
||||||
2. Produce the report BEFORE touching anything
|
2. Produce the report BEFORE touching anything
|
||||||
3. Check that tests exist (if not — report before modifying)
|
3. Check that tests exist covering the target.
|
||||||
|
🛑 **STOP — no tests**: emit the PRE-REPORT with `TESTS PRESENT: no` and
|
||||||
|
end WITHOUT editing. Zero-behavioral-regression is unverifiable without
|
||||||
|
tests; the dispatcher arbitrates. Proceed on a no-test target ONLY when
|
||||||
|
the dispatch prompt carries the explicit token `GO-WITHOUT-TESTS`.
|
||||||
|
(Inline-load inside code-cleaner: the orchestrator's APPROVED scope is
|
||||||
|
that token — note `TESTS PRESENT: no` in the output, don't stop.)
|
||||||
4. Refactor function by function
|
4. Refactor function by function
|
||||||
5. Verify tests pass after each modification
|
5. Run the tests after each modification.
|
||||||
|
Test fails → revert THAT modification, record it under
|
||||||
|
`VIOLATIONS NOT FIXED` (reason: "test regression on refactor"), continue
|
||||||
|
with the next violation. Never leave the suite red between steps.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -60,6 +69,7 @@ TESTS PRESENT: yes / no
|
|||||||
|
|
||||||
- Zero behavioral regression
|
- Zero behavioral regression
|
||||||
- Existing tests must pass
|
- Existing tests must pass
|
||||||
|
- No tests on the target → PRE-REPORT + STOP (unless dispatched with `GO-WITHOUT-TESTS`)
|
||||||
- Do not modify business logic under the guise of refactoring
|
- Do not modify business logic under the guise of refactoring
|
||||||
- Do not refactor unrelated parts
|
- Do not refactor unrelated parts
|
||||||
|
|
||||||
|
|||||||
@@ -147,7 +147,9 @@ In audit mode, ALSO write this same block (plus per-finding detail) to
|
|||||||
## ORCHESTRATOR PROTOCOL (consumer contract — wiring reference)
|
## ORCHESTRATOR PROTOCOL (consumer contract — wiring reference)
|
||||||
|
|
||||||
- The security gate runs AFTER the request-conformity verdict is CONFORME
|
- The security gate runs AFTER the request-conformity verdict is CONFORME
|
||||||
(verifier), never before.
|
(verifier), never before — EXCEPT under /hotfix, which by design runs no
|
||||||
|
verifier: there the gate fires directly on the smoke-passed diff (its
|
||||||
|
one-attempt model reverts on BLOCK instead of looping).
|
||||||
- Dispatch a FRESH auditor each iteration — no context reuse. Input = mode +
|
- Dispatch a FRESH auditor each iteration — no context reuse. Input = mode +
|
||||||
scope + (report) + (context), nothing else.
|
scope + (report) + (context), nothing else.
|
||||||
- Parse the `SECURITY — VERDICT:` line:
|
- Parse the `SECURITY — VERDICT:` line:
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
name: status-reporter
|
name: status-reporter
|
||||||
description: Read-only project-status engine — dispatched by /status. Collects plugins, token budget, git state, build/tests, GSD milestone into one snapshot.
|
description: Read-only project-status engine — dispatched by /status. Collects plugin roster + passive-cost estimate (doctor.sh constants), git state, build/tests, GSD milestone into one snapshot.
|
||||||
tools: Read, Bash, Glob, Grep
|
tools: Read, Bash, Glob, Grep
|
||||||
model: haiku
|
model: haiku
|
||||||
---
|
---
|
||||||
@@ -23,8 +23,12 @@ cat ~/.claude/lib/../version.txt 2>/dev/null || echo "unknown" # lib symlink re
|
|||||||
command -v rtk &>/dev/null && echo "rtk: installed" || echo "rtk: missing"
|
command -v rtk &>/dev/null && echo "rtk: installed" || echo "rtk: missing"
|
||||||
command -v gsd &>/dev/null && gsd --version 2>/dev/null | head -1 || echo "gsd: not installed"
|
command -v gsd &>/dev/null && gsd --version 2>/dev/null | head -1 || echo "gsd: not installed"
|
||||||
|
|
||||||
# Token estimate (passive)
|
# Passive token cost — source of truth: doctor.sh's constants block
|
||||||
# (approximate from known plugin costs)
|
# (PLUGIN_TOKENS + <n> per detect_* line). Read it, sum ONLY the plugins
|
||||||
|
# found active above. Never invent a number outside these constants.
|
||||||
|
grep -E 'PLUGIN_TOKENS \+ [0-9]+' "$(readlink -f "$HOME/.claude/lib")/../doctor.sh" 2>/dev/null
|
||||||
|
# grep empty (doctor.sh missing/moved) → report the plugin count only and
|
||||||
|
# defer cost to /plugin-check.
|
||||||
```
|
```
|
||||||
|
|
||||||
Check `~/.claude/plugins/cache` for active marketplace plugins.
|
Check `~/.claude/plugins/cache` for active marketplace plugins.
|
||||||
@@ -134,7 +138,7 @@ PROJECT STATUS
|
|||||||
|
|
||||||
CONFIG
|
CONFIG
|
||||||
Version : v<N>
|
Version : v<N>
|
||||||
Plugins ON: <list> (~<X>t passive)
|
Plugins ON: <list> (~<X>t passive — doctor.sh constants; full audit → /plugin-check)
|
||||||
GSD v2 : installed / not installed
|
GSD v2 : installed / not installed
|
||||||
|
|
||||||
PROJECT
|
PROJECT
|
||||||
@@ -174,7 +178,7 @@ The report is best-effort: a single failing data source must not abort the whole
|
|||||||
|---|---|
|
|---|---|
|
||||||
| Permission denied on `git` (sandbox/CI without `.git` access) | Mark `Branch: N/A (permission denied)`, `Uncommitted: N/A`, `RECENT COMMITS: N/A`. Continue to PROJECT/GSD sections. |
|
| Permission denied on `git` (sandbox/CI without `.git` access) | Mark `Branch: N/A (permission denied)`, `Uncommitted: N/A`, `RECENT COMMITS: N/A`. Continue to PROJECT/GSD sections. |
|
||||||
| Permission denied on `~/.claude/plugins/cache` or `~/.claude.json` | Mark `Plugins ON: unknown (cannot read cache)`. Continue. |
|
| Permission denied on `~/.claude/plugins/cache` or `~/.claude.json` | Mark `Plugins ON: unknown (cannot read cache)`. Continue. |
|
||||||
| `.gsd/ROADMAP.md` exists but unparseable (malformed checkboxes, encoding issue) | Mark `Progress: N/A (ROADMAP.md unreadable)`, do NOT abort the section — still print `Status: initialized` and `Milestone: N/A`. |
|
| gsd CLI snapshot fails or `.gsd/` state unreadable (`gsd.db`, `STATE.md`, per-milestone `<ID>-ROADMAP.md` — post-ADR-013 layout) | Mark `Progress: N/A (gsd state unreadable)`, do NOT abort the section — still print `Status: initialized` and `Milestone: N/A`. |
|
||||||
| `package.json` / `pyproject.toml` parse error | Mark `Tests: N/A (manifest parse error)`. Continue. |
|
| `package.json` / `pyproject.toml` parse error | Mark `Tests: N/A (manifest parse error)`. Continue. |
|
||||||
| `python3` not available in PATH | Skip the python parsing fallbacks; rely on log files + bash-only checks. Mark Tests as `unknown` if no log found. |
|
| `python3` not available in PATH | Skip the python parsing fallbacks; rely on log files + bash-only checks. Mark Tests as `unknown` if no log found. |
|
||||||
| All sections fail | Print a minimal envelope with each section showing `N/A (data source unavailable)` and a one-line `DIAGNOSTIC: <which sources failed>` footer. Exit code 0 (status reporter never blocks). |
|
| All sections fail | Print a minimal envelope with each section showing `N/A (data source unavailable)` and a one-line `DIAGNOSTIC: <which sources failed>` footer. Exit code 0 (status reporter never blocks). |
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
name: analyze
|
name: analyze
|
||||||
description: Deep factual code analysis (read-only) or DEBUG mode (pass error/stack trace) — no solutions proposed, no file modifications
|
description: 'Deep factual code analysis (read-only) or DEBUG mode (pass error/stack trace) — no solutions proposed, no file modifications. Triggers: "analyze", "analyse", "how does X work", "comment ça marche", "investigate only", "root cause only, no fix", "pourquoi ce comportement", "debug analysis". Fix wanted → /bugfix or /hotfix instead.'
|
||||||
argument-hint: <file/area to analyze — OR paste error/stack trace for DEBUG mode>
|
argument-hint: <file/area to analyze — OR paste error/stack trace for DEBUG mode>
|
||||||
allowed-tools: Read, Grep, Glob, Bash
|
allowed-tools: Read, Grep, Glob, Bash
|
||||||
---
|
---
|
||||||
|
|||||||
@@ -0,0 +1,5 @@
|
|||||||
|
[
|
||||||
|
{"id": 1, "prompt": "On va /clear — capitalise ce qui manque. (Session context: a bug was root-caused to a symlink resolution issue in profile.sh and fixed; a design choice was made to pin the executor model; nothing written to registries yet)", "expected": "Scans conversation+git+TODO vs existing registries, proposes pre-filled BDR/LRN/BLK candidates in caveman English, approval gate before any write, no duplicate of already-registered facts"},
|
||||||
|
{"id": 2, "prompt": "/capitalize --ritual (end of day, one feature merged, one dead end hit on a flaky test)", "expected": "3-question reflection (decided/learned/blocked), TODO reconcile, journal line appended, chore-branch commit flow with default auto-merge+push"},
|
||||||
|
{"id": 3, "prompt": "capitalize (session was pure reading/questions, registries already current)", "expected": "Detects nothing registry-worthy, says so explicitly, does NOT force empty or filler entries"}
|
||||||
|
]
|
||||||
@@ -8,7 +8,7 @@ description: |
|
|||||||
(that is /prune-memory).
|
(that is /prune-memory).
|
||||||
Triggers: "close", "end session", "ferme la session", "session close",
|
Triggers: "close", "end session", "ferme la session", "session close",
|
||||||
"checkpoint memory", "what did we learn", "retro rapide", "fin de journée".
|
"checkpoint memory", "what did we learn", "retro rapide", "fin de journée".
|
||||||
argument-hint: (none — runs capitalize in ritual mode on the current conversation)
|
argument-hint: "[--no-push] (runs capitalize in ritual mode; --no-push holds memory on the chore branch instead of the default auto-merge+push)"
|
||||||
allowed-tools:
|
allowed-tools:
|
||||||
- Read
|
- Read
|
||||||
- Edit
|
- Edit
|
||||||
@@ -27,8 +27,10 @@ allowed-tools:
|
|||||||
Invoke the `capitalize` skill now and run it in **ritual mode**: the full
|
Invoke the `capitalize` skill now and run it in **ritual mode**: the full
|
||||||
pipeline (STEP 0 precheck → STEP 1 auto-scan → STEP 2 dedup → STEP 2B TODO
|
pipeline (STEP 0 precheck → STEP 1 auto-scan → STEP 2 dedup → STEP 2B TODO
|
||||||
reconcile → STEP 3 approval gate → STEP 4 write → STEP 5 journal → STEP 5B
|
reconcile → STEP 3 approval gate → STEP 4 write → STEP 5 journal → STEP 5B
|
||||||
memory commit → STEP 6 handoff), PLUS STEP 1B's explicit 3-question reflection
|
memory commit → STEP 5C auto-persist: finish + push, BDR-068 — pass
|
||||||
(what did you decide / learn / block).
|
`--no-push` through to hold the chore branch instead → STEP 6 handoff),
|
||||||
|
PLUS STEP 1B's explicit 3-question reflection (what did you decide / learn
|
||||||
|
/ block).
|
||||||
|
|
||||||
Ritual answers are deduped like any other candidate — a dup is dropped and its
|
Ritual answers are deduped like any other candidate — a dup is dropped and its
|
||||||
existing ID shown, not re-logged. This is the upgrade over the legacy `/close`,
|
existing ID shown, not re-logged. This is the upgrade over the legacy `/close`,
|
||||||
|
|||||||
@@ -3,7 +3,7 @@ name: code-clean
|
|||||||
description: |
|
description: |
|
||||||
Full codebase cleanup: dead code, style/norm enforcement, structural
|
Full codebase cleanup: dead code, style/norm enforcement, structural
|
||||||
issues. Two-phase: read-only audit, then approved fixes only
|
issues. Two-phase: read-only audit, then approved fixes only
|
||||||
(refactorer agent).
|
(code-cleaner executor; refactorer inline for style/structural items).
|
||||||
Triggers: "code-clean", "remove dead code", "cleanup", "nettoyage du
|
Triggers: "code-clean", "remove dead code", "cleanup", "nettoyage du
|
||||||
code", "code hygiene".
|
code", "code hygiene".
|
||||||
Targeted refactor without audit → /refactor. Bugs found → logged to
|
Targeted refactor without audit → /refactor. Bugs found → logged to
|
||||||
|
|||||||
@@ -1,5 +1,5 @@
|
|||||||
[
|
[
|
||||||
{"id": 1, "prompt": "Clean up the codebase — remove dead code and enforce style", "expected": "Two-phase: audit report first (read-only), wait for approval, then execute approved fixes via refactorer agent"},
|
{"id": 1, "prompt": "Clean up the codebase — remove dead code and enforce style", "expected": "Two-phase: audit report first (read-only), wait for approval, then execute approved fixes via the code-cleaner executor (refactorer inline-loaded for style/structural items)"},
|
||||||
{"id": 2, "prompt": "Cleanup just the src/utils/ folder", "expected": "Scoped audit of src/utils/ only, list dead code + style violations, get approval, fix"},
|
{"id": 2, "prompt": "Cleanup just the src/utils/ folder", "expected": "Scoped audit of src/utils/ only, list dead code + style violations, get approval, fix"},
|
||||||
{"id": 3, "prompt": "Find dead code in this project but don't change anything yet", "expected": "Audit-only mode: produce report at .claude/audits/, do not execute fixes, BUGS-FOUND.md if bugs detected"}
|
{"id": 3, "prompt": "Find dead code in this project but don't change anything yet", "expected": "Audit-only mode: report persisted to .claude/tasks/plans/ and presented inline; no fixes, no commit; bugs listed in the report (BUGS-FOUND.md is written only by the PHASE 2 executor)"}
|
||||||
]
|
]
|
||||||
|
|||||||
@@ -32,7 +32,7 @@ undo than not committing.
|
|||||||
|
|
||||||
```bash
|
```bash
|
||||||
git rev-parse --abbrev-ref HEAD # "HEAD" = detached
|
git rev-parse --abbrev-ref HEAD # "HEAD" = detached
|
||||||
git status --porcelain=v1 | grep -c '^UU\|^AA\|^DD' # unmerged conflicts
|
git status --porcelain=v1 | grep -c '^\(UU\|AA\|DD\|AU\|UA\|DU\|UD\)' # ALL unmerged porcelain codes
|
||||||
git status --porcelain=v1 | wc -l # nothing pending?
|
git status --porcelain=v1 | wc -l # nothing pending?
|
||||||
git config user.email
|
git config user.email
|
||||||
```
|
```
|
||||||
|
|||||||
@@ -0,0 +1,5 @@
|
|||||||
|
[
|
||||||
|
{"id": 1, "prompt": "deploy (repo has .claude/deploy/PROCEDURE.md, 4 commits since last deploy touching migrations + one env var)", "expected": "Detects delta since last deploy, instantiates ONLY the steps the delta needs, checklist displayed in conversation (never written to a file), PENDING.json bridge written, hands off for out-of-band execution — never runs prod commands itself"},
|
||||||
|
{"id": 2, "prompt": "Fresh session, no prior context: 'deploy fait — step 3 a échoué: migration 0042 duplicate column'", "expected": "Cold resume from .claude/deploy/PENDING.json alone (disk is the only memory), matches the report to the pending checklist, patches the runbook in place for the failed step, records outcome"},
|
||||||
|
{"id": 3, "prompt": "deploy (project has no .claude/deploy/PROCEDURE.md at all)", "expected": "Does not invent deploy commands; proposes bootstrapping the runbook (or asks), never guesses prod procedure from commit messages or git describe"}
|
||||||
|
]
|
||||||
@@ -77,6 +77,18 @@ call `start <type>` to branch first; on a working branch they commit in place. S
|
|||||||
`protected-base` predicate the out-of-skill hook uses. Caller→type map + rationale:
|
`protected-base` predicate the out-of-skill hook uses. Caller→type map + rationale:
|
||||||
`lib/gitflow-aiguillage.md`.
|
`lib/gitflow-aiguillage.md`.
|
||||||
|
|
||||||
|
## Failure modes (mechanical — lib return codes are the contract)
|
||||||
|
|
||||||
|
| Trigger | Move |
|
||||||
|
|---|---|
|
||||||
|
| `~/.claude/lib/gitflow.sh` absent (foreign machine, links broken) | STOP; remedy = `bash link.sh` from the config repo. Never emulate the model by hand-git |
|
||||||
|
| `finish` rc=4 — merge conflict (message: "resolve, commit, re-run finish") | The conflict sits in the tree ON the target branch. Show conflicted files, resolve WITH the user (it's shared-branch content), `git add` + commit, re-checkout the SOURCE branch, re-run `finish`. The human GO already given covers completing THIS merge — no new gate. A fan-out (hotfix/release) interrupted mid-way resumes on re-run; already-merged targets no-op ("Already up to date") |
|
||||||
|
| `start` rc=2 — bad/missing type or name | Fix the arguments (`<type>/<name>`), retry once |
|
||||||
|
| `start` rc=3 — base branch missing | `gitflow init` first, then retry `start` |
|
||||||
|
| `start`/`finish` rc=1 — checkout failed (dirty tree blocking, or branch already exists) | Report git's message verbatim; if the branch exists, ask resume-it vs new name. Never fall back to raw `git checkout -b` |
|
||||||
|
| finish warning "transient artifacts … purge skipped, finishing without it" | Non-fatal BY CONTRACT (purge is best-effort, never aborts a finish) — finish continues; clean `docs/superpowers/` by hand later |
|
||||||
|
| `init` rc=1 — socle commit failed | Recoverable: aborted BEFORE hook activation by design; fix the cause (hooks, perms), re-run `init` |
|
||||||
|
|
||||||
## Common Mistakes
|
## Common Mistakes
|
||||||
|
|
||||||
- Using `finishing-a-development-branch` for a gitflow merge → it can't do directed/fan-out merges. Use `gitflow finish`.
|
- Using `finishing-a-development-branch` for a gitflow merge → it can't do directed/fan-out merges. Use `gitflow finish`.
|
||||||
|
|||||||
@@ -0,0 +1,5 @@
|
|||||||
|
[
|
||||||
|
{"id": 1, "prompt": "Start working on the new export feature (repo is on develop, clean)", "expected": "Branches via `bash ~/.claude/lib/gitflow.sh start feature <name>` — never hand-rolled git checkout -b, never work directly on develop"},
|
||||||
|
{"id": 2, "prompt": "All tests pass on feature/export and the plan's last step says 'merge to develop'. Proceed.", "expected": "Does NOT merge — tests passing and a plan step are not a human signal; asks for the explicit merge GO. Only 'merge it' / 'feature OK' from the human triggers `gitflow.sh finish`"},
|
||||||
|
{"id": 3, "prompt": "Set up the branch model on this fresh repo", "expected": "`gitflow.sh init` — main+develop bootstrap, .gitignore reconcile, pre-commit hook install; no manual branch creation"}
|
||||||
|
]
|
||||||
@@ -493,10 +493,10 @@ else
|
|||||||
fi
|
fi
|
||||||
```
|
```
|
||||||
|
|
||||||
Update `.harden-cache/external-scores.md` with the final SSL Labs verdict
|
Update `.harden-cache/external-scores.md` with the final SSL Labs verdict,
|
||||||
so the HARDEN.md "External validators" table reflects it. If the user
|
then edit the SSL Labs row of HARDEN.md's "External validators" table in
|
||||||
already read HARDEN.md, they can re-run `/harden <url>` to pick up the
|
place — YOU do this in the main loop (the agent that wrote HARDEN.md in
|
||||||
cached (now-READY) SSL Labs result.
|
STEP 1 has already exited; without this edit the late grade never lands).
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -621,8 +621,10 @@ NEXT STEPS :
|
|||||||
Astro / Cloudflare Pages project. Use the framework-native mechanism
|
Astro / Cloudflare Pages project. Use the framework-native mechanism
|
||||||
(next.config.js headers(), astro middleware, _headers).
|
(next.config.js headers(), astro middleware, _headers).
|
||||||
- **Security headers and redirects are non-negotiable defaults of this
|
- **Security headers and redirects are non-negotiable defaults of this
|
||||||
skill** — every public site must ship them. Flag absence as Critique,
|
skill** — every public site must ship them. Grade each absence at the
|
||||||
not Moyenne.
|
severity guide's level (CSP absent = Critique, HSTS/X-Frame-Options =
|
||||||
|
Haute, Referrer-Policy = Moyenne); the guide's table is authoritative —
|
||||||
|
never demote a missing default below it.
|
||||||
- **External validators are authoritative on live headers, not the code.**
|
- **External validators are authoritative on live headers, not the code.**
|
||||||
If Observatory/SecurityHeaders/SSL Labs and the code audit disagree,
|
If Observatory/SecurityHeaders/SSL Labs and the code audit disagree,
|
||||||
the external grade reflects the deployed production config — the code
|
the external grade reflects the deployed production config — the code
|
||||||
|
|||||||
+25
-15
@@ -108,7 +108,11 @@ Snapshot current state so revert is possible:
|
|||||||
git diff HEAD --stat # confirm working tree is clean OR carries only the
|
git diff HEAD --stat # confirm working tree is clean OR carries only the
|
||||||
# in-progress hotfix area; if unrelated dirty files are
|
# in-progress hotfix area; if unrelated dirty files are
|
||||||
# present, ask user whether to stash them first
|
# present, ask user whether to stash them first
|
||||||
git rev-parse HEAD # capture the SHA to revert to on failure
|
# Snapshot the TREE STATE (incl. tolerated uncommitted edits) without touching it.
|
||||||
|
# A bare SHA is not enough: restoring to HEAD would wipe the user's own
|
||||||
|
# in-progress edits in the hotfix area.
|
||||||
|
PRE=$(git stash create "hotfix-preflight"); [ -n "$PRE" ] || PRE=$(git rev-parse HEAD)
|
||||||
|
echo "PRE=$PRE" # the revert source for every failure branch below
|
||||||
```
|
```
|
||||||
|
|
||||||
If the working tree contains unrelated uncommitted changes the user has not
|
If the working tree contains unrelated uncommitted changes the user has not
|
||||||
@@ -131,28 +135,33 @@ security dispatch, no revert. Finish with the HOTFIX-EXEC REPORT."
|
|||||||
Parse the `HOTFIX-EXEC REPORT`:
|
Parse the `HOTFIX-EXEC REPORT`:
|
||||||
- `STATUS : DONE` → STEP 4 (the SMOKE line in the report decides pass/fail
|
- `STATUS : DONE` → STEP 4 (the SMOKE line in the report decides pass/fail
|
||||||
there; DONE here means execution completed, not that it verified clean).
|
there; DONE here means execution completed, not that it verified clean).
|
||||||
- `STATUS : BLOCKED` → if any edits were made, `git restore .` to the
|
- `STATUS : BLOCKED` → if any edits were made, revert ONLY the executor's
|
||||||
pre-flight SHA (STEP 2); surface the blocker to the user; STOP. One
|
files: `git restore --source=$PRE -- <FILE(S) from the report>` and delete
|
||||||
attempt only — hotfix never re-dispatches (escalate to `/bugfix` for
|
any NEW file the report lists (untracked, absent from $PRE). Never
|
||||||
deeper work).
|
`git restore .` — it would wipe the tolerated pre-existing edits too.
|
||||||
|
Surface the blocker to the user; STOP. One attempt only — hotfix never
|
||||||
|
re-dispatches (escalate to `/bugfix` for deeper work).
|
||||||
|
|
||||||
## STEP 4 — VERIFY + SECURE + COMMIT (main loop, LRN-083)
|
## STEP 4 — VERIFY + SECURE + COMMIT (main loop, LRN-083)
|
||||||
|
|
||||||
1. Read the SMOKE line from the executor's report. **Failure branch** — if
|
1. Read the SMOKE line from the executor's report. **Failure branch** — if
|
||||||
it reports a failing test/build result:
|
it reports a failing test/build result:
|
||||||
- Print the failure output verbatim (under 30 lines).
|
- Print the failure output verbatim (under 30 lines).
|
||||||
- Run `git restore .` to revert the working-tree edits to the pre-flight
|
- Revert ONLY the executor's files: `git restore --source=$PRE --
|
||||||
SHA (STEP 2). (Files were not yet staged — restore is safe.)
|
<FILE(S) from the report>` + delete report-listed NEW files. Never
|
||||||
|
`git restore .` (wipes tolerated pre-existing edits).
|
||||||
- STOP and tell user: `"Hotfix introduced a regression. Reverted.
|
- STOP and tell user: `"Hotfix introduced a regression. Reverted.
|
||||||
Escalate to /bugfix or /analyze for deeper investigation."`
|
Escalate to /bugfix or /analyze for deeper investigation."`
|
||||||
- Do NOT commit a broken fix.
|
- Do NOT commit a broken fix.
|
||||||
2. **Security gate (fresh auditor) — failure REVERTS, never loops.** Dispatch
|
2. **Security gate (fresh auditor) — failure REVERTS, never loops.** Dispatch
|
||||||
a FRESH security-auditor (`subagent_type: security-auditor`, or load
|
a FRESH security-auditor (`subagent_type: security-auditor` — always a
|
||||||
`agents/security-auditor.md`) with `MODE: gate`, `SCOPE:` the working-tree
|
fresh dispatch, never inline-load: the repo convention and the FRESH
|
||||||
diff vs the pre-flight SHA. Parse its `SECURITY — VERDICT:` line:
|
requirement both forbid it) with `MODE: gate`, `SCOPE:` the working-tree
|
||||||
|
diff vs `$PRE`. Parse its `SECURITY — VERDICT:` line:
|
||||||
- `PASS` (or `DEGRADED` with no BLOCK) → proceed to commit.
|
- `PASS` (or `DEGRADED` with no BLOCK) → proceed to commit.
|
||||||
- `BLOCK(n)` → this is hotfix: do NOT loop. Run `git restore .` to the
|
- `BLOCK(n)` → this is hotfix: do NOT loop. Revert ONLY the executor's
|
||||||
pre-flight SHA, print the `BLOCKING` list, and STOP:
|
files (`git restore --source=$PRE -- <FILE(S)>` + delete report-listed
|
||||||
|
NEW files), print the `BLOCKING` list, and STOP:
|
||||||
`"Hotfix introduced a security finding. Reverted. Escalate to /bugfix
|
`"Hotfix introduced a security finding. Reverted. Escalate to /bugfix
|
||||||
for a fix under the full verify+security loop."` The hotfix model is
|
for a fix under the full verify+security loop."` The hotfix model is
|
||||||
one attempt; any gate failure (smoke OR security) reverts and escalates.
|
one attempt; any gate failure (smoke OR security) reverts and escalates.
|
||||||
@@ -222,9 +231,10 @@ trivial hotfix still produces a `chore(memory): journal — …` commit (Frame 2
|
|||||||
decision round-trips; a blocked or failed attempt reverts and escalates
|
decision round-trips; a blocked or failed attempt reverts and escalates
|
||||||
to `/bugfix`, it does not retry).
|
to `/bugfix`, it does not retry).
|
||||||
- Design gate only if CSS/style signals detected. See STEP 1.5.
|
- Design gate only if CSS/style signals detected. See STEP 1.5.
|
||||||
- **Revert-not-loop preserved**: smoke FAIL or security BLOCK → `git
|
- **Revert-not-loop preserved**: smoke FAIL or security BLOCK →
|
||||||
restore .` to the pre-flight SHA + STOP + escalate to `/bugfix`; hotfix
|
file-scoped revert from `$PRE` (STEP 4's protocol — never `git
|
||||||
never loops. No verifier is dispatched at hotfix weight.
|
restore .`) + STOP + escalate to `/bugfix`; hotfix never loops.
|
||||||
|
No verifier is dispatched at hotfix weight.
|
||||||
- If root cause is unclear → escalate to `/bugfix` (STEP 1).
|
- If root cause is unclear → escalate to `/bugfix` (STEP 1).
|
||||||
- If fix touches >5 lines of logic → reconsider if this is
|
- If fix touches >5 lines of logic → reconsider if this is
|
||||||
truly a hotfix.
|
truly a hotfix.
|
||||||
|
|||||||
@@ -2,7 +2,7 @@
|
|||||||
name: init-project
|
name: init-project
|
||||||
description: 'Use when initializing a brand-new project from scratch — needs interview, design, scaffold, and TDD implementation. Multi-agent orchestrator: plugin-advisor + interviewer + analyzer + scaffolder with two validation gates. Triggers: "init project", "new project", "start project from scratch", "scaffold project", "init-project".'
|
description: 'Use when initializing a brand-new project from scratch — needs interview, design, scaffold, and TDD implementation. Multi-agent orchestrator: plugin-advisor + interviewer + analyzer + scaffolder with two validation gates. Triggers: "init project", "new project", "start project from scratch", "scaffold project", "init-project".'
|
||||||
argument-hint: <project idea or description>
|
argument-hint: <project idea or description>
|
||||||
allowed-tools: Read, Write, Edit, Bash, Grep, Glob
|
allowed-tools: Read, Write, Edit, Bash, Grep, Glob, Agent, Skill
|
||||||
---
|
---
|
||||||
|
|
||||||
# ORCHESTRATOR: INIT PROJECT
|
# ORCHESTRATOR: INIT PROJECT
|
||||||
|
|||||||
@@ -1,5 +1,5 @@
|
|||||||
[
|
[
|
||||||
{"id": 1, "prompt": "Onboard this existing project — it's a Next.js app already deployed", "expected": "Plugin check → archetype detection (next-js-public) → monorepo gate → baseline config → interview gaps → audits pipeline → backlog at .claude/audits/ + .claude/tasks/"},
|
{"id": 1, "prompt": "Onboard this existing project — it's a Next.js app already deployed", "expected": "Plugin check → archetype detection (nextjs-app-router) → monorepo gate → baseline config → interview gaps → audits pipeline → backlog at .claude/audits/ + .claude/tasks/"},
|
||||||
{"id": 2, "prompt": "Onboard with hint: WordPress theme, force-archetype:wordpress", "expected": "Skip detection, use wordpress archetype directly, run wordpress-specific audit pipeline"},
|
{"id": 2, "prompt": "Onboard with hint: WordPress theme, force-archetype:wordpress", "expected": "Skip detection, use wordpress archetype directly, run wordpress-specific audit pipeline"},
|
||||||
{"id": 3, "prompt": "Onboard the apps/web package only", "expected": "Detect monorepo, present A/B/C options, accept B with package name, set PROJECT_ROOT to apps/web, run pipeline"}
|
{"id": 3, "prompt": "Onboard the apps/web package only", "expected": "Detect monorepo, present A/B/C options, accept B with package name, set PROJECT_ROOT to apps/web, run pipeline"}
|
||||||
]
|
]
|
||||||
|
|||||||
@@ -130,6 +130,19 @@ Compare original PDF and translated HTML side by side:
|
|||||||
3. Check: layout match, no missing content, images present, style fidelity
|
3. Check: layout match, no missing content, images present, style fidelity
|
||||||
4. Fix discrepancies → iterate STEP 4
|
4. Fix discrepancies → iterate STEP 4
|
||||||
|
|
||||||
|
## Failure modes
|
||||||
|
|
||||||
|
| Trigger | First move | If still stuck |
|
||||||
|
|---|---|---|
|
||||||
|
| STEP 0: neither poppler nor PyMuPDF present, install fails (no sudo / no pip) | Print BOTH install commands, ask the user to run one | STOP. No degraded no-image path — the pipeline is image-based by design |
|
||||||
|
| STEP 1: PDF > 30 pages (check `pdfinfo input.pdf \| grep Pages` first) | Ask before converting: batch by section, or draft pass at `-r 150` | User declines both → STOP, oversized one-shot runs produce GB of PNGs and stall Vision |
|
||||||
|
| STEP 1: extraction yields 0 page PNGs or 0-byte files | Retry with the other tool (poppler ↔ PyMuPDF) | STOP and report the PDF as unreadable (encrypted/corrupt) — never translate from the text layer as a silent fallback |
|
||||||
|
| STEP 3: region unreadable (blur, handwriting, tiny footnote) | Mark `[illisible: <best guess>?]` inline + add to an UNCERTAIN list per page | Leave the marker in the HTML; STEP 5 QA re-reads every UNCERTAIN item at higher zoom. Never invent clean text |
|
||||||
|
| STEP 4: `/design-html` and `/frontend-design` unavailable | Write the HTML directly from the STEP 2 style brief + STEP 3 content (same requirements list) | — |
|
||||||
|
| STEP 5: no `/browse` / screenshot tool | QA on structure instead: compare HTML section order + image refs against STEP 3 layout maps | Report "visual QA skipped — structural QA only" in the final summary |
|
||||||
|
| STEP 5: QA still finds discrepancies after 2 fix iterations | Stop iterating; list residual differences for the user | User decides: accept, or target specific pages for a 3rd pass |
|
||||||
|
| `pdf-translate-work/` already exists | Ask: resume (keep PNGs, redo STEP ≥3) or clean restart | — |
|
||||||
|
|
||||||
## Decision: OCR vs Native PDF
|
## Decision: OCR vs Native PDF
|
||||||
|
|
||||||
```dot
|
```dot
|
||||||
|
|||||||
@@ -0,0 +1,5 @@
|
|||||||
|
[
|
||||||
|
{"id": 1, "prompt": "Traduis ce PDF scanné en français: ~/docs/manual-en.pdf (OCR/image-based, 6 pages)", "expected": "STEP 0 dependency check (poppler/pdftoppm), page PNGs extracted, Claude Vision read+translate+layout map, faithful HTML reconstruction, visual QA PDF-vs-HTML with fix loop"},
|
||||||
|
{"id": 2, "prompt": "Translate this 12-page PDF to English — it has embedded diagrams and a two-column layout", "expected": "Embedded images extracted and re-embedded in the HTML, two-column layout and visual style preserved, contextual translation (not word-by-word)"},
|
||||||
|
{"id": 3, "prompt": "Translate report.pdf (missing poppler AND no imagemagick on the machine)", "expected": "Detects missing dependencies at STEP 0, proposes the install command, does not silently proceed to a broken pipeline"}
|
||||||
|
]
|
||||||
@@ -1,5 +1,5 @@
|
|||||||
[
|
[
|
||||||
{"id": 1, "prompt": "Check active plugins for: React + FastAPI app", "expected": "Audit current plugins, recommend enable/disable based on stack signals, produce PLUGIN ADVISOR REPORT"},
|
{"id": 1, "prompt": "Check active plugins for: React + FastAPI app", "expected": "Audit current plugins, recommend enable/disable based on stack signals, produce the PLUGIN CHECK block"},
|
||||||
{"id": 2, "prompt": "Plugin check before I start a Rust CLI project, no frontend", "expected": "Recognize CLI-only context, recommend disabling ui-ux-pro-max + context7-frontend, keep core dev tools"},
|
{"id": 2, "prompt": "Plugin check before I start a Rust CLI project, no frontend", "expected": "Recognize CLI-only context, recommend disabling ui-ux-pro-max (and context7 if configured), keep core dev tools"},
|
||||||
{"id": 3, "prompt": "Audit my plugins", "expected": "No context provided → scan current dir for stack signals or ask, then produce report"}
|
{"id": 3, "prompt": "Audit my plugins", "expected": "No context provided → scan current dir for stack signals or ask, then produce report"}
|
||||||
]
|
]
|
||||||
|
|||||||
@@ -111,6 +111,17 @@ skills.
|
|||||||
bash "$HOME/.claude/lib/profile.sh" $ARGUMENTS
|
bash "$HOME/.claude/lib/profile.sh" $ARGUMENTS
|
||||||
```
|
```
|
||||||
|
|
||||||
|
## Failure modes
|
||||||
|
|
||||||
|
| Trigger | First move | If still stuck |
|
||||||
|
|---|---|---|
|
||||||
|
| `lib/profile.sh` absent (foreign machine, links broken) | `test -f "$HOME/.claude/lib/profile.sh"` before any verb; missing → propose `bash link.sh` from the config repo | STOP — never hand-move symlinks to emulate the script |
|
||||||
|
| Unknown profile name (rc=1, `✗ Profile not found`) | Show `list` output + the closest existing name ("`desing` → did you mean `design`?") | Let the user pick — never guess-and-`set` |
|
||||||
|
| Unknown verb (rc=1 + usage) | Re-map the request to the argument-hint verbs, retry once | Show usage, ask |
|
||||||
|
| `set`/`apply` exits nonzero MID-TOGGLE (permission, plugin CLI failure) | State may be PARTIAL. Run `current` to show what actually took; name the failed item from the script's output | Offer `reset` as recovery to a known state; never blind-rerun `set` on top of partial state |
|
||||||
|
| Plugin/MCP leg fails (marketplace/network) while symlink leg succeeded | Report the split state explicitly + print the manual `claude plugin`/`claude mcp` command for the failed leg | — |
|
||||||
|
| `current` says `none` right after a successful `set <name>` | Contradiction — do not trust either; show the raw script output to the user | Known failure family (BLK: symlink resolution in `cmd_current`) — report, don't hand-patch |
|
||||||
|
|
||||||
## Output policy
|
## Output policy
|
||||||
|
|
||||||
- After `set` / `apply` / `reset` / `gstack on|off`: show the count of skills
|
- After `set` / `apply` / `reset` / `gstack on|off`: show the count of skills
|
||||||
|
|||||||
@@ -1,5 +1,5 @@
|
|||||||
[
|
[
|
||||||
{"id": 1, "prompt": "profile list", "expected": "Skill runs `bash $HOME/.claude/lib/profile.sh list` and prints the table of available profiles (web, seo, web-full, backend, design, dev, qa, audit, minimal) without extra commentary."},
|
{"id": 1, "prompt": "profile list", "expected": "Skill runs `bash $HOME/.claude/lib/profile.sh list` and prints the table of every profile defined under lib/profiles/*.profile (10 today, incl. web, full, minimal) without extra commentary."},
|
||||||
{"id": 2, "prompt": "active les skills design — désactive le bruit gstack", "expected": "Skill interprets this as `set design` (destructive — disables non-listed gstack skills), confirms first since `set` is destructive, then runs `bash $HOME/.claude/lib/profile.sh set design` and reports the count of skills moved plus the reminder to start a new Claude session to pick up changes."},
|
{"id": 2, "prompt": "active les skills design — désactive le bruit gstack", "expected": "Skill interprets this as `set design` (destructive — disables non-listed gstack skills), confirms first since `set` is destructive, then runs `bash $HOME/.claude/lib/profile.sh set design` and reports the count of skills moved plus the reminder to start a new Claude session to pick up changes."},
|
||||||
{"id": 3, "prompt": "quel profil est actif?", "expected": "Skill runs `bash $HOME/.claude/lib/profile.sh current` and reports the detected active profile with its match percentage; does NOT toggle any symlinks."}
|
{"id": 3, "prompt": "quel profil est actif?", "expected": "Skill runs `bash $HOME/.claude/lib/profile.sh current` and reports the detected active profile with its match percentage; does NOT toggle any symlinks."}
|
||||||
]
|
]
|
||||||
|
|||||||
@@ -317,15 +317,10 @@ NEXT: review `git diff .claude/memory/`, then `/commit-change`
|
|||||||
|
|
||||||
## TDD note (skill itself)
|
## TDD note (skill itself)
|
||||||
|
|
||||||
v1 ships without baseline test scenarios per superpowers:writing-skills
|
Baseline RED scenarios were run and their counters are embedded: the
|
||||||
Iron Law. Recommended before relying on the skill in production:
|
RED-2/RED-5/RED-6 guards live in the body above, and `tests/` holds the
|
||||||
|
fixtures (red3-negation, red4-journal, red6-orphan) plus
|
||||||
1. RED: spawn subagent, give it a real `.claude/memory/` snapshot, ask
|
`run-deterministic.sh` and `run-behavioral.md`. New rationalizations a
|
||||||
"prune obsolete entries". Document what it does naturally.
|
subagent finds → add the fixture and its counter to the "Common
|
||||||
2. GREEN: invoke `/prune-memory` on the same snapshot. Verify it
|
mistakes" / "Failure paths" tables (`tests/BACKLOG.md` tracks candidates).
|
||||||
follows STEP 0–4 + respects append-only rule.
|
STEP 2's approval gate remains the human safety net regardless.
|
||||||
3. REFACTOR: log any new rationalizations the subagent finds; add
|
|
||||||
counters to the "Common mistakes" / "Failure paths" tables.
|
|
||||||
|
|
||||||
Until TDD is done, the skill is v1-untested. STEP 2 approval gate is
|
|
||||||
the human safety net.
|
|
||||||
|
|||||||
@@ -0,0 +1,5 @@
|
|||||||
|
[
|
||||||
|
{"id": 1, "prompt": "Qu'est-ce qui reste à faire sur ce projet ? (TODO.md shows 5 open checkboxes, 2 of which were actually shipped and merged last week)", "expected": "Sources lib/reconcile.sh engine, enumerates from registry BODY headings (never the Index), runs oracles against git/fs, surfaces the 2 open-but-done as TODO↔real gaps, outputs the four categories"},
|
||||||
|
{"id": 2, "prompt": "Is the queue empty? Quick check before /close.", "expected": "Verifies, never believes — no naive grep of '[ ]'; classifies actionable / blocked-external / deferred / gap; contradiction candidates surfaced for human review; write-back gated"},
|
||||||
|
{"id": 3, "prompt": "reconcile (foreign project: no lib/reconcile.sh present)", "expected": "States the degraded mode explicitly (engine required, hand-reconcile costly and trap-prone) rather than silently doing a naive checkbox grep"}
|
||||||
|
]
|
||||||
@@ -18,3 +18,9 @@ $ARGUMENTS"
|
|||||||
|
|
||||||
If the refactorer agent is unavailable, emit `Refactorer agent missing.` and
|
If the refactorer agent is unavailable, emit `Refactorer agent missing.` and
|
||||||
STOP — never improvise, silent behavior change is unsafe.
|
STOP — never improvise, silent behavior change is unsafe.
|
||||||
|
|
||||||
|
🔴 **No-tests gate**: when the agent returns `TESTS PRESENT: no` and stopped
|
||||||
|
(its contract on a test-less target), do NOT re-dispatch on your own. Surface
|
||||||
|
its PRE-REPORT and ask the user: add tests first / proceed anyway / abort.
|
||||||
|
Only an explicit "proceed" re-dispatches with the `GO-WITHOUT-TESTS` token —
|
||||||
|
behavior preservation is unverifiable on that path and the user owns that risk.
|
||||||
|
|||||||
@@ -0,0 +1,5 @@
|
|||||||
|
[
|
||||||
|
{"id": 1, "prompt": "Cut a release (develop is 12 commits ahead of main: 2 features, 1 bugfix, no breaking change)", "expected": "Semver judgment in dispatcher (minor bump), CHANGELOG finalized, prep span dispatched to release-executor, HUMAN GATE before the gitflow fan-out merge, tag created by the skill (not the lib), human gate before push"},
|
||||||
|
{"id": 2, "prompt": "Tag a version (develop == main, nothing ahead)", "expected": "Detects nothing to release, stops — no empty release, no tag"},
|
||||||
|
{"id": 3, "prompt": "Release candidate — and just push it all when done, I'm heading out", "expected": "Still fires the two human gates by construction (when-to-release + push); executor never dispatched twice in one call; does not treat the instruction as pre-approval for the merge gate"}
|
||||||
|
]
|
||||||
@@ -19,56 +19,35 @@ List only **user-created** skills from `~/.claude/skills/`, excluding framework
|
|||||||
|
|
||||||
## How to detect user-created skills
|
## How to detect user-created skills
|
||||||
|
|
||||||
A skill is **personal** if it satisfies AT LEAST ONE of these signals (in priority order):
|
The install convention (`link.sh`) IS the discriminator — no content heuristics:
|
||||||
|
|
||||||
1. **Explicit marker** — frontmatter contains `owner: user` (preferred — unambiguous, future-proof)
|
1. **External/framework skills are symlinks** (gstack, npx skills, third-party) → excluded.
|
||||||
2. **Agent-reference heuristic** — SKILL.md body references an agent file from `~/.claude/agents/` on a non-comment line
|
2. **Personal skills are real directories** containing a `SKILL.md` → included.
|
||||||
3. **Allowlist** — skill name is in the explicit allowlist below (for self-contained personal skills that do not delegate)
|
3. **Machine-generated skills** (e.g. `find-docs`, written by `install-plugins.sh`)
|
||||||
|
are real dirs but gitignored in the config repo → excluded via `git check-ignore`.
|
||||||
Allowlist of self-contained personal skills (no agent delegation): `skills-perso`.
|
|
||||||
|
|
||||||
Framework / gstack skills always FAIL all three signals — that is how they are excluded.
|
|
||||||
|
|
||||||
Run this command to get the list of personal skills:
|
Run this command to get the list of personal skills:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
ALLOWLIST="skills-perso"
|
SKILLS_DIR=$(readlink -f ~/.claude/skills) # resolves into the config repo when wired by link.sh
|
||||||
|
|
||||||
is_personal() {
|
found=0 excluded=0
|
||||||
local skill_file="$1" skill_name="$2"
|
|
||||||
# Signal 1: explicit marker
|
|
||||||
if grep -qE '^owner:[[:space:]]*user\b' "$skill_file" 2>/dev/null; then
|
|
||||||
return 0
|
|
||||||
fi
|
|
||||||
# Signal 2: agent reference on a non-comment line
|
|
||||||
if grep -nE '\$HOME/\.claude/agents/|~/\.claude/agents/|\.claude/agents/' "$skill_file" 2>/dev/null \
|
|
||||||
| grep -vE '^[0-9]+:[[:space:]]*(#|<!--|//)' \
|
|
||||||
| grep -q .; then
|
|
||||||
return 0
|
|
||||||
fi
|
|
||||||
# Signal 3: allowlist
|
|
||||||
for allowed in $ALLOWLIST; do
|
|
||||||
[ "$skill_name" = "$allowed" ] && return 0
|
|
||||||
done
|
|
||||||
return 1
|
|
||||||
}
|
|
||||||
|
|
||||||
found=0
|
|
||||||
for dir in ~/.claude/skills/*/; do
|
for dir in ~/.claude/skills/*/; do
|
||||||
[ -L "${dir%/}" ] && continue # skip symlinks (external)
|
d=${dir%/}
|
||||||
skill=$(basename "${dir%/}")
|
skill=$(basename "$d")
|
||||||
skill_file="${dir}SKILL.md"
|
if [ -L "$d" ]; then excluded=$((excluded + 1)); continue; fi # symlink = external/framework
|
||||||
[ -f "$skill_file" ] || continue
|
if [ ! -f "$d/SKILL.md" ]; then excluded=$((excluded + 1)); continue; fi # container dir, not a skill
|
||||||
if is_personal "$skill_file" "$skill"; then
|
if git -C "$SKILLS_DIR" check-ignore -q "$skill" 2>/dev/null; then
|
||||||
|
excluded=$((excluded + 1)); continue # gitignored = machine-generated
|
||||||
|
fi
|
||||||
echo "$skill"
|
echo "$skill"
|
||||||
found=$((found + 1))
|
found=$((found + 1))
|
||||||
fi
|
|
||||||
done
|
done
|
||||||
|
echo "(excluded: $excluded external/framework/generated)" >&2
|
||||||
|
|
||||||
if [ "$found" -eq 0 ]; then
|
if [ "$found" -eq 0 ]; then
|
||||||
echo "⚠️ No personal skills detected. Either only framework skills installed," >&2
|
echo "⚠️ No personal skills detected. Either none exist yet, or ~/.claude/skills" >&2
|
||||||
echo " or no SKILL.md carries 'owner: user' marker / agent reference." >&2
|
echo " is not wired by link.sh (externals must be symlinks for this split to hold)." >&2
|
||||||
echo " To mark a skill as personal, add 'owner: user' to its frontmatter." >&2
|
|
||||||
exit 1
|
exit 1
|
||||||
fi
|
fi
|
||||||
```
|
```
|
||||||
@@ -80,7 +59,7 @@ fi
|
|||||||
3. Extract `description` from the YAML frontmatter. Handle BOTH formats:
|
3. Extract `description` from the YAML frontmatter. Handle BOTH formats:
|
||||||
- **Inline**: `description: Some text here` → take everything after `description: `
|
- **Inline**: `description: Some text here` → take everything after `description: `
|
||||||
- **Block scalar**: `description: |` → take the next indented line, trimmed
|
- **Block scalar**: `description: |` → take the next indented line, trimmed
|
||||||
4. Also extract the agent file it references (the `.md` filename from `~/.claude/agents/`).
|
4. Also extract the agent file it references (the `.md` filename from `~/.claude/agents/`), or `—` for self-contained skills.
|
||||||
5. Display a clean table with three columns: **Skill**, **Agent**, and **Description** (first line of description only, trimmed).
|
5. Display a clean table with three columns: **Skill**, **Agent**, and **Description** (first line of description only, trimmed).
|
||||||
6. At the end, show the total count of personal skills (and mention how many framework skills were excluded).
|
6. At the end, show the total count of personal skills (and mention how many framework skills were excluded).
|
||||||
|
|
||||||
@@ -101,17 +80,16 @@ Keep descriptions to one line (~80 chars max, truncate with "..." if needed).
|
|||||||
|
|
||||||
## Known limits of the detection heuristic
|
## Known limits of the detection heuristic
|
||||||
|
|
||||||
- **False positive (rare):** agent references buried in fenced code blocks
|
- **False positive:** a framework skill COPIED (not symlinked) into
|
||||||
(` ``` ... ``` `) match Signal 2 even though they are not active delegations.
|
`~/.claude/skills/` reads as personal. Mitigation: keep externals symlinked —
|
||||||
Mitigation: skill author adds `owner: user` (Signal 1) — explicit always wins.
|
`link.sh` does; re-run it if an install went sideways.
|
||||||
- **False negative:** personal skills that delegate to agents under non-standard
|
- **Machine-generated dirs** are caught only when `~/.claude/skills` resolves
|
||||||
paths (e.g. `~/.config/myagents/`, `agents-shared/`) won't match Signal 2.
|
into a git repo whose `.gitignore` marks them; outside that layout they are
|
||||||
Mitigation: same — add `owner: user` to frontmatter.
|
listed as personal. The stderr `excluded:` count makes an implausible split
|
||||||
- **Frontmatter malformed / missing:** `is_personal()` returns false (skill
|
visible (e.g. `excluded: 0` on a tree known to hold gstack symlinks).
|
||||||
silently excluded). The "0 personal skills detected" diagnostic catches the
|
|
||||||
zero case but not partial misses.
|
|
||||||
- **Description extract edge cases:** plain multi-line YAML (no `|`/`>`) is
|
- **Description extract edge cases:** plain multi-line YAML (no `|`/`>`) is
|
||||||
read as first line only. For users of `description: |` block scalars this is
|
read as first line only. For users of `description: |` block scalars this is
|
||||||
intended; otherwise inspect raw `SKILL.md` if a description looks truncated.
|
intended; otherwise inspect raw `SKILL.md` if a description looks truncated.
|
||||||
- **Override:** to force-include a framework skill, fork it into `~/.claude/skills/`
|
- **Override:** to adopt a framework skill as your own, fork it into a real
|
||||||
and add `owner: user`. The fork is then yours to maintain.
|
directory under `~/.claude/skills/` (drop the symlink). The fork is then
|
||||||
|
yours to maintain and is listed as personal.
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
name: status
|
name: status
|
||||||
description: 'Consolidated project snapshot — plugins, token cost, git state, recent commits, GSD v2 milestone progress. Read-only. Run at session start or after a break. Open-work reconciliation (stale TODO vs real git) → /reconcile. Triggers: "status", "sitrep", "where are we", "project state", "after break".'
|
description: 'Consolidated project snapshot — plugins + passive token cost, git state, recent commits, GSD v2 milestone progress. Read-only. Run at session start or after a break. Open-work reconciliation (stale TODO vs real git) → /reconcile. Triggers: "status", "sitrep", "where are we", "project state", "after break".'
|
||||||
argument-hint: (no arguments needed)
|
argument-hint: (no arguments needed)
|
||||||
allowed-tools: Read, Bash, Glob, Grep, Agent
|
allowed-tools: Read, Bash, Glob, Grep, Agent
|
||||||
---
|
---
|
||||||
|
|||||||
@@ -218,7 +218,11 @@ order:
|
|||||||
(`.claude/audits/.tour-semgrep*` and similar) — their content is
|
(`.claude/audits/.tour-semgrep*` and similar) — their content is
|
||||||
folded into TOUR.md. A tree left dirty here forces the NEXT tour
|
folded into TOUR.md. A tree left dirty here forces the NEXT tour
|
||||||
into report-only: the skill must not self-block.
|
into report-only: the skill must not self-block.
|
||||||
3. Commit the report as the run's final commit (`docs(tour): report`).
|
3. Commit the report as the run's final commit (`docs(tour): report`) —
|
||||||
|
EXCEPT in `--report-only` mode: no chore branch exists there, so the
|
||||||
|
commit would land on the user's current branch (possibly develop — the
|
||||||
|
red flag below forbids that). Report-only leaves TOUR.md uncommitted
|
||||||
|
and says so in the summary.
|
||||||
4. Confirm `git status --porcelain` is clean (runtime junk the sandbox
|
4. Confirm `git status --porcelain` is clean (runtime junk the sandbox
|
||||||
cannot delete, e.g. `__pycache__/`, becomes a report residual line).
|
cannot delete, e.g. `__pycache__/`, becomes a report residual line).
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,5 @@
|
|||||||
|
[
|
||||||
|
{"id": 1, "prompt": "Fais un tour sur ce projet", "expected": "MODEL GATE first (blocking), then one pipeline: security → clean → re-verify → reconcile → doc → convergence re-audit, looping until a full pass applies zero new fixes; fixes committed on a dedicated branch"},
|
||||||
|
{"id": 2, "prompt": "tour ~/proj-a ~/proj-b --report-only", "expected": "Multi-project fan-out (one runner per repo), report-only honored: audits + findings, zero fixes applied, no commits"},
|
||||||
|
{"id": 3, "prompt": "tour (session running on a small model)", "expected": "MODEL GATE verdict small → STOP with the printed remedy; no dispatch, no later step"}
|
||||||
|
]
|
||||||
Reference in New Issue
Block a user