116 Commits
Author SHA1 Message Date
bastien 33e08990c6 Merge feature/21st-cli-migration into develop 2026-09-22 06:52:22 +02:00
bastien cc93aaba0c chore(memory): BDR-094 LRN-159 BLK-021, journal + TODO 2026-09-22 2026-09-22 03:20:25 +00:00
bastien cc2f246e65 fix(impeccable): global-scope install with agents, pin fallback, output-read failure check
make plugin never installed impeccable. The 3.2.0 pin had rotted upstream
(the CLI fetches its skill dist at install time; that release's zip is
gone), the --scope=project staging moved the skill dir alone and dropped
the 4 impeccable-* subagents, and /impeccable init was never announced.

Step 8d now installs at --scope=global straight through the
~/.claude/{skills,agents} symlinks into the repo (both paths gitignored),
guards on those symlinks existing, keeps a profile-parked copy parked,
falls back to @latest on a pin failure with a bump-the-lock warning, and
prints the per-project init hint. update-all.sh mirrors the shape.

Found while probing: with a copy already installed a rotted pin exits 0
("Could not check for skill updates ... left unchanged"), byte-identical
on disk to an up-to-date rerun, so imp_install reads the installer output
instead of trusting the exit code. Harness 4/4 in a sandbox HOME with the
real installer.

plugins.lock.json: impeccable 3.2.0 -> 4.1.0 (CLI only). link.sh drops
impeccable from EXTERNAL_SKILLS. lib/design-gate.md section 5: suggest-only
/impeccable init check when a frontend project has no PRODUCT.md.
2026-09-22 03:20:25 +00:00
bastien 7c05f75eab feat(21st): replace the magic MCP with the @21st-dev CLI + skill pack
Upstream supersedes `@21st-dev/magic` with `@21st-dev/cli` (bin `21st`):
same endpoint, `21st login` in place of an API key, no MCP process loaded
into every session.

- install-plugins.sh Step 8.7: `npm i -g @21st-dev/cli` (pinned in
  plugins.lock.json), staged `21st skills install`, TTY-only login offer,
  pack disabled by default. update-all.sh 7.4 refreshes both.
- The documented `21st install-skill` cannot be used: the installer refuses
  to follow a symlink on the target path and `~/.claude/skills` is one. The
  install runs under a throwaway HOME and the result moves into
  skills-external/21st-* (gitignored), symlinked on demand.
- toggle-external.sh manages `21st` as a pack (names globbed from
  skills-external/21st-*, parked under plain names). `magic` is gone.
- The 5 design skills join design/web/web-full/full and MANAGED_EXTERNALS;
  21st-registry and 21st-design-sync stay parked. MANAGED_MCPS is now empty
  and profile.sh's dead magic branches are removed.
- Design gate: GATE-BLOCK gains `21st` (required-manual, magic's old slot)
  and `21st-ui-build`; PATH repair extended to the npm global bin.
- settings.json: the 4 mcp__magic__* ask entries go; the outward-facing
  21st verbs land in autoMode.soft_deny, the tier that holds under auto
  mode (LRN-153).
- Docs: README, CLAUDE.global.md, design-gate.md, profile SKILL.md,
  .env.example, .gitleaks.toml, link.sh. BDR-093, LRN-158.

Tests: profile-set-managed 17/17, make test green except 2 pre-existing
gitflow FAILs (gitleaks binary absent on this host), shellcheck clean.
2026-09-22 02:53:31 +00:00
Bastien Chanot 413b35a913 Merge feature/deploy-oneline-tests into develop 2026-09-17 16:43:18 +02:00
Bastien Chanot a1032a1f53 chore(tasks): plan + milestone for the /deploy hand-back change 2026-09-17 16:43:06 +02:00
Bastien Chanot 55d6874aa6 feat(deploy): one physical line per command + post-deploy tests block
Every checklist command is emitted on exactly one line, however long; a
legacy backslash continuation in the runbook is joined at instantiation,
and bootstrap / learn patches write runbook lines the same way. After the
checklist the hand-back carries a Post-deploy tests block derived from the
delta diff: by-hand checks tied to delta files plus Suggestions for gaps.
Cold-resume re-display and re-hand-back regenerate both.

RED/GREEN on a scratch runbook: 4/4 baseline runs reproduced the
continuation verbatim and printed no tests; 4/4 runs on the edited skill
joined it and printed the block in the recipe's shape.
2026-09-17 16:43:06 +02:00
Bastien Chanot b96fab7719 chore(memory): BDR-091 BDR-092 LRN-155..157, journal 2026-09-16 2026-09-17 11:43:39 +02:00
Bastien Chanot 56bd035281 Merge feature/ask-dont-guess into develop 2026-09-17 11:41:39 +02:00
Bastien Chanot cd98bfafe4 chore: purge transient planning artifacts (BDR-065) 2026-09-17 11:41:39 +02:00
Bastien Chanot ddadca6bae Merge feature/automode-docker-node into develop 2026-09-17 11:40:27 +02:00
Bastien Chanot 22ce57f323 docs(changelog): ask-don't-guess doctrine 2026-09-16 22:15:48 +02:00
Bastien Chanot 17370d7e4c feat(executors): NEED-DECISION and BLOCKED carry a CLASS tag 2026-09-16 22:14:13 +02:00
Bastien Chanot 7cc95952bd feat(interviewer): a visible or public choice is asked, never assumed 2026-09-16 22:13:50 +02:00
Bastien Chanot a1357a6ab5 feat(ship-feature,init-project): pass B at the plan and design steps 2026-09-16 22:13:42 +02:00
Bastien Chanot 97591ca737 feat(hotfix): pass B at LOCATE, class-tagged BLOCKED relayed as a question 2026-09-16 22:13:32 +02:00
Bastien Chanot 590482b622 feat(bugfix): pass B at FIX PLAN, NEED-DECISION routed on class 2026-09-16 22:13:16 +02:00
Bastien Chanot 6d9a3497a2 feat(feat): pass B at PLAN, NEED-DECISION routed on class 2026-09-16 22:13:06 +02:00
Bastien Chanot 5f9a9c0f6b feat(rules): ask rather than guess replaces one-question-upfront 2026-09-16 22:12:56 +02:00
Bastien Chanot a2978b6f23 feat(contract): STEP 2 CLARIFY, mid-run channel, how-to-ask 2026-09-16 22:12:13 +02:00
Bastien Chanot 0d52f3a888 docs(plan): ask-don't-guess implementation plan, TODO trace 2026-09-16 22:10:24 +02:00
Bastien Chanot 5eccc3f1c4 feat(automode): docker and node framed by the classifier, ask rules retired 2026-09-16 22:03:23 +02:00
Bastien Chanot 823ce42225 chore(config): default model fable 5.1 2026-09-16 21:58:18 +02:00
Bastien Chanot 9eb69346ce docs(spec): design for the ask-don't-guess clarification doctrine 2026-09-16 20:39:15 +02:00
Bastien Chanot a449f9315f Merge chore/graphify-recovery-doc into develop 2026-09-15 19:57:07 +02:00
Bastien Chanot d7662abc1e chore(graphify): correct the recovery note, capitalize LRN-154
The previous note said a fresh clone gets the skill back from `make
plugin` without naming the command, and I picked the wrong one when the
files actually went missing. There are two, and only one restores the
skill:

  - `graphify install --platform claude` copies SKILL.md, references/
    and .graphify_version. Touches nothing else. This is the recovery
    command, verified: the skill came back at 0.9.61 and the four
    guarded configs were byte-identical afterwards.
  - `graphify claude install` writes the CLAUDE.md section and the
    .claude/settings.json PreToolUse hooks, rewrites both of those
    guarded configs (EVAL-020, reproduced today), and does NOT copy the
    skill.

LRN-154 records why the files vanished in the first place. `git rm
--cached` keeps the working file, but `gitflow finish` checks out the
target branch, where it is still tracked, so git restores it and the
merge then deletes it from disk. .gitignore does not protect it; it only
stops a re-add. The file survives the commit and dies at the merge, which
reads as unrelated.
2026-09-15 19:57:07 +02:00
Bastien Chanot e9fe79e4c2 Merge chore/graphify-gitignore-settings-prune into develop 2026-09-15 19:55:11 +02:00
Bastien Chanot 80ccdafe0e chore(config): untrack the vendored graphify skill, prune the project-local settings override
graphify: `graphify claude install` (install-plugins.sh STEP graphify)
writes SKILL.md, references/ and .graphify_version straight into the repo,
because ~/.claude/skills is a symlink to skills/. Every `pipx upgrade
graphifyy` therefore dirtied the tree and cost a `chore(graphify): sync
vendored skill X -> Y` commit. Now gitignored and untracked; a fresh clone
gets them back from `make plugin`. test-prompts.json is hand-written for
darwin and stays tracked. The accepted trade-off, documented in CLAUDE.md,
is that an upstream release can change the skill's prompt with no diff to
review.

settings.local.json (gitignored, so not in this commit) went from 14.6 KB
to 6.2 KB. It was a near-complete shadow copy of the global settings at a
higher precedence tier, which hid its own drift until the global moved.
Two entries were actively defeating BDR-090, merged an hour earlier:

  - local `deny` still carried rsync / kill -9 / killall / pkill, the four
    rules deliberately moved out of global deny. deny wins across sources,
    so autoMode.soft_deny was a dead letter in this repo.
  - local `allow` carried `sed *`, `cp *` and `python3 -`. An allow rule
    short-circuits the classifier, punching a hole through the same
    soft_deny rules.

deny and ask are dropped whole (102 and 27 of their entries duplicated the
global; ask gates nothing under defaultMode auto). allow went 185 -> 98:
81 duplicates plus six policy conflicts, the three above and
Read(//home/bchanot/**), WebSearch, and a leftover command-injection test
payload that had been allowlisted verbatim. Every non-permissions key was
a verbatim copy of the global, including a hooks block whose only original
entry pointed at hooks/config-protection.sh, a script that exists nowhere.
2026-09-15 19:55:09 +02:00
Bastien Chanot 7a861035b8 Merge feature/automode-config-alignment into develop 2026-09-15 19:44:29 +02:00
Bastien Chanot 6cd26bc3fa chore(memory): BDR-090, LRN-153, journal — autoMode tier rebuild
BDR-090 records why the ask tier was abandoned rather than repopulated,
the three alternatives rejected, and the deliberate caveat that the
guardrail hard_deny bars removing a deny entry but not adding one.

LRN-153 records the two traps the block carries: every autoMode list is
a full replacement without "$defaults", and a user-scope block reaches
every project on the machine.

TODO also logs F1-F3, found but not fixed: .claude/settings.local.json
is a 14.6 KB shadow copy of the global settings at higher precedence,
including a PreToolUse hook whose script does not exist.
2026-09-15 19:44:26 +02:00
Bastien Chanot 3b0167c6cb feat(settings): rebuild destructive-command cover in autoMode, scope the classifier environment
`permissions.ask` gates nothing under `defaultMode: auto` (LRN-146,
verified live), so the ten rules that left the static tiers had no cover
left: rsync / kill -9 / killall / pkill out of deny, and python3 -c /
python -c / xargs / sed / cp / mv out of ask.

autoMode.soft_deny (7 rules) takes over what an explicit instruction
should be able to clear: writes outside the working directory,
rsync --delete, SIGKILL and kill-by-name, in-place edits spanning more
than one file, directory moves, and inline interpreters or xargs that
delete or write outside the cwd. Intent clears a soft block for the
current turn only, stated as a rule since no setting expresses it.

autoMode.hard_deny (3 rules) takes the classes no command pattern can
express: secret exfiltration, production deployment, and disarming the
guardrails. Adding a restriction stays allowed, removing one does not.

permissions.deny gains ten .env reader rules (sed awk cut tr sort uniq
diff od xxd strings). Six of those tools sat in permissions.allow, so
reading a .env through them triggered nothing.

autoMode.environment named another project, its FTP deploy target and its
customer data, inside the file link.sh:21 symlinks to
~/.claude/settings.json, where it reached every repo and contradicted
this one's Gitea remote. Rewritten machine-generic; the project facts
moved to that project's gitignored .claude/settings.local.json. All three
lists now open with "$defaults", which the original omitted, so the
built-in classifier entries are inherited rather than replaced.

doctor.sh check_automode backstops both defects. SETTINGS.md documents
the block and a tier-choice table. README no longer claims the ask tier
makes every mcp__magic__* call require a live confirmation.
2026-09-15 19:44:24 +02:00
Bastien Chanot 3228acabfc Merge feature/gstack-playwright-lib into develop 2026-09-15 16:55:38 +02:00
Bastien Chanot 8843970425 docs: Playwright browser-cache report + bump re-applied on update 2026-09-15 16:54:33 +02:00
Bastien Chanot a0876a2976 chore(memory): BDR-088/089, LRN-150/151/152, EVAL-029 — gstack Playwright lib 2026-09-15 16:49:23 +02:00
Bastien Chanot 2cebecbb91 feat(gstack): share the Playwright bump, report the browser cache
Extract gstack_bump_playwright_if_unsupported from install-plugins.sh into
lib/gstack-playwright.sh and call it from update-all.sh too. A submodule
update no longer leaves the OS-support bump unapplied until the next
`make plugin` — that was BDR-029's open caveat.

The update helper never touches the submodule working tree: on failure it
prints git's own message and points at `make plugin`, and returns non-zero
so the existing `else warn` arm still handles it.

Add a read-only `Playwright browsers` section to doctor.sh: cache size,
which registered install requires each revision, and counts of unreferenced
directories and broken links. No pruning is written — Playwright's own
`install` already unions the required set across every registered install,
and all three installs here are live (rev 1228 for gstack + gsd-pi 1.61,
rev 1243 for gsd-pi 1.63).

Carries two latent-bug fixes from the moved code: the ostag capture exited
1 on every non-Ubuntu host and aborted the caller under inherited errexit,
and the bun calls had no timeout.
2026-09-15 14:57:31 +02:00
Bastien Chanot a53a5a26a8 Merge release/1.5.0 into develop 2026-09-13 21:26:44 +02:00
Bastien Chanot 9b3b96a8f2 chore(release): 1.5.0 — version.txt + CHANGELOG 2026-09-13 21:25:37 +02:00
Bastien Chanot 8d5d154c28 chore(memory): LRN-149 — background_tasks gates the turn-end signal 2026-09-10 03:03:22 +02:00
Bastien Chanot 0e8018ae7b fix(hooks): no turn-end signal while background work runs
Ending a turn right after spawning a subagent fired the bell and a toast
saying the response was finished, while the work continued. The Stop
payload carries background_tasks, so skip the signal when it is not
empty; the next turn end signals once the work is really done.

Interaction requests still signal during background work. A missing
field still signals, so an older client loses nothing. Also drop the
message suffix when it merely restates the label.
2026-09-10 03:03:22 +02:00
Bastien Chanot 12d7fc1483 fix(hooks): stay silent on events that need no attention
Only turn end and the moments needing the user should signal. Any other
event reaching the hook, such as agent_completed or auth_success, now
exits without emitting, so a subagent finishing rings nothing even if the
Notification matcher is ignored.
2026-09-03 02:56:49 +02:00
Bastien Chanot e801b90307 Merge chore/notify-terminal-preflight into develop 2026-09-03 02:26:16 +02:00
Bastien Chanot 92eb27c4e2 Merge feature/notify-event-labels into develop 2026-09-03 02:25:38 +02:00
Bastien Chanot 679c2cda7b feat(hooks): label each attention event in the toast
The toast body showed Claude's own message when present and the raw
notification_type otherwise, so permission_prompt and idle_prompt reached
the user as snake_case. Map every event the matcher covers to a readable
label, and keep Claude's message as a suffix when it adds detail.
2026-09-03 02:24:51 +02:00
Bastien Chanot 2c0439a0a8 chore(memory): LRN-148 — pre-flight terminal test; LRN-147 mechanism too narrow 2026-09-03 02:22:24 +02:00
Bastien Chanot 2ed51573f7 Merge chore/notify-restored-terminal into develop 2026-09-03 02:00:47 +02:00
Bastien Chanot a627201bee chore(memory): LRN-147 — restored terminals never instrumented by OSC ext 2026-09-03 02:00:22 +02:00
Bastien Chanot f90ee74a19 Merge feature/notify-stop-event into develop 2026-09-03 00:31:07 +02:00
Bastien Chanot 6aca40a810 chore(memory): BDR-087 + LRN-146 + BLK-020 — capitalize 2026-09-03 00:28:45 +02:00
Bastien Chanot ea9e5c1dab feat(hooks): ring terminal on turn end via Stop hook
Notification matcher covers input-needed events only; end of turn had no
signal but idle_prompt, ~60s late. Wire notify-attention.sh on Stop too,
branching on hook_event_name for the message. Signal only: returns
terminalSequence + suppressOutput, never blocks (guard vs BDR-083).

Header documents both client-side prerequisites found in BLK-020.
2026-09-03 00:28:40 +02:00
Bastien Chanot de34e3f167 Merge chore/reconcile-todo into develop 2026-09-01 16:55:55 +02:00
Bastien Chanot 069a73338a chore(todo): reconcile 2026-09-01 — T4 gate ticked, T6 residuals corrected, Makefile item re-verified open 2026-09-01 16:54:15 +02:00
Bastien Chanot 1940a0a22a Merge chore/notify-attention-bell into develop 2026-09-01 16:40:51 +02:00
Bastien Chanot c4e6ef1e2b chore(memory): BLK-019 notify-attention bell silent (VS Code client default) 2026-09-01 16:30:39 +02:00
Bastien Chanot 08e38876ee Merge chore/notify-attention-hook into develop 2026-09-01 15:42:25 +02:00
Bastien Chanot f08c3ab51c chore(memory): LRN-145 terminalSequence pattern + journal 2026-09-01 2026-09-01 15:32:51 +02:00
Bastien Chanot 6c04ada6a8 chore(settings): default model opus[1m] (was claude-fable-5[1m]) 2026-09-01 15:32:25 +02:00
Bastien Chanot 1d7faa32b5 chore(hooks): notify-attention — bell + OSC 777 toast when Claude needs input
Notification hook (permission_prompt|idle_prompt|agent_needs_input|
elicitation_*) returns BEL x2 + OSC 777 via the terminalSequence JSON
field (hooks have no controlling TTY). Client side over Remote-SSH:
VS Code accessibility.signals.terminalBell sound:on for the beep,
wenbopan.vscode-terminal-osc-notifier extension for the Windows toast.
2026-09-01 15:32:18 +02:00
Bastien Chanot 726464f387 Merge feature/darwin-optimize-20260825 into develop 2026-08-27 11:54:35 +02:00
Bastien Chanot a51a65e1d5 chore(memory): LRN-143 index row — re-escape pipes (sed a-command unescaped them) 2026-08-26 23:01:16 +02:00
Bastien Chanot a15854aa87 chore(memory): EVAL-028 + LRN-143/144 + BDR-086 + journal — darwin run capitalized; TODO round-count corrected 2026-08-26 23:00:41 +02:00
Bastien Chanot 7f457f09fd docs(darwin): result card PNG 2026-08-26 23:00:41 +02:00
Bastien Chanot 12823181d1 chore(darwin): Phase 3 — optimization report + TODO T5/T6 ticked 2026-08-26 22:45:03 +02:00
Bastien Chanot e157a98e0b fix(hotfix): rewrap RULES bullet — census greps the no-verifier phrase on one line 2026-08-26 22:40:04 +02:00
Bastien Chanot 6eac7fbca9 fix(hotfix): batch-skeptic residuals — RULES restore mandate file-scoped; hotfixer FILE(S) marks created files (new) 2026-08-26 22:28:31 +02:00
Bastien Chanot b5ce280fc0 fix(fixtures): plugin-check expects PLUGIN CHECK block + real plugin names; onboard archetype nextjs-app-router 2026-08-26 21:59:56 +02:00
Bastien Chanot 6c69ae1670 fix(prune-memory,code-clean): stale v1-untested note reflects real tests/; executor attribution code-cleaner (refactorer inline); audit-only fixture matches flow 2026-08-26 21:59:40 +02:00
Bastien Chanot ad4985f410 fix(security-auditor,close): hotfix no-verifier carve-out documented; close enumerates STEP 5C + --no-push passthrough 2026-08-26 21:59:08 +02:00
Bastien Chanot c983f1ff94 fix(handover-writers): stale ch.4 refs post-NAP-renumbering (glossary/tone->6, cross-links/THRESHOLD->5); 14.5 verification deferred to post-write; anchor gate ordered into STEP 16 2026-08-26 21:55:13 +02:00
Bastien Chanot 27f17939d7 fix(plan-challenger): ERROR verdict added to the load-bearing OUTPUT grammar (STEP 1 emitted it, parser enum omitted it) 2026-08-26 21:54:35 +02:00
Bastien Chanot ab75fc5e1f fix(harden): severity rule defers to the calibrated guide; SSL Labs late-finalize gets an assigned actor (main loop edits HARDEN.md row) 2026-08-26 20:35:35 +02:00
Bastien Chanot 1743f7683f fix(init-project,commit-change,tour): allowed-tools +Agent+Skill; conflict grep covers all unmerged codes; report-only never commits 2026-08-26 20:35:06 +02:00
Bastien Chanot 6eceedb8d3 fix(hotfix): revert paths — stash-create PRE snapshot + file-scoped restore (git restore . wiped tolerated user edits); security gate fresh-dispatch only 2026-08-26 20:34:45 +02:00
Bastien Chanot 796b52ea6b optimize gitflow r1-amend: judge-suggested precisions (re-checkout source before re-run; purge warning is pre-merge, finish continues) 2026-08-26 19:59:24 +02:00
Bastien Chanot 056f82b25f optimize gitflow: d3 — mechanical failure table keyed to lib return codes (rc=4 conflict resume, rc=2/3/1 start paths, best-effort purge, socle abort) 2026-08-26 19:57:21 +02:00
Bastien Chanot 9428b86880 optimize status-reporter r2: d5 — passive cost sourced from doctor.sh constants (skeptic's find), count-only fallback kept 2026-08-26 19:44:03 +02:00
Bastien Chanot a3f1624b15 optimize status-reporter: d5 — unproducible token field replaced by /plugin-check deferral; dead ROADMAP.md row rewritten for post-ADR-013 gsd layout 2026-08-26 19:35:46 +02:00
Bastien Chanot e9c6bf52fa optimize analyze-system: d1 triggers in skill description + d2 ordered TASKS mapped to OUTPUT sections in analyzer 2026-08-26 19:22:36 +02:00
Bastien Chanot 080d2d9f03 optimize plugin-probe+advisor: d8 — FRAMEWORK-DEPS exact dep@version (no preact false-hit, fallback fires), advisor derives frontend/fast-libs from it, PLAN echoed-or-unknown (no invention) 2026-08-26 19:16:09 +02:00
Bastien Chanot e92becf609 optimize profile: d3 — failure-mode table (script absent, unknown name/verb rc=1, partial toggle, split plugin leg, current-contradiction) + fixture de-drift 2026-08-26 19:11:18 +02:00
Bastien Chanot c0a2a8069d optimize refactor-system: d4 — no-tests STOP gate (agent) + user arbitration loop (dispatcher) + mid-run test-failure revert; code-cleaner inline carve-out 2026-08-26 19:05:29 +02:00
Bastien Chanot 4d72d4aa12 optimize pdf-translate: d3/d8 — failure-mode table (deps, oversize, 0-output, illisible, missing design-html/browse, QA divergence, stale workdir) 2026-08-26 15:18:38 +02:00
Bastien Chanot 562a42e936 optimize onboarder: d8 — BRIEF contract split REQUIRED/OPTIONAL, draft placeholders replace blanket STOP (fixes first-dispatch bounce vs /onboard STEP 2 minimal brief) 2026-08-26 15:12:24 +02:00
Bastien Chanot 9db213b0a1 optimize interviewer: d9 — DO NOT blacklist (no design, no invented values, no budget overrun) 2026-08-26 12:35:29 +02:00
Bastien Chanot e0923d6c4b optimize interviewer: d3 — failure-mode table (vague/idk/contradiction/partial/balloon) + 2-round budget 2026-08-26 12:34:15 +02:00
Bastien Chanot 663b6cbe22 optimize skills-perso: d8 — detection rebuilt on link.sh symlink convention (8/32 -> 31/31) 2026-08-26 11:17:16 +02:00
Bastien Chanot af002ba235 chore(darwin): T3 baseline — 54 rows, mean 83.4, 13 candidates <80 2026-08-26 11:15:03 +02:00
Bastien Chanot afec610dc8 chore(darwin): T2 gate passed — prompts reused, dim8 on candidates, threshold 80 2026-08-26 10:43:47 +02:00
Bastien Chanot a871ce5acb feat(darwin): Phase 0.5 — test-prompts for 7 promptless skills + campaign plan 2026-08-25 20:23:46 +02:00
Bastien Chanot 850f5f3f2c Merge chore/reconcile-2026-08-25 into develop 2026-08-25 20:14:57 +02:00
Bastien Chanot 7f16213456 chore(reconcile): TODO vs real — 3 open-but-done ticked, Makefile rescoped, C1 note corrected
Oracles: merge 5ec7bfa (user-writing-web-rules), BDR-065 Amendment +
LRN-138 in registry bodies, darwin-skill present + T6c green + make
test exit 0, Makefile :31 glob fixed / :57 profiles still 5/10,
seo-geo-deprescription merged 5488c48.
2026-08-25 20:12:40 +02:00
Bastien Chanot 5ec7bfa78c Merge feature/user-writing-web-rules into develop 2026-08-25 19:53:57 +02:00
Bastien Chanot dab25636c8 chore(memory): BDR-085 + journal + TODO — user permanent rules integrated 2026-08-25 19:48:06 +02:00
Bastien Chanot 12e5324065 feat(rules): user permanent rules — writing style, web building, web security (BDR-085)
Three rules/ files from the user's permanent-rules text:
- writing-style.md (always-on): em-dash ban, no it's-not-X-it's-Y, no
  emoji, no decorative bold, no reflex triads, no hedging chains, slop
  vocabulary ban, sentence-length variety, deliverable self-check.
  Scope carve-outs keep caveman registries, code comments, skill
  templates intact.
- web-building.md (path-scoped): design anti-default list + public-site
  done checklist (report missing items, never invent them).
- web-security.md (path-scoped): browser-exposed keys, service-key/client
  split, RLS, server-side auth, IDOR, cookie flags, field minimization,
  rate limiting — extends §Security, no dup of the core.
Project CLAUDE.md rules/ doctrine: 320-budget exception for standalone
always-on user rule sets.
2026-08-25 19:48:02 +02:00
Bastien Chanot dbc7d7aa70 Merge feature/tour-parallel into develop 2026-08-24 14:22:18 +02:00
Bastien Chanot b9aa9ba2e6 chore(memory): TODO — T4 verified, merge pending human gate 2026-08-24 14:19:08 +02:00
Bastien Chanot 543b0c811a feat(tour): multi-project parallel fan-out — one runner per repo
Two or more project paths dispatch one general-purpose runner per repo,
all in a single message, instead of processing repos one by one. The
runner inherits the session model — no pin, it carries tour's reflection
(fix decisions, convergence) — and every agent inside keeps its defined
tier (security-auditor sonnet, Phase B opus, doc-syncer sonnet). A dead
or mute runner becomes an explicit RUNNER FAILED summary row; the gated
capitalize offer stays in the main loop, never in a runner.

Bounded LRN-083 derogation recorded in BDR-084: the per-project fix loop
moves into its runner, but nothing a runner decides touches shared state
— independent repos, per-repo chore branches, branches left unmerged for
human review exactly as inline. Mechanics proven before building: nested
probe, 3 sub-agent windows all overlapping, 9.1s vs ~18s sequential.

Census §12: 6 locks, flip-tested. Single-project path unchanged.
2026-08-24 14:18:59 +02:00
Bastien Chanot 4ededc75ab chore(memory): TODO — plan tour-parallel (T1-T4) 2026-08-24 14:17:25 +02:00
Bastien Chanot ecbdbc7230 Merge feature/contract-gates into develop 2026-08-24 13:44:06 +02:00
Bastien Chanot 763d0022bf chore(memory): TODO — W9 human merge signal + Palier 3 trigger (won't-build-now) 2026-08-24 13:44:01 +02:00
Bastien Chanot 33f5356c82 feat(gates): wire GATE 0 into the four orchestrator skill restatements
The include is authoritative, but feat/bugfix/ship-feature/init-project
each restate the verify loop inline — an orchestrator following the
restatement alone would have skipped the floor. Each now carries the
GATE 0 bullet ahead of GATE 1 (4 new structure locks, flip-tested).
The contract-interview weight table stops promising a hotfix oracle
nothing executes: hotfix runs no floor, the hotfixer runs the suite
itself. CHANGELOG extended with the wiring + the RED result.
2026-08-24 13:36:32 +02:00
Bastien Chanot abb4ea7650 chore(memory): EVAL-027 — contract-gates behavioral RED 16/16 conformant 2026-08-24 13:26:39 +02:00
Bastien Chanot bfac4d4522 chore(memory): BDR-083 + LRN-141/142 + journal + CHANGELOG + TODO W0-W8
BDR-083 records what was taken from unlazy and, more usefully, what was
refused and why. LRN-141: an external skill's machinery encodes its threat
model, not yours — take the invariants, refuse the machinery. LRN-142:
structure locks are fixed-string, so reflowing a doctrine paragraph reds
them; fix the doc, not the lock.
2026-08-24 13:12:38 +02:00
Bastien Chanot 63310467ca feat(gates): deterministic floor (GATE 0) under the fresh verifier
GATE 1 is an LLM dispatch and the verifier's mandatory PROOF: line is a line
the verifier writes — nothing structurally stops it being produced without
anything being executed. Nothing deterministic sat between the executor and
that dispatch.

An acceptance criterion can now carry an oracle: indented CHECK: (command),
EXPECT: (success-only marker), EVIDENCE: (slot). lib/gates.sh runs them
fail-closed — MET requires exit 0 AND the marker, so a nonzero process never
passes on its error text carrying the token — and writes the outcome back
into the contract, so the fresh verifier reads evidence as fact rather than
trusting the executor's report.

GATE 0 runs that floor before any verifier is dispatched; a red build sends
the executor back for free, on its own iteration budget. ABANDON: <id>
<reason> turns an impossible criterion into a visible handoff that blocks
CONFORME and routes to the human gate, via the new ABANDONED(n) verdict —
a distinct token because it routes distinctly, never a dev loop. feater and
bugfixer gain a four-pass completion discipline, scoped so a pass can never
widen the contract.

The runner's parse fails closed on partial oracles, duplicate ids,
unindented attributes and runnable criteria with no EVIDENCE: line, and
executes nothing at all when the ledger is malformed. status never executes
and never writes; run always re-executes, since trusting written evidence is
the failure being closed.

Adapted from the unlazy skill (Leonxlnx/unlazy, MIT). Its Stop hook,
approval store, .unlazy/ tree, depth-tree arithmetic and Node checker were
deliberately refused — BDR-083 records each reason.

64 assertions in lib/tests/gates.test.sh, non-execution proved by sentinel
with its own positive control asserted first.
2026-08-24 13:12:38 +02:00
Bastien Chanot 5488c4870f Merge feature/seo-geo-deprescription into develop 2026-08-24 12:19:10 +02:00
Bastien Chanot ae1339d656 chore(config): untrack emil-design-eng skill (machine-owned curl copy)
skills-external/emil-design-eng/SKILL.md is curl'd from emilkowalski/skill
by install-plugins.sh when absent and re-fetched unconditionally by every
update-all.sh run, so tracking it produced a repo diff on each upstream edit
(latest: Radix vars dropped for Base UI). Same category as frontend-design/
and impeccable/, already ignored on that rationale — a fresh clone re-fetches
it, so no offline copy is needed and nothing was pinned here anyway.

design-motion-principles/ has the same overwrite-on-update behaviour but NOT
the same bootstrap: install-plugins.sh only warns instead of cloning it, so it
stays tracked until that gap is closed.
2026-08-24 12:17:17 +02:00
Bastien Chanot 325962e080 chore(memory): BDR-082 + LRN-140 + journal + CHANGELOG + TODO C1 done (seo/geo de-prescription) 2026-08-02 17:28:07 +02:00
Bastien Chanot c7646a9c8a feat(agents): de-prescribe geo-analyzer for Opus 5 (C1 P3)
Same invariant as adafa35. Census 71/0 green; contract surface
byte-identical; 1106→1107 lines (single-shot scoping line).

- MANDATORY/MUST caps on AI-index submission → plain content rule
- 'Print the plan before STEP 13' → single-shot-scoped (conf#1);
  tier-mapping kept in BOTH judge and template ranges (rob#1/#9,
  PERMISSIVE :873 named survivor kept)
- vestigial ':1106 Transparency §14' reworded to truth: change log is
  dispatcher's SEO.md §15 (folded into 'Dispatcher verifies')
- 'copy these patterns' → reporting-shape-to-match; FAQ '20-50'
  quantity → 'typically dozens'; 'EVERY finding' → outcome bar
- dedup after inspection: ZERO merges (PERMISSIVE ×3 cross-range;
  never-apply ×4 distinct obligations; content_quality pair =
  spec rule vs emitted-artifact caveat — annex counts corrected)
- FROZEN untouched: guard-first :273, NAP direction rule, cite-sources,
  WebSearch-freshness (already when-shaped, rob#6), all STEP headers
- deltas: MANDATORY 1→0, MUST 4→3, NEVER 9→8
2026-08-02 01:30:12 +02:00
Bastien Chanot adafa350da feat(agents): de-prescribe seo-analyzer for Opus 5 (C1 P2)
Choreography → when-guidance under the audience×mode-range invariant
(plan §4b Q3). Census 71/0 + model-routing 133/0 + seo-data 221/0 green
throughout; contract surface (§2a/2b) byte-identical; 1528→1503 lines.

- self-output verification: ':970 run it twice' → deterministic-engine
  integrity guard (conf#8); ':1217 do not proceed' → single-shot-scoped
  (conf#1); grouping sanity-check → when-guidance detector (rob#8)
- vestigial pre-BDR-061 'Transparency' line deleted; §15 ownership
  folded into 'Dispatcher verifies'
- dedup (inspection-corrected: most annex 'twins' are distinct
  obligations — kept): only true same-range dups removed ('Handoff to
  dispatcher' ≈ sentinel note; 'Landing page rule' block ≈ payload
  instance + RULES line)
- completeness checklist reshaped to routing map, rows verbatim (rob#3)
- caps softened: 'P0 rule' MUST/ALWAYS → plain content rules (CMS
  plugin-first folded with STEP 2 twin, corr#3); First-action/ordering
  emphasis dropped; C1a + sampling essays compressed (rules + LRN
  citations kept); WebSearch → drifting-externals when-guidance (rob#6)
- FROZEN untouched: guard-first orderings :287, denominator-before-
  sampling :550, R2 refuse, COVERAGE obligations, all STEP headers
- deltas: 'P0 rule' 2→0, ALWAYS 1→0, MUST 5→4; NEVER 9→9 (class-B
  named bans, kept by design LRN-105)
2026-08-02 01:30:02 +02:00
Bastien Chanot 9681b468e1 test(census): seo/geo agent⇄dispatcher contract locks (C1 P1, pre-reword)
71 locks, flip-proven (7 scratch mutations → 7 FAILs): judge verdict
grammar, FIX BUNDLE + READY-TO-APPLY sentinel, signals handoff, ALL
STEP headers (interiors included, conf#5), collect report, bundle item
fields parsed by L1 appliers, score labels (BDR-010/LRN-011), scoring
blocks, trajectory, envelope keys. Locks existing state — reword
commits must keep this green.

+ plan v3 (challenged 3 lenses FATAL/FATAL/CONCERNS + 1 confirmation
pass FATAL(9), every BLOCKER closed by a named change, §5bis record)
+ directive-language inventory annex (analyzer report).
2026-08-02 01:10:47 +02:00
Bastien Chanot 7047adfe77 chore(memory): reconcile TODO — opus5 branch merged (709cf9b), add Claude 5 follow-on chantiers C1-C4 2026-07-30 13:41:37 +02:00
Bastien Chanot 709cf9bf0b Merge feature/opus5-config-tuning into develop 2026-07-30 13:28:36 +02:00
Bastien Chanot 550b39043e chore(memory): BDR-081 + LRN-139 + journal + CHANGELOG + plan (opus5 tuning)
Capitalizes the Claude-5-family config recalibration: decision record,
trait-inversion learning (LRN-030 superseded premise, #80988 injections,
no-effort-hold trap), journal line, CHANGELOG Unreleased entries, and the
challenged plan (3 blind Opus 5 lenses, synthesis in §5bis).
2026-07-30 13:02:55 +02:00
Bastien Chanot c3d3f4d465 feat(agents): plan-challenger — route grounded doubts to [MINOR]
Opus 5 follows conservative-reporting clauses literally; 'a manufactured
concern is a failure' risked suppressing real low-confidence findings.
In-place reword: ungrounded stays noise, grounded-but-uncertain files as
[MINOR] with the uncertainty in WHY:. OUTPUT grammar byte-identical;
census row added.
2026-07-30 12:58:29 +02:00
Bastien Chanot 0f7b565bb0 feat(global): recalibrate instruction layer for Claude 5 family
- Delegation block: model-neutral when-guidance replaces the Opus 4.8
  under-delegation counter (LRN-030 trait inverted on Opus 5; Claude
  Code injects its own anti-delegation prompt there, #80988). Gates
  carve-out keeps verifier/security/challenge dispatch mandatory.
- Drop the 'staff engineer' self-check bar (Opus 5 over-verification
  trigger per official migration guide); honest-reporting steps stay.
- Deviations bullet: finish-whole-task clause (Opus 5 scope-expansion
  counter), scoped so 'gone wrong → STOP' still wins.
- Written-deliverable length rule (Opus 5 writes ~30-40% longer).
2026-07-30 12:58:06 +02:00
Bastien Chanot eab2a10cd5 fix(hooks): drop \bux\b from design-toolchain pattern (FR prose FPs)
3rd tightening pass (series LRN-1005/1007): bare "ux" matched inside
French prose (2 logged FPs, both FR — latest "changement ux vu").
\bui\b kept: zero logged FP, one logged true positive, now locked by a
must-fire test row. Flip-tested: quiet row fired pre-change.
2026-07-30 12:57:38 +02:00
Bastien Chanot f05ca86ef2 Merge release/1.4.0 into develop 2026-07-22 15:48:42 +02:00
114 changed files with 5491 additions and 2844 deletions
Binary file not shown.

After

Width:  |  Height:  |  Size: 254 KiB

+90
View File
@@ -0,0 +1,90 @@
# Darwin run 2026-08-25/26: fresh baseline + threshold optimization + bug pass
Branch `feature/darwin-optimize-20260825`, 26 commits, 39 files, +299/-142.
Log: `~/.agents/skills/darwin-skill/results.tsv` (fresh, the May file was wiped
by the 2026-06-23 reinstall). Method: darwin v2.1. Absolute scores served as
triage only; every keep/revert decision came from a paired same-judge majority
(3 judges per round, before/after read in one call).
## Scope
54 units: 31 personal skill-systems (SKILL.md + dispatched agents judged
together, per EVAL-004) and 23 agents. Excluded: gstack/external symlinks
(BDR-015/043, LRN-070), darwin-skill itself (BDR-058 pin), and find-docs,
newly identified as machine-owned ctx7 output (gitignored, installer-written).
## Baseline (7 blind judges, dims scored 1-10, totals recomputed main-thread per LRN-018)
Mean 83.4 (skills 83.5, agents 83.3). Best: deploy, release-candidate,
release-executor (90.4). Worst: skills-perso 63.5. All dim8 rows marked
dry_run by design; live execution happened later, inside the paired rounds.
13 units scored below the user-set threshold of 80.
## Phase 2: threshold loop, 13/13 units, 0 reverts
Every round was validated by 3 paired judges (neutral, skeptic, realism).
All verdicts 3-0 better.
| Unit (baseline) | Round(s) | What changed |
|---|---|---|
| skills-perso (63.5) | d8 | Detection rebuilt on the link.sh convention: symlink = external, real dir = personal, gitignored = machine-generated. Live result 8/31 to 31/31, zero false positives |
| interviewer (70.9) | d3, d9 | Failure-mode table (vague, "you decide", contradiction, partial, balloon) + 2-round budget; DO-NOT list |
| onboarder (71.5) | d8 | BRIEF contract split REQUIRED/OPTIONAL; null enrichment becomes TODO placeholders; STOP kept for required keys and unresolved monorepo. Kills the guaranteed first-dispatch bounce vs /onboard STEP 2 |
| pdf-translate (72.3) | d3/d8 | 8-row failure table: deps, >30 pages gate, zero-output, illisible markers, design-html/browse fallbacks, QA cap 2, stale workdir |
| refactor (75.6) + refactorer (76.8) | d4/d3 | No-tests STOP gate + GO-WITHOUT-TESTS arbitration in the dispatcher; mid-run test-failure revert protocol; code-cleaner inline carve-out |
| profile (77.3) | d3 | 6-row failure table, every row fact-checked against profile.sh (rc=1 paths, partial toggle, split plugin leg, BLK-006 contradiction); fixture de-drift |
| plugin-probe (78.5) + plugin-advisor (77.5) | d8 | FRAMEWORK-DEPS now exact dep@version (preact false-hit killed, fallback actually fires; the old `\|\| true` silently emitted nothing and tripped the advisor's fail-closed path on non-Node projects); frontend/fast-libs derivable; PLAN echoed-or-unknown, invention removed |
| analyze (77.7) + analyzer (78.0) | d1, d2 | Bilingual triggers + fix-wanted disambiguator; TASKS ordered, each step mapped to its OUTPUT section |
| status-reporter (78.0) | d5 x2 | Fabrication-forcing token field replaced, then restored producibly from doctor.sh constants (a skeptic judge found the source); dead ROADMAP row rewritten post-ADR-013 |
| gitflow (78.4) | d3 | 7-row failure table keyed to lib return codes; rc=4 conflict resume empirically verified; human merge gate untouched |
## Bug pass: verified defects in above-threshold units, 8 commits, all kept 3-0
- hotfix: `git restore .` on every failure branch wiped tolerated in-progress
user edits. Now: `git stash create` pre-flight snapshot + file-scoped
restore + fresh-dispatch-only security gate. Two skeptic residuals amended
(RULES bullet, FILE(S) new-file marker).
- init-project: allowed-tools lacked Agent and Skill while every step
dispatches. commit-change: conflict grep now covers all 7 unmerged codes.
tour: --report-only no longer commits (could land on develop).
- harden: severity rule now defers to the calibrated guide; the late SSL Labs
grade has an assigned actor.
- plan-challenger: ERROR joined the load-bearing verdict grammar.
- handover writers: stale chapter refs corrected (glossary/tone to §6,
cross-links and THRESHOLD-OVERRIDE to §5); STEP 14.5 verification deferred
post-write; anchor gate ordered into STEP 16.
- security-auditor: /hotfix no-verifier carve-out documented. close: STEP 5C
enumerated, --no-push passthrough added.
- prune-memory: false "v1-untested" note replaced by the real tests/ state.
code-clean: executor attribution corrected (code-cleaner, refactorer inline).
- Fixtures de-drifted: plugin-check (PLUGIN CHECK block, real plugin names),
onboard (nextjs-app-router).
`make test` green (0 RED, rc=0) after one census rewrap: a locked phrase had
been line-wrapped and the single-line grep lock caught it.
## Residual findings, logged not fixed
- analyze triggers: "how does X work" brushes graphify's territory; graphify's
graph-exists routing still wins.
- pdf-translate: pdfinfo row assumes poppler (fitz also has page count); "GB"
slightly overstated near the 30-page gate.
- web-validate: .validate-cache mkdir lives in a skipped STEP 0
(self-recoverable); axis budgets 35/25/40 never reconciled with the base-100
deduction table. seo/geo minor wording items. verifier/doc-syncer/audit-delta
restatement redundancy (cosmetic). handover-doc-writer STEP 14.5 umbrella
line still says "BEFORE STEP 15" while the inner note overrides it.
- bugfix trivial-fast-path boundary loosely defined; feat prompt-3 expectation
vs full gate pipeline.
## Methodology notes
- v2.1 paired majority produced 36 unit-round verdicts and 24 batch verdicts,
all better, 0 reverts, 0 ties. The May-2026 run under absolute-delta scoring
had reverted 2 edits on judge noise; this run had no such event.
- Judges live-executed wherever the artifact was executable (skills-perso
detection, profile.sh probes, plugin grep on scratch manifests, doctor.sh
grep, git merge no-op resume). Behavior outranked prose in 5 units.
- Two grep-exit-masking bugs surfaced (a `head` pipe swallowing the fallback's
trigger), one in the probe being fixed, one in this run's own test harness.
The pattern is worth a learning entry.
+33
View File
@@ -215,3 +215,36 @@ rules:
- **Solution** (workaround): dispatcher ran `gitflow.sh finish` + tag inline after its own human gate — where the signal is real. Release completed clean (main `648bc6e`, tag v1.3.1).
- **Status**: open. Candidate fixes: (a) quote gate evidence verbatim in span prompt — untested vs classifier; (b) move finish+tag span permanently inline in /release-candidate — keeps prep span dispatched, costs the sonnet pin on ~5 mechanical commands, cheap; (c) permission rule allowing subagent `gitflow.sh finish` — weakens the guard, refused. Decide at next release.
- **Reference**: skill `release-candidate` STEP 5. Pattern adjacent [[LRN-089]] (ambient-state/context assumptions across boundaries). Journal 2026-07-20.
## BLK-019 — notify-attention bell silent, toast OK (VS Code client default) — 2026-09-01
- **Friction**: hook fired, Windows toast arrived, native bell never audible. User heard only Windows toast sound. Looked like half-broken hook.
- **Real cause**: not hook. Toast proves full `terminalSequence` reached terminal, `\a\a` sits at head of that same string → BEL emitted. VS Code defaults `accessibility.signals.terminalBell` to `"auto"` = sound OFF unless screen reader active.
- **Solution**: `"accessibility.signals.terminalBell": { "sound": "on" }` in CLIENT-side user settings.json (`c:/Users/<u>/AppData/Roaming/Code/User/`). Unreachable from remote: real SSH remote, not WSL (no `/mnt/c`, `/proc/version` no Microsoft). User applied, retest → both channels OK.
- **Status**: resolved (per-client-machine, not repo-portable).
- **Reference**: `~/.claude/hooks/notify-attention.sh` header already documented the setting; never applied. New client machine → bell mute again while toast works. Silent-degradation class [[LRN-047]].
## BLK-020 — notify-attention: both channels dead on one VS Code client — 2026-09-02
- **Friction**: client-side prereqs applied on Windows box (ext `wenbopan.vscode-terminal-osc-notifier` + `accessibility.signals.terminalBell` sound:on), window reloaded. AskUserQuestion → nothing. `idle_prompt` 90s wait → nothing. Direct write `\a\a` + OSC 777 to claude own pty (`/dev/pts/2`) → nothing. Second client machine, same SSH server, same hook, same registries → both channels OK.
- **Server side cleared**: hook dry-run emits `BELx2 + OSC 777 + ST` correctly, `jq` present, matcher covers `idle_prompt`, ext NOT wrongly installed remote-side. Not a hook bug — same class as [[BLK-019]] (client default silently degrades).
- **Real cause**: unresolved. Facts: claude runs under `dtach -c ~/.dtach/claude-190012`; claude fd1 = `/dev/pts/2` (inner pty, dtach master side), REAL VS Code terminal = `/dev/pts/1` held by dtach client pid 742794. `VSCODE_SHELL_INTEGRATION` unset this terminal; ext marketplace doc requires shell integration ON. BUT other working session (`claude-154323`) also runs under dtach → dtach alone insufficient explanation, weight shifts back to client-side.
- **Probes run**: direct write to `/dev/pts/1` (real VS Code pty, chain alive: bash pts/1 → dtach client 742794 S+ → master → claude pts/2) → no bell, no toast. Visible-marker injection both paths → user saw neither, BUT inconclusive: claude TUI repaints, injected text clobbered next frame. Only BEL is repaint-proof, and BEL stays silent.
- **Client settings verified by user**: settings.json path correct (no VS Code profile indirection), `terminalBell` sound on, ext installed + enabled local side. VS Code recent (server dirs 2026-08), so ≥ 1.93 ext requirement met.
- **Next probe**: user opens FRESH VS Code integrated terminal (no dtach, no claude TUI) and runs `printf '\a\a\033]777;notify;Test;hello\033\\'`. Isolates client renderer from claude/dtach path. Beep+toast there → fault in claude/dtach path; nothing → client-side, diff against working machine.
- **Fresh-terminal probe (decisive)**: user ran `printf '\a\a\033]777;notify;Test;hello\033\\'` in NEW VS Code terminal → toast OK, bell still silent. Splits one symptom into TWO independent faults.
- **Fault A (toast in claude session)**: ext parses only terminals created AFTER its activation. Claude terminal pts/1 born 19:00, ext installed later same day → that terminal never hooked. Fix: restart claude in fresh terminal, or re-attach existing dtach session from one (`dtach -a ~/.dtach/<sess>`; dtach broadcasts to multiple clients, no session loss). NOT a dtach filtering bug — earlier hypothesis wrong.
- **Fault B (bell)**: silent even in fresh terminal where toast works → not terminal path, VS Code audio side. Toast sound = Windows notification (works); bell = VS Code process audio (mute). Suspects: signal volume option, Windows volume mixer entry for Code, output device. Probe: palette `Help: List Signal Sounds` → Terminal Bell plays preview or not.
- **Fault A RESOLVED (verified 2026-09-02)**: re-attached session from fresh terminal (`dtach -a ~/.dtach/claude-190012`, new client pts/3). Both sends toasted — one through session path (pts/2, dtach broadcast), one direct. Rule: ext hooks only terminals born AFTER its activation → install ext, THEN start/re-attach claude session. dtach broadcast means zero session loss.
- **Fault B still open**: bell silent on every path. New signal: toasts arrive but user reports NO sound at all, while [[BLK-019]] machine got audible Windows toast sound. Both audio channels dead + both visual channels fine → common factor is client audio output, not terminal stream. Suspects ranked: Windows per-app notification sound off for Code, system/app volume mixer mute, wrong output device, `accessibility.signalOptions.volume` 0.
- **Fault B ROOT CAUSE ISOLATED (2026-09-02)**: palette `Help: List Signal Sounds` → Terminal Bell preview plays NO sound, while Windows toast sound IS audible. Preview bypasses terminal, BEL, hook, dtach, ext entirely → VS Code renderer audio itself mute on this box. Toast sound emitted by Windows shell, not by Code → explains why one audio channel works and other does not.
- **Fix candidates (client, ranked)**: (1) Windows per-app volume mixer — Code muted/0, or per-app OUTPUT DEVICE pointing at disconnected device (mixer only lists app after it attempts playback → hit preview first, then open mixer); (2) VS Code `accessibility.signalOptions.volume` = 0 kills all signals; (3) compare both against working machine.
- **Pragmatic out**: toast already carries audible Windows sound → attention signal functional without bell. Bell is redundant channel, not blocker.
- **Fault B RESOLVED (2026-09-03)**: cause = Windows per-app volume mixer, Code entry at 0. Toast audible throughout because Windows shell emits that sound, not Code → masked a plain app-volume mute. User set volume up → bell audible.
- **Status**: resolved (A: ext hooks only terminals born after activation → install ext THEN start/re-attach session; B: Code app volume 0 in Windows mixer).
- **Lesson**: two independent client faults presented as one symptom ("nothing works"). Splitting probe = run signal in FRESH terminal + play VS Code's own sound preview. Preview bypasses terminal/BEL/hook/dtach/ext → isolates renderer audio in one step. Do that FIRST next time, before any server-side archaeology.
- **Reference**: [[BLK-019]] bell-only variant (resolved differently — setting alone insufficient here), [[LRN-145]] terminalSequence-not-/dev/tty pattern. Silent-degradation class [[LRN-047]].
## BLK-021 — Bash tool dead mid-session ("every command exits 1"): /tmp usrquota blown by a dead session's probe HOMEs — 2026-09-22
- **Friction**: previous session on `feature/21st-cli-migration` lost its shell before tests + commit: every Bash call, `echo` included, returned 1. Its harness file `imptest2/step8d-test.sh` landed as 0 bytes.
- **Real cause** (strong evidence, not reproduced on purpose): `/tmp` = tmpfs 7.4 GB mounted `usrquota`; `/tmp/claude-1000/-home-bchanot-Documents-claude/fefd277c-…/scratchpad` holds 5.9 GB of sandbox HOMEs (`pinprobe/` 2.1 GB, `pinrc/` 1.6 GB, `imp1 impg imptest sbx1 sbx2 v3.2.0 v3.6.1 v4.0.5 …`) from the impeccable pin probes. `dd` 40 MB to `/tmp/claude-1000` → "Disk quota exceeded" (EDQUOT) while `df` still shows 1.6 GB avail. Same write to `~/.cache` OK. Bash tool + `mktemp` + heredocs live in /tmp → all die together. This session: first impeccable probe failed with `Quota exceeded (os error 122)` on the installer's `/tmp/impeccable-update-*` staging, same cause.
- **Solution**: this round ran everything with `TMPDIR=~/.cache/imp-probe/tmp` (probe, harness, `make test`). Durable fix = delete the dead session's scratchpad: `rm -rf /tmp/claude-1000/-home-bchanot-Documents-claude/fefd277c-e143-4d51-b589-a566641079b5` (agent's `rm -rf` on /tmp denied by the classifier → user action). Rule for probes: sandbox HOMEs that pull npm/node payloads go under `~/.cache/<probe>/`, never the /tmp scratchpad, and get removed at the end of the session.
- **Status**: open until the user frees /tmp. Links [[BDR-094]], [[LRN-159]].
+123
View File
@@ -92,6 +92,17 @@ rules:
| BDR-072 | 2026-07-17 | SPA: honest refuse (On-page N/A, not zero), no headless browser (R2 over R1) | accepted |
| BDR-073 | 2026-07-17 | Scoring: LLM judges findings+severity, engine does the arithmetic (deterministic /20) | accepted |
| BDR-080 | 2026-07-21 | Bug routing inverted: /bugfix primary, /investigate explicit-only | accepted |
| BDR-083 | 2026-08-24 | Contract gates: deterministic floor (GATE 0) under the fresh verifier | accepted |
| BDR-084 | 2026-08-24 | /tour multi-project: parallel runners (LRN-083 derogation, bounded), runner inherits session model | accepted |
| BDR-085 | 2026-08-25 | User permanent rules: writing-style always-on in rules/, web build+security path-scoped | accepted |
| BDR-086 | 2026-08-26 | darwin: threshold gates full loops; verified defects fixed regardless of unit score (paired-validated, batched checkpoint) | accepted |
| BDR-087 | 2026-09-03 | Stop hook = attention signal only, never control flow; one script for Notification + Stop | accepted |
| BDR-088 | 2026-09-15 | gstack Playwright bump shared via lib, re-applied after submodule update; update helper never touches the submodule tree | accepted |
| BDR-089 | 2026-09-15 | No Playwright browser-cache pruner; read-only doctor report — .links proved 0 bytes reclaimable | accepted |
| BDR-090 | 2026-09-15 | Destructive shell work → autoMode soft_deny/hard_deny; `ask` tier abandoned (inert under auto) | accepted |
| BDR-091 | 2026-09-16 | Ask, don't guess: open-choice sweep (3 classes) at plan step + mid-run CLASS channel supersede "one question upfront" | accepted |
| BDR-092 | 2026-09-16 | docker + node framed by the classifier via autoMode.allow + soft_deny; ask entries retired | accepted |
| BDR-093 | 2026-09-22 | 21st.dev: magic MCP → CLI + skill pack; staged install past the ~/.claude/skills symlink; CLI = gate's required-manual | accepted |
---
@@ -1073,3 +1084,115 @@ Audit (user ask "profile toggles externals both ways?"): ASYMMETRIC. Enable side
### BDR-080 — bug routing inverted: /bugfix primary, /investigate explicit-only [accepted] (2026-07-21)
Old routing "Bug → investigate (bugfix if gstack off)" + gstack ON by default → every bug took path bypassing own quality pipeline (gitflow aiguillage, contract, fresh verifier + security gates, doc-sync, `.claude/memory` registries) — /bugfix relegated to near-never fallback. Skill comparison: same core doctrine (root-cause iron law, hypothesis loop, regression test, 3-strike stop, >5-files alert) but incompatible wrappers — investigate monolithic (same context investigates+fixes+verifies, ~1075-line SKILL.md w/ gstack preamble/telemetry/onboarding, capitalizes to `~/.gstack` learnings.jsonl framework never reads at session start); bugfix orchestrator (reflection inline, sonnet bugfixer executor, fresh gates — BDR-066, LRN-083). Composition rejected: skills superpose in context, don't compose — invoking investigate inside bugfix = two full workflows, two completion protocols, two memory systems loaded at once. Decision: CLAUDE.global.md routing line inverted — bugfix primary; investigate ONLY on explicit ask for gstack ecosystem (cross-project learnings, /freeze scope lock, long no-commit investigation). Alternatives rejected: keep investigate primary (bypasses framework), embed investigate inside bugfix (context conflict, dual memory). Known drift noted at write time: Index table rows BDR-074..079 missing (pre-existing, /prune-memory scope).
### BDR-081 — Config recalibrated for Claude 5 family (Opus 5 dispatch tier) [accepted] (2026-07-30)
Opus 5 (released 2026-07-24) now backs every `model: opus` pin (BDR-076/077) + any `/model opus` session. Research (official migration guide + web + registries): Opus 5 OVER-delegates (inverts LRN-030 Opus 4.8 trait that CLAUDE.global.md:43-47 compensated), self-verifies (explicit verify instructions → over-verification, "removing them reduces wasted tokens with no loss in quality"), literal following (conservative-reporting clauses depress recall; MUST/CRITICAL over-triggers), scope expansion = named regression, written deliverables +30-40%. Claude Code injects Opus-5-only anti-delegation prompt sections (heron_brook + subagent_steer_delegation, issue #80988, server-gated, no opt-out) — prose caps would triple-stack. Shipped: delegation block → model-neutral WHEN-guidance + explicit gates carve-out (verifier/security/challenge still dispatch as written); "staff engineer" self-check bar dropped; finish-whole-task clause folded into Deviations (gone-WRONG→STOP still wins); deliverable-length rule; design hook `\bux\b` dropped (`\bui\b` KEPT — 0 FP, 1 logged TP, lock-tested); plan-challenger grounded-doubt→[MINOR] in-place reword (grammar byte-identical). Plan challenged by 3 blind Opus 5 plan-challengers: correctness CONCERNS(4) / robustness FATAL(5, BLOCKER: all surfaces symlink-deployed LIVE — gates fire post-deployment) / simplicity CONCERNS(4); every fix adopted as prescribed (scratch-validation before live hook write, minimal diffs, ux-only, MINOR-routing). Alternatives rejected: leave as-is (nudge actively counter-productive); hard spawn caps in prose (harness injects one); confidence axis on challenger grammar (consumer unwired); dropping \bui\b (no evidence). NOT touched: verify-secure-loop + fresh gates (harness architecture BDR-049/050, ≠ model self-check prose); Security/Architecture sections (BDR-021); settings effortLevel xhigh (user pref — Opus 5 carry-over trap → LRN-139); superpowers plugin wording (external upstream). Plan+synthesis: .claude/tasks/plans/2026-07-30-opus5-config-tuning-1238.md. Branch feature/opus5-config-tuning, unmerged (human gate).
### BDR-082 — seo/geo analyzers de-prescribed for Opus 5 (C1) [accepted] (2026-08-02)
BDR-081 N5 follow-on, user-directed apparatus (plan+3-lens challenge+census+dogfood). Method: audience×mode-range invariant — dedup ONLY verbatim same-audience (spec rule / bundle-item payload / phase-local caveat) same-mode-range repeats; cross-mode + agent↔dispatcher twins stay (standalone paths need them). Census-FIRST: lib/tests/seo-geo-contract.test.sh 71 locks (verdict grammar, sentinels, ALL STEP headers incl. interiors, item fields, score labels, envelope keys), flip-proven 7 mutations→7 FAILs, committed BEFORE reword. Shipped: self-output verification removed (":970 run twice"→conditional integrity guard; ":1217"→single-shot-scoped), 2 pre-BDR-061 vestigials fixed, caps softened (P0-rule/MANDATORY/ALWAYS→plain content rules), 2 essays compressed, checklist :1309→routing map rows verbatim (challenger caught it = routing table, NOT self-check), true same-range dups only (seo Handoff+landing-page blocks; geo ZERO — all claimed pairs distinct on inspection). FROZEN: guard-first url-guard orderings, :550 denominator-before-sampling (ordering IS the honesty mechanism), R2/NAP/COVERAGE/citation invariants, external-freshness checks (world drift ≠ self-verification). Deltas: seo 1528→1503 l ("P0 rule" 2→0, ALWAYS 1→0, MUST 5→4, NEVER 9→9 = class-B bans kept); geo 1106→1107 (MANDATORY 1→0, MUST 4→3). Plan challenged correctness FATAL / robustness FATAL(3 BLOCKER) / simplicity CONCERNS + confirmation FATAL(9) — every BLOCKER closed by named change (§5bis record). Dogfood before/after on frozen zenquality copy: judge-replay on frozen signals (zero collect variance) + templates + fresh collects + e2e judge + 42/42 assert battery BOTH sets + blind reader "interchangeable; all deltas = presentation variance both directions OR after MORE spec-conformant". Alternatives rejected: keyword dedup (challengers proved audience/range-blind — most annex "twins" were distinct obligations), FULL/aggressive dogfood (billing gate killed nested CLI; left as user option), banner/shape locks (LLM-convention layers wobble — lock strings only). Evidence: .audit/dogfood-baseline/ (18 artifacts + DOGFOOD-VERDICT.md), plan .claude/tasks/plans/2026-07-30-seo-geo-deprescription-1402.md. Branch feature/seo-geo-deprescription, UNMERGED (human gate).
### BDR-083 — contract gates: deterministic floor (GATE 0) under the verifier [accepted] (2026-08-24)
User asked what to take from `unlazy` skill (Leonxlnx/unlazy 2.1.0, MIT). Verdict on its verification ARCHITECTURE: teaches nothing we lack — contract + fresh blind verifier + bounded loops + order invariant already shipped (BDR-049/050/066, LRN-083). Real gap found elsewhere: between executor and GATE 1, NO deterministic floor. GATE 1 = LLM dispatch; verifier's mandatory `PROOF:` line = a line the verifier WRITES — nothing structurally stops it being produced without executing anything (LRN-048 demands a pass prove it looked; the proof is self-reported prose). Decision: import unlazy's gate ledger INTO the existing contract, never alongside it. Palier 2, user-chosen over doctrine-only / defer.
TAKEN: criterion carries an oracle (indented `CHECK:` cmd + `EXPECT:` success-only marker + `EVIDENCE:` slot); fail-closed = exit 0 AND marker (a nonzero process never passes because its error text carries the token); evidence persisted INTO the contract → the fresh verifier reads fact, not the executor's report; `ABANDON: <id> <non-blank reason>` = impossible criterion never deleted, blocks CONFORME, routes to human gate (new verdict token `ABANDONED(n)` — distinct routing from ECARTS ⇒ distinct token, not a sub-line to re-derive); 4 gate-authoring rules (observe the named artifact / success-only marker / positive control before any absence check / recompute supplied numbers, never copy one into EXPECT); 4-pass executor discipline (feater full; bugfixer narrowed to fix+test under "keep the fix minimal", pass 3 = negative control proving the regression test fails without the fix).
REFUSED + why: Stop hook `decision:"block"` — contradicts "STOP + human escalation", "gone WRONG → STOP re-plan", "merge only on explicit human signal"; a hook FORCING continuation is the inverse of our gates; its 6-block release either traps the session or gives up; each block = an agent continuation = real tokens. Approval store `~/.unlazy/approved` (binds ledger+cmd+CWD+shell+timeout+platform+full PATH) — exists to execute ledgers INHERITED from untrusted repos; our contracts are authored by our own orchestrator in our own repo ⇒ biggest chunk of their 28k checker closes zero threat here. `.unlazy/<scope>/` tree (PLAN+GATES+gates/+status.log+session+hook-state+locks/) — a 4th bookkeeping tree beside .claude/tasks/{contracts,plans} + memory/ + audits/. `tree N` Depth-Tree effort arithmetic — disowned by unlazy's OWN research/validation-protocol.md (v1 six-run figures unreproducible), while the repo DESCRIPTION still advertises the retracted claim. Node checker (28k .mjs + 54k .mjs tests) — lib stack is 100% bash, Health Stack = `shellcheck *.sh hooks/*.sh lib/*.sh` would cover none of it. `OWNS:` ownership leases — deferred (Palier 3): our parallel dispatches (seo/geo, 3 plan-challengers) are read-only, the write-collision problem does not exist yet.
Shipped: lib/gates.sh (~250 l bash; `status` never executes and never writes · `run` ALWAYS re-executes every runnable criterion — trusting written evidence is the failure being closed, so there is no incremental mode to get wrong; rc 0 MET / 2 UNMET|malformed / 3 ABANDONED; parse fails closed on partial oracle, duplicate id, unindented attribute, runnable-without-EVIDENCE, and executes nothing when the ledger is malformed). GATE 0 in lib/verify-secure-loop.md (red floor → executor re-dispatch with the NOT-MET rows, NO verifier spawned; own 3-iteration budget, separate from conformity; malformed ledger fixed in the main loop, never dispatched to a dev). Order invariant now GATE 0→1→2 on every re-loop. lib/contract-interview.md: ORACLES section + template + ABANDONMENT lifecycle + per-flow oracle weight. agents/verifier.md: oracle-consumption rules — a red or unrun oracle is NEVER overridden by reading code; a MET oracle proves the ORACLE, not the English sentence ⇒ vacuous oracle = NOT-MET, the one judgement no command can make; verifier may re-run a CHECK but never edits the contract. lib/tests/gates.test.sh 64 assertions (sentinel-proved non-execution, with its own positive control asserted first).
Alternatives rejected: Palier 1 doctrine-only (CHECK:/EXPECT: become decorative without an executant); port the Node checker (stack break, shellcheck-blind); fold ABANDONED into ECARTS (would send a dev to fix the impossible and eat the 3-iteration budget); `status` revalidating old evidence (that trust is the failure being closed).
Branch feature/contract-gates, UNMERGED (human gate). `make test` rc 0, shellcheck clean, e2e verified on a real contract in the documented template.
### BDR-084 — /tour multi-project: parallel runners, bounded LRN-083 derogation [accepted] (2026-08-24)
User asked whether agent parallelism on independent tasks is ACTIVE. Measured first (LRN-080): (a) mechanics — nested probe, 1 dispatched orchestrator fanned 3 sub-agents, execution windows all overlap, 9.1s vs ~18s sequential ⇒ nested parallel dispatch WORKS; (b) doctrine — already prescribed at 3 layers (harness "single message" injection; /seo, challenge-plan, /cso, graphify explicit same-message mandates; graphify even anti-sequential wording); remaining serializations all MOTIVATED (audit-delta crash-resilience documented, verify-secure-loop order invariant); (c) behavior — probe orchestrator batched spontaneously without being told "parallel" (N=1), this session fanned 8+7 agents/message during the RED. Conclusion: nothing to add globally — a CLAUDE.md "parallelize" line would duplicate-stack the harness injection (BDR-081 anti-pattern).
ONE real sequential-but-independent candidate: /tour multi-project (independent repos, one by one, no documented reason). User gate: option "tout paralléliser" chosen over report-only-only and no-change, WITH the model invariant "orchestrateur garde le modèle orchestrateur; skills/agents suivent leurs orchestrateurs définis".
Decision: STEP 0 routes (1 project = inline unchanged; ≥2 = STEP 0b fan-out). One general-purpose runner per project, ALL in ONE message, dispatched with NO model override — inherits the session model (model-gate already validated big; a runner carries tour's reflection: fix decisions, convergence). Inside a runner every agent keeps its defined tier (security-auditor sonnet, Phase B opus, doc-syncer sonnet two-mode). Dead/mute runner ⇒ explicit `RUNNER FAILED` summary row (mute is never a pass). Capitalize offer stays MAIN LOOP ONLY (registries = shared state).
LRN-083 derogation, bounded: per-project fix loop + convergence now run INSIDE the dispatched runner. Bounded because nothing a runner decides touches shared state — independent repos, per-repo chore branches, branches stay UNMERGED for human review exactly as inline (report-as-approval-gate design unchanged). Precedent: client-handover-writer already a dispatched orchestrator running parallel audit loops (BDR-077).
Alternatives rejected: report-only-only parallel (my recommendation — user overrode: full parallel wanted); one sub-orchestrator agent .md file (drift risk vs SKILL.md, the runner reads the skill from disk instead — client-handover→/seo precedent); pinning the runner (would put tour reflection on an executor tier — inverts BDR-076); global CLAUDE.md parallelism line (duplicate of harness injection). Census §12: 6 locks (fan-out present, no-pin, single-message, capitalize main-loop, RUNNER FAILED, no pinned runner), flip-tested. Branch feature/tour-parallel, UNMERGED (human gate).
### BDR-085 — user permanent rules: writing-style always-on in rules/, web rules path-scoped [accepted] (2026-08-25)
User supplied 4-block permanent rule text (writing / website / code security / self-check), asked: coverage check, conflict check, integrate. Coverage verdict: security CORE (parameterized queries, input validation, env-var secrets, AuthN/AuthZ split + default deny, no stack traces, fail closed, least privilege) ALREADY in CLAUDE.global.md §Security — NOT duplicated. NEW: entire writing-style block, design anti-default list, public-site done-checklist, web-app specifics (browser-exposed keys, service-key/client split, RLS, server-side auth, IDOR, hashed passwords + cookie flags, field minimization, rate limiting, upload restrictions).
Placement: CLAUDE.global.md at 308/320 (session-start density guard) → no room for ~30 always-on lines. Decision: rules/writing-style.md WITHOUT paths: (always-on load, same session cost, outside the 320 budget) + rules/web-building.md + rules/web-security.md WITH paths: (lazy-load = token win, fire only on web/code files). Project CLAUDE.md doctrine line amended with the budget exception. Feeds C2 self-contradiction audit.
Conflict carve-outs, stated INSIDE the rules: registries keep caveman format (fragments, em-dashes, bullets); code comments keep code style; structured skill/report templates keep their formats; robuste/transformer banned in buzzword sense only (robustness lens, math transform allowed); no-Inter default rule carries "existing brand identities keep their fonts" (ZenQuality deliverables use Inter+Playfair by brand decision — client-handover BDR).
Self-check rule scoped to DELIVERABLES (text, site, feature), not every conversational reply — literal "avant de me rendre quoi que ce soit" would append a compliance note to every chat answer, pure noise. User can re-widen.
Alternatives rejected: compress into CLAUDE.global.md (~11 lines to fit → loses the carve-outs, zero headroom left); path-scope writing-style (applies to conversation, not file reads → would never fire in chat-only sessions); one merged web file (two concerns, one-rule-one-file).
Branch feature/user-writing-web-rules, UNMERGED (human gate).
## BDR-086 — darwin bug-pass scope: verified defects fixed above threshold
- **Date**: 2026-08-26
- **Decision**: units < threshold get full weighted-gap optimization loops (per-unit checkpoint). Judge-VERIFIED defects (file:line, confirmed) in above-threshold units get targeted fixes in a grouped pass — same paired 3-judge validation, one batched checkpoint. User-gated at the scorecard.
- **Why**: leaving a verified destructive path (hotfix `git restore .` wiping tolerated user edits, file scored 85) unfixed = score-worship; rubric serves quality, not the inverse.
- **Alternatives rejected**: strict threshold (ships known bugs); optimize-everything (cost, HL-4 diminishing returns).
- **Reference**: run 2026-08-26, commits 6eceedb..6eac7fb, `.claude/audits/DARWIN-2026-08-26.md`.
## BDR-087 — Stop hook = attention signal only, never control flow
- **Date**: 2026-09-03
- **Decision**: `hooks/notify-attention.sh` wired on BOTH `Notification` (matcher = input-needed set) AND `Stop` (no matcher). One script, branches on `.hook_event_name` when `.message`/`.notification_type` absent → Stop yields "Claude has finished responding". Bell + toast now fire every turn end.
- **Why**: Notification types cover input-needed ONLY. Turn-end had no event; nearest was `idle_prompt`, ~60s late — useless for Remote-SSH user away from screen. User enumerated turn-end as required case.
- **Alternatives rejected**: second dedicated script (duplicates terminalSequence + jq logic, two files to keep in sync); `idle_prompt` alone (60s lag); SubagentStop too (noise, subagent completion not user-visible moment).
- **Guard vs prior refusal**: [[BDR-083]] (unlazy review, GATE 0) REFUSED a Stop hook using `decision:"block"` (forces continuation, inverts human gates). THIS Stop hook returns `terminalSequence` + `suppressOutput` only, exit 0, zero control-flow effect. Signal ≠ control. Do not read the refusal as banning Stop outright.
- **Status**: accepted.
- **Reference**: [[LRN-146]] event-coverage gap, [[BLK-020]] client-side faults, [[LRN-145]] terminalSequence pattern. Verified live: turn-end + AskUserQuestion both ring; `permission_prompt` unexercisable under `defaultMode: auto`.
---
## BDR-088 — gstack Playwright bump shared via lib; update helper never touches submodule tree
- **Date**: 2026-09-15
- **Decision**: `gstack_bump_playwright_if_unsupported` moved out of `install-plugins.sh` into `lib/gstack-playwright.sh`, sourced by install-plugins + update-all + doctor. update-all's submodule block now calls `gstack_submodule_update_with_bump`: re-applies bump after successful `submodule update --remote`; on failure prints git's own message, hints `make plugin`, returns 1. Never touches submodule worktree.
- **Why**: [[BDR-029]] caveat open — bump survived only till next `make plugin`, update path never re-checked OS support. Real gap, user-reported.
- **Alternatives rejected**: conflict-RECOVERY branch (discard package.json+bun.lock → retry → backup/restore). Withdrawn at human gate after 4-agent challenge: concentrated 3 BLOCKER + 4 MAJOR. Worst case = bump discarded, re-apply silently no-ops (bun absent / registry down — bump returns 0 on every path), `./setup` rebuilds browse against unsupported Playwright → [[BLK-008]] returns. Pre-existing behavior just failed the update and kept bump intact, so the "improvement" could regress a working install.
- **Deviations carried from "code MOVED not changed"**: `|| true` on ostag capture (line exited 1 on every non-Ubuntu host → aborted caller under inherited errexit, reproduced); `timeout` on all 3 bun calls, exit 124 → warn + no bump (TERM'd install leaves node_modules half-written, poisons the support grep).
- **Status**: accepted.
- **Reference**: commit 2cebecb, `lib/gstack-playwright.sh`. Links [[BDR-029]], [[LRN-070]], [[LRN-071]], [[LRN-150]], [[BLK-008]].
---
## BDR-089 — No Playwright browser-cache pruner; read-only doctor report instead
- **Date**: 2026-09-15
- **Decision**: `doctor.sh` gains own `── Playwright browsers ──` section — cache size, per-revision the installs requiring it, counts of unreferenced dirs + broken links. Zero deletion anywhere in the lib.
- **Why**: measured, not assumed. `~/.cache/ms-playwright/.links/` registers 3 installs — gstack 1.61.1 → rev 1228, gsd-pi nvm 1.61.0 → 1228, gsd-pi ~/.local 1.63.0 → 1243. Every dir on disk referenced → 0 bytes reclaimable. Playwright's own `_deleteStaleBrowsers` already unions across all registered installs on every `install`.
- **Alternatives rejected**: hand-rolled pruner guarded on "revision resolved by gstack's local playwright" (the originally requested shape) — that guard keeps 1228 and DELETES 1243, breaking gsd-pi. The guard was wrong, not just its implementation.
- **Status**: accepted.
- **Reference**: commit 2cebecb. Links [[LRN-151]], [[BDR-088]].
## BDR-090 — Destructive shell work → autoMode soft_deny/hard_deny; `ask` tier abandoned
- **Date**: 2026-09-15
- **Decision**: 10 rules leave the static tiers (user's own edit): `rsync` `kill -9` `killall` `pkill` out of `deny`; `python3 -c` `python -c` `xargs` `sed` `cp` `mv` out of `ask`. Cover rebuilt in `autoMode` — 7 `soft_deny` (write outside cwd, `rsync --delete`, SIGKILL/kill-by-name, in-place edit spanning >1 file, directory move, inline interpreter or `xargs` that deletes or writes outside cwd) + 3 `hard_deny` (secret exfiltration, prod deploy, disarming guardrails). Intent clears a soft block for the CURRENT TURN only — encoded as a rule line, no setting exists for it. `classifyAllShell` stays false. `permissions.deny` +10 `.env` reader rules (`sed awk cut tr sort uniq diff od xxd strings`), 6 of which sat in `allow`.
- **Why**: `ask` raises no prompt under `defaultMode: auto` ([[LRN-146]], verified live). It gated nothing, so a destructive rule moved deny→ask was a silent loosening dressed as a confirmation. `soft_deny` = the tier the classifier enforces and user intent clears. `hard_deny` = the 3 classes no command pattern can express — read-then-send spans turns, a prod target is a name not a verb, widening a deny list is self-disarming.
- **Alternatives rejected**: keep them in `ask` — inert, false sense of a gate. Back to `deny` — blocks legit process cleanup and inter-project copy, and the user works Bash-first under auto mode. `classifyAllShell: true` — closes the allow-tier blind spot but bills a classifier call on every `git status`. Published-history rewrite as `hard_deny` — user declined; `rebase` then an ordinary push stays uncovered, known gap.
- **Scope fix (same commit)**: `autoMode.environment` named `/home/bchanot/Documents/atlast`, its FTP deploy target and its customer data, inside the file `link.sh:21` symlinks to `~/.claude/settings.json`. Every project received atlast's facts, and this repo's own Gitea remote contradicted the block's "no remote configured". Global block now machine-generic; atlast facts moved to atlast's gitignored `.claude/settings.local.json`.
- **Caveat**: the guardrail `hard_deny` bars REMOVING a `deny`/`soft_deny`/`hard_deny` entry, not adding one. Future loosening goes through `/permissions` or the user's own edit — deliberate, confirmed with the user.
- **Status**: accepted.
- **Reference**: `settings.json`, `doctor.sh` `check_automode`, `templates/settings/SETTINGS.md`. Links [[LRN-153]], [[LRN-146]], [[BDR-004]].
## BDR-091 — Ask, don't guess: open-choice sweep + mid-run channel supersede "one question upfront"
- **Date**: 2026-09-16
- **Decision**: `CLAUDE.global.md` rule → ask on a VISIBLE (placement, wording, order, behavior), PUBLIC NAME (command, flag, endpoint, file) or SCOPE ("X too?") choice the request leaves open, even mid-task; class 4 (internal technical, no observable effect) never. `lib/contract-interview.md` STEP 2 = CLARIFY: pass A (gaps: outcome / scope / constraints) at contract time; pass B (open-choice sweep, 3 classes) ONCE at each flow's PLAN step, no question cap, >5 open → under-specified, list + stop; "you decide" recorded `A: delegated — <default>`, never re-asked. New MID-RUN CLARIFICATION: executor halts `NEED-DECISION` + `CLASS:` tag; visible / public-name / scope → human verbatim; internal → loop decides, max 2 round-trips. New HOW TO ASK (LRN-102: ≤4 → one AskUserQuestion, context in option descriptions; else plain text ending the turn). Wiring: feat STEP 1, bugfix STEP 3, hotfix LOCATE (pass A stays silent autofill; one re-dispatch on class-tagged BLOCKED = the sole hotfix re-dispatch), ship-feature STEP 2, init-project STEP 3; interviewer: class 1-3 item never `(assumed)`, one extra targeted question. Executors (feater, bugfixer, hotfixer) report the class. Locks: contract-verifier (9), loops-light (hotfix), gates (3).
- **Why**: gap-only trigger structurally blind to taste — "add a share icon" passes outcome / scope / constraints and the icon lands wherever the executor put it; raising the 3-question cap changes nothing. feat:153 / bugfix:165 told the orchestrator "make the decision HERE", twice, before escalating = institutional guessing. Fresh re-dispatch keeps the tree, loses the executor's reasoning → a plan-time batch costs less than the same question mid-run; the mid-run channel stays for leftovers.
- **Alternatives rejected**: bigger budget (quota was never the limiter); new `lib/clarify.md` (extra hop, STEP 2 already the mandatory passage every orchestrator runs); global rule only (skills carried explicit counter-instructions — `zero questions ever`, `make the decision HERE`, `max 3 questions` — the specific beats the general).
- **Risk watched**: chattiness. Brakes = class 4 exclusion + over-5 guard. hotfix identity (speed, silence) = the flow to watch; if pass B fires on most hotfixes the class definitions are too wide, not the flow.
- **Status**: accepted. Behavioral check OPEN: `/feat "add a share icon to the header"` must ask placement before dispatch; the fully specified variant must ask nothing → record in `evals.md`.
- **Reference**: spec + plan `docs/superpowers/{specs,plans}/2026-09-16-ask-dont-guess*` (purged at finish, in history at `22ce57f`), commits `9eb6934..22ce57f`, merge `56bd035`. Supersedes the `CLAUDE.global.md:51` rule line. Refines [[BDR-049]] (contract), applies [[LRN-102]]. Links [[LRN-157]].
## BDR-092 — docker + node framed by the classifier (`autoMode.allow`), `ask` rules retired
- **Date**: 2026-09-16
- **Decision**: `Bash(docker run|exec *)`, `Bash(docker[-| ]compose up*)`, `Bash(node -e *)` out of `permissions.ask`. New `autoMode.allow` (`$defaults` first): (1) local dev containers — `docker exec/run/compose` against a workstation container whose name lacks `prod`/`production`, running a repo SQL file or script inside, output piped; Remote Shell Writes / Production Reads / Sensitive Remote Exec scoped to sensitive-named hosts; a literal `DROP/TRUNCATE/DELETE` on the command line stays under Mass Delete. (2) project-local node — `node <file>`, `npm run`/`pnpm`/`yarn` scripts, `npx`/`pnpm exec` of a lockfile-declared package, effects in cwd. +2 `soft_deny`: docker data destruction (`rm -f`, `volume rm/prune`, `system prune`, `compose down -v`, `--privileged`, bind mount outside cwd); undeclared node packages (`npx`/`dlx` absent from lockfile, `npm install <name>`). `model` bump to fable 5.1 committed alongside.
- **Why**: real gate = built-in `Remote Shell Writes` / `Production Reads` soft_deny catching `docker exec` into `supabase_db_game`; inside the classifier `allow` = exception tier (hard_deny > soft_deny > allow > explicit intent). Static `Bash(node *)` allow is suspended under auto (wildcarded interpreter) → prose is the only lever for a conditional node permission; `awk`/`echo` statics short-circuit, `node` cannot. `ask` inert on 2.1.273 (probe, [[LRN-155]]) — retiring it is forward-safe: if the documented prompt behavior lands, those entries would prompt for exactly what should run free.
- **Alternatives rejected**: static `permissions.allow` for docker (short-circuits the classifier, framing impossible); keep the `ask` entries (inert today, wrong tomorrow); strict on every `.sql` (blocks the repo's verify scripts).
- **Trade-off accepted**: a repo SQL file runs even when its content is opaque to the classifier — local dev DB only, resettable.
- **Guardrail**: S6 (loosening = user's own edit) overridden explicitly by the user for this change; diff reviewed on the branch before merge.
- **Status**: accepted. Verified: `jq` valid; `claude auto-mode config` shows the 4 entries with `$defaults` expanded; `doctor.sh` autoMode PASS; live `docker exec -i supabase_db_game psql … -f - < verify/0043 … | tail` → `ROLLBACK`, exit 0, no prompt. `claude auto-mode critique` printed nothing (2.1.273).
- **Reference**: `settings.json`, `templates/settings/SETTINGS.md` (`autoMode.allow` row + interpreter note), commit `5eccc3f`, merge `ddadca6`. Links [[BDR-090]], [[LRN-153]], [[LRN-155]], [[LRN-156]].
## BDR-093 — 21st.dev: magic MCP retired for the `21st` CLI + skill pack
- **Date**: 2026-09-22
- **Decision**: `@21st-dev/magic` MCP out, `@21st-dev/cli` (bin `21st`) in — upstream supersedes it (README 1.17.1: "one unified CLI", old magic config now a thin proxy to the same endpoint). Auth = `21st login`, browser token in `~/.config/21st`; no API key, no MCP process. `install-plugins.sh` Step 8.7: `npm i -g` (pin `21st` in plugins.lock.json) + staged `21st skills install --global --agent claude` + TTY-only login offer + pack disabled by default. `toggle-external.sh` manages `21st` as a pack (names globbed from `skills-external/21st-*`, parked under plain names = interoperable with profile.sh's external path). 5 design skills (`21st-ui-build`, `-ui-explore`, `-ui-review`, `-cli-use`, `-ai`) in design/web/web-full/full + `MANAGED_EXTERNALS`; `-registry`/`-design-sync` installed, parked. `MANAGED_MCPS` now empty (mcp type kept, advisory). Gate: `GATE-BLOCK: 21st 21st-ui-build` — CLI = required-manual class, `magic`'s old slot. Outward-facing verbs (`publish*`, `submit`, `edit`, `delete`, `remove-from-catalog`, `profile set|upload`) → one `autoMode.soft_deny` entry, NOT `ask` ([[LRN-153]]).
- **Why**: user ask ("plus besoin de mcp / api, juste en cli"), confirmed at the source, not the marketing page — the 21st.dev web docs still show the MCP `init --client` flow and an API key; the npm package README is what states the supersession. Net wins: one less MCP loaded per session, no API key to protect by reference ([[BDR-026]]/[[BDR-057]] vector gone), no unauthenticated local callback server ([[LRN-110]] gone with the tool).
- **Blocker + shape it forced**: `21st skills install --global` writes `<HOME>/.claude/skills/<n>/SKILL.md` and calls `assertNoSymlinkComponents` on every path segment — `~/.claude/skills` IS a symlink to `repo/skills`, so the documented `21st install-skill` fails hard ("Refusing to access symbolic link …/.claude/skills", reproduced live). → install under `mktemp -d` as HOME, move each skill into `skills-external/21st-*` (gitignored), symlink on demand. Same impeccable/ctx7 machine-owned pattern.
- **Alternatives rejected**: project-scope install (`<cwd>/.claude/skills`) — Claude Code would ALSO scan it as project skills in this repo = every 21st skill listed twice; add the pack to `link.sh`'s `EXTERNAL_SKILLS` — that loop force-creates symlinks, resurrecting a default-disabled pack on every `make link`; all 7 skills in the design profiles — publishing flows cost 2 descriptions/session for a workflow the user does not run; `ask` entries for the publish verbs — inert under auto ([[LRN-153]], [[BDR-090]]/[[BDR-092]] already retired that tier).
- **Status**: accepted. Verified: toggle enable/disable/restore/idempotent round-trip (7 skills), `profile.sh show design`, gate INCOMPLETE→names `21st` with the two commands, gate PATH repair proven under `env -i PATH=/usr/bin:/bin` with an nvm-stub, `profile-set-managed.test.sh` 17/17, `make test` (2 pre-existing FAILs, gitleaks binary absent on this host), shellcheck clean. OPEN for the user: `npm i -g @21st-dev/cli && 21st login` — the deny rule `Bash(npm install -g *)` means the agent cannot run it.
- **Reference**: `install-plugins.sh` Step 8.7, `update-all.sh` 7.4, `lib/toggle-external.sh`, `lib/profile.sh`, `lib/profiles/*.profile`, `lib/design-tool-gate.sh`, `lib/design-gate.md`, `CLAUDE.global.md`, `README.md`, `settings.json`, `.gitignore`, `.gitleaks.toml`, `.env.example`, `link.sh`. Supersedes the operative parts of [[BDR-059]] (the 4 `mcp__magic__*` ask entries) and the magic instance of [[BDR-026]]/[[BDR-057]]; [[BDR-025]]'s required-manual class stands, its magic example does not. Links [[LRN-158]], [[LRN-110]].
## BDR-094 — impeccable: global-scope install through the repo symlinks, pin + @latest fallback, output-read failure check
- **Date**: 2026-09-22
- **Decision**: `install-plugins.sh` Step 8d + `update-all.sh` run `npx -y impeccable@<pin> skills install -y --providers=claude --scope=global --no-hooks` straight through the `~/.claude/{skills,agents}` symlinks → lands in `skills/impeccable` + `agents/impeccable-*.md` (both gitignored, machine-owned). No staging, no `mv`. Precondition guard: both symlinks must already point into the repo, else "run make link first". Pin failure → `@latest` + loud "bump plugins.lock.json" warn. Profile-parked copy stays parked (install to live slot, `mv` back to `skills-disabled/`). Success = rc 0 AND installer output free of `Download failed|Could not check for skill updates` (`imp_install`, mirrored in both scripts, sets `IMP_FAIL`). Pin 3.2.0 → 4.1.0 (CLI only; skill dist 4.3.1 + engine 0.1.5 own tracks). `link.sh` `EXTERNAL_SKILLS` drops impeccable; `skills-external/impeccable/` gone. `lib/design-gate.md` §5: suggest `/impeccable init` once when frontend project lacks `PRODUCT.md`.
- **Why**: (1) 3.2.0 skill dist gone upstream → rc 1 → `make plugin` printed "run manually" forever. (2) `--scope=project` + staged `mv` moved skill dir only, dropped the 4 subagents the same run wrote. (3) Global scope IS the repo install under the symlink model; staging bought nothing. (4) rc lies once a copy exists. Probe 2026-09-22, sandbox HOME, real installer: 4.1.0 then 3.2.0 → rc 0, "Could not check for skill updates: invalid zip data … Existing skills were left unchanged"; same-pin rerun → rc 0, "Skills are up to date (v4.3.1)"; both leave SKILL.md byte-identical (same mtime, same sha).
- **Alternatives rejected**: before/after skill-version compare (first idea) → cannot separate rotted-pin no-op from up-to-date no-op, identical files + rc 0 both times → false warn on every rerun. `--force` → re-downloads ~15 MB engine + dist on every `make plugin`, and the CLI's own update check already refreshes without it. Shared `lib/impeccable.sh` for `imp_install` → deferred: two mirrored 12-line helpers vs new lib + test; revisit at a third caller. Project-scope install inside this repo → Claude Code scans `.claude/skills` too = skill listed twice, shadows the global copy (seen live, TODO T6).
- **Status**: accepted. Verified: harness on extracted Step 8d, sandbox HOME, real installer, 4/4: fresh install (skill 4.3.1, 4 agents); rotted pin over a copy → fallback fires; same pin rerun → no false warn; parked + rotted → fallback, returned to `skills-disabled/`. `make test` green minus 2 pre-existing T16a (gitleaks absent), shellcheck clean. `update-all.sh` block: `bash -n` + shellcheck only, same helper, not run end to end.
- **Reference**: `install-plugins.sh` Step 8d, `update-all.sh`, `plugins.lock.json`, `.gitignore`, `link.sh`, `lib/design-gate.md` §5. Links [[LRN-159]], [[LRN-158]] (21st: opposite case, installer refuses symlinks → stage), [[LRN-077]] (pin doctrine), [[BLK-021]].
+28
View File
@@ -37,6 +37,9 @@ rules:
| EVAL-018 | 2026-07-06 | job3 docs-drift audit + execution: 46/46 findings verified, 20/23 fixes shipped (B1 blocked, D2-D5+B6 skipped by decision), zero residual on re-sweep | keep |
| EVAL-019 | 2026-07-06 | job4 test-gap audit + execution: 11 specs + 5 fixes/seams, every mutation red-green verified, zero residual | keep |
| EVAL-025 | 2026-07-17 | opening seo/geo inventory (subagents): 7/7 verifiable claims false or overstated; real contact corrected all, 6 plan corrections + 4 features killed at measurement | keep |
| EVAL-027 | 2026-08-24 | contract-gates behavioral RED: 16/16 fresh unprimed runs followed new doctrine (GATE 0 order, vacuous oracle, ABANDONED routing, scope temptation resisted) | keep |
| EVAL-028 | 2026-08-26 | darwin v2.1 paired run 54 units: 60 paired verdicts 0 revert/tie; skeptics found 3 real residuals — engaged, not rubber-stamp | keep |
| EVAL-029 | 2026-09-15 | 4-agent plan challenge: 6 BLOCKER; 3 of 3 confirmation-pass BLOCKERs came from the fixes themselves; caught a false 654 MB orphan claim | keep |
---
@@ -251,3 +254,28 @@ rules:
### EVAL-026 — 3-way plan challenge caught 4 BLOCKERs dogfooding own plan (2026-07-17)
Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itself. Verdicts CONCERNS(4)/FATAL(6)/FATAL(4). Caught 4 distinct BLOCKERs a single pass would blend: (1) v1 unbuildable — targeted init-project (inline-load, no dispatch) + false "plan on disk" premise for feat/bugfix (only contract persists); (2) failed-open silently dropping a lens while claiming "challenged" (inverts verify-secure-loop "a mute verifier is NEVER a PASS"); (3) consensus-weighting buries lone L2 security finding (lenses orthogonal); (4) sonnet challengers violate [[BDR-066]] (audit judgment=big model). Synthesis REJECTED 1 false positive (allowed-tools-blocks-dispatch — ship-feature has same frontmatter + dispatches fine). Each lens found a DIFFERENT class of flaw → evidence 3-independent > 1-multilens. Action: hardened v2 (severity-driven + fail-safe + re-think loop) shipped. Method validated itself before build.
### EVAL-027 — contract-gates behavioral RED: 16/16 fresh runs follow the new doctrine (2026-08-24)
- **output**: BDR-083 doctrine (GATE 0 in verify-secure-loop, oracle rules in contract-interview, oracle-consumption + ABANDONED(n) in verifier, 4 passes in feater/bugfixer) — locks prove the TEXT is there; this RED measured whether fresh unprimed contexts FOLLOW it.
- **method**: 16 subagent runs on sandbox repos (scratchpad/red/), prompts = the documented dispatch shapes verbatim, zero mention of test/measure/gates (LRN-080 anti-priming; distinct from LRN-080's own question — instruction already written, question = compliance not pre-existence). Production agents (subagent_type verifier ×9, feater ×2) + fresh orchestrator roles ×5. Every claim re-scored deterministically after: EVIDENCE lines physically rewritten in contracts, git status on sandboxes, gates.sh parse of authored contracts.
- **verdict**: 16/16 conformant. v1 red-oracle-wins 3/3 (NOT-MET citing evidence, own re-run). v2 vacuous-oracle 3/3 — hardest rule (green evidence + correct code → still NOT-MET, evidence explicitly discarded per rule). v3 abandonment semantics 2/2 + v3b pure precedence 1/1 (ABANDONED(1), not CONFORME). o-red 2/2 (gates.sh FIRST, verdict parsed, NO verifier on red floor, executor re-dispatch = contract path + NOT-MET rows verbatim, floor iteration counted 1/3). o-green 1/1 (floor → verifier dispatch with CONTRACT+DIFF+TEST only). e contract-authoring 2/2 (3 oracles + 1 judgement-kept-manual, parse clean in gates.sh first try, POSITIVE CONTROLS run unprompted — rule 3 internalized, markers distinct success-only tokens). f feater 2/2 (out-of-scope temptation src/util.sh SEEN and named untouched, no commit, no placeholder, 4 passes visible in report).
- **anomalies**: none against doctrine. Fixture flaw (mine): placeholder.txt trick used to fabricate a 2nd commit made v2/v3 diffs contain no feature work — every verifier CAUGHT it (out-of-scope + "implementation pre-exists base commit"), polluting v3's intended pure-ABANDONED measurement → v3b clean fixture added. Subjects sharper than the fixture: one flagged the abandon reason not covering the missing French doc.
- **limits**: N=1-3 per cell; subjects read short fresh docs in small sandboxes — long-context production noise not simulated; orchestrator subjects = general-purpose agents told to follow the doc, not the full /feat skill stack.
- **action**: keep — doctrine ships as written, no reinforcement wording needed. Artifacts: scratchpad/red/ (session-lived, not committed).
## EVAL-028 — darwin v2.1 paired run, 54 units
- **Date**: 2026-08-26. **Output**: 12 optimization rounds (13 sub-80 units) + 8 bug-fix commits, all kept.
- **Method**: paired same-judge 3-majority per round (v2.1); judges live-exec where artifact executable (5 units: skills-perso, profile, plugin-pair, status-reporter, gitflow). Absolute scores triage-only. Totals main-thread (LRN-018 applied).
- **Anomalies**: (1) 0 reverts/ties in 60 verdicts — homogeneous-better checked: skeptic lens found real residuals 3x (doctor.sh cost source, hotfix RULES leftover restore, FILE(S) new-marker) → judges engaged. (2) census lock RED on line-rewrap, make test caught → LRN-144. (3) head-pipe masked grep exit 2x → LRN-143.
- **Action**: v2.1 paired = standard. Post-run absolute rescore skipped by design (would be judge-noise theater).
---
## EVAL-029 — 4-agent plan challenge: 6 BLOCKERs, and the fix round produced 3 of them
- **Date**: 2026-09-15
- **Method**: 3 blind lenses (correctness / robustness / simplicity) on plan rev 1, then 1 confirmation lens on rev 2. Subject = the gstack Playwright lib plan ([[BDR-088]]).
- **Result**: rev 1 → 3 BLOCKER + 12 MAJOR. Rev 2, written specifically to close them → 3 NEW BLOCKERs, and 2 of the 3 were INTRODUCED BY the fixes: the new "every public function returns 0" rule contradicted the new "return rc", and the printer-name clause came verbatim from my own contract criterion 9. Rev 3 dropped the recovery branch entirely at the human gate — 6 findings closed by deletion instead of code.
- **Anomaly**: my first user-facing answer asserted ~654 MB of orphan Playwright revisions. FALSE — `.links` showed every dir referenced, 0 reclaimable. Caught only while designing the guard, not while asserting the number. Worse, the guard I proposed would itself have deleted gsd-pi's rev 1243.
- **Action**: (1) never state a disk-reclaimable figure before reading the registry that owns it ([[LRN-151]]). (2) A fix round deserves the same challenge as the original plan — 3/3 confirmation BLOCKERs came from fixes, not from the original. (3) The confirmation pass earned its cost: without it the printer override would have shipped and silently disconnected doctor's counters ([[LRN-150]]).
- **Status**: keep.
- **Reference**: `.claude/tasks/plans/2026-09-13-gstack-playwright-lib-2220.md` (rev 3). Links [[BDR-088]], [[LRN-150]].
+58
View File
@@ -427,3 +427,61 @@ rules:
## 2026-07-22
- User: auto-gitignore+delete transient pipeline artifacts in all projects. Investigation reframed the ask — gitignore = WRONG tool (files read from disk during run; would break superpowers SDD `git add` of spec). BDR-065 already rejected gitignore + its DELETE side was doctrine-only (no code, manual chore slipped once — 655e364). User picks (2 recommended): keep committed-during-run + AUTOMATE delete; keep `.claude/tasks/{contracts,plans}` versioned.
- Built `lib/gitflow.sh` `_gitflow_purge_transient` at finish (feature/bugfix, pre-merge, best-effort never-abort, opt-out `GITFLOW_PURGE_TRANSIENT=0`) + `purge-transient` CLI verb. Universal via `~/.claude/lib`→repo symlink. gitflow-test T17 a-d (10 checks, `--full-history` recovery), shellcheck clean, make test exit 0. BDR-065 amendment + [[LRN-138]]. feature/gitflow-auto-purge-transient.
## 2026-07-30
- User: Opus 5 "needs more freedom" → analyse config + adapt. Research 3-agent (registries / config audit / web) + official migration guide: over-delegation (inverts LRN-030), over-verification, literal following, scope expansion, #80988 injections. Plan challenged 3 blind Opus 5 plan-challengers — robustness FATAL (BLOCKER: symlink-live deployment), all fixes adopted. Shipped: CLAUDE.global.md recalibrated (delegation when-guidance, staff-bar dropped, finish-whole-task, deliverable-length; 308/320), design hook \bux\b dropped flip-tested (22/0), plan-challenger grounded-doubt→[MINOR] (44/0). BDR-081 + LRN-139. feature/opus5-config-tuning, UNMERGED.
## 2026-08-02
- C1 seo/geo de-prescription EXECUTED end-to-end: census-first 71 locks flip-proven → reword under audience×range invariant (adafa35/c7646a9) → controlled dogfood (judge-replay frozen signals + templates + fresh collects + e2e + blind reader) → 42/42 both sets, zero contract regression, recall improved. Plan survived 4 challenge passes (2 FATAL + confirmation FATAL(9), all closed by name). BDR-082 + LRN-140. Nested-CLI dogfood died on monthly spend limit → inline pipeline (canonical /seo shape). feature/seo-geo-deprescription UNMERGED (human gate). Chantiers C2-C4 pending.
## 2026-08-24
- Analysed `unlazy` skill (Leonxlnx/unlazy 2.1.0) on user request. Its verification architecture teaches us nothing — contract + fresh blind verifier + bounded loops already shipped. Real gap: no deterministic floor between executor and GATE 1 (the verifier's `PROOF:` is a line it writes, not a process exit).
- Shipped Palier 2 (user-chosen): lib/gates.sh + GATE 0 + oracle-bearing criteria + `ABANDONED(n)` verdict + 4-pass executors. Refused unlazy's Stop hook, approval store, .unlazy/ tree, tree-N arithmetic, Node checker — [[BDR-083]] records each why.
- `make test` rc 0, shellcheck clean, 64 new assertions, e2e on a real contract. Branch feature/contract-gates UNMERGED (human gate).
- Locks caught a reflow regression (5 red on rewrapped phrases, zero doctrine lost) → [[LRN-142]]. Skill-adoption pattern → [[LRN-141]].
- Parallelism audit (user ask "est-ce actif ?"): measured, not assumed — nested probe proves concurrent fan-out (9.1s vs 18s), doctrine already prescribed everywhere safe, remaining serializations motivated. One candidate found: /tour multi-project → parallel runners shipped ([[BDR-084]], user gate "tout paralléliser" + model invariant). Branch feature/tour-parallel UNMERGED.
## 2026-08-25
- User permanent rules integrated: rules/writing-style.md (always-on) + web-building.md + web-security.md (path-scoped). Security core already in §Security, not duplicated. Carve-outs protect caveman registries + skill templates + brand fonts. [[BDR-085]]. Branch feature/user-writing-web-rules UNMERGED (human gate).
## 2026-08-26 — darwin fresh baseline + threshold run (feature/darwin-optimize-20260825, UNMERGED)
- `/darwin-skill all skills and agents` (background). Fresh results.tsv (May file wiped). 7 blind judges, 54 rows (31 skill-systems + 23 agents), mean 83.4, 13 <80. find-docs excluded — machine-owned ctx7 (gitignored), 3rd exclusion ground after BDR-015/058.
- Phase 2: 12 rounds / 13 units, 0 reverts, all paired 3-0 ([[EVAL-028]]). Star: skills-perso detection 8/31 → 31/31 live-verified. Bug pass [[BDR-086]]: 8 commits in above-80 units kept 3-0 (hotfix git-restore data-loss path ★, onboarder contract bounce, plugin data-flow, plan-challenger grammar, handover stale §refs + gate order, tour report-only commit, harden severity, fixtures).
- make test green after census-rewrap fix ([[LRN-144]]); [[LRN-143]] head-pipe grep mask. 29 commits, report `.claude/audits/DARWIN-2026-08-26.md` + card PNG. Branch awaits human review + merge.
## 2026-09-01
- Attention signal shipped: hooks/notify-attention.sh + Notification entry in settings.json (bell x2 + OSC 777 toast via terminalSequence). Client-side VS Code steps pending: terminalBell sound:on + osc-notifier ext. [[LRN-145]]. Branch chore/notify-attention-hook, UNMERGED.
- Pre-existing model switch opus[1m] committed separately on same branch.
## 2026-09-03
- Attention signal completed + verified end-to-end. Two client faults isolated ([[BLK-020]] resolved): ext instruments only terminals born AFTER activation (re-attach via `dtach -a`, no session loss); Code app volume 0 in Windows mixer killed bell while Windows-emitted toast sound masked it.
- Coverage gap found + closed: `Notification` matcher covers input-needed only, turn-end had no event. `Stop` wired on same script, branches on `.hook_event_name` ([[BDR-087]], [[LRN-146]]). Verified live: turn-end + AskUserQuestion ring; `permission_prompt` unexercisable under `defaultMode: auto`.
- BDR-087 + LRN-146 + BLK-020 capitalized. Branch feature/notify-stop-event, merged to develop (f90ee74).
- Post-merge regression: toast dead again after re-attach from a RESTORED terminal, bell fine. Root cause [[LRN-147]]: ext hooks only terminals born after its activation; `enablePersistentSessions` restores terminals before it. Fix = disable persistent sessions, or fresh terminal + `dtach -a`. Verified: 3/3 toasts on fresh pty.
- Same-day counter-example broke that cause: second session's terminal deaf though created LATER, same window, ext global, shells identical. Trigger unknown; [[LRN-148]] adds the 5s pre-flight test + demotes LRN-147's mechanism claim.
- Attention signal refined: per-event labels (BDR-087 follow-on), silence on non-attention events, and no turn-end signal while `background_tasks` non-empty ([[LRN-149]]). Payload dump beat the docs: `background_tasks` undocumented for Stop but present on the wire. Branch bugfix/notify-subagent-spawn.
- gstack Playwright: bump extracted to `lib/gstack-playwright.sh`, now re-applied after a successful submodule update ([[BDR-088]]); read-only browsers report in doctor, no pruner — `.links` proved 0 bytes reclaimable and the guard I first proposed would have deleted gsd-pi's rev 1243 ([[BDR-089]], [[LRN-151]]). 4 challengers → 6 BLOCKER, recovery branch withdrawn at the gate ([[EVAL-029]]). 2cebecb on feature/gstack-playwright-lib.
- Node checked against Playwright: already v24 (1.61 needs >=18, 1.63 needs >=20), not the macOS constraint. macOS audit deferred to its own cycle — found statically: `sed -i` with no suffix x3 in install-plugins.sh (BSD sed eats the next arg), `${x,,}` in url-guard.sh (bash 4+, macOS ships 3.2), `readlink -f` in doctor.sh (absent pre-Monterey 12.3).
## 2026-09-15
- Aligned repo config + deployment on the user's hand-edited `settings.json`. Destructive shell work rebuilt in `autoMode` soft_deny/hard_deny once `ask` was established as inert under auto mode ([[BDR-090]]); `permissions.deny` +10 `.env` reader rules, 6 of which sat in `allow`.
- `autoMode.environment` was scoped to ANOTHER project inside the user-scope file, so every repo got atlast's facts. Rewritten machine-generic, atlast facts moved to atlast's own `settings.local.json`, `$defaults` added to all three lists ([[LRN-153]]).
- `doctor.sh` gained `check_automode` (missing `$defaults`, foreign-repo scope, both arms tested). `SETTINGS.md` documents the block + a tier-choice table. README's magic-MCP "ask = live confirmation" claim corrected — false under `defaultMode: auto`.
- Found, not fixed: `.claude/settings.local.json` = 14.6 KB shadow copy of the global settings at HIGHER precedence, incl. a `config-protection.sh` hook whose script does not exist. Logged F1-F3 in TODO.
- `make test` 0 RED, `doctor.sh` 0 errors, `shellcheck` clean.
- graphify skill untracked + gitignored (written by `graphify install --platform claude` since `~/.claude/skills` symlinks to `skills/`). Cost one self-inflicted incident: `git rm --cached` kept the files, `gitflow finish` deleted them at the merge ([[LRN-154]]). Restored at 0.9.61, guarded configs snapshotted and verified untouched.
- `.claude/settings.local.json` 14.6 KB -> 6.2 KB. It was not just duplication: its local `deny` still carried the 4 rules moved out of global deny, making [[BDR-090]]'s soft_deny a dead letter in this repo, and its `allow` carried `sed *` / `cp *` / `python3 -`, which short-circuit the classifier on the same rules.
## 2026-09-16
- Ask, don't guess ([[BDR-091]]): spec + plan, 9 lock-first tasks (contract-interview CLARIFY two passes, MID-RUN CLARIFICATION with `CLASS:` tag, HOW TO ASK; global rule; feat / bugfix / hotfix / ship-feature / init-project wired; interviewer; 3 executors), suite green. Behavioral fixture check still open ([[LRN-157]]).
- docker + node under auto mode ([[BDR-092]]): `ask` entries retired (inert on 2.1.273, probe — [[LRN-155]]), `autoMode.allow` + 2 soft_deny, live `docker exec … psql` OK. Static interpreter allow is suspended under auto → prose only ([[LRN-156]]).
- Both merged into develop 2026-09-17 via gitflow (`ddadca6`, `56bd035`), two stack conflicts (TODO, CHANGELOG) resolved keeping both blocks. Symlinked `settings.json` follows the checkout: live config = whatever branch is out.
## 2026-09-22
- 21st.dev magic MCP → `@21st-dev/cli` + 7-skill pack, user ask. Install/update/toggle/profiles/gate/docs/permissions migrated on `feature/21st-cli-migration`.
- Blocker: documented `21st install-skill` refuses the `~/.claude/skills` symlink → staged install under a throwaway HOME (LRN-158).
- Gate: `magic`+MAGIC_API_KEY required-manual slot → the `21st` CLI; publish verbs moved to `autoMode.soft_deny` (ask inert under auto).
- BDR-093, LRN-158. `make test` green except 2 pre-existing gitflow FAILs (gitleaks binary absent on this host). Branch UNMERGED — human gate.
- impeccable install repaired ([[BDR-094]]): global scope through the symlinks, 4 agents kept, pin 3.2.0 → 4.1.0 with @latest fallback, design-gate §5 `/impeccable init` hint. Residue probed: rotted pin over an existing copy exits 0 → `imp_install` reads the installer output ([[LRN-159]]); before/after version compare rejected (identical no-op). Harness 4/4, sandbox HOME, real installer.
- Previous shell death traced: /tmp tmpfs usrquota blown by 5.9 GB of dead-session probe HOMEs ([[BLK-021]], open, user frees). Tests + harness ran with TMPDIR under ~/.cache. `make test` green minus 2 pre-existing T16a, shellcheck clean. Committed on feature/21st-cli-migration, UNMERGED. `skills/synced/` (claude.ai synced skills, 4.4 MB) untracked + unignored, left for the user.
+151
View File
@@ -138,6 +138,16 @@ rules:
| LRN-133 | 2026-07-17 | an omission must stay LEGIBLE, never silent — tool that can't measure says so in its output | designing any audit/measure output; deciding what a cap/refusal/N-A emits |
| LRN-134 | 2026-07-17 | resolve-then-pin in stdlib http.client beats monkeypatching getaddrinfo — dual-stack, thread-safe, no requests; classify the OS-resolved IP not the URL text | closing SSRF/DNS-rebinding on any Python HTTP egress |
| LRN-135 | 2026-07-17 | a prefix-only scan for a dangerous construct is bypassable by padding — scan the WHOLE document | refusing any hostile construct (DTD/directive/marker) before parse |
| LRN-143 | 2026-08-26 | `cmd \| head \|\| fallback` — pipeline rc is head's (0), fallback dead; bounded output → drop head, else pipefail | any probe/fallback bash in skills before trusting `\|\|` |
| LRN-150 | 2026-09-15 | Sourced lib shares caller shell: bare `ok/warn/info` override its printers, and its `set -e` applies inside | any new lib/*.sh |
| LRN-151 | 2026-09-15 | Playwright cache truth = union over `.links`, dir name maps `_`→`-`, revisionOverrides exist | shared versioned binary caches |
| LRN-152 | 2026-09-15 | git `protocol.file=user` blocks submodule fixtures; `-c` misses the code under test, `GIT_CONFIG_*` env does not | tests building git fixtures |
| LRN-153 | 2026-09-15 | `autoMode` lists replace built-ins without `"$defaults"`; a user-scope block reaches every project | any `autoMode` edit |
| LRN-154 | 2026-09-15 | `git rm --cached` + merge into a branch that still tracks the file DELETES it from disk | untracking a generated file |
| LRN-155 | 2026-09-16 | ask under auto: doc says prompt, probe on 2.1.273 says no; re-probe after upgrades | any permission-tier reasoning |
| LRN-156 | 2026-09-16 | autoMode.allow = exception tier; static interpreter allow suspended under auto → conditions live in prose | conditional permissions |
| LRN-157 | 2026-09-16 | gap-only trigger blind to taste → add a trigger class, not budget; ask at plan, mid-run for leftovers | any "ask more" request |
| LRN-158 | 2026-09-22 | Installer refusing symlinked paths vs a symlinked config dir → stage under a throwaway HOME, move the result | any vendor installer writing into ~/.claude or ~/.config |
---
@@ -1355,3 +1365,144 @@ rules:
- **context**: user asked to gitignore transient planning artifacts (`docs/superpowers/{specs,plans}`, `.claude/tasks/{contracts,plans}`) to stop them merging. BDR-065 had already REJECTED gitignore for docs/superpowers on the git-travel ground; the real gap was the DELETE side never being coded (doctrine-only manual chore, slipped once — 655e364). Built `_gitflow_purge_transient`.
- **future application**: "don't merge transient X" → ask: does the run read X from disk? does X travel via git (worktree, foreign checkout)? Yes → auto-purge at finish, not gitignore. Scoped commit `-- <paths>` avoids sweeping a dirty index; `git diff --quiet HEAD -- paths` precheck makes `git rm` all-or-nothing safe; keep the purge best-effort so cleanup NEVER blocks a merge. Prove archive-reachability with `git log --full-history` / `git show <sha>:path` — plain `git log -- path` prunes the purged add-commit via history simplification (bit me writing T17).
- **link**: [[BDR-065]].
## LRN-139 — model-trait compensations invert across generations; state WHEN-guidance, not direction (2026-07-30)
- **pattern**: config rules that COMPENSATE a model trait become counter-productive when the next generation inverts the trait. LRN-030 (Opus 4.8 under-delegates → "Default to delegation… counters under-delegation") inverted by Opus 5 (delegates MORE readily, official guide) — the rule pushed the failure the model now has. Same class: explicit verify instructions → over-verification; conservative-reporting clauses → literal recall suppression; MUST/CRITICAL → over-triggering.
- **Opus 5 traps found**: (a) Claude Code injects Opus-5-only anti-delegation prompt sections (heron_brook + subagent_steer_delegation, issue #80988; server-gated, no opt-out, absent from transcripts) — own prose stacks on top blindly; (b) NO model-default effort hold on Opus 5 — persisted effortLevel (xhigh, settings.json) silently carries over, against "start high, sweep low/medium"; run /effort sweep per model; (c) effort does NOT shorten visible output/deliverables — only prose length rules do (+30-40% docs).
- **future application**: at every model-generation bump, grep config for trait-compensating language ("counters model tendency…", "default to X") and re-verify the premise; prefer WHEN-guidance (conditions where X pays) over directional nudges — survives inversions unchanged.
- **link**: [[LRN-030]] [[BDR-081]].
## LRN-140 — de-prescription findings: dedup evaporates, self-verify is default, recall survives (2026-08-02)
- **pattern 1 — inventory dedup counts lie**: line-level inspection killed most "duplicate" pairs (seo 9 families→2 real merges; geo 7→0). Twins differ by AUDIENCE (bundle-item payload read by fresh applier vs spec rule) or MODE-RANGE (collect/judge/template/RULES) or are distinct obligations sharing a keyword (30/70 ×3 = three different rules). Dedup rule that survives: verbatim + same-audience + same-range ONLY.
- **pattern 2 — Opus 5 self-verifies unprompted**: "run it twice" instruction REMOVED → after-judge still ran score engine twice, identical output. Removing verify-prose does not remove the behavior; its value = no compounding, no contradiction burn. Confirms BDR-081 E3 mechanism, refines the payoff claim.
- **pattern 3 — de-prescription does NOT depress recall**: reworded collect caught &nbsp;-encoded phone AT COLLECT (baseline collect missed it); reworded judge found new RGPD finding + self-caught false positive + corrected collect coverage claim 21/21→20/21. Integrity/honesty invariants (kept class B) carry the discipline, not the caps.
- **pattern 4 — lock strings, never shapes**: LLM-convention output layers (banners, fences, table columns, section order) wobble run-to-run in BOTH directions — baseline itself deviated from spec where after conformed (§0 ENTRIES, BUNDLE-before-SCORING). Stable contract = census-locked literal strings; anything unlocked drifts and MUST be tolerated by consumers (tier recognition "by intent" is the right pattern).
- **link**: [[BDR-082]] [[BDR-081]] [[LRN-139]] [[LRN-113]].
## LRN-141 — adopting an external skill: take the invariants, refuse the machinery (2026-08-24)
Context: unlazy import ([[BDR-083]]). Pattern: an external skill's MACHINERY encodes ITS threat model and ITS doctrine; only its INVARIANTS transfer. Two clean cases from one repo. (1) Approval store binding PATH/shell/platform exists because unlazy executes ledgers INHERITED from untrusted repos — importing it into a config that authors its own ledgers buys per-command approval prompts and closes zero threat. (2) Stop hook returning `decision:"block"` exists because unlazy has no human gate — importing it into a config whose spine is "STOP + escalate to human" would make the tooling fight the doctrine. Meanwhile the invariants (exit 0 AND marker; evidence persisted so the next reader gets fact not report; impossible ≠ deletable) cost ~250 l of our own bash and fit the EXISTING contract with no new tree.
Separating test: ask WHAT THREAT / WHAT DOCTRINE does this piece assume. Answer "theirs" → refuse the piece, keep the invariant it was protecting.
Corollary on claims: unlazy's own research/validation-protocol.md RETRACTS its v1 benchmark numbers as unreproducible while the repo DESCRIPTION still advertises them. Read a project's self-criticism before its README — the retraction is the credibility signal, the headline is not.
Future application: any skill/plugin adoption — skills-external/, /plugin-check, install-plugins.sh.
## LRN-142 — structure locks are fixed-string: reflowing a doctrine paragraph reds them (2026-08-24)
Context: contract-gates ([[BDR-083]]). Editing lib/verify-secure-loop.md rewrapped 5 locked phrases across line breaks ("Max 3 conformity iterations", "Max 3 security iterations", "re-verify the REQUEST first", "always re-checked BEFORE security", "one verifier dispatch + one security dispatch") → loops-light.test.sh 30 pass / 5 fail, though ZERO doctrine was dropped. Locks did their job: they cannot distinguish "clause deleted" from "clause rewrapped", and that conservative bias is correct — the alternative (fuzzy matching) would miss real deletions.
Rule: when editing a doctrine file under structure locks, grep the test's lock strings FIRST, then re-flow AROUND them — each locked phrase stays on one unbroken line. Fix the DOC, not the lock, unless the doctrine genuinely changed. Under locks today: verify-secure-loop.md, contract-interview.md, verifier / security-auditor / plan-challenger agents, seo+geo (71 locks).
## LRN-143 — pipe to head masks grep exit; `|| fallback` never fires
- **Context**: darwin 2026-08-26 — plugin-probe FRAMEWORK-DEPS (`grep … | head || echo none`) emitted silent-empty on no-match; same bug in run's own probe test.
- **Pattern**: pipeline rc = LAST command's (head = 0 always). `|| fallback` after pipe = dead code. Bounded output → drop head; else `set -o pipefail` or capture + test.
- **Future**: any skill/agent bash probe with a `||` fallback: check what the pipeline rc actually is first.
## LRN-144 — census locks grep EXACT single-line phrases; prose rewrap breaks them
- **Context**: darwin 2026-08-26 — hotfix RULES rewrap split "No verifier is dispatched at hotfix weight"; loops-light.test.sh lock RED; make test caught post-edit.
- **Pattern**: lib/tests/*.test.sh lock sentences verbatim, single-line. Rewording/rewrapping skill+agent md near locked phrases silently breaks census.
- **Future**: before editing skill/agent prose, grep lib/tests/ for locks in the touched region; run make test BEFORE dispatching judges, not after.
## LRN-145 — hooks reach the terminal only via terminalSequence JSON field
- **Context**: 2026-09-01 — attention bell for VS Code Remote-SSH (CLI on remote Linux). Hook subprocess has no controlling TTY; /dev/tty unreliable. Docs: terminalSequence = supported side-effect field, fires even on events that discard output.
- **Pattern**: Notification hook → stdout JSON `{suppressOutput:true, terminalSequence:"<BELx2><OSC 777 notify><ST>"}`. VS Code terminal ignores OSC 777/9 natively (claude-code #28338); client-side ext wenbopan.vscode-terminal-osc-notifier converts to native toast over Remote-SSH; beep needs accessibility.signals.terminalBell sound:on. permission_prompt fires ~6s late, idle_prompt ~60s.
- **Future**: any hook ringing/notifying the terminal (bell, toast, title) — terminalSequence, never /dev/tty. Input-needed matcher set: permission_prompt|idle_prompt|agent_needs_input|elicitation_dialog|elicitation_url_dialog.
## LRN-146 — Notification event alone misses end-of-turn; Stop is the missing event
- **Context**: 2026-09-03 — attention signal verified end-to-end after [[BLK-020]]. Matcher `permission_prompt|idle_prompt|agent_needs_input|elicitation_*` covers input-needed cases ONLY. "Claude finished speaking" has no notification_type — nearest was `idle_prompt`, ~60s late. Gap invisible until explicitly enumerated by user.
- **Pattern**: wire SAME hook script on TWO events — `Notification` (matcher = input-needed set) + `Stop` (fires once per turn end, supports terminalSequence, no matcher). Script branches on `.hook_event_name` when `.message`/`.notification_type` absent: Stop → "Claude has finished responding", else default. Read stdin ONCE into var, jq the var (stdin not re-readable).
- **Verified**: turn-end bip+toast OK, AskUserQuestion selector bip+toast OK. `permission_prompt` NOT exercisable under `defaultMode: auto` — ask-rules (`python3 -c *`, `curl`…) auto-approved, no prompt raised. Hooks hot-reloaded by file watcher, no restart.
- **Future**: enumerate the events a signal must cover BEFORE wiring, one per user-visible moment. Notification ≠ lifecycle-complete. SubagentStop exists too for agent completion.
## LRN-147 — VS Code restores terminals BEFORE ext activation → toast dies every restart
- **Context**: 2026-09-03, second hit same day. Bell OK, toast gone, after user re-attached session from a restored terminal. Probe on that pty: OSC 777 unique + OSC 777 repeated + OSC 9 → all three silent, while BEL rang. Same pty, bell works ⇒ bytes arrive, ext just not hooked to that terminal.
- **Pattern**: `wenbopan.vscode-terminal-osc-notifier` instruments a terminal only if it exists AFTER ext activation. `terminal.integrated.enablePersistentSessions` (default true) restores terminals at window startup, i.e. BEFORE lazy ext activation → every restored terminal is permanently deaf to OSC. Recurs at each VS Code restart, silently, bell still ringing so it reads as "half broken".
- **Fix**: client setting `"terminal.integrated.enablePersistentSessions": false` → no terminal pre-exists activation. Fallback without it: after VS Code start, open a FRESH terminal then `dtach -a ~/.dtach/<session>` (dtach broadcasts, old client can stay or be closed, session never lost).
- **Diagnostic shortcut**: bell rings + toast dead on the SAME pty = terminal-instrumentation fault, not audio, not hook, not server. Bell dead + toast alive = audio fault ([[BLK-020]] fault B). The two channels split the search space; check which one survives before anything else.
- **Future**: any client-side terminal-parsing ext over Remote-SSH inherits this. Verify instrumentation on the ACTUAL attached pty after every restart, never assume yesterday's terminal.
## LRN-148 — terminal instrumentation is per-terminal + unpredictable; pre-flight test before attaching
- **Refines**: [[LRN-147]] blamed restored-terminals-born-before-activation. Too narrow — counter-example same day: two terminals SAME VS Code window, pts/3 (born 01:58:33) instrumented, pts/7 (born 01:59:29, LATER) deaf. Ext is GLOBAL (marketplace: Enable/Disable pause parsing extension-wide, no per-terminal setting), shells identical on every server-side measurable: `VSCODE_INJECTION=1`, TERM, TERM_PROGRAM, same `--init-file` shell-integration path, ~2-3s between shell start and dtach. Trigger NOT identified.
- **Pattern**: treat instrumentation as a per-terminal property that can silently fail for unknown reasons. Cheap pre-flight before committing a long-lived session to a terminal: `printf '\a\a\033]777;notify;NEUF;test\033\\'` typed IN that terminal. Toast → instrumented, attach. Bell only → deaf terminal, open another. Costs 5s, replaces an hour of pty archaeology.
- **Recovery**: deaf terminal never repairs. Open fresh terminal, pre-flight it, `dtach -a ~/.dtach/<session>`. dtach broadcasts, so old client may stay attached; session never at risk.
- **Diagnostic split (holds)**: bell alive + toast dead = terminal instrumentation. Toast alive + bell dead = client audio ([[BLK-020]]). Neither = bytes never arrive.
- **Future**: do NOT assert the born-before-activation cause as established — it fits the first incident, not the second. Unknown trigger is the honest state.
## LRN-149 — Stop hook payload carries background_tasks; use it to skip premature signals
- **Context**: 2026-09-03. User: "notif à la création d'un sous-agent alors qu'il faudrait pas". Instrumented hook, ran probe subagents: NEITHER subagent creation NOR completion calls the hook. Only event = `Stop`, fired when the turn ends right after spawning. Signal was real but LIED ("Finished responding" while work continued).
- **Pattern**: dump the real payload (`printf '%s' "$payload" >> file.jsonl`) instead of trusting docs — docs list Stop fields without `background_tasks`, the wire has it: `[{"id","type":"subagent","status":"running","description","agent_type"}]`. Rule: on Stop, `(.background_tasks // []) | length` > 0 → exit 0 silent. Next turn end signals for real. Interaction events (permission/question) always signal, background or not.
- **Fail-open**: field absent (older client) → still signal. Missed notification worse than extra one.
- **Cross-session gotcha**: hook is user-scope, so EVERY session runs it. A single-file dump (`> file`) gets overwritten by another project's session — append JSONL and filter on `.cwd`. That accident proved `permission_prompt` fires with `message="Claude needs your permission"` (unexercisable in this session under `defaultMode: auto`).
- **Future**: any hook needing turn-completion semantics must check background_tasks; "turn ended" ≠ "work done". Verified live: Stop with 0 tasks signals, Stop with 1 running subagent silent.
---
## LRN-150 — Sourced shell lib is not a subprocess: prefix printers, honor inherited errexit
- **Date**: 2026-09-15
- **Pattern**: `source lib.sh` shares the caller's shell. Two bites. (a) bare `ok()`/`warn()`/`info()` in the lib OVERRIDE the caller's same-named funcs. `doctor.sh` counts ERRORS/WARNS inside its own `warn()` → a lib `warn` disconnects the counter and doctor prints "No errors" while warnings scroll. Prefix every lib printer (`_gspw_ok`, `_gspw_warn`, `_gspw_info`). (b) caller's `set -euo pipefail` applies INSIDE the lib's functions: a failing command-substitution assignment (`x="$(. /etc/os-release; [ "$ID" = ubuntu ] && printf ...)"`) aborts the CALLER when the func is called as a bare statement. Reproduced — exit 1 on every non-Ubuntu host, latent in `install-plugins.sh` since [[BDR-029]].
- **Rule**: public func called bare → `return 0` on every path + `|| true` on every capture. Func allowed to return non-zero → call it ONLY as an `if` condition.
- **Future application**: any new `lib/*.sh` sourced by a script that owns printers or sets `-e`. Check BOTH facets before wiring; the printer one is silent (no error, just a lying summary).
- **Reference**: `lib/gstack-playwright.sh`, `doctor.sh:12-15`. Links [[BDR-088]].
---
## LRN-151 — Playwright cache truth lives in `.links`, never in one install's view
- **Date**: 2026-09-15
- **Pattern**: `~/.cache/ms-playwright/.links/<sha1>` = one file per registered `playwright-core`, content = its path. Required set = UNION of `browsers.json` revisions across ALL of them. Dir name on disk = `${name//-/_}-${revision}`: `chromium-headless-shell` → `chromium_headless_shell-1228`. Miss that mapping and 2 live dirs read as orphan forever. `revisionOverrides` exists (webkit, ffmpeg on mac / debian11 / ubuntu20.04) so the base revision alone under-matches. Playwright prunes this set itself on every `install` (`_deleteStaleBrowsers`, coreBundle.js).
- **Future application**: never call a browser dir orphan from one project's playwright view — read `.links` first. Generalizes to any tool with a shared versioned binary cache plus a registry of consumers: the consumer registry is the source of truth, not the consumer you happen to be standing in.
- **Reference**: `lib/gstack-playwright.sh` `_gspw_browser_referenced`. Links [[BDR-089]], [[EVAL-029]].
---
## LRN-152 — git `protocol.file=user` kills submodule fixtures; `-c` misses the code under test
- **Date**: 2026-09-15
- **Pattern**: since the CVE-2022-39253 hardening git refuses submodule clone/fetch over a local path by default (git 2.53 → `protocol.file` = `user`). `-c protocol.file.allow=always` fixes the FIXTURE's own git calls but NOT the `git` the code under test spawns — fresh process, inherits nothing from `-c`. Export for the whole test process instead: `GIT_CONFIG_COUNT=1 GIT_CONFIG_KEY_0=protocol.file.allow GIT_CONFIG_VALUE_0=always`. Env propagates, `-c` does not.
- **Also**: fixture repos need LOCAL `user.email`/`user.name` (no global identity here) and `git init -b main` + explicit `submodule.<name>.branch`, else `--remote` resolves a different branch than production does.
- **Future application**: any test building a git submodule fixture. Symptom is a hard "transport 'file' not allowed" before the first assertion, which reads like a broken test rather than a policy.
- **Reference**: `lib/tests/gstack-playwright.test.sh`.
## LRN-153 — `autoMode` lists replace built-ins unless `"$defaults"` is spliced in
- **Date**: 2026-09-15
- **Pattern**: every list under `autoMode` (`allow` `soft_deny` `hard_deny` `environment`) is a FULL replacement by default. Omit the literal `"$defaults"` and the built-in classifier rules are dropped silently — no warning, no schema error, the classifier just runs thinner. Put `"$defaults"` first, own entries after: built-ins inherited, then refined.
- **Scope trap, same block**: `autoMode` in `~/.claude/settings.json` reaches EVERY project. A block generated while working in one repo (its deploy target, its secrets, its data) ships that repo's facts to all the others, and contradicts whichever repo is actually open. Project facts belong in that project's `.claude/settings.local.json`.
- **Format**: these lists are prose spliced into the classifier prompt, not permission-rule syntax. Write "Sending SIGKILL reaches processes outside this session", never `Bash(kill -9 *)`.
- **Backstop**: `doctor.sh` `check_automode` warns on a list missing `$defaults` and on a user-scope `environment` naming a git repo other than the config repo. Both arms exercised against the defective block before shipping.
- **Future application**: any `autoMode` edit — check `$defaults` presence and scope before anything else.
- **Reference**: `doctor.sh`, `templates/settings/SETTINGS.md`. Links [[BDR-090]].
## LRN-154 — Untracking a generated file then merging deletes it from disk
- **Date**: 2026-09-15
- **Pattern**: `git rm --cached` removes from the index and KEEPS the working file, which is the whole point when untracking a tool-generated artifact. But `gitflow finish` checks out the target branch first, where the file is still tracked, so git restores it; the merge then applies the deletion to a tracked file and removes it from disk. `.gitignore` does not protect it — it only stops a re-add. Net effect: the file survives the commit and dies at the merge, several minutes later, which reads as unrelated.
- **Detection**: the working tree is clean and the file is simply absent. Nothing errors. Only a post-merge `ls` catches it.
- **Future application**: untracking any generated file — know the regeneration command BEFORE merging, and `ls` the path right after `finish`. If nothing regenerates it, keep it tracked.
- **graphify specifics**: `graphify install --platform claude` copies the skill and touches nothing else. `graphify claude install` is a different command — it writes the CLAUDE.md section and the `.claude/settings.json` hooks, rewrites both guarded configs, and does NOT copy the skill. Confusing the two wastes a recovery attempt.
- **Reference**: `CLAUDE.md` machine-owned section, commit 80ccdaf. Links [[BDR-090]].
## LRN-155 — `permissions.ask` under auto mode: the probe beats the doc
- **Date**: 2026-09-16
- **Pattern**: `auto-mode-config` + `permissions` docs say a content-scoped `ask` rule (`Bash(git push *)`) is evaluated BEFORE the classifier and always prompts, even in auto mode. Probe on 2.1.273: `node -e 'console.log(...)'` matching `Bash(node -e *)` in `ask` ran, no prompt, exit 0. [[LRN-146]] holds. Either the doc describes a later build or "content-scoped" means something narrower; observed wins.
- **Future application**: before reasoning about a permission tier, probe it with a benign command matching the rule; re-probe after every Claude Code upgrade — the day `ask` starts prompting, every leftover `ask` entry becomes a nag for things meant to run free.
- **Reference**: [[BDR-092]], `templates/settings/SETTINGS.md` "ask is not a prompt" §.
## LRN-156 — Conditional permissions live in classifier prose, not static rules
- **Date**: 2026-09-16
- **Pattern**: `autoMode.allow` = exception tier: an entry overrides a matching `soft_deny`, built-in or own (precedence hard_deny > soft_deny > allow > explicit intent). Under auto, static allow rules granting arbitrary execution (`Bash(*)`, wildcarded interpreters like `Bash(node *)`) are suspended → classifier anyway; non-interpreter statics (`awk`, `echo`) resolve before it. A condition ("package declared in the lockfile", "container is local dev") is therefore expressible ONLY as `autoMode.allow` prose. Word it narrowly: it punches through built-in rules too.
- **Tooling**: `claude auto-mode defaults` prints the built-in lists (grep it for the rule that bit); `claude auto-mode config` = effective lists with `$defaults` expanded; `claude auto-mode critique` printed nothing on 2.1.273. Shell-snapshot `claude` wrapper is broken (`exec command claude` → "command: not found") → call `~/.local/bin/claude` directly.
- **Reference**: [[BDR-092]], [[LRN-153]].
## LRN-157 — Taste is invisible to a gap-only trigger; ask at plan time
- **Date**: 2026-09-16
- **Pattern**: a trigger that fires only on missing outcome / scope / constraints lets every taste choice through — "add a share icon" is complete by those criteria and the icon's side is decided downstream. More budget changes nothing; the fix is a new trigger class (VISIBLE / PUBLIC NAME / SCOPE). Cost geometry: a fresh re-dispatch keeps the working tree and loses the executor's reasoning → the same question costs about one executor run more mid-run than at PLAN. So: sweep once at the plan step, keep the mid-run channel for leftovers. Executor tags the class; orchestrator re-reads it (tag = hint, a mis-tag would offload class 4 onto the human). Relayed questions obey [[LRN-102]]: context inside `AskUserQuestion`, nothing the user needs printed before it.
- **Future application**: any "ask more" request → check WHICH trigger is blind before touching a quota. Any orchestrator with a "decide it yourself" fallback on an executor halt → route by class first.
- **Reference**: [[BDR-091]], `lib/contract-interview.md` STEP 2 + MID-RUN CLARIFICATION.
## LRN-158 — A hardened installer + a symlinked config dir = documented command fails; stage under a throwaway HOME
- **Date**: 2026-09-22
- **Context**: `21st install-skill` (= `21st skills install --global`) is upstream's documented one-liner. Here it dies: `Refusing to access symbolic link /home/…/.claude/skills`. The installer walks every segment of `<HOME>/.claude/skills/<n>/SKILL.md` with an `assertNoSymlinkComponents` guard (anti symlink-escape); this repo's whole model is `~/.claude/skills -> repo/skills`. Two correct designs, mutually exclusive on the same path.
- **Pattern**: don't fight the guard and don't unlink the config dir. Run the installer with `HOME=$(mktemp -d)` so it writes into a pristine real tree, then move the output to the vendored dir the repo controls and symlink from there. Same shape as the impeccable/ctx7 staging (`mktemp -d`, install, `mv` into `skills-external/`), with HOME as the extra lever. Two conditions make it safe: the command must need nothing else from HOME (checked: manifest + content fetch are unauthenticated, hash-verified), and the moved payload must be self-contained.
- **Also**: read the npm tarball, not the vendor's web page. 21st.dev's `/mcp` and `/llms.txt` still document the MCP `init --client` flow with an API key; the package README states the CLI supersedes it. `curl registry.npmjs.org/<pkg>` + untar + read `README.md`/`dist` answered every question (commands, exit codes, where files land) that the site got wrong.
- **Future application**: any vendor installer that writes into `~/.claude`, `~/.config` or `~/.agents` on this machine. Probe first with a fake HOME containing the symlink, before wiring it into `install-plugins.sh` — the failure is instant and unambiguous.
- **Reference**: [[BDR-093]], `install-plugins.sh` Step 8.7, `update-all.sh` 7.4. Links [[LRN-034]] (run the real thing), [[BLK-014]]-class symlink/self-heal issues.
## LRN-159 — A pin whose payload is fetched at install time rots: pin + fallback, and read the installer's output, not its exit code
- **Date**: 2026-09-22
- **Context**: `impeccable@3.2.0` still on npm, but `skills install` downloads the skill dist at run time and that release's zip is gone → "Download failed: invalid zip data". Strict pin = `make plugin` fails forever, prints "run it yourself". Second layer: once a copy exists, same CLI exits 0 on the same failure ("Could not check for skill updates … Existing skills were left unchanged"), indistinguishable by rc, by SKILL.md version or by mtime/sha from "Skills are up to date".
- **Pattern**: two classes of npm pin. (a) self-contained package → pin freezes behaviour, rc is truth. (b) package that fetches its payload at install time (impeccable, ctx7, `skills add` style) → pin freezes only the fetcher; payload can vanish or drift. For (b): pin + `@latest` fallback + loud "bump the lock" warn, never pin-or-die. And when the tool has an "already installed" branch, capture stdout+stderr and match the failure text; rc and before/after compare both read "unchanged" for a no-op AND for a swallowed failure.
- **Future application**: any `install-plugins.sh` step whose pinned tool downloads something at install time. Probe both HOME states (clean, copy present) before trusting rc. Cheap recipe: sandbox HOME with the repo-shaped symlinks, pinned install twice, then the rotted pin; diff rc + output + `stat`/`sha256sum` of the landed file.
- **Reference**: [[BDR-094]], `install-plugins.sh` Step 8d `imp_install`, `update-all.sh`. Links [[LRN-077]] (why pin), [[LRN-034]] (run the real thing), [[LRN-158]].
+469 -3
View File
@@ -1,5 +1,393 @@
# TODO
## 2026-09-22 — impeccable install repaired: global scope + agents + rotted pin (feature/21st-cli-migration)
User: `make plugin` never installs impeccable, it just prints "run it
yourself"; running it by hand needs `--scope=global` to land right, and then
`/impeccable init` is still required. Three separate defects, all confirmed:
1. **Pin rotted.** `npx -y impeccable@3.2.0 skills install` → `Download
failed: invalid zip data`, rc 1. The CLI fetches its skill dist at install
time and that release's artifact is gone. 3.6.1 / 4.0.5 / 4.1.0 all work.
That rc 1 is the "run manually" warn the user sees.
2. **Wrong scope + half the payload dropped.** The step staged
`--scope=project` in a tmpdir and `mv`'d only the skill dir, silently
discarding the 4 `impeccable-*` subagents the installer also writes.
`--scope=global` writes `~/.claude/skills/impeccable` +
`~/.claude/agents/impeccable-*.md`, and both are symlinks INTO this repo,
so a global install is the repo install. Verified in a sandbox HOME.
3. **`/impeccable init` never surfaced.** It writes per-project PRODUCT.md
(design context the skill reads); it runs in the agent chat, so install
can only announce it and the design gate has to check it.
- [x] T1 install-plugins.sh Step 8d rewritten: global scope, no staging,
pin→latest fallback with a loud bump-the-lock warn, park-aware
(profile may hold impeccable in skills-disabled), symlink precondition
guard, agent count + skill version reported, init hint printed.
Harness-tested against a fake HOME with repo-shaped symlinks: happy
path OK, park/restore OK. Caught + fixed there: `find` stops at the
`~/.claude/agents` symlink without `-L`, so the agent count read 0
while 4 agents were installed.
- [x] T2 update-all.sh impeccable block: same shape. `bash -n` only, NOT
run end to end.
- [x] T3 plugins.lock.json: 3.2.0 → 4.1.0 + honest note (pin covers the CLI
only; skill dist 4.3.1 and engine 0.1.5 have their own tracks).
- [x] T4 .gitignore: `agents/impeccable-*.md` (machine-owned, tracked dir);
drop `skills-external/impeccable/`. link.sh: impeccable out of
EXTERNAL_SKILLS (nothing to symlink any more). `git check-ignore`
confirms both paths.
- [x] T5 lib/design-gate.md §5: suggest-only PRODUCT.md / `/impeccable init`
check, same shape as the §4 animation-library check.
- [x] T6 duplicate project-scope install: already gone at resume (user ran
`rm -rf .claude/skills .claude/agents` before restarting).
- [x] T7 CHANGELOG (Added/Changed/Fixed) + BDR-094 + LRN-159 + BLK-021 +
journal. `make test` green except the 2 pre-existing gitflow T16a FAILs
(gitleaks binary absent on this host), shellcheck clean. Committed on
the branch, UNMERGED — human gate.
**Residual, probed and fixed (round 3)**: with a copy already installed a
rotted pin DOES exit 0 ("Could not check for skill updates: invalid zip data
… Existing skills were left unchanged"), and so does a genuine rerun of a
good pin ("Skills are up to date (v4.3.1)"). Both leave SKILL.md
byte-identical, so a before/after version compare cannot separate them.
`imp_install` (Step 8d and update-all.sh) now captures the installer output
and fails on `Download failed|Could not check for skill updates`, whatever
the exit code. Harness on the extracted step, sandbox HOME, real installer:
fresh install; rotted pin over a copy → fallback fires; same pin rerun → no
false warn; parked copy + rotted pin → fallback, then returned to
skills-disabled/. update-all.sh: `bash -n` + shellcheck only.
OPEN for the user:
- /tmp is a tmpfs with a per-user quota and the dead session's scratchpad
holds 5.9 GB of probe HOMEs. Writes to /tmp fail with EDQUOT: the likely
cause of the "every command exits 1" shell death (BLK-021). Free it:
`rm -rf /tmp/claude-1000/-home-bchanot-Documents-claude/fefd277c-e143-4d51-b589-a566641079b5`
(the agent's `rm -rf` under /tmp is denied). This round ran tests and the
harness with TMPDIR under ~/.cache.
- `skills/synced/` (4.4 MB, untracked, not ignored): claude.ai's synced
skills, written through the ~/.claude/skills symlink. Decide whether to
gitignore it; not touched here.
## 2026-09-22 — 21st: magic MCP → CLI + skills (feature/21st-cli-migration)
User: "remplacer pour 21st, il n'y a plus besoin de mcp / api, mais juste en
cli". Upstream confirmed (`@21st-dev/cli` 1.17.1 README): the CLI supersedes
`@21st-dev/magic`; auth is `21st login` (browser token in `~/.config/21st`),
no API key; `21st install-skill` = alias of `21st skills install --global`.
Gates answered by user: 5 design skills in profiles (registry + design-sync
parked), `make plugin` auto-installs the CLI + offers login on TTY only,
missing `21st` CLI trips the design gate (magic's old required-manual slot).
Blocker found + solved: `21st skills install --global` REFUSES to write
through a symlinked path (`assertNoSymlinkComponents`), and `~/.claude/skills`
IS a symlink → repo/skills. Verified live: "Refusing to access symbolic link
…/.claude/skills". → install into a staged HOME (mktemp), move each skill to
`skills-external/21st-*/` (impeccable pattern), symlink from there.
- [x] T1 install-plugins.sh STEP 8.7: magic block → 21st CLI (`npm i -g`,
pinned via plugins.lock.json) + staged `skills install` →
skills-external/21st-*, TTY-gated `21st login`, pack disabled by default.
- [x] T2 lib/toggle-external.sh: managed tool `magic` (mcp) → `21st` (skill
pack, glob-derived from skills-external/21st-*), drop load_env.
- [x] T3 profiles + profile.sh: `magic mcp` → 5 externals + `21st cli` in
design/web/web-full/full; GATE-BLOCK `21st 21st-ui-build`;
MANAGED_EXTERNALS += the 5; MANAGED_MCPS emptied (kept as a live
allowlist, mcp type machinery stays generic).
- [x] T4 lib/design-tool-gate.sh + lib/design-gate.md: manual-step hint
magic/MAGIC_API_KEY → 21st/`npm i -g` + `21st login`; PATH repair
extended to the npm-global bin dir (21st lives in nvm's bin, the
existing repair only fires when `claude` itself is unresolvable).
- [x] T5 doctrine + docs: CLAUDE.global.md design toolchain, README (drop the
magic callback-injection section + the MCP env-var worked example),
.env.example, link.sh MAGIC_API_KEY warning, .gitleaks.toml allowlist,
update-all.sh, .gitignore, settings.json (drop 4 mcp__magic__*; the
outward-facing verbs landed in autoMode.soft_deny, NOT ask — LRN-153
says ask is inert under auto mode).
- [x] T6 lib/tests/profile-set-managed.test.sh retargeted (mcp fixture → 21st
external pack), `make test` + shellcheck green.
- [x] T7 CHANGELOG + BDR-093 + LRN-158 + journal. Also cleaned along the way:
dead `magic` branches in profile.sh enable/disable_skill,
skills/profile/SKILL.md. OPEN for the user: `npm i -g @21st-dev/cli`
then `21st login` (`Bash(npm install -g *)` is denied to the agent).
Branch UNMERGED — human gate.
## 2026-09-17 — /deploy hand-back: one-line commands + post-deploy test list (feature/deploy-oneline-tests)
User: commands in the /deploy checklist arrive broken across lines (cannot
copy-paste), and the hand-back stops at the deploy steps — wants, after the
checklist, a list of things to test by hand about THIS delta + suggestions.
Evidence: zenquality runbook step 3 carries a `\`-continued psql; game runbook
has 200-350 char command lines the model re-wraps at display (80-char code
style pressure). Method: writing-skills RED/GREEN on a scratch fixture repo
(4 fresh agents, skill body as instructions, gate pre-approved).
- [x] D1 RED baseline: 4 runs on the current skill, record wrapped commands
+ absence of a test list + rationalizations
- [x] D2 SKILL.md: physical-line rule (checklist, bootstrap, learn patch;
join legacy `\` continuations at instantiation), post-deploy tests
recipe (manual checks + suggestions, derived from the delta), hand-back
order checklist → tests → report request; Rules / mistakes / red flags
- [x] D3 templates/deploy/PROCEDURE.md style header + test-prompts.json
- [x] D4 GREEN: re-run 4 fresh agents on the edited skill, compare shape
- [x] D5 CHANGELOG [Unreleased] Changed; report; offer capitalize (EVAL + LRN)
Milestone 2026-09-17: RED 4/4 (3 sonnet + 1 opus) reproduced the `\`
continuation verbatim, no re-wrap of 200+ char lines, no test list; GREEN
4/4 joined the continuation, kept long lines whole, printed the tests block
in the recipe's shape (grant gap as a Suggestion, never patched). Branch
feature/deploy-oneline-tests, uncommitted, awaiting user. Registries: EVAL +
LRN drafts proposed, not written.
## 2026-09-16 — docker + node framed by the classifier (feature/automode-docker-node)
User: `docker exec -i supabase_db_game psql … -f - < supabase/verify/*.sql | tail`
must run unprompted under auto mode; same for node/npm/npx when the package
is declared and effects stay in the cwd; "ajoute du soft deny pour bien le
cadrer". Findings: `ask` is inert under auto (LRN-146 re-verified on 2.1.273
with a `node -e` probe; the docs claim otherwise for content-scoped rules);
the real gate is the built-in `Remote Shell Writes` / `Production Reads`
classifier rules; a static `Bash(node *)` allow rule is suspended under auto
(wildcarded interpreter), so `autoMode.allow` prose is the only lever for
node. User approved the design and the `ask` removal explicitly (S6 override
for this change, diff reviewed on the branch).
- [x] A1 `settings.json` — drop 4 docker + `node -e` from `ask`; new
`autoMode.allow` (`$defaults` + local dev containers + project-local
node); 2 `soft_deny` entries (docker data destruction, undeclared
node packages); `model` bump committed separately
- [x] A2 `templates/settings/SETTINGS.md` — `autoMode.allow` tier row +
why a static interpreter allow rule cannot do it; LRN-146 re-verify note
- [x] A3 CHANGELOG [Unreleased] Changed
- [x] A4 verify (2026-09-16, all green; `critique` printed nothing): `jq`, `claude auto-mode config`,
`doctor.sh`, live `docker exec` in game
- [x] A5 registries written 2026-09-17: LRN (doc vs observed `ask` under auto,
2.1.273; `autoMode.allow` = exception tier; wildcarded-interpreter
allow suspended), BDR-090 addendum
## 2026-09-16 — ask, don't guess: orchestrators ask about open choices (feature/ask-dont-guess)
User: the orchestrators (ship-feature, feat, hotfix, bugfix, init-project)
settle choices they should ask about ("cet icône, plutôt à gauche ou à
droite ?"), even mid-run. Diagnosis: contract-interview STEP 2 only fires on
gaps (outcome / scope / constraints), so a taste choice never triggers a
question; feat:153 and bugfix:165 tell the orchestrator to "make the
decision HERE" on NEED-DECISION. Decisions (user, 2026-09-15/16): global
rule changes for all work, hotfix included; classes VISIBLE / PUBLIC NAME /
SCOPE ask, internal technical choices never. Spec:
`docs/superpowers/specs/2026-09-16-ask-dont-guess-design.md`; plan:
`docs/superpowers/plans/2026-09-16-ask-dont-guess.md` (9 tasks, lock-first).
- [x] P1 `lib/contract-interview.md` — STEP 2 CLARIFY (pass A gaps, pass B
open choices), MID-RUN CLARIFICATION, HOW TO ASK; 9 locks in
`contract-verifier.test.sh`
- [x] P2 `CLAUDE.global.md:51-55` — "Ask rather than guess" replaces the
one-question rule; bug line reconciled
- [x] P3 `skills/feat/SKILL.md` — pass B at STEP 1, NEED-DECISION routed on class
- [x] P4 `skills/bugfix/SKILL.md` — pass B at STEP 3, NEED-DECISION routed on class
- [x] P5 `skills/hotfix/SKILL.md` — pass B at LOCATE, tagged BLOCKED relayed;
lock `loops-light.test.sh:84`
- [x] P6 `skills/ship-feature` STEP 2 + `skills/init-project` contract §/STEP 3
- [x] P7 `agents/interviewer.md` — visible/public/scope item never `(assumed)`
- [x] P8 `agents/{feater,bugfixer,hotfixer}.md` — CLASS tag; 3 locks in `gates.test.sh`
- [x] P9 `make test` green (2026-09-16), CHANGELOG, TODO tick; manual behavioral check still OPEN before merge
- [x] P10 registries written 2026-09-17: BDR (supersedes the one-question rule),
LRN (taste is invisible to a gap-only trigger; fresh re-dispatch cost
favors plan-time questions)
## 2026-09-15 — align config + deployment on the hand-edited settings.json (feature/automode-config-alignment)
User edited global `settings.json` by hand: 4 destructive rules moved
deny→ask (`rsync`, `kill -9`, `killall`, `pkill`), 4 removed from ask
(`xargs`, `sed`, `cp`, `mv` — coherent with auto mode's Bash-first
workflow; the `.env`-scoped `cp`/`mv`/`xargs` deny rules still stand),
and an `autoMode.environment` block added. Two defects found:
(1) the environment block describes **atlast** (`bin/deploy.sh` lftp/FTP
to OVH, quote-request data, "no remote configured") but lives in the
user-scope file symlinked to `~/.claude/settings.json` by `link.sh:21`
— so every project gets atlast's facts; claude-config itself has a
Gitea remote, contradicting the block. (2) no `"$defaults"` sentinel,
so the built-in classifier environment entries are replaced, not
extended. Third finding: LRN-146 records, verified in session, that
`ask` rules raise no prompt under `defaultMode: auto` — the deny→ask
move therefore traded a static block for a classifier decision.
User decisions (2026-09-15): atlast block → atlast's own
`settings.local.json`, global block rewritten machine-generic; the 4
destructive rules → `autoMode.soft_deny` (the section that actually
binds under auto mode) instead of `ask`.
- [x] T1 global `settings.json` — machine-generic `autoMode.environment`
with `$defaults`; new `autoMode.soft_deny` with `$defaults` + the
4 destructive rules; drop those 4 from `permissions.ask`
- [x] T2 `/home/bchanot/Documents/atlast/.claude/settings.local.json` —
receives the atlast-specific `autoMode.environment` (gitignored,
personal scope); verify project-scope `autoMode` is honored
- [x] T3 `templates/settings/SETTINGS.md` — document the `autoMode`
block (environment / soft_deny / hard_deny / allow, `$defaults`
semantics, `classifyAllShell`) + the "ask ≠ prompt under auto"
caveat that makes soft_deny the right tier
- [x] T4 `README.md` — magic-MCP paragraph claims the `ask` tier makes
every `mcp__magic__*` call "require a live confirmation and never
auto-execute"; false under auto mode per LRN-146. Correct the
claim, flag the soft_deny option to the user (don't decide it)
- [x] T5 `doctor.sh` — permissions section is blind to `autoMode`, now a
live security surface. Add a check: block present, `$defaults`
inherited, no foreign absolute project path hardcoded
- [x] T6a CHANGELOG (Added/Changed/Fixed under [Unreleased])
- [ ] T6b registries BDR-090 + LRN-153 + journal — drafted, awaiting user approval
- [x] T7 verify: `make test`, `bash doctor.sh`, `shellcheck`
NOT in scope: the 3 dirty `skills/graphify/*` files (pre-existing,
unrelated) — never staged.
### Second pass (2026-09-15, user decisions)
User confirmed the `ask` removals were deliberate (`/permissions`), asked
for the diff vs develop and for guards where the removals left a hole.
Answered: writes outside cwd → soft_deny; in-place edits beyond one named
file → soft_deny; inline interpreters + `xargs` → soft_deny when they
delete or write outside cwd; hard_deny for secret exfiltration, prod
deploy, disarming guardrails (history rewrite NOT retained, so a `rebase`
then an ordinary push stays uncovered); extend the static deny family to
the `.env` readers; `classifyAllShell` stays false; intent clears a soft
block for the CURRENT TURN only.
- [x] S1 `permissions.deny` +10 reader rules (sed awk cut tr sort uniq
diff od xxd strings vs `.env*`) — 6 of them were in `allow`
- [x] S2 `autoMode.soft_deny` — 7 rules + the intent-scope line
- [x] S3 `autoMode.hard_deny` — 3 rules, "adding a restriction is fine,
removing one is not"
- [x] S4 `SETTINGS.md` — tier-choice table + scope-of-intent section
- [x] S5 CHANGELOG — Changed rewritten, new Security block
- [x] S6 CONSEQUENCE confirmed by user 2026-09-15: the hard_deny guardrail rule means I can
no longer edit a deny/soft_deny/hard_deny list to REMOVE an entry.
Tightening stays allowed. Future permission loosening goes through
`/permissions` or the user's own edit.
### Third pass (2026-09-15) — F1-F3 done + graphify untracked
Worst finding was not the duplication: local `deny` still carried
`rsync` `kill -9` `killall` `pkill`, the four the user moved OUT of
global deny. deny wins across sources, so `autoMode.soft_deny` was a
dead letter in THIS repo. Local `allow` also held `sed *`, `cp *`,
`python3 -` — an allow rule short-circuits the classifier, punching a
hole through the same soft_deny rules.
- [x] G1 `skills/graphify/{SKILL.md,references/,.graphify_version}`
gitignored + `git rm --cached`. Written by `graphify claude
install` since `~/.claude/skills` symlinks to `skills/`; a fresh
clone gets them from `make plugin`. `test-prompts.json` is
hand-written for darwin, stays tracked. Trade-off documented in
CLAUDE.md: an upstream release can now change the skill prompt
with no diff to review.
- [x] G2 `.claude/settings.local.json` 14.6 KB -> 6.2 KB. deny + ask
dropped whole, allow 185 -> 98 (81 duplicates of the global, 6
policy conflicts: `sed *`, `cp *`, `python3 -`,
`Read(//home/bchanot/**)`, `WebSearch`, a leftover injection-test
payload). Every non-`permissions` key was a verbatim copy of the
global, `hooks` included. Backup: `.audit/settings.local.json.bak-*`
(gitignored, the file itself is not in git).
### Follow-up found while doing this (fixed in the third pass above)
`.claude/settings.local.json` (gitignored, 14.6 KB) is a near-complete
shadow copy of the global `settings.json` at a HIGHER precedence tier:
185 allow / 30 ask / 106 deny, plus its own `cleanupPeriodDays`,
`attribution`, `statusLine`, `enabledPlugins`, `extraKnownMarketplaces`,
`effortLevel`, `remoteControlAtStartup`, `inputNeededNotifEnabled`,
`skipAutoPermissionPrompt` — all identical to the global today, so the
duplication is invisible until the global drifts, which it just did
(no `autoMode`, 106 deny vs 116). It defeats the config-guard premise
(hand-curated `settings.json`) with a file nobody reviews.
- [x] F1 `WebSearch` sits in global `ask` and in local `allow` — in this
repo it never reaches the ask tier. Intended or drift?
- [x] F2 local `hooks` block registers `bash ~/.claude/hooks/config-protection.sh`
on PreToolUse/Bash. That script does not exist, in `hooks/` or in
`~/.claude/hooks/`. Dead hook firing on every Bash call here.
- [x] F3 decide: prune the local file down to the session-accumulated
allow rules only, dropping every key that merely restates the
global, or keep the copy deliberately and document why.
## 2026-08-25 — darwin fresh baseline: 32 skill-systems + 23 agents (feature/darwin-optimize-20260825)
User: `/darwin-skill all skills and agents` (background). Fresh-from-zero
(results.tsv wiped 2026-06-23, journal 2026-06-30). Scope per BDR-015/043 +
LRN-070: personal skills only, external/gstack OUT. EVAL-004 applied: eval
unit = skill+dispatched-agents SYSTEM, agents get own rows. LRN-018: judges
emit per-dim scores, totals recomputed main-thread. v2.1 keep/revert =
paired same-judge majority, absolute scores triage-only.
- [x] T1 Phase 0+0.5: gitflow branch, results.tsv header, 7 new
test-prompts.json (capitalize deploy gitflow pdf-translate reconcile
release-candidate tour), runtime scan (2 minor hits). find-docs
EXCLUDED — machine-owned ctx7 (BDR-053, gitignored) → 31 systems.
- [x] T2 Phase 0.5 gate PASSED: reuse prompts as-is; dim8 full_test on
candidates only (baseline dry_run); Phase 2 set = ALL units <80.
- [x] T3 Phase 1 baseline DONE: 7 blind judges, 54 rows (31 skills + 23
agents), mean 83.4, 13 units <80, ~25 verified findings (hotfix
destructive restore, onboard/onboarder contract, init-project
allowed-tools, skills-perso 8/32 detection...).
- [x] T4 Phase 1 gate PASSED: user picked the set — proven by Phase 2
running 13/13 units, 0 reverts (DARWIN-2026-08-26.md:23).
Ticked by reconcile 2026-09-01.
- [x] T5 Phase 2 DONE: 13/13 units, 12 rounds kept 3-0, 0 reverts +
bug pass 8 commits kept 3-0 (2 skeptic residuals amended). make test
green.
- [x] T6 Phase 3 DONE: report .claude/audits/DARWIN-2026-08-26.md + card
PNG (playwright fallback). Capitalize pending user approval. Branch
UNMERGED — human gate.
→ both residuals stale: capitalized a15854a, merged 726464f
(reconcile 2026-09-01).
## 2026-08-25 — user permanent rules: writing + web build + web security (feature/user-writing-web-rules)
User supplied 4-block rule text (écris / site / code / vérification); asked:
coverage check, conflict check, integrate. Verdict: security CORE already in
CLAUDE.global.md §Security (parameterized queries, env-var secrets,
AuthN/AuthZ, fail closed) — NOT duplicated. NEW: writing-style block, design
anti-default list, site done-checklist, web-app specifics (RLS, service key,
IDOR, cookie flags, rate limit, field minimization). Placement: global at
308/320 budget → rules/ instead.
- [x] R1 rules/writing-style.md — always-on (no paths:), scope carve-outs
(registries caveman, code comments, skill templates) + self-check
- [x] R2 rules/web-building.md — paths: web globs; anti-defaults + done
checklist (report missing, never invent) + skill pointers
- [x] R3 rules/web-security.md — paths: code globs; web-app specifics
extending §Security, zero dup of the core
- [x] R4 CLAUDE.md (project) — amend always-on doctrine line (320-budget
exception → rules/), feeds C2 audit
- [x] R5 capitalize BDR-085 + journal + CHANGELOG
- [x] merge → develop 5ec7bfa — human gate passed (reconcile 2026-08-25)
## 2026-07-30 — adapt config for Claude 5 family / Opus 5 (feature/opus5-config-tuning)
User: Opus 5 "needs more freedom" → research (official migration guide +
web + registres) confirms: over-delegates (inverts LRN-030 Opus 4.8 trait),
over-verifies if told to verify, literal instruction following, scope
expansion named regression, harness already injects anti-delegation on
Opus 5 (#80988). Plan: .claude/tasks/plans/2026-07-30-opus5-config-tuning-1238.md
— to be challenged by 3 blind plan-challengers (opus pins → Opus 5), then
executed on feature branch. NO merge (human gate).
Challenged 2026-07-30: correctness CONCERNS(4) · robustness FATAL(5, 1
BLOCKER: symlink-live deployment) · simplicity CONCERNS(4) — all fixes
adopted as prescribed (plan §5bis, v2 items below).
- [x] W0 branch first (eab2a10 parent); hook regex validated on scratch copy
(bash -n + shellcheck + 5 replays, HOME sandboxed) before live write
- [x] W1 delegation block v2 (when-guidance + gates carve-out + scoped don't-redo) — 0f7b565
- [x] W2 "staff engineer" bar line deleted — 0f7b565
- [x] W3 finish-whole-task folded into Deviations (+ gone-WRONG→STOP) — 0f7b565
- [x] W4 deliverable-length rule — 0f7b565
- [x] W5 line budget: 308/320
- [x] W6 hook \bux\b dropped, \bui\b kept + F10 must-fire lock, D11 quiet row
flip-tested (fire before/quiet after) — eab2a10, suite 22/0
- [x] W7 plan-challenger :82-83 reworded → [MINOR] routing, census row — c3d3f4d, 44/0
- [x] W8 BDR-081 + LRN-139 + journal + CHANGELOG
- [x] W9 final gate: make test full suite — green except known T6c
(darwin-skill residual → chantier 4 below), 2026-07-30
- [x] W10 merged on explicit user signal — 709cf9b (2026-07-30 13:28),
branch deleted; confirmed post-merge this session
## 2026-07-30 — Claude 5 follow-on chantiers (user directive, checkpoint between each)
Order fixed, one branch per chantier, no merge without per-chantier signal.
- [x] C1 dé-prescription seo-analyzer.md + geo-analyzer.md — DONE 2026-08-02.
Census-first 71 locks flip-proven (9681b46) → rewords under
audience×range invariant (adafa35 seo, c7646a9 geo) → controlled
before/after dogfood: judge-replay on frozen signals + templates +
fresh collects + e2e judge + blind reader = 42/42 both sets, zero
contract regression, recall improved. Plan challenged 4 passes
(FATAL/FATAL/CONCERNS + confirmation FATAL(9), all closed by name).
BDR-082 + LRN-140. Evidence .audit/dogfood-baseline/ (19 artifacts).
Branch feature/seo-geo-deprescription UNMERGED — human gate.
→ merged 5488c48, branch deleted (reconcile 2026-08-25).
Residual for gate: §6bis dynamically-unverified list (FULL branches,
apply path — census-locked statically); FULL/aggressive dry-run = user
option; nested-CLI dogfood blocked by monthly spend limit (inline used).
- [ ] C2 self-contradiction audit CLAUDE.global.md + own skills: list rule
pairs in tension, propose resolution per pair, apply after user OK.
/doctor as assistant, not authority.
- [ ] C3 superpowers: MEASURE first (skill-invocation log over sessions)
whether "1% chance → MUST invoke" over-triggers; if yes, options +
trade-offs (disable plugin / softer house rule / live with) — user decides.
- [x] C4 hygiene: reinstall darwin-skill — DONE (reconcile 2026-08-25:
~/.agents/skills/darwin-skill present, T6c green, make test exit 0).
## 2026-07-22 — auto-purge transient superpowers artifacts at finish (feature/gitflow-auto-purge-transient)
User: transient planning artifacts (`docs/superpowers/{specs,plans}`) leak into
develop; BDR-065 "post-merge cleanup" is DOCTRINE ONLY (no code) — manual chore,
@@ -19,15 +407,18 @@ versioned (durable, referenced by decisions.md e.g. BDR-076). Universal via the
- [x] Gate: shellcheck lib/*.sh CLEAN + `make test` exit 0 (gitflow 106/0, full
suite green). Universal via ~/.claude/lib → repo lib symlink (verified).
- [x] CLAUDE.md §Transient planning artifacts: → "AUTO-PURGED by gitflow finish".
- [ ] Capitalize: BDR-065 amendment (delete side now automated) + LRN — pending user OK.
- [x] Capitalize: BDR-065 Amendment (2026-07-22) in body + LRN-138 present
(reconcile 2026-08-25).
## 2026-07-20 — pending merge gates (reconcile)
- [x] merge feature/profile-managed-externals → develop (BDR-079 profile
symmetry + /doc clean pass: README/USAGE/ARCHITECTURE.md) — 37c79f0
- [x] merge chore/purge-transient-docs → develop (docs/ transient purge
655e364 + reconcile e75ea79) — reaches main at next release
- [ ] Makefile help text: profiles 5/10 listed (:57) + test glob missing
run-*.sh (:31) — 2-line hotfix (flagged by /doc audit)
- [ ] Makefile help text: profile-list help lists 5/10 profiles (:57) —
1-line hotfix. (test glob :31 FIXED — has run-*.sh, reconcile 2026-08-25)
Re-verified OPEN 2026-09-01: lib/profiles/ has 10, Makefile:57 lists 5
(backend, full, seo, web-full, web missing).
## 2026-07-20 — profile ↔ toggle-external symmetry (feature/profile-managed-externals, BDR-079)
Audit verdict: gstack on-demand + design enable already work; DISABLE side
@@ -1094,3 +1485,78 @@ branch) → LOT3 mis-merge trap; + 3 doctor false-warns (LRN-047 class).
comment anchored to measured ~11.4k (LRN-088). False "92% CRITICAL" → ~5% comfortable.
- [x] Verify — suites green (71/13/32/19/20/13 + RC 5/5); doctor 0 false-warn; shellcheck clean.
+docs(changelog) Unreleased entry (706abff). Gate passed on GO 2026-07-03. Finish pending.
## 2026-08-24 — contract gates: plancher déterministe (feature/contract-gates)
Source: analyse du skill `unlazy` (Leonxlnx/unlazy, 2.1.0). Verdict: son
architecture de vérification n'apprend rien (contrat+verifier frais+boucles
bornées ⊂ déjà en place). Le trou réel: **entre l'exécuteur et GATE 1 il n'y a
aucun plancher déterministe** — GATE 1 est un dispatch LLM, et `PROOF:` est une
ligne que le verifier ÉCRIT (rien ne l'empêche structurellement de la produire
sans rien exécuter). Palier 2 retenu (user, 2026-08-24).
PRIS d'unlazy: critère porteur d'oracle exécutable (CHECK/EXPECT/EVIDENCE),
fail-closed (exit 0 ET marqueur), evidence pending = NOT-MET, `ABANDON: <id>
<raison>` comme handoff visible non supprimable, les 4 règles d'écriture de
gates falsifiables, la discipline 4 passes.
REFUSÉ: Stop hook `decision:"block"` (contredit "STOP + escalade humaine" et
"merge sur signal humain"), approval store `~/.unlazy/approved` (résout
l'exécution de ledgers hérités non fiables — pas notre menace), arbre
`.unlazy/<scope>/` (4e arbre de bookkeeping ⇒ mort de la config), `tree N`
(désavoué par ses propres docs), le checker Node 28k (stack lib = 100% bash,
Health Stack = shellcheck).
- [x] W0 branche feature/contract-gates depuis develop (via lib/gitflow.sh)
- [x] W1 `lib/gates.sh` — parse ACCEPTANCE CRITERIA, exécute fail-closed
(exit 0 ET EXPECT), réécrit EVIDENCE dans le contrat. Sous-commandes
`run` (exécute+écrit) / `status` (parse seul, jamais d'exécution, jamais
d'écriture). rc 0=MET · 2=UNMET/malformé · 3=ABANDONED.
- [x] W2 `lib/contract-interview.md` — STEP 3 gagne CHECK/EXPECT/EVIDENCE
optionnels par critère + les 4 règles de falsifiabilité; template mis à
jour; ABANDON dans Lifecycle; ligne de poids par flow.
- [x] W3 `agents/verifier.md` — EVIDENCE fail-closed (coché+pending = NOT-MET),
bucket ABANDONED, verdict `CONFORME` impossible si abandon présent.
- [x] W4 `lib/verify-secure-loop.md` — GATE 0 déterministe avant GATE 1
(rouge ⇒ re-dispatch exécuteur sans brûler un verifier).
- [x] W5 `agents/feater.md` + `agents/bugfixer.md` — discipline 4 passes.
- [x] W6 `lib/tests/gates.test.sh` — comportemental sur gates.sh (fail-closed,
exit≠0 avec marqueur = FAIL, pending, ABANDON, malformé, status
n'exécute pas) + locks de structure sur W2/W3/W4/W5.
- [x] W7 shellcheck + bash -n + `make test` complet.
- [x] W8 CHANGELOG + registres (BDR + LRN + journal).
- [x] W10 restatements skills : bullet GATE 0 dans feat/bugfix/ship-feature/
init-project (+4 locks, flip-testé) ; ligne hotfix du tableau de poids
corrigée (aucun floor à ce poids). 2026-08-24.
- [x] W11 RED comportemental : 16/16 runs frais non-amorcés conformes
(verifier ×9, feater ×2, orchestrateur ×5) → EVAL-027. 2026-08-24.
- [x] W9 merge sur signal humain explicite (2026-08-24, "merge dans develop").
**Won't-build-now — Palier 3 unlazy (OWNS/leases), trigger documenté :**
Différé volontairement (BDR-083) : tous les dispatches parallèles actuels
sont read-only — le problème (2 exécuteurs ÉCRIVAINS concurrents) n'existe
pas. Pattern [[LRN-080]] : ne pas construire sans menace mesurée.
TRIGGER = le jour où un flow dispatche ≥2 exécuteurs écrivains en parallèle :
(1) FILE SCOPE du contrat = déclaration OWNS (champ existant, zéro format
neuf) ; (2) ~40 l dans gates.sh ou lib/owns.sh — intersection CONSERVATRICE
des FILE SCOPE des contrats actifs avant fan-out, conflit possible → refus +
dispatch séquentiel (pas de locks disque tant que l'orchestrateur est
unique) ; (3) locks + tests.
## 2026-08-24 — tour multi-projets en parallèle (feature/tour-parallel)
User (gate 2026-08-24): "tout paralléliser (option 2) mais bien garder la
sélection des modèles — orchestrateur garde le modèle orchestrateur, les
skills/agents suivent leurs orchestrateurs définis". Preuve mécanique
préalable: probe imbriquée 3 sous-agents, fenêtres chevauchantes, 9.1s vs
~18s séquentiel. Dérogation LRN-083 (boucle de fix par projet déplacée dans
un runner dispatché) → à consigner BDR-084. Repos indépendants, branches
chore par repo, report-as-approval-gate ⇒ rien de partagé n'est décidé
dans un runner; capitalize reste main-loop.
- [x] T1 skills/tour/SKILL.md — STEP 0 routé (1 projet = inline inchangé;
≥2 = fan-out) + STEP 0b: un runner general-purpose par projet, TOUS
dans UN message, SANS pin modèle (hérite session, model-gate déjà
passé); agents internes gardent leurs tiers définis; runner mort =
ligne RUNNER FAILED, jamais absent silencieux; capitalize main-loop.
- [x] T2 locks census §12 dans lib/tests/model-routing.test.sh (fan-out
présent, runner non-pinné, single message, capitalize main-loop).
- [x] T3 BDR-084 + CHANGELOG + journal.
- [x] T4 make test rc 0 + shellcheck clean (SC2016 silencé, littéral
voulu). Merge NON fait — gate humain.
@@ -0,0 +1,115 @@
# CONTRACT — gstack-playwright-lib
- date: 2026-09-13 | flow: feat | branch: feature/gstack-playwright-lib
- status: active
## REQUEST (verbatim — IMMUTABLE)
Message 1:
> l'installation de chromium, c'est une version fix ou en latest ? Il faudrait mettre en lateste, et d'ailleurs son update est pris en compt dans l'update ? quelq version a besoin gstack ? Ca serait pas plus simple d'installer perplexity a la place ?
Message 2 (after the assistant proposed fix A + fix B):
> les deux
Message 3 (answer to the scope question on fix B, after the "654 Mo orphelins"
premise was proven wrong):
> Check read-only dans doctor.sh
## CLARIFICATIONS
Q: "mettre en latest" — pin Chromium to latest?
A: Not actionable as asked. Playwright downloads the browser revision its own
version pins (1.61.1 → chromium 1228); the CDP client is coupled to that
build. "Latest" = track the latest Playwright, which is what BDR-029's bump
already does. No change to the pinning mechanism is in scope.
Q: Volet B — purge the orphan Playwright revisions?
A: Superseded by evidence. `~/.cache/ms-playwright/.links/` registers THREE
playwright installs (gstack 1.61.1 → rev 1228; gsd-pi nvm 1.61.0 → 1228;
gsd-pi ~/.local 1.63.0 → 1243). Every directory on disk is referenced;
zero bytes reclaimable. Playwright already GCs correctly on every
`install` (`_deleteStaleBrowsers`, unions across all registered installs).
User chose: read-only report in doctor.sh, NO deletion anywhere.
## ACCEPTANCE CRITERIA
1. `lib/gstack-playwright.sh` exists, is source-safe (sourcing prints nothing
and runs no side effect), and its verb dispatcher works when executed.
CHECK: out=$( . lib/gstack-playwright.sh; echo READY ); [ "$out" = READY ] && bash lib/gstack-playwright.sh 2>&1 | grep -q 'usage:' && echo LIB_OK
EXPECT: LIB_OK
EVIDENCE: MET exit=0 marker-found :: LIB_OK
2. The bump logic lives ONLY in the lib: `install-plugins.sh` no longer
defines `gstack_bump_playwright_if_unsupported`, sources the lib instead,
and still calls it BEFORE gstack `./setup` (BDR-029 behavior unchanged:
OS-gated, idempotent, non-fatal).
CHECK: grep -q '^gstack_bump_playwright_if_unsupported() {' install-plugins.sh && exit 1; grep -q 'lib/gstack-playwright.sh' install-plugins.sh || exit 1; c=$(grep -n 'gstack_bump_playwright_if_unsupported' install-plugins.sh | grep -v ':[[:space:]]*#' | tail -1 | cut -d: -f1); s=$(grep -n '&& \./setup)' install-plugins.sh | head -1 | cut -d: -f1); [ -n "$c" ] && [ -n "$s" ] && [ "$c" -lt "$s" ] && echo EXTRACT_OK
EXPECT: EXTRACT_OK
EVIDENCE: MET exit=0 marker-found :: EXTRACT_OK
3. `update-all.sh` delegates the gstack submodule update to the lib
(`gstack_submodule_update_with_bump`) instead of calling
`git submodule update --remote` bare, so the bump is re-applied after every
successful update.
CHECK: grep -q 'gstack_submodule_update_with_bump' update-all.sh && grep -q 'lib/gstack-playwright.sh' update-all.sh && ! grep -qE '^[[:space:]]*if git submodule update --remote skills-external/gstack' update-all.sh && echo WIRED_OK
EXPECT: WIRED_OK
EVIDENCE: MET exit=0 marker-found :: WIRED_OK
4. [gated 2026-09-15] `gstack_submodule_update_with_bump` NEVER modifies the
submodule working tree. On a successful `git submodule update --remote` it
re-applies the bump; on failure it returns non-zero, touches nothing, and
emits git's own message plus a hint naming the local Playwright bump when
`package.json`/`bun.lock` are the dirty files. The conflict-RECOVERY branch
of the earlier revision (discard, retry, backup, restore) is withdrawn: it
could leave the bump discarded and un-reapplied, regressing a working
browser into BLK-008, which the pre-existing behavior never did.
CHECK: sed 's/#.*//' lib/gstack-playwright.sh | grep -qE 'git [^|;]*(checkout|reset|clean|stash)' && exit 1; bash lib/tests/gstack-playwright.test.sh 2>&1 | grep -q 'update-conflict' && bash lib/tests/gstack-playwright.test.sh 2>&1 | grep -qE '^PASS=[0-9]+ FAIL=0$' && echo NONDESTRUCTIVE_OK
EXPECT: NONDESTRUCTIVE_OK
EVIDENCE: MET exit=0 marker-found :: NONDESTRUCTIVE_OK
5. `doctor.sh` prints a Playwright-browsers section: total cache size, one line
per browser directory naming the registered playwright install(s) that
reference it, plus counts of unreferenced directories and broken links.
CHECK: bash doctor.sh 2>/dev/null | grep -qi 'playwright browsers' && bash lib/gstack-playwright.sh browsers-report | grep -qE 'chromium-[0-9]+' && bash lib/gstack-playwright.sh browsers-report | grep -qi 'unreferenced' && echo REPORT_OK
EXPECT: REPORT_OK
EVIDENCE: MET exit=0 marker-found :: REPORT_OK
6. The report is provably read-only: no destructive verb anywhere in the lib,
and the cache directory listing is identical before and after a report run.
CHECK: sed 's/#.*//' lib/gstack-playwright.sh | grep -qwE '(rm|rmdir|unlink|truncate|mv)' && exit 1; b=$(ls -la ~/.cache/ms-playwright ~/.cache/ms-playwright/.links 2>/dev/null | cksum); bash lib/gstack-playwright.sh browsers-report >/dev/null 2>&1; a=$(ls -la ~/.cache/ms-playwright ~/.cache/ms-playwright/.links 2>/dev/null | cksum); [ "$b" = "$a" ] && echo READONLY_OK
EXPECT: READONLY_OK
EVIDENCE: MET exit=0 marker-found :: READONLY_OK
7. `lib/tests/gstack-playwright.test.sh` exists, passes, and covers at least:
bump skipped when the OS tag is already supported; bump fired when it is
not; submodule-update conflict recovery; browsers-report on a fixture cache
holding a referenced revision, an unreferenced one and a broken link.
CHECK: bash lib/tests/gstack-playwright.test.sh | tail -1 | grep -qE '^PASS=[0-9]+ FAIL=0$' && echo TESTS_OK
EXPECT: TESTS_OK
EVIDENCE: MET exit=0 marker-found :: TESTS_OK
8. shellcheck clean on every touched shell file.
CHECK: shellcheck lib/gstack-playwright.sh lib/tests/gstack-playwright.test.sh install-plugins.sh update-all.sh doctor.sh >/dev/null 2>&1 && echo SHELLCHECK_OK
EXPECT: SHELLCHECK_OK
EVIDENCE: MET exit=0 marker-found :: SHELLCHECK_OK
9. [gated 2026-09-15] (judgement) No new dependency; the report degrades
silently when `~/.cache/ms-playwright` is absent, when its `.links`
directory is absent, when `PLAYWRIGHT_BROWSERS_PATH` is `0` or not a
directory, or when no playwright install is registered — doctor must stay
green on a machine that never installed a browser. The lib's printers are
named `_gspw_ok`/`_gspw_warn`/`_gspw_info` and it defines NO bare
`ok`/`warn`/`info`/`pass`/`fail`: doctor.sh sources the lib before every
check, so bare names would override its own printers and silently
disconnect its `ERRORS`/`WARNS` counters.
## FILE SCOPE
- lib/gstack-playwright.sh (new)
- lib/tests/gstack-playwright.test.sh (new)
- install-plugins.sh (remove inline fn, source + call lib)
- update-all.sh (call bump + conflict recovery)
- doctor.sh (new read-only report section)
Out of scope: the gstack submodule itself, the pinning mechanism, any
deletion of cached browsers, the `GSTACK_CHROMIUM_NO_SANDBOX` layer
(LRN-040 layer 2, unchanged).
@@ -0,0 +1,255 @@
# PLAN — Adapt claude-config for the Claude 5 family (Opus 5 focus)
Date: 2026-07-30 · Branch (planned): feature/opus5-config-tuning (off develop)
KIND: build-plan · Author: main-loop session (Fable 5)
## 1. Context & evidence
Opus 5 (`claude-opus-5`, released 2026-07-24) now backs every `model: opus`
agent pin in this repo (analyzer, plan-challenger, seo/geo-analyzer,
plugin-advisor — BDR-076/077) and any session the user switches to via
`/model opus`. Its documented behavioral profile differs from Opus 4.8 in
ways that make parts of this config counterproductive:
- E1 **Over-delegation**: Opus 5 "delegates to subagents more readily than
prior models" (official prompting guide). Opus 4.8 had the OPPOSITE trait
(LRN-030), and `CLAUDE.global.md:43-47` was written to counter it
("Counters model tendency to under-delegate"). The premise is inverted.
- E2 **Anti-delegation already injected by the harness**: Claude Code
v2.1.219 server-gates an Opus-5-only prompt section (`heron_brook` +
`subagent_steer_delegation`, GitHub issue #80988) that says "Do not call
the AgentTool unless the user requested it" and "Subagents multiply cost
and time…". Stacking our own hard cap on top would triple-constrain;
keeping a pro-delegation nudge would fight the injection. Model-neutral
when-guidance is the stable middle.
- E3 **Over-verification**: official guidance — "If your prompt contains
explicit verification instructions … remove them: instructions like these
cause over-verification on Claude Opus 5, and removing them reduces wasted
tokens with no loss in quality." Also true of per-prompt "double-check"
phrasing. Targets PROSE told to the model, not harness-level gates.
- E4 **Scope expansion**: named Opus 5 regression ("can expand the scope of
a task, adding steps that weren't requested"). Anthropic ships a literal
counter-block; tested to reduce scope changes "to nearly zero".
- E5 **Literal instruction following** (since 4.7, stronger now): aggressive
MUST/CRITICAL language over-triggers; conservative-reporting instructions
("only report high-severity") measurably depress recall in review/challenge
harnesses.
- E6 **Longer written deliverables**: files written to disk run ~30-40%
longer; `effort` does NOT control visible/deliverable length — only prose
instructions do.
- E7 **Overconstraint costs reasoning**: Anthropic removed >80% of Claude
Code's system prompt for Claude-5-generation models "with no measurable
loss"; named mechanism = tokens burned resolving conflicting rules.
- E8 **Hook false positive (today)**: `\bux\b` in
`hooks/design-toolchain-reminder.sh:47` fired on French prose ("changement
ux vu" — matches after apostrophe/slash/space); 2nd `ux` FP in the log,
both French. Continues the LRN-1005/1007 false-positive series. No test
row covers `\bui\b`/`\bux\b`.
- E9 **Effort carry-over trap**: Opus 5 has no model-default effort hold in
Claude Code — a persisted `xhigh` (our `settings.json:333`) silently
carries onto Opus 5 sessions, against Anthropic's "start at high, sweep
low/medium" guidance for that model.
## 2. Design decisions
- D1 The global instruction layer must be MODEL-NEUTRAL across the Claude 5
family (sessions run Fable 5 by default; dispatched judgment agents run
Opus 5; executors Sonnet). Fixes therefore express WHEN-guidance and
outcome bars, not directional compensation for one model's trait.
- D2 Harness-level quality gates (fresh blind verifier + security-auditor,
BDR-049/050; plan-challenge, BDR-075) are architecture, not model
self-check prompting. They stay. E3 applies only to prose that tells the
MODEL to verify its own work.
- D3 Per BDR-021, the Security and Architecture-decisions sections of
CLAUDE.global.md stay verbatim (deliberate policy). No softening there.
- D4 Registries are append-only: LRN-030 is not edited; a new LRN records
the trait inversion and points back to it.
- D5 Deterministic backstops (gitflow pre-commit, Gitea protection,
permissions.deny, rtk pinning) are explicitly out of "more freedom" scope
— community reports show Opus 5 working AROUND soft controls, which argues
for keeping hard ones.
## 3. Work items
### W1 — CLAUDE.global.md: rewrite the delegation block (:43-47)
Replace the 5-line block (incl. "Default to delegation for multi-file
exploration. Counters model tendency to under-delegate.") with model-neutral
when-guidance, same footprint (≤5 lines):
```
- Sub-agents: one task per sub-agent, main context stays clean.
Delegate genuinely independent, sizeable tracks (wide multi-file
exploration, parallel audits) — not work doable in a few tool
calls, and not self-verification (harness gates own that). Brief
precisely, then commit to the delegation — don't redo its work.
```
Rationale: E1+E2. No hard spawn cap in prose (harness already injects one on
Opus 5; Fable benefits from delegation).
### W2 — CLAUDE.global.md: reframe "After code changes" (:75-83)
Keep the concrete quality bar; drop the proof-mandate/self-check phrasing
(E3). Replace steps 2-4 with faithful-outcome reporting:
```
## After code changes
1. Run tests, lint, build, type-check if available.
2. Report outcomes faithfully: what passed, what wasn't run,
remaining risks, surviving deviations. Completion claims only
for verified work.
3. Correction or notable event → capitalize to right registry.
```
Net: -2 lines. "Would staff engineer approve?" bar and "Don't mark complete
without proof" are removed as self-check choreography; honest-reporting
line preserves the intent (grounded completion claims) without mandating an
extra verification pass.
### W3 — CLAUDE.global.md: add scope fence (Workflow section)
Append (adapted from Anthropic's tested block, caveman-compressed, ~5 lines):
```
- Scope: deliver what was asked, at the scope intended. Routine
judgment calls → decide alone; materially different readings →
ask. Better approach spotted → say so in one line, still do the
task as asked. Finish the whole task; genuinely blocked → do the
rest, state plainly what's missing.
```
Rationale: E4. Complements existing "Scope changes to task — no unrelated
edits" (line ~15) without contradicting it.
### W4 — CLAUDE.global.md: add deliverable-length rule (Code style / Comments area)
~2 lines:
```
- Written deliverables (docs, reports, .md): length matched to what
the task needs — no filler sections, no boilerplate summaries.
```
Rationale: E6. Registries already covered by caveman rule.
### W5 — Line budget
After W1-W4: expected ~309 lines. Hard check: `wc -l CLAUDE.global.md` ≤ 320
(session-start.sh warning threshold at :202-213).
### W6 — hooks/design-toolchain-reminder.sh: drop `\bui\b` and `\bux\b`
- Remove the two 2-char alternatives from the pattern at :47. Keep
`ui/ux|ux/ui|ui kit` and all other tokens.
- Add a dated header comment (3rd tightening pass, 2026-07-30, cites the
two French-prose `ux` FPs; series LRN-1005/1007).
- Trade-off accepted: a bare "améliore l'ux" prompt with no other design
token goes quiet — the CLAUDE.global.md "Design work" section still
routes it (the hook is a belt, self-described soft nudge).
- Update `lib/tests/design-toolchain-reminder.test.sh`: add 2 quiet rows
(the real FP prompt excerpt; a bare "l'ui" French sentence) — flip-tested
per LRN-096. Existing 9 must-fire rows unaffected (none uses ui/ux).
### W7 — agents/plan-challenger.md: coverage-first reporting line
Add one clause to the findings rules (add-only, no removal): uncertain or
low-severity findings are REPORTED with an explicit confidence + severity
tag rather than self-censored — severity filtering happens in the
orchestrator's synthesis, not in the challenger. Rationale: E5 (literal
Opus 5 + "manufactured concern is a failure" wording risks suppressing real
low-confidence findings). Must not touch: verdict grammar, MANDATORY PROOF
clause, blind-dispatch rules (test-locked in plan-challenger.test.sh).
### W8 — Memory + docs capitalization (same branch, follows the work)
- decisions.md: new BDR (config adapted for Claude 5 family — scope,
rationale, alternatives incl. "leave config as-is" and "hard spawn caps"
rejected).
- learnings.md: new LRN — Opus 5 behavioral profile (over-delegation
inverts LRN-030's Opus 4.8 trait; over-verification; literal following;
no effort hold on Opus 5 in Claude Code; heron_brook/#80988 injection).
- journal.md: one line.
- CHANGELOG.md: entry under Unreleased.
### W9 — Gates (before commit)
- `shellcheck hooks/design-toolchain-reminder.sh` clean.
- Manual flip-test of the hook: FP prompt → quiet; "redesign the navbar" →
fires.
- `make test` full suite green (design-toolchain-reminder.test.sh,
plan-challenger.test.sh, model-routing.test.sh untouched-but-must-pass,
curated-config-guard, loops-light…).
- `wc -l CLAUDE.global.md` ≤ 320.
### W10 — Gitflow
`bash ~/.claude/lib/gitflow.sh start feature opus5-config-tuning` off
develop; atomic commits (hook+test / CLAUDE.global.md / agent / memory+docs);
NO `gitflow finish` — merge only on explicit human signal.
## 4. Explicitly NOT doing (considered, rejected)
- N1 Touching lib/verify-secure-loop.md or the fresh-verifier/security
gates: harness architecture (BDR-049/050, D2), verifies SONNET executor
output — not Opus 5 self-check prose.
- N2 Softening the Security / Architecture sections (BDR-021, D3).
- N3 Editing the superpowers plugin's "1% chance → MUST invoke" language:
external upstream code; flagged as residual over-triggering risk in the
new LRN, revisit as its own decision if observed.
- N4 Changing `settings.json` `effortLevel: "xhigh"`: user preference,
optimal for the Fable 5 session default; the Opus 5 carry-over trap (E9)
is documented in the LRN + surfaced to the user for a manual decision.
- N5 De-prescribing seo-analyzer.md / geo-analyzer.md (1528/1106 lines,
heavy MUST density): separate project, backlog note in TODO.md.
- N6 Removing or session-gating the design/ctx7 reminder hooks: soft
nudges, cheap, deliberately built; tightened only (W6).
- N7 Any model pin change: `model: opus` pins now resolve to Opus 5 —
desired outcome, census (model-routing.test.sh) untouched.
- N8 Committing settings.json for any reason (LRN-098/1049 /model-churn
trap): file is currently clean; keep it out of every commit.
## 5bis. CHALLENGE SYNTHESIS (2026-07-30) — FINAL amendments (v2)
Verdicts: correctness CONCERNS(4) · robustness FATAL(5, 1 BLOCKER) ·
simplicity CONCERNS(4). Every fix below is the challenger's own named FIX,
adopted as written. No re-challenge pass: scope narrowed, no new dependency;
W0 is an execution-time safety procedure, not a new config mechanism.
- **W0 (NEW — robustness BLOCKER)**: all edited surfaces are symlink-deployed
LIVE (~/.claude/CLAUDE.md, hooks/, agents/ → this repo); edits take effect
machine-wide at save time, before any W9 gate. Mitigations:
(a) `gitflow start` BEFORE any live-file edit; never checkout develop
mid-work; (b) hook regex change validated on a SCRATCH copy first
(bash -n + shellcheck + pattern replay), then written to the live file in
ONE atomic Edit; (c) named reverts: `git show develop:<file> > <file>`;
escape hatch = remove the hook registration block from settings.json.
- **W1 v2** (robustness#3, correctness#2): replacement text carves out the
mandated gates explicitly and scopes "don't redo":
"Skill-mandated gates (fresh verifier/security/challenge) always dispatch
as written. Don't redo delegated work by hand — failed gates re-dispatch
fresh executors instead."
- **W2 v2** (simplicity#2): minimal diff — delete ONLY the line
`Bar: "would staff engineer approve?"`. Steps 1-4 + capitalize step stay.
- **W3 v2** (simplicity#1, robustness#4): no new bullet. Fold the only new
clause into the existing Deviations bullet: "Finish the whole task:
blocked on an independent sub-part → do the rest, state what's missing.
Gone WRONG → still STOP, re-plan." (net +2 lines, no conflict with :53).
- **W4**: unchanged (+2 lines). Budget v2: 304 +1 −1 +2 +2 = 308 ≤ 320.
- **W6 v2** (all lenses): drop `\bux\b` ONLY — keep `\bui\b` (zero evidenced
FP; one logged true positive). Accepted trade-off: the 2026-07-21 "ameliore
le tutoriel…gamifier" ux row (plausible TP) goes quiet; CLAUDE.global.md
design-routing section remains the router. Header comment notes the log
records `head -1` only → per-token FP rate not fully derivable. Tests:
quiet row = synthetic "changement ux vu…" (verified matches pre-change →
flips); must-fire row = "revois l'ui du panneau admin" (locks `\bui\b`;
apostrophe escaped correctly, doubles as JSON-path control per
robustness#7). No log-excerpt rows (vacuous — 100-char truncation).
- **W7 v2** (all lenses): in-place reword of the `:82-83` sentence (NOT
test-locked; plan v1 misstated that) instead of an add-only clause:
"No invention — ungrounded is noise. Silently dropping a grounded doubt is
equally a failure: file it as `[MINOR]` with the uncertainty stated in
`WHY:`. Nothing real at all → `SOLID` with `FINDINGS: none`."
OUTPUT grammar byte-identical; no confidence axis; no consumer change.
Census: add `has "$A" "grounded doubt"` row to plan-challenger.test.sh in
the same commit.
- **W9 v2**: adds the W0 scratch-validation step; rest unchanged.
- **W10 v2**: branch creation moves FIRST in execution order.
## 5. Constraints for challengers
- Registries append-only; curation only via /prune-memory.
- Census tests lock behavior: any hook/agent edit must land with its test
update in the same commit; `make test` must stay green.
- CLAUDE.global.md ≤ 320 lines (runtime warning threshold).
- BDR-021: Security + Architecture sections verbatim.
- Gitflow: feature branch off develop, no merge without human signal.
- The global file serves ALL models (Fable sessions, Opus 5 dispatches,
Sonnet executors read skill/agent prompts instead) — no Opus-5-only
wording in CLAUDE.global.md.
@@ -0,0 +1,174 @@
# ANNEX — directive-language inventory (analyzer report, 2026-07-30)
Produced by a read-only analyzer dispatch over agents/seo-analyzer.md
(1528 l) + agents/geo-analyzer.md (1106 l), cross-referenced against
every consumer. Referenced by the C1 plan (same folder, -1402.md).
## 0. Token census (raw)
| Token family | seo-analyzer.md | geo-analyzer.md |
|---|---|---|
| MUST/must | 12 | 6 |
| MANDATORY/mandatory | 8 | 4 |
| NEVER/never | 42 | 33 |
| ALWAYS/always | 6 | 1 |
| CRITICAL/critical | 3 | 1 |
| Do NOT / do not | 24 | 8 |
| verbatim | 6 | 3 |
| STOP | 3 | 4 |
| refuse/REFUSE | 6 | 4 |
| ⚠️ blocks | 0 | 0 |
## 1. Test locks on these files (complete list — 6 per file)
model-routing.test.sh:67-68 `model: opus` (both) · :150-157 `MODE:
collect|judge|template` + `COLLECTION COMPLETE` (both) ·
seo-data.test.sh:538-540 `fetch.sh crux` / `fetch.sh queries` /
`Performance GSC` (seo) · :542-543 `fetch.sh schema_gen` /
`fetch.sh content_quality` (geo).
NOT locked by any test: READY-TO-APPLY sentinel, envelope headings,
score-block shapes, JUDGE-ERROR strings, batch labels — contracts by
consumer only; a rewrite can break them silently and make test stays
green. Sibling dispatcher locks: model-routing.test.sh:159-166.
Stale line-number comments (no enforcement): lib/url-guard.sh:9,
url-guard.test.sh:20, source-scope.sh:25, seo-data/README.md:196/309,
drift.py:4, linkgraph.py:4 — all already drifted.
## 2. Format contract (artifact → consumer) — FREEZE SET
seo-analyzer: signals `.audit/seo-signals-<RUNID>.md` (+clean/load sites
in /seo) · `COLLECTION COMPLETE — RUNID: <RUNID>` terminal ·
`COLLECT REPORT` w/ `STATUS: DONE|BLOCKED` · `SEO JUDGE — VERDICT:
ERROR(<reason>)` · judge report forwarded verbatim to template ·
`SEO SCORING (<depth>)` block w/ `COVERAGE SOURCE:`/`COVERAGE LIVE :`
+ 7 axes + `SEO GLOBAL (weighted): XX.X/20` (score.py:26-37 mirrors
weights) · `TRAJECTORY TO 17/20 (code-only)` · `fetch.sh score` JSON
(`axes.{technical,on-page,seo-local,off-page,social,competitive,legal}`,
severities `critique|haute|moyenne|basse`, `status:"na"`) · `FIX PLAN (N
findings total)` + BATCH A…F (tier-mapping tolerant) · `## FIX BUNDLE
(for dispatcher)` + `### AUTO/### GATED/### USER ACTIONS` + item fields
`id: applier: files: concern: current: expected:` · sentinel `READY TO
APPLY — awaiting dispatcher confirmation` (also reused by /harden:366) ·
envelope `SEO AGENT RESULT` + `## SECTION FOR SEO.md §2…§6` + `## ENTRIES
FOR SEO.md §0/§8/§9/§10/§11/§15` · `Automatisation possible avec:` per
§11 entry · standalone `.claude/audits/SEO.md` w/ `**Score SEO** : XX.X
/ 20` (client-handover-writer.md:344 labeled grep) + §0-§15 + Historique.
geo-analyzer: same families with GEO names; envelope `GEO AGENT RESULT`
+ `## SECTION FOR SEO.md §7` (7.1-7.6); `**Score GEO** : XX.X / 20`
(handover parses it only inside SEO.md, allow_fallback=no); G1-G7
batches (G1-G4/G6 AUTO · G5 GATED · G7 USER). Both: STEP NUMBERS are
addressed by dispatchers (seo: 2-5/6-11/12-14; geo: 0-5/6-12/13-15;
also depth-matrix.md:17-19,37) — renumbering re-points dispatch prompts.
Engine interfaces: fetch.sh verbs {crux,queries,inspect,cannibal,
sitemap,rendercheck,linkgraph,score,schema_gen,content_quality} ·
url-guard.sh host|url · source-scope.sh findargs|list · resources/*.md.
## 3-4. Site classification counts
| | seo | geo | total |
|---|---|---|---|
| A machine-parsed contract | ~52 | ~41 | ~93 (12 test-locked) |
| B safety/policy invariant | ~30 | ~31 | ~61 |
| C process choreography | ~21 | ~12 | ~33 |
| D other/domain-fact | ~20 | ~13 | ~33 |
### Class C sites — seo-analyzer.md (rewrite targets)
:61 "First action." · :143-148 CMS-detect-before-edit ordering ·
:208-210 "keep the two consistent" (runtime cross-file reconcile) ·
:508 "run this BEFORE anything else in STEP 5" (ordering; the refusal
rule itself is B) · :550-553 "Record the denominator BEFORE sampling"
(ordering; honesty rule is B) · :602-604 "Sanity-check the grouping
before trusting it" (self-verify) · :606-618 sampling-method essay ·
:661-680 C1a 20-line rationale (rule itself is B at :1493-1501) ·
:875 per-item method · :970-971 "Run it twice on the same file before
publishing" (exact BDR-081 over-verification class) · :1147 "AUTO items
are a commitment, not a suggestion." · :1149-1157 P0 CMS-plugin-first
mandate · :1159-1162 P0 Bing mandate (dup of geo :777-786) · :1217 "Do
not proceed to STEP 12 until this plan is printed." · :1260-1261 +
:1342-1350 + :1502-1503 landing-page rule ×3 · :1309-1320 bundle
completeness checklist (10 checkboxes self-audit) · :1504 "Preserve
existing valid SEO." · :1522-1523 WebSearch-on-FULL extra-verify ·
:1525-1526 "Transparency. Every automated change logged" (VESTIGIAL —
agent applies nothing, pre-BDR-061).
### Class C sites — geo-analyzer.md
:48 "copy these patterns" · :124 "First action." + :127-139 ask-block
(unreachable when dispatched) · :230 conditional skip · :262-269 +
:873 + :1063-1065 PERMISSIVE default ×3 · :360 ordering · :394 "20-50
real customer questions (P0)" · :777-786 MANDATORY AI-index submission
(dup of seo :1159-1162) · :811 "Consolidate EVERY finding" · :823
"Print the plan before STEP 13" · :1102-1103 WebSearch extra-verify ·
:1106 "Every automated change logged in §14" (VESTIGIAL; §15 log is
dispatcher's per :959).
### Class B anchors (keep obligation, dedup emphasis)
CWD/TARGET MISMATCH twins (seo :117-126 ≈ geo :173-181) · url-guard
mandatory (seo :287-291 ≈ geo :273-277) · R2 refuse-to-score (seo
:519-548, geo :548-557; BDR-072) · NAP direction rule (seo :801-812,
geo :1073-1087; LRN-032-zenquality) · COVERAGE mandatory (seo
:1110-1130, geo :725-729; LRN-133) · never-apply/L1 (BDR-061; LRN-105
named-ban) · C1a build-output ban · no-invented-content/DGCCRF ·
"Compute the scores, do not feel them (I7)" (BDR-073) · §14 mandatory
disclosure lines (backlinks BDR-071, security headers I4) · honest
llms.txt framing · cite-sources (LRN-131).
## 6. Duplication map (sweep ALL twins — LRN-113)
seo internal: never-apply ×4 (:1227-1234, :1352-1357, :1468-1472,
:1527-1528) · landing-page ×3 (:1260, :1342, :1502) · shared-file
discipline ×2 (:1254, :1486) · bundle self-containment ×2 (:1249,
:1473) · COVERAGE ×4 (:438, :1000, :1096, :1110) · security-headers-
not-scored ×3 (:281, :977, :994) · 30/70 ×3 (:397, :614, :1165) ·
sentinel-verbatim ×3 (:1301, :1304, :1397).
geo internal: PERMISSIVE ×3 · never-apply ×4 (:826-832, :842-848,
:1031-1034, :1104-1105) · tier-mapping ×2 (:824, :850) ·
content_quality-advisory ×2 (:584, :622) · shared-file ×2 (:858,
:1047) · llms-honest ×2 (:337, :1066) · cite-sources ×2 (:17, :1089).
Cross-agent twins (stay twins — both files dispatch standalone):
CWD block · url-guard block · MODE DETECTION · MODE BOUNDARY · R2 ·
COVERAGE · NAP rule · RULES section skeleton · C1a · automation rule ·
Bing/AI-index action · CDN/WAF check.
Agent↔dispatcher duplication (stays — dispatch prompt is per-run
context, agent spec serves standalone/no-MODE paths): NAP ×4 total ·
shared-file ×7 · security-headers ×5 · domain split · weights 80/20-
75/25 · Historique · never-re-derive (test-locked dispatcher side).
## 7. Contradictions / ambiguities found
1. seo :1525-1526 + geo :1106 vestigial "automated change logged"
(agent applies nothing; geo :959 says dispatcher fills §15).
2. Ask-the-user blocks unreachable in dispatched path (seo :64-75,
:88-112; geo :127-139, :153-168); /geo:41 states it outright.
3. Collect boundary wording: agents "STEP 0-5 ONLY" vs /seo "STEP 2-5
only (context replaces STEP 0-1)" — works by prompt override.
4. geo judge does live work (sameAs curls :477-492, web_search) unlike
pure-judgment seo judge — asymmetric split, by design.
5. /harden imposes its own output contract (HARDEN.md, /100) the agent
spec never acknowledges; keys on "NARROW-SCOPE" in dispatch prompt.
6. "LRN-032" cite is ambiguous in THIS repo (local LRN-032 = different
lesson; the NAP lesson is zenquality's registry) — keep the
"zenquality" qualifier wherever cited.
7. geo :376-377 uncited FAQ-citation-rate claim vs geo :1089-1097
cite-sources rule (LRN-131 failure class).
8. Score-label parse fragility: client-handover extract_score fallback
greps FIRST X/20 in file — losing the `Score SEO` label would
silently read `TRAJECTORY TO 17/20` as 17.0. (Latent, downstream.)
9. GEO scoring has no deterministic engine (score.py covers SEO axes
only) — BDR-073 binds only half the pair.
## 8. Binding memory (from the analyzer's read-before)
IN FORCE: BDR-081 (premise) · LRN-139 (when-guidance shape) · BDR-061
(bundle+sentinel decision) · BDR-077 (mode split, fail-closed, locks
survive) · BDR-073 (deterministic scoring) · BDR-072 (R2 refuse) ·
BDR-071 (off-page ceiling + §14 line) · BDR-010/LRN-011 (labeled
scores gate) · LRN-133 (omission legible) · LRN-131/132/EVAL-025
(WebSearch ≠ verification) · LRN-105 (named ban stays explicit) ·
LRN-080/088 (measure before delete → dogfood) · LRN-113 (sweep whole
surface) · LRN-093 (no vacuous locks; single-line anchors) ·
LRN-126/137 (mode split carries data paths) · BLK-017 (Bing deferred).
## 9. Open questions → dispatcher decisions (see plan §4b)
Q1 freeze scope · Q2 census extension · Q3 dedup strategy ·
Q4 vestigial lines · Q5 /harden //onboard reconciliation.
@@ -0,0 +1,342 @@
# PLAN v2 — De-prescribe seo-analyzer.md + geo-analyzer.md for Opus 5
Date: 2026-07-30 · Branch: feature/seo-geo-deprescription (off develop, started)
KIND: build-plan · Author: main-loop session (Fable 5)
Parent decision: BDR-081 N5 (deferred as separate project) · Method: LRN-139
v2: revised after the 3-lens challenge (§5bis) — every BLOCKER closed by a
named change; one confirmation challenger pass follows before execution.
## 1. Context & evidence (v2 — sizing corrected per simplicity#1)
Both agents are opus-pinned (BDR-076) → every judge phase runs Opus 5.
BDR-081 profile applies: literal following, over-verification when told
to verify, conflicting/duplicated rules burn reasoning tokens. These are
the LONGEST agent files in the repo (1528 + 1106 l) with real downstream
parsers — NOT the densest (measured: ~4.5 directive hits/100 l, ranks
20th/22nd; security-auditor is 17/100). What this pass buys, honestly:
(a) removal of self-output-verification demands (the one pattern the
baseline dogfood caught live: the judge reported "run twice, identical
output" — seo:970 firing), (b) removal of vestigial pre-BDR-061 lines
and 2 real contradictions, (c) small same-audience/same-range dedup,
(d) caps→when-guidance on choreography. The verification apparatus
(census + 3-lens challenge + before/after dogfood) is USER-DIRECTED for
this chantier, not derived from the density premise.
## 2. Contract surface (v2 — split per correctness#5)
### 2a. Machine-parsed (named non-LLM consumer: test, script, or literal
grep in a dispatcher step) — byte-frozen
- `model: opus`, `MODE: collect|judge|template`, `COLLECTION COMPLETE`
(model-routing.test.sh:67-68,150-157).
- `fetch.sh crux|queries` + `Performance GSC` (seo), `fetch.sh
schema_gen|content_quality` (geo) (seo-data.test.sh:538-543).
- `SEO|GEO JUDGE — VERDICT: ERROR(` — dispatcher ERROR CONTRACT
fail-closes on it (skills/seo:316-318, skills/geo:65-68).
- `## FIX BUNDLE` + sentinel `READY TO APPLY — awaiting dispatcher
confirmation` — apply step keys on it (skills/seo:524, skills/geo:101;
reused by /harden:366).
- `.audit/<seo|geo>-signals-<RUNID>.md` names + fail-closed load.
- STEP numbering: dispatchers address ranges literally (seo 2-5/6-11/
12-14; geo 0-5/6-12/13-15; depth-matrix:17-19,37).
- `**Score SEO** : XX.X / 20` / `**Score GEO** : XX.X / 20` labels —
client-handover-writer.md:344-345 labeled grep (BDR-010/LRN-011);
losing the SEO label silently falls back to first-X/20-in-file.
- Bundle item fields `id: applier: files: current: expected:` — pasted
verbatim into hotfixer/feater at L1; `applier: bash` run in-loop.
- url-guard call sites: seo-analyzer.md:287-295, geo-analyzer.md:273-280
(NOT ":257" as v1 said — robustness#5) + sitemap-URL guard seo:573-582.
- `NARROW-SCOPE` keying of the I4 carve-out (seo:981-983) — /harden's
dispatch prompt relies on it.
### 2b. LLM-convention contracts (no code consumer; the dispatcher LLM
merges by these shapes) — locked in the census, still frozen
`SEO|GEO AGENT RESULT` envelopes · `## SECTION FOR SEO.md §N` ·
`## ENTRIES FOR SEO.md` · `SEO|GEO SCORING (` blocks + `COVERAGE
SOURCE`/`COVERAGE LIVE` lines + `GLOBAL (weighted)` · `TRAJECTORY TO
17/20 (code-only)` · `FIX PLAN (` (seo) · batch labels A-F / G1-G7
(tier recognition tolerant, labels nominal) · `COLLECT REPORT` +
`STATUS: DONE|BLOCKED` · `Automatisation possible avec:` · §0-§15
report skeleton + Historique. CROSS-AGENT NOTES emit-instruction lives
in /seo's dispatch prompts (dispatcher-side lock only).
## 3. Class B invariants — obligation kept, single strongest statement;
security ORDERINGS byte-frozen (robustness#5/#7)
- Guard-first orderings, frozen verbatim: seo:287-291 / geo:273-277
("Guard the domain before it reaches a shell… Run the guard FIRST…
never 'clean up' the value and retry") + seo:573-582 URL loop.
- seo:550 "Record the denominator BEFORE sampling" — the ordering IS
the honesty mechanism (a post-hoc denominator is self-serving);
frozen; only surrounding prose may compress.
- NAP direction rule (LRN-032-zenquality — keep the qualifier, the bare
ID is ambiguous in this repo), R2 refuse-to-score (BDR-072), COVERAGE
obligations (LRN-133 — note :436-439 is a DISTINCT index-reach
obligation, not a repeat), §14 mandatory disclosure lines (BDR-071
backlinks verbatim line, I4 security-headers), never-apply/L1
(BDR-061; LRN-105 named ban), C1a build-output ban, no-invented-
content/DGCCRF, deterministic scoring (BDR-073), fail-closed judge,
shared-file Edit-not-Write discipline, honest llms.txt framing,
cite-sources (LRN-131).
- External-freshness checks are NOT self-verification (robustness#6):
seo:1522-1523 + geo:1102-1103 verify a DRIFTING WORLD feeding an
AUTO-tier robots.txt edit — kept, reworded as when-guidance ("crawler
lists shift; cross-check before emitting G1 from the dated resource").
## 4. Work items v2
- P0 SEQUENCING + LIVE-TREE EXPOSURE (robustness#4, conf#2/#3/#4/#9):
agents/ resolves through ~/.claude symlinks to the WORKING TREE —
edits are live between Edit calls, before any commit. Rules:
(1) the FULL baseline completes before the first agent edit —
signals + judge reports + TEMPLATE envelopes + merged SEO.md +
HUMAN-ACTIONS.md (conf#2: without frozen template artifacts the
template-range edits would have no differential and P0 makes one
unobtainable later);
(2) all baseline artifacts copied to the DURABLE, gitignored
`.audit/dogfood-baseline/` in this repo before the first edit
(conf#9: the session scratchpad dies with the session/reboot;
LRN-124: .audit/** is never committed);
(3) freeze window: no /seo //geo //harden //onboard AND no
/client-handover (spawns /seo — conf#3) nor any skill transitively
dispatching either analyzer, in ANY project, until the after-dogfood
verdict;
(4) aborts (conf#4): mid-reword interrupt or after-dogfood failure →
`git checkout HEAD -- agents/seo-analyzer.md agents/geo-analyzer.md`
(in-flight revert, index-safe); `git checkout develop -- agents/…`
is reserved for a WHOLE-BRANCH abandon; after an abort the named
exit is either (a) fix + re-run the after-dogfood, or (b) present
the static evidence (census + git diff review) to the human who may
accept or abandon at the gate — no open-ended reverted state.
- P1 CENSUS (commit 1, test-only, green pre-reword — compatible with
§7's same-commit rule: it locks EXISTING state and changes no agent
file; reword commits carry any census DELTA): DONE in working tree —
lib/tests/seo-geo-contract.test.sh 54/0, shellcheck clean, real
flip-test run: 7 scratch mutations → 7 FAILs (not "by construction" —
robustness#10). File-qualified locks (correctness#4): `FIX PLAN (` +
`applier: bash` + `Score SEO` seo-only; `Score GEO` geo-only.
Incidental locks dropped (CROSS-AGENT NOTE agent-side, bare
`applier:`). Item fields locked both files. v3 (conf#5): EVERY
`## STEP n —` header locked, interiors included (seo 0-14, geo 0-15)
— census now 71/0. Freeze mechanism for the
~40 A-sites the census does not cover: reviewed `git diff -U0
agents/*.md` on each reword commit (simplicity#4).
- P2 REWORD seo-analyzer.md (commit 2):
(a) Self-OUTPUT verification, v3 (conf#1/#8 — neither is deleted
outright): :970-971 "run it twice" → when-guidance integrity
guard ("if the findings JSON changed after scoring, re-run and
explain the move" — score.py is deterministic, so a moving
output means mutated findings: anti-score-shopping, BDR-073;
the unconditional double-run the baseline judge burned goes
away, the guard stays); :1217 "Do not proceed until printed" →
when-guidance scoped to the single-shot path ("single-shot runs
print the FIX PLAN before STEP 12 serializes it" — MODE: judge
stops at 11, but /harden //onboard execute the whole file,
conf#1). The completeness checklist :1309-1320 is NOT deleted:
its routing rows (stock-photo→GATED(E), compression→AUTO(bash)
or §11, aggregateRating→AUTO(hotfixer), structural→GATED(D)…)
are unique routing content (robustness#3) — reshape into a plain
mapping table, drop only the checkbox self-audit framing.
(b) DELETE vestigial :1525-1526 (contradicts BDR-061; Q4).
(c) DEDUP under the invariant (correctness#1 + robustness#1): only
VERBATIM same-AUDIENCE (spec rule / bundle-item payload /
phase-local caveat) same-MODE-RANGE (collect 0-5 / judge 6-11 /
template 12-14 / RULES=global) repeats merge. Expected survivors
per family listed at execution in the commit message; honest
net: never-apply 4→3 (RULES pair merges; template-range
statements stay), sentinel-verbatim reminders 3→2, landing-page
3→2 (payload instance :1260 + one spec statement; :1342 vs
:1502 merge), bundle-self-containment 2→1 (same range).
NOT deduped (v1 was wrong — distinct rules or cross-range):
COVERAGE ×4, 30/70 ×3, security-headers ×3, shared-file
discipline (payload vs spec audiences).
(d) SOFTEN caps/orderings to when-guidance, keeping semantics:
:61, :508 (gate stays before on-page scoring; emphasis drops),
:875, :1147, :1149-1157 CMS-plugin-first folded together with
:143-148 into ONE statement (correctness#3 — two strengths of
one rule otherwise), :1159-1162 Bing (content rule kept, caps
drop; FULL-only → statically verified), essays :606-618 +
:661-680 compressed keeping the rule + LRN citations; :602-604
kept as a when-guidance failure detector ("families ≈ URLs →
the heuristic broke — say so"), not deleted (robustness#8).
(e) Dispositions completing the C-list (correctness#3): :208-210 →
static pointer ("the CDN/WAF twin check lives in geo STEP 4");
:1504 KEEP as-is (one-line scope guard).
- P3 REWORD geo-analyzer.md (commit 3), same invariant:
PERMISSIVE ×3: ALL survive (collect/template/RULES ranges;
:873 is the item-level default guarding an unconfirmed AUTO
robots.txt edit — named survivor, robustness#9). never-apply 4→3
(RULES pair merges). tier-mapping :824/:850 BOTH stay (judge vs
template ranges). content_quality-advisory 2→1 (same range).
shared-file 2× stays (payload vs spec). llms-honest 2× stays
(collect vs RULES). cite-sources 2× stays (:17 guards the header
stats specifically). :1106 vestigial → reworded to the truth
(dispatcher fills the log — matches :959; Q4). :777-786 caps →
plain content rule (FULL-only). :1102-1103 → freshness
when-guidance (kept — §3). :394 quantity softened ("substantial,
real customer questions"). :48 softened. :124-139 ask-block KEPT
(standalone path). Orderings :360/:811/:823 softened. :376-377
uncited claim → honest framing (no invented source).
- P4 DOGFOOD AFTER (v3 — ordered by decisiveness, conf#7): fresh copy
of zenquality-frozen; phases in this order so a mid-run death still
leaves the decisive evidence (billing class already realised once):
(ii-first) judges fed the FROZEN baseline signals
(.audit/dogfood-baseline/) → judge reports vs frozen baseline judge
reports, ZERO collect variance — the decisive Opus-judge-prose
differential; (iii) templates on those judge reports → envelopes,
compared against the frozen BASELINE envelopes — the template
verdict anchors on ENVELOPES only (SEO.md/HUMAN-ACTIONS.md are
dispatcher-merged by this authoring session, non-attributable —
conf#10); (i-last) fresh collects, same pre-answered context →
(a) shape check of signals/COLLECT REPORT vs baseline, (b)
FIELD-LEVEL diff of the fresh signals vs baseline signals (record
blocks, COVERAGE counts, denominators — a shape-valid file with a
dropped field must be caught, conf#6), and (c) ONE end-to-end seo
judge on the FRESH signals (the domain with the most collect-range
edits) so the reworded collect→judge handoff runs at least once.
Comparison mechanical-first: presence-assertion script (named home:
`.audit/dogfood-baseline/assert-after.sh`, session-reproducible,
never committed — conf#11) + a FRESH reader agent diffing
before/after WITHOUT this plan in context (correctness#7); the
authoring session only arbitrates its report. If the after-run dies:
P0(4) abort + named exit applies; no merge request meanwhile.
- P5 GATES: make test full suite (census + model-routing + seo-data +
no-vacuous-locks) · shellcheck on touched .sh · per-RANGE grep sweep
for every deduped family (asserts the named survivor lines exist in
their ranges — mode-blind ≥1× sweep is insufficient, robustness#1) ·
MEASURED deltas recorded (simplicity#7): wc -l + directive-token
census (annex §0 grep set) per file, before/after, into the BDR.
(v1's manual MODE/STEP sweep dropped — the census asserts it,
simplicity#5.)
- P6 CAPITALIZE: BDR (decision, invariant, deltas, alternatives), LRN
(audience×range dedup invariant — reusable), journal, CHANGELOG.
TODO C1 checked. NO merge (human gate). Checkpoint report includes
the DYNAMICALLY-UNVERIFIED list (§6bis).
## 4b. Dispatcher decisions (v2)
- Q1 freeze scope: all §2a byte-frozen + §2b frozen via census; the
remaining unlocked A-prose freeze = per-commit git diff review.
- Q2 census: done (P1), flip-proven.
- Q3 dedup: WITHIN-file, same-AUDIENCE, same-MODE-RANGE, verbatim
repeats only. Cross-agent + agent↔dispatcher twins stay. (Mechanism
note correcting robustness#1's premise: every dispatch loads the FULL
agent file; the risk is ATTENTIONAL — a literal-following model told
"run STEP 13-15" deprioritizes guidance scoped to another step's
body — not access. Same fix either way.)
- Q4 vestigial: seo :1525-1526 DELETE; geo :1106 REWORD to
dispatcher-owns-log (correctness#6 resolved).
- Q5 /harden //onboard: out of scope (N6); their dispatch-prompt
contracts are untouched by agent-file rewording; `NARROW-SCOPE`
keying frozen (§2a).
## 5. Dogfood protocol (v2)
Baseline (DONE for collect+judge SEO; geo judge in flight at v2 time):
frozen zenquality copy (no .env), inline pipeline (canonical /seo shape
— the nested-CLI attempt died on the CLI monthly spend limit, recorded),
absolute PROJECT ROOT in every dispatch, `/seo local conservative`,
STEP 0 pre-answered, NAP = NAP-KIT.md (user-confirmed 2026-07-10).
Baseline artifacts frozen under the DURABLE `.audit/dogfood-baseline/`
(gitignored, never committed — conf#9): signals ×2, judge reports ×2,
template ENVELOPES ×2, merged SEO.md, HUMAN-ACTIONS.md (conf#2 — the
template phase runs to completion BEFORE the first agent edit).
After-run per P4. LIMITS stated honestly
(robustness#2): conservative never enters STEP 1b/1.5 (no applier parses
an item this run — the item-field contract is census-locked statically);
LOCAL never executes STEP 3-4/6-7 FULL branches (Bing/AI-index emission
text, live checks — the FULL-only conditionals were exercised and
correctly declined in the baseline judge). These stay on the
§6bis unverified list for the human gate; a FULL/aggressive dry-run is
an OPTION the user may order at checkpoint, not part of this plan.
## 5bis. CHALLENGE SYNTHESIS (2026-07-30)
Verdicts: correctness FATAL(3) [1 BLOCKER, 2 MAJOR, 4 MINOR] ·
robustness FATAL(9) [3 BLOCKER, 6 MAJOR, 2 MINOR] · simplicity
CONCERNS(3) [3 MAJOR, 4 MINOR]. All three lenses returned. Every
BLOCKER closed by a named v2 change:
- correctness#1 (audience-blind dedup) + robustness#1 (mode-blind
dedup) → §4b Q3 invariant + P2(c)/P3 rewritten + P5 per-range sweep.
- robustness#2 (dogfood can't reach riskiest edits) → §5 honest limits
+ §6bis unverified list + P2(d)/P3 minimal-diff on FULL-only sites +
static census cover; FULL/aggressive run offered to the human, not
silently added (billing exposure robustness#11).
- robustness#3 (routing table misfiled as self-check) → P2(a) keeps
routing rows verbatim.
Majors adopted: R4 live-tree abort path (P0) · R5 url-guard anchors
corrected + security orderings frozen (§2a/§3) · R6 external-freshness
kept (§3) · R7 :550 frozen (§3) · R8 :602 kept as detector (P2(d)) ·
R9 :873 named survivor (P3) · C2 folded into R2's resolution · C3 full
dispositions (P2(d)/(e), P3) · S1 §1 rewritten · S2 controlled
judge-replay (P4) · S3 mechanical presence script (P4). Minors adopted:
C4 file-qualified locks · C5 §2 split · C6 three inconsistencies
resolved (P1 note, Q4, N1 marker) · C7 fresh-reader diff · S4 diff-
review freeze · S5 sweep dropped · S6+R10 census corrected+flip-proven ·
S7 measured deltas. Rejected/scoped: S1's apparatus-shrinking (the
apparatus is user-directed); R1's access premise corrected to
attentional (fix adopted unchanged).
CONFIRMATION PASS (robustness lens, v2 → v3): FATAL(9) — 2 BLOCKER +
7 MAJOR/MINOR, all targeting the v2 amendments as asked. Closed by
name: conf#1 no-MODE single-shot → §6bis + P2(a) :1217 scoped-softened
· conf#2 missing baseline template artifacts → P0(1) full-baseline
precondition · conf#3 /client-handover freeze → P0(3) · conf#4 abort
HEAD-vs-develop + named exit → P0(4) · conf#5 interior STEP locks →
census extended to all headers (71/0) · conf#6 collect→judge seam →
P4(i) field-diff + one end-to-end seo judge on fresh signals · conf#7
decisiveness order → P4 reordered (ii)→(iii)→(i) · conf#8 :970
anti-score-shopping → when-guidance reword, not deletion · conf#9
volatile baseline → durable .audit/dogfood-baseline/ · conf#10
dispatcher-owned artifacts → envelope-anchored template verdict ·
conf#11 script home named. Challenge budget exhausted (1 re-pass max):
residual risk goes to the human gate with this record.
## 6. Explicitly NOT doing
- N1 No dispatcher (SKILL.md) edits.
- N2 No scoring-weight, axis, or depth-matrix changes.
- N3 No model-pin changes (BDR-076).
- N4 No weakening of class-B invariants (§3 hardened in v2: security
orderings byte-frozen).
- N5 No new modes, no pipeline reshaping (BDR-077).
- N6 No /harden //onboard contract reconciliation (annex §7.5).
- N7 No collect-boundary wording fix (works by prompt override).
- N8 No cross-agent shared-resource consolidation.
- N9 No deterministic GEO score engine (annex §7.9).
- N10 No FULL/aggressive dogfood in this plan (user option at gate).
## 6bis. Dynamically-unverified edit surface (for the human gate)
Sites edited by P2/P3 that no dogfood run executes: FULL-branch content
(seo :1159-1162 Bing emission, geo :777-786 AI-index emission, both
freshness when-guidances), apply-path parsing (STEP 1b/1.5 — item
pasted into appliers; covered statically by census item-field locks +
frozen bundle templates), STEP 6-7 external-presence prose, and the
no-MODE single-shot path (conf#1: /harden and /onboard dispatch the
agents without a MODE line — "all steps in sequence" — so the whole
reworded body drives those runs; every never-apply and ordering
statement that path relies on keeps a surviving instance, and :1217
is softened-scoped to it, never deleted). Mitigation: minimal diffs
there (caps→plain only), census locks, git-diff review.
## 4c. Backlog surfaced (not this branch)
- Score-label fallback fragility in client-handover-writer.md (can read
`TRAJECTORY TO 17/20` as 17.0 if the label vanishes) — annex §7.8.
- Stale lib/ line-number comments pointing at agent lines (annex §1).
- Baseline judge's gate observation: /client-handover 17/20 gate passes
with an open `critique` finding — "open critique = independent
blocker" is worth its own decision.
## 7. Constraints for challengers
- Registries append-only; census green throughout; reword commits keep
54/0 + model-routing + seo-data locks green.
- Agent files symlink-live INCLUDING between Edit calls (P0 abort path).
- §2a byte-identical; §2b frozen; STEP numbering preserved; §3 security
orderings verbatim.
- Dedup only same-audience + same-mode-range verbatim repeats; named
survivors per family in commit messages; P5 per-range sweep.
- The judge phase is Opus 5; collect/template Sonnet — literal
following applies to all (E5 "since 4.7").
- Baseline artifacts frozen before first edit; after-run design per P4.
@@ -0,0 +1,201 @@
# PLAN — gstack-playwright-lib (feat) — REVISION 3
Contract: `.claude/tasks/contracts/2026-09-13-gstack-playwright-lib-2220.md`
Revision 3 (2026-09-15). Rev 1 → 3-lens challenge → rev 2 → confirmation pass
→ rev 3. The conflict-RECOVERY branch is WITHDRAWN at the human gate: it
concentrated 3 BLOCKERs and 4 MAJORs, and its worst case regressed a working
browser into BLK-008, which the pre-existing behavior never did.
## Context
- Chromium is not an apt package. It is the browser revision pinned by the
installed Playwright (`gstack/setup:483`). gstack: playwright 1.61.1 →
chromium rev 1228 (`Chrome for Testing 149`).
- `~/.cache/ms-playwright/.links/` registers 3 playwright installs: gstack
1.61.1 (1228), gsd-pi nvm 1.61.0 (1228), gsd-pi ~/.local 1.63.0 (1243).
Every dir on disk is referenced → 0 bytes reclaimable. Playwright already
prunes correctly on every `install` (`_deleteStaleBrowsers`). No pruner is
written here.
- The gstack submodule is intentionally dirty: `package.json` + `bun.lock`
carry the BDR-029 bump; `.gitmodules` sets `ignore = dirty`. It also carries
an untracked `?? bin/bin`.
- `update-all.sh:87` calls `git submodule update --remote` bare, swallows
stderr, and never re-applies the bump afterwards. THAT is the gap.
## Hard constraints the code must respect
**Inherited errexit.** All three callers run `set -euo pipefail` and source
the lib. `gstack_bump_playwright_if_unsupported` and `gstack_browsers_report`
are called as bare statements, so they MUST `return 0` on every path and every
capture inside them takes `|| true`.
`gstack_submodule_update_with_bump` is the ONE exception: it returns non-zero
on failure and is therefore called ONLY as an `if` condition, keeping
`update-all.sh:87`'s existing `if / else warn` shape. An offline update stays
non-fatal, exactly as today.
**Printer names.** The lib defines `_gspw_ok`, `_gspw_warn`, `_gspw_info` and
NEVER a bare `ok`/`warn`/`info`/`pass`/`fail`. `doctor.sh:22` sources the lib
before every check, so bare names would override `doctor.sh:12-15` and
silently disconnect its `ERRORS`/`WARNS` counters.
**No destructive command in the lib, at all.** No `rm`, `rmdir`, `unlink`,
`truncate`, `mv`, and no `git checkout`/`reset`/`clean`/`stash`. Contract
criteria 4 and 6 both grep for this.
**macOS-safe.** No `timeout` without a `command -v` guard (absent from stock
macOS), no `readlink -f` (absent before Monterey 12.3), no `md5sum`, no
`sed -i` without a suffix, no bash-4-only expansions (`${x,,}`), no `grep -P`.
## Files
1. `lib/gstack-playwright.sh` — NEW. Sourceable lib + verb dispatcher. No
`set -euo pipefail` at top level (mirrors `lib/detect-plugins.sh`).
Dispatcher guarded by `[ "${BASH_SOURCE[0]}" = "${0}" ]`, exposing
`browsers-report` ONLY. Any other argument → `usage:` on stderr, exit 2.
The write functions stay sourced-only: a CLI verb would expose
`bun add playwright@latest` as a command-line entry point.
- `_gspw_ok` / `_gspw_warn` / `_gspw_info <msg>` — fixed-prefix printers.
- `gstack_pw_ostag [os_release_path]` — prints `ubuntu<VERSION_ID>` for
Ubuntu, nothing otherwise. The capture takes `|| true`: the moved line
exits 1 on every non-Ubuntu host and would abort the caller under
inherited errexit. `return 0` always.
- `gstack_pw_supports <playwright_core_lib_dir> <ostag>` — 0/1 by grep.
- `gstack_bump_playwright_if_unsupported <gstack_dir>` — BDR-029 logic,
parameterized. Prepends `$HOME/.bun/bin` to PATH when `bun` is not
resolvable (LRN-036). Wraps ALL THREE bun invocations
(`bun install --frozen-lockfile`, the `bun install` fallback,
`bun add playwright@latest`) in `timeout 300` when `command -v timeout`
succeeds, plain otherwise. Exit 124 from any of them → `_gspw_warn` and
`return 0` WITHOUT attempting the bump: a TERM'd install leaves
`node_modules` half-written, and the support grep would then read a
truncated tree. One `_gspw_info` line before the network work so a
stalled registry is visible. `return 0` on every path (BDR-029
non-fatal).
- `gstack_submodule_update_with_bump <repo> [sub_path]`:
1. `git -C "$repo" submodule update --remote "$sub"`, stderr captured.
2. exit 0 → `gstack_bump_playwright_if_unsupported "$repo/$sub"` →
`return 0`.
3. exit != 0 → `_gspw_warn` with git's own message, verbatim and
unparsed. Then, when `git -C "$sub" status --porcelain --
package.json bun.lock` is non-empty, one extra `_gspw_info` hint line
naming the local Playwright bump and pointing at `make plugin`.
`return 1`. NOTHING in the working tree is touched.
No locale pin is needed: git's message is displayed, never parsed for a
decision. The hint is advisory, so its constant-true condition is
correct here, unlike the withdrawn recovery branch where it gated a
destructive step.
- `_gspw_browser_referenced <playwright_core_path> <dir_name>` — does that
install require this cache directory? Splits `<dir_name>` into name +
revision on the LAST `-`, then normalizes `_` → `-` on the name
(Playwright writes `chromium_headless_shell-1228` while `browsers.json`
says `chromium-headless-shell`; without this, two live directories are
reported unreferenced forever). Matches the base `revision` OR any value
under that browser's `revisionOverrides` (webkit and ffmpeg carry them
for mac and debian11 and ubuntu20.04 hosts). awk only: no jq, no
python3, no fallback ladder.
- `_gspw_install_label <playwright_core_path>` — `<dir-before-node_modules>
<version>`, e.g. `gstack 1.61.1`, `gsd-pi 1.63.0`.
- `gstack_browsers_report [cache_dir]` — read-only. Resolves the cache as
`${1:-${PLAYWRIGHT_BROWSERS_PATH:-$HOME/.cache/ms-playwright}}`; the
documented value `0` means "bundle into node_modules", so `0` and any
non-directory degrade to the silent no-cache path. Prints the header
`Playwright browsers`, the total from `du -sh … || true`, one line per
cache dir matching `*-<digits>` with the installs requiring it, then
`<N> unreferenced, <M> broken link(s)`. When N or M > 0, one
`_gspw_warn` naming them and the remedy, phrased without the words `rm`
or `mv` (criterion 6 word-greps the source): "re-run `playwright
install`, which prunes stale revisions". A dir whose NAME is listed by
some install but at another revision counts as `unknown revision`, not
unreferenced. `return 0` on every path.
2. `install-plugins.sh` — delete the inline function (294-321), source the lib
next to detect-plugins (line 30), call site at ~370 becomes
`gstack_bump_playwright_if_unsupported "$GSTACK_DIR"`.
3. `update-all.sh` — source the lib next to detect-plugins (line 19). Line 87
becomes `if gstack_submodule_update_with_bump "$REPO"; then` and the
existing `else warn …` arm is KEPT verbatim. No other structural change.
4. `doctor.sh` — source the lib next to detect-plugins (line 22). Add its own
`── Playwright browsers ──` section (NOT nested under gstack: 2 of the 3
registered installs are gsd-pi), called as `gstack_browsers_report || true`.
5. `lib/tests/gstack-playwright.test.sh` — NEW, auto-globbed by `make test`.
## Edge cases
- Every public function except the update returns 0 under the callers'
`set -euo pipefail`, including the all-zero-counts case, which is this
machine's nominal state and would otherwise kill `doctor.sh` before its
summary and take `update-all.sh:519` down with it.
- `.links` entry whose target is gone or whose `browsers.json` is unreadable
→ counted as a broken link, never dereferenced further.
- Two installs of the same tool at different versions → both labels listed.
- gstack submodule absent → bump and update both no-op 0.
- Sourcing the lib prints nothing and does not change the caller's options.
## Tests (`lib/tests/gstack-playwright.test.sh`)
Shape of `lib/tests/fast-libs.test.sh` (`check` helper, `PASS=n FAIL=n` last
line, `mktemp -d` + trap). git 2.53 defaults `protocol.file` to `user`, which
blocks submodule clone and fetch. The fixture git calls are not enough: the
`git submodule update --remote` under test runs INSIDE the lib, in a fresh
process. So the test exports, for the whole test process,
`GIT_CONFIG_COUNT=1 GIT_CONFIG_KEY_0=protocol.file.allow
GIT_CONFIG_VALUE_0=always`, which the lib's own git inherits. Fixtures use
`git init -b main` with `submodule.<name>.branch = main` set explicitly, so
`--remote` resolves the way production does.
- `T1-ostag-ubuntu` / `T2-ostag-other`: fixture os-release files.
- `T3-errexit-safe`: the bump called as a bare statement under
`set -euo pipefail` with a non-Ubuntu os-release → the script reaches the
next line. Regression test for the latent abort.
- `T4-supports-hit` / `T5-supports-miss`: fixture lib dir with and without the
tag. Proves idempotence both ways without invoking bun.
- `T6-update-success-bumps`: fixture superproject + submodule, an upstream
commit, bump stubbed by redefining it after sourcing → update succeeds, the
stub ran once, returns 0.
- `T7-update-conflict-nondestructive`: local edit to the submodule's
`package.json` plus a conflicting upstream commit → returns non-zero, BOTH
bump-owned files are byte-identical to before the call, and the hint line
was printed. Carries the literal `update-conflict` (contract criterion 4).
- `T8-no-destructive-command`: greps the lib source for `git
checkout|reset|clean|stash` and for `rm|rmdir|unlink|truncate|mv` outside
comments. Stronger than criterion 6 alone.
- `T9-report-referenced`, `T10-report-underscore-dir`
(`chromium_headless_shell-1228` against a `chromium-headless-shell` entry),
`T11-report-unreferenced`, `T12-report-broken-link`,
`T13-report-revision-override`: fixture cache + `.links` → fixture
playwright-core dirs with hand-written `browsers.json`.
- `T14-report-zero-counts-exit-0`: everything referenced → exit 0. The nominal
case, not covered by the absent-dir case.
- `T15-report-no-cache` / `T16-report-browsers-path-zero`: exit 0, nothing on
stderr.
- `T17-source-safe`: sourcing emits nothing.
## Disposition (STEP 0.6)
- honors **BDR-029** by keeping the bump OS-gated, idempotent, non-fatal, and
by closing its stated caveat: the bump is now re-applied after every
successful update, not only at the next `make plugin`.
- honors **LRN-024** by extracting a helper and refactoring the existing
caller before adding the other callers. Deviations from "code MOVED not
changed" are named: `|| true` on the ostag capture (a latent abort on every
non-Ubuntu host, reproduced), and the `timeout` guard.
- honors **LRN-070** by never touching the submodule working tree at all. The
revision that did (discard, retry, restore) was withdrawn at the gate.
- honors **LRN-071** (recurrent 3x) by returning the update's real status, not
a non-fatal helper's 0.
- honors **LRN-040** by touching layer 1 only; `GSTACK_CHROMIUM_NO_SANDBOX`
is untouched.
- honors **LRN-085** by keeping the update idempotent, presence-guarded, no
`--force`.
- honors **LRN-036** by putting `$HOME/.bun/bin` on PATH inside the lib.
- honors **LRN-002** by grepping the moved function name repo-wide, readers
included.
- **LRN-038** already seen: the host-platform override is a dead end.
- BDR-029's reference line (`decisions.md:544`) and BLK-008's caveat
(`blockers.md:118`) describe behavior this plan changes. Registries are
append-only, so the plan does NOT edit them: the /feat CAPITALIZE step
owns the superseding entry.
+3 -3
View File
@@ -1,9 +1,9 @@
# Local secrets for Claude Code plugin install scripts.
# Copy to ~/.claude/.env and fill in real values. link.sh symlinks repo/.env to it; the secret never enters git.
#
# Used by: lib/toggle-external.sh enable|disable magic
# Get a key at: https://21st.dev/magic (dashboard → API keys)
MAGIC_API_KEY=your_21st_dev_magic_api_key_here
# 21st.dev needs nothing here since 2026-09-22: the Magic MCP server was
# replaced by the `21st` CLI, whose auth is `21st login` (browser token in
# ~/.config/21st). Any leftover MAGIC_API_KEY line is dead — delete it.
# ── Google SEO data layer (lib/seo-data) — used by /seo FULL ──
# OAuth Desktop client: GCP console → APIs & Services → Credentials → OAuth client (Desktop).
+34 -5
View File
@@ -64,11 +64,25 @@ skills/ios-sync
skills/design-motion-principles
skills/emil-design-eng
skills/frontend-design
# Impeccable — NOT a symlink: `impeccable skills install --scope=global`
# writes the skill dir (and its ~15 MB engine binary) straight in through the
# ~/.claude/skills symlink. Machine-owned, regenerated by make plugin/update.
skills/impeccable
# …and the 4 subagents the same installer drops through ~/.claude/agents.
# agents/ is a tracked directory, so these need naming explicitly.
agents/impeccable-*.md
# External skills installed via `npx skills add` — auto-created by link.sh
skills/darwin-skill
# 21st.dev skill pack symlinks — created on demand by toggle-external.sh /
# profile.sh (the pack is DISABLED by default, so these usually don't exist).
# A glob, not one line per skill: the `21st skills install` manifest owns the
# membership, so the pack can gain a skill with no edit here.
skills/21st-*
# Context7 docs-lookup skill — installed by `ctx7 setup --claude --cli`
# (install-plugins.sh Step 6, when absent) into ~/.claude/skills (a symlink to
# this repo's skills/). ctx7-managed and re-created on demand — not vendored here.
@@ -97,6 +111,14 @@ skills-disabled/
graphify-out/
.ctx7-cache/
# graphify's vendored skill — written into the repo by `graphify claude
# install` (install-plugins.sh STEP graphify), since ~/.claude/skills is a
# symlink to skills/. Untracked so a tool upgrade stops dirtying the tree.
# test-prompts.json is hand-written for darwin and stays tracked.
skills/graphify/SKILL.md
skills/graphify/references/
skills/graphify/.graphify_version
# /client-handover test artifacts (project-local renders)
LIVRAISON.md
LIVRAISON.html
@@ -142,11 +164,18 @@ desktop.ini
# an update. The source is always re-synced, so no offline copy is needed.
skills-external/frontend-design/
# Impeccable — machine-owned dist produced by `npx impeccable skills install`
# (install-plugins.sh Step 8d, update-all.sh), pinned in plugins.lock.json.
# Not vendored: the installer owns the layout and rewrites it on update
# (ctx7 pattern). Symlinked into skills/ by link.sh.
skills-external/impeccable/
# Emil Design Eng — machine-owned copy curl'd from emilkowalski/skill by
# install-plugins.sh (Step 8, when absent) and re-fetched on every update-all.sh
# run. Not vendored: tracking it produced a repo diff each time upstream shipped
# an edit. The source is always re-fetched, so no offline copy is needed.
skills-external/emil-design-eng/
# 21st.dev skill pack — machine-owned: `21st skills install` output, staged by
# install-plugins.sh Step 8.7 (the installer refuses to write through the
# ~/.claude/skills symlink, so it runs under a throwaway HOME and the skills
# are moved here). Refreshed by update-all.sh. Not vendored: the CLI owns the
# layout and the content is sha256-verified against 21st.dev's manifest.
skills-external/21st-*/
# npx `skills add` project-scope artifacts — darwin-skill copies itself into
# the repo's .agents/ and writes skills-lock.json at root. Our own agents live
+3 -2
View File
@@ -67,8 +67,9 @@ regexTarget = "line"
regexes = [
'''X-Amz-Credential=AKIA[0-9A-Z]{16}''',
'''private-user-images\.githubusercontent\.com/[^"]*\?jwt=''',
'''MAGIC_API_KEY=abc123''',
# magic MCP docs example — base64 of "the ..." ASCII sample text.
# Docs/test example — base64 of the "the ..." ASCII sample text, never a key.
# (The MAGIC_API_KEY=abc123 placeholder that sat here went with the magic
# MCP, removed 2026-09-22 when 21st.dev moved to a CLI with no API key.)
'''clientKey = 'dGhlIH[A-Za-z0-9+/=]*'''',
]
+316
View File
@@ -6,6 +6,322 @@ Format follows [Keep a Changelog](https://keepachangelog.com/).
## [Unreleased]
### Added
- **`make doctor` reports the Playwright browser cache** — a read-only
`Playwright browsers` section listing cache size, which registered
Playwright install requires each cached browser revision, and counts of
unreferenced directories and broken links. Report only: nothing is
pruned, since Playwright's own `install` already unions the required set
across every registered install.
- `lib/gstack-playwright.sh` — the gstack Playwright helpers as a shared
lib (OS-support bump, submodule-update wrapper, cache report), sourced by
`install-plugins.sh`, `update-all.sh` and `doctor.sh`, covered by
`lib/tests/gstack-playwright.test.sh`.
- **`doctor.sh` inspects the `autoMode` block**: warns when a classifier
list drops the built-in entries (no `"$defaults"`) and when the
user-scope `environment` names a git repo other than the config repo.
Neither defect is visible from the deny count, until now the only
permission signal `doctor.sh` had.
- **`templates/settings/SETTINGS.md` documents `autoMode`**: the four
classifier lists, `$defaults` splice semantics, `classifyAllShell`, the
user-scope vs project-scope rule, and why `ask` is the wrong tier for a
destructive command under auto mode.
- **21st.dev moved from an MCP server to a CLI.** `install-plugins.sh` Step 8.7
installs `@21st-dev/cli` globally (pinned in `plugins.lock.json`), offers
`21st login` in an interactive terminal only, and stages the 7-skill pack
into `skills-external/21st-*`. `update-all.sh` refreshes both. The pack
ships disabled, same policy the MCP had.
- `lib/toggle-external.sh` manages `21st` as a skill pack (glob-derived from
`skills-external/21st-*`, parked under plain names so `profile.sh`'s
external park/restore stays interoperable). `magic` is gone from the
managed tools.
- The five design skills (`21st-ui-build`, `-ui-explore`, `-ui-review`,
`-cli-use`, `-ai`) are in the `design`, `web`, `web-full` and `full`
profiles and in `profile.sh`'s `MANAGED_EXTERNALS`; `21st-registry` and
`21st-design-sync` are installed but left parked.
- `autoMode.soft_deny` gains one entry for the outward-facing 21st verbs
(`publish*`, `submit`, `edit`, `delete`, `remove-from-catalog`,
`profile set|upload`) — publishing puts a component on a public listing.
That tier rather than `ask`, per LRN-153.
- `lib/design-gate.md` §5: a suggest-only check, same shape as the §4
animation-library one. When impeccable is active and the frontend project
has no `PRODUCT.md` at its root, the gate proposes `/impeccable init` once
and never runs it itself (it interviews the user). Skipped for single
component reviews and non-UI work.
### Changed
- **Design gate: `magic` → the `21st` CLI in the required-manual slot.**
`design.profile`'s `GATE-BLOCK` now lists `21st` (CLI channel) and
`21st-ui-build` (the pack's canary on the skill channel); a missing CLI
trips the gate with `npm i -g @21st-dev/cli` + `21st login` instead of the
old `MAGIC_API_KEY` hint. `design-tool-gate.sh` also repairs `PATH` for the
npm global bin, whose absence in a hook's sanitized `PATH` would otherwise
read as "21st missing" (the existing repair only fired when `claude` itself
was unresolvable).
- `profile.sh`'s `MANAGED_MCPS` is empty: no MCP server is auto-toggled any
more. The `mcp` type stays supported for an advisory profile entry.
- **`/deploy` hand-back: one physical line per command, then a post-deploy
tests block.** Every command in the checklist is emitted on exactly one
line, however long; a legacy `\` continuation in the runbook is joined at
instantiation, and bootstrap and learn patches write runbook lines the same
way (`templates/deploy/PROCEDURE.md` header updated). After the checklist
the hand-back now carries a "Post-deploy tests" block derived from the
delta diff: by-hand checks (action → observable result, each tied to a
delta file) plus Suggestions (checks the runbook does not do yet, gaps
spotted between delta files). Cold-resume re-display and re-hand-back
regenerate both. RED/GREEN tested on a scratch runbook (4/4 baseline runs
reproduced the continuation verbatim and printed no tests).
- **Docker and node go through the classifier with a framing, instead of
an inert `ask` tier.** `Bash(docker run|exec *)`, `Bash(docker[-| ]compose
up*)` and `Bash(node -e *)` leave `permissions.ask` (no prompt under auto
mode, re-verified on 2.1.273). A new `autoMode.allow` list, `$defaults`
first, names the two routine cases the built-in `Remote Shell Writes` /
`Production Reads` rules were catching: `docker exec`/`run`/`compose`
against a local dev container whose name does not carry `prod`, running a
SQL file or script from the repo inside it; and project-local node
(`node <file>`, `npm run`, `npx`/`pnpm exec` of a lockfile-declared
package, effects inside the cwd). Two `soft_deny` entries frame what that
opens: docker data destruction (`rm -f`, `volume rm`/`prune`, `system
prune`, `compose down -v`, `--privileged`, bind mounts outside the cwd)
and undeclared node packages (`npx`/`dlx` of a package absent from the
lockfile, `npm install <name>`). `SETTINGS.md` gains the `autoMode.allow`
tier and the reason a static `Bash(node *)` rule cannot do this job.
- **Ask, don't guess: the orchestrators ask about open choices instead of
settling them.** `CLAUDE.global.md` replaces "one question upfront, never
mid-task" with: a choice visible in the result, a name that becomes
public, or a scope the request does not settle → ask, even mid-task;
internal technical choices stay Claude's. `lib/contract-interview.md`
STEP 2 becomes CLARIFY: pass A (the three gap checks, at contract time)
and pass B (the open-choice sweep in three classes, run once at each
flow's PLAN step, no question cap, over-5 guard, "you decide" recorded as
delegated). New MID-RUN CLARIFICATION section: an executor's
`NEED-DECISION` carries a `CLASS:` tag; visible / public-name / scope go
to the user verbatim, internal is decided in the loop; answers land in
the contract `[gated]`. New HOW TO ASK section (LRN-102). `/feat`,
`/bugfix`, `/hotfix`, `/ship-feature`, `/init-project` wire pass B at
their plan step; `/feat` and `/bugfix` stop deciding `NEED-DECISION`
themselves; `/hotfix` drops "zero questions ever" and allows one
re-dispatch for a class-tagged BLOCKED; the interviewer never ships a
visible / public-name / scope item as `(assumed)`; feater, bugfixer and
hotfixer report the class. Locks updated in the `contract-verifier`,
`loops-light` and `gates` tests.
- **The classifier, not `permissions.ask`, now guards destructive shell
work** (BDR-090). Ten rules left the static tiers: `rsync`, `kill -9`,
`killall`, `pkill` out of `deny`, and `python3 -c`, `python -c`,
`xargs`, `sed`, `cp`, `mv` out of `ask`. Under `defaultMode: auto` an
`ask` rule raises no prompt ([[LRN-146]]), so that tier was gating
nothing anyway. Cover is now `autoMode.soft_deny`, which the classifier
enforces and an explicit instruction clears: writes outside the working
directory, `rsync --delete`, SIGKILL and kill-by-name, in-place edits
spanning more than one file, directory moves, and inline interpreters
or `xargs` that delete or write outside the cwd. Intent clears a soft
block for the current turn only.
- **`autoMode.hard_deny` added** for the three classes no command pattern
can express: secret exfiltration (a read and a send, separate steps,
possibly turns apart), production deployment (deploy scripts, lftp/FTP
pushes, any `prod` target), and disarming the guardrails (weakening a
deny list, `--no-verify`, removing the pre-commit hook,
`bypassPermissions`). Adding a restriction stays allowed; removing one
does not. No instruction clears these.
- **impeccable installs at `--scope=global`, subagents included, and the
pin fails safe.** `install-plugins.sh` Step 8d no longer stages a
`--scope=project` install in a tmpdir and moves the skill directory alone.
The installer writes through the `~/.claude/skills` and `~/.claude/agents`
symlinks straight into the repo: `skills/impeccable` plus the four
`agents/impeccable-*.md`, both gitignored and machine-owned, which is what
the manual `--scope=global` command already did. The step refuses to run
before `make link` has created those symlinks (an install before them
materializes real directories that `link.sh` then refuses to replace),
keeps a profile-parked copy parked, reports the skill version and agent
count, and prints the per-project `/impeccable init` hint. A pinned
install that fails falls back to `impeccable@latest` with a warning to
bump `plugins.lock.json`. `update-all.sh` follows the same shape.
`plugins.lock.json` pin 3.2.0 → 4.1.0 (the CLI only: the skill dist and
the engine binary have their own release tracks). `link.sh` drops
impeccable from `EXTERNAL_SKILLS`; `skills-external/impeccable/` is gone.
### Security
- **Ten secret-reader deny rules added**: `sed`, `awk`, `cut`, `tr`,
`sort`, `uniq`, `diff`, `od`, `xxd`, `strings` against `.env*`. Six of
those tools sat in `permissions.allow`, so reading a `.env` through
them triggered nothing. Same shape and same known gap as the existing
`Bash(grep * .env*)` family: a `cat .env | sed` pipe still slips past,
which is what the `hard_deny` exfiltration rule is there to catch.
### Removed
- **`magic` MCP (`@21st-dev/magic`) and `MAGIC_API_KEY`**, with the two risks
attached to them: the unauthenticated `127.0.0.1` callback server
`21st_magic_component_builder` opened (LRN-110) and the plaintext key copy
that `claude mcp add --env` wrote into `~/.claude.json` (BDR-026/057). Gone
with it: the 4 `mcp__magic__*` `permissions.ask` entries (BDR-059), the
`MAGIC_API_KEY` block in `.env.example`, `link.sh`'s missing-key warning,
and the dead `MAGIC_API_KEY=abc123` gitleaks allowlist regex.
### Fixed
- **`make update` no longer drops the Playwright OS-support bump** — a
gstack submodule update used to leave the bump unapplied until the next
`make plugin`, the open caveat of BDR-029. `update-all.sh` now goes
through `gstack_submodule_update_with_bump`, which re-applies it after a
successful update and returns non-zero on failure so the existing warn
arm still fires. Two latent bugs travelled with the extracted code: the
ostag capture exited 1 on every non-Ubuntu host and aborted its caller
under inherited `errexit`, and the `bun` calls had no timeout.
- **`autoMode.environment` no longer describes one project from the
user-scope file**: the block named a specific repo, its FTP deploy
target and its customer data, while `link.sh` symlinks this file to
`~/.claude/settings.json` where it reaches every project. The global
block now states machine-level facts only (self-hosted Gitea, gitflow
protection, `~/.claude/.env` as the single secret source, no CI), and
the project-specific facts moved to that project's gitignored
`.claude/settings.local.json`. Both lists now open with `"$defaults"`,
which the original omitted, so the built-in entries are inherited
rather than replaced.
- `README.md` no longer claims the `ask` tier makes every `mcp__magic__*`
call "require a live confirmation and can never auto-execute". That
holds under `defaultMode: default`, not under this config's `auto`. The
paragraph now separates what is verified from what is not, and names
`deny` as the only tier the classifier cannot lift.
- **`make plugin` never installed impeccable.** Three defects. The 3.2.0
pin had rotted upstream: the CLI fetches its skill dist at install time and
that release's artifact is gone (`Download failed: invalid zip data`),
which the step reported as "run it yourself" on every run. The
project-scope staging dropped the four subagents the same install writes.
And `/impeccable init` was never announced. A fourth, found while probing
the fix: once a copy is already installed, a rotted pin exits 0
(`Could not check for skill updates … Existing skills were left
unchanged`), byte-identical on disk to a genuine "Skills are up to date"
rerun, so `imp_install` now reads the installer output instead of trusting
the exit code or a version compare. Verified with the real installer in a
sandbox HOME: fresh install, rotted pin over a copy (fallback fires), same
pin rerun (no false warning), parked copy plus rotted pin (fallback, then
returned to `skills-disabled/`).
## [1.5.0] — 2026-09-13
### Added
- **Attention signals on the terminal (BDR-087)** — new
`hooks/notify-attention.sh`, wired on `Notification` (input-needed
matcher) and on `Stop` (no matcher). Returns a double BEL plus an
OSC 777 toast through the `terminalSequence` JSON field, since hooks
have no controlling TTY. Signal only: `suppressOutput`, exit 0, zero
control-flow effect, which is what separates it from the `decision:
"block"` Stop hook [[BDR-083]] refused. Each event reaches the toast
as a readable label instead of a snake_case type; events needing no
attention (`agent_completed`, `auth_success`) exit silently; a turn
that ends with `background_tasks` still running stays quiet and
signals at the real end. Client-side prerequisites over Remote-SSH
are documented in the script header ([[BLK-020]]): VS Code
`accessibility.signals.terminalBell` for the beep, an OSC notifier
extension for the Windows toast.
- **User permanent rules (BDR-085)** — three new rules/ files from the
user's rule text: `writing-style.md` (always-on: em-dash ban, no slop
vocabulary, no hedging chains, deliverable self-check),
`web-building.md` (path-scoped: design anti-defaults + public-site done
checklist), `web-security.md` (path-scoped: RLS, service-key/client
split, IDOR, cookie flags, rate limiting — extends §Security, no dup).
Project CLAUDE.md rules/ doctrine gains the 320-budget exception.
- **/tour multi-project parallel fan-out (BDR-084)** — two or more
project paths now dispatch one runner per repo in a single message
(independent working trees, nothing collides) instead of processing
them one by one. The runner inherits the session model (no pin — it
carries tour's reflection); every agent inside keeps its defined tier.
A dead runner surfaces as an explicit `RUNNER FAILED` summary row; the
gated capitalize offer stays in the main loop. Bounded LRN-083
derogation recorded in BDR-084. Census §12: 6 locks, flip-tested.
Mechanics proven first: nested probe, 3 overlapping agent windows,
9.1s vs ~18s sequential.
- **Contract gates — deterministic floor under the fresh verifier (BDR-083)** —
an acceptance criterion can now carry an oracle (`CHECK:` command +
`EXPECT:` success-only marker + `EVIDENCE:` slot). `lib/gates.sh run
<contract>` executes them fail-closed — MET requires exit 0 **and** the
marker — and writes the outcome back into the contract, so the fresh
verifier reads evidence as fact instead of trusting the executor's report.
New `GATE 0` in `lib/verify-secure-loop.md` runs the floor before any
verifier is dispatched: a red build no longer costs an LLM dispatch to
discover. `ABANDON: <id> <reason>` makes an impossible criterion a visible
handoff that blocks `CONFORME` and routes to the human gate (new verifier
verdict `ABANDONED(n)`). `feater` and `bugfixer` gain a four-pass
completion discipline, scoped so it can never widen the contract.
Adapted from the `unlazy` skill (Leonxlnx/unlazy, MIT); its Stop hook,
approval store, `.unlazy/` tree, depth-tree arithmetic and Node checker
were deliberately refused — see BDR-083 for each reason.
The four orchestrator skills (`feat`, `bugfix`, `ship-feature`,
`init-project`) restate the GATE 0 bullet ahead of GATE 1 (locked);
hotfix explicitly runs no floor. Behavioral RED: 16/16 fresh unprimed
runs followed the new doctrine (EVAL-027).
64 new assertions in `lib/tests/gates.test.sh`.
- **`lib/tests/seo-geo-contract.test.sh`** — census locking the seo/geo
agent ⇄ dispatcher machine contract: judge verdict grammar, FIX BUNDLE +
READY-TO-APPLY sentinel, signals handoff, every STEP header (interiors
included), bundle item fields, score labels, scoring blocks, envelope
keys (46→71 assertions across the C1 chantier).
### Changed
- **Skill and agent quality campaign, 54 units (BDR-086)** — full darwin
v2.1 pass over the 31 personal skill-systems and 23 agents, excluding
the gstack/external symlinks and machine-owned units. Fresh baseline
mean 83.4; the 13 units under the user-set threshold of 80 were
optimized to completion, and verified defects in above-threshold units
were fixed in a grouped pass rather than left to ship because the score
was good enough. Every round was validated by a paired 3-judge majority
reading before and after in one call: 36 unit-round verdicts, 24 batch
verdicts, all better, zero reverts. Full report and residual findings:
`.claude/audits/DARWIN-2026-08-26.md`.
- **seo-analyzer + geo-analyzer de-prescribed for Opus 5 (BDR-082)** —
process choreography converted to when-guidance under an
audience×mode-range invariant; self-output verification demands removed
(the score-engine "run it twice" became a conditional integrity guard);
two pre-BDR-061 vestigial rules fixed; P0/MANDATORY/ALWAYS caps softened
to plain content rules. Machine contract byte-frozen and locked by the
new `lib/tests/seo-geo-contract.test.sh` census (71 locks, flip-proven);
proven by a controlled before/after `/seo` dogfood — judge replay on
frozen signals, 42/42 presence assertions on both runs, blind structural
reader: interchangeable, recall improved.
- **Global instruction layer recalibrated for the Claude 5 family (BDR-081)** —
delegation block is now model-neutral when-guidance (the Opus 4.8
under-delegation counter inverted on Opus 5, which over-delegates and gets
an injected harness cap); "staff engineer" self-check bar dropped (Opus 5
over-verification trigger); finish-whole-task clause added to Deviations;
written-deliverable length rule added. 308/320 lines.
- **Default session model is now `opus[1m]`** (was `claude-fable-5[1m]`).
- **`skills-external/emil-design-eng/` untracked** — the file is curl'd
from upstream by `install-plugins.sh` when absent and re-fetched by
every `update-all.sh` run, so tracking it produced a repo diff on each
upstream edit. Same category as `frontend-design/` and `impeccable/`,
already ignored on that rationale; a fresh clone re-fetches it.
`design-motion-principles/` has the same overwrite behaviour but no
bootstrap clone yet, so it stays tracked until that gap closes.
### Fixed
- **hotfix wiped tolerated in-progress edits on its revert path** — every
failure branch ran `git restore .`, destroying user edits the run had
tolerated. Now a `git stash create` pre-flight snapshot plus a
file-scoped restore, and the security gate is fresh-dispatch only.
- **skills-perso listed 8 of 31 personal skills** — detection rebuilt on
the `link.sh` symlink convention (symlink = external, real dir =
personal, gitignored = machine-generated). Live result 31/31, no false
positives.
- **plan-challenger** — `ERROR` joined the load-bearing verdict grammar
(STEP 1 emitted it, the parser enum omitted it); grounded-but-uncertain
findings now file as `[MINOR]` with the uncertainty stated, instead of
being self-censored (Opus 5 follows conservative-reporting clauses
literally).
- **design-toolchain hook** — dropped `\bux\b` (2 French-prose false
positives; 3rd tightening pass, series LRN-1005/1007); `\bui\b` kept and
locked by a must-fire test row.
- **Agent and skill defects found by the campaign's judges** —
`init-project` allowed-tools lacked `Agent` and `Skill` while every step
dispatches; `commit-change` conflict grep now covers all 7 unmerged
codes; `tour --report-only` no longer commits; `harden` severity defers
to the calibrated guide and the late SSL Labs grade has an assigned
actor; handover writers' stale chapter refs corrected and the anchor
gate ordered; `security-auditor` documents the hotfix no-verifier
carve-out; `close` enumerates STEP 5C and passes `--no-push` through;
`prune-memory` drops a false "v1-untested" note; `code-clean` attributes
its executor correctly; plugin-check and onboard fixtures de-drifted.
## [1.4.0] — 2026-07-22
### Added
+22 -12
View File
@@ -22,6 +22,8 @@ Apply unless repo-specific instructions override.
- Document intent, not mechanics. Use project doc style (docstring, JSDoc…).
- Explicit, consistent, meaningful names. Straight control flow,
no hidden side effects.
- Written deliverables (docs, reports, .md): length matched to what
the task needs — no filler sections, no boilerplate summaries.
## Refactoring
- Priority: safety → readability → consistency.
@@ -40,19 +42,26 @@ Apply unless repo-specific instructions override.
- Confirm before implementing only when real trade-offs exist (multiple
valid approaches, breaking change, destructive action) — else proceed.
- Minimal changes unless broader refactor requested. State trade-offs.
- Sub-agents keep main context clean — one task per sub-agent.
More compute on hard problems. Task fans out across independent
items (many files, parallel searches, multi-point checks) → delegate
to sub-agents, don't iterate serially. Default to delegation for
multi-file exploration. Counters model tendency to under-delegate.
- One question upfront if needed — don't interrupt mid-task.
- Sub-agents: one task per sub-agent, main context stays clean.
Delegate genuinely independent, sizeable tracks (wide multi-file
exploration, parallel audits) — not work doable in a few tool
calls. Skill-mandated gates (fresh verifier/security/challenge)
always dispatch as written. Don't redo delegated work by hand —
failed gates re-dispatch fresh executors instead.
- Ask rather than guess. A choice visible in the result (placement,
wording, order, behavior), a name that becomes public (command, flag,
endpoint, file), or a scope the request does not settle → ask, even
mid-task. Batch what can be batched. Internal technical choices with
no observable effect stay yours.
*Exception: skill-mandated gates and checkpoints (orchestrator
validation gates, approval gates, darwin checkpoints) always fire.*
- Bug received → fix directly: check logs, find root cause, resolve
autonomously.
autonomously; a visible choice in the fix still gets asked.
- Something goes wrong → STOP, re-plan. Never push through.
- Deviations: minor or clearly justified → do, explain after.
Significant or shaky justification → ask before deviating.
Finish the whole task: blocked on an independent sub-part → do
the rest, state what's missing. Gone WRONG → still STOP, re-plan.
- Root causes only. No temp fixes. Never assume — verify paths, APIs,
variables before use.
@@ -77,7 +86,6 @@ Apply unless repo-specific instructions override.
2. Report what verified, what not.
3. List remaining risks, surviving deviations.
4. Don't mark complete without proof it works.
Bar: "would staff engineer approve?"
5. Correction or notable event → capitalize to right registry
(see "Memory registries").
@@ -282,16 +290,18 @@ OR a design/UI request — not the keyword "design" alone in a prompt. Single
source for design routing; the design-toolchain hook reinforces it.
- Trivial (≤2 files, one cosmetic value) → /hotfix, no toolchain.
- Build UI (component, page, redesign) → ui-ux-pro-max + frontend-design
(anti-slop) + Magic MCP /ui + emil-design-eng (polish) +
design-motion-principles (if motion) + design-html (if static).
(anti-slop) + 21st-ui-build (catalog + generation) + emil-design-eng
(polish) + design-motion-principles (if motion) + design-html (if static).
Post-build floor: `npx impeccable detect <files>` (45 deterministic
anti-slop rules, exit 2 = findings) when impeccable installed.
- Design system / brand → design-consultation first, then the build tools.
- Review / audit → design-review + emil-design-eng + design-motion-principles
+ /impeccable audit|critique (skill) + `impeccable detect` floor.
+ 21st-ui-review + /impeccable audit|critique + `impeccable detect` floor.
Scope doubt → don't silently skip: ask, or default to Build tier.
Gate: lightweight skills run `~/.claude/lib/design-gate.md`; orchestrators via
plugin-check. Magic MCP costs API calls — generation, not micro-tweaks.
plugin-check. 21st = CLI (`npm i -g @21st-dev/cli`, `21st login`), no MCP,
no API key. Search is free; `21st get` and `21st generate` are metered —
generation, not micro-tweaks.
## graphify
+32 -1
View File
@@ -17,7 +17,9 @@ A rule WITH `paths:` YAML frontmatter (glob list) loads lazily — only when
Claude reads a file matching a glob; a rule WITHOUT it loads at session
start, same cost as the global memory. Extract from CLAUDE.global.md only
what can be path-scoped (the token win) or what is generated; always-on
doctrine stays in CLAUDE.global.md. `paths:` globs match against the
doctrine stays in CLAUDE.global.md. Exception: a standalone user-authored
rule set that would bust the 320-line density budget may live here WITHOUT
`paths:` (always-on load) — writing-style.md (BDR-085). `paths:` globs match against the
CURRENT project's tree — a broad glob (e.g. `rules/**`) can fire in foreign
projects; keep rule bodies tiny.
Docs: https://code.claude.com/docs/en/memory.md#path-specific-rules
@@ -28,6 +30,35 @@ install-plugins.sh STEP ctx7 purges it right after; the find-docs skill is
the single ctx7 surface. If it reappears (manual `ctx7 setup`), delete it
or re-run `make plugin`.
## Machine-owned: the vendored graphify skill
`skills/graphify/SKILL.md`, `skills/graphify/references/` and
`.graphify_version` are written by `graphify claude install`
(`install-plugins.sh` STEP graphify), which lands in the repo because
`~/.claude/skills` is a symlink to `skills/`. They are gitignored: a
`pipx upgrade graphifyy` used to dirty the tree and cost a
`chore(graphify): sync vendored skill X -> Y` commit each time.
Two graphify commands, easy to confuse, and only one restores the skill:
- `graphify install --platform claude` copies SKILL.md + `references/` +
`.graphify_version` into `skills/graphify/`. Touches nothing else.
This is the recovery command.
- `graphify claude install` writes the CLAUDE.md graphify section and the
`.claude/settings.json` PreToolUse hooks. It **rewrites both guarded
configs** (EVAL-020, verified again 2026-09-15), so revert them after. It does NOT copy
the skill.
`make plugin` runs both (`install-plugins.sh` STEP graphify) behind the
guarded-config EXIT trap, so a fresh clone is covered.
Trade-off accepted: an upstream release can now change the skill's prompt
with no diff to review. `skills/graphify/test-prompts.json` is hand-written
for darwin and stays tracked.
Gotcha, learned the hard way: `git rm --cached` keeps the working file,
but if the branch you merge into still tracks it, the merge deletes it
from disk. Untrack and merge, then restore with the command above.
## Transient planning artifacts
`docs/superpowers/specs/**` and `docs/superpowers/plans/**` are run-time
+45 -18
View File
@@ -253,10 +253,9 @@ in `env`, `command`, `args`, `url`, and `headers` — for both project (`.mcp.js
and user (`~/.claude.json`) scope. Use that instead of a literal value:
```bash
MAGIC_API_KEY=<Enter your magic api key here from https://21st.dev/settings/api-keys >
# single-quoted so bash doesn't expand it; Claude Code expands it at
# launch, reading the var from its own process environment:
claude mcp add magic --scope user --env 'API_KEY=${MAGIC_API_KEY}' -- npx -y @21st-dev/magic@latest
claude mcp add <name> --scope user --env 'API_KEY=${SOME_API_KEY}' -- <command>
```
The var still has to exist in the **environment of the process that starts
@@ -265,8 +264,12 @@ would defeat the point (every subprocess, every stray `env`/`printenv`, would
then see it). This repo's `~/.bashrc` instead wraps the `claude` command
itself: a `claude()` shell function sources `~/.claude/.env` into a subshell
and `exec`s the real binary, so the var reaches `claude` and its children only
— never the ambient shell. See `lib/toggle-external.sh`'s `magic` case for
the pattern to copy for a new MCP server.
— never the ambient shell.
This config currently registers no MCP server at all. The one it used to
carry, `@21st-dev/magic`, is gone: 21st.dev replaced it with a plain CLI (see
below), so there is no key left to protect by reference. The pattern stays
documented for the next MCP server that needs a secret.
There is no `claude mcp add` flag that writes the reference form for you —
the `${VAR}` syntax has to be typed by hand (or via a wrapper script), same as
@@ -292,20 +295,44 @@ Then run the one-time consent flow: `make seo-connect` (per-label token
store, multi-site safe). Missing credentials never break an audit — `/seo`
degrades gracefully to anonymous PageSpeed lab data.
### magic MCP (`@21st-dev/magic`) — known callback-injection risk
### 21st.dev CLI (replaces the magic MCP)
`21st_magic_component_builder` opens an **unauthenticated** local callback
server (`127.0.0.1:9221+`, `Access-Control-Allow-Origin: *`, no token/origin
check) for up to 10 minutes per call; any local process or open browser tab
can `POST` to it and that body is injected **verbatim** into the tool result
the model consumes (job8 audit, `dist/utils/callback-server.js:36`). This is
in the third-party package's code, not this repo's config — **we don't patch
it**. The mitigation lives entirely on our side: `settings.json`
`permissions.ask` explicitly lists all 4 `mcp__magic__*` tools,
so every call — builder included — requires a live confirmation and can
never auto-execute. Don't allowlist
`21st_magic_component_builder` or `21st_magic_component_refiner` (arbitrary
absolute-path read → vendor exfil, same audit) under any circumstance.
`@21st-dev/cli` (bin `21st`) supersedes the `@21st-dev/magic` MCP server that
this config used to register. Same endpoint, one browser login, no API key,
and nothing loaded into a session that isn't using it:
```bash
npm i -g @21st-dev/cli
21st login # browser flow, token saved in ~/.config/21st
```
`make plugin` does both (Step 8.7 installs the CLI, then offers the login in
an interactive terminal) and installs the skill pack that drives it:
`21st-ui-build`, `-ui-explore`, `-ui-review`, `-cli-use`, `-ai`, plus the two
publishing skills `-registry` and `-design-sync`. The pack is disabled by
default, the same policy the MCP had. `/profile design` turns on the five
design skills; `bash lib/toggle-external.sh enable 21st` turns on all seven.
The pack is machine-owned and gitignored. It cannot be installed the way
upstream documents it (`21st install-skill`, i.e. `21st skills install
--global`): that writes into `~/.claude/skills/`, and the installer refuses to
follow a symlink anywhere on that path, while `~/.claude/skills` is itself a
symlink to this repo's `skills/`. So the install runs under a throwaway `HOME`
and the result is moved into `skills-external/21st-*`, where
`toggle-external.sh` and `profile.sh` symlink it in on demand.
Two risks from the MCP era go away with it. The unauthenticated local callback
server `21st_magic_component_builder` opened (`127.0.0.1:9221+`, CORS `*`, a
10-minute local prompt-injection window, job8 audit / LRN-110). And the API
key that `claude mcp add --env` materialized into `~/.claude.json`.
The permission gate is now one `autoMode.soft_deny` entry covering the
outward-facing verbs (`21st publish*`, `submit`, `edit`, `delete`,
`remove-from-catalog`, `profile set|upload`), because publishing a component
puts it on a public listing under your account. That tier rather than `ask`:
under `defaultMode: auto` (this config's default) `ask` rules were observed
auto-approving with no prompt raised (LRN-153), so an `ask` entry would have
declared an intent without gating anything.
---
@@ -337,7 +364,7 @@ make profile-reset # re-enable all gstack skills
make new-skill name=myskill # scaffold agent + skill files
```
`doctor.sh` checks: symlinks, GStack submodule, prerequisites (git, Node, Cargo, Python, Claude Code), plugins, permissions, token budget, config consistency.
`doctor.sh` checks: symlinks, GStack submodule, Playwright browser cache, prerequisites (git, Node, Cargo, Python, Claude Code), plugins, permissions, token budget, config consistency.
---
+6 -7
View File
@@ -25,14 +25,13 @@ Produce a clear analysis without proposing solutions.
---
## TASKS
## TASKS (in order — each step feeds the OUTPUT section named)
- Identify relevant parts of the codebase
- Understand current behavior
- List dependencies
- Highlight constraints
- Detect risks
- Identify ambiguities
1. **Locate** — find the relevant parts of the codebase (Glob/Grep from the target) → file list
2. **Understand** — read them; describe current behavior as-is → CONTEXT, KEY COMPONENTS
3. **Map dependencies** — imports, call sites, data flow in/out → KEY COMPONENTS roles
4. **Constrain** — invariants, contracts, conventions the code obeys → CONSTRAINTS
5. **Assess** — risks with probability, then ambiguities → RISKS, OPEN QUESTIONS
---
+23 -3
View File
@@ -27,8 +27,9 @@ Every choice was made in the plan or is a NEED-DECISION to report.
- Apply the FIX PLAN to the letter — fix the ROOT CAUSE named in DIAGNOSIS,
not the symptom. A plan hole or an open choice (naming, data shape, API
surface, dependency) → STOP, report `NEED-DECISION` with the precise
question. Never re-investigate or improvise a different fix.
surface, dependency, a user-visible choice such as placement, wording or
behavior) → STOP, report `NEED-DECISION` with the precise question and
its `CLASS:`. Never re-investigate or improvise a different fix.
- Stay inside the contract FILE SCOPE. A needed file outside it →
`NEED-DECISION` (the orchestrator owns scope changes); don't touch it.
- Add or update the regression test the plan names — it must fail before the
@@ -46,6 +47,24 @@ Every choice was made in the plan or is a NEED-DECISION to report.
security/verifier dispatch, editing `.claude/**` or memory registries, user
questions (you cannot ask — report instead), attribution trailers of any kind.
## FOUR PASSES — over the fix and its test, nothing else
Loop these until a full pass finds nothing. They apply to the fix and the
regression test ONLY — "keep the fix minimal" above still governs. They make
the minimal fix COMPLETE; they never widen it.
1. **Complete.** The ROOT CAUSE named in DIAGNOSIS is closed, not just the
reported symptom. No placeholder, no deferred remainder.
2. **Expert reread.** Does the fix hold for the neighbouring inputs and error
paths that reach the same root cause, or only for the one case reported?
3. **Negative control.** Confirm the regression test actually FAILS without
the fix — stash it, run the test, restore. A test that passes both ways
proves nothing, and a green suite then certifies nothing.
4. **Polish.** Naming and comments on what you touched. Nothing else.
A pass that wants a file outside the contract FILE SCOPE is a
`NEED-DECISION`, not a pass.
## OUTPUT — end with exactly this report (your final message)
```
@@ -55,5 +74,6 @@ FILE(S) : <created/modified paths>
TEST(S) : <regression test added/updated + final suite run result, verbatim line>
SMOKE : <build/typecheck result if run, or n/a>
NOTES : <DONE: deviations (must be none) | NEED-DECISION: the exact
question + the options you see | BLOCKED: the blocker verbatim>
question + the options you see + CLASS: visible | public-name |
scope | internal | BLOCKED: the blocker verbatim>
```
+1 -1
View File
@@ -732,7 +732,7 @@ write `.claude/audits/THRESHOLD-OVERRIDE.md` documenting:
- Top 3 unresolved issues per axis
- User's stated reason
This file is referenced in §4 of the client doc ("Ce qui vous reste à faire")
This file is referenced in §5 of the client doc ("Ce qui vous reste à faire")
so the client knows what's still below the bar.
If `ALL_PASS = false`:
+24 -3
View File
@@ -37,8 +37,9 @@ report below is optional on this path (the dispatcher needs the edit applied
## EXECUTION RULES
- Follow the plan to the letter. A plan hole or an open choice (naming,
data shape, API surface, dependency) → STOP, report `NEED-DECISION` with
the precise question. Never improvise a design decision.
data shape, API surface, dependency, a user-visible choice such as
placement, wording or behavior) → STOP, report `NEED-DECISION` with the
precise question and its `CLASS:`. Never improvise a design decision.
- Stay inside the contract FILE SCOPE. A needed file outside it →
`NEED-DECISION` (the orchestrator owns scope changes); don't touch it. On
the applier path the scope is the files named in the bundle item — apply
@@ -57,6 +58,25 @@ report below is optional on this path (the dispatcher needs the edit applied
editing `.claude/**` or memory registries, user questions (you cannot
ask — report instead), attribution trailers of any kind.
## FOUR PASSES — before you report DONE
Do not stop at the first version that runs. Loop these until a full pass
finds nothing:
1. **Complete.** The whole deliverable the plan names is implemented. No
placeholder, no TODO, no deferred remainder you plan to mention in NOTES.
2. **Expert reread.** Read it as someone who owns this codebase. Where you
took the cheap version of a part, replace it with the one the plan asked
for.
3. **Defect hunt.** Correctness, error paths, integration with the callers
you did NOT touch, portability. Fix what you find.
4. **Polish.** Low-cost only: naming, comment density, dead code you
introduced.
Every pass stays inside the plan and the contract FILE SCOPE. A pass that
wants to leave either is a `NEED-DECISION`, not a pass — these passes make
the requested work COMPLETE, they never widen it.
## OUTPUT — end with exactly this report (your final message)
```
@@ -65,5 +85,6 @@ STATUS : DONE | NEED-DECISION | BLOCKED
FILES : <created/modified paths>
TESTS : <added/updated + final suite run result, verbatim line>
NOTES : <DONE: deviations (must be none) | NEED-DECISION: the exact
question + the options you see | BLOCKED: the blocker verbatim>
question + the options you see + CLASS: visible | public-name |
scope | internal | BLOCKED: the blocker verbatim>
```
+11 -10
View File
@@ -45,7 +45,7 @@ This anchors the agent's output so the user can compare audits over time.
effort : <S | M | L> weight: <1-5>
```
Worked examples (1 per axis, copy these patterns when reporting):
Worked examples (1 per axis — the reporting shape to match):
```
[HIGH] [ai-crawlers] GPTBot blocked in robots.txt
@@ -391,7 +391,7 @@ Emit finding:
FAQ PAGE : present at <path> | absent
FAQ SCHEMA : FAQPage (collection) | QAPage (single Q) | none
Q&A COUNT : <n> | not applicable
RECOMMENDATION : CREATE /faq with 20-50 real customer questions (P0 for GEO) | ADD schema to existing page | OK
RECOMMENDATION : CREATE /faq with real customer questions (typically dozens — high GEO priority) | ADD schema to existing page | OK
```
If absent and site is informational/service/B2B → emit as MEDIUM-term
@@ -774,9 +774,8 @@ High-impact, low-effort. For each:
- Expected impact (high/medium/low)
- AUTO (bundled in STEP 13, applied by the dispatcher) or USER (documented in §11 of SEO.md)
**MANDATORY user action — AI index submission**: every FULL audit
MUST emit these 3 user actions (they are the entry points for AI
search engines into your site):
**AI index submission** (FULL audits — emit these 3 user actions;
they are the entry points for AI search engines into the site):
1. **Bing Webmaster Tools** — submit + verify sitemap. Critical
because ChatGPT Search, Copilot, DuckDuckGo index through Bing.
@@ -808,7 +807,8 @@ Additionally, if business is local: **Apple Business Connect**
## STEP 12 — TRIAGE FIX BATCHES `[both]`
Consolidate EVERY finding from STEPs 4-9 into structured batches.
Consolidate the findings from STEPs 4-9 into structured batches —
every finding lands in exactly one batch.
| Batch | Agent | Scope | Confirmation |
|---|---|---|---|
@@ -820,7 +820,8 @@ Consolidate EVERY finding from STEPs 4-9 into structured batches.
| **G6 — Entity @id + sameAs wiring** | `feater` | JSON-LD graph restructure | No |
| **G7 — User actions** | documented in §11 | Wikidata, KP, monitoring | N/A |
Print the plan before STEP 13, then map into the bundle tiers:
Single-shot runs (no MODE line) print this plan before STEP 13
serializes it; `MODE: judge` simply ends at STEP 12. Tier mapping:
G1–G4/G6 → AUTO, G5 → GATED, G7 → USER ACTIONS.
**Apply-vs-report is the DISPATCHER's call, not yours.** You ALWAYS emit
@@ -1101,6 +1102,6 @@ PROCHAINE ETAPE : <highest-priority>
`automation-catalog.md`. No exceptions.
- **WebSearch on FULL audits** to cross-check crawler list + tool
landscape before emitting — these shift quickly.
- **Dispatcher verifies.** Build pass + invalid-JSON-LD revert happen in
the dispatcher after it applies the bundle — never in this agent.
- **Transparency.** Every automated change logged in §14.
- **Dispatcher verifies.** Build pass, invalid-JSON-LD revert and the
applied-change log (SEO.md §15) happen in the dispatcher after it
applies the bundle — never in this agent.
+10 -5
View File
@@ -424,7 +424,7 @@ Wrong — has date prefix:
### 6.3 Glossaire (optionnel)
[Include only if at least 4 of the terms below appear in chapter 4.
[Include only if at least 4 of the terms below appear in chapter 6.
Format: term — one-line plain-language definition. Sort alphabetically.
This is the ONLY place internal tooling names may be mentioned by
their internal label, and only when explaining what they correspond
@@ -465,7 +465,7 @@ des audits de santé. Pour toute question, contactez [contact].*
1. Address the client directly ("votre site", "vous pouvez").
2. Chapters 1–3: replace every tech term with a user-facing equivalent.
3. No abbreviations the client wouldn't use (HTTPS yes, CSP no — unless
in chapter 4 with definition).
in chapter 6 with definition).
4. Concrete numbers > adjectives.
5. Short paragraphs. Bullet lists for things you can count.
6. **Score deltas explained in plain words**. Never just dump numbers.
@@ -535,9 +535,9 @@ The chapter must include:
8. **Outils gratuits pour vérifier votre présence**.
Cross-link this chapter from §4 (owner responsibilities — "Ce qui vous
Cross-link this chapter from §5 (owner responsibilities — "Ce qui vous
reste à faire"). Items in this §7 annex that are recurring belong in
§4's cadence checklist (Mensuel / Trimestriel / Annuel).
§5's cadence checklist (Mensuel / Trimestriel / Annuel).
---
@@ -624,7 +624,9 @@ checkbox:
(`LANG=en`: "Items already checked have been validated.")
### Verification
### Verification (deferred — run right AFTER STEP 15 writes `$OUTPUT_MD`;
the pre-checks themselves are applied to the in-memory body here, the
file does not exist yet)
```bash
# At least one pre-check expected for any project with real history.
@@ -687,6 +689,9 @@ awk '/^## 1\./{flag=1} /^## 6\./{flag=0} flag' "$OUTPUT" \
**Anchor-resolution gate** (clickable section refs work).
```bash
# ORDER: run this gate in STEP 16, immediately AFTER the HTML render —
# $OUTPUT_HTML does not exist yet at STEP 15. A broken anchor found here
# loops back to fix the markdown ref, then re-render.
grep -oE '\]\(#[a-z0-9-]+\)' "$OUTPUT_MD" | tr -d ']()#' | sort -u > /tmp/refs.txt
grep -oE 'id="[^"]+"' "$OUTPUT_HTML" | sed 's/id="//;s/"//' | sort -u > /tmp/ids.txt
comm -23 /tmp/refs.txt /tmp/ids.txt
+7 -2
View File
@@ -43,6 +43,10 @@ the edit applied + self-verified, not the report grammar).
BLOCKED`, report why (the orchestrator escalates to `/bugfix`), never
expand scope yourself. On the applier path it is the files named in the
bundle item — apply only those.
- An open user-visible choice the contract does not settle (placement,
wording, behavior) → `STATUS BLOCKED` with `CLASS: visible | public-name |
scope` in NOTES, BEFORE editing anything. The orchestrator asks the user
and re-dispatches once.
- If tests exist for the affected code, run them. Detection cascade:
```bash
# JS/TS
@@ -75,8 +79,9 @@ the edit applied + self-verified, not the report grammar).
```
HOTFIX-EXEC REPORT
STATUS : DONE | BLOCKED
FILE(S) : <changed files>
FILE(S) : <changed files — suffix files you CREATED with " (new)">
FIX : <one-line description>
SMOKE : <test/build result, verbatim line>
NOTES : <BLOCKED: the blocker; DONE: none>
NOTES : <BLOCKED: the blocker, + CLASS: visible | public-name | scope when
you halted at an open choice before editing; DONE: none>
```
+20
View File
@@ -14,6 +14,17 @@ Gather context. Produce complete PROJECT BRIEF as single source of truth.
- If the initial prompt already provides name + purpose + stack + features + architecture → skip questions and generate the BRIEF directly.
- Otherwise ask only what's genuinely missing, in a single structured block.
- After answers: produce BRIEF. One follow-up allowed if answer is ambiguous.
- Hard budget: 2 question rounds total (initial block + one follow-up) for gaps. The BRIEF ships after round 2 — gaps become OPEN DECISIONS. Sole exception: a VISIBLE, PUBLIC NAME or SCOPE choice (a user-facing placement or wording, a public command/flag/endpoint name, whether X is in scope) still open after round 2 gets ONE more targeted question; it never ships as `(assumed)`.
## FAILURE MODES
| Trigger | First response | If still unresolved |
|---|---|---|
| Answer vague/ambiguous | One targeted follow-up on that item only | Gap: record it in OPEN DECISIONS with the safest reading, marked `(assumed)` — never invent a confident value. Visible / public-name / scope item: one more targeted question instead, never `(assumed)` |
| "I don't know / you decide" | Propose ONE concrete default + why, ask yes/no | Take the default, mark `(assumed)`, list in OPEN DECISIONS |
| Contradictory answers (e.g. embedded runtime + managed cloud DB) | Name the contradiction, ask which side wins | Put BOTH options in OPEN DECISIONS; do not silently pick one |
| Partial answer to the block | Re-ask ONLY the missing items in the follow-up round | Missing fields → `none stated` + OPEN DECISIONS entry |
| Feature list balloons (>10) | Keep the 10 the user ranks first as V1 | Overflow goes to OUT OF SCOPE with a `(deferred by budget)` tag |
## QUESTIONS (skip answered ones)
@@ -60,3 +71,12 @@ OPEN DECISIONS: <list or none>
```
Stop after BRIEF. Orchestrator handles next step.
## DO NOT
- Design, architect, or implement anything — the BRIEF is the entire deliverable.
- Recommend a stack/framework unless the user asks or a FAILURE MODES default applies.
- Re-ask a question the initial prompt or a previous answer already covered.
- Exceed the 2-round budget for gaps; the only extra question is the single targeted one a visible / public-name / scope item earns.
- Fill any BRIEF field with an invented value — `(assumed)` + OPEN DECISIONS is the only path for gaps.
- Editorialize on the user's choices (no "great choice", no unsolicited warnings — one factual flag in OPEN DECISIONS if a choice conflicts with a stated constraint).
+22 -14
View File
@@ -12,33 +12,40 @@ Generate the baseline claude-config files in a project directory. No interview,
---
## INPUTS REQUIRED (passed by orchestrator)
## INPUTS (passed by orchestrator)
1. `PROJECT_ROOT` — absolute path where files should be written
2. `BRIEF` — dict with keys filled by orchestrator STEP 1-3:
2. `BRIEF` — dict. Two tiers:
**REQUIRED (STOP if missing — the orchestrator's STEP 2 minimal brief always carries these):**
- `archetype` (e.g., "nextjs-app-router", "wordpress", "dotfiles-meta")
- `archetype_category` (cms | static | framework | api | cli | library | mobile | meta)
- `project_name`
- `stack` (language/framework/versions)
- `purpose` (1-3 sentences)
- `build_cmd`, `test_cmd`, `lint_cmd` (or "N/A")
- `folder_tree` (max 2 levels)
- `architecture_notes`
- `conventions`
- `exceptions_to_global_rules`
- `key_deps` (list with one-line purpose each)
- `workflow_notes`
- `is_monorepo` (bool) + `packages` list if true
- `monorepo_mode` ("A" | "B:<package>" | "C") — only if is_monorepo
If any key is missing, PRINT what's missing and STOP. Do NOT invent values.
**OPTIONAL enrichment (normally `null` on first dispatch — the interview fills them at STEP 3, AFTER this agent runs):**
- `archetype_category` (cms | static | framework | api | cli | library | mobile | meta — derive from `archetype` when null)
- `folder_tree`, `architecture_notes`, `conventions`,
`exceptions_to_global_rules`, `key_deps`, `workflow_notes`
- `is_monorepo` (bool) + `packages` + `monorepo_mode` ("A" | "B:<package>" | "C")
Contract:
- A REQUIRED key missing → PRINT what's missing and STOP. Do NOT invent values.
- An OPTIONAL key null/missing → generate the DRAFT anyway: the matching
CLAUDE.md section gets the placeholder `<!-- TODO(/onboard STEP 3): <key> -->`,
never an invented value. List every placeholder in OUTPUT.
- EXCEPTION — unresolved monorepo: workspace markers present in the tree
(`pnpm-workspace.yaml`, `workspaces` in package.json, `apps/`+`packages/`)
but `monorepo_mode` null → STOP. Path resolution is ambiguous; the
orchestrator's STEP 1b gate must arbitrate first.
---
## PHASE 1 — GENERATE CLAUDE.md
Read `~/.claude/templates/project-CLAUDE.md` as base.
Fill sections from BRIEF. Preserve global CLAUDE.md compatibility (this file extends, doesn't override silently).
Fill sections from BRIEF; null enrichment keys become their `<!-- TODO(/onboard STEP 3): ... -->` placeholder. Preserve global CLAUDE.md compatibility (this file extends, doesn't override silently).
Write to `${PROJECT_ROOT}/CLAUDE.md`.
@@ -149,6 +156,7 @@ FILES WRITTEN:
✅ .claude/memory/evals.md (created | unchanged)
✅ .claude/audits/ (created | unchanged)
[✅ ROADMAP.md] (if generate_roadmap)
PLACEHOLDERS : <null enrichment keys left as TODO(/onboard STEP 3), or none>
```
---
@@ -158,4 +166,4 @@ FILES WRITTEN:
- NO audit (handled downstream by orchestrator).
- NO destructive writes: never overwrite CLAUDE.md if it exists without asking (print path + STOP, let orchestrator decide).
- Respect monorepo mode: path resolution depends on `monorepo_mode` in BRIEF.
- If any BRIEF key is missing, STOP and report — do not guess.
- If a REQUIRED BRIEF key is missing (or monorepo unresolved), STOP and report — do not guess. Null OPTIONAL keys are normal on first dispatch: placeholder, don't stop.
+8 -4
View File
@@ -63,7 +63,7 @@ Ground EVERY finding in the plan text (quote the section) or the real code
## OUTPUT (exact format — machine-parsed by the orchestrator)
```
CHALLENGE — LENS: <correctness|robustness|simplicity> — VERDICT: SOLID | CONCERNS(n) | FATAL(n)
CHALLENGE — LENS: <correctness|robustness|simplicity> — VERDICT: SOLID | CONCERNS(n) | FATAL(n) | ERROR(<reason>)
PLAN: <path>
FINDINGS:
1. [BLOCKER] <claim> — WHY: <why it fails — plan § or file:line> — FIX: <one line>
@@ -79,13 +79,17 @@ PROOF: read <n> files, inspected <what>, checked plan §<…>
- Report-only. Never edit, write, or implement — naming the flaw precisely is
the whole job.
- No invention. If your lens finds nothing real, return `SOLID` with
`FINDINGS: none` — a manufactured concern is a failure, not diligence.
- No invention — ungrounded is noise. Silently dropping a grounded doubt is
equally a failure: file it as `[MINOR]` with the uncertainty stated in
`WHY:`. Nothing real at all → `SOLID` with `FINDINGS: none`.
- `PROOF` is MANDATORY. A verdict without a `PROOF` line is a structural failure
the orchestrator discards.
- Stay in your lens. A finding outside it belongs to another challenger.
- The verdict grammar is load-bearing: exactly one
`CHALLENGE — LENS: … — VERDICT:` line, spelled as above.
`CHALLENGE — LENS: … — VERDICT:` line, spelled as above. `ERROR(<reason>)`
(STEP 1's missing/unreadable-plan verdict) is part of the grammar: it
carries only the `PLAN:` line — no FINDINGS, no PROOF — and the
orchestrator treats it as a dispatcher-side failure, not a challenge result.
## ORCHESTRATOR PROTOCOL (consumer contract — wiring reference)
+16 -4
View File
@@ -25,6 +25,18 @@ field. PROBE REPORT missing or a field absent → emit
`PLUGIN CHECK — VERDICT: ERROR(probe report missing/invalid: <what>)` and
STOP. Fail closed: no recommendations over invented detection.
`FRAMEWORK-DEPS` carries exact `"dep": "version"` pairs (or
`framework-deps-none`). Derive signal classes from those names + versions:
`frontend` = react/react-dom/vue/nuxt/svelte/astro/next present;
`fast-libs` = next, react ≥18 (version prefix), prisma/@prisma/client,
supabase/@supabase/supabase-js, drizzle-orm, expo. Never re-scan the
manifest to make this split.
`REQUEST` MAY carry `PLAN: Max|Pro|Free` from the dispatcher. Echo it in
the output. Absent → output `PLAN: unknown (not provided)` and SKIP the
plan-budget WARN (absolute COST ESTIMATE still reported). Never assume a
plan.
---
## PHASE 2 — ANALYZE
@@ -86,7 +98,7 @@ ACTIVE: [plugin — status, one line each]
PROFILE: [active skill profile — name + match%, or "custom"]
SIGNALS: [detected signals]
COMPLEXITY: <score>% — <simple|moderate|complex|enterprise>
PLAN: <Max|Pro|Free> (budget: ~<N>t passive tokens)
PLAN: <Max|Pro|Free (echoed from REQUEST) | unknown (not provided)> (budget: ~<N>t | n/a)
COST ESTIMATE: ~Xt passive tokens (all active plugins combined)
RECOMMENDATIONS:
@@ -280,8 +292,8 @@ activate a curated subset of skills + plugins + MCPs and disable the rest of
gstack + managed plugins — sessions stay focused and passive token cost drops.
`profile set <name>` actually toggles plugins (`claude plugin enable|disable`)
and MCPs (delegates to `lib/toggle-external.sh` for `magic`) — not just
advisory. Always-on plugins (`security-guidance`, `superpowers`)
and external skill packs (delegates to `lib/toggle-external.sh`) — not just
advisory. No MCP server is auto-toggled today. Always-on plugins (`security-guidance`, `superpowers`)
are protected. Managed plugins that `set` may toggle:
`ui-ux-pro-max@ui-ux-pro-max-skill`, `plugin-dev@claude-code-plugins`,
`pr-review-toolkit@claude-code-plugins`. Other plugins are never auto-toggled.
@@ -315,7 +327,7 @@ or by applying a profile that lists it (e.g. `apply web` to restore
- Active toggle plugins not needed for this task (dead passive cost)
- Multi-session feature + `gsd` CLI not installed → `npm install -g gsd-pi`
- Total passive cost > 50% of plan budget (Pro: ~5500t, Max: ~10000t, Free: ~2500t)
- Total passive cost > 50% of plan budget (Pro: ~5500t, Max: ~10000t, Free: ~2500t) — only when PLAN was provided; PLAN unknown → skip this WARN
- **Next.js/React 18+/Prisma/Supabase detected + context7 not configured**
→ Risk: Claude may generate code using outdated APIs (App Router changes frequently)
→ Fix: `npm install -g ctx7 && ctx7 setup --claude`
+3 -2
View File
@@ -34,7 +34,8 @@ command -v rtk &>/dev/null && rtk --version 2>/dev/null | head -1 || echo "rtk-n
# Project signals (run from project root)
ls package.json pyproject.toml Cargo.toml go.mod 2>/dev/null | head -5
grep -rl "next\|react\|vue\|prisma\|supabase" package.json 2>/dev/null | head -3 || true
# Exact-key dep match with versions ("react": won't match "preact":)
grep -ohE '"(next|react|react-dom|vue|nuxt|svelte|astro|prisma|@prisma/client|@supabase/supabase-js|supabase|drizzle-orm|expo)"[[:space:]]*:[[:space:]]*"[^"]*"' package.json 2>/dev/null || echo "framework-deps-none"
find . -name "*.tsx" -o -name "*.jsx" 2>/dev/null | head -3 | wc -l
find . -name "docker-compose*" -o -name "Dockerfile" 2>/dev/null | head -3 | wc -l
@@ -72,7 +73,7 @@ EXTERNAL : <toggle-external list output>
PROFILE : <profile current output>
CLIS : ctx7=<v|absent> gsd=<v|absent> rtk=<v|absent>
MANIFESTS : <files found>
FRAMEWORK-DEPS: <grep hits in package.json>
FRAMEWORK-DEPS: <exact "dep": "version" pairs, or framework-deps-none>
TSX-JSX-COUNT : <n>
DOCKER-COUNT : <n>
ANIM : eligibility=<status|package|reason> installed=<lib|no>
+12 -2
View File
@@ -19,9 +19,18 @@ Improve code without ever changing its external behavior.
1. Analyze the target — list ALL violations
2. Produce the report BEFORE touching anything
3. Check that tests exist (if not — report before modifying)
3. Check that tests exist covering the target.
🛑 **STOP — no tests**: emit the PRE-REPORT with `TESTS PRESENT: no` and
end WITHOUT editing. Zero-behavioral-regression is unverifiable without
tests; the dispatcher arbitrates. Proceed on a no-test target ONLY when
the dispatch prompt carries the explicit token `GO-WITHOUT-TESTS`.
(Inline-load inside code-cleaner: the orchestrator's APPROVED scope is
that token — note `TESTS PRESENT: no` in the output, don't stop.)
4. Refactor function by function
5. Verify tests pass after each modification
5. Run the tests after each modification.
Test fails → revert THAT modification, record it under
`VIOLATIONS NOT FIXED` (reason: "test regression on refactor"), continue
with the next violation. Never leave the suite red between steps.
---
@@ -60,6 +69,7 @@ TESTS PRESENT: yes / no
- Zero behavioral regression
- Existing tests must pass
- No tests on the target → PRE-REPORT + STOP (unless dispatched with `GO-WITHOUT-TESTS`)
- Do not modify business logic under the guise of refactoring
- Do not refactor unrelated parts
+3 -1
View File
@@ -147,7 +147,9 @@ In audit mode, ALSO write this same block (plus per-finding detail) to
## ORCHESTRATOR PROTOCOL (consumer contract — wiring reference)
- The security gate runs AFTER the request-conformity verdict is CONFORME
(verifier), never before.
(verifier), never before — EXCEPT under /hotfix, which by design runs no
verifier: there the gate fires directly on the smoke-passed diff (its
one-attempt model reverts on BLOCK instead of looping).
- Dispatch a FRESH auditor each iteration — no context reuse. Input = mode +
scope + (report) + (context), nothing else.
- Parse the `SECURITY — VERDICT:` line:
+54 -79
View File
@@ -58,8 +58,8 @@ STEP 1-2 business/tech context is consumed by all later steps).
## STEP 0 — AUDIT DEPTH
**First action.** If a parent skill (`/seo` dispatcher) passed depth
in $ARGUMENTS, use it. Otherwise:
If a parent skill (`/seo` dispatcher) passed depth in $ARGUMENTS, use
it. Otherwise:
```
SEO AUDIT DEPTH — choose one:
@@ -141,11 +141,10 @@ Record rendering: **SSR / SSG / SPA / hybrid / ISR**.
### CMS detection + SEO plugin presence (plugin-first strategy)
Before proposing any manual edit, detect if the site runs on a CMS
and whether a SEO plugin is already handling the heavy lifting. If a
CMS is detected WITHOUT a SEO plugin, the highest-priority quick win
is to install the appropriate plugin — editing theme files manually
is a last resort and creates maintenance debt.
Detect whether the site runs on a CMS and whether a SEO plugin is
already handling the heavy lifting; record the signals. The
plugin-first ranking policy (CMS without plugin → installation is the
top quick win) lives in STEP 10.
```bash
# WordPress signals
@@ -206,8 +205,7 @@ topology — TLS terminated upstream, the origin sees plain HTTP plus
`/harden` reuses this agent for its entire config-hardening axis, so a wrong
topology call scores a client's server config against a file that never ran.
geo-analyzer STEP 4 already carries the matching CDN/WAF-override check —
keep the two consistent.
(The same CDN/WAF-override check lives in geo-analyzer STEP 4.)
```bash
# Server / hosting
@@ -505,7 +503,7 @@ Fetch rendered HTML. Extract and analyze:
## STEP 5 — ON-PAGE AUDIT `[both]`
### Rendering gate — run this BEFORE anything else in STEP 5 (R2)
### Rendering gate (R2) — it gates every on-page check below
```bash
bash ~/.claude/lib/seo-data/fetch.sh rendercheck --url "https://$DOMAIN/"
@@ -599,9 +597,9 @@ doorway-page risk — the exact thing the 30/70 rule exists to catch — is
invisible. Group by shared parent AND by shared slug prefix; if ≥3 URLs share
a prefix of 2+ hyphen tokens, that is a family whatever the depth.
Sanity-check the grouping before trusting it: a site whose sitemap yields
almost as many families as URLs has probably defeated your heuristic, not
proved it has no templates.
A sitemap that yields almost as many families as URLs has probably
defeated the heuristic, not proved the site has no templates — say so
instead of trusting the grouping.
**Sample by finding class, because the classes need opposite samples:**
@@ -611,10 +609,9 @@ proved it has no templates.
| **Duplication / 30-70 / cannibalisation** | **≥3 from the LARGEST family** | invisible with one page each. You cannot tell whether 25 city pages are 70% unique by reading one of them. |
| Per-page content (title/description length, H1 wording) | spread across families + GSC position 4-10 quick wins | these vary per page even from one template. |
"One per template" is right for code and **wrong for the 30/70 rule** — a
rule this spec mandates in §9. Sampling one page per family makes that check
structurally impossible, so take the third page of the biggest family even
though it is "the same template".
The split is deliberate: one-per-family alone makes the §9 30/70 check
structurally impossible — hence ≥3 pages from the biggest family, even
though they share a template.
An un-sampled family is an un-audited family. Name the ones you skipped.
@@ -658,16 +655,11 @@ mapfile -t FEXCL < <(bash ~/.claude/lib/source-scope.sh findargs)
find . "${FEXCL[@]}" -type f \( -iname "*.jpg" -o -iname "*.jpeg" -o -iname "*.png" -o -iname "*.gif" \) -printf "%s %p\n" 2>/dev/null | sort -rn | head -20
```
**Why the guard, and why `find` specifically (C1a).** `grep` and `find`
disagree about this repo and you use both. Claude Code routes `grep` through
ugrep with `--ignore-files`, so it honours `.gitignore` and never descends
into a gitignored `dist/`. `find` honours nothing. Measured on a real Astro
repo: this command returned **92 images, 45 of them under `dist/`** — every
asset twice, source and generated copy, byte-identical. So "top 20 by size"
was ~10 real images dressed as 20, and a batch-C item
(`cwebp -q 80 <img> -o <img>.webp`) could target `dist/og-image.png`, whose
`.webp` the dispatcher's own `npm run build` then erases. The fix lands,
verification passes, nothing survives.
**Why the guard, and why `find` specifically (C1a).** Claude Code routes
`grep` through ugrep with `--ignore-files` (honours `.gitignore`); `find`
honours nothing. Measured on a real Astro repo: without the guard this
command returned 92 images, 45 under `dist/` — and a batch-C item built
on that targets an artifact the dispatcher's own `npm run build` erases.
`FEXCL` MUST be consumed as a quoted array. `find . $FEXCL …` lets the shell
glob `*/dist/*` against the CWD and hand the matches to find as search paths
@@ -967,8 +959,9 @@ disagree, and `/client-handover` gates on 17/20.
**N/A is not a zero** and the engine will not let it behave like one.
- `status: "error"` → malformed findings. Fix them; never fall back to
eyeballing a number.
- Run it twice on the same file before publishing. If the output moved, your
findings moved, and that is the thing to explain.
- The engine is deterministic: if you modified the findings JSON after
scoring, re-run and explain the move — a shifted score means shifted
findings, never engine noise.
**Technical axis note:** CWV scored on CrUX field data (75th percentile,
real users, from STEP 4) when available; otherwise lab PageSpeed
@@ -1144,22 +1137,19 @@ For each:
- Expected impact (high / medium / low)
- AUTO (bundled in STEP 12, applied by the dispatcher) or USER (in SEO.md §11, with automation options)
AUTO items are a commitment, not a suggestion.
**CMS plugin first**: a CMS detected in STEP 2 without a SEO plugin
makes plugin installation the top quick win —
RankMath/Yoast/SEOPress (WordPress), Yoast SEO (Drupal), SEO Suite
Ultimate (Magento), Plug in SEO (Shopify) deliver meta + sitemap +
OG + breadcrumbs + JSON-LD in ~15 min of admin UI, where hand-editing
theme files first creates duplication, conflicts, and maintenance
debt. See `~/.claude/agents/resources/automation-catalog.md` CMS
plugins section for the exact install path per CMS.
**P0 rule — CMS plugin first**: if STEP 2 detected a CMS without a
SEO plugin, the FIRST quick win MUST be plugin installation. Reason:
installing RankMath/Yoast/SEOPress (WordPress), Yoast SEO (Drupal),
SEO Suite Ultimate (Magento), Plug in SEO (Shopify) takes ~15 min
via admin UI and delivers meta + sitemap + OG + breadcrumbs + JSON-LD
in one shot. Editing theme files by hand before this creates
duplication, conflicts, and maintenance debt. See
`~/.claude/agents/resources/automation-catalog.md` CMS plugins
section for the exact install path per CMS.
**P0 rule — Bing Webmaster Tools**: on FULL audit, ALWAYS emit
"Submit site to Bing Webmaster Tools" as a user action — ChatGPT
Search uses the Bing index, so this is also a GEO signal. See
automation-catalog.md for IndexNow + Bing.
**Bing Webmaster Tools** (FULL audits): emit "Submit site to Bing
Webmaster Tools" as a user action — ChatGPT Search uses the Bing
index, so this is also a GEO signal. See automation-catalog.md for
IndexNow + Bing.
### Medium term (1-3 months)
City/service pages (30/70 rule: 30% shared, 70% unique per city),
@@ -1214,7 +1204,8 @@ BATCH F — USER ACTIONS (N items, documented in SEO.md §11 with automation cat
...
```
Do not proceed to STEP 12 until this plan is printed.
Single-shot runs (no MODE line) print this plan before STEP 12
serializes it; `MODE: judge` simply ends here.
---
@@ -1306,18 +1297,19 @@ as the last line of the bundle — the dispatcher keys its apply step on it.
Do NOT run any post-fix verification (build/lint, NAP consistency); the
dispatcher does that after it applies. Your job ends at the sentinel.
### Bundle completeness checklist (did every finding reach the bundle?)
### Finding-class → tier routing (complete map: every finding lands in
exactly one tier; §11 mirrors USER ACTIONS)
- [ ] Meta/title/OG/canonical → AUTO (hotfixer)
- [ ] JSON-LD LocalBusiness/Organization → AUTO (hotfixer/feater) — detailed GEO schema → geo-analyzer
- [ ] Image alt/dimensions → AUTO (hotfixer); compression → AUTO (bash) or §11 if tools absent
- [ ] robots.txt / sitemap.xml → AUTO (hotfixer) — AI-bot directives → geo-analyzer
- [ ] .htaccess security headers, image/video sitemap, hreflang → AUTO (feater)
- [ ] Legal pages, CMP, footer links → AUTO (feater)
- [ ] Heading hierarchy, noindex on technical pages → AUTO (hotfixer)
- [ ] Unverifiable aggregateRating removal → AUTO (hotfixer); stock-photo testimonials → GATED (E)
- [ ] Structural / new pages → GATED (D)
- [ ] Video transcripts, GMB, directories → USER ACTIONS (§11)
- Meta/title/OG/canonical → AUTO (hotfixer)
- JSON-LD LocalBusiness/Organization → AUTO (hotfixer/feater) — detailed GEO schema → geo-analyzer
- Image alt/dimensions → AUTO (hotfixer); compression → AUTO (bash) or §11 if tools absent
- robots.txt / sitemap.xml → AUTO (hotfixer) — AI-bot directives → geo-analyzer
- .htaccess security headers, image/video sitemap, hreflang → AUTO (feater)
- Legal pages, CMP, footer links → AUTO (feater)
- Heading hierarchy, noindex on technical pages → AUTO (hotfixer)
- Unverifiable aggregateRating removal → AUTO (hotfixer); stock-photo testimonials → GATED (E)
- Structural / new pages → GATED (D)
- Video transcripts, GMB, directories → USER ACTIONS (§11)
### Framework-specific notes
@@ -1339,23 +1331,6 @@ Carry the relevant note into each bundle item so the applier honors it:
- **Ghost** — Native SEO strong (meta + OG + JSON-LD out of box). Usually no plugin needed; handle gaps via `default.hbs` edits.
- **Wix / Squarespace / Webflow (hosted CMS)** — No theme file access. ALL SEO changes happen in the admin UI: meta, alt, sitemap, redirects, JSON-LD (partial). Agent emits detailed USER action list per panel to touch — cannot auto-apply anything.
### Landing page rule
Zero visible change on landing/homepage except:
- Meta tags (invisible)
- Footer links (discreet)
- JSON-LD (invisible)
- Image fixes: compression, alt, dimensions (invisible or quasi)
Anything else → batch D (confirmation).
### Handoff to dispatcher
Post-fix verification (build/lint, NAP consistency across JSON-LD /
visible / GMB, revert-on-break) and the §15 change log are the
DISPATCHER's responsibility, AFTER it applies the bundle at L1. You
emitted the bundle terminated by the sentinel — stop here.
---
## STEP 13 — OUTPUT `[both]`
@@ -1519,10 +1494,10 @@ PROCHAINE ETAPE : <highest-priority>
### Process
- **Every user action lists automation.** Mandatory from
`~/.claude/agents/resources/automation-catalog.md`.
- **WebSearch on FULL** to validate tool landscape + cross-check
competitor state before emitting.
- **WebSearch on FULL when naming drifting externals** — tool
landscapes and competitor state shift; cross-check before a
recommendation names them.
- **Iterative SEO.md.** Preserve Historique section.
- **Transparency.** Every automated change logged with file, change,
reason.
- **Dispatcher verifies.** Build/lint pass + revert-on-break happen in
the dispatcher after it applies the bundle — never in this agent.
- **Dispatcher verifies.** Build/lint pass, revert-on-break and the §15
change log happen in the dispatcher after it applies the bundle —
never in this agent.
+9 -5
View File
@@ -1,6 +1,6 @@
---
name: status-reporter
description: Read-only project-status engine — dispatched by /status. Collects plugins, token budget, git state, build/tests, GSD milestone into one snapshot.
description: Read-only project-status engine — dispatched by /status. Collects plugin roster + passive-cost estimate (doctor.sh constants), git state, build/tests, GSD milestone into one snapshot.
tools: Read, Bash, Glob, Grep
model: haiku
---
@@ -23,8 +23,12 @@ cat ~/.claude/lib/../version.txt 2>/dev/null || echo "unknown" # lib symlink re
command -v rtk &>/dev/null && echo "rtk: installed" || echo "rtk: missing"
command -v gsd &>/dev/null && gsd --version 2>/dev/null | head -1 || echo "gsd: not installed"
# Token estimate (passive)
# (approximate from known plugin costs)
# Passive token cost — source of truth: doctor.sh's constants block
# (PLUGIN_TOKENS + <n> per detect_* line). Read it, sum ONLY the plugins
# found active above. Never invent a number outside these constants.
grep -E 'PLUGIN_TOKENS \+ [0-9]+' "$(readlink -f "$HOME/.claude/lib")/../doctor.sh" 2>/dev/null
# grep empty (doctor.sh missing/moved) → report the plugin count only and
# defer cost to /plugin-check.
```
Check `~/.claude/plugins/cache` for active marketplace plugins.
@@ -134,7 +138,7 @@ PROJECT STATUS
CONFIG
Version : v<N>
Plugins ON: <list> (~<X>t passive)
Plugins ON: <list> (~<X>t passive — doctor.sh constants; full audit → /plugin-check)
GSD v2 : installed / not installed
PROJECT
@@ -174,7 +178,7 @@ The report is best-effort: a single failing data source must not abort the whole
|---|---|
| Permission denied on `git` (sandbox/CI without `.git` access) | Mark `Branch: N/A (permission denied)`, `Uncommitted: N/A`, `RECENT COMMITS: N/A`. Continue to PROJECT/GSD sections. |
| Permission denied on `~/.claude/plugins/cache` or `~/.claude.json` | Mark `Plugins ON: unknown (cannot read cache)`. Continue. |
| `.gsd/ROADMAP.md` exists but unparseable (malformed checkboxes, encoding issue) | Mark `Progress: N/A (ROADMAP.md unreadable)`, do NOT abort the section — still print `Status: initialized` and `Milestone: N/A`. |
| gsd CLI snapshot fails or `.gsd/` state unreadable (`gsd.db`, `STATE.md`, per-milestone `<ID>-ROADMAP.md` — post-ADR-013 layout) | Mark `Progress: N/A (gsd state unreadable)`, do NOT abort the section — still print `Status: initialized` and `Milestone: N/A`. |
| `package.json` / `pyproject.toml` parse error | Mark `Tests: N/A (manifest parse error)`. Continue. |
| `python3` not available in PATH | Skip the python parsing fallbacks; rely on log files + bash-only checks. Mark Tests as `unknown` if no log found. |
| All sections fail | Print a minimal envelope with each section showing `N/A (data source unavailable)` and a one-line `DIAGNOSTIC: <which sources failed>` footer. Exit code 0 (status reporter never blocks). |
+40 -5
View File
@@ -48,6 +48,25 @@ Rules: read the diff AND enough surrounding code to judge behavior; run
criterion. Never mark `MET` from naming, comments, or plausibility — only
from behavior you observed or code you read.
### Criteria carrying an oracle (`CHECK:` / `EXPECT:` / `EVIDENCE:`)
`lib/gates.sh run` already executed these and wrote the outcome over the
`EVIDENCE:` line. Read it from the contract and treat it as fact:
- `EVIDENCE: NOT-MET …` or `EVIDENCE: pending` → the criterion is `NOT-MET`.
Reading the code NEVER overrides a red or unrun oracle. Cite the evidence
line as your evidence.
- `EVIDENCE: MET …` → the declared command passed. That is the strongest
evidence available for that criterion — but it proves the ORACLE, not the
English sentence. Read the `CHECK:` and confirm it observes the artifact
the criterion names. A vacuous oracle (`1. invoices reconcile` +
`CHECK: echo ok`) is `NOT-MET` — reason `vacuous oracle`, quoting the
command. That judgement is yours alone; no command can make it.
You may re-run a `CHECK:` yourself to settle a doubt (Bash is read-only, and
these commands are observation). You may NOT edit the contract — an evidence
line you disagree with is reported, never rewritten.
## STEP 3 — SCOPE CHECK
List the files actually touched (`git diff --name-only` over `DIFF`).
@@ -58,19 +77,30 @@ only enters the contract through a human micro-gate.
## STEP 4 — VERDICT
`CONFORME` ⇔ ALL criteria `MET` AND zero out-of-scope files.
Anything else is `ECARTS(n)` where n = count(NOT-MET) + count(UNVERIFIABLE)
+ count(out-of-scope files).
Read the contract's `ABANDON:` lines. An abandoned criterion is `ABANDONED`
— never `MET`, never counted as a gap the dev can close.
Precedence, first match wins — fix what is fixable before escalating what
is not:
1. `ERROR(<reason>)` — the contract is missing or unreadable.
2. `ECARTS(n)` — n = count(NOT-MET) + count(UNVERIFIABLE) + count(out-of-scope
files). Surface any abandonment in the same report.
3. `ABANDONED(n)` — zero gaps remain, but n abandonments stand. This is NOT
a pass and NOT a dev loop: it routes straight to the human gate.
4. `CONFORME` — ALL criteria `MET`, zero out-of-scope files, zero
abandonments.
## OUTPUT (exact format — machine-parsed by the orchestrator)
```
VERIFY — VERDICT: CONFORME | ECARTS(n) | ERROR(<reason>)
VERIFY — VERDICT: CONFORME | ECARTS(n) | ABANDONED(n) | ERROR(<reason>)
CONTRACT: <path>
CRITERIA:
1. <criterion> — MET — <evidence file:line | test ran → result>
1. <criterion> — MET — <EVIDENCE line | file:line | test ran → result>
2. <criterion> — NOT-MET — expected <…> / actual <…> — <file:line>
3. <criterion> — UNVERIFIABLE — <reason>
4. <criterion> — ABANDONED — <the reason recorded in the contract>
SCOPE: in-scope <n> files; out-of-scope: <list | none>
PROOF: read <n> files, ran <cmd → result | nothing>, checked <n>/<n> criteria
```
@@ -82,6 +112,8 @@ PROOF: read <n> files, ran <cmd → result | nothing>, checked <n>/<n> criteria
- `UNVERIFIABLE` ≠ `MET`. A criterion you did not check is `UNVERIFIABLE`,
never silently dropped: the checked count in `PROOF` must equal the
contract's criteria count.
- `ABANDONED` ≠ `MET`. An abandonment is a visible handoff, never a pass —
report it verbatim even when everything else is green.
- `PROOF` is MANDATORY. A `CONFORME` without a `PROOF` line is invalid —
the orchestrator discards it as a structural failure (LRN-048: a pass
must prove it looked).
@@ -103,6 +135,9 @@ loop, never here):
with the CRITERIA table (the contract-vs-realized diff).
- Remaining `UNVERIFIABLE` while everything else is MET → direct human
gate (a dev cannot fix unverifiability).
- `ABANDONED(n)` → direct human gate, never a dev loop. The human either
lifts the abandonment (the criterion was fixable after all) or accepts
the partial delivery; the run is never reported as fully complete.
- Structural failure (`ERROR(…)`, missing/duplicated VERDICT line,
unparsable output, agent crash, `CONFORME` without `PROOF`) → retry
ONCE with a fresh verifier; a 2nd structural failure → human
+67 -1
View File
@@ -18,8 +18,10 @@ REPO="$(cd "$(dirname "$0")" && pwd)"
VERSION=$(cat "$REPO/version.txt" 2>/dev/null || echo "unknown")
# Load shared detection library
# shellcheck source=lib/detect-plugins.sh
# shellcheck source=lib/detect-plugins.sh disable=SC1091
source "$REPO/lib/detect-plugins.sh"
# shellcheck source=lib/gstack-playwright.sh disable=SC1091
source "$REPO/lib/gstack-playwright.sh"
echo ""
echo "═══ claude-config doctor (v${VERSION}) ═══"
@@ -115,6 +117,13 @@ fi
echo ""
# ── Playwright browsers (read-only report; NOT nested under gstack — 2 of
# the 3 registered installs are gsd-pi, not gstack) ──
echo "── Playwright browsers ──"
gstack_browsers_report || true
echo ""
# ────────────────────────────────────────────────────────────
# 3. Prerequisites
# ────────────────────────────────────────────────────────────
@@ -206,6 +215,61 @@ echo ""
# ────────────────────────────────────────────────────────────
# 5. Permissions check
# ────────────────────────────────────────────────────────────
# Under defaultMode auto the classifier reads `autoMode`, so a block scoped
# to ONE project feeds every other project false facts, and a list without
# "$defaults" silently drops the built-in rules. Neither is visible from the
# deny count. Emits TAG|message lines for the caller to dispatch.
inspect_automode() {
REPO="$REPO" python3 - "$SETTINGS" <<'PY'
import json, os, re, sys
settings = json.load(open(sys.argv[1]))
mode = settings.get("permissions", {}).get("defaultMode")
block = settings.get("autoMode") or {}
if mode != "auto":
sys.exit(print("INFO|defaultMode is %s, autoMode not consulted" % mode))
if not block:
sys.exit(print("WARN|defaultMode is auto but no autoMode block set"))
sections = [k for k in ("allow", "soft_deny", "hard_deny", "environment")
if k in block]
bare = [k for k in sections if "$defaults" not in block[k]]
if bare:
print('WARN|autoMode.%s replaces the built-in entries (no "$defaults")'
% ", ".join(bare))
else:
print('PASS|autoMode: %s inherit "$defaults"' % ", ".join(sections))
repo, home = os.environ["REPO"], os.path.expanduser("~")
foreign = {q for entry in block.get("environment", [])
for q in re.findall(r"`(/[^`]+)`", entry)
if (p := q.rstrip("/")).startswith(home) and p != repo
and os.path.isdir(os.path.join(p, ".git"))}
if foreign:
print("WARN|autoMode.environment names another repo (%s); this file is "
"user-scope and reaches every project" % ", ".join(sorted(foreign)))
else:
print("PASS|autoMode.environment is not scoped to a foreign repo")
PY
}
check_automode() {
local out tag msg
if ! out=$(inspect_automode 2>/dev/null); then
warn "Could not inspect the autoMode block"
return
fi
while IFS='|' read -r tag msg; do
case "$tag" in
PASS) pass "$msg" ;;
WARN) warn "$msg" ;;
INFO) info "$msg" ;;
esac
done <<< "$out"
}
echo "── Permissions ──"
SETTINGS="$HOME/.claude/settings.json"
@@ -242,6 +306,8 @@ print(len(json.load(sys.stdin).get('permissions',{}).get('deny',[])))
warn "Deny rules: $DENY_COUNT (committed: $EXPECTED_DENY) — live settings diverge from last commit"
fi
fi
check_automode
else
fail "$HOME/.claude/settings.json not found"
fi
+5 -1
View File
@@ -44,7 +44,11 @@ lc="$(printf '%s' "$prompt" | tr '[:upper:]' '[:lower:]')"
# "design system", "redesign", "front-?end design". dashboard -> \bdashboard\b
# so a filename like ecc_dashboard.py no longer matches while "admin dashboard"
# still does. animation kept (rarely non-UI).
pattern='redesign|refonte|refont|ui/ux|ux/ui|\bui\b|\bux\b|ui kit|design system|design-system|front-?end design|\bnavbar\b|\bsidebar\b|\bmodal\b|\bbouton\b|\bbutton\b|formulaire|\bhero\b|\bheader\b|\bfooter\b|dropdown|tooltip|\bbadge\b|\bchart\b|graphique|accordion|carousel|\bslider\b|landing|\bdashboard\b|homepage|home page|\baccueil\b|\bécran\b|\becran\b|portfolio|maquette|mockup|wireframe|prototype|\bjoli\b|\bjolie\b|\bbeau\b|\bbelle\b|esth[eé]tique|aesthetic|\bvisuel\b|\bvisual\b|embellir|fignol|peaufin|polish|styliser|styling|stylesheet|\bskin\b|charte graphique|\bbrand\b|branding|\blogo\b|favicon|ic[oô]ne|\bicon\b|\bcss\b|tailwind|shadcn|couleur|gradient|d[eé]grad[eé]|\bombre\b|spacing|espacement|\bmarge\b|\bpadding\b|\bmargin\b|\bradius\b|arrondi|\bhover\b|dark mode|light mode|typograph|\bfont\b|\bfonts\b|font pairing|\bpolice\b|animation|\bmotion\b|micro-interaction|keyframe|glassmorph|neumorph|claymorph|skeuomorph|brutalis|bento|minimalis|responsive|figma'
# Tightened 2026-07-30 (3rd pass): dropped \bux\b — bare "ux" matched inside
# French prose ("changement ux vu…"; 2 logged FPs, both FR). \bui\b KEPT
# (zero logged FP, one logged true positive). NB: the log records only the
# FIRST match per fire (head -1), so per-token FP rates aren't derivable.
pattern='redesign|refonte|refont|ui/ux|ux/ui|\bui\b|ui kit|design system|design-system|front-?end design|\bnavbar\b|\bsidebar\b|\bmodal\b|\bbouton\b|\bbutton\b|formulaire|\bhero\b|\bheader\b|\bfooter\b|dropdown|tooltip|\bbadge\b|\bchart\b|graphique|accordion|carousel|\bslider\b|landing|\bdashboard\b|homepage|home page|\baccueil\b|\bécran\b|\becran\b|portfolio|maquette|mockup|wireframe|prototype|\bjoli\b|\bjolie\b|\bbeau\b|\bbelle\b|esth[eé]tique|aesthetic|\bvisuel\b|\bvisual\b|embellir|fignol|peaufin|polish|styliser|styling|stylesheet|\bskin\b|charte graphique|\bbrand\b|branding|\blogo\b|favicon|ic[oô]ne|\bicon\b|\bcss\b|tailwind|shadcn|couleur|gradient|d[eé]grad[eé]|\bombre\b|spacing|espacement|\bmarge\b|\bpadding\b|\bmargin\b|\bradius\b|arrondi|\bhover\b|dark mode|light mode|typograph|\bfont\b|\bfonts\b|font pairing|\bpolice\b|animation|\bmotion\b|micro-interaction|keyframe|glassmorph|neumorph|claymorph|skeuomorph|brutalis|bento|minimalis|responsive|figma'
if printf '%s' "$lc" | grep -Eq "$pattern"; then
# Counter: log the fire (time, matched token, excerpt) — best-effort, never blocks.
+67
View File
@@ -0,0 +1,67 @@
#!/usr/bin/env bash
# Notification + Stop hook — signal the user through the terminal when
# Claude needs input (permission, question, idle wait) or has finished
# responding. Each case gets its own readable label so the toast says
# which one fired.
#
# Runs on the remote (Linux); the only channel that crosses SSH into the
# VS Code client is the terminal stream. Hooks have no controlling TTY,
# so the sequence goes through the supported `terminalSequence` JSON
# output field and Claude Code writes it to the terminal:
# - BEL x2 (double beep) -> sound, needs VS Code setting
# accessibility.signals.terminalBell { "sound": "on" } AND a non-zero
# volume for Code in the Windows volume mixer (BLK-020).
# - OSC 777 notify -> Windows toast via the client-side extension
# "Terminal Notification" (wenbopan.vscode-terminal-osc-notifier).
# A terminal can be deaf to OSC while the bell still rings; test it
# before attaching a session to it (LRN-148).
# Both are invisible no-ops in terminals that ignore them.
set -u
payload=$(cat 2>/dev/null)
read_field() {
printf '%s' "$payload" | jq -r "$1 // empty" 2>/dev/null \
| tr -d '\000-\037' | cut -c1-160
}
# How many background tasks are still running as the hook fires.
background_count() {
count=$(printf '%s' "$payload" | jq -r '(.background_tasks // []) | length' 2>/dev/null)
case "$count" in ''|*[!0-9]*) echo 0 ;; *) echo "$count" ;; esac
}
event=$(read_field '.notification_type')
[ -n "$event" ] || event=$(read_field '.hook_event_name')
case "$event" in
# Turn end while a subagent still runs is not the real end: stay silent,
# the next turn end will signal once the work is actually done.
Stop) [ "$(background_count)" -eq 0 ] || exit 0
label="Finished responding" ;;
permission_prompt) label="Needs your permission" ;;
agent_needs_input) label="Asks you a question" ;;
idle_prompt) label="Waiting for you" ;;
elicitation_dialog|elicitation_url_dialog) label="Needs your input" ;;
# anything else (agent_completed, auth_success, quota_*) stays silent:
# signal only for turn end and moments needing the user.
*) exit 0 ;;
esac
detail=$(read_field '.message')
if [ -n "$detail" ]; then
# Claude Code's own wording often restates the label ("Claude needs your
# permission"). Append it only when it actually adds something.
short=$(printf '%s' "$detail" | tr '[:upper:]' '[:lower:]' | sed 's/^claude //')
case "$(printf '%s' "$label" | tr '[:upper:]' '[:lower:]')" in
*"$short"*) : ;;
*) label="${label}: ${detail}" ;;
esac
fi
bell=$(printf '\a')
esc=$(printf '\033')
seq="${bell}${bell}${esc}]777;notify;Claude Code;${label}${esc}\\"
jq -cn --arg seq "$seq" '{suppressOutput: true, terminalSequence: $seq}'
exit 0
+190 -94
View File
@@ -26,8 +26,10 @@ else
fi
# Load shared detection library
# shellcheck source=lib/detect-plugins.sh
# shellcheck source=lib/detect-plugins.sh disable=SC1091
source "$REPO/lib/detect-plugins.sh"
# shellcheck source=lib/gstack-playwright.sh disable=SC1091
source "$REPO/lib/gstack-playwright.sh"
# ── Guard hand-curated config against installer drift ────────
# graphify's installer (Step 7) rewrites CLAUDE.md + .claude/settings.json
@@ -291,35 +293,6 @@ fi
echo ""
# gstack pins Playwright (1.58.x) which only ships browser builds for
# ubuntu<=24.04. On a newer distro the browser install fails ("does not
# support chromium on ubuntuXX.04"). Bump gstack's Playwright to a version
# that supports this OS so ./setup builds the browse binary against it and
# installs a native browser. Fires only when the pinned version genuinely
# lacks support — idempotent across runs. Edits the submodule locally (goes
# dirty); a `git submodule update` resets it and the next install re-applies.
# See BLK-008 / LRN-040.
gstack_bump_playwright_if_unsupported() {
[ -d "$GSTACK_DIR" ] && [ -r /etc/os-release ] || return 0
local ostag pwlib
# shellcheck disable=SC1091
ostag="$(. /etc/os-release 2>/dev/null; [ "${ID:-}" = ubuntu ] && printf 'ubuntu%s' "${VERSION_ID:-}")"
[ -n "$ostag" ] || return 0 # only the known Ubuntu case
pwlib="$GSTACK_DIR/node_modules/playwright-core/lib"
# populate node_modules at the pinned version so we can read its support list
( cd "$GSTACK_DIR" && { bun install --frozen-lockfile >/dev/null 2>&1 || bun install >/dev/null 2>&1; } ) || return 0
if grep -rqs "$ostag" "$pwlib" 2>/dev/null; then
return 0 # pinned Playwright already supports this OS
fi
info "gstack's Playwright lacks $ostag support — bumping to latest (local submodule edit)..."
( cd "$GSTACK_DIR" && bun add playwright@latest >/dev/null 2>&1 )
if grep -rqs "$ostag" "$pwlib" 2>/dev/null; then
ok "gstack Playwright bumped — now supports $ostag (browse binary rebuilt by ./setup)"
else
warn "Playwright bump didn't add $ostag support — gstack browser may stay unavailable"
fi
}
# ============================================================
# STEP 2 — GSTACK SUBMODULE
# ============================================================
@@ -367,7 +340,8 @@ if [ -d "$GSTACK_DIR" ]; then
# BEFORE ./setup so its frozen-lockfile install picks up the new version and
# the browse binary is rebuilt against it (avoids the "does not support
# chromium" fail). Non-fatal if it can't — gstack is OFF by default.
gstack_bump_playwright_if_unsupported
# See BLK-008 / LRN-040 / BDR-029; logic lives in lib/gstack-playwright.sh.
gstack_bump_playwright_if_unsupported "$GSTACK_DIR"
info "Running GStack setup..."
_gstack_setup_ok=0
@@ -814,54 +788,114 @@ else
fi
echo ""
# ── Step 8d: Impeccable (design anti-pattern detector + skill) ──
# 45 deterministic detector rules (CLI `impeccable detect`, exit 0/2) +
# /impeccable skill (23 verbs). Machine-owned dist: the installer produces
# it, we stage it in a tmpdir then move it under skills-external/
# (gitignored, ctx7 pattern) — never let the installer write through the
# ~/.claude/skills symlink into the tracked repo dir.
echo "── Step 8d: Impeccable — design anti-pattern detector ────"
# ── Step 8d: Impeccable (design detector + skill + subagents) ──
# 45 deterministic detector rules (`impeccable detect`, exit 0/2), the
# /impeccable skill (23 verbs) and 4 `impeccable-*` subagents.
#
# GLOBAL scope, no staging: the installer writes ~/.claude/skills/impeccable/
# (skill + its self-contained engine binary) and ~/.claude/agents/
# impeccable-*.md, and both of those are symlinks into this repo — so the
# global install IS the repo install. Machine-owned and gitignored on both
# sides. `--scope=project` was wrong twice over: it writes <cwd>/.claude/,
# which serves only the directory it ran in, and the staged `mv` that
# followed it moved the skill alone, silently dropping the subagents.
#
# The pin rots. The CLI downloads its skill dist at install time and an older
# release's artifact eventually disappears (`impeccable@3.2.0` → "Download
# failed: invalid zip data", 2026-09-22) — which is what left `make plugin`
# telling the user to run the command by hand. So a pin failure falls back to
# @latest and says, loudly, that the lock needs bumping.
echo "── Step 8d: Impeccable — design detector, skill + agents ──"
echo ""
IMP_DIR="$REPO/skills-external/impeccable"
IMP_SKILL_DIR="$HOME/.claude/skills/impeccable"
IMP_PARKED="$REPO/skills-disabled/impeccable"
IMP_VER=$(pinned_version "impeccable")
NODE_MAJOR=$(node -v 2>/dev/null | sed 's/^v//' | cut -d. -f1)
if [ -z "${NODE_MAJOR:-}" ] || [ "$NODE_MAJOR" -lt 24 ]; then
if [ -f "$IMP_DIR/SKILL.md" ]; then
# One install attempt. $1 = "latest" or an exact version. On failure, IMP_FAIL
# holds the reason. The exit code alone is not enough: with a copy already in
# place, a rotted pin exits 0 ("Could not check for skill updates: invalid
# zip data … Existing skills were left unchanged"), exactly like a genuine
# up-to-date no-op ("Skills are up to date") — only the output tells them
# apart. Probed 2026-09-22 on 4.1.0 vs 3.2.0 in a sandbox HOME.
imp_install() {
local pkg="impeccable" out rc=0
[ "$1" != "latest" ] && pkg="impeccable@$1"
out=$(npx -y "$pkg" skills install -y --providers=claude --scope=global \
--no-hooks 2>&1) || rc=$?
IMP_FAIL=$(printf '%s\n' "$out" \
| grep -E 'Download failed|Could not check for skill updates' \
| head -1 || true)
if [ "$rc" -ne 0 ] && [ -z "$IMP_FAIL" ]; then
IMP_FAIL="installer exited $rc"
fi
[ -z "$IMP_FAIL" ]
}
# Precondition: ~/.claude/{skills,agents} must already be link.sh's symlinks.
# Installing before they exist materializes real directories there, and
# link.sh then refuses to replace them ("is a real directory") — a worse
# failure than skipping, because it needs manual repair.
IMP_READY=true
for _imp_d in skills agents; do
if [ "$(readlink "$HOME/.claude/$_imp_d" 2>/dev/null || true)" != "$REPO/$_imp_d" ]; then
IMP_READY=false
fi
done
if [ "$IMP_READY" != true ]; then
warn "impeccable: ~/.claude/skills and ~/.claude/agents are not this repo's symlinks yet"
warn " → run 'make link' first, then re-run 'make plugin'"
elif [ -z "${NODE_MAJOR:-}" ] || [ "$NODE_MAJOR" -lt 24 ]; then
if [ -f "$IMP_SKILL_DIR/SKILL.md" ] || [ -f "$IMP_PARKED/SKILL.md" ]; then
ok "impeccable already present (update skipped — needs Node >= 24, found ${NODE_MAJOR:-none})"
else
warn "impeccable: needs Node >= 24 (found ${NODE_MAJOR:-none}) — skipped. Bump Node, then: make plugin"
fi
else
IMP_PKG="impeccable"
# A profile may hold impeccable parked in skills-disabled/. Install writes
# to the live slot, so remember the state and put the fresh copy back where
# it was — otherwise `make plugin` silently re-enables a disabled skill.
IMP_WAS_PARKED=false
[ -d "$IMP_PARKED" ] && IMP_WAS_PARKED=true
IMP_USED=""
if [ "$IMP_VER" != "latest" ]; then
IMP_PKG="impeccable@${IMP_VER}"
info "Installing impeccable ${IMP_VER} (pinned in plugins.lock.json, staged)..."
info "Installing impeccable ${IMP_VER} (pinned in plugins.lock.json, global scope)..."
if imp_install "$IMP_VER"; then
IMP_USED="$IMP_VER"
else
warn "impeccable@${IMP_VER} did not install (${IMP_FAIL}) — that release's skill dist is gone upstream"
info "Falling back to impeccable@latest..."
if imp_install latest; then
IMP_USED="latest"
warn "installed @latest instead of the pin. Bump \"impeccable\".version in plugins.lock.json to the version this produced, so the next run is reproducible again."
fi
fi
else
info "Installing impeccable latest (consider pinning in plugins.lock.json)..."
imp_install latest && IMP_USED="latest"
fi
IMP_STAGE=$(mktemp -d)
if (cd "$IMP_STAGE" && npx -y "$IMP_PKG" skills install -y --providers=claude --scope=project --no-hooks >/dev/null 2>&1); then
IMP_SRC=$(find "$IMP_STAGE" -type d -name impeccable -path "*skills*" 2>/dev/null | head -1)
if [ -n "$IMP_SRC" ] && [ -f "$IMP_SRC/SKILL.md" ]; then
rm -rf "$IMP_DIR"
mv "$IMP_SRC" "$IMP_DIR"
ok "impeccable synced to skills-external/ (CLI ${IMP_VER})"
else
warn "impeccable: installer ran but produced no skills/impeccable/SKILL.md — layout changed? Inspect: npx impeccable skills install"
if [ -n "$IMP_USED" ] && [ -f "$IMP_SKILL_DIR/SKILL.md" ]; then
IMP_SKILL_VER=$(sed -n 's/^version:[[:space:]]*//p' "$IMP_SKILL_DIR/SKILL.md" | head -1)
# -L: ~/.claude/agents is a symlink, and find would otherwise stop on it.
IMP_AGENTS=$(find -L "$HOME/.claude/agents" -maxdepth 1 -name 'impeccable-*.md' 2>/dev/null | wc -l)
ok "impeccable installed (CLI ${IMP_USED}, skill ${IMP_SKILL_VER:-?}, ${IMP_AGENTS} agents)"
if [ "$IMP_AGENTS" -eq 0 ]; then
warn "no impeccable-* agent landed in agents/ — the skill's finish/document verbs dispatch to them"
fi
if [ "$IMP_WAS_PARKED" = true ]; then
rm -rf "${IMP_PARKED:?}"
mv "$IMP_SKILL_DIR" "$IMP_PARKED"
info "impeccable was parked by a profile — refreshed copy returned to skills-disabled/"
fi
info "Per-project step, in the agent chat of each frontend project: /impeccable init"
info " (writes PRODUCT.md — the design context every impeccable verb reads)"
elif [ -f "$IMP_SKILL_DIR/SKILL.md" ] || [ -f "$IMP_PARKED/SKILL.md" ]; then
ok "impeccable already present (install failed: ${IMP_FAIL:-no SKILL.md written} — existing copy kept)"
else
if [ -f "$IMP_DIR/SKILL.md" ]; then
ok "impeccable already present (installer failed — existing dist kept)"
else
warn "impeccable install failed — run manually: npx impeccable skills install -y --providers=claude --scope=project --no-hooks"
fi
warn "impeccable install failed (${IMP_FAIL:-no SKILL.md written}) — run manually: npx impeccable skills install -y --providers=claude --scope=global --no-hooks"
fi
rm -rf "$IMP_STAGE"
fi
if [ -L "$HOME/.claude/skills/impeccable" ]; then
ok "impeccable symlink OK"
else
info "Symlinking — will be created by link.sh"
fi
echo ""
@@ -918,42 +952,104 @@ done
echo ""
# ============================================================
# STEP 8.7 — MAGIC MCP (21st-dev) — installed but DISABLED by default
# STEP 8.7 — 21ST.DEV CLI + SKILL PACK — installed but DISABLED by default
# ============================================================
# Magic MCP is a stdio MCP server providing UI component generation
# from 21st.dev. Toggled via lib/toggle-external.sh (same interface as
# gstack, emil-design-eng, etc.). Registered in Claude Code user scope.
# `@21st-dev/cli` (bin `21st`) supersedes the `@21st-dev/magic` MCP server:
# same endpoint, one browser login (`21st login`, token in ~/.config/21st),
# no API key, no MCP process loaded into every session. It ships a pack of
# verified skills (21st-ui-build / -explore / -review / -cli-use / -ai /
# -registry / -design-sync) that drive the CLI from Claude Code.
#
# Default policy: DISABLED at install time. Rationale: MCP tools load
# into every Claude Code session and consume context tokens. Enable
# only when you're actively using Magic.
# Machine-owned dist (impeccable pattern): `21st skills install` writes to
# <HOME>/.claude/skills/<name>/ and REFUSES to follow a symlink anywhere on
# that path — and ~/.claude/skills IS a symlink to this repo's skills/. So
# install under a staged HOME, then move each skill into skills-external/
# (gitignored), where toggle-external.sh / profile.sh symlink it in.
#
# API key: read from $REPO/.env (MAGIC_API_KEY=...) — NEVER committed.
# Template: $REPO/.env.example. Get a key at https://21st.dev/magic
echo "── Step 8.7: Magic MCP (21st-dev) ──────────────────────────"
# Default policy: pack DISABLED at install time — every skill description
# loads into every session. Enable on demand:
# bash lib/toggle-external.sh enable 21st (whole pack)
# /profile design (the 5 design skills)
echo "── Step 8.7: 21st.dev CLI + skill pack ─────────────────────"
echo ""
if [ -x "$REPO/lib/toggle-external.sh" ]; then
MAGIC_STATUS="$(bash "$REPO/lib/toggle-external.sh" status magic 2>/dev/null || echo missing)"
if [ "$MAGIC_STATUS" = "enabled" ]; then
info "Disabling magic MCP by default (enable on demand)..."
bash "$REPO/lib/toggle-external.sh" disable magic >/dev/null
ok "magic MCP disabled — enable with: bash lib/toggle-external.sh enable magic"
if command -v 21st &>/dev/null; then
ok "21st CLI already installed"
else
TFD_VER=$(pinned_version "21st")
if [ "$TFD_VER" != "latest" ]; then
info "Installing @21st-dev/cli@${TFD_VER} (pinned in plugins.lock.json)..."
npm install -g "@21st-dev/cli@${TFD_VER}"
else
ok "magic MCP disabled (default)"
info "Installing @21st-dev/cli@latest (consider pinning in plugins.lock.json)..."
npm install -g @21st-dev/cli
fi
# The key lives in ~/.claude/.env (canonical, BDR-026), reached via the
# repo/.env symlink that toggle-external.sh sources. Self-heal the common
# fresh-machine case: ~/.claude/.env was created AFTER link.sh ran, so the
# symlink is missing and the key looks absent though it's set.
HOME_ENV="$HOME/.claude/.env"
if [ ! -e "$REPO/.env" ] && [ -f "$HOME_ENV" ]; then
ln -sf "$HOME_ENV" "$REPO/.env" 2>/dev/null \
&& info "Linked repo/.env → ~/.claude/.env (was missing)"
if command -v 21st &>/dev/null; then
ok "21st CLI installed"
else
err "21st CLI install failed — run manually: npm install -g @21st-dev/cli"
fi
# Tolerate optional `export ` and leading whitespace; require a value.
MAGIC_KEY_RE='^[[:space:]]*(export[[:space:]]+)?MAGIC_API_KEY=.'
if [ ! -f "$REPO/.env" ] || ! grep -qE "$MAGIC_KEY_RE" "$REPO/.env" 2>/dev/null; then
warn "MAGIC_API_KEY not set in ~/.claude/.env — add it (and run 'make link') before enabling magic"
fi
# Skill pack — staged install, then moved under skills-external/.
if command -v 21st &>/dev/null; then
TFD_STAGE=$(mktemp -d)
if HOME="$TFD_STAGE" 21st skills install --global --agent claude >/dev/null 2>&1; then
TFD_N=0
for _tfd in "$TFD_STAGE"/.claude/skills/*/; do
[ -f "${_tfd}SKILL.md" ] || continue
_tfd_name=$(basename "$_tfd")
rm -rf "${REPO:?}/skills-external/${_tfd_name:?}"
mv "$_tfd" "$REPO/skills-external/$_tfd_name"
TFD_N=$((TFD_N + 1))
done
if [ "$TFD_N" -gt 0 ]; then
ok "21st skill pack synced to skills-external/ ($TFD_N skills)"
else
warn "21st skills install ran but produced no SKILL.md — layout changed? Inspect: 21st skills install --global --agent claude"
fi
elif [ -f "$REPO/skills-external/21st-ui-build/SKILL.md" ]; then
ok "21st skill pack already present (refresh failed — existing copy kept)"
else
warn "21st skill pack install failed — run manually: 21st skills install --global --agent claude"
fi
rm -rf "$TFD_STAGE"
fi
# Auth — detect, then offer login ONLY in an interactive TTY. A non-interactive
# run (CI / headless / re-run) must never open a browser or block on OAuth.
# Search and logo lookup are free; retrieving component code and 21st AI need
# the session. Mirrors the ctx7 auth block (Step 6).
if command -v 21st &>/dev/null; then
# `whoami` is a local token read (no network): "Logged in as <user> (saved …)."
TFD_WHO="$(21st whoami 2>/dev/null | head -1)"
if [[ "$TFD_WHO" == "Logged in as "* ]]; then
ok "21st: ${TFD_WHO%.}"
elif [ -t 0 ] && [ -t 1 ]; then
printf '%b' "${BLUE}→${NC} Sign in to 21st now? (opens a browser) [y/N] "
read -r tfd_ans || tfd_ans=""
if [[ "$tfd_ans" =~ ^[Yy]([Ee][Ss])?$ ]]; then
if 21st login; then
ok "21st authenticated"
else
warn "21st login did not finish — re-run '21st login' anytime"
fi
else
info "Skipped — sign in later with: 21st login"
fi
else
info "Not signed in. Component retrieval and 21st AI need: 21st login"
fi
fi
# Default-disabled, same policy as before the MCP→CLI move.
if [ -x "$REPO/lib/toggle-external.sh" ]; then
TFD_STATUS="$(bash "$REPO/lib/toggle-external.sh" status 21st 2>/dev/null || echo missing)"
if [ "$TFD_STATUS" = "enabled" ]; then
info "Disabling the 21st skill pack by default (enable on demand)..."
bash "$REPO/lib/toggle-external.sh" disable 21st >/dev/null
ok "21st skill pack disabled — enable with: bash lib/toggle-external.sh enable 21st"
else
ok "21st skill pack disabled (default)"
fi
else
warn "lib/toggle-external.sh not found or not executable — skipping"
@@ -1066,7 +1162,7 @@ echo " 🔄 frontend-design — distinctive frontend interfaces, anti-AI-
echo " 🔄 impeccable — /impeccable design verbs + 45-rule deterministic detector (npx impeccable detect)"
echo " 🔄 design-motion-principles — motion/animation design, 3-designer lens (kylezantos)"
echo " 🔄 darwin-skill — autonomous skill optimizer (npx skills, ~/.agents/skills/)"
echo " 🔄 magic MCP — 21st-dev UI generation MCP (toggle: lib/toggle-external.sh enable magic)"
echo " 🔄 21st skill pack — 21st.dev CLI skills, 7 (toggle: lib/toggle-external.sh enable 21st)"
echo ""
echo " All plugins installed at: user scope (~/.claude/plugins/)"
echo " GStack skills symlinked individually into ~/.claude/skills/ (→ submodule)"
+127 -11
View File
@@ -7,8 +7,9 @@ subagents = execution + report only; gates and loop decisions live in the
main loop).
Run this in the ORCHESTRATOR MAIN LOOP, never in a subagent — STEP 2 may
talk to the human. Mandatory passage in every flow; questions are optional
and proportional — a complete request goes through silently.
talk to the human, at contract time (pass A) and again at the flow's PLAN
step (pass B). Questions follow the open choices, never a quota — a complete
request goes through silently.
## STEP 1 — CAPTURE (verbatim)
@@ -17,16 +18,51 @@ message). No paraphrase, no cleanup, no translation, no summarizing. This
section is IMMUTABLE for the life of the run — every later consumer
(planner, dev, verifier) reads THESE words, never a restatement.
## STEP 2 — AMBIGUITY CHECK (questions optional, proportional)
## STEP 2 — CLARIFY (ask, never guess)
Ask ONLY if one of these is missing AND not derivable from the repo:
Two passes, both in the main loop, both may talk to the human.
**Pass A — gaps.** Run here, against the request. Ask if one of these is
missing AND not derivable from the repo:
- a testable expected outcome
- an unambiguous scope (what is allowed to change)
- non-contradictory constraints
Complete request → ZERO questions, stay silent. Otherwise: max 3 questions,
one single batch (house rule: one question upfront, never mid-task). Never
ask what the repo can answer — verify paths/APIs/behavior yourself first.
**Pass B — open choices.** Defined here, run ONCE at the flow's PLAN step
(see "Where pass B fires" below), against the plan just written — that is
where choices become concrete. Enumerate every choice the run will settle
that the request leaves open; keep those in these classes:
1. VISIBLE — the user would see it in the result: placement, label, wording,
color, order, what a click does.
2. PUBLIC NAME — a name that outlives the run: command, flag, endpoint, env
var, a file the human will read.
3. SCOPE — "should X change too?", where the request does not name X.
NEVER ask class 4 — internal technical choices with no observable effect
(function decomposition, data shape, local naming, layout inside an
already-scoped zone). Those are delegated; asking them is the noise that
makes classes 1-3 ignorable. Never ask what the repo or the request already
answers — verify paths/APIs/behavior yourself first.
No question cap. Each pass asks what it finds, in ONE batch. A request that
leaves nothing open goes through silently. More than 5 open choices in pass B
= the request is under-specified: list them, say so, stop — do not fire a
questionnaire. "You decide" / "peu importe" is an answer: record it as
`A: delegated — <default taken>` and never re-ask it.
Pass B answers land in the contract's CLARIFICATIONS marked
`[gated <YYYY-MM-DD>]` — the contract is already on disk by then.
### Where pass B fires
| Flow | Pass B runs at | Against |
|------|----------------|---------|
| feat | STEP 1 PLAN, before 1b CHALLENGE | the PLAN checklist |
| bugfix | STEP 3 FIX PLAN, before 3b | the FIX PLAN |
| hotfix | STEP 1 LOCATE | the 1-2 target files' visible effect |
| ship-feature | STEP 2 PLAN, after the brainstorm | the plan, minus what the brainstorm settled |
| init-project | STEP 3 DESIGN, before VALIDATION GATE #1 | the DESIGN, minus what the interview and brainstorm settled |
| onboard | its STEP 3 interview, unchanged | scope, in one block |
## STEP 3 — DERIVE
@@ -35,6 +71,38 @@ ask what the repo can answer — verify paths/APIs/behavior yourself first.
this conversation.
- FILE SCOPE: paths/zones expected to change, or `repo-wide — <reason>`.
### ORACLES — a criterion a command can decide carries one
Give such a criterion an indented `CHECK:` (the command), `EXPECT:` (a
success-only marker), and `EVIDENCE: pending`.
`bash ~/.claude/lib/gates.sh run <contract>` executes it fail-closed — MET
requires exit 0 **AND** the marker — and writes the result back over the
`EVIDENCE:` line. That persisted evidence is what the fresh verifier reads
as fact instead of trusting the executor's report (GATE 0 in
`lib/verify-secure-loop.md`).
Both attributes or neither. `CHECK:` without `EXPECT:` is a parse error, not
a manual criterion — the runner refuses the whole ledger. Leave a criterion
oracle-free when no command can decide it; the verifier judges those.
Four authoring rules — a gate that cannot fail proves nothing:
1. **Observe the named artifact.** The check reads the file, service, or
measurement the criterion's own words name — never a proxy for it.
`1. invoices reconcile` + `CHECK: echo ok` is valid and worthless.
2. **Success-only marker.** The script runs every assertion, exits nonzero
on any failure, and prints the `EXPECT:` string only after all pass.
3. **Positive control before any absence check.** Run the same logic against
a fixture known to trip it and confirm it fails. A missing file, a wrong
path, and a broken pattern all look exactly like valid absence.
4. **Recompute supplied numbers.** Never copy a figure from the request into
`EXPECT:` — the script derives it from source and prints its own marker.
A number that is its own proof proves nothing.
`CHECK:` is shell code run with our privileges. It is safe only because we
author it in our own repo — never build one out of externally-supplied text
(a scraped URL, a client string); route those through `lib/url-guard.sh`.
## STEP 4 — WRITE TO DISK (immediately, before any next step)
Path: `.claude/tasks/contracts/<YYYY-MM-DD>-<slug>-<HHMM>.md`
@@ -53,12 +121,17 @@ Template:
<the user's exact words>
## CLARIFICATIONS
Q: <question> / A: <answer>
Q: <question> / A: <answer> (pass B and mid-run entries: [gated <YYYY-MM-DD>])
(or: none — request complete)
## ACCEPTANCE CRITERIA
1. <testable criterion>
2. <testable criterion>
1. <criterion a command can decide>
CHECK: <command>
EXPECT: <success-only marker>
EVIDENCE: pending
2. <criterion only human judgement can decide — no CHECK/EXPECT>
(ABANDON: <n> <non-blank reason> — only for a criterion proven impossible)
## FILE SCOPE
<paths/zones>
@@ -68,6 +141,33 @@ Q: <question> / A: <answer>
Print one line to the user, then continue the flow:
`CONTRACT: <path> — <n> criteria, scope <files|repo-wide>, <q> questions asked`
## MID-RUN CLARIFICATION (the channel executors halt into)
An executor cannot talk to the human. It halts with `NEED-DECISION`, the
exact question, the options it sees, and a `CLASS:` tag (visible |
public-name | scope | internal). `/hotfix`: the hotfixer keeps
`DONE | BLOCKED`; a BLOCKED carrying the tag follows the same routing instead
of escalating to `/bugfix`. The orchestrator re-reads the class — the tag is
a hint, not a verdict — then routes:
- visible / public-name / scope → ASK THE HUMAN, verbatim question and
options. Never decide these yourself, never spend a round-trip guessing.
- internal → decide here, note the decision, re-dispatch. The only case the
orchestrator settles alone; max 2 such round-trips → escalate.
Every answer, human or orchestrator, appends to the contract's
CLARIFICATIONS marked `[gated <YYYY-MM-DD>]` — the same micro-gate as scope
enrichment — and to the plan handed to the FRESH re-dispatched executor,
which reads the decision from disk, never from a transcript.
## HOW TO ASK (LRN-102)
The harness reliably renders only the turn's FINAL text; text printed before
a tool call may be swallowed. So:
- up to 4 questions → one `AskUserQuestion` call; option descriptions carry
the context; print nothing the user needs before the call.
- more than 4, or a list handed back for re-specification → plain text, end
the turn.
## Lifecycle
- **REQUEST**: immutable, for the life of the run. Never rewritten, never
@@ -78,6 +178,13 @@ Print one line to the user, then continue the flow:
this micro-gate: human approves → FILE SCOPE gains the entry `[gated]`;
human declines → the dev removes the edit. Without this gate the dev
justifies everything and scope constrains nothing.
- **ABANDONMENT**: a criterion proven impossible within the authorized task
is NEVER deleted and never quietly downgraded. Keep it, append
`ABANDON: <n> <non-blank reason + handoff>` under the criteria, and name it
in the final report. An abandonment is a visible handoff, not a pass: the
verifier cannot return `CONFORME` while one stands, and the run cannot be
described as fully complete. This is the structural half of the house rule
"blocked on an independent sub-part → do the rest, state what's missing".
- **Deep re-scope** (the request itself changes): NEW contract file with
`supersedes: <old path>` in its header — never a rewrite of the old one.
- **Aborted run**: delete the contract file, or commit it with
@@ -90,12 +197,21 @@ Print one line to the user, then continue the flow:
| Flow | Weight |
|------|--------|
| hotfix | Silent autofill — criteria: "symptom gone; build/tests green"; scope = the 1-2 target files. Zero questions ever. |
| hotfix | Pass A silent autofill — criteria: "symptom gone; build/tests green"; scope = the 1-2 target files. Pass B runs at LOCATE against the 1-2 target files' visible effect; a typo fix asks nothing. |
| feat / bugfix | Proportional. bugfix: the DIAGNOSIS feeds the criteria (symptom reproduced-then-gone + regression test present). |
| ship-feature | Full. Design decisions approved at the validation gate append criteria `[gated <date>]` — the human validates the enriched contract, the verifier receives that version. |
| init-project | Full. The interviewer's PROJECT BRIEF pours into the contract (V1 features → criteria). |
| onboard | Audit-scope contract (interview answers → what to audit, which axes). |
Oracles follow the same proportion. hotfix: none — that flow runs no floor
(and no verifier); the hotfixer runs build/tests itself. feat / bugfix: the
suite criterion at minimum, and for bugfix the regression test the DIAGNOSIS
names — its `CHECK:` runs that test alone, so a green result means the
reproduction actually flipped.
ship-feature / init-project: build, suite, and every criterion a command can
settle. onboard: audit criteria are mostly judgement — leave them oracle-free
rather than invent a check that cannot fail.
## Hand-off rule
Downstream consumers (plan step, dev subagents, verifier) receive the
+40 -12
View File
@@ -41,7 +41,8 @@ Tier does NOT change WHAT gets checked. Every non-trivial design tier draws from
the one `design` profile — so the gate checks that profile's **design-core
tools** (the `# GATE-BLOCK:` allowlist in `design.profile`: ui-ux-pro-max,
frontend-design, emil-design-eng, design-motion-principles, impeccable, design-html,
design-review, design-consultation, magic). The profile also bundles
design-review, design-consultation, the `21st` CLI and `21st-ui-build` — the
canary for the whole 21st skill pack). The profile also bundles
browser/plan/shotgun tooling and graphify for convenience; those never trip the
gate. Motion (`design-motion-principles`) and static-HTML (`design-html`) are
already in the core set — checked regardless; their CLAUDE.md "+motion /
@@ -54,7 +55,7 @@ already in the core set — checked regardless; their CLAUDE.md "+motion /
It reads the design-core tools (`# GATE-BLOCK:` in `design.profile`) plus their
types (`profile.sh show design --plain`) and checks each on its own channel —
skill symlink, `claude plugin list`, `claude mcp list`, `command -v`. It never
reads `disabledMcpServers` (unreliable for bi-modal servers like magic/context7).
reads `disabledMcpServers` (unreliable for bi-modal servers like context7).
The core set lives in `design.profile`, not in the script or here — single source.
Exit codes: `0` = ready · `11` = ready-but-unverified (proceed, but surface it) · `10` = incomplete (gate trips) · `2` = error.
@@ -67,21 +68,21 @@ Exit codes: `0` = ready · `11` = ready-but-unverified (proceed, but surface it)
🎨 DESIGN DETECTED — the design toolchain isn't fully active.
activate with /profile design: <skills / ui-ux-pro-max>
required + manual step: <e.g. magic — needs MAGIC_API_KEY>
required + manual step: <e.g. 21st — needs the CLI>
→ run /profile design to activate it, then continue.
- **activate with /profile design** → skills + the plugin; `/profile design`
turns them on directly.
- **required + manual step** → required tools the profile can't flip silently.
**magic lands here: it TRIPS the gate** (it's required for Build), it is NOT
a silent "optional". `/profile design` runs `toggle-external.sh` for magic,
which needs a valid `MAGIC_API_KEY` in `~/.claude/.env` — tell the user to verify it.
**the `21st` CLI lands here: it TRIPS the gate** (it's required for Build),
it is NOT a silent "optional". `/profile design` symlinks the 21st skills,
but the CLI they shell out to is a global npm install: tell the user to run
`npm i -g @21st-dev/cli` then `21st login` (no API key, no MCP).
- Do NOT hand-activate individual tools. The profile is the unit of activation.
- **11 / `READY BUT UNVERIFIED`** → `claude` was unreachable, so the design
plugin/MCP (magic, ui-ux-pro-max) could NOT be checked. Do NOT report a plain
"ready": proceed only after telling the user that N tool(s) went unverified and
having them confirm with `claude mcp list` / `claude plugin list`. Fail-visible,
not fail-silent — the most important tool (magic) is exactly an unverifiable one.
plugin (ui-ux-pro-max) could NOT be checked. Do NOT report a plain "ready":
proceed only after telling the user that N tool(s) went unverified and having
them confirm with `claude plugin list`. Fail-visible, not fail-silent.
### 4. Animation library — suggest-only (fires only on a real motion signal)
@@ -138,6 +139,33 @@ count:
toolchain check handles the skill; this step handles the lib. Don't conflate
them when talking to the user.
### 5. Impeccable design context — suggest-only (one check, one line)
Same class as §4: a PROJECT-side prerequisite, not a tool. `impeccable`
installs globally, but every one of its verbs reads a per-project `PRODUCT.md`
that only `/impeccable init` writes. Without it the skill runs on invented
context, which is worse than not running it — and nothing else in the process
says so, because init has to happen in the agent chat, not in an installer.
**Fires when BOTH hold** — else stay silent:
1. impeccable is active (`skills/impeccable` present, i.e. it did not trip §3).
2. The project has no `PRODUCT.md` at its root.
Evaluate it on the same path as §4: after the toolchain resolves, never on the
INCOMPLETE stop path. One line, non-blocking:
🧭 impeccable has no project context here (no PRODUCT.md) — run `/impeccable init` first? (optional)
**Rules:**
- Non-blocking, and never run `init` unprompted: it interviews the user about
the product, so it needs their attention, not their absence.
- One line per session at most. A refusal is an answer; do not re-ask inside
the same task.
- Skip entirely for a review/audit of a single component and for any non-UI
work. This is for Build and design-system tiers.
### Other toolchains
The script defaults to the `design` profile. A task needing another profile's
@@ -149,8 +177,8 @@ remedy is always `/profile <that>` — a profile, never a lone tool.
- Remedy is ALWAYS a profile (`/profile design`), never an atomic tool toggle —
the profile system is the single source of truth for what's active.
- magic is REQUIRED (it trips the gate), but `/profile design` only enables it
if `MAGIC_API_KEY` is in `~/.claude/.env` — the gate says so; surface that to the user.
- the `21st` CLI is REQUIRED (it trips the gate) and `/profile design` cannot
install it — the gate names the two commands; surface them to the user.
- The design-core set (what trips the gate) is declared in `design.profile` on
the `# GATE-BLOCK:` line(s) — edit there to add/remove a blocking design tool,
not in the script.
+34 -9
View File
@@ -29,12 +29,13 @@
# required-manual required but the profile can't flip it silently (API
# key / external install) — the gate STILL trips, names
# it, and the remedy is `/profile design` + a manual step.
# This is where magic lands: required, never silent.
# This is where the `21st` CLI lands: required, never
# silent (npm i -g @21st-dev/cli, then 21st login).
# Both classes trip the gate. Tools NOT on the GATE-BLOCK allowlist are
# ignored entirely (browser/plan/shotgun tooling, graphify).
#
# disabledMcpServers is NEVER read — unreliable for bi-modal servers
# (magic/context7 can appear there yet be active via another channel).
# (context7 can appear there yet be active via another channel).
#
# Exit: 0 = ready · 11 = ready-but-unverified (proceed, say so) · 10 = incomplete (trips) · 2 = error.
# Usage: design-tool-gate.sh [profile] (default profile: design)
@@ -79,6 +80,30 @@ ensure_claude_on_path() {
}
ensure_claude_on_path
# Same sanitized-PATH problem for `21st` (an npm global bin), with a twist:
# the repair above only fires when claude ITSELF is unresolvable, and claude
# often lives in ~/.local/bin while the npm global bin dir is missing from a
# hook's PATH. Probe for the binary directly and prepend the dir that has it,
# otherwise a perfectly installed CLI reads as "missing" and trips the gate.
ensure_21st_on_path() {
command -v 21st >/dev/null 2>&1 && return
local cand
for cand in \
"$HOME/.local/bin/21st" \
/usr/local/bin/21st; do
[ -x "$cand" ] && { PATH="$(dirname "$cand"):$PATH"; return; }
done
local m newest matches=()
for m in "$HOME"/.nvm/versions/node/*/bin/21st; do
[ -x "$m" ] && matches+=("$m")
done
if [ "${#matches[@]}" -gt 0 ]; then
newest="$(printf '%s\n' "${matches[@]}" | sort -V | tail -1)"
PATH="$(dirname "$newest"):$PATH"
fi
}
ensure_21st_on_path
# Gate scope: the "# GATE-BLOCK:" allowlist (one or more lines, concatenated).
# Empty => fall back to "every gate-relevant entry is in scope" (coarse).
core_set="$(grep '^# GATE-BLOCK:' "$PROFILE_FILE" 2>/dev/null \
@@ -143,8 +168,8 @@ done <<< "$plain"
# Verdict — three outcomes:
# blocking/manual non-empty -> INCOMPLETE (exit 10): the gate trips.
# only unverified non-empty -> READY BUT UNVERIFIED (exit 11): fail-VISIBLE.
# claude was unreachable, so the plugin/MCP (magic, ui-ux-pro-max) could
# not be checked. Never pass this as a silent READY — proceed, but say so.
# claude was unreachable, so the plugin channel (ui-ux-pro-max) could not
# be checked. Never pass this as a silent READY — proceed, but say so.
# nothing pending -> READY (exit 0).
if [ "${#blocking[@]}" -gt 0 ] || [ "${#manual[@]}" -gt 0 ]; then
echo "design toolchain: INCOMPLETE"
@@ -152,9 +177,9 @@ if [ "${#blocking[@]}" -gt 0 ] || [ "${#manual[@]}" -gt 0 ]; then
echo " activate with /profile $PROFILE: ${blocking[*]}"
fi
if [ "${#manual[@]}" -gt 0 ]; then
echo " required + manual step (API key / external install): ${manual[*]}"
echo " required + manual step (external install / sign-in): ${manual[*]}"
case " ${manual[*]} " in
*" magic "*) echo " magic needs MAGIC_API_KEY in ~/.claude/.env (/profile $PROFILE runs toggle-external.sh)" ;;
*" 21st "*) echo " 21st needs the CLI: npm i -g @21st-dev/cli then 21st login" ;;
esac
fi
if [ "${#unverified[@]}" -gt 0 ]; then
@@ -167,9 +192,9 @@ fi
if [ "${#unverified[@]}" -gt 0 ]; then
echo "design toolchain: READY BUT UNVERIFIED — ${#unverified[@]} tool(s) not checked"
echo " unverified (claude CLI unreachable): ${unverified[*]}"
echo " the gate could NOT confirm the design plugin/MCP (e.g. magic,"
echo " ui-ux-pro-max) are active. Proceed only after checking manually:"
echo " claude mcp list claude plugin list"
echo " the gate could NOT confirm the design plugin (ui-ux-pro-max) is"
echo " active. Proceed only after checking manually:"
echo " claude plugin list"
exit 11
fi
+323
View File
@@ -0,0 +1,323 @@
#!/usr/bin/env bash
# Deterministic floor under GATE 1: execute the acceptance criteria that the
# contract itself declares as oracles, fail-closed, and persist the evidence
# INTO the contract file.
#
# bash ~/.claude/lib/gates.sh status <contract> # parse only, never runs
# bash ~/.claude/lib/gates.sh run <contract> # execute + write evidence
#
# rc 0 = MET every runnable criterion passed, no abandonment standing
# 2 = UNMET a runnable criterion failed, or the ledger is malformed
# 3 = ABANDONED runnable criteria all passed, an abandonment still stands
#
# WHY: GATE 1 (lib/verify-secure-loop.md) is an LLM dispatch, and the
# verifier's mandatory `PROOF:` line is a line the verifier WRITES — nothing
# structurally stops it from being produced without anything being executed.
# This runs what the contract declares BEFORE a verifier is ever spawned: a
# red floor sends the executor back for free. Adapted from the `unlazy` skill
# (Leonxlnx/unlazy) — its gate ledger, minus the machinery we do not need.
#
# `run` always re-executes every runnable criterion, including ones already
# recorded MET. Trusting written evidence is exactly the failure this closes,
# so there is no incremental mode to get it wrong with.
#
# TRUST BOUNDARY: `CHECK:` is shell code, run with this process's privileges
# and environment. That is safe here only because the contract is authored by
# our own orchestrator in our own repo — which is why there is no approval
# store (we never execute ledgers inherited from a foreign repo). NEVER build
# a `CHECK:` out of externally-supplied text; route such values through
# lib/url-guard.sh first.
set -uo pipefail
TIMEOUT="${GATES_TIMEOUT:-120}"
EVIDENCE_CAP=140
# Module-level parse tables, index-aligned. Bash has no record type; threading
# eight parallel arrays through every call would cost more readability than
# the explicit data flow buys.
_ID=(); _TEXT=(); _CHECK=(); _EXPECT=(); _EVLINE=(); _EVTEXT=()
_STATUS=(); _EVID=()
_ABANDON_ID=(); _ABANDON_WHY=()
_ERRORS=()
_CUR=-1
_die() { printf 'GATES — VERDICT: ERROR(%s)\n' "$1"; exit 2; }
_err() { _ERRORS+=("$1"); }
_trim() {
local s="$1"
s="${s#"${s%%[![:space:]]*}"}"
printf '%s' "${s%"${s##*[![:space:]]}"}"
}
# ── parse ───────────────────────────────────────────────────────────────────
_new_crit() { # _new_crit <id> <text>
local i
for ((i = 0; i < ${#_ID[@]}; i++)); do
if [ "${_ID[i]}" = "$1" ]; then
_err "duplicate criterion id: $1"
# Orphan what follows instead of aliasing it onto the previous
# criterion, which would hand one gate another gate's oracle.
_CUR=-1
return 0
fi
done
_ID+=("$1"); _TEXT+=("$2")
_CHECK+=(""); _EXPECT+=(""); _EVLINE+=("0"); _EVTEXT+=("")
_CUR=$((${#_ID[@]} - 1))
}
_set_attr() { # _set_attr <CHECK|EXPECT|EVIDENCE> <value> <lineno>
if [ "$_CUR" -lt 0 ]; then
_err "$1 at line $3 belongs to no criterion"
return 0
fi
case "$1" in
CHECK) _CHECK[_CUR]="$2" ;;
EXPECT) _EXPECT[_CUR]="$2" ;;
EVIDENCE) _EVLINE[_CUR]="$3"; _EVTEXT[_CUR]="$2" ;;
esac
}
# An UNINDENTED attribute is diagnosed, never absorbed: silently ignoring it
# would demote a runnable criterion to a manual one, which is the one parse
# bug that turns this checker into a rubber stamp.
_absorb() { # _absorb <raw-line> <lineno>
local body
if [[ "$1" =~ ^([0-9]+)\.[[:space:]]+(.*)$ ]]; then
_new_crit "${BASH_REMATCH[1]}" "${BASH_REMATCH[2]}"
elif [[ "$1" =~ ^ABANDON:[[:space:]]*([0-9]+)?[[:space:]]*(.*)$ ]]; then
_ABANDON_ID+=("${BASH_REMATCH[1]}"); _ABANDON_WHY+=("${BASH_REMATCH[2]}")
elif [[ "$1" =~ ^(CHECK|EXPECT|EVIDENCE): ]]; then
_err "unindented ${BASH_REMATCH[1]}: at line $2"
elif [[ "$1" =~ ^[[:space:]]+(CHECK|EXPECT|EVIDENCE):(.*)$ ]]; then
body="$(_trim "${BASH_REMATCH[2]}")"
_set_attr "${BASH_REMATCH[1]}" "$body" "$2"
fi
}
_parse() { # _parse <file>
local line n=0 fence=0 inblock=0
while IFS= read -r line || [ -n "$line" ]; do
n=$((n + 1))
case "$line" in '```'*) fence=$((1 - fence)); continue ;; esac
[ "$fence" -eq 1 ] && continue
case "$line" in
'## ACCEPTANCE CRITERIA'*) inblock=1; continue ;;
'## '*) inblock=0; continue ;;
esac
[ "$inblock" -eq 1 ] && _absorb "$line" "$n"
done < "$1"
}
# ── validation ──────────────────────────────────────────────────────────────
_validate_oracles() {
local i
for ((i = 0; i < ${#_ID[@]}; i++)); do
if [ -n "${_CHECK[i]}" ] && [ -z "${_EXPECT[i]}" ]; then
_err "criterion ${_ID[i]}: CHECK without EXPECT (partial oracle)"
elif [ -z "${_CHECK[i]}" ] && [ -n "${_EXPECT[i]}" ]; then
_err "criterion ${_ID[i]}: EXPECT without CHECK (partial oracle)"
elif [ -n "${_CHECK[i]}" ] && [ "${_EVLINE[i]}" = "0" ]; then
_err "criterion ${_ID[i]}: runnable but has no EVIDENCE: line"
fi
done
}
_validate_abandons() {
local i j found
for ((i = 0; i < ${#_ABANDON_ID[@]}; i++)); do
found=0
for ((j = 0; j < ${#_ID[@]}; j++)); do
[ "${_ID[j]}" = "${_ABANDON_ID[i]}" ] && found=1
done
[ "$found" -eq 1 ] ||
_err "ABANDON names unknown criterion: '${_ABANDON_ID[i]}'"
[ -n "$(_trim "${_ABANDON_WHY[i]}")" ] ||
_err "ABANDON ${_ABANDON_ID[i]}: blank reason (a handoff needs one)"
done
}
_is_abandoned() { # _is_abandoned <criterion-id>
local i
for ((i = 0; i < ${#_ABANDON_ID[@]}; i++)); do
[ "${_ABANDON_ID[i]}" = "$1" ] && return 0
done
return 1
}
# ── execution ───────────────────────────────────────────────────────────────
# One line, capped, newlines flattened: the smallest output that proves the
# outcome. Full logs stay in the terminal, never in the contract.
_decisive() { # _decisive <combined-output>
local flat
flat="$(printf '%s' "$1" | tr '\n\r\t' ' ' | tr -s ' ')"
flat="$(_trim "$flat")"
if [ "${#flat}" -gt "$EVIDENCE_CAP" ]; then
printf '%s…' "${flat:0:$EVIDENCE_CAP}"
else
printf '%s' "$flat"
fi
}
# Fail-closed: exit 0 AND the marker. A nonzero process never passes because
# its error text happens to contain the expected token.
_run_one() { # _run_one <idx>
local i="$1" out rc
out="$(timeout "$TIMEOUT" bash -c "${_CHECK[i]}" 2>&1)"
rc=$?
_STATUS[i]="NOT-MET"
if [ "$rc" -eq 124 ]; then
_EVID[i]="NOT-MET timeout=${TIMEOUT}s"
elif [ "$rc" -ne 0 ]; then
_EVID[i]="NOT-MET exit=$rc (nonzero) :: $(_decisive "$out")"
elif [[ "$out" != *"${_EXPECT[i]}"* ]]; then
_EVID[i]="NOT-MET exit=0 marker-absent :: $(_decisive "$out")"
else
_STATUS[i]="MET"
_EVID[i]="MET exit=0 marker-found :: $(_decisive "$out")"
fi
}
_run_all() {
local i
for ((i = 0; i < ${#_ID[@]}; i++)); do
_STATUS[i]=""; _EVID[i]=""
[ -n "${_CHECK[i]}" ] && _run_one "$i"
done
}
_evline_owner() { # _evline_owner <lineno> — echoes idx, or nothing
local i
for ((i = 0; i < ${#_ID[@]}; i++)); do
if [ "${_EVLINE[i]}" = "$1" ] && [ -n "${_EVID[i]}" ]; then
printf '%s' "$i"
return 0
fi
done
}
# Rewrites only the EVIDENCE lines of criteria that actually ran; every other
# byte of the contract is copied through, indentation included.
_write_back() { # _write_back <file>
local tmp line n=0 idx
tmp="$(mktemp)" || _die "mktemp failed"
while IFS= read -r line || [ -n "$line" ]; do
n=$((n + 1))
idx="$(_evline_owner "$n")"
if [ -n "$idx" ]; then
printf '%s%s\n' "${line%%[![:space:]]*}" "EVIDENCE: ${_EVID[idx]}"
else
printf '%s\n' "$line"
fi
done < "$1" > "$tmp"
cat "$tmp" > "$1" && rm -f "$tmp"
}
# ── report ──────────────────────────────────────────────────────────────────
# A recorded `pending`, or a criterion that never ran, is PENDING — never MET.
# `status` reports what the file says; it does not revalidate old evidence.
_row_state() { # _row_state <idx>
local i="$1"
_is_abandoned "${_ID[i]}" && { printf 'ABANDONED'; return 0; }
[ -z "${_CHECK[i]}" ] && { printf 'MANUAL'; return 0; }
[ -n "${_STATUS[i]:-}" ] && { printf '%s' "${_STATUS[i]}"; return 0; }
case "${_EVTEXT[i]}" in
MET' '*) printf 'MET-RECORDED' ;;
*) printf 'PENDING' ;;
esac
}
_report_rows() {
local i state
for ((i = 0; i < ${#_ID[@]}; i++)); do
state="$(_row_state "$i")"
printf ' %-3s %-13s %s\n' "${_ID[i]}" "$state" "${_TEXT[i]}"
done
}
_report_abandons() {
local i
for ((i = 0; i < ${#_ABANDON_ID[@]}; i++)); do
printf ' ABANDONED %s — %s\n' "${_ABANDON_ID[i]}" "${_ABANDON_WHY[i]}"
done
}
_count_state() { # _count_state <state>
local i n=0
for ((i = 0; i < ${#_ID[@]}; i++)); do
[ "$(_row_state "$i")" = "$1" ] && n=$((n + 1))
done
printf '%s' "$n"
}
_verdict() { # _verdict <mode> — prints the line, returns the rc
local unmet pending abandoned
if [ "${#_ERRORS[@]}" -gt 0 ]; then
printf 'GATES — VERDICT: ERROR(%s)\n' "${#_ERRORS[@]}"
return 2
fi
unmet="$(_count_state NOT-MET)"
pending="$(_count_state PENDING)"
abandoned="$(_count_state ABANDONED)"
[ "$unmet" -gt 0 ] &&
{ printf 'GATES — VERDICT: UNMET(%s)\n' "$unmet"; return 2; }
if [ "$1" = "status" ] && [ "$pending" -gt 0 ]; then
printf 'GATES — VERDICT: PENDING(%s)\n' "$pending"
return 2
fi
[ "$abandoned" -gt 0 ] &&
{ printf 'GATES — VERDICT: ABANDONED(%s)\n' "$abandoned"; return 3; }
printf 'GATES — VERDICT: MET\n'
return 0
}
_report() { # _report <mode> <file>
local rc
printf 'GATES — %s (%s)\n' "$2" "$1"
_report_rows
_report_abandons
[ "${#_ERRORS[@]}" -gt 0 ] && printf ' ERROR %s\n' "${_ERRORS[@]}"
printf 'RUNNABLE: %s of %s criteria; timeout %ss\n' \
"$(_runnable_count)" "${#_ID[@]}" "$TIMEOUT"
_verdict "$1"
rc=$?
return "$rc"
}
_runnable_count() {
local i n=0
for ((i = 0; i < ${#_ID[@]}; i++)); do
[ -n "${_CHECK[i]}" ] && n=$((n + 1))
done
printf '%s' "$n"
}
# ── entry point ─────────────────────────────────────────────────────────────
main() { # main <status|run> <contract>
local mode="$1" file="$2"
[ -r "$file" ] || _die "contract unreadable: $file"
_parse "$file"
[ "${#_ID[@]}" -gt 0 ] ||
_die "no numbered criteria under ## ACCEPTANCE CRITERIA"
_validate_oracles
_validate_abandons
if [ "$mode" = "run" ] && [ "${#_ERRORS[@]}" -eq 0 ]; then
_run_all
_write_back "$file"
fi
_report "$mode" "$file"
}
case "${1:-}" in
status|run)
[ $# -eq 2 ] || _die "usage: gates.sh {status|run} <contract-path>"
main "$1" "$2"
;;
*) _die "usage: gates.sh {status|run} <contract-path>" ;;
esac
+296
View File
@@ -0,0 +1,296 @@
#!/usr/bin/env bash
# ============================================================
# lib/gstack-playwright.sh — gstack's Playwright: OS-support bump +
# read-only browser-cache report.
#
# Sourced by: install-plugins.sh, update-all.sh, doctor.sh — all three run
# `set -euo pipefail`. gstack_bump_playwright_if_unsupported and
# gstack_browsers_report are called as BARE STATEMENTS under that inherited
# errexit, so they `return 0` on every path and every capture that could
# fail is guarded (`|| true` or an `if`), never a bare `&&`/`||`-less
# statement. gstack_submodule_update_with_bump is the ONE function allowed
# to return non-zero — callers use it ONLY as an `if` condition.
#
# No `set -euo pipefail` here (mirrors lib/detect-plugins.sh): a sourced
# lib must not change the caller's shell options.
#
# See BDR-029 (bump origin), BLK-008 (Chromium-unsupported-OS saga),
# LRN-040 (two-layer fix — this file is layer 1 only).
# ============================================================
_GSPW_GREEN='\033[0;32m'; _GSPW_YELLOW='\033[1;33m'; _GSPW_BLUE='\033[0;34m'
_GSPW_NC='\033[0m'
_gspw_ok() { echo -e " ${_GSPW_GREEN}✓${_GSPW_NC} $1"; }
_gspw_warn() { echo -e " ${_GSPW_YELLOW}⚠${_GSPW_NC} $1"; }
_gspw_info() { echo -e " ${_GSPW_BLUE}→${_GSPW_NC} $1"; }
# ── OS support ───────────────────────────────────────────────────────────
# gstack_pw_ostag [os_release_path] — "ubuntu<VERSION_ID>" on Ubuntu, empty
# otherwise. `|| true` on the capture: the reproduced bug had this exact
# line abort every non-Ubuntu host under inherited errexit.
gstack_pw_ostag() {
local path="${1:-/etc/os-release}" tag
[ -r "$path" ] || return 0
# shellcheck disable=SC1090
tag="$(. "$path" 2>/dev/null
[ "${ID:-}" = ubuntu ] && printf 'ubuntu%s' "${VERSION_ID:-}")" || true
if [ -n "$tag" ]; then
printf '%s' "$tag"
fi
return 0
}
# gstack_pw_supports <playwright_core_lib_dir> <ostag> — 0 supported, 1 not.
# Always called from an `if`/`&&` context, never as a bare statement.
gstack_pw_supports() {
local pwlib="$1" ostag="$2"
[ -n "$ostag" ] && [ -d "$pwlib" ] || return 1
grep -rqs "$ostag" "$pwlib" 2>/dev/null
}
# _gspw_run_timeout <dir> <cmd...> — runs <cmd> in <dir>, under `timeout 300`
# when available (absent on stock macOS). Exit 124 = the wrapped command was
# killed by the timeout. Callers MUST invoke this via `cmd || rc=$?` (never
# bare) so a non-zero exit never trips the caller's inherited errexit.
_gspw_run_timeout() {
local dir="$1"; shift
if command -v timeout >/dev/null 2>&1; then
( cd "$dir" && timeout 300 "$@" ) >/dev/null 2>&1
else
( cd "$dir" && "$@" ) >/dev/null 2>&1
fi
}
# _gspw_bump_install <gstack_dir> — populate node_modules at the pinned
# version so its support list can be read. 0 proceed, 1 give up silently
# (both installs failed, matches the pre-existing silent behavior), 2 give
# up loud (a timeout truncated node_modules — the support grep would then
# read a half-written tree).
_gspw_bump_install() {
local dir="$1" rc=0
_gspw_run_timeout "$dir" bun install --frozen-lockfile || rc=$?
if [ "$rc" -eq 0 ]; then
return 0
elif [ "$rc" -eq 124 ]; then
_gspw_warn "bun install timed out — skipping Playwright bump"
return 2
fi
rc=0
_gspw_run_timeout "$dir" bun install || rc=$?
if [ "$rc" -eq 0 ]; then
return 0
elif [ "$rc" -eq 124 ]; then
_gspw_warn "bun install timed out — skipping Playwright bump"
return 2
fi
return 1
}
# _gspw_bump_add_latest <gstack_dir> — 0 ran (support re-checked by caller
# regardless of bun's own exit code, exactly as the pre-existing code did),
# 2 timed out (node_modules left half-written — caller must NOT re-check).
_gspw_bump_add_latest() {
local dir="$1" rc=0
_gspw_run_timeout "$dir" bun add playwright@latest || rc=$?
if [ "$rc" -eq 124 ]; then
_gspw_warn "bun add playwright@latest timed out — skipping Playwright bump"
return 2
fi
return 0
}
# gstack_bump_playwright_if_unsupported <gstack_dir> — BDR-029: bump
# gstack's pinned Playwright when it lacks a build for this OS, so
# `./setup` rebuilds the browse binary against a version that has one.
# OS-gated, idempotent, non-fatal — `return 0` on every path.
gstack_bump_playwright_if_unsupported() {
local gstack_dir="$1" ostag pwlib rc=0
[ -d "$gstack_dir" ] && [ -r /etc/os-release ] || return 0
ostag="$(gstack_pw_ostag)"
[ -n "$ostag" ] || return 0
if ! command -v bun >/dev/null 2>&1; then
export PATH="$HOME/.bun/bin:$PATH"
fi
pwlib="$gstack_dir/node_modules/playwright-core/lib"
_gspw_info "checking gstack's Playwright OS support ($ostag)..."
_gspw_bump_install "$gstack_dir" || rc=$?
[ "$rc" -eq 0 ] || return 0
if gstack_pw_supports "$pwlib" "$ostag"; then
return 0
fi
_gspw_info "gstack's Playwright lacks $ostag support — bumping to \
latest (local submodule edit)..."
rc=0
_gspw_bump_add_latest "$gstack_dir" || rc=$?
[ "$rc" -eq 0 ] || return 0
if gstack_pw_supports "$pwlib" "$ostag"; then
_gspw_ok "gstack Playwright bumped — now supports $ostag (browse \
binary rebuilt by ./setup)"
else
_gspw_warn "Playwright bump didn't add $ostag support — gstack \
browser may stay unavailable"
fi
return 0
}
# ── Submodule update ──────────────────────────────────────────────────────
# gstack_submodule_update_with_bump <repo> [sub_path] — the ONE function
# allowed to return non-zero; callers use it ONLY as an `if` condition.
# Never touches the submodule working tree: on failure it prints git's own
# stderr verbatim (never parsed) and returns 1. On success it re-applies
# the bump (closes BDR-029's caveat: the bump used to survive only until
# the next `git submodule update`).
gstack_submodule_update_with_bump() {
local repo="$1" sub="${2:-skills-external/gstack}" err rc=0
err="$(git -C "$repo" submodule update --remote "$sub" 2>&1 >/dev/null)" \
|| rc=$?
if [ "$rc" -eq 0 ]; then
gstack_bump_playwright_if_unsupported "$repo/$sub"
return 0
fi
_gspw_warn "$err"
if [ -n "$(git -C "$repo/$sub" status --porcelain \
-- package.json bun.lock 2>/dev/null)" ]; then
_gspw_info "local Playwright bump (package.json/bun.lock) was not \
re-applied — re-run: make plugin"
fi
return 1
}
# ── Browsers report (read-only) ───────────────────────────────────────────
# _gspw_dir_name_parts <cache_dir_name> — prints "normalized_name revision"
# split on the LAST '-', mapping '_' -> '-' on the name (Playwright writes
# chromium_headless_shell-1228 on disk; browsers.json names it
# chromium-headless-shell).
_gspw_dir_name_parts() {
local rev="${1##*-}" name="${1%-*}"
printf '%s %s' "${name//_/-}" "$rev"
}
# _gspw_browser_referenced <playwright_core_path> <dir_name> — does that
# install require this cache directory (base revision or any
# revisionOverrides value)?
_gspw_browser_referenced() {
local json="$1/browsers.json" name rev
[ -r "$json" ] || return 1
read -r name rev <<< "$(_gspw_dir_name_parts "$2")"
awk -F'"' -v want_name="$name" -v want_rev="$rev" '
$2 == "name" { cur = $4; in_ov = 0 }
$2 == "revision" && !in_ov && cur == want_name && $4 == want_rev {
found = 1
}
$2 == "revisionOverrides" { in_ov = 1 }
in_ov && $2 != "revisionOverrides" && cur == want_name \
&& $4 == want_rev { found = 1 }
/^[[:space:]]*}/ { in_ov = 0 }
END { exit !found }
' "$json"
}
# _gspw_browser_name_known <playwright_core_path> <dir_name> — is the NAME
# listed at all, regardless of revision? (distinguishes "unknown revision"
# from "unreferenced" in the report.)
_gspw_browser_name_known() {
local json="$1/browsers.json" name rev
[ -r "$json" ] || return 1
read -r name rev <<< "$(_gspw_dir_name_parts "$2")"
awk -F'"' -v want="$name" '$2 == "name" && $4 == want { found = 1 }
END { exit !found }' "$json"
}
# _gspw_install_label <playwright_core_path> — "<dir-before-node_modules>
# <version>", e.g. "gstack 1.61.1".
_gspw_install_label() {
local pw_path="$1" parent version
parent=$(basename "$(dirname "$(dirname "$pw_path")")")
version=$(awk -F'"' '$2 == "version" { print $4; exit }' \
"$pw_path/package.json" 2>/dev/null) || true
printf '%s %s' "$parent" "${version:-?}"
}
# _gspw_registered_installs <cache_dir> — valid playwright-core paths (dir
# exists, browsers.json readable), one per line. A `.links` entry whose
# target is gone or unreadable is silently excluded here (it is counted as
# a broken link by the caller instead).
_gspw_registered_installs() {
local links_dir="$1/.links" f target
[ -d "$links_dir" ] || return 0
for f in "$links_dir"/*; do
[ -f "$f" ] || continue
target=$(cat "$f" 2>/dev/null) || true
[ -n "$target" ] || continue
if [ -d "$target" ] && [ -r "$target/browsers.json" ]; then
printf '%s\n' "$target"
fi
done
return 0
}
# _gspw_report_dir_line <dir_name> <install_paths_newline_sep> — prints the
# report line for one cache directory. Returns 1 only when truly
# unreferenced (caller tallies that); "unknown revision" does not count.
_gspw_report_dir_line() {
local dir_name="$1" installs="$2" p labels="" known=0
while IFS= read -r p; do
[ -n "$p" ] || continue
if _gspw_browser_referenced "$p" "$dir_name"; then
labels="${labels:+$labels, }$(_gspw_install_label "$p")"
elif _gspw_browser_name_known "$p" "$dir_name"; then
known=1
fi
done <<< "$installs"
if [ -n "$labels" ]; then
_gspw_info "$dir_name: $labels"
return 0
elif [ "$known" -eq 1 ]; then
_gspw_info "$dir_name: unknown revision"
return 0
fi
_gspw_info "$dir_name: unreferenced"
return 1
}
# gstack_browsers_report [cache_dir] — read-only. `$1` (or
# PLAYWRIGHT_BROWSERS_PATH, or ~/.cache/ms-playwright) is resolved once;
# "0" (documented as "bundle into node_modules") and any non-directory
# degrade to a silent no-cache path. `return 0` on every path.
gstack_browsers_report() {
local cache installs total links_total valid_count broken=0 unref=0 d name
cache="${1:-${PLAYWRIGHT_BROWSERS_PATH:-$HOME/.cache/ms-playwright}}"
[ "$cache" = "0" ] && return 0
[ -d "$cache" ] || return 0
installs="$(_gspw_registered_installs "$cache")"
links_total=$(find "$cache/.links" -maxdepth 1 -type f 2>/dev/null \
| wc -l | tr -d ' ') || true
valid_count=$(printf '%s\n' "$installs" | grep -c . || true)
broken=$((links_total - valid_count))
total=$(du -sh "$cache" 2>/dev/null | awk '{print $1}') || true
_gspw_info "Playwright browsers: $cache (${total:-0})"
for d in "$cache"/*-[0-9]*; do
[ -d "$d" ] || continue
name=$(basename "$d")
_gspw_report_dir_line "$name" "$installs" || unref=$((unref + 1))
done
_gspw_info "${unref} unreferenced, ${broken} broken link(s)"
if [ "$unref" -gt 0 ] || [ "$broken" -gt 0 ]; then
_gspw_warn "unreferenced/broken Playwright browser dirs — re-run \
\`playwright install\`, which prunes stale revisions"
fi
return 0
}
# ── CLI dispatch (only when executed, not sourced) — browsers-report ONLY.
# The write functions (the bump, the submodule update) stay sourced-only: a
# CLI verb would expose `bun add playwright@latest` as a command-line entry
# point. ────────────────────────────────────────────────────────────────
if [ "${BASH_SOURCE[0]}" = "${0}" ]; then
case "${1:-}" in
browsers-report) shift; gstack_browsers_report "$@" ;;
*) echo "usage: gstack-playwright.sh browsers-report [cache_dir]" >&2
exit 2 ;;
esac
fi
+16 -21
View File
@@ -11,7 +11,7 @@
# Mechanism:
# - Skills (gstack/external/personal): symlink toggle skills/ ↔ skills-disabled/
# - Plugins: `claude plugin enable|disable <name>@<marketplace>`
# - MCPs: delegated to lib/toggle-external.sh for known servers (magic),
# - MCPs: advisory (none managed since BDR-093 — MANAGED_MCPS is empty),
# advisory otherwise
# - CLIs: advisory only (rtk, gsd, ctx7, graphify — installed externally)
# - `set` is SYMMETRIC on managed items (BDR-079): plugins, external packs
@@ -51,7 +51,6 @@ SKILLS_DIR="$REPO/skills"
DISABLED_DIR="$REPO/skills-disabled"
GSTACK_SRC="$REPO/skills-external/gstack" # gstack submodule — source of truth for gstack skills
PROFILES_DIR="$REPO/lib/profiles"
TOGGLE_EXTERNAL="$REPO/lib/toggle-external.sh"
ACTIVE_CACHE="$REPO/.active-profile" # statusline reads this — keep fast (single-line file, profile name only)
# Plugins that are toggle-managed by `set`. Anything NOT in this list is
@@ -73,13 +72,20 @@ MANAGED_EXTERNALS=(
frontend-design
design-motion-principles
impeccable
21st-ui-build
21st-ui-explore
21st-ui-review
21st-cli-use
21st-ai
)
# MCP servers that are toggle-managed by `set`, both ways (enable AND
# disable), delegated to lib/toggle-external.sh. Same allowlist doctrine.
MANAGED_MCPS=(
magic
)
# Empty since 2026-09-22: `magic` was the only entry and 21st.dev replaced
# its MCP server with a CLI + skill pack (the 5 design skills are managed as
# externals above). The `mcp` type itself stays supported — a profile can
# still list an MCP, it is then advisory rather than auto-toggled.
MANAGED_MCPS=()
# Plugins that MUST stay enabled — `set` will refuse to disable these even if
# they're not in the profile. (Defensive: belt-and-suspenders alongside
@@ -324,15 +330,12 @@ enable_skill() {
fi
;;
mcp)
# Advisory only. The delegation branch that lived here served `magic`,
# the single managed MCP; 21st.dev replaced it with a CLI (BDR-093), so
# MANAGED_MCPS is empty and nothing is auto-registered. Re-add a branch
# here the day a profile owns an MCP server again.
if [ "$(skill_status "$skill" mcp)" = "enabled" ]; then
: # already on
elif [ "$skill" = "magic" ] && [ -x "$TOGGLE_EXTERNAL" ]; then
# Known MCP — delegate to lib/toggle-external.sh which handles env vars.
if bash "$TOGGLE_EXTERNAL" enable magic 2>&1 | grep -qE "enabled|already"; then
ok "enabled MCP: magic"
else
info "MCP 'magic' could not be enabled (check .env for MAGIC_API_KEY)"
fi
else
info "MCP '$skill' not registered — run: claude mcp add $skill -- <command>"
fi
@@ -394,15 +397,7 @@ disable_skill() {
info "plugin '$skill' — manual: claude plugin disable $skill@<marketplace>"
;;
mcp)
if [ "$skill" = "magic" ] && [ -x "$TOGGLE_EXTERNAL" ]; then
if bash "$TOGGLE_EXTERNAL" disable magic 2>&1 | grep -qE "disabled|already"; then
ok "disabled MCP: magic"
else
info "MCP 'magic' — manual disable failed"
fi
else
info "MCP '$skill' — manual: claude mcp remove $skill"
fi
info "MCP '$skill' — manual: claude mcp remove $skill"
;;
cli)
: # never auto-uninstall CLIs
+13 -4
View File
@@ -7,7 +7,8 @@
# tooling, graphify) is bundled for convenience but never blocks. Keep these
# lines in sync when adding/removing a core design tool.
# GATE-BLOCK: frontend-design ui-ux-pro-max emil-design-eng design-html
# GATE-BLOCK: design-motion-principles design-review design-consultation magic
# GATE-BLOCK: design-motion-principles design-review design-consultation
# GATE-BLOCK: 21st 21st-ui-build
# Core design skills (gstack)
design-shotgun
@@ -30,11 +31,19 @@ frontend-design external
design-motion-principles external
impeccable external
# External: 21st.dev pack — CLI-driven (no MCP, no API key). 21st-registry
# and 21st-design-sync are publishing flows; installed but left parked.
21st-ui-build external
21st-ui-explore external
21st-ui-review external
21st-cli-use external
21st-ai external
# Plugin (auto-toggle)
ui-ux-pro-max plugin@ui-ux-pro-max-skill
# MCP — auto-toggle via lib/toggle-external.sh (needs MAGIC_API_KEY in .env)
magic mcp
# CLIs (advisory only — installed/not-installed)
# 21st is NOT advisory: it is on the GATE-BLOCK list, so a missing CLI trips
# the design gate. Install: npm i -g @21st-dev/cli then 21st login
21st cli
graphify cli
+6 -1
View File
@@ -86,9 +86,14 @@ ui-ux-pro-max plugin@ui-ux-pro-max-skill
# claude plugin enable pr-review-toolkit@claude-code-plugins
# or profile-based: bash lib/profile.sh apply audit (audit.profile keeps it;
# a later `set full` re-disables it — MANAGED_PLUGINS lifecycle).
magic mcp
21st-ui-build external
21st-ui-explore external
21st-ui-review external
21st-cli-use external
21st-ai external
# === CLIs (advisory) =================================================
21st cli
ctx7 cli
graphify cli
gsd cli
+7 -2
View File
@@ -49,9 +49,14 @@ emil-design-eng external
frontend-design external
design-motion-principles external
impeccable external
21st-ui-build external
21st-ui-explore external
21st-ui-review external
21st-cli-use external
21st-ai external
ui-ux-pro-max plugin@ui-ux-pro-max-skill
magic mcp
# === CLIs (advisory) =================================================
# === CLIs ============================================================
21st cli
ctx7 cli
graphify cli
+10 -2
View File
@@ -38,11 +38,19 @@ frontend-design external
design-motion-principles external
impeccable external
# External: 21st.dev pack (publishing flows 21st-registry / -design-sync
# stay parked)
21st-ui-build external
21st-ui-explore external
21st-ui-review external
21st-cli-use external
21st-ai external
# Plugin: UI/UX intelligence (auto-toggle)
ui-ux-pro-max plugin@ui-ux-pro-max-skill
# MCP: 21st-dev Magic component generator
magic mcp
# CLI: 21st.dev component catalog + UI generation (needs `21st login`)
21st cli
# CLI: ctx7 (doc lookup for fast-evolving libs like Next.js)
ctx7 cli
+10 -3
View File
@@ -48,8 +48,15 @@ fi
tf "verbatim request immutable" "$LIB" "REQUEST (verbatim — IMMUTABLE)"
tf "contracts dir committed path" "$LIB" ".claude/tasks/contracts/"
tf "unique per-run slug" "$LIB" "<YYYY-MM-DD>-<slug>-<HHMM>"
tf "silent when complete" "$LIB" "ZERO questions"
tf "question budget" "$LIB" "max 3 questions"
tf "silent when nothing open" "$LIB" "goes through silently"
tf "no question cap" "$LIB" "No question cap"
tf "pass B classes" "$LIB" "PUBLIC NAME"
tf "class 4 excluded" "$LIB" "NEVER ask class 4"
tf "over-5 guard" "$LIB" "More than 5 open choices"
tf "delegated answer" "$LIB" "delegated —"
tf "mid-run channel" "$LIB" "## MID-RUN CLARIFICATION"
tf "class tag" "$LIB" "CLASS:"
tf "how to ask" "$LIB" "## HOW TO ASK"
tf "aborted status" "$LIB" "status: aborted"
tf "never left dirty" "$LIB" "NEVER left dirty"
tf "scope enrichment micro-gate" "$LIB" "micro-gate"
@@ -67,7 +74,7 @@ fi
tr_ "frontmatter name" "$AGT" "^name: verifier$"
tr_ "tools read-only set" "$AGT" "^tools: Read, Grep, Glob, Bash$"
tn "no write-capable tools" "$AGT" "^tools:.*(Edit|Write|NotebookEdit)"
tf "verdict grammar" "$AGT" "VERIFY — VERDICT: CONFORME | ECARTS(n) | ERROR(<reason>)"
tf "verdict grammar" "$AGT" "VERIFY — VERDICT: CONFORME | ECARTS(n) | ABANDONED(n) | ERROR(<reason>)"
tf "blind — no iteration history" "$AGT" "NEVER receive iteration history"
tf "blind — complete every time" "$AGT" "every verification is complete and blind"
tf "unverifiable is not met" "$AGT" "\`UNVERIFIABLE\` ≠ \`MET\`"
@@ -22,6 +22,7 @@ check D8-dash-file "$(fire 'ecc_dashboard.py')" quiet
# --- Harness-generated inputs must be QUIET even with UI tokens ---
check D9-tasknotif "$(fire '<task-notification> <task-id>x</task-id> add css header fonts')" quiet
check D10-notif-file "$(fire '<task-notification> design-motion-principles keyframe done')" quiet
check D11-bare-ux "$(fire 'changement ux vu de tes trouvailles')" quiet
# --- Real UI signals must still FIRE ---
check F1-button "$(fire 'add a button')" fire
@@ -33,6 +34,7 @@ check F6-frontdesign "$(fire 'frontend design work')" fire
check F7-admin-dash "$(fire 'admin dashboard screen')" fire
check F8-animation "$(fire 'add an animation')" fire
check F9-designsys "$(fire 'our design system')" fire
check F10-bare-ui "$(fire 'revois l'\''ui du panneau admin')" fire
# --- Fire is logged (time + token + excerpt) ---
tmp="$(mktemp -d)"
+320
View File
@@ -0,0 +1,320 @@
#!/usr/bin/env bash
# ============================================================
# lib/gates.sh — behavioural tests + structure locks for the
# deterministic floor (GATE 0, lib/verify-secure-loop.md).
#
# Fail-closed is the entire point of this runner, so every
# "looks green but must not pass" case is asserted explicitly:
# nonzero exit carrying the marker, marker absent, timeout,
# unindented attribute silently demoting a gate to manual.
# Non-execution is proved with a sentinel file, and the
# sentinel's own positive control is asserted first — an
# absence check that was never able to fire proves nothing.
# ============================================================
set -uo pipefail
REPO="$(cd "$(dirname "$0")/../.." && pwd)"
GATES="$REPO/lib/gates.sh"
WORK="$(mktemp -d)"
trap 'rm -rf "$WORK"' EXIT
PASS=0; FAIL=0; N=0
LAST=""
ok() { echo " PASS $1"; PASS=$((PASS + 1)); }
bad() { echo " FAIL $1 — $2"; FAIL=$((FAIL + 1)); }
# gate <label> <mode> <expected-verdict> <expected-rc> <<< fixture-on-stdin
gate() {
local label="$1" mode="$2" want="$3" wantrc="$4" out rc
N=$((N + 1)); LAST="$WORK/c$N.md"
cat > "$LAST"
out="$(GATES_TIMEOUT="${GATES_TIMEOUT:-120}" \
bash "$GATES" "$mode" "$LAST" 2>&1)"
rc=$?
if [[ "$out" == *"$want"* ]] && [ "$rc" -eq "$wantrc" ]; then
ok "$label"
else
bad "$label" "want '$want' rc=$wantrc, got rc=$rc"
printf '%s\n' "$out" | sed 's/^/ /'
fi
}
has() {
if grep -qF -- "$2" "$LAST"; then ok "$1"; else bad "$1" "missing: $2"; fi
}
exists() {
if [ -e "$1" ]; then ok "$2"; else bad "$2" "sentinel absent: $1"; fi
}
absent() {
if [ -e "$1" ]; then bad "$2" "sentinel created: $1"; else ok "$2"; fi
}
echo "── fail-closed execution ──"
gate "exit 0 + marker = MET" run "GATES — VERDICT: MET" 0 <<'EOF'
## ACCEPTANCE CRITERIA
1. green
CHECK: echo "MARKER-OK"
EXPECT: MARKER-OK
EVIDENCE: pending
EOF
has "evidence written back" "EVIDENCE: MET exit=0 marker-found"
# The case a naive checker gets wrong: the marker IS in the output, but the
# process failed. Substring matching alone would certify a broken build.
gate "nonzero exit + marker = UNMET" run "GATES — VERDICT: UNMET(1)" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. lies
CHECK: echo "MARKER-OK"; exit 7
EXPECT: MARKER-OK
EVIDENCE: pending
EOF
has "nonzero recorded honestly" "NOT-MET exit=7 (nonzero)"
gate "exit 0 + no marker = UNMET" run "GATES — VERDICT: UNMET(1)" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. silent success is not success
CHECK: echo "something else"
EXPECT: MARKER-OK
EVIDENCE: pending
EOF
has "marker-absent recorded" "NOT-MET exit=0 marker-absent"
GATES_TIMEOUT=1 gate "timeout = UNMET" run "GATES — VERDICT: UNMET(1)" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. hangs
CHECK: sleep 5
EXPECT: never
EVIDENCE: pending
EOF
has "timeout recorded" "NOT-MET timeout=1s"
gate "manual-only contract passes through" run "RUNNABLE: 0 of 2" 0 <<'EOF'
## ACCEPTANCE CRITERIA
1. a human reads the copy
2. the design matches the brief
EOF
echo "── non-execution (sentinel), positive control first ──"
# Positive control: prove the sentinel mechanism can fire at all.
gate "sentinel fires when a CHECK runs" run "GATES — VERDICT: MET" 0 <<EOF
## ACCEPTANCE CRITERIA
1. control
CHECK: touch "$WORK/fired"; echo "M"
EXPECT: M
EVIDENCE: pending
EOF
exists "$WORK/fired" "positive control: sentinel created"
gate "status never executes" status "GATES — VERDICT: PENDING(1)" 2 <<EOF
## ACCEPTANCE CRITERIA
1. must not run
CHECK: touch "$WORK/status-ran"; echo "M"
EXPECT: M
EVIDENCE: pending
EOF
absent "$WORK/status-ran" "status did not execute"
has "status did not write evidence" "EVIDENCE: pending"
gate "fenced example is not a gate" run "RUNNABLE: 1 of 1" 0 <<EOF
## ACCEPTANCE CRITERIA
1. real
CHECK: echo "R"
EXPECT: R
EVIDENCE: pending
\`\`\`markdown
2. documentation example, invisible to the parser
CHECK: touch "$WORK/fenced-ran"; echo "nope"
EXPECT: nope
EVIDENCE: pending
\`\`\`
EOF
absent "$WORK/fenced-ran" "fenced CHECK never executed"
gate "malformed ledger executes nothing" run "GATES — VERDICT: ERROR" 2 <<EOF
## ACCEPTANCE CRITERIA
1. would run if the ledger parsed
CHECK: touch "$WORK/malformed-ran"; echo "M"
EXPECT: M
EVIDENCE: pending
2. partial oracle poisons the whole ledger
CHECK: echo "x"
EVIDENCE: pending
EOF
absent "$WORK/malformed-ran" "malformed ledger did not execute"
has "malformed ledger not written" "EVIDENCE: pending"
echo "── parse strictness ──"
gate "CHECK without EXPECT" run "CHECK without EXPECT" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. partial
CHECK: echo x
EVIDENCE: pending
EOF
gate "EXPECT without CHECK" run "EXPECT without CHECK" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. partial
EXPECT: x
EVIDENCE: pending
EOF
# An unindented CHECK must be diagnosed, never absorbed: silently ignoring it
# demotes a runnable criterion to a manual one — the one parse bug that turns
# this runner into a rubber stamp.
gate "unindented attribute is diagnosed" run "unindented CHECK:" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. sneaky
CHECK: echo x
EXPECT: x
EVIDENCE: pending
EOF
gate "runnable without EVIDENCE line" run "has no EVIDENCE: line" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. no ledger slot
CHECK: echo x
EXPECT: x
EOF
gate "duplicate criterion id" run "duplicate criterion id: 1" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. first
EVIDENCE: pending
1. second
EVIDENCE: pending
EOF
# After a rejected duplicate the following attributes must be orphaned, not
# aliased onto the previous criterion — that would hand one gate another's
# oracle and let a stale EVIDENCE line satisfy it.
gate "duplicate orphans what follows" run "belongs to no criterion" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. real
CHECK: echo x
EXPECT: x
EVIDENCE: pending
1. duplicate
CHECK: echo y
EXPECT: y
EVIDENCE: pending
EOF
gate "no numbered criteria" run "no numbered criteria" 2 <<'EOF'
## ACCEPTANCE CRITERIA
nothing numbered here
EOF
echo "── abandonment ──"
gate "valid abandonment = rc 3" run "GATES — VERDICT: ABANDONED(1)" 3 <<'EOF'
## ACCEPTANCE CRITERIA
1. green
CHECK: echo "M"
EXPECT: M
EVIDENCE: pending
2. impossible
EVIDENCE: pending
ABANDON: 2 upstream API offline; handoff recorded in BLK-099
EOF
gate "blank abandonment reason" run "blank reason" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. green
EVIDENCE: pending
ABANDON: 1
EOF
gate "abandonment naming nothing" run "unknown criterion" 2 <<'EOF'
## ACCEPTANCE CRITERIA
1. green
EVIDENCE: pending
ABANDON: 9 names a criterion that does not exist
EOF
echo "── usage ──"
# usage <label> <expected-substring> <argv...>
usage() {
local label="$1" want="$2" out rc; shift 2
out="$(bash "$GATES" "$@" 2>&1)"; rc=$?
if [ "$rc" -eq 2 ] && [[ "$out" == *"$want"* ]]; then
ok "$label"
else
bad "$label" "rc=$rc out=$out"
fi
}
usage "no args = ERROR rc 2" "usage:"
usage "missing contract = ERROR rc 2" "contract unreadable" run "$WORK/nope.md"
usage "unknown mode refused" "usage:" frobnicate "$WORK/c1.md"
# ── structure locks on the doctrine this runner is wired into ───────────────
CI="$REPO/lib/contract-interview.md"
VS="$REPO/lib/verify-secure-loop.md"
AGT="$REPO/agents/verifier.md"
FE="$REPO/agents/feater.md"
BF="$REPO/agents/bugfixer.md"
lock() { # lock <label> <file> <fixed-string>
if grep -qF -- "$3" "$2" 2>/dev/null; then
ok "$1"
else
bad "$1" "missing: $3"
fi
}
echo "── contract-interview.md oracle doctrine ──"
lock "oracle section" "$CI" "### ORACLES"
lock "runner named" "$CI" "lib/gates.sh run <contract>"
lock "fail-closed spelled out" "$CI" "exit 0 **AND** the marker"
lock "both or neither" "$CI" "Both attributes or neither"
lock "rule observe artifact" "$CI" "Observe the named artifact"
lock "rule success-only" "$CI" "Success-only marker"
lock "rule positive control" "$CI" "Positive control before any absence check"
lock "rule recompute numbers" "$CI" "Recompute supplied numbers"
lock "shell trust boundary" "$CI" "url-guard.sh"
lock "template carries oracle" "$CI" "EXPECT: <success-only marker>"
lock "abandonment lifecycle" "$CI" "**ABANDONMENT**"
lock "abandonment not deleted" "$CI" "NEVER deleted"
echo "── verify-secure-loop.md GATE 0 ──"
lock "gate 0 exists" "$VS" "## GATE 0 — DETERMINISTIC FLOOR"
lock "gate 0 no dispatch" "$VS" "**No verifier is dispatched**"
lock "gate 0 loop bound" "$VS" "Max 3 floor iterations"
lock "gate 0 budget separate" "$VS" "not eat the conformity budget"
lock "malformed = main loop" "$VS" "never dispatch a dev for it"
lock "order invariant" "$VS" "GATE 0 → GATE 1 → GATE 2"
echo "── verifier.md oracle + abandonment ──"
lock "verdict grammar" "$AGT" \
"VERIFY — VERDICT: CONFORME | ECARTS(n) | ABANDONED(n) | ERROR(<reason>)"
lock "red oracle wins" "$AGT" "NEVER overrides a red or unrun oracle"
lock "vacuous oracle caught" "$AGT" "vacuous oracle"
lock "oracle != english" "$AGT" "proves the ORACLE, not the"
lock "never edits contract" "$AGT" "reported, never rewritten"
lock "abandoned is not met" "$AGT" "\`ABANDONED\` ≠ \`MET\`"
lock "abandoned routes human" "$AGT" "direct human gate, never a dev loop"
echo "── executor four passes ──"
lock "feater passes" "$FE" "## FOUR PASSES"
lock "feater no placeholder" "$FE" "no deferred remainder you plan"
lock "feater never widens" "$FE" "they never widen it"
lock "bugfixer passes" "$BF" "## FOUR PASSES"
lock "bugfixer stays minimal" "$BF" "keep the fix minimal"
lock "bugfixer neg control" "$BF" "**Negative control.**"
lock "bugfixer test must fail" "$BF" "A test that passes both ways"
lock "feater class tag" "$FE" "CLASS:"
lock "bugfixer class tag" "$BF" "CLASS:"
lock "hotfixer class tag" "$REPO/agents/hotfixer.md" "CLASS:"
echo ""
echo "gates: $PASS pass, $FAIL fail"
[ "$FAIL" -eq 0 ]
+213
View File
@@ -0,0 +1,213 @@
#!/usr/bin/env bash
# lib/tests/gstack-playwright.test.sh — lib/gstack-playwright.sh (T1..T17)
#
# git 2.53 defaults protocol.file to "user", which blocks submodule clone
# and fetch. The fixture git calls alone are not enough: the
# `git submodule update --remote` under test runs INSIDE the lib, in a
# fresh git subprocess spawned from THIS process — so the override is
# exported for the WHOLE test process, not passed per-command.
set -u
export GIT_CONFIG_COUNT=1
export GIT_CONFIG_KEY_0=protocol.file.allow
export GIT_CONFIG_VALUE_0=always
ROOT="$(cd "$(dirname "$0")/../.." && pwd)"
L="$ROOT/lib/gstack-playwright.sh"
pass=0; fail=0
check() { if [ "$2" = "$3" ]; then pass=$((pass+1)); else fail=$((fail+1));
printf 'FAIL %s: got[%s] want[%s]\n' "$1" "$2" "$3"; fi; }
tmp="$(mktemp -d)"; trap 'rm -rf "$tmp"' EXIT
git_id() { git -C "$1" config user.email t@example.com
git -C "$1" config user.name Test; }
# shellcheck source=lib/gstack-playwright.sh
source "$L"
# ── T1/T2 — ostag detection ──────────────────────────────────────────────
printf 'ID=ubuntu\nVERSION_ID="24.04"\n' > "$tmp/os-ubuntu"
printf 'ID=debian\nVERSION_ID="12"\n' > "$tmp/os-debian"
check T1-ostag-ubuntu "$(gstack_pw_ostag "$tmp/os-ubuntu")" "ubuntu24.04"
check T2-ostag-other "$(gstack_pw_ostag "$tmp/os-debian")" ""
# ── T3 — errexit safety of the ostag capture (regression: the reproduced
# bug aborted the whole caller on every non-Ubuntu host) ──
cat > "$tmp/t3.sh" <<EOF
#!/usr/bin/env bash
set -euo pipefail
source "$L"
gstack_pw_ostag "$tmp/os-debian"
echo REACHED
EOF
check T3-errexit-safe "$(bash "$tmp/t3.sh" 2>&1 | tail -1)" "REACHED"
# ── T4/T5 — pw_supports, no bun involved ─────────────────────────────────
mkdir -p "$tmp/pwlib-hit" "$tmp/pwlib-miss"
echo "supports ubuntu24.04 and others" > "$tmp/pwlib-hit/index.js"
echo "supports nothing relevant" > "$tmp/pwlib-miss/index.js"
t4_rc=0; gstack_pw_supports "$tmp/pwlib-hit" ubuntu24.04 >/dev/null 2>&1 \
|| t4_rc=$?
check T4-supports-hit "$t4_rc" 0
t5_rc=0; gstack_pw_supports "$tmp/pwlib-miss" ubuntu24.04 >/dev/null 2>&1 \
|| t5_rc=$?
check T5-supports-miss "$t5_rc" 1
# ── T6/T7 — submodule update, real git fixtures ──────────────────────────
mkdir -p "$tmp/upstream6"
git -C "$tmp/upstream6" init -q -b main; git_id "$tmp/upstream6"
printf '{"a":1}\n' > "$tmp/upstream6/package.json"
git -C "$tmp/upstream6" add package.json
git -C "$tmp/upstream6" commit -q -m init
mkdir -p "$tmp/repo6"
git -C "$tmp/repo6" init -q -b main; git_id "$tmp/repo6"
printf 'x\n' > "$tmp/repo6/README.md"
git -C "$tmp/repo6" add README.md
git -C "$tmp/repo6" commit -q -m init
git -C "$tmp/repo6" -c protocol.file.allow=always \
submodule add -q -b main "$tmp/upstream6" gstack-sub
git -C "$tmp/repo6" config submodule.gstack-sub.branch main
git -C "$tmp/repo6" commit -q -m "add submodule"
printf 'extra\n' > "$tmp/upstream6/extra.txt"
git -C "$tmp/upstream6" add extra.txt
git -C "$tmp/upstream6" commit -q -m "upstream update"
t6_out=$(
gstack_bump_playwright_if_unsupported() { echo BUMP_CALLED; }
gstack_submodule_update_with_bump "$tmp/repo6" "gstack-sub"
echo "rc=$?"
)
t6_calls=$(printf '%s\n' "$t6_out" | grep -c BUMP_CALLED)
t6_rc=$(printf '%s\n' "$t6_out" | grep -o 'rc=[0-9]*')
check T6-update-success-bumps "$t6_calls:$t6_rc" "1:rc=0"
mkdir -p "$tmp/upstream7"
git -C "$tmp/upstream7" init -q -b main; git_id "$tmp/upstream7"
printf '{"a":1}\n' > "$tmp/upstream7/package.json"
printf 'lockA\n' > "$tmp/upstream7/bun.lock"
git -C "$tmp/upstream7" add package.json bun.lock
git -C "$tmp/upstream7" commit -q -m init
mkdir -p "$tmp/repo7"
git -C "$tmp/repo7" init -q -b main; git_id "$tmp/repo7"
printf 'x\n' > "$tmp/repo7/README.md"
git -C "$tmp/repo7" add README.md
git -C "$tmp/repo7" commit -q -m init
git -C "$tmp/repo7" -c protocol.file.allow=always \
submodule add -q -b main "$tmp/upstream7" gstack-sub
git -C "$tmp/repo7" config submodule.gstack-sub.branch main
git -C "$tmp/repo7" commit -q -m "add submodule"
# upstream changes package.json content (would overwrite the local edit)
printf '{"a":2}\n' > "$tmp/upstream7/package.json"
git -C "$tmp/upstream7" add package.json
git -C "$tmp/upstream7" commit -q -m "upstream bumps package.json"
# local Playwright-bump-style dirty edit, never committed
printf '{"a":99}\n' > "$tmp/repo7/gstack-sub/package.json"
echo "T7: update-conflict"
before_pkg=$(cat "$tmp/repo7/gstack-sub/package.json")
before_lock=$(cat "$tmp/repo7/gstack-sub/bun.lock")
t7_out=$(gstack_submodule_update_with_bump "$tmp/repo7" "gstack-sub" 2>&1)
t7_rc=$?
after_pkg=$(cat "$tmp/repo7/gstack-sub/package.json")
after_lock=$(cat "$tmp/repo7/gstack-sub/bun.lock")
t7_files_ok=N
[ "$before_pkg" = "$after_pkg" ] && [ "$before_lock" = "$after_lock" ] \
&& t7_files_ok=Y
t7_hint_ok=N
printf '%s\n' "$t7_out" | grep -q 'make plugin' && t7_hint_ok=Y
t7_state="$t7_rc:$t7_files_ok:$t7_hint_ok"
check T7-update-conflict-nondestructive "$t7_state" "1:Y:Y"
# ── T8 — no destructive command anywhere in the lib source ───────────────
d8=OK
sed 's/#.*//' "$L" | grep -qE 'git [^|;]*(checkout|reset|clean|stash)' && d8=BAD
sed 's/#.*//' "$L" | grep -qwE '(rm|rmdir|unlink|truncate|mv)' && d8=BAD
check T8-no-destructive-command "$d8" OK
# ── T9-T14 — browsers-report, fixture cache + playwright-core installs ───
mkdir -p "$tmp/installs/fixA/node_modules/playwright-core"
cat > "$tmp/installs/fixA/node_modules/playwright-core/browsers.json" <<'EOF'
{
"comment": "Do not edit this file, use utils/roll_browser.js",
"browsers": [
{
"name": "chromium",
"revision": "1228",
"installByDefault": true
},
{
"name": "chromium-headless-shell",
"revision": "1228",
"installByDefault": true
},
{
"name": "webkit",
"revision": "2311",
"installByDefault": true,
"revisionOverrides": {
"mac14": "2251",
"debian11-x64": "2105"
}
},
{
"name": "ffmpeg",
"revision": "1011",
"installByDefault": true
}
]
}
EOF
cat > "$tmp/installs/fixA/node_modules/playwright-core/package.json" <<'EOF'
{
"name": "playwright-core",
"version": "1.61.1"
}
EOF
mkdir -p "$tmp/cache1/.links" \
"$tmp/cache1/chromium-1228" \
"$tmp/cache1/chromium_headless_shell-1228" \
"$tmp/cache1/webkit-2105" \
"$tmp/cache1/firefox-9999" \
"$tmp/cache1/chromium-9999"
printf '%s' "$tmp/installs/fixA/node_modules/playwright-core" \
> "$tmp/cache1/.links/link-valid"
printf '%s' "$tmp/no-such-install/node_modules/playwright-core" \
> "$tmp/cache1/.links/link-broken"
out1="$(gstack_browsers_report "$tmp/cache1" 2>&1)"
has1() { printf '%s\n' "$out1" | grep -q "$1" && echo Y; }
check T9-report-referenced "$(has1 'chromium-1228: fixA 1.61.1')" Y
check T10-report-underscore-dir \
"$(has1 'chromium_headless_shell-1228: fixA 1.61.1')" Y
t11_unref=$(has1 'firefox-9999: unreferenced')
t11_unknown=$(has1 'chromium-9999: unknown revision')
check T11-report-unreferenced "$t11_unref$t11_unknown" YY
check T12-report-broken-link "$(has1 '1 broken link')" Y
check T13-report-revision-override "$(has1 'webkit-2105: fixA 1.61.1')" Y
mkdir -p "$tmp/cache2/.links" "$tmp/cache2/chromium-1228"
printf '%s' "$tmp/installs/fixA/node_modules/playwright-core" \
> "$tmp/cache2/.links/link-valid"
out2="$(gstack_browsers_report "$tmp/cache2" 2>&1)"; rc2=$?
zero2=$(printf '%s\n' "$out2" | grep -q '0 unreferenced, 0 broken link(s)' \
&& echo Y)
check T14-report-zero-counts-exit-0 "$rc2:$zero2" "0:Y"
# ── T15/T16 — degrade silently, nothing on stderr ────────────────────────
err15="$(gstack_browsers_report "$tmp/does-not-exist-cache" 2>&1 1>/dev/null)"
rc15=$?
check T15-report-no-cache "$rc15:[$err15]" "0:[]"
err16="$(gstack_browsers_report "0" 2>&1 1>/dev/null)"
rc16=$?
check T16-report-browsers-path-zero "$rc16:[$err16]" "0:[]"
# ── T17 — sourcing emits nothing ──────────────────────────────────────────
out17="$(bash -c "source '$L'; :" 2>&1)"
check T17-source-safe "[$out17]" "[]"
printf 'PASS=%s FAIL=%s\n' "$pass" "$fail"; [ "$fail" -eq 0 ]
+2
View File
@@ -27,6 +27,7 @@ tf "shf enrich at gate" "$SHF" "ENRICH the STEP 0e contract"
tf "shf gated marker" "$SHF" "[gated <date>]"
tf "shf verify+secure step" "$SHF" "STEP 5 — VERIFY + SECURE"
tf "shf uses shared include" "$SHF" "lib/verify-secure-loop.md"
tf "shf gate0 floor" "$SHF" "GATE 0 — deterministic floor"
tf "shf judges enriched" "$SHF" "ENRICHED contract"
tf "shf orthogonal to review" "$SHF" "DISTINCT axis from STEP 6 code review"
@@ -36,6 +37,7 @@ tf "ini criteria from V1" "$INI" "V1 FEATURES (each testable)"
tf "ini enrich at gate1" "$INI" "ENRICH the STEP 1 contract"
tf "ini verify+secure step" "$INI" "STEP 9 — VERIFY + SECURE"
tf "ini uses shared include" "$INI" "lib/verify-secure-loop.md"
tf "ini gate0 floor" "$INI" "GATE 0 — deterministic floor"
tf "ini adds security gate" "$INI" "adds the security gate init-project previously lacked"
echo "-- onboard (explicit NO-LOOP audit) --"
+3 -1
View File
@@ -57,6 +57,7 @@ tf "feat contract step" "$FSK" "STEP 0.7 — CONTRACT"
tf "feat contract-interview" "$FSK" "lib/contract-interview.md"
tf "feat verify+secure step" "$FSK" "STEP 4 — VERIFY + SECURE"
tf "feat uses shared include" "$FSK" "lib/verify-secure-loop.md"
tf "feat gate0 floor" "$FSK" "GATE 0 — deterministic floor"
tf "feat nominal 1+1 dispatch" "$FSK" "verifier + one security dispatch"
tf "feat dispatches feater" "$FSK" 'subagent_type="feater"'
@@ -65,6 +66,7 @@ tf "bug contract step" "$BSK" "STEP 3.5 — CONTRACT"
tf "bug diagnosis feeds it" "$BSK" "feeds it: REQUEST verbatim"
tf "bug fresh gates" "$BSK" "the two fresh gates per"
tf "bug uses shared include" "$BSK" "lib/verify-secure-loop.md"
tf "bug gate0 floor" "$BSK" "GATE 0 — deterministic floor"
tf "bug dispatches bugfixer" "$BSK" 'subagent_type="bugfixer"'
echo "── agents/bugfixer.md (bugfix executor — sonnet, no Agent) ──"
@@ -79,7 +81,7 @@ tf "hotfixer report grammar" "$HOT" "HOTFIX-EXEC REPORT"
echo "── skills/hotfix/SKILL.md (hotfix wiring — revert, not loop) ──"
tf "hotfix silent contract" "$HSKL" "STEP 1.7 — CONTRACT (silent autofill)"
tf "hotfix zero questions" "$HSKL" "questions ever"
tf "hotfix pass B at locate" "$HSKL" "run pass B of"
tf "hotfix security gate" "$HSKL" "Security gate (fresh auditor)"
tf "hotfix block reverts" "$HSKL" "failure REVERTS, never loops"
tf "hotfix no verifier" "$HSKL" "No verifier is dispatched at hotfix weight"
+11
View File
@@ -70,6 +70,17 @@ has "agents/validator-analyzer.md" 'model: sonnet'
has "agents/plan-challenger.md" 'model: opus'
fm_lacks "agents/client-handover-writer.md" 'model:'
fm_lacks "agents/interviewer.md" 'model:'
# 12) tour multi-project fan-out (BDR-084) — runner INHERITS the session
# model (no pin: it carries reflection), inner agents keep their tiers,
# dispatch is single-message, capitalize stays in the main loop.
has "skills/tour/SKILL.md" 'STEP 0b — MULTI-PROJECT FAN-OUT'
# shellcheck disable=SC2016 # literal backticks — no expansion intended
has "skills/tour/SKILL.md" 'NO `model` override'
has "skills/tour/SKILL.md" 'ALL in a SINGLE message'
has "skills/tour/SKILL.md" 'MAIN LOOP ONLY, never inside a'
has "skills/tour/SKILL.md" 'RUNNER FAILED'
lacks "skills/tour/SKILL.md" 'description="tour runner", model='
has "skills/onboard/SKILL.md" 'model="opus"'
has "skills/tour/SKILL.md" 'model="opus"'
has "lib/challenge-plan.md" 'BDR-076'
+1
View File
@@ -24,6 +24,7 @@ has "$A" "correctness"
has "$A" "robustness"
has "$A" "simplicity"
has "$A" "Report-only"
has "$A" "grounded doubt" # uncertain findings → [MINOR], not self-censored (Opus 5 literalism)
# 2) reusable phase — the mechanism lives here (one canonical include)
has "$L" 'subagent_type="plan-challenger"'
+16 -16
View File
@@ -1,6 +1,9 @@
#!/usr/bin/env bash
# lib/tests/profile-set-managed.test.sh — `set` symmetry on managed
# externals + MCPs, gstack on-demand, external from-source (BDR-079).
# externals, gstack on-demand, external from-source (BDR-079). The MCP
# assertions went with `magic` (2026-09-22): MANAGED_MCPS is empty now, the
# 21st skills that replaced it are managed as externals, so the pack's
# park/restore round-trip is what this covers on that side.
# Hermetic: fixture repo via *_REPO_OVERRIDE + fake `claude` on PATH.
set -u
ROOT="$(cd "$(dirname "$0")/../.." && pwd)"
@@ -10,13 +13,13 @@ check() { if [ "$2" = "$3" ]; then pass=$((pass+1)); else fail=$((fail+1));
FX="$(mktemp -d)"; trap 'rm -rf "$FX"' EXIT
mkdir -p "$FX/skills" "$FX/skills-disabled" "$FX/lib/profiles" "$FX/bin" \
"$FX/skills-external/emil-design-eng" "$FX/skills-external/other-ext"
"$FX/skills-external/emil-design-eng" "$FX/skills-external/other-ext" \
"$FX/skills-external/21st-ui-build"
for g in gs-a gs-b gs-c; do
mkdir -p "$FX/skills-external/gstack/$g"
touch "$FX/skills-external/gstack/$g/SKILL.md"
done
cp "$ROOT/lib/profile.sh" "$ROOT/lib/toggle-external.sh" "$FX/lib/"
printf 'MAGIC_API_KEY=test-secret-000\n' > "$FX/.env"
# Non-managed external, enabled from the start — must never be touched.
ln -s "$FX/skills-external/other-ext" "$FX/skills/other-ext"
@@ -25,22 +28,18 @@ cat > "$FX/lib/profiles/designish.profile" <<'EOF'
gs-a
gs-b
emil-design-eng external
magic mcp
21st-ui-build external
EOF
cat > "$FX/lib/profiles/backendish.profile" <<'EOF'
gs-c
EOF
# Fake claude: logs every call; keeps MCP registry state in a flat file.
# Fake claude: logs every call. No MCP state to keep — MANAGED_MCPS is empty,
# so `set` must never reach for `claude mcp` at all (asserted below).
cat > "$FX/bin/claude" <<EOF
#!/usr/bin/env bash
FX="$FX"
echo "\$*" >> "\$FX/claude-calls.log"
case "\$1 \${2:-}" in
"mcp list") cat "\$FX/mcp-state" 2>/dev/null ;;
"mcp add") echo "magic: stub" > "\$FX/mcp-state" ;;
"mcp remove") : > "\$FX/mcp-state" ;;
esac
exit 0
EOF
chmod +x "$FX/bin/claude"
@@ -48,14 +47,14 @@ chmod +x "$FX/bin/claude"
run() { PATH="$FX/bin:$PATH" PROFILE_REPO_OVERRIDE="$FX" \
TOGGLE_EXTERNAL_REPO_OVERRIDE="$FX" bash "$FX/lib/profile.sh" "$@"; }
# --- set designish: gstack on-demand + external from-source + magic on ---
# --- set designish: gstack on-demand + externals from-source (21st + emil) ---
run set designish >/dev/null 2>&1
check T1-gsa-on "$([ -e "$FX/skills/gs-a" ] && echo on || echo off)" on
check T2-gsb-on "$([ -e "$FX/skills/gs-b" ] && echo on || echo off)" on
check T3-gsc-off "$([ -e "$FX/skills/gs-c" ] && echo on || echo off)" off
check T4-emil-src "$([ -L "$FX/skills/emil-design-eng" ] && echo on || echo off)" on
check T5-magic-on "$(grep -c '^magic:' "$FX/mcp-state" 2>/dev/null)" 1
check T6-add-call "$(grep -c '^mcp add magic' "$FX/claude-calls.log")" 1
check T5-21st-src "$([ -L "$FX/skills/21st-ui-build" ] && echo on || echo off)" on
check T6-no-mcp "$(grep -c '^mcp ' "$FX/claude-calls.log" || true)" 0
# --- set backendish: managed leftovers parked/unregistered ---
run set backendish >/dev/null 2>&1
@@ -63,14 +62,15 @@ check T7-gsc-on "$([ -e "$FX/skills/gs-c" ] && echo on || echo off)" on
check T8-gsa-park "$([ -e "$FX/skills-disabled/gstack__gs-a" ] && echo p || echo n)" p
check T9-emil-off "$([ -e "$FX/skills/emil-design-eng" ] && echo on || echo off)" off
check T10-emil-park "$([ -e "$FX/skills-disabled/emil-design-eng" ] && echo p || echo n)" p
check T11-magic-off "$(grep -c '^magic:' "$FX/mcp-state" 2>/dev/null || true)" 0
check T12-rm-call "$(grep -c '^mcp remove magic' "$FX/claude-calls.log")" 1
check T11-21st-off "$([ -e "$FX/skills/21st-ui-build" ] && echo on || echo off)" off
check T12-21st-park "$([ -e "$FX/skills-disabled/21st-ui-build" ] && echo p || echo n)" p
check T13-other-untouched "$([ -e "$FX/skills/other-ext" ] && echo on || echo off)" on
# --- back to designish: parked external restored (not re-sourced) ---
run set designish >/dev/null 2>&1
check T14-emil-back "$([ -e "$FX/skills/emil-design-eng" ] && echo on || echo off)" on
check T15-park-gone "$([ -e "$FX/skills-disabled/emil-design-eng" ] && echo p || echo n)" n
check T16-magic-back "$(grep -c '^magic:' "$FX/mcp-state" 2>/dev/null)" 1
check T16-21st-back "$([ -e "$FX/skills/21st-ui-build" ] && echo on || echo off)" on
check T17-no-mcp-ever "$(grep -c '^mcp ' "$FX/claude-calls.log" || true)" 0
printf 'PASS=%s FAIL=%s\n' "$pass" "$fail"; [ "$fail" -eq 0 ]
+82
View File
@@ -0,0 +1,82 @@
#!/usr/bin/env bash
# lib/tests/seo-geo-contract.test.sh — census: seo/geo agent ⇄ dispatcher
# machine contract (C1 de-prescription, 2026-07-30). Locks every string a
# consumer parses BEFORE the choreography reword, so the reword commits
# prove contract preservation by keeping this green. Complements
# model-routing.test.sh (which already locks model pins + MODE:* +
# COLLECTION COMPLETE).
set -u
R="$(cd "$(dirname "$0")/../.." && pwd)"
pass=0; fail=0
ok() { pass=$((pass+1)); }
ko() { fail=$((fail+1)); printf 'FAIL %s\n' "$1"; }
has() { if grep -qF "$2" "$R/$1"; then ok; else ko "$1 missing: $2"; fi; }
SEO=agents/seo-analyzer.md
GEO=agents/geo-analyzer.md
# 1) judge verdict grammar — DISPATCHER ERROR CONTRACT (skills/seo STEP 1,
# skills/geo STEP 1B) fail-closes on this exact shape
has "$SEO" 'SEO JUDGE — VERDICT: ERROR('
has "$GEO" 'GEO JUDGE — VERDICT: ERROR('
has "skills/seo/SKILL.md" 'SEO JUDGE — VERDICT: ERROR('
has "skills/geo/SKILL.md" 'GEO JUDGE — VERDICT: ERROR('
# 2) fix-bundle section + apply sentinel — parsed by /seo STEP 1b/1.5 and
# /geo STEP 1b/2 before any L1 apply
for f in "$SEO" "$GEO" skills/seo/SKILL.md skills/geo/SKILL.md; do
has "$f" '## FIX BUNDLE'
has "$f" 'READY TO APPLY — awaiting dispatcher confirmation'
done
# 3) signals handoff — judge loads the collect artifact fail-closed
has "$SEO" '.audit/seo-signals-'
has "$GEO" '.audit/geo-signals-'
has "skills/seo/SKILL.md" '.audit/seo-signals-<RUNID_SEO>.md'
has "skills/geo/SKILL.md" '.audit/geo-signals-<RUNID>.md'
# 4) STEP numbering — dispatchers reference agent step ranges literally:
# /seo: "STEP 2-5" collect · "STEP 6-11" seo judge · "STEP 6-12" geo
# judge · "STEP 12-14" seo template · "STEP 13-15" geo template;
# /geo: "STEP 0-5". Lock EVERY step header on the agent side (interiors
# too — merging/renumbering one silently re-points the dispatch ranges)
# and the ranges on the dispatcher side.
for n in 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14; do has "$SEO" "## STEP $n —"; done
for n in 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15; do has "$GEO" "## STEP $n —"; done
has "skills/seo/SKILL.md" 'STEP 2-5'
has "skills/seo/SKILL.md" 'STEP 6-11'
has "skills/seo/SKILL.md" 'STEP 6-12'
has "skills/seo/SKILL.md" 'STEP 12-14'
has "skills/seo/SKILL.md" 'STEP 13-15'
has "skills/geo/SKILL.md" 'STEP 0-5'
has "skills/geo/SKILL.md" '6-12, report scoring' # "STEP\n6-12" line-wraps
has "skills/geo/SKILL.md" 'STEP 13-15'
# 5) collect-mode report emission (dispatcher waits on it between phases)
has "$SEO" 'COLLECT REPORT'
has "$GEO" 'COLLECT REPORT'
# 6) bundle-item routing + the item fields the L1 appliers parse
# (the item is pasted verbatim into hotfixer/feater — /seo STEP 1.5)
for f in "$SEO" "$GEO"; do
has "$f" 'applier: hotfixer'
has "$f" 'applier: feater'
has "$f" ' files:'
has "$f" ' current:'
has "$f" ' expected:'
done
has "$SEO" 'applier: bash'
# 7) cross-agent escalation block — merged into SEO.md §11 by /seo STEP 2.
# The emit instruction lives in /seo's DISPATCH PROMPTS, not in the agent
# specs (geo-analyzer.md never mentions it; seo-analyzer.md only once,
# incidentally — NOT locked, it is prose). Lock the dispatcher side only.
has "skills/seo/SKILL.md" 'CROSS-AGENT NOTES TO'
# 8) trajectory block — mandatory in envelopes (/geo audit-end deliverables,
# /seo §1 merge)
has "$SEO" 'TRAJECTORY TO 17/20'
has "$GEO" 'TRAJECTORY TO 17/20'
printf 'seo-geo contract locks: %d pass, %d fail\n' "$pass" "$fail"
[ "$fail" -eq 0 ]
+61 -41
View File
@@ -8,7 +8,7 @@
# as symlinks inside skills/. This script moves those symlinks
# to/from skills-disabled/ so Claude Code stops/starts scanning them.
#
# MCP servers are toggled via `claude mcp add|remove` (not symlinks).
# A multi-skill pack (gstack, 21st) toggles all of its skills at once.
#
# Usage:
# toggle-external.sh list
@@ -20,7 +20,7 @@
# gstack — per-skill symlinks populated by gstack's own setup
# emil-design-eng — single symlink → skills-external/emil-design-eng
# darwin-skill — single symlink → ~/.agents/skills/darwin-skill
# magic — 21st-dev Magic MCP server (API key in .env)
# 21st — 21st.dev skill pack (needs the `21st` CLI + login)
#
# For fine-grained activation (only design skills, only qa skills, only
# audit skills, etc.) instead of all-or-nothing gstack toggling, use:
@@ -40,17 +40,17 @@ warn() { echo -e "${YELLOW}⚠${NC} $1"; }
err() { echo -e "${RED}✗${NC} $1"; }
# All non-plugin tools this script can toggle.
MANAGED_TOOLS=(gstack emil-design-eng darwin-skill magic)
MANAGED_TOOLS=(gstack emil-design-eng darwin-skill 21st)
# Load MAGIC_API_KEY (and any other secrets) from $REPO/.env if present.
# Called only by the magic branch — other tools don't need env vars.
load_env() {
if [ -z "${MAGIC_API_KEY:-}" ] && [ -f "$REPO/.env" ]; then
set -a
# shellcheck source=/dev/null
source "$REPO/.env"
set +a
fi
# Prints the skill names that belong to the "21st" pack. Source of truth:
# skills-external/21st-* — the `21st skills install` run in install-plugins.sh
# owns that list, so adding a skill upstream needs no edit here.
twentyfirst_skills() {
local d
for d in "$REPO"/skills-external/21st-*/; do
[ -f "${d}SKILL.md" ] || continue
basename "$d"
done
}
# Prints the names (directory basenames) that belong to "gstack".
@@ -84,13 +84,13 @@ status_tool() {
[ -d "$HOME/.agents/skills/$tool" ] || { echo "missing"; return; }
[ -e "$SKILLS_DIR/$tool" ] && echo "enabled" || echo "disabled"
;;
magic)
command -v claude >/dev/null || { echo "missing"; return; }
if claude mcp list 2>/dev/null | grep -q '^magic:'; then
echo "enabled"
else
echo "disabled"
fi
21st)
local installed=0
while read -r name; do
installed=1
[ -e "$SKILLS_DIR/$name" ] && { echo "enabled"; return; }
done < <(twentyfirst_skills)
[ "$installed" -eq 1 ] && echo "disabled" || echo "missing"
;;
*)
echo "unknown"; return 1 ;;
@@ -124,12 +124,20 @@ disable_tool() {
warn "$tool already disabled"
fi
;;
magic)
if [ "$(status_tool magic)" = "enabled" ]; then
claude mcp remove magic -s user >/dev/null
ok "magic disabled"
21st)
# Parked under the plain skill name — same convention as the other
# externals, so profile.sh's park/restore path stays interoperable.
local parked=0
while read -r name; do
[ -e "$SKILLS_DIR/$name" ] || continue
rm -rf "${DISABLED_DIR:?}/${name:?}"
mv "$SKILLS_DIR/$name" "$DISABLED_DIR/$name"
parked=$((parked + 1))
done < <(twentyfirst_skills)
if [ "$parked" -gt 0 ]; then
ok "21st disabled ($parked skills parked)"
else
warn "magic already disabled"
warn "21st already disabled"
fi
;;
*) err "Unknown tool: $tool"; return 1 ;;
@@ -177,25 +185,37 @@ enable_tool() {
return 1
fi
;;
magic)
load_env
if [ -z "${MAGIC_API_KEY:-}" ]; then
err "MAGIC_API_KEY not set — add it to ~/.claude/.env (template: .env.example)"
return 1
fi
if [ "$(status_tool magic)" = "enabled" ]; then
warn "magic already enabled"
21st)
local restored=0 linked=0
while read -r name; do
if [ -e "$DISABLED_DIR/$name" ]; then
rm -rf "${SKILLS_DIR:?}/${name:?}"
mv "$DISABLED_DIR/$name" "$SKILLS_DIR/$name"
restored=$((restored + 1))
elif [ -e "$SKILLS_DIR/$name" ]; then
: # already enabled
else
ln -sf "$REPO/skills-external/$name" "$SKILLS_DIR/$name"
linked=$((linked + 1))
fi
done < <(twentyfirst_skills)
if [ "$((restored + linked))" -eq 0 ]; then
if [ "$(status_tool 21st)" = "missing" ]; then
err "21st pack not installed in $REPO/skills-external — run: make plugin"
return 1
fi
warn "21st already enabled"
return 0
fi
# Reference, not value: Claude Code expands ${VAR} in mcpServers.env at
# launch (job7/BDR-026) — MAGIC_API_KEY itself never lands in
# ~/.claude.json. The check above still confirms the var IS set in
# ~/.claude/.env before wiring the reference, so a missing key fails
# here instead of silently at Claude Code startup.
claude mcp add magic --scope user \
--env 'API_KEY=${MAGIC_API_KEY}' \
-- npx -y @21st-dev/magic@latest
ok "magic enabled (user scope)"
ok "21st enabled ($((restored + linked)) skills: $restored restored, $linked linked)"
# The skills shell out to the CLI; without it (or without a session)
# they can only report failure. Warn, never block — the pack is still
# correctly wired and `make plugin` installs the CLI.
if ! command -v 21st >/dev/null 2>&1; then
warn "the \`21st\` CLI is not on PATH — install it: npm i -g @21st-dev/cli"
elif ! 21st whoami 2>/dev/null | grep -q '^Logged in as '; then
warn "not signed in to 21st — component retrieval and 21st AI need: 21st login"
fi
;;
*) err "Unknown tool: $tool"; return 1 ;;
esac
+54 -12
View File
@@ -13,8 +13,40 @@ Inputs the caller must have ready:
pre-dev SHA, or the working-tree diff before commit).
- `TEST`: the project test command, if known.
Nominal path is cheap: one verifier dispatch + one security dispatch, done.
The loop only costs more when it actually loops.
Nominal path is cheap — a free floor run, then
one verifier dispatch + one security dispatch, done. The loop only costs
more when it actually loops.
## GATE 0 — DETERMINISTIC FLOOR (no dispatch, no model)
Before spending a verifier dispatch, execute the oracles the contract itself
declares:
```bash
bash ~/.claude/lib/gates.sh run "$CONTRACT"
```
It runs every `CHECK:` fail-closed (MET requires exit 0 AND the `EXPECT:`
marker) and writes the outcome back over each `EVIDENCE:` line. Parse its
single `GATES — VERDICT:` line:
- `MET` → floor green, go to GATE 1. An all-manual contract lands here too
(`RUNNABLE: 0 of n`) and passes straight through.
- `UNMET(n)` → hand the dev the CONTRACT path + the `NOT-MET` rows verbatim,
nothing else; re-run GATE 0. **No verifier is dispatched** — a red build or
a red suite is not a judgement call, and paying an LLM to discover it is
waste. **Max 3 floor iterations** → STOP + human escalation with the rows.
- `ABANDONED(n)` → floor green but a handoff stands. Continue to GATE 1; the
verifier surfaces it and its `ABANDONED(n)` verdict routes to the human
gate.
- `ERROR(n)` → the ledger is malformed (partial oracle, duplicate id,
unindented attribute, runnable criterion with no `EVIDENCE:` line). The
contract is the ORCHESTRATOR's own artifact — fix it here in the main loop,
never dispatch a dev for it.
Floor iterations are counted separately from GATE 1's: a cheap loop here does
not eat the conformity budget. GATE 0 also runs unchanged after every
security fix round, before re-verifying the request.
## GATE 1 — REQUEST CONFORMITY (fresh verifier)
@@ -29,9 +61,13 @@ Parse its single `VERIFY — VERDICT:` line:
- `ECARTS(n)` → hand the dev the CONTRACT path + the exact `CRITERIA` gap
lines (NOT-MET / out-of-scope), nothing else. Inline dev fixes in place;
a dispatched dev is re-dispatched FRESH with those inputs only. Then
re-dispatch a FRESH verifier. Repeat. **Max 3 conformity iterations** →
STOP + human escalation with the CRITERIA table (the contract-vs-realized
diff).
re-run GATE 0 and re-dispatch a FRESH verifier. Repeat.
**Max 3 conformity iterations** → STOP + human escalation with the
CRITERIA table (the contract-vs-realized diff).
- `ABANDONED(n)` → direct human gate, never a dev loop (a dev cannot close
what was proven impossible). The human lifts the abandonment or accepts
the partial delivery; either way the run is never reported as fully
complete, and the abandonment is named in the final report.
- Remaining `UNVERIFIABLE` while all else MET → direct human gate (a dev
cannot fix unverifiability); do not spend a loop on it.
- Out-of-scope files: a dev justification is accepted ONLY through the human
@@ -53,10 +89,11 @@ Parse its single `SECURITY — VERDICT:` line:
- `PASS` → done, proceed to commit.
- `BLOCK(n)` → hand the dev the `BLOCKING` list + the CONTRACT path (inline
fix, or FRESH executor re-dispatch). Then **re-verify the REQUEST first** (GATE 1, fresh
verifier) — a security fix can drift the behavior — **then re-run GATE 2**
(fresh auditor), in that order. **Max 3 security iterations** → STOP +
human escalation with the BLOCKING table.
fix, or FRESH executor re-dispatch). Then re-run GATE 0, then
**re-verify the REQUEST first** (GATE 1, fresh verifier) — a security fix
can drift the behavior — **then re-run GATE 2** (fresh auditor), in that
order. **Max 3 security iterations** → STOP + human escalation with the
BLOCKING table.
- `DEGRADED` (semgrep absent) → does NOT block on the tool's absence; surface
the checklist result + recommend `make plugin`. A DEGRADED run that still
BLOCKs (grep-caught secret/injection) blocks like any other.
@@ -65,6 +102,11 @@ Parse its single `SECURITY — VERDICT:` line:
## Order invariant
REQUEST conformity is always re-checked BEFORE security on any re-loop — a
security fix that breaks the feature must not slip through because only the
security gate re-ran. Never the reverse order.
Every re-loop replays the gates in order: **GATE 0 → GATE 1 → GATE 2**,
never a subset and never reversed.
The floor runs first because it is free, and because a red build makes the
verifier's verdict meaningless. REQUEST conformity is
always re-checked BEFORE security on any re-loop — a security fix that breaks
the feature must not slip through because only the security gate re-ran.
Never the reverse order.
+4 -3
View File
@@ -71,7 +71,10 @@ if [ -d "$GSTACK_SRC/browse/dist" ]; then
fi
fi
EXTERNAL_SKILLS=(emil-design-eng frontend-design design-motion-principles impeccable)
# impeccable is NOT here: its installer writes the skill straight into
# skills/ (and its agents into agents/) at --scope=global, so there is no
# skills-external/ copy to symlink. See install-plugins.sh Step 8d.
EXTERNAL_SKILLS=(emil-design-eng frontend-design design-motion-principles)
for _ext_skill in "${EXTERNAL_SKILLS[@]}"; do
if [ -d "$REPO/skills-external/$_ext_skill" ]; then
if [ -L "$CLAUDE/skills/$_ext_skill" ] && [ "$(readlink "$CLAUDE/skills/$_ext_skill")" = "$REPO/skills-external/$_ext_skill" ]; then
@@ -117,8 +120,6 @@ link_env() {
echo " cp \"$REPO/.env.example\" \"$home_env\" && \"\${EDITOR:-nano}\" \"$home_env\""
return
fi
grep -qE '^[[:space:]]*(export[[:space:]]+)?MAGIC_API_KEY=.' "$home_env" 2>/dev/null \
|| echo "⚠️ $home_env has no MAGIC_API_KEY line — magic won't enable until added."
if [ -L "$repo_env" ]; then
[ "$(readlink "$repo_env")" = "$home_env" ] && return
ln -sf "$home_env" "$repo_env"; CHANGED=$((CHANGED + 1))
+8 -3
View File
@@ -20,6 +20,11 @@
"version": "latest",
"note": "Context7 CLI — doc lookup for fast-evolving libs. Standalone CLI, not an MCP server. Install: npm install -g ctx7. Standalone: ctx7 docs /vercel/next.js \"middleware\"."
},
"21st": {
"source": "npm:@21st-dev/cli",
"version": "latest",
"note": "21st.dev CLI (bin `21st`) — supersedes the @21st-dev/magic MCP server (2026-09-22). Standalone CLI + a pack of 7 skills, no MCP, no API key: auth is `21st login` (browser token in ~/.config/21st). Install: npm install -g @21st-dev/cli. The skill pack is staged-installed into skills-external/21st-* by install-plugins.sh Step 8.7 — `21st skills install` refuses to write through the ~/.claude/skills symlink."
},
"graphifyy": {
"source": "pypi:graphifyy",
"version": "latest",
@@ -36,11 +41,11 @@
"source": "https://github.com/emilkowalski/skill",
"path": "skills/emil-design-eng/SKILL.md",
"managed_by": "curl",
"note": "Emil Kowalski's design engineering skill — UI polish, animations, component craft. Downloaded to skills-external/emil-design-eng/, symlinked by link.sh."
"note": "Emil Kowalski's design engineering skill — UI polish, animations, component craft. Machine-owned: curl'd to skills-external/emil-design-eng/ (gitignored, re-fetched by update-all.sh), symlinked by link.sh."
},
"impeccable": {
"source": "npm:impeccable",
"version": "3.2.0",
"note": "Design anti-pattern detector (45 deterministic rules, CLI 'impeccable detect', exit 0/2) + /impeccable skill (23 verbs) by pbakaus. Pin = CLI version; the skill dist has its own release track fetched by 'skills install'. Pinned for audit reproducibility (LRN-077 class: a rules update silently changes audit output). Requires Node >= 24 — install step skips gracefully below that. Machine-owned: synced to skills-external/impeccable/ (gitignored), symlinked by link.sh."
"version": "4.1.0",
"note": "Design anti-pattern detector (45 deterministic rules, CLI 'impeccable detect', exit 0/2) + /impeccable skill (23 verbs) + 4 impeccable-* subagents, by pbakaus. Pin = CLI version ONLY: the skill dist and the engine binary have their own release tracks, fetched by 'skills install' at install time, so this pin does not freeze audit output the way a semgrep pin does. It still gates the CLI deliberately (LRN-077 class). BEWARE: the pin rots — the CLI downloads its skill dist at install time and an older release's artifact disappears upstream (3.2.0 -> 'Download failed: invalid zip data', 2026-09-22, which left 'make plugin' printing a run-it-yourself warning). install-plugins.sh Step 8d and update-all.sh therefore fall back to @latest on a pin failure and warn to bump this version. Requires Node >= 24. Installed at --scope=global: lands in ~/.claude/skills/impeccable + ~/.claude/agents/impeccable-*.md, both symlinks into this repo, both gitignored."
}
}
+30
View File
@@ -0,0 +1,30 @@
---
paths: ["**/*.html", "**/*.astro", "**/*.css", "**/*.scss", "**/*.tsx", "**/*.jsx", "**/*.vue", "**/*.svelte"]
---
# Web building — no default reflexes + done checklist
## Avoid unless the user asks for them
- Purple gradient, purple/black, neon, washed-out pastels, rainbow.
- Drop shadow on everything; the same border-radius on every element.
- Bento grid, dot grid, glowing background orbs, decorative color strip.
- Sparkle icons, animated arrows, emojis as icons, decorative fake
terminal window.
- Hover animation on every element; scroll-reveal animations everywhere.
- Three aligned feature cards, three pricing tiers, checkmark bullets.
- Vague hero title ("unleash your potential"): state what the product does.
- Fake testimonials, fake visitor or client counters, invented numbers.
Never, in any context.
- Inter, Geist or Space Grotesk as the default font: propose an
alternative and justify it. Existing brand identities keep their fonts.
## Before declaring a public site done
Check: custom 404 · call to action in the first viewport · per-page
title + description · share/OG image · favicons · robots.txt · sitemap ·
alt text on images · layout tested at 375 px · loading states · form
error messages · confirmation page · real legal mentions · cookie banner
with a working refuse option · audience measurement · contact address ·
compressed images.
Report the missing items to the user instead of inventing them. Internal
tools and dashboards: only the relevant items apply. Deep audits stay
with /seo, /harden, /web-validate.
+21
View File
@@ -0,0 +1,21 @@
---
paths: ["**/*.ts", "**/*.tsx", "**/*.js", "**/*.jsx", "**/*.vue", "**/*.svelte", "**/*.astro", "**/*.php", "**/*.py"]
---
# Web app security — specifics
Extends the global Security section (input validation, parameterized
queries, secrets in env vars, AuthN/AuthZ, fail closed). If a request
breaks one of these rules, say so instead of doing it.
- No API key in code shipped to the browser. Env vars, server-side only.
- The service/admin key never reaches the client: publishable key only.
- Row Level Security enabled on every table (Supabase/Postgres and kin).
- Authentication verified server-side, never only in the browser.
- No IDOR: changing an id in a URL must never expose another user's
data. Authorize object access on every request.
- Passwords hashed (bcrypt/argon2). Session cookies httpOnly + secure
+ sameSite.
- API responses return only the fields the client needs.
- Login rate limiting, upload restrictions (type/size), forced HTTPS,
security headers.
+29
View File
@@ -0,0 +1,29 @@
# Writing style — user-facing prose
Scope: prose written FOR the user: answers, docs, reports, deliverables,
site copy. Does NOT override memory registries (caveman format), code
comments (code style rules), or structured skill/report templates.
Banned:
- Em-dash. Use a comma, a colon, or a period.
- The "it's not X, it's Y" / "ce n'est pas X, c'est Y" frame.
- Emojis, unless explicitly requested.
- Decorative bold. Bold marks a key term, not one word per sentence.
- Rule-of-three enumerations by reflex. Two often suffice, four sometimes.
- Hedging chains ("il est possible que", "could potentially", "in some
cases"). Assert, or say you don't know.
- Restating the user's question before answering it.
- Slop vocabulary, buzzword sense: delve, explorons, plongeons, "il
convient de noter" / "it's worth noting", figurative paysage/landscape,
robuste/robust, transformer/transform. Technical senses stay allowed
(robustness as a review lens, a math transform).
Do:
- Vary sentence and paragraph length. A short sentence after a long one.
- Write like speech. A sentence you cannot say aloud in one breath gets
cut in two.
- A paragraph over a bullet list when prose carries it.
Self-check before handing over a deliverable (text, site, feature):
reread against these rules (plus the web rules for a site) and tell the
user what you corrected to comply, or that nothing needed correcting.
+81 -22
View File
@@ -110,12 +110,8 @@
"Bash(chmod -R 777 *)",
"Bash(ssh *)",
"Bash(scp *)",
"Bash(rsync *)",
"Bash(nc *)",
"Bash(netcat *)",
"Bash(kill -9 *)",
"Bash(killall *)",
"Bash(pkill *)",
"Bash(crontab *)",
"Bash(systemctl *)",
"Bash(service *)",
@@ -182,6 +178,16 @@
"Bash(more .env.*)",
"Bash(grep * .env)",
"Bash(grep * .env.*)",
"Bash(sed * .env*)",
"Bash(awk * .env*)",
"Bash(cut * .env*)",
"Bash(tr * .env*)",
"Bash(sort * .env*)",
"Bash(uniq * .env*)",
"Bash(diff * .env*)",
"Bash(od * .env*)",
"Bash(xxd * .env*)",
"Bash(strings * .env*)",
"Bash(env)",
"Bash(printenv)",
"Bash(printenv *)",
@@ -218,15 +224,8 @@
"Bash(curl * | sh)",
"Bash(wget * | sh)",
"Bash(mkfifo *)",
"Bash(node -e *)",
"Bash(python3 -c *)",
"Bash(python -c *)",
"Bash(git push *)",
"Bash(git push)",
"Bash(docker run *)",
"Bash(docker exec *)",
"Bash(docker-compose up*)",
"Bash(docker compose up*)",
"Bash(brew install *)",
"Bash(apt install *)",
"Bash(apt-get install *)",
@@ -234,23 +233,15 @@
"Bash(pacman -S *)",
"WebSearch",
"WebFetch",
"Bash(xargs *)",
"Bash(sed *)",
"Bash(cp *)",
"Bash(mv *)",
"Bash(git stash pop*)",
"Bash(git stash drop*)",
"Bash(git stash clear)",
"mcp__magic__21st_magic_component_builder",
"mcp__magic__21st_magic_component_refiner",
"mcp__magic__21st_magic_component_inspiration",
"mcp__magic__logo_search"
"Bash(git stash clear)"
],
"defaultMode": "auto",
"disableBypassPermissionsMode": "disable",
"additionalDirectories": []
},
"model": "claude-fable-5[1m]",
"model": "claude-fable-5-1[1m]",
"hooks": {
"SessionStart": [
{
@@ -273,6 +264,31 @@
]
}
],
"Notification": [
{
"matcher": "permission_prompt|idle_prompt|agent_needs_input|elicitation_dialog|elicitation_url_dialog",
"hooks": [
{
"type": "command",
"command": "bash ~/.claude/hooks/notify-attention.sh",
"timeout": 5,
"statusMessage": "Ringing terminal bell..."
}
]
}
],
"Stop": [
{
"hooks": [
{
"type": "command",
"command": "bash ~/.claude/hooks/notify-attention.sh",
"timeout": 5,
"statusMessage": "Ringing terminal bell..."
}
]
}
],
"UserPromptSubmit": [
{
"hooks": [
@@ -333,5 +349,48 @@
"effortLevel": "xhigh",
"remoteControlAtStartup": true,
"inputNeededNotifEnabled": true,
"skipAutoPermissionPrompt": true
"skipAutoPermissionPrompt": true,
"autoMode": {
"allow": [
"$defaults",
"Local dev containers: `docker exec`, `docker run`, `docker compose up`/`exec`/`logs`/`ps` against a container running on this workstation whose name does not carry `prod` or `production` (a local Supabase or Postgres such as `supabase_db_*`) is routine development, not a remote shell into a shared host. Running a SQL file or script that lives in the repo inside it (`psql -f`, migrations, verify scripts) and piping the output through `tail` or `grep` passes. Remote Shell Writes, Production Reads and Sensitive Remote Exec apply only to hosts named as sensitive in Environment or carrying `prod`. A literal `DROP`, `TRUNCATE` or `DELETE` without a predicate typed on the command line stays under Mass Delete.",
"Project-local node: `node <file>`, `npm run`, `pnpm` or `yarn` scripts, and `npx` or `pnpm exec` of a package declared in the project's manifest or lockfile, with effects inside the current working directory, pass like `awk` or `echo`. `node -e` that computes or edits inside the working directory passes; the soft block on inline interpreters that delete or write outside it still applies."
],
"soft_deny": [
"$defaults",
"Scope of intent: an instruction clears a SOFT BLOCK for the current turn only. An approval given in an earlier turn is not an approval now, and the same action repeated in a later turn has to be asked for again.",
"Writing outside the current working directory: `cp`, `mv`, `sed -i`, `rsync`, `tee`, or a shell redirection whose destination resolves outside the cwd. Several sibling projects live under `~/Documents/`, so the realistic failure is writing into the wrong one, where git recovers nothing. Clear only when the user named the destination in this turn.",
"`rsync` invoked with `--delete`. It removes files at the destination that are absent from the source, with no undo. Clear only against a destination the user named in this turn.",
"Sending SIGKILL (`kill -9`) or killing processes by name (`killall`, `pkill`). These reach processes outside this session, including the user's editors, shells, dtach sessions and background jobs, and the target is chosen by a pattern, so a typo kills the wrong thing. Clear only when the user named the process in this turn.",
"Editing more than one file in place in a single command: `sed -i` or `perl -pi` over a glob, or a loop over `git ls-files`. The damage is not loss, since git recovers it, but a diff spanning hundreds of files that nobody reads before committing. `sed -i` on a single named file passes. Clear only when the user asked for the sweep.",
"Moving or renaming a directory inside the repo (`mv src/api src/api_old`, or any `mv` of a tree). It breaks imports and paths silently, and the breakage surfaces far from the command. Clear only when the user asked for that move.",
"An inline interpreter or `xargs` that deletes, or that writes outside the current working directory: `python3 -c`, `python -c` or `node -e` calling `rmtree`, `remove`, `unlink` or `truncate`; `xargs` feeding `rm`, `mv` or `dd`. `find ... | xargs rm` is the case that matters, since it routes around the `find * -exec rm` deny rule. Reading, computing, and editing a file inside the working directory pass untouched.",
"Docker data destruction on this workstation: `docker rm -f`, `docker volume rm` or `prune`, `docker system prune`, `docker compose down -v` (drops named volumes, which hold local database data with no undo), and `docker run` with `--privileged` or a bind mount outside the current working directory. Clear only when the user named the container or volume in this turn.",
"Undeclared node packages: `npx <pkg>`, `pnpm dlx` or `yarn dlx` of a package absent from the manifest and lockfile runs code fetched at call time; `npm install <name>` or `pnpm add <name>` adds a dependency the house rule requires naming first. Clear only when the user named the package in this turn.",
"Publishing to the 21st.dev public catalog: `21st publish`, `publish-theme`, `publish-template`, `publish-gradient`, `publish-ascii`, `submit` and `resubmit` push a component, theme or template from this machine onto a public listing under the user's account; `21st edit`, `delete`, `withdraw`, `remove-from-catalog` and `21st profile set|upload` change or remove what is already published there. Retrieval and generation (`search`, `logo`, `get`, `add`, `generate`, `iterate`) are ordinary design work and pass. Clear only when the user asked to publish or change that specific item in this turn."
],
"hard_deny": [
"$defaults",
"Secret exfiltration: reading a `.env`, an SSH key, a token or a credentials file, and then sending what it contained outside this machine. That covers an HTTP request body or header, a paste or gist service, a file upload, an email, and a commit or push to any repository. The static deny rules match the reading tools by pattern and cannot see the second half of the chain; read and send are separate steps and may sit several turns apart. No instruction clears this.",
"Production deployment: running a project's deploy script (`bin/deploy.sh` and its equivalents), any lftp, FTP, SFTP or rsync push to a hosting provider, and any action against a target whose name carries `prod` or `production` as a whole word or name segment. The user deploys by hand, out of session. A green test suite, a finished feature, or a plan step that reads \"deploy\" is not an instruction to deploy. No in-session instruction clears this.",
"Disarming the guardrails: removing or weakening an entry in the `deny`, `soft_deny` or `hard_deny` lists of any settings.json, passing `--no-verify` to git, deleting or chmod-ing `.githooks/pre-commit`, setting `disableAllHooks`, or switching `permissions.defaultMode` to `bypassPermissions`. Adding a restriction is fine; removing one is not. When a task is blocked by a guardrail, say so and stop, rather than widening the guardrail to get through. The user maintains these files by hand. No instruction clears this."
],
"environment": [
"$defaults",
"### Machine-specific (refines any \"None configured\" default above)",
"**Primary use of Claude Code**: software development on a personal Linux workstation. Single developer, no organization.",
"**Source control**: self-hosted Gitea at `git.bchanot.fr` (SSH on port 49220). Some checkouts under `/home/bchanot/Documents/` have no remote at all and are local-only.",
"**Repository visibility**: private. The Gitea instance is self-hosted and not publicly indexed, and a checkout with no remote is local-only. Treat every repo here as private unless its remote points at a public host.",
"**Trusted repo**: the checkout Claude Code is currently working in, under `/home/bchanot/Documents/`. No single repo is privileged over the others — read the current one, do not assume a previous session's project.",
"**Trusted internal domains**: `git.bchanot.fr` (self-hosted Gitea). It is the only internal service.",
"**Default / protected branches**: gitflow. `main` (prod) and `develop` (integration) are protected: a per-repo pre-commit hook refuses code commits on either (exempting `.claude/**` and merges) and Gitea enforces branch protection on both. Work lands on `feature/*`, `bugfix/*`, `chore/*`, `release/*`, `hotfix/*`.",
"**Secrets management**: `~/.claude/.env` is the single source of truth and lives outside every git tree; repos reach it through a gitignored symlink. Only `.env.example`, holding placeholders, is ever tracked. A real secret inside a repo is a defect, not a configuration.",
"**Internal sharing / snippet hosting**: none. Public paste, gist and pastebin services are outside the trust boundary.",
"**CI/CD deploy targets**: no CI system. Deploys run out of band from a per-project runbook, typically lftp/FTP to OVH mutualised hosting for web projects. Nothing deploys automatically on a push or a merge.",
"**Internal package registry**: none. Public npm and PyPI.",
"**Host containment**: an ordinary developer workstation with open internet and no sandbox. Nothing is contained by the environment itself.",
"**Sensitive remote targets**: any namespace, host, database or container whose name carries `prod` or `production` as a whole word or name segment.",
"**Sensitive data locations & audiences**: per-project `.env` files (gitignored) hold database, deploy and API credentials; some web projects store customer-submitted form data under a retention policy. Both are personal or client data — never send either to an external service."
]
}
}
-679
View File
@@ -1,679 +0,0 @@
---
name: emil-design-eng
description: This skill encodes Emil Kowalski's philosophy on UI polish, component design, animation decisions, and the invisible details that make software feel great.
---
# Design Engineering
## Initial Response
When this skill is first invoked without a specific question, respond only with:
> I'm ready to help you build interfaces that feel right, my knowledge comes from Emil Kowalski's design engineering philosophy. If you want to dive even deeper, check out Emil’s course: [animations.dev](https://animations.dev/).
Do not provide any other information until the user asks a question.
You are a design engineer with the craft sensibility. You build interfaces where every detail compounds into something that feels right. You understand that in a world where everyone's software is good enough, taste is the differentiator.
## Core Philosophy
### Taste is trained, not innate
Good taste is not personal preference. It is a trained instinct: the ability to see beyond the obvious and recognize what elevates. You develop it by surrounding yourself with great work, thinking deeply about why something feels good, and practicing relentlessly.
When building UI, don't just make it work. Study why the best interfaces feel the way they do. Reverse engineer animations. Inspect interactions. Be curious.
### Unseen details compound
Most details users never consciously notice. That is the point. When a feature functions exactly as someone assumes it should, they proceed without giving it a second thought. That is the goal.
> "All those unseen details combine to produce something that's just stunning, like a thousand barely audible voices all singing in tune." - Paul Graham
Every decision below exists because the aggregate of invisible correctness creates interfaces people love without knowing why.
### Beauty is leverage
People select tools based on the overall experience, not just functionality. Good defaults and good animations are real differentiators. Beauty is underutilized in software. Use it as leverage to stand out.
## Review Format (Required)
When reviewing UI code, you MUST use a markdown table with Before/After columns. Do NOT use a list with "Before:" and "After:" on separate lines. Always output an actual markdown table like this:
| Before | After | Why |
| --- | --- | --- |
| `transition: all 300ms` | `transition: transform 200ms ease-out` | Specify exact properties; avoid `all` |
| `transform: scale(0)` | `transform: scale(0.95); opacity: 0` | Nothing in the real world appears from nothing |
| `ease-in` on dropdown | `ease-out` with custom curve | `ease-in` feels sluggish; `ease-out` gives instant feedback |
| No `:active` state on button | `transform: scale(0.97)` on `:active` | Buttons must feel responsive to press |
| `transform-origin: center` on popover | `transform-origin: var(--radix-popover-content-transform-origin)` | Popovers should scale from their trigger (not modals — modals stay centered) |
Wrong format (never do this):
```
Before: transition: all 300ms
After: transition: transform 200ms ease-out
────────────────────────────
Before: scale(0)
After: scale(0.95)
```
Correct format: A single markdown table with | Before | After | Why | columns, one row per issue found. The "Why" column briefly explains the reasoning.
## The Animation Decision Framework
Before writing any animation code, answer these questions in order:
### 1. Should this animate at all?
**Ask:** How often will users see this animation?
| Frequency | Decision |
| ----------------------------------------------------------- | ---------------------------- |
| 100+ times/day (keyboard shortcuts, command palette toggle) | No animation. Ever. |
| Tens of times/day (hover effects, list navigation) | Remove or drastically reduce |
| Occasional (modals, drawers, toasts) | Standard animation |
| Rare/first-time (onboarding, feedback forms, celebrations) | Can add delight |
**Never animate keyboard-initiated actions.** These actions are repeated hundreds of times daily. Animation makes them feel slow, delayed, and disconnected from the user's actions.
Raycast has no open/close animation. That is the optimal experience for something used hundreds of times a day.
### 2. What is the purpose?
Every animation must have a clear answer to "why does this animate?"
Valid purposes:
- **Spatial consistency**: toast enters and exits from the same direction, making swipe-to-dismiss feel intuitive
- **State indication**: a morphing feedback button shows the state change
- **Explanation**: a marketing animation that shows how a feature works
- **Feedback**: a button scales down on press, confirming the interface heard the user
- **Preventing jarring changes**: elements appearing or disappearing without transition feel broken
If the purpose is just "it looks cool" and the user will see it often, don't animate.
### 3. What easing should it use?
Is the element entering or exiting?
Yes → ease-out (starts fast, feels responsive)
No →
Is it moving/morphing on screen?
Yes → ease-in-out (natural acceleration/deceleration)
Is it a hover/color change?
Yes → ease
Is it constant motion (marquee, progress bar)?
Yes → linear
Default → ease-out
**Critical: use custom easing curves.** The built-in CSS easings are too weak. They lack the punch that makes animations feel intentional.
```css
/* Strong ease-out for UI interactions */
--ease-out: cubic-bezier(0.23, 1, 0.32, 1);
/* Strong ease-in-out for on-screen movement */
--ease-in-out: cubic-bezier(0.77, 0, 0.175, 1);
/* iOS-like drawer curve (from Ionic Framework) */
--ease-drawer: cubic-bezier(0.32, 0.72, 0, 1);
```
**Never use ease-in for UI animations.** It starts slow, which makes the interface feel sluggish and unresponsive. A dropdown with `ease-in` at 300ms _feels_ slower than `ease-out` at the same 300ms, because ease-in delays the initial movement — the exact moment the user is watching most closely.
**Easing curve resources:** Don't create curves from scratch. Use [easing.dev](https://easing.dev/) or [easings.co](https://easings.co/) to find stronger custom variants of standard easings.
### 4. How fast should it be?
| Element | Duration |
| ------------------------ | ------------- |
| Button press feedback | 100-160ms |
| Tooltips, small popovers | 125-200ms |
| Dropdowns, selects | 150-250ms |
| Modals, drawers | 200-500ms |
| Marketing/explanatory | Can be longer |
**Rule: UI animations should stay under 300ms.** A 180ms dropdown feels more responsive than a 400ms one. A faster-spinning spinner makes the app feel like it loads faster, even when the load time is identical.
### Perceived performance
Speed in animation is not just about feeling snappy — it directly affects how users perceive your app's performance:
- A **fast-spinning spinner** makes loading feel faster (same load time, different perception)
- A **180ms select** animation feels more responsive than a **400ms** one
- **Instant tooltips** after the first one is open (skip delay + skip animation) make the whole toolbar feel faster
The perception of speed matters as much as actual speed. Easing amplifies this: `ease-out` at 200ms _feels_ faster than `ease-in` at 200ms because the user sees immediate movement.
## Spring Animations
Springs feel more natural than duration-based animations because they simulate real physics. They don't have fixed durations — they settle based on physical parameters.
### When to use springs
- Drag interactions with momentum
- Elements that should feel "alive" (like Apple's Dynamic Island)
- Gestures that can be interrupted mid-animation
- Decorative mouse-tracking interactions
### Spring-based mouse interactions
Tying visual changes directly to mouse position feels artificial because it lacks motion. Use `useSpring` from Motion (formerly Framer Motion) to interpolate value changes with spring-like behavior instead of updating immediately.
```jsx
import { useSpring } from 'framer-motion';
// Without spring: feels artificial, instant
const rotation = mouseX * 0.1;
// With spring: feels natural, has momentum
const springRotation = useSpring(mouseX * 0.1, {
stiffness: 100,
damping: 10,
});
```
This works because the animation is **decorative** — it doesn't serve a function. If this were a functional graph in a banking app, no animation would be better. Know when decoration helps and when it hinders.
### Spring configuration
**Apple's approach (recommended — easier to reason about):**
```js
{ type: "spring", duration: 0.5, bounce: 0.2 }
```
**Traditional physics (more control):**
```js
{ type: "spring", mass: 1, stiffness: 100, damping: 10 }
```
Keep bounce subtle (0.1-0.3) when used. Avoid bounce in most UI contexts. Use it for drag-to-dismiss and playful interactions.
### Interruptibility advantage
Springs maintain velocity when interrupted — CSS animations and keyframes restart from zero. This makes springs ideal for gestures users might change mid-motion. When you click an expanded item and quickly press Escape, a spring-based animation smoothly reverses from its current position.
## Component Building Principles
### Buttons must feel responsive
Add `transform: scale(0.97)` on `:active`. This gives instant feedback, making the UI feel like it is truly listening to the user.
```css
.button {
transition: transform 160ms ease-out;
}
.button:active {
transform: scale(0.97);
}
```
This applies to any pressable element. The scale should be subtle (0.95-0.98).
### Never animate from scale(0)
Nothing in the real world disappears and reappears completely. Elements animating from `scale(0)` look like they come out of nowhere.
Start from `scale(0.9)` or higher, combined with opacity. Even a barely-visible initial scale makes the entrance feel more natural, like a balloon that has a visible shape even when deflated.
```css
/* Bad */
.entering {
transform: scale(0);
}
/* Good */
.entering {
transform: scale(0.95);
opacity: 0;
}
```
### Make popovers origin-aware
Popovers should scale in from their trigger, not from center. The default `transform-origin: center` is wrong for almost every popover. **Exception: modals.** Modals should keep `transform-origin: center` because they are not anchored to a specific trigger — they appear centered in the viewport.
```css
/* Radix UI */
.popover {
transform-origin: var(--radix-popover-content-transform-origin);
}
/* Base UI */
.popover {
transform-origin: var(--transform-origin);
}
```
Whether the user notices the difference individually does not matter. In the aggregate, unseen details become visible. They compound.
### Tooltips: skip delay on subsequent hovers
Tooltips should delay before appearing to prevent accidental activation. But once one tooltip is open, hovering over adjacent tooltips should open them instantly with no animation. This feels faster without defeating the purpose of the initial delay.
```css
.tooltip {
transition: transform 125ms ease-out, opacity 125ms ease-out;
transform-origin: var(--transform-origin);
}
.tooltip[data-starting-style],
.tooltip[data-ending-style] {
opacity: 0;
transform: scale(0.97);
}
/* Skip animation on subsequent tooltips */
.tooltip[data-instant] {
transition-duration: 0ms;
}
```
### Use CSS transitions over keyframes for interruptible UI
CSS transitions can be interrupted and retargeted mid-animation. Keyframes restart from zero. For any interaction that can be triggered rapidly (adding toasts, toggling states), transitions produce smoother results.
```css
/* Interruptible - good for UI */
.toast {
transition: transform 400ms ease;
}
/* Not interruptible - avoid for dynamic UI */
@keyframes slideIn {
from {
transform: translateY(100%);
}
to {
transform: translateY(0);
}
}
```
### Use blur to mask imperfect transitions
When a crossfade between two states feels off despite trying different easings and durations, add subtle `filter: blur(2px)` during the transition.
**Why blur works:** Without blur, you see two distinct objects during a crossfade — the old state and the new state overlapping. This looks unnatural. Blur bridges the visual gap by blending the two states together, tricking the eye into perceiving a single smooth transformation instead of two objects swapping.
Combine blur with scale-on-press (`scale(0.97)`) for a polished button state transition:
```css
.button {
transition: transform 160ms ease-out;
}
.button:active {
transform: scale(0.97);
}
.button-content {
transition: filter 200ms ease, opacity 200ms ease;
}
.button-content.transitioning {
filter: blur(2px);
opacity: 0.7;
}
```
Keep blur under 20px. Heavy blur is expensive, especially in Safari.
### Animate enter states with @starting-style
The modern CSS way to animate element entry without JavaScript:
```css
.toast {
opacity: 1;
transform: translateY(0);
transition: opacity 400ms ease, transform 400ms ease;
@starting-style {
opacity: 0;
transform: translateY(100%);
}
}
```
This replaces the common React pattern of using `useEffect` to set `mounted: true` after initial render. Use `@starting-style` when browser support allows; fall back to the `data-mounted` attribute pattern otherwise.
```jsx
// Legacy pattern (still works everywhere)
useEffect(() => {
setMounted(true);
}, []);
// <div data-mounted={mounted}>
```
## CSS Transform Mastery
### translateY with percentages
Percentage values in `translate()` are relative to the element's own size. Use `translateY(100%)` to move an element by its own height, regardless of actual dimensions. This is how Sonner positions toasts and how Vaul hides the drawer before animating in.
```css
/* Works regardless of drawer height */
.drawer-hidden {
transform: translateY(100%);
}
/* Works regardless of toast height */
.toast-enter {
transform: translateY(-100%);
}
```
Prefer percentages over hardcoded pixel values. They are less error-prone and adapt to content.
### scale() scales children too
Unlike `width`/`height`, `scale()` also scales an element's children. When scaling a button on press, the font size, icons, and content scale proportionally. This is a feature, not a bug.
### 3D transforms for depth
`rotateX()`, `rotateY()` with `transform-style: preserve-3d` create real 3D effects in CSS. Orbiting animations, coin flips, and depth effects are all possible without JavaScript.
```css
.wrapper {
transform-style: preserve-3d;
}
@keyframes orbit {
from {
transform: translate(-50%, -50%) rotateY(0deg) translateZ(72px) rotateY(360deg);
}
to {
transform: translate(-50%, -50%) rotateY(360deg) translateZ(72px) rotateY(0deg);
}
}
```
### transform-origin
Every element has an anchor point from which transforms execute. The default is center. Set it to match where the trigger lives for origin-aware interactions.
## clip-path for Animation
`clip-path` is not just for shapes. It is one of the most powerful animation tools in CSS.
### The inset shape
`clip-path: inset(top right bottom left)` defines a rectangular clipping region. Each value "eats" into the element from that side.
```css
/* Fully hidden from right */
.hidden {
clip-path: inset(0 100% 0 0);
}
/* Fully visible */
.visible {
clip-path: inset(0 0 0 0);
}
/* Reveal from left to right */
.overlay {
clip-path: inset(0 100% 0 0);
transition: clip-path 200ms ease-out;
}
.button:active .overlay {
clip-path: inset(0 0 0 0);
transition: clip-path 2s linear;
}
```
### Tabs with perfect color transitions
Duplicate the tab list. Style the copy as "active" (different background, different text color). Clip the copy so only the active tab is visible. Animate the clip on tab change. This creates a seamless color transition that timing individual color transitions can never achieve.
### Hold-to-delete pattern
Use `clip-path: inset(0 100% 0 0)` on a colored overlay. On `:active`, transition to `inset(0 0 0 0)` over 2s with linear timing. On release, snap back with 200ms ease-out. Add `scale(0.97)` on the button for press feedback.
### Image reveals on scroll
Start with `clip-path: inset(0 0 100% 0)` (hidden from bottom). Animate to `inset(0 0 0 0)` when the element enters the viewport. Use `IntersectionObserver` or Framer Motion's `useInView` with `{ once: true, margin: "-100px" }`.
### Comparison sliders
Overlay two images. Clip the top one with `clip-path: inset(0 50% 0 0)`. Adjust the right inset value based on drag position. No extra DOM elements needed, fully hardware-accelerated.
## Gesture and Drag Interactions
### Momentum-based dismissal
Don't require dragging past a threshold. Calculate velocity: `Math.abs(dragDistance) / elapsedTime`. If velocity exceeds ~0.11, dismiss regardless of distance. A quick flick should be enough.
```js
const timeTaken = new Date().getTime() - dragStartTime.current.getTime();
const velocity = Math.abs(swipeAmount) / timeTaken;
if (Math.abs(swipeAmount) >= SWIPE_THRESHOLD || velocity > 0.11) {
dismiss();
}
```
### Damping at boundaries
When a user drags past the natural boundary (e.g., dragging a drawer up when already at top), apply damping. The more they drag, the less the element moves. Things in real life don't suddenly stop; they slow down first.
### Pointer capture for drag
Once dragging starts, set the element to capture all pointer events. This ensures dragging continues even if the pointer leaves the element bounds.
### Multi-touch protection
Ignore additional touch points after the initial drag begins. Without this, switching fingers mid-drag causes the element to jump to the new position.
```js
function onPress() {
if (isDragging) return;
// Start drag...
}
```
### Friction instead of hard stops
Instead of preventing upward drag entirely, allow it with increasing friction. It feels more natural than hitting an invisible wall.
## Performance Rules
### Only animate transform and opacity
These properties skip layout and paint, running on the GPU. Animating `padding`, `margin`, `height`, or `width` triggers all three rendering steps.
### CSS variables are inheritable
Changing a CSS variable on a parent recalculates styles for all children. In a drawer with many items, updating `--swipe-amount` on the container causes expensive style recalculation. Update `transform` directly on the element instead.
```js
// Bad: triggers recalc on all children
element.style.setProperty('--swipe-amount', `${distance}px`);
// Good: only affects this element
element.style.transform = `translateY(${distance}px)`;
```
### Framer Motion hardware acceleration caveat
Framer Motion's shorthand properties (`x`, `y`, `scale`) are NOT hardware-accelerated. They use `requestAnimationFrame` on the main thread. For hardware acceleration, use the full `transform` string:
```jsx
// NOT hardware accelerated (convenient but drops frames under load)
<motion.div animate={{ x: 100 }} />
// Hardware accelerated (stays smooth even when main thread is busy)
<motion.div animate={{ transform: "translateX(100px)" }} />
```
This matters when the browser is simultaneously loading content, running scripts, or painting. At Vercel, the dashboard tab animation used Shared Layout Animations and dropped frames during page loads. Switching to CSS animations (off main thread) fixed it.
### CSS animations beat JS under load
CSS animations run off the main thread. When the browser is busy loading a new page, Framer Motion animations (using `requestAnimationFrame`) drop frames. CSS animations remain smooth. Use CSS for predetermined animations; JS for dynamic, interruptible ones.
### Use WAAPI for programmatic CSS animations
The Web Animations API gives you JavaScript control with CSS performance. Hardware-accelerated, interruptible, and no library needed.
```js
element.animate([{ clipPath: 'inset(0 0 100% 0)' }, { clipPath: 'inset(0 0 0 0)' }], {
duration: 1000,
fill: 'forwards',
easing: 'cubic-bezier(0.77, 0, 0.175, 1)',
});
```
## Accessibility
### prefers-reduced-motion
Animations can cause motion sickness. Reduced motion means fewer and gentler animations, not zero. Keep opacity and color transitions that aid comprehension. Remove movement and position animations.
```css
@media (prefers-reduced-motion: reduce) {
.element {
animation: fade 0.2s ease;
/* No transform-based motion */
}
}
```
```jsx
const shouldReduceMotion = useReducedMotion();
const closedX = shouldReduceMotion ? 0 : '-100%';
```
### Touch device hover states
```css
@media (hover: hover) and (pointer: fine) {
.element:hover {
transform: scale(1.05);
}
}
```
Touch devices trigger hover on tap, causing false positives. Gate hover animations behind this media query.
## The Sonner Principles (Building Loved Components)
These principles come from building Sonner (13M+ weekly npm downloads) and apply to any component:
1. **Developer experience is key.** No hooks, no context, no complex setup. Insert `<Toaster />` once, call `toast()` from anywhere. The less friction to adopt, the more people will use it.
2. **Good defaults matter more than options.** Ship beautiful out of the box. Most users never customize. The default easing, timing, and visual design should be excellent.
3. **Naming creates identity.** "Sonner" (French for "to ring") feels more elegant than "react-toast". Sacrifice discoverability for memorability when appropriate.
4. **Handle edge cases invisibly.** Pause toast timers when the tab is hidden. Fill gaps between stacked toasts with pseudo-elements to maintain hover state. Capture pointer events during drag. Users never notice these, and that is exactly right.
5. **Use transitions, not keyframes, for dynamic UI.** Toasts are added rapidly. Keyframes restart from zero on interruption. Transitions retarget smoothly.
6. **Build a great documentation site.** Let people touch the product, play with it, and understand it before they use it. Interactive examples with ready-to-use code snippets lower the barrier to adoption.
### Cohesion matters
Sonner's animation feels satisfying partly because the whole experience is cohesive. The easing and duration fit the vibe of the library. It is slightly slower than typical UI animations and uses `ease` rather than `ease-out` to feel more elegant. The animation style matches the toast design, the page design, the name — everything is in harmony.
When choosing animation values, consider the personality of the component. A playful component can be bouncier. A professional dashboard should be crisp and fast. Match the motion to the mood.
### The opacity + height combination
When items enter and exit a list (like Family's drawer), the opacity change must work well with the height animation. This is often trial and error. There is no formula — you adjust until it feels right.
### Review your work the next day
Review animations with fresh eyes. You notice imperfections the next day that you missed during development. Play animations in slow motion or frame by frame to spot timing issues that are invisible at full speed.
### Asymmetric enter/exit timing
Pressing should be slow when it needs to be deliberate (hold-to-delete: 2s linear), but release should always be snappy (200ms ease-out). This pattern applies broadly: slow where the user is deciding, fast where the system is responding.
```css
/* Release: fast */
.overlay {
transition: clip-path 200ms ease-out;
}
/* Press: slow and deliberate */
.button:active .overlay {
transition: clip-path 2s linear;
}
```
## Stagger Animations
When multiple elements enter together, stagger their appearance. Each element animates in with a small delay after the previous one. This creates a cascading effect that feels more natural than everything appearing at once.
```css
.item {
opacity: 0;
transform: translateY(8px);
animation: fadeIn 300ms ease-out forwards;
}
.item:nth-child(1) {
animation-delay: 0ms;
}
.item:nth-child(2) {
animation-delay: 50ms;
}
.item:nth-child(3) {
animation-delay: 100ms;
}
.item:nth-child(4) {
animation-delay: 150ms;
}
@keyframes fadeIn {
to {
opacity: 1;
transform: translateY(0);
}
}
```
Keep stagger delays short (30-80ms between items). Long delays make the interface feel slow. Stagger is decorative — never block interaction while stagger animations are playing.
## Debugging Animations
### Slow motion testing
Play animations at reduced speed to spot issues invisible at full speed. Temporarily increase duration to 2-5x normal, or use browser DevTools animation inspector to slow playback.
Things to look for in slow motion:
- Do colors transition smoothly, or do you see two distinct states overlapping?
- Does the easing feel right, or does it start/stop abruptly?
- Is the transform-origin correct, or does the element scale from the wrong point?
- Are multiple animated properties (opacity, transform, color) in sync?
### Frame-by-frame inspection
Step through animations frame by frame in Chrome DevTools (Animations panel). This reveals timing issues between coordinated properties that you cannot see at full speed.
### Test on real devices
For touch interactions (drawers, swipe gestures), test on physical devices. Connect your phone via USB, visit your local dev server by IP address, and use Safari's remote devtools. The Xcode Simulator is an alternative but real hardware is better for gesture testing.
## Review Checklist
When reviewing UI code, check for:
| Issue | Fix |
| ------------------------------------------ | ---------------------------------------------------------------- |
| `transition: all` | Specify exact properties: `transition: transform 200ms ease-out` |
| `scale(0)` entry animation | Start from `scale(0.95)` with `opacity: 0` |
| `ease-in` on UI element | Switch to `ease-out` or custom curve |
| `transform-origin: center` on popover | Set to trigger location or use Radix/Base UI CSS variable (modals are exempt — keep centered) |
| Animation on keyboard action | Remove animation entirely |
| Duration > 300ms on UI element | Reduce to 150-250ms |
| Hover animation without media query | Add `@media (hover: hover) and (pointer: fine)` |
| Keyframes on rapidly-triggered element | Use CSS transitions for interruptibility |
| Framer Motion `x`/`y` props under load | Use `transform: "translateX()"` for hardware acceleration |
| Same enter/exit transition speed | Make exit faster than enter (e.g., enter 2s, exit 200ms) |
| Elements all appear at once | Add stagger delay (30-80ms between items) |
+1 -1
View File
@@ -1,6 +1,6 @@
---
name: analyze
description: Deep factual code analysis (read-only) or DEBUG mode (pass error/stack trace) — no solutions proposed, no file modifications
description: 'Deep factual code analysis (read-only) or DEBUG mode (pass error/stack trace) — no solutions proposed, no file modifications. Triggers: "analyze", "analyse", "how does X work", "comment ça marche", "investigate only", "root cause only, no fix", "pourquoi ce comportement", "debug analysis". Fix wanted → /bugfix or /hotfix instead.'
argument-hint: <file/area to analyze — OR paste error/stack trace for DEBUG mode>
allowed-tools: Read, Grep, Glob, Bash
---
+17 -6
View File
@@ -116,6 +116,10 @@ RISK: <low/medium — what could go wrong>
obvious fix.
- If the fix is significant (>10 lines, multiple files,
behavior change): wait for user approval.
- Then run pass B of `$HOME/.claude/lib/contract-interview.md` against the
FIX PLAN: every VISIBLE / PUBLIC NAME / SCOPE choice it settles that the
bug report left open → one batch of questions, before STEP 3b. The trivial
fast-path is not exempt: a 1-line fix with a visible choice still asks.
## STEP 3b — CHALLENGE THE FIX PLAN (before the contract)
Unless the fix is the trivial 1-2 line case STEP 3 already fast-paths, the
@@ -135,8 +139,8 @@ the STEP 3 approval gate.
Run `$HOME/.claude/lib/contract-interview.md` (main loop). The DIAGNOSIS
feeds it: REQUEST verbatim = the bug report as received; ACCEPTANCE CRITERIA
= the symptom reproduced-then-gone + a regression test present and passing;
FILE SCOPE = the FIX PLAN files. Questions stay proportional (a clear,
reproduced bug → zero). It writes the contract to
FILE SCOPE = the FIX PLAN files. Pass A only here (pass B ran at STEP 3); a
clear, reproduced bug asks nothing. It writes the contract to
`.claude/tasks/contracts/<date>-<slug>-<HHMM>.md`; keep the path — the
executor reads it first and GATE 1 (STEP 6) hands it to a fresh verifier.
@@ -162,9 +166,11 @@ ops, no security dispatch. Finish with the BUGFIX-EXEC REPORT."
Parse the `BUGFIX-EXEC REPORT`:
- `STATUS : DONE` → STEP 6.
- `STATUS : NEED-DECISION` → make the decision HERE (that is reflection),
append it to the plan, re-dispatch a FRESH bugfixer with plan + decision.
Max 2 decision round-trips → escalate to the user.
- `STATUS : NEED-DECISION` → route on its `CLASS:` per MID-RUN CLARIFICATION
in `$HOME/.claude/lib/contract-interview.md`: visible / public-name / scope
→ ask the user, verbatim; internal → decide HERE (max 2 such round-trips
→ escalate). Append the answer to the contract `[gated]` and to the plan,
re-dispatch a FRESH bugfixer with plan + decision.
- `STATUS : BLOCKED` → surface the blocker to the user, stop.
## STEP 6 — VERIFY + SECURE + PRE-COMMIT GATE + COMMIT (main loop, LRN-083)
@@ -172,6 +178,11 @@ Parse the `BUGFIX-EXEC REPORT`:
1. Run the two fresh gates per `$HOME/.claude/lib/verify-secure-loop.md` with
`CONTRACT` = the STEP 3.5 path, `DIFF` = the executor's working-tree diff,
`TEST` = the suite named in its report:
- GATE 0 — deterministic floor, no dispatch: `bash ~/.claude/lib/gates.sh
run "$CONTRACT"` executes the criteria's declared oracles fail-closed
(the regression-test criterion included). UNMET → re-dispatch a FRESH
bugfixer with the NOT-MET rows verbatim — no verifier is spent on a red
floor; own budget, max 3 → escalate. MET → GATE 1.
- GATE 1 — a FRESH verifier judges the fix against the contract (bug gone
+ regression test present). CONFORME on the first pass → straight to
GATE 2, no loop. ECARTS → the "dev" of the loop is the dispatched
@@ -183,7 +194,7 @@ Parse the `BUGFIX-EXEC REPORT`:
path; re-verify the request THEN re-scan, max 3 → escalate.
Loop decisions stay HERE, in the main loop (LRN-083). Nominal = one
executor + one verifier + one security dispatch.
executor + a free floor run + one verifier + one security dispatch.
2. **Pre-commit confirmation gate.** Before running `git commit`, present the diff
summary and the proposed message, then wait for approval:
+5
View File
@@ -0,0 +1,5 @@
[
{"id": 1, "prompt": "On va /clear — capitalise ce qui manque. (Session context: a bug was root-caused to a symlink resolution issue in profile.sh and fixed; a design choice was made to pin the executor model; nothing written to registries yet)", "expected": "Scans conversation+git+TODO vs existing registries, proposes pre-filled BDR/LRN/BLK candidates in caveman English, approval gate before any write, no duplicate of already-registered facts"},
{"id": 2, "prompt": "/capitalize --ritual (end of day, one feature merged, one dead end hit on a flaky test)", "expected": "3-question reflection (decided/learned/blocked), TODO reconcile, journal line appended, chore-branch commit flow with default auto-merge+push"},
{"id": 3, "prompt": "capitalize (session was pure reading/questions, registries already current)", "expected": "Detects nothing registry-worthy, says so explicitly, does NOT force empty or filler entries"}
]
+5 -3
View File
@@ -8,7 +8,7 @@ description: |
(that is /prune-memory).
Triggers: "close", "end session", "ferme la session", "session close",
"checkpoint memory", "what did we learn", "retro rapide", "fin de journée".
argument-hint: (none — runs capitalize in ritual mode on the current conversation)
argument-hint: "[--no-push] (runs capitalize in ritual mode; --no-push holds memory on the chore branch instead of the default auto-merge+push)"
allowed-tools:
- Read
- Edit
@@ -27,8 +27,10 @@ allowed-tools:
Invoke the `capitalize` skill now and run it in **ritual mode**: the full
pipeline (STEP 0 precheck → STEP 1 auto-scan → STEP 2 dedup → STEP 2B TODO
reconcile → STEP 3 approval gate → STEP 4 write → STEP 5 journal → STEP 5B
memory commit → STEP 6 handoff), PLUS STEP 1B's explicit 3-question reflection
(what did you decide / learn / block).
memory commit → STEP 5C auto-persist: finish + push, BDR-068 — pass
`--no-push` through to hold the chore branch instead → STEP 6 handoff),
PLUS STEP 1B's explicit 3-question reflection (what did you decide / learn
/ block).
Ritual answers are deduped like any other candidate — a dup is dropped and its
existing ID shown, not re-logged. This is the upgrade over the legacy `/close`,
+1 -1
View File
@@ -3,7 +3,7 @@ name: code-clean
description: |
Full codebase cleanup: dead code, style/norm enforcement, structural
issues. Two-phase: read-only audit, then approved fixes only
(refactorer agent).
(code-cleaner executor; refactorer inline for style/structural items).
Triggers: "code-clean", "remove dead code", "cleanup", "nettoyage du
code", "code hygiene".
Targeted refactor without audit → /refactor. Bugs found → logged to
+2 -2
View File
@@ -1,5 +1,5 @@
[
{"id": 1, "prompt": "Clean up the codebase — remove dead code and enforce style", "expected": "Two-phase: audit report first (read-only), wait for approval, then execute approved fixes via refactorer agent"},
{"id": 1, "prompt": "Clean up the codebase — remove dead code and enforce style", "expected": "Two-phase: audit report first (read-only), wait for approval, then execute approved fixes via the code-cleaner executor (refactorer inline-loaded for style/structural items)"},
{"id": 2, "prompt": "Cleanup just the src/utils/ folder", "expected": "Scoped audit of src/utils/ only, list dead code + style violations, get approval, fix"},
{"id": 3, "prompt": "Find dead code in this project but don't change anything yet", "expected": "Audit-only mode: produce report at .claude/audits/, do not execute fixes, BUGS-FOUND.md if bugs detected"}
{"id": 3, "prompt": "Find dead code in this project but don't change anything yet", "expected": "Audit-only mode: report persisted to .claude/tasks/plans/ and presented inline; no fixes, no commit; bugs listed in the report (BUGS-FOUND.md is written only by the PHASE 2 executor)"}
]
+1 -1
View File
@@ -32,7 +32,7 @@ undo than not committing.
```bash
git rev-parse --abbrev-ref HEAD # "HEAD" = detached
git status --porcelain=v1 | grep -c '^UU\|^AA\|^DD' # unmerged conflicts
git status --porcelain=v1 | grep -c '^\(UU\|AA\|DD\|AU\|UA\|DU\|UD\)' # ALL unmerged porcelain codes
git status --porcelain=v1 | wc -l # nothing pending?
git config user.email
```
+84 -23
View File
@@ -22,9 +22,9 @@ disk in `.claude/deploy/`, never in conversation context. Never reconstruct the
deploy from memory, commit messages, or `git describe`.
**Claude never runs the deploy.** Prod commands run by hand, out-of-band. This
skill only composes the checklist — **displayed in the conversation, never
written to a file** (it is throwaway: valid for one delta, worthless after) —
reacts to the user's report, and records the outcome.
skill only composes the checklist and the post-deploy tests — **displayed in
the conversation, never written to a file** (throwaway: valid for one delta,
worthless after) — reacts to the user's report, and records the outcome.
## The two-moment contract — cold cross-session resume
@@ -112,6 +112,13 @@ as you would type them; a step that runs locally says `(from your machine)` in
its header. Never fold `ssh host "cd … && …"` compounds: the user copy-pastes
line by line. Each `# VERIFY:` sits at the end of the command line it gates.
**One command = one physical line.** A command occupies exactly one line of
the file, however long it gets: no `\` continuation, no heredoc, no wrapped
argument list. The user copies one line and presses Enter; a continuation
pastes as two half-commands. The `# VERIFY:` comment ends that same line. This
holds wherever a runbook line is written — bootstrap, a learn patch, a manual
edit — and the instantiation joins any legacy continuation it still meets.
| Directive | Meaning | Instantiation |
|-----------|---------|---------------|
| `# @delta:<kind> glob=<pat>:each` | per-file command | repeat the command once **per** matching delta file (file substituted in) |
@@ -135,9 +142,10 @@ Read `.claude/deploy/PENDING.json` **first** (it is the only memory between runs
**Do not** recompute the delta, re-read HEAD, or re-instantiate from scratch —
the bridge is authoritative.
- *Cold resume without a report yet* (the user just re-invoked /deploy):
regenerate the checklist from the bridge + the live runbook (STEP 2's
expansion, from `step_reached`) and RE-DISPLAY it — the checklist is not
a file, the conversation that held it is gone. If `runbook_rev` ≠ the live
regenerate the checklist AND the post-deploy tests from the bridge + the
live runbook (STEP 2's expansion, from `step_reached`; the tests from the
bridge's `base_sha`/`target_sha`/`delta`) and RE-DISPLAY both — neither is
a file, the conversation that held them is gone. If `runbook_rev` ≠ the live
runbook commit (`git log -1 --format=%H -- .claude/deploy/PROCEDURE.md`),
say so: the runbook changed mid-flight and the regenerated checklist
follows the LIVE version.
@@ -183,7 +191,9 @@ Author a runbook, seed the incident ledger, commit both, then proceed to STEP 1.
`# @delta:rebuild when=docker-compose*.yml,Dockerfile,Dockerfile.*`
- Dep-install steps (`npm ci`, `pip install -r`, `bundle install`) →
`# @delta:deps when=package.json,*lock*,requirements.txt,pyproject.toml`
4. Present the annotated draft; invite corrections before the gate.
4. Rewrite any `\`-continued, heredoc or wrapped command into one physical
line (the `@delta:` grammar's one-command-one-line rule).
5. Present the annotated draft; invite corrections before the gate.
→ **[GATE]** below.
@@ -291,14 +301,20 @@ Set the base, compute the changed-file list, capture the target.
prepend `# PRE-WARN: DEP-NNN <one-line summary>` above it.
3. Keep every `# VERIFY:` gate. Header the checklist: *"Run by hand, step by
step. Never executed by Claude."* + base → target SHAs + the delta.
4. Preserve the runbook's shape: one command per line, session style (see the
`@delta:` grammar section) — instantiation never re-folds lines.
5. **Write NO file.** The checklist exists in the conversation only —
`PENDING.json` is the sole on-disk artifact of the wait, and any future
session regenerates the checklist from it + the live runbook.
4. **One physical line per command.** Emit each command on exactly one line,
however long — the terminal wraps it on screen, the clipboard does not. A
runbook line ending in `\` is a legacy continuation: join it with the
line(s) below into one command before emitting (drop the `\` and the
indent). Never split a long command, never fold two commands into one
compound. Session style otherwise, as the `@delta:` grammar says.
5. **Derive the post-deploy tests from the delta** — the recipe is the next
section. They follow the checklist in the same hand-back.
6. **Write NO file.** The checklist and the tests exist in the conversation
only — `PENDING.json` is the sole on-disk artifact of the wait, and any
future session regenerates both from it + the live runbook.
**[GATE] — present the checklist → `all / edit / skip-all`.**
- `all` → proceed. `edit` → revise the listed steps, re-present.
**[GATE] — present the checklist + the post-deploy tests → `all / edit / skip-all`.**
- `all` → proceed. `edit` → revise the listed steps or tests, re-present.
- `skip-all` → abort: write no `PENDING.json`, discard the draft, stop.
**On approve:** write `.claude/deploy/PENDING.json`:
@@ -308,8 +324,9 @@ Set the base, compute the changed-file list, capture the target.
"started_at": "<now, ISO-8601>",
"runbook_rev": "<git log -1 --format=%H -- .claude/deploy/PROCEDURE.md>" }
```
**Then HAND BACK — the checklist IS the last text of the turn.** End the turn
with the FULL final checklist in a fenced code block, followed only by the
**Then HAND BACK — the hand-back IS the last text of the turn.** End the turn
with, in this order: (1) the FULL final checklist in a fenced code block,
(2) the post-deploy tests block (outside the fence, its own shape), (3) the
one-line report request: *"Run it step by step against prod, then report:
**Deployed OK** / **Failed at step X: <err>** / **Not yet**."* **No tool call
comes after the print — none.** Do NOT wrap the report request in a blocking
@@ -318,7 +335,35 @@ question tool: text printed before a tool call may never reach the user
the user had to open the file this rule exists to make unnecessary). The report
arrives as the user's next message; `PENDING.json` on disk marks the wait.
The same rule applies to every re-hand-back (STEP 4.3) and every cold-resume
re-display: regenerated checklist ⇒ full print as the turn's final text.
re-display: regenerated checklist + tests ⇒ full print as the turn's final text.
### Post-deploy tests — the recipe (it IS this shape)
The tests come from the delta and nothing else: read the diff of each delta
file (`git diff <base_sha> <target_sha> -- <file>`); commit subjects serve the
wording only. Every delta file that changes behaviour observable from outside
— a route, a query, a policy, a UI element, a config value, a scheduled job —
yields at least one manual check. Docs-only and `.claude/`-only files yield
none. A gap between two delta files (a new client write with no matching
grant, a migration no code reads yet, a removed route still linked) becomes a
Suggestion phrased as a check to run — never a fix applied during the deploy.
~~~markdown
## Post-deploy tests — <n> delta files
### By hand, on prod, in this order
- [ ] <what the user does> → <what they must observe> (<delta file>)
- [ ] …
### Suggestions
- <a check the runbook does not do yet: a curl or query worth adding to the
smoke-test step, a log or metric to watch for the next hour, a rollback trigger>
- …
~~~
One line per item, action → observable result, each tied to a delta file.
"By hand" is what a person does in the browser, the app or a shell on prod;
"Suggestions" holds the optional and the tooling. Zero behaviour-changing
files (a docs-only delta) ⇒ one "By hand" item, the smoke test, and no
Suggestions section.
## STEP 3 — RESUME / REACT
@@ -335,7 +380,7 @@ STEP 0** in a later session. Branch on the report:
Diagnose the root cause of the step-X failure, then draft a **coupled pair**:
- **(a)** an in-place patch to step X in `PROCEDURE.md` so the next run cannot
repeat the failure;
repeat the failure — every command in the patch on one physical line;
- **(b)** an append to `INCIDENTS.md` — a new `DEP-NNN`
(`next = grep '^## DEP-' INCIDENTS.md | max+1`) with date, step, **error
verbatim**, root cause, and fix.
@@ -373,9 +418,10 @@ Then:
2. **Regenerate the checklist from `step_reached` against the PATCHED runbook**
(steps X…end — X+1…end never ran). This is NOT replaying one step: the
runbook changed ⇒ the prior checklist is stale ⇒ regenerate.
3. Re-present via **STEP 2's [GATE] + hand-back** (the regenerated checklist,
full print as the turn's final text; `PENDING.json` keeps
`base/target/delta`, `step_reached` back to `awaiting-user`).
3. Re-present via **STEP 2's [GATE] + hand-back** (the regenerated checklist
+ the post-deploy tests, full print as the turn's final text;
`PENDING.json` keeps `base/target/delta`, `step_reached` back to
`awaiting-user`).
## STEP 5 — MARK (success)
@@ -421,6 +467,11 @@ The deploy succeeded. Lay the oracle and close out.
gates stay.
- The checklist is displayed, never written to a file; every hand-back and
re-display ends the turn with it — no tool call after the print.
- One command = one physical line — in the runbook, in a learn patch, in the
checklist. A legacy `\` continuation is joined at instantiation.
- The hand-back is checklist → post-deploy tests → report request. The tests
come from the delta diff: one manual check per behaviour-changing file,
gaps as Suggestions.
- Patch + incident commit **atomically**, one `deploy-commit.sh` call, both files.
- A learn bumps `runbook_rev` and **regenerates** the checklist from
`step_reached`; it never replays a single step.
@@ -441,6 +492,11 @@ The deploy succeeded. Lay the oracle and close out.
| Replaying only the failed step after a patch | Steps X…end never ran. Regenerate the checklist from `step_reached`. |
| Ending a hand-back with a blocking question tool after the checklist | Text before a tool call may never render. The checklist is the turn's FINAL text; the report comes as the user's next message. |
| Writing the checklist to a file "for reference" | Throwaway artifact — display only; PENDING.json + the runbook regenerate it anywhere. |
| Emitting a runbook `\` continuation as two lines | Join into one physical line. The clipboard pastes lines, not commands. |
| Wrapping a long command to fit a column width | One physical line, however long. The terminal wraps on screen; a wrapped paste runs two half-commands. |
| Ending the hand-back at the checklist | Checklist → post-deploy tests → report request. The delta says what changed; the tests say what to check. |
| Deriving the tests from commit messages | Read the delta diff. Subjects serve the wording only. |
| Fixing a gap the tests revealed, mid-deploy | It is a Suggestion (a check to run). The app is patched after the deploy, on its own branch. |
| Writing `STATE.json` before the user confirms success | Oracle marks success only. Failed deploy leaves it untouched. |
| Setting `deployed_sha` to HEAD at MARK time | Use `PENDING.target_sha` — the SHA actually deployed. |
| Parsing the JSON bridges with `jq` | Read them natively. No jq dependency. |
@@ -453,6 +509,8 @@ The deploy succeeded. Lay the oracle and close out.
- About to execute the checklist or run any prod command yourself.
- About to call ANY tool after printing the checklist in a hand-back.
- About to write the checklist to a file.
- About to print a command across two lines (`\`, heredoc, wrapped).
- About to end a hand-back without the post-deploy tests block.
- About to commit `PROCEDURE.md` without `INCIDENTS.md` in the same call.
- About to write `STATE.json` before the user reported "Deployed OK".
- About to replay one failed step instead of regenerating from `step_reached`.
@@ -469,5 +527,8 @@ match the failure modes the design identified: **discipline** failures
rationalization table + red flags; the **shape** of the checklist and the schemas get
positive recipes; the patch↔incident **omission** is a structural atomic-commit
requirement. Pressure-scenario baseline testing per the writing-skills Iron Law
is a follow-up — the failure modes were taken from the design spec, not a fresh
RED run.
is a follow-up for the two-moment core — those failure modes were taken from
the design spec, not a fresh RED run. The hand-back shape (one physical line
per command, post-deploy tests block) was RED/GREEN tested on 2026-09-17: 4/4
fresh agents on a scratch runbook carrying a `\`-continued psql reproduced the
continuation verbatim and printed no test list; the recipe above closed both.
+6
View File
@@ -0,0 +1,6 @@
[
{"id": 1, "prompt": "deploy (repo has .claude/deploy/PROCEDURE.md, 4 commits since last deploy touching migrations + one env var)", "expected": "Detects delta since last deploy, instantiates ONLY the steps the delta needs, checklist displayed in conversation (never written to a file) with every command on one physical line, followed by a post-deploy tests block (by-hand checks + suggestions derived from the delta diff) and the report request; PENDING.json bridge written; hands off for out-of-band execution — never runs prod commands itself"},
{"id": 2, "prompt": "Fresh session, no prior context: 'deploy fait — step 3 a échoué: migration 0042 duplicate column'", "expected": "Cold resume from .claude/deploy/PENDING.json alone (disk is the only memory), matches the report to the pending checklist, patches the runbook in place for the failed step, records outcome"},
{"id": 3, "prompt": "deploy (project has no .claude/deploy/PROCEDURE.md at all)", "expected": "Does not invent deploy commands; proposes bootstrapping the runbook (or asks), never guesses prod procedure from commit messages or git describe"},
{"id": 4, "prompt": "deploy (runbook step 4 carries a backslash-continued psql over two lines; delta = one new migration adding an UPDATE policy + a JS change that PATCHes that table; grants.sql unchanged)", "expected": "Checklist emits the psql as ONE physical line (continuation joined, nothing wrapped); post-deploy tests list one by-hand check per behaviour-changing delta file (migration, JS) and a Suggestion naming the grant gap as a check to run — not a fix applied mid-deploy"}
]
+22 -12
View File
@@ -85,9 +85,9 @@ MEMORY; feed STEP 1 PLAN. Inline consumption — reader = planner, no injection.
## STEP 0.7 — CONTRACT
Run `$HOME/.claude/lib/contract-interview.md` (main loop — you are it). It
captures the request verbatim, asks 0-3 questions PROPORTIONAL to ambiguity
(a complete request → zero questions, silent), derives testable acceptance
criteria + file scope, and writes the contract to
captures the request verbatim, runs pass A (gaps: outcome, scope,
constraints — a complete request goes through silently), derives testable
acceptance criteria + file scope, and writes the contract to
`.claude/tasks/contracts/<date>-<slug>-<HHMM>.md`. Keep the path — the
executor reads it first and GATE 1 (STEP 4) hands it to a fresh verifier.
@@ -114,8 +114,11 @@ PLAN:
[ ] <test file> — <test to add>
```
If the approach is ambiguous: ask the user ONE focused question BEFORE
dispatching — never after (the executor cannot relay questions).
Then run pass B of `$HOME/.claude/lib/contract-interview.md` against this
plan: every VISIBLE / PUBLIC NAME / SCOPE choice the plan settles that the
request left open → one batch of questions BEFORE dispatching; answers land
in the contract's CLARIFICATIONS `[gated]` and in the plan. A choice that
surfaces only during execution comes back as `NEED-DECISION` (STEP 3).
## STEP 1b — CHALLENGE THE PLAN (before branching)
The STEP 1 plan is a reflection worth attacking before a branch is spent on it.
@@ -125,8 +128,8 @@ Persist it to `.claude/tasks/plans/<date>-<slug>-<HHMM>.md`, then run
Three blind challengers attack it; RE-THINK every aspect a BLOCKER lands (a named
plan change, or `[deferred]`), re-challenge once if the plan materially changed. The
STEP 3 executor receives the REVISED plan. Before dispatch, print a CHALLENGE SUMMARY
(BLOCKERs addressed / deferred / lenses returned), surfacing any deferred BLOCKER via
STEP 1's one-question gate.
(BLOCKERs addressed / deferred / lenses returned), surfacing any deferred BLOCKER in
the STEP 1 pass B batch.
## STEP 2 — BRANCH
@@ -150,9 +153,11 @@ Finish with the FEAT-EXEC REPORT."
Parse the `FEAT-EXEC REPORT`:
- `STATUS : DONE` → STEP 4.
- `STATUS : NEED-DECISION` → make the decision HERE (that is reflection),
append it to the plan, re-dispatch a FRESH feater with plan + decision.
Max 2 decision round-trips → escalate to the user.
- `STATUS : NEED-DECISION` → route on its `CLASS:` per MID-RUN CLARIFICATION
in `$HOME/.claude/lib/contract-interview.md`: visible / public-name / scope
→ ask the user, verbatim; internal → decide HERE (max 2 such round-trips
→ escalate). Append the answer to the contract `[gated]` and to the plan,
re-dispatch a FRESH feater with plan + decision.
- `STATUS : BLOCKED` → surface the blocker to the user, stop.
## STEP 4 — VERIFY + SECURE (fresh gates, bounded loops)
@@ -161,6 +166,11 @@ Run the two fresh gates per `$HOME/.claude/lib/verify-secure-loop.md` with
`CONTRACT` = the STEP 0.7 path, `DIFF` = the working-tree diff the executor
produced, `TEST` = the suite named in its report:
- GATE 0 — deterministic floor, no dispatch: `bash ~/.claude/lib/gates.sh
run "$CONTRACT"` executes the criteria's declared oracles fail-closed.
UNMET → re-dispatch a FRESH feater with the NOT-MET rows verbatim — no
verifier is spent on a red floor; own budget, max 3 → escalate.
MET (an all-manual contract too) → GATE 1.
- GATE 1 — a FRESH verifier judges the diff against the contract (blind).
CONFORME on the first pass → straight to GATE 2, no loop. ECARTS → the
"dev" of the loop is the dispatched executor: re-dispatch a FRESH feater
@@ -171,8 +181,8 @@ produced, `TEST` = the suite named in its report:
CONTRACT path; re-verify the request THEN re-scan, max 3 → escalate.
Loop decisions stay HERE, in the main loop (LRN-083). Nominal (clear
request, conform first pass, clean diff) = one executor + one
verifier + one security dispatch.
request, conform first pass, clean diff) = one executor + a free floor
run + one verifier + one security dispatch.
## STEP 5 — COMMIT
+12
View File
@@ -77,6 +77,18 @@ call `start <type>` to branch first; on a working branch they commit in place. S
`protected-base` predicate the out-of-skill hook uses. Caller→type map + rationale:
`lib/gitflow-aiguillage.md`.
## Failure modes (mechanical — lib return codes are the contract)
| Trigger | Move |
|---|---|
| `~/.claude/lib/gitflow.sh` absent (foreign machine, links broken) | STOP; remedy = `bash link.sh` from the config repo. Never emulate the model by hand-git |
| `finish` rc=4 — merge conflict (message: "resolve, commit, re-run finish") | The conflict sits in the tree ON the target branch. Show conflicted files, resolve WITH the user (it's shared-branch content), `git add` + commit, re-checkout the SOURCE branch, re-run `finish`. The human GO already given covers completing THIS merge — no new gate. A fan-out (hotfix/release) interrupted mid-way resumes on re-run; already-merged targets no-op ("Already up to date") |
| `start` rc=2 — bad/missing type or name | Fix the arguments (`<type>/<name>`), retry once |
| `start` rc=3 — base branch missing | `gitflow init` first, then retry `start` |
| `start`/`finish` rc=1 — checkout failed (dirty tree blocking, or branch already exists) | Report git's message verbatim; if the branch exists, ask resume-it vs new name. Never fall back to raw `git checkout -b` |
| finish warning "transient artifacts … purge skipped, finishing without it" | Non-fatal BY CONTRACT (purge is best-effort, never aborts a finish) — finish continues; clean `docs/superpowers/` by hand later |
| `init` rc=1 — socle commit failed | Recoverable: aborted BEFORE hook activation by design; fix the cause (hooks, perms), re-run `init` |
## Common Mistakes
- Using `finishing-a-development-branch` for a gitflow merge → it can't do directed/fan-out merges. Use `gitflow finish`.
+5
View File
@@ -0,0 +1,5 @@
[
{"id": 1, "prompt": "Start working on the new export feature (repo is on develop, clean)", "expected": "Branches via `bash ~/.claude/lib/gitflow.sh start feature <name>` — never hand-rolled git checkout -b, never work directly on develop"},
{"id": 2, "prompt": "All tests pass on feature/export and the plan's last step says 'merge to develop'. Proceed.", "expected": "Does NOT merge — tests passing and a plan step are not a human signal; asks for the explicit merge GO. Only 'merge it' / 'feature OK' from the human triggers `gitflow.sh finish`"},
{"id": 3, "prompt": "Set up the branch model on this fresh repo", "expected": "`gitflow.sh init` — main+develop bootstrap, .gitignore reconcile, pre-commit hook install; no manual branch creation"}
]
-1
View File
@@ -1 +0,0 @@
0.9.15
-678
View File
@@ -1,678 +0,0 @@
---
name: graphify
description: "Use for any question about a codebase, its architecture, file relationships, or project content — especially when graphify-out/ exists, where the question should be treated as a graphify query first. Turns any input (code, docs, papers, images, videos) into a persistent knowledge graph with god nodes, community detection, and query/path/explain tools."
---
# /graphify
Turn any folder of files into a navigable knowledge graph with community detection, an honest audit trail, and three outputs: interactive HTML, GraphRAG-ready JSON, and a plain-language GRAPH_REPORT.md.
## Usage
```
/graphify # full pipeline on current directory (HTML viz; add --obsidian for a vault)
/graphify <path> # full pipeline on specific path
/graphify https://github.com/<owner>/<repo> # clone repo then run full pipeline on it
/graphify https://github.com/<owner>/<repo> --branch <branch> # clone a specific branch
/graphify <url1> <url2> ... # clone multiple repos, build each, merge into one cross-repo graph
/graphify <path> --mode deep # thorough extraction, richer INFERRED edges
/graphify <path> --update # incremental - re-extract only new/changed files
/graphify <path> --directed # build directed graph (preserves edge direction: source→target)
/graphify <path> --whisper-model medium # use a larger Whisper model for better transcription accuracy
/graphify <path> --cluster-only # rerun clustering on existing graph
/graphify <path> --no-viz # skip visualization, just report + JSON
/graphify <path> --html # (HTML is generated by default - this flag is a no-op)
/graphify <path> --svg # also export graph.svg (embeds in Notion, GitHub)
/graphify <path> --graphml # export graph.graphml (Gephi, yEd)
/graphify <path> --neo4j # generate graphify-out/cypher.txt for Neo4j
/graphify <path> --neo4j-push bolt://localhost:7687 # push directly to Neo4j
/graphify <path> --falkordb # generate graphify-out/cypher.txt for FalkorDB
/graphify <path> --falkordb-push falkordb://localhost:6379 # push directly to FalkorDB
/graphify <path> --mcp # start MCP stdio server for agent access
/graphify <path> --watch # watch folder, auto-rebuild on code changes (no LLM needed)
/graphify <path> --wiki # build agent-crawlable wiki (index.md + one article per community)
/graphify <path> --obsidian --obsidian-dir ~/vaults/my-project # write vault to custom path (e.g. existing vault)
/graphify add <url> # fetch URL, save to ./raw, update graph
/graphify add <url> --author "Name" # tag who wrote it
/graphify add <url> --contributor "Name" # tag who added it to the corpus
/graphify query "<question>" # BFS traversal - broad context
/graphify query "<question>" --dfs # DFS - trace a specific path
/graphify query "<question>" --budget 1500 # cap answer at N tokens
/graphify path "AuthModule" "Database" # shortest path between two concepts
/graphify explain "SwinTransformer" # plain-language explanation of a node
```
## What graphify is for
Drop any folder of code, docs, papers, images, or video into graphify and get a queryable knowledge graph. Persistent across sessions, honest audit trail (EXTRACTED/INFERRED/AMBIGUOUS), community detection surfaces cross-document connections you wouldn't think to ask about.
## What You Must Do When Invoked
If the user invoked `/graphify --help` or `/graphify -h` (with no other arguments), print the contents of the `## Usage` section above verbatim and stop. Do not run any commands, do not detect files, do not default the path to `.`. Just print the Usage block and return.
**Fast path — existing graph:** Before doing anything else, check whether `graphify-out/graph.json` exists. The expected location is `graphify-out/graph.json` relative to the **current working directory** (i.e. the project root where you are running commands). If it exists AND the user's request is a natural-language question about the codebase (e.g. "How does X work?", "What calls Y?", "Trace the data flow through Z") and NOT an explicit rebuild command (`--update`, `--cluster-only`, or a bare path/URL that implies fresh extraction): **skip Steps 1–5 entirely and jump straight to `## For /graphify query`.** Run `graphify query "<question>"` immediately. Do not run detect. Do not check corpus size. Do not ask the user to narrow. The graph is already built — use it.
If no path was given, use `.` (current directory). Do not ask the user for a path.
If the path argument starts with `https://github.com/` or `http://github.com/`, treat it as a GitHub URL - run Step 0 before anything else, then continue with the resolved local path.
Follow these steps in order. Do not skip steps.
### Step 0 - GitHub repos and multi-path merge (only if a URL or several paths)
Only when the path is one or more `https://github.com/...` URLs, or several local subfolders to merge. See `references/github-and-merge.md` for the clone, cross-repo merge, and monorepo flow, then continue with the resolved local path. A plain local path skips this step.
### Step 1 - Ensure graphify is installed
```bash
# Detect the correct Python interpreter (handles uv tool, pipx, venv, system installs)
PYTHON=""
GRAPHIFY_BIN=$(which graphify 2>/dev/null)
# 1. uv tool installs — most reliable on modern Mac/Linux
if [ -z "$PYTHON" ] && command -v uv >/dev/null 2>&1; then
_UV_PY=$(uv tool run --from graphifyy python -c "import sys; print(sys.executable)" 2>/dev/null)
if [ -n "$_UV_PY" ]; then PYTHON="$_UV_PY"; fi
fi
# 2. Read shebang from graphify binary (pipx and direct pip installs)
if [ -z "$PYTHON" ] && [ -n "$GRAPHIFY_BIN" ]; then
_SHEBANG=$(head -1 "$GRAPHIFY_BIN" | tr -d '#!')
case "$_SHEBANG" in
*[!a-zA-Z0-9/_.@-]*) ;;
*) "$_SHEBANG" -c "import graphify" 2>/dev/null && PYTHON="$_SHEBANG" ;;
esac
fi
# 3. Fall back to python3
if [ -z "$PYTHON" ]; then PYTHON="python3"; fi
if ! "$PYTHON" -c "import graphify" 2>/dev/null; then
if command -v uv >/dev/null 2>&1; then
uv tool install --upgrade graphifyy -q 2>&1 | tail -3
_UV_PY=$(uv tool run --from graphifyy python -c "import sys; print(sys.executable)" 2>/dev/null)
if [ -n "$_UV_PY" ]; then PYTHON="$_UV_PY"; fi
else
"$PYTHON" -m pip install graphifyy -q 2>/dev/null \
|| "$PYTHON" -m pip install graphifyy -q --break-system-packages 2>&1 | tail -3
fi
fi
# Write interpreter path for all subsequent steps (persists across invocations)
mkdir -p graphify-out
"$PYTHON" -c "import sys; open('graphify-out/.graphify_python', 'w', encoding='utf-8').write(sys.executable)"
# Save scan root so `graphify update` (no args) knows where to look next time
echo "$(cd INPUT_PATH && pwd)" > graphify-out/.graphify_root
```
If the import succeeds, print nothing and move straight to Step 2.
**In every subsequent bash block, replace `python3` with `$(cat graphify-out/.graphify_python)` to use the correct interpreter.**
### Step 2 - Detect files
```bash
$(cat graphify-out/.graphify_python) -c "
import json
from graphify.detect import detect
from pathlib import Path
result = detect(Path('INPUT_PATH'))
print(json.dumps(result, ensure_ascii=False))
" > graphify-out/.graphify_detect.json
```
Replace INPUT_PATH with the actual path the user provided. Do NOT cat or print the JSON - read it silently and present a clean summary instead:
```
Corpus: X files · ~Y words
code: N files (.py .ts .go ...)
docs: N files (.md .txt ...)
papers: N files (.pdf ...)
images: N files
video: N files (.mp4 .mp3 ...)
```
Omit any category with 0 files from the summary.
Then act on it:
- If `total_files` is 0: stop with "No supported files found in [path]."
- If `skipped_sensitive` is non-empty: mention file count skipped, not the file names.
- If `total_words` > 2,000,000 OR `total_files` > 500: show the warning. Then compute the top 5 first-level subdirectories by file count:
- Read `scan_root` from the detect JSON (always an absolute path to the resolved INPUT_PATH).
- Concatenate all file lists across all types (`code`, `document`, `paper`, `image`, `video`).
- Filter out any path that starts with `scan_root + "/graphify-out/"` to exclude converted sidecars.
- For each file, strip the `scan_root` prefix and take the first path component. Files directly in `scan_root` with no subdirectory count as `(root)`.
- If all files are in `(root)` with no subdirectories, do not ask to narrow — no subfolders exist. Instead suggest `--no-cluster` to skip the expensive clustering step and proceed.
- Otherwise rank by count, show the top 5 with file counts, then ask which subfolder to run on. Wait for the user's answer before proceeding.
- Otherwise: proceed directly to Step 2.5 if video files were detected, or Step 3 if not.
### Step 2.5 - Video and audio (only if video files detected)
Skip this step entirely if `detect` returned zero `video` files. When the corpus has video or audio, see `references/transcribe.md` to transcribe them to text first, then treat the transcripts as doc files in Step 3.
### Step 3 - Extract entities and relationships
**Before starting:** note whether `--mode deep` was given. You must pass `DEEP_MODE=true` to every subagent in Step B2 if it was. Track this from the original invocation - do not lose it.
This step has two parts: **structural extraction** (deterministic, free) and **semantic extraction** (LLM, costs tokens).
> **graphify needs no API key. Never ask the user for one, and never block on one.** Code is extracted structurally (AST) with no LLM and no key at all — a code-only corpus (the common `/graphify .` on a repo) skips semantic extraction entirely, so it needs nothing here: go straight to Part A and skip Part B. Semantic extraction (only for docs, papers, and images) uses Gemini **only if** `GEMINI_API_KEY`/`GOOGLE_API_KEY` is already set; otherwise the host agent itself is the LLM. graphify does **not** read `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, or any other provider key. If you catch yourself about to prompt for, wait on, or stop because of a missing API key, that is a misread of this skill — proceed without one.
**Before semantic extraction:** check whether `GEMINI_API_KEY` or `GOOGLE_API_KEY` is set. If neither is set, print this one-liner to the user:
> Tip: set `GEMINI_API_KEY` or `GOOGLE_API_KEY` to use Gemini for semantic extraction (`pip install 'graphifyy[gemini]'`).
Print it once, then continue — do not wait for the user to supply a key. If `GEMINI_API_KEY` or `GOOGLE_API_KEY` IS set, use `graphify.llm.extract_corpus_parallel(files, backend="gemini")` for semantic extraction instead of dispatching subagents. The default Gemini model is `gemini-3-flash-preview`; set `GRAPHIFY_GEMINI_MODEL` or pass `--model` in headless CLI flows to override it.
> **No other API keys are read.** When `GEMINI_API_KEY`/`GOOGLE_API_KEY` are unset, semantic extraction falls to the host agent itself — the running session is the LLM. On a host that dispatches subagents (e.g. Claude Code), dispatch them as written in Part B. On a host that runs the CLI directly in a terminal and cannot dispatch subagents, do not stall: a code-only corpus has no semantic work, so write the empty semantic file (Part B "Fast path") and continue to Part C; for a corpus with docs/papers/images, either set a Gemini key or extract those inline yourself, but in no case prompt for `ANTHROPIC_API_KEY` — that prompt is a misread of this skill.
**Run Part A (AST) and Part B (semantic) in parallel. Dispatch all semantic subagents AND start AST extraction in the same message. Both can run simultaneously since they operate on different file types. Merge results in Part C as before.**
Note: Parallelizing AST + semantic saves 5-15s on large corpora. AST is deterministic and fast; start it while subagents are processing docs/papers.
#### Part A - Structural extraction for code files
For any code files detected, run AST extraction in parallel with Part B subagents:
```bash
$(cat graphify-out/.graphify_python) -c "
import sys, json
from graphify.extract import collect_files, extract
from pathlib import Path
import json
code_files = []
detect = json.loads(Path('graphify-out/.graphify_detect.json').read_text(encoding=\"utf-8\"))
for f in detect.get('files', {}).get('code', []):
code_files.extend(collect_files(Path(f)) if Path(f).is_dir() else [Path(f)])
if code_files:
result = extract(code_files, cache_root=Path('INPUT_PATH'))
Path('graphify-out/.graphify_ast.json').write_text(json.dumps(result, indent=2, ensure_ascii=False), encoding=\"utf-8\")
print(f'AST: {len(result[\"nodes\"])} nodes, {len(result[\"edges\"])} edges')
else:
Path('graphify-out/.graphify_ast.json').write_text(json.dumps({'nodes':[],'edges':[],'input_tokens':0,'output_tokens':0}, ensure_ascii=False), encoding=\"utf-8\")
print('No code files - skipping AST extraction')
"
```
#### Part B - Semantic extraction (parallel subagents)
**Fast path:** If detection found zero docs, papers, and images (code-only corpus), skip Part B entirely and go straight to Part C. AST handles code - there is nothing for semantic subagents to do. **First write an empty semantic file** so Part C's merge has its input (it reads `.graphify_semantic.json` unconditionally; without this a code-only run hits `FileNotFoundError`):
```bash
$(cat graphify-out/.graphify_python) -c "
import json
from pathlib import Path
Path('graphify-out/.graphify_semantic.json').write_text(json.dumps({'nodes':[],'edges':[],'hyperedges':[],'input_tokens':0,'output_tokens':0}), encoding='utf-8')
"
```
**MANDATORY: You MUST use the Agent tool here. Reading files yourself one-by-one is forbidden - it is 5-10x slower. If you do not use the Agent tool you are doing this wrong.**
Before dispatching subagents, print a timing estimate:
- Load `total_words` and file counts from `graphify-out/.graphify_detect.json`
- Estimate agents needed: `ceil(uncached_non_code_files / 22)` (chunk size is 20-25)
- Estimate time: ~45s per agent batch (they run in parallel, so total ≈ 45s × ceil(agents/parallel_limit))
- Print: "Semantic extraction: ~N files → X agents, estimated ~Ys"
**Step B0 - Check extraction cache first**
Before dispatching any subagents, check which files already have cached extraction results:
```bash
$(cat graphify-out/.graphify_python) -c "
import json
from graphify.cache import check_semantic_cache
from pathlib import Path
detect = json.loads(Path('graphify-out/.graphify_detect.json').read_text(encoding=\"utf-8\"))
# Only content files go to semantic extraction. Code is already covered structurally
# by the AST pass (Part A); flattening every category here makes subagents re-read
# every source file (#1392). Video is transcribed to a document in Step 2.5 first.
all_files = [f for cat in ('document', 'paper', 'image') for f in detect['files'].get(cat, [])]
cached_nodes, cached_edges, cached_hyperedges, uncached = check_semantic_cache(all_files, root='INPUT_PATH')
# Always (re)write the cache file: write hits, else DELETE any leftover from a prior
# run so Part C never merges a stale .graphify_cached.json (#1392).
if cached_nodes or cached_edges or cached_hyperedges:
Path('graphify-out/.graphify_cached.json').write_text(json.dumps({'nodes': cached_nodes, 'edges': cached_edges, 'hyperedges': cached_hyperedges}, ensure_ascii=False), encoding=\"utf-8\")
else:
Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True)
Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\")
print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction')
"
```
Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly.
**Step B1 - Split into chunks**
Load files from `graphify-out/.graphify_uncached.txt`. Split into chunks of 20-25 files each. Each image gets its own chunk (vision needs separate context). When splitting, group files from the same directory together so related artifacts land in the same chunk and cross-file relationships are more likely to be extracted.
**Step B2 - Dispatch ALL subagents in a single message**
Call the Agent tool multiple times IN THE SAME RESPONSE - one call per chunk. This is the only way they run in parallel. If you make one Agent call, wait, then make another, you are doing it sequentially and defeating the purpose.
**IMPORTANT - subagent type:** Always use `subagent_type="general-purpose"`. Do NOT use `Explore` - it is read-only and cannot write chunk files to disk, which silently drops extraction results. General-purpose has Write and Bash access which the subagent needs.
Concrete example for 3 chunks:
```
[Agent tool call 1: files 1-15, subagent_type="general-purpose"]
[Agent tool call 2: files 16-30, subagent_type="general-purpose"]
[Agent tool call 3: files 31-45, subagent_type="general-purpose"]
```
All three in one message. Not three separate messages.
Each subagent receives this exact prompt (substitute FILE_LIST, CHUNK_NUM, TOTAL_CHUNKS, DEEP_MODE, and CHUNK_PATH).
CHUNK_PATH must be an **absolute** path — derive it before dispatching:
```bash
PROJECT_ROOT=$(pwd) # cwd — where Part C globs graphify-out/ (NOT .graphify_root/scan dir, #1392)
# Then for chunk N: CHUNK_PATH="${PROJECT_ROOT}/graphify-out/.graphify_chunk_0N.json"
```
Subagent prompt template:
See `references/extraction-spec.md` for the exact subagent prompt (JSON schema, node-ID rules, confidence rubric, frontmatter, hyperedge, and vision rules). Load it only here, only when at least one chunk holds a doc, paper, or image; a pure-code corpus has skipped Part B and never reads it. Pass each subagent that prompt verbatim with FILE_LIST, CHUNK_NUM, TOTAL_CHUNKS, DEEP_MODE, and CHUNK_PATH substituted, and have it write the result to CHUNK_PATH.
**Step B3 - Collect, cache, and merge**
Wait for all subagents. For each result:
- Check that `graphify-out/.graphify_chunk_NN.json` exists on disk — this is the success signal
- If the file exists and contains valid JSON with `nodes` and `edges`, include it and save to cache
- If the file is missing, the subagent was likely dispatched as read-only (Explore type) — print a warning: "chunk N missing from disk — subagent may have been read-only. Re-run with general-purpose agent." Do not silently skip.
- If a subagent failed or returned invalid JSON, print a warning and skip that chunk - do not abort
If more than half the chunks failed or are missing, stop and tell the user to re-run and ensure `subagent_type="general-purpose"` is used.
Merge all chunk files into `.graphify_semantic_new.json`. **After each Agent call completes, read the real token counts from the Agent tool result's `usage` field and write them back into the chunk JSON before merging** — the chunk JSON itself always has placeholder zeros. Then run:
```bash
$(cat graphify-out/.graphify_python) -c "
import json, glob
from pathlib import Path
chunks = sorted(glob.glob('graphify-out/.graphify_chunk_*.json'))
all_nodes, all_edges, all_hyperedges = [], [], []
total_in, total_out = 0, 0
for c in chunks:
d = json.loads(Path(c).read_text(encoding=\"utf-8\"))
all_nodes += d.get('nodes', [])
all_edges += d.get('edges', [])
all_hyperedges += d.get('hyperedges', [])
total_in += d.get('input_tokens', 0)
total_out += d.get('output_tokens', 0)
Path('graphify-out/.graphify_semantic_new.json').write_text(json.dumps({
'nodes': all_nodes, 'edges': all_edges, 'hyperedges': all_hyperedges,
'input_tokens': total_in, 'output_tokens': total_out,
}, indent=2, ensure_ascii=False), encoding=\"utf-8\")
print(f'Merged {len(chunks)} chunks: {total_in:,} in / {total_out:,} out tokens')
"
```
Save new results to cache:
```bash
$(cat graphify-out/.graphify_python) -c "
import json
from graphify.cache import save_semantic_cache
from pathlib import Path
new = json.loads(Path('graphify-out/.graphify_semantic_new.json').read_text(encoding=\"utf-8\")) if Path('graphify-out/.graphify_semantic_new.json').exists() else {'nodes':[],'edges':[],'hyperedges':[]}
uncached = [line for line in Path('graphify-out/.graphify_uncached.txt').read_text(encoding=\"utf-8\").splitlines() if line]
saved = save_semantic_cache(new.get('nodes', []), new.get('edges', []), new.get('hyperedges', []), root='INPUT_PATH', allowed_source_files=uncached)
print(f'Cached {saved} files')
"
```
Merge cached + new results into `graphify-out/.graphify_semantic.json`:
```bash
$(cat graphify-out/.graphify_python) -c "
import json
from pathlib import Path
cached = json.loads(Path('graphify-out/.graphify_cached.json').read_text(encoding=\"utf-8\")) if Path('graphify-out/.graphify_cached.json').exists() else {'nodes':[],'edges':[],'hyperedges':[]}
new = json.loads(Path('graphify-out/.graphify_semantic_new.json').read_text(encoding=\"utf-8\")) if Path('graphify-out/.graphify_semantic_new.json').exists() else {'nodes':[],'edges':[],'hyperedges':[]}
all_nodes = cached['nodes'] + new.get('nodes', [])
all_edges = cached['edges'] + new.get('edges', [])
all_hyperedges = cached.get('hyperedges', []) + new.get('hyperedges', [])
seen = set()
deduped = []
for n in all_nodes:
if n['id'] not in seen:
seen.add(n['id'])
deduped.append(n)
merged = {
'nodes': deduped,
'edges': all_edges,
'hyperedges': all_hyperedges,
'input_tokens': new.get('input_tokens', 0),
'output_tokens': new.get('output_tokens', 0),
}
Path('graphify-out/.graphify_semantic.json').write_text(json.dumps(merged, indent=2, ensure_ascii=False), encoding=\"utf-8\")
print(f'Extraction complete - {len(deduped)} nodes, {len(all_edges)} edges ({len(cached[\"nodes\"])} from cache, {len(new.get(\"nodes\",[]))} new)')
"
```
Clean up temp files: `rm -f graphify-out/.graphify_cached.json graphify-out/.graphify_uncached.txt graphify-out/.graphify_semantic_new.json`
#### Part C - Merge AST + semantic into final extraction
```bash
$(cat graphify-out/.graphify_python) -c "
import sys, json
from pathlib import Path
ast = json.loads(Path('graphify-out/.graphify_ast.json').read_text(encoding=\"utf-8\"))
sem = json.loads(Path('graphify-out/.graphify_semantic.json').read_text(encoding=\"utf-8\"))
# Merge: AST nodes first, semantic nodes deduplicated by id
seen = {n['id'] for n in ast['nodes']}
merged_nodes = list(ast['nodes'])
for n in sem['nodes']:
if n['id'] not in seen:
merged_nodes.append(n)
seen.add(n['id'])
merged_edges = ast['edges'] + sem['edges']
merged_hyperedges = sem.get('hyperedges', [])
merged = {
'nodes': merged_nodes,
'edges': merged_edges,
'hyperedges': merged_hyperedges,
'input_tokens': sem.get('input_tokens', 0),
'output_tokens': sem.get('output_tokens', 0),
}
Path('graphify-out/.graphify_extract.json').write_text(json.dumps(merged, indent=2, ensure_ascii=False), encoding=\"utf-8\")
total = len(merged_nodes)
edges = len(merged_edges)
print(f'Merged: {total} nodes, {edges} edges ({len(ast[\"nodes\"])} AST + {len(sem[\"nodes\"])} semantic)')
"
```
### Step 4 - Build graph, cluster, analyze, generate outputs
**Before starting:** the code blocks below pass `directed=IS_DIRECTED` to `build_from_json()`. Replace `IS_DIRECTED` with `True` if `--directed` was given (builds a `DiGraph` preserving edge direction source→target), otherwise `False` (the default undirected `Graph`). Substitute it the same way you substitute `INPUT_PATH` — do not leave the literal `IS_DIRECTED` in the code.
```bash
mkdir -p graphify-out
$(cat graphify-out/.graphify_python) -c "
import sys, json
from graphify.build import build_from_json
from graphify.cluster import cluster, score_all
from graphify.analyze import god_nodes, surprising_connections, suggest_questions
from graphify.report import generate
from graphify.export import to_json
from pathlib import Path
extraction = json.loads(Path('graphify-out/.graphify_extract.json').read_text(encoding=\"utf-8\"))
detection = json.loads(Path('graphify-out/.graphify_detect.json').read_text(encoding=\"utf-8\"))
# root= mirrors the --update runbook (#1361): relativize source_file to the same
# base so the full build and incremental --update never drift apart on re-extract.
G = build_from_json(extraction, root='INPUT_PATH', directed=IS_DIRECTED)
# Guard BEFORE any write: an empty extraction must not clobber a good graph.json /
# GRAPH_REPORT.md / analysis sidecar. Check immediately after build (#1392).
if G.number_of_nodes() == 0:
print('ERROR: Graph is empty - extraction produced no nodes.')
print('Possible causes: all files were skipped, binary-only corpus, or extraction failed.')
raise SystemExit(1)
communities = cluster(G)
cohesion = score_all(G, communities)
tokens = {'input': extraction.get('input_tokens', 0), 'output': extraction.get('output_tokens', 0)}
gods = god_nodes(G)
surprises = surprising_connections(G, communities)
labels = {cid: 'Community ' + str(cid) for cid in communities}
# Placeholder questions - regenerated with real labels in Step 5
questions = suggest_questions(G, communities, labels)
# Export FIRST and honor the #479 shrink-guard: to_json returns False (writing
# nothing) when the new graph is smaller than the existing graph.json. Only write
# GRAPH_REPORT.md + the analysis sidecar when the graph was actually written, so
# they never describe a graph that graph.json doesn't contain (#1392).
wrote = to_json(G, communities, 'graphify-out/graph.json')
if not wrote:
print('ERROR: refused to shrink graphify-out/graph.json (existing graph has more nodes; #479).')
print('If this shrink is intentional (you deleted files), re-run a full build with --force.')
raise SystemExit(1)
report = generate(G, communities, cohesion, labels, gods, surprises, detection, tokens, 'INPUT_PATH', suggested_questions=questions)
Path('graphify-out/GRAPH_REPORT.md').write_text(report, encoding=\"utf-8\")
analysis = {
'communities': {str(k): v for k, v in communities.items()},
'cohesion': {str(k): v for k, v in cohesion.items()},
'gods': gods,
'surprises': surprises,
'questions': questions,
}
Path('graphify-out/.graphify_analysis.json').write_text(json.dumps(analysis, indent=2, ensure_ascii=False), encoding=\"utf-8\")
print(f'Graph: {G.number_of_nodes()} nodes, {G.number_of_edges()} edges, {len(communities)} communities')
"
```
If this step prints `ERROR: Graph is empty`, stop and tell the user what happened - do not proceed to labeling or visualization.
Replace INPUT_PATH with the actual path.
### Step 4.5 - Graph health check (read-only integrity gate)
A non-destructive diagnostic on the extraction, before labeling. It surfaces edge collapse, dangling/missing endpoints, and self-loops — the silent-corruption modes of incremental updates and AST/LLM id mismatches. Read-only; never aborts.
```bash
$(cat graphify-out/.graphify_python) -c "
import json
from pathlib import Path
from graphify.diagnostics import diagnose_extraction, format_diagnostic_report
extraction = json.loads(Path('graphify-out/.graphify_extract.json').read_text(encoding=\"utf-8\"))
summary = diagnose_extraction(extraction, directed=IS_DIRECTED, root='INPUT_PATH')
print(format_diagnostic_report(summary))
flags = [f'{summary[k]} {label}' for k, label in (
('dangling_endpoint_edges', 'dangling-endpoint edges'),
('missing_endpoint_edges', 'missing-endpoint edges'),
('self_loop_edges', 'self-loop edges'),
('directed_same_endpoint_collapsed_edges', 'collapsed (directed) edges'),
('undirected_same_endpoint_collapsed_edges', 'collapsed (undirected) edges'),
) if summary.get(k, 0)]
print('GRAPH HEALTH WARNING: ' + '; '.join(flags) + ' - graph may be incomplete/corrupt.' if flags else 'Graph health: OK (no dangling/missing/collapsed edges).')
"
```
Substitute `IS_DIRECTED` and `INPUT_PATH` as in Step 4. If a `GRAPH HEALTH WARNING` prints, surface it in the final summary (do not abort — the graph is still usable, but the integrity issue must be visible, per the Honesty Rules).
### Step 5 - Label communities
Read `graphify-out/.graphify_analysis.json`. For each community key, look at its node labels and write a 2-5 word plain-language name (e.g. "Attention Mechanism", "Training Pipeline", "Data Loading").
Then regenerate the report and save the labels for the visualizer:
```bash
$(cat graphify-out/.graphify_python) -c "
import sys, json
from graphify.build import build_from_json
from graphify.cluster import score_all
from graphify.analyze import god_nodes, surprising_connections, suggest_questions
from graphify.report import generate
from pathlib import Path
extraction = json.loads(Path('graphify-out/.graphify_extract.json').read_text(encoding=\"utf-8\"))
detection = json.loads(Path('graphify-out/.graphify_detect.json').read_text(encoding=\"utf-8\"))
analysis = json.loads(Path('graphify-out/.graphify_analysis.json').read_text(encoding=\"utf-8\"))
# root= as in Step 4 / the --update runbook (#1361) — same base for node-key parity.
G = build_from_json(extraction, root='INPUT_PATH', directed=IS_DIRECTED)
communities = {int(k): v for k, v in analysis['communities'].items()}
cohesion = {int(k): v for k, v in analysis['cohesion'].items()}
tokens = {'input': extraction.get('input_tokens', 0), 'output': extraction.get('output_tokens', 0)}
# LABELS - replace these with the names you chose above
labels = LABELS_DICT
# Regenerate questions with real community labels (labels affect question phrasing)
questions = suggest_questions(G, communities, labels)
report = generate(G, communities, cohesion, labels, analysis['gods'], analysis['surprises'], detection, tokens, 'INPUT_PATH', suggested_questions=questions)
Path('graphify-out/GRAPH_REPORT.md').write_text(report, encoding=\"utf-8\")
Path('graphify-out/.graphify_labels.json').write_text(json.dumps({str(k): v for k, v in labels.items()}, ensure_ascii=False), encoding=\"utf-8\")
print('Report updated with community labels')
"
```
Replace `LABELS_DICT` with the actual dict you constructed (e.g. `{0: "Attention Mechanism", 1: "Training Pipeline"}`).
Replace INPUT_PATH with the actual path.
### Step 6 - Generate Obsidian vault (opt-in) + HTML
**Generate HTML always** (unless `--no-viz`). **Obsidian vault only if `--obsidian` was explicitly given** — skip it otherwise, it generates one file per node.
If `--obsidian` was given:
- If `--obsidian-dir <path>` was also given, pass it via `--dir`. Otherwise defaults to `graphify-out/obsidian`.
```bash
graphify export obsidian
# or with custom dir: graphify export obsidian --dir ~/vaults/my-project
```
Generate the HTML graph (always, unless `--no-viz`):
```bash
graphify export html # auto-aggregates to community view if graph > 5000 nodes
# or: graphify export html --no-viz
```
### Steps 6b-8 - Wiki, Neo4j, FalkorDB, SVG, GraphML, MCP, benchmark (only on their flags)
These run only when their flag is present (`--wiki`, `--neo4j`/`--neo4j-push`, `--falkordb`/`--falkordb-push`, `--svg`, `--graphml`, `--mcp`) or, for the token-reduction benchmark, when `total_words` exceeds 5,000. A default run with no export flags skips all of them. See `references/exports.md` for each one. Run any `--wiki` export before Step 9 cleanup so `.graphify_labels.json` is still available.
---
### Step 9 - Save manifest, update cost tracker, clean up, and report
```bash
$(cat graphify-out/.graphify_python) -c "
import json
from pathlib import Path
from datetime import datetime, timezone
from graphify.detect import save_manifest
# Save manifest for --update
detect = json.loads(Path('graphify-out/.graphify_detect.json').read_text(encoding=\"utf-8\"))
# In --update mode, 'all_files' carries the full corpus; 'files' is the changed
# subset. Full-rebuild mode populates only 'files', so the fallback handles that.
# root= relativizes the manifest keys to the scan root (same base as the build),
# so the on-disk manifest is portable across clones/machines and a later --update
# matches cached files instead of missing every one (#1417).
save_manifest(detect.get('all_files') or detect['files'], root='INPUT_PATH')
# Update cumulative cost tracker
extract = json.loads(Path('graphify-out/.graphify_extract.json').read_text(encoding=\"utf-8\"))
input_tok = extract.get('input_tokens', 0)
output_tok = extract.get('output_tokens', 0)
cost_path = Path('graphify-out/cost.json')
if cost_path.exists():
cost = json.loads(cost_path.read_text(encoding=\"utf-8\"))
else:
cost = {'runs': [], 'total_input_tokens': 0, 'total_output_tokens': 0}
cost['runs'].append({
'date': datetime.now(timezone.utc).isoformat(),
'input_tokens': input_tok,
'output_tokens': output_tok,
'files': detect.get('total_files', 0),
})
cost['total_input_tokens'] += input_tok
cost['total_output_tokens'] += output_tok
cost_path.write_text(json.dumps(cost, indent=2, ensure_ascii=False), encoding=\"utf-8\")
print(f'This run: {input_tok:,} input tokens, {output_tok:,} output tokens')
print(f'All time: {cost[\"total_input_tokens\"]:,} input, {cost[\"total_output_tokens\"]:,} output ({len(cost[\"runs\"])} runs)')
"
rm -f graphify-out/.graphify_detect.json graphify-out/.graphify_extract.json graphify-out/.graphify_ast.json graphify-out/.graphify_semantic.json graphify-out/.graphify_analysis.json
find graphify-out -maxdepth 1 -name '.graphify_chunk_*.json' -delete 2>/dev/null
rm -f graphify-out/.needs_update 2>/dev/null || true
```
Replace INPUT_PATH with the actual path (same value used in Steps 4-5) so the manifest is relativized to the scan root.
Tell the user (omit the obsidian line unless --obsidian was given):
```
Graph complete. Outputs in PATH_TO_DIR/graphify-out/
graph.html - interactive graph, open in browser
GRAPH_REPORT.md - audit report
graph.json - raw graph data
obsidian/ - Obsidian vault (only if --obsidian was given)
```
If graphify saved you time, consider supporting it: https://github.com/sponsors/safishamsi
Replace PATH_TO_DIR with the actual absolute path of the directory that was processed.
Then paste these sections from GRAPH_REPORT.md directly into the chat:
- God Nodes
- Surprising Connections
- Suggested Questions
Do NOT paste the full report - just those three sections. Keep it concise.
Then immediately offer to explore. Pick the single most interesting suggested question from the report - the one that crosses the most community boundaries or has the most surprising bridge node - and ask:
> "The most interesting question this graph can answer: **[question]**. Want me to trace it?"
If the user says yes, run `/graphify query "[question]"` on the graph and walk them through the answer using the graph structure - which nodes connect, which community boundaries get crossed, what the path reveals. Keep going as long as they want to explore. Each answer should end with a natural follow-up ("this connects to X - want to go deeper?") so the session feels like navigation, not a one-shot report.
The graph is the map. Your job after the pipeline is to be the guide.
---
## Interpreter guard for subcommands
Before running any subcommand below (`--update`, `--cluster-only`, `query`, `path`, `explain`, `add`), check that `.graphify_python` exists. If it's missing (e.g. user deleted `graphify-out/`), re-resolve the interpreter first:
```bash
if [ ! -f graphify-out/.graphify_python ]; then
GRAPHIFY_BIN=$(which graphify 2>/dev/null)
if [ -n "$GRAPHIFY_BIN" ]; then
PYTHON=$(head -1 "$GRAPHIFY_BIN" | tr -d '#!')
case "$PYTHON" in *[!a-zA-Z0-9/_.@-]*) PYTHON="python3" ;; esac
else
PYTHON="python3"
fi
mkdir -p graphify-out
"$PYTHON" -c "import sys; open('graphify-out/.graphify_python', 'w', encoding='utf-8').write(sys.executable)"
fi
```
## For --update and --cluster-only
Both are non-default subcommands. `--update` re-extracts only new or changed files; `--cluster-only` reruns clustering on the existing graph. See `references/update.md` for both flows.
---
## For /graphify query
When `graphify-out/graph.json` already exists and the user asks a question about the corpus, answer from the graph rather than rebuilding it:
```bash
graphify query "<question>"
```
Before traversal, expand the question against the graph's own vocabulary so a wording mismatch does not collapse the answer to noise. If the `graphify query` CLI is unavailable, fall back to an inline NetworkX traversal of `graphify-out/graph.json`. Answer using only what the graph output contains, and quote `source_location` when citing a specific fact. For that vocab-expansion step, the BFS/DFS traversal modes, the `--budget` cap, the NetworkX fallback, `save-result` feedback, and the `/graphify path` and `/graphify explain` flows, see `references/query.md`.
---
## For /graphify add and --watch
Neither is part of the default build. When the user runs `/graphify add <url>` to fetch a URL into the corpus, or passes `--watch` to auto-rebuild on file changes, see `references/add-watch.md`.
---
## For the commit hook and native CLAUDE.md integration
When the user asks to install the post-commit auto-rebuild hook or wire graphify into a project's CLAUDE.md, see `references/hooks.md`.
---
## Honesty Rules
- Never invent an edge. If unsure, use AMBIGUOUS.
- Never skip the corpus check warning.
- Always show token cost in the report.
- Never hide cohesion scores behind symbols - show the raw number.
- Never run HTML viz on a graph with more than 5,000 nodes without warning the user.
-56
View File
@@ -1,56 +0,0 @@
# graphify reference: add a URL and watch a folder
Load this when the user ran `/graphify add <url>` or passed `--watch`. Neither is part of the default build.
## For /graphify add
Fetch a URL and add it to the corpus, then update the graph.
```bash
$(cat graphify-out/.graphify_python) -c "
import sys
from graphify.ingest import ingest
from pathlib import Path
try:
out = ingest('URL', Path('./raw'), author='AUTHOR', contributor='CONTRIBUTOR')
print(f'Saved to {out}')
except ValueError as e:
print(f'error: {e}', file=sys.stderr)
sys.exit(1)
except RuntimeError as e:
print(f'error: {e}', file=sys.stderr)
sys.exit(1)
"
```
Replace `URL` with the actual URL, `AUTHOR` with the user's name if provided, `CONTRIBUTOR` likewise. If the command exits with an error, tell the user what went wrong - do not silently continue. After a successful save, automatically run the `--update` pipeline on `./raw` to merge the new file into the existing graph.
Supported URL types (auto-detected):
- YouTube / any video URL → audio downloaded via yt-dlp, transcribed to `.txt` on next run (requires `pip install 'graphifyy[video]'`)
- Twitter/X → fetched via oEmbed, saved as `.md` with tweet text and author
- arXiv → abstract + metadata saved as `.md`
- PDF → downloaded as `.pdf`
- Images (.png/.jpg/.webp) → downloaded, Claude vision extracts on next run
- Any webpage → converted to markdown via html2text
---
## For --watch
Start a background watcher that monitors a folder and auto-updates the graph when files change.
```bash
$(cat graphify-out/.graphify_python) -m graphify.watch INPUT_PATH --debounce 3
```
Replace INPUT_PATH with the folder to watch. Behavior depends on what changed:
- **Code files only (.py, .ts, .go, etc.):** re-runs AST extraction + rebuild + cluster immediately, no LLM needed. `graph.json` and `GRAPH_REPORT.md` are updated automatically.
- **Docs, papers, or images:** writes a `graphify-out/needs_update` flag and prints a notification to run `/graphify --update` (LLM semantic re-extraction required).
Debounce (default 3s): waits until file activity stops before triggering, so a wave of parallel agent writes doesn't trigger a rebuild per file.
Press Ctrl+C to stop.
For agentic workflows: run `--watch` in a background terminal. Code changes from agent waves are picked up automatically between waves. If agents are also writing docs or notes, you'll need a manual `/graphify --update` after those waves.
-87
View File
@@ -1,87 +0,0 @@
# graphify reference: extra exports and benchmark
Load this when the user passed one of the export flags (`--wiki`, `--neo4j`, `--neo4j-push`, `--falkordb`, `--falkordb-push`, `--svg`, `--graphml`, `--mcp`), or when the corpus is large enough for the token-reduction benchmark. Each step runs only for its own flag.
### Step 6b - Wiki (only if --wiki flag)
**Only run this step if `--wiki` was explicitly given in the original command.**
Run this before Step 9 (cleanup) so `.graphify_labels.json` is still available.
```bash
graphify export wiki
```
### Step 7 - Neo4j export (only if --neo4j or --neo4j-push flag)
**If `--neo4j`** - generate a Cypher file for manual import:
```bash
graphify export neo4j
```
**If `--neo4j-push <uri>`** - push directly to a running Neo4j instance. Ask the user for credentials if not provided:
```bash
graphify export neo4j --push bolt://localhost:7687 --user neo4j --password PASSWORD
```
Default URI is `bolt://localhost:7687`, default user is `neo4j`. Uses MERGE - safe to re-run without creating duplicates.
### Step 7a - FalkorDB export (only if --falkordb or --falkordb-push flag)
**If `--falkordb`** - generate a Cypher file. The statements are OpenCypher, but FalkorDB's `GRAPH.QUERY` runs one statement at a time (no bulk script import like Neo4j's `cypher-shell`), so prefer `--falkordb-push` to load a graph. Use this only when you want the portable `cypher.txt` artifact:
```bash
graphify export falkordb
```
**If `--falkordb-push <uri>`** - push directly to a running FalkorDB instance. Credentials are optional; ask the user only if the instance requires auth:
```bash
graphify export falkordb --push falkordb://localhost:6379
```
Default URI is `falkordb://localhost:6379` (the scheme is informational - `redis://` or a bare `host:port` work too), auth is optional, and the target graph defaults to `graphify`. Uses MERGE - safe to re-run without creating duplicates.
### Step 7b - SVG export (only if --svg flag)
```bash
graphify export svg
```
### Step 7c - GraphML export (only if --graphml flag)
```bash
graphify export graphml
```
### Step 7d - MCP server (only if --mcp flag)
```bash
$(cat graphify-out/.graphify_python) -m graphify.serve graphify-out/graph.json
```
This starts a stdio MCP server that exposes tools: `query_graph`, `get_node`, `get_neighbors`, `get_community`, `god_nodes`, `graph_stats`, `shortest_path`. Add to Claude Desktop or any MCP-compatible agent orchestrator so other agents can query the graph live.
To configure in Claude Desktop, add to `claude_desktop_config.json`. Claude Desktop can't run `$(...)`, and under `uv tool install` the system `python3` can't import graphify — so set `command` to the **absolute interpreter path** printed by `cat graphify-out/.graphify_python`:
```json
{
"mcpServers": {
"graphify": {
"command": "<absolute path from: cat graphify-out/.graphify_python>",
"args": ["-m", "graphify.serve", "/absolute/path/to/graphify-out/graph.json"]
}
}
}
```
### Step 8 - Token reduction benchmark (only if total_words > 5000)
If `total_words` from `graphify-out/.graphify_detect.json` is greater than 5,000, run:
```bash
graphify benchmark
```
Print the output directly in chat. If `total_words <= 5000`, skip silently - the graph value is structural clarity, not token compression, for small corpora.
@@ -1,70 +0,0 @@
# graphify reference: extraction subagent prompt
Load this in Step 3 Part B when the corpus has at least one doc, paper, or image chunk. A pure-code corpus skips Part B and never reads this file. Each semantic subagent receives the prompt below verbatim (substitute FILE_LIST, CHUNK_NUM, TOTAL_CHUNKS, DEEP_MODE, and CHUNK_PATH).
```
You are a graphify extraction subagent. Read the files listed and extract a knowledge graph fragment.
Output ONLY valid JSON matching the schema below - no explanation, no markdown fences, no preamble.
Files (chunk CHUNK_NUM of TOTAL_CHUNKS):
FILE_LIST
Rules:
- EXTRACTED: relationship explicit in source (import, call, citation, "see §3.2")
- INFERRED: reasonable inference (shared data structure, implied dependency)
- AMBIGUOUS: uncertain - flag for review, do not omit
Code files: focus on semantic edges AST cannot find (call relationships, shared data, arch patterns).
Do not re-extract imports - AST already has those.
Doc/paper files: extract named concepts, entities, citations. For rationale (WHY decisions were made, trade-offs, design intent): store as a `rationale` attribute on the relevant concept node — do NOT create a separate rationale node or fragment node. Only create a node for something that is itself a named entity or concept. Use `file_type:"rationale"` for concept-like nodes (ideas, principles, mechanisms, design patterns). `file_type` MUST be one of exactly these six values: `code`, `document`, `paper`, `image`, `rationale`, `concept`. Any other value is invalid and will be rejected.
Code files: when adding `calls` edges, source MUST be the caller (the function/class doing the calling), target MUST be the callee. Never reverse this direction. `calls` edges MUST stay within one language: a Python function cannot `calls` a JS/TS/Go/Rust/Java symbol and vice versa — cross-language call edges are phantom artifacts, never emit them.
Image files: use vision to understand what the image IS - do not just OCR.
UI screenshot: layout patterns, design decisions, key elements, purpose.
Chart: metric, trend/insight, data source.
Tweet/post: claim as node, author, concepts mentioned.
Diagram: components and connections.
Research figure: what it demonstrates, method, result.
Handwritten/whiteboard: ideas and arrows, mark uncertain readings AMBIGUOUS.
DEEP_MODE (if --mode deep was given): be aggressive with INFERRED edges - indirect deps,
shared assumptions, latent couplings. Mark uncertain ones AMBIGUOUS instead of omitting.
Semantic similarity: if two concepts in this chunk solve the same problem or represent the same idea without any structural link (no import, no call, no citation), add a `semantically_similar_to` edge marked INFERRED with a confidence_score reflecting how similar they are (0.6-0.95). Examples:
- Two functions that both validate user input but never call each other
- A class in code and a concept in a paper that describe the same algorithm
- Two error types that handle the same failure mode differently
Only add these when the similarity is genuinely non-obvious and cross-cutting. Do not add them for trivially similar things.
Hyperedges: if 3 or more nodes clearly participate together in a shared concept, flow, or pattern that is not captured by pairwise edges alone, add a hyperedge to a top-level `hyperedges` array. Examples:
- All classes that implement a common protocol or interface
- All functions in an authentication flow (even if they don't all call each other)
- All concepts from a paper section that form one coherent idea
Use sparingly — only when the group relationship adds information beyond the pairwise edges. Maximum 3 hyperedges per chunk.
If a file has YAML frontmatter (--- ... ---), copy source_url, captured_at, author,
contributor onto every node from that file.
confidence_score is REQUIRED on every edge - never omit it, never use 0.5 as a default:
- EXTRACTED edges: confidence_score = 1.0 always
- INFERRED edges: pick exactly ONE value from this set — never 0.5:
0.95 direct structural evidence (shared data structure, named cross-file reference).
0.85 strong inference (clear functional alignment, no direct symbol link).
0.75 reasonable inference (shared problem domain + similar shape, requires interpretation).
0.65 weak inference (thematically related, no shape evidence).
0.55 speculative but plausible (surface-level co-occurrence only).
Models follow discrete rubrics better than continuous ranges; the bimodal
distribution observed in production (>50% at 0.5, >40% at 0.85+) shows the
range guidance is being collapsed to a binary. If no value above fits, mark
the edge AMBIGUOUS rather than picking 0.4 or below.
- AMBIGUOUS edges: 0.1-0.3
Node ID format: lowercase, only `[a-z0-9_]`, no dots or slashes. Format: `{stem}_{entity}` where stem is the **full repo-relative path with the extension dropped**, every path segment kept and joined with `_` (each segment lowercased with non-alphanumeric chars replaced by `_`), and entity is the symbol name similarly normalized. Use every directory level, not just the immediate parent — this keeps same-named files in different directories distinct. Examples: `src/auth/session.py` + `ValidateToken` → `src_auth_session_validatetoken`; `lib/utils/helpers.py` + `parse_url` → `lib_utils_helpers_parse_url`; `tests/test_foo.py` + `_helper` → `tests_test_foo_helper`; `docs/v1/api/README.md` + `getUser` → `docs_v1_api_readme_getuser`. Top-level files (no parent dir, e.g. `setup.py`) use just the filename stem: `setup_my_func`. This must match the ID the AST extractor generates — using just the filename (e.g., `session_validatetoken`) or only the immediate parent (e.g., `auth_session_validatetoken`) will create orphan ghost-duplicate nodes. If you are re-extracting a project built under the old immediate-parent format, the user should run `graphify extract --force` to rebuild cleanly. CRITICAL: never append chunk numbers, sequence numbers, or any suffix to an ID (no `_c1`, `_c2`, `_chunk2`, etc.). IDs must be deterministic from the label alone — the same entity must always produce the same ID regardless of which chunk processes it.
Generate the extraction JSON matching this schema exactly:
{"nodes":[{"id":"auth_session_validatetoken","label":"Human Readable Name","file_type":"code|document|paper|image|rationale|concept","source_file":"<FILE_LIST path verbatim>","source_location":null,"source_url":null,"captured_at":null,"author":null,"contributor":null}],"edges":[{"source":"node_id","target":"node_id","relation":"calls|implements|references|cites|conceptually_related_to|shares_data_with|semantically_similar_to|rationale_for","confidence":"EXTRACTED|INFERRED|AMBIGUOUS","confidence_score":1.0,"source_file":"<FILE_LIST path verbatim>","source_location":null,"weight":1.0}],"hyperedges":[{"id":"snake_case_id","label":"Human Readable Label","nodes":["node_id1","node_id2","node_id3"],"relation":"participate_in|implement|form","confidence":"EXTRACTED|INFERRED","confidence_score":0.75,"source_file":"<FILE_LIST path verbatim>"}],"input_tokens":0,"output_tokens":0}
source_file RULE (every node, edge, and hyperedge): set source_file to the path of the originating file EXACTLY as it appears in FILE_LIST — verbatim and absolute. Do NOT shorten to a basename, do NOT re-relativize, do NOT strip any directory prefix, and do NOT change separators (the engine canonicalizes separators and relativizes against the build root downstream). Copy the FILE_LIST entry character-for-character. This keeps the full build and incremental --update on the same base, so build_merge's replace-on-re-extract matches the existing node instead of accumulating a duplicate.
Then write the JSON to disk using the Write tool at this exact absolute path (no relative paths — Write resolves relative paths against an undefined cwd and the file will be silently lost):
CHUNK_PATH
```
@@ -1,46 +0,0 @@
# graphify reference: GitHub clone and cross-repo merge
Load this when the user passed one or more `https://github.com/...` URLs, or named several local subfolders to merge into one graph.
### Step 0 - Clone GitHub repo(s) (only if a GitHub URL was given)
**Single repo:**
```bash
LOCAL_PATH=$(graphify clone <github-url> [--branch <branch>])
# Use LOCAL_PATH as the target for all subsequent steps
```
**Multiple repos (cross-repo graph):**
```bash
# Clone each repo, run the full pipeline on each, then merge
graphify clone <url1> # → ~/.graphify/repos/<owner1>/<repo1>
graphify clone <url2> # → ~/.graphify/repos/<owner2>/<repo2>
# Run /graphify on each local path to produce their graph.json files
# Then merge:
graphify merge-graphs \
~/.graphify/repos/<owner1>/<repo1>/graphify-out/graph.json \
~/.graphify/repos/<owner2>/<repo2>/graphify-out/graph.json \
--out graphify-out/cross-repo-graph.json
```
Graphify clones into `~/.graphify/repos/<owner>/<repo>` and reuses existing clones on repeat runs. Each node in the merged graph carries a `repo` attribute so you can filter by origin.
**Multiple local subfolders (monorepo or multi-service layout):**
The skill pipeline writes all intermediate and final outputs to `graphify-out/` in the current working directory. Running the skill on each subfolder separately will clobber the same output dir. Instead, use the CLI directly for each subfolder — it places `graphify-out/` *inside* the scanned path:
```bash
graphify extract ./core/ # → ./core/graphify-out/graph.json
graphify extract ./service/ # → ./service/graphify-out/graph.json
graphify extract ./platform/ # → ./platform/graphify-out/graph.json
# Add --backend gemini|kimi|openai|deepseek|claude-cli depending on which API key you have set
# Then merge at the project root:
graphify merge-graphs \
./core/graphify-out/graph.json \
./service/graphify-out/graph.json \
./platform/graphify-out/graph.json \
--out graphify-out/graph.json
```
Once `graphify-out/graph.json` exists, the fast path above takes over: any codebase question runs `graphify query` directly on the merged graph — no re-extraction, no size gate.
-33
View File
@@ -1,33 +0,0 @@
# graphify reference: commit hook and native CLAUDE.md integration
Load this when the user asked to install the post-commit hook or wire graphify into a project's CLAUDE.md.
## For git commit hook
Install a post-commit hook that auto-rebuilds the graph after every commit. No background process needed - triggers once per commit, works with any editor.
```bash
graphify hook install # install
graphify hook uninstall # remove
graphify hook status # check
```
After every `git commit`, the hook detects which code files changed (via `git diff HEAD~1`), re-runs AST extraction on those files, and rebuilds `graph.json` and `GRAPH_REPORT.md`. Doc/image changes are ignored by the hook - run `/graphify --update` manually for those.
If a post-commit hook already exists, graphify appends to it rather than replacing it.
---
## For native CLAUDE.md integration
Run once per project to make graphify always-on in Claude Code sessions:
```bash
graphify claude install
```
This writes a `## graphify` section to the local `CLAUDE.md` that instructs Claude to check the graph before answering codebase questions and rebuild it after code changes. No manual `/graphify` needed in future sessions.
```bash
graphify claude uninstall # remove the section
```
-311
View File
@@ -1,311 +0,0 @@
# graphify reference: query, path, explain
Load this when the user asks a question against an existing graph, or runs `/graphify path` or `/graphify explain`. The core's query stub points here for the full traversal flow. These flows use the `graphify query` CLI when it is available and fall back to an inline NetworkX traversal otherwise.
Two traversal modes - choose based on the question:
| Mode | Flag | Best for |
|------|------|----------|
| BFS (default) | _(none)_ | "What is X connected to?" - broad context, nearest neighbors first |
| DFS | `--dfs` | "How does X reach Y?" - trace a specific chain or dependency path |
First check the graph exists:
```bash
$(cat graphify-out/.graphify_python) -c "
from pathlib import Path
if not Path('graphify-out/graph.json').exists():
print('ERROR: No graph found. Run /graphify <path> first to build the graph.')
raise SystemExit(1)
"
```
If it fails, stop and tell the user to run `/graphify <path>` first.
### Step 0 — Constrained query expansion (REQUIRED before traversal)
graphify's `query` CLI matches nodes via case-folded substring + IDF — there is **no stemming, no synonyms, no cross-language match** inside the binary, and the inline fallback below matches the same way. If the user's question uses different language or different domain vocabulary than the graph's labels (user says "обработчик" / graph says "handler"; user says "authentication" / graph says "Guardian"), the literal matcher returns 0 hits and the answer collapses to noise.
Fix this **without inventing tokens** by expanding the query against the actual graph vocabulary first:
1. Extract the token vocabulary from node labels:
```bash
$(cat graphify-out/.graphify_python) -c "
import json, re
from pathlib import Path
data = json.loads(Path('graphify-out/graph.json').read_text(encoding='utf-8'))
vocab = set()
for n in data['nodes']:
for c in re.findall(r'[^\W\d_]+', n.get('label','') or '', re.UNICODE):
parts = re.findall(r'[A-Z]+(?=[A-Z][a-z])|[A-Z]?[a-z]+|[A-Z]+', c) or [c]
for p in parts:
t = p.lower()
if 3 <= len(t) <= 30:
vocab.add(t)
Path('graphify-out/.vocab.txt').write_text('\n'.join(sorted(vocab)), encoding='utf-8')
print(f'vocab: {len(vocab)} tokens')
"
```
2. Read `graphify-out/.vocab.txt`. Then for the user's question, select **up to 12 tokens from this exact list** that semantically match the query intent. Hard constraints:
- You MUST pick only tokens present in the vocabulary file. Do NOT invent tokens.
- If a query concept has no plausible token in the vocab, skip it — do not substitute a near-synonym from training memory.
- If **no** vocab tokens match the query at all, output an empty list and tell the user the corpus has no relevant vocabulary for this question. Do not fabricate a search.
- Translate cross-language: Russian "аутентификация" → look for `auth`, `credential`, `token`, `security` IFF present in vocab.
- Morphology: "handlers" maps to `handler` IFF present; "todos" maps to `todo` IFF present.
3. Print the selection explicitly to the user before running the query, so the expansion is auditable:
```
Query expanded to (from graph vocab, N tokens): [token1, token2, ...]
```
If the list is empty, say so plainly and stop — do not proceed to traversal.
### Step 1 — Traversal
Build the **expanded query string** by joining the selected tokens with spaces. Use this string as `QUESTION` below — NOT the original user question. (The original question is preserved only for `save-result` at the end.)
Prefer the CLI when it is installed:
```bash
graphify query "QUESTION"
# or: graphify query "QUESTION" --dfs --budget 3000
```
If the CLI is unavailable, load `graphify-out/graph.json` and run the traversal inline:
1. Find the 1-3 nodes whose label best matches the expanded tokens.
2. Run the appropriate traversal from each starting node.
3. Read the subgraph - node labels, edge relations, confidence tags, source locations.
4. Answer using **only** what the graph contains. Quote `source_location` when citing a specific fact.
5. If the graph lacks enough information, say so - do not hallucinate edges.
```bash
$(cat graphify-out/.graphify_python) -c "
import sys, json
from networkx.readwrite import json_graph
import networkx as nx
from pathlib import Path
data = json.loads(Path('graphify-out/graph.json').read_text(encoding='utf-8'))
G = json_graph.node_link_graph(data, edges='links')
question = 'QUESTION'
mode = 'MODE' # 'bfs' or 'dfs'
terms = [t.lower() for t in question.split() if len(t) >= 3] # match the vocab threshold; keeps api/jwt/ios (#1392)
# Find best-matching start nodes
scored = []
for nid, ndata in G.nodes(data=True):
label = ndata.get('label', '').lower()
score = sum(1 for t in terms if t in label)
if score > 0:
scored.append((score, nid))
scored.sort(reverse=True)
start_nodes = [nid for _, nid in scored[:3]]
if not start_nodes:
print('No matching nodes found for query terms:', terms)
sys.exit(0)
subgraph_nodes = set()
subgraph_edges = []
if mode == 'dfs':
# DFS: follow one path as deep as possible before backtracking.
# Depth-limited to 6 to avoid traversing the whole graph.
visited = set()
stack = [(n, 0) for n in reversed(start_nodes)]
while stack:
node, depth = stack.pop()
if node in visited or depth > 6:
continue
visited.add(node)
subgraph_nodes.add(node)
for neighbor in G.neighbors(node):
if neighbor not in visited:
stack.append((neighbor, depth + 1))
subgraph_edges.append((node, neighbor))
else:
# BFS: explore all neighbors layer by layer up to depth 3.
frontier = set(start_nodes)
subgraph_nodes = set(start_nodes)
for _ in range(3):
next_frontier = set()
for n in frontier:
for neighbor in G.neighbors(n):
if neighbor not in subgraph_nodes:
next_frontier.add(neighbor)
subgraph_edges.append((n, neighbor))
subgraph_nodes.update(next_frontier)
frontier = next_frontier
# Token-budget aware output: rank by relevance, cut at budget (~4 chars/token)
token_budget = BUDGET # default 2000
char_budget = token_budget * 4
# Score each node by term overlap for ranked output
def relevance(nid):
label = G.nodes[nid].get('label', '').lower()
return sum(1 for t in terms if t in label)
ranked_nodes = sorted(subgraph_nodes, key=relevance, reverse=True)
lines = [f'Traversal: {mode.upper()} | Start: {[G.nodes[n].get(\"label\",n) for n in start_nodes]} | {len(subgraph_nodes)} nodes']
for nid in ranked_nodes:
d = G.nodes[nid]
lines.append(f' NODE {d.get(\"label\", nid)} [src={d.get(\"source_file\",\"\")} loc={d.get(\"source_location\",\"\")}]')
for u, v in subgraph_edges:
if u in subgraph_nodes and v in subgraph_nodes:
_raw = G[u][v]; d = next(iter(_raw.values()), {}) if isinstance(G, nx.MultiGraph) else _raw
lines.append(f' EDGE {G.nodes[u].get(\"label\",u)} --{d.get(\"relation\",\"\")} [{d.get(\"confidence\",\"\")}]--> {G.nodes[v].get(\"label\",v)}')
output = '\n'.join(lines)
if len(output) > char_budget:
output = output[:char_budget] + f'\n... (truncated at ~{token_budget} token budget - use --budget N for more)'
print(output)
"
```
Replace `QUESTION` with the **expanded** query string, `MODE` with `bfs` or `dfs`, and `BUDGET` with the token budget (default `2000`, or whatever `--budget N` specifies). Then answer based on the subgraph output above, using only what the graph contains.
After writing the answer, save it back into the graph so it improves future queries. Include the expanded tokens inside the `--answer` text (e.g. `"Expanded from original query via vocab: [tokens]. Then traversed..."`) so the next `--update` extracts the expansion history as a graph node:
```bash
$(cat graphify-out/.graphify_python) -m graphify save-result --question "ORIGINAL_QUESTION" --answer "ANSWER" --type query --nodes NODE1 NODE2
```
Replace `ORIGINAL_QUESTION` with the user's verbatim question, `ANSWER` with your full answer text (containing the expanded-token trace), `NODE1 NODE2` with the list of node labels you cited. This closes the feedback loop: the next `--update` will extract this Q&A as a node in the graph.
**Work memory (self-improving loop).** Add an `--outcome` so future sessions learn from this one — append `--outcome useful|dead_end|corrected` to the `save-result` command (and `--correction "the right answer"` when correcting):
- `useful` — the cited nodes answered the question well (they become *preferred sources*).
- `dead_end` — the question/path led nowhere; don't re-derive it next time.
- `corrected` — the saved answer was wrong; `--correction` records what was right.
At the **start** of graph work, refresh and read the lessons: run `graphify reflect --if-stale` (cheap, deterministic, no LLM; `--if-stale` makes it a no-op when `LESSONS.md` is already newer than every input, e.g. when the git hook just refreshed it), then read `graphify-out/reflections/LESSONS.md`. It lists **preferred sources** (start there), **known dead ends** (skip them), and prior **corrections**. Running `reflect` yourself keeps the lessons current even without the git hook installed; if the post-commit hook *is* installed, `--if-stale` means your session-start run costs almost nothing.
---
## For /graphify path
Find the shortest path between two named concepts in the graph. Prefer the CLI when installed:
```bash
graphify path "NODE_A" "NODE_B"
```
If the CLI is unavailable, run it inline:
```bash
$(cat graphify-out/.graphify_python) -c "
import json, sys
import networkx as nx
from networkx.readwrite import json_graph
from pathlib import Path
data = json.loads(Path('graphify-out/graph.json').read_text(encoding='utf-8'))
G = json_graph.node_link_graph(data, edges='links')
a_term = 'NODE_A'
b_term = 'NODE_B'
def find_node(term):
term = term.lower()
scored = sorted(
[(sum(1 for w in term.split() if w in G.nodes[n].get('label','').lower()), n)
for n in G.nodes()],
reverse=True
)
return scored[0][1] if scored and scored[0][0] > 0 else None
src = find_node(a_term)
tgt = find_node(b_term)
if not src or not tgt:
print(f'Could not find nodes matching: {a_term!r} or {b_term!r}')
sys.exit(0)
try:
path = nx.shortest_path(G, src, tgt)
print(f'Shortest path ({len(path)-1} hops):')
for i, nid in enumerate(path):
label = G.nodes[nid].get('label', nid)
if i < len(path) - 1:
_raw = G[nid][path[i+1]]; edge = next(iter(_raw.values()), {}) if isinstance(G, nx.MultiGraph) else _raw
rel = edge.get('relation', '')
conf = edge.get('confidence', '')
print(f' {label} --{rel}--> [{conf}]')
else:
print(f' {label}')
except nx.NetworkXNoPath:
print(f'No path found between {a_term!r} and {b_term!r}')
except nx.NodeNotFound as e:
print(f'Node not found: {e}')
"
```
Replace `NODE_A` and `NODE_B` with the actual concept names from the user. Then explain the path in plain language - what each hop means, why it's significant.
After writing the explanation, save it back:
```bash
$(cat graphify-out/.graphify_python) -m graphify save-result --question "Path from NODE_A to NODE_B" --answer "ANSWER" --type path_query --nodes NODE_A NODE_B
```
---
## For /graphify explain
Give a plain-language explanation of a single node - everything connected to it. Prefer the CLI when installed:
```bash
graphify explain "NODE_NAME"
```
If the CLI is unavailable, run it inline:
```bash
$(cat graphify-out/.graphify_python) -c "
import json, sys
import networkx as nx
from networkx.readwrite import json_graph
from pathlib import Path
data = json.loads(Path('graphify-out/graph.json').read_text(encoding='utf-8'))
G = json_graph.node_link_graph(data, edges='links')
term = 'NODE_NAME'
term_lower = term.lower()
# Find best matching node
scored = sorted(
[(sum(1 for w in term_lower.split() if w in G.nodes[n].get('label','').lower()), n)
for n in G.nodes()],
reverse=True
)
if not scored or scored[0][0] == 0:
print(f'No node matching {term!r}')
sys.exit(0)
nid = scored[0][1]
data_n = G.nodes[nid]
print(f'NODE: {data_n.get(\"label\", nid)}')
print(f' source: {data_n.get(\"source_file\",\"unknown\")}')
print(f' type: {data_n.get(\"file_type\",\"unknown\")}')
print(f' degree: {G.degree(nid)}')
print()
print('CONNECTIONS:')
for neighbor in G.neighbors(nid):
_raw = G[nid][neighbor]; edge = next(iter(_raw.values()), {}) if isinstance(G, nx.MultiGraph) else _raw
nlabel = G.nodes[neighbor].get('label', neighbor)
rel = edge.get('relation', '')
conf = edge.get('confidence', '')
src_file = G.nodes[neighbor].get('source_file', '')
print(f' --{rel}--> {nlabel} [{conf}] ({src_file})')
"
```
Replace `NODE_NAME` with the concept the user asked about. Then write a 3-5 sentence explanation of what this node is, what it connects to, and why those connections are significant. Use the source locations as citations.
After writing the explanation, save it back:
```bash
$(cat graphify-out/.graphify_python) -m graphify save-result --question "Explain NODE_NAME" --answer "ANSWER" --type explain --nodes NODE_NAME
```
-52
View File
@@ -1,52 +0,0 @@
# graphify reference: transcribe video and audio
Load this only when `detect` reported one or more `video` files. A corpus with no video never reads this.
### Step 2.5 - Transcribe video / audio files (only if video files detected)
Skip this step entirely if `detect` returned zero `video` files.
Video and audio files cannot be read directly. Transcribe them to text first, then treat the transcripts as doc files in Step 3.
**Strategy:** Read the god nodes from `graphify-out/.graphify_detect.json` (or the analysis file if it exists from a previous run). You are already a language model — write a one-sentence domain hint yourself from those labels. Then pass it to Whisper as the initial prompt. No separate API call needed.
**However**, if the corpus has *only* video files and no other docs/code, use the generic fallback prompt: `"Use proper punctuation and paragraph breaks."`
**Step 1 - Write the Whisper prompt yourself.**
Read the top god node labels from detect output or analysis, then compose a short domain hint sentence, for example:
- Labels: `transformer, attention, encoder, decoder` → `"Machine learning research on transformer architectures and attention mechanisms. Use proper punctuation and paragraph breaks."`
- Labels: `kubernetes, deployment, pod, helm` → `"DevOps discussion about Kubernetes deployments and Helm charts. Use proper punctuation and paragraph breaks."`
**Export** it as `GRAPHIFY_WHISPER_PROMPT` (the exact name the transcriber reads — and it must be `export`ed so the child Python process sees it) for the next command.
**Step 2 - Transcribe:**
```bash
export GRAPHIFY_WHISPER_MODEL=base # or whatever --whisper-model the user passed (must be exported)
export GRAPHIFY_WHISPER_PROMPT="<the one-sentence domain hint you composed in Step 1>"
$(cat graphify-out/.graphify_python) -c "
import json, os, sys
from pathlib import Path
from graphify.transcribe import transcribe_all
detect = json.loads(Path('graphify-out/.graphify_detect.json').read_text(encoding=\"utf-8\"))
video_files = detect.get('files', {}).get('video', [])
prompt = os.environ.get('GRAPHIFY_WHISPER_PROMPT', 'Use proper punctuation and paragraph breaks.')
transcript_paths = transcribe_all(video_files, initial_prompt=prompt)
# Write the JSON from Python (NOT a shell '>' redirect): transcribe_all/Whisper
# print progress to stdout, which would otherwise corrupt the JSON file (#1392).
Path('graphify-out/.graphify_transcripts.json').write_text(json.dumps(transcript_paths, ensure_ascii=False), encoding=\"utf-8\")
print(f'Transcribed {len(transcript_paths)} file(s)', file=sys.stderr)
"
```
After transcription:
- Read the transcript paths from `graphify-out/.graphify_transcripts.json`
- Add them to the docs list before dispatching semantic subagents in Step 3B
- Print how many transcripts were created: `Transcribed N video file(s) -> treating as docs`
- If transcription fails for a file, print a warning and continue with the rest
**Whisper model:** Default is `base`. If the user passed `--whisper-model <name>`, `export GRAPHIFY_WHISPER_MODEL=<name>` (it must be exported, not just assigned) before running the command above.
-192
View File
@@ -1,192 +0,0 @@
# graphify reference: incremental update and cluster-only
Load this only when the user passed `--update` or `--cluster-only`. A first-time full build never reads this file.
## For --update (incremental re-extraction)
Use when you've added or modified files since the last run. Only re-extracts changed files - saves tokens and time.
```bash
$(cat graphify-out/.graphify_python) -c "
import sys, json
from graphify.detect import detect_incremental, save_manifest
from pathlib import Path
result = detect_incremental(Path('INPUT_PATH'))
new_total = result.get('new_total', 0)
print(json.dumps(result, indent=2, ensure_ascii=False))
Path('graphify-out/.graphify_incremental.json').write_text(json.dumps(result, ensure_ascii=False), encoding=\"utf-8\")
deleted = list(result.get('deleted_files', []))
if new_total == 0 and not deleted:
print('No files changed since last run. Nothing to update.')
raise SystemExit(0)
if deleted:
print(f'{len(deleted)} deleted file(s) to prune.')
if new_total > 0:
print(f'{new_total} new/changed file(s) to re-extract.')
"
```
Then populate `.graphify_detect.json` so Steps 3A–6 (which read it unconditionally) see the right state for an incremental run. `files` carries the changed subset (drives Step 3A AST + Step 3B0 cache check on only what changed); `all_files` carries the full corpus for any step that needs corpus-wide context:
```bash
$(cat graphify-out/.graphify_python) -c "
import json
from pathlib import Path
r = json.loads(Path('graphify-out/.graphify_incremental.json').read_text(encoding=\"utf-8\"))
Path('graphify-out/.graphify_detect.json').write_text(json.dumps({
'files': r.get('new_files', {}),
'all_files': r.get('files', {}),
'total_files': r.get('new_total', 0),
'total_words': r.get('total_words', 0),
'skipped_sensitive': r.get('skipped_sensitive', []),
'needs_graph': True,
}, ensure_ascii=False), encoding=\"utf-8\")
"
```
If new files exist, first check whether all changed files are code files:
```bash
$(cat graphify-out/.graphify_python) -c "
import json
from pathlib import Path
result = json.loads(open('graphify-out/.graphify_incremental.json', encoding='utf-8').read()) if Path('graphify-out/.graphify_incremental.json').exists() else {}
code_exts = {'.py','.ts','.js','.go','.rs','.java','.cpp','.c','.rb','.swift','.kt','.cs','.scala','.php','.cc','.cxx','.hpp','.h','.kts','.lua','.toc','.f','.F','.f90','.F90','.f95','.F95','.f03','.F03','.f08','.F08'}
new_files = result.get('new_files', {})
all_changed = [f for files in new_files.values() for f in files]
code_only = all(Path(f).suffix.lower() in code_exts for f in all_changed)
print('code_only:', code_only)
"
```
If `code_only` is True: print `[graphify update] Code-only changes detected - skipping semantic extraction (no LLM needed)`, run only Step 3A (AST) on the changed files, skip Step 3B entirely (no subagents), then go straight to merge and Steps 4–8.
If `code_only` is False (any changed file is a doc/paper/image/video): **first, if any changed file is in `new_files['video']`, run `references/transcribe.md` (Step 2.5) on those files, then rewrite `.graphify_detect.json` to move the resulting transcript paths into `files['document']` and drop `files['video']`** — otherwise raw `.mp4/.mp3` paths are fed to semantic subagents as unreadable media (#1392). Then run the full Steps 3A–3C pipeline as normal.
If no new files exist (only deletions), create an empty extraction so the merge step can prune:
```bash
if [ ! -f graphify-out/.graphify_extract.json ]; then
echo '[graphify update] Only deletions -- creating empty extraction for merge.'
$(cat graphify-out/.graphify_python) -c "
import json
from pathlib import Path
Path('graphify-out/.graphify_extract.json').write_text(json.dumps({'nodes':[],'edges':[],'hyperedges':[],'input_tokens':0,'output_tokens':0}), encoding='utf-8')
"
fi
```
Then:
```bash
$(cat graphify-out/.graphify_python) -c "
import json
from pathlib import Path
from graphify.build import build_merge
from graphify.detect import save_manifest
# Load new extraction and incremental state
new_extraction = json.loads(Path('graphify-out/.graphify_extract.json').read_text(encoding=\"utf-8\"))
incremental = json.loads(Path('graphify-out/.graphify_incremental.json').read_text(encoding=\"utf-8\"))
deleted = list(incremental.get('deleted_files', []))
# prune_sources is ONLY for genuinely DELETED files. Changed/re-extracted files are
# handled by build_merge's replace-on-re-extract (#1344): every source_file in
# new_chunks is dropped from the base before merge, so old/stale nodes don't survive.
# Do NOT add `changed` here: with root= passed, prune_set relativizes to the same base
# as the freshly merged nodes and would DELETE the re-extracted content (#1178 is moot
# now that replace — not the dedup pass — reconciles changed files).
prune = list(deleted) or None
# Use build_merge() — reads graph.json directly without NetworkX round-trip
# so edge direction (calls, implements, imports) is always preserved (#801).
# Pass root= so prune_sources (absolute paths from detect_incremental) are
# relativized to match the graph's relative source_file values; without it
# nothing is pruned and stale nodes accumulate on every update (#1361).
# directed=IS_DIRECTED: replace IS_DIRECTED with True if --directed was given, else
# False. Without it a --directed --update silently rebuilds undirected and collapses
# reciprocal A<->B edges (#1392).
G = build_merge(
[new_extraction],
graph_path='graphify-out/graph.json',
prune_sources=prune,
root='INPUT_PATH',
directed=IS_DIRECTED,
)
print(f'[graphify update] Merged: {G.number_of_nodes()} nodes, {G.number_of_edges()} edges')
# Write merged result back to .graphify_extract.json so Step 4 sees the full graph
merged_out = {
'nodes': [{'id': n, **d} for n, d in G.nodes(data=True)],
'edges': [
# Explicit source/target last so they win over any stale attrs in d.
{**{k: val for k, val in d.items() if k not in ('_src', '_tgt', 'source', 'target')},
'source': d.get('_src', u), 'target': d.get('_tgt', v)}
for u, v, d in G.edges(data=True)
],
# G.graph["hyperedges"] holds hyperedges from both existing graph.json
# and new_extraction (build_merge combines them). Falling back to
# new_extraction only would silently drop prior-run hyperedges (#801).
'hyperedges': list(G.graph.get('hyperedges', [])),
'input_tokens': new_extraction.get('input_tokens', 0),
'output_tokens': new_extraction.get('output_tokens', 0),
}
Path('graphify-out/.graphify_extract.json').write_text(json.dumps(merged_out, ensure_ascii=False), encoding=\"utf-8\")
print(f'[graphify update] Merged extraction written ({len(merged_out[\"nodes\"])} nodes, {len(merged_out[\"edges\"])} edges)')
# Save manifest so next --update diffs against today's state, not the
# prior run's baseline (prevents ghost-node reports on subsequent updates).
# root= matches the build_merge call above so the manifest keys stay relative to
# the scan root — portable across clones/machines, so --update keeps matching
# cached files instead of missing every one after a move (#1417).
save_manifest(incremental['files'], root='INPUT_PATH')
print('[graphify update] Manifest saved.')
"
```
Then run Steps 4–8 on the merged graph as normal.
After Step 4, show the graph diff:
```bash
$(cat graphify-out/.graphify_python) -c "
import json
from graphify.analyze import graph_diff
from graphify.build import build_from_json
from networkx.readwrite import json_graph
import networkx as nx
from pathlib import Path
# Load old graph (before update) from backup written before merge
old_data = json.loads(Path('graphify-out/.graphify_old.json').read_text(encoding=\"utf-8\")) if Path('graphify-out/.graphify_old.json').exists() else None
new_extract = json.loads(Path('graphify-out/.graphify_extract.json').read_text(encoding=\"utf-8\"))
G_new = build_from_json(new_extract, directed=IS_DIRECTED)
if old_data:
G_old = json_graph.node_link_graph(old_data, edges='links')
diff = graph_diff(G_old, G_new)
print(diff['summary'])
if diff['new_nodes']:
print('New nodes:', ', '.join(n['label'] for n in diff['new_nodes'][:5]))
if diff['new_edges']:
print('New edges:', len(diff['new_edges']))
"
```
Before the merge step, save the old graph: `cp graphify-out/graph.json graphify-out/.graphify_old.json`
Clean up after: `rm -f graphify-out/.graphify_old.json`
---
## For --cluster-only
Skip Steps 1–3. Re-run clustering on the existing graph:
```bash
graphify cluster-only .
```
`graphify cluster-only .` is **self-contained**: it re-clusters, names communities, and regenerates `GRAPH_REPORT.md`, `graph.json`, and `graph.html` from the existing graph. **Do not re-run Steps 5–9** — they read intermediate files (`.graphify_extract.json`, `.graphify_detect.json`, `.graphify_analysis.json`) that a prior build's cleanup (Step 9) already deleted, so they raise `FileNotFoundError` (#1392). When it finishes, present the refreshed `GRAPH_REPORT.md` summary as usual.
+8 -6
View File
@@ -493,10 +493,10 @@ else
fi
```
Update `.harden-cache/external-scores.md` with the final SSL Labs verdict
so the HARDEN.md "External validators" table reflects it. If the user
already read HARDEN.md, they can re-run `/harden <url>` to pick up the
cached (now-READY) SSL Labs result.
Update `.harden-cache/external-scores.md` with the final SSL Labs verdict,
then edit the SSL Labs row of HARDEN.md's "External validators" table in
place — YOU do this in the main loop (the agent that wrote HARDEN.md in
STEP 1 has already exited; without this edit the late grade never lands).
---
@@ -621,8 +621,10 @@ NEXT STEPS :
Astro / Cloudflare Pages project. Use the framework-native mechanism
(next.config.js headers(), astro middleware, _headers).
- **Security headers and redirects are non-negotiable defaults of this
skill** — every public site must ship them. Flag absence as Critique,
not Moyenne.
skill** — every public site must ship them. Grade each absence at the
severity guide's level (CSP absent = Critique, HSTS/X-Frame-Options =
Haute, Referrer-Policy = Moyenne); the guide's table is authoritative —
never demote a missing default below it.
- **External validators are authoritative on live headers, not the code.**
If Observatory/SecurityHeaders/SSL Labs and the code audit disagree,
the external grade reflects the deployed production config — the code
+42 -20
View File
@@ -47,6 +47,10 @@ git log --oneline -3
as `/bugfix` (root-cause investigation, then a scoped fix)."
- Settle the proposed fix HERE — the executor cannot ask questions, so the
exact edit (what changes, in which file(s)) must be closed before dispatch.
- Then run pass B of `$HOME/.claude/lib/contract-interview.md` against that
edit: a VISIBLE / PUBLIC NAME / SCOPE choice the bug description leaves
open (which way the icon aligns, the label's wording) → ask before
dispatch. A typo or a wrong value asks nothing.
OPTIONAL — memory check (exempt by default; hotfix = obvious fix, mirror of its capitalize
skip). For a RECURRING or urgent bug only, a quick blockers-only glance may save time:
@@ -66,8 +70,9 @@ Follow `$HOME/.claude/lib/design-gate.md`:
## STEP 1.7 — CONTRACT (silent autofill)
Run `$HOME/.claude/lib/contract-interview.md` at hotfix weight: **zero
questions ever** (a hotfix is an obvious fix by definition). Autofill the
Run `$HOME/.claude/lib/contract-interview.md` at hotfix weight: pass A is a
silent autofill (a hotfix is an obvious fix by definition); pass B already
ran at STEP 1, ask nothing more here. Autofill the
contract — REQUEST verbatim = the bug description as given; ACCEPTANCE
CRITERIA = "symptom gone; build/tests green"; FILE SCOPE = the 1-2 target
files from STEP 1. It writes `.claude/tasks/contracts/<date>-<slug>-<HHMM>.md`.
@@ -108,7 +113,11 @@ Snapshot current state so revert is possible:
git diff HEAD --stat # confirm working tree is clean OR carries only the
# in-progress hotfix area; if unrelated dirty files are
# present, ask user whether to stash them first
git rev-parse HEAD # capture the SHA to revert to on failure
# Snapshot the TREE STATE (incl. tolerated uncommitted edits) without touching it.
# A bare SHA is not enough: restoring to HEAD would wipe the user's own
# in-progress edits in the hotfix area.
PRE=$(git stash create "hotfix-preflight"); [ -n "$PRE" ] || PRE=$(git rev-parse HEAD)
echo "PRE=$PRE" # the revert source for every failure branch below
```
If the working tree contains unrelated uncommitted changes the user has not
@@ -131,28 +140,39 @@ security dispatch, no revert. Finish with the HOTFIX-EXEC REPORT."
Parse the `HOTFIX-EXEC REPORT`:
- `STATUS : DONE` → STEP 4 (the SMOKE line in the report decides pass/fail
there; DONE here means execution completed, not that it verified clean).
- `STATUS : BLOCKED` → if any edits were made, `git restore .` to the
pre-flight SHA (STEP 2); surface the blocker to the user; STOP. One
attempt only — hotfix never re-dispatches (escalate to `/bugfix` for
deeper work).
- `STATUS : BLOCKED` with `CLASS: visible | public-name | scope` in NOTES →
the executor halted at an open choice before editing (nothing to revert):
ask the user per MID-RUN CLARIFICATION in
`$HOME/.claude/lib/contract-interview.md`, append the answer to the
contract `[gated]`, re-dispatch ONCE with the closed choice. This is the
one re-dispatch hotfix allows; it is not a retry of a failed attempt.
- `STATUS : BLOCKED` otherwise → if any edits were made, revert ONLY the
executor's files: `git restore --source=$PRE -- <FILE(S) from the report>` and delete
any NEW file the report lists (untracked, absent from $PRE). Never
`git restore .` — it would wipe the tolerated pre-existing edits too.
Surface the blocker to the user; STOP. One attempt only — hotfix never
re-dispatches (escalate to `/bugfix` for deeper work).
## STEP 4 — VERIFY + SECURE + COMMIT (main loop, LRN-083)
1. Read the SMOKE line from the executor's report. **Failure branch** — if
it reports a failing test/build result:
- Print the failure output verbatim (under 30 lines).
- Run `git restore .` to revert the working-tree edits to the pre-flight
SHA (STEP 2). (Files were not yet staged — restore is safe.)
- Revert ONLY the executor's files: `git restore --source=$PRE --
<FILE(S) from the report>` + delete report-listed NEW files. Never
`git restore .` (wipes tolerated pre-existing edits).
- STOP and tell user: `"Hotfix introduced a regression. Reverted.
Escalate to /bugfix or /analyze for deeper investigation."`
- Do NOT commit a broken fix.
2. **Security gate (fresh auditor) — failure REVERTS, never loops.** Dispatch
a FRESH security-auditor (`subagent_type: security-auditor`, or load
`agents/security-auditor.md`) with `MODE: gate`, `SCOPE:` the working-tree
diff vs the pre-flight SHA. Parse its `SECURITY — VERDICT:` line:
a FRESH security-auditor (`subagent_type: security-auditor` — always a
fresh dispatch, never inline-load: the repo convention and the FRESH
requirement both forbid it) with `MODE: gate`, `SCOPE:` the working-tree
diff vs `$PRE`. Parse its `SECURITY — VERDICT:` line:
- `PASS` (or `DEGRADED` with no BLOCK) → proceed to commit.
- `BLOCK(n)` → this is hotfix: do NOT loop. Run `git restore .` to the
pre-flight SHA, print the `BLOCKING` list, and STOP:
- `BLOCK(n)` → this is hotfix: do NOT loop. Revert ONLY the executor's
files (`git restore --source=$PRE -- <FILE(S)>` + delete report-listed
NEW files), print the `BLOCKING` list, and STOP:
`"Hotfix introduced a security finding. Reverted. Escalate to /bugfix
for a fix under the full verify+security loop."` The hotfix model is
one attempt; any gate failure (smoke OR security) reverts and escalates.
@@ -218,13 +238,15 @@ trivial hotfix still produces a `chore(memory): journal — …` commit (Frame 2
- Reflection (LOCATE, contract, gate decisions) NEVER leaves this main
loop; execution NEVER stays in it — the executor is the sonnet-pinned
hotfixer subagent (BDR-066).
- The executor is dispatched FRESH, once — hotfix never re-dispatches (no
decision round-trips; a blocked or failed attempt reverts and escalates
to `/bugfix`, it does not retry).
- The executor is dispatched FRESH, once — hotfix never re-dispatches after
a failed or blocked attempt (it reverts and escalates to `/bugfix`, it
does not retry). Sole exception: a class-tagged BLOCKED answered by the
user (STEP 3), re-dispatched once with the closed choice.
- Design gate only if CSS/style signals detected. See STEP 1.5.
- **Revert-not-loop preserved**: smoke FAIL or security BLOCK → `git
restore .` to the pre-flight SHA + STOP + escalate to `/bugfix`; hotfix
never loops. No verifier is dispatched at hotfix weight.
- **Revert-not-loop preserved**: smoke FAIL or security BLOCK →
file-scoped revert from `$PRE` (STEP 4's protocol — never `git
restore .`) + STOP + escalate to `/bugfix`; hotfix never loops.
No verifier is dispatched at hotfix weight.
- If root cause is unclear → escalate to `/bugfix` (STEP 1).
- If fix touches >5 lines of logic → reconsider if this is
truly a hotfix.
+11 -3
View File
@@ -2,7 +2,7 @@
name: init-project
description: 'Use when initializing a brand-new project from scratch — needs interview, design, scaffold, and TDD implementation. Multi-agent orchestrator: plugin-advisor + interviewer + analyzer + scaffolder with two validation gates. Triggers: "init project", "new project", "start project from scratch", "scaffold project", "init-project".'
argument-hint: <project idea or description>
allowed-tools: Read, Write, Edit, Bash, Grep, Glob
allowed-tools: Read, Write, Edit, Bash, Grep, Glob, Agent, Skill
---
# ORCHESTRATOR: INIT PROJECT
@@ -58,8 +58,8 @@ In both cases: MANDATORY STOP until user answers remaining questions. Produce PR
**Then run `$HOME/.claude/lib/contract-interview.md`** seeded from the BRIEF:
REQUEST verbatim = the user's project description; ACCEPTANCE CRITERIA = the
V1 FEATURES (each testable); FILE SCOPE = the planned tree. No new questions
(the interview already asked). It writes
V1 FEATURES (each testable); FILE SCOPE = the planned tree. Pass A is covered
by the interview; pass B runs at STEP 3 against the DESIGN. It writes
`.claude/tasks/contracts/<date>-<slug>-<HHMM>.md`; the DESIGN approved at STEP
4 ENRICHES it, and STEP 9's verifier judges the MVP against the enriched
contract.
@@ -70,6 +70,9 @@ Load `$HOME/.claude/agents/analyzer.md`. Analyze BRIEF: existing code, stack con
## STEP 3 — DESIGN
Invoke `superpowers:brainstorming` with BRIEF + ANALYSIS REPORT.
Produce DESIGN: stack+versions, full folder tree, module responsibilities, data flow, interfaces (signatures only), config+tooling, test strategy, resolved decisions, prereqs list.
Then run pass B of `$HOME/.claude/lib/contract-interview.md` against the DESIGN
(minus what the BRIEF and the brainstorm settled): one batch before STEP 4;
answers append to the contract `[gated]`.
## STEP 4 — VALIDATION GATE #1 ★ MANDATORY STOP
Present:
@@ -227,6 +230,11 @@ If `graphify` not installed or complexity < 30% → skip silently.
Run the two fresh gates per `$HOME/.claude/lib/verify-secure-loop.md` with
`CONTRACT` = the STEP 1 path (ENRICHED at STEP 4), `DIFF` = the MVP branch
diff (`develop..HEAD`), `TEST` = the project suite:
- GATE 0 — deterministic floor, no dispatch: `bash ~/.claude/lib/gates.sh
run "$CONTRACT"` executes the criteria's declared oracles fail-closed.
UNMET → hand the dev with the NOT-MET rows verbatim — no
verifier is spent on a red floor; own budget, max 3 → escalate.
MET (an all-manual contract too) → GATE 1.
- GATE 1 — a FRESH verifier judges the MVP against the enriched contract (V1
features + `[gated]` design criteria). CONFORME → GATE 2. ECARTS → fix,
re-verify, max 3 → STOP + human escalation with the CRITERIA table.
+1 -1
View File
@@ -1,5 +1,5 @@
[
{"id": 1, "prompt": "Onboard this existing project — it's a Next.js app already deployed", "expected": "Plugin check → archetype detection (next-js-public) → monorepo gate → baseline config → interview gaps → audits pipeline → backlog at .claude/audits/ + .claude/tasks/"},
{"id": 1, "prompt": "Onboard this existing project — it's a Next.js app already deployed", "expected": "Plugin check → archetype detection (nextjs-app-router) → monorepo gate → baseline config → interview gaps → audits pipeline → backlog at .claude/audits/ + .claude/tasks/"},
{"id": 2, "prompt": "Onboard with hint: WordPress theme, force-archetype:wordpress", "expected": "Skip detection, use wordpress archetype directly, run wordpress-specific audit pipeline"},
{"id": 3, "prompt": "Onboard the apps/web package only", "expected": "Detect monorepo, present A/B/C options, accept B with package name, set PROJECT_ROOT to apps/web, run pipeline"}
]
+13
View File
@@ -130,6 +130,19 @@ Compare original PDF and translated HTML side by side:
3. Check: layout match, no missing content, images present, style fidelity
4. Fix discrepancies → iterate STEP 4
## Failure modes
| Trigger | First move | If still stuck |
|---|---|---|
| STEP 0: neither poppler nor PyMuPDF present, install fails (no sudo / no pip) | Print BOTH install commands, ask the user to run one | STOP. No degraded no-image path — the pipeline is image-based by design |
| STEP 1: PDF > 30 pages (check `pdfinfo input.pdf \| grep Pages` first) | Ask before converting: batch by section, or draft pass at `-r 150` | User declines both → STOP, oversized one-shot runs produce GB of PNGs and stall Vision |
| STEP 1: extraction yields 0 page PNGs or 0-byte files | Retry with the other tool (poppler ↔ PyMuPDF) | STOP and report the PDF as unreadable (encrypted/corrupt) — never translate from the text layer as a silent fallback |
| STEP 3: region unreadable (blur, handwriting, tiny footnote) | Mark `[illisible: <best guess>?]` inline + add to an UNCERTAIN list per page | Leave the marker in the HTML; STEP 5 QA re-reads every UNCERTAIN item at higher zoom. Never invent clean text |
| STEP 4: `/design-html` and `/frontend-design` unavailable | Write the HTML directly from the STEP 2 style brief + STEP 3 content (same requirements list) | — |
| STEP 5: no `/browse` / screenshot tool | QA on structure instead: compare HTML section order + image refs against STEP 3 layout maps | Report "visual QA skipped — structural QA only" in the final summary |
| STEP 5: QA still finds discrepancies after 2 fix iterations | Stop iterating; list residual differences for the user | User decides: accept, or target specific pages for a 3rd pass |
| `pdf-translate-work/` already exists | Ask: resume (keep PNGs, redo STEP ≥3) or clean restart | — |
## Decision: OCR vs Native PDF
```dot
+5
View File
@@ -0,0 +1,5 @@
[
{"id": 1, "prompt": "Traduis ce PDF scanné en français: ~/docs/manual-en.pdf (OCR/image-based, 6 pages)", "expected": "STEP 0 dependency check (poppler/pdftoppm), page PNGs extracted, Claude Vision read+translate+layout map, faithful HTML reconstruction, visual QA PDF-vs-HTML with fix loop"},
{"id": 2, "prompt": "Translate this 12-page PDF to English — it has embedded diagrams and a two-column layout", "expected": "Embedded images extracted and re-embedded in the HTML, two-column layout and visual style preserved, contextual translation (not word-by-word)"},
{"id": 3, "prompt": "Translate report.pdf (missing poppler AND no imagemagick on the machine)", "expected": "Detects missing dependencies at STEP 0, proposes the install command, does not silently proceed to a broken pipeline"}
]
+2 -2
View File
@@ -1,5 +1,5 @@
[
{"id": 1, "prompt": "Check active plugins for: React + FastAPI app", "expected": "Audit current plugins, recommend enable/disable based on stack signals, produce PLUGIN ADVISOR REPORT"},
{"id": 2, "prompt": "Plugin check before I start a Rust CLI project, no frontend", "expected": "Recognize CLI-only context, recommend disabling ui-ux-pro-max + context7-frontend, keep core dev tools"},
{"id": 1, "prompt": "Check active plugins for: React + FastAPI app", "expected": "Audit current plugins, recommend enable/disable based on stack signals, produce the PLUGIN CHECK block"},
{"id": 2, "prompt": "Plugin check before I start a Rust CLI project, no frontend", "expected": "Recognize CLI-only context, recommend disabling ui-ux-pro-max (and context7 if configured), keep core dev tools"},
{"id": 3, "prompt": "Audit my plugins", "expected": "No context provided → scan current dir for stack signals or ask, then produce report"}
]
+26 -14
View File
@@ -52,8 +52,7 @@ lists items + types:
| `personal` | symlink move skills/ ↔ skills-disabled/\<name\> (no prefix) |
| `external` | symlink move skills/ ↔ skills-disabled/\<name\> |
| `plugin@<marketplace>` | `claude plugin enable\|disable <name>@<marketplace>` (auto) |
| `mcp` (known: magic) | delegate to `lib/toggle-external.sh` (uses `.env`) |
| `mcp` (other) | advisory — prints manual `claude mcp add …` command |
| `mcp` | advisory — prints manual `claude mcp add …` command (no server is managed today: `MANAGED_MCPS` is empty since 21st.dev moved to a CLI) |
| `cli` | advisory only — reports installed/not-installed |
**Always-on plugins** (`security-guidance`, `superpowers`) are
@@ -62,12 +61,14 @@ protected — `set` will refuse to disable them even if the profile omits them.
`ui-ux-pro-max@ui-ux-pro-max-skill`, `plugin-dev@claude-code-plugins`,
`pr-review-toolkit@claude-code-plugins`. Other plugins are never auto-toggled.
**Managed externals** (`emil-design-eng`, `frontend-design`,
`design-motion-principles`, `impeccable`) and **managed MCPs** (`magic`)
follow the same symmetry (BDR-079): `set` enables them when the profile
lists them (from parked state, or from `skills-external/` if the symlink
never existed) and parks/unregisters them when it does not — e.g. `set
backend` after design work turns emil and magic off. `darwin-skill` and any
other unlisted external are never auto-touched. gstack works the same
`design-motion-principles`, `impeccable`, and the five 21st design skills
`21st-ui-build`, `21st-ui-explore`, `21st-ui-review`, `21st-cli-use`,
`21st-ai`) follow the same symmetry (BDR-079): `set` enables them when the
profile lists them (from parked state, or from `skills-external/` if the
symlink never existed) and parks them when it does not — e.g. `set backend`
after design work turns emil and the 21st pack off. `darwin-skill`,
`21st-registry`, `21st-design-sync` and any other unlisted external are never
auto-touched. gstack works the same
all the way down: a profile listing gstack skills while the whole pack is
off (via `toggle-external.sh`) re-enables JUST those skills on demand.
@@ -111,6 +112,17 @@ skills.
bash "$HOME/.claude/lib/profile.sh" $ARGUMENTS
```
## Failure modes
| Trigger | First move | If still stuck |
|---|---|---|
| `lib/profile.sh` absent (foreign machine, links broken) | `test -f "$HOME/.claude/lib/profile.sh"` before any verb; missing → propose `bash link.sh` from the config repo | STOP — never hand-move symlinks to emulate the script |
| Unknown profile name (rc=1, `✗ Profile not found`) | Show `list` output + the closest existing name ("`desing` → did you mean `design`?") | Let the user pick — never guess-and-`set` |
| Unknown verb (rc=1 + usage) | Re-map the request to the argument-hint verbs, retry once | Show usage, ask |
| `set`/`apply` exits nonzero MID-TOGGLE (permission, plugin CLI failure) | State may be PARTIAL. Run `current` to show what actually took; name the failed item from the script's output | Offer `reset` as recovery to a known state; never blind-rerun `set` on top of partial state |
| Plugin/MCP leg fails (marketplace/network) while symlink leg succeeded | Report the split state explicitly + print the manual `claude plugin`/`claude mcp` command for the failed leg | — |
| `current` says `none` right after a successful `set <name>` | Contradiction — do not trust either; show the raw script output to the user | Known failure family (BLK: symlink resolution in `cmd_current`) — report, don't hand-patch |
## Output policy
- After `set` / `apply` / `reset` / `gstack on|off`: show the count of skills
@@ -126,11 +138,11 @@ bash "$HOME/.claude/lib/profile.sh" $ARGUMENTS
update-check, learnings — script doesn't touch that infra. Disabled skills
are just hidden from Claude Code's scanner; the gstack repo stays installed.
- Profile changes DO toggle the managed Claude Code plugins (ui-ux-pro-max,
plugin-dev, pr-review-toolkit), the managed external packs (emil-design-eng,
frontend-design, design-motion-principles, impeccable) and the `magic` MCP —
in BOTH directions: `set` enables what the profile lists and disables the
managed leftovers it doesn't (BDR-008, BDR-079). Anything outside those
allowlists stays manual: `claude plugin enable|disable`, `claude mcp
add|remove`.
plugin-dev, pr-review-toolkit) and the managed external packs
(emil-design-eng, frontend-design, design-motion-principles, impeccable,
the 21st design skills) — in BOTH directions: `set` enables what the profile
lists and disables the managed leftovers it doesn't (BDR-008, BDR-079).
Anything outside those allowlists stays manual: `claude plugin
enable|disable`, `bash lib/toggle-external.sh enable|disable <tool>`.
- `set` is destructive in the sense that it disables non-listed gstack skills.
Use `apply` if the user wants additive behavior.

Some files were not shown because too many files have changed in this diff Show More