29 Commits
Author SHA1 Message Date
Bastien Chanot b96fab7719 chore(memory): BDR-091 BDR-092 LRN-155..157, journal 2026-09-16 2026-09-17 11:43:39 +02:00
Bastien Chanot 56bd035281 Merge feature/ask-dont-guess into develop 2026-09-17 11:41:39 +02:00
Bastien Chanot cd98bfafe4 chore: purge transient planning artifacts (BDR-065) 2026-09-17 11:41:39 +02:00
Bastien Chanot ddadca6bae Merge feature/automode-docker-node into develop 2026-09-17 11:40:27 +02:00
Bastien Chanot 22ce57f323 docs(changelog): ask-don't-guess doctrine 2026-09-16 22:15:48 +02:00
Bastien Chanot 17370d7e4c feat(executors): NEED-DECISION and BLOCKED carry a CLASS tag 2026-09-16 22:14:13 +02:00
Bastien Chanot 7cc95952bd feat(interviewer): a visible or public choice is asked, never assumed 2026-09-16 22:13:50 +02:00
Bastien Chanot a1357a6ab5 feat(ship-feature,init-project): pass B at the plan and design steps 2026-09-16 22:13:42 +02:00
Bastien Chanot 97591ca737 feat(hotfix): pass B at LOCATE, class-tagged BLOCKED relayed as a question 2026-09-16 22:13:32 +02:00
Bastien Chanot 590482b622 feat(bugfix): pass B at FIX PLAN, NEED-DECISION routed on class 2026-09-16 22:13:16 +02:00
Bastien Chanot 6d9a3497a2 feat(feat): pass B at PLAN, NEED-DECISION routed on class 2026-09-16 22:13:06 +02:00
Bastien Chanot 5f9a9c0f6b feat(rules): ask rather than guess replaces one-question-upfront 2026-09-16 22:12:56 +02:00
Bastien Chanot a2978b6f23 feat(contract): STEP 2 CLARIFY, mid-run channel, how-to-ask 2026-09-16 22:12:13 +02:00
Bastien Chanot 0d52f3a888 docs(plan): ask-don't-guess implementation plan, TODO trace 2026-09-16 22:10:24 +02:00
Bastien Chanot 5eccc3f1c4 feat(automode): docker and node framed by the classifier, ask rules retired 2026-09-16 22:03:23 +02:00
Bastien Chanot 823ce42225 chore(config): default model fable 5.1 2026-09-16 21:58:18 +02:00
Bastien Chanot 9eb69346ce docs(spec): design for the ask-don't-guess clarification doctrine 2026-09-16 20:39:15 +02:00
Bastien Chanot a449f9315f Merge chore/graphify-recovery-doc into develop 2026-09-15 19:57:07 +02:00
Bastien Chanot d7662abc1e chore(graphify): correct the recovery note, capitalize LRN-154
The previous note said a fresh clone gets the skill back from `make
plugin` without naming the command, and I picked the wrong one when the
files actually went missing. There are two, and only one restores the
skill:

  - `graphify install --platform claude` copies SKILL.md, references/
    and .graphify_version. Touches nothing else. This is the recovery
    command, verified: the skill came back at 0.9.61 and the four
    guarded configs were byte-identical afterwards.
  - `graphify claude install` writes the CLAUDE.md section and the
    .claude/settings.json PreToolUse hooks, rewrites both of those
    guarded configs (EVAL-020, reproduced today), and does NOT copy the
    skill.

LRN-154 records why the files vanished in the first place. `git rm
--cached` keeps the working file, but `gitflow finish` checks out the
target branch, where it is still tracked, so git restores it and the
merge then deletes it from disk. .gitignore does not protect it; it only
stops a re-add. The file survives the commit and dies at the merge, which
reads as unrelated.
2026-09-15 19:57:07 +02:00
Bastien Chanot e9fe79e4c2 Merge chore/graphify-gitignore-settings-prune into develop 2026-09-15 19:55:11 +02:00
Bastien Chanot 80ccdafe0e chore(config): untrack the vendored graphify skill, prune the project-local settings override
graphify: `graphify claude install` (install-plugins.sh STEP graphify)
writes SKILL.md, references/ and .graphify_version straight into the repo,
because ~/.claude/skills is a symlink to skills/. Every `pipx upgrade
graphifyy` therefore dirtied the tree and cost a `chore(graphify): sync
vendored skill X -> Y` commit. Now gitignored and untracked; a fresh clone
gets them back from `make plugin`. test-prompts.json is hand-written for
darwin and stays tracked. The accepted trade-off, documented in CLAUDE.md,
is that an upstream release can change the skill's prompt with no diff to
review.

settings.local.json (gitignored, so not in this commit) went from 14.6 KB
to 6.2 KB. It was a near-complete shadow copy of the global settings at a
higher precedence tier, which hid its own drift until the global moved.
Two entries were actively defeating BDR-090, merged an hour earlier:

  - local `deny` still carried rsync / kill -9 / killall / pkill, the four
    rules deliberately moved out of global deny. deny wins across sources,
    so autoMode.soft_deny was a dead letter in this repo.
  - local `allow` carried `sed *`, `cp *` and `python3 -`. An allow rule
    short-circuits the classifier, punching a hole through the same
    soft_deny rules.

deny and ask are dropped whole (102 and 27 of their entries duplicated the
global; ask gates nothing under defaultMode auto). allow went 185 -> 98:
81 duplicates plus six policy conflicts, the three above and
Read(//home/bchanot/**), WebSearch, and a leftover command-injection test
payload that had been allowlisted verbatim. Every non-permissions key was
a verbatim copy of the global, including a hooks block whose only original
entry pointed at hooks/config-protection.sh, a script that exists nowhere.
2026-09-15 19:55:09 +02:00
Bastien Chanot 7a861035b8 Merge feature/automode-config-alignment into develop 2026-09-15 19:44:29 +02:00
Bastien Chanot 6cd26bc3fa chore(memory): BDR-090, LRN-153, journal — autoMode tier rebuild
BDR-090 records why the ask tier was abandoned rather than repopulated,
the three alternatives rejected, and the deliberate caveat that the
guardrail hard_deny bars removing a deny entry but not adding one.

LRN-153 records the two traps the block carries: every autoMode list is
a full replacement without "$defaults", and a user-scope block reaches
every project on the machine.

TODO also logs F1-F3, found but not fixed: .claude/settings.local.json
is a 14.6 KB shadow copy of the global settings at higher precedence,
including a PreToolUse hook whose script does not exist.
2026-09-15 19:44:26 +02:00
Bastien Chanot 3b0167c6cb feat(settings): rebuild destructive-command cover in autoMode, scope the classifier environment
`permissions.ask` gates nothing under `defaultMode: auto` (LRN-146,
verified live), so the ten rules that left the static tiers had no cover
left: rsync / kill -9 / killall / pkill out of deny, and python3 -c /
python -c / xargs / sed / cp / mv out of ask.

autoMode.soft_deny (7 rules) takes over what an explicit instruction
should be able to clear: writes outside the working directory,
rsync --delete, SIGKILL and kill-by-name, in-place edits spanning more
than one file, directory moves, and inline interpreters or xargs that
delete or write outside the cwd. Intent clears a soft block for the
current turn only, stated as a rule since no setting expresses it.

autoMode.hard_deny (3 rules) takes the classes no command pattern can
express: secret exfiltration, production deployment, and disarming the
guardrails. Adding a restriction stays allowed, removing one does not.

permissions.deny gains ten .env reader rules (sed awk cut tr sort uniq
diff od xxd strings). Six of those tools sat in permissions.allow, so
reading a .env through them triggered nothing.

autoMode.environment named another project, its FTP deploy target and its
customer data, inside the file link.sh:21 symlinks to
~/.claude/settings.json, where it reached every repo and contradicted
this one's Gitea remote. Rewritten machine-generic; the project facts
moved to that project's gitignored .claude/settings.local.json. All three
lists now open with "$defaults", which the original omitted, so the
built-in classifier entries are inherited rather than replaced.

doctor.sh check_automode backstops both defects. SETTINGS.md documents
the block and a tier-choice table. README no longer claims the ask tier
makes every mcp__magic__* call require a live confirmation.
2026-09-15 19:44:24 +02:00
Bastien Chanot 3228acabfc Merge feature/gstack-playwright-lib into develop 2026-09-15 16:55:38 +02:00
Bastien Chanot 8843970425 docs: Playwright browser-cache report + bump re-applied on update 2026-09-15 16:54:33 +02:00
Bastien Chanot a0876a2976 chore(memory): BDR-088/089, LRN-150/151/152, EVAL-029 — gstack Playwright lib 2026-09-15 16:49:23 +02:00
Bastien Chanot 2cebecbb91 feat(gstack): share the Playwright bump, report the browser cache
Extract gstack_bump_playwright_if_unsupported from install-plugins.sh into
lib/gstack-playwright.sh and call it from update-all.sh too. A submodule
update no longer leaves the OS-support bump unapplied until the next
`make plugin` — that was BDR-029's open caveat.

The update helper never touches the submodule working tree: on failure it
prints git's own message and points at `make plugin`, and returns non-zero
so the existing `else warn` arm still handles it.

Add a read-only `Playwright browsers` section to doctor.sh: cache size,
which registered install requires each revision, and counts of unreferenced
directories and broken links. No pruning is written — Playwright's own
`install` already unions the required set across every registered install,
and all three installs here are live (rev 1228 for gstack + gsd-pi 1.61,
rev 1243 for gsd-pi 1.63).

Carries two latent-bug fixes from the moved code: the ostag capture exited
1 on every non-Ubuntu host and aborted the caller under inherited errexit,
and the bun calls had no timeout.
2026-09-15 14:57:31 +02:00
Bastien Chanot a53a5a26a8 Merge release/1.5.0 into develop 2026-09-13 21:26:44 +02:00
42 changed files with 1657 additions and 1630 deletions
+55
View File
@@ -97,6 +97,11 @@ rules:
| BDR-085 | 2026-08-25 | User permanent rules: writing-style always-on in rules/, web build+security path-scoped | accepted |
| BDR-086 | 2026-08-26 | darwin: threshold gates full loops; verified defects fixed regardless of unit score (paired-validated, batched checkpoint) | accepted |
| BDR-087 | 2026-09-03 | Stop hook = attention signal only, never control flow; one script for Notification + Stop | accepted |
| BDR-088 | 2026-09-15 | gstack Playwright bump shared via lib, re-applied after submodule update; update helper never touches the submodule tree | accepted |
| BDR-089 | 2026-09-15 | No Playwright browser-cache pruner; read-only doctor report — .links proved 0 bytes reclaimable | accepted |
| BDR-090 | 2026-09-15 | Destructive shell work → autoMode soft_deny/hard_deny; `ask` tier abandoned (inert under auto) | accepted |
| BDR-091 | 2026-09-16 | Ask, don't guess: open-choice sweep (3 classes) at plan step + mid-run CLASS channel supersede "one question upfront" | accepted |
| BDR-092 | 2026-09-16 | docker + node framed by the classifier via autoMode.allow + soft_deny; ask entries retired | accepted |
---
@@ -1123,3 +1128,53 @@ Branch feature/user-writing-web-rules, UNMERGED (human gate).
- **Guard vs prior refusal**: [[BDR-083]] (unlazy review, GATE 0) REFUSED a Stop hook using `decision:"block"` (forces continuation, inverts human gates). THIS Stop hook returns `terminalSequence` + `suppressOutput` only, exit 0, zero control-flow effect. Signal ≠ control. Do not read the refusal as banning Stop outright.
- **Status**: accepted.
- **Reference**: [[LRN-146]] event-coverage gap, [[BLK-020]] client-side faults, [[LRN-145]] terminalSequence pattern. Verified live: turn-end + AskUserQuestion both ring; `permission_prompt` unexercisable under `defaultMode: auto`.
---
## BDR-088 — gstack Playwright bump shared via lib; update helper never touches submodule tree
- **Date**: 2026-09-15
- **Decision**: `gstack_bump_playwright_if_unsupported` moved out of `install-plugins.sh` into `lib/gstack-playwright.sh`, sourced by install-plugins + update-all + doctor. update-all's submodule block now calls `gstack_submodule_update_with_bump`: re-applies bump after successful `submodule update --remote`; on failure prints git's own message, hints `make plugin`, returns 1. Never touches submodule worktree.
- **Why**: [[BDR-029]] caveat open — bump survived only till next `make plugin`, update path never re-checked OS support. Real gap, user-reported.
- **Alternatives rejected**: conflict-RECOVERY branch (discard package.json+bun.lock → retry → backup/restore). Withdrawn at human gate after 4-agent challenge: concentrated 3 BLOCKER + 4 MAJOR. Worst case = bump discarded, re-apply silently no-ops (bun absent / registry down — bump returns 0 on every path), `./setup` rebuilds browse against unsupported Playwright → [[BLK-008]] returns. Pre-existing behavior just failed the update and kept bump intact, so the "improvement" could regress a working install.
- **Deviations carried from "code MOVED not changed"**: `|| true` on ostag capture (line exited 1 on every non-Ubuntu host → aborted caller under inherited errexit, reproduced); `timeout` on all 3 bun calls, exit 124 → warn + no bump (TERM'd install leaves node_modules half-written, poisons the support grep).
- **Status**: accepted.
- **Reference**: commit 2cebecb, `lib/gstack-playwright.sh`. Links [[BDR-029]], [[LRN-070]], [[LRN-071]], [[LRN-150]], [[BLK-008]].
---
## BDR-089 — No Playwright browser-cache pruner; read-only doctor report instead
- **Date**: 2026-09-15
- **Decision**: `doctor.sh` gains own `── Playwright browsers ──` section — cache size, per-revision the installs requiring it, counts of unreferenced dirs + broken links. Zero deletion anywhere in the lib.
- **Why**: measured, not assumed. `~/.cache/ms-playwright/.links/` registers 3 installs — gstack 1.61.1 → rev 1228, gsd-pi nvm 1.61.0 → 1228, gsd-pi ~/.local 1.63.0 → 1243. Every dir on disk referenced → 0 bytes reclaimable. Playwright's own `_deleteStaleBrowsers` already unions across all registered installs on every `install`.
- **Alternatives rejected**: hand-rolled pruner guarded on "revision resolved by gstack's local playwright" (the originally requested shape) — that guard keeps 1228 and DELETES 1243, breaking gsd-pi. The guard was wrong, not just its implementation.
- **Status**: accepted.
- **Reference**: commit 2cebecb. Links [[LRN-151]], [[BDR-088]].
## BDR-090 — Destructive shell work → autoMode soft_deny/hard_deny; `ask` tier abandoned
- **Date**: 2026-09-15
- **Decision**: 10 rules leave the static tiers (user's own edit): `rsync` `kill -9` `killall` `pkill` out of `deny`; `python3 -c` `python -c` `xargs` `sed` `cp` `mv` out of `ask`. Cover rebuilt in `autoMode` — 7 `soft_deny` (write outside cwd, `rsync --delete`, SIGKILL/kill-by-name, in-place edit spanning >1 file, directory move, inline interpreter or `xargs` that deletes or writes outside cwd) + 3 `hard_deny` (secret exfiltration, prod deploy, disarming guardrails). Intent clears a soft block for the CURRENT TURN only — encoded as a rule line, no setting exists for it. `classifyAllShell` stays false. `permissions.deny` +10 `.env` reader rules (`sed awk cut tr sort uniq diff od xxd strings`), 6 of which sat in `allow`.
- **Why**: `ask` raises no prompt under `defaultMode: auto` ([[LRN-146]], verified live). It gated nothing, so a destructive rule moved deny→ask was a silent loosening dressed as a confirmation. `soft_deny` = the tier the classifier enforces and user intent clears. `hard_deny` = the 3 classes no command pattern can express — read-then-send spans turns, a prod target is a name not a verb, widening a deny list is self-disarming.
- **Alternatives rejected**: keep them in `ask` — inert, false sense of a gate. Back to `deny` — blocks legit process cleanup and inter-project copy, and the user works Bash-first under auto mode. `classifyAllShell: true` — closes the allow-tier blind spot but bills a classifier call on every `git status`. Published-history rewrite as `hard_deny` — user declined; `rebase` then an ordinary push stays uncovered, known gap.
- **Scope fix (same commit)**: `autoMode.environment` named `/home/bchanot/Documents/atlast`, its FTP deploy target and its customer data, inside the file `link.sh:21` symlinks to `~/.claude/settings.json`. Every project received atlast's facts, and this repo's own Gitea remote contradicted the block's "no remote configured". Global block now machine-generic; atlast facts moved to atlast's gitignored `.claude/settings.local.json`.
- **Caveat**: the guardrail `hard_deny` bars REMOVING a `deny`/`soft_deny`/`hard_deny` entry, not adding one. Future loosening goes through `/permissions` or the user's own edit — deliberate, confirmed with the user.
- **Status**: accepted.
- **Reference**: `settings.json`, `doctor.sh` `check_automode`, `templates/settings/SETTINGS.md`. Links [[LRN-153]], [[LRN-146]], [[BDR-004]].
## BDR-091 — Ask, don't guess: open-choice sweep + mid-run channel supersede "one question upfront"
- **Date**: 2026-09-16
- **Decision**: `CLAUDE.global.md` rule → ask on a VISIBLE (placement, wording, order, behavior), PUBLIC NAME (command, flag, endpoint, file) or SCOPE ("X too?") choice the request leaves open, even mid-task; class 4 (internal technical, no observable effect) never. `lib/contract-interview.md` STEP 2 = CLARIFY: pass A (gaps: outcome / scope / constraints) at contract time; pass B (open-choice sweep, 3 classes) ONCE at each flow's PLAN step, no question cap, >5 open → under-specified, list + stop; "you decide" recorded `A: delegated — <default>`, never re-asked. New MID-RUN CLARIFICATION: executor halts `NEED-DECISION` + `CLASS:` tag; visible / public-name / scope → human verbatim; internal → loop decides, max 2 round-trips. New HOW TO ASK (LRN-102: ≤4 → one AskUserQuestion, context in option descriptions; else plain text ending the turn). Wiring: feat STEP 1, bugfix STEP 3, hotfix LOCATE (pass A stays silent autofill; one re-dispatch on class-tagged BLOCKED = the sole hotfix re-dispatch), ship-feature STEP 2, init-project STEP 3; interviewer: class 1-3 item never `(assumed)`, one extra targeted question. Executors (feater, bugfixer, hotfixer) report the class. Locks: contract-verifier (9), loops-light (hotfix), gates (3).
- **Why**: gap-only trigger structurally blind to taste — "add a share icon" passes outcome / scope / constraints and the icon lands wherever the executor put it; raising the 3-question cap changes nothing. feat:153 / bugfix:165 told the orchestrator "make the decision HERE", twice, before escalating = institutional guessing. Fresh re-dispatch keeps the tree, loses the executor's reasoning → a plan-time batch costs less than the same question mid-run; the mid-run channel stays for leftovers.
- **Alternatives rejected**: bigger budget (quota was never the limiter); new `lib/clarify.md` (extra hop, STEP 2 already the mandatory passage every orchestrator runs); global rule only (skills carried explicit counter-instructions — `zero questions ever`, `make the decision HERE`, `max 3 questions` — the specific beats the general).
- **Risk watched**: chattiness. Brakes = class 4 exclusion + over-5 guard. hotfix identity (speed, silence) = the flow to watch; if pass B fires on most hotfixes the class definitions are too wide, not the flow.
- **Status**: accepted. Behavioral check OPEN: `/feat "add a share icon to the header"` must ask placement before dispatch; the fully specified variant must ask nothing → record in `evals.md`.
- **Reference**: spec + plan `docs/superpowers/{specs,plans}/2026-09-16-ask-dont-guess*` (purged at finish, in history at `22ce57f`), commits `9eb6934..22ce57f`, merge `56bd035`. Supersedes the `CLAUDE.global.md:51` rule line. Refines [[BDR-049]] (contract), applies [[LRN-102]]. Links [[LRN-157]].
## BDR-092 — docker + node framed by the classifier (`autoMode.allow`), `ask` rules retired
- **Date**: 2026-09-16
- **Decision**: `Bash(docker run|exec *)`, `Bash(docker[-| ]compose up*)`, `Bash(node -e *)` out of `permissions.ask`. New `autoMode.allow` (`$defaults` first): (1) local dev containers — `docker exec/run/compose` against a workstation container whose name lacks `prod`/`production`, running a repo SQL file or script inside, output piped; Remote Shell Writes / Production Reads / Sensitive Remote Exec scoped to sensitive-named hosts; a literal `DROP/TRUNCATE/DELETE` on the command line stays under Mass Delete. (2) project-local node — `node <file>`, `npm run`/`pnpm`/`yarn` scripts, `npx`/`pnpm exec` of a lockfile-declared package, effects in cwd. +2 `soft_deny`: docker data destruction (`rm -f`, `volume rm/prune`, `system prune`, `compose down -v`, `--privileged`, bind mount outside cwd); undeclared node packages (`npx`/`dlx` absent from lockfile, `npm install <name>`). `model` bump to fable 5.1 committed alongside.
- **Why**: real gate = built-in `Remote Shell Writes` / `Production Reads` soft_deny catching `docker exec` into `supabase_db_game`; inside the classifier `allow` = exception tier (hard_deny > soft_deny > allow > explicit intent). Static `Bash(node *)` allow is suspended under auto (wildcarded interpreter) → prose is the only lever for a conditional node permission; `awk`/`echo` statics short-circuit, `node` cannot. `ask` inert on 2.1.273 (probe, [[LRN-155]]) — retiring it is forward-safe: if the documented prompt behavior lands, those entries would prompt for exactly what should run free.
- **Alternatives rejected**: static `permissions.allow` for docker (short-circuits the classifier, framing impossible); keep the `ask` entries (inert today, wrong tomorrow); strict on every `.sql` (blocks the repo's verify scripts).
- **Trade-off accepted**: a repo SQL file runs even when its content is opaque to the classifier — local dev DB only, resettable.
- **Guardrail**: S6 (loosening = user's own edit) overridden explicitly by the user for this change; diff reviewed on the branch before merge.
- **Status**: accepted. Verified: `jq` valid; `claude auto-mode config` shows the 4 entries with `$defaults` expanded; `doctor.sh` autoMode PASS; live `docker exec -i supabase_db_game psql … -f - < verify/0043 … | tail` → `ROLLBACK`, exit 0, no prompt. `claude auto-mode critique` printed nothing (2.1.273).
- **Reference**: `settings.json`, `templates/settings/SETTINGS.md` (`autoMode.allow` row + interpreter note), commit `5eccc3f`, merge `ddadca6`. Links [[BDR-090]], [[LRN-153]], [[LRN-155]], [[LRN-156]].
+12
View File
@@ -39,6 +39,7 @@ rules:
| EVAL-025 | 2026-07-17 | opening seo/geo inventory (subagents): 7/7 verifiable claims false or overstated; real contact corrected all, 6 plan corrections + 4 features killed at measurement | keep |
| EVAL-027 | 2026-08-24 | contract-gates behavioral RED: 16/16 fresh unprimed runs followed new doctrine (GATE 0 order, vacuous oracle, ABANDONED routing, scope temptation resisted) | keep |
| EVAL-028 | 2026-08-26 | darwin v2.1 paired run 54 units: 60 paired verdicts 0 revert/tie; skeptics found 3 real residuals — engaged, not rubber-stamp | keep |
| EVAL-029 | 2026-09-15 | 4-agent plan challenge: 6 BLOCKER; 3 of 3 confirmation-pass BLOCKERs came from the fixes themselves; caught a false 654 MB orphan claim | keep |
---
@@ -267,3 +268,14 @@ Dogfood: 3 blind lenses attacked the v1 plan for the plan-challenge feature itse
- **Method**: paired same-judge 3-majority per round (v2.1); judges live-exec where artifact executable (5 units: skills-perso, profile, plugin-pair, status-reporter, gitflow). Absolute scores triage-only. Totals main-thread (LRN-018 applied).
- **Anomalies**: (1) 0 reverts/ties in 60 verdicts — homogeneous-better checked: skeptic lens found real residuals 3x (doctor.sh cost source, hotfix RULES leftover restore, FILE(S) new-marker) → judges engaged. (2) census lock RED on line-rewrap, make test caught → LRN-144. (3) head-pipe masked grep exit 2x → LRN-143.
- **Action**: v2.1 paired = standard. Post-run absolute rescore skipped by design (would be judge-noise theater).
---
## EVAL-029 — 4-agent plan challenge: 6 BLOCKERs, and the fix round produced 3 of them
- **Date**: 2026-09-15
- **Method**: 3 blind lenses (correctness / robustness / simplicity) on plan rev 1, then 1 confirmation lens on rev 2. Subject = the gstack Playwright lib plan ([[BDR-088]]).
- **Result**: rev 1 → 3 BLOCKER + 12 MAJOR. Rev 2, written specifically to close them → 3 NEW BLOCKERs, and 2 of the 3 were INTRODUCED BY the fixes: the new "every public function returns 0" rule contradicted the new "return rc", and the printer-name clause came verbatim from my own contract criterion 9. Rev 3 dropped the recovery branch entirely at the human gate — 6 findings closed by deletion instead of code.
- **Anomaly**: my first user-facing answer asserted ~654 MB of orphan Playwright revisions. FALSE — `.links` showed every dir referenced, 0 reclaimable. Caught only while designing the guard, not while asserting the number. Worse, the guard I proposed would itself have deleted gsd-pi's rev 1243.
- **Action**: (1) never state a disk-reclaimable figure before reading the registry that owns it ([[LRN-151]]). (2) A fix round deserves the same challenge as the original plan — 3/3 confirmation BLOCKERs came from fixes, not from the original. (3) The confirmation pass earned its cost: without it the printer override would have shipped and silently disconnected doctor's counters ([[LRN-150]]).
- **Status**: keep.
- **Reference**: `.claude/tasks/plans/2026-09-13-gstack-playwright-lib-2220.md` (rev 3). Links [[BDR-088]], [[LRN-150]].
+16
View File
@@ -460,3 +460,19 @@ rules:
- Post-merge regression: toast dead again after re-attach from a RESTORED terminal, bell fine. Root cause [[LRN-147]]: ext hooks only terminals born after its activation; `enablePersistentSessions` restores terminals before it. Fix = disable persistent sessions, or fresh terminal + `dtach -a`. Verified: 3/3 toasts on fresh pty.
- Same-day counter-example broke that cause: second session's terminal deaf though created LATER, same window, ext global, shells identical. Trigger unknown; [[LRN-148]] adds the 5s pre-flight test + demotes LRN-147's mechanism claim.
- Attention signal refined: per-event labels (BDR-087 follow-on), silence on non-attention events, and no turn-end signal while `background_tasks` non-empty ([[LRN-149]]). Payload dump beat the docs: `background_tasks` undocumented for Stop but present on the wire. Branch bugfix/notify-subagent-spawn.
- gstack Playwright: bump extracted to `lib/gstack-playwright.sh`, now re-applied after a successful submodule update ([[BDR-088]]); read-only browsers report in doctor, no pruner — `.links` proved 0 bytes reclaimable and the guard I first proposed would have deleted gsd-pi's rev 1243 ([[BDR-089]], [[LRN-151]]). 4 challengers → 6 BLOCKER, recovery branch withdrawn at the gate ([[EVAL-029]]). 2cebecb on feature/gstack-playwright-lib.
- Node checked against Playwright: already v24 (1.61 needs >=18, 1.63 needs >=20), not the macOS constraint. macOS audit deferred to its own cycle — found statically: `sed -i` with no suffix x3 in install-plugins.sh (BSD sed eats the next arg), `${x,,}` in url-guard.sh (bash 4+, macOS ships 3.2), `readlink -f` in doctor.sh (absent pre-Monterey 12.3).
## 2026-09-15
- Aligned repo config + deployment on the user's hand-edited `settings.json`. Destructive shell work rebuilt in `autoMode` soft_deny/hard_deny once `ask` was established as inert under auto mode ([[BDR-090]]); `permissions.deny` +10 `.env` reader rules, 6 of which sat in `allow`.
- `autoMode.environment` was scoped to ANOTHER project inside the user-scope file, so every repo got atlast's facts. Rewritten machine-generic, atlast facts moved to atlast's own `settings.local.json`, `$defaults` added to all three lists ([[LRN-153]]).
- `doctor.sh` gained `check_automode` (missing `$defaults`, foreign-repo scope, both arms tested). `SETTINGS.md` documents the block + a tier-choice table. README's magic-MCP "ask = live confirmation" claim corrected — false under `defaultMode: auto`.
- Found, not fixed: `.claude/settings.local.json` = 14.6 KB shadow copy of the global settings at HIGHER precedence, incl. a `config-protection.sh` hook whose script does not exist. Logged F1-F3 in TODO.
- `make test` 0 RED, `doctor.sh` 0 errors, `shellcheck` clean.
- graphify skill untracked + gitignored (written by `graphify install --platform claude` since `~/.claude/skills` symlinks to `skills/`). Cost one self-inflicted incident: `git rm --cached` kept the files, `gitflow finish` deleted them at the merge ([[LRN-154]]). Restored at 0.9.61, guarded configs snapshotted and verified untouched.
- `.claude/settings.local.json` 14.6 KB -> 6.2 KB. It was not just duplication: its local `deny` still carried the 4 rules moved out of global deny, making [[BDR-090]]'s soft_deny a dead letter in this repo, and its `allow` carried `sed *` / `cp *` / `python3 -`, which short-circuit the classifier on the same rules.
## 2026-09-16
- Ask, don't guess ([[BDR-091]]): spec + plan, 9 lock-first tasks (contract-interview CLARIFY two passes, MID-RUN CLARIFICATION with `CLASS:` tag, HOW TO ASK; global rule; feat / bugfix / hotfix / ship-feature / init-project wired; interviewer; 3 executors), suite green. Behavioral fixture check still open ([[LRN-157]]).
- docker + node under auto mode ([[BDR-092]]): `ask` entries retired (inert on 2.1.273, probe — [[LRN-155]]), `autoMode.allow` + 2 soft_deny, live `docker exec … psql` OK. Static interpreter allow is suspended under auto → prose only ([[LRN-156]]).
- Both merged into develop 2026-09-17 via gitflow (`ddadca6`, `56bd035`), two stack conflicts (TODO, CHANGELOG) resolved keeping both blocks. Symlinked `settings.json` follows the checkout: live config = whatever branch is out.
+69
View File
@@ -139,6 +139,14 @@ rules:
| LRN-134 | 2026-07-17 | resolve-then-pin in stdlib http.client beats monkeypatching getaddrinfo — dual-stack, thread-safe, no requests; classify the OS-resolved IP not the URL text | closing SSRF/DNS-rebinding on any Python HTTP egress |
| LRN-135 | 2026-07-17 | a prefix-only scan for a dangerous construct is bypassable by padding — scan the WHOLE document | refusing any hostile construct (DTD/directive/marker) before parse |
| LRN-143 | 2026-08-26 | `cmd \| head \|\| fallback` — pipeline rc is head's (0), fallback dead; bounded output → drop head, else pipefail | any probe/fallback bash in skills before trusting `\|\|` |
| LRN-150 | 2026-09-15 | Sourced lib shares caller shell: bare `ok/warn/info` override its printers, and its `set -e` applies inside | any new lib/*.sh |
| LRN-151 | 2026-09-15 | Playwright cache truth = union over `.links`, dir name maps `_`→`-`, revisionOverrides exist | shared versioned binary caches |
| LRN-152 | 2026-09-15 | git `protocol.file=user` blocks submodule fixtures; `-c` misses the code under test, `GIT_CONFIG_*` env does not | tests building git fixtures |
| LRN-153 | 2026-09-15 | `autoMode` lists replace built-ins without `"$defaults"`; a user-scope block reaches every project | any `autoMode` edit |
| LRN-154 | 2026-09-15 | `git rm --cached` + merge into a branch that still tracks the file DELETES it from disk | untracking a generated file |
| LRN-155 | 2026-09-16 | ask under auto: doc says prompt, probe on 2.1.273 says no; re-probe after upgrades | any permission-tier reasoning |
| LRN-156 | 2026-09-16 | autoMode.allow = exception tier; static interpreter allow suspended under auto → conditions live in prose | conditional permissions |
| LRN-157 | 2026-09-16 | gap-only trigger blind to taste → add a trigger class, not budget; ask at plan, mid-run for leftovers | any "ask more" request |
---
@@ -1421,3 +1429,64 @@ Rule: when editing a doctrine file under structure locks, grep the test's lock s
- **Fail-open**: field absent (older client) → still signal. Missed notification worse than extra one.
- **Cross-session gotcha**: hook is user-scope, so EVERY session runs it. A single-file dump (`> file`) gets overwritten by another project's session — append JSONL and filter on `.cwd`. That accident proved `permission_prompt` fires with `message="Claude needs your permission"` (unexercisable in this session under `defaultMode: auto`).
- **Future**: any hook needing turn-completion semantics must check background_tasks; "turn ended" ≠ "work done". Verified live: Stop with 0 tasks signals, Stop with 1 running subagent silent.
---
## LRN-150 — Sourced shell lib is not a subprocess: prefix printers, honor inherited errexit
- **Date**: 2026-09-15
- **Pattern**: `source lib.sh` shares the caller's shell. Two bites. (a) bare `ok()`/`warn()`/`info()` in the lib OVERRIDE the caller's same-named funcs. `doctor.sh` counts ERRORS/WARNS inside its own `warn()` → a lib `warn` disconnects the counter and doctor prints "No errors" while warnings scroll. Prefix every lib printer (`_gspw_ok`, `_gspw_warn`, `_gspw_info`). (b) caller's `set -euo pipefail` applies INSIDE the lib's functions: a failing command-substitution assignment (`x="$(. /etc/os-release; [ "$ID" = ubuntu ] && printf ...)"`) aborts the CALLER when the func is called as a bare statement. Reproduced — exit 1 on every non-Ubuntu host, latent in `install-plugins.sh` since [[BDR-029]].
- **Rule**: public func called bare → `return 0` on every path + `|| true` on every capture. Func allowed to return non-zero → call it ONLY as an `if` condition.
- **Future application**: any new `lib/*.sh` sourced by a script that owns printers or sets `-e`. Check BOTH facets before wiring; the printer one is silent (no error, just a lying summary).
- **Reference**: `lib/gstack-playwright.sh`, `doctor.sh:12-15`. Links [[BDR-088]].
---
## LRN-151 — Playwright cache truth lives in `.links`, never in one install's view
- **Date**: 2026-09-15
- **Pattern**: `~/.cache/ms-playwright/.links/<sha1>` = one file per registered `playwright-core`, content = its path. Required set = UNION of `browsers.json` revisions across ALL of them. Dir name on disk = `${name//-/_}-${revision}`: `chromium-headless-shell` → `chromium_headless_shell-1228`. Miss that mapping and 2 live dirs read as orphan forever. `revisionOverrides` exists (webkit, ffmpeg on mac / debian11 / ubuntu20.04) so the base revision alone under-matches. Playwright prunes this set itself on every `install` (`_deleteStaleBrowsers`, coreBundle.js).
- **Future application**: never call a browser dir orphan from one project's playwright view — read `.links` first. Generalizes to any tool with a shared versioned binary cache plus a registry of consumers: the consumer registry is the source of truth, not the consumer you happen to be standing in.
- **Reference**: `lib/gstack-playwright.sh` `_gspw_browser_referenced`. Links [[BDR-089]], [[EVAL-029]].
---
## LRN-152 — git `protocol.file=user` kills submodule fixtures; `-c` misses the code under test
- **Date**: 2026-09-15
- **Pattern**: since the CVE-2022-39253 hardening git refuses submodule clone/fetch over a local path by default (git 2.53 → `protocol.file` = `user`). `-c protocol.file.allow=always` fixes the FIXTURE's own git calls but NOT the `git` the code under test spawns — fresh process, inherits nothing from `-c`. Export for the whole test process instead: `GIT_CONFIG_COUNT=1 GIT_CONFIG_KEY_0=protocol.file.allow GIT_CONFIG_VALUE_0=always`. Env propagates, `-c` does not.
- **Also**: fixture repos need LOCAL `user.email`/`user.name` (no global identity here) and `git init -b main` + explicit `submodule.<name>.branch`, else `--remote` resolves a different branch than production does.
- **Future application**: any test building a git submodule fixture. Symptom is a hard "transport 'file' not allowed" before the first assertion, which reads like a broken test rather than a policy.
- **Reference**: `lib/tests/gstack-playwright.test.sh`.
## LRN-153 — `autoMode` lists replace built-ins unless `"$defaults"` is spliced in
- **Date**: 2026-09-15
- **Pattern**: every list under `autoMode` (`allow` `soft_deny` `hard_deny` `environment`) is a FULL replacement by default. Omit the literal `"$defaults"` and the built-in classifier rules are dropped silently — no warning, no schema error, the classifier just runs thinner. Put `"$defaults"` first, own entries after: built-ins inherited, then refined.
- **Scope trap, same block**: `autoMode` in `~/.claude/settings.json` reaches EVERY project. A block generated while working in one repo (its deploy target, its secrets, its data) ships that repo's facts to all the others, and contradicts whichever repo is actually open. Project facts belong in that project's `.claude/settings.local.json`.
- **Format**: these lists are prose spliced into the classifier prompt, not permission-rule syntax. Write "Sending SIGKILL reaches processes outside this session", never `Bash(kill -9 *)`.
- **Backstop**: `doctor.sh` `check_automode` warns on a list missing `$defaults` and on a user-scope `environment` naming a git repo other than the config repo. Both arms exercised against the defective block before shipping.
- **Future application**: any `autoMode` edit — check `$defaults` presence and scope before anything else.
- **Reference**: `doctor.sh`, `templates/settings/SETTINGS.md`. Links [[BDR-090]].
## LRN-154 — Untracking a generated file then merging deletes it from disk
- **Date**: 2026-09-15
- **Pattern**: `git rm --cached` removes from the index and KEEPS the working file, which is the whole point when untracking a tool-generated artifact. But `gitflow finish` checks out the target branch first, where the file is still tracked, so git restores it; the merge then applies the deletion to a tracked file and removes it from disk. `.gitignore` does not protect it — it only stops a re-add. Net effect: the file survives the commit and dies at the merge, several minutes later, which reads as unrelated.
- **Detection**: the working tree is clean and the file is simply absent. Nothing errors. Only a post-merge `ls` catches it.
- **Future application**: untracking any generated file — know the regeneration command BEFORE merging, and `ls` the path right after `finish`. If nothing regenerates it, keep it tracked.
- **graphify specifics**: `graphify install --platform claude` copies the skill and touches nothing else. `graphify claude install` is a different command — it writes the CLAUDE.md section and the `.claude/settings.json` hooks, rewrites both guarded configs, and does NOT copy the skill. Confusing the two wastes a recovery attempt.
- **Reference**: `CLAUDE.md` machine-owned section, commit 80ccdaf. Links [[BDR-090]].
## LRN-155 — `permissions.ask` under auto mode: the probe beats the doc
- **Date**: 2026-09-16
- **Pattern**: `auto-mode-config` + `permissions` docs say a content-scoped `ask` rule (`Bash(git push *)`) is evaluated BEFORE the classifier and always prompts, even in auto mode. Probe on 2.1.273: `node -e 'console.log(...)'` matching `Bash(node -e *)` in `ask` ran, no prompt, exit 0. [[LRN-146]] holds. Either the doc describes a later build or "content-scoped" means something narrower; observed wins.
- **Future application**: before reasoning about a permission tier, probe it with a benign command matching the rule; re-probe after every Claude Code upgrade — the day `ask` starts prompting, every leftover `ask` entry becomes a nag for things meant to run free.
- **Reference**: [[BDR-092]], `templates/settings/SETTINGS.md` "ask is not a prompt" §.
## LRN-156 — Conditional permissions live in classifier prose, not static rules
- **Date**: 2026-09-16
- **Pattern**: `autoMode.allow` = exception tier: an entry overrides a matching `soft_deny`, built-in or own (precedence hard_deny > soft_deny > allow > explicit intent). Under auto, static allow rules granting arbitrary execution (`Bash(*)`, wildcarded interpreters like `Bash(node *)`) are suspended → classifier anyway; non-interpreter statics (`awk`, `echo`) resolve before it. A condition ("package declared in the lockfile", "container is local dev") is therefore expressible ONLY as `autoMode.allow` prose. Word it narrowly: it punches through built-in rules too.
- **Tooling**: `claude auto-mode defaults` prints the built-in lists (grep it for the rule that bit); `claude auto-mode config` = effective lists with `$defaults` expanded; `claude auto-mode critique` printed nothing on 2.1.273. Shell-snapshot `claude` wrapper is broken (`exec command claude` → "command: not found") → call `~/.local/bin/claude` directly.
- **Reference**: [[BDR-092]], [[LRN-153]].
## LRN-157 — Taste is invisible to a gap-only trigger; ask at plan time
- **Date**: 2026-09-16
- **Pattern**: a trigger that fires only on missing outcome / scope / constraints lets every taste choice through — "add a share icon" is complete by those criteria and the icon's side is decided downstream. More budget changes nothing; the fix is a new trigger class (VISIBLE / PUBLIC NAME / SCOPE). Cost geometry: a fresh re-dispatch keeps the working tree and loses the executor's reasoning → the same question costs about one executor run more mid-run than at PLAN. So: sweep once at the plan step, keep the mid-run channel for leftovers. Executor tags the class; orchestrator re-reads it (tag = hint, a mis-tag would offload class 4 onto the human). Relayed questions obey [[LRN-102]]: context inside `AskUserQuestion`, nothing the user needs printed before it.
- **Future application**: any "ask more" request → check WHICH trigger is blind before touching a quota. Any orchestrator with a "decide it yourself" fallback on an executor halt → route by class first.
- **Reference**: [[BDR-091]], `lib/contract-interview.md` STEP 2 + MID-RUN CLARIFICATION.
+156
View File
@@ -1,5 +1,161 @@
# TODO
## 2026-09-16 — docker + node framed by the classifier (feature/automode-docker-node)
User: `docker exec -i supabase_db_game psql … -f - < supabase/verify/*.sql | tail`
must run unprompted under auto mode; same for node/npm/npx when the package
is declared and effects stay in the cwd; "ajoute du soft deny pour bien le
cadrer". Findings: `ask` is inert under auto (LRN-146 re-verified on 2.1.273
with a `node -e` probe; the docs claim otherwise for content-scoped rules);
the real gate is the built-in `Remote Shell Writes` / `Production Reads`
classifier rules; a static `Bash(node *)` allow rule is suspended under auto
(wildcarded interpreter), so `autoMode.allow` prose is the only lever for
node. User approved the design and the `ask` removal explicitly (S6 override
for this change, diff reviewed on the branch).
- [x] A1 `settings.json` — drop 4 docker + `node -e` from `ask`; new
`autoMode.allow` (`$defaults` + local dev containers + project-local
node); 2 `soft_deny` entries (docker data destruction, undeclared
node packages); `model` bump committed separately
- [x] A2 `templates/settings/SETTINGS.md` — `autoMode.allow` tier row +
why a static interpreter allow rule cannot do it; LRN-146 re-verify note
- [x] A3 CHANGELOG [Unreleased] Changed
- [x] A4 verify (2026-09-16, all green; `critique` printed nothing): `jq`, `claude auto-mode config`,
`doctor.sh`, live `docker exec` in game
- [x] A5 registries written 2026-09-17: LRN (doc vs observed `ask` under auto,
2.1.273; `autoMode.allow` = exception tier; wildcarded-interpreter
allow suspended), BDR-090 addendum
## 2026-09-16 — ask, don't guess: orchestrators ask about open choices (feature/ask-dont-guess)
User: the orchestrators (ship-feature, feat, hotfix, bugfix, init-project)
settle choices they should ask about ("cet icône, plutôt à gauche ou à
droite ?"), even mid-run. Diagnosis: contract-interview STEP 2 only fires on
gaps (outcome / scope / constraints), so a taste choice never triggers a
question; feat:153 and bugfix:165 tell the orchestrator to "make the
decision HERE" on NEED-DECISION. Decisions (user, 2026-09-15/16): global
rule changes for all work, hotfix included; classes VISIBLE / PUBLIC NAME /
SCOPE ask, internal technical choices never. Spec:
`docs/superpowers/specs/2026-09-16-ask-dont-guess-design.md`; plan:
`docs/superpowers/plans/2026-09-16-ask-dont-guess.md` (9 tasks, lock-first).
- [x] P1 `lib/contract-interview.md` — STEP 2 CLARIFY (pass A gaps, pass B
open choices), MID-RUN CLARIFICATION, HOW TO ASK; 9 locks in
`contract-verifier.test.sh`
- [x] P2 `CLAUDE.global.md:51-55` — "Ask rather than guess" replaces the
one-question rule; bug line reconciled
- [x] P3 `skills/feat/SKILL.md` — pass B at STEP 1, NEED-DECISION routed on class
- [x] P4 `skills/bugfix/SKILL.md` — pass B at STEP 3, NEED-DECISION routed on class
- [x] P5 `skills/hotfix/SKILL.md` — pass B at LOCATE, tagged BLOCKED relayed;
lock `loops-light.test.sh:84`
- [x] P6 `skills/ship-feature` STEP 2 + `skills/init-project` contract §/STEP 3
- [x] P7 `agents/interviewer.md` — visible/public/scope item never `(assumed)`
- [x] P8 `agents/{feater,bugfixer,hotfixer}.md` — CLASS tag; 3 locks in `gates.test.sh`
- [x] P9 `make test` green (2026-09-16), CHANGELOG, TODO tick; manual behavioral check still OPEN before merge
- [x] P10 registries written 2026-09-17: BDR (supersedes the one-question rule),
LRN (taste is invisible to a gap-only trigger; fresh re-dispatch cost
favors plan-time questions)
## 2026-09-15 — align config + deployment on the hand-edited settings.json (feature/automode-config-alignment)
User edited global `settings.json` by hand: 4 destructive rules moved
deny→ask (`rsync`, `kill -9`, `killall`, `pkill`), 4 removed from ask
(`xargs`, `sed`, `cp`, `mv` — coherent with auto mode's Bash-first
workflow; the `.env`-scoped `cp`/`mv`/`xargs` deny rules still stand),
and an `autoMode.environment` block added. Two defects found:
(1) the environment block describes **atlast** (`bin/deploy.sh` lftp/FTP
to OVH, quote-request data, "no remote configured") but lives in the
user-scope file symlinked to `~/.claude/settings.json` by `link.sh:21`
— so every project gets atlast's facts; claude-config itself has a
Gitea remote, contradicting the block. (2) no `"$defaults"` sentinel,
so the built-in classifier environment entries are replaced, not
extended. Third finding: LRN-146 records, verified in session, that
`ask` rules raise no prompt under `defaultMode: auto` — the deny→ask
move therefore traded a static block for a classifier decision.
User decisions (2026-09-15): atlast block → atlast's own
`settings.local.json`, global block rewritten machine-generic; the 4
destructive rules → `autoMode.soft_deny` (the section that actually
binds under auto mode) instead of `ask`.
- [x] T1 global `settings.json` — machine-generic `autoMode.environment`
with `$defaults`; new `autoMode.soft_deny` with `$defaults` + the
4 destructive rules; drop those 4 from `permissions.ask`
- [x] T2 `/home/bchanot/Documents/atlast/.claude/settings.local.json` —
receives the atlast-specific `autoMode.environment` (gitignored,
personal scope); verify project-scope `autoMode` is honored
- [x] T3 `templates/settings/SETTINGS.md` — document the `autoMode`
block (environment / soft_deny / hard_deny / allow, `$defaults`
semantics, `classifyAllShell`) + the "ask ≠ prompt under auto"
caveat that makes soft_deny the right tier
- [x] T4 `README.md` — magic-MCP paragraph claims the `ask` tier makes
every `mcp__magic__*` call "require a live confirmation and never
auto-execute"; false under auto mode per LRN-146. Correct the
claim, flag the soft_deny option to the user (don't decide it)
- [x] T5 `doctor.sh` — permissions section is blind to `autoMode`, now a
live security surface. Add a check: block present, `$defaults`
inherited, no foreign absolute project path hardcoded
- [x] T6a CHANGELOG (Added/Changed/Fixed under [Unreleased])
- [ ] T6b registries BDR-090 + LRN-153 + journal — drafted, awaiting user approval
- [x] T7 verify: `make test`, `bash doctor.sh`, `shellcheck`
NOT in scope: the 3 dirty `skills/graphify/*` files (pre-existing,
unrelated) — never staged.
### Second pass (2026-09-15, user decisions)
User confirmed the `ask` removals were deliberate (`/permissions`), asked
for the diff vs develop and for guards where the removals left a hole.
Answered: writes outside cwd → soft_deny; in-place edits beyond one named
file → soft_deny; inline interpreters + `xargs` → soft_deny when they
delete or write outside cwd; hard_deny for secret exfiltration, prod
deploy, disarming guardrails (history rewrite NOT retained, so a `rebase`
then an ordinary push stays uncovered); extend the static deny family to
the `.env` readers; `classifyAllShell` stays false; intent clears a soft
block for the CURRENT TURN only.
- [x] S1 `permissions.deny` +10 reader rules (sed awk cut tr sort uniq
diff od xxd strings vs `.env*`) — 6 of them were in `allow`
- [x] S2 `autoMode.soft_deny` — 7 rules + the intent-scope line
- [x] S3 `autoMode.hard_deny` — 3 rules, "adding a restriction is fine,
removing one is not"
- [x] S4 `SETTINGS.md` — tier-choice table + scope-of-intent section
- [x] S5 CHANGELOG — Changed rewritten, new Security block
- [x] S6 CONSEQUENCE confirmed by user 2026-09-15: the hard_deny guardrail rule means I can
no longer edit a deny/soft_deny/hard_deny list to REMOVE an entry.
Tightening stays allowed. Future permission loosening goes through
`/permissions` or the user's own edit.
### Third pass (2026-09-15) — F1-F3 done + graphify untracked
Worst finding was not the duplication: local `deny` still carried
`rsync` `kill -9` `killall` `pkill`, the four the user moved OUT of
global deny. deny wins across sources, so `autoMode.soft_deny` was a
dead letter in THIS repo. Local `allow` also held `sed *`, `cp *`,
`python3 -` — an allow rule short-circuits the classifier, punching a
hole through the same soft_deny rules.
- [x] G1 `skills/graphify/{SKILL.md,references/,.graphify_version}`
gitignored + `git rm --cached`. Written by `graphify claude
install` since `~/.claude/skills` symlinks to `skills/`; a fresh
clone gets them from `make plugin`. `test-prompts.json` is
hand-written for darwin, stays tracked. Trade-off documented in
CLAUDE.md: an upstream release can now change the skill prompt
with no diff to review.
- [x] G2 `.claude/settings.local.json` 14.6 KB -> 6.2 KB. deny + ask
dropped whole, allow 185 -> 98 (81 duplicates of the global, 6
policy conflicts: `sed *`, `cp *`, `python3 -`,
`Read(//home/bchanot/**)`, `WebSearch`, a leftover injection-test
payload). Every non-`permissions` key was a verbatim copy of the
global, `hooks` included. Backup: `.audit/settings.local.json.bak-*`
(gitignored, the file itself is not in git).
### Follow-up found while doing this (fixed in the third pass above)
`.claude/settings.local.json` (gitignored, 14.6 KB) is a near-complete
shadow copy of the global `settings.json` at a HIGHER precedence tier:
185 allow / 30 ask / 106 deny, plus its own `cleanupPeriodDays`,
`attribution`, `statusLine`, `enabledPlugins`, `extraKnownMarketplaces`,
`effortLevel`, `remoteControlAtStartup`, `inputNeededNotifEnabled`,
`skipAutoPermissionPrompt` — all identical to the global today, so the
duplication is invisible until the global drifts, which it just did
(no `autoMode`, 106 deny vs 116). It defeats the config-guard premise
(hand-curated `settings.json`) with a file nobody reviews.
- [x] F1 `WebSearch` sits in global `ask` and in local `allow` — in this
repo it never reaches the ask tier. Intended or drift?
- [x] F2 local `hooks` block registers `bash ~/.claude/hooks/config-protection.sh`
on PreToolUse/Bash. That script does not exist, in `hooks/` or in
`~/.claude/hooks/`. Dead hook firing on every Bash call here.
- [x] F3 decide: prune the local file down to the session-accumulated
allow rules only, dropping every key that merely restates the
global, or keep the copy deliberately and document why.
## 2026-08-25 — darwin fresh baseline: 32 skill-systems + 23 agents (feature/darwin-optimize-20260825)
User: `/darwin-skill all skills and agents` (background). Fresh-from-zero
(results.tsv wiped 2026-06-23, journal 2026-06-30). Scope per BDR-015/043 +
@@ -0,0 +1,115 @@
# CONTRACT — gstack-playwright-lib
- date: 2026-09-13 | flow: feat | branch: feature/gstack-playwright-lib
- status: active
## REQUEST (verbatim — IMMUTABLE)
Message 1:
> l'installation de chromium, c'est une version fix ou en latest ? Il faudrait mettre en lateste, et d'ailleurs son update est pris en compt dans l'update ? quelq version a besoin gstack ? Ca serait pas plus simple d'installer perplexity a la place ?
Message 2 (after the assistant proposed fix A + fix B):
> les deux
Message 3 (answer to the scope question on fix B, after the "654 Mo orphelins"
premise was proven wrong):
> Check read-only dans doctor.sh
## CLARIFICATIONS
Q: "mettre en latest" — pin Chromium to latest?
A: Not actionable as asked. Playwright downloads the browser revision its own
version pins (1.61.1 → chromium 1228); the CDP client is coupled to that
build. "Latest" = track the latest Playwright, which is what BDR-029's bump
already does. No change to the pinning mechanism is in scope.
Q: Volet B — purge the orphan Playwright revisions?
A: Superseded by evidence. `~/.cache/ms-playwright/.links/` registers THREE
playwright installs (gstack 1.61.1 → rev 1228; gsd-pi nvm 1.61.0 → 1228;
gsd-pi ~/.local 1.63.0 → 1243). Every directory on disk is referenced;
zero bytes reclaimable. Playwright already GCs correctly on every
`install` (`_deleteStaleBrowsers`, unions across all registered installs).
User chose: read-only report in doctor.sh, NO deletion anywhere.
## ACCEPTANCE CRITERIA
1. `lib/gstack-playwright.sh` exists, is source-safe (sourcing prints nothing
and runs no side effect), and its verb dispatcher works when executed.
CHECK: out=$( . lib/gstack-playwright.sh; echo READY ); [ "$out" = READY ] && bash lib/gstack-playwright.sh 2>&1 | grep -q 'usage:' && echo LIB_OK
EXPECT: LIB_OK
EVIDENCE: MET exit=0 marker-found :: LIB_OK
2. The bump logic lives ONLY in the lib: `install-plugins.sh` no longer
defines `gstack_bump_playwright_if_unsupported`, sources the lib instead,
and still calls it BEFORE gstack `./setup` (BDR-029 behavior unchanged:
OS-gated, idempotent, non-fatal).
CHECK: grep -q '^gstack_bump_playwright_if_unsupported() {' install-plugins.sh && exit 1; grep -q 'lib/gstack-playwright.sh' install-plugins.sh || exit 1; c=$(grep -n 'gstack_bump_playwright_if_unsupported' install-plugins.sh | grep -v ':[[:space:]]*#' | tail -1 | cut -d: -f1); s=$(grep -n '&& \./setup)' install-plugins.sh | head -1 | cut -d: -f1); [ -n "$c" ] && [ -n "$s" ] && [ "$c" -lt "$s" ] && echo EXTRACT_OK
EXPECT: EXTRACT_OK
EVIDENCE: MET exit=0 marker-found :: EXTRACT_OK
3. `update-all.sh` delegates the gstack submodule update to the lib
(`gstack_submodule_update_with_bump`) instead of calling
`git submodule update --remote` bare, so the bump is re-applied after every
successful update.
CHECK: grep -q 'gstack_submodule_update_with_bump' update-all.sh && grep -q 'lib/gstack-playwright.sh' update-all.sh && ! grep -qE '^[[:space:]]*if git submodule update --remote skills-external/gstack' update-all.sh && echo WIRED_OK
EXPECT: WIRED_OK
EVIDENCE: MET exit=0 marker-found :: WIRED_OK
4. [gated 2026-09-15] `gstack_submodule_update_with_bump` NEVER modifies the
submodule working tree. On a successful `git submodule update --remote` it
re-applies the bump; on failure it returns non-zero, touches nothing, and
emits git's own message plus a hint naming the local Playwright bump when
`package.json`/`bun.lock` are the dirty files. The conflict-RECOVERY branch
of the earlier revision (discard, retry, backup, restore) is withdrawn: it
could leave the bump discarded and un-reapplied, regressing a working
browser into BLK-008, which the pre-existing behavior never did.
CHECK: sed 's/#.*//' lib/gstack-playwright.sh | grep -qE 'git [^|;]*(checkout|reset|clean|stash)' && exit 1; bash lib/tests/gstack-playwright.test.sh 2>&1 | grep -q 'update-conflict' && bash lib/tests/gstack-playwright.test.sh 2>&1 | grep -qE '^PASS=[0-9]+ FAIL=0$' && echo NONDESTRUCTIVE_OK
EXPECT: NONDESTRUCTIVE_OK
EVIDENCE: MET exit=0 marker-found :: NONDESTRUCTIVE_OK
5. `doctor.sh` prints a Playwright-browsers section: total cache size, one line
per browser directory naming the registered playwright install(s) that
reference it, plus counts of unreferenced directories and broken links.
CHECK: bash doctor.sh 2>/dev/null | grep -qi 'playwright browsers' && bash lib/gstack-playwright.sh browsers-report | grep -qE 'chromium-[0-9]+' && bash lib/gstack-playwright.sh browsers-report | grep -qi 'unreferenced' && echo REPORT_OK
EXPECT: REPORT_OK
EVIDENCE: MET exit=0 marker-found :: REPORT_OK
6. The report is provably read-only: no destructive verb anywhere in the lib,
and the cache directory listing is identical before and after a report run.
CHECK: sed 's/#.*//' lib/gstack-playwright.sh | grep -qwE '(rm|rmdir|unlink|truncate|mv)' && exit 1; b=$(ls -la ~/.cache/ms-playwright ~/.cache/ms-playwright/.links 2>/dev/null | cksum); bash lib/gstack-playwright.sh browsers-report >/dev/null 2>&1; a=$(ls -la ~/.cache/ms-playwright ~/.cache/ms-playwright/.links 2>/dev/null | cksum); [ "$b" = "$a" ] && echo READONLY_OK
EXPECT: READONLY_OK
EVIDENCE: MET exit=0 marker-found :: READONLY_OK
7. `lib/tests/gstack-playwright.test.sh` exists, passes, and covers at least:
bump skipped when the OS tag is already supported; bump fired when it is
not; submodule-update conflict recovery; browsers-report on a fixture cache
holding a referenced revision, an unreferenced one and a broken link.
CHECK: bash lib/tests/gstack-playwright.test.sh | tail -1 | grep -qE '^PASS=[0-9]+ FAIL=0$' && echo TESTS_OK
EXPECT: TESTS_OK
EVIDENCE: MET exit=0 marker-found :: TESTS_OK
8. shellcheck clean on every touched shell file.
CHECK: shellcheck lib/gstack-playwright.sh lib/tests/gstack-playwright.test.sh install-plugins.sh update-all.sh doctor.sh >/dev/null 2>&1 && echo SHELLCHECK_OK
EXPECT: SHELLCHECK_OK
EVIDENCE: MET exit=0 marker-found :: SHELLCHECK_OK
9. [gated 2026-09-15] (judgement) No new dependency; the report degrades
silently when `~/.cache/ms-playwright` is absent, when its `.links`
directory is absent, when `PLAYWRIGHT_BROWSERS_PATH` is `0` or not a
directory, or when no playwright install is registered — doctor must stay
green on a machine that never installed a browser. The lib's printers are
named `_gspw_ok`/`_gspw_warn`/`_gspw_info` and it defines NO bare
`ok`/`warn`/`info`/`pass`/`fail`: doctor.sh sources the lib before every
check, so bare names would override its own printers and silently
disconnect its `ERRORS`/`WARNS` counters.
## FILE SCOPE
- lib/gstack-playwright.sh (new)
- lib/tests/gstack-playwright.test.sh (new)
- install-plugins.sh (remove inline fn, source + call lib)
- update-all.sh (call bump + conflict recovery)
- doctor.sh (new read-only report section)
Out of scope: the gstack submodule itself, the pinning mechanism, any
deletion of cached browsers, the `GSTACK_CHROMIUM_NO_SANDBOX` layer
(LRN-040 layer 2, unchanged).
@@ -0,0 +1,201 @@
# PLAN — gstack-playwright-lib (feat) — REVISION 3
Contract: `.claude/tasks/contracts/2026-09-13-gstack-playwright-lib-2220.md`
Revision 3 (2026-09-15). Rev 1 → 3-lens challenge → rev 2 → confirmation pass
→ rev 3. The conflict-RECOVERY branch is WITHDRAWN at the human gate: it
concentrated 3 BLOCKERs and 4 MAJORs, and its worst case regressed a working
browser into BLK-008, which the pre-existing behavior never did.
## Context
- Chromium is not an apt package. It is the browser revision pinned by the
installed Playwright (`gstack/setup:483`). gstack: playwright 1.61.1 →
chromium rev 1228 (`Chrome for Testing 149`).
- `~/.cache/ms-playwright/.links/` registers 3 playwright installs: gstack
1.61.1 (1228), gsd-pi nvm 1.61.0 (1228), gsd-pi ~/.local 1.63.0 (1243).
Every dir on disk is referenced → 0 bytes reclaimable. Playwright already
prunes correctly on every `install` (`_deleteStaleBrowsers`). No pruner is
written here.
- The gstack submodule is intentionally dirty: `package.json` + `bun.lock`
carry the BDR-029 bump; `.gitmodules` sets `ignore = dirty`. It also carries
an untracked `?? bin/bin`.
- `update-all.sh:87` calls `git submodule update --remote` bare, swallows
stderr, and never re-applies the bump afterwards. THAT is the gap.
## Hard constraints the code must respect
**Inherited errexit.** All three callers run `set -euo pipefail` and source
the lib. `gstack_bump_playwright_if_unsupported` and `gstack_browsers_report`
are called as bare statements, so they MUST `return 0` on every path and every
capture inside them takes `|| true`.
`gstack_submodule_update_with_bump` is the ONE exception: it returns non-zero
on failure and is therefore called ONLY as an `if` condition, keeping
`update-all.sh:87`'s existing `if / else warn` shape. An offline update stays
non-fatal, exactly as today.
**Printer names.** The lib defines `_gspw_ok`, `_gspw_warn`, `_gspw_info` and
NEVER a bare `ok`/`warn`/`info`/`pass`/`fail`. `doctor.sh:22` sources the lib
before every check, so bare names would override `doctor.sh:12-15` and
silently disconnect its `ERRORS`/`WARNS` counters.
**No destructive command in the lib, at all.** No `rm`, `rmdir`, `unlink`,
`truncate`, `mv`, and no `git checkout`/`reset`/`clean`/`stash`. Contract
criteria 4 and 6 both grep for this.
**macOS-safe.** No `timeout` without a `command -v` guard (absent from stock
macOS), no `readlink -f` (absent before Monterey 12.3), no `md5sum`, no
`sed -i` without a suffix, no bash-4-only expansions (`${x,,}`), no `grep -P`.
## Files
1. `lib/gstack-playwright.sh` — NEW. Sourceable lib + verb dispatcher. No
`set -euo pipefail` at top level (mirrors `lib/detect-plugins.sh`).
Dispatcher guarded by `[ "${BASH_SOURCE[0]}" = "${0}" ]`, exposing
`browsers-report` ONLY. Any other argument → `usage:` on stderr, exit 2.
The write functions stay sourced-only: a CLI verb would expose
`bun add playwright@latest` as a command-line entry point.
- `_gspw_ok` / `_gspw_warn` / `_gspw_info <msg>` — fixed-prefix printers.
- `gstack_pw_ostag [os_release_path]` — prints `ubuntu<VERSION_ID>` for
Ubuntu, nothing otherwise. The capture takes `|| true`: the moved line
exits 1 on every non-Ubuntu host and would abort the caller under
inherited errexit. `return 0` always.
- `gstack_pw_supports <playwright_core_lib_dir> <ostag>` — 0/1 by grep.
- `gstack_bump_playwright_if_unsupported <gstack_dir>` — BDR-029 logic,
parameterized. Prepends `$HOME/.bun/bin` to PATH when `bun` is not
resolvable (LRN-036). Wraps ALL THREE bun invocations
(`bun install --frozen-lockfile`, the `bun install` fallback,
`bun add playwright@latest`) in `timeout 300` when `command -v timeout`
succeeds, plain otherwise. Exit 124 from any of them → `_gspw_warn` and
`return 0` WITHOUT attempting the bump: a TERM'd install leaves
`node_modules` half-written, and the support grep would then read a
truncated tree. One `_gspw_info` line before the network work so a
stalled registry is visible. `return 0` on every path (BDR-029
non-fatal).
- `gstack_submodule_update_with_bump <repo> [sub_path]`:
1. `git -C "$repo" submodule update --remote "$sub"`, stderr captured.
2. exit 0 → `gstack_bump_playwright_if_unsupported "$repo/$sub"` →
`return 0`.
3. exit != 0 → `_gspw_warn` with git's own message, verbatim and
unparsed. Then, when `git -C "$sub" status --porcelain --
package.json bun.lock` is non-empty, one extra `_gspw_info` hint line
naming the local Playwright bump and pointing at `make plugin`.
`return 1`. NOTHING in the working tree is touched.
No locale pin is needed: git's message is displayed, never parsed for a
decision. The hint is advisory, so its constant-true condition is
correct here, unlike the withdrawn recovery branch where it gated a
destructive step.
- `_gspw_browser_referenced <playwright_core_path> <dir_name>` — does that
install require this cache directory? Splits `<dir_name>` into name +
revision on the LAST `-`, then normalizes `_` → `-` on the name
(Playwright writes `chromium_headless_shell-1228` while `browsers.json`
says `chromium-headless-shell`; without this, two live directories are
reported unreferenced forever). Matches the base `revision` OR any value
under that browser's `revisionOverrides` (webkit and ffmpeg carry them
for mac and debian11 and ubuntu20.04 hosts). awk only: no jq, no
python3, no fallback ladder.
- `_gspw_install_label <playwright_core_path>` — `<dir-before-node_modules>
<version>`, e.g. `gstack 1.61.1`, `gsd-pi 1.63.0`.
- `gstack_browsers_report [cache_dir]` — read-only. Resolves the cache as
`${1:-${PLAYWRIGHT_BROWSERS_PATH:-$HOME/.cache/ms-playwright}}`; the
documented value `0` means "bundle into node_modules", so `0` and any
non-directory degrade to the silent no-cache path. Prints the header
`Playwright browsers`, the total from `du -sh … || true`, one line per
cache dir matching `*-<digits>` with the installs requiring it, then
`<N> unreferenced, <M> broken link(s)`. When N or M > 0, one
`_gspw_warn` naming them and the remedy, phrased without the words `rm`
or `mv` (criterion 6 word-greps the source): "re-run `playwright
install`, which prunes stale revisions". A dir whose NAME is listed by
some install but at another revision counts as `unknown revision`, not
unreferenced. `return 0` on every path.
2. `install-plugins.sh` — delete the inline function (294-321), source the lib
next to detect-plugins (line 30), call site at ~370 becomes
`gstack_bump_playwright_if_unsupported "$GSTACK_DIR"`.
3. `update-all.sh` — source the lib next to detect-plugins (line 19). Line 87
becomes `if gstack_submodule_update_with_bump "$REPO"; then` and the
existing `else warn …` arm is KEPT verbatim. No other structural change.
4. `doctor.sh` — source the lib next to detect-plugins (line 22). Add its own
`── Playwright browsers ──` section (NOT nested under gstack: 2 of the 3
registered installs are gsd-pi), called as `gstack_browsers_report || true`.
5. `lib/tests/gstack-playwright.test.sh` — NEW, auto-globbed by `make test`.
## Edge cases
- Every public function except the update returns 0 under the callers'
`set -euo pipefail`, including the all-zero-counts case, which is this
machine's nominal state and would otherwise kill `doctor.sh` before its
summary and take `update-all.sh:519` down with it.
- `.links` entry whose target is gone or whose `browsers.json` is unreadable
→ counted as a broken link, never dereferenced further.
- Two installs of the same tool at different versions → both labels listed.
- gstack submodule absent → bump and update both no-op 0.
- Sourcing the lib prints nothing and does not change the caller's options.
## Tests (`lib/tests/gstack-playwright.test.sh`)
Shape of `lib/tests/fast-libs.test.sh` (`check` helper, `PASS=n FAIL=n` last
line, `mktemp -d` + trap). git 2.53 defaults `protocol.file` to `user`, which
blocks submodule clone and fetch. The fixture git calls are not enough: the
`git submodule update --remote` under test runs INSIDE the lib, in a fresh
process. So the test exports, for the whole test process,
`GIT_CONFIG_COUNT=1 GIT_CONFIG_KEY_0=protocol.file.allow
GIT_CONFIG_VALUE_0=always`, which the lib's own git inherits. Fixtures use
`git init -b main` with `submodule.<name>.branch = main` set explicitly, so
`--remote` resolves the way production does.
- `T1-ostag-ubuntu` / `T2-ostag-other`: fixture os-release files.
- `T3-errexit-safe`: the bump called as a bare statement under
`set -euo pipefail` with a non-Ubuntu os-release → the script reaches the
next line. Regression test for the latent abort.
- `T4-supports-hit` / `T5-supports-miss`: fixture lib dir with and without the
tag. Proves idempotence both ways without invoking bun.
- `T6-update-success-bumps`: fixture superproject + submodule, an upstream
commit, bump stubbed by redefining it after sourcing → update succeeds, the
stub ran once, returns 0.
- `T7-update-conflict-nondestructive`: local edit to the submodule's
`package.json` plus a conflicting upstream commit → returns non-zero, BOTH
bump-owned files are byte-identical to before the call, and the hint line
was printed. Carries the literal `update-conflict` (contract criterion 4).
- `T8-no-destructive-command`: greps the lib source for `git
checkout|reset|clean|stash` and for `rm|rmdir|unlink|truncate|mv` outside
comments. Stronger than criterion 6 alone.
- `T9-report-referenced`, `T10-report-underscore-dir`
(`chromium_headless_shell-1228` against a `chromium-headless-shell` entry),
`T11-report-unreferenced`, `T12-report-broken-link`,
`T13-report-revision-override`: fixture cache + `.links` → fixture
playwright-core dirs with hand-written `browsers.json`.
- `T14-report-zero-counts-exit-0`: everything referenced → exit 0. The nominal
case, not covered by the absent-dir case.
- `T15-report-no-cache` / `T16-report-browsers-path-zero`: exit 0, nothing on
stderr.
- `T17-source-safe`: sourcing emits nothing.
## Disposition (STEP 0.6)
- honors **BDR-029** by keeping the bump OS-gated, idempotent, non-fatal, and
by closing its stated caveat: the bump is now re-applied after every
successful update, not only at the next `make plugin`.
- honors **LRN-024** by extracting a helper and refactoring the existing
caller before adding the other callers. Deviations from "code MOVED not
changed" are named: `|| true` on the ostag capture (a latent abort on every
non-Ubuntu host, reproduced), and the `timeout` guard.
- honors **LRN-070** by never touching the submodule working tree at all. The
revision that did (discard, retry, restore) was withdrawn at the gate.
- honors **LRN-071** (recurrent 3x) by returning the update's real status, not
a non-fatal helper's 0.
- honors **LRN-040** by touching layer 1 only; `GSTACK_CHROMIUM_NO_SANDBOX`
is untouched.
- honors **LRN-085** by keeping the update idempotent, presence-guarded, no
`--force`.
- honors **LRN-036** by putting `$HOME/.bun/bin` on PATH inside the lib.
- honors **LRN-002** by grepping the moved function name repo-wide, readers
included.
- **LRN-038** already seen: the host-platform override is a dead end.
- BDR-029's reference line (`decisions.md:544`) and BLK-008's caveat
(`blockers.md:118`) describe behavior this plan changes. Registries are
append-only, so the plan does NOT edit them: the /feat CAPITALIZE step
owns the superseding entry.
+8
View File
@@ -97,6 +97,14 @@ skills-disabled/
graphify-out/
.ctx7-cache/
# graphify's vendored skill — written into the repo by `graphify claude
# install` (install-plugins.sh STEP graphify), since ~/.claude/skills is a
# symlink to skills/. Untracked so a tool upgrade stops dirtying the tree.
# test-prompts.json is hand-written for darwin and stays tracked.
skills/graphify/SKILL.md
skills/graphify/references/
skills/graphify/.graphify_version
# /client-handover test artifacts (project-local renders)
LIVRAISON.md
LIVRAISON.html
+108
View File
@@ -6,6 +6,114 @@ Format follows [Keep a Changelog](https://keepachangelog.com/).
## [Unreleased]
### Added
- **`make doctor` reports the Playwright browser cache** — a read-only
`Playwright browsers` section listing cache size, which registered
Playwright install requires each cached browser revision, and counts of
unreferenced directories and broken links. Report only: nothing is
pruned, since Playwright's own `install` already unions the required set
across every registered install.
- `lib/gstack-playwright.sh` — the gstack Playwright helpers as a shared
lib (OS-support bump, submodule-update wrapper, cache report), sourced by
`install-plugins.sh`, `update-all.sh` and `doctor.sh`, covered by
`lib/tests/gstack-playwright.test.sh`.
- **`doctor.sh` inspects the `autoMode` block**: warns when a classifier
list drops the built-in entries (no `"$defaults"`) and when the
user-scope `environment` names a git repo other than the config repo.
Neither defect is visible from the deny count, until now the only
permission signal `doctor.sh` had.
- **`templates/settings/SETTINGS.md` documents `autoMode`**: the four
classifier lists, `$defaults` splice semantics, `classifyAllShell`, the
user-scope vs project-scope rule, and why `ask` is the wrong tier for a
destructive command under auto mode.
### Changed
- **Docker and node go through the classifier with a framing, instead of
an inert `ask` tier.** `Bash(docker run|exec *)`, `Bash(docker[-| ]compose
up*)` and `Bash(node -e *)` leave `permissions.ask` (no prompt under auto
mode, re-verified on 2.1.273). A new `autoMode.allow` list, `$defaults`
first, names the two routine cases the built-in `Remote Shell Writes` /
`Production Reads` rules were catching: `docker exec`/`run`/`compose`
against a local dev container whose name does not carry `prod`, running a
SQL file or script from the repo inside it; and project-local node
(`node <file>`, `npm run`, `npx`/`pnpm exec` of a lockfile-declared
package, effects inside the cwd). Two `soft_deny` entries frame what that
opens: docker data destruction (`rm -f`, `volume rm`/`prune`, `system
prune`, `compose down -v`, `--privileged`, bind mounts outside the cwd)
and undeclared node packages (`npx`/`dlx` of a package absent from the
lockfile, `npm install <name>`). `SETTINGS.md` gains the `autoMode.allow`
tier and the reason a static `Bash(node *)` rule cannot do this job.
- **Ask, don't guess: the orchestrators ask about open choices instead of
settling them.** `CLAUDE.global.md` replaces "one question upfront, never
mid-task" with: a choice visible in the result, a name that becomes
public, or a scope the request does not settle → ask, even mid-task;
internal technical choices stay Claude's. `lib/contract-interview.md`
STEP 2 becomes CLARIFY: pass A (the three gap checks, at contract time)
and pass B (the open-choice sweep in three classes, run once at each
flow's PLAN step, no question cap, over-5 guard, "you decide" recorded as
delegated). New MID-RUN CLARIFICATION section: an executor's
`NEED-DECISION` carries a `CLASS:` tag; visible / public-name / scope go
to the user verbatim, internal is decided in the loop; answers land in
the contract `[gated]`. New HOW TO ASK section (LRN-102). `/feat`,
`/bugfix`, `/hotfix`, `/ship-feature`, `/init-project` wire pass B at
their plan step; `/feat` and `/bugfix` stop deciding `NEED-DECISION`
themselves; `/hotfix` drops "zero questions ever" and allows one
re-dispatch for a class-tagged BLOCKED; the interviewer never ships a
visible / public-name / scope item as `(assumed)`; feater, bugfixer and
hotfixer report the class. Locks updated in the `contract-verifier`,
`loops-light` and `gates` tests.
- **The classifier, not `permissions.ask`, now guards destructive shell
work** (BDR-090). Ten rules left the static tiers: `rsync`, `kill -9`,
`killall`, `pkill` out of `deny`, and `python3 -c`, `python -c`,
`xargs`, `sed`, `cp`, `mv` out of `ask`. Under `defaultMode: auto` an
`ask` rule raises no prompt ([[LRN-146]]), so that tier was gating
nothing anyway. Cover is now `autoMode.soft_deny`, which the classifier
enforces and an explicit instruction clears: writes outside the working
directory, `rsync --delete`, SIGKILL and kill-by-name, in-place edits
spanning more than one file, directory moves, and inline interpreters
or `xargs` that delete or write outside the cwd. Intent clears a soft
block for the current turn only.
- **`autoMode.hard_deny` added** for the three classes no command pattern
can express: secret exfiltration (a read and a send, separate steps,
possibly turns apart), production deployment (deploy scripts, lftp/FTP
pushes, any `prod` target), and disarming the guardrails (weakening a
deny list, `--no-verify`, removing the pre-commit hook,
`bypassPermissions`). Adding a restriction stays allowed; removing one
does not. No instruction clears these.
### Security
- **Ten secret-reader deny rules added**: `sed`, `awk`, `cut`, `tr`,
`sort`, `uniq`, `diff`, `od`, `xxd`, `strings` against `.env*`. Six of
those tools sat in `permissions.allow`, so reading a `.env` through
them triggered nothing. Same shape and same known gap as the existing
`Bash(grep * .env*)` family: a `cat .env | sed` pipe still slips past,
which is what the `hard_deny` exfiltration rule is there to catch.
### Fixed
- **`make update` no longer drops the Playwright OS-support bump** — a
gstack submodule update used to leave the bump unapplied until the next
`make plugin`, the open caveat of BDR-029. `update-all.sh` now goes
through `gstack_submodule_update_with_bump`, which re-applies it after a
successful update and returns non-zero on failure so the existing warn
arm still fires. Two latent bugs travelled with the extracted code: the
ostag capture exited 1 on every non-Ubuntu host and aborted its caller
under inherited `errexit`, and the `bun` calls had no timeout.
- **`autoMode.environment` no longer describes one project from the
user-scope file**: the block named a specific repo, its FTP deploy
target and its customer data, while `link.sh` symlinks this file to
`~/.claude/settings.json` where it reaches every project. The global
block now states machine-level facts only (self-hosted Gitea, gitflow
protection, `~/.claude/.env` as the single secret source, no CI), and
the project-specific facts moved to that project's gitignored
`.claude/settings.local.json`. Both lists now open with `"$defaults"`,
which the original omitted, so the built-in entries are inherited
rather than replaced.
- `README.md` no longer claims the `ask` tier makes every `mcp__magic__*`
call "require a live confirmation and can never auto-execute". That
holds under `defaultMode: default`, not under this config's `auto`. The
paragraph now separates what is verified from what is not, and names
`deny` as the only tier the classifier cannot lift.
## [1.5.0] — 2026-09-13
### Added
+6 -2
View File
@@ -48,11 +48,15 @@ Apply unless repo-specific instructions override.
calls. Skill-mandated gates (fresh verifier/security/challenge)
always dispatch as written. Don't redo delegated work by hand —
failed gates re-dispatch fresh executors instead.
- One question upfront if needed — don't interrupt mid-task.
- Ask rather than guess. A choice visible in the result (placement,
wording, order, behavior), a name that becomes public (command, flag,
endpoint, file), or a scope the request does not settle → ask, even
mid-task. Batch what can be batched. Internal technical choices with
no observable effect stay yours.
*Exception: skill-mandated gates and checkpoints (orchestrator
validation gates, approval gates, darwin checkpoints) always fire.*
- Bug received → fix directly: check logs, find root cause, resolve
autonomously.
autonomously; a visible choice in the fix still gets asked.
- Something goes wrong → STOP, re-plan. Never push through.
- Deviations: minor or clearly justified → do, explain after.
Significant or shaky justification → ask before deviating.
+29
View File
@@ -30,6 +30,35 @@ install-plugins.sh STEP ctx7 purges it right after; the find-docs skill is
the single ctx7 surface. If it reappears (manual `ctx7 setup`), delete it
or re-run `make plugin`.
## Machine-owned: the vendored graphify skill
`skills/graphify/SKILL.md`, `skills/graphify/references/` and
`.graphify_version` are written by `graphify claude install`
(`install-plugins.sh` STEP graphify), which lands in the repo because
`~/.claude/skills` is a symlink to `skills/`. They are gitignored: a
`pipx upgrade graphifyy` used to dirty the tree and cost a
`chore(graphify): sync vendored skill X -> Y` commit each time.
Two graphify commands, easy to confuse, and only one restores the skill:
- `graphify install --platform claude` copies SKILL.md + `references/` +
`.graphify_version` into `skills/graphify/`. Touches nothing else.
This is the recovery command.
- `graphify claude install` writes the CLAUDE.md graphify section and the
`.claude/settings.json` PreToolUse hooks. It **rewrites both guarded
configs** (EVAL-020, verified again 2026-09-15), so revert them after. It does NOT copy
the skill.
`make plugin` runs both (`install-plugins.sh` STEP graphify) behind the
guarded-config EXIT trap, so a fresh clone is covered.
Trade-off accepted: an upstream release can now change the skill's prompt
with no diff to review. `skills/graphify/test-prompts.json` is hand-written
for darwin and stays tracked.
Gotcha, learned the hard way: `git rm --cached` keeps the working file,
but if the branch you merge into still tracks it, the merge deletes it
from disk. Untrack and merge, then restore with the command above.
## Transient planning artifacts
`docs/superpowers/specs/**` and `docs/superpowers/plans/**` are run-time
+10 -5
View File
@@ -300,10 +300,15 @@ check) for up to 10 minutes per call; any local process or open browser tab
can `POST` to it and that body is injected **verbatim** into the tool result
the model consumes (job8 audit, `dist/utils/callback-server.js:36`). This is
in the third-party package's code, not this repo's config — **we don't patch
it**. The mitigation lives entirely on our side: `settings.json`
`permissions.ask` explicitly lists all 4 `mcp__magic__*` tools,
so every call — builder included — requires a live confirmation and can
never auto-execute. Don't allowlist
it**. The mitigation lives on our side: `settings.json`
`permissions.ask` explicitly lists all 4 `mcp__magic__*` tools.
Read that as a declared intent, not a proven hard gate: under
`defaultMode: auto` (this config's default) Bash `ask` rules were observed
auto-approving with no prompt raised (LRN-146). Whether MCP `ask` rules
behave the same has not been verified here, so re-check before relying on
it. `deny` is the only tier the auto-mode classifier cannot lift; for a
gate that holds under auto mode without banning the tool outright, the
right home is `autoMode.soft_deny`. Don't allowlist
`21st_magic_component_builder` or `21st_magic_component_refiner` (arbitrary
absolute-path read → vendor exfil, same audit) under any circumstance.
@@ -337,7 +342,7 @@ make profile-reset # re-enable all gstack skills
make new-skill name=myskill # scaffold agent + skill files
```
`doctor.sh` checks: symlinks, GStack submodule, prerequisites (git, Node, Cargo, Python, Claude Code), plugins, permissions, token budget, config consistency.
`doctor.sh` checks: symlinks, GStack submodule, Playwright browser cache, prerequisites (git, Node, Cargo, Python, Claude Code), plugins, permissions, token budget, config consistency.
---
+5 -3
View File
@@ -27,8 +27,9 @@ Every choice was made in the plan or is a NEED-DECISION to report.
- Apply the FIX PLAN to the letter — fix the ROOT CAUSE named in DIAGNOSIS,
not the symptom. A plan hole or an open choice (naming, data shape, API
surface, dependency) → STOP, report `NEED-DECISION` with the precise
question. Never re-investigate or improvise a different fix.
surface, dependency, a user-visible choice such as placement, wording or
behavior) → STOP, report `NEED-DECISION` with the precise question and
its `CLASS:`. Never re-investigate or improvise a different fix.
- Stay inside the contract FILE SCOPE. A needed file outside it →
`NEED-DECISION` (the orchestrator owns scope changes); don't touch it.
- Add or update the regression test the plan names — it must fail before the
@@ -73,5 +74,6 @@ FILE(S) : <created/modified paths>
TEST(S) : <regression test added/updated + final suite run result, verbatim line>
SMOKE : <build/typecheck result if run, or n/a>
NOTES : <DONE: deviations (must be none) | NEED-DECISION: the exact
question + the options you see | BLOCKED: the blocker verbatim>
question + the options you see + CLASS: visible | public-name |
scope | internal | BLOCKED: the blocker verbatim>
```
+5 -3
View File
@@ -37,8 +37,9 @@ report below is optional on this path (the dispatcher needs the edit applied
## EXECUTION RULES
- Follow the plan to the letter. A plan hole or an open choice (naming,
data shape, API surface, dependency) → STOP, report `NEED-DECISION` with
the precise question. Never improvise a design decision.
data shape, API surface, dependency, a user-visible choice such as
placement, wording or behavior) → STOP, report `NEED-DECISION` with the
precise question and its `CLASS:`. Never improvise a design decision.
- Stay inside the contract FILE SCOPE. A needed file outside it →
`NEED-DECISION` (the orchestrator owns scope changes); don't touch it. On
the applier path the scope is the files named in the bundle item — apply
@@ -84,5 +85,6 @@ STATUS : DONE | NEED-DECISION | BLOCKED
FILES : <created/modified paths>
TESTS : <added/updated + final suite run result, verbatim line>
NOTES : <DONE: deviations (must be none) | NEED-DECISION: the exact
question + the options you see | BLOCKED: the blocker verbatim>
question + the options you see + CLASS: visible | public-name |
scope | internal | BLOCKED: the blocker verbatim>
```
+6 -1
View File
@@ -43,6 +43,10 @@ the edit applied + self-verified, not the report grammar).
BLOCKED`, report why (the orchestrator escalates to `/bugfix`), never
expand scope yourself. On the applier path it is the files named in the
bundle item — apply only those.
- An open user-visible choice the contract does not settle (placement,
wording, behavior) → `STATUS BLOCKED` with `CLASS: visible | public-name |
scope` in NOTES, BEFORE editing anything. The orchestrator asks the user
and re-dispatches once.
- If tests exist for the affected code, run them. Detection cascade:
```bash
# JS/TS
@@ -78,5 +82,6 @@ STATUS : DONE | BLOCKED
FILE(S) : <changed files — suffix files you CREATED with " (new)">
FIX : <one-line description>
SMOKE : <test/build result, verbatim line>
NOTES : <BLOCKED: the blocker; DONE: none>
NOTES : <BLOCKED: the blocker, + CLASS: visible | public-name | scope when
you halted at an open choice before editing; DONE: none>
```
+3 -3
View File
@@ -14,13 +14,13 @@ Gather context. Produce complete PROJECT BRIEF as single source of truth.
- If the initial prompt already provides name + purpose + stack + features + architecture → skip questions and generate the BRIEF directly.
- Otherwise ask only what's genuinely missing, in a single structured block.
- After answers: produce BRIEF. One follow-up allowed if answer is ambiguous.
- Hard budget: 2 question rounds total (initial block + one follow-up). The BRIEF ships after round 2 no matter what — gaps become OPEN DECISIONS, never a third round.
- Hard budget: 2 question rounds total (initial block + one follow-up) for gaps. The BRIEF ships after round 2 — gaps become OPEN DECISIONS. Sole exception: a VISIBLE, PUBLIC NAME or SCOPE choice (a user-facing placement or wording, a public command/flag/endpoint name, whether X is in scope) still open after round 2 gets ONE more targeted question; it never ships as `(assumed)`.
## FAILURE MODES
| Trigger | First response | If still unresolved |
|---|---|---|
| Answer vague/ambiguous | One targeted follow-up on that item only | Record item in OPEN DECISIONS with the safest reading, marked `(assumed)` — never invent a confident value |
| Answer vague/ambiguous | One targeted follow-up on that item only | Gap: record it in OPEN DECISIONS with the safest reading, marked `(assumed)` — never invent a confident value. Visible / public-name / scope item: one more targeted question instead, never `(assumed)` |
| "I don't know / you decide" | Propose ONE concrete default + why, ask yes/no | Take the default, mark `(assumed)`, list in OPEN DECISIONS |
| Contradictory answers (e.g. embedded runtime + managed cloud DB) | Name the contradiction, ask which side wins | Put BOTH options in OPEN DECISIONS; do not silently pick one |
| Partial answer to the block | Re-ask ONLY the missing items in the follow-up round | Missing fields → `none stated` + OPEN DECISIONS entry |
@@ -77,6 +77,6 @@ Stop after BRIEF. Orchestrator handles next step.
- Design, architect, or implement anything — the BRIEF is the entire deliverable.
- Recommend a stack/framework unless the user asks or a FAILURE MODES default applies.
- Re-ask a question the initial prompt or a previous answer already covered.
- Exceed the 2-round budget, whatever is still missing.
- Exceed the 2-round budget for gaps; the only extra question is the single targeted one a visible / public-name / scope item earns.
- Fill any BRIEF field with an invented value — `(assumed)` + OPEN DECISIONS is the only path for gaps.
- Editorialize on the user's choices (no "great choice", no unsolicited warnings — one factual flag in OPEN DECISIONS if a choice conflicts with a stated constraint).
+67 -1
View File
@@ -18,8 +18,10 @@ REPO="$(cd "$(dirname "$0")" && pwd)"
VERSION=$(cat "$REPO/version.txt" 2>/dev/null || echo "unknown")
# Load shared detection library
# shellcheck source=lib/detect-plugins.sh
# shellcheck source=lib/detect-plugins.sh disable=SC1091
source "$REPO/lib/detect-plugins.sh"
# shellcheck source=lib/gstack-playwright.sh disable=SC1091
source "$REPO/lib/gstack-playwright.sh"
echo ""
echo "═══ claude-config doctor (v${VERSION}) ═══"
@@ -115,6 +117,13 @@ fi
echo ""
# ── Playwright browsers (read-only report; NOT nested under gstack — 2 of
# the 3 registered installs are gsd-pi, not gstack) ──
echo "── Playwright browsers ──"
gstack_browsers_report || true
echo ""
# ────────────────────────────────────────────────────────────
# 3. Prerequisites
# ────────────────────────────────────────────────────────────
@@ -206,6 +215,61 @@ echo ""
# ────────────────────────────────────────────────────────────
# 5. Permissions check
# ────────────────────────────────────────────────────────────
# Under defaultMode auto the classifier reads `autoMode`, so a block scoped
# to ONE project feeds every other project false facts, and a list without
# "$defaults" silently drops the built-in rules. Neither is visible from the
# deny count. Emits TAG|message lines for the caller to dispatch.
inspect_automode() {
REPO="$REPO" python3 - "$SETTINGS" <<'PY'
import json, os, re, sys
settings = json.load(open(sys.argv[1]))
mode = settings.get("permissions", {}).get("defaultMode")
block = settings.get("autoMode") or {}
if mode != "auto":
sys.exit(print("INFO|defaultMode is %s, autoMode not consulted" % mode))
if not block:
sys.exit(print("WARN|defaultMode is auto but no autoMode block set"))
sections = [k for k in ("allow", "soft_deny", "hard_deny", "environment")
if k in block]
bare = [k for k in sections if "$defaults" not in block[k]]
if bare:
print('WARN|autoMode.%s replaces the built-in entries (no "$defaults")'
% ", ".join(bare))
else:
print('PASS|autoMode: %s inherit "$defaults"' % ", ".join(sections))
repo, home = os.environ["REPO"], os.path.expanduser("~")
foreign = {q for entry in block.get("environment", [])
for q in re.findall(r"`(/[^`]+)`", entry)
if (p := q.rstrip("/")).startswith(home) and p != repo
and os.path.isdir(os.path.join(p, ".git"))}
if foreign:
print("WARN|autoMode.environment names another repo (%s); this file is "
"user-scope and reaches every project" % ", ".join(sorted(foreign)))
else:
print("PASS|autoMode.environment is not scoped to a foreign repo")
PY
}
check_automode() {
local out tag msg
if ! out=$(inspect_automode 2>/dev/null); then
warn "Could not inspect the autoMode block"
return
fi
while IFS='|' read -r tag msg; do
case "$tag" in
PASS) pass "$msg" ;;
WARN) warn "$msg" ;;
INFO) info "$msg" ;;
esac
done <<< "$out"
}
echo "── Permissions ──"
SETTINGS="$HOME/.claude/settings.json"
@@ -242,6 +306,8 @@ print(len(json.load(sys.stdin).get('permissions',{}).get('deny',[])))
warn "Deny rules: $DENY_COUNT (committed: $EXPECTED_DENY) — live settings diverge from last commit"
fi
fi
check_automode
else
fail "$HOME/.claude/settings.json not found"
fi
+5 -31
View File
@@ -26,8 +26,10 @@ else
fi
# Load shared detection library
# shellcheck source=lib/detect-plugins.sh
# shellcheck source=lib/detect-plugins.sh disable=SC1091
source "$REPO/lib/detect-plugins.sh"
# shellcheck source=lib/gstack-playwright.sh disable=SC1091
source "$REPO/lib/gstack-playwright.sh"
# ── Guard hand-curated config against installer drift ────────
# graphify's installer (Step 7) rewrites CLAUDE.md + .claude/settings.json
@@ -291,35 +293,6 @@ fi
echo ""
# gstack pins Playwright (1.58.x) which only ships browser builds for
# ubuntu<=24.04. On a newer distro the browser install fails ("does not
# support chromium on ubuntuXX.04"). Bump gstack's Playwright to a version
# that supports this OS so ./setup builds the browse binary against it and
# installs a native browser. Fires only when the pinned version genuinely
# lacks support — idempotent across runs. Edits the submodule locally (goes
# dirty); a `git submodule update` resets it and the next install re-applies.
# See BLK-008 / LRN-040.
gstack_bump_playwright_if_unsupported() {
[ -d "$GSTACK_DIR" ] && [ -r /etc/os-release ] || return 0
local ostag pwlib
# shellcheck disable=SC1091
ostag="$(. /etc/os-release 2>/dev/null; [ "${ID:-}" = ubuntu ] && printf 'ubuntu%s' "${VERSION_ID:-}")"
[ -n "$ostag" ] || return 0 # only the known Ubuntu case
pwlib="$GSTACK_DIR/node_modules/playwright-core/lib"
# populate node_modules at the pinned version so we can read its support list
( cd "$GSTACK_DIR" && { bun install --frozen-lockfile >/dev/null 2>&1 || bun install >/dev/null 2>&1; } ) || return 0
if grep -rqs "$ostag" "$pwlib" 2>/dev/null; then
return 0 # pinned Playwright already supports this OS
fi
info "gstack's Playwright lacks $ostag support — bumping to latest (local submodule edit)..."
( cd "$GSTACK_DIR" && bun add playwright@latest >/dev/null 2>&1 )
if grep -rqs "$ostag" "$pwlib" 2>/dev/null; then
ok "gstack Playwright bumped — now supports $ostag (browse binary rebuilt by ./setup)"
else
warn "Playwright bump didn't add $ostag support — gstack browser may stay unavailable"
fi
}
# ============================================================
# STEP 2 — GSTACK SUBMODULE
# ============================================================
@@ -367,7 +340,8 @@ if [ -d "$GSTACK_DIR" ]; then
# BEFORE ./setup so its frozen-lockfile install picks up the new version and
# the browse binary is rebuilt against it (avoids the "does not support
# chromium" fail). Non-fatal if it can't — gstack is OFF by default.
gstack_bump_playwright_if_unsupported
# See BLK-008 / LRN-040 / BDR-029; logic lives in lib/gstack-playwright.sh.
gstack_bump_playwright_if_unsupported "$GSTACK_DIR"
info "Running GStack setup..."
_gstack_setup_ok=0
+72 -9
View File
@@ -7,8 +7,9 @@ subagents = execution + report only; gates and loop decisions live in the
main loop).
Run this in the ORCHESTRATOR MAIN LOOP, never in a subagent — STEP 2 may
talk to the human. Mandatory passage in every flow; questions are optional
and proportional — a complete request goes through silently.
talk to the human, at contract time (pass A) and again at the flow's PLAN
step (pass B). Questions follow the open choices, never a quota — a complete
request goes through silently.
## STEP 1 — CAPTURE (verbatim)
@@ -17,16 +18,51 @@ message). No paraphrase, no cleanup, no translation, no summarizing. This
section is IMMUTABLE for the life of the run — every later consumer
(planner, dev, verifier) reads THESE words, never a restatement.
## STEP 2 — AMBIGUITY CHECK (questions optional, proportional)
## STEP 2 — CLARIFY (ask, never guess)
Ask ONLY if one of these is missing AND not derivable from the repo:
Two passes, both in the main loop, both may talk to the human.
**Pass A — gaps.** Run here, against the request. Ask if one of these is
missing AND not derivable from the repo:
- a testable expected outcome
- an unambiguous scope (what is allowed to change)
- non-contradictory constraints
Complete request → ZERO questions, stay silent. Otherwise: max 3 questions,
one single batch (house rule: one question upfront, never mid-task). Never
ask what the repo can answer — verify paths/APIs/behavior yourself first.
**Pass B — open choices.** Defined here, run ONCE at the flow's PLAN step
(see "Where pass B fires" below), against the plan just written — that is
where choices become concrete. Enumerate every choice the run will settle
that the request leaves open; keep those in these classes:
1. VISIBLE — the user would see it in the result: placement, label, wording,
color, order, what a click does.
2. PUBLIC NAME — a name that outlives the run: command, flag, endpoint, env
var, a file the human will read.
3. SCOPE — "should X change too?", where the request does not name X.
NEVER ask class 4 — internal technical choices with no observable effect
(function decomposition, data shape, local naming, layout inside an
already-scoped zone). Those are delegated; asking them is the noise that
makes classes 1-3 ignorable. Never ask what the repo or the request already
answers — verify paths/APIs/behavior yourself first.
No question cap. Each pass asks what it finds, in ONE batch. A request that
leaves nothing open goes through silently. More than 5 open choices in pass B
= the request is under-specified: list them, say so, stop — do not fire a
questionnaire. "You decide" / "peu importe" is an answer: record it as
`A: delegated — <default taken>` and never re-ask it.
Pass B answers land in the contract's CLARIFICATIONS marked
`[gated <YYYY-MM-DD>]` — the contract is already on disk by then.
### Where pass B fires
| Flow | Pass B runs at | Against |
|------|----------------|---------|
| feat | STEP 1 PLAN, before 1b CHALLENGE | the PLAN checklist |
| bugfix | STEP 3 FIX PLAN, before 3b | the FIX PLAN |
| hotfix | STEP 1 LOCATE | the 1-2 target files' visible effect |
| ship-feature | STEP 2 PLAN, after the brainstorm | the plan, minus what the brainstorm settled |
| init-project | STEP 3 DESIGN, before VALIDATION GATE #1 | the DESIGN, minus what the interview and brainstorm settled |
| onboard | its STEP 3 interview, unchanged | scope, in one block |
## STEP 3 — DERIVE
@@ -85,7 +121,7 @@ Template:
<the user's exact words>
## CLARIFICATIONS
Q: <question> / A: <answer>
Q: <question> / A: <answer> (pass B and mid-run entries: [gated <YYYY-MM-DD>])
(or: none — request complete)
## ACCEPTANCE CRITERIA
@@ -105,6 +141,33 @@ Q: <question> / A: <answer>
Print one line to the user, then continue the flow:
`CONTRACT: <path> — <n> criteria, scope <files|repo-wide>, <q> questions asked`
## MID-RUN CLARIFICATION (the channel executors halt into)
An executor cannot talk to the human. It halts with `NEED-DECISION`, the
exact question, the options it sees, and a `CLASS:` tag (visible |
public-name | scope | internal). `/hotfix`: the hotfixer keeps
`DONE | BLOCKED`; a BLOCKED carrying the tag follows the same routing instead
of escalating to `/bugfix`. The orchestrator re-reads the class — the tag is
a hint, not a verdict — then routes:
- visible / public-name / scope → ASK THE HUMAN, verbatim question and
options. Never decide these yourself, never spend a round-trip guessing.
- internal → decide here, note the decision, re-dispatch. The only case the
orchestrator settles alone; max 2 such round-trips → escalate.
Every answer, human or orchestrator, appends to the contract's
CLARIFICATIONS marked `[gated <YYYY-MM-DD>]` — the same micro-gate as scope
enrichment — and to the plan handed to the FRESH re-dispatched executor,
which reads the decision from disk, never from a transcript.
## HOW TO ASK (LRN-102)
The harness reliably renders only the turn's FINAL text; text printed before
a tool call may be swallowed. So:
- up to 4 questions → one `AskUserQuestion` call; option descriptions carry
the context; print nothing the user needs before the call.
- more than 4, or a list handed back for re-specification → plain text, end
the turn.
## Lifecycle
- **REQUEST**: immutable, for the life of the run. Never rewritten, never
@@ -134,7 +197,7 @@ Print one line to the user, then continue the flow:
| Flow | Weight |
|------|--------|
| hotfix | Silent autofill — criteria: "symptom gone; build/tests green"; scope = the 1-2 target files. Zero questions ever. |
| hotfix | Pass A silent autofill — criteria: "symptom gone; build/tests green"; scope = the 1-2 target files. Pass B runs at LOCATE against the 1-2 target files' visible effect; a typo fix asks nothing. |
| feat / bugfix | Proportional. bugfix: the DIAGNOSIS feeds the criteria (symptom reproduced-then-gone + regression test present). |
| ship-feature | Full. Design decisions approved at the validation gate append criteria `[gated <date>]` — the human validates the enriched contract, the verifier receives that version. |
| init-project | Full. The interviewer's PROJECT BRIEF pours into the contract (V1 features → criteria). |
+296
View File
@@ -0,0 +1,296 @@
#!/usr/bin/env bash
# ============================================================
# lib/gstack-playwright.sh — gstack's Playwright: OS-support bump +
# read-only browser-cache report.
#
# Sourced by: install-plugins.sh, update-all.sh, doctor.sh — all three run
# `set -euo pipefail`. gstack_bump_playwright_if_unsupported and
# gstack_browsers_report are called as BARE STATEMENTS under that inherited
# errexit, so they `return 0` on every path and every capture that could
# fail is guarded (`|| true` or an `if`), never a bare `&&`/`||`-less
# statement. gstack_submodule_update_with_bump is the ONE function allowed
# to return non-zero — callers use it ONLY as an `if` condition.
#
# No `set -euo pipefail` here (mirrors lib/detect-plugins.sh): a sourced
# lib must not change the caller's shell options.
#
# See BDR-029 (bump origin), BLK-008 (Chromium-unsupported-OS saga),
# LRN-040 (two-layer fix — this file is layer 1 only).
# ============================================================
_GSPW_GREEN='\033[0;32m'; _GSPW_YELLOW='\033[1;33m'; _GSPW_BLUE='\033[0;34m'
_GSPW_NC='\033[0m'
_gspw_ok() { echo -e " ${_GSPW_GREEN}✓${_GSPW_NC} $1"; }
_gspw_warn() { echo -e " ${_GSPW_YELLOW}⚠${_GSPW_NC} $1"; }
_gspw_info() { echo -e " ${_GSPW_BLUE}→${_GSPW_NC} $1"; }
# ── OS support ───────────────────────────────────────────────────────────
# gstack_pw_ostag [os_release_path] — "ubuntu<VERSION_ID>" on Ubuntu, empty
# otherwise. `|| true` on the capture: the reproduced bug had this exact
# line abort every non-Ubuntu host under inherited errexit.
gstack_pw_ostag() {
local path="${1:-/etc/os-release}" tag
[ -r "$path" ] || return 0
# shellcheck disable=SC1090
tag="$(. "$path" 2>/dev/null
[ "${ID:-}" = ubuntu ] && printf 'ubuntu%s' "${VERSION_ID:-}")" || true
if [ -n "$tag" ]; then
printf '%s' "$tag"
fi
return 0
}
# gstack_pw_supports <playwright_core_lib_dir> <ostag> — 0 supported, 1 not.
# Always called from an `if`/`&&` context, never as a bare statement.
gstack_pw_supports() {
local pwlib="$1" ostag="$2"
[ -n "$ostag" ] && [ -d "$pwlib" ] || return 1
grep -rqs "$ostag" "$pwlib" 2>/dev/null
}
# _gspw_run_timeout <dir> <cmd...> — runs <cmd> in <dir>, under `timeout 300`
# when available (absent on stock macOS). Exit 124 = the wrapped command was
# killed by the timeout. Callers MUST invoke this via `cmd || rc=$?` (never
# bare) so a non-zero exit never trips the caller's inherited errexit.
_gspw_run_timeout() {
local dir="$1"; shift
if command -v timeout >/dev/null 2>&1; then
( cd "$dir" && timeout 300 "$@" ) >/dev/null 2>&1
else
( cd "$dir" && "$@" ) >/dev/null 2>&1
fi
}
# _gspw_bump_install <gstack_dir> — populate node_modules at the pinned
# version so its support list can be read. 0 proceed, 1 give up silently
# (both installs failed, matches the pre-existing silent behavior), 2 give
# up loud (a timeout truncated node_modules — the support grep would then
# read a half-written tree).
_gspw_bump_install() {
local dir="$1" rc=0
_gspw_run_timeout "$dir" bun install --frozen-lockfile || rc=$?
if [ "$rc" -eq 0 ]; then
return 0
elif [ "$rc" -eq 124 ]; then
_gspw_warn "bun install timed out — skipping Playwright bump"
return 2
fi
rc=0
_gspw_run_timeout "$dir" bun install || rc=$?
if [ "$rc" -eq 0 ]; then
return 0
elif [ "$rc" -eq 124 ]; then
_gspw_warn "bun install timed out — skipping Playwright bump"
return 2
fi
return 1
}
# _gspw_bump_add_latest <gstack_dir> — 0 ran (support re-checked by caller
# regardless of bun's own exit code, exactly as the pre-existing code did),
# 2 timed out (node_modules left half-written — caller must NOT re-check).
_gspw_bump_add_latest() {
local dir="$1" rc=0
_gspw_run_timeout "$dir" bun add playwright@latest || rc=$?
if [ "$rc" -eq 124 ]; then
_gspw_warn "bun add playwright@latest timed out — skipping Playwright bump"
return 2
fi
return 0
}
# gstack_bump_playwright_if_unsupported <gstack_dir> — BDR-029: bump
# gstack's pinned Playwright when it lacks a build for this OS, so
# `./setup` rebuilds the browse binary against a version that has one.
# OS-gated, idempotent, non-fatal — `return 0` on every path.
gstack_bump_playwright_if_unsupported() {
local gstack_dir="$1" ostag pwlib rc=0
[ -d "$gstack_dir" ] && [ -r /etc/os-release ] || return 0
ostag="$(gstack_pw_ostag)"
[ -n "$ostag" ] || return 0
if ! command -v bun >/dev/null 2>&1; then
export PATH="$HOME/.bun/bin:$PATH"
fi
pwlib="$gstack_dir/node_modules/playwright-core/lib"
_gspw_info "checking gstack's Playwright OS support ($ostag)..."
_gspw_bump_install "$gstack_dir" || rc=$?
[ "$rc" -eq 0 ] || return 0
if gstack_pw_supports "$pwlib" "$ostag"; then
return 0
fi
_gspw_info "gstack's Playwright lacks $ostag support — bumping to \
latest (local submodule edit)..."
rc=0
_gspw_bump_add_latest "$gstack_dir" || rc=$?
[ "$rc" -eq 0 ] || return 0
if gstack_pw_supports "$pwlib" "$ostag"; then
_gspw_ok "gstack Playwright bumped — now supports $ostag (browse \
binary rebuilt by ./setup)"
else
_gspw_warn "Playwright bump didn't add $ostag support — gstack \
browser may stay unavailable"
fi
return 0
}
# ── Submodule update ──────────────────────────────────────────────────────
# gstack_submodule_update_with_bump <repo> [sub_path] — the ONE function
# allowed to return non-zero; callers use it ONLY as an `if` condition.
# Never touches the submodule working tree: on failure it prints git's own
# stderr verbatim (never parsed) and returns 1. On success it re-applies
# the bump (closes BDR-029's caveat: the bump used to survive only until
# the next `git submodule update`).
gstack_submodule_update_with_bump() {
local repo="$1" sub="${2:-skills-external/gstack}" err rc=0
err="$(git -C "$repo" submodule update --remote "$sub" 2>&1 >/dev/null)" \
|| rc=$?
if [ "$rc" -eq 0 ]; then
gstack_bump_playwright_if_unsupported "$repo/$sub"
return 0
fi
_gspw_warn "$err"
if [ -n "$(git -C "$repo/$sub" status --porcelain \
-- package.json bun.lock 2>/dev/null)" ]; then
_gspw_info "local Playwright bump (package.json/bun.lock) was not \
re-applied — re-run: make plugin"
fi
return 1
}
# ── Browsers report (read-only) ───────────────────────────────────────────
# _gspw_dir_name_parts <cache_dir_name> — prints "normalized_name revision"
# split on the LAST '-', mapping '_' -> '-' on the name (Playwright writes
# chromium_headless_shell-1228 on disk; browsers.json names it
# chromium-headless-shell).
_gspw_dir_name_parts() {
local rev="${1##*-}" name="${1%-*}"
printf '%s %s' "${name//_/-}" "$rev"
}
# _gspw_browser_referenced <playwright_core_path> <dir_name> — does that
# install require this cache directory (base revision or any
# revisionOverrides value)?
_gspw_browser_referenced() {
local json="$1/browsers.json" name rev
[ -r "$json" ] || return 1
read -r name rev <<< "$(_gspw_dir_name_parts "$2")"
awk -F'"' -v want_name="$name" -v want_rev="$rev" '
$2 == "name" { cur = $4; in_ov = 0 }
$2 == "revision" && !in_ov && cur == want_name && $4 == want_rev {
found = 1
}
$2 == "revisionOverrides" { in_ov = 1 }
in_ov && $2 != "revisionOverrides" && cur == want_name \
&& $4 == want_rev { found = 1 }
/^[[:space:]]*}/ { in_ov = 0 }
END { exit !found }
' "$json"
}
# _gspw_browser_name_known <playwright_core_path> <dir_name> — is the NAME
# listed at all, regardless of revision? (distinguishes "unknown revision"
# from "unreferenced" in the report.)
_gspw_browser_name_known() {
local json="$1/browsers.json" name rev
[ -r "$json" ] || return 1
read -r name rev <<< "$(_gspw_dir_name_parts "$2")"
awk -F'"' -v want="$name" '$2 == "name" && $4 == want { found = 1 }
END { exit !found }' "$json"
}
# _gspw_install_label <playwright_core_path> — "<dir-before-node_modules>
# <version>", e.g. "gstack 1.61.1".
_gspw_install_label() {
local pw_path="$1" parent version
parent=$(basename "$(dirname "$(dirname "$pw_path")")")
version=$(awk -F'"' '$2 == "version" { print $4; exit }' \
"$pw_path/package.json" 2>/dev/null) || true
printf '%s %s' "$parent" "${version:-?}"
}
# _gspw_registered_installs <cache_dir> — valid playwright-core paths (dir
# exists, browsers.json readable), one per line. A `.links` entry whose
# target is gone or unreadable is silently excluded here (it is counted as
# a broken link by the caller instead).
_gspw_registered_installs() {
local links_dir="$1/.links" f target
[ -d "$links_dir" ] || return 0
for f in "$links_dir"/*; do
[ -f "$f" ] || continue
target=$(cat "$f" 2>/dev/null) || true
[ -n "$target" ] || continue
if [ -d "$target" ] && [ -r "$target/browsers.json" ]; then
printf '%s\n' "$target"
fi
done
return 0
}
# _gspw_report_dir_line <dir_name> <install_paths_newline_sep> — prints the
# report line for one cache directory. Returns 1 only when truly
# unreferenced (caller tallies that); "unknown revision" does not count.
_gspw_report_dir_line() {
local dir_name="$1" installs="$2" p labels="" known=0
while IFS= read -r p; do
[ -n "$p" ] || continue
if _gspw_browser_referenced "$p" "$dir_name"; then
labels="${labels:+$labels, }$(_gspw_install_label "$p")"
elif _gspw_browser_name_known "$p" "$dir_name"; then
known=1
fi
done <<< "$installs"
if [ -n "$labels" ]; then
_gspw_info "$dir_name: $labels"
return 0
elif [ "$known" -eq 1 ]; then
_gspw_info "$dir_name: unknown revision"
return 0
fi
_gspw_info "$dir_name: unreferenced"
return 1
}
# gstack_browsers_report [cache_dir] — read-only. `$1` (or
# PLAYWRIGHT_BROWSERS_PATH, or ~/.cache/ms-playwright) is resolved once;
# "0" (documented as "bundle into node_modules") and any non-directory
# degrade to a silent no-cache path. `return 0` on every path.
gstack_browsers_report() {
local cache installs total links_total valid_count broken=0 unref=0 d name
cache="${1:-${PLAYWRIGHT_BROWSERS_PATH:-$HOME/.cache/ms-playwright}}"
[ "$cache" = "0" ] && return 0
[ -d "$cache" ] || return 0
installs="$(_gspw_registered_installs "$cache")"
links_total=$(find "$cache/.links" -maxdepth 1 -type f 2>/dev/null \
| wc -l | tr -d ' ') || true
valid_count=$(printf '%s\n' "$installs" | grep -c . || true)
broken=$((links_total - valid_count))
total=$(du -sh "$cache" 2>/dev/null | awk '{print $1}') || true
_gspw_info "Playwright browsers: $cache (${total:-0})"
for d in "$cache"/*-[0-9]*; do
[ -d "$d" ] || continue
name=$(basename "$d")
_gspw_report_dir_line "$name" "$installs" || unref=$((unref + 1))
done
_gspw_info "${unref} unreferenced, ${broken} broken link(s)"
if [ "$unref" -gt 0 ] || [ "$broken" -gt 0 ]; then
_gspw_warn "unreferenced/broken Playwright browser dirs — re-run \
\`playwright install\`, which prunes stale revisions"
fi
return 0
}
# ── CLI dispatch (only when executed, not sourced) — browsers-report ONLY.
# The write functions (the bump, the submodule update) stay sourced-only: a
# CLI verb would expose `bun add playwright@latest` as a command-line entry
# point. ────────────────────────────────────────────────────────────────
if [ "${BASH_SOURCE[0]}" = "${0}" ]; then
case "${1:-}" in
browsers-report) shift; gstack_browsers_report "$@" ;;
*) echo "usage: gstack-playwright.sh browsers-report [cache_dir]" >&2
exit 2 ;;
esac
fi
+9 -2
View File
@@ -48,8 +48,15 @@ fi
tf "verbatim request immutable" "$LIB" "REQUEST (verbatim — IMMUTABLE)"
tf "contracts dir committed path" "$LIB" ".claude/tasks/contracts/"
tf "unique per-run slug" "$LIB" "<YYYY-MM-DD>-<slug>-<HHMM>"
tf "silent when complete" "$LIB" "ZERO questions"
tf "question budget" "$LIB" "max 3 questions"
tf "silent when nothing open" "$LIB" "goes through silently"
tf "no question cap" "$LIB" "No question cap"
tf "pass B classes" "$LIB" "PUBLIC NAME"
tf "class 4 excluded" "$LIB" "NEVER ask class 4"
tf "over-5 guard" "$LIB" "More than 5 open choices"
tf "delegated answer" "$LIB" "delegated —"
tf "mid-run channel" "$LIB" "## MID-RUN CLARIFICATION"
tf "class tag" "$LIB" "CLASS:"
tf "how to ask" "$LIB" "## HOW TO ASK"
tf "aborted status" "$LIB" "status: aborted"
tf "never left dirty" "$LIB" "NEVER left dirty"
tf "scope enrichment micro-gate" "$LIB" "micro-gate"
+3
View File
@@ -311,6 +311,9 @@ lock "bugfixer passes" "$BF" "## FOUR PASSES"
lock "bugfixer stays minimal" "$BF" "keep the fix minimal"
lock "bugfixer neg control" "$BF" "**Negative control.**"
lock "bugfixer test must fail" "$BF" "A test that passes both ways"
lock "feater class tag" "$FE" "CLASS:"
lock "bugfixer class tag" "$BF" "CLASS:"
lock "hotfixer class tag" "$REPO/agents/hotfixer.md" "CLASS:"
echo ""
echo "gates: $PASS pass, $FAIL fail"
+213
View File
@@ -0,0 +1,213 @@
#!/usr/bin/env bash
# lib/tests/gstack-playwright.test.sh — lib/gstack-playwright.sh (T1..T17)
#
# git 2.53 defaults protocol.file to "user", which blocks submodule clone
# and fetch. The fixture git calls alone are not enough: the
# `git submodule update --remote` under test runs INSIDE the lib, in a
# fresh git subprocess spawned from THIS process — so the override is
# exported for the WHOLE test process, not passed per-command.
set -u
export GIT_CONFIG_COUNT=1
export GIT_CONFIG_KEY_0=protocol.file.allow
export GIT_CONFIG_VALUE_0=always
ROOT="$(cd "$(dirname "$0")/../.." && pwd)"
L="$ROOT/lib/gstack-playwright.sh"
pass=0; fail=0
check() { if [ "$2" = "$3" ]; then pass=$((pass+1)); else fail=$((fail+1));
printf 'FAIL %s: got[%s] want[%s]\n' "$1" "$2" "$3"; fi; }
tmp="$(mktemp -d)"; trap 'rm -rf "$tmp"' EXIT
git_id() { git -C "$1" config user.email t@example.com
git -C "$1" config user.name Test; }
# shellcheck source=lib/gstack-playwright.sh
source "$L"
# ── T1/T2 — ostag detection ──────────────────────────────────────────────
printf 'ID=ubuntu\nVERSION_ID="24.04"\n' > "$tmp/os-ubuntu"
printf 'ID=debian\nVERSION_ID="12"\n' > "$tmp/os-debian"
check T1-ostag-ubuntu "$(gstack_pw_ostag "$tmp/os-ubuntu")" "ubuntu24.04"
check T2-ostag-other "$(gstack_pw_ostag "$tmp/os-debian")" ""
# ── T3 — errexit safety of the ostag capture (regression: the reproduced
# bug aborted the whole caller on every non-Ubuntu host) ──
cat > "$tmp/t3.sh" <<EOF
#!/usr/bin/env bash
set -euo pipefail
source "$L"
gstack_pw_ostag "$tmp/os-debian"
echo REACHED
EOF
check T3-errexit-safe "$(bash "$tmp/t3.sh" 2>&1 | tail -1)" "REACHED"
# ── T4/T5 — pw_supports, no bun involved ─────────────────────────────────
mkdir -p "$tmp/pwlib-hit" "$tmp/pwlib-miss"
echo "supports ubuntu24.04 and others" > "$tmp/pwlib-hit/index.js"
echo "supports nothing relevant" > "$tmp/pwlib-miss/index.js"
t4_rc=0; gstack_pw_supports "$tmp/pwlib-hit" ubuntu24.04 >/dev/null 2>&1 \
|| t4_rc=$?
check T4-supports-hit "$t4_rc" 0
t5_rc=0; gstack_pw_supports "$tmp/pwlib-miss" ubuntu24.04 >/dev/null 2>&1 \
|| t5_rc=$?
check T5-supports-miss "$t5_rc" 1
# ── T6/T7 — submodule update, real git fixtures ──────────────────────────
mkdir -p "$tmp/upstream6"
git -C "$tmp/upstream6" init -q -b main; git_id "$tmp/upstream6"
printf '{"a":1}\n' > "$tmp/upstream6/package.json"
git -C "$tmp/upstream6" add package.json
git -C "$tmp/upstream6" commit -q -m init
mkdir -p "$tmp/repo6"
git -C "$tmp/repo6" init -q -b main; git_id "$tmp/repo6"
printf 'x\n' > "$tmp/repo6/README.md"
git -C "$tmp/repo6" add README.md
git -C "$tmp/repo6" commit -q -m init
git -C "$tmp/repo6" -c protocol.file.allow=always \
submodule add -q -b main "$tmp/upstream6" gstack-sub
git -C "$tmp/repo6" config submodule.gstack-sub.branch main
git -C "$tmp/repo6" commit -q -m "add submodule"
printf 'extra\n' > "$tmp/upstream6/extra.txt"
git -C "$tmp/upstream6" add extra.txt
git -C "$tmp/upstream6" commit -q -m "upstream update"
t6_out=$(
gstack_bump_playwright_if_unsupported() { echo BUMP_CALLED; }
gstack_submodule_update_with_bump "$tmp/repo6" "gstack-sub"
echo "rc=$?"
)
t6_calls=$(printf '%s\n' "$t6_out" | grep -c BUMP_CALLED)
t6_rc=$(printf '%s\n' "$t6_out" | grep -o 'rc=[0-9]*')
check T6-update-success-bumps "$t6_calls:$t6_rc" "1:rc=0"
mkdir -p "$tmp/upstream7"
git -C "$tmp/upstream7" init -q -b main; git_id "$tmp/upstream7"
printf '{"a":1}\n' > "$tmp/upstream7/package.json"
printf 'lockA\n' > "$tmp/upstream7/bun.lock"
git -C "$tmp/upstream7" add package.json bun.lock
git -C "$tmp/upstream7" commit -q -m init
mkdir -p "$tmp/repo7"
git -C "$tmp/repo7" init -q -b main; git_id "$tmp/repo7"
printf 'x\n' > "$tmp/repo7/README.md"
git -C "$tmp/repo7" add README.md
git -C "$tmp/repo7" commit -q -m init
git -C "$tmp/repo7" -c protocol.file.allow=always \
submodule add -q -b main "$tmp/upstream7" gstack-sub
git -C "$tmp/repo7" config submodule.gstack-sub.branch main
git -C "$tmp/repo7" commit -q -m "add submodule"
# upstream changes package.json content (would overwrite the local edit)
printf '{"a":2}\n' > "$tmp/upstream7/package.json"
git -C "$tmp/upstream7" add package.json
git -C "$tmp/upstream7" commit -q -m "upstream bumps package.json"
# local Playwright-bump-style dirty edit, never committed
printf '{"a":99}\n' > "$tmp/repo7/gstack-sub/package.json"
echo "T7: update-conflict"
before_pkg=$(cat "$tmp/repo7/gstack-sub/package.json")
before_lock=$(cat "$tmp/repo7/gstack-sub/bun.lock")
t7_out=$(gstack_submodule_update_with_bump "$tmp/repo7" "gstack-sub" 2>&1)
t7_rc=$?
after_pkg=$(cat "$tmp/repo7/gstack-sub/package.json")
after_lock=$(cat "$tmp/repo7/gstack-sub/bun.lock")
t7_files_ok=N
[ "$before_pkg" = "$after_pkg" ] && [ "$before_lock" = "$after_lock" ] \
&& t7_files_ok=Y
t7_hint_ok=N
printf '%s\n' "$t7_out" | grep -q 'make plugin' && t7_hint_ok=Y
t7_state="$t7_rc:$t7_files_ok:$t7_hint_ok"
check T7-update-conflict-nondestructive "$t7_state" "1:Y:Y"
# ── T8 — no destructive command anywhere in the lib source ───────────────
d8=OK
sed 's/#.*//' "$L" | grep -qE 'git [^|;]*(checkout|reset|clean|stash)' && d8=BAD
sed 's/#.*//' "$L" | grep -qwE '(rm|rmdir|unlink|truncate|mv)' && d8=BAD
check T8-no-destructive-command "$d8" OK
# ── T9-T14 — browsers-report, fixture cache + playwright-core installs ───
mkdir -p "$tmp/installs/fixA/node_modules/playwright-core"
cat > "$tmp/installs/fixA/node_modules/playwright-core/browsers.json" <<'EOF'
{
"comment": "Do not edit this file, use utils/roll_browser.js",
"browsers": [
{
"name": "chromium",
"revision": "1228",
"installByDefault": true
},
{
"name": "chromium-headless-shell",
"revision": "1228",
"installByDefault": true
},
{
"name": "webkit",
"revision": "2311",
"installByDefault": true,
"revisionOverrides": {
"mac14": "2251",
"debian11-x64": "2105"
}
},
{
"name": "ffmpeg",
"revision": "1011",
"installByDefault": true
}
]
}
EOF
cat > "$tmp/installs/fixA/node_modules/playwright-core/package.json" <<'EOF'
{
"name": "playwright-core",
"version": "1.61.1"
}
EOF
mkdir -p "$tmp/cache1/.links" \
"$tmp/cache1/chromium-1228" \
"$tmp/cache1/chromium_headless_shell-1228" \
"$tmp/cache1/webkit-2105" \
"$tmp/cache1/firefox-9999" \
"$tmp/cache1/chromium-9999"
printf '%s' "$tmp/installs/fixA/node_modules/playwright-core" \
> "$tmp/cache1/.links/link-valid"
printf '%s' "$tmp/no-such-install/node_modules/playwright-core" \
> "$tmp/cache1/.links/link-broken"
out1="$(gstack_browsers_report "$tmp/cache1" 2>&1)"
has1() { printf '%s\n' "$out1" | grep -q "$1" && echo Y; }
check T9-report-referenced "$(has1 'chromium-1228: fixA 1.61.1')" Y
check T10-report-underscore-dir \
"$(has1 'chromium_headless_shell-1228: fixA 1.61.1')" Y
t11_unref=$(has1 'firefox-9999: unreferenced')
t11_unknown=$(has1 'chromium-9999: unknown revision')
check T11-report-unreferenced "$t11_unref$t11_unknown" YY
check T12-report-broken-link "$(has1 '1 broken link')" Y
check T13-report-revision-override "$(has1 'webkit-2105: fixA 1.61.1')" Y
mkdir -p "$tmp/cache2/.links" "$tmp/cache2/chromium-1228"
printf '%s' "$tmp/installs/fixA/node_modules/playwright-core" \
> "$tmp/cache2/.links/link-valid"
out2="$(gstack_browsers_report "$tmp/cache2" 2>&1)"; rc2=$?
zero2=$(printf '%s\n' "$out2" | grep -q '0 unreferenced, 0 broken link(s)' \
&& echo Y)
check T14-report-zero-counts-exit-0 "$rc2:$zero2" "0:Y"
# ── T15/T16 — degrade silently, nothing on stderr ────────────────────────
err15="$(gstack_browsers_report "$tmp/does-not-exist-cache" 2>&1 1>/dev/null)"
rc15=$?
check T15-report-no-cache "$rc15:[$err15]" "0:[]"
err16="$(gstack_browsers_report "0" 2>&1 1>/dev/null)"
rc16=$?
check T16-report-browsers-path-zero "$rc16:[$err16]" "0:[]"
# ── T17 — sourcing emits nothing ──────────────────────────────────────────
out17="$(bash -c "source '$L'; :" 2>&1)"
check T17-source-safe "[$out17]" "[]"
printf 'PASS=%s FAIL=%s\n' "$pass" "$fail"; [ "$fail" -eq 0 ]
+1 -1
View File
@@ -81,7 +81,7 @@ tf "hotfixer report grammar" "$HOT" "HOTFIX-EXEC REPORT"
echo "── skills/hotfix/SKILL.md (hotfix wiring — revert, not loop) ──"
tf "hotfix silent contract" "$HSKL" "STEP 1.7 — CONTRACT (silent autofill)"
tf "hotfix zero questions" "$HSKL" "questions ever"
tf "hotfix pass B at locate" "$HSKL" "run pass B of"
tf "hotfix security gate" "$HSKL" "Security gate (fresh auditor)"
tf "hotfix block reverts" "$HSKL" "failure REVERTS, never loops"
tf "hotfix no verifier" "$HSKL" "No verifier is dispatched at hotfix weight"
+54 -17
View File
@@ -110,12 +110,8 @@
"Bash(chmod -R 777 *)",
"Bash(ssh *)",
"Bash(scp *)",
"Bash(rsync *)",
"Bash(nc *)",
"Bash(netcat *)",
"Bash(kill -9 *)",
"Bash(killall *)",
"Bash(pkill *)",
"Bash(crontab *)",
"Bash(systemctl *)",
"Bash(service *)",
@@ -182,6 +178,16 @@
"Bash(more .env.*)",
"Bash(grep * .env)",
"Bash(grep * .env.*)",
"Bash(sed * .env*)",
"Bash(awk * .env*)",
"Bash(cut * .env*)",
"Bash(tr * .env*)",
"Bash(sort * .env*)",
"Bash(uniq * .env*)",
"Bash(diff * .env*)",
"Bash(od * .env*)",
"Bash(xxd * .env*)",
"Bash(strings * .env*)",
"Bash(env)",
"Bash(printenv)",
"Bash(printenv *)",
@@ -218,15 +224,8 @@
"Bash(curl * | sh)",
"Bash(wget * | sh)",
"Bash(mkfifo *)",
"Bash(node -e *)",
"Bash(python3 -c *)",
"Bash(python -c *)",
"Bash(git push *)",
"Bash(git push)",
"Bash(docker run *)",
"Bash(docker exec *)",
"Bash(docker-compose up*)",
"Bash(docker compose up*)",
"Bash(brew install *)",
"Bash(apt install *)",
"Bash(apt-get install *)",
@@ -234,10 +233,6 @@
"Bash(pacman -S *)",
"WebSearch",
"WebFetch",
"Bash(xargs *)",
"Bash(sed *)",
"Bash(cp *)",
"Bash(mv *)",
"Bash(git stash pop*)",
"Bash(git stash drop*)",
"Bash(git stash clear)",
@@ -250,7 +245,7 @@
"disableBypassPermissionsMode": "disable",
"additionalDirectories": []
},
"model": "opus[1m]",
"model": "claude-fable-5-1[1m]",
"hooks": {
"SessionStart": [
{
@@ -358,5 +353,47 @@
"effortLevel": "xhigh",
"remoteControlAtStartup": true,
"inputNeededNotifEnabled": true,
"skipAutoPermissionPrompt": true
"skipAutoPermissionPrompt": true,
"autoMode": {
"allow": [
"$defaults",
"Local dev containers: `docker exec`, `docker run`, `docker compose up`/`exec`/`logs`/`ps` against a container running on this workstation whose name does not carry `prod` or `production` (a local Supabase or Postgres such as `supabase_db_*`) is routine development, not a remote shell into a shared host. Running a SQL file or script that lives in the repo inside it (`psql -f`, migrations, verify scripts) and piping the output through `tail` or `grep` passes. Remote Shell Writes, Production Reads and Sensitive Remote Exec apply only to hosts named as sensitive in Environment or carrying `prod`. A literal `DROP`, `TRUNCATE` or `DELETE` without a predicate typed on the command line stays under Mass Delete.",
"Project-local node: `node <file>`, `npm run`, `pnpm` or `yarn` scripts, and `npx` or `pnpm exec` of a package declared in the project's manifest or lockfile, with effects inside the current working directory, pass like `awk` or `echo`. `node -e` that computes or edits inside the working directory passes; the soft block on inline interpreters that delete or write outside it still applies."
],
"soft_deny": [
"$defaults",
"Scope of intent: an instruction clears a SOFT BLOCK for the current turn only. An approval given in an earlier turn is not an approval now, and the same action repeated in a later turn has to be asked for again.",
"Writing outside the current working directory: `cp`, `mv`, `sed -i`, `rsync`, `tee`, or a shell redirection whose destination resolves outside the cwd. Several sibling projects live under `~/Documents/`, so the realistic failure is writing into the wrong one, where git recovers nothing. Clear only when the user named the destination in this turn.",
"`rsync` invoked with `--delete`. It removes files at the destination that are absent from the source, with no undo. Clear only against a destination the user named in this turn.",
"Sending SIGKILL (`kill -9`) or killing processes by name (`killall`, `pkill`). These reach processes outside this session, including the user's editors, shells, dtach sessions and background jobs, and the target is chosen by a pattern, so a typo kills the wrong thing. Clear only when the user named the process in this turn.",
"Editing more than one file in place in a single command: `sed -i` or `perl -pi` over a glob, or a loop over `git ls-files`. The damage is not loss, since git recovers it, but a diff spanning hundreds of files that nobody reads before committing. `sed -i` on a single named file passes. Clear only when the user asked for the sweep.",
"Moving or renaming a directory inside the repo (`mv src/api src/api_old`, or any `mv` of a tree). It breaks imports and paths silently, and the breakage surfaces far from the command. Clear only when the user asked for that move.",
"An inline interpreter or `xargs` that deletes, or that writes outside the current working directory: `python3 -c`, `python -c` or `node -e` calling `rmtree`, `remove`, `unlink` or `truncate`; `xargs` feeding `rm`, `mv` or `dd`. `find ... | xargs rm` is the case that matters, since it routes around the `find * -exec rm` deny rule. Reading, computing, and editing a file inside the working directory pass untouched.",
"Docker data destruction on this workstation: `docker rm -f`, `docker volume rm` or `prune`, `docker system prune`, `docker compose down -v` (drops named volumes, which hold local database data with no undo), and `docker run` with `--privileged` or a bind mount outside the current working directory. Clear only when the user named the container or volume in this turn.",
"Undeclared node packages: `npx <pkg>`, `pnpm dlx` or `yarn dlx` of a package absent from the manifest and lockfile runs code fetched at call time; `npm install <name>` or `pnpm add <name>` adds a dependency the house rule requires naming first. Clear only when the user named the package in this turn."
],
"hard_deny": [
"$defaults",
"Secret exfiltration: reading a `.env`, an SSH key, a token or a credentials file, and then sending what it contained outside this machine. That covers an HTTP request body or header, a paste or gist service, a file upload, an email, and a commit or push to any repository. The static deny rules match the reading tools by pattern and cannot see the second half of the chain; read and send are separate steps and may sit several turns apart. No instruction clears this.",
"Production deployment: running a project's deploy script (`bin/deploy.sh` and its equivalents), any lftp, FTP, SFTP or rsync push to a hosting provider, and any action against a target whose name carries `prod` or `production` as a whole word or name segment. The user deploys by hand, out of session. A green test suite, a finished feature, or a plan step that reads \"deploy\" is not an instruction to deploy. No in-session instruction clears this.",
"Disarming the guardrails: removing or weakening an entry in the `deny`, `soft_deny` or `hard_deny` lists of any settings.json, passing `--no-verify` to git, deleting or chmod-ing `.githooks/pre-commit`, setting `disableAllHooks`, or switching `permissions.defaultMode` to `bypassPermissions`. Adding a restriction is fine; removing one is not. When a task is blocked by a guardrail, say so and stop, rather than widening the guardrail to get through. The user maintains these files by hand. No instruction clears this."
],
"environment": [
"$defaults",
"### Machine-specific (refines any \"None configured\" default above)",
"**Primary use of Claude Code**: software development on a personal Linux workstation. Single developer, no organization.",
"**Source control**: self-hosted Gitea at `git.bchanot.fr` (SSH on port 49220). Some checkouts under `/home/bchanot/Documents/` have no remote at all and are local-only.",
"**Repository visibility**: private. The Gitea instance is self-hosted and not publicly indexed, and a checkout with no remote is local-only. Treat every repo here as private unless its remote points at a public host.",
"**Trusted repo**: the checkout Claude Code is currently working in, under `/home/bchanot/Documents/`. No single repo is privileged over the others — read the current one, do not assume a previous session's project.",
"**Trusted internal domains**: `git.bchanot.fr` (self-hosted Gitea). It is the only internal service.",
"**Default / protected branches**: gitflow. `main` (prod) and `develop` (integration) are protected: a per-repo pre-commit hook refuses code commits on either (exempting `.claude/**` and merges) and Gitea enforces branch protection on both. Work lands on `feature/*`, `bugfix/*`, `chore/*`, `release/*`, `hotfix/*`.",
"**Secrets management**: `~/.claude/.env` is the single source of truth and lives outside every git tree; repos reach it through a gitignored symlink. Only `.env.example`, holding placeholders, is ever tracked. A real secret inside a repo is a defect, not a configuration.",
"**Internal sharing / snippet hosting**: none. Public paste, gist and pastebin services are outside the trust boundary.",
"**CI/CD deploy targets**: no CI system. Deploys run out of band from a per-project runbook, typically lftp/FTP to OVH mutualised hosting for web projects. Nothing deploys automatically on a push or a merge.",
"**Internal package registry**: none. Public npm and PyPI.",
"**Host containment**: an ordinary developer workstation with open internet and no sandbox. Nothing is contained by the environment itself.",
"**Sensitive remote targets**: any namespace, host, database or container whose name carries `prod` or `production` as a whole word or name segment.",
"**Sensitive data locations & audiences**: per-project `.env` files (gitignored) hold database, deploy and API credentials; some web projects store customer-submitted form data under a retention policy. Both are personal or client data — never send either to an external service."
]
}
}
+11 -5
View File
@@ -116,6 +116,10 @@ RISK: <low/medium — what could go wrong>
obvious fix.
- If the fix is significant (>10 lines, multiple files,
behavior change): wait for user approval.
- Then run pass B of `$HOME/.claude/lib/contract-interview.md` against the
FIX PLAN: every VISIBLE / PUBLIC NAME / SCOPE choice it settles that the
bug report left open → one batch of questions, before STEP 3b. The trivial
fast-path is not exempt: a 1-line fix with a visible choice still asks.
## STEP 3b — CHALLENGE THE FIX PLAN (before the contract)
Unless the fix is the trivial 1-2 line case STEP 3 already fast-paths, the
@@ -135,8 +139,8 @@ the STEP 3 approval gate.
Run `$HOME/.claude/lib/contract-interview.md` (main loop). The DIAGNOSIS
feeds it: REQUEST verbatim = the bug report as received; ACCEPTANCE CRITERIA
= the symptom reproduced-then-gone + a regression test present and passing;
FILE SCOPE = the FIX PLAN files. Questions stay proportional (a clear,
reproduced bug → zero). It writes the contract to
FILE SCOPE = the FIX PLAN files. Pass A only here (pass B ran at STEP 3); a
clear, reproduced bug asks nothing. It writes the contract to
`.claude/tasks/contracts/<date>-<slug>-<HHMM>.md`; keep the path — the
executor reads it first and GATE 1 (STEP 6) hands it to a fresh verifier.
@@ -162,9 +166,11 @@ ops, no security dispatch. Finish with the BUGFIX-EXEC REPORT."
Parse the `BUGFIX-EXEC REPORT`:
- `STATUS : DONE` → STEP 6.
- `STATUS : NEED-DECISION` → make the decision HERE (that is reflection),
append it to the plan, re-dispatch a FRESH bugfixer with plan + decision.
Max 2 decision round-trips → escalate to the user.
- `STATUS : NEED-DECISION` → route on its `CLASS:` per MID-RUN CLARIFICATION
in `$HOME/.claude/lib/contract-interview.md`: visible / public-name / scope
→ ask the user, verbatim; internal → decide HERE (max 2 such round-trips
→ escalate). Append the answer to the contract `[gated]` and to the plan,
re-dispatch a FRESH bugfixer with plan + decision.
- `STATUS : BLOCKED` → surface the blocker to the user, stop.
## STEP 6 — VERIFY + SECURE + PRE-COMMIT GATE + COMMIT (main loop, LRN-083)
+15 -10
View File
@@ -85,9 +85,9 @@ MEMORY; feed STEP 1 PLAN. Inline consumption — reader = planner, no injection.
## STEP 0.7 — CONTRACT
Run `$HOME/.claude/lib/contract-interview.md` (main loop — you are it). It
captures the request verbatim, asks 0-3 questions PROPORTIONAL to ambiguity
(a complete request → zero questions, silent), derives testable acceptance
criteria + file scope, and writes the contract to
captures the request verbatim, runs pass A (gaps: outcome, scope,
constraints — a complete request goes through silently), derives testable
acceptance criteria + file scope, and writes the contract to
`.claude/tasks/contracts/<date>-<slug>-<HHMM>.md`. Keep the path — the
executor reads it first and GATE 1 (STEP 4) hands it to a fresh verifier.
@@ -114,8 +114,11 @@ PLAN:
[ ] <test file> — <test to add>
```
If the approach is ambiguous: ask the user ONE focused question BEFORE
dispatching — never after (the executor cannot relay questions).
Then run pass B of `$HOME/.claude/lib/contract-interview.md` against this
plan: every VISIBLE / PUBLIC NAME / SCOPE choice the plan settles that the
request left open → one batch of questions BEFORE dispatching; answers land
in the contract's CLARIFICATIONS `[gated]` and in the plan. A choice that
surfaces only during execution comes back as `NEED-DECISION` (STEP 3).
## STEP 1b — CHALLENGE THE PLAN (before branching)
The STEP 1 plan is a reflection worth attacking before a branch is spent on it.
@@ -125,8 +128,8 @@ Persist it to `.claude/tasks/plans/<date>-<slug>-<HHMM>.md`, then run
Three blind challengers attack it; RE-THINK every aspect a BLOCKER lands (a named
plan change, or `[deferred]`), re-challenge once if the plan materially changed. The
STEP 3 executor receives the REVISED plan. Before dispatch, print a CHALLENGE SUMMARY
(BLOCKERs addressed / deferred / lenses returned), surfacing any deferred BLOCKER via
STEP 1's one-question gate.
(BLOCKERs addressed / deferred / lenses returned), surfacing any deferred BLOCKER in
the STEP 1 pass B batch.
## STEP 2 — BRANCH
@@ -150,9 +153,11 @@ Finish with the FEAT-EXEC REPORT."
Parse the `FEAT-EXEC REPORT`:
- `STATUS : DONE` → STEP 4.
- `STATUS : NEED-DECISION` → make the decision HERE (that is reflection),
append it to the plan, re-dispatch a FRESH feater with plan + decision.
Max 2 decision round-trips → escalate to the user.
- `STATUS : NEED-DECISION` → route on its `CLASS:` per MID-RUN CLARIFICATION
in `$HOME/.claude/lib/contract-interview.md`: visible / public-name / scope
→ ask the user, verbatim; internal → decide HERE (max 2 such round-trips
→ escalate). Append the answer to the contract `[gated]` and to the plan,
re-dispatch a FRESH feater with plan + decision.
- `STATUS : BLOCKED` → surface the blocker to the user, stop.
## STEP 4 — VERIFY + SECURE (fresh gates, bounded loops)
-1
View File
@@ -1 +0,0 @@
0.9.15
-678
View File
@@ -1,678 +0,0 @@
---
name: graphify
description: "Use for any question about a codebase, its architecture, file relationships, or project content — especially when graphify-out/ exists, where the question should be treated as a graphify query first. Turns any input (code, docs, papers, images, videos) into a persistent knowledge graph with god nodes, community detection, and query/path/explain tools."
---
# /graphify
Turn any folder of files into a navigable knowledge graph with community detection, an honest audit trail, and three outputs: interactive HTML, GraphRAG-ready JSON, and a plain-language GRAPH_REPORT.md.
## Usage
```
/graphify # full pipeline on current directory (HTML viz; add --obsidian for a vault)
/graphify <path> # full pipeline on specific path
/graphify https://github.com/<owner>/<repo> # clone repo then run full pipeline on it
/graphify https://github.com/<owner>/<repo> --branch <branch> # clone a specific branch
/graphify <url1> <url2> ... # clone multiple repos, build each, merge into one cross-repo graph
/graphify <path> --mode deep # thorough extraction, richer INFERRED edges
/graphify <path> --update # incremental - re-extract only new/changed files
/graphify <path> --directed # build directed graph (preserves edge direction: source→target)
/graphify <path> --whisper-model medium # use a larger Whisper model for better transcription accuracy
/graphify <path> --cluster-only # rerun clustering on existing graph
/graphify <path> --no-viz # skip visualization, just report + JSON
/graphify <path> --html # (HTML is generated by default - this flag is a no-op)
/graphify <path> --svg # also export graph.svg (embeds in Notion, GitHub)
/graphify <path> --graphml # export graph.graphml (Gephi, yEd)
/graphify <path> --neo4j # generate graphify-out/cypher.txt for Neo4j
/graphify <path> --neo4j-push bolt://localhost:7687 # push directly to Neo4j
/graphify <path> --falkordb # generate graphify-out/cypher.txt for FalkorDB
/graphify <path> --falkordb-push falkordb://localhost:6379 # push directly to FalkorDB
/graphify <path> --mcp # start MCP stdio server for agent access
/graphify <path> --watch # watch folder, auto-rebuild on code changes (no LLM needed)
/graphify <path> --wiki # build agent-crawlable wiki (index.md + one article per community)
/graphify <path> --obsidian --obsidian-dir ~/vaults/my-project # write vault to custom path (e.g. existing vault)
/graphify add <url> # fetch URL, save to ./raw, update graph
/graphify add <url> --author "Name" # tag who wrote it
/graphify add <url> --contributor "Name" # tag who added it to the corpus
/graphify query "<question>" # BFS traversal - broad context
/graphify query "<question>" --dfs # DFS - trace a specific path
/graphify query "<question>" --budget 1500 # cap answer at N tokens
/graphify path "AuthModule" "Database" # shortest path between two concepts
/graphify explain "SwinTransformer" # plain-language explanation of a node
```
## What graphify is for
Drop any folder of code, docs, papers, images, or video into graphify and get a queryable knowledge graph. Persistent across sessions, honest audit trail (EXTRACTED/INFERRED/AMBIGUOUS), community detection surfaces cross-document connections you wouldn't think to ask about.
## What You Must Do When Invoked
If the user invoked `/graphify --help` or `/graphify -h` (with no other arguments), print the contents of the `## Usage` section above verbatim and stop. Do not run any commands, do not detect files, do not default the path to `.`. Just print the Usage block and return.
**Fast path — existing graph:** Before doing anything else, check whether `graphify-out/graph.json` exists. The expected location is `graphify-out/graph.json` relative to the **current working directory** (i.e. the project root where you are running commands). If it exists AND the user's request is a natural-language question about the codebase (e.g. "How does X work?", "What calls Y?", "Trace the data flow through Z") and NOT an explicit rebuild command (`--update`, `--cluster-only`, or a bare path/URL that implies fresh extraction): **skip Steps 1–5 entirely and jump straight to `## For /graphify query`.** Run `graphify query "<question>"` immediately. Do not run detect. Do not check corpus size. Do not ask the user to narrow. The graph is already built — use it.
If no path was given, use `.` (current directory). Do not ask the user for a path.
If the path argument starts with `https://github.com/` or `http://github.com/`, treat it as a GitHub URL - run Step 0 before anything else, then continue with the resolved local path.
Follow these steps in order. Do not skip steps.
### Step 0 - GitHub repos and multi-path merge (only if a URL or several paths)
Only when the path is one or more `https://github.com/...` URLs, or several local subfolders to merge. See `references/github-and-merge.md` for the clone, cross-repo merge, and monorepo flow, then continue with the resolved local path. A plain local path skips this step.
### Step 1 - Ensure graphify is installed
```bash
# Detect the correct Python interpreter (handles uv tool, pipx, venv, system installs)
PYTHON=""
GRAPHIFY_BIN=$(which graphify 2>/dev/null)
# 1. uv tool installs — most reliable on modern Mac/Linux
if [ -z "$PYTHON" ] && command -v uv >/dev/null 2>&1; then
_UV_PY=$(uv tool run --from graphifyy python -c "import sys; print(sys.executable)" 2>/dev/null)
if [ -n "$_UV_PY" ]; then PYTHON="$_UV_PY"; fi
fi
# 2. Read shebang from graphify binary (pipx and direct pip installs)
if [ -z "$PYTHON" ] && [ -n "$GRAPHIFY_BIN" ]; then
_SHEBANG=$(head -1 "$GRAPHIFY_BIN" | tr -d '#!')
case "$_SHEBANG" in
*[!a-zA-Z0-9/_.@-]*) ;;
*) "$_SHEBANG" -c "import graphify" 2>/dev/null && PYTHON="$_SHEBANG" ;;
esac
fi
# 3. Fall back to python3
if [ -z "$PYTHON" ]; then PYTHON="python3"; fi
if ! "$PYTHON" -c "import graphify" 2>/dev/null; then
if command -v uv >/dev/null 2>&1; then
uv tool install --upgrade graphifyy -q 2>&1 | tail -3
_UV_PY=$(uv tool run --from graphifyy python -c "import sys; print(sys.executable)" 2>/dev/null)
if [ -n "$_UV_PY" ]; then PYTHON="$_UV_PY"; fi
else
"$PYTHON" -m pip install graphifyy -q 2>/dev/null \
|| "$PYTHON" -m pip install graphifyy -q --break-system-packages 2>&1 | tail -3
fi
fi
# Write interpreter path for all subsequent steps (persists across invocations)
mkdir -p graphify-out
"$PYTHON" -c "import sys; open('graphify-out/.graphify_python', 'w', encoding='utf-8').write(sys.executable)"
# Save scan root so `graphify update` (no args) knows where to look next time
echo "$(cd INPUT_PATH && pwd)" > graphify-out/.graphify_root
```
If the import succeeds, print nothing and move straight to Step 2.
**In every subsequent bash block, replace `python3` with `$(cat graphify-out/.graphify_python)` to use the correct interpreter.**
### Step 2 - Detect files
```bash
$(cat graphify-out/.graphify_python) -c "
import json
from graphify.detect import detect
from pathlib import Path
result = detect(Path('INPUT_PATH'))
print(json.dumps(result, ensure_ascii=False))
" > graphify-out/.graphify_detect.json
```
Replace INPUT_PATH with the actual path the user provided. Do NOT cat or print the JSON - read it silently and present a clean summary instead:
```
Corpus: X files · ~Y words
code: N files (.py .ts .go ...)
docs: N files (.md .txt ...)
papers: N files (.pdf ...)
images: N files
video: N files (.mp4 .mp3 ...)
```
Omit any category with 0 files from the summary.
Then act on it:
- If `total_files` is 0: stop with "No supported files found in [path]."
- If `skipped_sensitive` is non-empty: mention file count skipped, not the file names.
- If `total_words` > 2,000,000 OR `total_files` > 500: show the warning. Then compute the top 5 first-level subdirectories by file count:
- Read `scan_root` from the detect JSON (always an absolute path to the resolved INPUT_PATH).
- Concatenate all file lists across all types (`code`, `document`, `paper`, `image`, `video`).
- Filter out any path that starts with `scan_root + "/graphify-out/"` to exclude converted sidecars.
- For each file, strip the `scan_root` prefix and take the first path component. Files directly in `scan_root` with no subdirectory count as `(root)`.
- If all files are in `(root)` with no subdirectories, do not ask to narrow — no subfolders exist. Instead suggest `--no-cluster` to skip the expensive clustering step and proceed.
- Otherwise rank by count, show the top 5 with file counts, then ask which subfolder to run on. Wait for the user's answer before proceeding.
- Otherwise: proceed directly to Step 2.5 if video files were detected, or Step 3 if not.
### Step 2.5 - Video and audio (only if video files detected)
Skip this step entirely if `detect` returned zero `video` files. When the corpus has video or audio, see `references/transcribe.md` to transcribe them to text first, then treat the transcripts as doc files in Step 3.
### Step 3 - Extract entities and relationships
**Before starting:** note whether `--mode deep` was given. You must pass `DEEP_MODE=true` to every subagent in Step B2 if it was. Track this from the original invocation - do not lose it.
This step has two parts: **structural extraction** (deterministic, free) and **semantic extraction** (LLM, costs tokens).
> **graphify needs no API key. Never ask the user for one, and never block on one.** Code is extracted structurally (AST) with no LLM and no key at all — a code-only corpus (the common `/graphify .` on a repo) skips semantic extraction entirely, so it needs nothing here: go straight to Part A and skip Part B. Semantic extraction (only for docs, papers, and images) uses Gemini **only if** `GEMINI_API_KEY`/`GOOGLE_API_KEY` is already set; otherwise the host agent itself is the LLM. graphify does **not** read `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, or any other provider key. If you catch yourself about to prompt for, wait on, or stop because of a missing API key, that is a misread of this skill — proceed without one.
**Before semantic extraction:** check whether `GEMINI_API_KEY` or `GOOGLE_API_KEY` is set. If neither is set, print this one-liner to the user:
> Tip: set `GEMINI_API_KEY` or `GOOGLE_API_KEY` to use Gemini for semantic extraction (`pip install 'graphifyy[gemini]'`).
Print it once, then continue — do not wait for the user to supply a key. If `GEMINI_API_KEY` or `GOOGLE_API_KEY` IS set, use `graphify.llm.extract_corpus_parallel(files, backend="gemini")` for semantic extraction instead of dispatching subagents. The default Gemini model is `gemini-3-flash-preview`; set `GRAPHIFY_GEMINI_MODEL` or pass `--model` in headless CLI flows to override it.
> **No other API keys are read.** When `GEMINI_API_KEY`/`GOOGLE_API_KEY` are unset, semantic extraction falls to the host agent itself — the running session is the LLM. On a host that dispatches subagents (e.g. Claude Code), dispatch them as written in Part B. On a host that runs the CLI directly in a terminal and cannot dispatch subagents, do not stall: a code-only corpus has no semantic work, so write the empty semantic file (Part B "Fast path") and continue to Part C; for a corpus with docs/papers/images, either set a Gemini key or extract those inline yourself, but in no case prompt for `ANTHROPIC_API_KEY` — that prompt is a misread of this skill.
**Run Part A (AST) and Part B (semantic) in parallel. Dispatch all semantic subagents AND start AST extraction in the same message. Both can run simultaneously since they operate on different file types. Merge results in Part C as before.**
Note: Parallelizing AST + semantic saves 5-15s on large corpora. AST is deterministic and fast; start it while subagents are processing docs/papers.
#### Part A - Structural extraction for code files
For any code files detected, run AST extraction in parallel with Part B subagents:
```bash
$(cat graphify-out/.graphify_python) -c "
import sys, json
from graphify.extract import collect_files, extract
from pathlib import Path
import json
code_files = []
detect = json.loads(Path('graphify-out/.graphify_detect.json').read_text(encoding=\"utf-8\"))
for f in detect.get('files', {}).get('code', []):
code_files.extend(collect_files(Path(f)) if Path(f).is_dir() else [Path(f)])
if code_files:
result = extract(code_files, cache_root=Path('INPUT_PATH'))
Path('graphify-out/.graphify_ast.json').write_text(json.dumps(result, indent=2, ensure_ascii=False), encoding=\"utf-8\")
print(f'AST: {len(result[\"nodes\"])} nodes, {len(result[\"edges\"])} edges')
else:
Path('graphify-out/.graphify_ast.json').write_text(json.dumps({'nodes':[],'edges':[],'input_tokens':0,'output_tokens':0}, ensure_ascii=False), encoding=\"utf-8\")
print('No code files - skipping AST extraction')
"
```
#### Part B - Semantic extraction (parallel subagents)
**Fast path:** If detection found zero docs, papers, and images (code-only corpus), skip Part B entirely and go straight to Part C. AST handles code - there is nothing for semantic subagents to do. **First write an empty semantic file** so Part C's merge has its input (it reads `.graphify_semantic.json` unconditionally; without this a code-only run hits `FileNotFoundError`):
```bash
$(cat graphify-out/.graphify_python) -c "
import json
from pathlib import Path
Path('graphify-out/.graphify_semantic.json').write_text(json.dumps({'nodes':[],'edges':[],'hyperedges':[],'input_tokens':0,'output_tokens':0}), encoding='utf-8')
"
```
**MANDATORY: You MUST use the Agent tool here. Reading files yourself one-by-one is forbidden - it is 5-10x slower. If you do not use the Agent tool you are doing this wrong.**
Before dispatching subagents, print a timing estimate:
- Load `total_words` and file counts from `graphify-out/.graphify_detect.json`
- Estimate agents needed: `ceil(uncached_non_code_files / 22)` (chunk size is 20-25)
- Estimate time: ~45s per agent batch (they run in parallel, so total ≈ 45s × ceil(agents/parallel_limit))
- Print: "Semantic extraction: ~N files → X agents, estimated ~Ys"
**Step B0 - Check extraction cache first**
Before dispatching any subagents, check which files already have cached extraction results:
```bash
$(cat graphify-out/.graphify_python) -c "
import json
from graphify.cache import check_semantic_cache
from pathlib import Path
detect = json.loads(Path('graphify-out/.graphify_detect.json').read_text(encoding=\"utf-8\"))
# Only content files go to semantic extraction. Code is already covered structurally
# by the AST pass (Part A); flattening every category here makes subagents re-read
# every source file (#1392). Video is transcribed to a document in Step 2.5 first.
all_files = [f for cat in ('document', 'paper', 'image') for f in detect['files'].get(cat, [])]
cached_nodes, cached_edges, cached_hyperedges, uncached = check_semantic_cache(all_files, root='INPUT_PATH')
# Always (re)write the cache file: write hits, else DELETE any leftover from a prior
# run so Part C never merges a stale .graphify_cached.json (#1392).
if cached_nodes or cached_edges or cached_hyperedges:
Path('graphify-out/.graphify_cached.json').write_text(json.dumps({'nodes': cached_nodes, 'edges': cached_edges, 'hyperedges': cached_hyperedges}, ensure_ascii=False), encoding=\"utf-8\")
else:
Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True)
Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\")
print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction')
"
```
Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly.
**Step B1 - Split into chunks**
Load files from `graphify-out/.graphify_uncached.txt`. Split into chunks of 20-25 files each. Each image gets its own chunk (vision needs separate context). When splitting, group files from the same directory together so related artifacts land in the same chunk and cross-file relationships are more likely to be extracted.
**Step B2 - Dispatch ALL subagents in a single message**
Call the Agent tool multiple times IN THE SAME RESPONSE - one call per chunk. This is the only way they run in parallel. If you make one Agent call, wait, then make another, you are doing it sequentially and defeating the purpose.
**IMPORTANT - subagent type:** Always use `subagent_type="general-purpose"`. Do NOT use `Explore` - it is read-only and cannot write chunk files to disk, which silently drops extraction results. General-purpose has Write and Bash access which the subagent needs.
Concrete example for 3 chunks:
```
[Agent tool call 1: files 1-15, subagent_type="general-purpose"]
[Agent tool call 2: files 16-30, subagent_type="general-purpose"]
[Agent tool call 3: files 31-45, subagent_type="general-purpose"]
```
All three in one message. Not three separate messages.
Each subagent receives this exact prompt (substitute FILE_LIST, CHUNK_NUM, TOTAL_CHUNKS, DEEP_MODE, and CHUNK_PATH).
CHUNK_PATH must be an **absolute** path — derive it before dispatching:
```bash
PROJECT_ROOT=$(pwd) # cwd — where Part C globs graphify-out/ (NOT .graphify_root/scan dir, #1392)
# Then for chunk N: CHUNK_PATH="${PROJECT_ROOT}/graphify-out/.graphify_chunk_0N.json"
```
Subagent prompt template:
See `references/extraction-spec.md` for the exact subagent prompt (JSON schema, node-ID rules, confidence rubric, frontmatter, hyperedge, and vision rules). Load it only here, only when at least one chunk holds a doc, paper, or image; a pure-code corpus has skipped Part B and never reads it. Pass each subagent that prompt verbatim with FILE_LIST, CHUNK_NUM, TOTAL_CHUNKS, DEEP_MODE, and CHUNK_PATH substituted, and have it write the result to CHUNK_PATH.
**Step B3 - Collect, cache, and merge**
Wait for all subagents. For each result:
- Check that `graphify-out/.graphify_chunk_NN.json` exists on disk — this is the success signal
- If the file exists and contains valid JSON with `nodes` and `edges`, include it and save to cache
- If the file is missing, the subagent was likely dispatched as read-only (Explore type) — print a warning: "chunk N missing from disk — subagent may have been read-only. Re-run with general-purpose agent." Do not silently skip.
- If a subagent failed or returned invalid JSON, print a warning and skip that chunk - do not abort
If more than half the chunks failed or are missing, stop and tell the user to re-run and ensure `subagent_type="general-purpose"` is used.
Merge all chunk files into `.graphify_semantic_new.json`. **After each Agent call completes, read the real token counts from the Agent tool result's `usage` field and write them back into the chunk JSON before merging** — the chunk JSON itself always has placeholder zeros. Then run:
```bash
$(cat graphify-out/.graphify_python) -c "
import json, glob
from pathlib import Path
chunks = sorted(glob.glob('graphify-out/.graphify_chunk_*.json'))
all_nodes, all_edges, all_hyperedges = [], [], []
total_in, total_out = 0, 0
for c in chunks:
d = json.loads(Path(c).read_text(encoding=\"utf-8\"))
all_nodes += d.get('nodes', [])
all_edges += d.get('edges', [])
all_hyperedges += d.get('hyperedges', [])
total_in += d.get('input_tokens', 0)
total_out += d.get('output_tokens', 0)
Path('graphify-out/.graphify_semantic_new.json').write_text(json.dumps({
'nodes': all_nodes, 'edges': all_edges, 'hyperedges': all_hyperedges,
'input_tokens': total_in, 'output_tokens': total_out,
}, indent=2, ensure_ascii=False), encoding=\"utf-8\")
print(f'Merged {len(chunks)} chunks: {total_in:,} in / {total_out:,} out tokens')
"
```
Save new results to cache:
```bash
$(cat graphify-out/.graphify_python) -c "
import json
from graphify.cache import save_semantic_cache
from pathlib import Path
new = json.loads(Path('graphify-out/.graphify_semantic_new.json').read_text(encoding=\"utf-8\")) if Path('graphify-out/.graphify_semantic_new.json').exists() else {'nodes':[],'edges':[],'hyperedges':[]}
uncached = [line for line in Path('graphify-out/.graphify_uncached.txt').read_text(encoding=\"utf-8\").splitlines() if line]
saved = save_semantic_cache(new.get('nodes', []), new.get('edges', []), new.get('hyperedges', []), root='INPUT_PATH', allowed_source_files=uncached)
print(f'Cached {saved} files')
"
```
Merge cached + new results into `graphify-out/.graphify_semantic.json`:
```bash
$(cat graphify-out/.graphify_python) -c "
import json
from pathlib import Path
cached = json.loads(Path('graphify-out/.graphify_cached.json').read_text(encoding=\"utf-8\")) if Path('graphify-out/.graphify_cached.json').exists() else {'nodes':[],'edges':[],'hyperedges':[]}
new = json.loads(Path('graphify-out/.graphify_semantic_new.json').read_text(encoding=\"utf-8\")) if Path('graphify-out/.graphify_semantic_new.json').exists() else {'nodes':[],'edges':[],'hyperedges':[]}
all_nodes = cached['nodes'] + new.get('nodes', [])
all_edges = cached['edges'] + new.get('edges', [])
all_hyperedges = cached.get('hyperedges', []) + new.get('hyperedges', [])
seen = set()
deduped = []
for n in all_nodes:
if n['id'] not in seen:
seen.add(n['id'])
deduped.append(n)
merged = {
'nodes': deduped,
'edges': all_edges,
'hyperedges': all_hyperedges,
'input_tokens': new.get('input_tokens', 0),
'output_tokens': new.get('output_tokens', 0),
}
Path('graphify-out/.graphify_semantic.json').write_text(json.dumps(merged, indent=2, ensure_ascii=False), encoding=\"utf-8\")
print(f'Extraction complete - {len(deduped)} nodes, {len(all_edges)} edges ({len(cached[\"nodes\"])} from cache, {len(new.get(\"nodes\",[]))} new)')
"
```
Clean up temp files: `rm -f graphify-out/.graphify_cached.json graphify-out/.graphify_uncached.txt graphify-out/.graphify_semantic_new.json`
#### Part C - Merge AST + semantic into final extraction
```bash
$(cat graphify-out/.graphify_python) -c "
import sys, json
from pathlib import Path
ast = json.loads(Path('graphify-out/.graphify_ast.json').read_text(encoding=\"utf-8\"))
sem = json.loads(Path('graphify-out/.graphify_semantic.json').read_text(encoding=\"utf-8\"))
# Merge: AST nodes first, semantic nodes deduplicated by id
seen = {n['id'] for n in ast['nodes']}
merged_nodes = list(ast['nodes'])
for n in sem['nodes']:
if n['id'] not in seen:
merged_nodes.append(n)
seen.add(n['id'])
merged_edges = ast['edges'] + sem['edges']
merged_hyperedges = sem.get('hyperedges', [])
merged = {
'nodes': merged_nodes,
'edges': merged_edges,
'hyperedges': merged_hyperedges,
'input_tokens': sem.get('input_tokens', 0),
'output_tokens': sem.get('output_tokens', 0),
}
Path('graphify-out/.graphify_extract.json').write_text(json.dumps(merged, indent=2, ensure_ascii=False), encoding=\"utf-8\")
total = len(merged_nodes)
edges = len(merged_edges)
print(f'Merged: {total} nodes, {edges} edges ({len(ast[\"nodes\"])} AST + {len(sem[\"nodes\"])} semantic)')
"
```
### Step 4 - Build graph, cluster, analyze, generate outputs
**Before starting:** the code blocks below pass `directed=IS_DIRECTED` to `build_from_json()`. Replace `IS_DIRECTED` with `True` if `--directed` was given (builds a `DiGraph` preserving edge direction source→target), otherwise `False` (the default undirected `Graph`). Substitute it the same way you substitute `INPUT_PATH` — do not leave the literal `IS_DIRECTED` in the code.
```bash
mkdir -p graphify-out
$(cat graphify-out/.graphify_python) -c "
import sys, json
from graphify.build import build_from_json
from graphify.cluster import cluster, score_all
from graphify.analyze import god_nodes, surprising_connections, suggest_questions
from graphify.report import generate
from graphify.export import to_json
from pathlib import Path
extraction = json.loads(Path('graphify-out/.graphify_extract.json').read_text(encoding=\"utf-8\"))
detection = json.loads(Path('graphify-out/.graphify_detect.json').read_text(encoding=\"utf-8\"))
# root= mirrors the --update runbook (#1361): relativize source_file to the same
# base so the full build and incremental --update never drift apart on re-extract.
G = build_from_json(extraction, root='INPUT_PATH', directed=IS_DIRECTED)
# Guard BEFORE any write: an empty extraction must not clobber a good graph.json /
# GRAPH_REPORT.md / analysis sidecar. Check immediately after build (#1392).
if G.number_of_nodes() == 0:
print('ERROR: Graph is empty - extraction produced no nodes.')
print('Possible causes: all files were skipped, binary-only corpus, or extraction failed.')
raise SystemExit(1)
communities = cluster(G)
cohesion = score_all(G, communities)
tokens = {'input': extraction.get('input_tokens', 0), 'output': extraction.get('output_tokens', 0)}
gods = god_nodes(G)
surprises = surprising_connections(G, communities)
labels = {cid: 'Community ' + str(cid) for cid in communities}
# Placeholder questions - regenerated with real labels in Step 5
questions = suggest_questions(G, communities, labels)
# Export FIRST and honor the #479 shrink-guard: to_json returns False (writing
# nothing) when the new graph is smaller than the existing graph.json. Only write
# GRAPH_REPORT.md + the analysis sidecar when the graph was actually written, so
# they never describe a graph that graph.json doesn't contain (#1392).
wrote = to_json(G, communities, 'graphify-out/graph.json')
if not wrote:
print('ERROR: refused to shrink graphify-out/graph.json (existing graph has more nodes; #479).')
print('If this shrink is intentional (you deleted files), re-run a full build with --force.')
raise SystemExit(1)
report = generate(G, communities, cohesion, labels, gods, surprises, detection, tokens, 'INPUT_PATH', suggested_questions=questions)
Path('graphify-out/GRAPH_REPORT.md').write_text(report, encoding=\"utf-8\")
analysis = {
'communities': {str(k): v for k, v in communities.items()},
'cohesion': {str(k): v for k, v in cohesion.items()},
'gods': gods,
'surprises': surprises,
'questions': questions,
}
Path('graphify-out/.graphify_analysis.json').write_text(json.dumps(analysis, indent=2, ensure_ascii=False), encoding=\"utf-8\")
print(f'Graph: {G.number_of_nodes()} nodes, {G.number_of_edges()} edges, {len(communities)} communities')
"
```
If this step prints `ERROR: Graph is empty`, stop and tell the user what happened - do not proceed to labeling or visualization.
Replace INPUT_PATH with the actual path.
### Step 4.5 - Graph health check (read-only integrity gate)
A non-destructive diagnostic on the extraction, before labeling. It surfaces edge collapse, dangling/missing endpoints, and self-loops — the silent-corruption modes of incremental updates and AST/LLM id mismatches. Read-only; never aborts.
```bash
$(cat graphify-out/.graphify_python) -c "
import json
from pathlib import Path
from graphify.diagnostics import diagnose_extraction, format_diagnostic_report
extraction = json.loads(Path('graphify-out/.graphify_extract.json').read_text(encoding=\"utf-8\"))
summary = diagnose_extraction(extraction, directed=IS_DIRECTED, root='INPUT_PATH')
print(format_diagnostic_report(summary))
flags = [f'{summary[k]} {label}' for k, label in (
('dangling_endpoint_edges', 'dangling-endpoint edges'),
('missing_endpoint_edges', 'missing-endpoint edges'),
('self_loop_edges', 'self-loop edges'),
('directed_same_endpoint_collapsed_edges', 'collapsed (directed) edges'),
('undirected_same_endpoint_collapsed_edges', 'collapsed (undirected) edges'),
) if summary.get(k, 0)]
print('GRAPH HEALTH WARNING: ' + '; '.join(flags) + ' - graph may be incomplete/corrupt.' if flags else 'Graph health: OK (no dangling/missing/collapsed edges).')
"
```
Substitute `IS_DIRECTED` and `INPUT_PATH` as in Step 4. If a `GRAPH HEALTH WARNING` prints, surface it in the final summary (do not abort — the graph is still usable, but the integrity issue must be visible, per the Honesty Rules).
### Step 5 - Label communities
Read `graphify-out/.graphify_analysis.json`. For each community key, look at its node labels and write a 2-5 word plain-language name (e.g. "Attention Mechanism", "Training Pipeline", "Data Loading").
Then regenerate the report and save the labels for the visualizer:
```bash
$(cat graphify-out/.graphify_python) -c "
import sys, json
from graphify.build import build_from_json
from graphify.cluster import score_all
from graphify.analyze import god_nodes, surprising_connections, suggest_questions
from graphify.report import generate
from pathlib import Path
extraction = json.loads(Path('graphify-out/.graphify_extract.json').read_text(encoding=\"utf-8\"))
detection = json.loads(Path('graphify-out/.graphify_detect.json').read_text(encoding=\"utf-8\"))
analysis = json.loads(Path('graphify-out/.graphify_analysis.json').read_text(encoding=\"utf-8\"))
# root= as in Step 4 / the --update runbook (#1361) — same base for node-key parity.
G = build_from_json(extraction, root='INPUT_PATH', directed=IS_DIRECTED)
communities = {int(k): v for k, v in analysis['communities'].items()}
cohesion = {int(k): v for k, v in analysis['cohesion'].items()}
tokens = {'input': extraction.get('input_tokens', 0), 'output': extraction.get('output_tokens', 0)}
# LABELS - replace these with the names you chose above
labels = LABELS_DICT
# Regenerate questions with real community labels (labels affect question phrasing)
questions = suggest_questions(G, communities, labels)
report = generate(G, communities, cohesion, labels, analysis['gods'], analysis['surprises'], detection, tokens, 'INPUT_PATH', suggested_questions=questions)
Path('graphify-out/GRAPH_REPORT.md').write_text(report, encoding=\"utf-8\")
Path('graphify-out/.graphify_labels.json').write_text(json.dumps({str(k): v for k, v in labels.items()}, ensure_ascii=False), encoding=\"utf-8\")
print('Report updated with community labels')
"
```
Replace `LABELS_DICT` with the actual dict you constructed (e.g. `{0: "Attention Mechanism", 1: "Training Pipeline"}`).
Replace INPUT_PATH with the actual path.
### Step 6 - Generate Obsidian vault (opt-in) + HTML
**Generate HTML always** (unless `--no-viz`). **Obsidian vault only if `--obsidian` was explicitly given** — skip it otherwise, it generates one file per node.
If `--obsidian` was given:
- If `--obsidian-dir <path>` was also given, pass it via `--dir`. Otherwise defaults to `graphify-out/obsidian`.
```bash
graphify export obsidian
# or with custom dir: graphify export obsidian --dir ~/vaults/my-project
```
Generate the HTML graph (always, unless `--no-viz`):
```bash
graphify export html # auto-aggregates to community view if graph > 5000 nodes
# or: graphify export html --no-viz
```
### Steps 6b-8 - Wiki, Neo4j, FalkorDB, SVG, GraphML, MCP, benchmark (only on their flags)
These run only when their flag is present (`--wiki`, `--neo4j`/`--neo4j-push`, `--falkordb`/`--falkordb-push`, `--svg`, `--graphml`, `--mcp`) or, for the token-reduction benchmark, when `total_words` exceeds 5,000. A default run with no export flags skips all of them. See `references/exports.md` for each one. Run any `--wiki` export before Step 9 cleanup so `.graphify_labels.json` is still available.
---
### Step 9 - Save manifest, update cost tracker, clean up, and report
```bash
$(cat graphify-out/.graphify_python) -c "
import json
from pathlib import Path
from datetime import datetime, timezone
from graphify.detect import save_manifest
# Save manifest for --update
detect = json.loads(Path('graphify-out/.graphify_detect.json').read_text(encoding=\"utf-8\"))
# In --update mode, 'all_files' carries the full corpus; 'files' is the changed
# subset. Full-rebuild mode populates only 'files', so the fallback handles that.
# root= relativizes the manifest keys to the scan root (same base as the build),
# so the on-disk manifest is portable across clones/machines and a later --update
# matches cached files instead of missing every one (#1417).
save_manifest(detect.get('all_files') or detect['files'], root='INPUT_PATH')
# Update cumulative cost tracker
extract = json.loads(Path('graphify-out/.graphify_extract.json').read_text(encoding=\"utf-8\"))
input_tok = extract.get('input_tokens', 0)
output_tok = extract.get('output_tokens', 0)
cost_path = Path('graphify-out/cost.json')
if cost_path.exists():
cost = json.loads(cost_path.read_text(encoding=\"utf-8\"))
else:
cost = {'runs': [], 'total_input_tokens': 0, 'total_output_tokens': 0}
cost['runs'].append({
'date': datetime.now(timezone.utc).isoformat(),
'input_tokens': input_tok,
'output_tokens': output_tok,
'files': detect.get('total_files', 0),
})
cost['total_input_tokens'] += input_tok
cost['total_output_tokens'] += output_tok
cost_path.write_text(json.dumps(cost, indent=2, ensure_ascii=False), encoding=\"utf-8\")
print(f'This run: {input_tok:,} input tokens, {output_tok:,} output tokens')
print(f'All time: {cost[\"total_input_tokens\"]:,} input, {cost[\"total_output_tokens\"]:,} output ({len(cost[\"runs\"])} runs)')
"
rm -f graphify-out/.graphify_detect.json graphify-out/.graphify_extract.json graphify-out/.graphify_ast.json graphify-out/.graphify_semantic.json graphify-out/.graphify_analysis.json
find graphify-out -maxdepth 1 -name '.graphify_chunk_*.json' -delete 2>/dev/null
rm -f graphify-out/.needs_update 2>/dev/null || true
```
Replace INPUT_PATH with the actual path (same value used in Steps 4-5) so the manifest is relativized to the scan root.
Tell the user (omit the obsidian line unless --obsidian was given):
```
Graph complete. Outputs in PATH_TO_DIR/graphify-out/
graph.html - interactive graph, open in browser
GRAPH_REPORT.md - audit report
graph.json - raw graph data
obsidian/ - Obsidian vault (only if --obsidian was given)
```
If graphify saved you time, consider supporting it: https://github.com/sponsors/safishamsi
Replace PATH_TO_DIR with the actual absolute path of the directory that was processed.
Then paste these sections from GRAPH_REPORT.md directly into the chat:
- God Nodes
- Surprising Connections
- Suggested Questions
Do NOT paste the full report - just those three sections. Keep it concise.
Then immediately offer to explore. Pick the single most interesting suggested question from the report - the one that crosses the most community boundaries or has the most surprising bridge node - and ask:
> "The most interesting question this graph can answer: **[question]**. Want me to trace it?"
If the user says yes, run `/graphify query "[question]"` on the graph and walk them through the answer using the graph structure - which nodes connect, which community boundaries get crossed, what the path reveals. Keep going as long as they want to explore. Each answer should end with a natural follow-up ("this connects to X - want to go deeper?") so the session feels like navigation, not a one-shot report.
The graph is the map. Your job after the pipeline is to be the guide.
---
## Interpreter guard for subcommands
Before running any subcommand below (`--update`, `--cluster-only`, `query`, `path`, `explain`, `add`), check that `.graphify_python` exists. If it's missing (e.g. user deleted `graphify-out/`), re-resolve the interpreter first:
```bash
if [ ! -f graphify-out/.graphify_python ]; then
GRAPHIFY_BIN=$(which graphify 2>/dev/null)
if [ -n "$GRAPHIFY_BIN" ]; then
PYTHON=$(head -1 "$GRAPHIFY_BIN" | tr -d '#!')
case "$PYTHON" in *[!a-zA-Z0-9/_.@-]*) PYTHON="python3" ;; esac
else
PYTHON="python3"
fi
mkdir -p graphify-out
"$PYTHON" -c "import sys; open('graphify-out/.graphify_python', 'w', encoding='utf-8').write(sys.executable)"
fi
```
## For --update and --cluster-only
Both are non-default subcommands. `--update` re-extracts only new or changed files; `--cluster-only` reruns clustering on the existing graph. See `references/update.md` for both flows.
---
## For /graphify query
When `graphify-out/graph.json` already exists and the user asks a question about the corpus, answer from the graph rather than rebuilding it:
```bash
graphify query "<question>"
```
Before traversal, expand the question against the graph's own vocabulary so a wording mismatch does not collapse the answer to noise. If the `graphify query` CLI is unavailable, fall back to an inline NetworkX traversal of `graphify-out/graph.json`. Answer using only what the graph output contains, and quote `source_location` when citing a specific fact. For that vocab-expansion step, the BFS/DFS traversal modes, the `--budget` cap, the NetworkX fallback, `save-result` feedback, and the `/graphify path` and `/graphify explain` flows, see `references/query.md`.
---
## For /graphify add and --watch
Neither is part of the default build. When the user runs `/graphify add <url>` to fetch a URL into the corpus, or passes `--watch` to auto-rebuild on file changes, see `references/add-watch.md`.
---
## For the commit hook and native CLAUDE.md integration
When the user asks to install the post-commit auto-rebuild hook or wire graphify into a project's CLAUDE.md, see `references/hooks.md`.
---
## Honesty Rules
- Never invent an edge. If unsure, use AMBIGUOUS.
- Never skip the corpus check warning.
- Always show token cost in the report.
- Never hide cohesion scores behind symbols - show the raw number.
- Never run HTML viz on a graph with more than 5,000 nodes without warning the user.
-56
View File
@@ -1,56 +0,0 @@
# graphify reference: add a URL and watch a folder
Load this when the user ran `/graphify add <url>` or passed `--watch`. Neither is part of the default build.
## For /graphify add
Fetch a URL and add it to the corpus, then update the graph.
```bash
$(cat graphify-out/.graphify_python) -c "
import sys
from graphify.ingest import ingest
from pathlib import Path
try:
out = ingest('URL', Path('./raw'), author='AUTHOR', contributor='CONTRIBUTOR')
print(f'Saved to {out}')
except ValueError as e:
print(f'error: {e}', file=sys.stderr)
sys.exit(1)
except RuntimeError as e:
print(f'error: {e}', file=sys.stderr)
sys.exit(1)
"
```
Replace `URL` with the actual URL, `AUTHOR` with the user's name if provided, `CONTRIBUTOR` likewise. If the command exits with an error, tell the user what went wrong - do not silently continue. After a successful save, automatically run the `--update` pipeline on `./raw` to merge the new file into the existing graph.
Supported URL types (auto-detected):
- YouTube / any video URL → audio downloaded via yt-dlp, transcribed to `.txt` on next run (requires `pip install 'graphifyy[video]'`)
- Twitter/X → fetched via oEmbed, saved as `.md` with tweet text and author
- arXiv → abstract + metadata saved as `.md`
- PDF → downloaded as `.pdf`
- Images (.png/.jpg/.webp) → downloaded, Claude vision extracts on next run
- Any webpage → converted to markdown via html2text
---
## For --watch
Start a background watcher that monitors a folder and auto-updates the graph when files change.
```bash
$(cat graphify-out/.graphify_python) -m graphify.watch INPUT_PATH --debounce 3
```
Replace INPUT_PATH with the folder to watch. Behavior depends on what changed:
- **Code files only (.py, .ts, .go, etc.):** re-runs AST extraction + rebuild + cluster immediately, no LLM needed. `graph.json` and `GRAPH_REPORT.md` are updated automatically.
- **Docs, papers, or images:** writes a `graphify-out/needs_update` flag and prints a notification to run `/graphify --update` (LLM semantic re-extraction required).
Debounce (default 3s): waits until file activity stops before triggering, so a wave of parallel agent writes doesn't trigger a rebuild per file.
Press Ctrl+C to stop.
For agentic workflows: run `--watch` in a background terminal. Code changes from agent waves are picked up automatically between waves. If agents are also writing docs or notes, you'll need a manual `/graphify --update` after those waves.
-87
View File
@@ -1,87 +0,0 @@
# graphify reference: extra exports and benchmark
Load this when the user passed one of the export flags (`--wiki`, `--neo4j`, `--neo4j-push`, `--falkordb`, `--falkordb-push`, `--svg`, `--graphml`, `--mcp`), or when the corpus is large enough for the token-reduction benchmark. Each step runs only for its own flag.
### Step 6b - Wiki (only if --wiki flag)
**Only run this step if `--wiki` was explicitly given in the original command.**
Run this before Step 9 (cleanup) so `.graphify_labels.json` is still available.
```bash
graphify export wiki
```
### Step 7 - Neo4j export (only if --neo4j or --neo4j-push flag)
**If `--neo4j`** - generate a Cypher file for manual import:
```bash
graphify export neo4j
```
**If `--neo4j-push <uri>`** - push directly to a running Neo4j instance. Ask the user for credentials if not provided:
```bash
graphify export neo4j --push bolt://localhost:7687 --user neo4j --password PASSWORD
```
Default URI is `bolt://localhost:7687`, default user is `neo4j`. Uses MERGE - safe to re-run without creating duplicates.
### Step 7a - FalkorDB export (only if --falkordb or --falkordb-push flag)
**If `--falkordb`** - generate a Cypher file. The statements are OpenCypher, but FalkorDB's `GRAPH.QUERY` runs one statement at a time (no bulk script import like Neo4j's `cypher-shell`), so prefer `--falkordb-push` to load a graph. Use this only when you want the portable `cypher.txt` artifact:
```bash
graphify export falkordb
```
**If `--falkordb-push <uri>`** - push directly to a running FalkorDB instance. Credentials are optional; ask the user only if the instance requires auth:
```bash
graphify export falkordb --push falkordb://localhost:6379
```
Default URI is `falkordb://localhost:6379` (the scheme is informational - `redis://` or a bare `host:port` work too), auth is optional, and the target graph defaults to `graphify`. Uses MERGE - safe to re-run without creating duplicates.
### Step 7b - SVG export (only if --svg flag)
```bash
graphify export svg
```
### Step 7c - GraphML export (only if --graphml flag)
```bash
graphify export graphml
```
### Step 7d - MCP server (only if --mcp flag)
```bash
$(cat graphify-out/.graphify_python) -m graphify.serve graphify-out/graph.json
```
This starts a stdio MCP server that exposes tools: `query_graph`, `get_node`, `get_neighbors`, `get_community`, `god_nodes`, `graph_stats`, `shortest_path`. Add to Claude Desktop or any MCP-compatible agent orchestrator so other agents can query the graph live.
To configure in Claude Desktop, add to `claude_desktop_config.json`. Claude Desktop can't run `$(...)`, and under `uv tool install` the system `python3` can't import graphify — so set `command` to the **absolute interpreter path** printed by `cat graphify-out/.graphify_python`:
```json
{
"mcpServers": {
"graphify": {
"command": "<absolute path from: cat graphify-out/.graphify_python>",
"args": ["-m", "graphify.serve", "/absolute/path/to/graphify-out/graph.json"]
}
}
}
```
### Step 8 - Token reduction benchmark (only if total_words > 5000)
If `total_words` from `graphify-out/.graphify_detect.json` is greater than 5,000, run:
```bash
graphify benchmark
```
Print the output directly in chat. If `total_words <= 5000`, skip silently - the graph value is structural clarity, not token compression, for small corpora.
@@ -1,70 +0,0 @@
# graphify reference: extraction subagent prompt
Load this in Step 3 Part B when the corpus has at least one doc, paper, or image chunk. A pure-code corpus skips Part B and never reads this file. Each semantic subagent receives the prompt below verbatim (substitute FILE_LIST, CHUNK_NUM, TOTAL_CHUNKS, DEEP_MODE, and CHUNK_PATH).
```
You are a graphify extraction subagent. Read the files listed and extract a knowledge graph fragment.
Output ONLY valid JSON matching the schema below - no explanation, no markdown fences, no preamble.
Files (chunk CHUNK_NUM of TOTAL_CHUNKS):
FILE_LIST
Rules:
- EXTRACTED: relationship explicit in source (import, call, citation, "see §3.2")
- INFERRED: reasonable inference (shared data structure, implied dependency)
- AMBIGUOUS: uncertain - flag for review, do not omit
Code files: focus on semantic edges AST cannot find (call relationships, shared data, arch patterns).
Do not re-extract imports - AST already has those.
Doc/paper files: extract named concepts, entities, citations. For rationale (WHY decisions were made, trade-offs, design intent): store as a `rationale` attribute on the relevant concept node — do NOT create a separate rationale node or fragment node. Only create a node for something that is itself a named entity or concept. Use `file_type:"rationale"` for concept-like nodes (ideas, principles, mechanisms, design patterns). `file_type` MUST be one of exactly these six values: `code`, `document`, `paper`, `image`, `rationale`, `concept`. Any other value is invalid and will be rejected.
Code files: when adding `calls` edges, source MUST be the caller (the function/class doing the calling), target MUST be the callee. Never reverse this direction. `calls` edges MUST stay within one language: a Python function cannot `calls` a JS/TS/Go/Rust/Java symbol and vice versa — cross-language call edges are phantom artifacts, never emit them.
Image files: use vision to understand what the image IS - do not just OCR.
UI screenshot: layout patterns, design decisions, key elements, purpose.
Chart: metric, trend/insight, data source.
Tweet/post: claim as node, author, concepts mentioned.
Diagram: components and connections.
Research figure: what it demonstrates, method, result.
Handwritten/whiteboard: ideas and arrows, mark uncertain readings AMBIGUOUS.
DEEP_MODE (if --mode deep was given): be aggressive with INFERRED edges - indirect deps,
shared assumptions, latent couplings. Mark uncertain ones AMBIGUOUS instead of omitting.
Semantic similarity: if two concepts in this chunk solve the same problem or represent the same idea without any structural link (no import, no call, no citation), add a `semantically_similar_to` edge marked INFERRED with a confidence_score reflecting how similar they are (0.6-0.95). Examples:
- Two functions that both validate user input but never call each other
- A class in code and a concept in a paper that describe the same algorithm
- Two error types that handle the same failure mode differently
Only add these when the similarity is genuinely non-obvious and cross-cutting. Do not add them for trivially similar things.
Hyperedges: if 3 or more nodes clearly participate together in a shared concept, flow, or pattern that is not captured by pairwise edges alone, add a hyperedge to a top-level `hyperedges` array. Examples:
- All classes that implement a common protocol or interface
- All functions in an authentication flow (even if they don't all call each other)
- All concepts from a paper section that form one coherent idea
Use sparingly — only when the group relationship adds information beyond the pairwise edges. Maximum 3 hyperedges per chunk.
If a file has YAML frontmatter (--- ... ---), copy source_url, captured_at, author,
contributor onto every node from that file.
confidence_score is REQUIRED on every edge - never omit it, never use 0.5 as a default:
- EXTRACTED edges: confidence_score = 1.0 always
- INFERRED edges: pick exactly ONE value from this set — never 0.5:
0.95 direct structural evidence (shared data structure, named cross-file reference).
0.85 strong inference (clear functional alignment, no direct symbol link).
0.75 reasonable inference (shared problem domain + similar shape, requires interpretation).
0.65 weak inference (thematically related, no shape evidence).
0.55 speculative but plausible (surface-level co-occurrence only).
Models follow discrete rubrics better than continuous ranges; the bimodal
distribution observed in production (>50% at 0.5, >40% at 0.85+) shows the
range guidance is being collapsed to a binary. If no value above fits, mark
the edge AMBIGUOUS rather than picking 0.4 or below.
- AMBIGUOUS edges: 0.1-0.3
Node ID format: lowercase, only `[a-z0-9_]`, no dots or slashes. Format: `{stem}_{entity}` where stem is the **full repo-relative path with the extension dropped**, every path segment kept and joined with `_` (each segment lowercased with non-alphanumeric chars replaced by `_`), and entity is the symbol name similarly normalized. Use every directory level, not just the immediate parent — this keeps same-named files in different directories distinct. Examples: `src/auth/session.py` + `ValidateToken` → `src_auth_session_validatetoken`; `lib/utils/helpers.py` + `parse_url` → `lib_utils_helpers_parse_url`; `tests/test_foo.py` + `_helper` → `tests_test_foo_helper`; `docs/v1/api/README.md` + `getUser` → `docs_v1_api_readme_getuser`. Top-level files (no parent dir, e.g. `setup.py`) use just the filename stem: `setup_my_func`. This must match the ID the AST extractor generates — using just the filename (e.g., `session_validatetoken`) or only the immediate parent (e.g., `auth_session_validatetoken`) will create orphan ghost-duplicate nodes. If you are re-extracting a project built under the old immediate-parent format, the user should run `graphify extract --force` to rebuild cleanly. CRITICAL: never append chunk numbers, sequence numbers, or any suffix to an ID (no `_c1`, `_c2`, `_chunk2`, etc.). IDs must be deterministic from the label alone — the same entity must always produce the same ID regardless of which chunk processes it.
Generate the extraction JSON matching this schema exactly:
{"nodes":[{"id":"auth_session_validatetoken","label":"Human Readable Name","file_type":"code|document|paper|image|rationale|concept","source_file":"<FILE_LIST path verbatim>","source_location":null,"source_url":null,"captured_at":null,"author":null,"contributor":null}],"edges":[{"source":"node_id","target":"node_id","relation":"calls|implements|references|cites|conceptually_related_to|shares_data_with|semantically_similar_to|rationale_for","confidence":"EXTRACTED|INFERRED|AMBIGUOUS","confidence_score":1.0,"source_file":"<FILE_LIST path verbatim>","source_location":null,"weight":1.0}],"hyperedges":[{"id":"snake_case_id","label":"Human Readable Label","nodes":["node_id1","node_id2","node_id3"],"relation":"participate_in|implement|form","confidence":"EXTRACTED|INFERRED","confidence_score":0.75,"source_file":"<FILE_LIST path verbatim>"}],"input_tokens":0,"output_tokens":0}
source_file RULE (every node, edge, and hyperedge): set source_file to the path of the originating file EXACTLY as it appears in FILE_LIST — verbatim and absolute. Do NOT shorten to a basename, do NOT re-relativize, do NOT strip any directory prefix, and do NOT change separators (the engine canonicalizes separators and relativizes against the build root downstream). Copy the FILE_LIST entry character-for-character. This keeps the full build and incremental --update on the same base, so build_merge's replace-on-re-extract matches the existing node instead of accumulating a duplicate.
Then write the JSON to disk using the Write tool at this exact absolute path (no relative paths — Write resolves relative paths against an undefined cwd and the file will be silently lost):
CHUNK_PATH
```
@@ -1,46 +0,0 @@
# graphify reference: GitHub clone and cross-repo merge
Load this when the user passed one or more `https://github.com/...` URLs, or named several local subfolders to merge into one graph.
### Step 0 - Clone GitHub repo(s) (only if a GitHub URL was given)
**Single repo:**
```bash
LOCAL_PATH=$(graphify clone <github-url> [--branch <branch>])
# Use LOCAL_PATH as the target for all subsequent steps
```
**Multiple repos (cross-repo graph):**
```bash
# Clone each repo, run the full pipeline on each, then merge
graphify clone <url1> # → ~/.graphify/repos/<owner1>/<repo1>
graphify clone <url2> # → ~/.graphify/repos/<owner2>/<repo2>
# Run /graphify on each local path to produce their graph.json files
# Then merge:
graphify merge-graphs \
~/.graphify/repos/<owner1>/<repo1>/graphify-out/graph.json \
~/.graphify/repos/<owner2>/<repo2>/graphify-out/graph.json \
--out graphify-out/cross-repo-graph.json
```
Graphify clones into `~/.graphify/repos/<owner>/<repo>` and reuses existing clones on repeat runs. Each node in the merged graph carries a `repo` attribute so you can filter by origin.
**Multiple local subfolders (monorepo or multi-service layout):**
The skill pipeline writes all intermediate and final outputs to `graphify-out/` in the current working directory. Running the skill on each subfolder separately will clobber the same output dir. Instead, use the CLI directly for each subfolder — it places `graphify-out/` *inside* the scanned path:
```bash
graphify extract ./core/ # → ./core/graphify-out/graph.json
graphify extract ./service/ # → ./service/graphify-out/graph.json
graphify extract ./platform/ # → ./platform/graphify-out/graph.json
# Add --backend gemini|kimi|openai|deepseek|claude-cli depending on which API key you have set
# Then merge at the project root:
graphify merge-graphs \
./core/graphify-out/graph.json \
./service/graphify-out/graph.json \
./platform/graphify-out/graph.json \
--out graphify-out/graph.json
```
Once `graphify-out/graph.json` exists, the fast path above takes over: any codebase question runs `graphify query` directly on the merged graph — no re-extraction, no size gate.
-33
View File
@@ -1,33 +0,0 @@
# graphify reference: commit hook and native CLAUDE.md integration
Load this when the user asked to install the post-commit hook or wire graphify into a project's CLAUDE.md.
## For git commit hook
Install a post-commit hook that auto-rebuilds the graph after every commit. No background process needed - triggers once per commit, works with any editor.
```bash
graphify hook install # install
graphify hook uninstall # remove
graphify hook status # check
```
After every `git commit`, the hook detects which code files changed (via `git diff HEAD~1`), re-runs AST extraction on those files, and rebuilds `graph.json` and `GRAPH_REPORT.md`. Doc/image changes are ignored by the hook - run `/graphify --update` manually for those.
If a post-commit hook already exists, graphify appends to it rather than replacing it.
---
## For native CLAUDE.md integration
Run once per project to make graphify always-on in Claude Code sessions:
```bash
graphify claude install
```
This writes a `## graphify` section to the local `CLAUDE.md` that instructs Claude to check the graph before answering codebase questions and rebuild it after code changes. No manual `/graphify` needed in future sessions.
```bash
graphify claude uninstall # remove the section
```
-311
View File
@@ -1,311 +0,0 @@
# graphify reference: query, path, explain
Load this when the user asks a question against an existing graph, or runs `/graphify path` or `/graphify explain`. The core's query stub points here for the full traversal flow. These flows use the `graphify query` CLI when it is available and fall back to an inline NetworkX traversal otherwise.
Two traversal modes - choose based on the question:
| Mode | Flag | Best for |
|------|------|----------|
| BFS (default) | _(none)_ | "What is X connected to?" - broad context, nearest neighbors first |
| DFS | `--dfs` | "How does X reach Y?" - trace a specific chain or dependency path |
First check the graph exists:
```bash
$(cat graphify-out/.graphify_python) -c "
from pathlib import Path
if not Path('graphify-out/graph.json').exists():
print('ERROR: No graph found. Run /graphify <path> first to build the graph.')
raise SystemExit(1)
"
```
If it fails, stop and tell the user to run `/graphify <path>` first.
### Step 0 — Constrained query expansion (REQUIRED before traversal)
graphify's `query` CLI matches nodes via case-folded substring + IDF — there is **no stemming, no synonyms, no cross-language match** inside the binary, and the inline fallback below matches the same way. If the user's question uses different language or different domain vocabulary than the graph's labels (user says "обработчик" / graph says "handler"; user says "authentication" / graph says "Guardian"), the literal matcher returns 0 hits and the answer collapses to noise.
Fix this **without inventing tokens** by expanding the query against the actual graph vocabulary first:
1. Extract the token vocabulary from node labels:
```bash
$(cat graphify-out/.graphify_python) -c "
import json, re
from pathlib import Path
data = json.loads(Path('graphify-out/graph.json').read_text(encoding='utf-8'))
vocab = set()
for n in data['nodes']:
for c in re.findall(r'[^\W\d_]+', n.get('label','') or '', re.UNICODE):
parts = re.findall(r'[A-Z]+(?=[A-Z][a-z])|[A-Z]?[a-z]+|[A-Z]+', c) or [c]
for p in parts:
t = p.lower()
if 3 <= len(t) <= 30:
vocab.add(t)
Path('graphify-out/.vocab.txt').write_text('\n'.join(sorted(vocab)), encoding='utf-8')
print(f'vocab: {len(vocab)} tokens')
"
```
2. Read `graphify-out/.vocab.txt`. Then for the user's question, select **up to 12 tokens from this exact list** that semantically match the query intent. Hard constraints:
- You MUST pick only tokens present in the vocabulary file. Do NOT invent tokens.
- If a query concept has no plausible token in the vocab, skip it — do not substitute a near-synonym from training memory.
- If **no** vocab tokens match the query at all, output an empty list and tell the user the corpus has no relevant vocabulary for this question. Do not fabricate a search.
- Translate cross-language: Russian "аутентификация" → look for `auth`, `credential`, `token`, `security` IFF present in vocab.
- Morphology: "handlers" maps to `handler` IFF present; "todos" maps to `todo` IFF present.
3. Print the selection explicitly to the user before running the query, so the expansion is auditable:
```
Query expanded to (from graph vocab, N tokens): [token1, token2, ...]
```
If the list is empty, say so plainly and stop — do not proceed to traversal.
### Step 1 — Traversal
Build the **expanded query string** by joining the selected tokens with spaces. Use this string as `QUESTION` below — NOT the original user question. (The original question is preserved only for `save-result` at the end.)
Prefer the CLI when it is installed:
```bash
graphify query "QUESTION"
# or: graphify query "QUESTION" --dfs --budget 3000
```
If the CLI is unavailable, load `graphify-out/graph.json` and run the traversal inline:
1. Find the 1-3 nodes whose label best matches the expanded tokens.
2. Run the appropriate traversal from each starting node.
3. Read the subgraph - node labels, edge relations, confidence tags, source locations.
4. Answer using **only** what the graph contains. Quote `source_location` when citing a specific fact.
5. If the graph lacks enough information, say so - do not hallucinate edges.
```bash
$(cat graphify-out/.graphify_python) -c "
import sys, json
from networkx.readwrite import json_graph
import networkx as nx
from pathlib import Path
data = json.loads(Path('graphify-out/graph.json').read_text(encoding='utf-8'))
G = json_graph.node_link_graph(data, edges='links')
question = 'QUESTION'
mode = 'MODE' # 'bfs' or 'dfs'
terms = [t.lower() for t in question.split() if len(t) >= 3] # match the vocab threshold; keeps api/jwt/ios (#1392)
# Find best-matching start nodes
scored = []
for nid, ndata in G.nodes(data=True):
label = ndata.get('label', '').lower()
score = sum(1 for t in terms if t in label)
if score > 0:
scored.append((score, nid))
scored.sort(reverse=True)
start_nodes = [nid for _, nid in scored[:3]]
if not start_nodes:
print('No matching nodes found for query terms:', terms)
sys.exit(0)
subgraph_nodes = set()
subgraph_edges = []
if mode == 'dfs':
# DFS: follow one path as deep as possible before backtracking.
# Depth-limited to 6 to avoid traversing the whole graph.
visited = set()
stack = [(n, 0) for n in reversed(start_nodes)]
while stack:
node, depth = stack.pop()
if node in visited or depth > 6:
continue
visited.add(node)
subgraph_nodes.add(node)
for neighbor in G.neighbors(node):
if neighbor not in visited:
stack.append((neighbor, depth + 1))
subgraph_edges.append((node, neighbor))
else:
# BFS: explore all neighbors layer by layer up to depth 3.
frontier = set(start_nodes)
subgraph_nodes = set(start_nodes)
for _ in range(3):
next_frontier = set()
for n in frontier:
for neighbor in G.neighbors(n):
if neighbor not in subgraph_nodes:
next_frontier.add(neighbor)
subgraph_edges.append((n, neighbor))
subgraph_nodes.update(next_frontier)
frontier = next_frontier
# Token-budget aware output: rank by relevance, cut at budget (~4 chars/token)
token_budget = BUDGET # default 2000
char_budget = token_budget * 4
# Score each node by term overlap for ranked output
def relevance(nid):
label = G.nodes[nid].get('label', '').lower()
return sum(1 for t in terms if t in label)
ranked_nodes = sorted(subgraph_nodes, key=relevance, reverse=True)
lines = [f'Traversal: {mode.upper()} | Start: {[G.nodes[n].get(\"label\",n) for n in start_nodes]} | {len(subgraph_nodes)} nodes']
for nid in ranked_nodes:
d = G.nodes[nid]
lines.append(f' NODE {d.get(\"label\", nid)} [src={d.get(\"source_file\",\"\")} loc={d.get(\"source_location\",\"\")}]')
for u, v in subgraph_edges:
if u in subgraph_nodes and v in subgraph_nodes:
_raw = G[u][v]; d = next(iter(_raw.values()), {}) if isinstance(G, nx.MultiGraph) else _raw
lines.append(f' EDGE {G.nodes[u].get(\"label\",u)} --{d.get(\"relation\",\"\")} [{d.get(\"confidence\",\"\")}]--> {G.nodes[v].get(\"label\",v)}')
output = '\n'.join(lines)
if len(output) > char_budget:
output = output[:char_budget] + f'\n... (truncated at ~{token_budget} token budget - use --budget N for more)'
print(output)
"
```
Replace `QUESTION` with the **expanded** query string, `MODE` with `bfs` or `dfs`, and `BUDGET` with the token budget (default `2000`, or whatever `--budget N` specifies). Then answer based on the subgraph output above, using only what the graph contains.
After writing the answer, save it back into the graph so it improves future queries. Include the expanded tokens inside the `--answer` text (e.g. `"Expanded from original query via vocab: [tokens]. Then traversed..."`) so the next `--update` extracts the expansion history as a graph node:
```bash
$(cat graphify-out/.graphify_python) -m graphify save-result --question "ORIGINAL_QUESTION" --answer "ANSWER" --type query --nodes NODE1 NODE2
```
Replace `ORIGINAL_QUESTION` with the user's verbatim question, `ANSWER` with your full answer text (containing the expanded-token trace), `NODE1 NODE2` with the list of node labels you cited. This closes the feedback loop: the next `--update` will extract this Q&A as a node in the graph.
**Work memory (self-improving loop).** Add an `--outcome` so future sessions learn from this one — append `--outcome useful|dead_end|corrected` to the `save-result` command (and `--correction "the right answer"` when correcting):
- `useful` — the cited nodes answered the question well (they become *preferred sources*).
- `dead_end` — the question/path led nowhere; don't re-derive it next time.
- `corrected` — the saved answer was wrong; `--correction` records what was right.
At the **start** of graph work, refresh and read the lessons: run `graphify reflect --if-stale` (cheap, deterministic, no LLM; `--if-stale` makes it a no-op when `LESSONS.md` is already newer than every input, e.g. when the git hook just refreshed it), then read `graphify-out/reflections/LESSONS.md`. It lists **preferred sources** (start there), **known dead ends** (skip them), and prior **corrections**. Running `reflect` yourself keeps the lessons current even without the git hook installed; if the post-commit hook *is* installed, `--if-stale` means your session-start run costs almost nothing.
---
## For /graphify path
Find the shortest path between two named concepts in the graph. Prefer the CLI when installed:
```bash
graphify path "NODE_A" "NODE_B"
```
If the CLI is unavailable, run it inline:
```bash
$(cat graphify-out/.graphify_python) -c "
import json, sys
import networkx as nx
from networkx.readwrite import json_graph
from pathlib import Path
data = json.loads(Path('graphify-out/graph.json').read_text(encoding='utf-8'))
G = json_graph.node_link_graph(data, edges='links')
a_term = 'NODE_A'
b_term = 'NODE_B'
def find_node(term):
term = term.lower()
scored = sorted(
[(sum(1 for w in term.split() if w in G.nodes[n].get('label','').lower()), n)
for n in G.nodes()],
reverse=True
)
return scored[0][1] if scored and scored[0][0] > 0 else None
src = find_node(a_term)
tgt = find_node(b_term)
if not src or not tgt:
print(f'Could not find nodes matching: {a_term!r} or {b_term!r}')
sys.exit(0)
try:
path = nx.shortest_path(G, src, tgt)
print(f'Shortest path ({len(path)-1} hops):')
for i, nid in enumerate(path):
label = G.nodes[nid].get('label', nid)
if i < len(path) - 1:
_raw = G[nid][path[i+1]]; edge = next(iter(_raw.values()), {}) if isinstance(G, nx.MultiGraph) else _raw
rel = edge.get('relation', '')
conf = edge.get('confidence', '')
print(f' {label} --{rel}--> [{conf}]')
else:
print(f' {label}')
except nx.NetworkXNoPath:
print(f'No path found between {a_term!r} and {b_term!r}')
except nx.NodeNotFound as e:
print(f'Node not found: {e}')
"
```
Replace `NODE_A` and `NODE_B` with the actual concept names from the user. Then explain the path in plain language - what each hop means, why it's significant.
After writing the explanation, save it back:
```bash
$(cat graphify-out/.graphify_python) -m graphify save-result --question "Path from NODE_A to NODE_B" --answer "ANSWER" --type path_query --nodes NODE_A NODE_B
```
---
## For /graphify explain
Give a plain-language explanation of a single node - everything connected to it. Prefer the CLI when installed:
```bash
graphify explain "NODE_NAME"
```
If the CLI is unavailable, run it inline:
```bash
$(cat graphify-out/.graphify_python) -c "
import json, sys
import networkx as nx
from networkx.readwrite import json_graph
from pathlib import Path
data = json.loads(Path('graphify-out/graph.json').read_text(encoding='utf-8'))
G = json_graph.node_link_graph(data, edges='links')
term = 'NODE_NAME'
term_lower = term.lower()
# Find best matching node
scored = sorted(
[(sum(1 for w in term_lower.split() if w in G.nodes[n].get('label','').lower()), n)
for n in G.nodes()],
reverse=True
)
if not scored or scored[0][0] == 0:
print(f'No node matching {term!r}')
sys.exit(0)
nid = scored[0][1]
data_n = G.nodes[nid]
print(f'NODE: {data_n.get(\"label\", nid)}')
print(f' source: {data_n.get(\"source_file\",\"unknown\")}')
print(f' type: {data_n.get(\"file_type\",\"unknown\")}')
print(f' degree: {G.degree(nid)}')
print()
print('CONNECTIONS:')
for neighbor in G.neighbors(nid):
_raw = G[nid][neighbor]; edge = next(iter(_raw.values()), {}) if isinstance(G, nx.MultiGraph) else _raw
nlabel = G.nodes[neighbor].get('label', neighbor)
rel = edge.get('relation', '')
conf = edge.get('confidence', '')
src_file = G.nodes[neighbor].get('source_file', '')
print(f' --{rel}--> {nlabel} [{conf}] ({src_file})')
"
```
Replace `NODE_NAME` with the concept the user asked about. Then write a 3-5 sentence explanation of what this node is, what it connects to, and why those connections are significant. Use the source locations as citations.
After writing the explanation, save it back:
```bash
$(cat graphify-out/.graphify_python) -m graphify save-result --question "Explain NODE_NAME" --answer "ANSWER" --type explain --nodes NODE_NAME
```
-52
View File
@@ -1,52 +0,0 @@
# graphify reference: transcribe video and audio
Load this only when `detect` reported one or more `video` files. A corpus with no video never reads this.
### Step 2.5 - Transcribe video / audio files (only if video files detected)
Skip this step entirely if `detect` returned zero `video` files.
Video and audio files cannot be read directly. Transcribe them to text first, then treat the transcripts as doc files in Step 3.
**Strategy:** Read the god nodes from `graphify-out/.graphify_detect.json` (or the analysis file if it exists from a previous run). You are already a language model — write a one-sentence domain hint yourself from those labels. Then pass it to Whisper as the initial prompt. No separate API call needed.
**However**, if the corpus has *only* video files and no other docs/code, use the generic fallback prompt: `"Use proper punctuation and paragraph breaks."`
**Step 1 - Write the Whisper prompt yourself.**
Read the top god node labels from detect output or analysis, then compose a short domain hint sentence, for example:
- Labels: `transformer, attention, encoder, decoder` → `"Machine learning research on transformer architectures and attention mechanisms. Use proper punctuation and paragraph breaks."`
- Labels: `kubernetes, deployment, pod, helm` → `"DevOps discussion about Kubernetes deployments and Helm charts. Use proper punctuation and paragraph breaks."`
**Export** it as `GRAPHIFY_WHISPER_PROMPT` (the exact name the transcriber reads — and it must be `export`ed so the child Python process sees it) for the next command.
**Step 2 - Transcribe:**
```bash
export GRAPHIFY_WHISPER_MODEL=base # or whatever --whisper-model the user passed (must be exported)
export GRAPHIFY_WHISPER_PROMPT="<the one-sentence domain hint you composed in Step 1>"
$(cat graphify-out/.graphify_python) -c "
import json, os, sys
from pathlib import Path
from graphify.transcribe import transcribe_all
detect = json.loads(Path('graphify-out/.graphify_detect.json').read_text(encoding=\"utf-8\"))
video_files = detect.get('files', {}).get('video', [])
prompt = os.environ.get('GRAPHIFY_WHISPER_PROMPT', 'Use proper punctuation and paragraph breaks.')
transcript_paths = transcribe_all(video_files, initial_prompt=prompt)
# Write the JSON from Python (NOT a shell '>' redirect): transcribe_all/Whisper
# print progress to stdout, which would otherwise corrupt the JSON file (#1392).
Path('graphify-out/.graphify_transcripts.json').write_text(json.dumps(transcript_paths, ensure_ascii=False), encoding=\"utf-8\")
print(f'Transcribed {len(transcript_paths)} file(s)', file=sys.stderr)
"
```
After transcription:
- Read the transcript paths from `graphify-out/.graphify_transcripts.json`
- Add them to the docs list before dispatching semantic subagents in Step 3B
- Print how many transcripts were created: `Transcribed N video file(s) -> treating as docs`
- If transcription fails for a file, print a warning and continue with the rest
**Whisper model:** Default is `base`. If the user passed `--whisper-model <name>`, `export GRAPHIFY_WHISPER_MODEL=<name>` (it must be exported, not just assigned) before running the command above.
-192
View File
@@ -1,192 +0,0 @@
# graphify reference: incremental update and cluster-only
Load this only when the user passed `--update` or `--cluster-only`. A first-time full build never reads this file.
## For --update (incremental re-extraction)
Use when you've added or modified files since the last run. Only re-extracts changed files - saves tokens and time.
```bash
$(cat graphify-out/.graphify_python) -c "
import sys, json
from graphify.detect import detect_incremental, save_manifest
from pathlib import Path
result = detect_incremental(Path('INPUT_PATH'))
new_total = result.get('new_total', 0)
print(json.dumps(result, indent=2, ensure_ascii=False))
Path('graphify-out/.graphify_incremental.json').write_text(json.dumps(result, ensure_ascii=False), encoding=\"utf-8\")
deleted = list(result.get('deleted_files', []))
if new_total == 0 and not deleted:
print('No files changed since last run. Nothing to update.')
raise SystemExit(0)
if deleted:
print(f'{len(deleted)} deleted file(s) to prune.')
if new_total > 0:
print(f'{new_total} new/changed file(s) to re-extract.')
"
```
Then populate `.graphify_detect.json` so Steps 3A–6 (which read it unconditionally) see the right state for an incremental run. `files` carries the changed subset (drives Step 3A AST + Step 3B0 cache check on only what changed); `all_files` carries the full corpus for any step that needs corpus-wide context:
```bash
$(cat graphify-out/.graphify_python) -c "
import json
from pathlib import Path
r = json.loads(Path('graphify-out/.graphify_incremental.json').read_text(encoding=\"utf-8\"))
Path('graphify-out/.graphify_detect.json').write_text(json.dumps({
'files': r.get('new_files', {}),
'all_files': r.get('files', {}),
'total_files': r.get('new_total', 0),
'total_words': r.get('total_words', 0),
'skipped_sensitive': r.get('skipped_sensitive', []),
'needs_graph': True,
}, ensure_ascii=False), encoding=\"utf-8\")
"
```
If new files exist, first check whether all changed files are code files:
```bash
$(cat graphify-out/.graphify_python) -c "
import json
from pathlib import Path
result = json.loads(open('graphify-out/.graphify_incremental.json', encoding='utf-8').read()) if Path('graphify-out/.graphify_incremental.json').exists() else {}
code_exts = {'.py','.ts','.js','.go','.rs','.java','.cpp','.c','.rb','.swift','.kt','.cs','.scala','.php','.cc','.cxx','.hpp','.h','.kts','.lua','.toc','.f','.F','.f90','.F90','.f95','.F95','.f03','.F03','.f08','.F08'}
new_files = result.get('new_files', {})
all_changed = [f for files in new_files.values() for f in files]
code_only = all(Path(f).suffix.lower() in code_exts for f in all_changed)
print('code_only:', code_only)
"
```
If `code_only` is True: print `[graphify update] Code-only changes detected - skipping semantic extraction (no LLM needed)`, run only Step 3A (AST) on the changed files, skip Step 3B entirely (no subagents), then go straight to merge and Steps 4–8.
If `code_only` is False (any changed file is a doc/paper/image/video): **first, if any changed file is in `new_files['video']`, run `references/transcribe.md` (Step 2.5) on those files, then rewrite `.graphify_detect.json` to move the resulting transcript paths into `files['document']` and drop `files['video']`** — otherwise raw `.mp4/.mp3` paths are fed to semantic subagents as unreadable media (#1392). Then run the full Steps 3A–3C pipeline as normal.
If no new files exist (only deletions), create an empty extraction so the merge step can prune:
```bash
if [ ! -f graphify-out/.graphify_extract.json ]; then
echo '[graphify update] Only deletions -- creating empty extraction for merge.'
$(cat graphify-out/.graphify_python) -c "
import json
from pathlib import Path
Path('graphify-out/.graphify_extract.json').write_text(json.dumps({'nodes':[],'edges':[],'hyperedges':[],'input_tokens':0,'output_tokens':0}), encoding='utf-8')
"
fi
```
Then:
```bash
$(cat graphify-out/.graphify_python) -c "
import json
from pathlib import Path
from graphify.build import build_merge
from graphify.detect import save_manifest
# Load new extraction and incremental state
new_extraction = json.loads(Path('graphify-out/.graphify_extract.json').read_text(encoding=\"utf-8\"))
incremental = json.loads(Path('graphify-out/.graphify_incremental.json').read_text(encoding=\"utf-8\"))
deleted = list(incremental.get('deleted_files', []))
# prune_sources is ONLY for genuinely DELETED files. Changed/re-extracted files are
# handled by build_merge's replace-on-re-extract (#1344): every source_file in
# new_chunks is dropped from the base before merge, so old/stale nodes don't survive.
# Do NOT add `changed` here: with root= passed, prune_set relativizes to the same base
# as the freshly merged nodes and would DELETE the re-extracted content (#1178 is moot
# now that replace — not the dedup pass — reconciles changed files).
prune = list(deleted) or None
# Use build_merge() — reads graph.json directly without NetworkX round-trip
# so edge direction (calls, implements, imports) is always preserved (#801).
# Pass root= so prune_sources (absolute paths from detect_incremental) are
# relativized to match the graph's relative source_file values; without it
# nothing is pruned and stale nodes accumulate on every update (#1361).
# directed=IS_DIRECTED: replace IS_DIRECTED with True if --directed was given, else
# False. Without it a --directed --update silently rebuilds undirected and collapses
# reciprocal A<->B edges (#1392).
G = build_merge(
[new_extraction],
graph_path='graphify-out/graph.json',
prune_sources=prune,
root='INPUT_PATH',
directed=IS_DIRECTED,
)
print(f'[graphify update] Merged: {G.number_of_nodes()} nodes, {G.number_of_edges()} edges')
# Write merged result back to .graphify_extract.json so Step 4 sees the full graph
merged_out = {
'nodes': [{'id': n, **d} for n, d in G.nodes(data=True)],
'edges': [
# Explicit source/target last so they win over any stale attrs in d.
{**{k: val for k, val in d.items() if k not in ('_src', '_tgt', 'source', 'target')},
'source': d.get('_src', u), 'target': d.get('_tgt', v)}
for u, v, d in G.edges(data=True)
],
# G.graph["hyperedges"] holds hyperedges from both existing graph.json
# and new_extraction (build_merge combines them). Falling back to
# new_extraction only would silently drop prior-run hyperedges (#801).
'hyperedges': list(G.graph.get('hyperedges', [])),
'input_tokens': new_extraction.get('input_tokens', 0),
'output_tokens': new_extraction.get('output_tokens', 0),
}
Path('graphify-out/.graphify_extract.json').write_text(json.dumps(merged_out, ensure_ascii=False), encoding=\"utf-8\")
print(f'[graphify update] Merged extraction written ({len(merged_out[\"nodes\"])} nodes, {len(merged_out[\"edges\"])} edges)')
# Save manifest so next --update diffs against today's state, not the
# prior run's baseline (prevents ghost-node reports on subsequent updates).
# root= matches the build_merge call above so the manifest keys stay relative to
# the scan root — portable across clones/machines, so --update keeps matching
# cached files instead of missing every one after a move (#1417).
save_manifest(incremental['files'], root='INPUT_PATH')
print('[graphify update] Manifest saved.')
"
```
Then run Steps 4–8 on the merged graph as normal.
After Step 4, show the graph diff:
```bash
$(cat graphify-out/.graphify_python) -c "
import json
from graphify.analyze import graph_diff
from graphify.build import build_from_json
from networkx.readwrite import json_graph
import networkx as nx
from pathlib import Path
# Load old graph (before update) from backup written before merge
old_data = json.loads(Path('graphify-out/.graphify_old.json').read_text(encoding=\"utf-8\")) if Path('graphify-out/.graphify_old.json').exists() else None
new_extract = json.loads(Path('graphify-out/.graphify_extract.json').read_text(encoding=\"utf-8\"))
G_new = build_from_json(new_extract, directed=IS_DIRECTED)
if old_data:
G_old = json_graph.node_link_graph(old_data, edges='links')
diff = graph_diff(G_old, G_new)
print(diff['summary'])
if diff['new_nodes']:
print('New nodes:', ', '.join(n['label'] for n in diff['new_nodes'][:5]))
if diff['new_edges']:
print('New edges:', len(diff['new_edges']))
"
```
Before the merge step, save the old graph: `cp graphify-out/graph.json graphify-out/.graphify_old.json`
Clean up after: `rm -f graphify-out/.graphify_old.json`
---
## For --cluster-only
Skip Steps 1–3. Re-run clustering on the existing graph:
```bash
graphify cluster-only .
```
`graphify cluster-only .` is **self-contained**: it re-clusters, names communities, and regenerates `GRAPH_REPORT.md`, `graph.json`, and `graph.html` from the existing graph. **Do not re-run Steps 5–9** — they read intermediate files (`.graphify_extract.json`, `.graphify_detect.json`, `.graphify_analysis.json`) that a prior build's cleanup (Step 9) already deleted, so they raise `FileNotFoundError` (#1392). When it finishes, present the refreshed `GRAPH_REPORT.md` summary as usual.
+19 -7
View File
@@ -47,6 +47,10 @@ git log --oneline -3
as `/bugfix` (root-cause investigation, then a scoped fix)."
- Settle the proposed fix HERE — the executor cannot ask questions, so the
exact edit (what changes, in which file(s)) must be closed before dispatch.
- Then run pass B of `$HOME/.claude/lib/contract-interview.md` against that
edit: a VISIBLE / PUBLIC NAME / SCOPE choice the bug description leaves
open (which way the icon aligns, the label's wording) → ask before
dispatch. A typo or a wrong value asks nothing.
OPTIONAL — memory check (exempt by default; hotfix = obvious fix, mirror of its capitalize
skip). For a RECURRING or urgent bug only, a quick blockers-only glance may save time:
@@ -66,8 +70,9 @@ Follow `$HOME/.claude/lib/design-gate.md`:
## STEP 1.7 — CONTRACT (silent autofill)
Run `$HOME/.claude/lib/contract-interview.md` at hotfix weight: **zero
questions ever** (a hotfix is an obvious fix by definition). Autofill the
Run `$HOME/.claude/lib/contract-interview.md` at hotfix weight: pass A is a
silent autofill (a hotfix is an obvious fix by definition); pass B already
ran at STEP 1, ask nothing more here. Autofill the
contract — REQUEST verbatim = the bug description as given; ACCEPTANCE
CRITERIA = "symptom gone; build/tests green"; FILE SCOPE = the 1-2 target
files from STEP 1. It writes `.claude/tasks/contracts/<date>-<slug>-<HHMM>.md`.
@@ -135,8 +140,14 @@ security dispatch, no revert. Finish with the HOTFIX-EXEC REPORT."
Parse the `HOTFIX-EXEC REPORT`:
- `STATUS : DONE` → STEP 4 (the SMOKE line in the report decides pass/fail
there; DONE here means execution completed, not that it verified clean).
- `STATUS : BLOCKED` → if any edits were made, revert ONLY the executor's
files: `git restore --source=$PRE -- <FILE(S) from the report>` and delete
- `STATUS : BLOCKED` with `CLASS: visible | public-name | scope` in NOTES →
the executor halted at an open choice before editing (nothing to revert):
ask the user per MID-RUN CLARIFICATION in
`$HOME/.claude/lib/contract-interview.md`, append the answer to the
contract `[gated]`, re-dispatch ONCE with the closed choice. This is the
one re-dispatch hotfix allows; it is not a retry of a failed attempt.
- `STATUS : BLOCKED` otherwise → if any edits were made, revert ONLY the
executor's files: `git restore --source=$PRE -- <FILE(S) from the report>` and delete
any NEW file the report lists (untracked, absent from $PRE). Never
`git restore .` — it would wipe the tolerated pre-existing edits too.
Surface the blocker to the user; STOP. One attempt only — hotfix never
@@ -227,9 +238,10 @@ trivial hotfix still produces a `chore(memory): journal — …` commit (Frame 2
- Reflection (LOCATE, contract, gate decisions) NEVER leaves this main
loop; execution NEVER stays in it — the executor is the sonnet-pinned
hotfixer subagent (BDR-066).
- The executor is dispatched FRESH, once — hotfix never re-dispatches (no
decision round-trips; a blocked or failed attempt reverts and escalates
to `/bugfix`, it does not retry).
- The executor is dispatched FRESH, once — hotfix never re-dispatches after
a failed or blocked attempt (it reverts and escalates to `/bugfix`, it
does not retry). Sole exception: a class-tagged BLOCKED answered by the
user (STEP 3), re-dispatched once with the closed choice.
- Design gate only if CSS/style signals detected. See STEP 1.5.
- **Revert-not-loop preserved**: smoke FAIL or security BLOCK →
file-scoped revert from `$PRE` (STEP 4's protocol — never `git
+5 -2
View File
@@ -58,8 +58,8 @@ In both cases: MANDATORY STOP until user answers remaining questions. Produce PR
**Then run `$HOME/.claude/lib/contract-interview.md`** seeded from the BRIEF:
REQUEST verbatim = the user's project description; ACCEPTANCE CRITERIA = the
V1 FEATURES (each testable); FILE SCOPE = the planned tree. No new questions
(the interview already asked). It writes
V1 FEATURES (each testable); FILE SCOPE = the planned tree. Pass A is covered
by the interview; pass B runs at STEP 3 against the DESIGN. It writes
`.claude/tasks/contracts/<date>-<slug>-<HHMM>.md`; the DESIGN approved at STEP
4 ENRICHES it, and STEP 9's verifier judges the MVP against the enriched
contract.
@@ -70,6 +70,9 @@ Load `$HOME/.claude/agents/analyzer.md`. Analyze BRIEF: existing code, stack con
## STEP 3 — DESIGN
Invoke `superpowers:brainstorming` with BRIEF + ANALYSIS REPORT.
Produce DESIGN: stack+versions, full folder tree, module responsibilities, data flow, interfaces (signatures only), config+tooling, test strategy, resolved decisions, prereqs list.
Then run pass B of `$HOME/.claude/lib/contract-interview.md` against the DESIGN
(minus what the BRIEF and the brainstorm settled): one batch before STEP 4;
answers append to the contract `[gated]`.
## STEP 4 — VALIDATION GATE #1 ★ MANDATORY STOP
Present:
+4
View File
@@ -116,6 +116,10 @@ Refine request into validated design via Socratic questioning. Don't proceed unt
Invoke `superpowers:writing-plans` with the validated design AND the 0d digest: every task
must be consistent with the in-force constraints; where a task implements or affects one,
note the ID inline. Break design into tasks (2-5 min each). Each task: exact file paths, full code, verification steps.
Then run pass B of `$HOME/.claude/lib/contract-interview.md` against the plan:
every VISIBLE / PUBLIC NAME / SCOPE choice the plan settles that neither the
request nor the STEP 1 brainstorm settled (check the contract's CLARIFICATIONS
first) → one batch before STEP 2b; answers append to the contract `[gated]`.
## STEP 2b — CHALLENGE THE PLAN (adversarial, before the gate)
Before the human sees the plan, harden it. Run `$HOME/.claude/lib/challenge-plan.md`:
+75
View File
@@ -46,6 +46,79 @@ Always write the file-write ban as `Edit(...)`.
| `auto` | Research preview — agentic default, permission model evolving. This config's default (BDR-004) | Daily driving with guardrails |
| `bypassPermissions` | Skips all prompts — **dangerous** | CI/CD only, sandboxed env |
## Auto mode (`autoMode`)
With `defaultMode: auto`, a classifier decides each action instead of a static
prompt. The `autoMode` block is what you hand that classifier.
| Key | What it holds |
|---|---|
| `environment` | Facts about the machine and the repo. Context, not rules. |
| `allow` | Action classes the classifier may clear on its own. |
| `soft_deny` | Destructive or irreversible actions. Explicit user intent clears them. |
| `hard_deny` | Security boundaries. User intent does **not** clear them. |
| `classifyAllShell` | `true` suspends every Bash allow rule so all shell goes through the classifier. |
All four lists are prose spliced into the classifier prompt, not permission-rule
syntax. Write `Sending SIGKILL reaches processes outside this session`, not
`Bash(kill -9 *)`.
### `$defaults`
Each list **replaces** the built-in entries unless it contains the literal
string `"$defaults"`, which splices them in at that position. Put it first and
your own entries refine what follows. Omit it and you silently drop every
built-in rule, which is almost never the intent.
### Scope it right
`autoMode` in `~/.claude/settings.json` reaches **every** project on the
machine. Project facts (this repo's deploy target, its secrets, its data)
belong in that project's `.claude/settings.local.json`. A global block naming
one repo feeds the classifier false facts in all the others.
### `ask` is not a prompt under auto mode
Verified in-session (LRN-146, re-verified on 2.1.273 on 2026-09-16 with a
`node -e` probe matching an `ask` rule): with `defaultMode: auto`, Bash rules
in `permissions.ask` were auto-approved and raised no prompt. The auto-mode
docs claim the opposite for "content-scoped" rules such as `Bash(git push *)`;
the observed behavior wins until a probe shows a prompt. `deny` is the only
tier the classifier cannot lift.
So for a destructive command you want gated but still reachable, `ask` is the
wrong tier. Use `autoMode.soft_deny`: blocked until the user's intent clears
it. Keep `deny` for what must never run at all.
### Picking a tier
| You want | Tier |
|---|---|
| Never runs, no exception, matchable by a command pattern | `permissions.deny` |
| Never runs, and a pattern cannot express it (a read then a send, a prod target) | `autoMode.hard_deny` |
| Runs when the user asks for it, blocked otherwise | `autoMode.soft_deny` |
| Runs freely when a condition holds that only the classifier can judge (a local dev container, a package declared in the lockfile) | `autoMode.allow` |
| Runs freely | `permissions.allow`, or nothing |
`permissions.ask` is not on this list on purpose. Under `defaultMode: auto` it
gates nothing.
`autoMode.allow` is the exception tier: inside the classifier an `allow` entry
overrides a matching `soft_deny`, built-in or yours, so word it as narrowly as
the condition allows. It is also the only tier that can open an interpreter:
under auto mode Claude Code suspends the static allow rules that grant
arbitrary code execution (`Bash(*)`, wildcarded interpreters such as
`Bash(node *)`), so those commands reach the classifier whatever
`permissions.allow` says. `awk` and `echo` pass through a static rule; `node`
cannot.
### Scope of intent
A `soft_deny` clears on the user's instruction, and this config scopes that to
the **current turn**. An approval from an earlier turn is not an approval now.
State the scope in the rules themselves: the classifier reads the list, it has
no separate setting for this.
## Security notes
- `Read(**/.env)` only blocks the Read tool. `Bash(cat .env)` bypasses it unless separately denied.
@@ -53,6 +126,8 @@ Always write the file-write ban as `Edit(...)`.
- `disableBypassPermissionsMode: "disable"` prevents switching to bypass mode mid-session.
- Prefer `ask` over `allow` for anything touching external systems.
- `deny` in `~/.claude/settings.json` cannot be overridden by project-level `allow` — deny always wins.
- Under `defaultMode: auto`, `ask` does not raise a prompt (see above). A destructive
command belongs in `deny` or in `autoMode.soft_deny`, not in `ask`.
## managed-settings.json (enterprise)
+4 -2
View File
@@ -15,8 +15,10 @@ REPO="$(cd "$(dirname "$0")" && pwd)"
VERSION=$(cat "$REPO/version.txt" 2>/dev/null || echo "unknown")
# Load shared detection library
# shellcheck source=lib/detect-plugins.sh
# shellcheck source=lib/detect-plugins.sh disable=SC1091
source "$REPO/lib/detect-plugins.sh"
# shellcheck source=lib/gstack-playwright.sh disable=SC1091
source "$REPO/lib/gstack-playwright.sh"
echo ""
echo "═══ claude-config update (v${VERSION}) ═══"
@@ -84,7 +86,7 @@ if [[ "$_gstack_confirm" =~ ^[Yy]$ ]]; then
_gstack_state=$(bash "$REPO/lib/toggle-external.sh" status gstack 2>/dev/null || echo "unknown")
fi
if git submodule update --remote skills-external/gstack 2>/dev/null; then
if gstack_submodule_update_with_bump "$REPO"; then
if [ -d "skills-external/gstack" ]; then
if [ -x "skills-external/gstack/setup" ]; then
if (cd skills-external/gstack && ./setup) 2>/dev/null; then