diff --git a/CHANGELOG.md b/CHANGELOG.md index 4e8e6ef..ab4a49b 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -33,6 +33,7 @@ Format follows [Keep a Changelog](https://keepachangelog.com/). ### Changed - **Versioning reset to 1.0.0** — see the note above; the pre-reset `v4.0.0` tag is retired and the lineage restarts here. +- graphify skill dist refreshed 0.8.45 → 0.9.6 (out-of-band `make plugin`; SKILL.md + query/extraction references updated by the generator). - `/deploy` NEXT.sh reshaped on first-real-run feedback: runbook steps are **one command per line, interactive-session style** (an early step opens the ssh session; later lines run on the box; local steps say "from your machine") instead of folded `ssh host "cd … && …"` one-liners, and the **hand-back prints the full checklist inline** in the conversation (also on every re-hand-back) so the user never has to open `NEXT.sh` to know what to run. Step = comment header + command lines up to the next blank line; a `@delta:` directive governs the whole block. Template `templates/deploy/PROCEDURE.md` restyled to match. - `settings.json`: the default model is pinned to **Opus 4.8 (1M context)** (`claude-opus-4-8[1m]`), and `inputNeededNotifEnabled: true` is adopted (harness notification toggle) — committed layout otherwise unchanged. - **AI attribution trailers disabled** (`attribution` in `settings.json`): `Co-Authored-By` + `Claude-Session` lines are no longer emitted on commits and PRs — they leaked session URLs and cluttered messages. diff --git a/skills/graphify/.graphify_version b/skills/graphify/.graphify_version index 827dae8..9cf0386 100644 --- a/skills/graphify/.graphify_version +++ b/skills/graphify/.graphify_version @@ -1 +1 @@ -0.8.45 \ No newline at end of file +0.9.6 \ No newline at end of file diff --git a/skills/graphify/SKILL.md b/skills/graphify/SKILL.md index 6c7060a..b354243 100644 --- a/skills/graphify/SKILL.md +++ b/skills/graphify/SKILL.md @@ -77,7 +77,7 @@ fi if [ -z "$PYTHON" ] && [ -n "$GRAPHIFY_BIN" ]; then _SHEBANG=$(head -1 "$GRAPHIFY_BIN" | tr -d '#!') case "$_SHEBANG" in - *[!a-zA-Z0-9/_.-]*) ;; + *[!a-zA-Z0-9/_.@-]*) ;; *) "$_SHEBANG" -c "import graphify" 2>/dev/null && PYTHON="$_SHEBANG" ;; esac fi @@ -151,12 +151,14 @@ Skip this step entirely if `detect` returned zero `video` files. When the corpus This step has two parts: **structural extraction** (deterministic, free) and **semantic extraction** (LLM, costs tokens). -**Before dispatching subagents:** check whether `GEMINI_API_KEY` or `GOOGLE_API_KEY` is set. If neither is set, print this one-liner to the user: +> **graphify needs no API key. Never ask the user for one, and never block on one.** Code is extracted structurally (AST) with no LLM and no key at all — a code-only corpus (the common `/graphify .` on a repo) skips semantic extraction entirely, so it needs nothing here: go straight to Part A and skip Part B. Semantic extraction (only for docs, papers, and images) uses Gemini **only if** `GEMINI_API_KEY`/`GOOGLE_API_KEY` is already set; otherwise the host agent itself is the LLM. graphify does **not** read `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, or any other provider key. If you catch yourself about to prompt for, wait on, or stop because of a missing API key, that is a misread of this skill — proceed without one. + +**Before semantic extraction:** check whether `GEMINI_API_KEY` or `GOOGLE_API_KEY` is set. If neither is set, print this one-liner to the user: > Tip: set `GEMINI_API_KEY` or `GOOGLE_API_KEY` to use Gemini for semantic extraction (`pip install 'graphifyy[gemini]'`). -Print it once, then continue. If `GEMINI_API_KEY` or `GOOGLE_API_KEY` IS set, use `graphify.llm.extract_corpus_parallel(files, backend="gemini")` for semantic extraction instead of dispatching Claude subagents. The default Gemini model is `gemini-3-flash-preview`; set `GRAPHIFY_GEMINI_MODEL` or pass `--model` in headless CLI flows to override it. +Print it once, then continue — do not wait for the user to supply a key. If `GEMINI_API_KEY` or `GOOGLE_API_KEY` IS set, use `graphify.llm.extract_corpus_parallel(files, backend="gemini")` for semantic extraction instead of dispatching subagents. The default Gemini model is `gemini-3-flash-preview`; set `GRAPHIFY_GEMINI_MODEL` or pass `--model` in headless CLI flows to override it. -> **No other API keys are read.** If `GEMINI_API_KEY`/`GOOGLE_API_KEY` are unset, fall straight through to Claude Code subagent dispatch (Part B below) — the host session itself is the LLM. graphify does **not** read `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, or any other provider key from the environment. If a host agent prompts the user for `ANTHROPIC_API_KEY` to run extraction, that prompt is a misread of this skill — ignore it and dispatch subagents as written. +> **No other API keys are read.** When `GEMINI_API_KEY`/`GOOGLE_API_KEY` are unset, semantic extraction falls to the host agent itself — the running session is the LLM. On a host that dispatches subagents (e.g. Claude Code), dispatch them as written in Part B. On a host that runs the CLI directly in a terminal and cannot dispatch subagents, do not stall: a code-only corpus has no semantic work, so write the empty semantic file (Part B "Fast path") and continue to Part C; for a corpus with docs/papers/images, either set a Gemini key or extract those inline yourself, but in no case prompt for `ANTHROPIC_API_KEY` — that prompt is a misread of this skill. **Run Part A (AST) and Part B (semantic) in parallel. Dispatch all semantic subagents AND start AST extraction in the same message. Both can run simultaneously since they operate on different file types. Merge results in Part C as before.** @@ -179,7 +181,7 @@ for f in detect.get('files', {}).get('code', []): code_files.extend(collect_files(Path(f)) if Path(f).is_dir() else [Path(f)]) if code_files: - result = extract(code_files, cache_root=Path('.')) + result = extract(code_files, cache_root=Path('INPUT_PATH')) Path('graphify-out/.graphify_ast.json').write_text(json.dumps(result, indent=2, ensure_ascii=False), encoding=\"utf-8\") print(f'AST: {len(result[\"nodes\"])} nodes, {len(result[\"edges\"])} edges') else: @@ -224,7 +226,7 @@ detect = json.loads(Path('graphify-out/.graphify_detect.json').read_text(encodin # every source file (#1392). Video is transcribed to a document in Step 2.5 first. all_files = [f for cat in ('document', 'paper', 'image') for f in detect['files'].get(cat, [])] -cached_nodes, cached_edges, cached_hyperedges, uncached = check_semantic_cache(all_files) +cached_nodes, cached_edges, cached_hyperedges, uncached = check_semantic_cache(all_files, root='INPUT_PATH') # Always (re)write the cache file: write hits, else DELETE any leftover from a prior # run so Part C never merges a stale .graphify_cached.json (#1392). @@ -311,7 +313,7 @@ from graphify.cache import save_semantic_cache from pathlib import Path new = json.loads(Path('graphify-out/.graphify_semantic_new.json').read_text(encoding=\"utf-8\")) if Path('graphify-out/.graphify_semantic_new.json').exists() else {'nodes':[],'edges':[],'hyperedges':[]} -saved = save_semantic_cache(new.get('nodes', []), new.get('edges', []), new.get('hyperedges', [])) +saved = save_semantic_cache(new.get('nodes', []), new.get('edges', []), new.get('hyperedges', []), root='INPUT_PATH') print(f'Cached {saved} files') " ``` @@ -445,6 +447,32 @@ If this step prints `ERROR: Graph is empty`, stop and tell the user what happene Replace INPUT_PATH with the actual path. +### Step 4.5 - Graph health check (read-only integrity gate) + +A non-destructive diagnostic on the extraction, before labeling. It surfaces edge collapse, dangling/missing endpoints, and self-loops — the silent-corruption modes of incremental updates and AST/LLM id mismatches. Read-only; never aborts. + +```bash +$(cat graphify-out/.graphify_python) -c " +import json +from pathlib import Path +from graphify.diagnostics import diagnose_extraction, format_diagnostic_report + +extraction = json.loads(Path('graphify-out/.graphify_extract.json').read_text(encoding=\"utf-8\")) +summary = diagnose_extraction(extraction, directed=IS_DIRECTED, root='INPUT_PATH') +print(format_diagnostic_report(summary)) +flags = [f'{summary[k]} {label}' for k, label in ( + ('dangling_endpoint_edges', 'dangling-endpoint edges'), + ('missing_endpoint_edges', 'missing-endpoint edges'), + ('self_loop_edges', 'self-loop edges'), + ('directed_same_endpoint_collapsed_edges', 'collapsed (directed) edges'), + ('undirected_same_endpoint_collapsed_edges', 'collapsed (undirected) edges'), +) if summary.get(k, 0)] +print('GRAPH HEALTH WARNING: ' + '; '.join(flags) + ' - graph may be incomplete/corrupt.' if flags else 'Graph health: OK (no dangling/missing/collapsed edges).') +" +``` + +Substitute `IS_DIRECTED` and `INPUT_PATH` as in Step 4. If a `GRAPH HEALTH WARNING` prints, surface it in the final summary (do not abort — the graph is still usable, but the integrity issue must be visible, per the Honesty Rules). + ### Step 5 - Label communities Read `graphify-out/.graphify_analysis.json`. For each community key, look at its node labels and write a 2-5 word plain-language name (e.g. "Attention Mechanism", "Training Pipeline", "Data Loading"). @@ -601,7 +629,7 @@ if [ ! -f graphify-out/.graphify_python ]; then GRAPHIFY_BIN=$(which graphify 2>/dev/null) if [ -n "$GRAPHIFY_BIN" ]; then PYTHON=$(head -1 "$GRAPHIFY_BIN" | tr -d '#!') - case "$PYTHON" in *[!a-zA-Z0-9/_.-]*) PYTHON="python3" ;; esac + case "$PYTHON" in *[!a-zA-Z0-9/_.@-]*) PYTHON="python3" ;; esac else PYTHON="python3" fi diff --git a/skills/graphify/references/extraction-spec.md b/skills/graphify/references/extraction-spec.md index 2cc1919..388df76 100644 --- a/skills/graphify/references/extraction-spec.md +++ b/skills/graphify/references/extraction-spec.md @@ -58,7 +58,7 @@ confidence_score is REQUIRED on every edge - never omit it, never use 0.5 as a d the edge AMBIGUOUS rather than picking 0.4 or below. - AMBIGUOUS edges: 0.1-0.3 -Node ID format: lowercase, only `[a-z0-9_]`, no dots or slashes. Format: `{stem}_{entity}` where stem is `{parent_dir}_{filename_without_ext}` (the **immediate** parent directory name + the filename stem, both lowercased with non-alphanumeric chars replaced by `_`) and entity is the symbol name similarly normalized. Only one level of parent is used — not the full path. Examples: `src/auth/session.py` + `ValidateToken` → `auth_session_validatetoken`; `lib/utils/helpers.py` + `parse_url` → `utils_helpers_parse_url`; `tests/test_foo.py` + `_helper` → `tests_test_foo_helper`. Top-level files (no parent dir, e.g. `setup.py`) use just the filename stem: `setup_my_func`. This must match the ID the AST extractor generates — using just the filename (e.g., `session_validatetoken`) or the full path (e.g., `src_auth_session_validatetoken`) will create orphan ghost-duplicate nodes. If you are re-extracting a project that had ghost duplicates under the old format, the user should run `graphify extract --force` to rebuild cleanly. CRITICAL: never append chunk numbers, sequence numbers, or any suffix to an ID (no `_c1`, `_c2`, `_chunk2`, etc.). IDs must be deterministic from the label alone — the same entity must always produce the same ID regardless of which chunk processes it. +Node ID format: lowercase, only `[a-z0-9_]`, no dots or slashes. Format: `{stem}_{entity}` where stem is the **full repo-relative path with the extension dropped**, every path segment kept and joined with `_` (each segment lowercased with non-alphanumeric chars replaced by `_`), and entity is the symbol name similarly normalized. Use every directory level, not just the immediate parent — this keeps same-named files in different directories distinct. Examples: `src/auth/session.py` + `ValidateToken` → `src_auth_session_validatetoken`; `lib/utils/helpers.py` + `parse_url` → `lib_utils_helpers_parse_url`; `tests/test_foo.py` + `_helper` → `tests_test_foo_helper`; `docs/v1/api/README.md` + `getUser` → `docs_v1_api_readme_getuser`. Top-level files (no parent dir, e.g. `setup.py`) use just the filename stem: `setup_my_func`. This must match the ID the AST extractor generates — using just the filename (e.g., `session_validatetoken`) or only the immediate parent (e.g., `auth_session_validatetoken`) will create orphan ghost-duplicate nodes. If you are re-extracting a project built under the old immediate-parent format, the user should run `graphify extract --force` to rebuild cleanly. CRITICAL: never append chunk numbers, sequence numbers, or any suffix to an ID (no `_c1`, `_c2`, `_chunk2`, etc.). IDs must be deterministic from the label alone — the same entity must always produce the same ID regardless of which chunk processes it. Generate the extraction JSON matching this schema exactly: {"nodes":[{"id":"auth_session_validatetoken","label":"Human Readable Name","file_type":"code|document|paper|image|rationale|concept","source_file":"","source_location":null,"source_url":null,"captured_at":null,"author":null,"contributor":null}],"edges":[{"source":"node_id","target":"node_id","relation":"calls|implements|references|cites|conceptually_related_to|shares_data_with|semantically_similar_to|rationale_for","confidence":"EXTRACTED|INFERRED|AMBIGUOUS","confidence_score":1.0,"source_file":"","source_location":null,"weight":1.0}],"hyperedges":[{"id":"snake_case_id","label":"Human Readable Label","nodes":["node_id1","node_id2","node_id3"],"relation":"participate_in|implement|form","confidence":"EXTRACTED|INFERRED","confidence_score":0.75,"source_file":""}],"input_tokens":0,"output_tokens":0} diff --git a/skills/graphify/references/query.md b/skills/graphify/references/query.md index 3ed5f65..56565eb 100644 --- a/skills/graphify/references/query.md +++ b/skills/graphify/references/query.md @@ -31,7 +31,7 @@ Fix this **without inventing tokens** by expanding the query against the actual $(cat graphify-out/.graphify_python) -c " import json, re from pathlib import Path -data = json.loads(Path('graphify-out/graph.json').read_text()) +data = json.loads(Path('graphify-out/graph.json').read_text(encoding='utf-8')) vocab = set() for n in data['nodes']: for c in re.findall(r'[^\W\d_]+', n.get('label','') or '', re.UNICODE): @@ -40,7 +40,7 @@ for n in data['nodes']: t = p.lower() if 3 <= len(t) <= 30: vocab.add(t) -Path('graphify-out/.vocab.txt').write_text('\n'.join(sorted(vocab))) +Path('graphify-out/.vocab.txt').write_text('\n'.join(sorted(vocab)), encoding='utf-8') print(f'vocab: {len(vocab)} tokens') " ``` @@ -83,7 +83,7 @@ from networkx.readwrite import json_graph import networkx as nx from pathlib import Path -data = json.loads(Path('graphify-out/graph.json').read_text()) +data = json.loads(Path('graphify-out/graph.json').read_text(encoding='utf-8')) G = json_graph.node_link_graph(data, edges='links') question = 'QUESTION' @@ -173,6 +173,14 @@ $(cat graphify-out/.graphify_python) -m graphify save-result --question "ORIGINA Replace `ORIGINAL_QUESTION` with the user's verbatim question, `ANSWER` with your full answer text (containing the expanded-token trace), `NODE1 NODE2` with the list of node labels you cited. This closes the feedback loop: the next `--update` will extract this Q&A as a node in the graph. +**Work memory (self-improving loop).** Add an `--outcome` so future sessions learn from this one — append `--outcome useful|dead_end|corrected` to the `save-result` command (and `--correction "the right answer"` when correcting): + +- `useful` — the cited nodes answered the question well (they become *preferred sources*). +- `dead_end` — the question/path led nowhere; don't re-derive it next time. +- `corrected` — the saved answer was wrong; `--correction` records what was right. + +At the **start** of graph work, refresh and read the lessons: run `graphify reflect --if-stale` (cheap, deterministic, no LLM; `--if-stale` makes it a no-op when `LESSONS.md` is already newer than every input, e.g. when the git hook just refreshed it), then read `graphify-out/reflections/LESSONS.md`. It lists **preferred sources** (start there), **known dead ends** (skip them), and prior **corrections**. Running `reflect` yourself keeps the lessons current even without the git hook installed; if the post-commit hook *is* installed, `--if-stale` means your session-start run costs almost nothing. + --- ## For /graphify path @@ -192,7 +200,7 @@ import networkx as nx from networkx.readwrite import json_graph from pathlib import Path -data = json.loads(Path('graphify-out/graph.json').read_text()) +data = json.loads(Path('graphify-out/graph.json').read_text(encoding='utf-8')) G = json_graph.node_link_graph(data, edges='links') a_term = 'NODE_A' @@ -260,7 +268,7 @@ import networkx as nx from networkx.readwrite import json_graph from pathlib import Path -data = json.loads(Path('graphify-out/graph.json').read_text()) +data = json.loads(Path('graphify-out/graph.json').read_text(encoding='utf-8')) G = json_graph.node_link_graph(data, edges='links') term = 'NODE_NAME'