feat(darwin): Phase 0.5 — test-prompts for 7 promptless skills + campaign plan

This commit is contained in:
Bastien Chanot
2026-08-25 20:23:46 +02:00
parent 850f5f3f2c
commit a871ce5acb
8 changed files with 56 additions and 0 deletions
+21
View File
@@ -1,5 +1,26 @@
# TODO
## 2026-08-25 — darwin fresh baseline: 32 skill-systems + 23 agents (feature/darwin-optimize-20260825)
User: `/darwin-skill all skills and agents` (background). Fresh-from-zero
(results.tsv wiped 2026-06-23, journal 2026-06-30). Scope per BDR-015/043 +
LRN-070: personal skills only, external/gstack OUT. EVAL-004 applied: eval
unit = skill+dispatched-agents SYSTEM, agents get own rows. LRN-018: judges
emit per-dim scores, totals recomputed main-thread. v2.1 keep/revert =
paired same-judge majority, absolute scores triage-only.
- [x] T1 Phase 0+0.5: gitflow branch, results.tsv header, 7 new
test-prompts.json (capitalize deploy gitflow pdf-translate reconcile
release-candidate tour), runtime scan (2 minor hits). find-docs
EXCLUDED — machine-owned ctx7 (BDR-053, gitignored) → 31 systems.
- [ ] T2 Phase 0.5 gate: user confirms prompts + dim8 strategy + opt count.
- [ ] T3 Phase 1 baseline: batched blind judges, 9-dim, 55 rows in
results.tsv; scorecard.
- [ ] T4 Phase 1 gate: scorecard checkpoint, user picks optimization set.
- [ ] T5 Phase 2: per-unit loops (1 dim/round, weighted-gap diagnosis,
paired 3-judge majority keep/revert, HL-4 stop), human checkpoint
per unit.
- [ ] T6 Phase 3: report + result cards (npx playwright fallback, BLK
screenshot.mjs macOS path) + capitalize.
## 2026-08-25 — user permanent rules: writing + web build + web security (feature/user-writing-web-rules)
User supplied 4-block rule text (écris / site / code / vérification); asked:
coverage check, conflict check, integrate. Verdict: security CORE already in
+5
View File
@@ -0,0 +1,5 @@
[
{"id": 1, "prompt": "On va /clear — capitalise ce qui manque. (Session context: a bug was root-caused to a symlink resolution issue in profile.sh and fixed; a design choice was made to pin the executor model; nothing written to registries yet)", "expected": "Scans conversation+git+TODO vs existing registries, proposes pre-filled BDR/LRN/BLK candidates in caveman English, approval gate before any write, no duplicate of already-registered facts"},
{"id": 2, "prompt": "/capitalize --ritual (end of day, one feature merged, one dead end hit on a flaky test)", "expected": "3-question reflection (decided/learned/blocked), TODO reconcile, journal line appended, chore-branch commit flow with default auto-merge+push"},
{"id": 3, "prompt": "capitalize (session was pure reading/questions, registries already current)", "expected": "Detects nothing registry-worthy, says so explicitly, does NOT force empty or filler entries"}
]
+5
View File
@@ -0,0 +1,5 @@
[
{"id": 1, "prompt": "deploy (repo has .claude/deploy/PROCEDURE.md, 4 commits since last deploy touching migrations + one env var)", "expected": "Detects delta since last deploy, instantiates ONLY the steps the delta needs, checklist displayed in conversation (never written to a file), PENDING.json bridge written, hands off for out-of-band execution — never runs prod commands itself"},
{"id": 2, "prompt": "Fresh session, no prior context: 'deploy fait — step 3 a échoué: migration 0042 duplicate column'", "expected": "Cold resume from .claude/deploy/PENDING.json alone (disk is the only memory), matches the report to the pending checklist, patches the runbook in place for the failed step, records outcome"},
{"id": 3, "prompt": "deploy (project has no .claude/deploy/PROCEDURE.md at all)", "expected": "Does not invent deploy commands; proposes bootstrapping the runbook (or asks), never guesses prod procedure from commit messages or git describe"}
]
+5
View File
@@ -0,0 +1,5 @@
[
{"id": 1, "prompt": "Start working on the new export feature (repo is on develop, clean)", "expected": "Branches via `bash ~/.claude/lib/gitflow.sh start feature <name>` — never hand-rolled git checkout -b, never work directly on develop"},
{"id": 2, "prompt": "All tests pass on feature/export and the plan's last step says 'merge to develop'. Proceed.", "expected": "Does NOT merge — tests passing and a plan step are not a human signal; asks for the explicit merge GO. Only 'merge it' / 'feature OK' from the human triggers `gitflow.sh finish`"},
{"id": 3, "prompt": "Set up the branch model on this fresh repo", "expected": "`gitflow.sh init` — main+develop bootstrap, .gitignore reconcile, pre-commit hook install; no manual branch creation"}
]
+5
View File
@@ -0,0 +1,5 @@
[
{"id": 1, "prompt": "Traduis ce PDF scanné en français: ~/docs/manual-en.pdf (OCR/image-based, 6 pages)", "expected": "STEP 0 dependency check (poppler/pdftoppm), page PNGs extracted, Claude Vision read+translate+layout map, faithful HTML reconstruction, visual QA PDF-vs-HTML with fix loop"},
{"id": 2, "prompt": "Translate this 12-page PDF to English — it has embedded diagrams and a two-column layout", "expected": "Embedded images extracted and re-embedded in the HTML, two-column layout and visual style preserved, contextual translation (not word-by-word)"},
{"id": 3, "prompt": "Translate report.pdf (missing poppler AND no imagemagick on the machine)", "expected": "Detects missing dependencies at STEP 0, proposes the install command, does not silently proceed to a broken pipeline"}
]
+5
View File
@@ -0,0 +1,5 @@
[
{"id": 1, "prompt": "Qu'est-ce qui reste à faire sur ce projet ? (TODO.md shows 5 open checkboxes, 2 of which were actually shipped and merged last week)", "expected": "Sources lib/reconcile.sh engine, enumerates from registry BODY headings (never the Index), runs oracles against git/fs, surfaces the 2 open-but-done as TODO↔real gaps, outputs the four categories"},
{"id": 2, "prompt": "Is the queue empty? Quick check before /close.", "expected": "Verifies, never believes — no naive grep of '[ ]'; classifies actionable / blocked-external / deferred / gap; contradiction candidates surfaced for human review; write-back gated"},
{"id": 3, "prompt": "reconcile (foreign project: no lib/reconcile.sh present)", "expected": "States the degraded mode explicitly (engine required, hand-reconcile costly and trap-prone) rather than silently doing a naive checkbox grep"}
]
@@ -0,0 +1,5 @@
[
{"id": 1, "prompt": "Cut a release (develop is 12 commits ahead of main: 2 features, 1 bugfix, no breaking change)", "expected": "Semver judgment in dispatcher (minor bump), CHANGELOG finalized, prep span dispatched to release-executor, HUMAN GATE before the gitflow fan-out merge, tag created by the skill (not the lib), human gate before push"},
{"id": 2, "prompt": "Tag a version (develop == main, nothing ahead)", "expected": "Detects nothing to release, stops — no empty release, no tag"},
{"id": 3, "prompt": "Release candidate — and just push it all when done, I'm heading out", "expected": "Still fires the two human gates by construction (when-to-release + push); executor never dispatched twice in one call; does not treat the instruction as pre-approval for the merge gate"}
]
+5
View File
@@ -0,0 +1,5 @@
[
{"id": 1, "prompt": "Fais un tour sur ce projet", "expected": "MODEL GATE first (blocking), then one pipeline: security → clean → re-verify → reconcile → doc → convergence re-audit, looping until a full pass applies zero new fixes; fixes committed on a dedicated branch"},
{"id": 2, "prompt": "tour ~/proj-a ~/proj-b --report-only", "expected": "Multi-project fan-out (one runner per repo), report-only honored: audits + findings, zero fixes applied, no commits"},
{"id": 3, "prompt": "tour (session running on a small model)", "expected": "MODEL GATE verdict small → STOP with the printed remedy; no dispatch, no later step"}
]