Files
claude/.claude/audits/DARWIN-2026-08-26.md
T

6.0 KiB

Darwin run 2026-08-25/26: fresh baseline + threshold optimization + bug pass

Branch feature/darwin-optimize-20260825, 26 commits, 39 files, +299/-142. Log: ~/.agents/skills/darwin-skill/results.tsv (fresh, the May file was wiped by the 2026-06-23 reinstall). Method: darwin v2.1. Absolute scores served as triage only; every keep/revert decision came from a paired same-judge majority (3 judges per round, before/after read in one call).

Scope

54 units: 31 personal skill-systems (SKILL.md + dispatched agents judged together, per EVAL-004) and 23 agents. Excluded: gstack/external symlinks (BDR-015/043, LRN-070), darwin-skill itself (BDR-058 pin), and find-docs, newly identified as machine-owned ctx7 output (gitignored, installer-written).

Baseline (7 blind judges, dims scored 1-10, totals recomputed main-thread per LRN-018)

Mean 83.4 (skills 83.5, agents 83.3). Best: deploy, release-candidate, release-executor (90.4). Worst: skills-perso 63.5. All dim8 rows marked dry_run by design; live execution happened later, inside the paired rounds. 13 units scored below the user-set threshold of 80.

Phase 2: threshold loop, 13/13 units, 0 reverts

Every round was validated by 3 paired judges (neutral, skeptic, realism). All verdicts 3-0 better.

Unit (baseline) Round(s) What changed
skills-perso (63.5) d8 Detection rebuilt on the link.sh convention: symlink = external, real dir = personal, gitignored = machine-generated. Live result 8/31 to 31/31, zero false positives
interviewer (70.9) d3, d9 Failure-mode table (vague, "you decide", contradiction, partial, balloon) + 2-round budget; DO-NOT list
onboarder (71.5) d8 BRIEF contract split REQUIRED/OPTIONAL; null enrichment becomes TODO placeholders; STOP kept for required keys and unresolved monorepo. Kills the guaranteed first-dispatch bounce vs /onboard STEP 2
pdf-translate (72.3) d3/d8 8-row failure table: deps, >30 pages gate, zero-output, illisible markers, design-html/browse fallbacks, QA cap 2, stale workdir
refactor (75.6) + refactorer (76.8) d4/d3 No-tests STOP gate + GO-WITHOUT-TESTS arbitration in the dispatcher; mid-run test-failure revert protocol; code-cleaner inline carve-out
profile (77.3) d3 6-row failure table, every row fact-checked against profile.sh (rc=1 paths, partial toggle, split plugin leg, BLK-006 contradiction); fixture de-drift
plugin-probe (78.5) + plugin-advisor (77.5) d8 FRAMEWORK-DEPS now exact dep@version (preact false-hit killed, fallback actually fires; the old || true silently emitted nothing and tripped the advisor's fail-closed path on non-Node projects); frontend/fast-libs derivable; PLAN echoed-or-unknown, invention removed
analyze (77.7) + analyzer (78.0) d1, d2 Bilingual triggers + fix-wanted disambiguator; TASKS ordered, each step mapped to its OUTPUT section
status-reporter (78.0) d5 x2 Fabrication-forcing token field replaced, then restored producibly from doctor.sh constants (a skeptic judge found the source); dead ROADMAP row rewritten post-ADR-013
gitflow (78.4) d3 7-row failure table keyed to lib return codes; rc=4 conflict resume empirically verified; human merge gate untouched

Bug pass: verified defects in above-threshold units, 8 commits, all kept 3-0

  • hotfix: git restore . on every failure branch wiped tolerated in-progress user edits. Now: git stash create pre-flight snapshot + file-scoped restore + fresh-dispatch-only security gate. Two skeptic residuals amended (RULES bullet, FILE(S) new-file marker).
  • init-project: allowed-tools lacked Agent and Skill while every step dispatches. commit-change: conflict grep now covers all 7 unmerged codes. tour: --report-only no longer commits (could land on develop).
  • harden: severity rule now defers to the calibrated guide; the late SSL Labs grade has an assigned actor.
  • plan-challenger: ERROR joined the load-bearing verdict grammar.
  • handover writers: stale chapter refs corrected (glossary/tone to §6, cross-links and THRESHOLD-OVERRIDE to §5); STEP 14.5 verification deferred post-write; anchor gate ordered into STEP 16.
  • security-auditor: /hotfix no-verifier carve-out documented. close: STEP 5C enumerated, --no-push passthrough added.
  • prune-memory: false "v1-untested" note replaced by the real tests/ state. code-clean: executor attribution corrected (code-cleaner, refactorer inline).
  • Fixtures de-drifted: plugin-check (PLUGIN CHECK block, real plugin names), onboard (nextjs-app-router).

make test green (0 RED, rc=0) after one census rewrap: a locked phrase had been line-wrapped and the single-line grep lock caught it.

Residual findings, logged not fixed

  • analyze triggers: "how does X work" brushes graphify's territory; graphify's graph-exists routing still wins.
  • pdf-translate: pdfinfo row assumes poppler (fitz also has page count); "GB" slightly overstated near the 30-page gate.
  • web-validate: .validate-cache mkdir lives in a skipped STEP 0 (self-recoverable); axis budgets 35/25/40 never reconciled with the base-100 deduction table. seo/geo minor wording items. verifier/doc-syncer/audit-delta restatement redundancy (cosmetic). handover-doc-writer STEP 14.5 umbrella line still says "BEFORE STEP 15" while the inner note overrides it.
  • bugfix trivial-fast-path boundary loosely defined; feat prompt-3 expectation vs full gate pipeline.

Methodology notes

  • v2.1 paired majority produced 36 unit-round verdicts and 24 batch verdicts, all better, 0 reverts, 0 ties. The May-2026 run under absolute-delta scoring had reverted 2 edits on judge noise; this run had no such event.
  • Judges live-executed wherever the artifact was executable (skills-perso detection, profile.sh probes, plugin grep on scratch manifests, doctor.sh grep, git merge no-op resume). Behavior outranked prose in 5 units.
  • Two grep-exit-masking bugs surfaced (a head pipe swallowing the fallback's trigger), one in the probe being fixed, one in this run's own test harness. The pattern is worth a learning entry.