# Darwin run 2026-08-25/26: fresh baseline + threshold optimization + bug pass Branch `feature/darwin-optimize-20260825`, 26 commits, 39 files, +299/-142. Log: `~/.agents/skills/darwin-skill/results.tsv` (fresh, the May file was wiped by the 2026-06-23 reinstall). Method: darwin v2.1. Absolute scores served as triage only; every keep/revert decision came from a paired same-judge majority (3 judges per round, before/after read in one call). ## Scope 54 units: 31 personal skill-systems (SKILL.md + dispatched agents judged together, per EVAL-004) and 23 agents. Excluded: gstack/external symlinks (BDR-015/043, LRN-070), darwin-skill itself (BDR-058 pin), and find-docs, newly identified as machine-owned ctx7 output (gitignored, installer-written). ## Baseline (7 blind judges, dims scored 1-10, totals recomputed main-thread per LRN-018) Mean 83.4 (skills 83.5, agents 83.3). Best: deploy, release-candidate, release-executor (90.4). Worst: skills-perso 63.5. All dim8 rows marked dry_run by design; live execution happened later, inside the paired rounds. 13 units scored below the user-set threshold of 80. ## Phase 2: threshold loop, 13/13 units, 0 reverts Every round was validated by 3 paired judges (neutral, skeptic, realism). All verdicts 3-0 better. | Unit (baseline) | Round(s) | What changed | |---|---|---| | skills-perso (63.5) | d8 | Detection rebuilt on the link.sh convention: symlink = external, real dir = personal, gitignored = machine-generated. Live result 8/31 to 31/31, zero false positives | | interviewer (70.9) | d3, d9 | Failure-mode table (vague, "you decide", contradiction, partial, balloon) + 2-round budget; DO-NOT list | | onboarder (71.5) | d8 | BRIEF contract split REQUIRED/OPTIONAL; null enrichment becomes TODO placeholders; STOP kept for required keys and unresolved monorepo. Kills the guaranteed first-dispatch bounce vs /onboard STEP 2 | | pdf-translate (72.3) | d3/d8 | 8-row failure table: deps, >30 pages gate, zero-output, illisible markers, design-html/browse fallbacks, QA cap 2, stale workdir | | refactor (75.6) + refactorer (76.8) | d4/d3 | No-tests STOP gate + GO-WITHOUT-TESTS arbitration in the dispatcher; mid-run test-failure revert protocol; code-cleaner inline carve-out | | profile (77.3) | d3 | 6-row failure table, every row fact-checked against profile.sh (rc=1 paths, partial toggle, split plugin leg, BLK-006 contradiction); fixture de-drift | | plugin-probe (78.5) + plugin-advisor (77.5) | d8 | FRAMEWORK-DEPS now exact dep@version (preact false-hit killed, fallback actually fires; the old `\|\| true` silently emitted nothing and tripped the advisor's fail-closed path on non-Node projects); frontend/fast-libs derivable; PLAN echoed-or-unknown, invention removed | | analyze (77.7) + analyzer (78.0) | d1, d2 | Bilingual triggers + fix-wanted disambiguator; TASKS ordered, each step mapped to its OUTPUT section | | status-reporter (78.0) | d5 x2 | Fabrication-forcing token field replaced, then restored producibly from doctor.sh constants (a skeptic judge found the source); dead ROADMAP row rewritten post-ADR-013 | | gitflow (78.4) | d3 | 7-row failure table keyed to lib return codes; rc=4 conflict resume empirically verified; human merge gate untouched | ## Bug pass: verified defects in above-threshold units, 8 commits, all kept 3-0 - hotfix: `git restore .` on every failure branch wiped tolerated in-progress user edits. Now: `git stash create` pre-flight snapshot + file-scoped restore + fresh-dispatch-only security gate. Two skeptic residuals amended (RULES bullet, FILE(S) new-file marker). - init-project: allowed-tools lacked Agent and Skill while every step dispatches. commit-change: conflict grep now covers all 7 unmerged codes. tour: --report-only no longer commits (could land on develop). - harden: severity rule now defers to the calibrated guide; the late SSL Labs grade has an assigned actor. - plan-challenger: ERROR joined the load-bearing verdict grammar. - handover writers: stale chapter refs corrected (glossary/tone to §6, cross-links and THRESHOLD-OVERRIDE to §5); STEP 14.5 verification deferred post-write; anchor gate ordered into STEP 16. - security-auditor: /hotfix no-verifier carve-out documented. close: STEP 5C enumerated, --no-push passthrough added. - prune-memory: false "v1-untested" note replaced by the real tests/ state. code-clean: executor attribution corrected (code-cleaner, refactorer inline). - Fixtures de-drifted: plugin-check (PLUGIN CHECK block, real plugin names), onboard (nextjs-app-router). `make test` green (0 RED, rc=0) after one census rewrap: a locked phrase had been line-wrapped and the single-line grep lock caught it. ## Residual findings, logged not fixed - analyze triggers: "how does X work" brushes graphify's territory; graphify's graph-exists routing still wins. - pdf-translate: pdfinfo row assumes poppler (fitz also has page count); "GB" slightly overstated near the 30-page gate. - web-validate: .validate-cache mkdir lives in a skipped STEP 0 (self-recoverable); axis budgets 35/25/40 never reconciled with the base-100 deduction table. seo/geo minor wording items. verifier/doc-syncer/audit-delta restatement redundancy (cosmetic). handover-doc-writer STEP 14.5 umbrella line still says "BEFORE STEP 15" while the inner note overrides it. - bugfix trivial-fast-path boundary loosely defined; feat prompt-3 expectation vs full gate pipeline. ## Methodology notes - v2.1 paired majority produced 36 unit-round verdicts and 24 batch verdicts, all better, 0 reverts, 0 ties. The May-2026 run under absolute-delta scoring had reverted 2 edits on judge noise; this run had no such event. - Judges live-executed wherever the artifact was executable (skills-perso detection, profile.sh probes, plugin grep on scratch manifests, doctor.sh grep, git merge no-op resume). Behavior outranked prose in 5 units. - Two grep-exit-masking bugs surfaced (a `head` pipe swallowing the fallback's trigger), one in the probe being fixed, one in this run's own test harness. The pattern is worth a learning entry.