6.0 KiB
Darwin run 2026-08-25/26: fresh baseline + threshold optimization + bug pass
Branch feature/darwin-optimize-20260825, 26 commits, 39 files, +299/-142.
Log: ~/.agents/skills/darwin-skill/results.tsv (fresh, the May file was wiped
by the 2026-06-23 reinstall). Method: darwin v2.1. Absolute scores served as
triage only; every keep/revert decision came from a paired same-judge majority
(3 judges per round, before/after read in one call).
Scope
54 units: 31 personal skill-systems (SKILL.md + dispatched agents judged together, per EVAL-004) and 23 agents. Excluded: gstack/external symlinks (BDR-015/043, LRN-070), darwin-skill itself (BDR-058 pin), and find-docs, newly identified as machine-owned ctx7 output (gitignored, installer-written).
Baseline (7 blind judges, dims scored 1-10, totals recomputed main-thread per LRN-018)
Mean 83.4 (skills 83.5, agents 83.3). Best: deploy, release-candidate, release-executor (90.4). Worst: skills-perso 63.5. All dim8 rows marked dry_run by design; live execution happened later, inside the paired rounds. 13 units scored below the user-set threshold of 80.
Phase 2: threshold loop, 13/13 units, 0 reverts
Every round was validated by 3 paired judges (neutral, skeptic, realism). All verdicts 3-0 better.
| Unit (baseline) | Round(s) | What changed |
|---|---|---|
| skills-perso (63.5) | d8 | Detection rebuilt on the link.sh convention: symlink = external, real dir = personal, gitignored = machine-generated. Live result 8/31 to 31/31, zero false positives |
| interviewer (70.9) | d3, d9 | Failure-mode table (vague, "you decide", contradiction, partial, balloon) + 2-round budget; DO-NOT list |
| onboarder (71.5) | d8 | BRIEF contract split REQUIRED/OPTIONAL; null enrichment becomes TODO placeholders; STOP kept for required keys and unresolved monorepo. Kills the guaranteed first-dispatch bounce vs /onboard STEP 2 |
| pdf-translate (72.3) | d3/d8 | 8-row failure table: deps, >30 pages gate, zero-output, illisible markers, design-html/browse fallbacks, QA cap 2, stale workdir |
| refactor (75.6) + refactorer (76.8) | d4/d3 | No-tests STOP gate + GO-WITHOUT-TESTS arbitration in the dispatcher; mid-run test-failure revert protocol; code-cleaner inline carve-out |
| profile (77.3) | d3 | 6-row failure table, every row fact-checked against profile.sh (rc=1 paths, partial toggle, split plugin leg, BLK-006 contradiction); fixture de-drift |
| plugin-probe (78.5) + plugin-advisor (77.5) | d8 | FRAMEWORK-DEPS now exact dep@version (preact false-hit killed, fallback actually fires; the old || true silently emitted nothing and tripped the advisor's fail-closed path on non-Node projects); frontend/fast-libs derivable; PLAN echoed-or-unknown, invention removed |
| analyze (77.7) + analyzer (78.0) | d1, d2 | Bilingual triggers + fix-wanted disambiguator; TASKS ordered, each step mapped to its OUTPUT section |
| status-reporter (78.0) | d5 x2 | Fabrication-forcing token field replaced, then restored producibly from doctor.sh constants (a skeptic judge found the source); dead ROADMAP row rewritten post-ADR-013 |
| gitflow (78.4) | d3 | 7-row failure table keyed to lib return codes; rc=4 conflict resume empirically verified; human merge gate untouched |
Bug pass: verified defects in above-threshold units, 8 commits, all kept 3-0
- hotfix:
git restore .on every failure branch wiped tolerated in-progress user edits. Now:git stash createpre-flight snapshot + file-scoped restore + fresh-dispatch-only security gate. Two skeptic residuals amended (RULES bullet, FILE(S) new-file marker). - init-project: allowed-tools lacked Agent and Skill while every step dispatches. commit-change: conflict grep now covers all 7 unmerged codes. tour: --report-only no longer commits (could land on develop).
- harden: severity rule now defers to the calibrated guide; the late SSL Labs grade has an assigned actor.
- plan-challenger: ERROR joined the load-bearing verdict grammar.
- handover writers: stale chapter refs corrected (glossary/tone to §6, cross-links and THRESHOLD-OVERRIDE to §5); STEP 14.5 verification deferred post-write; anchor gate ordered into STEP 16.
- security-auditor: /hotfix no-verifier carve-out documented. close: STEP 5C enumerated, --no-push passthrough added.
- prune-memory: false "v1-untested" note replaced by the real tests/ state. code-clean: executor attribution corrected (code-cleaner, refactorer inline).
- Fixtures de-drifted: plugin-check (PLUGIN CHECK block, real plugin names), onboard (nextjs-app-router).
make test green (0 RED, rc=0) after one census rewrap: a locked phrase had
been line-wrapped and the single-line grep lock caught it.
Residual findings, logged not fixed
- analyze triggers: "how does X work" brushes graphify's territory; graphify's graph-exists routing still wins.
- pdf-translate: pdfinfo row assumes poppler (fitz also has page count); "GB" slightly overstated near the 30-page gate.
- web-validate: .validate-cache mkdir lives in a skipped STEP 0 (self-recoverable); axis budgets 35/25/40 never reconciled with the base-100 deduction table. seo/geo minor wording items. verifier/doc-syncer/audit-delta restatement redundancy (cosmetic). handover-doc-writer STEP 14.5 umbrella line still says "BEFORE STEP 15" while the inner note overrides it.
- bugfix trivial-fast-path boundary loosely defined; feat prompt-3 expectation vs full gate pipeline.
Methodology notes
- v2.1 paired majority produced 36 unit-round verdicts and 24 batch verdicts, all better, 0 reverts, 0 ties. The May-2026 run under absolute-delta scoring had reverted 2 edits on judge noise; this run had no such event.
- Judges live-executed wherever the artifact was executable (skills-perso detection, profile.sh probes, plugin grep on scratch manifests, doctor.sh grep, git merge no-op resume). Behavior outranked prose in 5 units.
- Two grep-exit-masking bugs surfaced (a
headpipe swallowing the fallback's trigger), one in the probe being fixed, one in this run's own test harness. The pattern is worth a learning entry.