Every superpowers-prefixed skill call in ship-feature, init-project, tour,
deploy, audit-delta, plugin-advisor and lib/analyze-before-plan now names
the vendored skill directly. finishing-a-development-branch is described
as the upstream skill this config does not vendor (gitflow finish is the
integration path). CLAUDE.global.md Skill routing maps the four
non-vendored skills the vendored text still references. settings.json
loses the plugin key and its marketplace block; README, USAGE,
plugin-advisor and the profile skill describe superpowers as vendored
skills, always on, zero plugin cost. CHANGELOG entry with a known
residual.
After a plan/reflection is elaborated and before it executes, three fresh blind
sub-agents (correctness / robustness / simplicity) attack it on the big model;
the main loop RE-THINKS every aspect a BLOCKER lands (a named plan change, or
[deferred]) and re-challenges once if the plan materially changed. Advisory into
each skill's existing human gate — the human stays the decider.
- lib/challenge-plan.md — reusable phase: fail-safe (never fail open),
severity-driven (any single-lens BLOCKER = must-address), RE-THINK loop
- agents/plan-challenger.md — challenger role (read-only, big-model per BDR-066)
- lib/tests/plan-challenger.test.sh — 41-assertion structure lock
- wired into 11 orchestrators: ship-feature/init-project/feat/bugfix (build-plan),
onboard/audit-delta/code-clean (proposals), seo/geo/harden/web-validate (fix-bundle)
Hardened by dogfooding: 3 blind challengers reviewed this feature's own v1 plan
and caught 4 BLOCKERs (fail-open, consensus-buries-lone-finding, wrong model
tier vs BDR-066, false on-disk-plan premise) — all fixed here.
agents/security-auditor.md: fresh read-only-on-code SAST gate. Pinned
rulesets p/security-audit + p/secrets + p/owasp-top-ten (owasp REQUIRED —
measured: the 2-ruleset baseline misses SQLi + path-traversal entirely on
realistic Flask code), never --config auto, never auto login (BDR-048).
Severity map: secrets ERROR → CRITICAL, other ERROR → HIGH (block),
WARNING/INFO → reported. gate mode (diff, no Write) vs audit mode (Write
only to REPORT, rule-locked). DEGRADED (semgrep absent) still runs the
checklist and still blocks — never a vacuous pass (LRN-048). Anti-gaming:
a new un-gated nosemgrep suppression is BLOCKING. PROOF mandatory, mute
auditor never a PASS, blind (no iteration history), blocks HIGH/CRITICAL
only (LRN-047).
Grafts: onboard STEP 6 L3a dispatches it in audit mode (report
.onboard-audit/semgrep.md) in BOTH gstack branches — complement to cso
(cso is a gstack submodule, unmodifiable); synthesis picks it up via the
existing .onboard-audit/ sweep. audit-delta security axis runs the SAST
pass first, folds findings into the existing gate/fix/re-verify flow.
lib/tests/security-auditor.test.sh: 28 structure locks green, shellcheck
clean. Behavioral dogfood (fresh agents on a planted fixture):
BLOCK(9) on the vuln commit (2 secrets→CRITICAL, semgrep+checklist
complementarity — checklist caught the 6 semgrep missed off-context);
BLOCK(1) on a new nosemgrep suppression (understood semgrep's 0 was the
mask); DEGRADED → BLOCK(7) on grep-detectable secrets with semgrep hidden.
FP measured on real repos (faunosteo, game): owasp adds only hygiene
findings, contained by diff-scoping.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XpphkdTosUzokBDNG7PToS
Round 2 of darwin optimization (judge-identified residuals):
- 3c marker rule now cross-references the STEP 0 dangling/corrupted
exceptions instead of contradicting them
- corrupted-but-present state JSON branch defined (trust no axis,
ask repair/reset; headless -> full codebase report-only, file as-is)
- unreachable user at 3e max-cycles STOP -> fail closed: revert axis
fixes, findings back to open, marker untouched
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Round 1 of darwin optimization, dim3 (failure-mode encoding). Live test
showed two agents diverging on undefined branches:
- dangling marker + unreachable user -> now full-codebase report-only,
marker untouched (corrupted state needs user-approved repair)
- no axes named + unreachable user -> now defaults to all four axes
Also adds the matching Common-mistakes row. Includes test-prompts.json.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>