Commit Graph
13 Commits
Author SHA1 Message Date
bastien 58c3a3e9b7 fix(effort): re-raise judgment dispatches, planning re-asserts, pairing caveat, dedupe audit script (final review I1-I3) 2026-09-28 20:48:27 +02:00
bastien 98ef991958 docs(effort): BDR-107 id, CHANGELOG entry, spec corrected for the rulings (vendored pins, exclusions, pairing rule) 2026-09-28 20:20:05 +02:00
bastien a117e7ed67 fix(effort): shifts are sent with the step's first tool call (harness pairing rule); challenge shifts under their heading; fence indentation 2026-09-28 19:55:47 +02:00
bastien 3c58160d0c feat(effort): wire phase shifts in the 13 orchestrators and the handover writer 2026-09-28 19:44:22 +02:00
bastien 94189adbc6 feat(effort): entry effort level on the 31 user-invoked skills (spec D3)
A/B /reconcile headless — BEFORE requests=18 output=12374 thinking=3135 effort={'high'} duration_ms=96518 / AFTER requests=15 output=9038 thinking=2248 effort={'low'} duration_ms=78410
2026-09-28 19:11:21 +02:00
bastien ddea411491 chore(config): superpowers citers by bare name, routing map, docs, settings
Every superpowers-prefixed skill call in ship-feature, init-project, tour,
deploy, audit-delta, plugin-advisor and lib/analyze-before-plan now names
the vendored skill directly. finishing-a-development-branch is described
as the upstream skill this config does not vendor (gitflow finish is the
integration path). CLAUDE.global.md Skill routing maps the four
non-vendored skills the vendored text still references. settings.json
loses the plugin key and its marketplace block; README, USAGE,
plugin-advisor and the profile skill describe superpowers as vendored
skills, always on, zero plugin cost. CHANGELOG entry with a known
residual.
2026-09-28 14:54:53 +02:00
Bastien Chanot 6bfc0543e5 feat(skills): add 3-way adversarial plan-challenge phase to reflection orchestrators
After a plan/reflection is elaborated and before it executes, three fresh blind
sub-agents (correctness / robustness / simplicity) attack it on the big model;
the main loop RE-THINKS every aspect a BLOCKER lands (a named plan change, or
[deferred]) and re-challenges once if the plan materially changed. Advisory into
each skill's existing human gate — the human stays the decider.

- lib/challenge-plan.md — reusable phase: fail-safe (never fail open),
  severity-driven (any single-lens BLOCKER = must-address), RE-THINK loop
- agents/plan-challenger.md — challenger role (read-only, big-model per BDR-066)
- lib/tests/plan-challenger.test.sh — 41-assertion structure lock
- wired into 11 orchestrators: ship-feature/init-project/feat/bugfix (build-plan),
  onboard/audit-delta/code-clean (proposals), seo/geo/harden/web-validate (fix-bundle)

Hardened by dogfooding: 3 blind challengers reviewed this feature's own v1 plan
and caught 4 BLOCKERs (fail-open, consensus-buries-lone-finding, wrong model
tier vs BDR-066, false on-disk-plan premise) — all fixed here.
2026-07-17 22:51:50 +02:00
Bastien Chanot 06413d9eb5 feat(model-routing): wire blocking model gate into 12 reflection orchestrators 2026-07-15 11:11:29 +02:00
Bastien ChanotandClaude Opus 4.8 2b297bd44a feat(agents): security-auditor SAST gate + onboard/audit-delta grafts (verify-loops lot 3)
agents/security-auditor.md: fresh read-only-on-code SAST gate. Pinned
rulesets p/security-audit + p/secrets + p/owasp-top-ten (owasp REQUIRED —
measured: the 2-ruleset baseline misses SQLi + path-traversal entirely on
realistic Flask code), never --config auto, never auto login (BDR-048).
Severity map: secrets ERROR → CRITICAL, other ERROR → HIGH (block),
WARNING/INFO → reported. gate mode (diff, no Write) vs audit mode (Write
only to REPORT, rule-locked). DEGRADED (semgrep absent) still runs the
checklist and still blocks — never a vacuous pass (LRN-048). Anti-gaming:
a new un-gated nosemgrep suppression is BLOCKING. PROOF mandatory, mute
auditor never a PASS, blind (no iteration history), blocks HIGH/CRITICAL
only (LRN-047).

Grafts: onboard STEP 6 L3a dispatches it in audit mode (report
.onboard-audit/semgrep.md) in BOTH gstack branches — complement to cso
(cso is a gstack submodule, unmodifiable); synthesis picks it up via the
existing .onboard-audit/ sweep. audit-delta security axis runs the SAST
pass first, folds findings into the existing gate/fix/re-verify flow.

lib/tests/security-auditor.test.sh: 28 structure locks green, shellcheck
clean. Behavioral dogfood (fresh agents on a planted fixture):
BLOCK(9) on the vuln commit (2 secrets→CRITICAL, semgrep+checklist
complementarity — checklist caught the 6 semgrep missed off-context);
BLOCK(1) on a new nosemgrep suppression (understood semgrep's 0 was the
mask); DEGRADED → BLOCK(7) on grep-detectable secrets with semgrep hidden.
FP measured on real repos (faunosteo, game): owasp adds only hygiene
findings, contained by diff-scoping.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XpphkdTosUzokBDNG7PToS
2026-07-03 19:13:02 +02:00
Bastien ChanotandClaude Fable 5 1aa9afe669 feat(tokens): compress the 10 fattest personal skill descriptions
6,416 → 4,243 chars (≈ −540 tokens/session, catalog loaded every
session). Kept per BDR-014/LRN-043: 'Use when' pattern, discriminating
FR+EN triggers, all 'For X → /Y' disambiguation lines. Cut: redundant
trigger synonyms, header/engine enumerations, prose the model derives.
find-docs excluded (ctx7-owned, regen clobbers — LRN-086); doc + geo
taken instead. All ≤ ~505 chars body (BDR-014 aspirational ceiling).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016zA3Qh2Q1QpcGXzXxKeDHR
2026-07-02 14:31:38 +02:00
Bastien ChanotandClaude Fable 5 9fc93fabd2 optimize audit-delta: close 3 judge-flagged consistency gaps
Round 2 of darwin optimization (judge-identified residuals):
- 3c marker rule now cross-references the STEP 0 dangling/corrupted
  exceptions instead of contradicting them
- corrupted-but-present state JSON branch defined (trust no axis,
  ask repair/reset; headless -> full codebase report-only, file as-is)
- unreachable user at 3e max-cycles STOP -> fail closed: revert axis
  fixes, findings back to open, marker untouched

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-11 17:54:35 +02:00
Bastien ChanotandClaude Fable 5 0d2ece757e optimize audit-delta: define unreachable-user branches (dangling marker, axis default)
Round 1 of darwin optimization, dim3 (failure-mode encoding). Live test
showed two agents diverging on undefined branches:
- dangling marker + unreachable user -> now full-codebase report-only,
  marker untouched (corrupted state needs user-approved repair)
- no axes named + unreachable user -> now defaults to all four axes
Also adds the matching Common-mistakes row. Includes test-prompts.json.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-11 17:48:41 +02:00
Bastien ChanotandClaude Fable 5 e12f8243e5 baseline: add audit-delta skill (pre-darwin-optimization snapshot)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-11 15:41:49 +02:00