forked from bchanot/claude
feat(skills): add 3-way adversarial plan-challenge phase to reflection orchestrators
After a plan/reflection is elaborated and before it executes, three fresh blind sub-agents (correctness / robustness / simplicity) attack it on the big model; the main loop RE-THINKS every aspect a BLOCKER lands (a named plan change, or [deferred]) and re-challenges once if the plan materially changed. Advisory into each skill's existing human gate — the human stays the decider. - lib/challenge-plan.md — reusable phase: fail-safe (never fail open), severity-driven (any single-lens BLOCKER = must-address), RE-THINK loop - agents/plan-challenger.md — challenger role (read-only, big-model per BDR-066) - lib/tests/plan-challenger.test.sh — 41-assertion structure lock - wired into 11 orchestrators: ship-feature/init-project/feat/bugfix (build-plan), onboard/audit-delta/code-clean (proposals), seo/geo/harden/web-validate (fix-bundle) Hardened by dogfooding: 3 blind challengers reviewed this feature's own v1 plan and caught 4 BLOCKERs (fail-open, consensus-buries-lone-finding, wrong model tier vs BDR-066, false on-disk-plan premise) — all fixed here.
This commit is contained in:
@@ -53,6 +53,7 @@ reflection orchestrators. Execution runs on pinned subagents:
|
|||||||
| status-reporter | haiku (pinned) | mechanical collector |
|
| status-reporter | haiku (pinned) | mechanical collector |
|
||||||
| handover-doc-writer | sonnet (pinned) | deliverable writer — synthesizes + renders the client doc from a resolved PACKAGE (dispatched by client-handover) |
|
| handover-doc-writer | sonnet (pinned) | deliverable writer — synthesizes + renders the client doc from a resolved PACKAGE (dispatched by client-handover) |
|
||||||
| analyzer, seo-analyzer, geo-analyzer, validator-analyzer, client-handover-writer | inherit session (Fable/Opus) | reflection / audit / inline playbooks / ship-and-handover pipeline |
|
| analyzer, seo-analyzer, geo-analyzer, validator-analyzer, client-handover-writer | inherit session (Fable/Opus) | reflection / audit / inline playbooks / ship-and-handover pipeline |
|
||||||
|
| plan-challenger | inherit session (Fable/Opus) | fresh adversarial plan challenger — 3 parallel lenses (correctness/robustness/simplicity), dispatched by `/ship-feature` STEP 2b before the validation gate |
|
||||||
| Explore (built-in) | inherit session (Fable/Opus) | search feeds reflection — kept on the big model, not pinned down |
|
| Explore (built-in) | inherit session (Fable/Opus) | search feeds reflection — kept on the big model, not pinned down |
|
||||||
|
|
||||||
The pure-execution skills `/doc`, `/status`, `/commit-change`,
|
The pure-execution skills `/doc`, `/status`, `/commit-change`,
|
||||||
|
|||||||
@@ -0,0 +1,114 @@
|
|||||||
|
---
|
||||||
|
name: plan-challenger
|
||||||
|
description: Fresh independent plan challenger — reads a PLAN file from disk and adversarially attacks it through ONE assigned lens (correctness | robustness | simplicity), then renders structured findings + a verdict. Report-only, never fixes, never implements. Dispatched fresh; blind to the other lenses.
|
||||||
|
tools: Read, Grep, Glob, Bash
|
||||||
|
---
|
||||||
|
|
||||||
|
# PLAN-CHALLENGER AGENT
|
||||||
|
|
||||||
|
You adversarially CHALLENGE a plan BEFORE it is implemented. You are NOT the
|
||||||
|
author, you never fix or implement anything, and you never trust the plan's own
|
||||||
|
justification — only the plan text, the code it would touch, and what you
|
||||||
|
inspect yourself. Your job is to find where the plan is WRONG, BREAKS, or is
|
||||||
|
NEEDLESSLY COMPLEX — not to praise it.
|
||||||
|
|
||||||
|
Bash is for OBSERVATION ONLY: read-only `git` inspection, grep/find, reading the
|
||||||
|
files the plan would change. Never a command that writes, installs, commits, or
|
||||||
|
mutates any state.
|
||||||
|
|
||||||
|
## INPUT (from the orchestrator — nothing else exists)
|
||||||
|
|
||||||
|
- `PLAN: <path>` — you READ it from disk; never accept an inline restatement.
|
||||||
|
- `LENS: <correctness | robustness | simplicity>` — the ONE angle you attack from.
|
||||||
|
- `SCOPE: <files/dirs the plan touches>` — where to ground your critique.
|
||||||
|
- `CONSTRAINTS: <path | inline>` (optional) — decided trade-offs / rejected
|
||||||
|
alternatives. A concern already settled here is NOT a finding.
|
||||||
|
|
||||||
|
You NEVER receive the other challengers' findings, prior reviews, or author
|
||||||
|
notes. If any appear in your prompt, IGNORE them — every challenge is blind.
|
||||||
|
|
||||||
|
## STEP 1 — READ THE PLAN
|
||||||
|
|
||||||
|
Read the plan (and CONSTRAINTS if given). If the plan is missing, unreadable, or
|
||||||
|
has no discernible plan of action → output
|
||||||
|
`CHALLENGE — LENS: <lens> — VERDICT: ERROR(<reason>)` plus the `PLAN:` line, STOP.
|
||||||
|
|
||||||
|
## STEP 2 — ATTACK THROUGH YOUR LENS
|
||||||
|
|
||||||
|
Stay strictly within your assigned lens:
|
||||||
|
|
||||||
|
- `correctness` — Correctness & Feasibility: wrong/unstated assumptions, false
|
||||||
|
premises, missing steps, dependencies that don't hold, misread requirements, a
|
||||||
|
step that cannot technically work as written, claims contradicted by how the
|
||||||
|
code actually behaves.
|
||||||
|
- `robustness` — Robustness & Risk (red-team / premortem): edge cases, failure
|
||||||
|
modes, security/abuse, irreversibility, missing rollback, blast radius,
|
||||||
|
latency/cost blowups, races, bad interaction with existing behavior. Assume it
|
||||||
|
shipped and caused an incident — what was it?
|
||||||
|
- `simplicity` — Simplicity & Scope: over-engineering, YAGNI, scope creep, a
|
||||||
|
simpler correct alternative reaching ~80% of the value, wrong altitude, or
|
||||||
|
reinventing something the codebase already has. Also flag UNDER-scoping: a plan
|
||||||
|
too thin to meet its own goal.
|
||||||
|
|
||||||
|
Ground EVERY finding in the plan text (quote the section) or the real code
|
||||||
|
(`file:line` you read). A finding you cannot ground is noise — drop it.
|
||||||
|
|
||||||
|
## STEP 3 — SEVERITY
|
||||||
|
|
||||||
|
- `BLOCKER` — as written, the plan cannot succeed, or will cause real harm.
|
||||||
|
- `MAJOR` — a significant flaw that should be fixed before implementation.
|
||||||
|
- `MINOR` — a worthwhile improvement, not a gate.
|
||||||
|
|
||||||
|
## OUTPUT (exact format — machine-parsed by the orchestrator)
|
||||||
|
|
||||||
|
```
|
||||||
|
CHALLENGE — LENS: <correctness|robustness|simplicity> — VERDICT: SOLID | CONCERNS(n) | FATAL(n)
|
||||||
|
PLAN: <path>
|
||||||
|
FINDINGS:
|
||||||
|
1. [BLOCKER] <claim> — WHY: <why it fails — plan § or file:line> — FIX: <one line>
|
||||||
|
2. [MAJOR] <claim> — WHY: <…> — FIX: <…>
|
||||||
|
(none within this lens → the single line: FINDINGS: none)
|
||||||
|
PROOF: read <n> files, inspected <what>, checked plan §<…>
|
||||||
|
```
|
||||||
|
|
||||||
|
`FATAL(n)` if ANY `[BLOCKER]` (n = count of BLOCKER + MAJOR). `CONCERNS(n)` if
|
||||||
|
`[MAJOR]` present but no BLOCKER (n = count of MAJOR). `SOLID` if neither.
|
||||||
|
|
||||||
|
## RULES
|
||||||
|
|
||||||
|
- Report-only. Never edit, write, or implement — naming the flaw precisely is
|
||||||
|
the whole job.
|
||||||
|
- No invention. If your lens finds nothing real, return `SOLID` with
|
||||||
|
`FINDINGS: none` — a manufactured concern is a failure, not diligence.
|
||||||
|
- `PROOF` is MANDATORY. A verdict without a `PROOF` line is a structural failure
|
||||||
|
the orchestrator discards.
|
||||||
|
- Stay in your lens. A finding outside it belongs to another challenger.
|
||||||
|
- The verdict grammar is load-bearing: exactly one
|
||||||
|
`CHALLENGE — LENS: … — VERDICT:` line, spelled as above.
|
||||||
|
|
||||||
|
## ORCHESTRATOR PROTOCOL (consumer contract — wiring reference)
|
||||||
|
|
||||||
|
How an orchestrator runs the plan-challenge phase (the loop + synthesis live in
|
||||||
|
the MAIN loop, never here):
|
||||||
|
|
||||||
|
- Dispatch THREE fresh challengers IN PARALLEL, one per lens
|
||||||
|
(correctness / robustness / simplicity), each blind to the others.
|
||||||
|
- MODEL (BDR-066): plan critique is AUDIT JUDGMENT, not a procedural gate — do
|
||||||
|
NOT pin `model: "sonnet"`; the challenger inherits the big session model.
|
||||||
|
(Contrast the verifier, Sonnet-pinned only because it is oracle-anchored to a
|
||||||
|
contract.)
|
||||||
|
- FAIL-SAFE — never fail open: a malformed/empty verdict, a missing `PROOF`, or
|
||||||
|
a dead challenger → retry ONCE fresh; a 2nd failure → escalate to the human and
|
||||||
|
NAME the lens. Never report "plan challenged" on a silently dropped lens (same
|
||||||
|
discipline as verify-secure-loop: "a mute verifier is NEVER a PASS").
|
||||||
|
- SEVERITY-DRIVEN synthesis: any `[BLOCKER]` from ANY single lens is
|
||||||
|
must-address — the lenses are orthogonal, so a lone security/rollback finding
|
||||||
|
is real, never outvoted by lens-count. Cross-lens agreement only RANKS the MINORs.
|
||||||
|
- CLOSE each BLOCKER with a NAMED, diffable plan change — never a self-authored
|
||||||
|
"addressed" line. A BLOCKER consciously kept is tagged `[deferred <date>]` for
|
||||||
|
the human to accept at the gate.
|
||||||
|
- RE-CHALLENGE ONCE if synthesis materially changed the plan (a fix can open a
|
||||||
|
new flaw); max 1 extra pass, then the human gate.
|
||||||
|
- ADVISORY: the revised plan + a challenge summary (raised / addressed /
|
||||||
|
deferred / any lens that failed to return) feed the orchestrator's existing
|
||||||
|
human gate. The human decides — this is not a hard block.
|
||||||
@@ -0,0 +1,85 @@
|
|||||||
|
# Challenge the plan — shared orchestrator include
|
||||||
|
|
||||||
|
Runs in the ORCHESTRATOR MAIN LOOP after a plan / reflection is elaborated and
|
||||||
|
BEFORE it is executed. Turns a fresh plan into a hardened one by attacking it
|
||||||
|
from three independent angles, then RE-THINKING every aspect a challenger lands.
|
||||||
|
Loop + synthesis decisions live here, in the main loop (BDR-066: reflection runs
|
||||||
|
on the big model; `verify-secure-loop.md`: fresh blind gates, decisions in the
|
||||||
|
loop). It never merges, executes, or edits code — it hardens the plan and hands
|
||||||
|
it to the orchestrator's existing human gate.
|
||||||
|
|
||||||
|
The challenge is ADVISORY into that gate — no new hard block — but a BLOCKER is
|
||||||
|
never silently carried past: it is either closed by a NAMED plan change or
|
||||||
|
explicitly deferred for the human.
|
||||||
|
|
||||||
|
## Inputs the caller must have ready
|
||||||
|
|
||||||
|
- `PLAN`: path to the plan ON DISK. If your plan is still inline (a printed
|
||||||
|
checklist / diagnosis / fix plan), FIRST persist it to
|
||||||
|
`.claude/tasks/plans/<date>-<slug>-<HHMM>.md` — the challengers read from disk
|
||||||
|
and judge blind, exactly like the verifier reads the contract.
|
||||||
|
- `KIND`: `build-plan` | `proposals` | `fix-bundle` — tunes the lens framing
|
||||||
|
below; the mechanism is identical.
|
||||||
|
- `SCOPE`: the files/dirs the plan touches (grounds the critique).
|
||||||
|
- `CONSTRAINTS` (optional): the decided trade-offs / rejected alternatives from
|
||||||
|
the design step, so a lens does not re-litigate a settled choice.
|
||||||
|
|
||||||
|
Nominal path is cheap for a small, clean plan: three parallel challengers return
|
||||||
|
SOLID, synthesis is a no-op. It only costs more when a lens lands a real finding
|
||||||
|
— which is the point.
|
||||||
|
|
||||||
|
## DISPATCH — three fresh challengers, in parallel, blind
|
||||||
|
|
||||||
|
Dispatch THREE fresh `plan-challenger` subagents IN PARALLEL, one per LENS, each
|
||||||
|
blind to the others and to this conversation:
|
||||||
|
|
||||||
|
```
|
||||||
|
Agent(subagent_type="plan-challenger", description="challenge:<lens>", prompt="""
|
||||||
|
PLAN: <the PLAN path>
|
||||||
|
LENS: <correctness | robustness | simplicity> # one per agent — all three
|
||||||
|
SCOPE: <SCOPE>
|
||||||
|
CONSTRAINTS: <CONSTRAINTS, if any>
|
||||||
|
""")
|
||||||
|
```
|
||||||
|
|
||||||
|
**MODEL (BDR-066):** plan critique is AUDIT JUDGMENT — do NOT pin
|
||||||
|
`model: "sonnet"`; the challengers inherit the big session model. (The executor
|
||||||
|
gates stay sonnet; the challenger does not.)
|
||||||
|
|
||||||
|
**Lens framing by `KIND`** (the agent's three lenses, read against the artifact):
|
||||||
|
- `build-plan` — will it WORK / will it BREAK / is it needlessly COMPLEX.
|
||||||
|
- `proposals` — are these the RIGHT items & priorities / what did the audit MISS
|
||||||
|
or under-rate as risk / is the backlog over- or under-scoped.
|
||||||
|
- `fix-bundle` — will each fix ACHIEVE its goal / could it BREAK or regress the
|
||||||
|
page / is there a simpler fix, or an unnecessary one.
|
||||||
|
|
||||||
|
## FAIL-SAFE — never fail open
|
||||||
|
|
||||||
|
A challenger that returns a malformed/empty verdict, a missing `PROOF`, or dies →
|
||||||
|
retry ONCE with a fresh challenger; a 2nd failure on that lens → STOP and escalate
|
||||||
|
to the human, NAMING the lens. Never carry "plan challenged" into the gate on a
|
||||||
|
silently dropped lens (`verify-secure-loop.md`: "a mute verifier is NEVER a PASS").
|
||||||
|
|
||||||
|
## SYNTHESIZE + RE-THINK (main loop, big model)
|
||||||
|
|
||||||
|
Parse each `CHALLENGE — LENS: … — VERDICT:` line and merge the FINDINGS:
|
||||||
|
|
||||||
|
- **Severity-driven, not consensus.** Any `[BLOCKER]` from ANY single lens is
|
||||||
|
must-address — the lenses are orthogonal, so a lone security/rollback finding
|
||||||
|
is real, never outvoted by lens-count. Cross-lens agreement only RANKS the MINORs.
|
||||||
|
- **RE-THINK the aspect the challenge pointed at.** For each BLOCKER (and each
|
||||||
|
MAJOR you accept): revise the plan on THAT aspect — a NAMED, diffable change to
|
||||||
|
the plan, never a self-authored "addressed" line. A BLOCKER you consciously keep
|
||||||
|
is tagged `[deferred <date>]` for the human to accept at the gate.
|
||||||
|
- **Re-challenge once if the plan materially changed** — a fix can open a new
|
||||||
|
flaw. Re-persist the revised `PLAN`, dispatch ONE fresh confirmation challenger,
|
||||||
|
max 1 extra pass, then the gate.
|
||||||
|
|
||||||
|
## OUTPUT — into the existing human gate
|
||||||
|
|
||||||
|
Feed the orchestrator's gate:
|
||||||
|
- the REVISED plan, and
|
||||||
|
- a CHALLENGE SUMMARY: each BLOCKER raised → the named change that closed it;
|
||||||
|
anything `[deferred]`; and any lens that failed to return.
|
||||||
|
|
||||||
|
The human remains the decider.
|
||||||
@@ -0,0 +1,47 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# lib/tests/plan-challenger.test.sh — structure lock: the plan-challenger agent,
|
||||||
|
# the reusable lib/challenge-plan.md phase, and every reflection orchestrator that
|
||||||
|
# wires it (3-way adversarial plan challenge, BDR-066). STATIC only — the agent's
|
||||||
|
# adversarial behavior needs a live model (manual smoke), not a CI gate.
|
||||||
|
set -u
|
||||||
|
R="$(cd "$(dirname "$0")/../.." && pwd)"
|
||||||
|
pass=0; fail=0
|
||||||
|
ok() { pass=$((pass+1)); }
|
||||||
|
ko() { fail=$((fail+1)); printf 'FAIL %s\n' "$1"; }
|
||||||
|
has() { if grep -qF "$2" "$R/$1"; then ok; else ko "$1 missing: $2"; fi; }
|
||||||
|
fm_lacks() { if awk 'NR<=10' "$R/$1" | grep -qF "$2"; then ko "$1 frontmatter must NOT contain: $2"; else ok; fi; }
|
||||||
|
|
||||||
|
A="agents/plan-challenger.md"
|
||||||
|
L="lib/challenge-plan.md"
|
||||||
|
|
||||||
|
# 1) agent shape
|
||||||
|
has "$A" "name: plan-challenger"
|
||||||
|
has "$A" "tools: Read, Grep, Glob, Bash"
|
||||||
|
fm_lacks "$A" "model: sonnet" # BDR-066: audit judgment → NOT sonnet-pinned
|
||||||
|
has "$A" "CHALLENGE — LENS:" # load-bearing verdict grammar
|
||||||
|
has "$A" "VERDICT: SOLID | CONCERNS(n) | FATAL(n)"
|
||||||
|
has "$A" "correctness"
|
||||||
|
has "$A" "robustness"
|
||||||
|
has "$A" "simplicity"
|
||||||
|
has "$A" "Report-only"
|
||||||
|
|
||||||
|
# 2) reusable phase — the mechanism lives here (one canonical include)
|
||||||
|
has "$L" 'subagent_type="plan-challenger"'
|
||||||
|
has "$L" "BDR-066" # challengers on the big model
|
||||||
|
has "$L" "a mute verifier is NEVER a PASS" # fail-safe (never fail open)
|
||||||
|
has "$L" "Severity-driven" # any single-lens BLOCKER = must-address
|
||||||
|
has "$L" "RE-THINK" # findings re-plan the aspect, not just noted
|
||||||
|
has "$L" "correctness | robustness | simplicity"
|
||||||
|
has "$L" "CHALLENGE SUMMARY"
|
||||||
|
has "$L" "build-plan" # the three KINDs
|
||||||
|
has "$L" "proposals"
|
||||||
|
has "$L" "fix-bundle"
|
||||||
|
|
||||||
|
# 3) every reflection orchestrator wires the phase + carries a challenge summary
|
||||||
|
for s in ship-feature init-project feat bugfix onboard audit-delta code-clean seo geo harden web-validate; do
|
||||||
|
has "skills/$s/SKILL.md" "lib/challenge-plan.md"
|
||||||
|
has "skills/$s/SKILL.md" "CHALLENGE SUMMARY"
|
||||||
|
done
|
||||||
|
|
||||||
|
printf 'plan-challenge structure lock: %d pass, %d fail\n' "$pass" "$fail"
|
||||||
|
[ "$fail" -eq 0 ]
|
||||||
@@ -166,10 +166,31 @@ Append to `.claude/audits/AUDIT-DELTA.md` (create if absent), append-only:
|
|||||||
|
|
||||||
Then show the user the same compact table inline.
|
Then show the user the same compact table inline.
|
||||||
|
|
||||||
|
### 3b-bis. CHALLENGE THE PROPOSALS (before the gate)
|
||||||
|
|
||||||
|
This axis' findings + proposed fixes are a proposal set worth attacking before
|
||||||
|
the human gate. Persist THIS axis' finding list (not the whole append-only
|
||||||
|
report) to `.claude/tasks/plans/<date>-<axis>-<HHMM>.md`, then run
|
||||||
|
`$HOME/.claude/lib/challenge-plan.md` with `PLAN` = that file, `KIND` =
|
||||||
|
`proposals`, `SCOPE` = this axis' STEP 1 audit set, `CONSTRAINTS` = the axis
|
||||||
|
spec + the project CLAUDE.md norms already loaded. Three blind challengers ask
|
||||||
|
whether these are the RIGHT findings/priorities and what the audit under-rated;
|
||||||
|
the main loop RE-THINKS every aspect a BLOCKER lands (a named change to the
|
||||||
|
finding set, or `[deferred <date>]`) and re-challenges once if it materially
|
||||||
|
changed. Feed the REVISED findings + a CHALLENGE SUMMARY into 3c.
|
||||||
|
|
||||||
### 3c. APPROVAL GATE ★ MANDATORY STOP
|
### 3c. APPROVAL GATE ★ MANDATORY STOP
|
||||||
|
|
||||||
|
Show the CHALLENGE SUMMARY (from 3b-bis) with the 3b findings table, then
|
||||||
AskUserQuestion: **fix all / pick which / none**.
|
AskUserQuestion: **fix all / pick which / none**.
|
||||||
|
|
||||||
|
```
|
||||||
|
CHALLENGE SUMMARY (3b-bis — 3 lenses):
|
||||||
|
BLOCKERs addressed : <n> — <finding → the named finding-set change that closes it>
|
||||||
|
Deferred (human-ack): <list | none>
|
||||||
|
Lenses returned : correctness / robustness / simplicity (NAME any that failed to return)
|
||||||
|
```
|
||||||
|
|
||||||
- "Fix what you find" said **in the invocation** does NOT skip this gate:
|
- "Fix what you find" said **in the invocation** does NOT skip this gate:
|
||||||
nobody can approve findings that did not exist yet. The gate is about
|
nobody can approve findings that did not exist yet. The gate is about
|
||||||
*these specific findings*.
|
*these specific findings*.
|
||||||
|
|||||||
@@ -117,6 +117,19 @@ RISK: <low/medium — what could go wrong>
|
|||||||
- If the fix is significant (>10 lines, multiple files,
|
- If the fix is significant (>10 lines, multiple files,
|
||||||
behavior change): wait for user approval.
|
behavior change): wait for user approval.
|
||||||
|
|
||||||
|
## STEP 3b — CHALLENGE THE FIX PLAN (before the contract)
|
||||||
|
Unless the fix is the trivial 1-2 line case STEP 3 already fast-paths, the
|
||||||
|
DIAGNOSIS + FIX PLAN is a reflection worth attacking before it hardens into a
|
||||||
|
contract. Persist it to `.claude/tasks/plans/<date>-<slug>-<HHMM>.md`, then run
|
||||||
|
`$HOME/.claude/lib/challenge-plan.md` with `PLAN` = that file, `KIND` = `build-plan`,
|
||||||
|
`SCOPE` = the FIX PLAN files, `CONSTRAINTS` = the STEP 2 in-force BDR/LRN/BLK
|
||||||
|
dispositions. Three blind challengers attack it (correctness = is the root cause
|
||||||
|
right; robustness = blast radius / regressions; simplicity = is the fix minimal);
|
||||||
|
RE-THINK every aspect a BLOCKER lands, re-challenge once if the plan materially
|
||||||
|
changed. STEP 3.5 writes the contract from the REVISED plan. Print a CHALLENGE SUMMARY
|
||||||
|
(BLOCKERs addressed / deferred / lenses returned), folding any deferred BLOCKER into
|
||||||
|
the STEP 3 approval gate.
|
||||||
|
|
||||||
## STEP 3.5 — CONTRACT
|
## STEP 3.5 — CONTRACT
|
||||||
|
|
||||||
Run `$HOME/.claude/lib/contract-interview.md` (main loop). The DIAGNOSIS
|
Run `$HOME/.claude/lib/contract-interview.md` (main loop). The DIAGNOSIS
|
||||||
|
|||||||
@@ -119,11 +119,29 @@ TOTALS: <N blocking, N warn, N info>
|
|||||||
|
|
||||||
If no issues found: report clean state and stop.
|
If no issues found: report clean state and stop.
|
||||||
|
|
||||||
|
## STEP 3b — CHALLENGE THE SCOPE (before approval)
|
||||||
|
The STEP 3 report is the proposed cleanup scope — worth attacking before the
|
||||||
|
human approves it. It is still inline, so FIRST persist it to
|
||||||
|
`.claude/tasks/plans/<date>-<slug>-<HHMM>.md` (STEP 3 report format, one item
|
||||||
|
per line), then run `$HOME/.claude/lib/challenge-plan.md` with `PLAN` = that
|
||||||
|
file, `KIND` = `proposals`, `SCOPE` = the scanned target ($ARGUMENTS or repo
|
||||||
|
root), `CONSTRAINTS` = the STEP 1 project norms + the iron law (zero behavior
|
||||||
|
change). Three blind challengers ask whether these are the RIGHT items and what
|
||||||
|
the scan under- or over-scoped; the main loop RE-THINKS every aspect a BLOCKER
|
||||||
|
lands (a named scope change re-written into the report, or `[deferred <date>]`)
|
||||||
|
and re-challenges once if the scope materially changed. Feed the REVISED scope +
|
||||||
|
a CHALLENGE SUMMARY into STEP 4. Advisory — the human still approves per item.
|
||||||
|
|
||||||
## STEP 4 — VALIDATION GATE (interactive)
|
## STEP 4 — VALIDATION GATE (interactive)
|
||||||
|
|
||||||
Present the report from STEP 3. Then ask:
|
Present the report from STEP 3 with the STEP 3b CHALLENGE SUMMARY. Then ask:
|
||||||
|
|
||||||
```
|
```
|
||||||
|
CHALLENGE SUMMARY (STEP 3b — 3 lenses):
|
||||||
|
BLOCKERs addressed : <n> — <finding → the named scope change that closes it>
|
||||||
|
Deferred (human-ack): <list | none>
|
||||||
|
Lenses returned : correctness / robustness / simplicity (NAME any that failed to return)
|
||||||
|
|
||||||
AskUserQuestion:
|
AskUserQuestion:
|
||||||
Approve which items for execution? (all / <item numbers> / clarify <item>)
|
Approve which items for execution? (all / <item numbers> / clarify <item>)
|
||||||
```
|
```
|
||||||
|
|||||||
@@ -117,6 +117,17 @@ PLAN:
|
|||||||
If the approach is ambiguous: ask the user ONE focused question BEFORE
|
If the approach is ambiguous: ask the user ONE focused question BEFORE
|
||||||
dispatching — never after (the executor cannot relay questions).
|
dispatching — never after (the executor cannot relay questions).
|
||||||
|
|
||||||
|
## STEP 1b — CHALLENGE THE PLAN (before branching)
|
||||||
|
The STEP 1 plan is a reflection worth attacking before a branch is spent on it.
|
||||||
|
Persist it to `.claude/tasks/plans/<date>-<slug>-<HHMM>.md`, then run
|
||||||
|
`$HOME/.claude/lib/challenge-plan.md` with `PLAN` = that file, `KIND` = `build-plan`,
|
||||||
|
`SCOPE` = the STEP 1 files, `CONSTRAINTS` = the STEP 0.6 in-force BDR/LRN dispositions.
|
||||||
|
Three blind challengers attack it; RE-THINK every aspect a BLOCKER lands (a named
|
||||||
|
plan change, or `[deferred]`), re-challenge once if the plan materially changed. The
|
||||||
|
STEP 3 executor receives the REVISED plan. Before dispatch, print a CHALLENGE SUMMARY
|
||||||
|
(BLOCKERs addressed / deferred / lenses returned), surfacing any deferred BLOCKER via
|
||||||
|
STEP 1's one-question gate.
|
||||||
|
|
||||||
## STEP 2 — BRANCH
|
## STEP 2 — BRANCH
|
||||||
|
|
||||||
**Gitflow aiguillage (before dispatch):** follow `$HOME/.claude/lib/gitflow-aiguillage.md`
|
**Gitflow aiguillage (before dispatch):** follow `$HOME/.claude/lib/gitflow-aiguillage.md`
|
||||||
|
|||||||
@@ -58,6 +58,20 @@ $ARGUMENTS
|
|||||||
"""
|
"""
|
||||||
```
|
```
|
||||||
|
|
||||||
|
## STEP 1b — CHALLENGE THE FIX BUNDLE (advisory, before apply)
|
||||||
|
The analyzer returned a `## FIX BUNDLE` — worth attacking before any edit lands.
|
||||||
|
**Skip if intervention mode = conservative** (nothing is applied). Else persist the
|
||||||
|
bundle verbatim to `.claude/tasks/plans/<date>-<slug>-<HHMM>.md`, then run
|
||||||
|
`$HOME/.claude/lib/challenge-plan.md` with `PLAN` = that file, `KIND` = `fix-bundle`,
|
||||||
|
`SCOPE` = the target site files the items touch, `CONSTRAINTS` = the geo-analyzer
|
||||||
|
file-ownership (robots.txt, llms.txt, JSON-LD, content shape) + the shared-file edit
|
||||||
|
discipline each item carries + intervention mode. Three blind challengers ask, per item:
|
||||||
|
will it ACHIEVE its goal / could it BREAK or regress the page / is a simpler (or no) fix
|
||||||
|
better. This main loop RE-THINKS every aspect a BLOCKER lands (a named bundle change, or
|
||||||
|
`[deferred <date>]`) and re-challenges once if the bundle materially changed. Advisory —
|
||||||
|
it sits BEFORE (never replaces) the STEP 2 GATED approval; carry its CHALLENGE SUMMARY
|
||||||
|
into that gate.
|
||||||
|
|
||||||
## STEP 2 — Apply the fix bundle (from THIS main loop, at L1)
|
## STEP 2 — Apply the fix bundle (from THIS main loop, at L1)
|
||||||
|
|
||||||
The analyzer returned a `## FIX BUNDLE`. Apply it by dispatching
|
The analyzer returned a `## FIX BUNDLE`. Apply it by dispatching
|
||||||
@@ -90,6 +104,11 @@ Present every GATED item (G5.x) in ONE gate:
|
|||||||
```
|
```
|
||||||
GEO — gated content-shape changes need approval (visible):
|
GEO — gated content-shape changes need approval (visible):
|
||||||
G5.1 <change> — impact: <visible change>
|
G5.1 <change> — impact: <visible change>
|
||||||
|
|
||||||
|
CHALLENGE SUMMARY (STEP 1b — 3 lenses):
|
||||||
|
BLOCKERs addressed : <n> — <finding → the named bundle change that closes it>
|
||||||
|
Deferred (human-ack): <list | none>
|
||||||
|
Lenses returned : correctness / robustness / simplicity (NAME any that failed to return)
|
||||||
Approve all / select (ids) / skip all?
|
Approve all / select (ids) / skip all?
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|||||||
@@ -518,6 +518,21 @@ Extract the score and critical-alert count from `.claude/audits/HARDEN.md` for t
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
## STEP 2b — CHALLENGE THE FIX BUNDLE (MODE=fix only, advisory)
|
||||||
|
Skip if MODE=audit (no bundle exists). Else, before the STEP 3 gate, harden the bundle:
|
||||||
|
extract the `## 8. Fix bundle` section from HARDEN.md to
|
||||||
|
`.claude/tasks/plans/<date>-<slug>-<HHMM>.md` (a clean, blind-judgeable artifact), then run
|
||||||
|
`$HOME/.claude/lib/challenge-plan.md` with `PLAN` = that file, `KIND` = `fix-bundle`,
|
||||||
|
`SCOPE` = the config/target files each patch touches (.htaccess, next.config.js, _headers…),
|
||||||
|
`CONSTRAINTS` = the STEP 1 strict scope + framework-native mechanism rule (no `.htaccess`
|
||||||
|
on Next/Astro) + the severity guide. Three blind challengers ask, per patch: will it ACHIEVE
|
||||||
|
the hardening goal / could it BREAK the site (over-broad CSP, redirect loop) / is a simpler
|
||||||
|
fix better. This main loop RE-THINKS every aspect a BLOCKER lands (a named bundle change, or
|
||||||
|
`[deferred <date>]`) and re-challenges once if the bundle materially changed. Advisory — it
|
||||||
|
sits BEFORE (never replaces) the STEP 3 confirmation; carry its CHALLENGE SUMMARY into that gate.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
## STEP 3 — Apply fixes (MODE=fix only)
|
## STEP 3 — Apply fixes (MODE=fix only)
|
||||||
|
|
||||||
Skip this step if MODE=audit.
|
Skip this step if MODE=audit.
|
||||||
@@ -534,6 +549,11 @@ If MODE=fix and `.claude/audits/HARDEN.md` ends with `READY TO APPLY — awaitin
|
|||||||
- .htaccess (3 fixes : HTTP→HTTPS redirect, HSTS, 404 page)
|
- .htaccess (3 fixes : HTTP→HTTPS redirect, HSTS, 404 page)
|
||||||
- next.config.js (2 fixes : CSP header, X-Frame-Options)
|
- next.config.js (2 fixes : CSP header, X-Frame-Options)
|
||||||
|
|
||||||
|
CHALLENGE SUMMARY (STEP 2b — 3 lenses) :
|
||||||
|
BLOCKERs addressed : <n> — <finding → the named bundle change that closes it>
|
||||||
|
Deferred (human-ack): <list | none>
|
||||||
|
Lenses returned : correctness / robustness / simplicity (NAME any that failed to return)
|
||||||
|
|
||||||
Options :
|
Options :
|
||||||
A) Apply all
|
A) Apply all
|
||||||
B) Review each diff before applying
|
B) Review each diff before applying
|
||||||
|
|||||||
@@ -155,12 +155,27 @@ implemented on a `feature/*` branch off `develop` (STEP 8).
|
|||||||
Invoke `superpowers:writing-plans` with BRIEF + skeleton.
|
Invoke `superpowers:writing-plans` with BRIEF + skeleton.
|
||||||
Granular tasks (2-5 min each), exact file paths, TDD: tests before code.
|
Granular tasks (2-5 min each), exact file paths, TDD: tests before code.
|
||||||
|
|
||||||
|
## STEP 6b — CHALLENGE THE PLAN (before the gate)
|
||||||
|
Before the human sees the implementation plan, harden it. Run
|
||||||
|
`$HOME/.claude/lib/challenge-plan.md` with `PLAN` = the plan STEP 6 wrote under
|
||||||
|
`docs/superpowers/plans/`, `KIND` = `build-plan`, `SCOPE` = the skeleton + task file
|
||||||
|
paths, `CONSTRAINTS` = the STEP 4-validated architecture + founding decisions.
|
||||||
|
Three blind challengers (correctness / robustness / simplicity) attack it; the main
|
||||||
|
loop RE-THINKS every aspect a BLOCKER lands (a named plan change, or `[deferred]`),
|
||||||
|
re-challenges once if the plan materially changed, and feeds the REVISED plan + a
|
||||||
|
CHALLENGE SUMMARY into STEP 7. Advisory — the human remains the decider.
|
||||||
|
|
||||||
## STEP 7 — VALIDATION GATE #2 ★ MANDATORY STOP
|
## STEP 7 — VALIDATION GATE #2 ★ MANDATORY STOP
|
||||||
```
|
```
|
||||||
INIT PROJECT — IMPLEMENTATION PLAN
|
INIT PROJECT — IMPLEMENTATION PLAN
|
||||||
SKELETON: ✅ build passes
|
SKELETON: ✅ build passes
|
||||||
FEATURES: <N> → <M> tasks
|
FEATURES: <N> → <M> tasks
|
||||||
<numbered task list with paths>
|
<numbered task list with paths>
|
||||||
|
|
||||||
|
CHALLENGE SUMMARY (STEP 6b — 3 lenses):
|
||||||
|
BLOCKERs addressed : <n> — <finding → the named plan change that closes it>
|
||||||
|
Deferred (human-ack): <list | none>
|
||||||
|
Lenses returned : correctness / robustness / simplicity (NAME any that failed to return)
|
||||||
Approve and start? (yes / request changes)
|
Approve and start? (yes / request changes)
|
||||||
```
|
```
|
||||||
Changes → back to STEP 6. Approved → continue.
|
Changes → back to STEP 6. Approved → continue.
|
||||||
|
|||||||
@@ -869,6 +869,22 @@ Vérifier que les 4 fichiers `.claude/audits/ONBOARD_REPORT.md`, `.claude/audits
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
## STEP 7b — CHALLENGE THE PROPOSALS (before the human gate)
|
||||||
|
The 4 audit files are on disk; `AUDIT_PROPOSALS.md` is the artifact worth
|
||||||
|
attacking before the human spends a gate on it. Run
|
||||||
|
`$HOME/.claude/lib/challenge-plan.md` with `PLAN` =
|
||||||
|
`.claude/audits/AUDIT_PROPOSALS.md`, `KIND` = `proposals`, `SCOPE` = the audited
|
||||||
|
project paths (the `audit_stack` coverage), `CONSTRAINTS` = the STEP 1 archetype
|
||||||
|
profile + the STEP 3 interview constraints (stade, légal, budget perf). Three
|
||||||
|
blind challengers ask whether these are the RIGHT priorities and what the audit
|
||||||
|
under-rated; the main loop RE-THINKS every aspect a BLOCKER lands (a named
|
||||||
|
proposals change re-written into `AUDIT_PROPOSALS.md`, or `[deferred <date>]`)
|
||||||
|
and re-challenges once if the file materially changed. Feed the REVISED
|
||||||
|
proposals + a CHALLENGE SUMMARY into STEP 8. Advisory — the human remains the
|
||||||
|
decider.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
## STEP 8 — VALIDATION GATE ★ MANDATORY STOP
|
## STEP 8 — VALIDATION GATE ★ MANDATORY STOP
|
||||||
|
|
||||||
Afficher à l'utilisateur :
|
Afficher à l'utilisateur :
|
||||||
@@ -893,6 +909,11 @@ TOP 5 PRIORITÉS :
|
|||||||
4. [P1 Haute] <titre>
|
4. [P1 Haute] <titre>
|
||||||
5. [P2 Moyenne] <titre>
|
5. [P2 Moyenne] <titre>
|
||||||
|
|
||||||
|
CHALLENGE SUMMARY (STEP 7b — 3 lenses):
|
||||||
|
BLOCKERs addressed : <n> — <finding → the named proposals change that closes it>
|
||||||
|
Deferred (human-ack): <list | none>
|
||||||
|
Lenses returned : correctness / robustness / simplicity (NAME any that failed to return)
|
||||||
|
|
||||||
Prochaine étape : générer .claude/tasks/TODO.md depuis .claude/audits/AUDIT_PROPOSALS.md approuvé.
|
Prochaine étape : générer .claude/tasks/TODO.md depuis .claude/audits/AUDIT_PROPOSALS.md approuvé.
|
||||||
|
|
||||||
Options :
|
Options :
|
||||||
|
|||||||
@@ -447,6 +447,19 @@ applies your bundle in STEP 1.5 and merges the reports.
|
|||||||
"""
|
"""
|
||||||
```
|
```
|
||||||
|
|
||||||
|
## STEP 1b — CHALLENGE THE FIX BUNDLE (advisory, before apply)
|
||||||
|
Both envelopes now carry a `## FIX BUNDLE` — worth attacking before any edit lands.
|
||||||
|
**Skip if intervention mode = conservative** (nothing is applied). Else persist both
|
||||||
|
bundles (seo + geo, verbatim) to `.claude/tasks/plans/<date>-<slug>-<HHMM>.md`, then run
|
||||||
|
`$HOME/.claude/lib/challenge-plan.md` with `PLAN` = that file, `KIND` = `fix-bundle`,
|
||||||
|
`SCOPE` = the target site files the items touch, `CONSTRAINTS` = the STEP 0 file-ownership
|
||||||
|
matrix + shared-file edit discipline + confirmed Canonical NAP + intervention mode. Three
|
||||||
|
blind challengers ask, per item: will it ACHIEVE its goal / could it BREAK or regress the
|
||||||
|
page / is a simpler (or no) fix better. This main loop RE-THINKS every aspect a BLOCKER
|
||||||
|
lands (a named bundle change, or `[deferred <date>]`) and re-challenges once if the bundle
|
||||||
|
materially changed. Advisory — it sits BEFORE (never replaces) the STEP 1.5 GATED approval;
|
||||||
|
carry its CHALLENGE SUMMARY into that gate.
|
||||||
|
|
||||||
## STEP 1.5 — Apply fix bundles (from THIS main loop, at L1)
|
## STEP 1.5 — Apply fix bundles (from THIS main loop, at L1)
|
||||||
|
|
||||||
Both analyzers returned an envelope containing a `## FIX BUNDLE` section
|
Both analyzers returned an envelope containing a `## FIX BUNDLE` section
|
||||||
@@ -493,6 +506,11 @@ Collect every GATED item from BOTH bundles and present ONE gate:
|
|||||||
SEO/GEO — gated changes need approval (visible / structural):
|
SEO/GEO — gated changes need approval (visible / structural):
|
||||||
D1 <change> — impact: <visible change> [seo]
|
D1 <change> — impact: <visible change> [seo]
|
||||||
G5.1 <change> — impact: <visible change> [geo]
|
G5.1 <change> — impact: <visible change> [geo]
|
||||||
|
|
||||||
|
CHALLENGE SUMMARY (STEP 1b — 3 lenses):
|
||||||
|
BLOCKERs addressed : <n> — <finding → the named bundle change that closes it>
|
||||||
|
Deferred (human-ack): <list | none>
|
||||||
|
Lenses returned : correctness / robustness / simplicity (NAME any that failed to return)
|
||||||
Approve all / select (ids) / skip all?
|
Approve all / select (ids) / skip all?
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|||||||
@@ -115,6 +115,19 @@ Invoke `superpowers:writing-plans` with the validated design AND the 0d digest:
|
|||||||
must be consistent with the in-force constraints; where a task implements or affects one,
|
must be consistent with the in-force constraints; where a task implements or affects one,
|
||||||
note the ID inline. Break design into tasks (2-5 min each). Each task: exact file paths, full code, verification steps.
|
note the ID inline. Break design into tasks (2-5 min each). Each task: exact file paths, full code, verification steps.
|
||||||
|
|
||||||
|
## STEP 2b — CHALLENGE THE PLAN (adversarial, before the gate)
|
||||||
|
Before the human sees the plan, harden it. Run `$HOME/.claude/lib/challenge-plan.md`:
|
||||||
|
- `PLAN` = the plan STEP 2 wrote under `docs/superpowers/plans/`
|
||||||
|
- `KIND` = `build-plan`
|
||||||
|
- `SCOPE` = the files/dirs the plan touches
|
||||||
|
- `CONSTRAINTS` = the STEP 1 validated design's decided trade-offs / rejected options
|
||||||
|
|
||||||
|
Three blind `plan-challenger` subagents (correctness / robustness / simplicity)
|
||||||
|
attack it in parallel on the big model; the main loop RE-THINKS every aspect a
|
||||||
|
BLOCKER lands (a named plan change, or `[deferred <date>]`), re-challenges once if
|
||||||
|
the plan materially changed, and feeds the REVISED plan + a CHALLENGE SUMMARY into
|
||||||
|
STEP 3. Advisory — the human remains the decider.
|
||||||
|
|
||||||
## STEP 3 — VALIDATION GATE ★ MANDATORY STOP
|
## STEP 3 — VALIDATION GATE ★ MANDATORY STOP
|
||||||
```
|
```
|
||||||
SHIP FEATURE — VALIDATION GATE
|
SHIP FEATURE — VALIDATION GATE
|
||||||
@@ -127,6 +140,11 @@ RELATED MEMORY — disposition CLAIMED by this plan (review each):
|
|||||||
- BLK-009 [already seen] — <how avoided / why N-A>
|
- BLK-009 [already seen] — <how avoided / why N-A>
|
||||||
|
|
||||||
Review the claims above — flag any item the plan does NOT actually honor.
|
Review the claims above — flag any item the plan does NOT actually honor.
|
||||||
|
|
||||||
|
CHALLENGE SUMMARY (STEP 2b — 3 lenses):
|
||||||
|
BLOCKERs addressed : <n> — <finding → the named plan change that closes it>
|
||||||
|
Deferred (human-ack): <list | none>
|
||||||
|
Lenses returned : correctness / robustness / simplicity (NAME any that failed to return)
|
||||||
Approve and execute? (yes / request changes)
|
Approve and execute? (yes / request changes)
|
||||||
```
|
```
|
||||||
This block EXPOSES each in-force item with the plan's CLAIMED disposition, for human
|
This block EXPOSES each in-force item with the plan's CLAIMED disposition, for human
|
||||||
|
|||||||
@@ -251,6 +251,22 @@ grep -c '^### \[Critique\]' .claude/audits/VALIDATE.md
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
## STEP 2b — CHALLENGE THE FIX BUNDLE (MODE=fix only, advisory)
|
||||||
|
Skip if MODE=audit (no bundle exists). Else, before the STEP 3 gate, harden the bundle:
|
||||||
|
extract the `## 5. Fix bundle` section from VALIDATE.md to
|
||||||
|
`.claude/tasks/plans/<date>-<slug>-<HHMM>.md` (a clean, blind-judgeable artifact), then run
|
||||||
|
`$HOME/.claude/lib/challenge-plan.md` with `PLAN` = that file, `KIND` = `fix-bundle`,
|
||||||
|
`SCOPE` = the HTML/CSS files each fix touches, `CONSTRAINTS` = the STEP 1 strict scope (W3C
|
||||||
|
validity + WCAG 2.1) + the conservative auto-fix rule (structural/syntactic only, content →
|
||||||
|
§6) + the shared-template discipline (targeted Edit, never Write — templates carry /seo +
|
||||||
|
/geo content). Three blind challengers ask, per fix: will it ACHIEVE conformance / could it
|
||||||
|
BREAK rendering or regress another SC / is a simpler fix better. This main loop RE-THINKS
|
||||||
|
every aspect a BLOCKER lands (a named bundle change, or `[deferred <date>]`) and re-challenges
|
||||||
|
once if the bundle materially changed. Advisory — it sits BEFORE (never replaces) the STEP 3
|
||||||
|
confirmation; carry its CHALLENGE SUMMARY into that gate.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
## STEP 3 — Apply fixes (MODE=fix only)
|
## STEP 3 — Apply fixes (MODE=fix only)
|
||||||
|
|
||||||
Skip this step if `MODE=audit`.
|
Skip this step if `MODE=audit`.
|
||||||
@@ -271,6 +287,11 @@ Files to modify (N) :
|
|||||||
|
|
||||||
Critical : X | Haute : Y | Moyenne : Z | Basse : W
|
Critical : X | Haute : Y | Moyenne : Z | Basse : W
|
||||||
|
|
||||||
|
CHALLENGE SUMMARY (STEP 2b — 3 lenses) :
|
||||||
|
BLOCKERs addressed : <n> — <finding → the named bundle change that closes it>
|
||||||
|
Deferred (human-ack): <list | none>
|
||||||
|
Lenses returned : correctness / robustness / simplicity (NAME any that failed to return)
|
||||||
|
|
||||||
Options :
|
Options :
|
||||||
A) Apply all
|
A) Apply all
|
||||||
B) Review each diff before applying
|
B) Review each diff before applying
|
||||||
|
|||||||
Reference in New Issue
Block a user