Merge feature/model-tiering-w1-no-inherit into develop
This commit is contained in:
@@ -28,7 +28,7 @@ on audits). User approved: opus for judgment agents, drop local opus pin.
|
||||
- [x] Memory: BDR-076 append + journal line. Commit (feat + chore),
|
||||
NO merge (human gate).
|
||||
|
||||
## 2026-07-17 — STATUS seo/geo parity (branch bugfix/seo-geo-integrity, 10 commits, UNMERGED)
|
||||
## 2026-07-17 — STATUS seo/geo parity (branch bugfix/seo-geo-integrity — MERGED to develop, 92301fe; "UNMERGED" note was stale, corrected 2026-07-19 W0)
|
||||
PHASE 1 — integrity: **DONE 7/7**. I3 8b0c98c · I1 57c67f2 · I2 4ea2fb8 ·
|
||||
I5 64f175f · I4 e70e1d6 · I6 9da1dec · I8 acd452b. Plus 9cd7b51 (A1+A2, two
|
||||
process anomalies surfaced by dogfooding /harden at zenquality.fr from the
|
||||
|
||||
@@ -0,0 +1,277 @@
|
||||
# ANALYSIS: model-tiering v2 — Fable = orchestration + plan/solution reflection only; dispatched fleet tiered opus/sonnet/haiku by task complexity; split mixed-tier agents
|
||||
|
||||
Produced by /analyze (main loop, Fable) + 4 subagent sweeps (2× agent-body
|
||||
classification, dispatch map, test-lock inventory), 2026-07-19. Facts verified
|
||||
against: model-routing.test.sh, challenge-plan.md, verify-secure-loop.md,
|
||||
model-gate.md, BDR-050/061/066/076, LRN-113/125/126 (read in full inline).
|
||||
Subagent-reported details not re-verified inline are marked (sub) — LRN-132
|
||||
applies: re-verify load-bearing ones before cutting code.
|
||||
|
||||
## CONTEXT
|
||||
|
||||
- Current state (branch `feature/opus-pin-audit-agents`, 2 commits, UNMERGED):
|
||||
main loop = session model (Fable; model-gate blocks small models in 15
|
||||
reflection skills). Dispatched pins: opus = analyzer, plan-challenger,
|
||||
seo-analyzer, geo-analyzer, validator-analyzer (BDR-076); sonnet = 14
|
||||
executors; haiku = status-reporter. Unpinned = interviewer,
|
||||
client-handover-writer (inline-load only).
|
||||
- Two execution modes with OPPOSITE tier semantics: Agent() dispatch →
|
||||
frontmatter pin applies; inline-load ("you become it") → pin INERT, runs on
|
||||
session model. 20 inline-load sites exist.
|
||||
- Target policy (user directive): Fable does ONLY main-loop orchestration +
|
||||
reflection on plan/solution. Everything dispatched runs opus (deep judgment)
|
||||
/ sonnet (standard execution) / haiku (mechanical) by ACTUAL task
|
||||
complexity. Agents mixing classes get split. Skills adapted. Zero loss, zero
|
||||
regression.
|
||||
|
||||
## KEY COMPONENTS — per-agent verdict vs target
|
||||
|
||||
### Fits, no change
|
||||
| agent | tier | note |
|
||||
|---|---|---|
|
||||
| plan-challenger | opus | coherent monolith; verdict grammar + PROOF load-bearing |
|
||||
| feater / bugfixer / hotfixer | sonnet | closed-plan executors; NEED-DECISION / BLOCKED valves |
|
||||
| security-auditor | sonnet | deterministic SAST gate; `SECURITY — VERDICT:` grammar |
|
||||
| scaffolder | sonnet (effort: high) | but see INERT-PIN below — never dispatched today |
|
||||
| status-reporter | haiku | exemplar mechanical |
|
||||
| client-handover-writer | none (inline orchestrator) | one haiku-able seam: STEP 1-2 git/context preflight |
|
||||
| interviewer | none (inline) | INTERACTIVE — asks user inline; a dispatched agent cannot ask (uniform ban). Structurally main-loop. |
|
||||
|
||||
### Tier-down candidates (no split)
|
||||
| agent | current → candidate | evidence |
|
||||
|---|---|---|
|
||||
| validator-analyzer | opus → sonnet | NOT mixed: runs external validators (authoritative), fixed severity tables, base-100 deduction scoring, allowlist-driven fix bundle; ambiguity punted to user §6. No deep judgment present. (sub) |
|
||||
| onboarder | sonnet → haiku candidate | template-fill + conditional writes; only light stack-block filtering. (sub) Also inert-pin today. |
|
||||
| release-executor | sonnet (keep, borderline) | mostly script runs + CHANGELOG templating, but carries a NEED-DECISION judgment valve (MAJOR-bump wording). (sub) |
|
||||
|
||||
### Split candidates (mixed classes inside one body)
|
||||
| agent | geometry (factual boundary) | complication |
|
||||
|---|---|---|
|
||||
| seo-analyzer | collection (STEP 2-5 curls/CWV/GSC/greps → haiku-class) / judgment (STEP 6-11 sampling, competitive, scoring, triage → opus) / templating (STEP 12-14 bundle+report → sonnet/haiku) | BDR-061: no Agent tool in analyzers (single-dispatch doctrine) → a split must be ORCHESTRATED BY THE SKILL at L1 with disk handoffs, or BDR-061 revised (nesting works ≥2.1.172 per BDR-060, but version-robust-by-design was chosen). seo-data.test.sh locks `fetch.sh` wiring strings IN the agent body (6 locks). STEP 1-2 context feeds every later step → large LRN-126 contract surface. |
|
||||
| geo-analyzer | identical 3-way geometry | same complications; shares severity vocab + sentinel |
|
||||
| commit-changer | MODE propose (narrative reconstruction + capitalize routing = deep) / MODE apply (stage+commit = mechanical) — boundary ALREADY exists as dispatch modes | 2 dispatch sites in /commit-change; per-dispatch `model=` override is an available lighter mechanism than a file split |
|
||||
| doc-syncer | drift detection + semantic doc-type analysis + MINOR/SIGNIFICANT calls (deep) / discovery + template render + PATCHED_FILES emit (mechanical) | 9 consumers on BOTH modes: dispatched ×2 (/doc, onboard) + inline-load ×7 (bugfix, hotfix, feat, init-project ×2, ship-feature, scaffolder) — LRN-125 dual-use-across-tiers hazard; runs its own user validation gate (STEP 8) → gate must be hoisted before any dispatch conversion |
|
||||
| handover-doc-writer | synthesis/vulgarization STEP 10-12 (deep) / render+deterministic gates STEP 13-16 (mechanical) | skill-leak ban list + `HANDOVER-DOC REPORT` grammar must survive |
|
||||
| plugin-advisor | detection PHASE 1 (mechanical) / complexity scoring + decision-table reasoning PHASE 2.5 (deep) | INERT PIN: inline-loaded ×4 (plugin-check, onboard, init-project, ship-feature), NEVER dispatched — sonnet pin is dead config; PHASE 4 asks the user (inline-only capability) |
|
||||
| verifier | STEP 2 evidence adjudication = deep judgment inside a sonnet procedural gate | BDR-066 kept sonnet DELIBERATELY (oracle-anchored to contract, ≤3×/loop). Tier-up = design arbitrage, not a mechanical fix. contract-verifier.test.sh locks name/tools/body (33 asserts). |
|
||||
|
||||
### INERT-PIN finding (structural gap vs target)
|
||||
scaffolder, onboarder, plugin-advisor are pinned sonnet but NEVER dispatched —
|
||||
inline-load only → they run on Fable today. doc-syncer's doc-commit steps
|
||||
(bugfix/hotfix/feat/init-project/ship-feature/scaffolder) also run inline on
|
||||
Fable. Under the target policy these are EXECUTION tasks burning Fable — a
|
||||
bigger real gap than any pin value. Each inline→dispatch conversion must hoist
|
||||
its user gates into the dispatcher first (dispatched agents cannot ask).
|
||||
|
||||
## CONSUMER MAP (summary; full tables in the dispatch-map sweep)
|
||||
|
||||
- ~50 Agent() dispatch sites across 20 skills + 2 lib includes +
|
||||
client-handover-writer (9 internal dispatches, incl. skills-via-general-purpose).
|
||||
- 20 inline-load sites (7× doc-syncer, 4× plugin-advisor, 3× analyzer, 2×
|
||||
interviewer, 1× each onboarder/scaffolder/client-handover-writer/refactorer).
|
||||
- Includes: model-gate.md ×15 skills (+5 locked EXCLUDED), challenge-plan.md
|
||||
×12, verify-secure-loop.md ×5, contract-interview ×5, capitalize-commit ×6,
|
||||
doc-commit ×6.
|
||||
- ~30 prose refs claim current tiers (sonnet-pinned X, opus-pinned Y, BDR-066/
|
||||
BDR-076 citations) → all go stale on tier changes (LRN-113 sweep required).
|
||||
- Only onboard uses explicit `model="opus"` dispatch params (7 sites); every
|
||||
typed agent relies on frontmatter pin; ship-feature/init-project mandate
|
||||
`model: "sonnet"` on SDD subagents by prose.
|
||||
|
||||
## CONSTRAINTS (zero-loss bar)
|
||||
|
||||
1. Verbatim machine-parsed grammars must survive verbatim: `VERIFY — VERDICT:
|
||||
CONFORME | ECARTS(n) | ERROR(<reason>)`, `SECURITY — VERDICT: PASS |
|
||||
BLOCK(n) | ERROR(<reason>)`, `CHALLENGE — LENS: … — VERDICT: SOLID |
|
||||
CONCERNS(n) | FATAL(n)`, mandatory `PROOF:` lines, sentinel `READY TO APPLY
|
||||
— awaiting dispatcher confirmation`, `<NAME>-EXEC REPORT` + `STATUS : DONE
|
||||
| NEED-DECISION | BLOCKED`, `PATCHED_FILES:`, `COMMIT PLAN`, labeled score
|
||||
lines parsed by client-handover extractors, `HANDOVER-DOC REPORT`.
|
||||
2. BDR-050 + LRN-083: loops + decisions live in the MAIN loop; gates dispatched
|
||||
fresh, blind, zero iteration history. Splits must not move loop decisions
|
||||
into children.
|
||||
3. BDR-061: seo/geo/validator have no Agent tool by doctrine (version-robust
|
||||
single dispatch level). Any intra-audit split is skill-orchestrated at L1
|
||||
unless BDR-061 is explicitly revised.
|
||||
4. LRN-126: every implicit data path (ARGUMENTS flags, detected vars, STEP-N
|
||||
side outputs) must cross the new handoff contracts explicitly; census-style
|
||||
tests will NOT catch severed wires — a data-flow read per split is required.
|
||||
5. LRN-125: no dual-use agent across tiers; audit consumer routes to the
|
||||
judgment agent, execution consumer to the executor.
|
||||
6. Interactivity: dispatched agents cannot ask the user. All human gates
|
||||
(AskUserQuestion / inline approval) stay in main loop or inline-loaded
|
||||
orchestrators. doc-syncer STEP 8 + plugin-advisor PHASE 4 gates must be
|
||||
hoisted before dispatch conversion.
|
||||
7. Test locks (fire on this refactor): model-routing (~61, epicenter — pins,
|
||||
dispatch strings, gate wiring loops, `model="opus"` literals, BDR-076 token),
|
||||
plan-challenger (~43 — frontmatter, grammar, challenge-plan doctrine
|
||||
sentences incl. BDR-066 token), loops-light (40 — verify-secure-loop 10
|
||||
sentences, sonnet pins, report grammars, "Agent" ABSENT from
|
||||
bugfixer/hotfixer — substring-fragile), contract-verifier (33),
|
||||
security-auditor (31), seo-data (6 body-wiring locks on seo/geo bodies),
|
||||
loops-heavy (19 skill prose), review-guards G3 (strict YAML on every agent
|
||||
file incl. new ones), no-vacuous-locks (no `\n` in new lock patterns —
|
||||
LRN-093), model-check (10 — tier vocabulary big/small; a new tier taxonomy
|
||||
must co-evolve witness + test). Census `for`-loops (model-routing:13-19,
|
||||
plan-challenger:42) must be edited for any new/renamed gated skill.
|
||||
8. model-gate.md prose has NO deterministic lock (include-path only) — free to
|
||||
rewrite, but behavioral-only verification.
|
||||
9. Gitflow: feature branch(es) via gitflow.sh; no merge without human signal.
|
||||
Unmerged branches in flight: `feature/opus-pin-audit-agents` (this refactor
|
||||
supersedes/absorbs it), `bugfix/seo-geo-integrity` (10 commits touching the
|
||||
seo surface → sequencing/conflict risk with a seo-analyzer split).
|
||||
10. BDR-076 survival: opus tier for judgment agents survives as baseline;
|
||||
validator-analyzer's opus pin would be superseded (tier-down); seo/geo pins
|
||||
refined by splits; challenge-plan/plan-challenger doctrine text + census
|
||||
§11 rewritten again.
|
||||
|
||||
## RISKS
|
||||
|
||||
- Severed implicit data paths on splits (LRN-126 precedent: 2 silent input
|
||||
losses caught only by whole-branch review) — probability: HIGH without a
|
||||
per-split data-flow pass.
|
||||
- Consumer staleness (LRN-113): ~30 prose refs + 9 identical gate preambles +
|
||||
2 census loops — partial sweep leaves contradictory doctrine — probability:
|
||||
HIGH without whole-surface grep + new guards.
|
||||
- Lost human gates on inline→dispatch conversions (doc-syncer STEP 8,
|
||||
plugin-advisor PHASE 4) — probability: MEDIUM-HIGH; hoist-first pattern
|
||||
exists (BDR-066 wave 4 did exactly this for client-handover).
|
||||
- Census under-coverage: NEW agent files are silently unlocked unless
|
||||
model-routing/census extended per agent (worse than a red) — MEDIUM.
|
||||
- haiku reliability on long tool chains (seo/geo collection legs: GSC, CWV,
|
||||
curl loops, retry policies): only haiku precedent is status-reporter
|
||||
(short, deterministic) — MEDIUM; unproven.
|
||||
- Split overhead: 3-dispatch audit pipeline re-serializes STEP 1-2 context per
|
||||
child; latency + token duplication vs today's monolith — MEDIUM.
|
||||
- Merge sequencing with `bugfix/seo-geo-integrity` (10 commits on seo surface)
|
||||
— MEDIUM.
|
||||
- Subagent-report trust (LRN-132): (sub)-marked classifications need spot
|
||||
re-verification during design — MEDIUM.
|
||||
|
||||
## OPEN QUESTIONS (design arbitrage needed)
|
||||
|
||||
1. verifier: keep sonnet (BDR-066 oracle-anchored rationale) or lift to opus
|
||||
(STEP 2 adjudication is the correctness gate)?
|
||||
2. seo/geo split mechanics: skill-orchestrated L1 pipeline (BDR-061-compatible)
|
||||
vs nested dispatch inside the analyzer (requires revising BDR-061;
|
||||
version floor OK per BDR-060)?
|
||||
3. Which inline-loads convert to dispatches (scaffolder, onboarder, doc-syncer
|
||||
doc-commit steps, plugin-advisor detection) vs stay inline as reflection?
|
||||
4. commit-changer: file split vs per-mode `model=` override at the 2 existing
|
||||
dispatch sites?
|
||||
5. haiku scope: which mechanical halves actually go haiku vs sonnet, given the
|
||||
reliability unknown on long tool chains?
|
||||
6. Gate taxonomy: keep binary big/small model-gate (guards main loop only) or
|
||||
extend model-check.sh to the full 4-tier vocabulary?
|
||||
7. Sequencing: land/absorb `feature/opus-pin-audit-agents` and
|
||||
`bugfix/seo-geo-integrity` before or during this refactor?
|
||||
|
||||
## DESIGN AMENDMENT (2026-07-19, user arbitrage — supersedes open questions)
|
||||
|
||||
User approved all 7 recommendations, PLUS one addition:
|
||||
|
||||
**No-inherit rule + fable pins.** No dispatched agent may inherit the session
|
||||
model anywhere. Every dispatch site carries an explicit tier: typed agents via
|
||||
frontmatter pin (`model: fable|opus|sonnet|haiku`), built-ins
|
||||
(general-purpose / Explore / Plan) via a `model=` param at EVERY call site.
|
||||
Rationale: sessions may run on another model (gate admits Opus; user may
|
||||
launch anything) — inheritance would silently mis-tier dispatched work.
|
||||
`model="fable"` lands where a dispatched child performs REFLECTION /
|
||||
ORCHESTRATION on behalf of the main loop:
|
||||
- client-handover-writer's 8 internal general-purpose skill-runner dispatches
|
||||
(/seo, /harden, /cso, /commit-change, /web-validate runs) — today they
|
||||
inherit; they host gated orchestration → `model="fable"`.
|
||||
- Doctrine line (model-gate.md or routing doctrine): ad-hoc reflection
|
||||
dispatches from the main loop (Explore digest, Plan, general-purpose) carry
|
||||
`model="fable"`; non-reflection ad-hoc dispatches carry their complexity
|
||||
tier. New census locks accordingly.
|
||||
- No TYPED agent moves to fable tier (plan-challenger/analyzer stay opus per
|
||||
approved verdicts). Inline-loads that remain (interviewer,
|
||||
client-handover-writer, analyzer-in-/analyze + DEBUG, init STEP 2) ARE the
|
||||
main loop — covered by model-gate, not pins.
|
||||
- External/gstack skills with inheriting general-purpose dispatches
|
||||
(design-shotgun, review, graphify) — external ownership (BDR-015 class):
|
||||
covered by doctrine, not edited, unless owned locally. Verify ownership at
|
||||
implementation.
|
||||
|
||||
## TARGET MODEL MAP — ship-feature (example, per-step)
|
||||
|
||||
| Step | What runs | Where | Model (target) | Δ vs today |
|
||||
|---|---|---|---|---|
|
||||
| MODEL GATE | witness + self-check | main loop | session (Fable; Opus admitted) | — |
|
||||
| 0 plugin check | detection probes | dispatched (plugin-advisor detection half) | haiku | today inline on session |
|
||||
| 0 plugin check | complexity scoring + reco | dispatched (advisor judgment half) | opus | today inline on session |
|
||||
| 0 plugin check | apply gate (user) | main loop | Fable | — |
|
||||
| 0b/0c context + ctx7 | trivial bash probes | main loop | Fable (trivial) | — |
|
||||
| 0d read-before digest | analyzer | dispatched | opus | pinned (BDR-076) |
|
||||
| 0e contract | contract-interview + micro-gates | main loop | Fable | — |
|
||||
| 1 brainstorm | superpowers:brainstorming | main loop | Fable | — |
|
||||
| 2 plan | superpowers:writing-plans | main loop | Fable | — |
|
||||
| 2b challenge | 3× plan-challenger | dispatched | opus | pinned |
|
||||
| 2b synthesis + RE-THINK | severity merge, plan revision | main loop | Fable | — |
|
||||
| 3 validation gate | human gate | main loop | Fable | — |
|
||||
| 4 SDD implement | per-task implementers + reviewers | dispatched | sonnet (explicit `model:"sonnet"`) | — |
|
||||
| 4 task decomposition / verdict arbitration | SDD driver | main loop | Fable | — |
|
||||
| 4b error diagnosis | analyzer DEBUG (inline) | main loop | Fable (reflection on the solution) | — |
|
||||
| 5 verify + secure | verifier, security-auditor (fresh) | dispatched | sonnet | — |
|
||||
| 5 loop decisions | ECARTS/BLOCK routing | main loop | Fable | — |
|
||||
| 6 code review | reviewer (superpowers) | dispatched | **opus explicit** | today INHERITS (leak) |
|
||||
| 7 capitalize | registry gate + commit | main loop | Fable | — |
|
||||
| 8 doc sync | doc-syncer | dispatched | sonnet | today INLINE on session |
|
||||
| 9 finish | gitflow + human go | main loop | Fable | — |
|
||||
|
||||
## TARGET MODEL MAP — init-project (example, per-step)
|
||||
|
||||
| Step | What runs | Where | Model (target) | Δ vs today |
|
||||
|---|---|---|---|---|
|
||||
| MODEL GATE | witness + self-check | main loop | session (Fable; Opus admitted) | — |
|
||||
| 0 plugin check | detection / scoring / gate | dispatched haiku / dispatched opus / main loop Fable | (as ship-feature) | today inline |
|
||||
| 1 interview | interviewer (interactive Q&A) | main loop (inline — a dispatched agent cannot ask) | Fable | structural |
|
||||
| 1 contract | contract-interview | main loop | Fable | — |
|
||||
| 2 analyze brief | analyzer (inline — greenfield design reflection) | main loop | Fable | stays inline |
|
||||
| 3 design | superpowers:brainstorming | main loop | Fable | — |
|
||||
| 4 gate #1 + contract enrich | human gate | main loop | Fable | — |
|
||||
| 5 scaffold | scaffolder | **dispatched** | sonnet (effort: high) | today INLINE on session — pin inert |
|
||||
| 5b readme bootstrap | doc-syncer | **dispatched** | sonnet | today INLINE |
|
||||
| 5c/5e/5f ctx7 + anim + gitflow init | deterministic bash | main loop | Fable (trivial) | — |
|
||||
| 6 plan | superpowers:writing-plans | main loop | Fable | — |
|
||||
| 6b challenge + synthesis | 3× plan-challenger / merge | dispatched opus / main loop Fable | — | pinned |
|
||||
| 7 gate #2 | human gate | main loop | Fable | — |
|
||||
| 8 SDD implement | implementers + reviewers | dispatched | sonnet | — |
|
||||
| 8b graphify | bash | main loop | Fable (trivial) | — |
|
||||
| 9 verify + secure | verifier, security-auditor | dispatched | sonnet | — |
|
||||
| 10 code review | reviewer | dispatched | **opus explicit** | today INHERITS (leak) |
|
||||
| 10b capitalize founding BDRs | registry gate | main loop | Fable | — |
|
||||
| 10c doc sync | doc-syncer | **dispatched** | sonnet | today INLINE |
|
||||
| 11 finish | gitflow + human go | main loop | Fable | — |
|
||||
|
||||
## RELATED MEMORY
|
||||
|
||||
- IN FORCE: BDR-066 — model routing waves 1-4 — the architecture being
|
||||
re-tiered; its rationale table is the baseline [accepted]. BDR-076 — opus
|
||||
pins on dispatched judgment — starting state, partially superseded by the
|
||||
new target [accepted, this branch]. BDR-050 — verify+secure loops in main
|
||||
loop, gates fresh [accepted]. BDR-049 — verifier fresh+blind+disk-contract
|
||||
[accepted]. BDR-048 — pinned semgrep gate [accepted]. BDR-061 — fix-bundle
|
||||
→ L1 apply, analyzers have no Agent tool [accepted]. BDR-060 — nested
|
||||
dispatch floor v2.1.172 [accepted]. BDR-075+amendment — challenge phase in
|
||||
12 orchestrators [accepted]. BDR-025 — unknown never silently passes
|
||||
[accepted]. BDR-022 — doc-syncer never touches .claude/ [accepted].
|
||||
LRN-125 — no dual-use across tiers. LRN-126 — splits sever implicit data
|
||||
paths; forward every consumed field. LRN-113 — whole-surface sweep + guard.
|
||||
LRN-083 — loops in main loop. LRN-093 — no `\n` in grep locks. LRN-096 —
|
||||
flip-test new guards. LRN-112 — nesting supported. LRN-105/107 — explicit
|
||||
tool bans in read-only mandates. LRN-011 — one subagent, N gated scores
|
||||
(alternative to 3-way split). LRN-057 — match mechanism to consumer.
|
||||
LRN-102 — final-text-only rendering guarantee. LRN-132 — subagent claims
|
||||
need verification.
|
||||
- ALREADY SEEN: BLK-004 — renamed/deleted agent files broke a consumer wrapper
|
||||
[resolved] (rename sweep discipline). EVAL-023 — BDR-066 post-merge ronde
|
||||
found 5 edge gaps [done] (plan a ronde here too). EVAL-026 — 3-way plan
|
||||
challenge caught 4 real BLOCKERs on its own plan [done] (run it on this
|
||||
refactor's plan).
|
||||
- NON-BINDING: ~200 remaining headings surfaced nothing binding beyond the
|
||||
above — BDR-067/068/069 (release/permissions), LRN-first-100 (tooling),
|
||||
BLK-005..017 (env) — counted, not detailed.
|
||||
- SELECTION: scanned ~230 headings — surfaced 28 = in-force 22 + seen 3 +
|
||||
non-binding (counted).
|
||||
@@ -0,0 +1,275 @@
|
||||
# PLAN: model-tiering v2 — full framework re-tier + splits
|
||||
|
||||
Input: `.claude/tasks/plans/2026-07-19-model-tiering-v2-analysis.md` (read it
|
||||
first — consumer map, test locks, LRN/BDR constraints live there).
|
||||
User arbitrage (2026-07-19): 7 recos approved + no-inherit/fable-pin amendment
|
||||
+ Fable scope = REFLECTION / ORCHESTRATION / PLANNING / LOGIC only.
|
||||
|
||||
## D0 — DOCTRINE (end state)
|
||||
|
||||
1. Main loop (session model, gated big by model-gate) keeps ONLY: brainstorm,
|
||||
plan, contract, loop decisions, gate arbitration, human interaction,
|
||||
conversation-context work (capitalize), trivial glue bash (<~1k tokens).
|
||||
Retention criteria (any suffices): interactive | needs conversation context
|
||||
| orchestration decision | dispatch overhead > step cost.
|
||||
2. NOTHING dispatched inherits. Typed agents: frontmatter pin. Built-ins
|
||||
(general-purpose/Explore/Plan): explicit `model=` at EVERY call site.
|
||||
VERIFIED (2026-07-19 spike, closes robustness BLOCKER): `model: "fable"`
|
||||
on a dispatch resolves to claude-fable-5 at runtime (echo spike via
|
||||
general-purpose); the harness enum-validates the `model` param — an
|
||||
invalid value fails LOUDLY (InputValidationError), no silent fallback.
|
||||
Call-site `model=` takes precedence over a typed agent's frontmatter pin
|
||||
(documented Agent-tool contract); fallback direction if a call site omits
|
||||
it = the frontmatter pin, i.e. today's behavior — fail-safe, never worse.
|
||||
3. Tiers: fable = dispatched reflection-on-behalf-of-main-loop (skill-runner
|
||||
children ONLY); opus = deep judgment (audit scoring, plan critique, drift
|
||||
semantics, review, synthesis); sonnet = standard execution from closed
|
||||
instructions + collectors AND probes (wave-1 prudence — robustness MAJOR:
|
||||
plugin PHASE 1 is a ~26-call branching bash chain, not a short probe);
|
||||
haiku = status-reporter ONLY in wave 1; haiku expansion = wave 2 after
|
||||
reliability proven per candidate.
|
||||
4. Grammars/sentinels/valves survive VERBATIM (list in analysis §CONSTRAINTS).
|
||||
Loops/gates stay in main loop (BDR-050/LRN-083). Fix-bundle → L1 apply
|
||||
(BDR-061) preserved: audit agents never get the Agent tool.
|
||||
5. Every split: LRN-126 data-flow pass (enumerate child-read fields vs
|
||||
parent-set; explicit handoff contract on disk or in prompt) PLUS an
|
||||
IN-WAVE planted-input smoke proving the fields cross the dispatch boundary
|
||||
at runtime — the smoke GATES that wave's merge (confirmation MAJOR:
|
||||
enumeration is design-time reading; census can't catch severed wires; a
|
||||
split must never reach develop empirically unproven). Every change:
|
||||
LRN-113 whole-surface sweep + census lock + flip-test (LRN-096, no `\n` in
|
||||
patterns LRN-093, strict YAML G3).
|
||||
|
||||
## D1 — AGENT END STATE
|
||||
|
||||
Pins (frontmatter):
|
||||
- opus: analyzer, plan-challenger, seo-judge*, geo-judge*, doc-auditor*,
|
||||
plugin-reasoner*, handover-synthesizer* (*new, from splits)
|
||||
- sonnet: feater, bugfixer, hotfixer, code-cleaner, refactorer, verifier,
|
||||
security-auditor, scaffolder (effort high), onboarder, release-executor,
|
||||
commit-changer, doc-syncer (patcher half), validator-analyzer (TIER-DOWN
|
||||
from opus), seo-worker*, geo-worker* (2-way split per domain — simplicity
|
||||
MAJOR: collector+templater both sonnet in wave 1 → one worker file with
|
||||
`MODE: collect | template`, no cross-domain share: domain bodies genuinely
|
||||
diverge), handover-renderer* (renamed handover-doc-writer render half),
|
||||
plugin-probe* (wave-1 prudence; haiku candidate wave 2)
|
||||
- haiku: status-reporter (only)
|
||||
- none (inline-only, main loop, gate-protected): interviewer,
|
||||
client-handover-writer
|
||||
Per-dispatch `model=` overrides (no new file): commit-changer propose=opus /
|
||||
apply=sonnet (2 sites in /commit-change — precedence over the sonnet
|
||||
frontmatter pin is the documented Agent-tool contract, verified direction
|
||||
D0.2; the pin stays as the no-inherit fallback = today's behavior; both
|
||||
call-site strings census-locked + W3 behavioral smoke); SDD
|
||||
implementers+reviewers
|
||||
sonnet (already prose-mandated → make it a census lock); code-review steps
|
||||
(ship-feature 6, init-project 10) = opus explicit; client-handover-writer's 8
|
||||
general-purpose skill-runners = fable; onboard's 7 general-purpose = opus
|
||||
(keep); any Explore/Plan ad-hoc reflection dispatch = fable (doctrine line in
|
||||
model-gate.md + CLAUDE.global routing note).
|
||||
|
||||
Splits (each = new agent file(s) + handoff contract + census + consumers):
|
||||
S1 plugin-advisor → plugin-probe (SONNET wave 1; PHASE 1 CLI probes → PROBE
|
||||
REPORT) + plugin-reasoner (opus; PHASE 2/2.5 scoring + reco → PLUGIN CHECK
|
||||
block). PHASE 3-4 report+apply-gate HOISTED into ONE shared include
|
||||
`lib/plugin-gate.md` (simplicity MINOR — doc-commit.md ×6 pattern, never
|
||||
4 hand-copies), referenced by the 4 consumers (plugin-check, onboard
|
||||
STEP 0, init-project STEP 0, ship-feature STEP 0) — main loop. The
|
||||
pre-recommendation validation checkpoint (advisor :201-212, straddles the
|
||||
seam, can skip PHASE 4) runs IN THE CONSUMER between the two dispatches
|
||||
(correctness MINOR); its inputs (toggle-external availability,
|
||||
project-signal presence) are PROBE REPORT fields. Handoff: PROBE REPORT
|
||||
fields = plugin list, toggle state, profile, CLI/anim/monorepo/embedded
|
||||
signals + checkpoint inputs (enumerate ALL PHASE-2-read fields).
|
||||
S2 doc-syncer → doc-auditor (opus; STEP 3-4 drift + semantic analysis + A3
|
||||
MINOR/SIGNIFICANT call w/ doc-shape.sh oracle → DRIFT REPORT [AUTO]/
|
||||
[HUMAN] items) + doc-syncer (sonnet; render/patch half, keeps
|
||||
PATCHED_FILES: grammar + BDR-022 bans). Validation gate stays in
|
||||
DISPATCHER (/doc skill, orchestrator steps) — auto-mode flows: auditor →
|
||||
dispatcher applies AUTO via doc-syncer → SIGNIFICANT escalates inline.
|
||||
Consumers rerouted: /doc, onboard, + doc-commit steps in bugfix/hotfix/
|
||||
feat/init-project(×2)/ship-feature (inline→dispatch conversion) +
|
||||
scaffolder PHASE 6 (scaffolder DISPATCHES nothing — it has no Agent tool:
|
||||
README bootstrap moves to init-project STEP 5b dispatch of doc-syncer).
|
||||
PLUS (robustness MAJOR): rework `lib/doc-commit.md`'s in-thread contract
|
||||
BEFORE converting any doc-commit site — it requires the orchestrator to
|
||||
"hold the patch context" to compose the rc-0 CHANGE SUMMARY (the review
|
||||
surface that replaced the removed MINOR gate). Dispatched doc-syncer adds
|
||||
a `CHANGE SUMMARY` block to its report grammar (per patched file: what
|
||||
changed and why, ≤1 line each); doc-commit.md's composer consumes THAT
|
||||
instead of in-thread context; census-locks the new field + a planted-input
|
||||
smoke proves the summary crosses the dispatch boundary.
|
||||
S3 seo-analyzer → 2-WAY (simplicity MAJOR — 3-way was YAGNI while collector
|
||||
and templater share the sonnet tier; commit-changer mode-precedent):
|
||||
seo-worker (sonnet; `MODE: collect` = STEP 2-5 signals → SIGNALS file;
|
||||
`MODE: template` = STEP 12-14 FIX BUNDLE + sentinel + SEO.md + envelope)
|
||||
+ seo-judge (opus; STEP 6-11 sampling judgment, competitive, scoring /20,
|
||||
trajectory, triage → FINDINGS+PLAN). Orchestrated by /seo at L1 (BDR-061
|
||||
conserved: no Agent tool in either). Wave-2 option: carve `MODE: collect`
|
||||
into a haiku file once proven — the mode boundary IS the future cut line.
|
||||
HANDOFF (robustness MAJOR — freshness/atomicity): run-scoped paths
|
||||
`.audit/seo-signals-<RUNID>.md` / `.audit/geo-signals-<RUNID>.md` —
|
||||
`.audit/` is the GITIGNORED derived-artifact tree (confirmation MINOR,
|
||||
LRN-124: a crash-stranded transient with scraped GSC/competitor content
|
||||
must never be committable; `.claude/audits/` keeps only the SEO.md/GEO.md
|
||||
deliverables). RUNID minted by the dispatcher per run, passed to every
|
||||
stage; the file ENDS with `COLLECTION COMPLETE — RUNID: <id>` and the
|
||||
judge FAILS CLOSED (report ERROR, never score) if the file is absent,
|
||||
RUNID mismatches, or the completeness sentinel is missing; dispatcher
|
||||
cleans the file post-run.
|
||||
DISPATCHER CONTRACT (confirmation MAJOR — fail-closed at the judge must
|
||||
not fail OPEN at the pipeline): on a judge ERROR the orchestrator
|
||||
(/seo /geo /harden /onboard) STOPS — no template dispatch, no L1 apply —
|
||||
surfaces the ERROR verbatim, retries ONCE with a fresh collect+judge,
|
||||
then escalates to the human. A mute or ERROR judge is NEVER carried into
|
||||
templating (verify-secure-loop discipline). This handler is part of the
|
||||
W5 skill rewrites, census-locked.
|
||||
Explicit field list per LRN-126 (STEP 1-2 business+tech context consumed
|
||||
by ALL later steps — full enumeration REQUIRED before cutting).
|
||||
seo-data.test.sh locks (fetch.sh wiring) move with the worker body —
|
||||
update suite same commit.
|
||||
S4 geo-analyzer → geo-worker (sonnet, 2 modes) + geo-judge (opus) — mirror of
|
||||
S3 incl. run-scoped `.audit/geo-signals-<RUNID>.md` + the same dispatcher
|
||||
ERROR contract. No cross-domain file share:
|
||||
seo vs geo bodies genuinely diverge (different checks, scoring blocks,
|
||||
envelopes) — that divergence, not LRN-125, is the reason.
|
||||
S5 handover-doc-writer → handover-synthesizer (opus; STEP 9 memory-registry
|
||||
load + STEP 10 phase clustering + STEP 12 6-chapter synthesis — STEP 9
|
||||
allocated here, it feeds the synthesis; correctness MINOR) +
|
||||
handover-renderer (sonnet; STEP 13-16 annex render, precheck apply,
|
||||
deterministic gates, HTML/PDF). client-handover-writer dispatches
|
||||
synthesizer then renderer; PACKAGE contract split per LRN-126
|
||||
(re-enumerate DEPLOY_HINTS/--skip-seo class fields — the EXACT prior
|
||||
failure). W4 MUST same-commit relock model-routing.test.sh:52-55 (the
|
||||
handover-doc-writer name + dispatch-string locks break on the rename;
|
||||
"make test green per wave" D4 invariant — correctness MINOR).
|
||||
Tier-downs (no split): validator-analyzer opus→sonnet (deterministic
|
||||
validators+tables). onboarder stays sonnet wave 1 (haiku candidate wave 2).
|
||||
release-executor stays sonnet (NEED-DECISION valve).
|
||||
Verifier: STAYS sonnet (approved — oracle-anchored gate).
|
||||
|
||||
## D2 — SKILL MAP (main loop = session model; every dispatch tier explicit)
|
||||
|
||||
Gated reflection skills (model-gate kept, 15):
|
||||
- ship-feature / init-project: per the two example maps in the analysis file
|
||||
(amendment section) + S1 gate hoist at STEP 0 + doc-commit conversions.
|
||||
- feat: scope/plan/contract/loop = main; challenge 3× plan-challenger opus;
|
||||
feater sonnet; verifier+security sonnet; doc-commit → doc-auditor opus +
|
||||
doc-syncer sonnet dispatch; commit via /commit-change (propose opus / apply
|
||||
sonnet).
|
||||
- bugfix: investigation/diagnosis/contract = main (reflection); challenge
|
||||
opus (3b); bugfixer sonnet; verifier+security sonnet; doc-commit as feat.
|
||||
- hotfix: LOCATE + guard = main (logic); challenge opus when guard fires;
|
||||
hotfixer sonnet; security gate sonnet (revert-not-loop conserved);
|
||||
doc-commit as feat.
|
||||
- analyze: analyzer INLINE = main loop (it IS the reflection) — unchanged.
|
||||
- code-clean: PHASE 1 audit inline = main (audit judgment feeding a human
|
||||
gate); code-cleaner sonnet PHASE 2 (hosts refactorer inline at SAME tier —
|
||||
LRN-125 OK); re-audit sonnet inside executor.
|
||||
- seo / geo: skill = orchestration + GATED arbitrage (main); pipeline
|
||||
collector sonnet → judge opus → templater sonnet (L1 serial); appliers
|
||||
hotfixer/feater sonnet at L1; build-verify inline.
|
||||
- web-validate: validator-analyzer sonnet; hotfixer applier sonnet; loop main.
|
||||
- harden: audit dispatch follows S3 narrow-scope path (seo-judge opus on
|
||||
harden axes w/ collector reuse); direct-Edit apply stays inline (tiny
|
||||
scope, BDR-061 carve-out conserved).
|
||||
- audit-delta: axis audits dispatched opus (delta judgment); security-auditor
|
||||
sonnet; fix gate + markers = main.
|
||||
- tour: orchestration main; security-auditor sonnet; cleanup audit = analyzer
|
||||
opus (or general-purpose model="opus"); fixes via sonnet appliers; doc axis
|
||||
→ S2 pipeline; reconcile axis = deterministic bash (main).
|
||||
- onboard: onboarder DISPATCHED sonnet (was inline); plugin S1 pipeline;
|
||||
analyzer opus; general-purpose audits model="opus" (kept); seo/geo → S3/S4
|
||||
pipelines; security-auditor + doc pipeline as above; synthesis
|
||||
general-purpose model="opus"; backlog arbitration = main.
|
||||
- client-handover: writer INLINE (orchestrator, main); its 8 skill-runner
|
||||
children model="fable"; handover S5 split (synth opus → render sonnet);
|
||||
gates all main.
|
||||
Excluded-from-gate skills (5, stay ungated): commit-change (propose opus /
|
||||
apply sonnet via model=; approval gates main); doc (S2: auditor opus →
|
||||
gate main → patcher sonnet); status (haiku); release-candidate (executor
|
||||
sonnet; version/when/push decisions main); refactor (refactorer sonnet).
|
||||
Memory/util skills (capitalize, close, prune-memory, reconcile, learn,
|
||||
profile, skills-perso, gitflow, deploy, plugin-check(S1), status): main
|
||||
loop by nature (conversation context, human gates, deterministic bash) —
|
||||
no dispatch changes except plugin-check S1.
|
||||
External/gstack skills (graphify, design-*, review, qa, ship, investigate…):
|
||||
NOT edited (external ownership, BDR-015 class) — covered by doctrine line;
|
||||
local wrapper skills only if locally owned. Verify ownership per file
|
||||
before touching (symlink → skip).
|
||||
|
||||
## D3 — WAVES (each = gitflow feature branch, tests green, census extended)
|
||||
|
||||
W0 SEQUENCING: merge `feature/opus-pin-audit-agents` → develop (baseline,
|
||||
human gate). `bugfix/seo-geo-integrity` is ALREADY MERGED (correctness
|
||||
MAJOR — the TODO.md "UNMERGED" note was stale; verified `92301fe` is an
|
||||
ancestor of develop AND this branch): no arbitrage, no W5 wait — one-line
|
||||
ancestry re-check in W0 + fix the stale TODO.md entry (reconcile-class
|
||||
correction). Absorb the analysis+plan files into the new feature branch.
|
||||
W1 NO-INHERIT ENFORCEMENT (small, high-value): code-review model= opus
|
||||
(ship-feature 6, init-project 10); client-handover-writer 8× model="fable";
|
||||
doctrine line in model-gate.md + census locks (`model="fable"`,
|
||||
`model=` presence per site); SDD sonnet prose → census lock. Prose sweep
|
||||
of stale BDR-066/076 claims touched by W1.
|
||||
W2 INLINE→DISPATCH CONVERSIONS: scaffolder (init 5 — liveness pings move to
|
||||
orchestrator; scaffolder loses PHASE 6 inline-load → init 5b owns README
|
||||
via S2), onboarder (onboard), doc-commit steps ×5 flows → S2 pipeline
|
||||
(gate hoist FIRST: /doc + flows own the validation gate; doc-syncer body
|
||||
loses its inline gate → census re-lock), S1 plugin split + gate hoist ×4
|
||||
consumers. Data-flow pass per LRN-126 on each (fields enumerated in the
|
||||
wave's contract file before edits).
|
||||
W3 TIER MOVES: validator-analyzer → sonnet (pin + prose + census flip);
|
||||
commit-changer per-mode model= (2 sites + prose + census).
|
||||
W4 S5 handover split (synth opus / render sonnet) + PACKAGE re-enumeration.
|
||||
W5 S3/S4 seo/geo pipelines: worker(2-mode)/judge ×2, /seo /geo /harden
|
||||
/onboard rerouted, seo-data.test.sh moved locks, run-scoped signals
|
||||
handoff (RUNID + completeness sentinel + fail-closed judge),
|
||||
envelope/sentinel/score grammars verbatim, COVERAGE lines preserved.
|
||||
W6 DOCTRINE + CLOSE-OUT: model-gate.md rewrite (protects main loop; tier
|
||||
table; fable-dispatch doctrine), challenge-plan.md + plan-challenger
|
||||
ORCHESTRATOR PROTOCOL text (keep BDR-066+BDR-076 tokens per census, add
|
||||
BDR-077), census consolidation (model-routing new sections; every new
|
||||
agent: YAML G3, pin lock, dispatch-string lock, AskUserQuestion/Agent
|
||||
bans), LRN-113 whole-surface prose sweep (~30 refs list in analysis),
|
||||
BDR-077 + LRN entries + journal, EVAL-023-style post-merge ronde.
|
||||
Per-split planted-input smokes run IN their own waves (W2/W4/W5, merge
|
||||
gates) — W6 is the consolidated ronde only, never the first empirical
|
||||
proof of a split.
|
||||
|
||||
## D4 — ZERO-REGRESSION PROTOCOL (every wave)
|
||||
|
||||
- Before edits: wave contract file (.claude/tasks/contracts/) with FILE SCOPE
|
||||
+ acceptance criteria; challenge-plan on THIS plan (done once, below);
|
||||
verify-secure-loop on each wave's diff (verifier sonnet + security sonnet).
|
||||
- Grammar diff-guard: `grep -F` each verbatim marker (analysis §CONSTRAINTS
|
||||
list) pre/post per wave — zero drift.
|
||||
- Census: flip-test every NEW lock (plant violation → RED) before trusting.
|
||||
- `make test` green per wave; no wave merges without human signal (gitflow).
|
||||
- Rollback story (robustness MINOR — waves are textually interdependent, an
|
||||
early wave is NOT independently revertible after later merges): revert in
|
||||
REVERSE merge order, or revert the whole stack; never a mid-stack single
|
||||
revert. Pre-merge, the rollback unit is the wave branch.
|
||||
|
||||
## CHALLENGE LOG (2026-07-19 — 3 blind lenses on plan v1)
|
||||
|
||||
- correctness: CONCERNS(2) — seo-geo-integrity phantom sequencing (fixed W0);
|
||||
commit-changer precedence ambiguity (fixed D1 + D0.2 citation + W3 smoke);
|
||||
3 MINORs (S5 STEP 9 + W4 relock; plugin checkpoint seam; templater label)
|
||||
— all fixed in place.
|
||||
- robustness: FATAL(4) — BLOCKER fable-dispatch unverified → CLOSED by spike
|
||||
(D0.2: resolves to claude-fable-5, enum-validated, loud failure); doc-commit
|
||||
in-thread contract (fixed S2: CHANGE SUMMARY crosses the report grammar);
|
||||
plugin-probe haiku contradiction (fixed: sonnet wave 1); signals handoff
|
||||
freshness (fixed S3: RUNID + sentinel + fail-closed); rollback claim
|
||||
(fixed D4).
|
||||
- simplicity: CONCERNS(1) — 3-way seo/geo YAGNI → 2-way worker/judge (fixed
|
||||
S3/S4); twin-templater share (dissolved by 2-way; divergence stated);
|
||||
plugin gate ×4 copies → lib/plugin-gate.md include (fixed S1).
|
||||
- Confirmation pass (fresh robustness challenger on v2): CONCERNS(2) — v1
|
||||
fixes HOLD (doc-commit CHANGE SUMMARY, plugin-probe sonnet, rollback order,
|
||||
fable spike, RUNID); 2 new MAJORs + 1 MINOR opened by the revisions, all
|
||||
fixed in v3: (a) per-split planted-input smokes moved IN-WAVE as merge
|
||||
gates (W6 = ronde only); (b) dispatcher ERROR contract on judge failure
|
||||
(STOP, no templating/apply, retry once, escalate — pipeline never fails
|
||||
open); (c) transient signals files relocated to gitignored `.audit/`
|
||||
(LRN-124). Protocol cap reached (1 re-challenge) → to the human gate.
|
||||
@@ -257,7 +257,12 @@ pipeline is reduced: only run /cso (single audit, single fix loop), skip
|
||||
STEP 6 deploy pause and STEP 7 /web-validate. Treat /cso as the only score for
|
||||
the gate.
|
||||
|
||||
For web projects, dispatch in **a single message with two parallel Agent calls**:
|
||||
**Model routing (BDR-077):** EVERY `general-purpose` skill-runner dispatch in
|
||||
this pipeline (initial audits, fix-loop re-dispatches, commit-change,
|
||||
web-validate) carries `model: "fable"` — the child hosts gated orchestration
|
||||
on the pipeline's behalf; it must never inherit the session model.
|
||||
|
||||
For web projects, dispatch in **a single message with two parallel Agent calls** (each with `model: "fable"`):
|
||||
|
||||
| Audit (web) | Subagent | Prompt template |
|
||||
|---------------|-------------------|-----------------|
|
||||
@@ -383,7 +388,7 @@ console). If no projected line is parseable, treat projected = 17
|
||||
|
||||
### Re-dispatch prompt template (SEO + GEO loop)
|
||||
|
||||
Send to `general-purpose` subagent:
|
||||
Send to `general-purpose` subagent (`model: "fable"`):
|
||||
|
||||
> Read `~/.claude/skills/seo/SKILL.md` and re-run it on this project.
|
||||
> Previous scores:
|
||||
@@ -413,7 +418,7 @@ Send to `general-purpose` subagent:
|
||||
|
||||
### Re-dispatch prompt template (HARDEN loop)
|
||||
|
||||
Send to `general-purpose` subagent:
|
||||
Send to `general-purpose` subagent (`model: "fable"`):
|
||||
|
||||
> Read `~/.claude/skills/harden/SKILL.md` and re-run it. Previous score:
|
||||
> **`<SCORE_HARDEN_PREVIOUS>`/20** — below threshold. Iteration `<N>` of
|
||||
@@ -424,7 +429,7 @@ Send to `general-purpose` subagent:
|
||||
|
||||
### Re-dispatch prompt template (CSO loop — non-web only)
|
||||
|
||||
Send to `general-purpose` subagent:
|
||||
Send to `general-purpose` subagent (`model: "fable"`):
|
||||
|
||||
> Read `~/.claude/skills/cso/SKILL.md` and re-run it in **daily mode**.
|
||||
> Previous score: **`<SCORE_CSO_PREVIOUS>`/20** — below threshold.
|
||||
@@ -510,7 +515,7 @@ listed changes manually before deploy." Continue to STEP 6.
|
||||
|
||||
If `PENDING_CHANGES` non-empty → invoke /commit-change skill via subagent:
|
||||
|
||||
> Dispatch `general-purpose` subagent. Prompt:
|
||||
> Dispatch `general-purpose` subagent (`model: "fable"`). Prompt:
|
||||
>
|
||||
> "Read `~/.claude/skills/commit-change/SKILL.md` and execute. All pending
|
||||
> changes were produced by the client-handover ship pipeline during the
|
||||
@@ -617,7 +622,7 @@ Skip if `VALIDATE_SKIPPED=true` or `PROJECT_TYPE != web` (in either case
|
||||
ensure `VALIDATE_SKIPPED=true` is set so the gate logic in STEP 8 treats
|
||||
VALIDATE as not-applicable rather than failed).
|
||||
|
||||
Dispatch `general-purpose` subagent:
|
||||
Dispatch `general-purpose` subagent (`model: "fable"`):
|
||||
|
||||
> Read `~/.claude/skills/web-validate/SKILL.md` and execute against the
|
||||
> deployed URL: `<DEPLOYED_URL>`. Audit W3C HTML validity (validator.nu),
|
||||
|
||||
@@ -35,3 +35,13 @@ yet rewritten) — that is why the self-check exists alongside it.
|
||||
|
||||
then end the turn. No later step runs, no agent is dispatched, nothing is
|
||||
edited.
|
||||
|
||||
## 4. Dispatch tiers (BDR-077 — no inherit)
|
||||
|
||||
The gate guards the MAIN loop only. Dispatched work NEVER inherits the
|
||||
session model: typed agents run on their frontmatter pin; built-ins
|
||||
(general-purpose / Explore / Plan) carry an explicit `model=` at every call
|
||||
site — `model: "fable"` when the child performs reflection/orchestration on
|
||||
the main loop's behalf (skill-runners), otherwise its complexity tier
|
||||
(opus = dispatched judgment, sonnet = execution/collection, haiku = short
|
||||
mechanical probes).
|
||||
|
||||
@@ -72,6 +72,15 @@ fm_lacks "agents/interviewer.md" 'model:'
|
||||
has "skills/onboard/SKILL.md" 'model="opus"'
|
||||
has "skills/tour/SKILL.md" 'model="opus"'
|
||||
has "lib/challenge-plan.md" 'BDR-076'
|
||||
# 12) BDR-077 W1 — no-inherit: skill-runner children pinned fable at every
|
||||
# call site; code-review dispatches carry opus; doctrine in model-gate.
|
||||
# (fable dispatch alias spike-verified 2026-07-19: resolves
|
||||
# claude-fable-5, enum-validated, loud failure — never silent fallback)
|
||||
has "agents/client-handover-writer.md" 'model: "fable"'
|
||||
has "skills/ship-feature/SKILL.md" 'model: "opus"'
|
||||
has "skills/init-project/SKILL.md" 'model: "opus"'
|
||||
has "lib/model-gate.md" 'model: "fable"'
|
||||
lacks "lib/model-gate.md" 'model: "sonnet" in the Agent call'
|
||||
|
||||
printf 'model-routing census: %d pass, %d fail\n' "$pass" "$fail"
|
||||
[ "$fail" -eq 0 ]
|
||||
|
||||
@@ -223,7 +223,10 @@ against the founding contract. Distinct axis from STEP 10 code review
|
||||
([[LRN-095]]) — both run.
|
||||
|
||||
## STEP 10 — CODE REVIEW
|
||||
Invoke `superpowers:requesting-code-review`. Fix all CRITICAL before proceeding.
|
||||
Invoke `superpowers:requesting-code-review`. **Model routing (BDR-077):** the
|
||||
review subagent it dispatches MUST carry `model: "opus"` in the Agent call —
|
||||
craft review is dispatched judgment, never inherited from the session. Fix
|
||||
all CRITICAL before proceeding.
|
||||
|
||||
## STEP 10b — CAPITALIZE FOUNDING DECISIONS (memory registries)
|
||||
A greenfield's founding architecture decisions are the highest-value BDRs — the
|
||||
|
||||
@@ -222,7 +222,10 @@ conformity + security vs. craft/design) — both run, neither subsumes the
|
||||
other ([[LRN-095]]).
|
||||
|
||||
## STEP 6 — CODE REVIEW
|
||||
Invoke `superpowers:requesting-code-review`. Fix all CRITICAL before proceeding.
|
||||
Invoke `superpowers:requesting-code-review`. **Model routing (BDR-077):** the
|
||||
review subagent it dispatches MUST carry `model: "opus"` in the Agent call —
|
||||
craft review is dispatched judgment, never inherited from the session. Fix
|
||||
all CRITICAL before proceeding.
|
||||
|
||||
## STEP 7 — CAPITALIZE (memory registries)
|
||||
Feature shipped implies at least one design decision worth capturing. Run this BEFORE STEP 9 FINISH — the implementation commits (STEP 4) already exist, so the entries' hash references are valid, and the memory commit lands on the branch that FINISH integrates (otherwise it strands outside the PR):
|
||||
|
||||
Reference in New Issue
Block a user