Commit Graph
4 Commits
Author SHA1 Message Date
bchanot 1f2d33b7a6 feat(model-router): wave 2-B — orchestrators declare phases, shifters and pins removed, frontmatter = off-state floor
The 15 Skill(effort-*) citers now call mcp__model-router__route per phase
(orchestrate at a dispatch span, reflect/plan for the skill's own level,
apply at the bookkeeping tail, escalate at the verify-secure caps); built-in
judgment dispatches carry an explicit effort= param. lib/effort-shift.md is
the route doctrine, lib/model-gate.md the mod rule (route answer = witness,
/route on as remedy). Deleted: skills/effort-*, lib/effort-pins.txt/.sh,
lib/model-check.sh, their tests, the installers' re-apply blocks. The mod
drops its Skill(effort-*) bridge. The tracked model:/effort: frontmatter
stays as the off-state floor, census-locked equal to the rows
(lib/tests/effort-routing.test.sh rewritten, 140 checks; analyzer → xhigh).

Contract .claude/tasks/contracts/2026-10-10-model-router-w2b-1045.md, plan
r4 § W2-B: GATE 0 MET, verifier ECARTS(7) then CONFORME 10/10, security
PASS, full make test green (design-tool-gate env red only).
2026-10-10 11:24:54 +02:00
bastien 557e4cc317 feat(effort): max at the verify-secure caps and ship-feature 4b; STOP texts suggest /effort-max 2026-09-28 20:02:56 +02:00
Bastien Chanot 354ff2644f feat(agents): pin dispatched judgment agents to opus — Fable = inline reflection only (BDR-076)
Reverses the BDR-066 rejected alternative (opus pins on audit agents):
session default is now Fable, so inherit burned Fable quota on every
dispatched audit/challenge. analyzer, plan-challenger, seo/geo/
validator-analyzer pinned model: opus; onboard's 6 general-purpose
audit dispatches carry model="opus"; tour Phase B repointed.
interviewer + client-handover-writer stay unpinned (inline-load only,
a pin there is inert). settings.json default: claude-fable-5[1m].
Census flipped: model-routing §3 + new §11 (61 pass), loops-light 35,
full make test green.
2026-07-19 17:38:55 +02:00
Bastien Chanot 6bfc0543e5 feat(skills): add 3-way adversarial plan-challenge phase to reflection orchestrators
After a plan/reflection is elaborated and before it executes, three fresh blind
sub-agents (correctness / robustness / simplicity) attack it on the big model;
the main loop RE-THINKS every aspect a BLOCKER lands (a named plan change, or
[deferred]) and re-challenges once if the plan materially changed. Advisory into
each skill's existing human gate — the human stays the decider.

- lib/challenge-plan.md — reusable phase: fail-safe (never fail open),
  severity-driven (any single-lens BLOCKER = must-address), RE-THINK loop
- agents/plan-challenger.md — challenger role (read-only, big-model per BDR-066)
- lib/tests/plan-challenger.test.sh — 41-assertion structure lock
- wired into 11 orchestrators: ship-feature/init-project/feat/bugfix (build-plan),
  onboard/audit-delta/code-clean (proposals), seo/geo/harden/web-validate (fix-bundle)

Hardened by dogfooding: 3 blind challengers reviewed this feature's own v1 plan
and caught 4 BLOCKERs (fail-open, consensus-buries-lone-finding, wrong model
tier vs BDR-066, false on-disk-plan premise) — all fixed here.
2026-07-17 22:51:50 +02:00