feat(routing): add tiers.yaml + rename spec/review-arch templates to airouter

W1 stage-routing work:
- model-registry/tiers.yaml: defines planning/spec-test/coding tier floors
  (complexity≥9/creativity≥9/context≥9 | spec_adherence≥9/test_pass_rate≥9 |
  spec_adherence≥7/test_pass_rate≥7) with qualifies_today lists
- spec-draft-airouter@2.yaml: honest name for what was spec-draft-opus@1
  (always ran Qwen3.6/airouter, not Opus)
- review-spec-arch-airouter@1.yaml: honest name for review-spec-arch-opus@1

Old @1 files kept for in-flight task backward-compatibility.

Claude-Session: https://claude.ai/code/session_01B35bPAKv5uyW1F9gzMcRN7
This commit is contained in:
Paul O'Reilly
2026-07-12 19:03:51 +12:00
parent 681f0ce4fa
commit 78229c5dba
3 changed files with 152 additions and 0 deletions

56
model-registry/tiers.yaml Normal file
View File

@@ -0,0 +1,56 @@
# Model capability tiers for typed-pipeline routing.
#
# Each tier defines score floors (against model-registry/*.yaml `scores` dimensions).
# The CP picks the cheapest registry entry meeting ALL floors (MR-9/MR-11 — cheapest-wins).
#
# Stages that need a specific tier set `tier:` on their `kind: agent` node (W2).
# Until W2 lands, templates hardcode `model:` directly and these floors serve as
# documentation of the intended constraint.
#
# See claude/stage-routing.md for the authoritative pipeline stage map.
tiers:
planning:
description: >
Judgment-heavy work: plan generation, spec re-draft (replan-spec-opus),
architecture review escalation (review-spec-opus). Requires sustained
reasoning over large context windows.
min_scores:
complexity: 9
creativity: 9
context_utilisation: 9
qualifies_today:
- deepseek-v4flash # benchmark-inflated 9s; hold to allowed_models ceiling
- minimax
- claude-opus-4
note: >
WF-SCH-14 allowed_models ceiling remains authoritative until M7 observed-score
recalibration corrects benchmark-inflated registry entries.
spec-test:
description: >
Spec-adherent structured work: spec drafting, spec arch review, scope/decompose,
test writing, impl review, merge verdict. Needs reliable spec adherence and
test execution.
min_scores:
spec_adherence: 9
test_pass_rate: 9
qualifies_today:
- minimax
- claude-sonnet-4
- claude-opus-4
coding:
description: >
Narrow TDD implementation against pre-validated tests. Lower floors because
tests already constrain the solution space.
min_scores:
spec_adherence: 7
test_pass_rate: 7
qualifies_today:
- airouter-qwen3
- airouter-deepseekv4flash
- minimax
- claude-sonnet-4
- claude-opus-4