# Model capability tiers for typed-pipeline routing. # # Each tier defines score floors (against model-registry/*.yaml `scores` dimensions). # The CP picks the cheapest registry entry meeting ALL floors (MR-9/MR-11 — cheapest-wins). # # Stages that need a specific tier set `tier:` on their `kind: agent` node (W2). # Until W2 lands, templates hardcode `model:` directly and these floors serve as # documentation of the intended constraint. # # See claude/stage-routing.md for the authoritative pipeline stage map. tiers: planning: description: > Judgment-heavy work: plan generation, spec re-draft (replan-spec-opus), architecture review escalation (review-spec-opus). Requires sustained reasoning over large context windows. min_scores: complexity: 9 creativity: 9 context_utilisation: 9 qualifies_today: - deepseek-v4flash # benchmark-inflated 9s; hold to allowed_models ceiling - minimax - claude-opus-4 note: > WF-SCH-14 allowed_models ceiling remains authoritative until M7 observed-score recalibration corrects benchmark-inflated registry entries. spec-test: description: > Spec-adherent structured work: spec drafting, spec arch review, scope/decompose, test writing, impl review, merge verdict. Needs reliable spec adherence and test execution. min_scores: spec_adherence: 9 test_pass_rate: 9 qualifies_today: - minimax - claude-sonnet-4 - claude-opus-4 coding: description: > Narrow TDD implementation against pre-validated tests. Lower floors because tests already constrain the solution space. min_scores: spec_adherence: 7 test_pass_rate: 7 qualifies_today: - airouter-qwen3 - airouter-deepseekv4flash - minimax - claude-sonnet-4 - claude-opus-4