Skip to content

Subagent model tiering

Cost-aware routing of Claude Code work to different model tiers — Haiku for searches, Sonnet for implementation, Opus for planning — so that token budget is spent on the steps where reasoning actually pays off.

Subagent model tiering is the discipline of matching the model to the work. A grep across a 200-file solution does not need Opus; a planning pass that decides which migration path to commit to does. The CLAUDE_CODE_SUBAGENT_MODEL environment variable and opusplan mode are the official knobs; this entry catalogues how Code Majesty configures them.

The configuration is one environment variable plus per-agent overrides:

# default all subagent traffic to Sonnet unless an agent specifies otherwise
export CLAUDE_CODE_SUBAGENT_MODEL=claude-sonnet-5

combined with model: frontmatter in .claude/agents/*.md for individual agents, and opusplan mode so Opus runs the planning pass while implementation continues on Sonnet. What we send where:

TierWorkWhy
HaikuSearches, file inventory, symbol lookups, log scansVolume is high, reasoning need is near zero
SonnetImplementation, test writing, refactors inside a verified planThe plan already carries the hard decisions
OpusPlanning passes, migration-path decisions, failure post-mortemsErrors here are the expensive ones

The observable effect: per-PR token cost becomes a function of plan size rather than session length, because the long mechanical middle of a session runs on the cheap tiers.

When to use

  • Any agent setup where token spend is non-trivial and the team wants predictable per-PR cost.
  • Sessions long enough that the marginal reasoning gain of Opus on every tool call is not worth the budget hit.

Distinguish from

A flat single-model setup is simpler to operate but spends Opus tokens on Haiku-grade work. Tiering needs the upfront config but typically pays back within the first long session.

Related terms