Mid-range is the production spend band: expensive enough that you should measure it, cheap enough that it can carry daily work if the model fits.
Use this page as a routing guide, not a leaderboard. The right answer depends on the tool path, data policy, context size, and cost per successful task.
Current Shortlist
| Lane | Pricing anchor | Best use | Caveat |
|---|---|---|---|
| Claude Sonnet 5 | $2 / $10 current provider anchor | First Claude cost/performance test | New tokenizer can produce about 30% more tokens for the same text |
| GLM-5.2 | $1.40 / $4.40 per 1M | July value pick to test for long-context supported-tool coding | AA and Z.AI benchmark signals still need repo-level checks |
| Kimi K3 API | $0.30 cache-hit / $3.00 cache-miss input, $15.00 output | Latest Kimi flagship and 1M-context eval lane | Weights released under the Kimi K3 License; check serving terms; CAR is not verified |
| Kimi K2.7 Code API | $0.19 cache-hit / $0.95 cache-miss input, $4.00 output | Cheaper routine Kimi coding API lane | 256K context; thinking is required; HighSpeed costs more |
| MiniMax M3 Standard | $0.60 / $2.40 per 1M up to 512K input | Value coding-agent lane to test | Standard vs Priority pricing and Token Plan quota are separate |
| GPT-6 Luna | See current guide | OpenAI-native bounded work and high-volume evaluation | Separate API pricing, plan entitlements, retries, and review cost |
| Kimi K2.6 API | $0.16 cache-hit / $0.95 cache-miss input, $4.00 output | Product/integration-specific compatibility lane | Keep only where non-thinking mode or an existing product contract still requires K2.6 |
| Gemini Flash/Pro lanes | model-specific | High-context Google workflow tests | Verify the exact Gemini model, tier, and data terms |
How To Choose
Start With The Tool Fit
If your team lives in Claude Code or claude.ai, start with Sonnet 5 and escalate to Opus 5.5 only when Sonnet misses something important.
If your workflow supports GLM, MiniMax, Kimi, Gemini, or GPT-6 through official or BYOK routes, run the same bug fix, refactor, and review task through those lanes before paying premium-model rates. GLM-5.2 remains a value pick to test because it combines 1M context, published API pricing, independent evidence, and supported-tool subscription routing. Kimi K3 is the current Kimi flagship to evaluate; its output price still makes CAR testing mandatory. Use GLM-5.2 vs Kimi K2.6/K2.7 for the cheaper Z.AI vs Kimi decision.
Measure Cost Per Success
Token price is only part of the bill. Track:
- prompt and output tokens,
- retries,
- failed patches,
- human review time,
- tests passed on first attempt,
- data policy and retention fit.
Escalate Deliberately
Move to Premium Models when the task proves it needs a stronger model:
- architecture or multi-file reasoning fails,
- long-horizon agent work loops,
- code review misses expensive defects,
- safety or compliance review needs a second pass,
- OpenAI or Claude tooling fit matters more than raw token price.
Production Routing Pattern
| If the task is… | Start with | Escalate to |
|---|---|---|
| Routine coding or summarization | Sonnet 5, GLM-5.2, or Kimi K2.7 Code | Kimi K3 or Opus 5.5 if quality matters |
| Supported-tool budget coding | GLM-5.2 or GPT-6 Luna | GPT-6 Sol or Opus 5.5 when the value lane misses |
| Long-context agent work | MiniMax M3 or Gemini lane | Opus 5.5, then revisit Fable 5 only if access, compliance, and guardrail tests pass |
| OpenAI-native work | GPT-6 Luna via the premium guide | GPT-6 Sol when the harder tier changes the result |
| Guardrail-sensitive cyber/bio/chem work | Approved access path first | Avoid Fable unless fallback/refusal is acceptable |
Sources
- Claude API docs: Pricing (Archive)
- OpenAI API docs: Pricing (Archive)
- Smart Spend Guide — current low-cost lane and subscription routing notes
- Premium LLM Comparison — Sonnet, Opus 5.5, Fable, and GPT-6 escalation ladder
- GLM-5.2 vs Kimi K2.6/K2.7 — current cheap coding-model comparison
- Claude Opus 5.5 — current premium Claude route
- Claude Opus 4.8 — historical premium Claude baseline
Related Resources
- /compare/models/ — Model comparison hub
- /compare/models/premium/ — Premium escalation guide
- /value/smart-spend/ — Current paid-stack strategy
- /tools/zai/ — GLM Coding Plan guide
- /models/claude-opus-5-5/ — current Opus 5.5 guide
- /models/claude-opus-4-8/ — historical Opus 4.8 guide
- /models/claude-sonnet-5/ — Sonnet 5 launch, migration, and tokenizer guide
- /models/kimi-k2.7-code/ — Kimi K2.7 Code guide
- /models/kimi-k3/ — Kimi K3 launch guide
- /models/minimax-m3/ — MiniMax M3 guide
GPT-6 and Opus 5.5 pointers were refreshed September 27, 2026. GLM, Kimi, MiniMax, Gemini, and other provider evidence retain their source dates; verify live pricing, availability, and open-weight status before committing large workloads.