Selection Shortcuts

GPT-6 and Opus 5.5 pointers were checked September 27, 2026; DeepSeek and GLM routing labels were reconciled October 2. October 3 adds Sol 6.1 identity, economics and access only; Sol 6’s benchmark rows retain their original revision and date. This is a selected comparison set, not a complete latest-release catalog. Anthropic’s model overview also lists Sonnet 5.5 and Fable 5.1; the Sonnet 5 and Fable 5 evidence below does not evaluate those successors. Other rows retain their own checked dates. Jump to model status, selected evidence, or history.

NeedStart withEscalate or compare with
Daily Claude productionSonnet 5Opus 5.5
Premium Claude reviewOpus 5.5Fable only after access, cost, and compliance checks
Z.AI model evaluationGLM-5.3 documentationGLM-5.2 below is a historical July comparison, not 5.3 evidence
Latest Kimi model testKimi K3GPT-6 Sol or Opus 5.5 for arbitration
Cheaper Kimi coding APIKimi K2.7 CodeGPT-6 Luna when 1M context matters
OpenAI evaluationGPT-6 Luna for focused workSol 6.1 for complex work where available; Astra for planning evaluations
DeepSeek direct API or MIT weightsDeepSeek V4.1 FlashGPT-6 Luna or a stronger reviewed lane when acceptance fails; 0731 evidence stays historical
Local or private inferenceGemma 4Compare hardware and quantization fit

See Budget Tier, Mid-Range, Premium, and Smart Spend for workload-specific decisions.

September evaluation candidates

The older model summaries below retain dated benchmark evidence. Do not compare scores across model revisions or benchmark variants without checking the source date and harness.

New to model releases? Read open weights vs open source and how to read an RL training dashboard for the difference between downloadable artifacts, training signals and useful results.

Use this section to answer two separate questions: what is newest, and what can be used normally today. An announced model can be newer without being generally available.

Active Models

These selected lanes retain their source-specific review dates. “Active” describes availability, not newest-in-family status or an AIHackers accepted-result evaluation.

Claude Sonnet 5

The daily Claude production lane. Current provider pricing is $2/$10 input/output per million tokens; keep plan and route details tied to their source dates, then escalate to Opus 5.5 only when the task needs a stronger second pass.

Claude Opus 5.5

The current premium Claude route for architecture, consequential agentic coding, knowledge work, and final arbitration. The current provider price anchor is $4/$20 input/output per million tokens with $0.20 cache read; use the Opus 5.5 guide for access, benchmark provenance, and the extended independent review. The Opus 5 page remains a dated record in historical guides.

Kimi K3

Moonshot’s newest Kimi flagship: 2.8T vendor-stated parameters, 1M context, native vision, $0.30/$3/$15 cache-hit/input/output pricing, and Artificial Analysis Index 57. Official weights and serving materials are released under the Kimi K3 License; self-hosting cost and AIHackers CAR remain unverified.

Kimi K2.7 Code

Moonshot’s cheaper routine Kimi coding API lane, with base and HighSpeed model IDs, 256K context, multimodal input, required thinking mode, and $0.95/$4.00 cache-miss input/output pricing.

GPT-6.1 Sol, Sol 6, Astra and Luna

Use Sol 6.1 for complex-work evaluation where available: Standard API rates are $2/$10 ordinary input/output with $0.10 cached reads and $2.50 writes per million tokens. Sol 6 remains a separate model with $0.20 reads and dated benchmark evidence; its scores do not evaluate 6.1. Luna remains the focused-work candidate, and the Astra planning evaluation retains its own evidence. Use the pricing owner for context tiers, credits, plan access and source dates. Sol 6.1’s AIHackers controlled evaluation and CAR are not-run.

DeepSeek V4.1 Flash

Current direct-API value candidate with native vision and MIT weights. Use deepseek-flash; legacy Flash aliases now route to V4.1. The 0731 article preserves historical evidence.

Active supporting lanes

Guarded and restricted: Fable and Mythos

Claude Fable 5 and Mythos 5

Keep these models out of default routing for different reasons. Anthropic restored Fable globally on native Claude surfaces July 1, but included usage is plan-specific, cloud rollout is incomplete, and safeguards and retention still matter. Mythos remains approved-organization access.

Selected Evidence Table

benchmark artifact

Selected Model Status, Cost, and Dated Evidence

ModelProviderStatusContextInput priceOutput priceCoding signalTool-use signalBenchmark evidenceSpeedVerdictSourcesChecked
GPT-6 AstraOpenAI active
API ID gpt-6-astra; eligible Codex access. October 3 plan check: Pro 200 subscriptions have reopened; allowance and feature access depend on plan and account.
1.05M$10.00 / 1M
$10 input / $50 output per 1M; higher long-context rates above 272K input
$50.00 / 1Mnot verifiednot verified
  • AA Intelligence Index v4.3 (Max): 53 (independent)
  • AIHackers CAR: not-run (site-owned)
not verifiedEvaluate Low for planning; raise effort after failed acceptance checks. Max benchmark evidence does not establish Low performance.OpenAI Astra model documentation, Artificial Analysis Astra evaluation, September 9, OpenAI Pro subscription access, checked October 32026-09-14
GPT-6.1 SolOpenAI active
API gpt-6.1-sol; Plus/Pro/Business/Enterprise/Edu Work/Codex rollout; Enterprise/Edu off until admin enablement. Standard/Fast where available; Ultrafast forthcoming.
1.05M$2.00 / 1M
Standard $2 ordinary / $0.10 read / $2.50 write / $10 output per 1M; >272K: 2x input/cache, 1.5x output; credits separate
$10.00 / 1MOpenAI describes near-Astra performance; vendor positioning, not a site-owned quality finding.API ID gpt-6.1-sol; Responses tool calling; compare accepted work before replacing a tested route.
  • AIHackers controlled task evaluation and CAR: not-run; no Sol 6 benchmark scores inherited (site-owned)
not verifiedComplex-work candidate with lower cache-read pricing than Sol 6; ordinary input/output rates unchanged.OpenAI GPT-6.1 Sol identity and API pricing, OpenAI Standard API prices, checked October 3, OpenAI GPT-6.1 Sol Work/Codex rollout2026-10-03
GPT-6 SolOpenAI active
API gpt-6-sol; paid Work/Codex rollout, separate from Chat. Client/workspace access varies.
1.05M$2.00 / 1M
$2 input / $0.20 cache / $2.50 cache write / $10 output; above 272K: 2x input/cache, 1.5x output
$10.00 / 1MAA Coding Agent Index 57 at max in Codex harness; predecessor 55 in same report.OpenAI starting effort Medium; API dollars separate from Work/Codex credits.
  • AA Coding Agent Index (September 22, max): 57; $2.99 per benchmark task, not AIHackers accepted-result cost (independent)
  • AIHackers CAR: not-run (site-owned)
not verifiedEveryday and complex coding candidate; lower prices do not remove quality and review costs.OpenAI GPT-6 Sol, Artificial Analysis GPT-6 Sol and Luna evaluation2026-09-27
GPT-6 LunaOpenAI active
API gpt-6-luna; paid Work/Codex rollout and Free/Go desktop access where available. Not Chat.
1.05M$0.10 / 1M
$0.10 input / $0.01 cache / $0.125 cache write / $0.50 output; above 272K: 2x input/cache, 1.5x output
$0.50 / 1MAA Coding Agent Index 41 at max in Codex harness; predecessor 43 in same report.OpenAI starting effort High for focused work; evaluate review burden before routing.
  • AA Coding Agent Index (September 22, max): 41; cheaper but lower score than predecessor in this evaluation (independent)
  • AIHackers CAR: not-run (site-owned)
not verifiedFocused high-volume candidate with acceptance checks; cheaper output is not a universal capability upgrade.OpenAI GPT-6 Luna, Artificial Analysis GPT-6 Sol and Luna evaluation2026-09-27
MiMo V2.6 FlashXiaomi MiMo active
API mimo-v2.6-flash; MIT-labeled Flash-RL weights. Account access untested.
1M$0.14 / 1M
Overseas real-time API list; excludes tools; Token Plan and Batch are separate.
$0.28 / 1Mnot verifiednot verified
  • DeepSWE v1.1 (release card): 67.9 (vendor)
  • AIHackers CAR: not-run (site-owned)
not verifiedValue candidate; test accepted results and review cost. Card score is distinct from live-RL checkpoint evaluations.MiMo V2.6 Flash release card, MiMo API prices2026-09-27
MiMo V2.6 ProXiaomi MiMo active
API mimo-v2.6-pro; MIT-labeled Pro-RL weights. Account access untested.
1M$0.435 / 1M
Overseas real-time API list; excludes tools; Token Plan and Batch are separate.
$0.87 / 1Mnot verifiednot verified
  • DeepSWE v1.1 (release card): 71.9 (vendor)
  • AIHackers CAR: not-run (site-owned)
not verifiedHarder-task candidate; compare accepted work with Flash. Card score is distinct from live-RL checkpoint evaluations.MiMo V2.6 Pro release card, MiMo API prices2026-09-27
DeepSeek V4.1 FlashDeepSeek active
Direct deepseek-flash API; native vision; MIT weights. Legacy Flash names now route here.
1M$0.15 / 1M
Off-peak $0.15 input / $0.60 output; peak $0.30 / $1.20; cache reads $0.003 / $0.006 per 1M
$0.60 / 1Mnot verifiednot verified
  • AA Intelligence Index v4.3 (Max): 40 (independent)
  • AIHackers CAR: not-run (site-owned)
not verifiedStrong API value candidate; evaluate accepted results, tool behavior, and vision. Historical 0731/preview tests do not evaluate V4.1.DeepSeek V4.1 Flash model card, DeepSeek September pricing and aliases, Artificial Analysis V4.1 Flash Max profile2026-09-14
Claude Sonnet 5Anthropic active
Generally available across Claude plans, Claude Code, the Claude API, GitHub Copilot, and supported AWS paths.
1M$2.00 / 1M
$2.00 input / $10.00 output checked September 27; earlier launch schedule superseded
$10.00 / 1MAnthropic reports substantial coding and agentic gains over Sonnet 4.6; independent normalized results are pending.Available in Claude Code and the Claude API; adaptive thinking is on by default.
  • Cross-model benchmark evidence: vendor-reported; updated chart and system card preferred (vendor)
  • Historical July Artificial Analysis task cost: $1.53 per Intelligence Index task at max (independent)
  • AIHackers repo eval: not verified (site-owned)
No site-owned normalized latency result is verified.First Claude cost/performance test before Opus 5.5; escalate only when the premium pass changes the accepted result.Claude Sonnet 5 current specifications, Anthropic Claude Sonnet 5 launch [archive], Claude Sonnet 5 migration guide [archive], GitHub Copilot Claude Sonnet 5 launch [archive], Claude Sonnet 5 on AWS [archive], Artificial Analysis: Claude Opus 5 [archive]2026-09-27
Claude Opus 5.5Anthropic active
September 22 release; claude-opus-5-5 on Claude API and documented cloud routes. Verify plan and region.
1M$4.00 / 1M
$4 input / $0.20 cache read / $20 output per 1M; 5m cache write $5; 1h write $8; Fast separate
$20.00 / 1MAA Terminal-Bench 4.0: 59.6% at max with default fallback; level with Astra xhigh in that run.Always-on adaptive thinking; medium default. API migration has breaking changes.
  • AA Intelligence Index (September 22, max): 58; highest measured at release, not directly comparable with July index scores (independent)
  • AIHackers CAR: not-run (site-owned)
Anthropic reports over 30% faster output generation than Opus 5; not an AIHackers measurement.Premium coding and knowledge-work candidate; start medium and measure the gain from higher effort.Claude Opus 5.5 specifications and pricing, Anthropic Opus 5.5 launch, Artificial Analysis Opus 5.5 evaluation2026-09-27
Claude Fable 5Anthropic active
Generally available; temporary subscription allowances ended July 7 and current subscription use is through usage credits.
1M$10.00 / 1M
$10.00 input / $1.00 cache hit / $50.00 output per 1M tokens
$50.00 / 1MAnthropic reports frontier launch results; independent reproducible ranking is pending.Guarded-domain requests can refuse or fall back; verify account behavior before routing.
  • Artificial Analysis Intelligence Index: 60 at max (independent)
  • AA-Briefcase: 1574 Elo / $22.30 per task (independent)
  • AIHackers repo eval: not verified (site-owned)
Task latency varies; compare complete-task runtime before escalation.Dated Fable 5 evidence, not a Fable 5.1 evaluation. High-cost guarded escalation only; use Opus 5.5 as the practical Claude premium baseline.Claude models overview [archive], Claude API pricing [archive], Anthropic Fable 5 and Mythos 5 [archive], Anthropic Fable/Mythos access statement [archive], Anthropic Fable 5 redeployment [archive], Artificial Analysis: Claude Opus 5 [archive], Artificial Analysis: Claude Opus 5 on AA-Briefcase [archive]2026-07-25
Claude Mythos 5Anthropic restricted
Restored to a set of approved US organizations; broader Glasswing access remains restricted.
1M$10.00 / 1M
$10.00 input / $1.00 cache hit / $50.00 output per 1M tokens
$50.00 / 1MGeneral coding quality is not independently verified for an accessible production route.Invitation-only research access; account and compliance approval required.
  • Independent cross-model evaluation: not verified (independent)
  • AIHackers repo eval: not verified (site-owned)
not verifiedRestricted research context, not a normal production or buying recommendation.Claude models overview [archive], Claude API pricing [archive], Anthropic Fable 5 and Mythos 5 [archive], Anthropic Fable/Mythos access statement [archive], Anthropic Fable 5 redeployment [archive]2026-07-01
Kimi K3Moonshot AI active
Kimi API/product flagship; Kimi K3 License weights. NVIDIA also lists K3 under trial terms as checked September 14; account access untested.
1M$3.00 / 1M
$0.30 cache-hit / $3.00 cache-miss input / $15.00 output per 1M tokens
$15.00 / 1MMoonshot reports strong max-reasoning launch-suite coding and agent results; AIHackers repo eval is not verified.Kimi API and Kimi Code support K3; Kimi docs warn to preserve full assistant history and avoid mid-session model switching.
  • Artificial Analysis Intelligence Index v4.1: 57 (independent)
  • Artificial Analysis output speed: 62 tokens/s (independent)
  • Moonshot launch benchmark suite: vendor-reported max-reasoning table (vendor)
  • AIHackers repo eval: not verified (site-owned)
Artificial Analysis reports 62 output tokens/s and flags high verbosity; measure total output cost per accepted task.Test as Kimi's newest 1M-context frontier-adjacent lane; keep K2.7 Code for cheaper routine Kimi coding until K3 passes local CAR tests.NVIDIA Kimi K3 model card, Kimi K3 launch blog [archive], Kimi K3 quickstart [archive], Kimi K3 API pricing [archive], Kimi K3 weights and license [archive], Kimi current model list [archive], Kimi Code model configuration, Artificial Analysis: Kimi K3 [archive]2026-08-01
Kimi K2.7 CodeMoonshot AI active
Cheaper routine Kimi coding API lane; HighSpeed is the same model at higher token prices.
256K$0.95 / 1M
$0.19 cache-hit / $0.95 cache-miss input / $4.00 output per 1M tokens; HighSpeed doubles those rates
$4.00 / 1MKimi K2.7 Code remains the lower-cost 256K coding lane after K3; independent normalized benchmarks are not imported.OpenAI-compatible API; thinking mode required in the documented K2.7 Code quickstart.
  • Program-Bench improvement vs K2.6: +10.4% (vendor)
  • MCP Mark Verified improvement vs K2.6: +11.4% (vendor)
  • SWE Marathon improvement vs K2.6: +76.2% (vendor)
  • Reasoning-token use vs K2.6: 30% lower (vendor)
  • AIHackers repo eval: not verified (site-owned)
HighSpeed model ID exists at a higher token price; latency not independently measured here.Cheaper routine Kimi coding API lane when Kimi routing fits and 256K context is enough.Kimi K2.7 Code quickstart [archive], Kimi K2.7 Code pricing [archive], Kimi Code K2.7 release notes [archive], SWE-bench, Berkeley Function Calling Leaderboard2026-06-28

Selected active and restricted models, not an exhaustive latest-release list. Independent, vendor, and site-owned not-run evidence remain separate. Read the price qualifications and each row’s checked date; availability does not establish cost per accepted result.

Do not compare scores without matching the benchmark variant and harness. SWE-bench Verified, SWE-Bench Pro, Terminal-Bench, Artificial Analysis, vendor preference tests, and site-owned repository tests answer different questions. Use How to Read AI Benchmarks to turn those results into a shortlist before running your own tasks.

Historical Guides

These pages remain indexable for older integrations, pricing history, and search intent. They are not current recommendations. Historical comparison status does not by itself mean an API has been retired.

benchmark artifact

Historical DeepSeek and GLM Evidence

ModelProviderStatusContextInput priceOutput priceCoding signalTool-use signalBenchmark evidenceSpeedVerdictSourcesChecked
DeepSeek V4 Flash 0731DeepSeek historical
Historical 0731 scores and prices. The direct legacy Flash name now serves V4.1 Flash; 0731 MIT weights remain a separate artifact.
1M$0.14 / 1M
Historical July 31 pricing: $0.14 cache-miss input / $0.0028 cache-hit input / $0.28 output per 1M tokens
$0.28 / 1MArtificial Analysis reports Intelligence Index 50, one point behind GPT-5.6 Luna at max effort; this is shortlist evidence, not a universal coding-quality claim.DeepSeek documents tool calls, Responses API, and Anthropic-compatible access; real harness behavior remains workload-specific.
  • Artificial Analysis Intelligence Index: 50 (independent)
  • AA-Omniscience hallucination rate: 84% (independent)
  • Kasra Rahjerdi security field test (June 2026): 0/10; pre-0731 preview only (third-party historical)
  • AIHackers 0731 repo eval: not-run (site-owned)
Independent write-up reports about 206M total output tokens for the Intelligence Index; production speed and accepted-task efficiency are not verified.Historical 0731 evidence and separate MIT weights; use the V4.1 Flash owner guide for current direct-API routing. Neither the preview test nor 0731 benchmarks evaluate V4.1.DeepSeek V4 Flash 0731 update [archive], DeepSeek V4 Flash 0731 model card [archive], DeepSeek API pricing [archive], Artificial Analysis DeepSeek V4 Flash 0731, OpenCode Zen pricing [archive], Kasra Rahjerdi’s original June 3 field test, AIHackers analysis of Kasra Rahjerdi’s field test2026-08-01
GLM-5.2Z.AI historical
Historical July comparison; Z.AI now documents GLM-5.3. These scores and prices do not evaluate 5.3. Historical status does not imply API retirement.
1M$1.40 / 1M
$1.40 input / $0.26 cached input / $4.40 output per 1M tokens
$4.40 / 1MZ.AI reports 62.1 on SWE-Bench Pro and 81.0 on Terminal-Bench 2.1.Supported-tool coding lane; BFCL score not imported.
  • Artificial Analysis Intelligence Index v4.1: 51 (independent)
  • SWE-Bench Pro: 62.1 (vendor)
  • Terminal-Bench 2.1: 81.0 (vendor)
Artificial Analysis flags higher output-token use; measure total cost per successful task.Dated July value candidate; retain for pinned integrations and historical comparisons. Check the newer GLM revision separately before new routing decisions.Z.AI GLM-5.3 successor documentation, Z.AI GLM-5.2 overview [archive], Z.AI pricing [archive], Artificial Analysis: GLM-5.2 article [archive], Artificial Analysis Intelligence Index v4.1, SWE-bench, Berkeley Function Calling Leaderboard2026-06-28

0731 and GLM-5.2 prices and benchmark results remain dated evidence. The pre-0731 security experiment was run by Kasra Rahjerdi; AIHackers supplies analysis, not an independent rerun. These results do not evaluate V4.1 Flash or GLM-5.3.

Older model names can still be correct when a tool genuinely exposes that version or a dated article describes the period. They should not be relabeled as current leaders.


October 2 correction: DeepSeek/GLM routing, historical separation and evidence provenance. This is not a complete provider refresh; other model summaries retain their stated evidence dates, and known Anthropic successors still need separate evaluation.

Models

Claude Sonnet 5 Guide

Claude Sonnet 5 pricing, 1M context, API migration changes, tokenizer cost implications, availability, and source-labeled launch evidence.

Models

GPT-5.6 Sol, Terra, and Luna Guide

Historical GPT-5.6 Sol, Terra, and Luna access, short- and long-context pricing, independent benchmark signals, and deployment controls.

Models

Claude Opus 4.8 Historical Guide

Claude Opus 4.8 historical pricing and benchmark context, retained for older integrations after Opus 5 became the premium Claude baseline.

Models

GLM-5.2: July Value Coding Pick

GLM-5.2 is the July best value coding model to test: 1M context, AA Index 51, $1.40/$4.40 API pricing, and Opus 4.8 comparison math.

Models

Kimi K2.7 Code: Coding Model Guide

Kimi K2.7 Code is Moonshot's cheaper routine coding lane after K3: 256K context, base/HighSpeed pricing, source-labeled benchmark deltas, and eval caveats.

Models

MiniMax M3: Value Coding Model Guide

MiniMax M3 is a value coding model candidate with 1M context, multimodal input, Token Plan economics, and vendor-reported benchmark strength. Test it before replacing premium lanes.

Models

Kimi K2.5 Historical Guide

Historical Kimi K2.5 capabilities, benchmarks, and older free-access routes. Use Kimi K3 for newest-Kimi decisions and K2.7 Code for cheaper coding decisions.

Models

GLM 4.7: Legacy Z.AI Model Guide

Legacy guide for Z.AI GLM 4.7. GLM-5.2 is now the current Z.AI coding-plan model; use this page for historical context and GLM-4.7 fallback routing.

Models

Claude Opus 4.5 Historical Guide

Historical Claude Opus 4.5 guide with current routing notes. Compare Sonnet 5, Opus 5, restored Fable 5, and GPT-5.6 before buying premium capacity.