Before buying or renewing a seat, run the cost-saving playbook’s audit. Compare accepted work and review time, then use AI Value for the current published buying guide.

Choose the coding tool for its workflow, controls, and billing model. Choose the underlying model separately. Tool subscriptions, API token prices, and public benchmark scores are not interchangeable.

Quick Decision

PriorityStart withWhyVerify before committing
OpenAI-native cloud agents and parallel tasksCodexOpenAI account integration and isolated task workflowsPlan limits, current model picker, workspace controls, and data terms
Claude-native terminal work and premium reviewClaude CodeSonnet 5 daily lane and Opus 5.5 premium escalationSubscription/API boundary, model availability, and retention requirements
Kimi-native codingKimi Code / K3 / K2.7 CodeK3 for newest flagship tests; K2.7 Code for cheaper routine Kimi codingMembership routing, quota, exact model ID, K3 entitlement, and checkout offer

Current recommendation: use the tool that fits your repository controls, then run the same real task through the model lanes you can actually access. Do not select a tool from an old SWE-bench row alone.

Current Model Map

Tool or providerActive model contextGuarded or pending contextStatus rule
OpenAI CodexGPT-6 Astra, Sol, and LunaTrusted Access for less-restricted cyber workConfirm the exact route and product access in the target plan
Claude CodeSonnet 5 for daily work; Opus 5.5 for premium reviewFable 5 and Mythos 5Fable is restored but guarded and high-cost; Mythos remains trusted-access only
Kimi Code / APIKimi K3, Kimi K2.7 Code, and K2.7 Code HighSpeedK3 weight release and serving route require separate checksK3 is the newest Kimi flagship; K2.7 is the cheaper routine coding lane; K2.5/K2.6 are historical or compatibility context

The model exposed by a subscription or tool-facing alias can differ from the public API model discussed in a benchmark. Confirm the exact model ID or account UI instead of inferring it from the product name.

Model Evidence

These are selected model records with individual checked dates, not a complete latest-release catalog. Anthropic’s overview also lists Sonnet 5.5 and Fable 5.1; the older Sonnet 5 and Fable 5 results below do not evaluate those successors.

benchmark artifact

Selected Coding Tool Model Evidence

ModelProviderStatusContextInput priceOutput priceCoding signalTool-use signalBenchmark evidenceSpeedVerdictSourcesChecked
GPT-6 SolOpenAI active
API gpt-6-sol; paid Work/Codex rollout, separate from Chat. Client/workspace access varies.
1.05M$2.00 / 1M
$2 input / $0.20 cache / $2.50 cache write / $10 output; above 272K: 2x input/cache, 1.5x output
$10.00 / 1MAA Coding Agent Index 57 at max in Codex harness; predecessor 55 in same report.OpenAI starting effort Medium; API dollars separate from Work/Codex credits.
  • AA Coding Agent Index (September 22, max): 57; $2.99 per benchmark task, not AIHackers accepted-result cost (independent)
  • AIHackers CAR: not-run (site-owned)
not verifiedEveryday and complex coding candidate; lower prices do not remove quality and review costs.OpenAI GPT-6 Sol, Artificial Analysis GPT-6 Sol and Luna evaluation2026-09-27
GPT-6 LunaOpenAI active
API gpt-6-luna; paid Work/Codex rollout and Free/Go desktop access where available. Not Chat.
1.05M$0.10 / 1M
$0.10 input / $0.01 cache / $0.125 cache write / $0.50 output; above 272K: 2x input/cache, 1.5x output
$0.50 / 1MAA Coding Agent Index 41 at max in Codex harness; predecessor 43 in same report.OpenAI starting effort High for focused work; evaluate review burden before routing.
  • AA Coding Agent Index (September 22, max): 41; cheaper but lower score than predecessor in this evaluation (independent)
  • AIHackers CAR: not-run (site-owned)
not verifiedFocused high-volume candidate with acceptance checks; cheaper output is not a universal capability upgrade.OpenAI GPT-6 Luna, Artificial Analysis GPT-6 Sol and Luna evaluation2026-09-27
Claude Sonnet 5Anthropic active
Generally available across Claude plans, Claude Code, the Claude API, GitHub Copilot, and supported AWS paths.
1M$2.00 / 1M
$2.00 input / $10.00 output checked September 27; earlier launch schedule superseded
$10.00 / 1MAnthropic reports substantial coding and agentic gains over Sonnet 4.6; independent normalized results are pending.Available in Claude Code and the Claude API; adaptive thinking is on by default.
  • Cross-model benchmark evidence: vendor-reported; updated chart and system card preferred (vendor)
  • Historical July Artificial Analysis task cost: $1.53 per Intelligence Index task at max (independent)
  • AIHackers repo eval: not verified (site-owned)
No site-owned normalized latency result is verified.First Claude cost/performance test before Opus 5.5; escalate only when the premium pass changes the accepted result.Claude Sonnet 5 current specifications, Anthropic Claude Sonnet 5 launch [archive], Claude Sonnet 5 migration guide [archive], GitHub Copilot Claude Sonnet 5 launch [archive], Claude Sonnet 5 on AWS [archive], Artificial Analysis: Claude Opus 5 [archive]2026-09-27
Claude Opus 5.5Anthropic active
September 22 release; claude-opus-5-5 on Claude API and documented cloud routes. Verify plan and region.
1M$4.00 / 1M
$4 input / $0.20 cache read / $20 output per 1M; 5m cache write $5; 1h write $8; Fast separate
$20.00 / 1MAA Terminal-Bench 4.0: 59.6% at max with default fallback; level with Astra xhigh in that run.Always-on adaptive thinking; medium default. API migration has breaking changes.
  • AA Intelligence Index (September 22, max): 58; highest measured at release, not directly comparable with July index scores (independent)
  • AIHackers CAR: not-run (site-owned)
Anthropic reports over 30% faster output generation than Opus 5; not an AIHackers measurement.Premium coding and knowledge-work candidate; start medium and measure the gain from higher effort.Claude Opus 5.5 specifications and pricing, Anthropic Opus 5.5 launch, Artificial Analysis Opus 5.5 evaluation2026-09-27
Claude Fable 5Anthropic active
Generally available; temporary subscription allowances ended July 7 and current subscription use is through usage credits.
1M$10.00 / 1M
$10.00 input / $1.00 cache hit / $50.00 output per 1M tokens
$50.00 / 1MAnthropic reports frontier launch results; independent reproducible ranking is pending.Guarded-domain requests can refuse or fall back; verify account behavior before routing.
  • Artificial Analysis Intelligence Index: 60 at max (independent)
  • AA-Briefcase: 1574 Elo / $22.30 per task (independent)
  • AIHackers repo eval: not verified (site-owned)
Task latency varies; compare complete-task runtime before escalation.Dated Fable 5 evidence, not a Fable 5.1 evaluation. High-cost guarded escalation only; use Opus 5.5 as the practical Claude premium baseline.Claude models overview [archive], Claude API pricing [archive], Anthropic Fable 5 and Mythos 5 [archive], Anthropic Fable/Mythos access statement [archive], Anthropic Fable 5 redeployment [archive], Artificial Analysis: Claude Opus 5 [archive], Artificial Analysis: Claude Opus 5 on AA-Briefcase [archive]2026-07-25
Kimi K3Moonshot AI active
Kimi API/product flagship; Kimi K3 License weights. NVIDIA also lists K3 under trial terms as checked September 14; account access untested.
1M$3.00 / 1M
$0.30 cache-hit / $3.00 cache-miss input / $15.00 output per 1M tokens
$15.00 / 1MMoonshot reports strong max-reasoning launch-suite coding and agent results; AIHackers repo eval is not verified.Kimi API and Kimi Code support K3; Kimi docs warn to preserve full assistant history and avoid mid-session model switching.
  • Artificial Analysis Intelligence Index v4.1: 57 (independent)
  • Artificial Analysis output speed: 62 tokens/s (independent)
  • Moonshot launch benchmark suite: vendor-reported max-reasoning table (vendor)
  • AIHackers repo eval: not verified (site-owned)
Artificial Analysis reports 62 output tokens/s and flags high verbosity; measure total output cost per accepted task.Test as Kimi's newest 1M-context frontier-adjacent lane; keep K2.7 Code for cheaper routine Kimi coding until K3 passes local CAR tests.NVIDIA Kimi K3 model card, Kimi K3 launch blog [archive], Kimi K3 quickstart [archive], Kimi K3 API pricing [archive], Kimi K3 weights and license [archive], Kimi current model list [archive], Kimi Code model configuration, Artificial Analysis: Kimi K3 [archive]2026-08-01
Kimi K2.7 CodeMoonshot AI active
Cheaper routine Kimi coding API lane; HighSpeed is the same model at higher token prices.
256K$0.95 / 1M
$0.19 cache-hit / $0.95 cache-miss input / $4.00 output per 1M tokens; HighSpeed doubles those rates
$4.00 / 1MKimi K2.7 Code remains the lower-cost 256K coding lane after K3; independent normalized benchmarks are not imported.OpenAI-compatible API; thinking mode required in the documented K2.7 Code quickstart.
  • Program-Bench improvement vs K2.6: +10.4% (vendor)
  • MCP Mark Verified improvement vs K2.6: +11.4% (vendor)
  • SWE Marathon improvement vs K2.6: +76.2% (vendor)
  • Reasoning-token use vs K2.6: 30% lower (vendor)
  • AIHackers repo eval: not verified (site-owned)
HighSpeed model ID exists at a higher token price; latency not independently measured here.Cheaper routine Kimi coding API lane when Kimi routing fits and 256K context is enough.Kimi K2.7 Code quickstart [archive], Kimi K2.7 Code pricing [archive], Kimi Code K2.7 release notes [archive], SWE-bench, Berkeley Function Calling Leaderboard2026-06-28

This table separates current model status and evidence provenance. It does not rank the surrounding coding tools or claim that one benchmark predicts repository productivity.

benchmark artifact

Historical Coding Model Evidence

ModelProviderStatusContextInput priceOutput priceCoding signalTool-use signalBenchmark evidenceSpeedVerdictSourcesChecked
GPT-5.5OpenAI active
Prior generation; October 14 retirement announced for ChatGPT/Work/Codex sign-in, not API. Dated metrics retained.
1.05M API; 400K Codex$5.00 / 1M
$5.00 input / $30.00 output per 1M tokens
$30.00 / 1Mnot verifiednot verifiednot verifiednot verifiedPrimary coding seat while ChatGPT/Codex limits fit the workload.OpenAI GPT-5.5 API model page, OpenAI GPT-5.5 ChatGPT limits, Artificial Analysis: GPT-5.5, LMArena leaderboard dataset2026-06-28
GPT-5.6 SolOpenAI historical
Previous generation; dated scores and prices retained. Current routing: GPT-6 Astra, Sol and Luna. Historical status here does not imply API retirement.
1.05M$5.00 / 1M
$5.00 input / $0.50 cache read / $30.00 output per 1M tokens
$30.00 / 1MArtificial Analysis reports 80 on its Coding Agent Index at max effort; OpenAI reports 64.6% on SWE-bench Pro.Generally available in API and paid Codex plans; max and ultra modes are vendor-documented.
  • Artificial Analysis Intelligence Index: 59 at max effort (independent)
  • Artificial Analysis Coding Agent Index: 80 at max effort (independent)
  • SWE-bench Pro: 64.6% (vendor)
  • AIHackers repo eval: not-run (site-owned)
OpenAI announced a selected-customer Cerebras preview for July; production latency is not verified.Historical comparison record; use GPT-6 Astra, Sol and Luna for current evaluation candidates.OpenAI GPT-5.6 general availability [archive], OpenAI API pricing [archive], Artificial Analysis GPT-5.6 evaluation [archive], Agent Arena leaderboard, OpenAI GPT-5.6 availability [archive], OpenAI GPT-5.6 system card2026-08-01
Claude Opus 5Anthropic historical
Previous generation; dated scores and prices retained. Current routing: Opus 5.5. Historical status here does not imply API retirement.
1M$5.00 / 1M
$5.00 input / $0.50 cache hit / $25.00 output per 1M tokens; Fast mode $10.00 / $50.00
$25.00 / 1MAnthropic reports major agentic-coding gains; Artificial Analysis reports joint first on its Coding Agent Index at xhigh.Thinking is on by default; five effort settings materially change cost, latency, and task performance.
  • Artificial Analysis Intelligence Index: 61 at max effort; $2.03 per task (independent)
  • AA-Briefcase: 1720 Elo / $17.79 max; 1606 Elo / $10.41 high (independent)
  • Anthropic launch evaluations: vendor-reported; configuration varies by evaluation (vendor)
  • AIHackers repo eval: not verified (site-owned)
Artificial Analysis reports high/xhigh/max AA-Briefcase runtimes of 25.7/34.3/36.2 minutes per task; Fast mode is a separate API research preview.Historical comparison record; use Opus 5.5 for current evaluation candidates.Anthropic Claude Opus 5 launch [archive], What's new in Claude Opus 5 [archive], Claude Opus 5 system card [archive], Claude models overview [archive], Claude API pricing [archive], Artificial Analysis: Claude Opus 5 [archive], Artificial Analysis: Claude Opus 5 on AA-Briefcase [archive]2026-07-25
Claude Opus 4.8Anthropic historical
Historical comparison; use Opus 5.5 for current premium evaluation.
1M$5.00 / 1M
$5.00 input / $25.00 output per 1M tokens
$25.00 / 1MHistorical premium Claude baseline; use Opus 5 for new task-level comparisons.Still available for pinned integrations; new Claude premium routing should test Opus 5.
  • Artificial Analysis Intelligence Index v4.1: 56 (independent)
  • Artificial Analysis output speed: 57.3 tokens/s (independent)
Artificial Analysis measured 57.3 output tokens/s; provider and workload latency vary.Historical premium baseline. Use Claude Opus 5.5 for current Claude premium routing.Claude models overview [archive], Claude API pricing [archive], Artificial Analysis: Claude Opus 4.8 [archive], Artificial Analysis Intelligence Index v4.1, LMArena leaderboard dataset, Berkeley Function Calling Leaderboard2026-07-25
GLM-5.2Z.AI historical
Historical July comparison; Z.AI now documents GLM-5.3. These scores and prices do not evaluate 5.3. Historical status does not imply API retirement.
1M$1.40 / 1M
$1.40 input / $0.26 cached input / $4.40 output per 1M tokens
$4.40 / 1MZ.AI reports 62.1 on SWE-Bench Pro and 81.0 on Terminal-Bench 2.1.Supported-tool coding lane; BFCL score not imported.
  • Artificial Analysis Intelligence Index v4.1: 51 (independent)
  • SWE-Bench Pro: 62.1 (vendor)
  • Terminal-Bench 2.1: 81.0 (vendor)
Artificial Analysis flags higher output-token use; measure total cost per successful task.Dated July value candidate; retain for pinned integrations and historical comparisons. Check the newer GLM revision separately before new routing decisions.Z.AI GLM-5.3 successor documentation, Z.AI GLM-5.2 overview [archive], Z.AI pricing [archive], Artificial Analysis: GLM-5.2 article [archive], Artificial Analysis Intelligence Index v4.1, SWE-bench, Berkeley Function Calling Leaderboard2026-06-28

These rows preserve earlier GPT-5.5/GPT-5.6, Opus and GLM-5.2 scores for migration and reproducibility. They are not current GPT-6 or Opus 5.5 rankings.

Useful evidence has three levels:

  1. Independent: a named third party publishes a methodology and model-specific result.
  2. Vendor: the provider publishes a result or relative improvement; useful for deciding what to test, not for adopting the claim as an AIHackers ranking.
  3. Site-owned: the same repository task, harness, completion rules, and artifacts are available for review.

AIHackers does not yet have a controlled cross-tool repository evaluation for current Codex, Claude Code, and Kimi Code. Claims such as “95% of another model’s capability,” “best quality,” or “8x better value” are therefore not supported here.

Pricing: Keep Three Lanes Separate

Tool subscriptions

Codex, Claude Code, and Kimi Code expose plan limits, quotas, credits, or checkout-controlled offers. Those entitlements can change independently of API list prices. Verify the live plan page and account before purchase.

API token prices

ModelInputCache readOutputStatus
Kimi K3$3.00$0.30$15.00Active API; weights released under the Kimi K3 License
Kimi K2.7 Code$0.95$0.19$4.00Active
GLM-5.2$1.40$0.26$4.40Active alternative
Claude Sonnet 5$2.00verify current docs$10.00Current daily Claude route
Claude Opus 5.5$4.00$0.20$20.00Current premium Claude route
GPT-6 Sol$2.00$0.20$10.00Current hard-work OpenAI route
GPT-6 Luna$0.10$0.01$0.50Current bounded OpenAI route

Prices are per 1 million tokens. They do not include tool subscription fees, retries, cache writes, failed patches, or human review.

Historical API price examples

These rows preserve the earlier model-price evidence for pinned integrations and dated comparisons. They are not current GPT-6 or Opus 5.5 prices.

Historical modelInputCache readOutputStatus
Claude Opus 5$5.00$0.50$25.00Historical premium record
GPT-5.5$5.00check historical OpenAI docs$30.00Existing integrations and historical comparisons
GPT-5.6 Sol$5.00$0.50$30.00Historical OpenAI record

Cost per successful task

This is the metric that matters. Record:

  • input, cache writes/reads, and output tokens;
  • wall-clock time and retries;
  • whether tests pass;
  • whether the patch is accepted without repair;
  • human review and cleanup time.

A cheaper token price can lose if a model produces longer outputs, retries more often, or requires expensive review. A premium model can be rational when it prevents a costly mistake.

Workflow Comparison

Workflow questionCodexClaude CodeKimi Code
Parallel cloud task workflowStrong product fitVerify current Claude workflowVerify current Kimi workflow
Terminal-first local orchestrationSupported product paths varyCore Claude Code workflowKimi CLI and compatible tools
Premium second-pass reviewUse active OpenAI modelOpus 5.5 is the current Claude routeRoute to another provider if needed
Low API token priceCompare current OpenAI APISonnet/Opus cost moreK2.7 Code is the low-cost Kimi lane
Restricted or sensitive workGPT-6 access route after eligibility checksFable/Mythos only after access/compliance checksNot applicable in this comparison

No tool makes an autonomous agent safe by default. Apply least-privilege credentials, scoped worktrees, confirmation gates for destructive actions, and independent test/diff review.

Evaluation Protocol

Run the same four tasks:

TaskPass condition
Repository mapCorrect modules, contracts, and commands; no invented files
Real bug fixMinimal correct patch with relevant tests
Cross-file refactorPreserves behavior and repository conventions
Patch reviewFile-grounded findings with no generic filler

Pin the tool version, model ID, reasoning setting, repository commit, prompt, and acceptance rules. If a tool hides the exact model, record that as an evaluation limitation.

Current Verdict

  • Choose Codex for OpenAI-native cloud-agent workflows after confirming the current model and plan limits.
  • Choose Claude Code when Claude-native terminal work and an Opus 5.5 premium review lane matter.
  • Choose Kimi Code/K3 when you want the newest Kimi flagship and can measure CAR; choose K2.7 Code when low Kimi API pricing fits.
  • Route GPT-6 Astra, Sol, and Luna by workload and plan; use restricted access only for eligible work. Keep Mythos 5 behind its access checks and test restored Fable only as a measured escalation after Opus 5.5.
  • Keep Opus 5, Opus 4.5, Kimi K2.5/K2.6, GPT-5.5, and GPT-5.6 only as historical or product-specific context.

Sources


OpenAI GPT-6 and Claude Opus 5.5 pointers were reviewed September 27, 2026. Kimi, other provider, and tool-plan evidence retain their source dates; model pickers, access, quotas, open-weight status, and API prices can change independently.