Before buying or renewing a seat, run the cost-saving playbook’s audit. Compare accepted work and review time, then use AI Value for the current published buying guide.
Choose the coding tool for its workflow, controls, and billing model. Choose the underlying model separately. Tool subscriptions, API token prices, and public benchmark scores are not interchangeable.
Quick Decision
| Priority | Start with | Why | Verify before committing |
|---|---|---|---|
| OpenAI-native cloud agents and parallel tasks | Codex | OpenAI account integration and isolated task workflows | Plan limits, current model picker, workspace controls, and data terms |
| Claude-native terminal work and premium review | Claude Code | Sonnet 5 daily lane and Opus 5.5 premium escalation | Subscription/API boundary, model availability, and retention requirements |
| Kimi-native coding | Kimi Code / K3 / K2.7 Code | K3 for newest flagship tests; K2.7 Code for cheaper routine Kimi coding | Membership routing, quota, exact model ID, K3 entitlement, and checkout offer |
Current recommendation: use the tool that fits your repository controls, then run the same real task through the model lanes you can actually access. Do not select a tool from an old SWE-bench row alone.
Current Model Map
| Tool or provider | Active model context | Guarded or pending context | Status rule |
|---|---|---|---|
| OpenAI Codex | GPT-6 Astra, Sol, and Luna | Trusted Access for less-restricted cyber work | Confirm the exact route and product access in the target plan |
| Claude Code | Sonnet 5 for daily work; Opus 5.5 for premium review | Fable 5 and Mythos 5 | Fable is restored but guarded and high-cost; Mythos remains trusted-access only |
| Kimi Code / API | Kimi K3, Kimi K2.7 Code, and K2.7 Code HighSpeed | K3 weight release and serving route require separate checks | K3 is the newest Kimi flagship; K2.7 is the cheaper routine coding lane; K2.5/K2.6 are historical or compatibility context |
The model exposed by a subscription or tool-facing alias can differ from the public API model discussed in a benchmark. Confirm the exact model ID or account UI instead of inferring it from the product name.
Model Evidence
These are selected model records with individual checked dates, not a complete latest-release catalog. Anthropic’s overview also lists Sonnet 5.5 and Fable 5.1; the older Sonnet 5 and Fable 5 results below do not evaluate those successors.
benchmark artifact
Selected Coding Tool Model Evidence
| Model | Provider | Status | Context | Input price | Output price | Coding signal | Tool-use signal | Benchmark evidence | Speed | Verdict | Sources | Checked |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-6 Sol | OpenAI |
active API gpt-6-sol; paid Work/Codex rollout, separate from Chat. Client/workspace access varies. | 1.05M | $2.00 / 1M $2 input / $0.20 cache / $2.50 cache write / $10 output; above 272K: 2x input/cache, 1.5x output | $10.00 / 1M | AA Coding Agent Index 57 at max in Codex harness; predecessor 55 in same report. | OpenAI starting effort Medium; API dollars separate from Work/Codex credits. |
| not verified | Everyday and complex coding candidate; lower prices do not remove quality and review costs. | OpenAI GPT-6 Sol, Artificial Analysis GPT-6 Sol and Luna evaluation | 2026-09-27 |
| GPT-6 Luna | OpenAI |
active API gpt-6-luna; paid Work/Codex rollout and Free/Go desktop access where available. Not Chat. | 1.05M | $0.10 / 1M $0.10 input / $0.01 cache / $0.125 cache write / $0.50 output; above 272K: 2x input/cache, 1.5x output | $0.50 / 1M | AA Coding Agent Index 41 at max in Codex harness; predecessor 43 in same report. | OpenAI starting effort High for focused work; evaluate review burden before routing. |
| not verified | Focused high-volume candidate with acceptance checks; cheaper output is not a universal capability upgrade. | OpenAI GPT-6 Luna, Artificial Analysis GPT-6 Sol and Luna evaluation | 2026-09-27 |
| Claude Sonnet 5 | Anthropic |
active Generally available across Claude plans, Claude Code, the Claude API, GitHub Copilot, and supported AWS paths. | 1M | $2.00 / 1M $2.00 input / $10.00 output checked September 27; earlier launch schedule superseded | $10.00 / 1M | Anthropic reports substantial coding and agentic gains over Sonnet 4.6; independent normalized results are pending. | Available in Claude Code and the Claude API; adaptive thinking is on by default. |
| No site-owned normalized latency result is verified. | First Claude cost/performance test before Opus 5.5; escalate only when the premium pass changes the accepted result. | Claude Sonnet 5 current specifications, Anthropic Claude Sonnet 5 launch [archive], Claude Sonnet 5 migration guide [archive], GitHub Copilot Claude Sonnet 5 launch [archive], Claude Sonnet 5 on AWS [archive], Artificial Analysis: Claude Opus 5 [archive] | 2026-09-27 |
| Claude Opus 5.5 | Anthropic |
active September 22 release; claude-opus-5-5 on Claude API and documented cloud routes. Verify plan and region. | 1M | $4.00 / 1M $4 input / $0.20 cache read / $20 output per 1M; 5m cache write $5; 1h write $8; Fast separate | $20.00 / 1M | AA Terminal-Bench 4.0: 59.6% at max with default fallback; level with Astra xhigh in that run. | Always-on adaptive thinking; medium default. API migration has breaking changes. |
| Anthropic reports over 30% faster output generation than Opus 5; not an AIHackers measurement. | Premium coding and knowledge-work candidate; start medium and measure the gain from higher effort. | Claude Opus 5.5 specifications and pricing, Anthropic Opus 5.5 launch, Artificial Analysis Opus 5.5 evaluation | 2026-09-27 |
| Claude Fable 5 | Anthropic |
active Generally available; temporary subscription allowances ended July 7 and current subscription use is through usage credits. | 1M | $10.00 / 1M $10.00 input / $1.00 cache hit / $50.00 output per 1M tokens | $50.00 / 1M | Anthropic reports frontier launch results; independent reproducible ranking is pending. | Guarded-domain requests can refuse or fall back; verify account behavior before routing. |
| Task latency varies; compare complete-task runtime before escalation. | Dated Fable 5 evidence, not a Fable 5.1 evaluation. High-cost guarded escalation only; use Opus 5.5 as the practical Claude premium baseline. | Claude models overview [archive], Claude API pricing [archive], Anthropic Fable 5 and Mythos 5 [archive], Anthropic Fable/Mythos access statement [archive], Anthropic Fable 5 redeployment [archive], Artificial Analysis: Claude Opus 5 [archive], Artificial Analysis: Claude Opus 5 on AA-Briefcase [archive] | 2026-07-25 |
| Kimi K3 | Moonshot AI |
active Kimi API/product flagship; Kimi K3 License weights. NVIDIA also lists K3 under trial terms as checked September 14; account access untested. | 1M | $3.00 / 1M $0.30 cache-hit / $3.00 cache-miss input / $15.00 output per 1M tokens | $15.00 / 1M | Moonshot reports strong max-reasoning launch-suite coding and agent results; AIHackers repo eval is not verified. | Kimi API and Kimi Code support K3; Kimi docs warn to preserve full assistant history and avoid mid-session model switching. |
| Artificial Analysis reports 62 output tokens/s and flags high verbosity; measure total output cost per accepted task. | Test as Kimi's newest 1M-context frontier-adjacent lane; keep K2.7 Code for cheaper routine Kimi coding until K3 passes local CAR tests. | NVIDIA Kimi K3 model card, Kimi K3 launch blog [archive], Kimi K3 quickstart [archive], Kimi K3 API pricing [archive], Kimi K3 weights and license [archive], Kimi current model list [archive], Kimi Code model configuration, Artificial Analysis: Kimi K3 [archive] | 2026-08-01 |
| Kimi K2.7 Code | Moonshot AI |
active Cheaper routine Kimi coding API lane; HighSpeed is the same model at higher token prices. | 256K | $0.95 / 1M $0.19 cache-hit / $0.95 cache-miss input / $4.00 output per 1M tokens; HighSpeed doubles those rates | $4.00 / 1M | Kimi K2.7 Code remains the lower-cost 256K coding lane after K3; independent normalized benchmarks are not imported. | OpenAI-compatible API; thinking mode required in the documented K2.7 Code quickstart. |
| HighSpeed model ID exists at a higher token price; latency not independently measured here. | Cheaper routine Kimi coding API lane when Kimi routing fits and 256K context is enough. | Kimi K2.7 Code quickstart [archive], Kimi K2.7 Code pricing [archive], Kimi Code K2.7 release notes [archive], SWE-bench, Berkeley Function Calling Leaderboard | 2026-06-28 |
This table separates current model status and evidence provenance. It does not rank the surrounding coding tools or claim that one benchmark predicts repository productivity.
benchmark artifact
Historical Coding Model Evidence
| Model | Provider | Status | Context | Input price | Output price | Coding signal | Tool-use signal | Benchmark evidence | Speed | Verdict | Sources | Checked |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-5.5 | OpenAI |
active Prior generation; October 14 retirement announced for ChatGPT/Work/Codex sign-in, not API. Dated metrics retained. | 1.05M API; 400K Codex | $5.00 / 1M $5.00 input / $30.00 output per 1M tokens | $30.00 / 1M | not verified | not verified | not verified | not verified | Primary coding seat while ChatGPT/Codex limits fit the workload. | OpenAI GPT-5.5 API model page, OpenAI GPT-5.5 ChatGPT limits, Artificial Analysis: GPT-5.5, LMArena leaderboard dataset | 2026-06-28 |
| GPT-5.6 Sol | OpenAI |
historical Previous generation; dated scores and prices retained. Current routing: GPT-6 Astra, Sol and Luna. Historical status here does not imply API retirement. | 1.05M | $5.00 / 1M $5.00 input / $0.50 cache read / $30.00 output per 1M tokens | $30.00 / 1M | Artificial Analysis reports 80 on its Coding Agent Index at max effort; OpenAI reports 64.6% on SWE-bench Pro. | Generally available in API and paid Codex plans; max and ultra modes are vendor-documented. |
| OpenAI announced a selected-customer Cerebras preview for July; production latency is not verified. | Historical comparison record; use GPT-6 Astra, Sol and Luna for current evaluation candidates. | OpenAI GPT-5.6 general availability [archive], OpenAI API pricing [archive], Artificial Analysis GPT-5.6 evaluation [archive], Agent Arena leaderboard, OpenAI GPT-5.6 availability [archive], OpenAI GPT-5.6 system card | 2026-08-01 |
| Claude Opus 5 | Anthropic |
historical Previous generation; dated scores and prices retained. Current routing: Opus 5.5. Historical status here does not imply API retirement. | 1M | $5.00 / 1M $5.00 input / $0.50 cache hit / $25.00 output per 1M tokens; Fast mode $10.00 / $50.00 | $25.00 / 1M | Anthropic reports major agentic-coding gains; Artificial Analysis reports joint first on its Coding Agent Index at xhigh. | Thinking is on by default; five effort settings materially change cost, latency, and task performance. |
| Artificial Analysis reports high/xhigh/max AA-Briefcase runtimes of 25.7/34.3/36.2 minutes per task; Fast mode is a separate API research preview. | Historical comparison record; use Opus 5.5 for current evaluation candidates. | Anthropic Claude Opus 5 launch [archive], What's new in Claude Opus 5 [archive], Claude Opus 5 system card [archive], Claude models overview [archive], Claude API pricing [archive], Artificial Analysis: Claude Opus 5 [archive], Artificial Analysis: Claude Opus 5 on AA-Briefcase [archive] | 2026-07-25 |
| Claude Opus 4.8 | Anthropic |
historical Historical comparison; use Opus 5.5 for current premium evaluation. | 1M | $5.00 / 1M $5.00 input / $25.00 output per 1M tokens | $25.00 / 1M | Historical premium Claude baseline; use Opus 5 for new task-level comparisons. | Still available for pinned integrations; new Claude premium routing should test Opus 5. |
| Artificial Analysis measured 57.3 output tokens/s; provider and workload latency vary. | Historical premium baseline. Use Claude Opus 5.5 for current Claude premium routing. | Claude models overview [archive], Claude API pricing [archive], Artificial Analysis: Claude Opus 4.8 [archive], Artificial Analysis Intelligence Index v4.1, LMArena leaderboard dataset, Berkeley Function Calling Leaderboard | 2026-07-25 |
| GLM-5.2 | Z.AI |
historical Historical July comparison; Z.AI now documents GLM-5.3. These scores and prices do not evaluate 5.3. Historical status does not imply API retirement. | 1M | $1.40 / 1M $1.40 input / $0.26 cached input / $4.40 output per 1M tokens | $4.40 / 1M | Z.AI reports 62.1 on SWE-Bench Pro and 81.0 on Terminal-Bench 2.1. | Supported-tool coding lane; BFCL score not imported. |
| Artificial Analysis flags higher output-token use; measure total cost per successful task. | Dated July value candidate; retain for pinned integrations and historical comparisons. Check the newer GLM revision separately before new routing decisions. | Z.AI GLM-5.3 successor documentation, Z.AI GLM-5.2 overview [archive], Z.AI pricing [archive], Artificial Analysis: GLM-5.2 article [archive], Artificial Analysis Intelligence Index v4.1, SWE-bench, Berkeley Function Calling Leaderboard | 2026-06-28 |
These rows preserve earlier GPT-5.5/GPT-5.6, Opus and GLM-5.2 scores for migration and reproducibility. They are not current GPT-6 or Opus 5.5 rankings.
Useful evidence has three levels:
- Independent: a named third party publishes a methodology and model-specific result.
- Vendor: the provider publishes a result or relative improvement; useful for deciding what to test, not for adopting the claim as an AIHackers ranking.
- Site-owned: the same repository task, harness, completion rules, and artifacts are available for review.
AIHackers does not yet have a controlled cross-tool repository evaluation for current Codex, Claude Code, and Kimi Code. Claims such as “95% of another model’s capability,” “best quality,” or “8x better value” are therefore not supported here.
Pricing: Keep Three Lanes Separate
Tool subscriptions
Codex, Claude Code, and Kimi Code expose plan limits, quotas, credits, or checkout-controlled offers. Those entitlements can change independently of API list prices. Verify the live plan page and account before purchase.
API token prices
| Model | Input | Cache read | Output | Status |
|---|---|---|---|---|
| Kimi K3 | $3.00 | $0.30 | $15.00 | Active API; weights released under the Kimi K3 License |
| Kimi K2.7 Code | $0.95 | $0.19 | $4.00 | Active |
| GLM-5.2 | $1.40 | $0.26 | $4.40 | Active alternative |
| Claude Sonnet 5 | $2.00 | verify current docs | $10.00 | Current daily Claude route |
| Claude Opus 5.5 | $4.00 | $0.20 | $20.00 | Current premium Claude route |
| GPT-6 Sol | $2.00 | $0.20 | $10.00 | Current hard-work OpenAI route |
| GPT-6 Luna | $0.10 | $0.01 | $0.50 | Current bounded OpenAI route |
Prices are per 1 million tokens. They do not include tool subscription fees, retries, cache writes, failed patches, or human review.
Historical API price examples
These rows preserve the earlier model-price evidence for pinned integrations and dated comparisons. They are not current GPT-6 or Opus 5.5 prices.
| Historical model | Input | Cache read | Output | Status |
|---|---|---|---|---|
| Claude Opus 5 | $5.00 | $0.50 | $25.00 | Historical premium record |
| GPT-5.5 | $5.00 | check historical OpenAI docs | $30.00 | Existing integrations and historical comparisons |
| GPT-5.6 Sol | $5.00 | $0.50 | $30.00 | Historical OpenAI record |
Cost per successful task
This is the metric that matters. Record:
- input, cache writes/reads, and output tokens;
- wall-clock time and retries;
- whether tests pass;
- whether the patch is accepted without repair;
- human review and cleanup time.
A cheaper token price can lose if a model produces longer outputs, retries more often, or requires expensive review. A premium model can be rational when it prevents a costly mistake.
Workflow Comparison
| Workflow question | Codex | Claude Code | Kimi Code |
|---|---|---|---|
| Parallel cloud task workflow | Strong product fit | Verify current Claude workflow | Verify current Kimi workflow |
| Terminal-first local orchestration | Supported product paths vary | Core Claude Code workflow | Kimi CLI and compatible tools |
| Premium second-pass review | Use active OpenAI model | Opus 5.5 is the current Claude route | Route to another provider if needed |
| Low API token price | Compare current OpenAI API | Sonnet/Opus cost more | K2.7 Code is the low-cost Kimi lane |
| Restricted or sensitive work | GPT-6 access route after eligibility checks | Fable/Mythos only after access/compliance checks | Not applicable in this comparison |
No tool makes an autonomous agent safe by default. Apply least-privilege credentials, scoped worktrees, confirmation gates for destructive actions, and independent test/diff review.
Evaluation Protocol
Run the same four tasks:
| Task | Pass condition |
|---|---|
| Repository map | Correct modules, contracts, and commands; no invented files |
| Real bug fix | Minimal correct patch with relevant tests |
| Cross-file refactor | Preserves behavior and repository conventions |
| Patch review | File-grounded findings with no generic filler |
Pin the tool version, model ID, reasoning setting, repository commit, prompt, and acceptance rules. If a tool hides the exact model, record that as an evaluation limitation.
Current Verdict
- Choose Codex for OpenAI-native cloud-agent workflows after confirming the current model and plan limits.
- Choose Claude Code when Claude-native terminal work and an Opus 5.5 premium review lane matter.
- Choose Kimi Code/K3 when you want the newest Kimi flagship and can measure CAR; choose K2.7 Code when low Kimi API pricing fits.
- Route GPT-6 Astra, Sol, and Luna by workload and plan; use restricted access only for eligible work. Keep Mythos 5 behind its access checks and test restored Fable only as a measured escalation after Opus 5.5.
- Keep Opus 5, Opus 4.5, Kimi K2.5/K2.6, GPT-5.5, and GPT-5.6 only as historical or product-specific context.
Sources
- OpenAI: GPT-6 Sol/Luna pricing owner and GPT-6 Astra planning evidence
- OpenAI: historical GPT-5.6 general availability and historical system card
- OpenAI: Codex
- Anthropic: Claude model overview and pricing
- Kimi: K3 launch blog, K3 quickstart, K2.7 Code quickstart, and K2.7 pricing
- Anthropic: Claude Opus 5.5 guide
- Artificial Analysis: historical Claude Opus 5, Claude Opus 4.8, and GLM-5.2
Related links
- /value/gpt-6-sol-luna-pricing/
- /posts/gpt-6-astra-planning-value/
- /models/gpt-5-6/ — historical GPT-5.6 guide
- /models/claude-opus-5-5/
- /models/claude-opus-5/ — historical Opus 5 guide
- /models/claude-opus-4-8/ — historical Opus 4.8 guide
- /models/kimi-k3/
- /models/kimi-k2.7-code/
- /tools/kimi-code/
- /value/smart-spend/
- /risks/codex/cloud-dependency-risks/
OpenAI GPT-6 and Claude Opus 5.5 pointers were reviewed September 27, 2026. Kimi, other provider, and tool-plan evidence retain their source dates; model pickers, access, quotas, open-weight status, and API prices can change independently.