Codex, Claude Code, and Cursor overlap, but they optimize different workflows. Compare their controls, execution model, repository fit, and live plan terms before comparing the models they can route.

Quick Decision

NeedStart withReason
OpenAI-native cloud-agent tasksCodexManaged task execution and OpenAI account integration
Terminal-native repository workClaude CodeClaude-native coding workflow with Sonnet and Opus lanes
IDE-first autocomplete and interactive editingCursorEditor-integrated completion, chat, and agent workflows
Premium Claude reviewClaude Code with Opus 5.5Opus 5.5 is the active premium Claude route
GPT-6 codingCodexAstra, Sol, and Luna routes require current plan and model-picker checks

There is no defensible universal winner. A tool that reduces interaction time can matter more than a small benchmark difference, while a premium model can matter when it prevents an expensive mistake.

Current Model Status

LaneCurrent read
OpenAIGPT-6 Astra/Sol/Luna are the current routes in this guide; verify exact product access
AnthropicSonnet 5 is the daily production lane and Opus 5.5 is the premium route
Fable/MythosFable restored but guarded and high-cost; Mythos remains trusted-access only
Older rowsOpus 4.5 and earlier GPT results are historical, not June 2026 rankings

Tool model pickers and aliases change. Confirm the exact model available in the target account instead of treating this page as an entitlement list.

These are selected model records with individual checked dates, not a complete latest-release catalog. Anthropic’s overview also lists Sonnet 5.5 and Fable 5.1; the older Sonnet 5 and Fable 5 results below do not evaluate those successors.

benchmark artifact

Selected Model Evidence

ModelProviderStatusContextInput priceOutput priceCoding signalTool-use signalBenchmark evidenceSpeedVerdictSourcesChecked
GPT-6 SolOpenAI active
API gpt-6-sol; paid Work/Codex rollout, separate from Chat. Client/workspace access varies.
1.05M$2.00 / 1M
$2 input / $0.20 cache / $2.50 cache write / $10 output; above 272K: 2x input/cache, 1.5x output
$10.00 / 1MAA Coding Agent Index 57 at max in Codex harness; predecessor 55 in same report.OpenAI starting effort Medium; API dollars separate from Work/Codex credits.
  • AA Coding Agent Index (September 22, max): 57; $2.99 per benchmark task, not AIHackers accepted-result cost (independent)
  • AIHackers CAR: not-run (site-owned)
not verifiedEveryday and complex coding candidate; lower prices do not remove quality and review costs.OpenAI GPT-6 Sol, Artificial Analysis GPT-6 Sol and Luna evaluation2026-09-27
GPT-6 LunaOpenAI active
API gpt-6-luna; paid Work/Codex rollout and Free/Go desktop access where available. Not Chat.
1.05M$0.10 / 1M
$0.10 input / $0.01 cache / $0.125 cache write / $0.50 output; above 272K: 2x input/cache, 1.5x output
$0.50 / 1MAA Coding Agent Index 41 at max in Codex harness; predecessor 43 in same report.OpenAI starting effort High for focused work; evaluate review burden before routing.
  • AA Coding Agent Index (September 22, max): 41; cheaper but lower score than predecessor in this evaluation (independent)
  • AIHackers CAR: not-run (site-owned)
not verifiedFocused high-volume candidate with acceptance checks; cheaper output is not a universal capability upgrade.OpenAI GPT-6 Luna, Artificial Analysis GPT-6 Sol and Luna evaluation2026-09-27
Claude Sonnet 5Anthropic active
Generally available across Claude plans, Claude Code, the Claude API, GitHub Copilot, and supported AWS paths.
1M$2.00 / 1M
$2.00 input / $10.00 output checked September 27; earlier launch schedule superseded
$10.00 / 1MAnthropic reports substantial coding and agentic gains over Sonnet 4.6; independent normalized results are pending.Available in Claude Code and the Claude API; adaptive thinking is on by default.
  • Cross-model benchmark evidence: vendor-reported; updated chart and system card preferred (vendor)
  • Historical July Artificial Analysis task cost: $1.53 per Intelligence Index task at max (independent)
  • AIHackers repo eval: not verified (site-owned)
No site-owned normalized latency result is verified.First Claude cost/performance test before Opus 5.5; escalate only when the premium pass changes the accepted result.Claude Sonnet 5 current specifications, Anthropic Claude Sonnet 5 launch [archive], Claude Sonnet 5 migration guide [archive], GitHub Copilot Claude Sonnet 5 launch [archive], Claude Sonnet 5 on AWS [archive], Artificial Analysis: Claude Opus 5 [archive]2026-09-27
Claude Opus 5.5Anthropic active
September 22 release; claude-opus-5-5 on Claude API and documented cloud routes. Verify plan and region.
1M$4.00 / 1M
$4 input / $0.20 cache read / $20 output per 1M; 5m cache write $5; 1h write $8; Fast separate
$20.00 / 1MAA Terminal-Bench 4.0: 59.6% at max with default fallback; level with Astra xhigh in that run.Always-on adaptive thinking; medium default. API migration has breaking changes.
  • AA Intelligence Index (September 22, max): 58; highest measured at release, not directly comparable with July index scores (independent)
  • AIHackers CAR: not-run (site-owned)
Anthropic reports over 30% faster output generation than Opus 5; not an AIHackers measurement.Premium coding and knowledge-work candidate; start medium and measure the gain from higher effort.Claude Opus 5.5 specifications and pricing, Anthropic Opus 5.5 launch, Artificial Analysis Opus 5.5 evaluation2026-09-27
Claude Fable 5Anthropic active
Generally available; temporary subscription allowances ended July 7 and current subscription use is through usage credits.
1M$10.00 / 1M
$10.00 input / $1.00 cache hit / $50.00 output per 1M tokens
$50.00 / 1MAnthropic reports frontier launch results; independent reproducible ranking is pending.Guarded-domain requests can refuse or fall back; verify account behavior before routing.
  • Artificial Analysis Intelligence Index: 60 at max (independent)
  • AA-Briefcase: 1574 Elo / $22.30 per task (independent)
  • AIHackers repo eval: not verified (site-owned)
Task latency varies; compare complete-task runtime before escalation.Dated Fable 5 evidence, not a Fable 5.1 evaluation. High-cost guarded escalation only; use Opus 5.5 as the practical Claude premium baseline.Claude models overview [archive], Claude API pricing [archive], Anthropic Fable 5 and Mythos 5 [archive], Anthropic Fable/Mythos access statement [archive], Anthropic Fable 5 redeployment [archive], Artificial Analysis: Claude Opus 5 [archive], Artificial Analysis: Claude Opus 5 on AA-Briefcase [archive]2026-07-25

This table evaluates current model evidence, not the surrounding coding tools. Tool productivity still requires the same repository task and acceptance rules.

benchmark artifact

Historical Model Evidence

ModelProviderStatusContextInput priceOutput priceCoding signalTool-use signalBenchmark evidenceSpeedVerdictSourcesChecked
GPT-5.5OpenAI active
Prior generation; October 14 retirement announced for ChatGPT/Work/Codex sign-in, not API. Dated metrics retained.
1.05M API; 400K Codex$5.00 / 1M
$5.00 input / $30.00 output per 1M tokens
$30.00 / 1Mnot verifiednot verifiednot verifiednot verifiedPrimary coding seat while ChatGPT/Codex limits fit the workload.OpenAI GPT-5.5 API model page, OpenAI GPT-5.5 ChatGPT limits, Artificial Analysis: GPT-5.5, LMArena leaderboard dataset2026-06-28
GPT-5.6 SolOpenAI historical
Previous generation; dated scores and prices retained. Current routing: GPT-6 Astra, Sol and Luna. Historical status here does not imply API retirement.
1.05M$5.00 / 1M
$5.00 input / $0.50 cache read / $30.00 output per 1M tokens
$30.00 / 1MArtificial Analysis reports 80 on its Coding Agent Index at max effort; OpenAI reports 64.6% on SWE-bench Pro.Generally available in API and paid Codex plans; max and ultra modes are vendor-documented.
  • Artificial Analysis Intelligence Index: 59 at max effort (independent)
  • Artificial Analysis Coding Agent Index: 80 at max effort (independent)
  • SWE-bench Pro: 64.6% (vendor)
  • AIHackers repo eval: not-run (site-owned)
OpenAI announced a selected-customer Cerebras preview for July; production latency is not verified.Historical comparison record; use GPT-6 Astra, Sol and Luna for current evaluation candidates.OpenAI GPT-5.6 general availability [archive], OpenAI API pricing [archive], Artificial Analysis GPT-5.6 evaluation [archive], Agent Arena leaderboard, OpenAI GPT-5.6 availability [archive], OpenAI GPT-5.6 system card2026-08-01
Claude Opus 5Anthropic historical
Previous generation; dated scores and prices retained. Current routing: Opus 5.5. Historical status here does not imply API retirement.
1M$5.00 / 1M
$5.00 input / $0.50 cache hit / $25.00 output per 1M tokens; Fast mode $10.00 / $50.00
$25.00 / 1MAnthropic reports major agentic-coding gains; Artificial Analysis reports joint first on its Coding Agent Index at xhigh.Thinking is on by default; five effort settings materially change cost, latency, and task performance.
  • Artificial Analysis Intelligence Index: 61 at max effort; $2.03 per task (independent)
  • AA-Briefcase: 1720 Elo / $17.79 max; 1606 Elo / $10.41 high (independent)
  • Anthropic launch evaluations: vendor-reported; configuration varies by evaluation (vendor)
  • AIHackers repo eval: not verified (site-owned)
Artificial Analysis reports high/xhigh/max AA-Briefcase runtimes of 25.7/34.3/36.2 minutes per task; Fast mode is a separate API research preview.Historical comparison record; use Opus 5.5 for current evaluation candidates.Anthropic Claude Opus 5 launch [archive], What's new in Claude Opus 5 [archive], Claude Opus 5 system card [archive], Claude models overview [archive], Claude API pricing [archive], Artificial Analysis: Claude Opus 5 [archive], Artificial Analysis: Claude Opus 5 on AA-Briefcase [archive]2026-07-25
Claude Opus 4.8Anthropic historical
Historical comparison; use Opus 5.5 for current premium evaluation.
1M$5.00 / 1M
$5.00 input / $25.00 output per 1M tokens
$25.00 / 1MHistorical premium Claude baseline; use Opus 5 for new task-level comparisons.Still available for pinned integrations; new Claude premium routing should test Opus 5.
  • Artificial Analysis Intelligence Index v4.1: 56 (independent)
  • Artificial Analysis output speed: 57.3 tokens/s (independent)
Artificial Analysis measured 57.3 output tokens/s; provider and workload latency vary.Historical premium baseline. Use Claude Opus 5.5 for current Claude premium routing.Claude models overview [archive], Claude API pricing [archive], Artificial Analysis: Claude Opus 4.8 [archive], Artificial Analysis Intelligence Index v4.1, LMArena leaderboard dataset, Berkeley Function Calling Leaderboard2026-07-25

These rows preserve prior GPT-5.5/GPT-5.6 and Opus 5 evidence for dated comparisons. They are not current GPT-6 or Opus 5.5 rankings.

Workflow Differences

Codex

Use Codex when the OpenAI-native agent workflow, managed execution environment, and task delegation fit the repository. Verify current workspace permissions, network access, model selection, rate limits, and data controls.

GPT-6 access varies by Codex plan and product surface. Confirm whether the account exposes Astra, Sol, or Luna, then record the exact model ID used for the task. Verified defenders can request restricted access for eligible cyber work.

Claude Code

Use Claude Code when a terminal-native workflow and Claude model routing fit. Start routine work with Sonnet 5 and escalate difficult review, architecture, or debugging to Opus 5.5.

Do not describe Claude Code as exposing private chain-of-thought. Evaluate the visible plan, tool calls, diffs, tests, and final explanation instead.

Cursor

Use Cursor when editor-integrated completion, interactive changes, and visual diff review matter most. Its available models, quotas, modes, and prices can change independently of provider API list prices, so check the current product and billing pages.

What to Compare

DimensionEvidence to collect
Repository controlAllowed paths, confirmation gates, worktree behavior, and diff review
Model identityExact model ID or a recorded “provider-managed/undisclosed” limitation
Completion qualityTests, accepted patches, regressions, and repair work
CostSubscription, API usage, premium requests, retries, and review time
LatencyTime to first useful edit and time to accepted completion
SecurityCredential scope, network access, retention, logs, and administrative controls

Avoid hard-coded concurrency, latency, or message-limit claims unless the current provider page documents them. Account tiers and rollout cohorts can produce materially different behavior.

Evaluation Protocol

Pin one repository commit and run:

  1. A repository architecture map.
  2. One real failing-test fix.
  3. A cross-file refactor with explicit boundaries.
  4. Review of another tool’s patch.

For every run, record the tool version, model, settings, prompt, time, tokens or credits, retries, tests, final diff, and reviewer disposition. A run without the exact model identity should not be merged into a model leaderboard.

Pricing Rules

  • Keep tool subscriptions separate from provider API pricing.
  • Treat checkout, credits, premium requests, and rate limits as live account facts.
  • Compare cost per accepted task, not only input-token list price.
  • Do not infer a tool’s monthly cost from a model API example.

Current API reference points are Sonnet 5 at $2/$10, Opus 5.5 at $4/$20 with $0.20 cache read, GPT-6 Sol at $2/$10 with $0.20 cache read, and GPT-6 Luna at $0.10/$0.50 with $0.01 cache read. Those numbers do not describe Cursor, Claude Code, or Codex subscription entitlements; verify the guide and checkout before purchase.

Security Rules

All three tools can act on valuable repositories and credentials. Use least privilege, separate development and production credentials, require confirmation for destructive or external actions, and verify completion from repository and external-state evidence.

Air-gapped or local-model support must be verified for the exact configuration. A terminal interface alone does not make a cloud-model workflow local or offline.

Verdict

  • Choose Codex for OpenAI-native managed agent tasks after validating workspace controls and the current model.
  • Choose Claude Code for terminal-native Claude workflows, with Sonnet 5 as the routine lane and Opus 5.5 as premium escalation.
  • Choose Cursor for IDE-first interaction after validating its current model roster, quotas, and privacy controls.
  • Use multiple tools only when the repository policy, commit boundaries, and review process keep their changes attributable.

Sources


OpenAI GPT-6 and Claude Opus 5.5 pointers were reviewed September 27, 2026. Cursor, other provider, and tool-plan evidence retain their source dates; model rosters, quotas, security controls, and preview access change independently.