The decision: evaluate GPT-6.1 Sol for complex coding and repeated professional work when it is available to your account. Use Luna for bounded extraction and repeatable transformations; reserve Astra for work where its additional capability passes your acceptance test. OpenAI recommends the client’s default effort for Sol 6.1, High for Luna and Light for Astra. Its “near-Astra performance” description is a vendor claim, not an AIHackers quality finding. (OpenAI model guidance)
October 3, 2026 economics/access check: this refresh adds Sol 6.1 and rechecks the retained price table, credit rates and speed modifiers. The Sol 6/Luna benchmark evidence below retains its September 27 review; no scores have been reassigned to Sol 6.1. AIHackers’ Sol 6.1 controlled task evaluation and CAR are not-run.
Sol 6.1 is a separate revision
The API ID is gpt-6.1-sol. Its model card lists a 1,050,000-token context, 128,000 maximum output, text/image input and text output. API reasoning supports low, medium (default), high, xhigh and max; it does not support none or minimal. Use Responses for tool calling; Chat Completions supports this model without tool calling. (Sol 6.1 model card)
The launch rollout includes Plus, Pro, Business, Enterprise and Edu across the documented Work/Codex clients; Enterprise/Edu require administrator enablement. Free and Go are excluded at launch. Standard and Fast are available where the client, account and workspace support them; Sol 6.1 Ultrafast is forthcoming. Ultra reasoning uses subagents and is separate from Ultrafast speed. A listing does not prove signed-in access. (Work/Codex model access, speed modes)
Sol 6 and Luna: retained identity
OpenAI offers the models through the API as gpt-6-sol and gpt-6-luna, and through ChatGPT Work and Codex. They are not available in ordinary ChatGPT chats in the current rollout. Plus, Pro, Business, Enterprise, and Edu users can receive both in Work and Codex; Free and Go users can try Luna in the desktop app. Workspace settings and gradual rollout still control what an account can select. Both models accept text and image input, support none, low, medium, high, xhigh, and max reasoning effort, and have a 1.05 million token context window with 128,000 maximum output tokens. (OpenAI model catalog, release notes)
For finite labels versus explanations, see Jev, Clef or OpenAI Decisions. Ordinary Luna structured generation and OpenAI Decisions have separate contracts; this page’s rates do not price Decisions.
API prices are the cleanest starting point
The Standard API card, checked October 3, is quoted in USD per 1 million tokens. Ordinary input, cache reads and cache writes are separate input buckets: a cache write replaces the ordinary-input charge for those tokens. These dollars are separate from any ChatGPT subscription or purchased product credits.
| Model | Input | Cached input | Cache write | Output | Long prompt input / cached / write / output* |
|---|---|---|---|---|---|
| GPT-6.1 Sol | $2.00 | $0.10 | $2.50 | $10.00 | $4.00 / $0.20 / $5.00 / $15.00 |
| GPT-6 Sol | $2.00 | $0.20 | $2.50 | $10.00 | $4.00 / $0.40 / $5.00 / $15.00 |
| GPT-6 Luna | $0.10 | $0.01 | $0.125 | $0.50 | $0.20 / $0.02 / $0.25 / $0.75 |
| GPT-5.6 Sol | $4.00 | $0.40 | $5.00 | $20.00 | $8.00 / $0.80 / $10.00 / $30.00 |
| GPT-5.6 Luna | $0.20 | $0.02 | $0.25 | $1.20 | $0.40 / $0.04 / $0.50 / $1.80 |
*Short-context rates apply to an individual request with at most 272,000 input tokens. Above that threshold, the full request uses 2× input/read/write rates and 1.5× output rates. Aggregating many short requests above 272K does not trigger the multiplier. GPT-5.6 rows use current listed comparison rates, not their launch prices; OpenAI labels Sol’s pricing promotional at least through November 21, 2026. (API pricing)
For the models in this table, Batch and Flex rates are 50% of Standard and Fast rates are 2× Standard, where supported. Batch suits deferred jobs; Flex trades speed and availability for price. Eligible regional processing adds 10%. Sol 6.1 supports US/EU residency, but Fast is unavailable with EU residency. Sol 6 and Luna’s current cards list EU residency with Standard, Flex and Batch. Check route eligibility before combining a discount with a regional setting. Cache writes cost 1.25× ordinary input; reads cost 5% on Sol 6.1 and 10% on the other rows. (API pricing, Sol 6.1, Sol 6, Luna)
Half-price cache reads do not halve every task
Sol 6.1’s $0.10 cached-input rate halves Sol 6’s $0.20; ordinary input, cache writes and output prices are unchanged. As an illustration, aggregate ten short-context Standard requests, each with 90K cached reads, 10K cache writes, no ordinary input and 10K output. Across the sequence, 900K reads + 100K writes + 100K output cost $0.09 + $0.25 + $1.00 = $1.34 on Sol 6.1, versus $1.43 on Sol 6. The token bill falls 6.3%, not 50%.
The same 1M input and 100K output processed without caching at ordinary rates would cost $3.00: that comparator gives 55.3% whole-token-cost reduction on Sol 6.1. This chosen 90% token-weighted read share is a scenario, not a typical workload or guaranteed hit rate. It excludes earlier cache preparation, tools, retries, human review, taxes and processing/regional modifiers. For cold starts, other token mixes and measurement, use the cache-cost scenarios and logging recipe.
At list rates, a 100,000-input/10,000-output turn costs about $0.30 on Sol and $0.015 on Luna before cache or processing discounts. A single 300,000-input/30,000-output codebase review crosses the 272,000-token threshold: it costs about $1.65 on Sol ($1.20 input + $0.45 output) and $0.0825 on Luna ($0.06 input + $0.0225 output), before cache writes. A ten-million-input/one-million-output day costs about $30 on Sol and $1.50 on Luna when it is an aggregate of individual requests that each stay at or below 272,000 input tokens. These are arithmetic illustrations, not AIHackers task results: tool calls, retries, cache misses, and review time can dominate the token bill.
Credits are a different ledger
For purchased-credit billing, ChatGPT Work and Codex share a ledger distinct from API dollars. The October 3 Standard-speed card quotes credits per 1 million tokens; it has no separate cache-write charge:
| Model | Input credits | Cached input credits | Output credits |
|---|---|---|---|
| GPT-6 Astra | 250 | 25 | 1,250 |
| GPT-6.1 Sol | 50 | 2.5 | 250 |
| GPT-6 Sol | 50 | 5 | 250 |
| GPT-6 Luna | 2.5 | 0.25 | 12.5 |
| GPT-5.6 Sol | 100 | 10 | 500 |
| GPT-5.6 Luna | 5 | 0.5 | 30 |
Fast consumes included subscription limits at 2.5× Standard, but purchased credits at 2×. Astra Ultrafast consumes included limits at 8× and purchased credits at 6×; its up-to-8× output-generation speed is not an elapsed-task guarantee. Sol 6.1 does not yet offer Ultrafast. API-key usage instead follows API prices. Included plan usage, purchased credits and promotional benefits remain separate balances; none supplies a fixed task count. Check the signed-in dashboard for current limits and reset times. Pro plans currently have no five-hour limit; weekly and model allowances still matter. (Pricing and credit rules, speed, Pro transition terms)
What the benchmark record actually says
Retained September 27 evidence: GPT-6 Sol and GPT-6 Luna only. Sol 6.1 is a distinct revision with no controlled AIHackers benchmark or CAR result in this refresh; the following scores do not evaluate it.
OpenAI’s launch page reports strong vendor evaluations. Sol scored 33.2% on AutomationBench at xhigh for a stated $0.27 per task, versus 26.9% for Claude Opus 5 at max; Luna improved 5.4 percentage points over GPT-5.6 Luna at high effort with 58% lower cost per task. OpenAI reports 68.8% for Sol and 66.6% for Luna on DeepSWE v1.1, and 60.5% for Sol on OSWorld 2.0 offline. The same page says its GPT evaluations ran in OpenAI’s research environment or API and that competitor scores came from public reports. The cited competitor is the older Opus 5; this row does not compare Sol with Opus 5.5. These are vendor launch results, with different harnesses across some comparisons. (OpenAI launch evaluation table)
The independent record is more specific than a single leaderboard. Artificial Analysis’ v4.3.2 comparison uses ten evaluations at max effort: Sol scores 48 versus GPT-5.6 Sol at 47 on the Intelligence Index, while Luna scores 37 versus GPT-5.6 Luna at 37. Its Coding Agent Index article reports Sol at 57 versus 55 and Luna at 41 versus 43. On the separate Intelligence Index task set, it reports costs of $1.06 versus $1.99 for Sol and $0.07 versus $0.18 for Luna. These are point measurements from Artificial Analysis’ harness, with its own token and provider assumptions. (Artificial Analysis Sol comparison, Luna comparison, AA Sol and Luna analysis)
LiveBench’s June 25 release publishes the task rows and category map. Recomputing the unweighted mean of its seven category means gives Sol 79.2 versus GPT-5.6 Sol 81.1, and Luna 72.0 versus GPT-5.6 Luna 73.6. That is a dated CSV snapshot, not a confidence interval. ARC Prize’s verified Luna result reports 86.7% on ARC-AGI-1 and 59.3% on ARC-AGI-2 at max reasoning, with the harness and effort shown on the result page. Those scores describe specific task families; they do not settle the model choice for a different workflow. (LiveBench task table, LiveBench category map, ARC Prize verified result)
Our Opus 5.5 review compares its independent evidence and effort tradeoffs against the launch enthusiasm. AIHackers has not run a controlled comparison of accepted results for Sol or Luna; CAR is not-run.
That leads to a practical routing rule:
| Workload | Start with | Reconsider when |
|---|---|---|
| High-volume extraction, triage, routine transformations | Luna at high effort | A representative acceptance check passes at a lower effort, or review time erases the saving |
| Multi-file coding, agent workflows, ambiguous implementation | Sol at medium effort | The task needs deeper verification or Astra’s computer-use and judgment |
| Ambiguous, high-value, end-to-end work | Astra at light effort | The task needs more depth, tool use, or sustained review |
| Long prompts | Either model below 272K input tokens | The full-request long-context multiplier changes the budget |
| Offline evaluations and enrichment | Batch or Flex | The delayed completion or occasional unavailability misses the deadline |
This routing table preserves the Sol 6/Luna comparison. Evaluate Sol 6.1 separately with the same task set, effort, tools, provider route and retry budget before changing a production default. The launch/offers owner now records released Pro 500 and dots, Pro 200’s allowance transition and dated reset events.
For a small cross-provider evaluation, use the September value shortlist: Luna alongside MiMo V2.6 Flash, DeepSeek V4.1 Flash and Opus 5.5. Apply the cornerstone cost method to the work you accept, including retries and review time.
Review scopes: benchmarks reviewed September 27, 2026; OpenAI identity, API prices/modifiers, product credit rates and speed/access terms checked October 3. No account checkout, quota measurement or paid task evaluation was run. Recheck the linked owners before committing spend.