Use this guide to set up Codex, choose a model, and check the controls that affect your work.

  • Start locally: install the CLI, choose how to sign in, and add repository instructions.
  • Check access and costs: verify your plan, usage meters, model availability, and billing route.
  • Run bounded tasks: distinguish permissions from task authorization, inspect failures, and keep one watcher for long jobs.

October 3 guide update

Use GPT-6.1 Sol for complex coding and Luna High for focused work; evaluate Astra Low for difficult planning and orchestration. For Sol 6.1, start with the effort available by default in your client, then adjust against acceptance checks. These are starting points from OpenAI’s model guidance, subject to your client, plan, rollout and workspace settings.

Pro $200 purchases have reopened. Eligible existing subscribers retain their previous allowance only through October 29, 2026, then move to a lower allowance at the same $200 monthly price. The launch offers guide details the September eligibility window and same-account return conditions; the current access check supersedes our purchase-pause wording. Pro 200 remains a distinct tier from Pro 500 and does not include Ultrafast. GPT-5.6 remains available during rollout.

OpenAI Codex is available through CLI, IDE, desktop and cloud workflows. GPT-6.1 Sol, GPT-6 Sol and Luna are in Work and Codex rather than ordinary Chat. GPT-5.4 and GPT-5.4-mini retired from ChatGPT-authenticated Codex on August 31. GPT-5.5 is scheduled to retire from ChatGPT, Work and Codex on October 14, 2026; that notice does not retire it from the API.

This refresh also covers cheap cache reuse, GitHub networking, and long-job polling. Current documentation, dated user reports and our local observations have different evidence scopes; earlier reset and benchmark records retain their dates.

Install

1
2
3
bun install -g @openai/codex
codex --version
codex

Codex can authenticate with ChatGPT or an API key:

  • ChatGPT authentication includes Codex according to the user’s Free, Go, Plus, Pro, Business, Edu, or Enterprise plan.
  • API-key authentication uses models available to that API key and API token billing.
  • API-key authentication does not provide Codex cloud features such as hosted code review or Slack integration.

Do not hard-code an expected CLI version in an evergreen setup guide. Verify the installed version and current release notes.

Current Model Selection

Scroll sideways to compare columns
NeedModel and starting effortAccess boundary
Complex coding and agent workgpt-6.1-sol, client default; API MediumPlus/Pro/Business/Enterprise/Edu rollout; Enterprise/Edu require administrator enablement; API separately billed
Existing Sol 6 integration or dated evaluationgpt-6-solKeep its own results; Sol 6.1 has separate identity and economics
Focused, repeatable workgpt-6-luna, HighPaid Work/Codex plans; Free/Go desktop rollout; API separately billed
Hard planning, research and multi-step judgmentgpt-6-astra, Light (low)Check your account and client
Prior family integrationGPT-5.6 Sol, Terra, LunaStill available during rollout; retain only when your workflow justifies it
Retiring ChatGPT-authenticated configurationgpt-5.5Replace before October 14; API unaffected by this notice

Set the local default in ~/.codex/config.toml:

1
model = "gpt-6.1-sol"

Choose temporarily:

1
codex -m gpt-6.1-sol

Or use /model in the CLI and the model selector in the IDE extension. For hosted tasks, check the controls in that surface and its environment; access still depends on the account and workspace.

Use the GPT-6 comparison for current model access, API prices and product credits. The GPT-5.6 guide retains its dated model and system-card evidence. The Sol preview investigation is retained as June history.

AGENTS.md

AGENTS.md is Markdown guidance, not an invented YAML agent registry. Put repository conventions, commands, boundaries, and verification requirements in the repository root:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
# Repository guidance

## Commands
- Install: `bun install`
- Test: `bun test`
- Build: `bun run build`

## Change rules
- Keep patches scoped to the requested feature.
- Do not add dependencies without approval.
- Never commit secrets or generated credentials.

## Verification
- Run the focused test first.
- Run the build before handoff.
- Report any skipped checks and why.

Closer nested AGENTS.md files can provide subtree-specific instructions. Model, provider, approval, and sandbox defaults belong in Codex configuration, while task-specific constraints belong in the prompt.

Approvals and Sandbox

Task authorization, sandbox networking, command rules, approval reviewers, gh, and browser runtimes are separate controls. A request to publish a change can authorize the work while the sandbox still blocks the required destination.

Scroll sideways to compare columns
ControlWhat it governs
Prompt and AGENTS.mdTask scope, conventions and acceptance checks; they do not grant OS access
Permission profileFilesystem roots and local command networking
Approval policyWhich boundary-crossing requests can ask for approval
Approval reviewerWhether a person or Auto-review evaluates eligible requests
Command rulesMatching command arguments for execution outside the sandbox
ghIts installed version, authentication, proxy, cache and log rendering
Browser, MCP and connectorsTheir own runtime, transport, credentials and applicable approvals

Permission profiles are beta. Choose either default_permissions with [permissions] or legacy sandbox_mode with [sandbox_workspace_write]. A loaded legacy setting or --sandbox flag can select the legacy path; managed requirements can constrain the choice. Check configuration precedence and the active session rather than assuming the saved default won.

For scoped GitHub access, extend :workspace, enable command networking, and allow the exact required hosts. Domain restrictions require an active proxy: network.enabled = true alone permits networking without enforcing the domain table. Enable features.network_proxy = true or the administrator-managed proxy. Managed denies take precedence; request an approved exception where needed. Approved destinations can work inside the sandbox without repeated escalation. (Network enforcement)

approval_policy = "on-request" with approvals_reviewer = "auto_review" changes who reviews eligible requests. It does not enable networking, expand writable roots or remove managed restrictions. Computer Use app approvals remain separate. Routine allowed actions do not invoke the reviewer. (Auto-review)

Keep credentials scoped to the task, use isolated worktrees, and review the patch and tests. A command allow rule can make an outside-sandbox retry succeed while ordinary sandbox networking remains blocked. Redirection or complex shell syntax can change rule matching; inspect the exact invocation and prefer narrow prefixes. (Command rules)

Why gh works in a terminal but fails in Codex

Diagnose the failing layer before changing permissions:

  1. Identify local versus hosted execution, the installed and running Codex versions, gh --version, command location and active profile. Use /status and codex doctor --summary; a host connectivity pass does not prove sandbox connectivity.
  2. Try a small authenticated repository request inside the ordinary sandbox. Separate socket/DNS/proxy failures from HTTP authentication errors, cache-path failures and empty log output.
  3. Check the destination used by the actual log download. An API allowlist can succeed while a redirected log host remains blocked. Preserve managed policy and add only an authorized exact host.
  4. Validate a nonempty uncached log with a task-owned temporary cache. Confirm that an unlisted destination is denied and protected paths remain read-only. Retain the result and remove the temporary cache.

Our October 3 Debian case: Codex 0.160.0 and gh 2.102.0 retrieved a 114,048-byte CI log in a fresh default sandbox. The observed log host was results-receiver.actions.githubusercontent.com; an unlisted HTTPS destination received proxy 403, and protected paths stayed read-only. The old gh 2.46.0 produced empty rendered output in the comparison. This does not establish a minimum fixed version or a complete GitHub allowlist. (Local validation and limits, GitHub log fallback)

An unwritable inherited cache path and a stale browser-launcher environment were specific to that machine. Browser operation remained unverified: the running MCP process retained its earlier environment, and Chrome’s extension/native bridge were absent. Reload the affected runtime when shared sessions are idle and follow supported plugin setup. Working gh, web search or a connector does not prove browser startup; a new terminal may still attach to an existing daemon. Prevalence and time saved are unmeasured.

Local and Cloud Work

Local Codex surfaces work with the checked-out repository and the configured sandbox. Cloud tasks use a configured cloud environment and hosted execution. Check environment setup, secrets, internet access, and repository state separately for each surface.

Do not promise fixed concurrency, microVM startup time, message counts, or credit consumption unless the current official plan or account UI documents it. Limits vary by plan, task size, model, and rollout.

Current Codex releases support subagents; ask directly or request delegation in applicable instructions. Give independent tasks clear ownership and acceptance checks. Each subagent adds model and tool work. In eligible ChatGPT Work accounts, Ultra permits proactive parallel delegation; Max increases reasoning effort, while Ultrafast is a speed mode with separate access and rates. Responses Multi-agent is a separate beta API feature. These names do not promise cheaper task completion.

OpenAI’s September 28–October 2 update also introduces gradually rolling-out Dots, which carry ongoing responsibilities between conversations and can delegate to Work or Codex. Check eligible plans, regions and workspace controls before designing around that access.

Building an application with OpenAI APIs

If you are integrating agents into your own application, choose the API separately from the Codex client you use to work on its code.

The Agents API runs an OpenAI-managed Codex harness with saved sessions; the Agents SDK runs the loop in your application; the Responses API provides direct model integration. The Agents API launched in public beta September 10 and added hosted browser computer use September 29, with website approvals and sign-in handled by the application. That browser is a different environment from a local Browser plugin.

Cache reuse and Sol 6.1

Sol 6.1 makes stable, repeated context inexpensive to reuse: Standard API cached input is $0.10 per million tokens, versus $2 ordinary input, $2.50 cache writes and $10 output for requests with at most 272K input tokens. Reads cost 5% of ordinary input and half Sol 6’s read price. A cache write replaces the ordinary-input charge for those tokens. Longer requests and other tiers use different rates. (API pricing, worked cost examples)

High reuse is observable, not guaranteed. One completed local Sol 6.1 session on October 3 recorded 8,325,760 cached input tokens out of 8,540,577 total input: 97.5%. That is a token-weighted share from client counters, not the percentage of calls that hit cache, an API invoice, a quota-saving estimate or a quality benchmark. (Observation and method)

Preserve stable instructions, tool definitions and prompt prefixes; use targeted reads and bounded logs. For direct Responses integrations, the Prompt Caching dashboard and diagnostics help measure reuse and explain misses. Work/Codex’s Standard purchased-credit card separately lists 2.5 cached-input credits versus 50 ordinary-input credits per million, with no separate cache-write fee. Included subscription allowances need their own meter. Cheap cached input still occupies context and does not make repeated polling useful. (Credit rules, cache measurement)

Wait for long jobs without repeated model checks

Repeated empty status checks can return control to the model and add another inference turn. A March 6 CLI report describes this during builds; a July 9 report attributes repeated polling to a 60-second blocking-wait instruction. A Reddit discussion checked October 3 reports similar behavior and inconsistent improvement from AGENTS.md. These are user reports, not a measured platform-wide failure rate or audited savings estimate.

60 seconds is not a documented universal polling default. The current CLI configuration reference defines background_terminal_max_timeout = 300000 as the default maximum empty write_stdin poll window: five minutes. This ceiling does not make the model choose five-minute waits or override a surface’s tool limits and higher-priority instructions. Keep command refresh, tool wait duration, model re-entry and user updates separate.

For GitHub CI, select the run for the reviewed commit and keep one watcher:

1
gh run watch RUN_ID --repo OWNER/REPO --compact --exit-status --interval 60

Replace the placeholders with the actual run and repository. gh run watch polls GitHub inside that process; its own default refresh is three seconds, and the example selects 60. This reduces GitHub requests but does not control Codex’s model calls. Retain the exec session/process ID and use the longest suitable output/completion wait supported by the active surface. Use completion notifications where supported; generic wake-on-exit remains an open CLI feature request, not a portable flag. Verify the run’s SHA and final conclusion before merging.

Put a scoped convention in AGENTS.md:

1
2
3
4
5
6
7
For healthy long-running jobs, retain one process or CI watcher and its ID.
Prefer supported completion notifications; otherwise use long waits within
the active tool and instruction limits. Check sooner for actionable output,
a relevant deadline, or evidence of a stall. Report meaningful state changes.
Do not restart the job or add duplicate watchers after an empty wait.
After completion, verify the exit status and relevant result, then clean up
the task-owned watcher and temporary output.

Measure empty checks, model turns, cached/uncached input, output and elapsed time before claiming savings. A warm prefix can discount repeated input, but output, reasoning and any cache misses still cost. No controlled polling-cost comparison was run for this refresh.

Plan Overview

For current prices, regional charges and included usage, use OpenAI’s pricing page and the signed-in account. The GPT-6 comparison separates purchased credits from API token rates and included subscription limits. Pro $200 reopening and the allowance transition are summarized in the access record, with detailed terms in the launch offers guide. Model availability does not imply unlimited use or a fixed monthly task count.

Usage limits and resets

Earlier reviews recorded differing five-hour and weekly displays across accounts. Use the current usage dashboard and pricing documentation for the windows that apply to your plan; old screenshots do not establish your current allowance. Model choice, reasoning effort, context, tool calls, searches, caching, retries, and subagents can change how quickly a task consumes the active allowance.

Check the signed-in usage dashboard and /status before purchasing overflow. Current CLI releases also expose /usage analytics; available views depend on the client and account. Keep these paths separate:

  • scheduled recovery ends the current allowance window;
  • provider-wide or incident resets refresh the named group;
  • banked resets are expiring, eligible user-controlled grants;
  • purchased full resets are documented for eligible personal Plus/Pro accounts and start a new weekly period rather than stack a bonus week; displayed prices remain account-specific; and
  • purchased credits fund supported work after included limits without a documented counter refill.

The September 27 verification record documents the source-checked reset mechanics and exclusions. Earlier $8 Plus/$80 Pro observations are historical, not universal prices; check your current offer and applicable usage windows. The current pricing page says Pro has no five-hour limit.

Read AI Subscription Capacity Is Perishable for the scheduling rule, the reset chronology for dated scope, and the cost-saving playbook before choosing between waiting, a displayed reset, credits, or another route.

Current Model Evidence

benchmark artifact

Codex Model Context and Alternatives

ModelProviderStatusContextInput priceOutput priceCoding signalTool-use signalBenchmark evidenceSpeedVerdictSourcesChecked
GPT-6 AstraOpenAI active
API ID gpt-6-astra; eligible Codex access. October 3 plan check: Pro 200 subscriptions have reopened; allowance and feature access depend on plan and account.
1.05M$10.00 / 1M
$10 input / $50 output per 1M; higher long-context rates above 272K input
$50.00 / 1Mnot verifiednot verified
  • AA Intelligence Index v4.3 (Max): 53 (independent)
  • AIHackers CAR: not-run (site-owned)
not verifiedEvaluate Low for planning; raise effort after failed acceptance checks. Max benchmark evidence does not establish Low performance.OpenAI Astra model documentation, Artificial Analysis Astra evaluation, September 9, OpenAI Pro subscription access, checked October 32026-09-14
GPT-6.1 SolOpenAI active
API gpt-6.1-sol; Plus/Pro/Business/Enterprise/Edu Work/Codex rollout; Enterprise/Edu off until admin enablement. Standard/Fast where available; Ultrafast forthcoming.
1.05M$2.00 / 1M
Standard $2 ordinary / $0.10 read / $2.50 write / $10 output per 1M; >272K: 2x input/cache, 1.5x output; credits separate
$10.00 / 1MOpenAI describes near-Astra performance; vendor positioning, not a site-owned quality finding.API ID gpt-6.1-sol; Responses tool calling; compare accepted work before replacing a tested route.
  • AIHackers controlled task evaluation and CAR: not-run; no Sol 6 benchmark scores inherited (site-owned)
not verifiedComplex-work candidate with lower cache-read pricing than Sol 6; ordinary input/output rates unchanged.OpenAI GPT-6.1 Sol identity and API pricing, OpenAI Standard API prices, checked October 3, OpenAI GPT-6.1 Sol Work/Codex rollout2026-10-03
GPT-6 SolOpenAI active
API gpt-6-sol; paid Work/Codex rollout, separate from Chat. Client/workspace access varies.
1.05M$2.00 / 1M
$2 input / $0.20 cache / $2.50 cache write / $10 output; above 272K: 2x input/cache, 1.5x output
$10.00 / 1MAA Coding Agent Index 57 at max in Codex harness; predecessor 55 in same report.OpenAI starting effort Medium; API dollars separate from Work/Codex credits.
  • AA Coding Agent Index (September 22, max): 57; $2.99 per benchmark task, not AIHackers accepted-result cost (independent)
  • AIHackers CAR: not-run (site-owned)
not verifiedEveryday and complex coding candidate; lower prices do not remove quality and review costs.OpenAI GPT-6 Sol, Artificial Analysis GPT-6 Sol and Luna evaluation2026-09-27
GPT-6 LunaOpenAI active
API gpt-6-luna; paid Work/Codex rollout and Free/Go desktop access where available. Not Chat.
1.05M$0.10 / 1M
$0.10 input / $0.01 cache / $0.125 cache write / $0.50 output; above 272K: 2x input/cache, 1.5x output
$0.50 / 1MAA Coding Agent Index 41 at max in Codex harness; predecessor 43 in same report.OpenAI starting effort High for focused work; evaluate review burden before routing.
  • AA Coding Agent Index (September 22, max): 41; cheaper but lower score than predecessor in this evaluation (independent)
  • AIHackers CAR: not-run (site-owned)
not verifiedFocused high-volume candidate with acceptance checks; cheaper output is not a universal capability upgrade.OpenAI GPT-6 Luna, Artificial Analysis GPT-6 Sol and Luna evaluation2026-09-27
Claude Opus 5.5Anthropic active
September 22 release; claude-opus-5-5 on Claude API and documented cloud routes. Verify plan and region.
1M$4.00 / 1M
$4 input / $0.20 cache read / $20 output per 1M; 5m cache write $5; 1h write $8; Fast separate
$20.00 / 1MAA Terminal-Bench 4.0: 59.6% at max with default fallback; level with Astra xhigh in that run.Always-on adaptive thinking; medium default. API migration has breaking changes.
  • AA Intelligence Index (September 22, max): 58; highest measured at release, not directly comparable with July index scores (independent)
  • AIHackers CAR: not-run (site-owned)
Anthropic reports over 30% faster output generation than Opus 5; not an AIHackers measurement.Premium coding and knowledge-work candidate; start medium and measure the gain from higher effort.Claude Opus 5.5 specifications and pricing, Anthropic Opus 5.5 launch, Artificial Analysis Opus 5.5 evaluation2026-09-27
Kimi K2.7 CodeMoonshot AI active
Cheaper routine Kimi coding API lane; HighSpeed is the same model at higher token prices.
256K$0.95 / 1M
$0.19 cache-hit / $0.95 cache-miss input / $4.00 output per 1M tokens; HighSpeed doubles those rates
$4.00 / 1MKimi K2.7 Code remains the lower-cost 256K coding lane after K3; independent normalized benchmarks are not imported.OpenAI-compatible API; thinking mode required in the documented K2.7 Code quickstart.
  • Program-Bench improvement vs K2.6: +10.4% (vendor)
  • MCP Mark Verified improvement vs K2.6: +11.4% (vendor)
  • SWE Marathon improvement vs K2.6: +76.2% (vendor)
  • Reasoning-token use vs K2.6: 30% lower (vendor)
  • AIHackers repo eval: not verified (site-owned)
HighSpeed model ID exists at a higher token price; latency not independently measured here.Cheaper routine Kimi coding API lane when Kimi routing fits and 256K context is enough.Kimi K2.7 Code quickstart [archive], Kimi K2.7 Code pricing [archive], Kimi Code K2.7 release notes [archive], SWE-bench, Berkeley Function Calling Leaderboard2026-06-28

October 3 Sol 6.1 identity/economics added; other benchmark evidence retains its own date. Access varies by plan and client; no Sol 6 scores transferred to 6.1. These rows are not a substitute for repository tests.

Benchmarks do not evaluate the entire Codex product. Measure accepted patches, tests, retries, wall-clock time, and review effort in the actual sandbox and repository.

Migration From Old Configurations

Search configuration and scripts for deprecated models:

1
rg 'gpt-5\.2|gpt-5\.3-codex|gpt-5\.4|gpt-5\.5' ~/.codex . --glob '*.toml' --glob '*.json' --glob '*.md'

For ChatGPT-authenticated Codex, replace retired defaults and prepare GPT-5.5 settings for October 14 with a GPT-6 model available to the account and suited to the workload. Leave historical articles and fixed eval baselines unchanged.

Sources


Model selection, prices/access, permissions, platform updates and polling sources checked October 3, 2026. The local networking receipt and cache counters describe one machine/session; browser repair, prevalence and polling savings remain unverified or unmeasured. Historical resets and benchmarks retain their dates. Plans, clients and account behavior can change independently.