Use this guide to set up Codex, choose a model, and check the controls that affect your work.
- Start locally: install the CLI, choose how to sign in, and add repository instructions.
- Check access and costs: verify your plan, usage meters, model availability, and billing route.
- Run bounded tasks: distinguish permissions from task authorization, inspect failures, and keep one watcher for long jobs.
October 3 guide update
Use GPT-6.1 Sol for complex coding and Luna High for focused work; evaluate Astra Low for difficult planning and orchestration. For Sol 6.1, start with the effort available by default in your client, then adjust against acceptance checks. These are starting points from OpenAI’s model guidance, subject to your client, plan, rollout and workspace settings.
Pro $200 purchases have reopened. Eligible existing subscribers retain their previous allowance only through October 29, 2026, then move to a lower allowance at the same $200 monthly price. The launch offers guide details the September eligibility window and same-account return conditions; the current access check supersedes our purchase-pause wording. Pro 200 remains a distinct tier from Pro 500 and does not include Ultrafast. GPT-5.6 remains available during rollout.
OpenAI Codex is available through CLI, IDE, desktop and cloud workflows. GPT-6.1 Sol, GPT-6 Sol and Luna are in Work and Codex rather than ordinary Chat. GPT-5.4 and GPT-5.4-mini retired from ChatGPT-authenticated Codex on August 31. GPT-5.5 is scheduled to retire from ChatGPT, Work and Codex on October 14, 2026; that notice does not retire it from the API.
This refresh also covers cheap cache reuse, GitHub networking, and long-job polling. Current documentation, dated user reports and our local observations have different evidence scopes; earlier reset and benchmark records retain their dates.
Install
| |
Codex can authenticate with ChatGPT or an API key:
- ChatGPT authentication includes Codex according to the user’s Free, Go, Plus, Pro, Business, Edu, or Enterprise plan.
- API-key authentication uses models available to that API key and API token billing.
- API-key authentication does not provide Codex cloud features such as hosted code review or Slack integration.
Do not hard-code an expected CLI version in an evergreen setup guide. Verify the installed version and current release notes.
Current Model Selection
| Need | Model and starting effort | Access boundary |
|---|---|---|
| Complex coding and agent work | gpt-6.1-sol, client default; API Medium | Plus/Pro/Business/Enterprise/Edu rollout; Enterprise/Edu require administrator enablement; API separately billed |
| Existing Sol 6 integration or dated evaluation | gpt-6-sol | Keep its own results; Sol 6.1 has separate identity and economics |
| Focused, repeatable work | gpt-6-luna, High | Paid Work/Codex plans; Free/Go desktop rollout; API separately billed |
| Hard planning, research and multi-step judgment | gpt-6-astra, Light (low) | Check your account and client |
| Prior family integration | GPT-5.6 Sol, Terra, Luna | Still available during rollout; retain only when your workflow justifies it |
| Retiring ChatGPT-authenticated configuration | gpt-5.5 | Replace before October 14; API unaffected by this notice |
Set the local default in ~/.codex/config.toml:
| |
Choose temporarily:
| |
Or use /model in the CLI and the model selector in the IDE extension. For hosted tasks, check the controls in that surface and its environment; access still depends on the account and workspace.
Use the GPT-6 comparison for current model access, API prices and product credits. The GPT-5.6 guide retains its dated model and system-card evidence. The Sol preview investigation is retained as June history.
AGENTS.md
AGENTS.md is Markdown guidance, not an invented YAML agent registry. Put repository conventions, commands, boundaries, and verification requirements in the repository root:
| |
Closer nested AGENTS.md files can provide subtree-specific instructions. Model, provider, approval, and sandbox defaults belong in Codex configuration, while task-specific constraints belong in the prompt.
Approvals and Sandbox
Task authorization, sandbox networking, command rules, approval reviewers, gh, and browser runtimes are separate controls. A request to publish a change can authorize the work while the sandbox still blocks the required destination.
| Control | What it governs |
|---|---|
Prompt and AGENTS.md | Task scope, conventions and acceptance checks; they do not grant OS access |
| Permission profile | Filesystem roots and local command networking |
| Approval policy | Which boundary-crossing requests can ask for approval |
| Approval reviewer | Whether a person or Auto-review evaluates eligible requests |
| Command rules | Matching command arguments for execution outside the sandbox |
gh | Its installed version, authentication, proxy, cache and log rendering |
| Browser, MCP and connectors | Their own runtime, transport, credentials and applicable approvals |
Permission profiles are beta. Choose either default_permissions with [permissions] or legacy sandbox_mode with [sandbox_workspace_write]. A loaded legacy setting or --sandbox flag can select the legacy path; managed requirements can constrain the choice. Check configuration precedence and the active session rather than assuming the saved default won.
For scoped GitHub access, extend :workspace, enable command networking, and allow the exact required hosts. Domain restrictions require an active proxy: network.enabled = true alone permits networking without enforcing the domain table. Enable features.network_proxy = true or the administrator-managed proxy. Managed denies take precedence; request an approved exception where needed. Approved destinations can work inside the sandbox without repeated escalation. (Network enforcement)
approval_policy = "on-request" with approvals_reviewer = "auto_review" changes who reviews eligible requests. It does not enable networking, expand writable roots or remove managed restrictions. Computer Use app approvals remain separate. Routine allowed actions do not invoke the reviewer. (Auto-review)
Keep credentials scoped to the task, use isolated worktrees, and review the patch and tests. A command allow rule can make an outside-sandbox retry succeed while ordinary sandbox networking remains blocked. Redirection or complex shell syntax can change rule matching; inspect the exact invocation and prefer narrow prefixes. (Command rules)
Why gh works in a terminal but fails in Codex
Diagnose the failing layer before changing permissions:
- Identify local versus hosted execution, the installed and running Codex versions,
gh --version, command location and active profile. Use/statusandcodex doctor --summary; a host connectivity pass does not prove sandbox connectivity. - Try a small authenticated repository request inside the ordinary sandbox. Separate socket/DNS/proxy failures from HTTP authentication errors, cache-path failures and empty log output.
- Check the destination used by the actual log download. An API allowlist can succeed while a redirected log host remains blocked. Preserve managed policy and add only an authorized exact host.
- Validate a nonempty uncached log with a task-owned temporary cache. Confirm that an unlisted destination is denied and protected paths remain read-only. Retain the result and remove the temporary cache.
Our October 3 Debian case: Codex 0.160.0 and gh 2.102.0 retrieved a 114,048-byte CI log in a fresh default sandbox. The observed log host was results-receiver.actions.githubusercontent.com; an unlisted HTTPS destination received proxy 403, and protected paths stayed read-only. The old gh 2.46.0 produced empty rendered output in the comparison. This does not establish a minimum fixed version or a complete GitHub allowlist. (Local validation and limits, GitHub log fallback)
An unwritable inherited cache path and a stale browser-launcher environment were specific to that machine. Browser operation remained unverified: the running MCP process retained its earlier environment, and Chrome’s extension/native bridge were absent. Reload the affected runtime when shared sessions are idle and follow supported plugin setup. Working gh, web search or a connector does not prove browser startup; a new terminal may still attach to an existing daemon. Prevalence and time saved are unmeasured.
Local and Cloud Work
Local Codex surfaces work with the checked-out repository and the configured sandbox. Cloud tasks use a configured cloud environment and hosted execution. Check environment setup, secrets, internet access, and repository state separately for each surface.
Do not promise fixed concurrency, microVM startup time, message counts, or credit consumption unless the current official plan or account UI documents it. Limits vary by plan, task size, model, and rollout.
Current Codex releases support subagents; ask directly or request delegation in applicable instructions. Give independent tasks clear ownership and acceptance checks. Each subagent adds model and tool work. In eligible ChatGPT Work accounts, Ultra permits proactive parallel delegation; Max increases reasoning effort, while Ultrafast is a speed mode with separate access and rates. Responses Multi-agent is a separate beta API feature. These names do not promise cheaper task completion.
OpenAI’s September 28–October 2 update also introduces gradually rolling-out Dots, which carry ongoing responsibilities between conversations and can delegate to Work or Codex. Check eligible plans, regions and workspace controls before designing around that access.
Building an application with OpenAI APIs
If you are integrating agents into your own application, choose the API separately from the Codex client you use to work on its code.
The Agents API runs an OpenAI-managed Codex harness with saved sessions; the Agents SDK runs the loop in your application; the Responses API provides direct model integration. The Agents API launched in public beta September 10 and added hosted browser computer use September 29, with website approvals and sign-in handled by the application. That browser is a different environment from a local Browser plugin.
Cache reuse and Sol 6.1
Sol 6.1 makes stable, repeated context inexpensive to reuse: Standard API cached input is $0.10 per million tokens, versus $2 ordinary input, $2.50 cache writes and $10 output for requests with at most 272K input tokens. Reads cost 5% of ordinary input and half Sol 6’s read price. A cache write replaces the ordinary-input charge for those tokens. Longer requests and other tiers use different rates. (API pricing, worked cost examples)
High reuse is observable, not guaranteed. One completed local Sol 6.1 session on October 3 recorded 8,325,760 cached input tokens out of 8,540,577 total input: 97.5%. That is a token-weighted share from client counters, not the percentage of calls that hit cache, an API invoice, a quota-saving estimate or a quality benchmark. (Observation and method)
Preserve stable instructions, tool definitions and prompt prefixes; use targeted reads and bounded logs. For direct Responses integrations, the Prompt Caching dashboard and diagnostics help measure reuse and explain misses. Work/Codex’s Standard purchased-credit card separately lists 2.5 cached-input credits versus 50 ordinary-input credits per million, with no separate cache-write fee. Included subscription allowances need their own meter. Cheap cached input still occupies context and does not make repeated polling useful. (Credit rules, cache measurement)
Wait for long jobs without repeated model checks
Repeated empty status checks can return control to the model and add another inference turn. A March 6 CLI report describes this during builds; a July 9 report attributes repeated polling to a 60-second blocking-wait instruction. A Reddit discussion checked October 3 reports similar behavior and inconsistent improvement from AGENTS.md. These are user reports, not a measured platform-wide failure rate or audited savings estimate.
60 seconds is not a documented universal polling default. The current CLI configuration reference defines background_terminal_max_timeout = 300000 as the default maximum empty write_stdin poll window: five minutes. This ceiling does not make the model choose five-minute waits or override a surface’s tool limits and higher-priority instructions. Keep command refresh, tool wait duration, model re-entry and user updates separate.
For GitHub CI, select the run for the reviewed commit and keep one watcher:
| |
Replace the placeholders with the actual run and repository. gh run watch polls GitHub inside that process; its own default refresh is three seconds, and the example selects 60. This reduces GitHub requests but does not control Codex’s model calls. Retain the exec session/process ID and use the longest suitable output/completion wait supported by the active surface. Use completion notifications where supported; generic wake-on-exit remains an open CLI feature request, not a portable flag. Verify the run’s SHA and final conclusion before merging.
Put a scoped convention in AGENTS.md:
| |
Measure empty checks, model turns, cached/uncached input, output and elapsed time before claiming savings. A warm prefix can discount repeated input, but output, reasoning and any cache misses still cost. No controlled polling-cost comparison was run for this refresh.
Plan Overview
For current prices, regional charges and included usage, use OpenAI’s pricing page and the signed-in account. The GPT-6 comparison separates purchased credits from API token rates and included subscription limits. Pro $200 reopening and the allowance transition are summarized in the access record, with detailed terms in the launch offers guide. Model availability does not imply unlimited use or a fixed monthly task count.
Usage limits and resets
Earlier reviews recorded differing five-hour and weekly displays across accounts. Use the current usage dashboard and pricing documentation for the windows that apply to your plan; old screenshots do not establish your current allowance. Model choice, reasoning effort, context, tool calls, searches, caching, retries, and subagents can change how quickly a task consumes the active allowance.
Check the signed-in usage dashboard and /status before purchasing overflow. Current CLI releases also expose /usage analytics; available views depend on the client and account. Keep these paths separate:
- scheduled recovery ends the current allowance window;
- provider-wide or incident resets refresh the named group;
- banked resets are expiring, eligible user-controlled grants;
- purchased full resets are documented for eligible personal Plus/Pro accounts and start a new weekly period rather than stack a bonus week; displayed prices remain account-specific; and
- purchased credits fund supported work after included limits without a documented counter refill.
The September 27 verification record documents the source-checked reset mechanics and exclusions. Earlier $8 Plus/$80 Pro observations are historical, not universal prices; check your current offer and applicable usage windows. The current pricing page says Pro has no five-hour limit.
Read AI Subscription Capacity Is Perishable for the scheduling rule, the reset chronology for dated scope, and the cost-saving playbook before choosing between waiting, a displayed reset, credits, or another route.
Current Model Evidence
benchmark artifact
Codex Model Context and Alternatives
| Model | Provider | Status | Context | Input price | Output price | Coding signal | Tool-use signal | Benchmark evidence | Speed | Verdict | Sources | Checked |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-6 Astra | OpenAI |
active API ID gpt-6-astra; eligible Codex access. October 3 plan check: Pro 200 subscriptions have reopened; allowance and feature access depend on plan and account. | 1.05M | $10.00 / 1M $10 input / $50 output per 1M; higher long-context rates above 272K input | $50.00 / 1M | not verified | not verified |
| not verified | Evaluate Low for planning; raise effort after failed acceptance checks. Max benchmark evidence does not establish Low performance. | OpenAI Astra model documentation, Artificial Analysis Astra evaluation, September 9, OpenAI Pro subscription access, checked October 3 | 2026-09-14 |
| GPT-6.1 Sol | OpenAI |
active API gpt-6.1-sol; Plus/Pro/Business/Enterprise/Edu Work/Codex rollout; Enterprise/Edu off until admin enablement. Standard/Fast where available; Ultrafast forthcoming. | 1.05M | $2.00 / 1M Standard $2 ordinary / $0.10 read / $2.50 write / $10 output per 1M; >272K: 2x input/cache, 1.5x output; credits separate | $10.00 / 1M | OpenAI describes near-Astra performance; vendor positioning, not a site-owned quality finding. | API ID gpt-6.1-sol; Responses tool calling; compare accepted work before replacing a tested route. |
| not verified | Complex-work candidate with lower cache-read pricing than Sol 6; ordinary input/output rates unchanged. | OpenAI GPT-6.1 Sol identity and API pricing, OpenAI Standard API prices, checked October 3, OpenAI GPT-6.1 Sol Work/Codex rollout | 2026-10-03 |
| GPT-6 Sol | OpenAI |
active API gpt-6-sol; paid Work/Codex rollout, separate from Chat. Client/workspace access varies. | 1.05M | $2.00 / 1M $2 input / $0.20 cache / $2.50 cache write / $10 output; above 272K: 2x input/cache, 1.5x output | $10.00 / 1M | AA Coding Agent Index 57 at max in Codex harness; predecessor 55 in same report. | OpenAI starting effort Medium; API dollars separate from Work/Codex credits. |
| not verified | Everyday and complex coding candidate; lower prices do not remove quality and review costs. | OpenAI GPT-6 Sol, Artificial Analysis GPT-6 Sol and Luna evaluation | 2026-09-27 |
| GPT-6 Luna | OpenAI |
active API gpt-6-luna; paid Work/Codex rollout and Free/Go desktop access where available. Not Chat. | 1.05M | $0.10 / 1M $0.10 input / $0.01 cache / $0.125 cache write / $0.50 output; above 272K: 2x input/cache, 1.5x output | $0.50 / 1M | AA Coding Agent Index 41 at max in Codex harness; predecessor 43 in same report. | OpenAI starting effort High for focused work; evaluate review burden before routing. |
| not verified | Focused high-volume candidate with acceptance checks; cheaper output is not a universal capability upgrade. | OpenAI GPT-6 Luna, Artificial Analysis GPT-6 Sol and Luna evaluation | 2026-09-27 |
| Claude Opus 5.5 | Anthropic |
active September 22 release; claude-opus-5-5 on Claude API and documented cloud routes. Verify plan and region. | 1M | $4.00 / 1M $4 input / $0.20 cache read / $20 output per 1M; 5m cache write $5; 1h write $8; Fast separate | $20.00 / 1M | AA Terminal-Bench 4.0: 59.6% at max with default fallback; level with Astra xhigh in that run. | Always-on adaptive thinking; medium default. API migration has breaking changes. |
| Anthropic reports over 30% faster output generation than Opus 5; not an AIHackers measurement. | Premium coding and knowledge-work candidate; start medium and measure the gain from higher effort. | Claude Opus 5.5 specifications and pricing, Anthropic Opus 5.5 launch, Artificial Analysis Opus 5.5 evaluation | 2026-09-27 |
| Kimi K2.7 Code | Moonshot AI |
active Cheaper routine Kimi coding API lane; HighSpeed is the same model at higher token prices. | 256K | $0.95 / 1M $0.19 cache-hit / $0.95 cache-miss input / $4.00 output per 1M tokens; HighSpeed doubles those rates | $4.00 / 1M | Kimi K2.7 Code remains the lower-cost 256K coding lane after K3; independent normalized benchmarks are not imported. | OpenAI-compatible API; thinking mode required in the documented K2.7 Code quickstart. |
| HighSpeed model ID exists at a higher token price; latency not independently measured here. | Cheaper routine Kimi coding API lane when Kimi routing fits and 256K context is enough. | Kimi K2.7 Code quickstart [archive], Kimi K2.7 Code pricing [archive], Kimi Code K2.7 release notes [archive], SWE-bench, Berkeley Function Calling Leaderboard | 2026-06-28 |
October 3 Sol 6.1 identity/economics added; other benchmark evidence retains its own date. Access varies by plan and client; no Sol 6 scores transferred to 6.1. These rows are not a substitute for repository tests.
Benchmarks do not evaluate the entire Codex product. Measure accepted patches, tests, retries, wall-clock time, and review effort in the actual sandbox and repository.
Migration From Old Configurations
Search configuration and scripts for deprecated models:
| |
For ChatGPT-authenticated Codex, replace retired defaults and prepare GPT-5.5 settings for October 14 with a GPT-6 model available to the account and suited to the workload. Leave historical articles and fixed eval baselines unchanged.
Sources
- OpenAI: permissions, Auto-review, command rules, subagents, configuration reference, and current updates — checked October 3, 2026; current-source archives
archive-pending - OpenAI Codex manual: Overview, models (Archive), pricing (Archive), and configuration
- OpenAI: GPT-5.6 general availability
- OpenAI: GPT-5.6 System Card
- Community paid-reset observations: $8 Plus, $80 Pro 20x and shifted recovery, and uneven rollout
Related links
- GPT-5.6 model evidence and dated context
- Prompt caching and coding-agent costs
- Plan work around subscription limits
- Codex reset reports and their dates
- The ExploitGym incident and access boundaries
- Compare Codex, Claude and Kimi coding workflows
- Compare Codex, Claude and Cursor coding workflows
- Codex cloud-dependency risks
- Write useful AGENTS.md instructions
- Reduce LLM costs with an acceptance-first playbook
- Choose Dots, Work or Codex for a coordination task
- Kimi Code setup and access boundaries
- What the documented AI-agent security incidents show
Model selection, prices/access, permissions, platform updates and polling sources checked October 3, 2026. The local networking receipt and cache counters describe one machine/session; browser repair, prevalence and polling savings remain unverified or unmeasured. Historical resets and benchmarks retain their dates. Plans, clients and account behavior can change independently.