The direct answer: coding-agent loops structurally use more requests, context, and output than a single model answer. Prompt caching can sharply discount repeated prefixes, but it does not erase context-window occupancy or make history free. There is no source-backed universal Codex multiplier, no documented “five recently edited files / 50K tokens” rule, and no controlled evidence that OpenCode is always more efficient.

October 3, 2026 cache/economics check: the mechanics, Sol 6.1 arithmetic and logging recipe below use current provider documentation. The Claude Code workflow subsection and historical incident, issue and harness observations retain their stated dates. All cost scenarios are illustrations; AIHackers has run no paid cache or accepted-result comparison for this refresh.

Prompt caching is reusable prefix computation

Prompt caching reuses the provider-side computation for an exact prompt prefix: stable instructions, tool definitions, examples, files, and earlier turns placed before changing content. The model still generates a new answer.

It is not:

  • response caching, which returns a previously saved answer;
  • memory, which stores selected facts across sessions;
  • RAG, which retrieves documents for the current request;
  • context compaction, which replaces older history with a summary; or
  • a local package, container, or build cache.

OpenAI’s minimum cacheable length is model-dependent; GPT-5.6 and later require at least 1,024 tokens. Exact prefix order matters: put stable material first and changing task data last. A stateful Responses chain can simplify request construction, but OpenAI still bills previous input in the chain as input. Eligible reuse may receive a cache discount; statefulness alone does not make it free.

Seven numbers that should not be mixed

Scroll sideways to compare columns
MeasureWhat it answersWhat it does not prove
Total input tokensHow much input crossed model context across requestsHow much was recomputed at full price
Cached readsHow much repeated prefix used an existing cache entryThat those tokens left the context window
Cache writesHow much prefix was stored for possible later reuseThat a later request actually hit it
Output and reasoningHow much new model work was generatedWhether the result was accepted
Context occupancyHow full the current model window isAPI cost or subscription quota
API billingDollar cost under model, cache, tier, and long-context rulesChatGPT/Codex plan consumption
Subscription quota or creditsProduct-specific allowance consumedA portable API token price

Why agent loops multiply work

A coding harness repeatedly asks a model to inspect, decide, act, and verify. A model response, client-managed tool result, retry, compaction recovery, or subagent turn can add another model request and another growing prefix. A provider-managed loop can be different: OpenAI’s Responses API can execute several built-in or custom tool calls within one client request.

Either way, more loop steps mean more model work. Codex’s current documentation explicitly says subagent workflows consume more tokens than comparable single-agent runs because each subagent performs its own model and tool work. That supports the structural claim—not universal figures such as “3–5×,” “8–12 calls,” or “40–80% waste.”

The widely repeated 40–80% number points the other way. OpenAI says Responses produced a 40–80% improvement in cache utilization versus Chat Completions in internal tests. It is not a measured Codex inflation rate.

Caching and compaction solve different problems

Caching discounts a repeated prefix while leaving it in the request and context. Compaction replaces older history with a shorter representation so the conversation can continue within its window.

Compaction can be beneficial and still have a cost. If its summary drops the exact file already inspected, the failed hypothesis, a test result, or the next edit, the agent may rediscover that information. The new compacted prefix may also need a new cache write before later turns can reuse it.

This is a failure mode to measure, not a universal diagnosis. A closed Codex v0.118 report initially blamed a lower compaction threshold; the reporter’s controlled test reversed that claim, and the issue was closed as not a bug. A separate July 24 Desktop report remains open and describes repeated compaction and rereads in one session. It has no maintainer-confirmed root cause.

OpenAI and Anthropic price caches differently

OpenAI prompt caching, checked October 3, is enabled by default for supported models. GPT-5.6 and later require at least 1,024 visible input tokens and support implicit or explicit breakpoints. prompt_cache_options.ttl currently supports only 30m, also the default minimum lifetime after the last write or reuse; OpenAI may retain entries longer. Reuse refreshes that lifetime without another write charge. Earlier models have different retention rules.

For GPT-5.6 and later, OpenAI cache writes cost 1.25× ordinary input, replacing that ordinary-input charge for the written tokens. Sol 6.1 reads cost 5% of ordinary input; most other GPT-5.6-and-later models use 10%. Neither rate is a whole-workload discount. Cached tokens remain in context and count toward token-per-minute limits. Use the pricing owner for model-specific rates, service tiers and the full-request multiplier above 272K input.

Anthropic prompt caching, checked October 3, offers automatic caching or explicit breakpoints. Five-minute writes cost 1.25× base input, one-hour writes 2×. Read ratios are model-specific: 0.05× for Opus 5.5, 0.025× for Fable 5.1/Mythos 5.1, and 0.1× for the other rows in its current table. Reuse refreshes the applicable duration, measured from the start of the request. Cacheable minimums differ by model; these are not OpenAI’s retention or pricing rules. Anthropic’s rate-limit rules exclude reads from input-token-per-minute usage except on Haiku 3.5; ordinary input and writes still count.

For independent requests with changing suffixes, a matching prefix alone may not be enough. OpenAI needs a matching eligible lookup boundary; put an explicit breakpoint after shared material when the implicit boundary would fall after changing data. In explicit-only mode, content after the last breakpoint uses ordinary input pricing, and a request without explicit breakpoints neither reads nor writes a cache. A request can create up to four writes. Top-level instructions cannot hold an explicit breakpoint; use a supported content block inside a developer message. Check usage fields before claiming reuse.

Worked example: traffic is not billed cost

Assume four GPT-6.1 Sol requests, each with an unchanged 8,000-token prefix, 1,000 changing tokens and 500 billable output tokens. Use explicit-only caching with a breakpoint after the prefix, before the changing suffix. At Standard short-context rates, ordinary input is $2/M, cached reads $0.10/M, writes $2.50/M and output $10/M. (Sol 6.1 pricing)

Scroll sideways to compare columns
PortionRaw trafficBilling treatmentCost
First stable prefix8,000Cache write$0.0200
Four changing inputs4,000Uncached input$0.0080
Three reused prefixes24,000Cached read$0.0024
Four outputs2,000Output$0.0200
Total38,000 tokensMixed$0.0504

Without caching, 36,000 ordinary input tokens cost $0.0720; the same output brings the total to $0.0920. The illustrated reduction is 45.2%, with raw traffic and context unchanged. This sequence includes its cold write and assumes all three subsequent prefixes match within the cache lifetime. It excludes tool charges, retries, human review and long-context/processing/regional modifiers.

Cold versus warm continuation

In that same setup, the cold request costs $0.0200 write + $0.0020 ordinary input + $0.0050 output = $0.0270. A matching warm continuation costs $0.0008 read + $0.0020 ordinary input + $0.0050 output = $0.0078. The uncached request costs $0.0230, so the cold write is initially more expensive. Continued useful reuse can recover that overhead; opening another session or changing the reusable prefix requires checking it again.

Cache reuse scenarios: show the denominator

Three ratios answer different questions: requests with any cached read / all requests counts hit requests; total cached-read tokens / total input tokens measures token-weighted reuse; 1 − workload cost / stated no-cache cost measures cost reduction. Aggregate token counts first. A hit on a short request does not carry the same weight as a hit on a long one.

For these hypothetical scenarios, aggregate 1M input + 100K output across individual Standard Sol 6.1 requests, each at or below 272K input. The no-cache comparator bills all input at $2/M and output at $10/M: $3.00. Reads, writes and ordinary input partition the same 1M input; writes are never charged again as ordinary input.

Scroll sideways to compare columns
Token-weighted cached-read shareReadsUnread inputTotal: unread input ordinary, no writesTotal: all unread input writtenReduction vs $3.00: ordinary / write case
0%01M$3.000$3.5000% / −16.7%
50%500K500K$2.050$2.30031.7% / 23.3%
90%900K100K$1.290$1.34057.0% / 55.3%
99%990K10K$1.119$1.12462.7% / 62.5%

The no-write column represents ordinary misses or changing input after the last explicit breakpoint; its reused entries were prepared outside this accounting window. The write column charges every unread token as a write within the window. Real traces can mix ordinary input and writes, and need their earlier preparation cost included for a full-lifecycle comparison. These percentages are chosen scenarios, not typical hit rates. Tools, retries, human time and price modifiers are excluded.

Output-heavy example: hold 1M input at 90% reads, 100K writes and no ordinary input, but generate 1M output. The token cost is $0.09 + $0.25 + $10 = $10.34, versus $12 uncached: 13.8% reduction. High reuse does not discount output, tools, retries or review time. Use the accepted-result method before calling a workflow cheaper.

Diagnose a cache miss

Compare a cold request with the next intended warm request before changing the workflow:

  • Expiry and routing: check time since the last write/reuse, the model’s lifetime, cache key and processing region. A matching prefix must reach a machine holding its entry; overflow routing can reduce reuse.
  • Model or prompt change: preserve the exact model revision, instructions, file order and reusable text. Check which earlier content changed before the matched boundary.
  • Tools and schemas: changes to tool order/definitions, output schemas or relevant settings can alter rendered instructions. They do not prove every setting change always invalidates all context.
  • Boundary placement: changing suffixes can miss a whole-message implicit boundary. Explicit-only mode needs a matching explicit breakpoint; extending an existing message can hide an old endpoint. Append a new message or preserve a supported block boundary where appropriate.
  • Effort updates: OpenAI documents a supported configuration-update route for changing reasoning effort while preserving earlier context. Use its model-specific procedure rather than assuming a top-level effort change is harmless.
  • Compaction: a shorter summary can change the prefix. Compare total input cost and acceptance afterward; fewer tokens can still help even when the hit percentage falls.

Use OpenAI’s cache diagnostics or the provider’s actual usage fields. Diagnose a miss, continue useful warm work, or make a compact handoff when quality and task boundaries justify it; do not keep an unhelpful thread alive just to improve a ratio.

Measure cache reuse and accepted work

For direct Responses API calls, the current usage recipe uses usage.input_tokens, usage.input_tokens_details.cached_tokens and usage.input_tokens_details.cache_write_tokens. Ordinary input is total input − reads − writes. Missing fields are unknown, not zero; a negative remainder signals invalid or mismatched telemetry. usage.output_tokens includes reasoning tokens, so do not add output_tokens_details.reasoning_tokens again. Other APIs and product clients may expose different fields; do not assume a Codex subscription meter is this API schema.

Copy one row per request; keep raw usage and receipts outside public logs:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
Request ID / date / task ID / attempt:
Provider route / exact requested and returned model / API or client version:
Requested and actual returned service tier / processing region:
Request input length / <=272K or >272K:
Total input / cached reads / cache writes / ordinary remainder:
Billable output (includes reasoning) / reasoning subset if exposed:
Cache mode / breakpoint locations / TTL / key identifier:
Prompt, tools, schemas or effort changed / compaction since prior request:
Tool charges / elapsed time / retry or failed attempt:
Estimated token cost / realized billed cost or unknown:
Accepted result / review and cleanup minutes:
Separate included-limit or purchased-credit meter, if applicable:

Within equal model, tier, region and context-price groups, calculate read share = sum(reads) / sum(input); count hit requests separately. Estimate token cost as (ordinary × input rate + reads × read rate + writes × write rate + output × output rate) / 1M, then reconcile actual billing and tools. Split long-context requests before aggregating; apply the full-request multiplier to each eligible request. The pricing owner supplies the rate card. Report cash, elapsed time, all attempts, accepted/submitted tasks and review minutes separately. Product credits have no separate cache-write fee and included subscription limits require their own meter; this formula does not predict either allowance.

Reduce cost without starving the agent

  • Keep project instructions concise, scoped, and stable.
  • Put stable instructions, tools, and examples before changing task content.
  • Search first; read targeted ranges instead of dumping whole repositories.
  • Bound logs, test output, and MCP responses.
  • Expose only the tools and MCP schemas needed for the task.
  • Set retry, subagent, time, and token budgets explicitly.
  • Compact at natural phase boundaries with decisions, evidence, diff state, and the exact next action.
  • Stop when acceptance criteria pass—or when the next step requires user authority.

Claude Code: reduce session context overhead

Anthropic’s session-efficiency guide, rechecked September 14, 2026, recommends keeping context relevant and choosing model and effort at a natural boundary. This is vendor guidance, not a measured AIHackers saving or evidence about Codex entitlements.

  • Set model and effort before starting the work.
  • Use /clear between unrelated tasks; preserve a useful checkpoint first.
  • Use /compact when an earlier phase of the same task is done, and specify the decisions and evidence to keep.
  • Inspect /context to identify session overhead.
  • @-mention the needed file instead of repeatedly rediscovering it.
  • Filter noisy command output before it enters context.
  • Use subagents only when isolation is worth duplicated setup and context.

Return to the cost-saving playbook to compare these interventions against accepted results and review time. Context efficiency and subscription capacity are separate measurements.

Codex versus OpenCode: run the same test

Do not compare Codex on GPT-5.6 Sol with OpenCode on a cheaper model and call the result harness efficiency. Use this card for several representative tasks:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
Model and provider:
Harness and version:
Repository commit:
Prompt and project instructions:
Tools/MCP schemas:
Reasoning effort:
Timeout and retry/subagent budgets:
Acceptance tests:

Accepted result (yes/no):
Input / cache-write / cache-read / output tokens:
Subscription quota, if applicable:
Tool calls, retries, compactions, and rereads:
Human review and cleanup minutes:
Total cost:
Cost per accepted result:

Hold model, provider, commit, prompt, tools, effort, timeout, and acceptance criteria constant. Report distributions and failure notes, not a permanent winner.

Frequently asked questions

What is prompt caching?

It reuses provider-side computation for an exact prompt prefix. It is not response caching, memory, RAG, or a local dependency cache.

Do cached tokens still cost money?

Usually yes. A cache hit discounts eligible repeated input under provider-specific rules; it does not make the whole conversation free.

Do cached tokens count toward the context window?

Yes. Caching changes serving and billing, not the amount of context the model attends to.

What does compaction do?

It replaces older history with a shorter summary. Missing execution detail can lead to cache reconstruction or targeted rereads.

Why do coding agents reread files?

Possible causes include compaction, omitted or truncated tool output, weak state preservation, and poor search strategy. A reread is behavior to measure, not proof of one universal defect.

Is OpenCode more token-efficient than Codex?

Not proven. Run the same model, provider, commit, prompt, tools, effort, timeout, and acceptance test, then compare cost per accepted result.

Did OpenAI reset Codex usage to compensate for caching defects?

No. OpenAI attributed the July 25 broad reset to outage recovery. The July 28 and July 29 announcements establish their scope but do not identify caching or compaction as a cause.

Sources and archive state


Global historical source/issue review: August 1, 2026. Cache mechanics, Sol 6.1 economics and usage logging rechecked October 3; the Claude Code workflow subsection retains its September 14 check. Earlier archive captures preserve earlier revisions; current-source archives are archive-pending. No paid measurements or traffic/SEO outcome were established.